Project information
MULTIPANDA: Multiscale PANgenomics for Data-driven Applications (MULTIPANDA)

Population-scale human genome assembly projects are becoming common. The SARS-CoV-2 pandemic has shown the importance of sequencing and classifying viral genomes. While both are instances of pangenomes, that is an ensemble of complete genomes, their sizes and properties are wildly different. A human genome is orders of magnitude longer than that of SARS-CoV-2, with a sequencing cost that exhibits the same difference.
Moreover, the analyses that are performed on those pangenomes are completely different; for example large structural variations are considered only on human genomes, especially to highlight the fundamental genomic differences between healthy and tumor cells. But even restricting our attention of human pangenomes, we can see that finding single nucleotide polymorphisms (by far the most common kind of mutations) requires a different approach than finding large structural variation (for example, the loss of an entire gene that is common in tumors).
This project builds upon the idea of representing pangenomes as a labeled graph, where each vertex is a portion of a genome and a path is a complete genome (or a complete haplotype) to design multi-scale algorithms that are able to adapt to the specific problem, building and considering a graph with the most appropriate resolution. This will allow to couple high accuracy with quick turnaround time, limiting the computational resources needed and making feasible to attack some of the most challenging problems in bioinformatics where massive pangenome graphs are able to remove redundancies without having to discard individual differences, such that the reliable identification of mutations causing rare diseases.

You are running an old browser version. We recommend updating your browser to its latest version.

More info