Project information
MULTIPANDA: Multiscale PANgenomics for Data-driven Applications
(MULTIPANDA)
- Project Identification
- 101344560
- Project Period
- 10/2026 - 9/2030
- Investor / Pogramme / Project type
-
European Union
- Horizon Europe
- Marie Skłodowska-Curie Staff Exchanges (MSCA SE)
- MU Faculty or unit
- Faculty of Informatics
- Cooperating Organization
-
Comenius University in Bratislava
The University of Tokyo
Uniwersytet Warszawski
Universita di Pisa
Institut Pasteur
University of Bielefeld
Geneton, s.r.o.
The Cyprus Institute
University of Milano-Bicocca (UNIMIB)
University of Helsinki
The Pennsylvania State University
Simon Fraser University
Cornell University
The Regents of the University of California
Population-scale human genome assembly projects are becoming common. The SARS-CoV-2 pandemic has shown the importance of sequencing and classifying viral genomes. While both are instances of pangenomes, that is an ensemble of complete genomes, their sizes and properties are wildly different. A human genome is orders of magnitude longer than that of SARS-CoV-2, with a sequencing cost that exhibits the same difference.
Moreover, the analyses that are performed on those pangenomes are completely different; for example large structural variations are considered only on human genomes, especially to highlight the fundamental genomic differences between healthy and tumor cells. But even restricting our attention of human pangenomes, we can see that finding single nucleotide polymorphisms (by far the most common kind of mutations) requires a different approach than finding large structural variation (for example, the loss of an entire gene that is common in tumors).
This project builds upon the idea of representing pangenomes as a labeled graph, where each vertex is a portion of a genome and a path is a complete genome (or a complete haplotype) to design multi-scale algorithms that are able to adapt to the specific problem, building and considering a graph with the most appropriate resolution. This will allow to couple high accuracy with quick turnaround time, limiting the computational resources needed and making feasible to attack some of the most challenging problems in bioinformatics where massive pangenome graphs are able to remove redundancies without having to discard individual differences, such that the reliable identification of mutations causing rare diseases.