Project description: Next generation sequencing strongly impacted our knowledge of the complement of genomic differences that reside in natural populations of plant species. But still, the common way to represent the “genome of a species” is one single reference sequence that misses out any kind of diversity.
As reference sequences are used as backbone (alignment targets) for resequencing analyses, all resequencing is biased towards the regions that are similar between the focal genome and the reference sequence. Regions that are absent in the reference sequence will not be included in the analysis.
Performing multiple resequencing studies against all known genomes would reduce this bias though comes with huge computational costs.
Novel computational structures (so-called genome graphs) can efficiently integrate many related genomes (e.g. all individuals of a species), but they cannot be used in combination with common alignment tools. We would like to develop graph theory-based methods that untangle the genome graphs into a linear sequence without loosing the information about the genomic differences. Such linear sequences could then be used with essentially all short and long read alignment algorithms - allowing for efficient resequencing against all known genomes of a species.
Funding and location: The PhD position is included in the International Max Planck Research School (IMPRS). The project will be performed in the Bioinformatics Group of the Department for Plant Developmental Biology, which is located at the Max Planck Institute for Plant Breeding Research (MPIPZ) in Cologne, Germany.
PhD in Bioinformatics