QFAB is looking for one PhD candidate to develop statistical methodologies and bioinformatics tools to analyse high-throughput "Omics" biological data sets.
Project Description
Recent advances in ‘omics’ technologies now enable quantitative monitoring of the abundances of various biological molecules in a high-throughput manner and thereby enable the determination of the variation between different biological states on a genomic scale.
Popular ‘omics’ platforms include transcriptomics which measures mRNA transcript levels, proteomics which quantifies protein abundances, and metabolomics which determine abundances of small cellular metabolites, amongst others. However, no single ‘omics’ analysis can fully unravel the complexities of molecular biology. The integration of multiple layers of information is therefore crucial to better understand biological systems.
Data integration is a challenging task due to the high dimensionality of the data as well as their heterogeneous nature. Variable selection, i.e. co-jointly selecting the relevant and informative genes, metabolites, proteins, etc., is therefore essential to obtaining biologically meaningful information from the data. Some attempts have been made recently to perform variable selection and integrate heterogeneous data sets (L Cao, Boitard, & Besse, 2011; L Cao, Martin, Robert-Grani, & Besse, 2009; L Cao, Rossouw, Robert-Grani, & Besse, 2008). Approaches such as the ones developed in the mixOmics R package (L Cao, Gonzlez, & Djean, 2009), for example, have demonstrated the important role of such tools in assisting biologists to characterize complex biological systems (Le Cao et al, 2008, 2009a).
The decreasing costs in high-throughput platforms now enable repeated measured experiments on the same individuals or biological samples. The transcriptome, proteome and metabolome are dynamic entities, with the presence, abundance and function of each transcript, protein and metabolite being critically dependent on its temporal and spatial location. While current efforts have been made towards developing methodologies to integrate two omics data sets and understand the relationship between two types of biological entities at a given time point, or to analyse a single data set measured across several time points to identify correlated profiles, no attempt has been made so far to integrate data from more than two platforms and take the time dependency into account.
One PhD Scholarship in Applied Statistics or Bioinformatics available