New Methods to Improve Phylogenomic Inference
Abstract
A complete set of DNA - known as a genome - provides a powerful resource for inferring evolutionary relationships among species, which are often represented as a bifurcating (phylogenetic) tree. However, it is increasingly appreciated that a single tree is insufficient to capture the full evolutionary history of a study system due to processes such as incomplete lineage sorting (ILS), hybridisation, introgression, and recombination that could create a mosaic of histories along the genomes. Estimation of this variation of local histories - often referred to as gene tree discordance - thus depends on the accuracy of the sequence data and the robustness of the inference methods being used.
In Chapter I, I assessed the non-overlapping window method that is commonly-used to address gene tree discordance from a set of aligned genomes, but often with an arbitrary selection of fixed window size. Using simulated chromosomes with different degrees of recombination and ILS, I showed that the Akaike Information Criterion (AIC) provides a less arbitrary approach to select the best window size given the alignment. I then applied this approach to empirical datasets from Heliconius butterflies and great apes, and showed that the best window sizes for these groups range from 125-250bp (for Heliconius butterflies) and 500-1,000bp (for great apes) across chromosomes.
In Chapter II, I extended the information-theory-based approach proposed in Chapter I to accommodate variable window sizes, because the size of non-recombining blocks could vary along the chromosomes. To do this, I developed an iterative splitting-and-merging approach that evaluates local improvements in AIC. I showed that using variable window sizes has consistently better accuracy than using fixed window sizes in recovering the 'true' topologies from simulated alignments, with at least 80% accuracy across simulations. I then applied this approach on empirical datasets from Heliconius butterflies and great apes, and further showed that the best window sizes varied substantially across chromosomes.
In Chapter III, I leveraged the availability of multiple reference genomes and proposed a phylogenetic method to detect reference bias at individual loci, assuming that in the absence of reference bias, reconstructed sequences of a single locus from the same sample should be identical regardless of the reference being used. Across empirical datasets of nine Eucalyptus species, I found that more than one-quarter of the reconstructed BUSCO loci showed strong evidence of reference bias, which consequently affected the species tree inference. Excluding these putatively biased loci, coupled with using a closely-related reference genome during mapping, resulted in species tree topologies that were more consistent with the published tree.
In Chapter IV, I assessed the role of phylogenetic distance as a barrier to introgression in Eucalyptus globulus. To do this, I estimated the proportion of introgression between E. globulus and 56 other Eucalypts using QuIBL (Quantifying Introgression via Branch Lengths). Based on ~1,000 BUSCO loci, I found significant correlations between the proportion of historical introgression, present-day crossability, and pairwise phylogenetic distance, which suggested that the accumulation of genetic changes over evolutionary time has played a central role in shaping hybridisation and introgression patterns of the group.
Overall, these chapters highlight two key challenges in phylogenomic analyses: inferring gene trees and assessing reference bias. In response, I proposed two phylogenetic methods and presented one case study that together provide a useful framework for future phylogenomic inference.
Description
Keywords
Citation
Collections
Source
Type
Book Title
Entity type
Access Statement
License Rights
Restricted until
Downloads
File
Description
Thesis Material