Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

SimTeller: Improving the Realism of Simulated DNA Alignments and Machine-Learning Detection for Simulated Alignments

Loading...
Thumbnail Image

Date

Authors

Zeng, Manyuan

Journal Title

Journal ISSN

Volume Title

Publisher

Abstract

Phylogenetics infers evolutionary relationships from DNA or protein sequence alignments. Simulated alignments are widely used in phylogenetics because they provide known ground truth and large-scale reproducible data. However, original simulation workflows often simplify the evolutionary process, reducing the realism of simulated datasets and limiting the performance of downstream machine-learning methods. This thesis addresses two interconnected research questions: how to improve the realism of simulated DNA alignments, and whether more realistic simulated datasets can improve machine-learning–based alignment classifiers. The first part developed an improved simulation workflow by incorporating ancestral sequence reconstruction (ASR) and mixture-model approach into the original simula- tion pipeline. To support large-scale mixture-model simulation, this project constructed the first large-scale database of mixture-model parameters and root sequences for 30,773 empirical DNA alignments. The existing Seqsharp classifiers generally failed to effectively classify alignments generated under the improved workflow, suggesting a smaller gap between the improved simulated data and empirical alignments. The second part investigated whether the improved simulated datasets could produce better alignment classifiers. I trained SimTeller, a set of new alignment classifiers us- ing convolutional neural network (CNN) and gradient boosting tree (GBT) approaches. Compared with Seqsharp, SimTeller demonstrated stronger robustness, practical ap- plicability, and generalisation ability across independent and cross-condition testing datasets. These results further confirmed that the improved simulation workflow from the first part may also benefit a broader range of phylogenetic machine-learning applications. Overall, this thesis developed an improved simulation workflow for DNA alignments and proposed SimTeller, a set of machine-learning–based classifiers that may serve as useful tools for evaluating the realism of simulated alignments

Description

the author deposted 21.07.2026

Citation

Source

Book Title

Entity type

Access Statement

License Rights

DOI

Restricted until