About this Event
2461 SW Campus Way, Corvallis, OR 97331
TITLE: Towards Foundational models for 16S Microbiome data
ABSTRACT: Microbiome sequencing provides useful signals for studying host phenotype and disease, and the large amount of publicly available 16S rRNA sequencing data creates an opportunity to learn reusable microbiome representations through large-scale pretraining. However, learning across studies is challenging because the observed amplicon sequence variants (ASVs) often differ substantially between cohorts, making fixed-vocabulary approaches difficult to transfer to new datasets. This thesis presents a sequence-aware representation learning framework for 16S V4 microbiome data that avoids requiring a fixed ASV vocabulary for input representation by encoding ASVs directly from their nucleotide sequences using a pretrained DNA language model and combining these embeddings with abundance information. A transformer encoder then learns contextualized microbial community representations, which are used to produce sample-level embeddings for downstream prediction tasks. Across a curated set of disease and phenotype prediction tasks, including COVID-19 detection, abdominal conditions, and antibiotic exposure, the proposed approach improves over fixed-vocabulary and sequence-only baselines, particularly when training and testing data come from different studies. These results suggest that representing ASVs through nucleotide sequence while modeling abundance and community context can improve the flexibility and cross-study generalizability of microbiome-based prediction models.
MAJOR ADVISOR: Xiaoli Fern
COMMITTEE: Prasad Tadepalli
COMMITTEE: Stephen Ramsey
MINOR ADVISOR: Pradad Tadepalli
GCR: Pankaj Jaiswal