Research
My current research sits at the intersection of multi-omics—focusing on exposomics, metabolomics, and proteomics—and data science, with the goal of uncovering complex biological mechanisms.
A central theme I’ve been developing in my work is stability. High-dimensional omics data are rich enough that almost any method will find something, but sampling variability means many findings don’t replicate — particularly in smaller cohorts. I’m interested in how we can make high-dimensional biological data more reproducible and interpretable — not just finding signals, but understanding whether those signals would hold up in a new sample. In a recent manuscript, I’ve been working on a stability-driven framework for multi-omics integration that combines sparse generalized canonical correlation analysis with repeated cross-validation and systematic evaluation of component- and feature-level reproducibility. Applying this to metabolomic and proteomic profiles in a thyroid cancer cohort, the framework identified latent structures that were consistent out-of-sample and tracked temporally with time to diagnosis.
My other work examines the relationship between per- and polyfluoroalkyl substances (PFAS) exposure and SARS-CoV-2 IgG antibody levels in pregnant individuals. Using untargeted metabolomic profiling and high-dimensional mediation analysis, we identified key metabolites mediating this association, offering insights into the biological pathways involved.
I believe that leveraging data-driven insights is key to improving public health outcomes and driving advancements in personalized medicine.