Sign in

Na Cai

@caina89.bsky.social
929 followers 191 following 84 posts

Assistant Prof at D-BSSE, ETH Zurich, studying genetics of complex traits, with a focus on psychiatric disorders www.nacailab.com

PostsRepliesMedia
Reposted by Na Cai
Yun S. Song @yun-s-song.bsky.social · 09/09/2026
We are thrilled to share that our GPN-Star manuscript is now published and freely available: doi.org/10.1038/s415... (1/n)
doi.org
Predicting genome-wide functional constraints with GPN-Star - Nature
GPN-Star, a genomic language model with a phylogeny-aware architecture for whole-genome alignment data, is shown to be a scalable and flexible tool for genetic variant effect prediction across species...
112247
Reposted by Na Cai
D-BSSE, ETH Zurich @handle.invalid · 26/05/2026
🧠Grant from #HortenHealthFoundation powers #AI-based research in computational genetics led by Na Cai @caina89.bsky.social @ethz.ch to uncover molecular subtypes of major depressive disorder 👉🏾 tinyurl.com/2wyrjxcs
021
Reposted by Na Cai
Mount Sinai Institute for Genomic Health @sinaigenhealth.bsky.social · 20/04/2026
🌟 Rising Stars Seminar 🧬 Learning Lifetime Disease Liability Reveals and Removes Genetic Confounding in Electronic Health Records 🎤 Na Cai, DPhil 📅 May 12 | 9–10am ET @sinaigenetics.bsky.social @caina89.bsky.social #Genomics #GWAS #AI
042
Na Cai @caina89.bsky.social · 13/03/2026
📢 Call for submissions: The RECOMB-Genetics satellite workshop will take place in Thessaloniki, Greece 🇬🇷 on May 25, 2026, just before the RECOMB conference (May 26–29) @recombconf.bsky.social 📅 Submission deadline: April 2 (AoE) 👉 recomb.org/recomb2026/r... Submit your work and join us!
recomb.org
052
Na Cai @caina89.bsky.social · 22/02/2026
We have released summary stats here zenodo.org/records/1840... and code here github.com/yazhengdi/ED.... Feedback/criticism/comments welcome! 19/19
github.com
000
Na Cai @caina89.bsky.social · 22/02/2026
Both Yazheng and I have learnt a lot from working on this project. We thank the participants of UKB and other cohorts for enabling our work, and many friends for giving us valuable feedback. Yazheng will be presenting this work at #recomb26 @recombconf.bsky.social 18/n
110
Na Cai @caina89.bsky.social · 22/02/2026
We hope this work provides a proof of concept and blueprint for improving the specificity and interpretability of EHR-based genetic studies, as well as their downstream utility in drug target identification and risk stratification. 17/n
100
Na Cai @caina89.bsky.social · 22/02/2026
Overall, our EDGAR framework enables the prediction of disease-specific liabilities from EHR events, disease-specific measures and deep phenotype labels, and the separation of genetic effects on disease-specific liability from heritable biases that influence EHR events. 16/n
110
Na Cai @caina89.bsky.social · 22/02/2026
When we did that, we find that high rGs between external EHR GWAS with confounding traits disappeared. This implies that when we have well-predicted disease liabilities, enabled by EDGAR, we can identify biases in one EHR, and remove it from existing GWAS on another. 15/n
100
Na Cai @caina89.bsky.social · 22/02/2026
Finally, we ask if the Common Bias identified in UKB is likely generalizable between EHRs: consistent with our hypothesis, we find that it has high rGs with external EHR GWAS, but not external deep phenotype GWAS. This makes us attempt removing it from external EHR GWAS. 14/n
110
Na Cai @caina89.bsky.social · 22/02/2026
We then identify a Common Bias factor across diseases with genomic SEM, and find that it has high rGs with socioeconomic and behavioral traits, many previously shown to affect UKB participation www.nature.com/articles/s41... 13/n
100
Na Cai @caina89.bsky.social · 22/02/2026
This drives us to first isolate the disease-specific heritable confounder affecting EHR codes of each disease, using GWAS-by-subtraction, and find that genetic effects contributing to these biases in different diseases have high rGs (i.e. they are shared) with each other. 12/n
100
Na Cai @caina89.bsky.social · 22/02/2026
This is corroborated by our cross-disease rG analysis in UKB: while deep phenotypes and EDGAR liabilities show similar cross-disease rGs, raw EHR codes show highly inflated rGs. This is likely due to a common heritable confounder www.nature.com/articles/s41... 11/n
100
Na Cai @caina89.bsky.social · 22/02/2026
In contrast, raw UKB EHR codes have lower rGs with external deep phenotype GWAS, but higher rGs and hits replication in external EHR GWAS - this shows replicability between EHRs goes beyond etiological nature of diseases - it is likely an EHR-systemic factor drives this. 10/n
110
Na Cai @caina89.bsky.social · 22/02/2026
We validate our findings in external GWAS: compared to EHR phenotypes, EDGAR liabilities show higher GWAS hit replication and rGs with external deep phenotype GWAS - this suggests EDGAR liabilities capture replicable, disease-specific genetic effects. 9/n
100
Na Cai @caina89.bsky.social · 22/02/2026
We then perform and compare GWAS on raw EHR codes, deep disease labels and EDGAR predicted liabilities: for most diseases, EDGAR predictions give more GWAS hits, higher rGs with deep phenotype labels, and higher PRS predictability and specificity. 8/n
101
Na Cai @caina89.bsky.social · 22/02/2026
We first show that raw EHR codes have low correlations with deep disease labels for nine diseases in UKB, and EDGAR liability predictions do significantly better, especially when incorporating disease-relevant measures as inputs. 7/n
100
Na Cai @caina89.bsky.social · 22/02/2026
We augment EDGAR with an active learning model that prioritizes individuals for obtaining deep disease labels, which reduces the N labels needed by >50% - substantially reducing costs for obtaining disease labels through patient recall in realistic EHR settings. 6/n
110
Na Cai @caina89.bsky.social · 22/02/2026
EDGAR predicts lifetime disease liabilities through aligning counts of EHR events with disease-specific measures realistically available in EHRs (e.g. blood biochemistry or spirometry) and independently ascertained, clinically validated disease labels (“deep phenotypes”). 5/n
100
Na Cai @caina89.bsky.social · 22/02/2026
In our new paper, we propose EDGAR (EHR Disease liability prediction for Genetic Architecture Recovery), a prediction framework that combines the scale of EHR with the disease-relevance of deep phenotypes to break this circularity of bias. 4/n
100
Na Cai @caina89.bsky.social · 22/02/2026
These influences, many of which hertiable, are difficult to disentangle from disease-specific genetics. They may also be replicable across EHRs, making confounded genetic findings look robust. Finally, these biases are likely to propagate in next-event prediction models. 3/n
100
Na Cai @caina89.bsky.social · 22/02/2026
First some background: using EHR codes in GWAS drives increase in statistical power in GWAS meta-analyses. Its potential, however, is undermined by systemic factors like coding practices and healthcare-access disparities between demographics: www.nature.com/articles/s41... 2/n
nature.com
Harnessing EHR data for health research - Nature Medicine
Electronic health records hold immense potential for providing clinically useful insights for populations and individuals; this Review summarizes the opportunities and challenges, with an emphasis on ...
110
Na Cai @caina89.bsky.social · 22/02/2026
Our new preprint “Learning lifetime disease liability reveals and removes genetic confounding in electronic health records” is now online! Link to paper: This work is led by my postdoc Yazheng Di and it’s our first project at @bsse.ethz.ch :) medrxiv.org/cgi/content/... Thread 1/n
medrxiv.org
Learning lifetime disease liability reveals and removes genetic confounding in electronic health records
Electronic health records (EHRs) have become the cornerstone of population-scale genetic studies1, but factors including patterns of healthcare use shape which and how diagnoses are recorded, leading ...
2176
Na Cai @caina89.bsky.social · 16/02/2026
My department @bsse.ethz.ch is inviting applications for a new assistant professor (tenure-track) in computational immunology, interested candidates please see advert and apply! Deadline 15 April 2026 ethz.ch/en/the-eth-z...
ethz.ch
Assistant Professor (Tenure Track) of Computational Immunology
012
Reposted by Na Cai
Jonathan Pritchard @jkpritch.bsky.social · 06/01/2026
New preprint alert: we use sign errors as a test of how well TWAS works. Very worryingly we find that TWAS gets the sign wrong around 1/3 of the time (compared to 50% for pure guessing). You can read more about our analysis here, and what we think is going on 👇
56728
Na Cai @caina89.bsky.social · 20/12/2025
And www.nature.com/articles/s41... in a more cross disorder setting - this one discusses the impact of phenotyping on cross disorder analyses
nature.com
Assessment and ascertainment in psychiatric molecular genetics: challenges and opportunities for cross-disorder research - Molecular Psychiatry
Molecular Psychiatry - Assessment and ascertainment in psychiatric molecular genetics: challenges and opportunities for cross-disorder research
020
Na Cai @caina89.bsky.social · 20/12/2025
We have two more commentary/review like papers on this www.nature.com/articles/s41...
nature.com
The genetic basis of major depressive disorder - Molecular Psychiatry
Molecular Psychiatry - The genetic basis of major depressive disorder
010
Na Cai @caina89.bsky.social · 20/12/2025
Importantly in this paper we derive a new metric PRS pleiotropy which we use to show shallow phenotypes indeed give non specific gwas signal that leads to non specific PRS predictions
210
Na Cai @caina89.bsky.social · 20/12/2025
We then worked out we can improve shallow phenotypes in biobank settings where there’s a small subset of individuals with high quality phenotypes through imputation www.nature.com/articles/s41...
nature.com
Phenotype integration improves power and preserves specificity in biobank-based genetic studies of major depressive disorder - Nature Genetics
Phenotype imputation increases the effective sample size of major depressive disorder cases in UK Biobank, enhancing study power and polygenic risk score (PRS) accuracy. A new pleiotropy metric enable...
250
Na Cai @caina89.bsky.social · 20/12/2025
Hi @michelnivard.bsky.social we have written quite a lot on sample size vs phenotyping quality. The earliest was www.nature.com/articles/s41... on depression where we show shallow phenotyping, despite giving higher gwas power, gives non specific gwas signal
nature.com
Minimal phenotyping yields genome-wide association signals of low specificity for major depression - Nature Genetics
Genetic analyses of depression based on minimal phenotyping identify nonspecific genetic risk factors shared between major depressive disorder (MDD) and other psychiatric conditions, suggesting that t...
240
Reposted by Na Cai
Sasha Gusev @sashagusevposts.bsky.social · 13/12/2025
I wrote about the bizarre case of Herasight, the embryo selection company going all in on eugenics.
open.substack.com
Embryo selection company Herasight goes all in on eugenics
...
612380
Reposted by Na Cai
Anna Docherty @annadocherty.bsky.social · 19/10/2025
The PGC Suicide Working Group will provide a symposium at #WCPG2025 🇲🇽 covering our latest multi-ancestry GWAS and CNV meta-analyses, sex-specific meta-analyses, and GxEHR analyses. 🥳 @lcstoshio.bsky.social @andreyshabalin.bsky.social @sarahcolbert.bsky.social @caina89.bsky.social
2185
Reposted by Na Cai
Yun S. Song @yun-s-song.bsky.social · 22/09/2025
We are excited to share GPN-Star, a cost-effective, biologically grounded genomic language modeling framework that achieves state-of-the-art performance across a wide range of variant effect prediction tasks relevant to human genetics. www.biorxiv.org/content/10.1... (1/n)
417690
Reposted by Na Cai
Rayan Chikhi @rayanchikhi.bsky.social · 03/09/2025
🌎👩‍🔬 For 15+ years biology has accumulated petabytes (million gigabytes) of🧬DNA sequencing data🧬 from the far reaches of our planet.🦠🍄🌵 Logan now democratizes efficient access to the world’s most comprehensive genetics dataset. Free and open. doi.org/10.1101/2024...
3218118
Reposted by Na Cai
Peter Sudmant @psudmant.bsky.social · 13/08/2025
The Sudmant lab at UC Berkeley is seeking a postdoc to work on a fully funded NIH project to understand differences in DNA repair and somatic mutation across the primate tree of life. Please spread widely to those who may be interested aprecruit.berkeley.edu/JPF05052
aprecruit.berkeley.edu
Postdoctoral Scholar – Genomics, Aging, Somatic Mutation, Structural Variation, Evolution , Cancer – Integrative Biology
University of California, Berkeley is hiring. Apply now!
05555
Reposted by Na Cai
Yun S. Song @yun-s-song.bsky.social · 06/06/2025
The 2026 Probabilistic Modeling in Genomics (ProbGen) meeting will be held at UC Berkeley, March 25-28, 2026. We have an amazing list of keynote speakers and session chairs: probgen2026.github.io Please help spread the news.
probgen2026.github.io
Home - ProbGen 2026
Your Site Description
27036
Na Cai @caina89.bsky.social · 26/07/2025
Congratulations, very exciting! 🎉
010
Reposted by Na Cai
Jolien Rietkerk @jolienrietkerk.bsky.social · 24/07/2025
‘Genetic Risk Effects on Psychiatric Disorders Act in Sets’, the title of our new preprint on @medrxivpreprint.bsky.social! This huge collaborative effort advances our understanding of psychiatric genetic architecture and emphasizes the importance of looking beyond additive effects. 🧵1/n; 🧪👩🏽‍🔬🧬
1118
Na Cai @caina89.bsky.social · 24/07/2025
Link to GitHub page: github.com/caina89/psyc...
github.com
GitHub - caina89/psychCE: Code used to implement coordinated epistasis on psychiatric disorders
Code used to implement coordinated epistasis on psychiatric disorders - caina89/psychCE
110
Na Cai @caina89.bsky.social · 24/07/2025
Link to the preprint: www.medrxiv.org/lookup/conte...
medrxiv.org
Genetic risk effects on psychiatric disorders act in sets
Genetic studies of psychiatric disorders have typically assumed that all genetic effects contribute additively to disease liability. However, it is likely that psychiatric disorders have unrecognized ...
173
Na Cai @caina89.bsky.social · 24/07/2025
This work would not have been possible if not for the persistent efforts of @jolienrietkerk.bsky.social Andy Dahl, Jonathan Flint, Andrew Schork and many others, as well as the data from participants in @ukbiobank.bsky.social and IPSYCH. We hope it will be informative to the field. 12/n
120
Na Cai @caina89.bsky.social · 24/07/2025
Though our investigations and findings are centered on psychiatric disorders, the implications are generalizable to all complex traits and diseases, especially those with heterogeneous architectures and unclear diagnostic boundaries. 11/n
130
Na Cai @caina89.bsky.social · 24/07/2025
Overall, our work provides a novel metric, the CE test, for informing diagnostic boundaries. Our results show that genetic effects on psych disorders act in sets, and calls for a re-evaluation of current approaches and assumptions. 10/n
130
Na Cai @caina89.bsky.social · 24/07/2025
Finally, we find that common genetic effects across all five psych disorders, expected to capture common etiological axes among them, forms a cross-order set most plausibly explained by common confounders external to each disorder’s etiology. 9/n
480
Na Cai @caina89.bsky.social · 24/07/2025
We further show that disorder-specific sets lead to comorbidity between two disorders, refuting the assumption that comorbidity is the result of additive effects of genetic effects on both, exemplified by the expectation that high rGs would lead to high comorbidity. 8/n
130
Na Cai @caina89.bsky.social · 24/07/2025
These findings imply that rGs are insufficient to inform nosology, and that the CE test provides a way to determine whether putative disorders or subtypes should be merged or remain separate diagnostically. 7/n
150
Na Cai @caina89.bsky.social · 24/07/2025
We also show that sets of genetic effects are disorder-specific, despite high genetic correlations (rG) between them. This opposes a well-established conjecture, based on high rG, that clinical boundaries do not reflect their pathogenic processes. 6/n
151
Na Cai @caina89.bsky.social · 24/07/2025
In our paper, we test for CE in five psychiatric disorders in the @ukbiobank.bsky.social and the iPSYCH dataset using both polygenic risk scores (PRS) and family-based genetic risk scores (FGRS). We find negative CE for 3 out of 5 disorders, demonstrating there are unrecognized subtypes. 5/n
120
Na Cai @caina89.bsky.social · 24/07/2025
The existence of these sets induces a structured form of statistical interactions called coordinated epistasis (CE), detailed in previous publications by Andy Dahl, Noah Zaitlen www.pnas.org/doi/10.1073/...; negative CE is seen when there are different sets of genetic effects. 4/n
pnas.org
A model and test for coordinated polygenic epistasis in complex traits | PNAS
Interactions between genetic variants—epistasis—is pervasive in model systems and can profoundly impact evolutionary adaption, population disease d...
140
Na Cai @caina89.bsky.social · 24/07/2025
These sets, should they exist, would each drive an unrecognized etiological subtype. Our paper tests for the existence of these sets in five psychiatric disorders. 3/n
120