Sign in

David Kelley

@drkbio.bsky.social
171 followers 57 following 54 posts

Making sophisticated guesses at how DNA will behave.

PostsRepliesMedia
David Kelley @drkbio.bsky.social · 14/10/2025
Will this recipe work for other organisms? We think it depends on genome size and proportion of nucleotides under selection, which drives the value of the self-supervised stage and training data scale. An exciting question for future work!
020
David Kelley @drkbio.bsky.social · 14/10/2025
This was a massive effort, driven by the incredible work of Calico intern Kuan-Hao Chao (@kuanhaochao.bsky.social ). Huge thanks to him, Majed Mohamed Magzoub, and Johannes Linder!
100
David Kelley @drkbio.bsky.social · 14/10/2025
My take: While MPRAs are powerful, they lose vital genomic context like local chromatin and post-transcriptional regulation. For modeling complex gene regulation in vivo, models trained on endogenous sequences are essential.
100
David Kelley @drkbio.bsky.social · 14/10/2025
Each wins on its “home field”: * MPRA-trained models excel at predicting MPRA data, including variant sequences. * Shorkie excels at predicting expression from promoters in their natural genomic context and eQTLs.
100
David Kelley @drkbio.bsky.social · 14/10/2025
How does Shorkie compare to models trained on massively parallel reporter assays (MPRAs)?
100
David Kelley @drkbio.bsky.social · 14/10/2025
This translates to variant effect prediction where Shorkie accurately predicts the impact of cis-eQTLs, outperforming alternative models at classifying influential regulatory variants.
100
David Kelley @drkbio.bsky.social · 14/10/2025
Shorkie also captures dynamic regulatory changes. Using new time-course RNA-seq data from TF inductions, we showed Shorkie can track how the importance of specific TF motifs changes over time.
100
David Kelley @drkbio.bsky.social · 14/10/2025
This pre-training strategy makes a huge difference. Shorkie substantially outperforms the same model trained from scratch, boosting gene-level expression prediction from a Pearson's R of 0.74 to 0.88.
110
David Kelley @drkbio.bsky.social · 14/10/2025
But which genomes work best? We trained on different phylogenetic levels, from close S. cerevisiae strains to the fungal kingdom. The Saccharomycetales order was the sweet spot, providing the right balance of diversity and conserved regulatory grammar for the model to learn from.
100
David Kelley @drkbio.bsky.social · 14/10/2025
Our hypothesis: Jumpstart supervised learning with self-supervision--before predicting chromatin and expression, we first asked our model to predict masked-out nucleotides across many related genomes, so it learns conserved elements like genes and their promoters.
100
David Kelley @drkbio.bsky.social · 14/10/2025
However, yeast's small genome provides limited data, making it tough for deep learning models to learn complex regulatory rules from scratch.
100
David Kelley @drkbio.bsky.social · 14/10/2025
At Calico, we've been studying S. cerevisiae for years to understand replicative aging. Along the way, we've generated rich datasets to probe its regulatory networks, which helped make this work possible.
100
David Kelley @drkbio.bsky.social · 14/10/2025
Excited to share our new paper on predicting gene expression in yeast! We introduce "Shorkie," a supervised ML model that builds off a self-supervised foundation to interpret regulatory DNA. Preprint: www.biorxiv.org/content/10.1...
biorxiv.org
Predicting dynamic expression patterns in budding yeast with a fungal DNA language model
Predicting gene expression from DNA sequence remains challenging due to complex regulatory codes. We introduce a masked DNA language model pretrained on 165 fungal genomes closely related to budding y...
196
David Kelley @drkbio.bsky.social · 04/08/2025
The poster abstract deadline for the @keystonesymposia.bsky.social AI in Molecular Biology meeting in Santa Fe is coming up on August 21st, so get your submissions in! www.keystonesymposia.org/conferences/...
keystonesymposia.org
AI in Molecular Biology | Keystone Symposia
Join us at the Keystone Symposia on AI in Molecular Biology, September 2025, in Santa Fe, with field leaders!
021
David Kelley @drkbio.bsky.social · 23/07/2025
We’ve done some experiments, but the metrics aren’t conclusive, so choose your own adventure! We’ve released these models open source, open weight for all to use. github.com/calico/borzo...
github.com
borzoi-paper/extensions/prime at main · calico/borzoi-paper
Analyses related to the Borzoi paper. Contribute to calico/borzoi-paper development by creating an account on GitHub.
020
David Kelley @drkbio.bsky.social · 23/07/2025
We hypothesized that training with cell-type-specific and 3' data might make these models particularly effective for transfer to datasets with similar properties.
110
David Kelley @drkbio.bsky.social · 23/07/2025
Transfer learning has emerged as a key application for multitask sequence models like these. For more, check out another recent paper from Han Yuan, whose analysis explores various transfer strategies and shows how powerful this approach can be. www.biorxiv.org/content/10.1...
biorxiv.org
Parameter-Efficient Fine-Tuning of a Supervised Regulatory Sequence Model
DNA sequence deep learning models accurately predict epigenetic and transcriptional profiles, enabling analysis of gene regulation and genetic variant effects. While large-scale training models like E...
100
David Kelley @drkbio.bsky.social · 23/07/2025
Hence the name: Borzoi Prime to emphasize their 3’ expertise!
100
David Kelley @drkbio.bsky.social · 23/07/2025
Indeed, he discovered the new models better predict alternative polyadenylation and QTL variants that affect where transcripts get cleaved and polyadenylated. This key regulatory layer influences cell type-specific protein production.
100
David Kelley @drkbio.bsky.social · 23/07/2025
Drawing on his expertise and interest in isoform regulation, Johannes hypothesized that single-cell RNA-seq’s 3’ sequencing protocols might reveal additional capabilities in these models.
100
David Kelley @drkbio.bsky.social · 23/07/2025
Using single cell eQTL studies, he evaluated the cell type specific variant effect predictions and found good concordance.
100
David Kelley @drkbio.bsky.social · 23/07/2025
As cell-type-specific applications emerged, Johannes Linder took a fresh look.
100
David Kelley @drkbio.bsky.social · 23/07/2025
We trained these models in early 2023 (which is why they’re algorithmically similar to the originals), but initial metrics were underwhelming, so we shelved them.
100
David Kelley @drkbio.bsky.social · 23/07/2025
Side note—want your amazing data included in future training runs of open source, open weight models? Make and release BigWig tracks!
100
David Kelley @drkbio.bsky.social · 23/07/2025
We curated several cell atlas collections to produce pseudobulk coverage tracks. Thank you to the CZI Tabula projects and the BICCN Brain Cell Atlas for making this possible!
100
David Kelley @drkbio.bsky.social · 23/07/2025
A limitation of the first Borzoi training run was the absence of cell type specific RNA-seq tracks; most are heterogeneous bulk samples.
100
David Kelley @drkbio.bsky.social · 23/07/2025
We’re excited to share a follow-up Borzoi training run and an analysis of the capabilities that emerged. www.biorxiv.org/content/10.1...
biorxiv.org
Predicting cell type-specific coverage profiles from DNA sequence
Predicting expression profiles from RNA-seq experiments provides a powerful approach for universal sequence-based variant effect prediction, enabling researchers to score variants that affect total ge...
130
David Kelley @drkbio.bsky.social · 21/07/2025
Alongside the manuscript and analysis, we released Borzoi predictions for 19.5 million common and low-frequency UK Biobank variants. Code for scoring additional variants with Borzoi is available here: github.com/calico/baske...
020
David Kelley @drkbio.bsky.social · 21/07/2025
Moving forward, we suspect there are further improvements available. The Borzoi predictions cover most body tissues, but they aren’t yet zoomed into specific cell types. Alternative nonlinear heritability models may usurp S-LDSC for fitting variant priors.
120
David Kelley @drkbio.bsky.social · 21/07/2025
Generally, we found that Borzoi predictions improve fine-mapping clarity and gene prioritization. We’re using Sniff to better analyze aging-related trait GWAS at Calico.
110
David Kelley @drkbio.bsky.social · 21/07/2025
In our paper, we used Borzoi (our latest regulatory sequence activity predictor) for this approach, calling the overall workflow Sniff (a quintessential pastime for all working hounds!)
130
David Kelley @drkbio.bsky.social · 21/07/2025
We can use these importance scores as priors in Bayesian fine-mapping methods like SuSiE. Variants that both look functional AND show statistical association get prioritized.
110
David Kelley @drkbio.bsky.social · 21/07/2025
The S-LDSC and PolyFun frameworks address these questions by asking: are variants with specific functional signatures more likely to drive trait associations? This lets us score each variant’s predicted importance for the trait.
110
David Kelley @drkbio.bsky.social · 21/07/2025
Can we learn which aspects of these fingerprints matter for each disease? Does a GWAS provide enough signal to figure out if liver function matters more than brain function for cholesterol levels?
110
David Kelley @drkbio.bsky.social · 21/07/2025
Promisingly, machine learning approaches to learn sequence to regulatory function are advancing and now capable of usefully predicting how genetic variants perturb gene regulation across the body. Think of it as creating a detailed “functional fingerprint” for each variant.
120
David Kelley @drkbio.bsky.social · 21/07/2025
However, gene regulation is complicated, varying across cell types of the body and responding to their environment. So predicting variant effects requires understanding this rich, multidimensional landscape.
110
David Kelley @drkbio.bsky.social · 21/07/2025
To affect a trait, a variant must first change how genes work—either by altering proteins or gene regulation. If we can predict these functional effects, we can be smarter about which variants to focus on.
110
David Kelley @drkbio.bsky.social · 21/07/2025
Genetic association studies are incredibly powerful for finding DNA regions linked to traits, but they often implicate dozens of variants due to linkage disequilibrium. Which one is the real culprit?
120
David Kelley @drkbio.bsky.social · 21/07/2025
I'm excited to share work on a research direction my team has been advancing: connecting machine learning derived genetic variant embeddings to downstream tasks in human genetics. This work was led by the amazing Divyanshi Srivastava! www.biorxiv.org/content/10.1...
biorxiv.org
Borzoi-informed fine mapping improves causal variant prioritization in complex trait GWAS
Genome-wide association studies (GWAS) have identified thousands of trait-associated loci. Prioritizing causal variants within these loci is critical for characterizing trait biology. Statistical fine...
23317
Reposted by David Kelley
Han Yuan @hy395.bsky.social · 02/06/2025
1/ DNA sequence models like Borzoi predict gene expression and variant effects across tissues — but how can someone adapt the model to a custom experiment? @drkbio.bsky.social, Johannes Linder and I propose a solution via parameter-efficient fine-tuning (PEFT). www.biorxiv.org/content/10.1...
biorxiv.org
Parameter-Efficient Fine-Tuning of a Supervised Regulatory Sequence Model
DNA sequence deep learning models accurately predict epigenetic and transcriptional profiles, enabling analysis of gene regulation and genetic variant effects. While large-scale training models like E...
13312
David Kelley @drkbio.bsky.social · 01/06/2025
The short talk and scholarship deadlines for the @keystonesymposia.bsky.social AI in Molecular Biology meeting in September were extended to June 3, so get your last minute submissions in! www.keystonesymposia.org/conferences/...
keystonesymposia.org
AI in Molecular Biology | Keystone Symposia
Join us at the Keystone Symposia on AI in Molecular Biology, September 2025, in Santa Fe, with field leaders!
021
David Kelley @drkbio.bsky.social · 11/05/2025
The short talk and scholarship deadlines for the @keystonesymposia.bsky.social AI in Molecular Biology meeting in September are coming up fast, May 20th. Looking forward to seeing the submissions! www.keystonesymposia.org/conferences/...
0174
David Kelley @drkbio.bsky.social · 20/04/2025
We're one month away from the short talk abstract and scholarship deadlines. Short Talk Abstract Deadline: May 20th, 2025 Scholarship Deadline: May 20th, 2025 Early Registration Deadline: July 17th, 2025 Poster Abstract Deadline: August 21st, 2025
022
David Kelley @drkbio.bsky.social · 20/04/2025
Join us to explore this rapidly emerging field in Santa Fe! keysym.us/KSAIBio26 #KSAIBio26
keysym.us
AI in Molecular Biology | Keystone Symposia
Join us at the Keystone Symposia on AI in Molecular Biology, September 2025, in Santa Fe, with field leaders!
110
David Kelley @drkbio.bsky.social · 20/04/2025
Working with a great team to organize @keystonesymposia.bsky.social AI in MolecularBiology, this September! We aimed for a wide range of biological topics, and emphasized speakers who blend sophisticated machine learning with compelling biological questions and analysis.
1156
David Kelley @drkbio.bsky.social · 17/04/2025
We implemented these techniques into our code repository github.com/calico/baske..., which the Borzoi repository github.com/calico/borzoi makes use of, so try it out!
github.com
GitHub - calico/baskerville: Machine learning methods for DNA sequence analysis.
Machine learning methods for DNA sequence analysis. - calico/baskerville
000
David Kelley @drkbio.bsky.social · 17/04/2025
It turns out that this stitching technique works well no matter how large the indel is, so you can go ahead and predict larger scale structural variation, too. For small indels, SVs, and tandem repeats, we demonstrated that Borzoi makes predictions that correspond well to GTEx eQTL results.
100
David Kelley @drkbio.bsky.social · 17/04/2025
This is a little tricky to explain, but we tried to provide intuition and diagrams in the paper, so check it out for more details!
100
David Kelley @drkbio.bsky.social · 17/04/2025
Briefly, we predict the indel effect twice, introducing it to shift the sequence leftward and rightward, and “stitch” the result together in order to minimize shifted boundaries throughout the model.
210
David Kelley @drkbio.bsky.social · 17/04/2025
@anyakors.bsky.social and I came up with several ideas to alleviate the extra variance, and one emerged as the best option for our models.
100