Sign in

Gonzalo Benegas

@gonzalobenegas.bsky.social
294 followers 829 following 25 posts

Research Scientist @ Open Athena | AI for Science gonzalobenegas.github.io

PostsRepliesMedia
Reposted by Gonzalo Benegas
Yun S. Song @yun-s-song.bsky.social · 09/09/2026
We are thrilled to share that our GPN-Star manuscript is now published and freely available: doi.org/10.1038/s415... (1/n)
doi.org
Predicting genome-wide functional constraints with GPN-Star - Nature
GPN-Star, a genomic language model with a phylogeny-aware architecture for whole-genome alignment data, is shown to be a scalable and flexible tool for genetic variant effect prediction across species...
112247
Reposted by Gonzalo Benegas
Open Athena @openathena.ai · 02/09/2026
With the generous support of The Jen-Hsun and Lori Huang Foundation, we have launched the largest live-streamed pretraining run in history, and Marin’s largest model yet: a fully open source 535B total parameter MoE model, with 18T tokens of data. Learn more at: openathena.ai/blog/huang-f...
openathena.ai
Marin's 535 billion parameter model training run launched with support of The Jen-Hsun and Lori Huang Foundation GPU gift
Open Athena announces the launch of Marin's largest training run yet—a 535B parameter large language model—with a generous gift of compute from The Jen-Hsun and Lori Huang Foundation.
0135
Gonzalo Benegas @gonzalobenegas.bsky.social · 06/08/2026
Excited to share MarinDNA, a 1B gLM that rivals Evo 2 40B on variant effect prediction while being 2,330x faster. With @eczech0.bsky.social, we built around a standard Transformer so we could reuse LLM infra and methods while focusing on data curation and scaling. openathena.ai/blog/marin-d... 🧵
1125
Reposted by Gonzalo Benegas
Al Merose @al.merose.com · 30/07/2026
Here's the tale of how @jder.bsky.social and I scaled Samudra, a neural ocean emulator capable of predicting 8 years of the ocean on a single GPU, to operate at a full 1/4° resolution (16x the size in bytes). It was quite a humbling process.
1183
Reposted by Gonzalo Benegas
Joana L. Rocha @joanocha.bsky.social · 08/06/2026
We are excited to share our pan-pangenome paper! This is a long time coming work of my postdoc with the incredible @psudmant.bsky.social, and a stellar group of people! I hope you enjoy reading it! And I am excited to talk more more about it this week at #PEQG26!! www.biorxiv.org/content/10.6...
biorxiv.org
A Pan-pangenome illuminates complex structural variation and selection in humans, chimpanzees, and bonobos
Complete, haplotype-resolved genome assemblies have provided unprecedented insight into the evolution of structurally complex, rapidly evolving regions of human genomes; however, population-scale pang...
15825
Reposted by Gonzalo Benegas
Yun S. Song @yun-s-song.bsky.social · 22/09/2025
We are excited to share GPN-Star, a cost-effective, biologically grounded genomic language modeling framework that achieves state-of-the-art performance across a wide range of variant effect prediction tasks relevant to human genetics. www.biorxiv.org/content/10.1... (1/n)
417690
Reposted by Gonzalo Benegas
Joana L. Rocha @joanocha.bsky.social · 25/06/2025
I am thrilled to announce that in January 2026 I will be starting my own lab at NYU Biology! Soon enough I will be recruiting postdocs and students! Please reach out if you are interested with a CV and description of your research interests, or if you know of people who could be interested! 🧬🗽 🦊
78318
Reposted by Gonzalo Benegas
Yun S. Song @yun-s-song.bsky.social · 23/05/2025
How can one efficiently simulate phylodynamics for populations with billions of individuals, as is typical in many applications, e.g., viral evolution and cancer genomics? In this work with M. Celentano, @wsdewitt.github.io , & S. Prillo, we provide a solution. doi.org/10.1073/pnas... 1/n
13715
Reposted by Gonzalo Benegas
Yun S. Song @yun-s-song.bsky.social · 07/04/2025
Thrilled to see my digital art on the cover of Trends Genet. The two binary strings represent reverse-complementary DNA sequences (00=A, 01=C, 10=G, 11=T) and the connecting rectangles represent “embeddings” learned by DNA language models. Pls check out our article as well: doi.org/10.1016/j.ti...
06913
Reposted by Gonzalo Benegas
Yun S. Song @yun-s-song.bsky.social · 04/03/2025
In our updated TraitGym preprint (w/ @gonzalobenegas.bsky.social & Gökcen Eraslan), we evaluate Evo 2 on regulatory variants associated with human traits. We see marked performance gains with scale on Mendelian traits, although still a bit behind alignment-based methods. doi.org/10.1101/2025... 1/n
13213
Gonzalo Benegas @gonzalobenegas.bsky.social · 13/02/2025
Can DNA sequence models predict mutations affecting human traits? We introduce TraitGym, a curated benchmark of causal regulatory variants for 113 Mendelian & 83 complex traits, and evaluate functional genomics and DNA language models. Joint work w/ Gökcen Eraslan and @yun-s-song.bsky.social 🧵👇
12815
Reposted by Gonzalo Benegas
bioRxiv Genetics @biorxiv-genetic.bsky.social · 13/02/2025
Benchmarking DNA Sequence Models for Causal Regulatory Variant Prediction in Human Genetics www.biorxiv.org/content/10.1101/202…
081
Reposted by Gonzalo Benegas
Yun S. Song @yun-s-song.bsky.social · 08/01/2025
Our work, which shows statistical issues with the previous claim of a severe ancient bottleneck in the ancestry of African populations, has been selected as a Featured article in Genetics. doi.org/10.1093/gene...
doi.org
A previously reported bottleneck in human ancestry 900 kya is likely a statistical artifact
Hu et al. (Science, 2023) recently inferred a severe ancient bottleneck around 900 thousand years (kya) ago in African ancestry but found no similar eviden
0154
Reposted by Gonzalo Benegas
Yun S. Song @yun-s-song.bsky.social · 03/01/2025
Coincidentally, another article from my lab on DNA language models got published on the same day as GPN-MSA. It's freely available for 50 days from this link: authors.elsevier.com/a/1kNCscQbJB... Genomic language models: opportunities and challenges Please share with your colleagues.
authors.elsevier.com
1102
Reposted by Gonzalo Benegas
Yun S. Song @yun-s-song.bsky.social · 02/01/2025
Happy New Year! Our GPN-MSA paper is finally published, under a slightly different title from the preprint. Please check it out and share it with your colleagues: doi.org/10.1038/s415... 1/4
doi.org
A DNA language model based on multispecies alignment predicts the effects of genome-wide variants - Nature Biotechnology
A language model predicts the effects of genetic variants in the human genome.
1167
Reposted by Gonzalo Benegas
Nature Biotechnology @natbiotech.nature.com · 02/01/2025
A DNA language model based on multispecies alignment predicts the effects of genome-wide variants - @yun-s-song.bsky.social go.nature.com/4gWppWg
go.nature.com
A DNA language model based on multispecies alignment predicts the effects of genome-wide variants - Nature Biotechnology
A language model predicts the effects of genetic variants in the human genome.
03113
Reposted by Gonzalo Benegas
Yun S. Song @yun-s-song.bsky.social · 16/11/2024
Large protein language models can learn complex epistatic interactions, but how much does that help with predicting variant effects? In this NeurIPS article, we show that classical independent-sites phylogenetic models can outperform pLMs on this task. 1/7 openreview.net/forum?id=H7m...
openreview.net
Ultrafast classical phylogenetic method beats large protein...
Amino acid substitution rate matrices are fundamental to statistical phylogenetics and evolutionary biology. Estimating them typically requires reconstructed trees for massive amounts of aligned...
29044