Sign in

Christian Dallago

@machine.learning.bio
1.3K followers 305 following 61 posts

🏳️‍🌈 NVIDIA & Duke. Was Allianz, VantAI, TUM. BioCS+ML dude. Lab page: machine.learning.bio GScholar: scholar.google.com/citations?user=4…

PostsRepliesMedia
Christian Dallago @machine.learning.bio · 24/09/2026
NVIDIA blog: blogs.nvidia.com/blog/open-pr... EMBL-EBI: embl.org/news/science... Featured in Nature: nature.com/articles/d41... Explore the research and build on it: Paper: research.nvidia.com/labs/dbr Pipeline: github.com/NVIDIA-BioNe...
010
Christian Dallago @machine.learning.bio · 24/09/2026
A special shoutout to Riya Narain, who spent the summer interning on my team and did so much of the work behind this release. I’m proud of what she’s contributed.
110
Christian Dallago @machine.learning.bio · 24/09/2026
A huge team effort with Google DeepMind (Risha Patel), EMBL-EBI (Jennifer Fleming, Sameer Velankar) and SNU (Yewon Han, Martin Steinegger), University of Glasgow (Joe Grove), SIB (Philippe Le Mercier), CEPI (Newton Wahome) and Sungkyunkwan University (Milot Mirdita).
110
Christian Dallago @machine.learning.bio · 24/09/2026
From viral assemblies to how proteases recognise their substrates: starting points for vaccine and antiviral experiments. The predictions are open through AFDB, so researchers can use them without repeating the computation: alphafold.ebi.ac.uk
120
Christian Dallago @machine.learning.bio · 24/09/2026
Building on the AlphaFold Database protein-complex release at GTC, we explored ~1.7M protein pairs across 2,800+ viral proteomes. That scale lets us compare how viral proteins assemble and identify conserved interfaces.
110
Christian Dallago @machine.learning.bio · 24/09/2026
Today I’m joining collaborators at a WEF/CEPI roundtable in New York to share our latest @nvidiabot.bsky.social AI for Life Sciences work. Featured in Nature: nature.com/articles/d41...
100
Reposted by Christian Dallago
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 17/03/2026
AlphaFold database has entered the era of complexes. Together with NVIDIA, DeepMind and EBI, we use ColabFold, OpenFold and MMseqs2-GPU to predict ~31 million complexes (homo & hetro-dimers) resulting in 1.8 million high-quality predictions 📄 research.nvidia.com/labs/dbr/ass... 🌐 alphafold.ebi.ac.uk
8263110
Reposted by Christian Dallago
EMBL-EBI @ebi.embl.org · 16/03/2026
You asked, we listened. Millions of AI-predicted protein complex structures are now available in the #AlphaFold Database. This spans homodimers from 20 of the most studied species, including humans, as well as the World Health Organization’s priority pathogens list. www.ebi.ac.uk/about/news/t...
816087
Christian Dallago @machine.learning.bio · 26/02/2026
Find out more: flip.protein.properties
flip.protein.properties
FLIP2: Expanding Protein Fitness Landscape Benchmarks
FLIP2: A comprehensive benchmark for protein fitness prediction with 7 datasets, 16 splits, and real-world engineering scenarios
021
Christian Dallago @machine.learning.bio · 26/02/2026
I'm especially happy about continuing to work with an amazing group of scientists. Thanks @kevinkaichuang.bsky.social , @kdidi.bsky.social , Bruce Wittmann, @kadinaj.bsky.social, Maya Czeneszew, @sarahalamdari.bsky.social, @alexijie.bsky.social, @thisismadani.bsky.social, ++
110
Christian Dallago @machine.learning.bio · 26/02/2026
I love this gritty work here; there are no new architectures, no leaderboard-topping number to screenshot. However, it's how we (and hopefully the community) can measure whether the models we're all building and using are getting better where it counts — at the bench.
110
Christian Dallago @machine.learning.bio · 26/02/2026
That's not a criticism to any method. New benchmarks are precisely needed to see where we are at, and set some target of where we could go from here.
110
Christian Dallago @machine.learning.bio · 26/02/2026
Especially on the wild-type and position splits, current transfer learning doesn't consistently win. No single pLM architecture dominates. Scaling hasn't closed the gap yet.
110
Christian Dallago @machine.learning.bio · 26/02/2026
The answer in 2026 is largely the same. Simple ridge regression on one-hot sequences, optionally supplemented with zero-shot pLM likelihoods, often matches or outperforms fine-tuned protein language models.
121
Christian Dallago @machine.learning.bio · 26/02/2026
FLIP2, adds seven new sequence-fitness landscapes - industrial enzymes, nucleases, rhodopsins, protein-protein interactions - and 16 splits that test the generalization axes protein engineers really hit: more mutations, new positions, higher fitness, different wild-types.
110
Christian Dallago @machine.learning.bio · 26/02/2026
We were interested in how things had changed 5 years after our first release. So, we built FLIP2 on select datasets from great labs across the world, many of which have gracefully agreed to make their data freely available.
110
Christian Dallago @machine.learning.bio · 26/02/2026
FLIP spawned fast development of several different benchmarking efforts across protein design, engineering, and variant effect assesment. The answer in 2021 was: sometimes, but simpler models hold up surprisingly well.
110
Christian Dallago @machine.learning.bio · 26/02/2026
Five years ago, we released FLIP. The core question was: can ML models for protein fitness prediction generalize in the ways that actually matter for protein engineering, i.e. low data, extrapolation to more mutations, out-of-distribution sequences?
245
Reposted by Christian Dallago
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 25/02/2026
We made FLIP2, a protein fitness benchmark spanning seven new datasets, including enzymes, protein-protein interactions, and light-sensitive proteins, as well as splits that measure generalization relevant to real-world protein engineering campaigns.
15215
Christian Dallago @machine.learning.bio · 22/12/2025
You can use the model right now to freely generate families for single sequence inputs (i.e., diversification conditioned by intrinsic representations of evolution), or to engineer proteins based on family promts (diversification by conditioning on particular evoluationary trajectories).
011
Christian Dallago @machine.learning.bio · 22/12/2025
In essence, we probed the model's ability to ricapitulate family statistics, bootstrap protein structure prediction, and assess mutation effect, demonstrating excellent performance across all tasks, especially using test-time-scaling via prompt conditioning.
100
Christian Dallago @machine.learning.bio · 22/12/2025
With ProFam-1, we scaled learning from single sequence to protein family definitions of different kinds, curating a large protein family corpus, ProtFam-atlas. I'm particularily stoked about the idea of inference-time-compute. This contribution laid out a very exciting path for future work.
100
Christian Dallago @machine.learning.bio · 22/12/2025
Our latest protein family-based GenAI collection of tools and datasets, ProFam, is out now. Everything -- from data, training and inference code, to a 215M llama-based ProFam-1 are fully open sourced. 🧵
251
Christian Dallago @machine.learning.bio · 27/10/2025
Another exciting opportunity, this time as a colleague at Duke! Join as tenure track assistant prof. in Cell Bio & let’s work on closing the gap between in-silico and in-vivo: www.nature.com/naturecareer... Important: application closes Nov 1st!!!
nature.com
Tenure-Track Assistant Professor Position –AI/ML for Cell Biology - Durham, North Carolina (US) job with Duke University School of Medicine | 12844591
Tenure-Track Assistant Professor Position –AI/ML for Cell Biology
031
Christian Dallago @machine.learning.bio · 17/10/2025
Another opening: Senior Multiscale Biology Applied Research Scientist! nvidia.eightfold.ai/careers/job/... Are fascinated by fundamental data modalities across biology like RNA-seq, mass spec & want to build computational tools that harnessing data to build intelligence? Come: join the team!
nvidia.eightfold.ai
Senior Applied Research Scientist, Multiscale Biology | NVIDIA Corporation
Apply your expertise in engineering biology through algorithms and tools for genes, tissues, organisms, and populations. Conduct collaborative applied research in multiscale biology using deep learnin...
041
Christian Dallago @machine.learning.bio · 15/10/2025
I don’t dare question the HR gods about their designs :)
100
Christian Dallago @machine.learning.bio · 15/10/2025
Are you passionate about leading collaborative, fast moving, applied bioinformatics research projects that help the entire community move forward? Apply to work in my team at NVIDIA: nvidia.eightfold.ai/careers/job/...
nvidia.eightfold.ai
Senior Applied Research Scientist, Bioinformatics | NVIDIA Corporation
Lead applied and collaborative research programs using bioinformatics, high performance computing, and deep learning for biological advancements. Develop and accelerate bioinformatics software and alg...
150
Christian Dallago @machine.learning.bio · 15/10/2025
I should add: structure prediction inference is INCREDIBLY efficient for the form factor and power profile.
000
Christian Dallago @machine.learning.bio · 15/10/2025
It's an inference monster. Structure prediction on it works. Design to come. This will be updated later today... research.nvidia.com/labs/dbr/ass...
research.nvidia.com
130
Reposted by Christian Dallago
Stephan Hacker @stephanhacker2.bsky.social · 30/09/2025
Great talk by @machine.learning.bio at the 5th Virtual @chembiotalks.bsky.social. He talked about the use of #MachineLearning and #BigData to address biological questions. Cool insights into both predicting functions and designing proteins ieeexplore.ieee.org/document/947... arxiv.org/abs/2503.00710
ieeexplore.ieee.org
ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning
Computational biology and bioinformatics provide vast data gold-mines from protein sequences, ideal for Language Models (LMs) taken from Natural Language Processing (NLP). These LMs reach for new pred...
062
Reposted by Christian Dallago
Nature Methods @natmethods.nature.com · 18/09/2025
GPU-accelerated MMseqs2 offers tremendous speedup for homology retrieval, protein structure prediction with ColabFold, and protein structure search with Foldseek. @martinsteinegger.bsky.social @milot.bsky.social @machine.learning.bio www.nature.com/articles/s41...
nature.com
GPU-accelerated homology search with MMseqs2 - Nature Methods
Graphics processing unit-accelerated MMseqs2 offers tremendous speedups for homology retrieval from metagenomic databases, query-centered multiple sequence alignment generation for structure predictio...
08221
Reposted by Christian Dallago
Arne Elofsson @handle.invalid · 11/09/2025
podcasts.apple.com/us/podcast/f...
podcasts.apple.com
From AlphaFold to MMseqs2-GPU: How AI is Accelerating Protein Science
Podcast Episode · NVIDIA AI Podcast · 09/10/2025 · 35m
0103
Reposted by Christian Dallago
Stephan Hacker @stephanhacker2.bsky.social · 16/07/2025
Looking forward to hearing about the potential of machine learning for #Biology and #DrugDiscovery from an industry perspective. Register for the Virtual @chembiotalks.bsky.social to hear the perspective of Chris Dallago (@machine.learning.bio) from Nvidia. #ChemBio #Chemsky #ML #MachineLearning
0105
Christian Dallago @machine.learning.bio · 17/07/2025
I still feel criminal for the handle but the manuscript embodies it well.
020
Christian Dallago @machine.learning.bio · 15/07/2025
Moore’s law applied to speed not accuracy. I don’t think fundamentally the discoveries we are after are entirely dependent on speed. I think the better law here is garbage in garbage out. In that sense, you can wait for better data/curation, but it’s also fun to take destiny in your own hands :)
120
Reposted by Christian Dallago
César de la Fuente @delafuentelab.bsky.social · 14/07/2025
(1/5) Venoms are a vast, largely untapped library of bioactive molecules—and our new paper in @natcomms.nature.com ‬ @natprot.nature.com reveals just how powerful they can be. 🐍⚡️
nature.com
Computational exploration of global venoms for antimicrobial discovery with Venomics artificial intelligence - Nature Communications
Researchers used artificial intelligence to mine global venom proteomes and discovered novel peptides with antimicrobial activity. Several candidates showed efficacy against drug-resistant bacteria in...
143
Reposted by Christian Dallago
César de la Fuente @delafuentelab.bsky.social · 11/07/2025
Excited to have participated in the 2025 Symposium on Generative AI in Molecule Discovery in beautiful Munich, along with amazing scientists and colleagues @machine.learning.bio, Francesca Grisoni, ‪@ewaszczurek.bsky.social‬‬, @fabiantheis.bsky.social and more... 🔬🤖 events.hifis.net/event/2015/
031
Reposted by Christian Dallago
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 07/07/2025
Folddisco finds similar (dis)continuous 3D motifs in large protein structure databases. Its efficient index enables fast uncharacterized active site annotation, protein conformational state analysis and PPI interface comparison. 1/9🧶🧬 📄 www.biorxiv.org/content/10.1... 🌐 search.foldseek.com/folddisco
816271
Christian Dallago @machine.learning.bio · 15/06/2025
Thanks to @pkoo562.bsky.social , @kevinkaichuang.bsky.social and Ananthan, and all of CSHL' staff for the help bringing this volume out. Hopefully, we could inspire all those fantastic biologists to reach out to their computational friends to take the next leap in their scientific discoveries.
030
Christian Dallago @machine.learning.bio · 15/06/2025
With contributions from fantastic colleagues @martinsteinegger.bsky.social , @mikeinouye.bsky.social, @jlistgarten.bsky.social , @ideasbyjin.bsky.social, @michael-heinzinger.bsky.social, and many more, the first CSHL volume on ML for Protein Science and Engineering is out: lnkd.in/dQdgGPpp
1154
Christian Dallago @machine.learning.bio · 15/06/2025
If only there was something like try/catch in every programming/scripting language ever :P
010
Christian Dallago @machine.learning.bio · 15/06/2025
The maddening bit is UniProt offers a handy URL that resolves whatever query you pass to it to a list of hits www.uniprot.org/help/api_que... Whenever I used to implement search bars I’d use my dbs as autocomplete but one hop to that API endpoint to catch whatever the user inputs 🤷‍♂️
210
Reposted by Christian Dallago
Stephan Hacker @stephanhacker2.bsky.social · 10/06/2025
The final program is now online for the 5th Virtual @chembiotalks.bsky.social: web.cvent.com/event/60e9f3... Looking forward to talks by Sarah O'Connor, Chengqi Yi, @machine.learning.bio, @cathleenzeymer.bsky.social, @kellychibale.bsky.social, Jennifer Prescher and @craigmcrews.bsky.social.
022
Christian Dallago @machine.learning.bio · 03/06/2025
@michael-heinzinger.bsky.social and I are seeking talented postdocs to support for the Marie Skłodowska-Curie Fellowship! Join our international AI+biology team, collaborate on protein design, and access top labs in the US & EU. Interested? Apply by July 15! Details: machine.learning.bio/news/msca
machine.learning.bio
Marie Skłodowska-Curie Fellowship
PostDoc Opportunity in AI for Protein Science Marie Skłodowska-Curie Fellowship
031
Reposted by Christian Dallago
Ewa Szczurek @ewaszczurek.bsky.social · 07/04/2025
Save the date for the Helmholtz Munich AI for Health Symposium 2025, devoted to the topic of Generative AI in Molecule Discovery! July 4, 2025 Helmholtz Munich Campus in Neuherberg, Germany What now? Register and submit abstracts (deadline May 2, 2025!) events.hifis.net/event/2015/r...
174
Reposted by Christian Dallago
Karsten Kreis @karstenkreis.bsky.social · 04/03/2025
📢📢 "Proteina: Scaling Flow-based Protein Structure Generative Models" #ICLR2025 (Oral Presentation) 🔥 Project page: research.nvidia.com/labs/genair/... 📜 Paper: arxiv.org/abs/2503.00710 🛠️ Code and weights: github.com/NVIDIA-Digit... 🧵Details in thread... (1/n)
13910
Reposted by Christian Dallago
Karsten Kreis @karstenkreis.bsky.social · 04/03/2025
🔸Proteina is a fantastic collaboration with wonderful colleagues at NVIDIA: 🔥 Tomas Geffner*, @kdidi.bsky.social*, Zuobai Zhang*, Danny Reidenbach, Zhonglin Cao, @jyim.bsky.social , Mario Geiger, @machine.learning.bio, Emine Kucukbenli, @arashv.bsky.social, @karstenkreis.bsky.social* 🔥 (10/n)
171
Christian Dallago @machine.learning.bio · 12/01/2025
Big time
030
Christian Dallago @machine.learning.bio · 08/01/2025
Whelp
030