Sign in

Tristan Bepler

@tbepler.bsky.social
414 followers 394 following 34 posts

Scientist and Group Leader of the Simons Machine Learning Center @SEMC_NYSBC. Co-founder and CEO of OpenProtein.AI. Opinions are my own.

PostsRepliesMedia
Reposted by Tristan Bepler
openprotein.bsky.social @openprotein.bsky.social · 26/08/2026
Antibody discovery workflow, streamlined: CDR annotation, germline calls, and liability flags, clustering by sequence similarity, and developability and activity scoring, all in one dataset view. Sign up link in the comments!
121
Reposted by Tristan Bepler
openprotein.bsky.social @openprotein.bsky.social · 23/06/2026
New on OpenProtein.AI: esmfold2, esmfold2-fast, esmc-300m/600m/6b, esm-if1, and Protenix-v2, plus full multichain support across all workflows. Sign up link in the comments!
121
Reposted by Tristan Bepler
openprotein.bsky.social @openprotein.bsky.social · 31/03/2026
We’re excited to announce our expanded partnership with Boehringer Ingelheim. Together, we are building the future of AI‑driven antibody discovery and optimization. www.openprotein.ai/strategic-partnership-with-boehringer-ingelheim
011
Tristan Bepler @tbepler.bsky.social · 10/02/2026
🤩
130
Tristan Bepler @tbepler.bsky.social · 16/12/2025
New job openings @openprotein.bsky.social across protein foundation model research, computational protein design, and cloud platform engineering www.openprotein.ai/careers
openprotein.ai
Careers
At OpenProtein.AI, we are building tools to democratize protein engineering. Join us.
020
Tristan Bepler @tbepler.bsky.social · 26/08/2025
Our preprint on sequence-to-property learning and zero-shot fitness prediction with PoET-2 is live: arxiv.org/abs/2508.04724 PoET-2 is also open sourced on github: github.com/OpenProteinA... Thanks to the @openprotein.bsky.social team!
arxiv.org
Understanding protein function with a multimodal retrieval-augmented foundation model
Protein language models (PLMs) learn probability distributions over natural protein sequences. By learning from hundreds of millions of natural protein sequences, protein understanding and design capa...
150
Reposted by Tristan Bepler
openprotein.bsky.social @openprotein.bsky.social · 25/07/2025
Boltz-1 & Boltz-2 now live via GUI & APIs! Predict protein, protein–RNA/DNA/ligand structures with confidence scores & binding affinity metrics for virtual screening. Compare finetuned models in the new overview page to find your best performer fast. www.openprotein.ai/early-access...
072
Reposted by Tristan Bepler
openprotein.bsky.social @openprotein.bsky.social · 25/06/2025
Product update: Indel Analysis lets you score insertions/deletions across your sequence using PoET-2. You can now also compare multiple 3D structures in Mol* to evaluate design alternatives. Sign up now: www.openprotein.ai/early-access...
011
Tristan Bepler @tbepler.bsky.social · 20/06/2025
How would you use a tool like this? Do you design or screen indels in your work? 4/4
000
Tristan Bepler @tbepler.bsky.social · 20/06/2025
Indels are still a major challenge for variant effect prediction and protein design. PoET-2 has significantly improved the state-of-the-art for functional and clinical indel variant effect prediction. 3/4
PoET-2 is state-of-the-art for zero-shot variant effect prediction of both DMS indels and clinical indels in ProteinGym.
100
Tristan Bepler @tbepler.bsky.social · 20/06/2025
It supports screening deletions, insertion sites, and replacement sites. Explore viable shortened proteins, or insert new structural or functional sequences like localization signals or structural tags. 2/4
100
Tristan Bepler @tbepler.bsky.social · 20/06/2025
Why does no one in AI protein engineering work on indels? We’re solving this at OpenProtein.AI. Check out our upcoming indel design tool! 🤩 1/4 @openprotein.bsky.social
141
Reposted by Tristan Bepler
Pascal Notin @pascalnotin.bsky.social · 08/05/2025
Have we hit a "scaling wall" for protein language models? 🤔 Our latest ProteinGym v1.3 release suggests that for zero-shot fitness prediction, simply making pLMs bigger isn't better beyond 1-4B parameters. The winning strategy? Combining MSAs & structure in multimodal models!
1247
Tristan Bepler @tbepler.bsky.social · 14/05/2025
Great to see this comparison with genome language models. The hype around these models seems to have strongly outstripped where they actually are in comparison with protein models.
000
Reposted by Tristan Bepler
openprotein.bsky.social @openprotein.bsky.social · 13/05/2025
Product update: PoET-2 now supports structure inputs for enhanced prediction and design via Python APIs. Check out our new inverse folding tutorial to see it in action. 🔗 docs.openprotein.ai/walkthroughs... Sign up for OpenProtein.AI: www.openprotein.ai/early-access...
docs.openprotein.ai
Inverse Folding with PoET-2 for Generation of Novel Luciferases — OpenProtein-Docs documentation
011
Tristan Bepler @tbepler.bsky.social · 12/05/2025
Huge thanks to the @openprotein.bsky.social team! We've got more exciting PoET-2 updates to come 🚀
000
Tristan Bepler @tbepler.bsky.social · 12/05/2025
Learn more about PoET-2 in our whitepaper: www.openprotein.ai/a-multimodal...
openprotein.ai
A multimodal foundation model for controllable protein generation and representation learning
PoET-2 is a next generation protein language model that transforms our ability to engineer proteins by learning from nature’s design principles. Through its unique multimodal architecture and training...
110
Tristan Bepler @tbepler.bsky.social · 12/05/2025
Sign up for OpenProtein.AI (free for academic use): www.openprotein.ai/early-access... and install the python client to get started: github.com/OpenProteinA...
openprotein.ai
Sign up for early access | OpenProtein.AI
Join the revolution in protein research with early access to our cutting-edge Open Protein AI platform. Sign up now to explore the future of protein analysis and discovery.
100
Tristan Bepler @tbepler.bsky.social · 12/05/2025
Generative protein sequence design, variant effect prediction, and fine-tuning are now fully supported for PoET-2 with structure and sequence prompts in the @openprotein.bsky.social python client and APIs! Check out our new walkthrough on inverse folding: docs.openprotein.ai/walkthroughs...
docs.openprotein.ai
Inverse Folding with PoET-2 for Generation of Novel Luciferases — OpenProtein-Docs documentation
242
Reposted by Tristan Bepler
SynBioBeta @synbiobeta.bsky.social · 21/02/2025
🧬 Protein Revolution: The Tiny Model Making a Massive Impact! PoET-2 is changing the game in computational protein design, slashing experimental data needs by 30x! 🚀 learn more: www.synbiobeta.com/read/protein... #ProteinDesign #BiotechInnovation #AIRevolution
synbiobeta.com
Protein Revolution: The Tiny Model Making a Massive Impact - SynBioBeta
012
Tristan Bepler @tbepler.bsky.social · 11/02/2025
Huge thanks to our incredible team @openprotein.bsky.social, especially Tim Truong. This is just the beginning of AI systems that truly understand protein biology. I can’t wait to see what the community can do with these models! 13/13
020
Tristan Bepler @tbepler.bsky.social · 11/02/2025
Read more in our white paper: www.openprotein.ai/a-multimodal... 12/13
openprotein.ai
A multimodal foundation model for controllable protein generation and representation learning
PoET-2 is a next generation protein language model that transforms our ability to engineer proteins by learning from nature’s design principles. Through its unique multimodal architecture and training...
121
Tristan Bepler @tbepler.bsky.social · 11/02/2025
This has huge implications for protein engineering - from more efficient directed evolution and multiproperty optimization to de novo protein design. 11/13
110
Tristan Bepler @tbepler.bsky.social · 11/02/2025
Most importantly, PoET-2 gets us closer to understanding the sequence-structure-function relationship - learning from just a handful of examples to predict properties and design new sequences. 10/13
110
Tristan Bepler @tbepler.bsky.social · 11/02/2025
Beyond predictions, PoET-2 introduces a powerful prompt grammar for protein generation. One model for: free sequence generation, inverse folding, motif scaffolding, and more! 9/13
110
Tristan Bepler @tbepler.bsky.social · 11/02/2025
The results show PoET-2 has learned fundamental principles: * Improves sequence and structure understanding * Accurate zero-shot function prediction, especially for insertions and deletions * 30x less data needed for transfer learning 8/13
110
Tristan Bepler @tbepler.bsky.social · 11/02/2025
This lets us break conventional scaling laws. PoET-2 achieves with 182M parameters what would require trillion-parameter models using standard architectures. 7/13
110
Tristan Bepler @tbepler.bsky.social · 11/02/2025
Enabled by our tiered attention structure, PoET-2 processes sequence families with order equivariance while preserving long-range dependencies within and between sequences, enabling processing of large sequence families with optional structural data. 6/13
110
Tristan Bepler @tbepler.bsky.social · 11/02/2025
Rather than regurgitating databases, PoET-2 "meta-learns" evolutionary principles through in-context learning - inferring structural and functional constraints at inference time from small numbers of examples. 5/13
110
Tristan Bepler @tbepler.bsky.social · 11/02/2025
PoET-2 takes a different approach. Instead of massive scale, we developed a multimodal architecture that learns to reason about sequences, structures, and evolutionary relationships simultaneously. 4/13
131
Tristan Bepler @tbepler.bsky.social · 11/02/2025
But memorizing sequences isn't enough. The real challenge: can we build models that learn the fundamental principles that govern how proteins evolve and function? 3/13
110
Tristan Bepler @tbepler.bsky.social · 11/02/2025
Since our first protein language models in 2019 (in Bonnie's lab!), the field has focused on scale - building ever-larger models up to 100B parameters to extract information from natural sequence databases. 2/13
110
Tristan Bepler @tbepler.bsky.social · 11/02/2025
Excited to share PoET-2, our next breakthrough in protein language modeling. It represents a fundamental shift in how AI learns from evolutionary sequences. 🧵 1/13
13215
Reposted by Tristan Bepler
openprotein.bsky.social @openprotein.bsky.social · 11/02/2025
🧬 Announcing PoET-2: A breakthrough protein language model that achieves trillion-parameter performance with just 182M parameters, transforming our ability to understand proteins.
142
Reposted by Tristan Bepler
Etowah Adams @etowah0.bsky.social · 10/02/2025
Can we learn protein biology from a language model? In new work led by @liambai.bsky.social and me, we explore how sparse autoencoders can help us understand biology—going from mechanistic interpretability to mechanistic biology.
24524
Reposted by Tristan Bepler
Etowah Adams @etowah0.bsky.social · 10/02/2025
We’re excited about the potential of SAEs in biology and would love to hear your ideas. Our preprint: www.biorxiv.org/content/10.1... Visualizer: interprot.com Github: github.com/etowahadams/... HuggingFace: huggingface.co/liambai/Inte...
biorxiv.org
From Mechanistic Interpretability to Mechanistic Biology: Training, Evaluating, and Interpreting Sparse Autoencoders on Protein Language Models
Protein language models (pLMs) are powerful predictors of protein structure and function, learning through unsupervised training on millions of protein sequences. pLMs are thought to capture common mo...
121
Tristan Bepler @tbepler.bsky.social · 21/12/2024
TARDIS is open source on github (github.com/SMLC-NYSBC/T...) and can also be installed via pip
github.com
GitHub - SMLC-NYSBC/TARDIS: Transformer And Rapid Dimensionless Instance Segmentation
Transformer And Rapid Dimensionless Instance Segmentation - SMLC-NYSBC/TARDIS
110
Tristan Bepler @tbepler.bsky.social · 21/12/2024
We applied TARDIS to annotate about 13,000 tomograms in the CZII data portal @chanzuckerberg.bsky.social cryoetdataportal.czscience.com
cryoetdataportal.czscience.com
CryoET Data Portal
130
Tristan Bepler @tbepler.bsky.social · 21/12/2024
As a flexible framework, new semantic segmenation networks can be plugged in for new biomolecules and imaging modalities, which we show in application to actin and TIRF microscopy.
100
Tristan Bepler @tbepler.bsky.social · 21/12/2024
We provide pre-trained networks for microtubule and membrane semantic segmenation and instance segmentation of general surface and linear-like structures.
110
Tristan Bepler @tbepler.bsky.social · 21/12/2024
A flexible framework for fast and accurate segmentation of filaments and membranes in tomograms and micrographs - the TARDIS manuscript is now live @biorxivpreprint.bsky.social ! Thanks to hard work by Robert Kiewisz and our many collaborators! www.biorxiv.org/content/10.1...
biorxiv.org
Accurate and fast segmentation of filaments and membranes in micrographs and tomograms with TARDIS
It is now possible to generate large volumes of high-quality images of biomolecules at near-atomic resolution and in near-native states using cryogenic electron microscopy/electron tomography (Cryo-EM...
2177
Tristan Bepler @tbepler.bsky.social · 01/12/2024
I'll be talking about shrinking protein language models with #PoET and protein engineering at openprotein.ai tomorrow at A*STAR's Bioinformatics Institute. If you can't make it, I'll also be presenting at the Berger Lab seminar @mitofficial.bsky.social on Wednesday!
030
Reposted by Tristan Bepler
Diego del Alamo @delalamo.xyz · 25/11/2024
Yet more evidence that transfer learning of sequence-only PLMs does not benefit from scale beyond 650M params 🧵
Fig. 3. Impact of ESM Model size on transfer learning. LassoCV regression results using ESM mean embeddings for 35 DMS datasets. The x-axis represents different ESM model sizes: ESM1v 650M (orange) and ESM2 8M, 35M, 150M, 650M, 3B, and 15B (blues). The y-axis displays R^2 for each task.
77820
Tristan Bepler @tbepler.bsky.social · 25/11/2024
It is interesting that scaling to billions of parameters only really seems to help with structure prediction and that this does not translate to transfer learning for property/function prediction
000
Tristan Bepler @tbepler.bsky.social · 25/11/2024
It also appears to be true for unsupervised variant effect prediction that increasing model size hurts performance
100
Tristan Bepler @tbepler.bsky.social · 25/11/2024
We also saw this when we looked at transfer learning with PoET embeddings compared with ESM - www.openprotein.ai/poet-foundat...
110