Sign in

Burstein lab

@bursteinlab.bsky.social
120 followers 63 following 30 posts

Having fun with microbial genomics and machine learning

PostsRepliesMedia
Burstein lab @bursteinlab.bsky.social · 24/09/2026
After fine-tuning on potentially malicious insertions, the model reached 0.95 AUROC and 0.95 AUPRC. This indicates that such edits do leave detectable genomic-context signatures, without relying on marker genes or predefined databases. Congrats Edan!!! 👏👏👏 4/4
000
Burstein lab @bursteinlab.bsky.social · 24/09/2026
No unified large dataset of engineered bacterial genomes. So we built scalable simulations of random and potentially harmful gene insertions. We trained a classifier on those, along with natural sequences and natural occurrences of the inserted genes as hard negatives. 3/4
100
Burstein lab @bursteinlab.bsky.social · 24/09/2026
Edan used gene-level language models, treating genes as words and genomic regions as sentences. A language model solution made sense since, in natural language, it is easy to recognize a word out of context put in the wrong cranberry 🤖 2/4
110
Burstein lab @bursteinlab.bsky.social · 24/09/2026
So many bacterial genomes are being edited… Could we spot them in the wild even without obvious markers??? Worry not! Our @EdanGabay has you covered. In our new preprint, she pinpoints such genes disrupting the natural genomic "grammar": www.biorxiv.org/content/10.6... 1/4
241
Burstein lab @bursteinlab.bsky.social · 12/07/2026
That's FAMUS: quick, modular, and honest about what it doesn't know. Grab it on the web, through conda, or from source. Huge kudos to 🏆 @guyshur.bsky.social 🏆 , who put the entire thing together. Questions, bugs, or just want to say hi? famus@bursteinlab.org 8/8
000
Burstein lab @bursteinlab.bsky.social · 12/07/2026
Prefer just to run it from your browser? We've got you covered with our web server at app.famus.bursteinlab.org. 7/8
app.famus.bursteinlab.org
FAMUS
Functional Annotation Method Using Supervised contrastive learning
100
Burstein lab @bursteinlab.bsky.social · 12/07/2026
FAMUS is modular by design. Got your own HMM database? We built a preprocessing pipeline to turn it into a FAMUS database too. It is all, of course, open source: github.com/burstein-lab..., also on conda: `conda install -c conda-forge -c bioconda famus` 6/8
github.com
GitHub - burstein-lab/famus
Contribute to burstein-lab/famus development by creating an account on GitHub.
110
Burstein lab @bursteinlab.bsky.social · 12/07/2026
Is it fast? We also built a lightweight version of FAMUS that runs close to the speed of a plain hmmsearch. Both the full and lightweight versions can run on CPU or (a bit faster) on GPU. 5/8
200
Burstein lab @bursteinlab.bsky.social · 12/07/2026
Benchmarking on KEGG and PANTHER, FAMUS improves on KofamScan and InterProScan when most sequences are unannotated, a realistic case for metagenomics and non-model organisms. This led us to compile FAMUS databases from KEGG, InterPro, eggNOG, and OrthoDB. 4/8
100
Burstein lab @bursteinlab.bsky.social · 12/07/2026
Instead of taking a "winner-take-all" approach, FAMUS transforms the similarity scores against the entire HMM database into a condensed vector space, pulling proteins from the same family close together. 3/8
100
Burstein lab @bursteinlab.bsky.social · 12/07/2026
FAMUS uses supervised contrastive learning (SupCon) to explicitly separate known protein families from out-of-scope proteins. SupCon also handles sparsity, giving good classification even for small protein families. 2/8
100
Burstein lab @bursteinlab.bsky.social · 12/07/2026
It annoyed us that annotation tools aren't great at saying "I just don't know". That's a real problem for metagenomes and non-model organisms, where it usually means picking some threshold and hoping for the best. We built FAMUS to try and fix this. doi.org/10.64898/202... 1/8
doi.org
FAMUS: A Few-Shot Learning Framework for Large-Scale Protein Annotation
Predicting gene function is a pivotal and challenging step in genomic and metagenomic data analysis. Current automatic annotation tools typically rely on the single most similar sequence from the quer...
1124
Burstein lab @bursteinlab.bsky.social · 09/07/2026
👇 doi.org/10.1093/bioi...
doi.org
Optimizing protein tokenization: reduced amino acid alphabets for efficient and accurate protein language models
AbstractMotivation. Protein language models (pLMs) typically tokenize sequences at the single-amino-acid level using a 20-residue alphabet, resulting in lo
000
Burstein lab @bursteinlab.bsky.social · 09/07/2026
Now available in bioinformatics! Congrats Ella 🎉
140
Reposted by Burstein lab
Burstein lab @bursteinlab.bsky.social · 16/02/2026
🧬 New preprint alert! Protein language models have transformed biology - but what about the tokens they read? In our new preprint, 👑@EllaRannon👑 studies how tokenization choices shape pLM performance and efficiency. 🧵 (1/5) www.biorxiv.org/content/10.6...
biorxiv.org
192
Reposted by Burstein lab
Modell lab @jwmodell.bsky.social · 27/03/2026
How do you kill a MRSA superbug armed with 15 different anti-phage defense systems? You make a smarter phage. Check out our latest preprint on overcoming bacterial immunity using defense-guided engineering to build durable therapeutic phage cocktails! Led by Sarah Voss. doi.org/10.64898/202...
38143
Burstein lab @bursteinlab.bsky.social · 16/02/2026
🚀 Our results highlight a promising direction for making protein language models more efficient and scalable. Read all about it! www.biorxiv.org/content/10.6... 5/5
biorxiv.org
000
Burstein lab @bursteinlab.bsky.social · 16/02/2026
⚡ Reduced alphabets yield shorter inputs and major runtime gains, while maintaining comparable, and sometimes improved, predictive performance. (4/5)
100
Burstein lab @bursteinlab.bsky.social · 16/02/2026
🔤 By combining Byte Pair Encoding (BPE) with reduced amino acid alphabets based on residue properties, we train new pLMs and evaluate them across diverse biological tasks, like solubility, enzyme, PPI, and stability prediction. (3/5)
120
Burstein lab @bursteinlab.bsky.social · 16/02/2026
⚖️ Unlike natural languages, proteins aren’t clearly separated into “words”, making tokenization tricky. Short tokens create long sentences, while long tokens lead to a sparse vocabulary that is hard to learn. But reducing the alphabet size might help! (2/5)
110
Burstein lab @bursteinlab.bsky.social · 16/02/2026
🧬 New preprint alert! Protein language models have transformed biology - but what about the tokens they read? In our new preprint, 👑@EllaRannon👑 studies how tokenization choices shape pLM performance and efficiency. 🧵 (1/5) www.biorxiv.org/content/10.6...
biorxiv.org
192
Burstein lab @bursteinlab.bsky.social · 07/01/2026
4/4 To allow standardized benchmarking for future tool development, we also released B-PPI-DB, a curated bacterial PPI database derived from STRING: doi.org/10.5281/zeno... We hope B-PPI is just the first of many efficient bacterial PPI predictors!
doi.org
B-PPI-DB: A Benchmarking Dataset for Bacterial Protein-Protein Interactions
B-PPI-DB is a benchmarking dataset designed for the training and evaluation of bacterial protein-protein interaction (PPI) prediction models. The dataset is derived from the STRING database (version 1...
000
Burstein lab @bursteinlab.bsky.social · 07/01/2026
3/4 B-PPI outperforms other rapid methods for bacterial PPI prediction without the high cost of structural folding and generalizes to unseen interactions with minimal fine-tuning.
100
Burstein lab @bursteinlab.bsky.social · 07/01/2026
2/4 B-PPI is a cross-attention model for bacterial PPI prediction at scale. Given protein pairs, it leverages ProstT5, a structure-aware protein language model, to generate embeddings, and outputs the interaction probability.
100
Burstein lab @bursteinlab.bsky.social · 07/01/2026
1/4 Ever wanted to predict bacterial protein-protein interactions (PPI) on a large scale? We wanted to, but realized there’s no such algorithm that is both rapid and optimized for bacterial protein analysis. This led our ⭐️Chen Agassy⭐️ to develop B-PPI: doi.org/10.64898/202...
153
Burstein lab @bursteinlab.bsky.social · 07/06/2025
6/6 🔮 What's next for NLP in biology? We discuss future directions as well. Join us in exploring the future of this exciting field! arxiv.org/abs/2506.02212
arxiv.org
Leveraging Natural Language Processing to Unravel the Mystery of Life: A Review of NLP Approaches in Genomics, Transcriptomics, and Proteomics
Natural Language Processing (NLP) has transformed various fields beyond linguistics by applying techniques originally developed for human language to the analysis of biological sequences. This review ...
010
Burstein lab @bursteinlab.bsky.social · 07/06/2025
5/6 💡 Discover how NLP is being applied to: • Protein structure prediction 🏗️ • Taxonomic classification 🌳 • Mutational effect prediction 🔀 • Gene expression prediction 📈 And much more!
120
Burstein lab @bursteinlab.bsky.social · 07/06/2025
4/6 🧩 Tokenization challenges? We've got that covered too! Explore different approaches to breaking down biological sequences and their impact on model performance.
110
Burstein lab @bursteinlab.bsky.social · 07/06/2025
3/6 📚 We break down the evolution of NLP models in biology, from classic word2vec to cutting-edge transformers and hyena operators. Understand their strengths, limitations, and exciting applications!
100
Burstein lab @bursteinlab.bsky.social · 07/06/2025
2/6 🔬 We dive deep into how NLP techniques are revolutionizing the analysis of biological 'languages': • DNA 🧬 • RNA 🧬 • Proteins 💪 • Entire genomes 🔍 Learn how these methods are unlocking new insights in genomics!
110
Burstein lab @bursteinlab.bsky.social · 07/06/2025
1/6 🧬📊 Curious about the buzz around NLP in biology? Feeling overwhelmed by the rapid developments? We've got you covered! Our review on NLP applications in genomics, by the wonderful Ella Rannon, is now out as a pre-print! #NLP #Bioinformatics arxiv.org/abs/2506.02212
14715
Burstein lab @bursteinlab.bsky.social · 07/06/2025
Hello Bluesky! Burstein lab just joined - are we too late for the party 👀 ??
1162