Sign in

Tatta Bio

@tattabio.bsky.social
208 followers 908 following 117 posts

Building genomic intelligence Metagenomic datasets, genomic language models, SeqHub Our research: tatta.bio Analyze your sequences on seqhub.org

PostsRepliesMedia
Tatta Bio @tattabio.bsky.social · 23/09/2026
And sustainable chemistry. Redesigning biosynthetic assembly lines to produce materials and commodity chemicals from biology rather than petrochemical feedstocks could make chemical manufacturing more sustainable.
000
Tatta Bio @tattabio.bsky.social · 23/09/2026
This could enable: New therapeutics. More than half of FDA-approved small-molecule drugs over the past 4 decades are natural products, their derivatives, or synthetic mimics. Redirecting biosynthetic assembly lines toward new molecular targets could substantially expand accessible chemical space.
100
Tatta Bio @tattabio.bsky.social · 23/09/2026
In our preprint released last Friday, we show that evolutionary sequence information learned by our genomic language model, gLM2, can be used to design complex, multi-domain enzymes and expand their chemistry beyond their natural repertoire. 🧵 Preprint: www.biorxiv.org/content/10.6...
biorxiv.org
Generative Design of New-to-nature Biosynthetic Assembly Lines with Genomic Language Modeling
Reprogramming biosynthetic assembly lines can extend biosynthesis beyond the chemical space explored by nature. However, this remains difficult because assembly-line function depends on coordinated in...
143
Tatta Bio @tattabio.bsky.social · 18/09/2026
Authors: @nathanlanclos.bsky.social, Kyrellos Ibrahim, Andre Cornman, Marco Huang, Vikram Gill, Aalini Jiang, Jonathan Abraham, Jennifer Gin, Yan Chen, Christopher Petzold, Justin Baerwald, Tanja Kortemme (@kortemmelab.bsky.social), @jaykeasling.bsky.social, @microyunha.bsky.social
000
Tatta Bio @tattabio.bsky.social · 18/09/2026
This work was an incredible collaborative effort with @jaykeasling.bsky.social and the JBEI team, the @kortemmelab.bsky.social, and an amazing group of coauthors. Huge thanks to everyone who made it possible.
100
Tatta Bio @tattabio.bsky.social · 18/09/2026
We think this is an exciting step for genomic language modeling: from learning the organization of natural biological systems to controllably designing biosynthetic machinery that can access chemical space beyond nature’s native repertoire.
100
Tatta Bio @tattabio.bsky.social · 18/09/2026
Experimentally, gLM2-designed variants increased production of δ-valerolactam, a molecule not naturally synthesized by PKSs, by up to 9.4× over the starting enzyme.
100
Tatta Bio @tattabio.bsky.social · 18/09/2026
The PKS is a 2,488-amino-acid molecular assembly line composed of eight protein domains that must work together to carry out coordinated, multi-step catalysis. Instead of optimizing one component at a time, gLM2 designs sequences in the context of the entire biosynthetic system.
100
Tatta Bio @tattabio.bsky.social · 18/09/2026
We're excited to share our new preprint! In collaboration with the @keaslinglab.bsky.social at UC Berkeley, we made gLM2 generative and used it to redesign a chimeric Type I polyketide synthase in the context of the full biosynthetic assembly line. 🧵 Preprint: www.biorxiv.org/content/10.6...
196
Tatta Bio @tattabio.bsky.social · 16/09/2026
Research: www.tatta.bio/intergenic-sae Here is an example query for identifying novel selenoproteins and SECIS elements not captured by known Rfams: seqhub.org/search?featu...
000
Tatta Bio @tattabio.bsky.social · 16/09/2026
Alongside our launch of SAE features (patterns gLM2 identifies in intergenic regions), we've expanded SeqHub search to incorporate them. CoSearch previously let you search up to 5 proteins co-occurring across genomes. It now also searches SAE features, Rfams, and Pfams in the same query with logic.
120
Tatta Bio @tattabio.bsky.social · 15/09/2026
Co-authors: Nicolo Zulaybar, Matt Tranzillo, Rachel Silverstein, @microyunha.bsky.social, Andre Cornman This work is supported by @moorefound.bsky.social and Schmidt Futures
000
Tatta Bio @tattabio.bsky.social · 15/09/2026
The goal is to make microbial noncoding sequence space systematically searchable: discover an intergenic pattern, trace where it occurs across evolution, and use its conserved genomic context to generate hypotheses about function. Explore: seqhub.org Paper: tatta.bio/intergenic-sae
seqhub.org
SeqHub - The Home for Biological Sequences
SeqHub is a platform for exploring, annotating, and sharing biological sequences.
111
Tatta Bio @tattabio.bsky.social · 15/09/2026
Using this framework, we identified divergent members of known RNA families, and found previously uncharacterized structured RNAs and candidate regulatory elements beyond existing annotation models.
110
Tatta Bio @tattabio.bsky.social · 15/09/2026
...without requiring a predefined motif, RNA family, or annotation. Those features are now integrated into SeqHub multimodal search, where they can be searched with proteins, Pfam domains, Rfam families, taxonomy, and genomic context.
110
Tatta Bio @tattabio.bsky.social · 15/09/2026
Microbial intergenic regions encode much of the regulatory and functional logic of the genome, but they remain systematically difficult to discover and annotate. We trained a sparse autoencoder on gLM2 to identify recurring patterns across hundreds of millions of microbial intergenic regions... 🧵
1183
Tatta Bio @tattabio.bsky.social · 10/09/2026
The proteins and context are retrieved from the OpenGenome database, using our genomic language model, gLM2. We wrote more about how the panel works and why genomic context matters for functional annotation on our blog and on UniProt's help page. UniProt refernce: www.uniprot.org/help/genomic...
uniprot.org
UniProt
UniProt is the world's leading high-quality, comprehensive and freely accessible resource of protein sequence and functional information.
010
Tatta Bio @tattabio.bsky.social · 10/09/2026
SeqHub is now integrated into UniProt, surfacing genomic context directly on prokaryotic protein entries. The panel (in the Sequence section of UniProt) shows similar proteins to the current entry, each displayed with neighboring genes from its source genome. Blog: seqhub.org/blog/seqhub-...
seqhub.org
SeqHub Genomic Neighborhoods Now Integrated in UniProt - SeqHub
UniProt entries for prokaryotic proteins now show a genomic context panel powered by SeqHub, surfacing functionally similar proteins and their genomic neighborhoods directly on the entry page.
160
Tatta Bio @tattabio.bsky.social · 04/09/2026
Big thanks to the team at UniProt for the collaboration, esp. @alexbateman1.bsky.social, Maria-Jesus Martin, Minjoon Kim, Daniel Rice, Conny WH Yu! Check out this example entry: lnkd.in/gbKtXiTc
010
Tatta Bio @tattabio.bsky.social · 04/09/2026
The SeqHub panel surfaces the genomic neighborhoods of similar proteins to the one searched. That context adds another layer of evidence for function, particularly for proteins still listed as hypothetical. We're excited to bring that signal into researchers' existing workflow.
120
Tatta Bio @tattabio.bsky.social · 04/09/2026
UniProt is where most researchers start when they want to understand a protein. As of today, prokaryotic entries there will show something new: genomic context, powered by SeqHub, right in the Sequence section of the page.
13811
Tatta Bio @tattabio.bsky.social · 25/08/2026
Explore the intergenic space for promoters, terminators, and other regulatory elements using SeqHub's DNA view. Click into any region of a contig to view and copy the nucleotide sequence directly.
000
Tatta Bio @tattabio.bsky.social · 13/08/2026
You can now explore DNA in SeqHub. Click into any gene or intergenic region within a contig to retrieve and copy nucleotides.
033
Tatta Bio @tattabio.bsky.social · 10/08/2026
One day until our SeqHub webinar. Tomorrow at 11am ET we'll cover what's new, including non-coding RNA annotations and protein-protein interaction predictions in search, plus time for your questions. Sign up: docs.google.com/forms/d/e/1F... Can't make it? DM us to find time to chat.
000
Tatta Bio @tattabio.bsky.social · 06/08/2026
By localizing these signals in 3D alongside SeqHub's functional annotations and genomic context, they enable more precise functional predictions for uncharacterized proteins.
010
Tatta Bio @tattabio.bsky.social · 06/08/2026
Biohub used sparse autoencoders (SAEs) to decompose the internal representations of their protein language model, ESMC, into 16,000+ distinct features. Each feature can correspond to a conserved structural motif, or a functional pattern that recurs across diverse proteins.
110
Tatta Bio @tattabio.bsky.social · 06/08/2026
SeqHub protein searches now show which regions of a predicted structure are tied to specific biological functions or evolutionary patterns, powered by @biohub.org's SAE features. 🧵
1177
Tatta Bio @tattabio.bsky.social · 05/08/2026
SeqHub's Diversity Search widens your results to surface distant homologs that other search algorithms rank low or miss entirely. When you find a hit with low sequence identity, you can run a structural alignment within SeqHub to validate it before committing time to the candidate.
020
Tatta Bio @tattabio.bsky.social · 30/07/2026
SeqHub has an updated interface and new features! Join a live walkthrough on August 11 at 11am ET to see SeqHub in action, whether you're just getting started or want to catch up on what's changed since our last update. Sign up: docs.google.com/forms/d/e/1F...
000
Tatta Bio @tattabio.bsky.social · 29/07/2026
You can test it out for free with a SeqHub account. docs.sequhub.org
docs.sequhub.org
000
Tatta Bio @tattabio.bsky.social · 29/07/2026
Retrieving gene neighborhoods is still one of the harder parts of working with sequence data, often needing large downloads and manual parsing of genomic locations. The SeqHub API returns functional annotations and surrounding genes for a query protein in one call, with locations already resolved.
100
Tatta Bio @tattabio.bsky.social · 28/07/2026
Sign up for webinar here and we'll send you an invite: docs.google.com/forms/d/e/1F...
000
Tatta Bio @tattabio.bsky.social · 28/07/2026
SeqHub, now in dark mode. Curious about the new UI or want to get up to speed on the latest SeqHub features? Join us for a webinar on August 11 at 11am EST. Sign up link below.
100
Tatta Bio @tattabio.bsky.social · 23/07/2026
We'll be at the AI x Bio Summit today at the NYSE, hosted by Decoding Bio. If you're at the event and curious to learn more about our models or platform, SeqHub, send us a DM or find Steph Flamen on the floor.
000
Tatta Bio @tattabio.bsky.social · 22/07/2026
We just launched a new UI for SeqHub so things might look a little different next time you log in 👀 With your feedback, we rebuilt navigation across the platform, with most tools now accessible from one location.
011
Tatta Bio @tattabio.bsky.social · 09/07/2026
We've arrived at ICML! @microyunha.bsky.social will be speaking at the GenBio Workshop on July 10th at 1:30pm local time: "Genomic Language Modeling for Context-Aware Biological Discovery." Hope to see you there! #ICML2026
010
Tatta Bio @tattabio.bsky.social · 07/07/2026
This week we are at ICML in Seoul, and next week we're headed to Washington, D.C. for ISMB. Our ICML (GenBio Workshop) talk: July 10th at 1:30pm Our ISMB talk: July 15 at 12:30pm We'd love to connect. Send a DM or email team@tatta.bio and we'll find time!
000
Tatta Bio @tattabio.bsky.social · 02/07/2026
In SeqHub, you can now search a protein to find FlashPPI2-predicted interaction partners across all 400 million proteins and 132K microbial genomes in our database (OpenGenome). You can still upload full genomes to find interactions within and between genomes.
020
Tatta Bio @tattabio.bsky.social · 01/07/2026
It's now available on @hf.co and in SeqHub (seqhub.org), where we've also run it across our database of over 130,000 microbial genomes, enabling immediate discovery of protein interactions. huggingface.co/tattabio/fla...
huggingface.co
tattabio/flashppi2 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
000
Tatta Bio @tattabio.bsky.social · 01/07/2026
Two weeks ago, our FlashPPI paper was published in @pnas.org. Today, we introduce our updated model, FlashPPI2. Fine-tuned on new AlphaFold structures, this new model achieves a 17% improvement over FlashPPI on the E. coli protein interaction benchmark. seqhub.org/blog/flashppi2
seqhub.org
FlashPPI2: Enhanced Model, Scaled Across 130,000+ Genomes - SeqHub
FlashPPI2 drives up PPI prediction performance (AUPRC) by 17% while maintaining inference speed at minutes per genome — now deployed across SeqHub's database of over 130,000 microbial genomes.
153
Tatta Bio @tattabio.bsky.social · 25/06/2026
If you're attending and want to connect live, reply here, send us a DM, or reach out at team@tatta.bio.
000
Tatta Bio @tattabio.bsky.social · 25/06/2026
We'll be at ICML in Seoul July 6 - 11! Our Chief Scientist, @microyunha.bsky.social, will be speaking at the GenBio Workshop on July 10 at 1:30pm local time, presenting "Genomic Language Modeling for Context-Aware Biological Discovery." #ICML2026
121
Tatta Bio @tattabio.bsky.social · 23/06/2026
In at least one case, a function predicted in SeqHub was subsequently confirmed at the bench. Grateful to the Rock Lab for their use and feedback that continues to shape the platform.
000
Tatta Bio @tattabio.bsky.social · 23/06/2026
When a genetic screen returns a hit, it often points to a gene with no known function. Annotation that previously required multiple tools can now start in one place, with predictions that outperform what was previously available.
100
Tatta Bio @tattabio.bsky.social · 23/06/2026
The @rocklabtb.bsky.social studies Mycobacterium tuberculosis and M. abscessus, two clinically significant mycobacteria with large stretches of unannotated genome.
100
Tatta Bio @tattabio.bsky.social · 23/06/2026
"SeqHub has become an integral part of our workflow...[it's] typically the first place we go to begin understanding what a gene might be doing and to identify its genomic neighbors across bacterial genomes." - Jeremy Rock, Rockefeller University seqhub.org/blog/rock-la...
seqhub.org
How the Rock Lab Uses SeqHub to Accelerate Discovery in Mtb and Mabs - SeqHub
The Rock Lab at Rockefeller University studies Mycobacterium tuberculosis and M. abscessus — two pathogens with large stretches of unannotated genome. SeqHub has become an integral part of their workf...
101
Tatta Bio @tattabio.bsky.social · 17/06/2026
That speed makes proteome-scale interaction screening practical in microbial systems for the first time. An improved model and expanded features are coming soon to SeqHub. In the meantime, check out FlashPPI run on a sample genome in SeqHub: seqhub.org/tattabio/eco...
seqhub.org
ecoli_k12 by tattabio - SeqHub
View and analyze 4,317 sequences on SeqHub
000
Tatta Bio @tattabio.bsky.social · 17/06/2026
Our FlashPPI paper is out in @pnas.org! FlashPPI is a model for proteome-wide protein-protein interaction prediction. 🔹4x better predictive performance & 2,400x faster than existing sequence-based methods 🔹20,000x faster than leading structure-based approaches. www.pnas.org/doi/10.1073/...
pnas.org
PNAS
Proceedings of the National Academy of Sciences (PNAS), a peer reviewed journal of the National Academy of Sciences (NAS) - an authoritative source of high-impact, original research that broadly spans...
121
Tatta Bio @tattabio.bsky.social · 16/06/2026
Search a protein in SeqHub and relevant literature appears alongside your genomic context results. Using the PaperBLAST database (Price & Arkin 2024), SeqHub retrieves papers, with specific passages highlighted, citing the sequence you searched and closely related ones.
020
Tatta Bio @tattabio.bsky.social · 09/06/2026
SeqHub is free for academic use. If you're designing a course around sequencing or genome annotation, reach out at team@tatta.bio to talk through how to integrate it.
000