Tatta Bio @tattabio.bsky.social · 23/09/2026And sustainable chemistry. Redesigning biosynthetic assembly lines to produce materials and commodity chemicals from biology rather than petrochemical feedstocks could make chemical manufacturing more sustainable. 000
Tatta Bio @tattabio.bsky.social · 23/09/2026This could enable: New therapeutics. More than half of FDA-approved small-molecule drugs over the past 4 decades are natural products, their derivatives, or synthetic mimics. Redirecting biosynthetic assembly lines toward new molecular targets could substantially expand accessible chemical space. 100
Tatta Bio @tattabio.bsky.social · 23/09/2026In our preprint released last Friday, we show that evolutionary sequence information learned by our genomic language model, gLM2, can be used to design complex, multi-domain enzymes and expand their chemistry beyond their natural repertoire. 🧵 Preprint: www.biorxiv.org/content/10.6...biorxiv.orgGenerative Design of New-to-nature Biosynthetic Assembly Lines with Genomic Language ModelingReprogramming biosynthetic assembly lines can extend biosynthesis beyond the chemical space explored by nature. However, this remains difficult because assembly-line function depends on coordinated in... 143
Tatta Bio @tattabio.bsky.social · 18/09/2026Authors: @nathanlanclos.bsky.social, Kyrellos Ibrahim, Andre Cornman, Marco Huang, Vikram Gill, Aalini Jiang, Jonathan Abraham, Jennifer Gin, Yan Chen, Christopher Petzold, Justin Baerwald, Tanja Kortemme (@kortemmelab.bsky.social), @jaykeasling.bsky.social, @microyunha.bsky.social 000
Tatta Bio @tattabio.bsky.social · 18/09/2026This work was an incredible collaborative effort with @jaykeasling.bsky.social and the JBEI team, the @kortemmelab.bsky.social, and an amazing group of coauthors. Huge thanks to everyone who made it possible. 100
Tatta Bio @tattabio.bsky.social · 18/09/2026We think this is an exciting step for genomic language modeling: from learning the organization of natural biological systems to controllably designing biosynthetic machinery that can access chemical space beyond nature’s native repertoire. 100
Tatta Bio @tattabio.bsky.social · 18/09/2026Experimentally, gLM2-designed variants increased production of δ-valerolactam, a molecule not naturally synthesized by PKSs, by up to 9.4× over the starting enzyme. 100
Tatta Bio @tattabio.bsky.social · 18/09/2026The PKS is a 2,488-amino-acid molecular assembly line composed of eight protein domains that must work together to carry out coordinated, multi-step catalysis. Instead of optimizing one component at a time, gLM2 designs sequences in the context of the entire biosynthetic system. 100
Tatta Bio @tattabio.bsky.social · 18/09/2026We're excited to share our new preprint! In collaboration with the @keaslinglab.bsky.social at UC Berkeley, we made gLM2 generative and used it to redesign a chimeric Type I polyketide synthase in the context of the full biosynthetic assembly line. 🧵 Preprint: www.biorxiv.org/content/10.6... 196
Tatta Bio @tattabio.bsky.social · 16/09/2026Research: www.tatta.bio/intergenic-sae Here is an example query for identifying novel selenoproteins and SECIS elements not captured by known Rfams: seqhub.org/search?featu... 000
Tatta Bio @tattabio.bsky.social · 16/09/2026Alongside our launch of SAE features (patterns gLM2 identifies in intergenic regions), we've expanded SeqHub search to incorporate them. CoSearch previously let you search up to 5 proteins co-occurring across genomes. It now also searches SAE features, Rfams, and Pfams in the same query with logic. 120
Tatta Bio @tattabio.bsky.social · 15/09/2026Co-authors: Nicolo Zulaybar, Matt Tranzillo, Rachel Silverstein, @microyunha.bsky.social, Andre Cornman This work is supported by @moorefound.bsky.social and Schmidt Futures 000
Tatta Bio @tattabio.bsky.social · 15/09/2026The goal is to make microbial noncoding sequence space systematically searchable: discover an intergenic pattern, trace where it occurs across evolution, and use its conserved genomic context to generate hypotheses about function. Explore: seqhub.org Paper: tatta.bio/intergenic-saeseqhub.orgSeqHub - The Home for Biological SequencesSeqHub is a platform for exploring, annotating, and sharing biological sequences. 111
Tatta Bio @tattabio.bsky.social · 15/09/2026Using this framework, we identified divergent members of known RNA families, and found previously uncharacterized structured RNAs and candidate regulatory elements beyond existing annotation models. 110
Tatta Bio @tattabio.bsky.social · 15/09/2026...without requiring a predefined motif, RNA family, or annotation. Those features are now integrated into SeqHub multimodal search, where they can be searched with proteins, Pfam domains, Rfam families, taxonomy, and genomic context. 110
Tatta Bio @tattabio.bsky.social · 15/09/2026Microbial intergenic regions encode much of the regulatory and functional logic of the genome, but they remain systematically difficult to discover and annotate. We trained a sparse autoencoder on gLM2 to identify recurring patterns across hundreds of millions of microbial intergenic regions... 🧵 1183
Tatta Bio @tattabio.bsky.social · 10/09/2026The proteins and context are retrieved from the OpenGenome database, using our genomic language model, gLM2. We wrote more about how the panel works and why genomic context matters for functional annotation on our blog and on UniProt's help page. UniProt refernce: www.uniprot.org/help/genomic...uniprot.orgUniProtUniProt is the world's leading high-quality, comprehensive and freely accessible resource of protein sequence and functional information. 010
Tatta Bio @tattabio.bsky.social · 10/09/2026SeqHub is now integrated into UniProt, surfacing genomic context directly on prokaryotic protein entries. The panel (in the Sequence section of UniProt) shows similar proteins to the current entry, each displayed with neighboring genes from its source genome. Blog: seqhub.org/blog/seqhub-...seqhub.orgSeqHub Genomic Neighborhoods Now Integrated in UniProt - SeqHubUniProt entries for prokaryotic proteins now show a genomic context panel powered by SeqHub, surfacing functionally similar proteins and their genomic neighborhoods directly on the entry page. 160
Tatta Bio @tattabio.bsky.social · 04/09/2026Big thanks to the team at UniProt for the collaboration, esp. @alexbateman1.bsky.social, Maria-Jesus Martin, Minjoon Kim, Daniel Rice, Conny WH Yu! Check out this example entry: lnkd.in/gbKtXiTc 010
Tatta Bio @tattabio.bsky.social · 04/09/2026The SeqHub panel surfaces the genomic neighborhoods of similar proteins to the one searched. That context adds another layer of evidence for function, particularly for proteins still listed as hypothetical. We're excited to bring that signal into researchers' existing workflow. 120
Tatta Bio @tattabio.bsky.social · 04/09/2026UniProt is where most researchers start when they want to understand a protein. As of today, prokaryotic entries there will show something new: genomic context, powered by SeqHub, right in the Sequence section of the page. 13811
Tatta Bio @tattabio.bsky.social · 25/08/2026Explore the intergenic space for promoters, terminators, and other regulatory elements using SeqHub's DNA view. Click into any region of a contig to view and copy the nucleotide sequence directly. 000
Tatta Bio @tattabio.bsky.social · 13/08/2026You can now explore DNA in SeqHub. Click into any gene or intergenic region within a contig to retrieve and copy nucleotides. 033
Tatta Bio @tattabio.bsky.social · 10/08/2026One day until our SeqHub webinar. Tomorrow at 11am ET we'll cover what's new, including non-coding RNA annotations and protein-protein interaction predictions in search, plus time for your questions. Sign up: docs.google.com/forms/d/e/1F... Can't make it? DM us to find time to chat. 000
Tatta Bio @tattabio.bsky.social · 06/08/2026By localizing these signals in 3D alongside SeqHub's functional annotations and genomic context, they enable more precise functional predictions for uncharacterized proteins. 010
Tatta Bio @tattabio.bsky.social · 06/08/2026Biohub used sparse autoencoders (SAEs) to decompose the internal representations of their protein language model, ESMC, into 16,000+ distinct features. Each feature can correspond to a conserved structural motif, or a functional pattern that recurs across diverse proteins. 110
Tatta Bio @tattabio.bsky.social · 06/08/2026SeqHub protein searches now show which regions of a predicted structure are tied to specific biological functions or evolutionary patterns, powered by @biohub.org's SAE features. 🧵 1177
Tatta Bio @tattabio.bsky.social · 05/08/2026SeqHub's Diversity Search widens your results to surface distant homologs that other search algorithms rank low or miss entirely. When you find a hit with low sequence identity, you can run a structural alignment within SeqHub to validate it before committing time to the candidate. 020
Tatta Bio @tattabio.bsky.social · 30/07/2026SeqHub has an updated interface and new features! Join a live walkthrough on August 11 at 11am ET to see SeqHub in action, whether you're just getting started or want to catch up on what's changed since our last update. Sign up: docs.google.com/forms/d/e/1F... 000
Tatta Bio @tattabio.bsky.social · 29/07/2026You can test it out for free with a SeqHub account. docs.sequhub.orgdocs.sequhub.org 000
Tatta Bio @tattabio.bsky.social · 29/07/2026Retrieving gene neighborhoods is still one of the harder parts of working with sequence data, often needing large downloads and manual parsing of genomic locations. The SeqHub API returns functional annotations and surrounding genes for a query protein in one call, with locations already resolved. 100
Tatta Bio @tattabio.bsky.social · 28/07/2026Sign up for webinar here and we'll send you an invite: docs.google.com/forms/d/e/1F... 000
Tatta Bio @tattabio.bsky.social · 28/07/2026SeqHub, now in dark mode. Curious about the new UI or want to get up to speed on the latest SeqHub features? Join us for a webinar on August 11 at 11am EST. Sign up link below. 100
Tatta Bio @tattabio.bsky.social · 23/07/2026We'll be at the AI x Bio Summit today at the NYSE, hosted by Decoding Bio. If you're at the event and curious to learn more about our models or platform, SeqHub, send us a DM or find Steph Flamen on the floor. 000
Tatta Bio @tattabio.bsky.social · 22/07/2026We just launched a new UI for SeqHub so things might look a little different next time you log in 👀 With your feedback, we rebuilt navigation across the platform, with most tools now accessible from one location. 011
Tatta Bio @tattabio.bsky.social · 09/07/2026We've arrived at ICML! @microyunha.bsky.social will be speaking at the GenBio Workshop on July 10th at 1:30pm local time: "Genomic Language Modeling for Context-Aware Biological Discovery." Hope to see you there! #ICML2026 010
Tatta Bio @tattabio.bsky.social · 07/07/2026This week we are at ICML in Seoul, and next week we're headed to Washington, D.C. for ISMB. Our ICML (GenBio Workshop) talk: July 10th at 1:30pm Our ISMB talk: July 15 at 12:30pm We'd love to connect. Send a DM or email team@tatta.bio and we'll find time! 000
Tatta Bio @tattabio.bsky.social · 02/07/2026In SeqHub, you can now search a protein to find FlashPPI2-predicted interaction partners across all 400 million proteins and 132K microbial genomes in our database (OpenGenome). You can still upload full genomes to find interactions within and between genomes. 020
Tatta Bio @tattabio.bsky.social · 01/07/2026It's now available on @hf.co and in SeqHub (seqhub.org), where we've also run it across our database of over 130,000 microbial genomes, enabling immediate discovery of protein interactions. huggingface.co/tattabio/fla...huggingface.cotattabio/flashppi2 · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 000
Tatta Bio @tattabio.bsky.social · 01/07/2026Two weeks ago, our FlashPPI paper was published in @pnas.org. Today, we introduce our updated model, FlashPPI2. Fine-tuned on new AlphaFold structures, this new model achieves a 17% improvement over FlashPPI on the E. coli protein interaction benchmark. seqhub.org/blog/flashppi2seqhub.orgFlashPPI2: Enhanced Model, Scaled Across 130,000+ Genomes - SeqHubFlashPPI2 drives up PPI prediction performance (AUPRC) by 17% while maintaining inference speed at minutes per genome — now deployed across SeqHub's database of over 130,000 microbial genomes. 153
Tatta Bio @tattabio.bsky.social · 25/06/2026If you're attending and want to connect live, reply here, send us a DM, or reach out at team@tatta.bio. 000
Tatta Bio @tattabio.bsky.social · 25/06/2026We'll be at ICML in Seoul July 6 - 11! Our Chief Scientist, @microyunha.bsky.social, will be speaking at the GenBio Workshop on July 10 at 1:30pm local time, presenting "Genomic Language Modeling for Context-Aware Biological Discovery." #ICML2026 121
Tatta Bio @tattabio.bsky.social · 23/06/2026In at least one case, a function predicted in SeqHub was subsequently confirmed at the bench. Grateful to the Rock Lab for their use and feedback that continues to shape the platform. 000
Tatta Bio @tattabio.bsky.social · 23/06/2026When a genetic screen returns a hit, it often points to a gene with no known function. Annotation that previously required multiple tools can now start in one place, with predictions that outperform what was previously available. 100
Tatta Bio @tattabio.bsky.social · 23/06/2026The @rocklabtb.bsky.social studies Mycobacterium tuberculosis and M. abscessus, two clinically significant mycobacteria with large stretches of unannotated genome. 100
Tatta Bio @tattabio.bsky.social · 23/06/2026"SeqHub has become an integral part of our workflow...[it's] typically the first place we go to begin understanding what a gene might be doing and to identify its genomic neighbors across bacterial genomes." - Jeremy Rock, Rockefeller University seqhub.org/blog/rock-la...seqhub.orgHow the Rock Lab Uses SeqHub to Accelerate Discovery in Mtb and Mabs - SeqHubThe Rock Lab at Rockefeller University studies Mycobacterium tuberculosis and M. abscessus — two pathogens with large stretches of unannotated genome. SeqHub has become an integral part of their workf... 101
Tatta Bio @tattabio.bsky.social · 17/06/2026That speed makes proteome-scale interaction screening practical in microbial systems for the first time. An improved model and expanded features are coming soon to SeqHub. In the meantime, check out FlashPPI run on a sample genome in SeqHub: seqhub.org/tattabio/eco...seqhub.orgecoli_k12 by tattabio - SeqHubView and analyze 4,317 sequences on SeqHub 000
Tatta Bio @tattabio.bsky.social · 17/06/2026Our FlashPPI paper is out in @pnas.org! FlashPPI is a model for proteome-wide protein-protein interaction prediction. 🔹4x better predictive performance & 2,400x faster than existing sequence-based methods 🔹20,000x faster than leading structure-based approaches. www.pnas.org/doi/10.1073/...pnas.orgPNASProceedings of the National Academy of Sciences (PNAS), a peer reviewed journal of the National Academy of Sciences (NAS) - an authoritative source of high-impact, original research that broadly spans... 121
Tatta Bio @tattabio.bsky.social · 16/06/2026Search a protein in SeqHub and relevant literature appears alongside your genomic context results. Using the PaperBLAST database (Price & Arkin 2024), SeqHub retrieves papers, with specific passages highlighted, citing the sequence you searched and closely related ones. 020
Tatta Bio @tattabio.bsky.social · 09/06/2026SeqHub is free for academic use. If you're designing a course around sequencing or genome annotation, reach out at team@tatta.bio to talk through how to integrate it. 000