Sign in

Yunha Hwang

@microyunha.bsky.social
1.4K followers 1.1K following 59 posts

Building genomic intelligence @ Tatta Bio

PostsRepliesMedia
Reposted by Yunha Hwang
Tatta Bio @tattabio.bsky.social · 18/09/2026
We're excited to share our new preprint! In collaboration with the @keaslinglab.bsky.social at UC Berkeley, we made gLM2 generative and used it to redesign a chimeric Type I polyketide synthase in the context of the full biosynthetic assembly line. 🧵 Preprint: www.biorxiv.org/content/10.6...
196
Yunha Hwang @microyunha.bsky.social · 15/09/2026
Microbial genomes have been extensively mined for new proteins, but the noncoding sequences between genes remain largely unexplored despite encoding important regulatory and functional information. In our new work, we asked whether gLM2 could help systematically discover these intergenic features.
0164
Yunha Hwang @microyunha.bsky.social · 04/09/2026
This one is really special to me! As someone who probably spends more time on UniProt than on Google, I’m incredibly proud that SeqHub search is now integrated with UniProt, bringing together UniProt’s rich protein annotations with SeqHub’s view of native genomic context across diverse microbes!
0201
Reposted by Yunha Hwang
Tatta Bio @tattabio.bsky.social · 13/08/2026
You can now explore DNA in SeqHub. Click into any gene or intergenic region within a contig to retrieve and copy nucleotides.
033
Reposted by Yunha Hwang
Tatta Bio @tattabio.bsky.social · 06/08/2026
SeqHub protein searches now show which regions of a predicted structure are tied to specific biological functions or evolutionary patterns, powered by @biohub.org's SAE features. 🧵
1177
Reposted by Yunha Hwang
Tatta Bio @tattabio.bsky.social · 22/07/2026
We just launched a new UI for SeqHub so things might look a little different next time you log in 👀 With your feedback, we rebuilt navigation across the platform, with most tools now accessible from one location.
011
Yunha Hwang @microyunha.bsky.social · 02/07/2026
With pre-calculated FlashPPI2 interactions in SeqHub, you can immediately discover predicted physical interaction partners in the native genomic context AND examine how conserved these patterns are across diverse organisms. Give your favorite protein a try, and let us know what you find!🪄
0101
Reposted by Yunha Hwang
Tatta Bio @tattabio.bsky.social · 01/07/2026
Two weeks ago, our FlashPPI paper was published in @pnas.org. Today, we introduce our updated model, FlashPPI2. Fine-tuned on new AlphaFold structures, this new model achieves a 17% improvement over FlashPPI on the E. coli protein interaction benchmark. seqhub.org/blog/flashppi2
seqhub.org
FlashPPI2: Enhanced Model, Scaled Across 130,000+ Genomes - SeqHub
FlashPPI2 drives up PPI prediction performance (AUPRC) by 17% while maintaining inference speed at minutes per genome — now deployed across SeqHub's database of over 130,000 microbial genomes.
153
Reposted by Yunha Hwang
Tatta Bio @tattabio.bsky.social · 25/06/2026
We'll be at ICML in Seoul July 6 - 11! Our Chief Scientist, @microyunha.bsky.social, will be speaking at the GenBio Workshop on July 10 at 1:30pm local time, presenting "Genomic Language Modeling for Context-Aware Biological Discovery." #ICML2026
121
Reposted by Yunha Hwang
Tatta Bio @tattabio.bsky.social · 23/06/2026
"SeqHub has become an integral part of our workflow...[it's] typically the first place we go to begin understanding what a gene might be doing and to identify its genomic neighbors across bacterial genomes." - Jeremy Rock, Rockefeller University seqhub.org/blog/rock-la...
seqhub.org
How the Rock Lab Uses SeqHub to Accelerate Discovery in Mtb and Mabs - SeqHub
The Rock Lab at Rockefeller University studies Mycobacterium tuberculosis and M. abscessus — two pathogens with large stretches of unannotated genome. SeqHub has become an integral part of their workf...
101
Yunha Hwang @microyunha.bsky.social · 17/06/2026
Excited to share the latest version of FlashPPI published with @pnas.org ! And stay tuned for updates soon…👀
0123
Yunha Hwang @microyunha.bsky.social · 26/05/2026
SeqHub API Beta now live - genomic neighborhood retrieval and functional annotation just became instantaneous🤗
040
Yunha Hwang @microyunha.bsky.social · 08/04/2026
SeqHub MSA live, integrated with structure vis! Another user requested feature🤝
042
Yunha Hwang @microyunha.bsky.social · 07/04/2026
We’re incredibly grateful to have Alex Bateman on our advisory board. Biology today would look very different without UniProt. Scientific data infrastructure is the bedrock of innovation, and we’re excited to learn from Alex’s experience helping building such a foundational resource.
020
Reposted by Yunha Hwang
Tatta Bio @tattabio.bsky.social · 02/04/2026
Most protein-protein interaction tools work on protein pairs. FlashPPI runs at proteome scale and now across two proteomes at once. Upload any two datasets (full genomes, partial genomes, or custom protein sets) and get back a predicted interaction network spanning both.
042
Reposted by Yunha Hwang
Tatta Bio @tattabio.bsky.social · 01/04/2026
We're hosting a live walkthrough of FlashPPI in SeqHub on April 15 at 11am EST. We'll briefly discuss our protein-protein interaction model then walk through how you can use it in SeqHub. Register here: forms.gle/iBQrpYnLeiF1...
forms.gle
SeqHub Platform Walk-Through Webinar Registration
Register to attend the live FlashPPI-focused walk-through of the SeqHub Platform. We'll send you an invite to this email once you submit the form. See you soon!
001
Yunha Hwang @microyunha.bsky.social · 24/03/2026
My group at MIT is seeking a research scientist with a strong *experimental* background to lead and help shape the lab’s experimental infrastructure, supporting efforts to advance AI-driven enzyme discovery and characterization. See the full JD here: acrobat.adobe.com/id/urn:aaid:...
acrobat.adobe.com
Adobe Acrobat
11616
Yunha Hwang @microyunha.bsky.social · 18/03/2026
Applications for MIT Novo-Nordisk AI postdoc fellowships are due Apr 15. Focus area lists AI and Biology topics, apply to work on this exciting field with amazing peers! engineering.mit.edu/novo-nordisk
engineering.mit.edu
Novo Nordisk Fellowship
Potential areas of focus for postdoctoral fellows participating in the MIT-Novo Nordisk Postdoc Program include, but are not limited to, the following: Your complete application should include: In the...
011
Yunha Hwang @microyunha.bsky.social · 05/03/2026
We thought a lot about how to deploy 𝑭𝒍𝒂𝒔𝒉𝑷𝑷𝑰, and we are very proud of this implementation that integrates annotation+context+CoSearch+agent with FlashPPI on SeqHub!
052
Yunha Hwang @microyunha.bsky.social · 03/03/2026
Step-by-step how to run FlashPPI on your favorite genomes!
091
Reposted by Yunha Hwang
Andre Cornman @ancornman1.bsky.social · 03/03/2026
Predicting protein-protein interactions (PPIs) at proteome scale can take months with co-folding models due to the massive all-vs-all comparisons required. We are excited to announce FlashPPI, a contrastive learning framework that predicts proteome wide physical interfaces in minutes. 1/🧵
16727
Yunha Hwang @microyunha.bsky.social · 03/03/2026
Protein–protein interactions (PPIs) are key to discovering and interpreting new biological functions. We’re excited to introduce 𝑭𝒍𝒂𝒔𝒉𝑷𝑷𝑰: a new application of gLM2 that uses genomic language modeling to predict proteome-wide PPIs in microbial genomes in minutes.
24222
Reposted by Yunha Hwang
Tatta Bio @tattabio.bsky.social · 12/02/2026
We’d love to join your lab meeting! We’ve been meeting with research groups to share how scientists are using SeqHub for sequence and genome analysis, and the conversations have been highly interactive and grounded in real workflows. Booking info below.
101
Reposted by Yunha Hwang
Tatta Bio @tattabio.bsky.social · 10/02/2026
We’re excited to welcome Daniela Bourges-Waldegg to the SeqHub Advisory Board! Daniela is EVP + Chief Digital & Technology Officer at @addgene.bsky.social. She will help shape our approach to building researcher-centered digital infrastructure with an eye toward long-term scientific impact.
052
Yunha Hwang @microyunha.bsky.social · 04/02/2026
First, @tattabio.bsky.social is now on Bluesky!💙 and second, we launched mult-sequence CoSearch on SeqHub!
072
Reposted by Yunha Hwang
Lizzy Wilbanks @lizzywilbanks.bsky.social · 05/11/2025
This. Is. So. Cool. 🤯
131
Reposted by Yunha Hwang
ISCB News @iscb.bsky.social · 28/10/2025
Released today from Tatta Bio: SeqHub! A place to explore, annotate, and share sequence data with functional insights.  Over 1,000 scientists worldwide have already used SeqHub to annotate more than 550,000 proteins, uncovering new insights and accelerating discovery.
201
Yunha Hwang @microyunha.bsky.social · 28/10/2025
We're thrilled to announce SeqHub, an AI-enabled platform for biological sequence analysis. SeqHub brings together sequence search, genome annotation, and data sharing in one place.
34919
Reposted by Yunha Hwang
Axel Visel @axelvisel.bsky.social · 25/08/2025
Ready to explore New Lineages of Life with @jgi.doe.gov ? 🧬🦠 Registration for our 2025 NeLLi Symposium is now open. For the first time in collaboration with @unlv.edu Mark the date: November 6-7 in Las Vegas, NV
163
Yunha Hwang @microyunha.bsky.social · 02/06/2025
At Tatta Bio, we have been thinking deeply about the sequence-to-function problem. We believe that before AI can power functional prediction, we first need to rethink how we curate, manage, and share sequence data. Here, we share our initial ideas on what we are building next:
tattabio.substack.com
Today's sequence data infrastructure is set up for failure in the age of AI.
Building an open and collaborative sequence platform for both Human and AI scientists.
184
Reposted by Yunha Hwang
Florian Trigodet @floriantrigodet.bsky.social · 28/04/2025
I am very happy (and anxious) to share with you our most recent work in which we evaluated four of the most popular long-read assemblers, www.biorxiv.org/content/10.1... and tell you just a little bit about it in the following 🧵
biorxiv.org
Assemblies of long-read metagenomes suffer from diverse errors
Genomes from metagenomes have revolutionised our understanding of microbial diversity, ecology, and evolution, propelling advances in basic science, biomedicine, and biotechnology. Assembly algorithms...
513674
Yunha Hwang @microyunha.bsky.social · 28/04/2025
It’s official! 🎉 I’m thrilled to announce that I will be joining MIT as an assistant professor in a shared appointment between Biology, EECS and Schwarzman College of Computing this fall.
9663
Yunha Hwang @microyunha.bsky.social · 24/03/2025
Tatta Bio is growing! We are hiring *two positions* in Business Development and Software Engineering to lead the development of AI-enabled scientific software for open science and biological sequence interpretation. Please check out the job postings at www.tatta.bio/careers and share widely!
tatta.bio
Job Board | Notion
Overview
052
Yunha Hwang @microyunha.bsky.social · 17/12/2024
Can LLM agents discover novel protein functions? Introducing Gaia Agent 🌎 🤖: an AI biologist capable of reasoning across genomic contexts to predict functions of proteins! Gaia Agent is now integrated with Gaia Search at gaia.tatta.bio
23813
Yunha Hwang @microyunha.bsky.social · 15/12/2024
If you are at #NeurIPS2024 don't miss @ancornman1.bsky.social's talk on OMG/gLM2 at 9AM! @workshopmlsb.bsky.social East meeting room 11,12
0123
Yunha Hwang @microyunha.bsky.social · 10/12/2024
Excited to be at #NeurIPS this week. @ancornman1.bsky.social will give a spotlight talk at the @workshopmlsb.bsky.social on gLM2/OMG! Please reach out if you want to chat about gLM2/OMG/Gaia and our latest projects😇 www.biorxiv.org/content/10.1...
biorxiv.org
The OMG dataset: An Open MetaGenomic corpus for mixed-modality genomic language modeling
Biological language model performance depends heavily on pretraining data quality, diversity, and size. While metagenomic datasets feature enormous biological diversity, their utilization as pretraini...
093
Reposted by Yunha Hwang
Mitja M. Zdouc @mmzdouc.bsky.social · 10/12/2024
Are you working on natural products? We’ve just released version 4.0 of the MIBiG data standard and repository! It now includes 3059 biosynthetic gene clusters, thanks to the combined efforts of 288 expert contributors. A thread: (1/8) academic.oup.com/nar/advance-...
academic.oup.com
MIBiG 4.0: advancing biosynthetic gene cluster curation through global collaboration
Abstract. Specialized or secondary metabolites are small molecules of biological origin, often showing potent biological activities with applications in ag
49153
Reposted by Yunha Hwang
Amy Lu @amyxlu.bsky.social · 06/12/2024
1/🧬 Excited to share PLAID, our new approach for co-generating sequence and all-atom protein structures by sampling from the latent space of ESMFold. This requires only sequences during training, which unlocks more data and annotations: bit.ly/plaid-proteins 🧵
overview of results for PLAID!
112437
Reposted by Yunha Hwang
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 23/11/2024
Our Big Fantastic Virus Database (BFVD) is now published NAR! It contains protein structure predictions of major viral clades, enhanced by petabase-scale homology search and it's explorable on the web. 🌐 bfvd.foldseek.com 💾 bfvd.steineggerlab.workers.dev 📄 academic.oup.com/nar/advance-...
6339126
Yunha Hwang @microyunha.bsky.social · 19/11/2024
Hello 🦋 #protein / #microbio / #BioML community! We are excited to release Gaia🌎, a context-aware protein search tool, extending protein search and discovery capabilities beyond sequence and structure, to include *genomic context*. Search your favorite protein sequences with on gaia.tatta.bio
1023875