Sign in

Clay Kosonocky

@kosonocky.bsky.social
783 followers 285 following 142 posts

ML + Biochemistry PhD Candidate at UT Austin. BioML Society Founder. All problems are solvable, so let's solve some biomlsociety.org

PostsRepliesMedia
Reposted by Clay Kosonocky
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 24/08/2026
Curate a molecule <-> text dataset, train a model, and use it to discover molecules that fight antibiotic resistance by deactivating beta-lactamase enzymes. @kosonocky.bsky.social www.biorxiv.org/content/10.6...
1132
Clay Kosonocky @kosonocky.bsky.social · 18/08/2026
Here is a link to the website: pubchef.org Hope you all enjoy surfing chemical space!
pubchef.org
PubCheF Explorer
011
Clay Kosonocky @kosonocky.bsky.social · 18/08/2026
Compound 61, inhibited serine beta-lactamases: pubchef.org/share/169810... Personally I find the ring structure of this molecule to be quite beautiful
pubchef.org
3-(6-methyl-4,8-dioxo-1,3,6,2-dioxazaborocan-2-yl)benzaldehyde — PubCheF Explorer
ZINC ID: 169810265 · C12H12BNO5
110
Clay Kosonocky @kosonocky.bsky.social · 18/08/2026
Compound 42, inhibited L1 metallo-beta-lactamase: pubchef.org/share/154546... (This molecule is quite interesting, it's a stereoisomer of a penicillin degradation product that has been used to chelate copper to treat Wilson's disease, yet it also inhibits metallo-beta-lactamses!)
pubchef.org
(2R)-2-acetamido-3-methyl-3-sulfanylbutanoic acid — PubCheF Explorer
ZINC ID: 154546 · C7H13NO3S
110
Clay Kosonocky @kosonocky.bsky.social · 18/08/2026
Additionally, we've made share links so you can send interesting molecules to your friends and collaborators. Here are a few from the paper! Nirmatrelvir, viral protease inhibitor: pubchef.org/share/191192...
pubchef.org
(1R,2S,5S)-N-[(1S)-1-cyano-2-[(3S)-2-oxopyrrolidin-3-yl]ethyl]-3-[(2S)-3,3-dimethyl-2-[(2,2,2-trifluoroacetyl)amino]butanoyl]-6,6-dimethyl-3-azabicyclo[3.1.0]hexane-2-carboxamide — PubCheF Explorer
ZINC ID: 1911923694 · C23H32F3N5O4
110
Clay Kosonocky @kosonocky.bsky.social · 18/08/2026
I've really enjoyed browsing molecules by doing a "Wikipedia rabbit hole"-style approach where I navigate by clicking on the predicted functions It feels like a very good way to explore chemical space and gain intuition on chemical structure-function relationships
110
Clay Kosonocky @kosonocky.bsky.social · 18/08/2026
We also added a very fast Morgan fingerprint structure search in case you want to find molecules close to, or far away from, other molecules!
110
Clay Kosonocky @kosonocky.bsky.social · 18/08/2026
Here, users can search purchasable, in-stock molecules from ZINC20 based on their predicted biological functions
110
Clay Kosonocky @kosonocky.bsky.social · 18/08/2026
One part of this project I'm very excited about is the website we made to explore the PubCheF-1 predictions over 13.2 million molecules! 🧵
151
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
This project was a multi-disciplinary effort between many labs and collaborators who put in months-to-years of work. Huge huge shoutout to everyone involved :) @edwardmarcotte.bsky.social
000
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
You can find a link to our preprint here: biorxiv.org/content/10.6... Enjoy!
biorxiv.org
Learning from human and chemical languages to predict biological function
Understanding how molecular structure encodes biological function remains a grand challenge in drug discovery. Here, we present PubCheF-1, a deep learning model that predicts literature-derived biolog...
110
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
We have also created a website where you can explore the predictions over ZINC20 for use in your own research! All of the predictions are completely open to the public and we hope that it is a useful (and fun) resource for the community pubchef.org
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
Overall, we believe that PubCheF-1 demonstrates that leveraging the relationships in scientific literature is a promising avenue to discover bioactive molecules. Furthermore, combining modalities such as structure, omics, and literature will only further improve discovery
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
Furthermore, we tested the same hits in neutropenic mouse models and found, once again, that they reduce the bacterial infection burden back to baseline
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
Since this alone doesn't answer questions of in vivo efficacy, we next tested whether the hits can rescue animals from infection. In a wax moth larvae model system, we found a marked increase in survival when our hits were administered with β-lactam antibiotics
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
The previous assays were in model systems, so we tested whether the hits work on various clinical isolates. Four of the inhibitors hits that we tested successfully sensitized numerous isolates to β-lactam antibiotics, further demonstrating their efficacy
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
To confirm the mechanisms of these inhibitors, we obtained crystal structures for two of the hits in complex with KPC-3. This showed that they bind covalently to the active-site serine residue, in agreement with existing boronate β-lactamase inhibitors
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
Given that the PubCheF dataset contained papers describing cell-based screens, we reasoned that these hits might have favorable bacterial entry properties too. We thus screened these in E. coli expressing β-lactamases and found that 8/110 hits caused a sig reduction in survival
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
13 of the 110 inhibitors caused a significant reduction in beta-lactamase activity in a pure protein screen, testing both serine- and metallo-lactamases, corresponding to a screening hit rate of 11.8%
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
We predicted functional profiles for all 13.2M in-stock molecules from ZINC20 (which we have made publicly available, see pubchef.org), and filtered them down for structurally novel β-lactamase inhibitors. In total, we spent a mere $4K on 110 potential inhibitors
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
This was promising, but the real test is whether it works in actual biology We thus decided to focus on discovering inhibitors targeting β-lactamases, a class of enzymes that cause antibiotic resistance and will cause millions of deaths in the next few decades if unaddressed
110
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
Importantly, PubCheF-1 can be used to predict novel inhibitors. As an in silico test of this we used Boltz-2, a structure-based affinity predictor, and found moderate concordance with it across 5 different protein targets, even for novel inhibitor structures
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
Layer-wise relevance propagation showed that PubCheF-1 even converged to reasonable mechanistic understandings of drug-target interactions in some cases 1. Warhead in nirmatrelvir that cov bonds to Mpro 2. Hydroxyl group in the prodrug zidovudine 3. Tryptophan group in LSD
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
Explainability analyses showed that multi-label training (subpanel B) allowed PubCheF-1 to better predict function by learning the semantic structure of the functional labels, which is further evidenced by its correlations being found directly in the model weights (subpanel C)
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
PubCheF-1 was found to be state-of-the-art in the task of function prediction, outperforming similar function annotators as well as the same PubCheF-1 model trained on a dataset derived from ChEMBL
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
We used LLMs to mine PubMed to create a dataset of 1.2M chemicals and annotations of their biological functions. We dubbed this the PubMed-derived Chemical Function (PubCheF) dataset, and used it to train PubCheF-1, which predicts functional profiles given molecular structures
100
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
Can we use human language to predict novel bioactive molecules that work in vivo? In our new paper, we trained a deep learning model on a dataset derived from PubMed text and used it to discover novel antibiotic co-inhibitors that reduced infection burden, even in mouse models!
143
Reposted by Clay Kosonocky
bioRxiv Bioinfo @biorxiv-bioinfo.bsky.social · 14/08/2026
Learning from human and chemical languages to predict biological function www.biorxiv.org/content/10.64898/20…
121
Clay Kosonocky @kosonocky.bsky.social · 17/08/2026
Great question! We intentionally removed molecules from high-throughput screens since they were definitely polluting our dataset. Check out Figure S1 in the paper
010
Clay Kosonocky @kosonocky.bsky.social · 02/07/2026
I'll be in Seoul next week for ICML 2026! 🇰🇷 Message me if you'd like to meet up for food, coffee, or drinks! Would love to talk all things proteins, small molecules, and biology in general 🧬
000
Clay Kosonocky @kosonocky.bsky.social · 19/05/2026
Apple Podcasts: podcasts.apple.com/us/podcast/a... Spotify: open.spotify.com/episode/1uMQ... Youtube: www.youtube.com/watch?v=jQR8...
podcasts.apple.com
AI Is Designing the Next Cancer Fighter | EP.53
Podcast Episode · Hidden Layers: AI and the People Behind It · May 14 · 42m
010
Clay Kosonocky @kosonocky.bsky.social · 19/05/2026
I also recently discussed the Bits to Binders competition with my co-organizers and Ron Green on the KUNGFU.AI Hidden Layers podcast! It was a very fun conversation 🤠🧬 (links below)
120
Clay Kosonocky @kosonocky.bsky.social · 19/05/2026
Read the full story here: www.conjectureblog.com/p/how-i-ende...
conjectureblog.com
How I ended up running a worldwide protein design competition
The origin story of the BioML Society
010
Clay Kosonocky @kosonocky.bsky.social · 19/05/2026
I’ve been asked a few times how I got started running a worldwide protein design competition. The short answer is that I never intended to start this project at all: it just sort of happened. The more correct answer requires me to tell the origin story of the BioML Society [1/2]
140
Clay Kosonocky @kosonocky.bsky.social · 21/04/2026
We hope this resource helps inform you on your future protein design campaigns! Enjoy! Link to the open access paper: sciencedirect.com/science/arti...
sciencedirect.com
Closing the loop: Experimentally validated methods in artificial intelligence–driven protein design
Artificial intelligence (AI) has reshaped protein design by enabling models trained on large-scale sequence and structure data to generate proteins wi…
030
Clay Kosonocky @kosonocky.bsky.social · 21/04/2026
We then dive deeper into binders, antibodies, and enzymes. In these sections, we discuss the nuances of each task and provide tables of experimental outcomes from all of the models released alongside wet lab validation that we could find
120
Clay Kosonocky @kosonocky.bsky.social · 21/04/2026
We believe that wet lab validation is the foundation for improving AI-driven protein design models. In our review, we walk through the protein design stack from data collection and modeling through to wet lab validation
111
Clay Kosonocky @kosonocky.bsky.social · 21/04/2026
Have you wondered what the wet lab success rates are for current AI-driven protein design models? Look no further! In our new review, @kevinkaichuang.bsky.social @avapamini.bsky.social, @sarahalamdari.bsky.social, and I report wet lab success rates for *over 200* different protein design tasks 🧬💻
13214
Clay Kosonocky @kosonocky.bsky.social · 09/03/2026
Huge shoutout to Twist for making the competition possible!
031
Reposted by Clay Kosonocky
Edward Marcotte @edwardmarcotte.bsky.social · 05/03/2026
Amazing work from Clay @kosonocky.bsky.social, Alex Abel and LEAH labs, and all the collaborators and participants in the Bits2Binders AI CAR-T therapy protein design competition. The results are in! bsky.app/profile/koso...
032
Reposted by Clay Kosonocky
Claus Wilke @clauswilke.com · 04/03/2026
Results from an impressive world-wide binder design competition.
093
Clay Kosonocky @kosonocky.bsky.social · 04/03/2026
If you want, you can check out the data yourself! We made it as accessible as possible :) github.com/kosonocky/bi...
github.com
GitHub - kosonocky/bits-to-binders
Contribute to kosonocky/bits-to-binders development by creating an account on GitHub.
010
Clay Kosonocky @kosonocky.bsky.social · 04/03/2026
Link to the preprint below! www.biorxiv.org/content/10.6...
biorxiv.org
131
Clay Kosonocky @kosonocky.bsky.social · 04/03/2026
And finally, huge thank you to all of participating teams and competitors whose designs were the foundation of this effort! ❤️
100
Clay Kosonocky @kosonocky.bsky.social · 04/03/2026
And huge shoutout to my amazing co-authors Alex Abel, Aaron Feller, Amanda Cifuentes Rieffer, Phillip Woolley, Jakub Lala (@jakublala.bsky.social), Daryl Barth, Ty Gardner, Prof Steve Ekker, Prof Andy Ellington, Wes Wierson, and Prof Edward Marcotte (@edwardmarcotte.bsky.social)
110
Clay Kosonocky @kosonocky.bsky.social · 04/03/2026
This competition was made possibly by many fantastic collaborators and industry partners. Huge thanks to LEAH Labs, @adaptyv.bio, @twistbioscience.com, TACC, @modal-labs.bsky.social, Lonza, ScaleReady, VWR, KUNGFU.AI, Maker Clinic, Synthia, Nucleate AI in BIotech, and the BioML Society
130
Clay Kosonocky @kosonocky.bsky.social · 04/03/2026
We hope this is useful for the protein design community! We also provide extensive detail on how the competition was organized, a list of all 400 metrics, and all competitor methods in detail
120
Clay Kosonocky @kosonocky.bsky.social · 04/03/2026
In summary, we find that most of the failures seem to have been caused by impaired translation and protein expression. We believe that optimizing sequence-level properties for expression is just as important as the structure-centric task of binding
121
Clay Kosonocky @kosonocky.bsky.social · 04/03/2026
In contrast, filtering out <0.50 Boltz-1 ipTM only marginally increased the recovery by 0.7% and proliferation by 0.3%. This removes two non-functional designs from the top 10, but also removes two broadly functional designs with binding affinity
110
Clay Kosonocky @kosonocky.bsky.social · 04/03/2026
If we removed seqs with: ≥60% GC and <1.9 DNA entropy ≥45 AA repeats (DNA) ≥8 EE repeats ≥30% K+E alpha helix We would remove 4,600 designs, increase recovery from 57% to 81%, and CD20-specific proliferation from 5.9% to 7.6% while removing 2/3 non-functional top 10 designs
131