Sign in

David Ryan Koes

@dkoes.compstruct.org
1.2K followers 760 following 84 posts

Removing barriers to computational drug discovery one bit at a time. Associate Professor in Computational and Systems Biology at the University of Pittsburgh. bits.csb.pitt.edu

PostsRepliesMedia
David Ryan Koes @dkoes.compstruct.org · 09/12/2025
Had an informative and enjoyable time at NeurIPS and MLSB. It was great to see my students present their work and catch up with other CPCB students, both past and present.
080
David Ryan Koes @dkoes.compstruct.org · 07/12/2025
Excited to be at the first independent @workshopmlsb.bsky.social
061
David Ryan Koes @dkoes.compstruct.org · 05/12/2025
If you like to sample from the Boltzmann distribution and are in San Diego for NeurIPS, be sure to check out Rishal's (@rishalchich.bsky.social) poster (#2110). Great work with Nick Boffi (@nmboffi.bsky.social) and Jacky Chen. neurips.cc/virtual/2025... arxiv.org/abs/2507.00846
0123
David Ryan Koes @dkoes.compstruct.org · 19/05/2025
Was proud and honored to hood Dr. Drew McNutt at the Pitt School of Medicine Diploma Ceremony. Here we are rocking both the blue and gold and tartan colors representing our joint Pitt-CMU CompBio PhD program. Congratulations to Drew and the other @cmupittcompbio.bsky.social graduates!
0203
David Ryan Koes @dkoes.compstruct.org · 25/03/2025
Interested in generative modeling and pharmacophores search for SBDD? Check out our talks at #ACSSpring2025
041
David Ryan Koes @dkoes.compstruct.org · 07/03/2025
@standupforscience.bsky.social
071
David Ryan Koes @dkoes.compstruct.org · 15/12/2024
Dang - forget BlueSky makes transparent backgrounds black. There's a hidden message.
150
David Ryan Koes @dkoes.compstruct.org · 15/12/2024
I think this is less important for the train set compared to getting a pristine, possibly manually curated test. A data loader that downsampled duplicated ligands would go a long way to smoothing out problematic systems.
NAG is present in >54,000 PLINDER systems but isn't particularly interesting for predicting protein-ligand binding.
120
David Ryan Koes @dkoes.compstruct.org · 15/12/2024
It turns out that the picomolar ligand BMP (www.rcsb.org/ligand/BMP) is present 30 times in the test set. It isn't clear where this affinity was sourced from and the structures include mutant proteins that presumably should have different binding affinities (but they are all labeled the same).
120
David Ryan Koes @dkoes.compstruct.org · 15/12/2024
Now onto the less good... When @alutsky.bsky.social trained a binding affinity prediction model on PLINDER, he got a negative correlation, which was surprising looking at the predictions (note only 182 test set systems having an affinity annotation).
150
David Ryan Koes @dkoes.compstruct.org · 15/12/2024
Using GraphPocket descriptors (developed in @vratin.bsky.social's MS thesis) there is some separation, but there remains a significant amount of overlap.
140
David Ryan Koes @dkoes.compstruct.org · 15/12/2024
Interestingly, using distances between learned descriptors gets different results. Using DeeplyTough descriptors the test set is as close to the training set as to the removed set.
130
David Ryan Koes @dkoes.compstruct.org · 15/12/2024
Here are persistent homology descriptors calculated by @alutsky.bsky.social (from his MS thesis). This descriptor space has a different scale.
140
David Ryan Koes @dkoes.compstruct.org · 15/12/2024
Using geometric shape descriptors of the pocket, the test set is separated from training set. Using Zernike descriptors calculated by @katztyler.bsky.social we see that the closest system to the test examples have been removed from training and the train set is not close to the test set.
160
David Ryan Koes @dkoes.compstruct.org · 15/12/2024
@workshopmlsb.bsky.social
Huge turnout for MLSB
1140
David Ryan Koes @dkoes.compstruct.org · 06/12/2024
My wife gave this to me today. Good news - it is empty!
090
David Ryan Koes @dkoes.compstruct.org · 29/11/2024
You lost me with the use of the singular…
Five pies, all made by me.
110
David Ryan Koes @dkoes.compstruct.org · 26/11/2024
Unlike previous approaches, SPRINT uses attention to learn from embeddings, resulting in some level of interpretability. It turns out the attention mask is sparse with some preference for binding site residues but more for less conserved residues.
110
David Ryan Koes @dkoes.compstruct.org · 26/11/2024
Results on LIT-PCBA (an admittedly somewhat problematic benchmark) are quite good, although I personally don't see this replacing more traditional techniques so much as supplementing them. We're looking into using it to initialize "deep docking" runs.
200
David Ryan Koes @dkoes.compstruct.org · 26/11/2024
We're looking forward to presenting this @workshopmlsb.bsky.social . This work was conceived and led by a group CPCB (@cmupittcompbio.bsky.social) students and learns a co-embedding of proteins and ligands to support ultra fast virtual screening. arxiv.org/pdf/2411.15418
Dimensionality reduced embedding space showing that ligand and their targets colocate and genomes self organize.SPRINT Enables Interpretable and Ultra-Fast Virtual
Screening against Thousands of Proteomes
Andrew T. McNutt ∗
University of Pittsburgh
anm329@pitt.edu
Abhinav K. Adduri∗
CMU, Arc Institute
abhinav.adduri@arcinstitute.org
Caleb N. Ellington∗
Carnegie Mellon University
cellingt@cs.cmu.edu
Monica T. Dayao
Carnegie Mellon University
mdayao@cs.cmu.edu
Eric P. Xing
CMU, MBZUAI, Petuum Inc.
epxing@cs.cmu.edu
Hosein Mohimani
Carnegie Mellon University
hoseinm@cs.cmu.edu
David R. Koes
University of Pittsburgh
dkoes@pitt.edu
1164
David Ryan Koes @dkoes.compstruct.org · 25/11/2024
This is why I've been looking at profilin - there are ALS linked mutations that open up a destabilizing pocket (yellow). Boltz-1 gets the pocket right (unsurprising - it is in the training set), but not a single ligand gets docked there. Maybe this means it is "undruggable", but I doubt it.
7292
David Ryan Koes @dkoes.compstruct.org · 22/11/2024
Here is how Boltz-1 (green), DynamicBind (magenta), and GNINA (blue) dock a collection of random molecules. GNINA, using a classical sampling algorithm (MCMC) hits all concave regions while the ML samplers have distinct preferences. Boltz is the most likely to induce a fit.
04215
David Ryan Koes @dkoes.compstruct.org · 21/11/2024
Here's how the ligands dock. Clearly there are some preferred binding sites.
020
David Ryan Koes @dkoes.compstruct.org · 21/11/2024
Some fun with Boltz-1 (www.biorxiv.org/content/10.1...). I generated 1000 samples profilin (green) and compare them to the NMR structures (cyan). Samples were generated by docking 1000 random ligands. NMR structures show more conformational diversity.
34110
David Ryan Koes @dkoes.compstruct.org · 01/05/2024
Excellent work from Ian Dunn x.com/ian_dunn_/st...
040