Sign in

Anthony Gitter

@anthonygitter.bsky.social
96 followers 44 following 55 posts

Computational biologist; Associate Prof. at University of Wisconsin-Madison; Jeanne M. Rowe Chair at Morgridge Institute gitterlab.org

PostsRepliesMedia
Anthony Gitter @anthonygitter.bsky.social · 02/10/2026
It's appealing to use the typically discarded directed evolution data as a resource for enzyme engineering. This approach isn't a panacea though. We discuss limitations of the pooled sequencing and have a partner study coming soon all about the machine learning challenges. 4/
010
Anthony Gitter @anthonygitter.bsky.social · 02/10/2026
The machine learning model has to account for different meanings of binned activity per directed evolution round and the binned, ordinal labels. We use ordinal regression with a evolutionary score as an additional feature. That model nominates new SadX mutations that we test. 3/
Recommendation of single mutants from the MLP model. A) Model architecture: input features in the MLP are the variant amino acid sequence and a scalar evolutionary density score derived from homologs of SadA (only promoted positions of the original directed evolution are shown). A learnable bias term per library adjusts predictions to account for different parent contexts. B) Top 10 single mutant predictions ranked by predicted probability of the High bin in the 3-VRL parent. The single observation column notes if the mutation was observed as a single mutation in any library and its corresponding activity.
100
Anthony Gitter @anthonygitter.bsky.social · 02/10/2026
The main idea is to rearray the existing sequence variants from each round of directed evolution based on enzyme activity bins. Then use long read sequencing to map sequences to function bins per round. That gives a training dataset for machine learning. 2/
A) Conventional directed evolution workflow illustrated for conversion of SadX, which favors C-H hydroxylation, to 4-IC, which favors C-H azidation. B) Latent sequence-function data from binned library variants (Dead, Low, Parent, High) is used to train an ordinal regression model to predict the top single mutants for experimental validation. C) Model reaction of 1a used for engineering efforts in A and B.Workflow to obtain sequencing data from three generations of directed evolution data using a single PacBio SMRT cell. A) SadX variants were rearrayed according to reaction conversions determined via B) LCMS analysis; high (red), parent (blue), low (yellow), and no (green) activity bins were determined relative to the standard deviation of the parent enzyme in each round. C) Product yields obtained from enzyme pools obtained from the pooled sequences mirror the initial LCMS high throughput (HT) screening trends. D) PacBio reads for the 3-VRL library variants, ranked by read count on a log-log scale, where each variant's position reflects its relative abundance. Marks on the sequence rank axis indicate the number of variants placed in each bin initially. Line thickness and color indicate variants supported by at least 100 read counts. Bottom rows mark the sequence rank of DNA variants which encode amino acid parents and early stop codons (Figure S9). E) Variants placed via LCMS screening and recovered after PacBio sequencing, stratified by library and bin. Recovered sequences are the subset of placed sequences that were identified by PacBio sequencing, have corresponding activity labels, and pass all quality filters.
100
Anthony Gitter @anthonygitter.bsky.social · 02/10/2026
Our new preprint with the @philromero.bsky.social and @jclewislab.bsky.social labs explores how to extract additional information for protein engineering from a previous directed evolution campaign. We test it with SadX for site-selective C-H azidation. 1/ chemrxiv.org/doi/full/10....
chemrxiv.org
Extracting Latent Sequence-Function Information from Directed Evolution to Design Improved C-H Azidases | ChemRxiv
Directed evolution is a powerful approach for enzyme engineering, but conventional workflows use only a fraction of the information generated during experimental screening. Typically, only top-perform...
111
Anthony Gitter @anthonygitter.bsky.social · 01/10/2026
A reviewer raised the question of how this sequence watermark would provide biosecurity advantages in a DNA synthesis setting over existing digital signatures. I wasn't convinced by the Discussion about that.
000
Anthony Gitter @anthonygitter.bsky.social · 26/08/2026
Didn't you wear this shirt yesterday? No, that was my ortho steric shirt.
011
Anthony Gitter @anthonygitter.bsky.social · 19/08/2026
I also like this report. Stuff like this where a lot of expertise and detail went into the prompt makes me curious about the control experiment where you give Claude Science the same Modal resources and time limit but a minimal prompt: "make me a <target X> binder".
110
Reposted by Anthony Gitter
Philip Romero @philromero.bsky.social · 18/08/2026
What if AI could interact directly with biology? Congrats to Coban, who gave AI the ability to experiment and learn through feedback. Over 25 autonomous rounds, it uncovered the determinants of enzyme specificity. Give AI the ability to experiment, then get out of the way. doi.org/10.64898/202...
doi.org
Learning protein function through autonomous experimental interaction
Biological AI learns primarily from existing observations, but many questions cannot be answered from available data alone. Here we show that AI can instead acquire knowledge by acting directly on biological systems and learning from the consequences. We developed a closed-loop framework in which autonomous agents design protein variants, construct and characterize them in a robotic laboratory, learn from the resulting experimental feedback, and decide what experiments to perform next. We then allowed the system to operate continuously and without human intervention for approximately one month, during which multiple agents independently explored protein sequence space while learning from shared experimental experience. Applied to glycoside hydrolases, the agents discovered enzymes with substantially altered substrate specificity toward non-native sugars and progressively learned the structure of the underlying sequence-function landscape. The resulting experimental experience also revealed determinants of substrate specificity and protein expression that were not specified as learning objectives. These results demonstrate that AI can autonomously interact with biology over extended periods to acquire knowledge through experience, establishing a framework for biological discovery driven by continuous experimental interaction. ### Competing Interest Statement The authors have declared no competing interest. National Institute of General Medical Sciences, 5R01GM150929
1144
Anthony Gitter @anthonygitter.bsky.social · 15/08/2026
@kosonocky.bsky.social this is a very interesting approach. What's your overall assessment of the PMID mappings in PubChem BioAssays? We've been finding high-throughput screening assays that link out to related, incorrect targets for that HTS and HTS without linked PMIDs.
110
Anthony Gitter @anthonygitter.bsky.social · 11/08/2026
This manuscript is a registered report presenting the SPRAS software and our evaluation plans. We are actively looking for feedback before we execute those plans. Email to me and Anna or GitHub issues are welcome: github.com/Reed-CompBio... /5
github.com
GitHub - Reed-CompBio/spras: Signaling Pathway Reconstruction Analysis Streamliner (SPRAS)
Signaling Pathway Reconstruction Analysis Streamliner (SPRAS) - Reed-CompBio/spras
020
Anthony Gitter @anthonygitter.bsky.social · 11/08/2026
Related pathway reconstruction frameworks show growing interest in this problem, and we compare our design to others: - nf-core/diseasemodulediscovery: doi.org/10.1093/bioi... - NetworkCommons: doi.org/10.1093/bioi... /4
110
Anthony Gitter @anthonygitter.bsky.social · 11/08/2026
One core challenge we had to address was setting parameters for different algorithms. They can substantially affect output networks. Machine learning hyperparameter tuning strategies don't work well in this network setting. 3/
Two stage parameter tuning.
110
Anthony Gitter @anthonygitter.bsky.social · 11/08/2026
Over the past decades, dozens of graph algorithms have been applied to this general problem under different formulations. SPRAS unites them and abstracts running many tools with the same Snakemake workflow. Containers manage the conflicting software dependencies. 2/
Distinguishing features of pathway reconstruction algorithms SPRAS supports.
110
Anthony Gitter @anthonygitter.bsky.social · 11/08/2026
Our new preprint with @annamritz.bsky.social's lab presents SPRAS, a framework for running and benchmarking pathway reconstruction algorithms. These are network biology tools that identify relevant subnetworks given omics data and biomolecule interactions. doi.org/10.64898/202... 1/
SPRAS framework for running and evaluating containerized pathway reconstruction algorithms.
130
Anthony Gitter @anthonygitter.bsky.social · 02/08/2026
At the time we concluded that deep learning hadn't yet transformed biomedicine. This was pre-Transformer, pre-AlphaFold, pre-GPTs.
050
Anthony Gitter @anthonygitter.bsky.social · 02/08/2026
10 years ago today @casey.greenelab.com launched our review "Opportunities and obstacles for deep learning in biology and medicine" greenelab.github.io/deep-review/. It was a widely collaborative project written on GitHub, which lead to the creation of manubot.org.
https://github.com/greenelab/deep-review/commit/e1529c48fe2dd83c81cc91a09d3b80fdf40e16bbWe examine applications of deep learning to a variety of biomedical problems—patient classification, fundamental biological processes, and treatment of patients—and discuss whether deep learning will be able to transform these tasks or if the biomedical sphere poses unique challenges. Following from an extensive literature review, we find that deep learning has yet to revolutionize biomedicine or definitively resolve any of the most pressing challenges in the field, but promising advances have been made on the prior state of the art.
192
Anthony Gitter @anthonygitter.bsky.social · 17/07/2026
Wow, that's fantastic
000
Anthony Gitter @anthonygitter.bsky.social · 17/07/2026
Who pays for the CyVerse data storage? The data uploader or someone else? Does MDRepo already solve most of the problems discussed in this commentary www.nature.com/articles/s41...?
nature.com
The need to implement FAIR principles in biomolecular simulations - Nature Methods
In the Big Data era, a change of paradigm in the use of molecular dynamics is required. Trajectories should be stored under FAIR (findable, accessible, interoperable and reusable) requirements to favo...
100
Reposted by Anthony Gitter
Travis Wheeler @wheelerlab.org · 17/07/2026
Thanks for the advert, @martinsteinegger.bsky.social. If you're reading this and you're sitting on a pile of molecular dynamics simulations, please consider contributing them to MDRepo. This is the path to AI for dynamics (And if you're wondering: yes, there was only 1 person in the audience! 🤥)
393
Reposted by Anthony Gitter
Pedro Beltrao @pedrobeltrao.bsky.social · 08/07/2026
The @qedscience.bsky.social "impact" score generated a lot of discussion on ranking preprints, including ideas on multi-dimensional rankings that are user specific. Besides describing the ideas, we can now prototype them (with Claude in this case). Here is Salient salient-sgeu6fs5ua-uc.a.run.app
salient-sgeu6fs5ua-uc.a.run.app
Rising — Salient
1176
Reposted by Anthony Gitter
Morgridge Institute for Research @morgridgeinstitute.bsky.social · 01/07/2026
A great story stemming from collaboration between @anthonygitter.bsky.social and Nate Wlodarchak, now in Colorado at the Rocky Mountain Regional VA Medical Center. ⬇️
011
Anthony Gitter @anthonygitter.bsky.social · 28/06/2026
Genomic and other biological data are in scope for this data scientist position if anyone in that area is looking.
000
Anthony Gitter @anthonygitter.bsky.social · 22/06/2026
Similarly, OpenRouter if you want flexibility as new open weights models come out, don't want to self host, and are okay with their pricing model. The commercial providers theoretically have ways to provide more access to research organizations, but I haven't figured out how to get it.
120
Reposted by Anthony Gitter
Hannah Wayment-Steele @hkws.bsky.social · 01/06/2026
In the W-S lab's first preprint, we describe how genomic language models know something about RNA thermodynamics. Though we think this is cool, things get tricky! A growing practice for interpreting LMs is to perturb input tokens, often called "Categorical Jacobian": 👇
13314
Reposted by Anthony Gitter
Peter Škrinjar @peterskrinjar.bsky.social · 11/05/2026
Now published in NSMB! Paper: doi.org/10.1038/s415... Full PDF: rdcu.be/fhBtI Overview of additions since the preprint👇 (1/5)
doi.org
Evaluating generalization in protein–ligand cofolding methods - Nature Structural & Molecular Biology
This work introduces the Runs N’ Poses dataset for benchmarking deep learning methods on the protein–ligand complex prediction task. It shows that current methods rely on memorization, challenging the...
23915
Reposted by Anthony Gitter
Martin Pacesa @martinpacesa.bsky.social · 29/04/2026
I am happy to share a review I recently wrote on the design of peptide binders. It gives an overview of experimentally validated tools and discusses the challenges of why peptide design is more difficult than the design of classical protein binders. www.chimia.ch/chimia/artic...
17522
Anthony Gitter @anthonygitter.bsky.social · 04/04/2026
Fantastic analysis from the OpenADMET team (Maria Castellanos, Hugo MacDermott-Opeskin) showing that the zero-shot ADMET models ADMETlab 3.0 and ADMET-AI generalize poorly to their recent OpenADMET-ExpansionRx Blind Challenge data openadmet.ghost.io/zero-shot-ex...
openadmet.ghost.io
Lessons Learned from the OpenADMET-ExpansionRx Blind Challenge: Can We Trust Zero-Shot ADMET Predictions?
Maria Castellanos Hugo MacDermott-Opeskin It’s been more than a month since the OpenADMET-ExpansionRx challenge wrapped up, but the conversation is just getting started. Launched on October 27, 2025...
000
Reposted by Anthony Gitter
Torsten Schwede @torstenschwede.bsky.social · 13/03/2026
Is #AI hitting a plateau in structure prediction? Help us find out at CASP17! 🧪🧬 Calling for Targets: Immune Complexes, protein - ligand complexes, RNA/DNA, conformational ensembles, membrane proteins, viral origins, and large complexes. The Rule of Thumb: If AF3 can’t model it, we want it.
The Critical Assessment of Structure Prediction (CASP) experiment is calling for prediction targets: Immune Complexes, Organic Ligand-Protein Complexes, Nucleic Acids and Complexes, Conformational Ensembles, Difficult Protein Structures and Complexes. 
Rule of Thumb: If AlphaFold3 can generate a high-quality model, it is likely not a CASP-grade challenge. If it struggles, we want it.
24935
Reposted by Anthony Gitter
Pedro Beltrao @pedrobeltrao.bsky.social · 04/03/2026
We have started a project trying to predic the interactions/structures of all yeast protein pairs using an AlphaFold pooling approach. We are making the current dataset open and we welcome collaborations. www.evocellnet.com/2026/03/mapp...
evocellnet.com
Mapping the yeast atructural interactome with AlphaFold3: an open call for collaboration
We are excited to announce the early-stage release of our S. cerevisiae  structural interactome mapping project. Using AlphaFold3 (AF3), w...
69853
Reposted by Anthony Gitter
Yun S. Song @yun-s-song.bsky.social · 21/02/2026
Can we simulate realistic evolutionary trajectories and “replay the tape of life”? In this work, we propose a flexible, generalizable deep learning framework for modeling how the entire protein sequence evolves over time while capturing complex interactions across sites. 1/n doi.org/10.64898/202...
doi.org
38735
Anthony Gitter @anthonygitter.bsky.social · 12/02/2026
👋 from the Nexus I still haven't built up my network here so my following patterns are a narrow slice of my interests.
110
Reposted by Anthony Gitter
Klara Hlouchova lab @hlouchova-lab.bsky.social · 03/11/2025
Can proteins fold and function with half of the amino acid alphabet? Using only 10 residues, we designed stable, mutation-resilient structures—no aromatics or basics involved. A minimalist foundation for ancient biology and synthetic design. tinyurl.com/37t8br4v #ProteinDesign #OriginsOfLife
tinyurl.com
Ancient amino acid sets enable stable protein folds
Early proteins likely arose from a chemically limited set of amino acids available through prebiotic chemistry, raising a central question in molecular evolution: could such primitive compositions yie...
12410
Anthony Gitter @anthonygitter.bsky.social · 23/01/2026
Mingchen replied to me on Twitter that it's also on bioRxiv now www.biorxiv.org/content/10.6...
biorxiv.org
020
Reposted by Anthony Gitter
Milot Mirdita @milot.bsky.social · 20/01/2026
My time in @martinsteinegger.bsky.social's group is ending, but I’m staying in Korea to build a lab at Sungkyunkwan University School of Medicine. If you or someone you know is interested in molecular machine learning and open-source bioinformatics, please reach out. I am hiring! mirdita.org
mirdita.org
Mirdita Lab - Laboratory for Computational Biology & Molecular Machine Learning
Mirdita Lab builds scalable bioinformatics methods.
710756
Reposted by Anthony Gitter
James Fraser @fraserlab.com · 29/12/2025
I'm really excited to break up the holiday relaxation time with a new preprint that benchmarks AlphaFold3 (AF3)/“co-folding” methods with 2 new stringent performance tests. Thread below - but first some links: A longer take: fraserlab.com/2025/12/29/k... Preprint: www.biorxiv.org/content/10.6...
fraserlab.com
Know when to co-fold'em
This is the official web page for the James Fraser Lab at UCSF.
57230
Reposted by Anthony Gitter
Max Fürst @maxfus.bsky.social · 16/12/2025
New preprint🚨 Imagine (re)designing a protein via inverse folding. AF2 predicts the designed sequence to a structure with pLDDT 94 & you get 1.8 Å RMSD to the input. Perfect design? What if I told u that the structure has 4 solvent-exposed Trp and 3 Pro where a Gly should be? Why to be wary🧵👇
46524
Anthony Gitter @anthonygitter.bsky.social · 15/12/2025
Cody also put in a ton of extra work to make the code organized and usable in the GitHub repo: github.com/Anantharaman... It links to a Colab notebook for model inference, training data, and pretrained models.
github.com
GitHub - AnantharamanLab/protein_set_transformer: Protein Set Transformer (PST) framework for training protein-language-model-based genome language models. Inference is possible for viral genomes usin...
Protein Set Transformer (PST) framework for training protein-language-model-based genome language models. Inference is possible for viral genomes using our pretrained viral foundation model. - Anan...
011
Reposted by Anthony Gitter
Karthik Anantharaman @karthik-a.bsky.social · 15/12/2025
Excited for our new paper on a genome language model for viruses in @natcomms.nature.com: "Protein Set Transformer: a protein-based genome language model to power high-diversity viromics"! Led by PhD student Cody Martin in collaboration with @anthonygitter.bsky.social doi.org/10.1038/s414...
doi.org
Protein Set Transformer: a protein-based genome language model to power high-diversity viromics - Nature Communications
A genome language model, Protein Set Transformer, trained on viral datasets, uncovers evolutionary rules of protein content and organization driving precise virus identification, host prediction, and ...
1114
Anthony Gitter @anthonygitter.bsky.social · 12/12/2025
Thanks, I didn't realize Rogue Scholar minted DOIs
000
Anthony Gitter @anthonygitter.bsky.social · 12/12/2025
Use @prereview.bsky.social for preprints and something else for other manuscripts?
010
Anthony Gitter @anthonygitter.bsky.social · 12/12/2025
What are good places to post an unsolicited manuscript peer review these days? I don't have a blog. I read manuscripts across arXiv, bioRxiv, ChemRxiv, OpenReview, random white papers, journals, etc. Do I dump it on Zenodo, post it here, and send it to the authors?
221
Anthony Gitter @anthonygitter.bsky.social · 21/11/2025
Our Assay2Mol manuscript was published at EMNLP 2025 doi.org/10.18653/v1/... See the preprint thread below for a summary of the methodology, results, and code. We added more control experiments in this version related to protein sequence identity and generated molecule size.
000
Anthony Gitter @anthonygitter.bsky.social · 20/11/2025
@hkws.bsky.social and I are creating the Madison AI for Proteins (MAIP) group to discuss early-stage research at monthly meetups, share computational resources, and grow this local community. Visit mad-ai-proteins.github.io to sign up for announcements and watch for our 2026 events.
mad-ai-proteins.github.io
MAIP
Madison AI for Proteins
010
Reposted by Anthony Gitter
Pedro Beltrao @pedrobeltrao.bsky.social · 19/11/2025
This looks like a fantastic resource to study human kinase signalling. So much MS instrument time.
0133
Anthony Gitter @anthonygitter.bsky.social · 14/11/2025
Something fun and sciencey is coming soon to Madison
000
Anthony Gitter @anthonygitter.bsky.social · 23/10/2025
Looks very interesting. Can I think of this like a more extreme form of the evotuning from UniRep or doi.org/10.1101/2024... except it uses one sequence instead of the sequence plus homologs?
doi.org
Protein Language Model Fitness Is a Matter of Preference
Leveraging billions of years of evolution, scientists have trained protein language models (pLMs) to understand the sequence and structure space of proteins aiding in the design of more functional pro...
130
Anthony Gitter @anthonygitter.bsky.social · 10/10/2025
Bioconductor R package: bioconductor.org/packages/MPAC Shiny app to explore results in manuscript: connect.doit.wisc.edu/content/122/
bioconductor.org
MPAC
Multi-omic Pathway Analysis of Cells (MPAC), integrates multi-omic data for understanding cellular mechanisms. It predicts novel patient groups with distinct pathway profiles as well as identifying ke...
000
Anthony Gitter @anthonygitter.bsky.social · 10/10/2025
MPAC uses PARADIGM as the probabilistic model but makes many improvements: - data-driven omic data discretization - permutation testing to eliminate spurious predictions - full workflow and downstream analyses in an R package - Shiny app for interactive visualization
100
Anthony Gitter @anthonygitter.bsky.social · 10/10/2025
The journal version of our Multi-omic Pathway Analysis of Cells (MPAC) software is now out: doi.org/10.1093/bioi... MPAC uses biological pathway graphs to model DNA copy number and gene expression changes and infer activity states of all pathway members.
Overview of the MPAC workflow. MPAC calculates inferred pathway levels (IPLs) from real and permuted CNA and RNA data. It filters real IPLs using the permuted IPLs to remove spurious IPLs. Then, MPAC focuses on the largest pathway subset network with filtered IPLs to compute GO term enrichment, predict patient groups, and identify key group-specific proteins.
121