Sign in

Samuel Sledzieski

@samsl.io
779 followers 256 following 73 posts

Research Fellow @flatironinstitute.org @simonsfoundation.org Formerly @csail.mit.edu @msftresearch.bsky.social @uconn.bsky.social Computational systems x structure biology | he/him | samsl.io | 👨🏼‍💻

PostsRepliesMedia
Samuel Sledzieski @samsl.io · 28/04/2026
MIMIC also treats experimental context as a first-class modality. Given assay + cellular context in natural language, it predicts condition-specific RNA reactivity better than sequence-only baselines, improving downstream RNA structure modeling.
Three-panel comparison of RNA 2D structure prediction. Top (orange): RNA sequence alone fed to ViennaRNA yields a branched 2D structure with F1 = 0.404. Middle (gray): RNA sequence plus experimental chemical reactivity data fed to ViennaRNA yields the reference structure. Bottom (blue): RNA sequence plus MIMIC-predicted reactivity fed to ViennaRNA yields a structure with F1 = 0.987, closely matching the experimental reference.
110
Samuel Sledzieski @samsl.io · 28/04/2026
Protein design makes the multimodal advantage clear. Backbone geometry and surface chemistry provide different, complementary constraints on protein function. Conditioning on both, MIMIC generates diverse, high-confidence sequences with strong in silico binding support.
Left: three strip plots showing TM-score vs. WT (median 0.89), MaSIF surface similarity vs. WT (median 0.91), and AF3 cofolding iPTM (median 0.81) for 37 MIMIC-designed PD-L1 sequences. Right: PyMOL-style structural visualization overlaying a MIMIC-designed protein (blue) on the PD-L1 binding partner (gray), with the interface region highlighted in red.
100
Samuel Sledzieski @samsl.io · 28/04/2026
Multimodal conditioning also improves design. MIMIC doesn't just predict aberrant splicing, it can design around it. For a pathogenic mutation, it proposes corrective edits that suppress cryptic exon inclusion while keeping the disease-causing mutation fixed.
Left panel: schematic of the MIMIC RNA design pipeline, showing a mutated RNA sequence conditioned on wild-type splice pattern and wild-type phylogenetic conservation scores to produce a designed sequence. Right panel: four stacked line plots showing PhyloP scores, MIMIC phyloP VEP scores (C>T splice-altering vs. C>A non-splice-altering), and SpliceAI acceptor and donor probabilities (unconditioned vs. PhyloP-conditioned) across positions relative to the HBB IVS-II-654 C>T pathogenic mutation, illustrating that conditioning on conservation suppresses the cryptic splice site.
110
Samuel Sledzieski @samsl.io · 28/04/2026
Another example is splicing. MIMIC does something most splice models can't: isoform-aware splice prediction. By conditioning on transcript boundaries, it can recover the full transcript-specific splice structure, not just score donor/acceptor sites in isolation.
Left: four horizontal bar charts showing AUPR for gene-level and transcript-level splice site prediction (coding and non-coding) comparing MIMIC, AlphaGenome, SpliceAI, and NT3. MIMIC (dark blue) leads all comparisons; a light blue bar shows MIMIC with TSS+TES conditioning. Right: line plot of splice site probability versus transcript position for SPRY1, comparing ground truth, unconditioned MIMIC prediction, and conditioned MIMIC prediction, with donor and acceptor site calls marked. After TSS+TES conditioning, false positive predictions decrease.
100
Samuel Sledzieski @samsl.io · 28/04/2026
What does this training paradigm buy you? MIMIC learns representations that are SOTA across both RNA and protein downstream benchmarks, and multimodal conditioning consistently improves sequence reconstruction in both nucleic-acid and amino-acid settings.
Two dot-plot benchmark comparisons. Left (PFMBench): MIMIC (dark blue) versus ~13 protein language model baselines (gray) across 11 tasks spanning function, structure, interaction, and developability. MIMIC leads or is competitive on most tasks. Bottom bar chart shows win rates of each baseline against MIMIC; all are below 50%. Right (mRNABench): Same format for 7 RNA/multimodal tasks. MIMIC leads on most; Evo2 and Orthrus show the highest win rates but remain below 50%.
100
Samuel Sledzieski @samsl.io · 28/04/2026
Biological data is complex, and training MIMIC required a new substrate. We built LORE: an aligned multimodal dataset connecting nucleic acid, protein, evolutionary, structural, regulatory, and experimental/context signals within shared biomolecular states.
Diagram titled "LORE: Multi-modal and multi-source data alignment." Left: a table mapping transcript IDs to UniProt IDs with checkmarks and X marks indicating data availability across modalities (PhyloP, RNA chemical probing, splice pattern, backbone structure). Right: expanded view of a single entry (ENST000012345 / P04637) showing linked RNA and protein data including sequence, PhyloP conservation track, splice pattern diagram, and AlphaFold structure.
132
Samuel Sledzieski @samsl.io · 28/04/2026
MIMIC is built so any subset of modalities can be observed, and any subset can be generated Sequence → prediction is only one case You can also go the other way: use structure, splicing, or assay context to constrain the sequences compatible with a biological state.
Schematic of MIMIC's split-track input encoding and encoder-decoder architecture. Left panel shows two input tracks: nucleic acid (DNA sequence + conservation scores summed per token) and protein (amino acid residues + backbone structure summed per token), plus cellular context and gene taxonomy tokens. Right panel shows selected tokens fed into an encoder, producing a multimodal embedding passed to a decoder that outputs predictions.
111
Samuel Sledzieski @samsl.io · 28/04/2026
Introducing MIMIC: a new foundation model trained natively across DNA, RNA and proteins. MIMIC is multimodal and generative: it can use structure, regulation, evolution, and experimental context to infer missing biology or design new sequences. 🧵⬇️
1209
Samuel Sledzieski @samsl.io · 22/07/2025
In v0.3.0, we solve *both* of these issues, enabling efficient inference on both personal computers and multi-GPU HPC systems. The secret? Our new Blocked Multi-GPU Parallel Inference (BMPI) procedure, led by Daniel Schaffer (github.com/schafferde).
100
Samuel Sledzieski @samsl.io · 23/06/2025
Our key hypothesis is that these motions carry information about allosteric networks, despite not explicitly measuring larger conformational change. Inspired by terrific work from Federica Maschietto, we show that in silico predictions of these networks closely matches DMS results in KRAS.
(a) Workflow showing use of RocketSHP to predict correlations between residue motion, network construction and clustering, and finally those clusters on the structure of a protein (b) Scatter plot showing betweenness centrality of each residue in KRAS, colored by variance in centrality when that residue is measured as part of an in silico deep mutational scan. (c) Corresponding scatter plot showing the true folding DDG of KRAS in a deep mutational scan, colored by variance of each residue. The peaks in this plot correspond strongly with the peaks in the previous plot.
100
Samuel Sledzieski @samsl.io · 23/06/2025
And we think we can! Thanks to excellent data curation efforts from @hkws.bsky.social, we show that RocketSHP-predicted fluctuations correlate well with experimental hetNOE measurements that capture fast + local movement.
(a) Scatter plot showing RocketSHP predicted RMSF and hetNOE, with a correlation of -0.67 and a confidence interval of [-0.68, -0.66]. (b) Line plot showing per-residue hetNOE or ReLU(1 - RMSF predicted) for the protein ACRIIA4, showing agreement between the two. (c) Two structures of ACRIIA4, colored by hetNOE or ReLU(1 - RMSF predicted), showing agreement between the two.
100
Samuel Sledzieski @samsl.io · 05/04/2025
I constantly reference their Supp. Fig. 3 when I’m trying to describe the method to people— maybe not exactly what you’re looking for but imo extremely informative.
1110
Samuel Sledzieski @samsl.io · 12/03/2025
Finally, we show how users can use MINT through two case studies: MINT predictions align with 23/24 experimentally validated oncogenic PPIs impacted by cancer mutations, and MINT estimates SARS-CoV-2 antibody cross-neutralization with high accuracy.
A cartoon shows how MINT can be used to predict whether mutations lead to a loss or conservation of PPI, compared with oncogenic PPIs discovered in Cheng et al. The bottom panel shows several example mutations, where MINT correctly predicts for all except one whether the mutation will conserve or disrupt binding of this complex.A cartoon shows how MINT can be used to predict SARS-CoV-2 cross-neutralization with different variant antibodies against different variants. The bottom panel shows that for multiple Omicron variants, MINT achieves high AUPRC at separating neutralizing and non-neutralizing antibodies.
110
Samuel Sledzieski @samsl.io · 12/03/2025
We show MINT works for diverse and challenging interaction tasks! It outperforms IgBert & IgT5 in predicting antibody binding affinity and estimating antibody expression. Fine-tuning MINT beats TITAN, PISTE and other TCR-specific models on
TCR–Epitope and TCR–Epitope–MHC interaction prediction.
Cartoons showing MINT applied for antibody interaction prediction and TR-epitope prediction. Box plots show that MINT outperforms other baseline methods on all data sets.
110
Samuel Sledzieski @samsl.io · 12/03/2025
MINT sets new benchmarks! It outperforms existing PLMs in:
 ✅ Binary PPI classification
 ✅ Binding affinity prediction
 ✅ Mutational impact assessment Across yeast, human, & complex PPIs, we see up to 29% gains vs. baselines! 📈
Diagram showing MINT used to predict binary binding, binding affinity, or variant effect. On the SKEMPI and gold-standard (Bernett) PPI data sets, box plots show that MINT achieves the highest PCC and AUPRC respectively.
110
Samuel Sledzieski @samsl.io · 12/03/2025
MINT is built on ESM-2 but adds a cross-chain attention mechanism to preserve inter-sequence relationships. We trained MINT on 96 million high-quality PPIs (from STRING-db). Instead of masked language modeling on single sequences, we now capture interaction-specific signals.
Three panel figure showing our approach (entitled MINT), the implementation as a modification of ESM with cross-chain attention, and several applications including general complex prediction, antibody binding, TCR-epitope-MHC interactions, mutational effect on cancer PPI, and cross-neutralization of SARS-CoV-2 variant effects
110
Samuel Sledzieski @samsl.io · 12/03/2025
Traditional PLMs struggle with PPIs since they model proteins independently. Previous approaches concatenated embeddings or sequences—leading to lost inter-residue context. We fix this with 🌿 MINT, which allows multiple interacting sequences as input.
Diagram showing traditional approaches for applying protein language models to multiple sequences, either by concatenating output embeddings or input sequences
120
Samuel Sledzieski @samsl.io · 03/03/2025
In my personal favorite figure, we annotate a diverse set of ~3.5 million sequences across kingdoms to see the distribution of symmetry use across life -- another example of the cool things you can do with PLM fine-tuning + lightweight classification + genome scale inference!
Percentage of proteins predicted to adopt different homo-oligomer symmetries across several different kingdoms. Icosahedral symmetry is overrepresented in viruses, and C4/C5 is overrepresented in invertebrates.
000
Samuel Sledzieski @samsl.io · 15/02/2025
I'm at @biophysicalsoc.bsky.social #BPS2025 with a bunch of folks from the Structural and Molecular Biophysics @flatironinstitute.org group-- come say hi and check out our posters/talks! @sonyahanson.bsky.social @pilarcossio.bsky.social @miroastore.bsky.social
095
Samuel Sledzieski @samsl.io · 14/12/2024
If you're at @neuripsconf.bsky.social @workshopmlsb.bsky.social tomorrow, make sure you stop by our poster presenting a paired-protein language model for modeling protein-protein interactions. Varun did a great job spearheading this work! #NeurIPS2024
Poster for "Learning the language of protein-protein interactions with ESM-Multimer"
180