Samuel Sledzieski @samsl.io · 28/04/2026MIMIC also treats experimental context as a first-class modality. Given assay + cellular context in natural language, it predicts condition-specific RNA reactivity better than sequence-only baselines, improving downstream RNA structure modeling. 110
Samuel Sledzieski @samsl.io · 28/04/2026Protein design makes the multimodal advantage clear. Backbone geometry and surface chemistry provide different, complementary constraints on protein function. Conditioning on both, MIMIC generates diverse, high-confidence sequences with strong in silico binding support. 100
Samuel Sledzieski @samsl.io · 28/04/2026Multimodal conditioning also improves design. MIMIC doesn't just predict aberrant splicing, it can design around it. For a pathogenic mutation, it proposes corrective edits that suppress cryptic exon inclusion while keeping the disease-causing mutation fixed. 110
Samuel Sledzieski @samsl.io · 28/04/2026Another example is splicing. MIMIC does something most splice models can't: isoform-aware splice prediction. By conditioning on transcript boundaries, it can recover the full transcript-specific splice structure, not just score donor/acceptor sites in isolation. 100
Samuel Sledzieski @samsl.io · 28/04/2026What does this training paradigm buy you? MIMIC learns representations that are SOTA across both RNA and protein downstream benchmarks, and multimodal conditioning consistently improves sequence reconstruction in both nucleic-acid and amino-acid settings. 100
Samuel Sledzieski @samsl.io · 28/04/2026Biological data is complex, and training MIMIC required a new substrate. We built LORE: an aligned multimodal dataset connecting nucleic acid, protein, evolutionary, structural, regulatory, and experimental/context signals within shared biomolecular states. 132
Samuel Sledzieski @samsl.io · 28/04/2026MIMIC is built so any subset of modalities can be observed, and any subset can be generated Sequence → prediction is only one case You can also go the other way: use structure, splicing, or assay context to constrain the sequences compatible with a biological state. 111
Samuel Sledzieski @samsl.io · 28/04/2026Introducing MIMIC: a new foundation model trained natively across DNA, RNA and proteins. MIMIC is multimodal and generative: it can use structure, regulation, evolution, and experimental context to infer missing biology or design new sequences. 🧵⬇️ 1209
Samuel Sledzieski @samsl.io · 22/07/2025In v0.3.0, we solve *both* of these issues, enabling efficient inference on both personal computers and multi-GPU HPC systems. The secret? Our new Blocked Multi-GPU Parallel Inference (BMPI) procedure, led by Daniel Schaffer (github.com/schafferde). 100
Samuel Sledzieski @samsl.io · 23/06/2025Our key hypothesis is that these motions carry information about allosteric networks, despite not explicitly measuring larger conformational change. Inspired by terrific work from Federica Maschietto, we show that in silico predictions of these networks closely matches DMS results in KRAS. 100
Samuel Sledzieski @samsl.io · 23/06/2025And we think we can! Thanks to excellent data curation efforts from @hkws.bsky.social, we show that RocketSHP-predicted fluctuations correlate well with experimental hetNOE measurements that capture fast + local movement. 100
Samuel Sledzieski @samsl.io · 05/04/2025I constantly reference their Supp. Fig. 3 when I’m trying to describe the method to people— maybe not exactly what you’re looking for but imo extremely informative. 1110
Samuel Sledzieski @samsl.io · 12/03/2025Finally, we show how users can use MINT through two case studies: MINT predictions align with 23/24 experimentally validated oncogenic PPIs impacted by cancer mutations, and MINT estimates SARS-CoV-2 antibody cross-neutralization with high accuracy. 110
Samuel Sledzieski @samsl.io · 12/03/2025We show MINT works for diverse and challenging interaction tasks! It outperforms IgBert & IgT5 in predicting antibody binding affinity and estimating antibody expression. Fine-tuning MINT beats TITAN, PISTE and other TCR-specific models on TCR–Epitope and TCR–Epitope–MHC interaction prediction. 110
Samuel Sledzieski @samsl.io · 12/03/2025MINT sets new benchmarks! It outperforms existing PLMs in: ✅ Binary PPI classification ✅ Binding affinity prediction ✅ Mutational impact assessment Across yeast, human, & complex PPIs, we see up to 29% gains vs. baselines! 📈 110
Samuel Sledzieski @samsl.io · 12/03/2025MINT is built on ESM-2 but adds a cross-chain attention mechanism to preserve inter-sequence relationships. We trained MINT on 96 million high-quality PPIs (from STRING-db). Instead of masked language modeling on single sequences, we now capture interaction-specific signals. 110
Samuel Sledzieski @samsl.io · 12/03/2025Traditional PLMs struggle with PPIs since they model proteins independently. Previous approaches concatenated embeddings or sequences—leading to lost inter-residue context. We fix this with 🌿 MINT, which allows multiple interacting sequences as input. 120
Samuel Sledzieski @samsl.io · 03/03/2025In my personal favorite figure, we annotate a diverse set of ~3.5 million sequences across kingdoms to see the distribution of symmetry use across life -- another example of the cool things you can do with PLM fine-tuning + lightweight classification + genome scale inference! 000
Samuel Sledzieski @samsl.io · 15/02/2025I'm at @biophysicalsoc.bsky.social #BPS2025 with a bunch of folks from the Structural and Molecular Biophysics @flatironinstitute.org group-- come say hi and check out our posters/talks! @sonyahanson.bsky.social @pilarcossio.bsky.social @miroastore.bsky.social 095
Samuel Sledzieski @samsl.io · 14/12/2024If you're at @neuripsconf.bsky.social @workshopmlsb.bsky.social tomorrow, make sure you stop by our poster presenting a paired-protein language model for modeling protein-protein interactions. Varun did a great job spearheading this work! #NeurIPS2024 180