Sign in

Fabien Plisson

@fabienplisson.bsky.social
1.3K followers 1.3K following 443 posts

BioDesign, Machine Learning, Drug Discovery | Rosenkranz Award 2021 | Dad | Polyglot | Capybarist | plissonf.github.io Founding ingeniebio.com ORCID 0000-0003-224

PostsRepliesMedia
Fabien Plisson @fabienplisson.bsky.social · 02/07/2026
ML is transforming protein design, but hidden biases remain overlooked. @danielaldas.bsky.social & I found that predictors trained on global datasets, such as AMPs, are influenced by uneven structural representation and data leakage. #ProteinDesign #ML #Bias #ExplainableAI doi.org/10.64898/202...
doi.org
Structural bias in machine learning-guided peptide design
Machine learning continues to accelerate peptide and protein design through the rapid prediction and generation of sequences with desired characteristics. Many applications focus on predicting propert...
070
Reposted by Fabien Plisson
Drew Jolly @amjolly.bsky.social · 08/08/2025
NSF Invests Nearly $32M to Accelerate Novel AI-Driven Approaches in Protein Design ow.ly/NJzp50WCc00 #NSF #AIwire
ow.ly
NSF Invests Nearly $32M to Accelerate Novel AI-Driven Approaches in Protein Design
Aug. 8, 2025 -- The U.S. National Science Foundation Directorate for Technology, Innovation and Partnerships (NSF TIP) announced an inaugural investment
041
Reposted by Fabien Plisson
Keith Fitzgerald @keithfitzgerald.bsky.social · 26/06/2025
Worth a watch: Head of Signal, Meredith Whittaker, on so-called "agentic AI" and the difference between how it's described in the marketing and what access and control it would actually require to work as advertised.
204109324360
Reposted by Fabien Plisson
Daniel Hurdiss @danielhurdiss.bsky.social · 23/05/2025
Think coronavirus spikes have run out of surprises? Think again. Our latest preprint dives into the highly unusual spikes of marine mammal coronaviruses. www.biorxiv.org/content/10.1... This #cryoEM study was led by @viralfusion.bsky.social, with key contributions from an amazing team.
38026
Fabien Plisson @fabienplisson.bsky.social · 05/05/2025
Never too late to publish, I am proud to see this work finally published: Tagitinin C, a Sesquiterpene Lactone, and Derivatives as Proteasome Inhibitors doi.org/10.1002/ejoc...
doi.org
Tagitinin C, a Sesquiterpene Lactone, and Derivatives as Proteasome Inhibitors
Tagitinin C, a germacranolide, isolated from Tithonia diversifolia was shown to have an interesting level of activity on the proteasome pathway. It is however a particularly unstable molecule, sensit...
110
Reposted by Fabien Plisson
Colin Jackson @cjjackson.bsky.social · 22/04/2025
Choosing ML architectures for protein engineering is often challenging. Our “new” updated preprint provides a rational framework to match ML models to protein fitness tasks, showing landscape ruggedness influences prediction accuracy. Mahakaran dana Adam et al www.biorxiv.org/content/10.1...
biorxiv.org
Investigating the determinants of performance in machine learning for protein fitness prediction
Machine learning (ML) has revolutionized protein biology, solving long-standing problems in protein folding, scaffold generation and function design tasks. A range of architectures have shown success ...
0143
Reposted by Fabien Plisson
Martin Steinegger 🇺🇦 @martinsteinegger.bsky.social · 29/03/2025
Run BioEmu in Colab - just click "Runtime → Run all"! Our notebook uses ColabFold to generate MSAs, BioEmu to predict trajectories, and Foldseek to cluster conformations. Thanks @jjimenezluna.bsky.social for the help! 🌐 colab.research.google.com/github/sokry... 📄 www.biorxiv.org/content/10.1...
colab.research.google.com
Google Colab
110141
Reposted by Fabien Plisson
Gina El Nesr @ginaelnesr.bsky.social · 20/03/2025
Protein function often depends on protein dynamics. To design proteins that function like natural ones, how do we predict their dynamics? @hkws.bsky.social and I are thrilled to share the first big, experimental datasets on protein dynamics and our new model: Dyna-1! 🧵
610538
Reposted by Fabien Plisson
Irina Bezsonova @irinabezsonova.bsky.social · 15/03/2025
You can download my protein structure-inspired artwork from pdb webpage: pdb101.rcsb.org/sci-art/bezs... @rcsbpdb.bsky.social @pdbeurope.bsky.social #sciart
pdb101.rcsb.org
PDB101: Irina Bezsonova Gallery
PDB-101: Training, Outreach, and Education portal of RCSB PDB
110728
Reposted by Fabien Plisson
Alicia Cintron, PhD @cintronrevised.com · 19/02/2025
YouTube is the world's 2nd-largest search engine. So why aren't more conference keynotes and presentations there? 🤔
273
Reposted by Fabien Plisson
Carl T. Bergstrom @carlbergstrom.com · 04/02/2025
Modern-Day Oracles or Bullshit Machines? Jevin West (@jevinwest.bsky.social) and I have spent the last eight months developing the course on large language models (LLMs) that we think every college freshman needs to take. thebullshitmachines.com
thebullshitmachines.com
INTRODUCTION
17528061006
Reposted by Fabien Plisson
Diego del Alamo @delalamo.xyz · 17/02/2025
Short thread about this interesting preprint that explores antibody design by combining MD with inverse folding and active learning. It's a bit rough around the edges but it introduces a cool idea I hope is further fleshed out www.biorxiv.org/content/10.1...
Fig. 1 Inverse Folding Molecular Dynamics (IF-MD) protocol. (A) Free energy landscape
of the process of protein-protein dissociation in the case of SARS-CoV-2 RBD (red) in complex with
the nanobody H11 (blue) visiting different metastable states along the unbinding mechanism along
two collective variables. Two different unbinding trajectories from the bound state are shown in black
and red. (B) IF-MD protocol to sample protein sequence space by constraining k ensemble averaged
observables q
target
k
and using Bayesian Optimization to propose new sequences. (C) Specific example
of IF-MD constrained to a target the unbinding rate constant, and using Infrequent Metadynamics
to sample the prior ensemble. Unbinding example trajectories sampled by Infrequent Metadynamics
are shon in blue, along the RMSD from the cryo-EM structure of H11 bound to SARS-CoV-2 RBD
as a function of time.
1356
Fabien Plisson @fabienplisson.bsky.social · 13/02/2025
Are protein language models the universal key? doi.org/10.1016/j.sb... Brilliant and thoughtful piece.
doi.org
Redirecting
031
Fabien Plisson @fabienplisson.bsky.social · 12/02/2025
Brilliant video on the development of protein structure prediction featuring #AF2 and #Rosetta series youtu.be/P_fHJIYENdI?...
youtu.be
What if all the world's biggest problems have the same solution?
YouTube video by Veritasium
030
Reposted by Fabien Plisson
César de la Fuente @delafuentelab.bsky.social · 05/02/2025
(1/4)In our new paper in @cp-trendsgenetics.bsky.social @cellpress.bsky.social, we highlight some fascinating tiny proteins—microproteins—which appear to be widespread in nature. www.cell.com/trends/genet...
cell.com
Microproteins: emerging roles as antibiotics
Recent advances in computational prediction and experimental techniques have detected previously unknown microproteins, particularly in the human microbiome. These small proteins, produced by diverse ...
184
Fabien Plisson @fabienplisson.bsky.social · 30/01/2025
Great editorial from @evotec.bsky.social using AI+ML tools to assist drug discovery campaigns in the small and middle chemical spaces pubs.acs.org/doi/10.1021/...
pubs.acs.org
Real-World Applications and Experiences of AI/ML Deployment for Drug Discovery
OR SEARCH CITATIONS
040
Fabien Plisson @fabienplisson.bsky.social · 29/01/2025
Closer to my experience in #CompBiology - the development of AI algorithms to discover and design antimicrobial peptides against #AMR, where predictive ML models are employed to predict the antimicrobial nature (classification) or activity (regression), primarily from sequences.
110
Fabien Plisson @fabienplisson.bsky.social · 29/01/2025
Drug discovery and biotechnology are multi-objective optimisation challenges. Machine learning models integrate well into an enzyme engineering pipeline to predict properties and functions. Parallelising these models allows multi-objective optimisation.
111
Reposted by Fabien Plisson
César de la Fuente @delafuentelab.bsky.social · 15/01/2025
(1/5)Venoms are an underexplored treasure trove of bioactive molecules. Our latest research taps into these evolutionary powerhouses to find a new arsenal against drug-resistant bacteria. We introduce Venomics AI. www.biorxiv.org/content/10.1...
152
Reposted by Fabien Plisson
Saloni @scientificdiscovery.dev · 23/12/2024
Many people talk about the "Golden Age of Antibiotics", but I hadn't seen it visualized properly. Just how many types of antibiotics were discovered during that time? So, I visualized it myself!
A timeline titled "The Golden Age of Antibiotics" shows when each antibiotic drug class was first available for medical use, with example antibiotics labeled. Classes are color-coded by their source: actinomycetes, other bacteria, fungi, or synthetic. Milestones include the first antibiotics (arsphenamines in 1910), as well as the discovery of many actinomycetes-derived antibiotics, such as streptomycin, and sulfonamides, penicillins, and tetracyclines. Data: Hutchings, Truman, Wilkinson (2019). Created by Saloni Dattani for Our World in Data.
231110382
Reposted by Fabien Plisson
AI x Bio Discovery @aixbiobot.bsky.social · 18/12/2024
scMusketeers: Addressing imbalanced cell type annotation and batch effect reduction with a modular autoencoder [new]
scMusketeers: Addressing imbalanced cell type annotation and batch effect reduction with a modular autoencoderFigure 1Figure 2Figure 3
011
Fabien Plisson @fabienplisson.bsky.social · 17/12/2024
What a stunning view #SydneyOpera #Xmas
000
Reposted by Fabien Plisson
Kresten Lindorff-Larsen @lindorfflarsen.bsky.social · 27/11/2024
This should go chiral
715018
Reposted by Fabien Plisson
Selene FeRNAndez @selfdz.bsky.social · 09/12/2024
If you want to support my voyage, you can still do so here: www.chuffed.org/project/help...
chuffed.org
Help Selene get to Antarctica
Hi! I'm Selene Fernandez, a scientist, a daughter and a mother.
031
Reposted by Fabien Plisson
Bristol BioDesign Institute @bristolbiodesign.bsky.social · 15/11/2024
Applications are open for PhD places on the Engineering Biology CDT programme, beginning Sept 2025. The programme is run jointly by the Universities of Bristol and Oxford. Apply to Bristol (by 13 Jan 2025): bristol.ac.uk/study/postgr... Apply to Oxford (by 8 Jan 2025): ox.ac.uk/admissions/g...
0128
Fabien Plisson @fabienplisson.bsky.social · 15/11/2024
Thanks to @sunnalab & @MQ_BDRC for inviting me to the symposium. I was very pleased to present our @CinvestavIra & @IngenieBio work in the #AI-driven peptide discovery & design. It stimulated some interesting questions over bias, representation learning, data scarcity and hype.
000
Reposted by Fabien Plisson
Anne Carpenter @drannecarpenter.bsky.social · 14/11/2024
Looks like Bryan Dickinson isn’t here yet so I’m cross posting his challenge: “if you or someone you know thinks they can actually predict PPIs, prove it. Here is a link to our blinded protein sequence data” dickinsonlab.uchicago.edu/ppi-challenge
dickinsonlab.uchicago.edu
PPI Challenge | dickinson-group
35026
Fabien Plisson @fabienplisson.bsky.social · 10/11/2024
Similar challenges - data scarcity, #bias, noise - also limit #ML-guided #proteindesign. We need to generate new data through computational (e.g., de novo design) & experimental (e.g., SPPS, #syntheticbio) approaches are essential doi.org/10.1038/s435...
doi.org
The future of machine learning for small-molecule drug discovery will be driven by data - Nature Computational Science
The application of machine learning techniques to small-molecule drug discovery has not yet yielded a true leap forward in the field. This Perspective discusses how a renewed focus on data and validat...
010
Fabien Plisson @fabienplisson.bsky.social · 10/11/2024
Happy to launch Iɴɢᴇɴɪᴇ Bɪᴏ, a consulting firm offering data-driven solutions for biomolecular discovery and design! We specialize in AI-driven molecular design, computational chemistry, and more. Let's collaborate! Visit ingeniebio.com #AI #ML #drugdiscovery
ingeniebio.com
Ingenie Bio | biomolecular design
Ingenie Bio integrates data-driven solutions to support biomolecular design across drug discovery and beyond.
140
Fabien Plisson @fabienplisson.bsky.social · 04/11/2024
Pleased to attend #ABACBS24 Symposium on Bioinformatics Excellence and Innovation #SBEI24 to see challenges in cloud computing, data mgt, #AI technology to #drugdiscovery research programs from @AWS, @googlecloud & bioinformatics facilities like @AusBiocommons @PeterMacCC
000
Fabien Plisson @fabienplisson.bsky.social · 01/11/2024
A big thanks to @ersiliaio for incorporating our #MachineLearning models into the Ersilia Model Hub! 🚀 Our models now help predict blood-brain barrier permeability for small molecules. 🧠 Check out the repository here:#DrugDiscovery #AI #ML (1/3) github.com/ersilia-os/eos…
github.com
GitHub - ersilia-os/eos3mk2
Contribute to ersilia-os/eos3mk2 development by creating an account on GitHub.
100
Fabien Plisson @fabienplisson.bsky.social · 18/10/2024
Similar challenges - data scarcity, #bias, noise - also limit #ML-guided #proteindesign. We need to generate new data through computational (e.g., de novo design) & experimental (e.g., SPPS, #syntheticbio) approaches are essential doi.org/10.1038/s43588…
doi.org
The future of machine learning for small-molecule drug discovery will be driven by data
Nature Computational Science - The application of machine learning techniques to small-molecule drug discovery has not yet yielded a true leap forward in the field. This Perspective discusses how a...
000
Fabien Plisson @fabienplisson.bsky.social · 16/10/2024
Fantastic view over Sydney CBD to attend @ProtoAxiom #ChallengerSummit2024 and meet like-minded people pushing forward the Australian #Biotechnology industry
000
Fabien Plisson @fabienplisson.bsky.social · 16/10/2024
I tried @notebooklm_pods with two of our recent peer-reviewed articles; it generated a 10-minute audio on structural bias in #ML-guided #peptide design with great accuracy. Enjoy! substack.com/home/post/p-15…
000
Fabien Plisson @fabienplisson.bsky.social · 14/10/2024
Our article, "Benchmarking protein structure predictors to assist #MachineLearning-guided peptide discovery," has been included in the @digital_rsc themed collection "Computational protein design and structure prediction: Celebrating the 2024 #NobelPrize in Chemistry"
100
Fabien Plisson @fabienplisson.bsky.social · 08/10/2024
Happy to launch Iɴɢᴇɴɪᴇ Bɪᴏ, a consulting firm offering data-driven solutions for biomolecular discovery and design! We specialize in AI-driven molecular design, computational chemistry, and more. Let's collaborate! Visit#AI #ML #drugdiscovery ingeniebio.com
ingeniebio.com
Ingenie Bio | biomolecular design
Ingenie Bio integrates data-driven solutions to support biomolecular design across drug discovery and beyond.
000
Fabien Plisson @fabienplisson.bsky.social · 22/09/2024
I'm excited to see some computational biology & drug discovery research down at Melbourne's @bmh2024Melb #BMH2024
100
Fabien Plisson @fabienplisson.bsky.social · 20/09/2024
Brilliant video by @ColdFusion_TV about the benefits of #AI across sectors - healthcare (prosthesis), environment (kelp ID), pharma (drug repurposing), and energy (batteries) - beyond productivity enhancement. youtu.be/wAgDicfEWFY?fe…
youtu.be
Are we all wrong about AI?
Go to https://80000hours.org/coldfusion, to learn about fulfilling, high-impact careers. It's no secret that AI is controversial today. Judging by some of the chaos it's caused, there's good reason to think AI seems to ruin everything it touches. But what about the flipside? What good is AI actually doing in the world? It's a question I don't hear asked much so today we'll find out. Note: Reinforcement learning etc are all included in this conversation. Sources: https://docs.google.com/document/d/1RpPbUyj3jK7n56aT45n1R9jGYl4w9tIR6GxbcvMVoUs/edit?usp=sharing ColdFusion Podcast: https://www.youtube.com/@ThroughTheWeb ColdFusion Music: https://www.youtube.com/@ColdFusionmusic http://burnwater.bandcamp.com Get my book: http://bit.ly/NewThinkingbook ColdFusion Socials: https://discord.gg/coldfusion https://facebook.com/ColdFusionTV https://twitter.com/ColdFusion_TV https://instagram.com/coldfusiontv ColdFusion Created by Dagogo Altraide Producer: Dagogo Altraide, Tawsif Akkas Editor: Brayden Laffrey, Dagogo Altraide Writers: Meehan Kathan, Dagogo Altraide
000
Fabien Plisson @fabienplisson.bsky.social · 18/09/2024
I am very pleased to have presented @peptidicos research work about #MachineLearning #peptidedesign at Macquarie Dementia Research Centre @MQ_DRC @Macquarie_Uni. Thanks @OleTietz. Beautiful campus.
000
Fabien Plisson @fabienplisson.bsky.social · 17/09/2024
Exciting times ahead - Australia's growing #synbio ecosystem, like advanced fermentation platforms @cauldronferm combined with #AI protein design, will provide sustainable and economical solutions to health issues like #AMR synbiobeta.com/read/building-…
synbiobeta.com
Building a Synthetic Biology Future in the Land Down Under - SynBioBeta
Government policy support for domestic manufacturing and sustainability is fueling Australia’s synthetic biology ecosystem
010
Fabien Plisson @fabienplisson.bsky.social · 15/09/2024
Celebrando #VivaMexico con @SelFdz en Sídney, Australia.
000
Fabien Plisson @fabienplisson.bsky.social · 14/09/2024
#NewProfilePic
000
Fabien Plisson @fabienplisson.bsky.social · 13/09/2024
Insightful paper on scaling laws applied to #proteins: Optimising pLMs through diverse training datasets is critical to enhancing performance and robustness and using resources more efficiently while avoiding overfitting. #AI #LLM #machinelearning #CompBio doi.org/10.1101/2024.0…
doi.org
Training Compute-Optimal Protein Language Models
We explore optimally training protein language models, an area of significant interest in biological research where guidance on best practices is limited. Most models are trained with extensive compute resources until performance gains plateau, focusing primarily on increasing model sizes rather than optimizing the efficient compute frontier that balances performance and compute budgets. Our investigation is grounded in a massive dataset consisting of 939 million protein sequences. We trained over 300 models ranging from 3.5 million to 10.7 billion parameters on 5 to 200 billion unique tokens, to investigate the relations between model sizes, training token numbers, and objectives. First, we observed the effect of diminishing returns for the Causal Language Model (CLM) and that of overfitting for the Masked Language Model (MLM) when repeating the commonly used Uniref database. To address this, we included metagenomic protein sequences in the training set to increase the diversity and avoid the plateau or overfitting effects. Second, we obtained the scaling laws of CLM and MLM on Transformer, tailored to the specific characteristics of protein sequence data. Third, we observe a transfer scaling phenomenon from CLM to MLM, further demonstrating the effectiveness of transfer through scaling behaviors based on estimated Effectively Transferred Tokens. Finally, to validate our scaling laws, we compare the large-scale versions of ESM-2 and PROGEN2 on downstream tasks, encompassing evaluations of protein generation as well as structure- and function-related tasks, all within less or equivalent pre-training compute budgets. ### Competing Interest Statement The authors have declared no competing interest.
010
Fabien Plisson @fabienplisson.bsky.social · 07/09/2024
Thanks to @RoySocChem + @LatinXChem for adding our seminal review, co-authored by @fersaldivar11, @DanielAldas, and @difacquim, to this themed collection.#LatinXChem #naturalproducts #drugdiscovery #machinelearning #artificial_intelligence shorturl.at/zJIaX
000
Fabien Plisson @fabienplisson.bsky.social · 22/07/2024
This Thursday, I will present @UNSWBABS our research on #MachineLearning guided peptide design and how we monitor existing #bias in biological data and tackle it to develop more informed ML/DL models. #XAI #AMR #peptide2drug
000
Fabien Plisson @fabienplisson.bsky.social · 14/07/2024
Traversée des rues de Montréal en découvrant avec mini-moi les statues du #Mignonisme de @Katerin55468086. Le premier Monsieur Rose est dans l'hôtel Fairmont #artwork #contemporain
000
Fabien Plisson @fabienplisson.bsky.social · 12/07/2024
After graduating the last @peptidicos members, I finally closed my research lab in Mexico @cinvestav. Up for a new venture Down Under. I will miss colleagues and friends #VivaMexico
010
Fabien Plisson @fabienplisson.bsky.social · 10/07/2024
Integrating predictive and generative #MachineLearning models in DMTA cycles enhances hit discovery and hit-to-lead optimisation in early-stage #drugdiscovery. This @NatureComms is a great resource for my next venture in AI-driven #BiomolecularDesign doi.org/10.1038/s41467…
doi.org
A data science roadmap for open science organizations engaged in early-stage drug discovery
Nature Communications - Artificial intelligence is greatly accelerating research in drug discovery, but its development is still hindered by the lack of available data. Here the authors present...
000
Fabien Plisson @fabienplisson.bsky.social · 28/06/2024
Thanks to @cinvestav, @Conahcyt_Mex, Mercedes Lopez-Perez, Luis González de la Vara, Agustino @bioingenios, @karma_twitt for their valuable support, as well as many professors and technicians from @CinvestavIra
010
Fabien Plisson @fabienplisson.bsky.social · 19/06/2024
Most predictive/generative #machinelearning models are built on #protein sequences. With @DanielAldas we developed a fast + parallelizable strategy to map the structural landscape of training set(s), giving you insights on structural constraints #XAI #AF2 doi.org/10.1039/D3DD00…
000