Sign in

Miquel Anglada-Girotto

@m1quelag.bsky.social
377 followers 1.7K following 56 posts

Love predicting genomic things. Postdoc @crgenomica.bsky.social at the Probabilistic Machine Learning and Genomics group. Creator of @splicingnews.bsky.social

PostsRepliesMedia
Reposted by Miquel Anglada-Girotto
Ezequiel Galpern @eag91.bsky.social · 25/09/2026
🧬 What do deep-learning models for proteins actually learn? Our new review looks at how model predictions relate to fitness, folding stability and function. With @cwjpugh.bsky.social, Mafalda Dias & @jonnyfrazer.bsky.social 🔗 chemrxiv.org/doi/full/10....
chemrxiv.org
From sequences and structures to fitness, folding and function: challenges in the age of AI | ChemRxiv
Deep-learning models have transformed our ability to predict the phenotypic effects of sequence perturbations. Protein language models (pLMs) score the evolutionary propensity of any amino-acid substi...
177
Reposted by Miquel Anglada-Girotto
Ezequiel Galpern @eag91.bsky.social · 05/08/2026
1/ New preprint! with @solersanchisx.bsky.social @cwjpugh.bsky.social @federicobilleci.bsky.social @jonnyfrazer.bsky.social and Mafalda Dias, we introduce a scalable framework to improve stability prediction and separate folding from functional constraints. www.biorxiv.org/content/10.6...
biorxiv.org
Blending physics-based and inverse folding models to disentangle variant effects on stability and function
Protein sequences are constrained not only by the need to fold into stable structures, but also by specific functional requirements imposed by natural selection. Yet predictions of how amino-acid chan...
1139
Miquel Anglada-Girotto @m1quelag.bsky.social · 28/07/2026
8/ Hopefully, this work makes it easier for others to explore new biological questions with AlphaGenome. There is much more research to come!
000
Miquel Anglada-Girotto @m1quelag.bsky.social · 28/07/2026
7/ I’m grateful to @al-murphy.bsky.social for his editorial feedback; Alejandro Buendia, @xinmingtu.bsky.social , Danila Bredikhin, and Adam He for their support; and my mentors, @jonnyfrazer.bsky.social and Mafalda Dias at @crg.eu , for their guidance.
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 28/07/2026
6/ This project means a lot to me personally. I have learned from open-source implementations of research papers for years, and contributing to @lucidrains.bsky.social ’s PyTorch implementation felt like the perfect opportunity to learn by doing.
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 28/07/2026
5/ The work is available here: github.com/MiqG/alphage... It is still under active development and has not yet been fully integrated into the main repositories or the JAX fine-tuning codebase.
github.com
GitHub - MiqG/alphagenome_finetuning_rna at v1.0.2
Contribute to MiqG/alphagenome_finetuning_rna development by creating an account on GitHub.
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 28/07/2026
4/ The model’s predictive performance is very promising, but memory remains a practical limitation. Training on just two samples required an entire H100 GPU, so there is still a lot of room to improve accessibility and efficiency.
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 28/07/2026
3/ The result is a functional framework to preprocess, load, and train on multiple RNA-seq modalities: gene expression, splice-site usage, and splice junctions.
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 28/07/2026
2/ What I thought would take a month became a much longer debugging journey. We encountered everything from conceptual challenges to simple bugs, and with help from the AlphaGenome authors, we worked through them.
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 28/07/2026
1/ We’ve been working on making it easier to fine-tune AlphaGenome on new RNA-seq samples, with a focus on one of my favourite regulatory layers: splicing! Blog post: genomicsxai.github.io/blogs/2026-0...
genomicsxai.github.io
Beyond coverage tracks: fine-tuning AlphaGenome's splicing heads from scratch
Summary DeepMind has released AlphaGenome’s code and model weights, and the community has since developed alphagenome_ft and alphagenome-pytorch to enable seamless fine-tuning in both JAX and PyTorch....
1164
Reposted by Miquel Anglada-Girotto
bioRxiv Genomics @biorxiv-genomic.bsky.social · 24/05/2026
OpenSplice: the impact of half a million mutations on the alternative splicing of 600 human exons www.biorxiv.org/content/10.64898/20…
093
Miquel Anglada-Girotto @m1quelag.bsky.social · 01/04/2026
Special thanks to all authors and founders!! Carolina Segura-Morales Dan F. Moakley @chaolinzhang.bsky.social @samuelmiver.bsky.social Andrea Califano Luis Serrano @crg.eu @columbiacancer.bsky.social @boehringerglobal.bsky.social
000
Miquel Anglada-Girotto @m1quelag.bsky.social · 01/04/2026
Many regulatory layers modulate splicing factors at the same time impacting their activity. How can we quantify it? Apparently "functional" target exons give us a hint and uncover two cancer programs. Have a look at our solution: www.nature.com/articles/s41...
nature.com
Exon inclusion signatures enable accurate estimation of splicing factor activity - Nature Communications
Splicing factors shape how genes are stitched into RNA, but their activity is hard to measure. Here, the authors benchmark network methods and show exon-inclusion signatures infer splicing factor acti...
143
Reposted by Miquel Anglada-Girotto
Ezequiel Galpern @eag91.bsky.social · 01/04/2026
Why are there 20 amino acids and 4 nucleotides? Combining Energy Landscape and Molecular Information theories provides constraints to the alphabet size of an evolving biopolymer, given its physico-chemical properties... Read more in our new article: www.nature.com/articles/s41...
nature.com
An information-theoretic argument for the restriction of the current biological alphabets to 4 nucleotides and 20 amino acids - Scientific Reports
Life as we know it is based on foldable biopolymers encoded with just 4 nucleotides or 20 amino acids. Evolution of these biopolymers requires effective and fast search of both the conformational spac...
133
Reposted by Miquel Anglada-Girotto
Universitat de Barcelona @ub.edu · 03/03/2026
#UBalsMitjans | 👌 @elpuntavui.cat entrevista Raúl Ruiz, estudiant de Bioquímica i professor de llengua de signes catalana, que ha coordinat un vocabulari de termes científics a la #UniBarcelona. «És una llengua pròpia, totalment vàlida per crear terminologia en àmbits especialitzats», afirma Ruiz.
elpuntavui.cat
"La llengua de signes s'hauria d'estudiar a totes les escoles"
"El 2010 la llengua de signes catalana es va reconèixer a través d'una llei però això és teòric, falta portar-ho a la pràctica" "És important veure la llengua de signes catalana des de la perspectiva...
061
Miquel Anglada-Girotto @m1quelag.bsky.social · 12/12/2025
Oh! I was not aware of this literature, thanks! Yes, I was thinking of evaluating how often using existing seq2func models with a tool like ledidi would recover the genotype of the person. I suppose as model personalization improves we'll hit a point we cannot share model weights...
030
Miquel Anglada-Girotto @m1quelag.bsky.social · 08/12/2025
Did you try doing your privacy benchmark with other models that predict ATAC? Should we be concerned about privacy with RNA coverage models too? It was shown how bad models are at personalized predictions, but it is the first time I see a benchmark on how good they can be at identifying people!
100
Reposted by Miquel Anglada-Girotto
Jonathan Frazer @jonnyfrazer.bsky.social · 24/11/2025
popEVE is out in Nature Genetics! 🎉 We built a proteome-wide model that combines cross-species and human population variation to rank missense variants by disease severity and help diagnose rare genetic disorders. rdcu.be/eRu7K
rdcu.be
Proteome-wide model for human disease genetics
Nature Genetics - popEVE is a proteome-wide deep generative model to identify and predict pathogenicity of missense mutations causing genetic disorders.
24919
Reposted by Miquel Anglada-Girotto
Jonathan Frazer @jonnyfrazer.bsky.social · 25/11/2025
LFB is NeurIPS-bound! 🎉 Mafalda, @cwjpugh.bsky.social and I will be in San Diego next week for NeurIPS -- happy to chat variant effect prediction (or just say hi). “From Likelihood to Fitness: Improving Variant Effect Prediction in Protein and Genome Language Models” openreview.net/pdf/a151f62e...
openreview.net
1101
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/11/2025
Great initiative! I have used it uploading computational papers. However there's a message saying that you are not very confident on the platform's feedback for this types of paper. Why is that? What would give you more confidence? My N is low but it was fine!
110
Reposted by Miquel Anglada-Girotto
Centre de Regulació Genòmica (CRG) @crg.eu · 25/11/2025
Our annual PhD call is closing at the end of this week on 30 November. If you're interested in carrying out world-class scientific research in Barcelona, you still have a few days left to submit your application! www.crg.eu/en/content/t...
01318
Miquel Anglada-Girotto @m1quelag.bsky.social · 21/11/2025
Very nice approach! Is the code (and pretrained weights) available? Thanks!!
110
Miquel Anglada-Girotto @m1quelag.bsky.social · 21/11/2025
Very nice!
010
Miquel Anglada-Girotto @m1quelag.bsky.social · 26/10/2025
Thank you!
010
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/10/2025
We would also like to thank @narjournal.bsky.social 's editorial team and reviewers for their feedback and support!
000
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/10/2025
Let us know what you think! We’re very excited to see how our approach can lead to new insights for you! This work would not have been possible without my super supervisors, Samuel Miravet-Verde & Luis Serrano and the Serrano Lab team, at the wonderful @crg.eu
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/10/2025
Although more validation will be needed, we believe our work will enable studying the state of splicing factors in widely available and single-cell atlases, contributing to providing a more complete picture of splicing regulation in data-scarce but experimentally very rich settings.
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/10/2025
Interestingly, during embryogenesis, the regulation of splicing factors follows the opposite trend from that observed during carcinogenesis. MYC, G2M, and E2F prioritized pathways are downregulated during human embryogenesis, supporting their role as regulators of the carcinogenic switch of SFs.
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/10/2025
Because our prioritization involved Perturb-seq experiments, we could ask which other pathways had also strong evidence as mediators. These were: G2M checkpoint, E2F targets, and spermatogenesis.
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/10/2025
Long story short, the MYC pathway was the top candidate, further supporting the known importance of MYC in regulating splicing (great references: Leclair et al. 2022 ( @olgaanczukow.bsky.social lab) and Koh et al. 2015 ( @guccionelab.bsky.social lab)).
110
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/10/2025
But most cancer-driver mutations don’t involve splicing factors, so how does cancer induce this aberrant regulation in splicing factors? We came up with a strategy to isolate the best candidate pathways connecting cancer-driver mutations and carcinogenic splicing factor regulation.
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/10/2025
A nice insight was to see that not one but many splicing factors can drive their own aberrant regulation. We found evidence that they do so through their splicing factor-exon and protein-protein interactions.
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/10/2025
Through Perturb-seq datasets in RPE1 pre-cancerous cells (Replogle et al. 2022 ( @weissmanlab.bsky.social )), we could dissect systematically which genes drive carcinogenesis regulation of splicing factors.
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/10/2025
This enabled us to use existing datasets to explore how splicing factors are regulated during carcinogenesis in bulk (Danielsson et al. 2013 (Emma Lundberg lab)) and single-cell models (Hodis et al. 2022 (Aviv Regev lab)).
100
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/10/2025
Building on our method for splicing factor activity analysis (doi.org/10.1101/2024...), we expand our database of experiments that perturb SFs and show that adjusting SF activities with a “shallow” neural net does well at recapitulating exon-inclusion-based activities from only gene expression.
doi.org
Exon inclusion signatures enable accurate estimation of splicing factor activity
Splicing factors control exon inclusion in messenger RNAs, shaping transcriptome and proteome diversity. Their catalytic activity is regulated by multiple layers, making single-omic measurements on th...
110
Miquel Anglada-Girotto @m1quelag.bsky.social · 25/10/2025
Wouldn’t it be cool to leverage the throughput of single-cell data to study splicing regulation even when we lack exon resolution? 😀 Here’s the peer-reviewed version of our paper on how we can measure changes in splicing factor activity in virtually any single-cell dataset: doi.org/10.1093/nar/...
doi.org
Using single-cell perturbation screens to decode the regulatory architecture of splicing factor programs
Abstract. Splicing factors shape the isoform pool of most transcribed genes, playing a critical role in cellular physiology. Their dysregulation is a hallm
1103
Miquel Anglada-Girotto @m1quelag.bsky.social · 23/10/2025
Couldn't think of a better place to make models! Come join us!
010
Miquel Anglada-Girotto @m1quelag.bsky.social · 15/07/2025
Es Castell
010
Miquel Anglada-Girotto @m1quelag.bsky.social · 14/07/2025
@splicingnews.bsky.social
000
Reposted by Miquel Anglada-Girotto
bioRxivpreprint @biorxivpreprint.bsky.social · 07/07/2025
An organoid model of the menstrual cycle reveals the role of the luminal epithelium in regeneration of the human endometrium www.biorxiv.org/content/10.1101/202…
084
Reposted by Miquel Anglada-Girotto
Anamaria Elek @aelek.bsky.social · 06/07/2025
I am very happy to have posted my first bioRxiv preprint. A long time in the making - and still adding a few final touches to it - but we're excited to finally have it out there in the wild: www.biorxiv.org/content/10.1... Read below for a few highlights...
biorxiv.org
Decoding cnidarian cell type gene regulation
Animal cell types are defined by differential access to genomic information, a process orchestrated by the combinatorial activity of transcription factors that bind to cis -regulatory elements (CREs) to control gene expression. However, the regulatory logic and specific gene networks that define cell identities remain poorly resolved across the animal tree of life. As early-branching metazoans, cnidarians can offer insights into the early evolution of cell type-specific genome regulation. Here, we profiled chromatin accessibility in 60,000 cells from whole adults and gastrula-stage embryos of the sea anemone Nematostella vectensis. We identified 112,728 CREs and quantified their activity across cell types, revealing pervasive combinatorial enhancer usage and distinct promoter architectures. To decode the underlying regulatory grammar, we trained sequence-based models predicting CRE accessibility and used these models to infer ontogenetic relationships among cell types. By integrating sequence motifs, transcription factor expression, and CRE accessibility, we systematically reconstructed the gene regulatory networks that define cnidarian cell types. Our results reveal the regulatory complexity underlying cell differentiation in a morphologically simple animal and highlight conserved principles in animal gene regulation. This work provides a foundation for comparative regulatory genomics to understand the evolutionary emergence of animal cell type diversity. ### Competing Interest Statement The authors have declared no competing interest. European Research Council, https://ror.org/0472cxd90, ERC-StG 851647 Ministerio de Ciencia e Innovación, https://ror.org/05r0vyz12, PID2021-124757NB-I00, FPI Severo Ochoa PhD fellowship European Union, https://ror.org/019w4f821, Marie Skłodowska-Curie INTREPiD co-fund agreement 75442, Marie Skłodowska-Curie grant agreement 101031767
15824
Miquel Anglada-Girotto @m1quelag.bsky.social · 02/07/2025
@splicingnews.bsky.social
000
Reposted by Miquel Anglada-Girotto
Jacob Schreiber @jmschreiber91.bsky.social · 18/06/2025
Last week I released bpnet-lite v0.5.0. BPNet/ChromBPNet are powerful models for understanding regulatory genomics from @anshulkundaje.bsky.social's group, and now it's way easier to go from raw data to trained models and analysis + results in PyTorch Try it out with `pip install bpnet-lite`
13811
Reposted by Miquel Anglada-Girotto
Jacob Schreiber @jmschreiber91.bsky.social · 03/06/2025
I wrote a quick application note on Tomtom-lite, a Python implementation of the Tomtom algorithm for comparing PWMs against each other. This implementation can be 10-1000x faster and, as a Python function, can be integrated into your workflows easier. www.biorxiv.org/content/10.1...
biorxiv.org
Tomtom-lite: Accelerating Tomtom enables large-scale and real-time motif similarity scoring
Summary Pairwise sequence similarity is a core operation in genomic analysis, yet most attention has been given to sequences made up of discrete characters. With the growing prevalence of machine lear...
25818
Reposted by Miquel Anglada-Girotto
EMBL-EBI @ebi.embl.org · 09/06/2025
Polygenic scores (PGS) offer insights into a person’s inherited risk of disease. GeneticScores.org is a new platform that enables secure, cloud-based calculation of polygenic scores to make genomic risk prediction more accessible. www.ebi.ac.uk/about/news/u... 🖥️🧬
1198
Miquel Anglada-Girotto @m1quelag.bsky.social · 07/06/2025
Today I learned artists study primitive art to understand how art was made out of the art business context. This made me wonder how science would be made nowadays out of the journal publishing context. Would we try to answer different questions?
040
Miquel Anglada-Girotto @m1quelag.bsky.social · 26/05/2025
Leveraging evolution to make fitness estimation scale with model size again! Great experiencing the making of this one behind the scenes 🙌
040
Reposted by Miquel Anglada-Girotto
Stephen Turner @stephenturner.us · 26/03/2025
polars-bio - fast, scalable and out-of-core operations on large genomic interval datasets www.biorxiv.org/content/10.1... 🧬🖥️🧪 github.com/biodatageeks...
0123
Miquel Anglada-Girotto @m1quelag.bsky.social · 21/02/2025
I would like to thank Samuel Miravet-Verde and Luis Serrano for their supervision, and the support of @crg.eu !
000
Miquel Anglada-Girotto @m1quelag.bsky.social · 21/02/2025
Our carcinogenesis use case is just an example of how single-cell perturbation screens can be leveraged. This goes along the lines of a recent study by Ota et al bsky.app/profile/jkpr... We believe our approach is flexible enough to study the architecture of many other splicing factor programs.
100