Sign in

GAMA Miguel Angel

@miangoar.bsky.social
261 followers 187 following 137 posts

Biologist that navigate in the oceans of diversity through space-time Protein evolution, metagenomics, AI/ML/DL Website miangoaren.github.io

PostsRepliesMedia
GAMA Miguel Angel @miangoar.bsky.social · 20/05/2026
Today, the AlphaFold2 paper reached the milestone of 50k citations according to Google Scholar! And AlphaFold3 will likely reach 15k citations tomorrow. Congratulations to the entire AlphaFold team, as well as to all the scientists who helped democratize protein structure prediction 🥳
261
GAMA Miguel Angel @miangoar.bsky.social · 20/04/2026
11/12 All the course information, the link to the slides, and more can be found on my site in GithubPages. In addition, YouTube has automatically dubbed the course into 18 other languages to make learning more accessible. I hope you find it useful :) miangoaren.github.io/teaching/pro...
021
GAMA Miguel Angel @miangoar.bsky.social · 20/04/2026
10/12 In the final lecture "Data & biases" (4h 16min), we discuss the main biological databases (eg. PDB, UniProt, NCBI), data cleaning strategies, data leakage, and the biases that can silently affect the generalization capabilities of your models youtu.be/bEt7tZKvfiI
100
GAMA Miguel Angel @miangoar.bsky.social · 20/04/2026
9/12 In the 9th lecture "AI-driven protein design" (8h 27min), we cover the design toolkit: from directed evolution and rational design, to protein language models (representation learning) and generative AI to create both protein sequences and structures youtu.be/PvMNlxZv_Bg
100
GAMA Miguel Angel @miangoar.bsky.social · 20/04/2026
8/12 In the 8th lecture "AlphaFold" (7h 39mins), we review in detail the AF2 and AF3 architectures, how they revolutionized structural biology, their strengths and weaknesses and what the post-AlphaFold era looks like for protein design youtu.be/4K8SDxk85a0
100
GAMA Miguel Angel @miangoar.bsky.social · 20/04/2026
7/12 In the 7th lecture "Protein evolution" (2h 38mins, and my personal favorite), we trace how proteins originated from simple peptides, how mutations shape the evolutionary paths and how epistasis drives the evolution of proteins youtu.be/rkmWSR8BUms
100
GAMA Miguel Angel @miangoar.bsky.social · 20/04/2026
6/12 In the 6th lecture "Protein function" (1h 58mins), we cover how proteins fold inside the cell, how enzymes work and how function is regulated through distinct mechanisms like allostery, post-translational modifications, and proteostasis youtu.be/Un6QaTM412A
100
GAMA Miguel Angel @miangoar.bsky.social · 20/04/2026
5/12 In the 5th lecture "Protein structure" (2h 55mins), we explore the principles of structural biology: from amino acids and secondary structure to fold classification schemes and the uneven shape of the protein universe youtu.be/7GmPNVhJhw0
100
GAMA Miguel Angel @miangoar.bsky.social · 20/04/2026
4/12 In the 4th lecture "Transformers & language models" (3h 42mins), we break down how the original Transformers architecture work, the differences between BERT and GPT, scaling laws, modern LLMs and how to work with them youtu.be/tNAKnz_tDIc
100
GAMA Miguel Angel @miangoar.bsky.social · 20/04/2026
3/12 In the 3rd lecture "Deep learning" (1h 54mins), we’ll review how neural networks work, from neurons and backpropagation to modern architectures. Then we explore the main DL frameworks used to build models youtu.be/YiEmCQuW-xc
110
GAMA Miguel Angel @miangoar.bsky.social · 20/04/2026
3/12 In the 2nd lecture "Machine learning" (1h 39mins), we’ll review what artificial intelligence is and its subfields, the current capabilities of the algorithms, and how a model is trained in general youtu.be/9fEl5RsLKJs
100
GAMA Miguel Angel @miangoar.bsky.social · 20/04/2026
2/12 In the first lecture "Basic computing concepts" (1h 58mins), we’ll review how CPUs and GPUs work, as well as essential software for data analysis like GNU/Linux and the python ecosystem for bioinformatics youtu.be/RddVvvYRpTc
100
GAMA Miguel Angel @miangoar.bsky.social · 20/04/2026
1/12 🧵 Do you want to learn how to design proteins using AI but don’t know anything about biology? I created a free 10-lesson course on YouTube. It’s now available in Spanish (original) and English (autodubbing w/Kokoro 82M). Here’s an overview of the topics covered in each lecture :)
120
GAMA Miguel Angel @miangoar.bsky.social · 12/02/2026
I am in the "Life Sciences Super Cluster" and quite far away from many colleagues in protein design who are in the cluster called "Computational Chemistry Nexus" 😭
100
GAMA Miguel Angel @miangoar.bsky.social · 04/02/2026
I recorded ~4h where we cover the main bio databases, data processing methods, many sources of bias and topics like generalization and data leakage :) youtu.be/SKpHaHgvCKE Slides drive.google.com/file/d/1jpEwDBncJCRviG_DaWs2EpzCL_1BfB9t/view English is available only via auto-translated subtitles
010
GAMA Miguel Angel @miangoar.bsky.social · 03/02/2026
I recorded ~8h introducing the main algorithms for protein design: from classical approaches to protein language models, AlphaFold, ESMFold, MPNN, diffusion models and more :) youtu.be/wKUYtAt87d4T... Slides drive.google.com/file/d/1EPLj... English is available only via auto-translated subtitles
011
GAMA Miguel Angel @miangoar.bsky.social · 02/02/2026
I’ve recorded ~8h explaining the architectures of AlphaFold, AF2 & AF3, as well as the context needed to understand their development, applications and limitations :) youtu.be/_jDRr5BcTaY Slides drive.google.com/file/d/1i4QE... English is available only via auto-translated subtitles
095
GAMA Miguel Angel @miangoar.bsky.social · 01/02/2026
The 7th lecture is available on YouTube :) We will review how proteins emerge and diversify throughout evolution, considering mutations and molecular interactions youtu.be/qaypRS8SX5M Slides drive.google.com/file/d/1BfQd... English is available only via auto-translated subtitles
120
GAMA Miguel Angel @miangoar.bsky.social · 31/01/2026
The 6th lecture is now available on YouTube :) We’ll review how proteins adopt their 3D shape, how they perform their functions and how their activity is regulated youtu.be/cZs8XtVYa5A Slides drive.google.com/file/d/1TpPj... English is available only via auto-translated subtitles
000
GAMA Miguel Angel @miangoar.bsky.social · 30/01/2026
The fifth lecture of the course is now available on YouTube :) We’ll review amino acid chemistry and how we organize and classify proteins youtu.be/gE6qXwpBP_s Slides drive.google.com/file/d/1F99V... For now, the English version is only available through the automatic translation of the subtitles
000
GAMA Miguel Angel @miangoar.bsky.social · 29/01/2026
The fourth lecture of the course is now available on YouTube :) We will review how Transformers and modern LLMs work youtu.be/vUpb6O6T2yQ Slides drive.google.com/file/d/1y2Vj... For now, the English version is only available through the automatic translation of the subtitles.
021
GAMA Miguel Angel @miangoar.bsky.social · 28/01/2026
The third lecture of the course is now available on YouTube :) We will review how neural networks work. youtu.be/pAgL7NsCUMU Slides drive.google.com/file/d/1cazt... For now, the English version is only available through the automatic translation of the subtitles.
010
GAMA Miguel Angel @miangoar.bsky.social · 27/01/2026
The second lecture of the course is now available on YouTube :) We will review what AI is, its subfields and how to train a model. youtu.be/Xx80O85-5rI Slides drive.google.com/file/d/1i-Jo... For now, the English version is only available through the automatic translation of the subtitles.
000
GAMA Miguel Angel @miangoar.bsky.social · 27/01/2026
The first lecture of the course is now available on YouTube :) youtu.be/uMkZzKbnoJI Slides drive.google.com/file/d/1uDwe... For now, the English version is only available through the automatic translation of the subtitles.
010
GAMA Miguel Angel @miangoar.bsky.social · 22/01/2026
2/3 The course includes +800 freely available slides, and starting next monday, I will publish one video per day. For example, the AlphaFold lecture is ~7.4 hours long and includes 148 slides, in which I cover the architectures of AF1, AF2 and AF3 as well as their applications.
100
GAMA Miguel Angel @miangoar.bsky.social · 22/01/2026
🧵1/3 I created this free 37-hour course, distributed across 10 lectures, to introduce AI-based protein design. For more information about the course and its specific topics, please visit the official course page:
2102
GAMA Miguel Angel @miangoar.bsky.social · 16/01/2026
Even the five most abundant folds account for ~31% of all domains in the PDB. For more information on these superfolds check out Protein superfamilies and domain superfolds pubmed.ncbi.nlm.nih.gov/7990952/
010
GAMA Miguel Angel @miangoar.bsky.social · 16/01/2026
1/2 If you think that the Protein Data Bank is a representative DB, it is not. The data is highly biased. The CATH suggests that there are 1,472 protein folds, yet among the ~600k domains present in the PDB, ~39% are represented by the 10 most abundant folds (AKA superfolds).
121
GAMA Miguel Angel @miangoar.bsky.social · 24/10/2025
I strongly recommend making cat-based diagrams to illustrate complex topics in protein science: "Figure 4 considers [...] invariance and equivariance with respect to translations and rotations in 3D. For illustration purposes, the figure includes a series of cat cartoons in 2D."
100
GAMA Miguel Angel @miangoar.bsky.social · 09/10/2025
I just want to create hype and say that I made a 10-class course to introduce people to AI-driven protein design. It’s around 750 slides and will be freely available for anyone who wants to use them and, most importantly, improve them. Stay tuned :)
042
GAMA Miguel Angel @miangoar.bsky.social · 05/09/2025
This is a breakthrough for protein science🔥AFAIK this is the largest protein DB, with >100B seqs (3B clustered at 50%). New biology will come from LOGAN: new folds, topologies, etc. You can also improve your AlphaFold models by building better MSAs. Future AI models will also use LOGAN for training
041
GAMA Miguel Angel @miangoar.bsky.social · 27/08/2025
12/13 Bindcraft started as a binder design tutorial for the Boston Protein Design and Modeling Club, and it evolved into one of the most promising tools in AI-based protein design. And Importantly, it is open-source!🤗 Congrats to all the authors!
110
GAMA Miguel Angel @miangoar.bsky.social · 27/08/2025
11/13 The authors have gone a step further and are currently developing BoltzDesign1, which instead of designing binders, focuses on biomolecular interactions between proteins and small molecules. However, one of the main limitations of both AIs is their high computational cost.
100
GAMA Miguel Angel @miangoar.bsky.social · 27/08/2025
9/13 the most important results IMO was the determination of atomic structures of four binders, where in all cases, the computational designs were highly consistent with the experimentally determined ones.
100
GAMA Miguel Angel @miangoar.bsky.social · 27/08/2025
8/13 They designed binders targeting: *proteins with no known binding sites *membrane proteins , which are much harder than intra/extra-cellular proteins *proteins lacking evolutionary information *proteins that interact with DNA/RNA *medically relevant proteins such as those causing allergies
100
GAMA Miguel Angel @miangoar.bsky.social · 27/08/2025
7/13 Then it uses ProteinMPNN to optimize for solubility, increasing the chances of experimental success. Finally, uses AF2 to predict the structure. To demonstrate Bindcraft’s utility, the authors carried out many wet-lab experiments, something not as common as I would like.
100
GAMA Miguel Angel @miangoar.bsky.social · 27/08/2025
6/13 Bindcraft takes advantage of this by first proposing a random seq and predicting its structure to assess how well it interacts with the target protein. It then uses info from each interaction, successful or not, to optimize the seqs until it arrives at a credible interaction
100
GAMA Miguel Angel @miangoar.bsky.social · 27/08/2025
5/13 Bindcraft is an improved version of AlphaFold2, specifically AF-Multimer, which predicts the structure of protein complexes. Having been trained on thousands of structures, AF-Multimer learned to identify which sites are most likely to form protein–protein interactions.
120
GAMA Miguel Angel @miangoar.bsky.social · 27/08/2025
4/13 Bindcraft designs both the sequence and structure of binders, achieving a success rate between 10-100%, since designing large or complex binders is more challenging. This is enormous, considering that our previous best physics/biochemistry-based methods reached a 0.1%.
110
GAMA Miguel Angel @miangoar.bsky.social · 27/08/2025
3/13 We have learned how to design PPI so that one protein, called a binder, can bind to another and regulate it. e.g., cancer drugs are binders. However, designing binders requires yrs of research and detailed biomolecular knowledge. So, what if we teach an AI to design binders?
100
GAMA Miguel Angel @miangoar.bsky.social · 27/08/2025
2/13 Proteins carry out many functions on their own, but when they interact with each other, they generate a diversity of mechanisms that expand and regulate those functions. PPI arose over millions of years of evolution, giving rise to processes as complex as metabolism.
100
GAMA Miguel Angel @miangoar.bsky.social · 27/08/2025
1/13 🧵 Today, Bindcraft was published in @nature.com , one of the most famous AIs in biology for designing protein–protein interactions (PPI). In my opinion. Bindcraft represents one of the most important advances in the post–AlphaFold2 era.
1102
GAMA Miguel Angel @miangoar.bsky.social · 27/08/2025
Does anyone know of a recent comparison of the main structural classification schemes of proteins and guidance on when to choose one? Something like this but including ECOD and perhaps seq-based schemes like Pfam, SUPERFAMILY and CDD. Img source (2020) pubmed.ncbi.nlm.nih.gov/32302382/
151
GAMA Miguel Angel @miangoar.bsky.social · 21/07/2025
I think that this pic from the original paper illustrates very well how fast Diamond2 is. I'm also waiting for the final publication of the Diamond Deepclust DB that is the largest protein seq DB AFAIK Sensitive protein alignments at tree-of-life scale using DIAMOND www.nature.com/articles/s41...
042
GAMA Miguel Angel @miangoar.bsky.social · 09/07/2025
9/9 Perhaps the ~1.6k de novo designs collected by The PDA should be a good choice to evaluate further models, which also need to take into consideration ways to prevent data leakage. OOD appears to be harder than ID generalization and each task requieres a particular split bsky.app/profile/bris...
060
GAMA Miguel Angel @miangoar.bsky.social · 09/07/2025
8/9 This is the only paper AFAIK that shows experimental validation for generated seqs below a 30% seqID threshold. They also demonstrate how their protocol based on ESM2 generates proteins with motifs surrounded by different contexts (D), as well as completely new motifs (E).
110
GAMA Miguel Angel @miangoar.bsky.social · 09/07/2025
7/9 So, what does OOD generalization mean? I don't know, but perhaps the answer lies in de novo proteins. In this paper, they applied fixed-backbone generation for various de novo designs and experimentally validated them, and their distribution is different from natural ones.
120
GAMA Miguel Angel @miangoar.bsky.social · 09/07/2025
6/9 So, the problem seems to be the seqs. What if we split by structures? Perhaps it gets better, but there are still issues like convergence (e.g. pockets of serinproteases or carbonic anhydrases) and fold switching (estimated to occur in ~4% of the PDB)
120
GAMA Miguel Angel @miangoar.bsky.social · 09/07/2025
5/9 Data leakage is not the only problem. We have others such as: 1 historical contingency (i.e., today’s proteins are not truly representative of the full diversity). 2 Proteins are not IID (due to factors like superfolds and shared motifs between non-homologous proteins)
110
GAMA Miguel Angel @miangoar.bsky.social · 09/07/2025
4/9 However, the commonly used seqID thresholds lie in the twilight zone (between 25-30% bsky.app/profile/mian...). This means that not all homologous seqs are divided correctly, leading to data leakage. e.g. Although these betalactamases share 12% seqID, they are homologous.
150