Sign in

Kevin K. Yang 楊凱筌

@kevinkaichuang.bsky.social
6.1K followers 3.4K following 411 posts

Principal Researcher in BioML at Microsoft Research. He/him/他. 🇹🇼 yangkky.github.io

PostsRepliesMedia
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 05/10/2026
Combine a chemical knowledge, generative modeling, and structure prediction into a pipeline that can design enzymes for new-to-nature chemistry. @zacharywu.bsky.social @alexechu.bsky.social @francesarnold.bsky.social @pushmeet.bsky.social @juewang.bsky.social www.biorxiv.org/content/10.6...
0124
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 23/09/2026
A genome language model reveals the interactions hidden in non-coding DNA, pointing to new putative RNAs and repetitive elements in bacterial genomes. @garykbrixi.bsky.social @maxewilkinson.bsky.social @mfgrp.bsky.social @brianhie.bsky.social www.biorxiv.org/content/10.6...
1177
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 22/09/2026
Use reward-guided search in a generative model's latent space to design high-affinity protein binders to proteins and carbohydrates. @karstenkreis.bsky.social @pierceogden.bsky.social www.biorxiv.org/content/10.6...
0157
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 18/09/2026
Use a genomic language model to optimize a polyketide synthase go produce a non-natural product. @nathanlanclos.bsky.social @ancornman1.bsky.social @kortemmelab.bsky.social @keaslinglab.bsky.social @microyunha.bsky.social www.biorxiv.org/content/10.6...
0111
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 16/09/2026
FoldDir: A flow-matching structure-conditioned protein sequence design model with excellent computational results and some functional nanobody redesigns. @rbaltman.bsky.social www.biorxiv.org/content/10.6...
072
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 14/09/2026
A nice review on ways to discover and engineer proteins that use fewer amino acids while maintaining structure and function. @FilipBuchel @KlaraH_lab www.sciencedirect.com/science/arti...
051
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 10/09/2026
The cost of gene synthesis prevents machine-learning guided directed evolution from reaching it's full potential. A very nice perspective by the inimitable Bruce Wittmann. arxiv.org/abs/2609.03046
072
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 08/09/2026
A fast method for building phylogenetic trees with millions of sequences and then a model that leverages the tree structure to predict viral protein evolution! @seyonec.bsky.social @claudiadriscoll.bsky.social @garykbrixi.bsky.social @brianhie.bsky.social www.biorxiv.org/content/10.6...
0153
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 31/08/2026
The largest experimental protein sequence-fitness dataset: 13k measurements across 3 proteins reveals that fitness landscapes are rugged, there are many functional sequences between orthologs, and ML models can't distinguish functional and non-functional sequences. www.biorxiv.org/content/10.6...
0174
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 28/08/2026
Use machine learning guided directed evolution to make a red-shifted fluorescent protein 4x brighter. www.biorxiv.org/content/10.6...
0101
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 24/08/2026
Curate a molecule <-> text dataset, train a model, and use it to discover molecules that fight antibiotic resistance by deactivating beta-lactamase enzymes. @kosonocky.bsky.social www.biorxiv.org/content/10.6...
1132
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 20/08/2026
Binder design is nice and all, but here we have three agents sharing data and autonomously controlling a lab to design enzymes with shifted substrate scopes and high activity! @cobanbrooks.bsky.social @pascalnotin.bsky.social @philromero.bsky.social www.biorxiv.org/content/10.6...
0205
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 12/08/2026
ApexFold: condition predictions of peptide secondary structure on the chemical environment @delafuentelab.bsky.social www.biorxiv.org/content/10.6...
0112
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 10/08/2026
Use machine learning to design libraries of enzymes with broad substrate scope as starting points for engineering. @francesarnold.bsky.social www.biorxiv.org/content/10.6...
1182
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 06/08/2026
Use a rank regression model to engineer improved enzymes www.nature.com/articles/s41...
091
Reposted by Kevin K. Yang 楊凱筌
Luca Pinello @lucapinello.bsky.social · 17/07/2026
1/ I'm excited to share that we're launching NECB 2026, the inaugural New England Computational Biology Symposium. Oct 1-2 at Microsoft Research New England, Cambridge. Two days of keynotes,talks, and posters to bring our community together across institutions.Space is limited. newenglandcompbio.org
11710
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 10/07/2026
Gaussian process kernels that encode evolutionary information and local linearity improve protein sequence-fitness predictions. arxiv.org/abs/2606.11057
091
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 01/07/2026
Proto: A probabilistic programming language for designing biomolecules. Absolutely beautiful paper from top to bottom. @adititm.bsky.social @bviggiano.bsky.social @brianhie.bsky.social www.biorxiv.org/content/10.6...
0102
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 19/06/2026
TadA-Bench: a 31-round, million-variant wet-lab enzyme engineering dataset. arxiv.org/abs/2606.02624
0121
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 15/06/2026
Aggregating many experiments into one zero-shot protein language model score obscures that current models cannot meaningfully rank a set of fit mutations or prioritize new-to-nature functions @clauswilke.com www.biorxiv.org/content/10.6...
0216
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 29/05/2026
Learn which structural domains are compatible with each other, then generate protein sequences conditioned on the desired structural domains. www.biorxiv.org/content/10.6...
070
Reposted by Kevin K. Yang 楊凱筌
Jase Gehring @skyjase.bsky.social · 15/05/2026
very very cool! i had the privilege of asking Mohammed AlQuraishi's opinion on this topic a couple years ago
061
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 15/05/2026
Screen 1M random protein sequences to discover that biology-like folds are accessible from random sequences with surprising frequency www.biorxiv.org/content/10.6...
1368
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 11/05/2026
Genie3: all-atom SE3-equivariance for fast and performant protein design. @yeqinglin.bsky.social @moalquraishi.bsky.social www.biorxiv.org/content/10.6...
0102
Reposted by Kevin K. Yang 楊凱筌
Clay Kosonocky @kosonocky.bsky.social · 21/04/2026
Have you wondered what the wet lab success rates are for current AI-driven protein design models? Look no further! In our new review, @kevinkaichuang.bsky.social @avapamini.bsky.social, @sarahalamdari.bsky.social, and I report wet lab success rates for *over 200* different protein design tasks 🧬💻
13214
Reposted by Kevin K. Yang 楊凱筌
Alex Lu @alexijie.bsky.social · 18/04/2026
Check out how we exploit observations of homology from evolution to design enhancers, even when we don't have prior knowledge or ability to specify function!
032
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 17/04/2026
EnhancAR: Use evolution and deep learning to design enhancers with desired expression profiles! Had a lot of fun working with Andrew Duncan, Micaela Consens @alexijie.bsky.social @lcrawford.bsky.social Jennifer Mitchell and Alan Moses! www.biorxiv.org/content/10.6...
0133
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 14/04/2026
A very nice review of AI in protein engineering from @jlistgarten.bsky.social www.science.org/doi/10.1126/...
0143
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 10/04/2026
DISCO: a steerable, multimodal protein diffusion model that can generate enzymes for new-to-nature chemistry. arxiv.org/abs/2604.05181
2173
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 08/04/2026
Sequence Display generates large-scale protein sequence–activity datasets by simultaneously reading out the sequence of interest and a downstream editing efficiency - perfect for data hungry ML methods! Happy to have played a small part in this!
1114
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 06/04/2026
Design enzymes using a model that generates sequence and structure given the substrate as a ligand. @lileics.bsky.social www.biorxiv.org/content/10.6...
0153
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 03/04/2026
Genomic context and protein sequence can be used together to discover bacterial phage defense systems. @emordret.bsky.social @alexhv.bsky.social @audeber.bsky.social www.science.org/doi/10.1126/...
1111
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 02/04/2026
Distilling the most confident mutation effect scores across different versions of the same protein language model results in a superior sequence-only variant effect predictor. @vntranos.bsky.social www.nature.com/articles/s41...
0133
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 13/03/2026
It's great when a big lab expands a subfigure from your paper into its own paper without any acknowledgment
0101
Reposted by Kevin K. Yang 楊凱筌
Diego del Alamo @delalamo.xyz · 26/02/2026
This plot is quite the indictment of fine-tuned PLMs, showing how performance is entirely data-dependent and, at the upper end of performance, equally achievable with randomized model weights
0132
Reposted by Kevin K. Yang 楊凱筌
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 25/02/2026
We made FLIP2, a protein fitness benchmark spanning seven new datasets, including enzymes, protein-protein interactions, and light-sensitive proteins, as well as splits that measure generalization relevant to real-world protein engineering campaigns.
15215
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 25/02/2026
We made FLIP2, a protein fitness benchmark spanning seven new datasets, including enzymes, protein-protein interactions, and light-sensitive proteins, as well as splits that measure generalization relevant to real-world protein engineering campaigns.
15215
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 24/02/2026
Train a protein language model to predict one homolog from another given the amount of evolutionary time separating them. @antoinekoehl.bsky.social @junhaobearxiong.bsky.social www.biorxiv.org/content/10.6...
0120
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 15/02/2026
🇹🇼 🇹🇼 🇹🇼
080
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 04/02/2026
Excellent review on using generative models to design enzymes @noeliaferruz.bsky.social @lassemiddendorf.bsky.social arxiv.org/abs/2602.03779
0122
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 24/01/2026
Hot take: it's 2026 we don't need to be manually doing citations or asking llms to do them overleaf should just autogenerate the citations from dois
2274
Reposted by Kevin K. Yang 楊凱筌
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 16/01/2026
For no reason, I remembered today that I too once got to take a picture holding Nobel Prize that I didn't earn
21219
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 16/01/2026
For no reason, I remembered today that I too once got to take a picture holding Nobel Prize that I didn't earn
21219
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 13/01/2026
EDEN: a family of genomic language models trained on up to 9.7 trillion nucleotides from @basecamp-research.bsky.social's BaseData can design large serine recombinases, bridge recombinases, and antimicrobial peptides. www.biorxiv.org/content/10.6... Happy to have played a small part in this!
0195
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 07/01/2026
A dataset of 40 million protein families and an autoregressive model of protein families. Great to see other protein Atlases popping up after Dayhoff!! @judewells.bsky.social @dmmiller597.bsky.social www.biorxiv.org/content/10.6...
0298
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 16/12/2025
Finetune a codon-level language model with 30k tryptophan synthases, then generate diverse, functional, enzymes with broad substrate scopes. Théophile Lambert @jsunn-y.bsky.social @francesarnold.bsky.social www.biorxiv.org/content/10.1...
0224
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 12/12/2025
Enzyme Engineering Database (EnzEngDB): a platform for sharing and interpreting sequence–function relationships across protein engineering campaigns @francescazfl.bsky.social @jsunn-y.bsky.social @francesarnold.bsky.social @arianemora.bsky.social Paper: doi.org/10.1093/nar/... DB: enzengdb.org
1236
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 10/12/2025
An energy-based model of protein conformational space can be used to predict structure from sequence, sample from the conformational landscape, rank structures, and predict mutation effects. @sokrypton.org www.biorxiv.org/content/10.6...
04914
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 06/12/2025
Becoming a real Asian by making my kid practice Chinese characters and math while he waits for his Saturday cello class.
0110
Kevin K. Yang 楊凱筌 @kevinkaichuang.bsky.social · 05/12/2025
Train a model to identify circularly-permuted structural homologs, then use it to discover novel pairs of related proteins! @aidenosinetrip1 @abulnaga.bsky.social @sokrypton.org www.biorxiv.org/content/10.1...
0203