Sign in

Calvin McCarter

@calvinmccarter.bsky.social
358 followers 650 following 131 posts

calvinmccarter.com

PostsRepliesMedia
Reposted by Calvin McCarter
arXiv q-bio.GN Genomics @qbiogn-bot.bsky.social · 01/09/2026
Calvin McCarter: Confounder-Aware Feature Correction for Single-Cell Batch Integration arxiv.org/abs/2608.28849 arxiv.org/pdf/2608.28849 arxiv.org/html/2608.28849
031
Calvin McCarter @calvinmccarter.bsky.social · 09/07/2026
A video of my talk at the ML4PE seminar, "How to Make the Most of Your Masked Language Model for Protein Engineering", is now online: youtube.com/watch?v=Lcmo...
m.youtube.com
How to Make the Most of Your Masked Language Model for Protein Engineering - Calvin McCarter
YouTube video by ML for protein engineering seminar series
000
Calvin McCarter @calvinmccarter.bsky.social · 30/04/2026
should we be surprised at this point?
000
Calvin McCarter @calvinmccarter.bsky.social · 23/04/2026
🌎
081
Calvin McCarter @calvinmccarter.bsky.social · 26/03/2026
Of all sad words of LLM, The saddest are these: "Failed to fetch arxiv.org" again
000
Calvin McCarter @calvinmccarter.bsky.social · 17/03/2026
a metaphor for x (formerly twitter) @norvid-studies.bsky.social
152
Calvin McCarter @calvinmccarter.bsky.social · 13/03/2026
We've been investing heavily in better protein language models (PLMs), but relatively little work addresses how to best generate with them. We present a new search-based method for PLMs and exhaustively benchmark models and methods, including with in vitro data from antibody therapeutics campaigns.🧵
2203
Calvin McCarter @calvinmccarter.bsky.social · 10/03/2026
The relationship between Spearman correlation and AUROC computed on the same data (one continuous variable, one binary variable) is surprising. On evenly balanced data, Spearman=0.75 leads to AUROC=0.93. As the binary variable gets more imbalanced, the discrepancy gets bigger.
020
Calvin McCarter @calvinmccarter.bsky.social · 10/03/2026
The Mann-Whitney U statistic and rank-biserial correlation are simply rescalings of AUROC: U = AUC * n_pos * n_neg r_rb = 2 * AUC - 1
020
Calvin McCarter @calvinmccarter.bsky.social · 19/08/2025
What makes tabular data unique (and interesting!) is not merely that it's arranged into rows and columns. New blogpost: calvinmccarter.substack.com/p/the-idiosy...
calvinmccarter.substack.com
The idiosyncrasies of tabular data
The things that make tabular data different
030
Calvin McCarter @calvinmccarter.bsky.social · 11/03/2025
Has anyone tried far-UVC in their home? It's now dropped into the ~$300 price range where I'm interested in trying it for myself. substack.com/home/post/p-...
substack.com
Flipping the switch on far-UVC
We’ve known about far-UVC’s promise for a decade. Why isn't it everywhere?
031
Calvin McCarter @calvinmccarter.bsky.social · 12/01/2025
Here's a link to the report: cleanlabelproject.org/wp-content/u... (TLDR heuristics: whey is better than plant-based, non-organic is better than organic, unflavored is better than chocolate-flavored)
cleanlabelproject.org
100
Calvin McCarter @calvinmccarter.bsky.social · 08/01/2025
Ultra exciting! And it's gratifying to see that this method uses the kernel density integral preprocessing method that I published in @tmlr-pub.bsky.social (2023). (One takeaway: even if your ML research focus isn't deep learning, pursue directions that complement rather than compete with it.)
020
Reposted by Calvin McCarter
Alan Amin @handle.invalid · 17/12/2024
How do you go from a hit in your antibody screen to a suitable drug? Now introducing CloneBO: we optimize antibodies in the lab by teaching a generative model how we optimize them in our bodies! w/ Nat Gruver, Yilun Kuang, Lily Li, @andrewgwils.bsky.social and the team at Big Hat! 1/7
1111
Calvin McCarter @calvinmccarter.bsky.social · 27/11/2024
Will neural networks achieve AGI before they figure out how to do tokenization internally? Or, on the way to AGI, will they invent a tokenizer that "just works"? Another way of framing it: is tokenization AGI-complete?
010
Calvin McCarter @calvinmccarter.bsky.social · 24/11/2024
JMLR and TMLR provide better reviewing, but worse publicity for accepted papers. Their only real platform, their accounts on X, get very little engagement these days. They should be on 🦋. Better yet, *MLR should have a biweekly or monthly arXiv-style email newsletter. @thegautamkamath.bsky.social
040
Calvin McCarter @calvinmccarter.bsky.social · 22/11/2024
Pretty interesting new paper in TMLR: "Controlling the Fidelity and Diversity of Deep Generative Models via Pseudo Density" openreview.net/forum?id=8Vk...
openreview.net
Controlling the Fidelity and Diversity of Deep Generative Models...
We introduce an approach to bias deep generative models, such as GANs and diffusion models, towards generating data with either enhanced fidelity or increased diversity. Our approach involves...
100
Calvin McCarter @calvinmccarter.bsky.social · 17/11/2024
Twitter is still a better place to share ML papers than Bluesky:
041
Calvin McCarter @calvinmccarter.bsky.social · 17/11/2024
Excited to share a new paper published in TMLR, "Towards Backwards-Compatible Data with Confounded Domain Adaptation", solving a key problem in using AI for biology. How can you combine datasets across settings, when "what you measure" and "how you measure it" are confounded? 1/7
130
Calvin McCarter @calvinmccarter.bsky.social · 15/11/2024
Suppose you collect tissue samples with varying degrees of freshness, and want to correct for this. But what if freshness is also correlated with true biological differences (eg healthy vs cancer tissue)? My new paper out in TMLR addresses this problem: openreview.net/forum?id=GSp...
openreview.net
Towards Backwards-Compatible Data with Confounded Domain Adaptation
Most current domain adaptation methods address either covariate shift or label shift, but are not applicable where they occur simultaneously and are confounded with each other. Domain adaptation...
021