Reposted by Calvin McCarterarXiv q-bio.GN Genomics @qbiogn-bot.bsky.social · 01/09/2026Calvin McCarter: Confounder-Aware Feature Correction for Single-Cell Batch Integration arxiv.org/abs/2608.28849 arxiv.org/pdf/2608.28849 arxiv.org/html/2608.28849 031
Calvin McCarter @calvinmccarter.bsky.social · 09/07/2026A video of my talk at the ML4PE seminar, "How to Make the Most of Your Masked Language Model for Protein Engineering", is now online: youtube.com/watch?v=Lcmo...m.youtube.comHow to Make the Most of Your Masked Language Model for Protein Engineering - Calvin McCarterYouTube video by ML for protein engineering seminar series 000
Calvin McCarter @calvinmccarter.bsky.social · 26/03/2026Of all sad words of LLM, The saddest are these: "Failed to fetch arxiv.org" again 000
Calvin McCarter @calvinmccarter.bsky.social · 17/03/2026a metaphor for x (formerly twitter) @norvid-studies.bsky.social 152
Calvin McCarter @calvinmccarter.bsky.social · 13/03/2026We've been investing heavily in better protein language models (PLMs), but relatively little work addresses how to best generate with them. We present a new search-based method for PLMs and exhaustively benchmark models and methods, including with in vitro data from antibody therapeutics campaigns.🧵 2203
Calvin McCarter @calvinmccarter.bsky.social · 10/03/2026The relationship between Spearman correlation and AUROC computed on the same data (one continuous variable, one binary variable) is surprising. On evenly balanced data, Spearman=0.75 leads to AUROC=0.93. As the binary variable gets more imbalanced, the discrepancy gets bigger. 020
Calvin McCarter @calvinmccarter.bsky.social · 10/03/2026The Mann-Whitney U statistic and rank-biserial correlation are simply rescalings of AUROC: U = AUC * n_pos * n_neg r_rb = 2 * AUC - 1 020
Calvin McCarter @calvinmccarter.bsky.social · 19/08/2025What makes tabular data unique (and interesting!) is not merely that it's arranged into rows and columns. New blogpost: calvinmccarter.substack.com/p/the-idiosy...calvinmccarter.substack.comThe idiosyncrasies of tabular dataThe things that make tabular data different 030
Calvin McCarter @calvinmccarter.bsky.social · 11/03/2025Has anyone tried far-UVC in their home? It's now dropped into the ~$300 price range where I'm interested in trying it for myself. substack.com/home/post/p-...substack.comFlipping the switch on far-UVCWe’ve known about far-UVC’s promise for a decade. Why isn't it everywhere? 031
Calvin McCarter @calvinmccarter.bsky.social · 12/01/2025Here's a link to the report: cleanlabelproject.org/wp-content/u... (TLDR heuristics: whey is better than plant-based, non-organic is better than organic, unflavored is better than chocolate-flavored)cleanlabelproject.org 100
Calvin McCarter @calvinmccarter.bsky.social · 08/01/2025Ultra exciting! And it's gratifying to see that this method uses the kernel density integral preprocessing method that I published in @tmlr-pub.bsky.social (2023). (One takeaway: even if your ML research focus isn't deep learning, pursue directions that complement rather than compete with it.) 020
Reposted by Calvin McCarterAlan Amin @handle.invalid · 17/12/2024How do you go from a hit in your antibody screen to a suitable drug? Now introducing CloneBO: we optimize antibodies in the lab by teaching a generative model how we optimize them in our bodies! w/ Nat Gruver, Yilun Kuang, Lily Li, @andrewgwils.bsky.social and the team at Big Hat! 1/7 1111
Calvin McCarter @calvinmccarter.bsky.social · 27/11/2024Will neural networks achieve AGI before they figure out how to do tokenization internally? Or, on the way to AGI, will they invent a tokenizer that "just works"? Another way of framing it: is tokenization AGI-complete? 010
Calvin McCarter @calvinmccarter.bsky.social · 24/11/2024JMLR and TMLR provide better reviewing, but worse publicity for accepted papers. Their only real platform, their accounts on X, get very little engagement these days. They should be on 🦋. Better yet, *MLR should have a biweekly or monthly arXiv-style email newsletter. @thegautamkamath.bsky.social 040
Calvin McCarter @calvinmccarter.bsky.social · 22/11/2024Pretty interesting new paper in TMLR: "Controlling the Fidelity and Diversity of Deep Generative Models via Pseudo Density" openreview.net/forum?id=8Vk...openreview.netControlling the Fidelity and Diversity of Deep Generative Models...We introduce an approach to bias deep generative models, such as GANs and diffusion models, towards generating data with either enhanced fidelity or increased diversity. Our approach involves... 100
Calvin McCarter @calvinmccarter.bsky.social · 17/11/2024Twitter is still a better place to share ML papers than Bluesky: 041
Calvin McCarter @calvinmccarter.bsky.social · 17/11/2024Excited to share a new paper published in TMLR, "Towards Backwards-Compatible Data with Confounded Domain Adaptation", solving a key problem in using AI for biology. How can you combine datasets across settings, when "what you measure" and "how you measure it" are confounded? 1/7 130
Calvin McCarter @calvinmccarter.bsky.social · 15/11/2024Suppose you collect tissue samples with varying degrees of freshness, and want to correct for this. But what if freshness is also correlated with true biological differences (eg healthy vs cancer tissue)? My new paper out in TMLR addresses this problem: openreview.net/forum?id=GSp...openreview.netTowards Backwards-Compatible Data with Confounded Domain AdaptationMost current domain adaptation methods address either covariate shift or label shift, but are not applicable where they occur simultaneously and are confounded with each other. Domain adaptation... 021