Sign in

Raj Movva

@rajmovva.bsky.social
257 followers 131 following 47 posts

NLP, ML & society, healthcare. PhD student at Berkeley, previously CS at MIT. rajivmovva.com

PostsRepliesMedia
Reposted by Raj Movva
Kenny Peng @kennypeng.bsky.social · 24/04/2026
We made traversle.io, a new daily word game! The goal is to traverse from a start word to a target word through a network of related words. (Our motivating question: is it possible to construct a network that allows human navigation?)
3164
Reposted by Raj Movva
Kenny Peng @kennypeng.bsky.social · 17/02/2026
New paper! The Linear Representation Hypothesis is a powerful intuition for how language models work, but lacks formalization. We give a mathematical framework in which we can ask and answer a basic question: how many features can be stored under the hypothesis? 🧵 arxiv.org/abs/2602.11246
14514
Reposted by Raj Movva
Isabel Silva Corpus @isabelcorpus.bsky.social · 01/12/2025
Excited to share a new working paper! What happened when Change.org integrated an AI writing tool into their platform? We provide causal evidence that petition text changed significantly while outcomes did not improve. 1/ arxiv.org/abs/2511.13949
arxiv.org
Introducing AI to an Online Petition Platform Changed Outputs but not Outcomes
The rapid integration of AI writing tools into online platforms raises critical questions about their impact on content production and outcomes. We leverage a unique natural experiment on Change$.$org...
45518
Raj Movva @rajmovva.bsky.social · 28/10/2025
another banger from @louisathomas.bsky.social www.newyorker.com/sports/sport...
newyorker.com
Why Can’t the N.B.A. Move On from Its Old Stars?
Even as the league drastically evolves, the narratives around it are still orbiting its aging icons.
020
Reposted by Raj Movva
Divya Shanmugam @dmshanmugam.bsky.social · 17/10/2025
New #NeurIPS2025 paper: how should we evaluate machine learning models without a large, labeled dataset? We introduce Semi-Supervised Model Evaluation (SSME), which uses labeled and unlabeled data to estimate performance! We find SSME is far more accurate than standard methods.
1217
Reposted by Raj Movva
Divya Shanmugam @dmshanmugam.bsky.social · 14/10/2025
I am on the job market this year! My research advances methods for reliable machine learning from real-world data, with a focus on healthcare. Happy to chat if this is of interest to you or your department/team.
22812
Reposted by Raj Movva
Vauhini Vara @vauhinivara.bsky.social · 02/09/2025
I've been working for many months on this article on Silicon Valley's under-the-radar role in bringing AI into schools across the US. I really hope you'll read it — here's a gift link — but I'll tell you some of the highlights in this thread. (1/x)
bloomberg.com
How Chatbots and AI Are Already Transforming Kids' Classrooms
Educators across the country are bringing chatbots into their lesson plans. Will it help kids learn or is it just another doomed ed-tech fad?
511158
Reposted by Raj Movva
Emma Pierson @emmapierson.bsky.social · 22/08/2025
🚨 New postdoc position in our lab at Berkeley EECS! 🚨 (please reshare) We seek applicants with experience in language modeling who are excited about high-impact applications in the health and social sciences! More info in thread 1/3
12112
Raj Movva @rajmovva.bsky.social · 18/08/2025
This is great, & there's clear analogy to the burgeoning mechanism design community for AI alignment: who is providing RLHF votes? Do their preferences reflect yours? Discussions about social choice and collective constitutions are interesting, but "what and who is in the data" is just as important.
040
Raj Movva @rajmovva.bsky.social · 05/08/2025
📢New POSITION PAPER: Use Sparse Autoencoders to Discover Unknown Concepts, Not to Act on Known Concepts Despite recent results, SAEs aren't dead! They can still be useful to mech interp, and also much more broadly: across FAccT, computational social science, and ML4H. 🧵
1404
Reposted by Raj Movva
Kenny Peng @kennypeng.bsky.social · 03/07/2025
Are LLMs correlated when they make mistakes? In our new ICML paper, we answer this question using responses of >350 LLMs. We find substantial correlation. On one dataset, LLMs agree on the wrong answer ~2x more than they would at random. 🧵(1/7) arxiv.org/abs/2506.07962
Heat map showing that more accurate models have more correlated errors.
1497
Reposted by Raj Movva
Ben Recht @beenwrekt.bsky.social · 24/06/2025
@jessica.bsky.social on individual reporting as a means to build collective knowledge.
argmin.net
Individual experiences and collective evidence
Jessica Dai on theory for the world as it could be
182
Raj Movva @rajmovva.bsky.social · 17/06/2025
ARR question: If I submit to a cycle, how long do those reviews "last"? e.g. if I submit to the July cycle but can't go to AACL, can I commit my July reviews to the conference associated with the next (October) cycle? @aclrollingreview.bsky.social
121
Reposted by Raj Movva
Divya Shanmugam @dmshanmugam.bsky.social · 14/06/2025
New work 🎉: conformal classifiers return sets of classes for each example, with a probabilistic guarantee the true class is included. But these sets can be too large to be useful. In our #CVPR2025 paper, we propose a method to make them more compact without sacrificing coverage.
A gif explaining the value of test-time augmentation to conformal classification. The video begins with an illustration of TTA reducing the size of the  predicted set of classes for a dog image, and goes on to explain that this is because TTA promotes the true class's predicted probability to be higher, even when it's predicted to be unlikely.
3226
Raj Movva @rajmovva.bsky.social · 06/06/2025
I would like to spend up to 5-10 hours to learn about basic macroeconomics (I know it's maybe fake, but setting that aside for a moment...). Does anyone have any recommendations?
000
Raj Movva @rajmovva.bsky.social · 10/05/2025
People love to hate on the transition 3-pointer as evidence of how the 3 has ruined basketball, but I think it's usually just the right play... if you have numbers in transition, your teammate can easily get a putback off a miss, so might as well try the 3
010
Raj Movva @rajmovva.bsky.social · 05/05/2025
We'll present HypotheSAEs at ICML this summer! 🎉 Draft: arxiv.org/abs/2502.04382 We're continuing to cook up new updates for our Python package: github.com/rmovva/Hypot... (Recently, "Matryoshka SAEs", which help extract coarse and granular concepts without as much hyperparameter fiddling.)
github.com
GitHub - rmovva/HypotheSAEs: Hypothesizing interpretable relationships in text datasets using sparse autoencoders.
Hypothesizing interpretable relationships in text datasets using sparse autoencoders. - rmovva/HypotheSAEs
1102
Raj Movva @rajmovva.bsky.social · 03/05/2025
Yesterday's Game 6 was depressing, and this article precisely delineated the reasons why. And sometimes, a precise retelling of what you're feeling is all you need to feel better. www.nytimes.com/athletic/633... @thompsonscribe.bsky.social
nytimes.com
These Warriors are old, tired and in trouble as Game 7 looms against Rockets
They're not done yet. Maybe a legendary performance awaits on Sunday. But the Warriors look like they're out of gas and out of answers.
020
Raj Movva @rajmovva.bsky.social · 01/05/2025
Check out Erica's nice work. They not only develop a well-grounded model for disparities in disease progression, but also conduct experiments with real NYP cardiology data! (Anyone who works in healthcare knows how much of a feat it is to use data other than MIMIC)
040
Reposted by Raj Movva
Emma Pierson @emmapierson.bsky.social · 25/04/2025
The US government recently flagged my scientific grant in its "woke DEI database". Many people have asked me what I will do. My answer today in Nature. We will not be cowed. We will keep using AI to build a fairer, healthier world. www.nature.com/articles/d41...
nature.com
My ‘woke DEI’ grant has been flagged for scrutiny. Where do I go from here?
My work in making artificial intelligence fair has been noticed by US officials intent on ending ‘class warfare propaganda’.
13813
Raj Movva @rajmovva.bsky.social · 22/04/2025
Good reading for PhD students on why meeting scheduling might be more important than you think: paulgraham.com/makersschedu...
paulgraham.com
Maker's Schedule, Manager's Schedule
000
Reposted by Raj Movva
Serina Chang @serinachang5.bsky.social · 11/04/2025
1st post on bsky! What happens when a static benchmark comes to life? ✨ Introducing ChatBench, a large-scale user study where we *converted* MMLU questions into thousands of user-AI conversations. Then, we trained a user simulator on ChatBench to generate user-AI outcomes on unseen questions. 1/ 🧵
151
Raj Movva @rajmovva.bsky.social · 10/04/2025
A weird (and seemingly fixable quirk) with ChatGPT Deep Research is hallucinations even when there is a linked citation with a highlighted passage: "X et al find Y (arXiv link)" and you click on the arXiv link, which is neither written by X nor do they find Y...
000
Reposted by Raj Movva
Mor Naaman @informor.bsky.social · 02/04/2025
Stop everything you are doing and review the best Open Data chart you will see all year. Credit: @nkgarg.bsky.social's lab
2299
Raj Movva @rajmovva.bsky.social · 02/04/2025
My labmates really cooked with this one
090
Reposted by Raj Movva
Gabriel Agostini @gsagostini.bsky.social · 28/03/2025
Migration data lets us study responses to environmental disasters, social change patterns, policy impacts, etc. But public data is too coarse, obscuring these important phenomena! We build MIGRATE: a dataset of yearly flows between 47 billion pairs of US Census Block Groups. 1/5
54118
Raj Movva @rajmovva.bsky.social · 21/03/2025
I've learned MCMC 5+ different times in the last 10 years from courses, videos, blogposts, etc, and it never clicked. 30 minutes chatting with Claude this morning, and I finally feel like I've figured it out...
010
Reposted by Raj Movva
Nikhil Garg @nkgarg.bsky.social · 18/03/2025
Really proud of @rajmovva.bsky.social and @kennypeng.bsky.social for this work! We hope that it's useful, and are already using it for many followup projects Preprint: arxiv.org/abs/2502.04382 Python package: github.com/rmovva/Hypot... Demo: hypothesaes.org
1182
Reposted by Raj Movva
Kenny Peng @kennypeng.bsky.social · 18/03/2025
(1/n) New paper/code! Sparse Autoencoders for Hypothesis Generation HypotheSAEs generates interpretable features of text data that predict a target variable: What features predict clicks from headlines / party from congressional speech / rating from Yelp review? arxiv.org/abs/2502.04382
1145
Raj Movva @rajmovva.bsky.social · 18/03/2025
💡New preprint & Python package: We use sparse autoencoders to generate hypotheses from large text datasets. Our method, HypotheSAEs, produces interpretable text features that predict a target variable, e.g. features in news headlines that predict engagement. 🧵1/
14013
Raj Movva @rajmovva.bsky.social · 15/03/2025
An interesting dynamic on Google reviews is that I tend to trust businesses which write custom replies to negative reviews, even if the reply doesn't really resolve the issue (e.g. see below). I distrust businesses which have stock replies, even if those replies offer to resolve the issue somehow.
110
Reposted by Raj Movva
Nikhil Garg @nkgarg.bsky.social · 10/03/2025
*Please repost* @sjgreenwood.bsky.social and I just launched a new personalized feed (*please pin*) that we hope will become a "must use" for #academicsky. The feed shows posts about papers filtered by *your* follower network. It's become my default Bluesky experience bsky.app/profile/pape...
23535299
Reposted by Raj Movva
Nikhil Garg @nkgarg.bsky.social · 27/02/2025
Now online @pnasnexus.org! Many discrimination auditing and electoral tasks use ML to predict race/ethnicity – by discretizing continuous scores. Can the discretization process cause bias in labels and downstream tasks? Yes! Led by @evandyx.bsky.social academic.oup.com/pnasnexus/ar...
Paper screenshot. Title: Addressing discretization-induced bias in demographic prediction 


Abstract: Racial and other demographic imputation is necessary for many applications, especially in auditing disparities and outreach targeting in political campaigns. The canonical approach is to construct continuous predictions—e.g. based on name and geography—and then to often discretize the predictions by selecting the most likely class (argmax), potentially with a minimum threshold (thresholding). We study how this practice produces discretization bias. For example, we show that argmax labeling, as used by a prominent commercial voter file vendor to impute race/ethnicity, results in a substantial under-count of Black voters, e.g. by 28.2% points in North Carolina. This bias can have substantial implications in downstream tasks that use such labels. We then introduce a joint optimization approach—and a tractable data-driven threshold heuristic—that can eliminate this bias, with negligible individual-level accuracy loss. Finally, we theoretically analyze discretization bias, show that calibrated continuous models are insufficient to eliminate it, and that an approach such as ours is necessary. Broadly, we warn researchers and practitioners against discretizing continuous demographic predictions without considering downstream consequences.
1285
Reposted by Raj Movva
Kenny Peng @kennypeng.bsky.social · 27/02/2025
In new work, we show a "No Free Lunch Theorem" for human-AI Collaboration (w/ @nkgarg.bsky.social and Jon Kleinberg). (And if you're at #AAAI, I'm presenting at 11:15am today in the Humans and AI session. Poster 12:30-2:30.) arxiv.org/abs/2411.15230
0123
Raj Movva @rajmovva.bsky.social · 18/02/2025
Is there a canoncial version of modernBERT for computing sentence/document embeddings? Primarily for classification & clustering (rather than retrieval), but any are fine. If not, what's the best practice to finetune for this? cc @benjaminwarner.dev @howard.fm
100
Raj Movva @rajmovva.bsky.social · 13/02/2025
Does anyone have a rule-of-thumb for how many words/sentences you can represent in a 3072-dim text embedding (using openai's large embs)? Like, there might be a token limit (here, 8192), but perhaps the embeddings are quite lossy beyond some smaller number of tokens?
030
Reposted by Raj Movva
dan bateyko @dbateyko.bsky.social · 30/01/2025
A coming flood of federal grant proposals will claim AI makes projects cheaper & better. Many will be empty promises. This admin needs to place the best bets it can on AI. For the Federation of American Scientists' Day One Project, my memo on AI grant governance fas.org/publication/...
fas.org
Bring AI Governance to Competitive Grants
To tackle AI risks in grant spending, grant-making agencies should adopt trustworthy AI practices in their grant competitions and start enforcing them against reckless grantees.
0103
Reposted by Raj Movva
Naomi Saphra @nsaphra.bsky.social · 27/01/2025
One of my grand interpretability goals is to improve human scientific understanding by analyzing scientific discovery models, but this is the most convincing case yet that we CAN learn from model interpretation: Chess grandmasters learned new play concepts from AlphaZero's internal representations.
arxiv.org
Bridging the Human-AI Knowledge Gap: Concept Discovery and Transfer in AlphaZero
Artificial Intelligence (AI) systems have made remarkable progress, attaining super-human performance across various domains. This presents us with an opportunity to further human knowledge and improv...
210923
Reposted by Raj Movva
Emma Pierson @emmapierson.bsky.social · 13/01/2025
Our article on using LLMs to promote health equity is out in New England Journal of Medicine AI! 85% of equity-related LLM papers focus on *harms*. But also vital are the equity-related *opportunities* LLMs create: detecting bias, extracting structured data, and improving access to health info.
1277
Reposted by Raj Movva
Divya Shanmugam @dmshanmugam.bsky.social · 18/12/2024
We have a new review on generative AI in medicine, to appear in the Annual Review of Biomedical Data Science! We cover over 250 papers in the recent literature to provide an updated overview of use cases and challenges for generative AI in medicine.
1238
Reposted by Raj Movva
Jonathan Frankle @jfrankle.com · 13/12/2024
Reflections on NeurIPS: There's always a big theme people seem to be preoccupied with. This year, it was the continuation of scaling/progress. Will it continue? What will the next generation of models hold? I even got to sass Dylan Patel (not on bsky) over it. Here are my personal thoughts 🧵
35911
Raj Movva @rajmovva.bsky.social · 30/11/2024
genAI has made us more suspicious that emails, cover letters, artworks, etc. are produced by AI. this shift forces us to change our behavior in order to prove our human-ness: a "burden of authenticity". waking my account up to share a recent blog post on the subject: rajivmovva.com/2024/11/08/g...
161
Reposted by Raj Movva
Lucy Li @lucy3.bsky.social · 18/11/2024
The soul-searching journey for figuring out what research area is right for you is tricky since so many papers are cool. I tell my early career students that they should try to differentiate papers that they'd like to read 📖, implement 🔨, *and* write 📝 from papers that they'd only like to read 📖.
46711