Gaurav Kamath @grvkamath.bsky.social · 05/06/2026Super cool project that I really enjoyed being part of! tl;dr - when a human or model encounters new visual stimuli, how closely is it mapped to other, previously encountered concepts? (Come for weird dog-monster, stay for the science 🙂 ) 061
Gaurav Kamath @grvkamath.bsky.social · 01/04/2026Will always be this one for me :)). Sets an amazing standard for high-level clarity, direct but accessible writing, and methodological rigour direct.mit.edu/tacl/article...direct.mit.eduInherent Disagreements in Human Textual InferencesAbstract. We analyze human’s disagreements about the validity of natural language inferences. We show that, very often, disagreements are not dismissible as annotation “noise”, but rather persist as w... 020
Gaurav Kamath @grvkamath.bsky.social · 07/03/2026Just read the paper -- super cool, thanks for sharing!! Extremely relevant to + in line with our findings. Seem like a common thread here about the negative effects of RLVR on output variation... 010
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026Paper + data + code: grvkamath.github.io/probcopa-dem... | Major thanks to my co-authors, without whom this work would not have been possible: Sreenath Madathil, @sebschu.bsky.social, Marie-Catherine de Marneffe and @sivareddyg.bsky.social !(10/10)grvkamath.github.ioProbCOPA Interactive Explorer 061
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026Takeaway: reasoning LLMs are getting better and better on math and code—deterministic reasoning tasks. But we should also evaluate them on open-ended, inherently uncertain everyday reasoning! (9/10) 182
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026Ensembling all 8 models helps close the gap with human response distributions — but still doesn't reach human-human baseline similarity. (8/10) 140
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026We also looked inside reasoning chains. 90/100 sampled chains showed models explicitly enumerating alternative scenarios — a consistent reasoning pattern. Longer reasoning chains also correlate with more human judgment variation. (7/10) 120
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026More strikingly: for almost every item in our dataset, humans showed more response variation than models. Increasing temperature is not enough; models devolve into outputting random tokens before reaching human-level variation. (6/10) 160
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026Models do okay at the extremes — for inferences that humans deem very likely or very unlikely, model responses cluster in similar regions. But where humans are more uncertain, this similarity breaks down. (5/10) 110
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026We tested 8 contemporary reasoning LLMs (GPT-5, Gemini-3, Kimi-K2-Thinking, Claude Sonnet-4.5, and more). They show a clear pattern: unlike humans, models almost never return medium likelihood scores. (4/10) 150
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026Human ratings are graded and varied — people do not always commit to hard judgments, and show variation on their exact judgments of inference likelihood. (3/10) 110
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026Humans reason in non-deterministic settings all the time. "There was an accident on the highway → traffic was worse than usual" is likely, but not certain. We built ProbCOPA, a dataset of 210 such inferences each rated by 25–30 people, to study this. (2/10) 130
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026 🚨New Paper!🚨 How do reasoning LLMs handle inferences that have no deterministic answer? We find that they diverge from humans in some significant ways, and fail to reflect human uncertainty… 🧵(1/10) 35820
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026More strikingly: for almost every item in our dataset, humans showed more response variation than models. Increasing temperature is not enough; models devolve into outputting random tokens before reaching human-level variation. (6/10) 000
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026Models do okay at the extremes — for inferences that humans deem very likely or very unlikely, model responses cluster in similar regions. But where humans are more uncertain, this similarity breaks down. (5/10) 100
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026We tested 8 contemporary reasoning LLMs (GPT-5, Gemini-3, Kimi-K2-Thinking, Claude Sonnet-4.5, and more). The pattern is striking: unlike humans, models almost never return medium likelihood scores. (4/10) 100
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026Human ratings are graded and varied — people do not always commit to hard judgments, and show variation on their exact judgments of inference likelihood. (3/10) 100
Gaurav Kamath @grvkamath.bsky.social · 04/03/2026Humans reason in non-deterministic settings all the time. "There was an accident on the highway → traffic was worse than usual" is likely, but not certain. We built ProbCOPA, a dataset of 210 such inferences each rated by 25–30 people, to study this. (2/10) 100
Gaurav Kamath @grvkamath.bsky.social · 11/02/2026Super cool interpretability work from @bennokrojer.bsky.social , that I think is also relevant to anyone interested in how word meanings are represented in LLMs! 010
Reposted by Gaurav KamathTom McCoy @rtommccoy.bsky.social · 15/08/2025🤖 🧠 NEW PAPER ON COGSCI & AI 🧠 🤖 Recent neural networks capture properties long thought to require symbols: compositionality, productivity, rapid learning So what role should symbols play in theories of the mind? For our answer...read on! Paper: arxiv.org/abs/2508.05776 1/n 810117
Reposted by Gaurav KamathProceedings of the National Academy of Sciences @pnas.org · 11/08/2025Using congressional speeches as a corpus, researchers quantify how younger and older adults adopt new meanings for words as language changes. Older people may be a bit slower to change, but can show considerable linguistic flexibility. In PNAS: www.pnas.org/doi/10.1073/... 041
Reposted by Gaurav KamathPhilip Ball @philipcball.bsky.social · 30/07/2025My latest column for @thenewworldmag.bsky.social looks at the question of how new meanings for words spread in the population. www.thenewworld.co.uk/philip-ball-...thenewworld.co.ukWhy we need to be more chill about language changeIt appears that our vocabulary is entrained with the Zeitgeist, whether we like it or not 4195
Gaurav Kamath @grvkamath.bsky.social · 30/07/2025What's most likely is that this IS a factor for a portion of our more recent data, but not enough to affect the main finding here (across a range of words and decades). Tyvm for the interest in this!! Cool article that's relevant: www.newyorker.com/magazine/200... 010
Gaurav Kamath @grvkamath.bsky.social · 30/07/2025Very valid q! It's likely a confound for some of the more recent data, but not most. (i) lots of the "speeches" are in fact shorter replies and remarks; (ii) the professionalization of speech-writing evolved over the 20th century, but we see no change in speakers' adoption behavior over time. 120
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025Ultimately, we hope the insights from this work spur more work that uses tools from NLP to answer questions about human language. Massive thanks to co-auths: Michelle Yang, @sivareddyg.bsky.social, @msonderegger.bsky.social and @dallascard.bsky.social! Paper: bit.ly/4fcWfma. (12/12)bit.lyPNASProceedings of the National Academy of Sciences (PNAS), a peer reviewed journal of the National Academy of Sciences (NAS) - an authoritative source of high-impact, original research that broadly spans... 010
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025Limitations: Congressional speech is time-annotated linguistic data, from thousands of speakers whose ages are known, across over a century—rare, required properties for this study. But Congress was and is not socially representative. Plus: what about other languages and societies? (11/12) 110
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025…while at a methodological level, they suggest that sociolinguists should avoid relying too much on apparent time differences—i.e. using older speakers as a window into the past—to identify ongoing semantic shifts. (10/12) 120
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025Our findings have both conceptual and methodological implications. At a conceptual level, they suggest that the social dynamics of word meaning change are generation-agnostic, and that speakers are capable of adapting their lexicon well into adulthood (unlike, e.g., their phonology)... (9/12) 120
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025These findings extend to the level of the individual: members of Congress that gave speeches over a long enough period of time showed significant changes in how they used some of our target words, mimicking population-level trends in word meaning change. (8/12) 130
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025Overall, we find that age has very little effect—older speakers lag slightly behind younger ones, but match their word usage within just a few years; in some cases, they even lead change. Semantic change appears driven almost purely by time, with only minor inter-generational differences. (7/12) 120
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025Finally, we use Generalized Additive Mixed-effect Models (GAMMs) to model the likelihood of a word being used in a specific sense, given the year of its use and a speaker’s age at the time, while accounting for other inter-speaker variation. (6/12) 130
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025We identify >100 words suspected to have undergone meaning change in our corpus. We then use a Masked Language Model to induce several distinct, interpretable senses of each of these words, by clustering the MLM’s substitution predictions for the target word given different usage contexts. (5/12) 120
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025To answer these questions, we conduct the first large-scale investigation of semantic change across both time and speaker age. We look at ~7.9M speeches from the U.S. Congress from 1873-2010, to ask whether semantic changes are led only by specific generations, or if everyone joins in. (4/12) 120
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025Example: “workshop” used to solely refer to a physical place of work; now it refers to a type of conference/seminar. As this change occurred, did older speakers learn to use “workshop” in its newer sense? Or did the dominant meaning of the word change only because those generations died out? (3/12) 130
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025A major question in linguistics is how languages evolve within our lifetimes. Change could be purely down to inter-generational turnover; or it could involve people of old and new generations alike participating in ongoing changes of language use. (2/12) 120
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025Our new paper in #PNAS (bit.ly/4fcWfma) presents a surprising finding—when words change meaning, older speakers rapidly adopt the new usage; inter-generational differences are often minor. w/ Michelle Yang, @sivareddyg.bsky.social , @msonderegger.bsky.social and @dallascard.bsky.social👇(1/12) 33317
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025…while at a methodological level, they suggest that sociolinguists should avoid relying too much on "apparent time" differences—i.e. using older speakers as a window into the past—to identify ongoing semantic shifts. (10/12) 000
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025Our findings have both conceptual and methodological implications. At a conceptual level, they suggest that the social dynamics of word meaning change are generation-agnostic, and that speakers are capable of adapting their lexicon well into adulthood (unlike, e.g., their phonology)... (9/12) 100
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025These findings extend to the level of the individual: members of Congress that gave speeches over a long enough period of time showed significant changes in how they used some of our target words, mimicking population-level trends in word meaning change. (8/12) 120
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025Overall, we find that age has very little effect—older speakers lag slightly behind younger ones, but match their word usage within just a few years; in some cases, they even lead change. Semantic change appears driven almost purely by time, with only minor inter-generational differences. (7/12) 100
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025Finally, we use Generalized Additive Mixed-effect Models (GAMMs) to model the likelihood of a word being used in a specific sense, given the year of its use and a speaker’s age at the time, while accounting for other inter-speaker variation. (6/12) 100
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025We identify >100 words suspected to have undergone meaning change in our corpus. We then use a Masked Language Model to induce several distinct, interpretable senses of each of these words, by clustering the MLM’s substitution predictions for the target word given different usage contexts. (5/12) 100
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025To answer these questions, we conduct the first large-scale investigation of semantic change across both time and speaker age. We look at ~7.9M speeches from the U.S. Congress from 1873-2010, to ask whether semantic changes are led only by specific generations, or if everyone joins in. (4/12) 100
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025Example: “workshop” used to solely refer to a physical place of work; now it also refers to a type of seminar. As this change occurred, did older speakers learn to use “workshop” in its newer sense? Or did the dominant meaning of the word change only because those generations died out? (3/12) 100
Gaurav Kamath @grvkamath.bsky.social · 29/07/2025A major question in linguistics is how languages evolve within our lifetimes. Change could be purely down to inter-generational turnover; or it could involve people of old and new generations alike participating in ongoing changes of language use. (2/12) 100