Sign in

Andrea de Varda

@andreadevarda.bsky.social
394 followers 400 following 52 posts

Postdoc at MIT BCS, interested in language(s) in humans and LMs andrea-de-varda.github.io

PostsRepliesMedia
Reposted by Andrea de Varda
Greta Tuckute @gretatuckute.bsky.social · 17/09/2026
Thanks so much for the fun conversation @wiair.bsky.social ! A pleasure to chat about language, LLMs, and memory--covering some work with @bkhmsi.bsky.social @mschrimpf.bsky.social @michael-lepori.bsky.social @klemenkotar.bsky.social @evfedorenko.bsky.social @thomashikaru.bsky.social, among others!
1245
Reposted by Andrea de Varda
Ev Fedorenko @evfedorenko.bsky.social · 01/09/2026
Go, @andreadevarda.bsky.social! A beautiful and comprehensive study! 🔑 findng: behav. measures are ~fully reducible to simple predictors of processing effort (surprisal, word length+frequency), but for 🧠 measures, LLM embeddings carry additional predictive power, likely capturing aspects of meaning.
094
Reposted by Andrea de Varda
Yevgeni Berzak @whylikethis.bsky.social · 02/09/2026
Wonderful work by @andreadevarda.bsky.social linking behavioral and neural manifestations of language comprehension using language models. With @rplevy.bsky.social and @evfedorenko.bsky.social
162
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
👀!=🧠 A unified theory needs both kinds of data with a clear understanding of which levels of representation each measure reflects. (10/10) Pre-print 🔗 tinyurl.com/mr3cre4b
tinyurl.com
000
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
A similar asymmetry emerges within the effort measures. Eye movements are driven mostly by length and frequency (context-independent). Brain responses are driven mostly by surprisal (context-dependent). (9/10)
100
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
👀 Behavior: effort dominates. SFL (three numbers) performs better than 1600-dim embeddings, esp. in eye-tracking. Adding EMB to SFL gives only small gains. 🧠 Brain: the opposite. Embeddings predict fMRI/N400 responses far better than effort. (8/10)
100
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
We test this on 13 datasets: 👀 4 eye-tracking, 3 self-paced reading, 1 Maze, 🧠 1 ERP (N400), and 4 fMRI. All responses are averaged across participants and brain responses are averaged across voxels/electrodes so brain and behavior are comparable (1-dimensional). (7/10)
100
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
Our hypothesis: behavior shows a bottleneck. Rich representations get compressed into effort dimensions (SFL) before they can influence processing times. Brain responses have more direct access to the high-dimensional representations. (6/10)
100
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
We use LMs to operationalize effort and meaning in one framework. From the same GPT/GPT-2 models we get (i) surprisal (+freq and len; SFL) for effort, and (ii) contextual embeddings (EMB), high-dimensional vectors that encode form and meaning. (5/10)
100
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
But language is about transmitting meaning, and effort abstracts away from much of it. "The chef cooked the meal" and "The wolf caught the deer" have ~identical effort profiles but mean very different things. And brain studies show sensitivity to meaning besides effort. (4/10)
100
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
One influential view is that responses to language are driven by processing effort, mostly captured by three word-level predictors: surprisal, frequency, and length (SFL). These are the "Big Three" in reading research and they also predict brain responses. (3/10)
100
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
Psycholinguists study behavior (👀 eye movements, reading times). Neuroscientists study brain activity (🧠 fMRI, ERPs, etc.). Both make inferences about the human language system, but they are rarely studied together. Do they reflect the same information? (2/10)
100
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
New preprint! 🧠👀🤖 Behavioral and brain responses to language reflect different levels of linguistic representation w/ @whylikethis.bsky.social , @evfedorenko.bsky.social , and @rplevy.bsky.social (1/10)
1356
Reposted by Andrea de Varda
Chiara Saponaro @chiarasaponaro.bsky.social · 20/08/2026
New paper out! 🎉 Can preverbal logical inferences scaffold the early acquisition of logical words? We asked this question with Mahham Fayyaz, Grace Pavalko and Nicolò Cesana-Arlotti, focusing on disjunction: escholarship.org/uc/item/3pf7....
escholarship.org
1124
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
So what have we learned? Human effort allocation is largely predictable with large-scale pre-training on text + correctness pressure under computational constraints. This finding supports resource-rational accounts of behavior.
000
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
Lastly, there is a fundamental tension in B&al’s pieces: the correlation cannot be both trivially guaranteed by training and fragile and easily dissociated.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
However, one refined version of their interpretation is not too different from ours: systems that learn to reason from human data at first, and then optimize correctness under computational limits converge with human effort. This is what we argue for in the original paper.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
Imitation also predicts the wrong direction of the inter-model difference. RL moves models away from imitating the pre-training corpus, so alignment should weaken in LRMs. Instead, it strengthens. Optimizing for correctness with no human data in the reward increases the similarity to humans.
110
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
B&al conjecture that models were trained to produce more tokens on problems humans find hard. But training data contain no reaction times, RLVR rewards correctness only, and 4 of our 7 datasets were published after model training. Difficulty must be computed from problem structure, not memorized.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
The relevant manipulations exist in our data and all point the same way. Varying problem structure (digits, operands, carry, operation-type) affects tokens and RTs similarly. Varying training (standard LLMs vs. reasoning models) changes the correlations between tokens and RTs.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
Our claim is (c). B&al’s manipulation acts on the token budget and they find that accuracy survives on easy tasks. Under (c), that is exactly what should happen, so their experiment is not informative.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
B&al ask for causal evidence, but causal between what and what? There are three possibilities. (a) Tokens cause model accuracy (b) Tokens cause human RTs, which nobody would hold (c) A common cause: the computational demands of the problem drive both token counts and RTs.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
B&al also claim more time would help humans on 100% of the tasks. At ceiling, it cannot, by definition. Beyond that, this is false: for some tasks, thinking makes humans and models worse (Liu et al. 2024), and time pressure mainly hurts hard problems.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
Above low effort, accuracy is indeed near flat. But e.g., a flat response above the effective dose is not evidence that a drug doesn’t work. It is what a causally effective resource looks like when the demand is small.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
As a side note, B&al claim that they found “that the LRM gpt-oss-120b performed at ceiling without outputting any tokens in most tasks”, but if you look at their fig 1, this is only true of 2 out of 6 tasks.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
Similarly, 60% of our participants obtained 100% accuracy in our dataset, but this does not make their RTs less indicative of effort. In fact, good RT studies try to keep accuracy at ceiling and analyze correct trials only, so that time measures are not contaminated by different processes.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
E.g., in our own human data: the problems 3+5 and 76-84 both have 100% accuracy, but mean RTs are 1,421 vs. 9,276 ms, respectively (6.5x increase). By the logic of B&al, this difference would not be meaningful because it is not linked to accuracy.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
This is the core logic of reaction time (RT) research. Humans solve simple math problems at ceiling no matter how much time you give them but their RTs still reflect problem difficulty, and that variance is used to build theories of arithmetic reasoning.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
The first comes down to a misconception about what a ceiling is. Their argument: gpt-oss-120b solves most of our tasks at ceiling with zero reasoning tokens, so token counts tell us nothing about effort. But when a task is easy, accuracy saturates; cost measures do not.
100
Andrea de Varda @andreadevarda.bsky.social · 07/08/2026
Bowers and colleagues (B&al) have a new response to our paper on the cost of thinking in reasoning models and humans. Here we address the core disagreements.🧵
142
Reposted by Andrea de Varda
Cognitive Science Society @cogscisociety.bsky.social · 25/07/2026
First up in the Glushko Dissertation Prize Symposium: Andrea de Varda @andreadevarda.bsky.social examines what multilingual neural language models can reveal about language and cognition. #CogSci2026
1302
Reposted by Andrea de Varda
Damián Blasi @damianblasi.bsky.social · 23/07/2026
How many languages have existed over the Holocene—and what does that reveal about the design space of languages and cultures? Now out in @science.org www.science.org/eprint/QDZNY....
science.org
The rise and fall of language diversity through the Holocene
Characterizing the factors that have shaped linguistic diversity is fundamental for understanding human history, culture, and cognition. In this study, we combined statistical and social computational...
213154
Reposted by Andrea de Varda
micha heilbron @mheilbron.bsky.social · 16/07/2026
What makes some stimuli more memorable than others? In a new paper w/ @davogelsang.bsky.social, we show that the magnitude of a stimulus's ANN representation predicts both image and word memorability Stimuli that activate more features, more strongly, leave a stronger memory trace Out now in JML⬇️
34411
Reposted by Andrea de Varda
MilaNLP Lab @milanlp.bsky.social · 13/07/2026
🧠🤖 It was a pleasure to host @andreadevarda.bsky.social for his talk, "Large Language Models as Models of Human Language(s) and Higher-Level Cognition." A truly inspiring talk! #NLProc
093
Reposted by Andrea de Varda
Ev Fedorenko @evfedorenko.bsky.social · 30/06/2026
I am so excited about this finding from @pengrui-han.bsky.social and @andreadevarda.bsky.social, also with Jacob Andreas! Perhaps modularity is inevitable in intelligent systems, biological or in silico. :)
0336
Reposted by Andrea de Varda
pengrui-han.bsky.social @pengrui-han.bsky.social · 30/06/2026
The human brain is strikingly modular: distinct networks for language, formal reasoning, social reasoning, physical reasoning. Is this fundamental to intelligent systems, or an accident of evolution? In our new preprint, we find the same modular organization emerges in LLMs.
129320
Andrea de Varda @andreadevarda.bsky.social · 01/07/2026
Like the human brain, LLMs use separate sets of units for language, formal reasoning, social reasoning, and intuitive physical reasoning. A modular organization of cognition may be a fundamental principle of intelligence!
0133
Reposted by Andrea de Varda
Anna (Anya) Ivanova @neuranna.bsky.social · 12/06/2026
Thanks to the Weber School for inviting me to give a #TEDx talk! I discuss how much people vary in their inner thought — from thinking mainly in words to thinking mostly abstractly — and the implications it has for understanding AI cognition. youtu.be/WAm0XQIRBMw
youtu.be
Do We Think In Words? Does AI? | Anna Ivanova | TEDxWeber School Youth
YouTube video by TEDx Talks
1245
Reposted by Andrea de Varda
Tom McCoy @rtommccoy.bsky.social · 22/05/2026
🤖🧠NEW PAPER🧠🤖 Children & neural networks can learn syntax from linear strings of words. How do they do it? Our hypothesis: Word co-occurrence statistics provide cues to syntax! (I.e., a new type of bootstrapping to consider!) Paper: arxiv.org/abs/2605.20529 1/n
Paper overview.
Title: "Collocational bootstrapping: A hypothesis about the learning of subject-verb agreement in humans and neural networks"
Authors: Claire Hobbs and Tom McCoy
Method: We trained many neural nets, varying how predictable a subject is given its verb. We tested them on subject-verb agreement
Findings: With the right level of predictability, neural networks robustly generalize. The predictability of child-directed language is near the neural net optimum.
Conclusion: Statistical regularities in word co-occurrence can support the learning of abstract syntactic rules
The text is accompanied by a graph showing neural-network accuracy as a function of the level of variability; the accuracy peaks at an in-between level of variability
2395
Reposted by Andrea de Varda
Adele Goldberg @adelegoldberg.bsky.social · 27/02/2026
Idan Blank (UCLA, psych) makes the complex intuitive if you want to learn how LLMs work, watch👇 newly posted to YouTube (no ads) www.youtube.com/watch?v=cGMn...
youtu.be
How Transformers Work: A Detailed, Conceptual Explanation (No Coding / Math)
YouTube video by IbanDlank
06815
Reposted by Andrea de Varda
Chiara Saponaro @chiarasaponaro.bsky.social · 17/03/2026
Can we process meaning unconsciously? Our new study suggests: not really… unless language has a way to express it!🧵 New paper out with Andrea Nadalini, Daniel Casasanto, @davidecrepaldi.bsky.social and Roberto Bottini
1131
Reposted by Andrea de Varda
Catherine Arnett @catherinearnett.bsky.social · 09/03/2026
@tylerachang.bsky.social and I will be presenting the Goldfish as an oral at #LREC2026 in Mallorca! 🌴
1204
Reposted by Andrea de Varda
Badr AlKhamissi @bkhmsi.bsky.social · 27/01/2026
Happy to share that our paper “Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like Specialization” (aka MiCRo) has been accepted to #ICLR2026!! 🎉 See you in Rio 🇧🇷 🏝️
072
Reposted by Andrea de Varda
CIMeC_UniTrento @cimecunitrento.bsky.social · 27/01/2026
Bridge AI and linguistics with the Computational and Theoretical Modelling of Language and Cognition (CLC) track at @cimecunitrento.bsky.social! Apply to our MSc in Cognitive Science First-call deadline for non-EU applicants: March 4, 2026. ℹ️ corsi.unitn.it/en/cognitive-science #cimec_unitrento #AI
032
Reposted by Andrea de Varda
Anna (Anya) Ivanova @neuranna.bsky.social · 11/12/2025
The last chapter of my PhD (expanded) is finally out as a preprint! “Semantic reasoning takes place largely outside the language network” 🧠🧐 www.biorxiv.org/content/10.6... What is semantic reasoning? Read on! 🧵👇
biorxiv.org
Semantic reasoning takes place largely outside the language network
The brain's language network is often implicated in the representation and manipulation of abstract semantic knowledge. However, this view is inconsistent with a large body of evidence suggesting that...
29126
Andrea de Varda @andreadevarda.bsky.social · 10/12/2025
In collaboration with @tomlamarra.bsky.social Andrea Amelio Ravelli @chiarasaponaro.bsky.social @beatricegiustolisi.bsky.social @mariannabolog.bsky.social
020
Andrea de Varda @andreadevarda.bsky.social · 10/12/2025
Some words sound like what they mean. In IconicITA we show that the (psycho)linguistic factors that modulate which words are most iconic are similar between English and Italian. Lots more details in the paper!
151
Andrea de Varda @andreadevarda.bsky.social · 10/12/2025
Great work led by Daria & Greta showing that diverse agreement types draw on shared units (even across languages)!
093
Reposted by Andrea de Varda
Colton Casto @coltoncasto.bsky.social · 26/11/2025
What does it mean to understand language? We argue that the brain’s core language system is limited, and that *deeply* understanding language requires EXPORTING info to other brain regions. w/ @neuranna.bsky.social @evfedorenko.bsky.social @nancykanwisher.bsky.social arxiv.org/abs/2511.19757 1/n🧵👇
arxiv.org
What does it mean to understand language?
Language understanding entails not just extracting the surface-level meaning of the linguistic input, but constructing rich mental models of the situation it describes. Here we propose that because pr...
28233
Andrea de Varda @andreadevarda.bsky.social · 21/11/2025
I'd love to watch this, is there a recording?
100