Sign in

Andrea de Varda

@andreadevarda.bsky.social
394 followers 401 following 52 posts

Postdoc at MIT BCS, interested in language(s) in humans and LMs andrea-de-varda.github.io

PostsRepliesMedia
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
👀 Behavior: effort dominates. SFL (three numbers) performs better than 1600-dim embeddings, esp. in eye-tracking. Adding EMB to SFL gives only small gains. 🧠 Brain: the opposite. Embeddings predict fMRI/N400 responses far better than effort. (8/10)
100
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
We test this on 13 datasets: 👀 4 eye-tracking, 3 self-paced reading, 1 Maze, 🧠 1 ERP (N400), and 4 fMRI. All responses are averaged across participants and brain responses are averaged across voxels/electrodes so brain and behavior are comparable (1-dimensional). (7/10)
100
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
Our hypothesis: behavior shows a bottleneck. Rich representations get compressed into effort dimensions (SFL) before they can influence processing times. Brain responses have more direct access to the high-dimensional representations. (6/10)
100
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
We use LMs to operationalize effort and meaning in one framework. From the same GPT/GPT-2 models we get (i) surprisal (+freq and len; SFL) for effort, and (ii) contextual embeddings (EMB), high-dimensional vectors that encode form and meaning. (5/10)
100
Andrea de Varda @andreadevarda.bsky.social · 01/09/2026
New preprint! 🧠👀🤖 Behavioral and brain responses to language reflect different levels of linguistic representation w/ @whylikethis.bsky.social , @evfedorenko.bsky.social , and @rplevy.bsky.social (1/10)
1356
Andrea de Varda @andreadevarda.bsky.social · 19/11/2025
Token count also captures differences across tasks. Avg. token count predicts avg. RT across domains (r = 0.97, left), and even item-level RTs across all tasks (r = 0.92 (!!), right). (5/6)
100
Andrea de Varda @andreadevarda.bsky.social · 19/11/2025
We found that the number of reasoning tokens generated by the model reliably correlates with human RTs within each task (mean r = 0.57, all ps < .001). (4/6)
110
Andrea de Varda @andreadevarda.bsky.social · 19/11/2025
Large reasoning models can solve many reasoning problems, but do their computations reflect how humans think? We compared human RTs to DeepSeek-R1’s CoT length across seven tasks: arithmetic (numeric & verbal), logic (syllogisms & ALE), relational reasoning, intuitive reasoning, and ARC (3/6)
100
Andrea de Varda @andreadevarda.bsky.social · 04/02/2025
Encoding models trained on existing fMRI datasets successfully predicted responses in new languages, generalizing across stimuli types and modalities (11/)
100
Andrea de Varda @andreadevarda.bsky.social · 04/02/2025
In the “across” condition, performance improves for models with stronger cross-lingual semantic alignment (where translations cluster together in the embedding space) (9/)
100
Andrea de Varda @andreadevarda.bsky.social · 04/02/2025
But what kind of model properties influence LM-to-brain alignment across languages? In the “within” condition, encoding performance is highest for models with good next-word prediction abilities (8/)
100
Andrea de Varda @andreadevarda.bsky.social · 04/02/2025
We also replicated in a cross-lingual setting the finding that the best fit to brain responses is obtained in intermediate-to-deep layers (for each subplot pair, the left one is “within”, the right one “across”) (7/)
100
Andrea de Varda @andreadevarda.bsky.social · 04/02/2025
We evaluated 20 multilingual LMs with different architectures and training objectives, and all of them were able to predict brain responses in the various languages (“within”) and critically, generalized zero-shot to unseen languages (“across”) (6/)
100
Andrea de Varda @andreadevarda.bsky.social · 04/02/2025
Critically, we fit two kinds of encoding models: 1️⃣ “within” encoding models, training and testing on data from a single language with cross-validation 2️⃣ “across” encoding models, training in N-1 languages and testing in the left-out language (5/)
100
Andrea de Varda @andreadevarda.bsky.social · 04/02/2025
In Study I, we: 1️⃣ Present participants with auditory passages and record their brain responses in the language network 2️⃣ Extract contextualized word embeddings from multilingual LMs 3️⃣ Fit encoding models predicting brain activity from the embeddings (4/)
100