Sign in

James Michaelov

@jamichaelov.bsky.social
4.3K followers 540 following 48 posts

Postdoc at Oxford. Research: language, the brain, NLP. jmichaelov.com

PostsRepliesMedia
James Michaelov @jamichaelov.bsky.social · 28/09/2026
Looking forward to #SNL2026! I’ll be presenting the work published in our recent JML paper about language model scaling and the N400. Find me (or email/message) if you want to chat about predictive coding and NLP in the study of human language comprehension! Paper link: doi.org/10.1016/j.jm...
Title: Better language models better model the N400, but not reading time

Abstract: The probability of a word in context, as captured by large language models, is predictive of both behavioral and neural measures of human language processing. Intuitively, language models that are better at next-word prediction might better model predictability effects in human language comprehension. Yet recent work suggests that language models can become too good at next-word prediction to model reading time, implying that the aspects of human comprehension indexed by reading time do not track perfectly with predictability from language statistics alone. However, it is unknown whether this decoupling is true of reading time only, or whether it is intrinsic to online measures of comprehension more generally. To address this question, we turn to another robust and well-studied measure of online processing, the N400 component of the event-related brain potential. We compare how a language model’s size, number of training tokens, and performance on natural language benchmarks correlate with its ability to predict both reading time and N400 amplitude. Based on an analysis of 4 reading time datasets and 9 N400 datasets, we replicate past results for reading time, but find that larger language models, models that are trained on more data, and models that perform better at next-word prediction and other more complex natural language tasks are better able to predict N400 amplitude. We interpret this difference between the N400 and reading time measures as potentially revealing the comparative importance of semantic prediction in the neurocognitive processes indexed by the N400.
092
Reposted by James Michaelov
Sean Trott @seantrott.bsky.social · 31/08/2026
New paper out in @openmindjournal.bsky.social! "Large Language Models as Distributional Baselines for Language Tasks". With @jamichaelov.bsky.social , @camrobjones.bsky.social , Tyler Chang, and Ben Bergen. direct.mit.edu/opmi/article...
direct.mit.edu
Large Language Models as Distributional Baselines for Language Tasks
Abstract. In order to ask questions about the mechanisms underpinning human cognition, researchers must control for properties of stimuli that could confound detected effects. In experiments involving...
1126
James Michaelov @jamichaelov.bsky.social · 01/07/2026
I’ll be in San Diego for #ACL2026 and presenting this work at @conll-conf.bsky.social #CoNLL2026 on July 3rd! Feel free to reach out if you want to chat!
0110
James Michaelov @jamichaelov.bsky.social · 11/06/2026
Here’s a summary of our main conclusions, and another link to the paper: arxiv.org/abs/2603.26539
000
James Michaelov @jamichaelov.bsky.social · 11/06/2026
We also discuss other nuances, including factors to consider in safety, socio-technical, and HCI research; approaches to mitigating the problems associated with closed-weight models; and the limits of what open weights alone can provide
100
James Michaelov @jamichaelov.bsky.social · 11/06/2026
Key conclusions: For most types of generalizable claims about model behavior and capabilities, we need open-weight models. For existence proofs and verifiable solution candidate generation, closed-weight models may be appropriate
110
James Michaelov @jamichaelov.bsky.social · 11/06/2026
We address a key question for any research involving language models: what kinds of information do we need about the models used in order to make reliable scientific inferences?
110
James Michaelov @jamichaelov.bsky.social · 11/06/2026
Seems like a good time to share our new preprint about model openness! (with @catherinearnett.bsky.social @tylerachang.bsky.social Pamela D. Rivière, Samuel M. Taylor @camrobjones.bsky.social @seantrott.bsky.social @rplevy.bsky.social Ben Bergen, and Micah Altman): arxiv.org/abs/2603.26539
1193
James Michaelov @jamichaelov.bsky.social · 27/03/2026
Had a great first day at #HSP2026 yesterday! Looking forward to presenting on the relationship between reading time, n-grams, and language model scaling at the 12.10-2pm poster session today!
040
James Michaelov @jamichaelov.bsky.social · 04/12/2025
Presenting this at the poster session this morning (11-2pm) at #5109
020
James Michaelov @jamichaelov.bsky.social · 01/12/2025
Looking forward to #NeurIPS25 this week 🏝️! I'll be presenting at Poster Session 3 (11-2 on Thursday). Feel free to reach out!
0103
James Michaelov @jamichaelov.bsky.social · 25/11/2025
I'll also be presenting this paper with @catherinearnett.bsky.social at #CogInterp!
030
James Michaelov @jamichaelov.bsky.social · 25/11/2025
Preprint: www.arxiv.org/abs/2510.24963
arxiv.org
Language Model Behavioral Phases are Consistent Across Architecture, Training Data, and Scale
We show that across architecture (Transformer vs. Mamba vs. RWKV), training dataset (OpenWebText vs. The Pile), and scale (14 million parameters to 12 billion parameters), autoregressive language mode...
160
James Michaelov @jamichaelov.bsky.social · 25/11/2025
Excited to announce that I’ll be presenting a paper at #NeurIPS this year! Reach out if you’re interested in chatting about LM training dynamics, architectural differences, shortcuts/heuristics, or anything at the CogSci/NLP/AI interface in general! #Neurips2025
2262
Reposted by James Michaelov
Catherine Arnett @catherinearnett.bsky.social · 27/07/2025
I’m in Vienna all week for @aclmeeting.bsky.social and I’ll be presenting this paper on Wednesday at 11am (Poster Session 4 in HALL X4 X5)! Reach out if you want to chat about multilingual NLP, tokenizers, and open models!
0171
James Michaelov @jamichaelov.bsky.social · 12/06/2025
See the full paper here: arxiv.org/abs/2506.06808 3/3
arxiv.org
Not quite Sherlock Holmes: Language model predictions do not reliably differentiate impossible from improbable events
Can language models reliably predict that possible events are more likely than merely improbable ones? By teasing apart possibility, typicality, and contextual relatedness, we show that despite the re...
020
James Michaelov @jamichaelov.bsky.social · 12/06/2025
In the most extreme case, LMs assign sentences such as ‘the car was given a parking ticket by the explorer’ (unlikely but possible event) a lower probability than ‘the car was given a parking ticket by the brake’ (animacy-violating event, semantically-related final word) over half of the time. 2/3
110
James Michaelov @jamichaelov.bsky.social · 12/06/2025
New paper accepted at ACL Findings! TL;DR: While language models generally predict sentences describing possible events to have a higher probability than impossible (animacy-violating) ones, this is not robust for generally unlikely events and is impacted by semantic relatedness. 1/3
1213
Reposted by James Michaelov
Catherine Arnett @catherinearnett.bsky.social · 05/06/2025
My paper with @tylerachang.bsky.social and @jamichaelov.bsky.social will appear at #ACL2025NLP! The updated preprint is available on arxiv. I look forward to chatting about bilingual models in Vienna!
182
Reposted by James Michaelov
Catherine Arnett @catherinearnett.bsky.social · 07/03/2025
✨New pre-print✨ Crosslingual transfer allows models to leverage their representations for one language to improve performance on another language. We characterize the acquisition of shared representations in order to better understand how and when crosslingual transfer happens.
2367
James Michaelov @jamichaelov.bsky.social · 08/02/2025
I’ve had success using the infini-gram API for this (though it can get overloaded with user requests at times): infini-gram.io
infini-gram.io
Home
010
James Michaelov @jamichaelov.bsky.social · 03/12/2024
I don’t think this is quite what you’re looking for, but @camrobjones.bsky.social recently ran some Turing-test-style studies and found that some people believed ELIZA to be a human (and participants were asked to give reasons for their responses)
150
Reposted by James Michaelov
James Michaelov @jamichaelov.bsky.social · 10/11/2024
With all the new people here on Bluesky, I think it’s a good time to (re-)introduce myself. I’m a postdoc at MIT carrying out research at the intersection of the cognitive science of language and AI. Here are some of the things I’ve worked on in the last year 🧵:
1242
James Michaelov @jamichaelov.bsky.social · 19/11/2024
Seems like a great initiative to have some of these location-based ones! I’d love to be added if possible!
010
James Michaelov @jamichaelov.bsky.social · 11/11/2024
Excited to be at #EMNLP #EMNLP2024 this year! Especially interested in chatting about the intersection of cognitive science/psycholinguistics and AI/NLP, training dynamics, robustness/reliability, meaning, and evaluation
1110
James Michaelov @jamichaelov.bsky.social · 11/11/2024
If there’s still space (and you accept postdocs), could I be added?
010
James Michaelov @jamichaelov.bsky.social · 11/11/2024
Thanks for creating this list - looks great! I’d love to be added if there’s still room
010
James Michaelov @jamichaelov.bsky.social · 11/11/2024
Thank you!
010
James Michaelov @jamichaelov.bsky.social · 11/11/2024
If there’s still room, is there any chance you could add me to this list?
110
James Michaelov @jamichaelov.bsky.social · 10/11/2024
Also, I’m going to be attending EMNLP next week - reach out if you want to meet/chat
040
James Michaelov @jamichaelov.bsky.social · 10/11/2024
Anyway, excited to learn and chat about about research along these lines and beyond here on Bluesky!
130
James Michaelov @jamichaelov.bsky.social · 10/11/2024
Of course, none of this work would have been possible without my amazing PhD advisor Ben Bergen, and my other great collaborators: Seana Coulson, @catherinearnett.bsky.social, Tyler Chang, Cyma Van Petten, and Megan Bardolph!
130
James Michaelov @jamichaelov.bsky.social · 10/11/2024
5: Recurrent models like RWKV and Mamba have recently emerged as viable alternatives to transformers. While they are intuitively more cognitively plausible, when used to model human language processing, how do they compare transformers? We find that they perform about the same overall:
openreview.net
Revenge of the Fallen? Recurrent Models Match Transformers at...
Transformers have generally supplanted recurrent neural networks as the dominant architecture for both natural language processing tasks and for modelling the effect of predictability on online...
140
James Michaelov @jamichaelov.bsky.social · 10/11/2024
4: Is the N400 sensitive only to the predicted probability of the stimuli encountered, or also the predicted probability of alternatives? We revisit this question with state-of-the-art NLP methods, with the results supporting the former hypothesis:
sciencedirect.com
Ignoring the alternatives: The N400 is sensitive to stimulus preactivation alone
The N400 component of the event-related brain potential is a neural signal of processing difficulty. In the language domain, it is widely believed to …
130
James Michaelov @jamichaelov.bsky.social · 10/11/2024
3: The N400, a neural index of language processing, is highly sensitive to the contextual probability of words. But to what extent can lexical prediction explain other N400 phenomena? Using GPT-3, we show that it can implicitly account for both semantic similarity and plausibility effects:
doi.org
Strong Prediction: Language Model Surprisal Explains Multiple N400 Effects
Abstract. Theoretical accounts of the N400 are divided as to whether the amplitude of the N400 response to a stimulus reflects the extent to which the stimulus was predicted, the extent to which the s...
130
James Michaelov @jamichaelov.bsky.social · 10/11/2024
2: Do multilingual language models learn that different languages can have the same grammatical structures? We use the structural priming paradigm from psycholinguistics to provide evidence that they do:
aclanthology.org
Structural Priming Demonstrates Abstract Grammatical Representations in Multilingual Language Models
James Michaelov, Catherine Arnett, Tyler Chang, Ben Bergen. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023.
140
James Michaelov @jamichaelov.bsky.social · 10/11/2024
If you’re interested in hearing more of my thoughts on this topic, check out this article at Communications of the ACM by Sandrine Ceurstemont that includes quotes from an interview with me and my co-author Ben Bergen:
cacmb4.acm.org
Bigger, Not Necessarily Better
The inverse scaling issue means larger LLMs sometimes handle things less well.
130
James Michaelov @jamichaelov.bsky.social · 10/11/2024
1. Training language models on more data generally improves their performance, but is this always the case? We show that inverse scaling can occur not just across models of different sizes, but also in individual models over the course of training:
aclanthology.org
Emergent Inabilities? Inverse Scaling Over the Course of Pretraining
James Michaelov, Ben Bergen. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023.
140
James Michaelov @jamichaelov.bsky.social · 10/11/2024
With all the new people here on Bluesky, I think it’s a good time to (re-)introduce myself. I’m a postdoc at MIT carrying out research at the intersection of the cognitive science of language and AI. Here are some of the things I’ve worked on in the last year 🧵:
1242
James Michaelov @jamichaelov.bsky.social · 02/04/2024
We instead show that the next-word prediction is sufficient to get both effects, at least qualitatively.
000
James Michaelov @jamichaelov.bsky.social · 02/04/2024
The study also has implications for psycholinguistics research. Some have suggested that the two kinds of related anomaly effect might require different mechanisms to each other and/or to other predictability effects.
100
James Michaelov @jamichaelov.bsky.social · 02/04/2024
While this may have benefits, predicting related but anomalous continuations to be more likely than unrelated anomalous continuations is a way in which language model predictions and human judgments are misaligned
110
James Michaelov @jamichaelov.bsky.social · 02/04/2024
Why do these effects occur? There are many possibilities, but one of the most likely is that despite being trained on lexical prediction only, the language models are making predictions at a semantic level.
100
James Michaelov @jamichaelov.bsky.social · 02/04/2024
But we also predict words that are RELATED to the word “full” such as “half” more than UNRELATED words like “mild”. Again, language models also show this effect:
100
James Michaelov @jamichaelov.bsky.social · 02/04/2024
We also see a similar effect where humans predict words that are related to the most likely continuation. Given a context such as “Lydia cannot eat anymore as she is so ___”, human participants predict “full” most strongly.
100
James Michaelov @jamichaelov.bsky.social · 02/04/2024
But we are also more likely to predict a word RELATED to the described event (mountain biking) like “dirt” than an UNRELATED word like “table”, even though neither makes sense in context. We see the same effect in language models (a lower surprisal indicates a stronger prediction):
100
James Michaelov @jamichaelov.bsky.social · 02/04/2024
Following a narrative such as “My friend Mike went mountain biking recently. He lost control for a moment and ran right into a tree. It’s a good thing he was wearing his ___”, human participants are most likely to predict the word “helmet”.
100
James Michaelov @jamichaelov.bsky.social · 02/04/2024
We look at one specific effect observed in the human N400 response (widely considered a neural index of prediction) called the ‘related anomaly effect’.
100
James Michaelov @jamichaelov.bsky.social · 02/04/2024
In the interest of actually posting about my research on here: We know that the predictions that language models make are similar to those that humans make as we process language, but how similar? aclanthology.org/2022.conll-1... 🧵:
100
James Michaelov @jamichaelov.bsky.social · 15/01/2024
🤖🤖🤖
010