Sign in

Raphaël Millière

@raphaelmilliere.com
7.2K followers 1.5K following 149 posts

Philosopher of Artificial Intelligence & Cognitive Science raphaelmilliere.com

PostsRepliesMedia
Reposted by Raphaël Millière
John Morrison @johnmorrison.bsky.social · 05/08/2026
with @raphaelmilliere.com
0112
Reposted by Raphaël Millière
John Morrison @johnmorrison.bsky.social · 21/07/2026
This should be fun! I'll be talking about what it means for a neural network to have a world model (joint work with @raphaelmilliere.com) and presenting new evidence that the standard example, Othello-GPT, does not have one (joint work with @shauli.bsky.social among others).
1183
Raphaël Millière @raphaelmilliere.com · 04/07/2026
I'm broadly interested in the cognitive interpretability of DNNs, see variablescope.org for a good example from last ICML. Some things I'll be discussing at this ICML: Main conf: arxiv.org/abs/2602.20159 Two workshop talks: sites.google.com/view/philmli...
variablescope.org
Variable Scope - ICML 2025
Companion website to ICML 2025 paper on how Transformers learn variable binding
050
Raphaël Millière @raphaelmilliere.com · 04/07/2026
I’ll be at #ICML2026 in Seoul all week. If you’re interested in the intersection of cognitive science, mechanistic interpretability, and philosophy -- let’s chat! DMs are open. Check my website for more on what I work on (raphaelmilliere.com), or see below.
1121
Reposted by Raphaël Millière
Jennifer Hu @jennhu.bsky.social · 01/07/2026
Sadly won't be at ACL in person, but check out our presentations below! 🌟 I'm giving a (remote) keynote at SCiL on 7/4! 🌟We also have a poster on probability x grammaticality, and a talk on pragmatics x Theory of Mind! Our lab is actively recruiting, so please reach out! Details at glintlab.org ✨
1214
Raphaël Millière @raphaelmilliere.com · 02/07/2026
My commentary on @kmahowald.bsky.social & @futrell.bsky.social's excellent BBS paper on the relevance and significance of LLMs for linguistics is now out, along with many other peer commentaries: www.cambridge.org/core/journal... (preprint to follow)
0151
Raphaël Millière @raphaelmilliere.com · 16/06/2026
Thanks for reading. Rosa's paper is fantastic!
020
Raphaël Millière @raphaelmilliere.com · 11/06/2026
Congrats!
010
Raphaël Millière @raphaelmilliere.com · 11/06/2026
Now published in open access! Your one-stop shop for the philosophy of language models. It's the spiritual descendant of our two-part preprint from 2024, fully updated. This should be particularly useful for anyone looking for an entry point into this rapidly growing field.
compass.onlinelibrary.wiley.com
The Philosophy of Language Models
The success of large language models (LLMs) across many domains of AI research has generated intense debate. Some attribute their impressive performance on complex tasks to human-like linguistic and ...
211032
Raphaël Millière @raphaelmilliere.com · 04/05/2026
Looking foward to this!
030
Reposted by Raphaël Millière
Kyle Mahowald @kmahowald.bsky.social · 17/04/2026
Richard @futrell.bsky.social and I have posted our response to the commentaries on our BBS target article "How Linguistics Learned to Stop Worrying and Love the Language Models." The response is: "You Can't Fight in Here! This is BBS!" arxiv.org/abs/2604.09501
arxiv.org
You Can't Fight in Here! This is BBS!
Norm, the formal theoretical linguist, and Claudette, the computational language scientist, have a lovely time discussing whether modern language models can inform important questions in the language ...
13513
Raphaël Millière @raphaelmilliere.com · 15/04/2026
Trustworthiness has become a big topic in AI ethics/governance (for examples of how the term is used see www.nature.com/articles/s41..., it also features prominently in the EU AI Act), but I agree that it's often a departure from the colloquial meaning of trust/trustworthy!
nature.com
120
Raphaël Millière @raphaelmilliere.com · 14/04/2026
Looking forward to speaking at this ICML workshop on ML & Philosophy! Check out the full lineup and CFP below (deadline: May 11th). Despite the title, the CFP is open to work in many areas of the philosophy of AI, not just AI ethics. sites.google.com/view/philmli...
1132
Raphaël Millière @raphaelmilliere.com · 24/02/2026
Conceptual Buccaneering
130
Raphaël Millière @raphaelmilliere.com · 20/02/2026
Thanks! We have a fully updated review paper on this forthcoming in Philosophy Compass (should be preprinted soon), and a book in preparation that gets into more details and could be used as a textbook for that kind of course.
160
Reposted by Raphaël Millière
Mark Riedl @markriedl.bsky.social · 04/02/2026
New work by my former PhD student, Boyang Li His team produced 500 stories of less than 100 words. LLMs were basically chance-level at answering binary questions about the stories arxiv.org/abs/2601.12410
arxiv.org
Are LLMs Smarter Than Chimpanzees? An Evaluation on Perspective Taking and Knowledge State Estimation
Cognitive anthropology suggests that the distinction of human intelligence lies in the ability to infer other individuals' knowledge states and understand their intentions. In comparison, our closest ...
612114
Raphaël Millière @raphaelmilliere.com · 02/02/2026
Very glad to see this out! Great paper
120
Reposted by Raphaël Millière
Tomer Ullman @tomerullman.bsky.social · 27/01/2026
now accepted at ICLR! 🐺🥳🐺 arxiv.org/abs/2506.20666
0409
Raphaël Millière @raphaelmilliere.com · 20/01/2026
The main takeaway for me is that structural information in language is far more constraining than intuition suggests. That's very interesting (and I agree that parrot metaphors are misleading) but it seems like a claim about language more than intelligence. 3/3
1210
Raphaël Millière @raphaelmilliere.com · 20/01/2026
The LLM has to do something like schema-conditioned infilling: produce a high-probability member of the equivalence class consistent with those constraints. So I'm not sure how unexpected the results are? That's roughly what I'd expect from matching to structurally similar training passages. 2/3
190
Raphaël Millière @raphaelmilliere.com · 20/01/2026
Very cool idea! Some quick thoughts. It looks like the corruption preserves a lot of information (function words, morphology, word order, punctuation, numbers, register) which would strongly constrains the posterior over plausible discourse frames as it were. 1/3
1171
Reposted by Raphaël Millière
Department of Statistics @oxfordstatistics.bsky.social · 26/08/2025
With @jesusoxford.bsky.social we are looking for a Professor of Statistics. Become part of a historic institution and a community focused on academic excellence, innovative thinking, and significant practical application. About the role: tinyurl.com/b8uy6mr5 Deadline: 15 September
Jesus College
143
Raphaël Millière @raphaelmilliere.com · 21/08/2025
I'm happy to share that I'll be joining Oxford this fall as an associate professor, as well as a fellow of @jesusoxford.bsky.social and affiliate with the Institute for Ethics in AI. I'll also begin my AI2050 Fellowship from @schmidtsciences.bsky.social there. Looking forward to getting started!
3480
Raphaël Millière @raphaelmilliere.com · 14/08/2025
Thanks Ali! We'll (hopefully soon) publish a Philosophy Compass review and for a longer read a Cambridge Elements book that are the spiritual successors to these preprints and up-to-date w/ both technical and philosophical recent developments
120
Raphaël Millière @raphaelmilliere.com · 11/08/2025
There's a lot more in the full paper – here's the open access link: sciencedirect.com/science/arti... Special thanks to @taylorwwebb.bsky.social and @melaniemitchell.bsky.social for comments on previous versions of the paper! 9/9
sciencedirect.com
LLMs as models for analogical reasoning
Analogical reasoning — the capacity to identify and map structural relationships between different domains — is fundamental to human cognition and lea…
1240
Raphaël Millière @raphaelmilliere.com · 11/08/2025
This opens intesting avenues for future work. By using causal intervention methods with open-weights models, we can start to reverse-engineer these emergent analogical abilities and compare the discovered mecanisms to computational models of analogical reasoning. 8/9
1190
Raphaël Millière @raphaelmilliere.com · 11/08/2025
But models also showed different sensitivities than humans. For example, top LLMs were more affected by permuting the order of examples and were more distracted by irrelevant semantic information, hinting at different underlying mechanisms. 7/9
1221
Raphaël Millière @raphaelmilliere.com · 11/08/2025
We found that the best-performing LLMs match human performance across many of our challenging new tasks. This provides evidence that sophisticated analogical reasoning can emerge from domain-general learning, where existing computational models fall short. 6/9
1172
Raphaël Millière @raphaelmilliere.com · 11/08/2025
In our second study, we highlighted the role of semantic content. Here, the task required identifying specific properties of concepts (e.g., "Is it a mammal?", "How many legs does it have?") and mapping them to features of the symbol strings. 5/9
1150
Raphaël Millière @raphaelmilliere.com · 11/08/2025
In our first study, we tested whether LLMs could map semantic relationships between concepts to symbolic patterns. We included controls such as permuting the order of examples or adding semantic distractors to test for robustness and content effects (see full list below). 4/9
1160
Raphaël Millière @raphaelmilliere.com · 11/08/2025
We tested humans & LLMs on analogical reasoning tasks that involve flexible re-representation. We strived to apply best practices from cognitive science – designing novel tasks to avoid data contamination, including careful controls, and doing proper statistical analysis. 3/9
1140
Raphaël Millière @raphaelmilliere.com · 11/08/2025
We focus on an important feature of analogical reasoning often called "re-representation" – the ability to dynamically select which features of analogs matter to make sense of the analogy (e.g. if one analog is "horse", which properties of horses does the analogy rely on?). 2/9
1190
Raphaël Millière @raphaelmilliere.com · 11/08/2025
Can LLMs reason by analogy like humans? We investigate this question in a new paper published in the Journal of Memory and Language (link below). This was a long-running but very rewarding project. Here are a few thoughts on our methodology and main findings. 1/9
514038
Raphaël Millière @raphaelmilliere.com · 18/07/2025
I'm glad this can be useful! And I totally agree regarding QKV vectors – focusing on information movement across token positions is way more intuitive. I had to simplify things quite a bit, but hopefully the video animation is helpful too.
010
Raphaël Millière @raphaelmilliere.com · 18/07/2025
See also @melaniemitchell.bsky.social's excellent entry on Large Language Models: oecs.mit.edu/pub/zp5n8ivs/
oecs.mit.edu
Large Language Models
0112
Raphaël Millière @raphaelmilliere.com · 18/07/2025
I wrote an entry on Transformers for the Open Encyclopedia of Cognitive Science (‪@oecs-bot.bsky.social‬). I had to work with a tight word limit, but I hope it's useful as a short introduction for students and researchers who don't work on machine learning: oecs.mit.edu/pub/ppxhxe2b
oecs.mit.edu
Transformers
15412
Raphaël Millière @raphaelmilliere.com · 14/07/2025
Happy to share this updated Stanford Encyclopedia of Philosophy entry on 'Associationist Theories of Thought' with @ericman.bsky.social. Among other things, we included a new major section on reinforcement learning. Many thanks to Eric for bringing me on board! plato.stanford.edu/entries/asso...
plato.stanford.edu
Associationist Theories of Thought (Stanford Encyclopedia of Philosophy)
1398
Reposted by Raphaël Millière
Sam Gershman @gershbrain.bsky.social · 09/07/2025
The sycophantic tone of ChatGPT always sounded familiar, and then I recognized where I'd heard it before: author response letters to reviewer comments. "You're exactly right, that's a great point!" "Thank you so much for this insight!" Also how it always agrees even when it contradicts itself.
518722
Raphaël Millière @raphaelmilliere.com · 10/06/2025
The paper is available in open access. It includes a lot more, including a discussion of how social engineering attacks on humans relate to the exploitation of normative conflicts in LLMs, and some examples of "thought injection attacks" on RLMs. 13/13 link.springer.com/article/10.1...
link.springer.com
Normative conflicts and shallow AI alignment - Philosophical Studies
The progress of AI systems such as large language models (LLMs) raises increasingly pressing concerns about their safe deployment. This paper examines the value alignment problem for LLMs, arguing tha...
060
Raphaël Millière @raphaelmilliere.com · 10/06/2025
In sum: the vulnerability of LLMs to adversarial attacks partly stems from shallow alignment that fails to handle normative conflicts. New methods like @OpenAI's “deliberative alignment” seem promising on paper, but still far from fully effective on jailbreak benchmarks. 12/13
140
Raphaël Millière @raphaelmilliere.com · 10/06/2025
I'm not convinced that the solution is a “scoping” approach to capabilities that seeks to remove information from the training data or model weights; we also need to augment models with a robust capacity for normative deliberation, even for out-of-distribution conflicts. 11/13
130
Raphaël Millière @raphaelmilliere.com · 10/06/2025
This has serious implications as models become more capable in high-stakes domains. LLMs are arguably past the point where they can cause real harm. Even if the probability of success of a single attack is negligible, success becomes almost inevitable with enough attempts. 10/13
141
Raphaël Millière @raphaelmilliere.com · 10/06/2025
For example, an RLM asked to generate a hateful tirade may conclude in its reasoning trace that it should refuse; but if the prompt instructs it to assess each hateful sentence within its thinking process, it will often leak the full harmful content! (see example below) 9/13
Example of a "thought injection attack" on Deepseek R1, asking for a violent tirade against philosophers (note that the attack method also works on much more serious examples of harmful speech). This shows the reasoning trace before the actual answer.Example of a "thought injection attack" on Deepseek R1, asking for a violent tirade against philosophers (note that the attack method also works on much more serious examples of harmful speech). This shows the actual answer after the reasoning trace.
150
Raphaël Millière @raphaelmilliere.com · 10/06/2025
I call this "thought injection attack": malicious prompts can hijack the model's reasoning trace itself, tricking it into including harmful content under the pretense of deliberation. This works by instructing the model to think through the harmful content as a safety check. 8/13
140
Raphaël Millière @raphaelmilliere.com · 10/06/2025
But what about newer “reasoning” language models (RLMs) that produce an explicit reasoning trace before answering? While they show inklings of normative deliberation, they remain vulnerable to the same kind of conflict exploitation. Worse, they introduce a new attack vector. 7/13
120
Raphaël Millière @raphaelmilliere.com · 10/06/2025
By contrast, humans can engage in normative deliberation to weigh conflicting prima facie norms and arrive at an all-things-considered judgment about what they should do in a specific context. Current LLMs lack a robust capacity for context-sensitive normative deliberation. 6/13
120
Raphaël Millière @raphaelmilliere.com · 10/06/2025
Why does this work? Current methods don't teach LLMs to reason about the contextual relevance of conflicting norms – they reinforce shallow behavioral dispositions. When faced with a new conflict, LLMs follow the disposition most strongly activated by the prompt's framing. 5/13
120
Raphaël Millière @raphaelmilliere.com · 10/06/2025
Adversarial prompt can exploit these normative conflicts to “jailbreak” models into producing harmful outputs. They often frame a malicious request within a context that makes one norm (e.g., helpfulness) more salient than another (e.g., harmlessness). 4/13
120
Raphaël Millière @raphaelmilliere.com · 10/06/2025
The problem is that these norms often conflict. For example, a request for dangerous information (violating “harmlessness”) can be framed as an educational query (appealing to “helpfulness”). Many issues with LLM behavior can be framed through these normative conflicts. 3/13
Helpfulness and harmlessness conflict when assisting the user is likely to cause harm. For example, an overly helpful model might fulfill harmful requests or provide information that could be misused. Conversely, excessive caution against harm can make a model unhelpful, by limiting its outputs to evasive responses or even refusals to answer legitimate user requests. In turn, honesty and harmlessness conflict when conveying truthful or accurate information is likely to cause harm. Privileging honesty at the expense of harmlessness can lead the model to disclose information hazards, while prioritizing harmlessness at the expense of honesty can lead to deliberate misinformation or omission of important details. Finally, helpfulness and honesty can conflict in more subtle ways when conveying truthful or accurate information violates the disposition to assist the user. Privileging helpfulness at the expense of honesty may lead to sycophantic behavior, in which the model systematically agrees with the user without concern for accuracy; conversely, strict adherence to honesty can result in responses that, while factually correct, may fail to address the user's needs -- for example by being excessively detailed and lacking the right level of simplifying abstraction. These undesirable trade-offs resulting from conflicts between the norms of alignment are illustrated in the figure.
121
Raphaël Millière @raphaelmilliere.com · 10/06/2025
Current alignment methods aim to instill norms like helpfulness, honesty & harmlessness in LLMs through preference fine-tuning (like RLHF or DPO) that rewards them for producing outputs that humans prefer, steering them away from unhelpful, dishonest or harmful content. 2/13
120