Sign in

Sean Trott

@seantrott.bsky.social
97 followers 47 following 75 posts
PostsRepliesMedia
Reposted by Sean Trott
Benjamin Riley @benjaminjriley.bsky.social · 05/10/2026
I'm again grateful to @theverge.com for publishing my newest long-form essay on the limits of the computational model of the mind, the benefits of understanding both biological and cultural evolution to explain human thinking, and the danger of cognitive hot dogs in education.
theverge.com
Our minds aren’t equipped to handle AI
AI is junk food for the mind: easy, tempting and ultimately very bad for you.
719648
Sean Trott @seantrott.bsky.social · 01/10/2026
And I'm entirely sympathetic to the argument that we shouldn't call that distributed system an LLM ("compound system" is sometimes used; or, nowadays, "AI agent"), and that it's still worthwhile understanding what the LLM part of it can and can't do. But again, that's not a deflationary claim.
010
Sean Trott @seantrott.bsky.social · 01/10/2026
Understanding this is important for thinking about what the current generation of AI can do (also very relevant to discussions around safety). These are systems explicitly trained (through RLVR/CoT) to decompose tasks into sub-tasks, and they use external tools to do help accomplish them.
proceedings.neurips.cc
Toolformer: Language Models Can Teach Themselves to Use Tools
110
Sean Trott @seantrott.bsky.social · 01/10/2026
Right. Maybe it's because I'm already partial to the extended mind thesis, but the idea that LLMs can be hooked up to external tools, output API calls, then use the result of those API calls in their subsequent generations seems to me quite impressive and interesting.
proceedings.neurips.cc
Toolformer: Language Models Can Teach Themselves to Use Tools
110
Sean Trott @seantrott.bsky.social · 01/10/2026
Though I am sympathetic to the point about closed models. I think it's still useful scientifically to know which parts of a system are responsible for a given behavior insofar as we want to make generalizable inferences about that system, or parts of that system.
010
Sean Trott @seantrott.bsky.social · 01/10/2026
Yeah, to me, it is entirely reasonable to speak of the LLM + calculator (and other tools) as a kind of "coupled system". It does not feel deflationary to me to point out that LLMs can be trained to output API calls to external tools. That makes them more capable!
210
Reposted by Sean Trott
Mike Frank @mcxfrank.bsky.social · 25/09/2026
We've just released a new version of childes-db, my lab's interface to the CHILDES database of child language transcripts. It lets you work with CHILDES data from R through a versioned, reproducible API. A few updates 🧵 childes-db.stanford.edu/
Screenshot of the childes-db website: 'A flexible and reproducible interface to CHILDES.' childes-db 2026.1 contains 56,579 transcripts from 9,151 children across 437 corpora, with panels for an R API tutorial and interactive visualizations of mean length of utterance by child age.
12913
Sean Trott @seantrott.bsky.social · 25/09/2026
This is all great advice!
011
Reposted by Sean Trott
Kensy Cooperrider @kensycoop.bsky.social · 24/09/2026
It's "advice to grad students" season! Here's a post I wrote several geological epochs ago, in 2019. My advice gets harder and harder to follow every year (even for me). But I think that means it gets better and better? kensycooperrider.com/blog/advice-...
kensycooperrider.com
Advice to a young scholar — Kensy Cooperrider
Ten tips for young scholars on doing good work (and living well at the same time).
22010
Reposted by Sean Trott
Kanishka Misra @kanishka.bsky.social · 09/09/2026
Understanding how learners conclude “X laughed Y” is incorrect is an age-old question, with several hypotheses, some of which have been ~impossible to disentangle! @tomyxw.bsky.social, @fredashi.bsky.social, and I use controlled rearing to shed light on this in our new EMNLP paper: 1/n
Title slide for “Disentangling Statistical Preemption from Entrenchment in Language Models’ Avoidance of Overgeneralization,” by Yixuan Wang, Freda Shi, and Kanishka Misra. Includes the main results plot and experimental design.
2179
Sean Trott @seantrott.bsky.social · 31/08/2026
We've been working on this paper for a number of years now, and I think that's for the better, as it allowed us to integrate more of the growing body of empirical and theoretical work on LMs. Link again here for those who are interested: direct.mit.edu/opmi/article...
direct.mit.edu
Large Language Models as Distributional Baselines for Language Tasks
Abstract. In order to ask questions about the mechanisms underpinning human cognition, researchers must control for properties of stimuli that could confound detected effects. In experiments involving...
020
Sean Trott @seantrott.bsky.social · 31/08/2026
Our answer is no: LMs are existence proofs that a certain causal route is viable; the inferential value of such a result depends on the theoretical context, e.g., the alternative hypothesis at stake.
120
Sean Trott @seantrott.bsky.social · 31/08/2026
Finally, we discuss a range of best practices, as well as objections. For example: does a "successful" baselines result for a given task entail that LM-like mechanisms best account for the construct of interest in humans?
110
Sean Trott @seantrott.bsky.social · 31/08/2026
We then articulate the conditions under which distributional predictability is most potentially relevant: 1) when a study uses linguistic stimuli; and 2) it's investigating a claim about the *necessity* of some other factor or construct presumed to be inaccessible to LMs.
110
Sean Trott @seantrott.bsky.social · 31/08/2026
We first argue why distributional predictability matters from both theoretical and empirical perspectives. E.g., there's now a large body of work showing that measures like LM-derived surprisal are *sensitive* to various manipulations, and also *predictive* of human DVs on tasks.
110
Sean Trott @seantrott.bsky.social · 31/08/2026
Psychologists control for confounds in the design or analysis of their experiments (e.g., word frequency). We argue that we're in a position to also control for *distributional linguistic predictability* using contemporary LMs, and describe the cases in which we should do so (and how).
120
Sean Trott @seantrott.bsky.social · 31/08/2026
New paper out in @openmindjournal.bsky.social! "Large Language Models as Distributional Baselines for Language Tasks". With @jamichaelov.bsky.social , @camrobjones.bsky.social , Tyler Chang, and Ben Bergen. direct.mit.edu/opmi/article...
direct.mit.edu
Large Language Models as Distributional Baselines for Language Tasks
Abstract. In order to ask questions about the mechanisms underpinning human cognition, researchers must control for properties of stimuli that could confound detected effects. In experiments involving...
1126
Sean Trott @seantrott.bsky.social · 31/08/2026
Finally, we discuss potential best practices, as well as objections to our argument. For example: does a "successful" baselines result for a given task entail that LM-like mechanisms are the best account of the corresponding construct of interest?
000
Sean Trott @seantrott.bsky.social · 31/08/2026
We then articulate the conditions under which it is most potentially relevant: 1) when a study uses linguistic stimuli; and 2) when it's investigating a claim about the *necessity* of some other factor presumed to be inaccessible to an LM (say, physical grounding).
100
Sean Trott @seantrott.bsky.social · 31/08/2026
We first argue why distributional predictability matters from both theoretical and empirical perspectives. E.g., there's now a large body of work showing that measures like LM-derived surprisal are *sensitive* to various manipulations, and also *predictive* of human DVs on tasks.
100
Sean Trott @seantrott.bsky.social · 31/08/2026
Psychologists control for confounds in the design or anlaysis of their experiments (e.g., word frequency). We argue that we're now in a position to account for *distributional linguistic predictability* using contemporary LMs, and describe the cases in which we should do so (and how).
100
Sean Trott @seantrott.bsky.social · 30/08/2026
Already on my list!
020
Reposted by Sean Trott
Gary Lupyan @glupyan.bsky.social · 29/08/2026
Yes, exactly. Platonic representations are one of the many erroneous conclusions reached under the assumption that intelligence is a one-dimensional scalar.
0195
Sean Trott @seantrott.bsky.social · 27/08/2026
Would love to see something like that! I guess the problem is that if the other models are still available and not-watermarked, there's just an incentive to use those ones.
100
Sean Trott @seantrott.bsky.social · 27/08/2026
The thing I'm more pessimistic about is more wrt continued success of detection tools. Pangram is quite reliable, but that feels to me like a contingent, rather than necessary, fact about the current generation of LLMs and their empirical distinctiveness from most human writing.
100
Sean Trott @seantrott.bsky.social · 27/08/2026
Based on my limited understanding I think it's probably the most reliable approach? But of course its success/implementation depends on good governance—otherwise we're just dependent on the good faith implementations at the AI companies.
200
Sean Trott @seantrott.bsky.social · 27/08/2026
This is my intuition as well. Would be interesting to see if there's a way to estimate differences in things like lexical or syntactic diversity across modes as well (in addition to num. word tokens encountered).
120
Sean Trott @seantrott.bsky.social · 27/08/2026
This is exactly what makes me quite sad about the impact of AI writing on online discourse. I'm also just less optimistic than some about our continued ability (either as humans, or with tools like Pangram) to distinguish human-authored writing from AI-generated text. I don't know the solution!
100
Sean Trott @seantrott.bsky.social · 19/08/2026
Yeah, esp. wrt things like homework, once "legible signals" of learning/understanding are now effectively decoupled from the thing itself (a construct validity problem!). So it's harder to know what you don't know. I don't think this is inevitable *in theory* but it is (I think) the *modal* outcome.
110
Sean Trott @seantrott.bsky.social · 11/08/2026
Totally! Wordbank is a fantastic resource. External validity / generalizability is just something I've been puzzling over philosophically so it was great to see the chapter on this.
010
Reposted by Sean Trott
Mike Frank @mcxfrank.bsky.social · 11/08/2026
Almost no psychology research is based on a true probability sample. The question is whether that undermines your research! Ch 10 of Experimentology argues that it depends on whether the effect you're studying is heterogeneous. 🧵 experimentology.io
1204
Sean Trott @seantrott.bsky.social · 11/08/2026
(And ofc if you could measure the target, you wouldn't need the model.) My understanding is that approaches like "comparative process tracing" are helpful here, which I guess comes back to your point about developing sufficiently precise theories that allow us to predict heterogeneity.
110
Sean Trott @seantrott.bsky.social · 11/08/2026
I agree with this, but one challenge is how to go about doing this without doing more heterogeneous sampling in the first place. There's a similar problem in work with model organisms, when it's often hard to know whether results on a model will generalize to the target w/out measuring the target.
120
Sean Trott @seantrott.bsky.social · 11/08/2026
Yeah that's fair!
010
Sean Trott @seantrott.bsky.social · 10/08/2026
Yeah I also worry that might be true. The "self-renewal" I refer to is really just my optimistic read for a best-case scenario but I'm not sure it's the most likely one.
110
Sean Trott @seantrott.bsky.social · 10/08/2026
Wrote up my own thoughts here if you are interested: seantrott.substack.com/p/a-field-in...
seantrott.substack.com
A field in flux?
Reflections on ACL 2026.
120
Sean Trott @seantrott.bsky.social · 10/08/2026
Resonate with a lot of this, especially the shift towards journals. Like you I've had better experiences both submitting and reviewing for journals. The workshops are great too!
110
Sean Trott @seantrott.bsky.social · 28/07/2026
Latent constructs or "capabilities" that provide useful explanatory purchase for some swath of behavior (just as they do for human behavior). So, e.g., "Theory of Mind", which I've done research on wrt LLMs. I'm more bearish on the intentional stance but it probably has a place too.
010
Sean Trott @seantrott.bsky.social · 27/07/2026
Yeah definitely. My epistemological view is something like: 1. Focus on empirical characterizations of behavior, not axiomatic assumptions about ANNs/LMs. 2. Where possible, do (1) using causal-mechanist language. 3. Where necessary, invoke "capacities" to explain behavior.
110
Sean Trott @seantrott.bsky.social · 27/07/2026
Agreed. But I do think the vocabulary matters insofar as it leads us to make different predictions about the future circumstances in which analogous behavior occurs.
110
Sean Trott @seantrott.bsky.social · 09/07/2026
At a high level I very much agree that the field could learn a lot from the history and philosophy of science.
000
Sean Trott @seantrott.bsky.social · 09/07/2026
I'd personally love to see more *actually incremental* work in AI/ML that clearly articulated how and why the work is in the service of some set of explanatory goals. Without some sense of what counts as progress I worry that Chang's notion of epistemic iteration isn't really possible.
100
Sean Trott @seantrott.bsky.social · 09/07/2026
Great book. I wonder if the implicit critique is actually (often) that the work is *not* incremental (nor revelatory) in the sense that it is unclear what kind of theory/paradigm/research program it is moving forward or operating under?
100
Sean Trott @seantrott.bsky.social · 08/07/2026
Great chapter on a very important topic. Construct validity continues to be one of the biggest challenges in research characterizing the behavior of LLMs as well.
0132
Sean Trott @seantrott.bsky.social · 05/07/2026
@jamichaelov.bsky.social @camrobjones.bsky.social + Sam Taylor + Pam Rivière
000
Sean Trott @seantrott.bsky.social · 05/07/2026
Will be presenting this work measuring behavioral sensitivity to false belief manipulations in a sample of open-weight LMs at tomorrow's poster session (11am). Stop by if you're interested! #ACL2026 aclanthology.org/2026.acl-lon...
160
Reposted by Sean Trott
Mike Frank @mcxfrank.bsky.social · 29/06/2026
Most of the statistical tests you learned in stats class — t-test, ANOVA, correlation — are actually special cases of a single thing: linear regression. Ch 7 of Experimentology argues that thinking in models, not tests, is more flexible and a better foundation for theory. 🧵 experimentology.io
330761
Sean Trott @seantrott.bsky.social · 21/06/2026
Great points throughout! This is basically my view as well.
020
Reposted by Sean Trott
James Michaelov @jamichaelov.bsky.social · 11/06/2026
Seems like a good time to share our new preprint about model openness! (with @catherinearnett.bsky.social @tylerachang.bsky.social Pamela D. Rivière, Samuel M. Taylor @camrobjones.bsky.social @seantrott.bsky.social @rplevy.bsky.social Ben Bergen, and Micah Altman): arxiv.org/abs/2603.26539
1193
Reposted by Sean Trott
Stella Biderman @stellaathena.bsky.social · 10/06/2026
In film, "we'll fix it in post" is what you say when something went wrong on set and you don't want to redo it. AI research has made it our entire methodology: train the model, then patch whatever comes out. Our new ICML oral argues this can't be the basis of a science of AI. 🧵
310823