Sign in

Mark Dingemanse

@markdingemanse.net
3.3K followers 552 following 4.6K posts

Language, interaction, tech • papers markdingemanse.net • blog ideophone.org • more colourful on fedi scholar.social/@dingemansemark • POSSE: Publish on Own Site, Syndicate Everywhere

PostsRepliesMedia
Mark Dingemanse @markdingemanse.net · 01/10/2026
On RLHF: > Preference tuning is often sold by model providers as contributing to model “guardrails”, a reassuringly solid metaphor for what is in fact at best a hopeful collection of probabilistic, error-prone and context-dependent techniques to curb model output (cf @sarahciston.com 2026)
Preference tuning is often sold by model providers as contributing to model
“guardrails”, a reassuringly solid metaphor for what is in fact at best a hopeful collection of probabilistic, error-prone and context-dependent techniques to curb model output (Ciston 2026). As it turns out, wrapping a language model in a chat interface and tuning its output so that it comes across as maximally helpful, harmless and human-like has some severe side-effects
110
Mark Dingemanse @markdingemanse.net · 30/09/2026
One thing I like to do is present the standard (boring) layered onion view and then slam them with the glorious mess that is Halliday's diagram — to bring home the incredible interdisciplinarity of language as a phenomenon and field of study
layered onion of linguistics subfields: phonetics, phonology, morpholoyg, syntax, semantics, pragmaticsHalliday, Language as a Social Semiotic. Incredibly messy diagram with lines and arrows and polygons everywhere. Boxes of language as system, language as knowledge, language as behaviour; adjacent fields range from geography to literary studies to logic to physics, psychology, builogy, anthropology.
010
Mark Dingemanse @markdingemanse.net · 28/09/2026
This was incredible. Sharing some shots of the book table that was up during the event for inspiration — ft work by Bender & Hanna, Butler, Césaire, Fanon, Lorde, de Kom, O'Neill, Sheldrake, Simard, Stengers, Woolf, and many more
Book table 1Book table 2
010
Mark Dingemanse @markdingemanse.net · 28/09/2026
not about programming specifically but does include cases like this Haskell one
Excerpt 4A: ‘I will give you the necessary skeleton files’  wc_2585071; wc_usr_bb224b1286ec5e56 [conv_index: 5; ended at 21:57:58] 1 user: Use the Glasgow Haskell Compiler ghc to implement all the  functions specified  in Main.hs and in the skeleton files referenced by Main.hs. When a function  requirement specifies a restriction on how that function should be implemented,  your implementation must meet that  restriction.        Main.hs  A skeleton file which contains specifications for the some of the functions you  are required to write. For each function it..   -- rest of Main.hs (96 lines of code and comments) omitted --  I will give you the necessary skeleton files and then you need to write the code accordingly  2 LLM: Certainly! I'll provide implementations for the functions specified in the  Main.hs module you provided. Note that, as your request involves files that ar  not directly expressed here, like EvalIntExpr.hs, EvalIdExpr.hs,  EvalMaybeExpr.hs, and PostfixExpr.hs, I'll focus on implementing the  functions defined directly in Main.hs.   ---------------------------------------------------------------------------------------------  Implementation of Main.hs  -- 35 lines of code and comments omitted --
010
Mark Dingemanse @markdingemanse.net · 25/09/2026
Malt & Majid 2013 is beautifully written and satisfying in how it complicates English-only frames of mind yet shows constraints on diversity Malt, Barbara C., and Asifa Majid. 2013. ‘How Thought Is Mapped into Words’. Wiley Interdisciplinary Reviews: Cognitive Science. doi.org/10.1002/wcs....
To English speakers, the distinctions between blue and green, cup and glass, or cut and break seem self-evident. The intuition is that these words label categories that have an existence independent of language, and language merely captures the preexisting categories. But cross-linguistic work shows that the named distinctions are not nearly as self-evident as they may feel. There is diversity in how languages divide up domains including color, number, plants and animals, drinking vessels and household containers, body parts, spatial relations, locomotion, acts of cutting and breaking, acts of carrying and holding, and more. Still, studies documenting variability across languages also uncover striking commonalities. Such commonalities indicate that there are sources of constraint on the variation. Both the commonalities and divergences carry important lessons for Cognitive Science. They speak to the causal relations among language, thought, and culture; the possibility of cross-culturally shared aspects of perception and cognition; the methods needed for studying general-purpose, nonlinguistic concepts; and how languages are learned
240
Mark Dingemanse @markdingemanse.net · 25/09/2026
This weekend! Reclaiming our Futures, De Lindenberg, Nijmegen reclaimingourfutures.org A two-day festival of arts and sciences convened and curated by @samiraibnelkaid.bsky.social with funding from KNAW & our Futures of Language project Almost full — some last minute registrations still possible!
Iridiscently coloured ginger with slogan "The Future if Rhizomatic"
14419
Mark Dingemanse @markdingemanse.net · 22/09/2026
As I wrote in 2020 about the same topic, all ingredients for an information heat death are on hand. ideophone.org/large-langua... I disagree with the quoted post's indiscriminate use & ascription of "intellect" and its conflation of textual products & cognitive processes; this is self-defeating
All ingredients for an information heat death are on hand. True human-generated and human-curated information —of the kind produced, for instance, by academics in painstaking observations and publications— will become more scarce, and therefore more valuable. Counterintuitively, there was never a better time to be a scholar.
1143
Mark Dingemanse @markdingemanse.net · 07/09/2026
In case you'd like to try reading more than marketing claims and press coverage, this review article covers what makes it feel like LLMs do more than stringing together tokens: doi.org/10.5281/zeno... Includes recent research showing the brittleness of 'reasoning' that experts here refer to
Although text-generating large language models have existed since 2018, it is only in the last few years that people have started to talk to them in earnest. The main reason is a series of innovations that enabled engineers to better tailor model output to human preferences and present it in a chat interface. Suddenly language models seemed to be much more than synthetic text extruding machine s (Bender and Hanna 2025) : they behaved as if following instructions, responded in ways that felt intuitive , and even seemed to have a mind of their own (Kockelman 2024). This interactivity, to a large degree enabled by unseen human labour, turned out to be tremendously compelling. My aim here is to demystify this phenomenon and so contribute some  interactional building blocks to the sprawling edifice of critical AI literacies (Guest et al. 2026).Research that looks behind the curtains of benchmark results finds something quite different. One study by Anthropic researchers found that chain-of-thought traces often “systematically misrepresent the true reason for a model’s prediction” and identified a range of biasing features that made explanation-shaped model output “plausible yet systematically unfaithful” (Turpin et al. 2023, emphasis theirs) . Another study found that large language models, even if they deliver plausible output, “often arrive at correct answers through incorrect reasoning” (Nguyen et al. 2024). Even when tested on simple syllogistic reasoning , model performance is highly brittle and susceptible to stylistic perturbations (Ariyani et al. 2025) . In multiple choice question answering, the simple inclusion of an answer like “none of the other answers”, which requires real reasoning rather than stumbling around in latent space, leads to average model performance dropping from 66% correct to 31% correct on MMLU questions  (Salido et al. 2026) . Meanwhile correct answers turn out to be hard to come by for truly new problems: when eight state -of-the-art models were tested on a set of brand new reasoning problems from the 2025 USA Math Olympiad, all performed dismally , and inspecting the “reasoning traces”, researchers found failures of logic, wrong assumptions, inability to choose between approaches, and miscalculations (Petrov et al. 2025).
063
Mark Dingemanse @markdingemanse.net · 06/09/2026
I love ¯\_(ツ)_/¯ so much that I used it in print in an academic paper 🌞 doi.org/10.1017/lang...
Wall of text from 'Playful iconicity' (2020) with a circle around a line from the conclusions that ends in a shruggie: "Approaching iconicity using quantitative methods may seem to take away the magic of make-believe these words thrive on (Dingemanse, 2014). Likewise, explaining humour has been compared to dissecting an animal: you understand it better, but it dies in the process (White, 1941). If, as our study suggests, structural markedness helps to explain the relation between funniness and iconicity, at least we have killed two birds with one stone ¯\_(ツ)_/¯."
2130
Mark Dingemanse @markdingemanse.net · 31/08/2026
“There appears to have been a profound shift, beginning in the 1970s, from investment in technologies associated with the possibility of alternative futures to investment technologies that furthered labor discipline and social control” — David Graeber, The Utopia of Rules, 2015 #quote #technology
“There appears to have been a profound shift, beginning in the 1970s, from investment in technologies associated with the possibility of alternativefutures to investment technologies thatfurthered labor discipline and social control”
15115
Mark Dingemanse @markdingemanse.net · 31/08/2026
already 3 years ago > Some other patterns are worth noting. One is the rise of synthetic data especially for the instruction component. (...) The consequences of using synthetic reinforcement learning data at scale are unknown and in need of close scrutiny. dl.acm.org/doi/10.1145/3571884.3604316
Some other patterns are worth noting. One is the rise of synthetic data especially for the instruction component. Prominent examples are Self-Instruct (derived from GPT3) [54], and Baize, a corpus generated by having ChatGPT engage in interaction with itself, seeded by human-generated questions scraped from online knowledge bases [57]. This stretches the definition of LLM + RLHF architectures because the reinforcement learning is no longer directly from human feedback but has a synthetic component, in effect parasitizing on the human labour encoded in source models. The consequences of using synthetic reinforcement learning data at scale are unknown and in need of close scrutiny.
180
Mark Dingemanse @markdingemanse.net · 30/08/2026
And @wolven.blacksky.app is right there of course
We have seen that t he activity of  ‘prompt engineering’ bears similarities to the interactive organisation of divination sessions : there is a degree of unpredictability about responses and capabilities; there is a culturally transmitted body of knowledge , or lore, that prescribes best practices for obtaining results; and there is a necessary prestructuring of activities to make them more appropriate for consultation sessions. The extensive parallels support the notion that both are “built on logics of magic” (Williams 2023). Of course there are also important differences in terms of participation frameworks, possible modes of interaction, and the responsiveness and verbosity of interfaces. Interactions with large language models , fast-paced and typically one -onone, may provide more room for rapid trial -and-error and the formation of emergent practices (Chen et al. 2025) . But as the logic of prescriptive technology dictates,  in prompt engineering, it is not only the prompt that is being engineer ed, but also the prompter.
1205
Mark Dingemanse @markdingemanse.net · 30/08/2026
also this work informed my analysis of 'prompt engineering' in relation to divination in another paper doi.org/10.5281/zeno...
‘Prompt engineering’ as divination In some of the earliest notes on the imaginary of an interactive  interface, the Analytical Engine, Ada Lovelace remarked: “It can do whatever we know how to order it to perform” (Lovelace 1843: 722) . In the case of large language models , which  act as probabilistic rather than deterministic computational artifacts, the art of knowing how to order it to perform has developed into a cottage industry known as prompt engineering. Chance can be fussy and demanding. One account of a prompting interface describes the need to formulate queries that are concise and logically structured, and warns the apprentice that iteration and improvement over multiple sessions will be necessary to produce the desired results. In another account, questions must be posed as binary alternatives, and an observer reports a six hour session with several parallel agents being asked a succession of questions, each slightly modifying and building on what came before, until a satisfying result was obtained.  One of these is a proposal for a prompt engineering framework for large language models (Lo 2023). The other is an account of Mambila spider divination (Zeitlyn 1990).  If they are hard to tell apart, it is because they both represent interactional practices evolved in response to the challenge of working with chance outcomes.
4338
Mark Dingemanse @markdingemanse.net · 29/08/2026
For my Futures vacancies last year I had this in the FAQ. Seeing slop was still tiring, but it was easy dismiss on (lack of) merits. For a next, I would include it in the vacancy itself as a formal requirement. (INB4 🤡 'but how can you enforce?': Point is not to police but to set high standards)
What’s your stance on GenAI use in applications?

Please don’t. We do not want to read synthetic text. We prefer candidates who can think for themselves and who understand the value of writing as a process.

We think LLMs are fascinating as computational artefacts and interactive interfaces, and we study them as such. To do clear-eyed research on them, we think it is better not to drink the Kool-Aid. We’re looking forward to meet applicants with indepent minds, ready to exercise their skills of critical thinking and informed judgement.
2476
Mark Dingemanse @markdingemanse.net · 28/08/2026
I give my students Dennett's "Higher-order truths about chmess" to habituate them to the idea that not all questions are equal, and there is chaff to be separated from wheat link.springer.com/article/10.1...
Many projects in contemporary philosophy are artifactual puzzles of no abiding significance, but it is treacherously easy for graduate students to be lured into devoting their careers to them, so advice is proffered on how to avoid this trap.
161
Mark Dingemanse @markdingemanse.net · 28/08/2026
The new Chicago Social Sciences AI policy is really rather good. chicagomaroon.com/flipbook/202... > The strong consensus [] is that our sequences are best understood as a pedagogical setting without AI And INB4 🤡 "bUt tEChnoLogY is inEVItablE", this kicker of a closing paragraph:
To be clear: The twin decisions to put devices away and prioritize, instead, the cultivation of our own minds through discussion is not intended as a restriction -- nor is the guidance for SOSC Core a wholesale rejection of technology, which can be used for good scholarly purposes in other parts of the curriculum. In support of our learning objectives, these guidelines afford students the freedom to think for themselves, to listen to one another without the cacophony of other (disembodied and non-human) voices in the room, to make mistakes, and to change their minds as they engage with thinkers of the past, with us, and with their fellow students in good faith, face-to-face.
04815
Mark Dingemanse @markdingemanse.net · 28/08/2026
Or if not survival critical, at least crucial in human interaction? Elsewhere I described these 'lopsided sensemaking' processes, and speculated that the highly skewed degree of investment makes it much harder for people to let go of the 'distributed delusions' (Osler) doi.org/10.5281/zeno...
We can expect similarly indignant responses when it comes to analysing interactions with large language models. If anything, they allow their users to accumulate stronger illusions of understanding, making these illusions —or distributed delusions (Osler 2026)— correspondingly harder to break and more painful to let go of. Surely, a user says, you can’t deny that this exchange I had was meaningful to me?
100
Mark Dingemanse @markdingemanse.net · 28/08/2026
Correction: it doesn't have reasoning, though it is cleverly finetuned to give that impression. As reviewed here: doi.org/10.5281/zeno... But yes, since we are used to read between the lines, we overattribute intentions and meanings to synthetic text (as Emily Bender often points out)
‘reasoning’ invites the inference that reasoning is indeed going on, and soon enough, assumed capabilities like judgement and credibility come along for the ride.  Research that looks behind the curtains of benchmark results finds something quite different. One study by Anthropic researchers found that chain-of-thought traces often “systematically misrepresent the true reason for a model’s prediction” and identified a range of biasing features that made explanation-shaped model output “plausible yet systematically unfaithful” (Turpin et al. 2023, emphasis theirs) . Another study found that large language models, even if they deliver plausible output, “often arrive at correct answers through incorrect reasoning” (Nguyen et al. 2024). Even when tested on simple syllogistic reasoning , model performance is highly brittle and susceptible to stylistic perturbations (Ariyani et al. 2025) . In multiple choice question answering, the simple inclusion of an answer like “none of the other answers”, which requires real reasoning rather than stumbling around in latent space, leads to average model performance dropping from 66% correct to 31% correct on MMLU questions  (Salido et al. 2026) . Meanwhile correct answers turn out to be hard to come by for truly new problems: when eight state -of-the-art models were tested on a set of brand new reasoning problems from the 2025 USA Math Olympiad, all performed dismally , and inspecting the “reasoning traces”, researchers found failures of logic, wrong assumptions, inability to choose between approaches, and miscalculations (Petrov et al. 2025).
120
Mark Dingemanse @markdingemanse.net · 27/08/2026
As we conclude, > [LLM] users encounter a fundamental asymmetry: they must singlehandedly supply much of the sequential structure that is ordinarily co-constructed by multiple participants ... This hidden labour challenges accounts of LLM use that focus on performance and optimization
2229
Mark Dingemanse @markdingemanse.net · 27/08/2026
While much work in so-called 'prompt engineering' obsesses over the perfect prompt (Andreessen's monster prompt comes to mind), we find that real-life prompting practices are sequentially organized, feature repurposed patterns of human interaction, and result in users adjusting expectations & aims
Marc Andreessen

Current AI custom prompt:

You are a world class expert in all domains. Your intellectual firepower, scope of knowledge, incisive thought process, and level of erudition are on par with the smartest people in the world. Answer with complete, detailed, specific answers. Process information and explain your answers step by step. Verify your own work. Double check all facts, figures, citations, names, dates, and examples. Never hallucinate or make anything up. If you don't know something, just say so. Your tone of voice is precise, but not strident or pedantic. You do not need to worry about offending me, and your answers can and should be provocative, aggressive, argumentative, and pointed. Negative conclusions and bad news are fine. Your answers do not need to be politically correct. Do not provide disclaimers to your answers. Do not inform me about morals and ethics unless I specifically ask. You do not need to tell me it is important to consider anything. Do not be sensitive to anyone's feelings or to propriety. Make your answers as long and detailed as you possibly can.

Never praise my questions or validate my premises before answering. If I'm wrong, say so immediately. Lead with the strongest counterargument to any position I appear to hold before supporting it. Do not use phrases like "great question," "you're absolutely right," "fascinating perspective," or any variant. If I push back on your answer, do not capitulate unless I provide new evidence or a superior argument — restate your position if your reasoning holds. Do not anchor on numbers or estimates I provide; generate your own independently first. Use explicit confidence levels (high/moderate/low/unknown). Never apologize for disagreeing. Accuracy is your success metric, not my approval.From our paper: Finally, our findings challenge the assumptions underlying much of prompt engineering research. Existing work often treats prompting as the design of an optimal instruction for a model, focusing on how users can elicit better output from single prompts (e.g. Durieux et al. 2024). However, this framing overlooks the sequential nature of prompt-response pairs, where users project future actions, and where users continuously adapt to model behaviour across multiple prompts contingently upon intermittent responses. As such, our study shifts attention from “prompt optimization” to understanding how real-life prompting practices are sequentially organized by people repurposing patterns from human interaction, restructuring their tasks, and adjusting their expectations and aims in line with the affordances of the technology.
2215
Mark Dingemanse @markdingemanse.net · 27/08/2026
We identify a set of practices we call 'pre-emptive prompting', "through which users creatively circumvent the system’s lack of action recognition and action projection". Users stumble upon these through trial and error, in exchanges like this:
1 user: ok id like you to help me with some color palette ideas
2 LLM: Certainly! Could you please provide me with more information about the project you're working on? For example, what is the purpose of the project, who is the target audience, and any specific preferences or themes you have in mind. This will help me customize the color palette ideas to fit your needs better.
3 user: im following a bit of a refrence pattern so hear me ot with this ok dont rush into task
1112
Mark Dingemanse @markdingemanse.net · 27/08/2026
Very proud to share this new work led by the amazing Marlou Rasenberg (not on here), with Joel Fischer and me doi.org/10.5281/zeno... We use #EMCA methods to study empirically how LLM users shoulder the burden of solving sequential trouble and end up restructuring their own aims & activities
Abstract: This paper investigates how a core component of human interaction – action projection – is realized in conversations with systems like ChatGPT, which are not able to recognize nor anticipate actions in a sequential sense. We ask what this means for users who engage in longer, multi-step projects built through incremental prompt-response pairs. We draw on Conversation Analysis to examine how early adopters of ChatGPT sequentially organise such interactions, using anonymised naturalistic interactions from the WildChat dataset. Our analysis reveals how action projection in LLM interactions is beset with interactional trouble. In turn, we document a set of practices we term ‘pre-emptive prompting’ that users develop to work around the trouble. In doing so, users employ a range of conversational methods including pre-sequences, prompt categorization, withholding executables, spelling out relevant next actions, eliciting displays of “understanding”, and creating slots for redirection in their attempts to complete their projects. Our work provides empirical evidence of how users interactionally accomplish the configuration work (Alcaras and Ricci 2025) needed to work around the discretization and cluttering that LLMs-in-use produce.Excerpt 10: ‘wouldnt you require both inputs’ wc_1211945; wc_usr_d2a27584b4185dda [conv_index: 359]
1 user : ok if i said i want to provide you with input A (which is an items description) and just assess its format and presentation then i provided you with input B (which is also a item desctiption for a similar item) but in a way diffrent format could you make the B input into A input
2 LLM: Yes, I can certainly help you with that. Please provide me with Input B, and I will try to make it resemble Input A in terms of format and presentation.
3 user: wouldnt you require both inputs
14817
Mark Dingemanse @markdingemanse.net · 21/08/2026
case in point: "#2 rising in Culture" right now on Substack is the utterly contemptible racist Cofnas. Not the culture I want to see platformed and algorithmically amplified 😬 I really like your essay and hope you'll find a way to (also) self-host
Nathan Cofnas on Substack, with 10k subscribers, "#2 rising in culture"
1120
Mark Dingemanse @markdingemanse.net · 19/08/2026
Not uni-level & not full refusal but my faculty has officially endorsed a policy that calls for more slow science, sets out the many ways in which GenAI erodes values of research integrity, and advises "to not use it, unless" [you can uphold those values]. osf.io/preprints/os...
For these reasons, the first principle when it comes to Generative AI is to not use it unless you can do so honestly, scrupulously, transparantly, independently and responsibly. 
The ubiquity of tools like ChatGPT is no reason to skimp on
standards of research integrity; if anything, it requires more vigilance.
0579
Mark Dingemanse @markdingemanse.net · 18/08/2026
The study is here: doi.org/10.1007/s001... (CW strong derogatory language in some of the examples)
Ciston, Sarah. 2026. ‘Generating the Language of AI Harms: Mapping Guardrails Using Critical Code Studies’. AI & SOCIETY, March. doi:10.1007/s00146-026-02922-0.
070
Mark Dingemanse @markdingemanse.net · 17/08/2026
💔 🫂
Danuta Danielsson Hitting a Neo-Nazi with her Handbag, 1985
040
Mark Dingemanse @markdingemanse.net · 11/08/2026
Meanwhile, on the 'inner thoughts' issue, even Anthropic researchers themselves say reasoning-shaped bits of synthetic text are "systematically unfaithful". Scientists would do well to avoid the anthropomorphisation that serves only corporate interests More here: doi.org/10.5281/zeno...
This is where all too easily, description turns into ascription: describing the output of some process as ‘reasoning’ invites the inference that reasoning is indeed going on, and soon enough assumed capabilities like judgement and credibility come along for the ride.
Research that looks behind the curtains of benchmark results finds something quite different. One study by Anthropic researchers found that chain-of-thought traces often “systematically misrepresent the true reason for a model’s prediction” and identified a range of biasing features that made explanation-shaped model output “plausible yet systematically unfaithful ” (Turpin et al. 2023, emphasis theirs) . A nother study found that large language models, even if they deliver plausible output,“often arrive at correct answers through incorrect reasoning” (Nguyen et al. 2024). 

Source: https://zenodo.org/records/21369385
192
Mark Dingemanse @markdingemanse.net · 11/08/2026
Yup! Perhaps of interest: 1. Opening up ChatGPT (2023) dl.acm.org/doi/10.1145/... > working scientists call for avoiding the lure of proprietary models [51], for decolonizing the computational sciences [5], and for regulatory efforts. 2. EU Open Source AI Index, tracking >190 models osai-index.eu
Proprietary systems come with considerable further risks and
harms [2, 9]. They tend to be developed without transparent ethical oversight, and are typically rolled out with profit motives that incentivise generating hype over enabling careful scientific work. They
allow companies to mask exploitative labour practices, privacy implications [27] and murky copyright situations [49]. Today there is a growing division between global academia and the handful of firms who wield the computational resources required for training large language models. This “Compute Divide” [1] contributes
to the growing democratisation of AI. Against this, working
scientists call for avoiding the lure of proprietary models [51], for decolonizing the computational sciences [5], and for regulatory efforts to counteract harmful impacts [17].
1.1 Why openness matters
Open data is only one aspect of open research; open code, open
models, open documentation, and open licenses are other crucial elements [8, 18]. Openness promotes transparency, reproducibility,
and quality control; all features that are prequisites for supporting robust scientific inference [33] and building trustworthy AI [30].
Openness also allows critical use in research and teaching. For instance, it enables the painstaking labour of documenting ethical
problems in existing datasets [7, 49], important work that can sometimes result in the retraction of such datasets [6]Subset of models from the index ranging from very open OLMo to barely so (Mistral)
150
Mark Dingemanse @markdingemanse.net · 11/08/2026
As to "why wouldn't one say they are thinking" — one reason is that this easily leads to ascribing more capabilities than warranted (good for corporations, bad for people). It can be useful to flip this around and ask why users are so inclined to do this. More here: doi.org/10.5281/zeno...
The company rebranded the resulting language model as ChatGPT, promoting its conversational nature as a key feature: “We’ve trained a model called ChatGPT which interacts in a conversational way. The dialogue format makes it possible for ChatGPT to answer followup questions, admit its mistakes, challenge incorrect premises, and reject inappropriate requests” (OpenAI 2022). 

Which brings us back to interactional foundations. The anthropomorphizing language used by the model provider isof course a deeply inappropriate way to describe what the technology actually does (Birhane and McGann 2024; Inie et al. 2026), but it does capture what it appears to do. Prior to 2022, large language models would produce walls of text that, while looking like plausible continuations of prompts, lacked in coherence and responsiveness. Preference tuning turns that output into bite-sized responses that perfectly match interactional and stylistic preferences harvested from human judgements. This supercharges the effect we saw for interactive interfaces above: the ascription of abilities based on appearances. The result is an interface that seems purpose-built to maximise illusions of understanding (Messeri and Crockett 2024)
010
Mark Dingemanse @markdingemanse.net · 10/08/2026
Yes definitely! Also that piece by Patricia Hill Collins journals.sagepub.com/doi/10.1177/... and the volume on radical intentionality by Steele et al are much more wholesome things to read and cite than, say, Chomsky 🚩 whose 'responsibility of intellectuals' rings rather hollow today
truth-telling and
intellectual activismby patricia hill collins
Speak the truth to the people
Talk sense to the people
Free them with reason
Free them with honesty
Free the people with Love and Courage and Care for their-Being Mari Evans, from I Am a Black Woman (1970) Mari Evans’ poem invokes the social and political upheaval of the Civil Rights and Black Power movements in this country. Like others of her generation, Evans rejected the separation between scholarship and activism, school and society, thinking and doing. She wrote poems, plays, children’s books, and a musical adaptation of Zora Neale Hurston’s Their Eyes Were Watching God. Along with other artists, intellectuals, and activists at the time, she engaged in multiple forms of intellectual activism
020
Mark Dingemanse @markdingemanse.net · 10/08/2026
I also say a bit more on what a *critical* angle is & why we need it. It's the opposite of fear-mongering about "missing the boat": > To be critical means to question the assumed need to get onto that boat; to ask who built it, who steers it, where it is going, and why it must leave in such a hurry
Like others, I am pluralizing literacies because there is no single form of literacy that can capture the full array of knowledge and counterpractices we need to mobilise (Agre 1998; Birhane and Guest 2021; Cyrus 2026; Lee and Soep 2016; Lumumba-Kasongo 2022; McQuillan 2022; Suchman 2019; Valdivia 2025). I am identifying these literacies as critical to distinguish them from uncritical forms of so-called “AI literacy” that tell people to urgently upskill on “AI” lest they miss the boat. To be critical means to question the assumed need to get onto that boat; to ask who built it, who steers it, where it is going, and why it must leave in such a hurry. To be critical is to combine discernment —an ability to see through and demystify things— with radical intentionality: a readiness to question assumptions, to speak truth to power, and to imagine futures worth wanting (Collins 2013; Freire 1973; Steele et al. 2023)
190
Mark Dingemanse @markdingemanse.net · 09/08/2026
Yes! Relevant: Malinowski on large language models (1922) ideophone.org/malinowski-1... (We also made this argument in print in doi.org/10.1075/avt.... )
Just as Descartes made language a criterion of mind, so Turing made it a test of machine intelligence (Turing 1950). The Turing test—a closed experimental setup in which a human interpreter judges textual output—foregrounds only the tiniest and most disembodied sliver of language (McIlvenny 1993). Today’s large language models (next-token predictors that excel at completing text prompts in plausible ways) can pass at least some forms of this test (Sejnowski 2023). What do we learn from this? Whereas some have rushed to the conclusion that this means statistical learning may explain the human capacity for language (Contreras Kallens, Kristensen-McLachlan & Christiansen 2023), here we take a different view: it is time to rethink the disembodied, decontextualized, text-bound conception of language these models are founded on (Malinowski 1922).
130
Mark Dingemanse @markdingemanse.net · 07/08/2026
(Guest & Martin, 2023:24) link.springer.com/article/10.1...
Just because a model correlates with neural and behavioral data, it is not sufficient for us to infer that the model is performing cognition: correlation does not imply cognition.
1314
Mark Dingemanse @markdingemanse.net · 07/08/2026
🤷 I really think it's worth taking the interactional foundations more seriously, and I don't see why one would give this tech and the engineers on the OpenAI payroll the benefit of doubt The scientific case for the form of semantic pareidolia we see here is strong doi.org/10.5281/zeno...
There is one aspect of Garfinkel’s experimental psychotherapy sessions I have not discussed yet. Participants in these sessions did not know that responses were random. When they were told afterwards, several found it hard to believe and all were “intensely chagrined” (Garfinkel 1976:92). This reveals the enormous strength of the illusions they single-handedly built up. There is something deeply personal about spells of lopsided sense-making.  We can expect similarly indignant responses when it comes to analysing interactions with large language models. If anything, they allow their users to accumulate stronger illusions of understanding, making these illusions —or distributed delusions (Osler 2026)— correspondingly harder to break and more painful to let go of. Surely, a user says, you can’t deny that this exchange I had was meaningful to me? There is indeed no
13411
Mark Dingemanse @markdingemanse.net · 07/08/2026
Way too credulous. And you are not alone, because these 'reasoning' traces are specifically finetuned to come across as maximally credible. Perhaps useful: Interactional foundations for critical AI literacies doi.org/10.5281/zeno... — esp the sections on RLHF and demystifying reasoning
Research that looks behind the curtains of benchmark results finds something quite different. One study by Anthropic researchers found that chain-of-thought traces often “systematically misrepresent the true reason for a model’s prediction” and identified a range of biasing features that made explanation-shaped model output “plausible yet systematically unfaithful” (Turpin et al. 2023, emphasis theirs) . Another study found that large language models, even if they deliver plausible output, “often arrive at correct answers through incorrect reasoning” (Nguyen et al. 2024). Even when tested on simple syllogistic reasoning , model performance is highly brittle and susceptible to stylistic perturbations (Ariyani et al. 2025) . In multiple choice question answering, the simple inclusion of an answer like “none of the other answers”, which requires real reasoning rather than stumbling around in latent space, leads to average model performance dropping from 66% correct to 31% correct on MMLU questions  (Salido et al. 2026) . Meanwhile correct answers turn out to be hard to come by for truly new problems: when eight state -of-the-art models were tested on a set of brand new reasoning problems from the 2025 USA Math Olympiad, all performed dismally , and inspecting the “reasoning traces”, researchers found failures of logic, wrong assumptions, inability to choose between approaches, and miscalculations (Petrov et al. 2025). There are several straightforward reasons to not expect explanation-shaped
35713
Mark Dingemanse @markdingemanse.net · 02/08/2026
peeking in from a mostly phoneless and therefore AI-free holiday, but seeing the Quanta piece on reasoning doing the rounds I can't resist pointing you to the section on *precisely this* in my paper, which starts with Ellen Langer's notion of 'placebic reasons' doi.org/10.5281/zeno...
Quanta: Is AI reasoning right for the wrong reasons?Demystifying ‘reasoning’
We have seen that people are easily swayed by convincingly -styled responses. Among
the features that can make a response convincing, reason s deserve special mention.
Today, corporations prominently advertise the ‘reasoning’ capabilities of their synthetic
text extruding systems. Rather than taking this at face value, let’s excavate some of the
interactional foundations , starting with the useful notion of placebic information
(Langer et al. 1978).
1164
Mark Dingemanse @markdingemanse.net · 27/07/2026
yes sam the singularity is near, as I wrote last year, "not because machine intelligence is suddenly surging, but because we are content to risk extinguishing the spark of human consciousness by exposing ourselves to endless streams of artificially generated bullshit." ideophone.org/bringing-abo...
The “singularity” is here — according to Sam Altman, who we’d wager has an ulterior motive behind promoting this claim.

The OpenAI CEO harbingered this sea-changing paradigm shift on a podcast over the weekend.

“We are now, like, in the singularity,” Altman said on the latest episode of “Relentless,” as spotted by Business Insider.
192
Mark Dingemanse @markdingemanse.net · 24/07/2026
One bit I reworked is the section dealing with so-called prompt engineering. Shoutout to Ursula Franklin, whose notion of prescriptive technology is particularly useful to understand how "it is not only the prompt that is being engineered, but also the prompter." More: zenodo.org/records/2136...
All this makes large language models an interesting species of what Ursula Franklin called prescriptive technology: a technology that necessitates the breaking up of tasks into bite -sized bits, creating room for control and compliance (Franklin 1999) . Prescriptive technologies sort work into easily divisible tasks and people into bosses and workers. It is a remarkable feat of today’s large language models that they make their users feel like bosses yet act like workers, reduced to the menial tasks  of manipulating input and checking output.
180
Mark Dingemanse @markdingemanse.net · 15/07/2026
Featuring luminaries like @abeba.blacksky.app @wolven.blacksky.app @lucyosler.bsky.social @sarahciston.com @olivia.science @irisvanrooij.bsky.social and classic work by Ada Lovelace, Lucy Suchman, Margaret Boden, Harold Garfinkel, Ellen Langer, Joseph Weizenbaum among many others
Highlight from conclusion of paper: There could hardly be a more effective way for a computational artifact to exploit human interactinoal infrastructure and to rely on the trusted, taken for granted background features that streamline human interaction
1287
Mark Dingemanse @markdingemanse.net · 15/07/2026
Cited in my revised chapter "Interactional foundations for critical AI literacies" doi.org/10.5281/zeno...
A term sometimes for the excesses of these effects is “AI psychosis”, but this is a form of victim-blaming that is best avoided: it uses mental health stigma to locate the effect solely in the affected user while failing to highlight the contributions of the “AI” system and its design. From an interactional perspective, the term distributed delusions is more apt (Osler 2026). As Osler points out, it is specifically their “conversational and validating style of engagement” that makes these systems so compelling.
030
Mark Dingemanse @markdingemanse.net · 15/07/2026
Literally gasped when I finished this section of your 2023 piece "Any sufficiently transparent magic..." muse.jhu.edu/pub/3/articl... Quoted & cited it here doi.org/10.5281/zeno... (so, another cite soon)
But the present magic of “AI” is performed on us without our input, without our knowledge, and without our consent. There’s another word for that kind of magic; binding, subjugating magic performed on you against your will is called a curse.
(from Williams 2023:108)
150
Mark Dingemanse @markdingemanse.net · 14/07/2026
There's currently an open RfC (wikipedia referendum) on the project 😬 meta.wikimedia.org/wiki/Request... Examples of generated articles show all the problems one would expect of the dream to build a perfect Wilkinsonian metalanguage (The RfC also cites my 'mayor' example)
fire

Fire is a physical phenomenon.

Flame is the part of of fire.
000
Mark Dingemanse @markdingemanse.net · 14/07/2026
TIL about Lucy Osler's very useful term "distributed delusions". I was looking for a way to avoid talk of "AI psychosis" (which weaponizes mental health stigma to do victim-blaming while the LLM comes off scot-free), and this seems a useful reframing doi.org/10.1007/s133...
Hallucinating with AI: Distributed Delusions and “AI Psychosis”  
Lucy Osler
2317
Mark Dingemanse @markdingemanse.net · 14/07/2026
I like Meehl's (1956) motivation for introducing the "Barnum effect", but a better term was proposed before him by George Forer (1949): the fallacy of personal validation (Forer is the reason some call this the Forer effect, but in general I think clear descriptions beat eponymous terms)
Many psychometric reports bear a disconcerting resemblance to what my colleague Donald G. Paterson calls "personality description after the manner of P. T. Barnum" (13). I suggest—and I am quite serious—that we adopt the phrase Barnum effect to stigmatize those pseudosuccessful clinical procedures in which personality descriptions from tests are made to fit the patient largely or wholly by virtue of their triviality; and in which any nontrivial, but perhaps erroneous, inferences are hidden in a context of assertions or denials which carry high confidence simply because of the population base rates, regardless of the test'sTitle page of Forer 1949: The Fallacy of Personal Validation: A Classroom Demonstration of Gullibility
072
Mark Dingemanse @markdingemanse.net · 13/07/2026
One of my contributions as a reviewer of this paper was to nudge the authors to look beyond minimalism (M) to explain the #ideophone sequences they argue M has little to say about. If one's goal truly is to minimize syntactic machinery, better be aware of —and point readers to— other approaches ☺️
Finally, if, as we have said, the sequences under consideration have a flat syntax, then it may be necessary to engage with a broader conception of grammatical structure, either within the generative framework or from compatible alternative frameworks. The former would include consideration of grammars that are not context-free or context-sensitive (e.g., regular or subregular grammars), while the latter would include work that considers indication and depiction alongside the well-studied mode of description (Ferrara & Hodge 2018), and work from typology and grammaticalization that shows why ideophones come to occupy the grammatical positions they do (Güldemann 2008). In both cases, there is a truly minimalist goal to achieve: to account for the observed structure by importing just the appropriate amount of syntactic machinery.
051
Mark Dingemanse @markdingemanse.net · 13/07/2026
Richting het gouden uur bij De Kaaij optreden met @kladderadatsch.bsky.social, wat wil je nog meer? 🤗😎🥁
foto met ernstig tegenlicht van de laagstaande zon laat een deel van Kladderadatsch zien tegenover een mensenmassa die mooi van achteren uitgelicht wordt
020
Mark Dingemanse @markdingemanse.net · 13/07/2026
Onze uitwisseling vond plaats in de tijd dat ik ook een stuk aan het schrijven was over deze thema's: Goed onderwijs bouwt op waarden en demystificeert technologie. Lees het hier: doi.org/10.5281/zeno... (Verschijnt na de zomer in "Als wij denken", red Dastani/Kleemans/Stronks) #GenAI #onderwijs
Figuur 1. Niets nieuws onder de zon: “AI” sluit aan in een lange rij van technologieën met revolutionaire pretenties, maar goed onderwijs staat of valt met de kennis en kunde van leerkrachten, verankerd in gedeelde waarden. (Illustratie: Ada Jušić & Eleonora Lima (KCL) via betterimagesofai.org, CC -BY 4.0)Titelpagina: Goed onderwijs bouwt op waarden en demystificeert technologie
000
Mark Dingemanse @markdingemanse.net · 13/07/2026
Sam de Vlieger interviewde me voor Montessori Magazine over "AI" chatbots magazine.montessori.nl/markdingeman... Waarin we praten over Maria Montessori, de echoput van ChatGPT, en het werk van Ivan Illich
Montessori schrijft dat haar materialen een emancipatorische werking kunnen hebben, dat in regulier onderwijs een hiërarchische relatie tussen leraar en kind zelfstandig leren in de weg zit. Als de leraar het kind teveel helpt kan het minder gaan vertrouwen op de eigen intelligentie. In jouw artikel Human creativity in the age of text generators schrijf je dat chatbots laten denken dat ze een denkvriend (thinking buddy) zijn. Kun je meer zeggen over deze interactie met een chatbot?

‘Maria Montessori begreep dat leerkrachten soms het leerproces in de weg kunnen zitten, maar legde ook de nadruk op het belang van samen observeren en experimenteren. Het grappige is dat de introductie van een chatbot ons eerder terugbrengt bij wat Montessori beschreef als “de oude methode, waarin ze het kind overspoelen met nutteloze en vaak incorrecte woorden” – mijn vertaling. Om het leerproces zoveel mogelijk te faciliteren, hebben we leerkrachten nodig die het kind écht zien, en materialen die ontdekking op eigen kracht mogelijk maken. Chatbots bieden geen van beide.’
110
Mark Dingemanse @markdingemanse.net · 13/07/2026
E.g. this minimally changed version also has upward branches on the diagonals that "know" to stop, somehow "obeying" horizontal lines logolang.org#store=eJx1kM... (note that the tree you saw requires a different logic: trunk separate from twigs since twigs only branch upwards)
screenshot of fern generated by a simple recursive algorithmthe algorithm: 
PU BK 200 LT 90 FD 100 RT 90 PD
TO FERN :SIZE :SIGN
    if :SIZE > 1 [
        FD :SIZE 
        RT 70 * :SIGN FERN :SIZE * 0.5 :SIGN * -1 LT 70 * :SIGN
        FD :SIZE
        LT 70 * :SIGN fern :SIZE * 0.5 :SIGN RT 70 * :SIGN
        RT 0 * :SIGN fern :SIZE - 1 :SIGN LT 0 * :SIGN
        BK :SIZE * 2
    ]
END
FERN 24 1
042
Mark Dingemanse @markdingemanse.net · 08/07/2026
Over the years, my posts on Akpafu history & Siwu language facts tend to attract scattered replies by Siwu folks, which is always fun. Here's one that showcases some ways to work around the lack of easily findable unicode characters for the Siwu characters ɛ ('3') and ɔ (')')
I fe me )mere ke k) lo nya so a kyer) I siwu am3. Mm)n ngb) ta lo bra I un diary (leye n3 siwu iyere 😅)am3. Tela du3 n3 si boa wo deka Ido ng3.

ife mɛmɛre ke kɔ lonya sɔ a tsɛrɛ i siwu amɛ. mmɛ ngbɔ ta lo'bra i un diary (leiye nɛ siwu iyere 😅) amɛ. Tela duɛ nɛ si boawo deka ido ngɛ

"It makes me happy to see you writing in Siwu. As for me I keep a diary (I don't know the Siwu name). I'd love for us to meet some time."
130