Sign in

Marianne de Heer Kloots

@mdhk.net
1.4K followers 583 following 147 posts

Linguist in AI & CogSci 🧠👩‍💻🤖 PhD student @illc-uva.bsky.social 🌐 mdhk.net 🐘 scholar.social/@mdhk 🐦 twitter.com/mariannedhk

PostsRepliesMedia
Marianne de Heer Kloots @mdhk.net · 22h
I’m presenting this today at #SNL2026 in Geneva! See you at poster A5 in the morning session 🎉
091
Reposted by Marianne de Heer Kloots
Laura Gwilliams @lauragwilliams.bsky.social · 28/09/2026
the Gwilliams Lab and friends are at SNL 2026!! if you're into speech comprehension, intracranial and non-invasive time series, from single neurons to the whole cortex, across ages, using both classic linguistic approaches and advanced ML techniques - we've got something for you! 🧠 🌀 ✨
0187
Reposted by Marianne de Heer Kloots
Simon Fisher @profsimonfisher.bsky.social · 25/09/2026
To make sense of speech, listeners must segment sound streams into pieces. Across languages, people do so based more on consonants than vowels. A new @science.org study investigates the phenomenon via EEG recording in humans & dogs, finding that dogs also show consonant bias in word segmentation.🗣️🐶🧪
science.org
Neural evidence that dogs segment the speech they hear with a humanlike consonant bias
Across many human languages, consonants carry more lexical information than vowels. During speech segmentation, humans, unlike nonhuman primates, rely more on consonant than vowel patterns, despite vo...
23017
Reposted by Marianne de Heer Kloots
Dirk Gütlin @gutlin.bsky.social · 25/09/2026
📣📣📣 Preprint Alert! 📣📣📣 Do predictive processes shape brain representations during learning? We investigated the effect of Predictive Coding by training identical recurrent networks with different optimization procedures and then comparing them to human EEG under the same visual learning task.
Overview figure of our learning paradigm, architecture, and analysis procedure.
15113
Reposted by Marianne de Heer Kloots
Vlad Ayzenberg @vayzenb.bsky.social · 22/07/2026
Excited to share our review in @cp-neuron.bsky.social with @lauriebayet.bsky.social and @mickbonner.bsky.social! We describe how implementing principles from child development can advance the mechanistic plausibility and capacities of AI models We packed A LOT into this review, here's a quick 🧵
16829
Reposted by Marianne de Heer Kloots
Jane Li 🦖 @janeli.bsky.social · 17/09/2026
🦀New preprint! (w/ @najoung.bsky.social)🦞 Is grammaticality a major organizing principle of NLM representations? We show that many NLMs exhibit abstract rep. separation for grammaticality. We believe this work addresses debates about confounds in measuring model gram. knowledge. [1/10]
12010
Marianne de Heer Kloots @mdhk.net · 17/09/2026
LLMs are everywhere, as are companies advertising LLM-powered products by their “beyond-human” performance levels. But how do LLMs *actually* fare against humans across the board? Join our effort to organize empirical evidence from the field into a systematic and publicly accessible overview ⬇️
0102
Reposted by Marianne de Heer Kloots
Lukas Edman @lukasnlp.bsky.social · 17/09/2026
Ever feel like it's too hard to keep track of what LLMs cannot do as well as humans? We're making your life easier over at: what-llms-can-not-do.github.io We're compiling a list of papers testing the abilities of LLMs against humans. Check it out! And you can help contribute too!
what-llms-can-not-do.github.io
What LLMs Can(not) Do
A living survey of benchmarks that compare large language models with humans.
24516
Reposted by Marianne de Heer Kloots
The Transmitter @thetransmitter.bsky.social · 07/09/2026
I fear that placing too much emphasis on a specific interpretation of dimensionality, or treating dimensionality as an end-all quantification of some aspect of neural computation, may lead us down the wrong path, writes @mattperich.bsky.social. #neuroskyence www.thetransmitter.org/neural-dynam...
thetransmitter.org
Dimensionality—neuroscience’s red herring?
Placing too much emphasis on a specific interpretation of dimensionality may lead neuroscience down the wrong path.
07421
Marianne de Heer Kloots @mdhk.net · 14/09/2026
We meet monthly at 3pm Amsterdam time; anyone interested in the topic is welcome to join. Request access to the reading group mailing list for announcements & zoom links here: groups.google.com/g/dnn-speech/ (please include a short description of who you are if not obvious from your e-mail address)
010
Marianne de Heer Kloots @mdhk.net · 14/09/2026
The reading group I've been organizing is entering its 4th academic year of existence! We'll have our first meeting of this year this Thursday, with @gretatuckute.bsky.social and @klemenkotar.bsky.social presenting recent work on their AuriStream model 🧠🔊 dnn-speech.github.io
dnn-speech.github.io
(D)NN-Speech Reading Group
We are a reading group meeting regularly to discuss papers on speech-based deep learning models and their use in modelling human speech processing and acquisition!
1131
Reposted by Marianne de Heer Kloots
Lisa Bylinina @bylinina.bsky.social · 11/09/2026
look at my new babylm paper! arxiv.org/abs/2609.11870 basically, i initialize token embeddings with representations from an image encoder rather than randomly and then train text-only as usual. kind of a visual demonstration to start off the word learning process
1155
Reposted by Marianne de Heer Kloots
Terence Tao @teorth.bsky.social · 11/09/2026
A group of 25 Fields Medalists, including myself, have made a joint declaration on Math and AI: mathandai.org . We welcome additional signatories. See also this article in the Economist announcing the declaration: www.economist.com/science-and-...
mathandai.org
Declaration — Math and AI
Read the declaration and add your name.
422052926
Marianne de Heer Kloots @mdhk.net · 11/09/2026
I’m presenting this today at #CLIN36 in Brussels! See you at poster 25 in the afternoon session 🎉
071
Marianne de Heer Kloots @mdhk.net · 10/09/2026
Yeah I’ve just noticed that voice resynthesis is starting to get used instead of earlier distortion techniques for this purpose. The resulting signal (to me) requires less effort to understand, which I guess for producers is an advantage when including longer recordings of this kind.
010
Marianne de Heer Kloots @mdhk.net · 10/09/2026
Anonymization for protecting people’s identity in audiovisual media? Like in documentaries/interviews on sensitive topics
100
Reposted by Marianne de Heer Kloots
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
🤖🧠NEW PAPER🧠🤖 (The result of an 8-year project!) LLMs seem very different from symbolic systems. Yet LLMs excel in symbolic domains (e.g., language/code/math). How do they do it? Our finding: LLM representations have implicit symbolic structure! Link in thread ⬇️ 1/n
Overview of the paper. 
Title: The Emergent Symbolic Structure of Artificial Neural Networks
Authors: Tom McCoy, Paul Soulos, Tal Linzen, Paul Smolensky
Left: Neural networks encode information in vectors (there is then an image of a vector), yet they excel at tasks long thought to require symbolic structure (there is then an image of a symbolic representation, specifically a syntax tree). How do LLMs do it?
Right: We find that LLM representations can be closely approximated with symbolic structures. This approximation lets us edit the structure of an LLM’s output by editing the structure of its internal representations, as shown. There is then an image of two edits to LLMs. In the first one, the original input is 3 + 6 * 8, with an answer of 51. But if we swap the positions of the 3 and the 6, the output becomes 30. In the second one, the original input is a Python command repeating the list [Z, U] three times, producing [Z, U, Z, U, Z, U]. But if we edit the input in a way that adds a Q at the end of the input, the output becomes [Z, U, Q, Z, U, Q, Z, U, Q].
431788
Reposted by Marianne de Heer Kloots
Victoria Bosch @initself.bsky.social · 28/08/2026
Are brains and artificial neural networks converging onto universal representations? There is a seductive idea making the rounds in NeuroAI / machine learning: train systems well enough, and they all converge on the same representation of reality (i.e. a unique world model). We have thoughts™ 1/n
cell.com
The Umwelt Representation Hypothesis: rethinking Universality
Recent studies reveal striking representational alignment between artificial neural networks (ANNs) and biological brains, leading to proposals that all sufficiently capable systems converge on univer...
818176
Reposted by Marianne de Heer Kloots
Grzegorz Chrupała @grzegorz.chrupala.me · 28/08/2026
New preprint 🐝🧵 arxiv.org/abs/2608.25779 Bee species signal food via their waggle dance in different ways. Species with horizontal combs point straight at food. Species with vertical combs dance relative to gravity instead. How could a population evolve from one system to the other?
Diagram comparing two honeybee waggle-dance styles. Left, "Horizontal comb": a bee on a flat, round comb dances in a straight line pointing directly at a flower, with the sun shown off to one side and the angle α marked between the sun's direction and the dance direction. Right, "Vertical comb": a bee on an upright square comb dances at an angle from a vertical gravity line; a separate arrow labeled "up = toward sun" shows that upward now stands in for the sun's direction, with the dance angle α measured from vertical instead of from the sun directly.
1102
Marianne de Heer Kloots @mdhk.net · 26/08/2026
I’ve wondered the same but the other way around! This seems a really weird way to use “what”, I don’t understand why native English speakers insist on doing that!
220
Reposted by Marianne de Heer Kloots
Sam Nastase @samnastase.bsky.social · 18/08/2026
New perspective piece out in @cp-neuron.bsky.social with Zaid Zada, @adelegoldberg.bsky.social, and Uri Hasson! We try to articulate some of our excitement about LLMs and discuss what kinds of insights they might provide into the neural computations supporting natural language in the human brain.
18625
Reposted by Marianne de Heer Kloots
Daniel Lakens @lakens.bsky.social · 15/08/2026
New Blog post: Which Data Repository Should you Use? In light of OSF closing down, I compare Zenodo, Dataverse, ResearchBox, PsychArchive, and local repositories on six important dimensions. If you want to know which to pick: It depends! daniellakens.blogspot.com/2026/08/whic...
daniellakens.blogspot.com
Which Data Repository Should You Use?
The Center for Open Science has announced that from November 16, 2026, no new projects can be created on the Open Science Framework. After F...
4197111
Reposted by Marianne de Heer Kloots
Naomi Saphra @nsaphra.bsky.social · 09/08/2026
I've been unsettled lately when reading messages and papers. It feels like I'm dissociating. Everything seems a bit alien, even if it's completely human. I've had a realization: When our simulations finally exited the Uncanny Valley, they brought the Uncanny with them.
nsaphra.net
Life on the Uncanny Precipice | Naomi Saphra
We were wrong about the Uncanny Valley.
1026362
Reposted by Marianne de Heer Kloots
Lisa DeBruine @debruine.bsky.social · 04/08/2026
Today is the only day you can experience Ray Bradbury’s “There Will Come Soft Rains” on the day it is set. Audio: archive.org/details/brad... PDF: thephilosopher.net/bredberi/wp-...
16431350
Reposted by Marianne de Heer Kloots
Marianne de Heer Kloots @mdhk.net · 30/06/2026
Let’s study learning trajectories in self-supervised speech models! 🔊 Do they reflect the hierarchical organization of spoken language? We have analyzed a lot of training checkpoints to find out 🌠 Preprint: arxiv.org/abs/2604.02043 ⬇️
12811
Reposted by Marianne de Heer Kloots
Dota Tianai Dong @dotadotadota.bsky.social · 21/07/2026
1/5 Over a decade of comparing deep neural networks to the human brain—but what have we actually learned? Our new @cp-trendscognsci.bsky.social Feature Review synthesizes a decade of brain–DNN comparisons, asking what they reveal about brain function across vision and language.
14421
Reposted by Marianne de Heer Kloots
Simon Fisher @profsimonfisher.bsky.social · 16/07/2026
Our ability to speak involves remarkable feats of motor sequencing that most of us take completely for granted. The new issue of @science.org has a fascinating essay by Sergey Stavisky, covering his team's use of brain-computer interfaces to restore speech for people with neurological injuries. 🧠🗣️🧪👇
science.org
Regaining your voice
AI speech neuroprostheses can restore day-to-day communication after neurological injury
2296
Reposted by Marianne de Heer Kloots
Shuqi Wang @shuqiw.bsky.social · 15/07/2026
Our paper "Rarely categorical, highly separable representations along the cortical hierarchy" is now out in Nature! www.nature.com/articles/s41... In this paper, we studied whether the brain is well-organized and interpretable, both globally and locally. (See thread below.)
nature.com
1228
Reposted by Marianne de Heer Kloots
Jennifer Hu @jennhu.bsky.social · 02/07/2026
What's more nonsensical: smashing a pumpkin using a number, or growing flowers inside a sneeze? Our paper on graded inconceivability is out now in Cognition! Come for the cognitive science 🧠🔍, stay for the whimsy 🌼🧚! 🔗Journal link: bit.ly/gradedInconCog
1416
Marianne de Heer Kloots @mdhk.net · 09/07/2026
120
Marianne de Heer Kloots @mdhk.net · 04/07/2026
Busy being awesome!
030
Marianne de Heer Kloots @mdhk.net · 03/07/2026
130
Marianne de Heer Kloots @mdhk.net · 03/07/2026
120
Marianne de Heer Kloots @mdhk.net · 03/07/2026
120
Marianne de Heer Kloots @mdhk.net · 03/07/2026
140
Marianne de Heer Kloots @mdhk.net · 03/07/2026
(I’m going to attach some of my favourite other commentaries below)
130
Marianne de Heer Kloots @mdhk.net · 03/07/2026
I can highly recommend the full treatment, incl. the target article and all other commentaries, for a great set of current perspectives on everything Linguistics ∩ LLMs!
130
Marianne de Heer Kloots @mdhk.net · 03/07/2026
^ I hope the link above gives access to everyone, otherwise find the preprint here ⬇️
120
Marianne de Heer Kloots @mdhk.net · 03/07/2026
Now out in BBS, as commentary on @futrell.bsky.social & @kmahowald.bsky.social's "How linguistics learned to stop worrying and love the language models"! Humans learn much of spoken language structure from speech (not text), & we can study models that do the same. www.cambridge.org/core/journal...
Cover page of our commentary.

Title: Linguists should learn to love speech-based deep learning models
Authors: Marianne de Heer Kloots, Paul Boersma, Willem Zuidema

Abstract: Futrell and Mahowald present a useful framework bridging technology-oriented deep learning systems and explanation-oriented linguistic theories. Unfortunately, the target article's focus on generative text-based Large Language Models (LLMs) fundamentally limits fruitful interactions with linguistics, as many interesting questions on human language fall outside what is captured by written text. We argue that audio-based deep learning models can and should play a crucial role.
1286
Reposted by Marianne de Heer Kloots
Kyle Mahowald @kmahowald.bsky.social · 02/07/2026
The full BBS treatment from me and @futrell.bsky.social on "How linguistics learned to stop worrying and love the LMs" is now out, with all the commentaries and our response. If you "Save PDF", it will give you the whole target article + commentary + response pdf: www.cambridge.org/core/journal...
cambridge.org
How linguistics learned to stop worrying and love the language models | Behavioral and Brain Sciences | Cambridge Core
How linguistics learned to stop worrying and love the language models - Volume 49
2359
Reposted by Marianne de Heer Kloots
Auditory-Visual Speech Association (AVISA) @avsp.bsky.social · 01/07/2026
Linguists should learn to love speech-based deep learning models www.cambridge.org/core/journal... "Once the bottleneck of text can be replaced...modelling more human-like linguistic processes..[& handle] the vast majority of local & regional language varieties that are rarely or never written down"
cambridge.org
142
Marianne de Heer Kloots @mdhk.net · 01/07/2026
Thanks for sharing 😄 I have this author link, hopefully that should make it readable for everyone! www.cambridge.org/core/journal...
cambridge.org
Linguists should learn to love speech-based deep learning models | Behavioral and Brain Sciences | Cambridge Core
Linguists should learn to love speech-based deep learning models - Volume 49
010
Marianne de Heer Kloots @mdhk.net · 01/07/2026
All related commentaries are out now! www.cambridge.org/core/journal...
cambridge.org
How linguistics learned to stop worrying and love the language models | Behavioral and Brain Sciences | Cambridge Core
How linguistics learned to stop worrying and love the language models - Volume 49
280
Marianne de Heer Kloots @mdhk.net · 30/06/2026
(I’m sorry to miss out on attending this great new workshop myself! Many thanks to the CDL reviewers for their helpful comments, and to @bbunzeck.bsky.social & @dnnslmr.bsky.social for their on-site poster help 🙌)
020
Marianne de Heer Kloots @mdhk.net · 30/06/2026
Our full paper is still under review, but if you're attending ACL in San Diego 🌊 🌴 🦭 this week, there'll be a non-archival poster on this project at the Computational Developmental Linguistics workshop's poster session: 📍 Sat July 4th, 4.30-6pm PDT 🔗 comp-dev-ling.github.io #ACL2026NLP #NLProc
100
Marianne de Heer Kloots @mdhk.net · 30/06/2026
➡️ Find more analyses & details on all of the above in our preprint! arxiv.org/abs/2604.02043 🙏 to all co-authors (Martijn Bentum @hmohebbi.bsky.social @cpouw.bsky.social @gaofeishen.com @wzuidema.bsky.social) & to @itcooperativesurf.bsky.social for granting me the resources that enabled this work 👩‍💻
arxiv.org
Tracking the emergence of linguistic structure in self-supervised models learning from speech
Self-supervised speech models learn effective representations of spoken language, which have been shown to reflect various aspects of linguistic structure. But when does such structure emerge in model...
100
Marianne de Heer Kloots @mdhk.net · 30/06/2026
💡 The relative order of learning curves is generally consistent between different seeds of the same architecture, and follows a similar pattern across architectures: acoustics first, syntax last. HuBERT's 2nd it. models shows greater parallelism, echoing observed differences in layerwise patterns.
Sigmoid curves fitted to the best-layer scores for all six models (two seeds for each architecture). Different levels of linguistic structure consistently show distinct learning dynamics across architectures and model seeds, with increased parallelism between levels for HuBERT's second iteration as compared to HuBERT's first iteration and Wav2Vec2 models.
110
Marianne de Heer Kloots @mdhk.net · 30/06/2026
💡 Across model training, layerwise patterns are relatively stable, emerging between 10k-50k steps and showing no major shifts afterwards. Linguistic probe results for the speech-trained model start outperforming those of a non-speech baseline around 10k training steps.
Figure displaying learning trajectory results (i.e. the evolution of probe scores across training checkpoints) and a single model (Wav2Vec2, seed 2). The left column shows the best-layer scores across training steps, and the right column shows heatmaps of all layerwise scores by training steps.
100
Marianne de Heer Kloots @mdhk.net · 30/06/2026
⚙️ Across layers we observe a sequential pattern of peaks for acoustic, phonetic, syllabic, and lexical/syntactic structure. HuBERT's 2nd iteration diverges from Wav2Vec2 and HuBERT it. 1, as found in other work on iterative refinement & layerwise organization (www.isca-archive.org/interspeech_...).
Figure displaying layerwise probe results for 3 models of different architectures (Wav2Vec2, HuBERT's first iteration, HuBERT's second iteration).
100