Sign in

Marianne de Heer Kloots

@mdhk.net
1.4K followers 585 following 147 posts

Linguist in AI & CogSci 🧠👩‍💻🤖 PhD student @illc-uva.bsky.social 🌐 mdhk.net 🐘 scholar.social/@mdhk 🐦 twitter.com/mariannedhk

PostsRepliesMedia
Marianne de Heer Kloots @mdhk.net · 04/07/2026
Busy being awesome!
030
Marianne de Heer Kloots @mdhk.net · 03/07/2026
Now out in BBS, as commentary on @futrell.bsky.social & @kmahowald.bsky.social's "How linguistics learned to stop worrying and love the language models"! Humans learn much of spoken language structure from speech (not text), & we can study models that do the same. www.cambridge.org/core/journal...
Cover page of our commentary.

Title: Linguists should learn to love speech-based deep learning models
Authors: Marianne de Heer Kloots, Paul Boersma, Willem Zuidema

Abstract: Futrell and Mahowald present a useful framework bridging technology-oriented deep learning systems and explanation-oriented linguistic theories. Unfortunately, the target article's focus on generative text-based Large Language Models (LLMs) fundamentally limits fruitful interactions with linguistics, as many interesting questions on human language fall outside what is captured by written text. We argue that audio-based deep learning models can and should play a crucial role.
1286
Marianne de Heer Kloots @mdhk.net · 30/06/2026
💡 The relative order of learning curves is generally consistent between different seeds of the same architecture, and follows a similar pattern across architectures: acoustics first, syntax last. HuBERT's 2nd it. models shows greater parallelism, echoing observed differences in layerwise patterns.
Sigmoid curves fitted to the best-layer scores for all six models (two seeds for each architecture). Different levels of linguistic structure consistently show distinct learning dynamics across architectures and model seeds, with increased parallelism between levels for HuBERT's second iteration as compared to HuBERT's first iteration and Wav2Vec2 models.
110
Marianne de Heer Kloots @mdhk.net · 30/06/2026
💡 Across model training, layerwise patterns are relatively stable, emerging between 10k-50k steps and showing no major shifts afterwards. Linguistic probe results for the speech-trained model start outperforming those of a non-speech baseline around 10k training steps.
Figure displaying learning trajectory results (i.e. the evolution of probe scores across training checkpoints) and a single model (Wav2Vec2, seed 2). The left column shows the best-layer scores across training steps, and the right column shows heatmaps of all layerwise scores by training steps.
100
Marianne de Heer Kloots @mdhk.net · 30/06/2026
⚙️ Across layers we observe a sequential pattern of peaks for acoustic, phonetic, syllabic, and lexical/syntactic structure. HuBERT's 2nd iteration diverges from Wav2Vec2 and HuBERT it. 1, as found in other work on iterative refinement & layerwise organization (www.isca-archive.org/interspeech_...).
Figure displaying layerwise probe results for 3 models of different architectures (Wav2Vec2, HuBERT's first iteration, HuBERT's second iteration).
100
Marianne de Heer Kloots @mdhk.net · 30/06/2026
Our probes target linguistic structuring in model representation (sub)spaces, primarily by clustering & distance metrics. For example, is there a subspace encoding distinctions between syllable types (consonant-vowel patterns)? Here's a visualization over training checkpoints of a Wav2Vec2 model.
100
Marianne de Heer Kloots @mdhk.net · 30/06/2026
What about self-supervised Transformer models learning from audio recordings of speech? To find out, we trained 6 new Wav2Vec2 and HuBERT models on 831 hours of spoken Dutch, and probed for 9 levels of linguistic structure across layers and intermediate checkpoints (up to 100k training steps). 🔍
Figure illustrating the study's methods: We train a set of six models on the same dataset, consisting of 831 hours of Dutch speech recordings (sourced from the Spoken Dutch Corpus, Multilingual LibriSpeech, and CommonVoice). We explore how results vary between model architectures with minimal differences in training set-up (Wav2Vec2, HuBERT's first iteration, HuBERT's second iteration). We probe each model's internal representations for nine types of linguistic structure, which differ in their degrees of abstraction from the acoustic signal and in their timescales of information integration. We compare results for each structure across model layers as well as training steps.
100
Marianne de Heer Kloots @mdhk.net · 30/06/2026
Let’s study learning trajectories in self-supervised speech models! 🔊 Do they reflect the hierarchical organization of spoken language? We have analyzed a lot of training checkpoints to find out 🌠 Preprint: arxiv.org/abs/2604.02043 ⬇️
12811
Marianne de Heer Kloots @mdhk.net · 19/12/2025
Cool posters from day 2! @sashakenjeeva.bsky.social openreview.net/forum?id=Vtd... github.com/markvandenho... openreview.net/forum?id=rX3... @nina-nusbaumer.bsky.social openreview.net/forum?id=GRz... www.ru.nl/personen/sui... openreview.net/forum?id=NcJ...
Poster title: Does multimodal pre-activation influence linguistic expectations in LLMs and humans?

Authors: Sasha Kenjeeva, Giovanni Cassani, Noortje Venhuizen, Afra AlishahiPoster title: Generalizing Without Evidence: How Transformer Models Infer Syntactic Rules From Sparse Input

Authors: Mark van den Hoorn, Raquel G. AlhamaPoster title: Dependency Length, Syntactic Complexity & Memory: A Reading Time Benchmark for Sentence Processing Modeling

Authors: Nina Nusbaumer, Corentin Bel, Iria de-Dios-Flores, Guillaume Wisniewski, Benoit CrabbéPoster title: 
The success of Neural Language Models on syntactic island effects is not universal: strong wh-island sensitivity in English but not in Dutch

Authors: Michelle Suijkerbuijk, Naomi Tachikawa Shapiro, Peter de Swart, Stefan L. Frank
141
Marianne de Heer Kloots @mdhk.net · 19/12/2025
Today we are time-travelling with @stefanfrank.bsky.social 😎
120
Marianne de Heer Kloots @mdhk.net · 17/12/2025
'Tis the season to preprint BBS commentaries; I'm happy to share ours too! 🎄✨ The textual basis of current LLMs causes trouble, but linguistically relevant insights *can* be found in systems modelling the more natural form of human spoken language: the speech signal itself. arxiv.org/abs/2512.14506
Commentary title: 
Linguists should learn to love speech-based deep learning models 

Authors: 
Marianne de Heer Kloots, Paul Boersma, Willem Zuidema

Abstract: 
Futrell and Mahowald present a useful framework bridging technology-oriented deep learning systems and explanation-oriented linguistic theories. Unfortunately, the target article's focus on generative text-based LLMs fundamentally limits fruitful interactions with linguistics, as many interesting questions on human language fall outside what is captured by written text. We argue that audio-based deep learning models can and should play a crucial role.
12810
Marianne de Heer Kloots @mdhk.net · 27/08/2025
Finally, downstream performance on Dutch speech-to-text transcription reflects the language-specific advantage for Dutch linguistic feature encoding in model-internal representations: on average, Wav2Vec2-NL has a 27% lower word error rate than the multilingual model.
Word Error Rate results for models fine-tuned for Dutch ASR (speech-to-text transcription), across 4 models and 5 evaluation datasets.
100
Marianne de Heer Kloots @mdhk.net · 27/08/2025
We find that language-specific advantages are well-detected by trained clustering or classification probes, and partially observable using zero-shot metrics. I.e. the encoding of Dutch linguistic features is enhanced in the Dutch model, as compared to models trained on English and multilingual data.
Layerwise phonetic and lexical analyses, across a read speech (MLS, top row) and a dialogue (IFADV, bottom row) dataset of spoken Dutch. Measures marked * involve optimized linear transforms, whereas others are computed zero-shot; shading indicates 95% confidence intervals. The Dutch Wav2Vec2-NL model achieves highest scores across most analyses of Dutch phone and word encoding, though the size of this language-specific advantage varies considerably across analyses.
100
Marianne de Heer Kloots @mdhk.net · 27/08/2025
But they also used different analysis techniques. We designed the SSL-NL dataset to test the encoding of Dutch phonetic and lexical features in SSL speech representations, while allowing for comparisons across different analysis methods. We compare both trained probes(*) and zero-shot metrics:
The model comparison set includes Wav2Vec2-NL and 3 other existing Wav2Vec2-base models: facebook's multilingual voxpopuli model, facebook's English base model, and another model trained on nonspeech acoustics. 

The set of analysis techniques includes probing classifiers (logistic regression), ABX similarities, PCA clustering, LDA clustering, and representational similarity analysis (RSA).

Word- and phone-level embeddings were created by mean-pooling model frame embeddings within words and phones respectively.

The SSL-NL evaluation dataset is a curated dataset of Dutch speech recordings and accompanying forced alignments, across two domains: audiobooks (MLS) and face-to-face conversations (IFADV).
100
Marianne de Heer Kloots @mdhk.net · 27/08/2025
✨ Do self-supervised speech models learn to encode language-specific linguistic features from their training data, or only more language-general acoustic correlates? At #Interspeech2025 we presented our new Wav2Vec2-NL model and SSL-NL evaluation dataset to test this! 📄 arxiv.org/abs/2506.00981 ⬇️
Interspeech paper title: What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training

Authors: Marianne de Heer Kloots, Hosein Mohebbi, Charlotte Pouw, Gaofei Shen, Willem Zuidema, Martijn Bentum
1196
Marianne de Heer Kloots @mdhk.net · 12/08/2025
Last but not least, I personally can’t wait for the social event on Thursday night that we’ve been planning for the past year ✨ It features a *live brain-controlled music act* by the AIAR collective 🧠🎶 2025.ccneuro.org/social-event/ Get one of the last remaining tickets at the registration desk now!
031
Marianne de Heer Kloots @mdhk.net · 12/08/2025
So exciting, #CCN2025 in Amsterdam started today! We have stroopwafels!! Catch me at my poster on Friday to chat about the role of context in neural representational alignment to spoken language systems (C34) 🙌

 🔗 2025.ccneuro.org/poster/?id=K...
170
Marianne de Heer Kloots @mdhk.net · 28/07/2025
We are having an impromptu overflow room around the corner 😅
130
Marianne de Heer Kloots @mdhk.net · 24/07/2025
Next week I’ll be in Vienna for my first *ACL conference! 🇦🇹✨ I will present our new BLiMP-NL dataset for evaluating language models on Dutch syntactic minimal pairs and human acceptability judgments ⬇️ 🗓️ Tuesday, July 29th, 16:00-17:30, Hall X4 / X5 (Austria Center Vienna)
The BLiMP-NL dataset consists of 84 Dutch minimal pair paradigms covering 22 syntactic phenomena, and comes with graded human acceptability ratings & self-paced reading times. 

An example minimal pair:
A. Ik bekijk de foto van mezelf in de kamer (I watch the photograph of myself in the room; grammatical)
B. Wij bekijken de foto van mezelf in de kamer (We watch the photograph of myself in the room; ungrammatical)

Differences in human acceptability ratings between sentences correlate with differences in model syntactic log-odds ratio scores.
2284
Marianne de Heer Kloots @mdhk.net · 29/12/2024
If you see this, post a concert picture you took this year.
040
Marianne de Heer Kloots @mdhk.net · 04/12/2024
red square against budget cuts in higher education
041
Marianne de Heer Kloots @mdhk.net · 19/11/2024
In a bizarre undemocratic turn of events, the massive national protest against our government's plans for higher education was cancelled last week. We'll be back stronger next Monday in The Hague! 🟥
poster announcing the protest on November 25th, 1pm, Malieveld, The Hague
131
Marianne de Heer Kloots @mdhk.net · 29/10/2024
Or this from within The Netherlands: campagnes.degoedezaak.org/campaigns/st... and come to Utrecht on Nov 14th! www.fnv.nl/cao-sector/o...
010
Marianne de Heer Kloots @mdhk.net · 18/09/2024
Bluesky now has over 10 million users, and I was #497,227! 😎
000
Marianne de Heer Kloots @mdhk.net · 04/09/2024
I will be presenting this work tomorrow (Thursday) at #INTERSPEECH2024, 10.00-10.40 in the Acesso room! Looking forward to discuss how we can learn from human speech science to interpret end-to-end neural speech models 💡 The paper is here: www.isca-archive.org/interspeech_...
paper title: Human-like Linguistic Biases in Neural Speech Models: Phonetic Categorization and Phonotactic Constraints in Wav2Vec2.0
000
Marianne de Heer Kloots @mdhk.net · 23/07/2024
It turns out the accuracy of dependency structures decoded from LM hidden layers (measured by Labelled Attachment Score) strongly correlates with similarity to brain activity in sentence reading! 🧠 This correlation disappears in a control condition with scrambled inputs.
Figure with two scatterplots, showing a strong correlation between dependency accuracy (Labelled Attachment Score) and brain alignment (Representational Similarity Score) on the left, and no correlation in a scrambled control condition on the right.
100
Marianne de Heer Kloots @mdhk.net · 23/07/2024
Language model internal states show surprising similarity to human brain activity in language comprehension — but how does this relate to their accurate representation of structured linguistic information, like syntactic dependencies? (i.e. links between words in a sentence)
Dependency parse for the sentence "De overtreder die de smeris ontvlucht was is een kronkelig paadje ingerend" (Dutch for: "The offender who had escaped from the cop ran into a winding path"). 
The picture shows all dependency links between the words of the sentence, such as the links between verbs and their subjects (Subj) and nouns and their determiners (Det).
130
Marianne de Heer Kloots @mdhk.net · 23/07/2024
Excited for #CogSci2024 this week! In session T.24 on Friday morning (10.30-12), Bram will present our work on representational alignment between LMs, brains, and syntactic structure 🤖🧠💬 w/ Rochelle Choenni, @mheilbron.bsky.social & @wzuidema.bsky.social 📑 escholarship.org/uc/item/1fp7... ⬇️
Paper title ('Language Models That Accurately Represent Syntactic Structure Exhibit Higher Representational Similarity To Brain Activity') and overview figure
151
Marianne de Heer Kloots @mdhk.net · 08/07/2024
📏 We also compare three analysis methods for decoding phoneme preference from model internals, and find interesting differences between them! ➡️ Read more in the paper: arxiv.org/abs/2407.03005
110
Marianne de Heer Kloots @mdhk.net · 08/07/2024
💡 We find similar adaptation to phonotactic context in Wav2Vec2 models, emerging around the 4th layer of their Transformer module. This effect is amplified by finetuning for text transcription, but also present in fully self-supervised models (when trained on English speech).
110
Marianne de Heer Kloots @mdhk.net · 08/07/2024
One case of such contextual biasing effects comes from phonotactic constraints. For example in English: TL << TR, SL >> SR This has been demonstrated in human listeners a while ago! (doi.org/10.3758/BF03...)
100
Marianne de Heer Kloots @mdhk.net · 13/06/2024
Feeling very inspired about ✨Using ANNs for Studying Human Language Learning and Processing (ann-humlang.github.io )✨ after the workshop that Tamar Johnson and I organized this week at the ILLC in Amsterdam! Many thanks to all our speakers and participants for such a great event,
120
Marianne de Heer Kloots @mdhk.net · 30/11/2023
A nice session at KNAW tonight looking back on the year since the launch of ChatGPT — Katia is giving a short technical glimpse behind the curtains (🥁) of LLMs right now, that I made some illustrations for! Livestream: www.youtube.com/live/Nn41XWA...
020
Marianne de Heer Kloots @mdhk.net · 22/10/2023
Finally, there's some useful settings you can tune to make things better on your home feed as well! I currently have this in Home Feed and Thread Preferences settings (forgot which ones are different from the defaults)
000
Marianne de Heer Kloots @mdhk.net · 22/08/2023
I'm a bird!
110