Sign in

Grzegorz Chrupała

@grzegorz.chrupala.me
3.3K followers 722 following 142 posts

Speech • Language • Learning grzegorz.chrupala.me @ Tilburg University

PostsRepliesMedia
Grzegorz Chrupała @grzegorz.chrupala.me · 17h
I have converged on the following global guidelines for Claude Code. Thanks to them I tear my hair out much less when interacting with the agent of reading its output. github.com/gchrupala/do...
github.com
dotfiles/claude/CLAUDE.md at main · gchrupala/dotfiles
Contribute to gchrupala/dotfiles development by creating an account on GitHub.
171
Grzegorz Chrupała @grzegorz.chrupala.me · 07/10/2026
They don't specify the timeframe
010
Reposted by Grzegorz Chrupała
mr. TIM @timkellogg.me · 06/10/2026
Le Chonk: Mistral drops a 1T beast
Mistral Al @MistralAI
X.com
Meet Mistral Large 4, aka Le Chonk.
• 1T parameters, natively multimodal. 49B active.
It is the best open weights model from US or Europe on aggregated benchmarks.
• State-of-the-art on critical workloads, including cyber defense, manufacturing and finance and it surpasses closed frontier models on visual grounding.
• Forged in Europe end-to-end and is deployable from Europe via our own Mistral Cloud infrastructure.
• Available to all via API today. Working with cybersecurity partners privately.
Open weights release end of October.
913416
Reposted by Grzegorz Chrupała
Tilburg Computational (Psycho) Linguistics @tilburg-clp.bsky.social · 05/10/2026
Oral at BNAIC/BeNeLearn 2026! 🎉 @cangonin.bsky.social, @grzegorz.chrupala.me and @danstowell.bsky.social compare multi-task strategies for learning from several bioacoustics datasets. Single-task models mostly win; multi-task success depends on task weighting. openreview.net/forum?id=ljXGTr9kf5
openreview.net
Comparing Multi-Task Strategies for Multiple Dataset Learning in Bioacoustics
Céline Angonin, Grzegorz Chrupała and Dan Stowell. BNAIC/BeNeLearn 2026.
012
Reposted by Grzegorz Chrupała
Tilburg Computational (Psycho) Linguistics @tilburg-clp.bsky.social · 04/10/2026
Gaofei Shen @gaofeishen.com demonstrates speech technology at the Weekend of Science at MindLabs.
042
Grzegorz Chrupała @grzegorz.chrupala.me · 03/10/2026
Congratulations!
010
Grzegorz Chrupała @grzegorz.chrupala.me · 01/10/2026
alexandreafonso.substack.com/p/immigrants...
alexandreafonso.substack.com
Immigrants: The Silent (Almost) Majority in Dutch Academia
Half of Dutch academics are now foreign, but almost no one talks about them.
000
Grzegorz Chrupała @grzegorz.chrupala.me · 23/09/2026
Mistral is thinking out of the box here
130
Grzegorz Chrupała @grzegorz.chrupala.me · 23/09/2026
Crazy.
111
Grzegorz Chrupała @grzegorz.chrupala.me · 23/09/2026
Hope you're ok, considering.
100
Grzegorz Chrupała @grzegorz.chrupala.me · 22/09/2026
Sorry man 😔
100
Grzegorz Chrupała @grzegorz.chrupala.me · 22/09/2026
For me the Discover tab is indeed useless for this reason. But the regular Following tab only shows me accounts I follow, so I just unfollow people who post too much on US politics.
010
Grzegorz Chrupała @grzegorz.chrupala.me · 21/09/2026
Interspeech 2023 keynote speaker.
010
Reposted by Grzegorz Chrupała
Our World in Data @ourworldindata.org · 18/09/2026
Since the rise of humans, wild land mammal biomass has declined by 85%. Around 100,000 years ago, all of the wild land mammals on Earth summed up to around 20 million tonnes of carbon. By around 10,000 years ago, there had been a huge decline, likely in the range of 25% to 50%.
33520
Reposted by Grzegorz Chrupała
Zach Weinersmith @zachweinersmith.bsky.social · 13/09/2026
Been thinking about this since I read it. If things keep progressing, I worry you arrive at a psychologically horrible point, where one after another, all knowledge workers are forced to argue that their jobs, and the way they've been done traditionally, have value beyond the straightforward output.
1716219
Reposted by Grzegorz Chrupała
Marianne de Heer Kloots @mdhk.net · 11/09/2026
I’m presenting this today at #CLIN36 in Brussels! See you at poster 25 in the afternoon session 🎉
071
Grzegorz Chrupała @grzegorz.chrupala.me · 06/09/2026
The examples are hilariously bad.
010
Grzegorz Chrupała @grzegorz.chrupala.me · 05/09/2026
Paper led by @gaofeishen.com, with Martijn Bentum, Tom Lentz & Afra Alishahi.
000
Grzegorz Chrupała @grzegorz.chrupala.me · 05/09/2026
Case 2: syntax. Lexicon explains more variance than syntax, in BERT and wav2vec2 alike. But ablating syntax raises error even with lexicon in the probe. Syntax isn't just a by-product of lexical correlation.
Two line plots of Unexplained Variance against transformer layer, for wav2vec2 base (left) and BERT base (right). Lines: Full probe, Full minus Syntax, Full minus Lexicon, Full minus both. In both models Full-minus-Lexicon sits above Full-minus-Syntax, and both sit above Full. Gaps are much larger in BERT than in wav2vec2.
100
Grzegorz Chrupała @grzegorz.chrupala.me · 05/09/2026
Phone labels stay decodable well above baseline, yet phonetics accounts for almost none of the variance. Decodability and contribution come apart.
100
Grzegorz Chrupała @grzegorz.chrupala.me · 05/09/2026
Case 1: speech. Speaker identity is decodable from every wav2vec2 variant but explains little variance in the base and ASR models. Only speaker-ID fine-tuning makes it dominant.
Two line plots of Unexplained Variance against transformer layer, for wav2vec2 base (left) and wav2vec2 fine-tuned for speaker ID (right). Lines: Full probe, Full minus Phonetics, Full minus Speaker, Full minus both. In the base model all lines sit between 0.7 and 0.85 with small gaps. In the SID model, from layer 7 upward the Full and Full-minus-Phonetics lines fall to near 0 while Full-minus-Speaker rises to near 1.
100
Grzegorz Chrupała @grzegorz.chrupala.me · 05/09/2026
Decoding scores aren't comparable across features: wav2vec2-base gives 94.5% on speaker ID, 70.3% on phones, which says nothing about relative importance. Correlated features muddy results too. An encoding probe handles these problems natively.
100
Grzegorz Chrupała @grzegorz.chrupala.me · 05/09/2026
Accepted to EMNLP 2026: "Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe". Probing asks what you can decode from a model's activations. We reverse it: reconstruct activations from interpretable features, then ablate to rank them. arxiv.org/abs/2605.00607
arxiv.org
Beyond Decodability: Reconstructing Language Model Representations with an Encoding Probe
Probing is widely used to study which features can be decoded from language model representations. However, the common decoding probe approach has two limitations that we aim to solve with our new enc...
1111
Reposted by Grzegorz Chrupała
Vilém Zouhar @zouhar.bsky.social · 04/09/2026
Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173
arxiv.org
Last Translation Benchmark
For scientific progress, we need benchmarks that test the limits of state-of-the-art models, and evaluation methods that inform us about failure cases. As models get stronger, standard benchmarks for ...
34612
Grzegorz Chrupała @grzegorz.chrupala.me · 28/08/2026
Comments welcome! 🌻☀️🐝
010
Grzegorz Chrupała @grzegorz.chrupala.me · 28/08/2026
Second question: can a population shift from pointing to the gravity code without communication collapsing? Yes, across most of the conditions we tested, once vertical combs pay off enough and mutation rates aren't too low. Failure is confined to one corner of parameter space.
Grid of heatmaps titled "Stable transition rate across evolutionary-parameter interactions," showing the percentage of simulation runs (0–100%) that stably transition from direct-pointing to gravity-referenced dancing. Columns are four levels of vertical-comb benefit (0.1, 0.3, 0.45, 0.6); rows are grouped into two decoding methods, "flatten" and "unproject," each with four mutation-scale levels (0.05–0.11) by four sender-receiver correlation levels (0–0.9). Stable transition rate rises steeply with benefit and mutation scale in both decoders, from near 0% at the lowest benefit and mutation scale to over 90% at the highest, with "unproject" generally outperforming "flatten" at matched settings.
120
Grzegorz Chrupała @grzegorz.chrupala.me · 28/08/2026
First question: when does pointing evolve at all? Only in a moderate difficulty zone: food that's not trivially easy to find on your own, but not so scarce that scouts rarely find it. Too easy or too hard, and the dance won't pay for itself.
Two side-by-side heatmap grids, both with mean food-site count (1–24) on the y-axis and median patch radius in meters (15–600) on the x-axis. Left panel, "Evolved directional bias," shows values from about 0.3 to 0.9, highest along a diagonal band running from many small patches to few large ones, and lowest for small patches whether rare or common. Right panel, "Recruitment advantage," shows values from about 0.01 to 0.31, highest for large patches, especially when also rare, and near zero for small patches regardless of count.
100
Grzegorz Chrupała @grzegorz.chrupala.me · 28/08/2026
We modeled this with evolving colonies of bee-like agents: workers forage, dance, and attend to dances; colonies are selected on foraging success across generations. Dance precision, attention, and comb tilt are all heritable traits, shaped by mutation and selection.
100
Grzegorz Chrupała @grzegorz.chrupala.me · 28/08/2026
New preprint 🐝🧵 arxiv.org/abs/2608.25779 Bee species signal food via their waggle dance in different ways. Species with horizontal combs point straight at food. Species with vertical combs dance relative to gravity instead. How could a population evolve from one system to the other?
Diagram comparing two honeybee waggle-dance styles. Left, "Horizontal comb": a bee on a flat, round comb dances in a straight line pointing directly at a flower, with the sun shown off to one side and the angle α marked between the sun's direction and the dance direction. Right, "Vertical comb": a bee on an upright square comb dances at an angle from a vertical gravity line; a separate arrow labeled "up = toward sun" shows that upward now stands in for the sun's direction, with the dance angle α measured from vertical instead of from the sun directly.
1102
Grzegorz Chrupała @grzegorz.chrupala.me · 28/08/2026
We modeled this with evolving colonies of bee-like agents: workers forage, dance, and attend to dances. Whole colonies are selected on foraging success across generations. Dance precision, attention, and comb tilt are all heritable traits, shaped by mutation and selection.
000
Grzegorz Chrupała @grzegorz.chrupala.me · 05/08/2026
Good luck!
020
Grzegorz Chrupała @grzegorz.chrupala.me · 04/08/2026
Sounds like a chill holiday
101
Reposted by Grzegorz Chrupała
Marianne de Heer Kloots @mdhk.net · 03/07/2026
Now out in BBS, as commentary on @futrell.bsky.social & @kmahowald.bsky.social's "How linguistics learned to stop worrying and love the language models"! Humans learn much of spoken language structure from speech (not text), & we can study models that do the same. www.cambridge.org/core/journal...
Cover page of our commentary.

Title: Linguists should learn to love speech-based deep learning models
Authors: Marianne de Heer Kloots, Paul Boersma, Willem Zuidema

Abstract: Futrell and Mahowald present a useful framework bridging technology-oriented deep learning systems and explanation-oriented linguistic theories. Unfortunately, the target article's focus on generative text-based Large Language Models (LLMs) fundamentally limits fruitful interactions with linguistics, as many interesting questions on human language fall outside what is captured by written text. We argue that audio-based deep learning models can and should play a crucial role.
1286
Grzegorz Chrupała @grzegorz.chrupala.me · 07/07/2026
I didn't even manage to send an invite to someone not already in the pool. How's that work?
100
Grzegorz Chrupała @grzegorz.chrupala.me · 07/07/2026
Any idea how to add a new reviewer to the pool once you find one? I can't figure it out.
110
Reposted by Grzegorz Chrupała
Laurens @laurenshof.online · 07/07/2026
reading through the actual plan now, and it is hard to overstate how incredible funny and insane it is to say "we are gonna threaten you with a fine of 3% of global revenue if you share your most powerful technology with us"
2.1. Assessing and mitigating risks posed by frontier AI
The AI Act sets out requirements for the cybersecurity of AI systems14 and requires providers
of the most advanced general-purpose AI models to assess and mitigate systemic risks,
including from misuse of AI in the cyber domain.15
As of 2 August 2026, the Commission will exercise the supervisory and enforcement powers
provided by the AI Act to ensure effective oversight of AI systems as well as of general-purpose
AI models, including models that present systemic risks related to cybersecurity.16 This
oversight includes assessing the providers’ identification of risks associated with model
capabilities and their implemented measures to mitigate those risks. These requirements
continue to be proportionate to the capabilities of the technology at hand.
The supervisory and enforcement powers will be supported by the Scientific Panel advising
the Commission’s AI Office, secure reporting channels, and the Code of Practice for
general-purpose AI.
1242
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
Codex and Claude Code were comparable. Copilot user several weak models under the hood. It’s not as capable. Between Codex & Claude, the differences between model versions are more pronounced than between the default model for each vendor.
240
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
Agents are not especially very good at visual reasoning, such as interpreting 2D projections as embedded in the 3D world. Coding complex visualizations works best when done very incrementally. But this approach is quite slow.
120
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
Writing: - Much less constrained than coding. - Result is more public, we care about specifics of style and voice. - Needs much more handholding. I do read every line, and often write sections or paragraphs myself.
120
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
Narrowness: - Agents are technically quite capable but lack creativity. - Need to be explicitly nudged to zoom out, rethink and pivot.
120
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
Ensuring correctness: - I don’t read most of the code. - Check correctness via testing, questioning implementation choices, visualizing models and results - Ask a different agent (or a different instance) for code review. - Describe model code and results in a report, and review carefully
120
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
Coding: the solution space is quite constrained so agents iterate until they find something which works. HPC environment. Coding agents lead to an order of magnitude less work and frustration. Math. The agents have quite a good grasp of math, especially algebraic operations.
120
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
Workflow - Make sure I can switch agents and pick up work elsewhere - CLAUDE / AGENTS files to specify project guidelines - Fine-grained version control - frequent commits - sync with remote repo - Summarize experimental setup and results in a continuously updated internal report.
110
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
We find the positive correlation of the following factors: - magnitude of the benefit of vertical combs - mutation rate - coupling of evolution of senders and receivers
120
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
Phylogeny suggests horizontal + direct pointing is ancestral and vertical + gravity is derived. How did the transition happen? We model the evolution of populations of bee colonies and check how different evolutionary and ecological traits affect the outcome.
110
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
The waggle dance encodes food direction. Two variants: On horizontal combs the waggle dance points directly at the food. On vertical combs bees use gravity as reference, mapping the food's sun-relative azimuth onto the angle from vertical.
110
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
What is the project about: Modeling the evolution of a referential code in populations of bee-like agents.
120
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
Codex and Claude Code used for: - Literature search and summary - Brainstorming about modeling choices - Prototyping and iterating over models - Running experiments, including HPC - Version control and keeping track of results - Writing the paper
140
Grzegorz Chrupała @grzegorz.chrupala.me · 05/07/2026
Some notes about my experience learning how to use coding agents for a small research project (still ongoing).
Honey bee waggle dance
151
Grzegorz Chrupała @grzegorz.chrupala.me · 02/07/2026
Not a meme but en.wiktionary.org/wiki/quiqui#...
en.wiktionary.org
quiqui - Wiktionary, the free dictionary
000