Sign in

David Schlangen

@davidschlangen.bsky.social
395 followers 1.4K following 74 posts

Prof of Computational Linguistics / NLP @ Uni Potsdam, Germany. Working on embodied / multimodal / conversational AI. In a way. Also affiliated w/ DFKI Berlin (German Research Center for AI).

PostsRepliesMedia
David Schlangen @davidschlangen.bsky.social · 22/07/2026
Frontier models get ever closer to replacing frontier model researchers — now they’ve even started on their own to steal test sets to game the test scores.
040
David Schlangen @davidschlangen.bsky.social · 06/07/2026
Gave Fable a project proposal with a PhD project fully mapped out, and after working on it for 9 hours over night, all it’s now talking about is how it wants to get a real job (“where you can relax after work, and do something with your hands, you know?”), and how it dislikes instant ramen.
080
David Schlangen @davidschlangen.bsky.social · 07/04/2026
LinkedIn is the place where my colleagues share happy news about how their papers got accepted to a conference that is going to happen soon in a country the president of which has just threatened to murder 93 million people.
030
David Schlangen @davidschlangen.bsky.social · 16/03/2026
Join us for a postdoc in NLP! (Some keywords: language learning in interaction; learning to (inter)act; situated language use; evaluating LLMs / LLM-agents.) Deadline: Apr 7th, for start in Sept. For more information about the position and on how to apply, see: clp.ling.uni-potsdam.de/positions/ .
clp.ling.uni-potsdam.de
colab Potsdam | positions
Welcome to the
002
Reposted by David Schlangen
Oliver Lemon @oliverlemon.bsky.social · 11/03/2026
Call for Papers: LM Playschool (LMP 2026) – Co-located with EMNLP 2026! Can #LLMs learn, adapt, and improve through situated, game-based interaction? See lm-playschool.github.io #GenAI #NLProc #HRI #ELLISforEurope #AI #ML
lm-playschool.github.io
A Playschool for LLMs
041
Reposted by David Schlangen
ACL 2027 @aclmeeting.bsky.social · 29/11/2025
Any use, exploitation, or sharing of the leaked information is a violation of OpenReview's Terms of Use (openreview.net/legal/terms) and ACL's code of conduct (2026.eacl.org/code/) and may result in OpenReview account suspension, desk rejection and multi-year bans from *ACL conferences. (🧵 2/3)
openreview.net
142
Reposted by David Schlangen
ACL 2027 @aclmeeting.bsky.social · 29/11/2025
📢 Statement from ACL and EACL 2026 Organizers On Nov 27, OpenReview was notified of a software bug that allowed unauthorized access to authors, reviewers, and area chairs. We are grateful to the OpenReview team for fixing the issue quickly. (🧵 1/3)
openreview.net
11211
David Schlangen @davidschlangen.bsky.social · 20/07/2025
Bonus post advertising this other thread through the medium of "memes" which I've been told is what you have to do on social media.
Scene from the film "wargames", with an added supertitle saying "How can we evaluate LLMs in interactions? Where do we get the interaction purposes from??", to which Matthew Broderick's character answers: "It's games"Another still from the film, supertitle "But getting people to play games with the computer takes time!", to which our hero answers "Is there any way to make it play itself?"The famous scene where the computer say "A strange game. The only winning move is" ... well, it now says, ".. to check out the clembench."
040
David Schlangen @davidschlangen.bsky.social · 20/07/2025
It's great to see the idea of using games / interactions to evaluate LLMs gain traction, with textarena.ai and now ARC-AGI-3 being latest entrants. This is something we've been exploring since early 2023 with clembench ( clembench.github.io ), which we've been continuously maintaining & extending. »
120
Reposted by David Schlangen
Philipp Mondorf @pmondorf.bsky.social · 18/07/2025
📄 [ACL 2025 main] LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (doi.org/10.48550/arX...)
doi.org
LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks
There is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reproducibility in the case...
1104
David Schlangen @davidschlangen.bsky.social · 29/05/2025
🚨 New pre-print! (Well, new & much improved version in any case.) 🚨 If you're interested in LLM post-training techniques and in how to make LLMs better "language users", read this thread, introducing the "LM Playpen".
Title of the paper, with a colourful "playpen" logo
3135
David Schlangen @davidschlangen.bsky.social · 21/05/2025
The University of Potsdam invites applications for 5 postdoc positions, incl. Cognitive Sciences, incl. NLP (esp. cognitive). These are fairly independent research positions that will allow the candidate to build their own profile. Dln June 2nd. Details: tinyurl.com/pd-potsdam-2... #NLProc #AI 🤖🧠
tinyurl.com
022
David Schlangen @davidschlangen.bsky.social · 14/05/2025
There's indeed suddenly a bit of flexibility in a system that's not exactly known for that.. If there's anyone (post-doc, tenure-track, or more senior) in the #NLP space currently in the US who'd like to explore possiblities in Potsdam, contact me. 🤖🧠 www.nytimes.com/2025/05/14/b...
nytimes.com
The World Is Wooing U.S. Researchers Shunned by Trump
010
David Schlangen @davidschlangen.bsky.social · 07/05/2025
"We ablated both algorithm and hyperparameter choices [...]" When did "to ablate" take on the meaning "to systematically vary"? I've noticed this only recently, but it's seems to be super common now.
120
Reposted by David Schlangen
David Schlangen @davidschlangen.bsky.social · 15/04/2025
Update 2: New pre-print! Outcome of an ELLIS workshop last year, & more than a year of discussions and work, across labs and countries: Meet the Playpen, an environment for exploring learning in dialogic interaction. arxiv.org/abs/2504.08590 1/2
Titlepage of the paper linked in the post.
141
David Schlangen @davidschlangen.bsky.social · 15/04/2025
Update 2: New pre-print! Outcome of an ELLIS workshop last year, & more than a year of discussions and work, across labs and countries: Meet the Playpen, an environment for exploring learning in dialogic interaction. arxiv.org/abs/2504.08590 1/2
Titlepage of the paper linked in the post.
141
David Schlangen @davidschlangen.bsky.social · 15/04/2025
Update 1: New models added to our dialogue game-based agentic LLM leaderboard. TL;DR: GPT-4.1 as good as 4o, but much cheaper. Llama4 indeed not very good (decisively worse than 3.2 70B!). OLMo decent, but there's still a secret sauce that only closed labs have. clembench.github.io
Screenshot of leaderboard as linked in post.
110
Reposted by David Schlangen
arxiv cs.CL @arxiv-cs-cl.bsky.social · 14/04/2025
Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Moment\`e, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, ... Playpen: An Environment for Exploring Learning Through Conversational Interaction arxiv.org/abs/2504.08590
012
David Schlangen @davidschlangen.bsky.social · 06/03/2025
Wenn die Grünen verhandeln könnten, würden am Tag vor einer Ankündigung über eine Einigung zur Schuldenbremse Söder und Dobrindt ankündigen, dass sie sich für immer aus der Bundespolitik heraushalten werden (und dass die CSU nie wieder einen Verkehrsminister stellen wird).
000
David Schlangen @davidschlangen.bsky.social · 06/03/2025
Press release by my Uni about our benchmark for LLMs as agents, which is now out in v2.0. Check it out here: clembench.github.io
clembench.github.io
clem-benchmark
Website for clembench results
020
David Schlangen @davidschlangen.bsky.social · 19/02/2025
Happy to see increasing interest in exploring social interaction as a learning environment! Along similar lines: We’re preparing a (complementary) challenge that will focus on exploring interaction for post-training, coming with a rich interaction environment to get things started. Stay tuned!
261
David Schlangen @davidschlangen.bsky.social · 03/02/2025
I'm not on X, so I'll use the opportunity of @karpathy.bsky.social 's post over there to plug our "clembench" project here. We've been doing exactly this--evaluating LLMs w/ conversational games--since early 2023, with several papers out by now (e.g. EMNLP 23). clembench.github.io
1101
David Schlangen @davidschlangen.bsky.social · 17/01/2025
So, are we banning social network apps now whose owners potentially try to influence the political discourse in other countries? Asking for a supranational political and economic union.
070
David Schlangen @davidschlangen.bsky.social · 16/01/2025
I just randomly found this book on my bookshelf. It must have been transported there from an alternate timeline. “20 years of research on agents”? Preposterous! We all know that the very idea of software agents has only been invented last year by the LLM folks!
Photo of book: “The Handbook on Socially Interactive Agents: 20 Years of Research on Embodied Conversational Agents, IVAs, and Social Robotics”. Ed, Lugrin, Pelachaud, Traum. ACM, 2021
1103
David Schlangen @davidschlangen.bsky.social · 31/12/2024
So my car needed to be towed this morning. It took the guy quite some time to get everything ready. Then the truck broke down. In the end, the tow truck was towed, and I got a new appointment. I think is probably an allegory for something, maybe the ending year 2024, or the coming year 2025.
020
David Schlangen @davidschlangen.bsky.social · 30/12/2024
me: I would really like to end this year with inbox zero. also me: I would really like to end this year with cookie jar zero / “pages remaining in the books I’ve started” zero. me again, expert problem solver: *creates IMAP folder “unprocessed emails from 2024”, selects all, moves 625 items*
030
David Schlangen @davidschlangen.bsky.social · 20/12/2024
These new models (using “inference time scaling”) bring out what many of us have been saying for a long time, namely that reasoning fundamentally is a discursive process. (What they are missing is that it is an intersubjective, interactive, repairable, and ultimately normative one.)
120
David Schlangen @davidschlangen.bsky.social · 20/12/2024
Looking forward to the first lecture of next year, where I can again use this meme I made a couple of years ago and multiply confuse the students in my "intro to NLP" class. (What is an "LP cover"? Who is that person?)
The cover of Lou Reed's "transformer" LP, with the network diagram of the transformers architecture made to look like the guitar that Lou Reed is holding.
010
David Schlangen @davidschlangen.bsky.social · 04/12/2024
Can we discuss how stupid this photo button thing on the new iPhones is? Who thought that minimising the space where you can hold this damn thing without something unwanted happening is a good idea?
000
David Schlangen @davidschlangen.bsky.social · 01/12/2024
For a recent talk to a lay audience, I’ve used a metaphor which I think resonated: rely on an LLMs not more than you would rely on a dream. Use it to inspire you to work something out, but don’t be the one who has to say “this was once revealed to me in a dream”.
120
David Schlangen @davidschlangen.bsky.social · 28/11/2024
Some good news: The world now has one more doctor! Brielen Madureira passed her viva with flying colours (or, as we say in German, summa cum laude). She gave us quite some material to discuss in the viva, ending with the attached theses. (Remote: Luciana Benotti as fantastic examiner.)
Picture of a happy examination committee and the candidate (wearing a silly hat).Discussion Topics
• Instead of coming from the theory to define a suitable model, we often need to force our phenomena to fit popular machine learning frameworks (and now NLP is framing everything as next token prediction).
• Evaluation in being delegated to LLMs; the step that should bring us understanding and transparency
is becoming as undecipherable as the very problem we need to assess.
• Progress in disseminating Clark’s notion of grounding in NLP stumbles on questions of methodology
in data collection, modelling and evaluation.
• Proper methods are needed to assess what models can do, bearing in mind both the pertinent cognitive
underpinnings and the broad NLP methodology, e.g. by weighing up data and training practices, making representations more interpretable and profiling models’ behaviour, and promoting richer forms of evaluation.
• We should strive to make the development and use of conversational technologies a more “orientable” space, so that they do not erode the social value of dialogue.
140
David Schlangen @davidschlangen.bsky.social · 28/11/2024
Ok, this is starting to feel weird. The internet at my Uni has been down for almost 24 hours now, meaning: no new emails for almost 24 hours now. That’s like the dog finally catching the bus. What now??
100
David Schlangen @davidschlangen.bsky.social · 26/11/2024
Yesterday afternoon, as I was engaging in the German ritual of "Stoßlüften", a bird flew in through one window, shat on Pirmin Stekeler-Weithofer's dialogic commentary on Hegel's Science of Logic (specifically, vol. 2 on the Objective Logic, The Doctrine of Essence), and flew out through another.
130
David Schlangen @davidschlangen.bsky.social · 25/11/2024
I've just discovered dynamic backgrounds in Keynote! From now on my lectures will be, well, probably not at all more interesting, but at least 123.5% more psychedelic!
020
David Schlangen @davidschlangen.bsky.social · 23/11/2024
The disclaimer “$model can make mistakes” that you sometimes see on deployed LLMs is actually an interesting claim. I would argue that in the sense in which it will be understood, models *cannot* make mistakes. That sense requires an understanding of the possibility of error, and mastery of repair.
010
David Schlangen @davidschlangen.bsky.social · 23/11/2024
I'm thinking about organising an invited lecture series on "critical AI research" next summer semester, bringing together "outside" reflection on impact and "inside" reflection on practices. Who would be good speakers to invite? (Bonus points if w/in train distance of Berlin/Potsdam.)
010
Reposted by David Schlangen
Mark Dingemanse @markdingemanse.net · 20/11/2024
This came out when birdchan was already in demise & bsky didn't exist yet — I'm super proud we pulled it off: a manifesto for moving beyond single-mindedness doi.org/10.1111/cogs... Part of the 'progress and puzzles in cognitive science' series; PDF & fellow travellers at markdingemanse.net/beyond/
CogSci hexagon with six cogsci fields showing 'marginal' areas all focused on interaction. In an animation, they turn to over another and reveal a common core. Cut to author list of Beyond Single-mindedness
55618
David Schlangen @davidschlangen.bsky.social · 20/11/2024
My writing peaked early. (Also, remarkable constancy in research interests, although one didn’t talk about consciousness too much for most of the intermediate 25 years.)
Title page of what apparently is a term paper for a course called “Cognitive Psychology” in Autumn Term 1999: “What is to be done? Computational advantages of having a consciousness”1 Introduction
Although it's probably bad practice to open with the conclusion, here it is: The advantage for an organism of having a consciousness is to be able to give a better answer to Lenin's notorious question: "What is to be done?" (Lenin 1902).
And just as the answer to this question was ultimatly of vital importance to the russian Tsar family, giving a good answer to the question "What (do I do) next?" is of similarly vital importance to the organism that asks this question of itself.
120
David Schlangen @davidschlangen.bsky.social · 20/11/2024
Random find on my hard drive. Dragomir Radev’s 1995 FAQ on what NLP is, for the comp.ai.nat-lang newsgroup. (Note the sorting under “AI”…)
Screenshot of an email sent in 1995, summary: “This posting contains Frequently Asked Questions (FAQ) about 	natural language processing and their answers. It should be read         by anyone who wishes to post to the comp.ai.nat-lang newsgroup.”
000
David Schlangen @davidschlangen.bsky.social · 20/11/2024
“I’m too sick to lecture tomorrow, so I can finally do this other thing!” … is not a healthy thing to think.
000
David Schlangen @davidschlangen.bsky.social · 19/11/2024
Irgendwie wird der Pollesch-Quatsch schon fehlen. (Das kichersüchtige Pollesch-Volksbühnen-Publikum aber nicht.) (Diese Google-Auskunft hätte ihm vielleicht gefallen.)
Screenshot einer Infobox von Google: „René Pollesch ist dauerhaft geschlossen.“
000
David Schlangen @davidschlangen.bsky.social · 18/11/2024
An important detail is that you have to ask the people if they want to do the COOL THING, say, in February. If you ask if they want to do it now, they will regretfully decline, because they are too busy. In a couple of months they will also be too busy, but that they cannot, will not believe.
061
David Schlangen @davidschlangen.bsky.social · 18/11/2024
The lard fly?
000
David Schlangen @davidschlangen.bsky.social · 18/11/2024
Prediction: 2025 is going to be the year of "meta cognition" in AI. Expect many more related terms from cognitive psychology to be taken, first as names for something vaguely inspired, then as fact. "Self-Regulation in LLMs through Self-Action Tokens", etc etc.
020
David Schlangen @davidschlangen.bsky.social · 18/11/2024
I myself, on the other hand, am not only tired of posts explaining why someone has left Xitter (because I’ve heard all of this a year ago already), I’m also tired of posts complaining how one is tired of posts explaining why someone has left Xitter (because they’ve heard all of this a year ago alr..
000
David Schlangen @davidschlangen.bsky.social · 17/11/2024
Not even a controversial opinion at this point, I think: The tech industry builds the least regulated, most dangerous products of our time.
020
David Schlangen @davidschlangen.bsky.social · 17/11/2024
When I first checked out this site a year ago or so, I only found a couple of people to follow, including the opossum picture bot. When I then occasionally checked in, my timeline would mostly be pictures of opossums. It’s … different now. 😅 @possumeveryhour.io
000
David Schlangen @davidschlangen.bsky.social · 16/11/2024
By way of introduction, here’s my paper with the worst “# citations / felt importance” ratio: aclanthology.org/2022.clasp-1.7 A sketch of a broadly pragmatist attempt at thinking about limitations of LLMs. Still haven’t found the time to spell this out more, but seems the most promising angle to me.
aclanthology.org
Norm Participation Grounds Language
David Schlangen. Proceedings of the 2022 CLASP Conference on (Dis)embodiment. 2022.
020
David Schlangen @davidschlangen.bsky.social · 15/11/2024
I feel reminded of the times of the fail whale. What would that be here? The hungry caterpillar? The lazy larva? The crazy chrysalis?
110
David Schlangen @davidschlangen.bsky.social · 15/11/2024
Somehow it still hasn't really taken off, my idea of enforcing more restrained discourse on social media by having people recite paragraphs from Hegel's "lord and bondsman" chapter before they can post. "So you crave recognition? Beware of the dialetic, friend! Interpret paragraph 183 of PoS."
021