David Schlangen @davidschlangen.bsky.social · 22/07/2026Frontier models get ever closer to replacing frontier model researchers — now they’ve even started on their own to steal test sets to game the test scores. 040
David Schlangen @davidschlangen.bsky.social · 06/07/2026Gave Fable a project proposal with a PhD project fully mapped out, and after working on it for 9 hours over night, all it’s now talking about is how it wants to get a real job (“where you can relax after work, and do something with your hands, you know?”), and how it dislikes instant ramen. 080
David Schlangen @davidschlangen.bsky.social · 07/04/2026LinkedIn is the place where my colleagues share happy news about how their papers got accepted to a conference that is going to happen soon in a country the president of which has just threatened to murder 93 million people. 030
David Schlangen @davidschlangen.bsky.social · 16/03/2026Join us for a postdoc in NLP! (Some keywords: language learning in interaction; learning to (inter)act; situated language use; evaluating LLMs / LLM-agents.) Deadline: Apr 7th, for start in Sept. For more information about the position and on how to apply, see: clp.ling.uni-potsdam.de/positions/ .clp.ling.uni-potsdam.decolab Potsdam | positionsWelcome to the 002
Reposted by David SchlangenOliver Lemon @oliverlemon.bsky.social · 11/03/2026Call for Papers: LM Playschool (LMP 2026) – Co-located with EMNLP 2026! Can #LLMs learn, adapt, and improve through situated, game-based interaction? See lm-playschool.github.io #GenAI #NLProc #HRI #ELLISforEurope #AI #MLlm-playschool.github.ioA Playschool for LLMs 041
Reposted by David SchlangenACL 2027 @aclmeeting.bsky.social · 29/11/2025Any use, exploitation, or sharing of the leaked information is a violation of OpenReview's Terms of Use (openreview.net/legal/terms) and ACL's code of conduct (2026.eacl.org/code/) and may result in OpenReview account suspension, desk rejection and multi-year bans from *ACL conferences. (🧵 2/3)openreview.net 142
Reposted by David SchlangenACL 2027 @aclmeeting.bsky.social · 29/11/2025📢 Statement from ACL and EACL 2026 Organizers On Nov 27, OpenReview was notified of a software bug that allowed unauthorized access to authors, reviewers, and area chairs. We are grateful to the OpenReview team for fixing the issue quickly. (🧵 1/3)openreview.net 11211
David Schlangen @davidschlangen.bsky.social · 20/07/2025Bonus post advertising this other thread through the medium of "memes" which I've been told is what you have to do on social media. 040
David Schlangen @davidschlangen.bsky.social · 20/07/2025It's great to see the idea of using games / interactions to evaluate LLMs gain traction, with textarena.ai and now ARC-AGI-3 being latest entrants. This is something we've been exploring since early 2023 with clembench ( clembench.github.io ), which we've been continuously maintaining & extending. » 120
Reposted by David SchlangenPhilipp Mondorf @pmondorf.bsky.social · 18/07/2025📄 [ACL 2025 main] LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (doi.org/10.48550/arX...)doi.orgLLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation TasksThere is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reproducibility in the case... 1104
David Schlangen @davidschlangen.bsky.social · 29/05/2025🚨 New pre-print! (Well, new & much improved version in any case.) 🚨 If you're interested in LLM post-training techniques and in how to make LLMs better "language users", read this thread, introducing the "LM Playpen". 3135
David Schlangen @davidschlangen.bsky.social · 21/05/2025The University of Potsdam invites applications for 5 postdoc positions, incl. Cognitive Sciences, incl. NLP (esp. cognitive). These are fairly independent research positions that will allow the candidate to build their own profile. Dln June 2nd. Details: tinyurl.com/pd-potsdam-2... #NLProc #AI 🤖🧠tinyurl.com 022
David Schlangen @davidschlangen.bsky.social · 14/05/2025There's indeed suddenly a bit of flexibility in a system that's not exactly known for that.. If there's anyone (post-doc, tenure-track, or more senior) in the #NLP space currently in the US who'd like to explore possiblities in Potsdam, contact me. 🤖🧠 www.nytimes.com/2025/05/14/b...nytimes.comThe World Is Wooing U.S. Researchers Shunned by Trump 010
David Schlangen @davidschlangen.bsky.social · 07/05/2025"We ablated both algorithm and hyperparameter choices [...]" When did "to ablate" take on the meaning "to systematically vary"? I've noticed this only recently, but it's seems to be super common now. 120
Reposted by David SchlangenDavid Schlangen @davidschlangen.bsky.social · 15/04/2025Update 2: New pre-print! Outcome of an ELLIS workshop last year, & more than a year of discussions and work, across labs and countries: Meet the Playpen, an environment for exploring learning in dialogic interaction. arxiv.org/abs/2504.08590 1/2 141
David Schlangen @davidschlangen.bsky.social · 15/04/2025Update 2: New pre-print! Outcome of an ELLIS workshop last year, & more than a year of discussions and work, across labs and countries: Meet the Playpen, an environment for exploring learning in dialogic interaction. arxiv.org/abs/2504.08590 1/2 141
David Schlangen @davidschlangen.bsky.social · 15/04/2025Update 1: New models added to our dialogue game-based agentic LLM leaderboard. TL;DR: GPT-4.1 as good as 4o, but much cheaper. Llama4 indeed not very good (decisively worse than 3.2 70B!). OLMo decent, but there's still a secret sauce that only closed labs have. clembench.github.io 110
Reposted by David Schlangenarxiv cs.CL @arxiv-cs-cl.bsky.social · 14/04/2025Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Moment\`e, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, ... Playpen: An Environment for Exploring Learning Through Conversational Interaction arxiv.org/abs/2504.08590 012
David Schlangen @davidschlangen.bsky.social · 06/03/2025Wenn die Grünen verhandeln könnten, würden am Tag vor einer Ankündigung über eine Einigung zur Schuldenbremse Söder und Dobrindt ankündigen, dass sie sich für immer aus der Bundespolitik heraushalten werden (und dass die CSU nie wieder einen Verkehrsminister stellen wird). 000
David Schlangen @davidschlangen.bsky.social · 06/03/2025Press release by my Uni about our benchmark for LLMs as agents, which is now out in v2.0. Check it out here: clembench.github.ioclembench.github.ioclem-benchmarkWebsite for clembench results 020
David Schlangen @davidschlangen.bsky.social · 19/02/2025Happy to see increasing interest in exploring social interaction as a learning environment! Along similar lines: We’re preparing a (complementary) challenge that will focus on exploring interaction for post-training, coming with a rich interaction environment to get things started. Stay tuned! 261
David Schlangen @davidschlangen.bsky.social · 03/02/2025I'm not on X, so I'll use the opportunity of @karpathy.bsky.social 's post over there to plug our "clembench" project here. We've been doing exactly this--evaluating LLMs w/ conversational games--since early 2023, with several papers out by now (e.g. EMNLP 23). clembench.github.io 1101
David Schlangen @davidschlangen.bsky.social · 17/01/2025So, are we banning social network apps now whose owners potentially try to influence the political discourse in other countries? Asking for a supranational political and economic union. 070
David Schlangen @davidschlangen.bsky.social · 16/01/2025I just randomly found this book on my bookshelf. It must have been transported there from an alternate timeline. “20 years of research on agents”? Preposterous! We all know that the very idea of software agents has only been invented last year by the LLM folks! 1103
David Schlangen @davidschlangen.bsky.social · 31/12/2024So my car needed to be towed this morning. It took the guy quite some time to get everything ready. Then the truck broke down. In the end, the tow truck was towed, and I got a new appointment. I think is probably an allegory for something, maybe the ending year 2024, or the coming year 2025. 020
David Schlangen @davidschlangen.bsky.social · 30/12/2024me: I would really like to end this year with inbox zero. also me: I would really like to end this year with cookie jar zero / “pages remaining in the books I’ve started” zero. me again, expert problem solver: *creates IMAP folder “unprocessed emails from 2024”, selects all, moves 625 items* 030
David Schlangen @davidschlangen.bsky.social · 20/12/2024These new models (using “inference time scaling”) bring out what many of us have been saying for a long time, namely that reasoning fundamentally is a discursive process. (What they are missing is that it is an intersubjective, interactive, repairable, and ultimately normative one.) 120
David Schlangen @davidschlangen.bsky.social · 20/12/2024Looking forward to the first lecture of next year, where I can again use this meme I made a couple of years ago and multiply confuse the students in my "intro to NLP" class. (What is an "LP cover"? Who is that person?) 010
David Schlangen @davidschlangen.bsky.social · 04/12/2024Can we discuss how stupid this photo button thing on the new iPhones is? Who thought that minimising the space where you can hold this damn thing without something unwanted happening is a good idea? 000
David Schlangen @davidschlangen.bsky.social · 01/12/2024For a recent talk to a lay audience, I’ve used a metaphor which I think resonated: rely on an LLMs not more than you would rely on a dream. Use it to inspire you to work something out, but don’t be the one who has to say “this was once revealed to me in a dream”. 120
David Schlangen @davidschlangen.bsky.social · 28/11/2024Some good news: The world now has one more doctor! Brielen Madureira passed her viva with flying colours (or, as we say in German, summa cum laude). She gave us quite some material to discuss in the viva, ending with the attached theses. (Remote: Luciana Benotti as fantastic examiner.) 140
David Schlangen @davidschlangen.bsky.social · 28/11/2024Ok, this is starting to feel weird. The internet at my Uni has been down for almost 24 hours now, meaning: no new emails for almost 24 hours now. That’s like the dog finally catching the bus. What now?? 100
David Schlangen @davidschlangen.bsky.social · 26/11/2024Yesterday afternoon, as I was engaging in the German ritual of "Stoßlüften", a bird flew in through one window, shat on Pirmin Stekeler-Weithofer's dialogic commentary on Hegel's Science of Logic (specifically, vol. 2 on the Objective Logic, The Doctrine of Essence), and flew out through another. 130
David Schlangen @davidschlangen.bsky.social · 25/11/2024I've just discovered dynamic backgrounds in Keynote! From now on my lectures will be, well, probably not at all more interesting, but at least 123.5% more psychedelic! 020
David Schlangen @davidschlangen.bsky.social · 23/11/2024The disclaimer “$model can make mistakes” that you sometimes see on deployed LLMs is actually an interesting claim. I would argue that in the sense in which it will be understood, models *cannot* make mistakes. That sense requires an understanding of the possibility of error, and mastery of repair. 010
David Schlangen @davidschlangen.bsky.social · 23/11/2024I'm thinking about organising an invited lecture series on "critical AI research" next summer semester, bringing together "outside" reflection on impact and "inside" reflection on practices. Who would be good speakers to invite? (Bonus points if w/in train distance of Berlin/Potsdam.) 010
Reposted by David SchlangenMark Dingemanse @markdingemanse.net · 20/11/2024This came out when birdchan was already in demise & bsky didn't exist yet — I'm super proud we pulled it off: a manifesto for moving beyond single-mindedness doi.org/10.1111/cogs... Part of the 'progress and puzzles in cognitive science' series; PDF & fellow travellers at markdingemanse.net/beyond/ 55618
David Schlangen @davidschlangen.bsky.social · 20/11/2024My writing peaked early. (Also, remarkable constancy in research interests, although one didn’t talk about consciousness too much for most of the intermediate 25 years.) 120
David Schlangen @davidschlangen.bsky.social · 20/11/2024Random find on my hard drive. Dragomir Radev’s 1995 FAQ on what NLP is, for the comp.ai.nat-lang newsgroup. (Note the sorting under “AI”…) 000
David Schlangen @davidschlangen.bsky.social · 20/11/2024“I’m too sick to lecture tomorrow, so I can finally do this other thing!” … is not a healthy thing to think. 000
David Schlangen @davidschlangen.bsky.social · 19/11/2024Irgendwie wird der Pollesch-Quatsch schon fehlen. (Das kichersüchtige Pollesch-Volksbühnen-Publikum aber nicht.) (Diese Google-Auskunft hätte ihm vielleicht gefallen.) 000
David Schlangen @davidschlangen.bsky.social · 18/11/2024An important detail is that you have to ask the people if they want to do the COOL THING, say, in February. If you ask if they want to do it now, they will regretfully decline, because they are too busy. In a couple of months they will also be too busy, but that they cannot, will not believe. 061
David Schlangen @davidschlangen.bsky.social · 18/11/2024Prediction: 2025 is going to be the year of "meta cognition" in AI. Expect many more related terms from cognitive psychology to be taken, first as names for something vaguely inspired, then as fact. "Self-Regulation in LLMs through Self-Action Tokens", etc etc. 020
David Schlangen @davidschlangen.bsky.social · 18/11/2024I myself, on the other hand, am not only tired of posts explaining why someone has left Xitter (because I’ve heard all of this a year ago already), I’m also tired of posts complaining how one is tired of posts explaining why someone has left Xitter (because they’ve heard all of this a year ago alr.. 000
David Schlangen @davidschlangen.bsky.social · 17/11/2024Not even a controversial opinion at this point, I think: The tech industry builds the least regulated, most dangerous products of our time. 020
David Schlangen @davidschlangen.bsky.social · 17/11/2024When I first checked out this site a year ago or so, I only found a couple of people to follow, including the opossum picture bot. When I then occasionally checked in, my timeline would mostly be pictures of opossums. It’s … different now. 😅 @possumeveryhour.io 000
David Schlangen @davidschlangen.bsky.social · 16/11/2024By way of introduction, here’s my paper with the worst “# citations / felt importance” ratio: aclanthology.org/2022.clasp-1.7 A sketch of a broadly pragmatist attempt at thinking about limitations of LLMs. Still haven’t found the time to spell this out more, but seems the most promising angle to me.aclanthology.orgNorm Participation Grounds LanguageDavid Schlangen. Proceedings of the 2022 CLASP Conference on (Dis)embodiment. 2022. 020
David Schlangen @davidschlangen.bsky.social · 15/11/2024I feel reminded of the times of the fail whale. What would that be here? The hungry caterpillar? The lazy larva? The crazy chrysalis? 110
David Schlangen @davidschlangen.bsky.social · 15/11/2024Somehow it still hasn't really taken off, my idea of enforcing more restrained discourse on social media by having people recite paragraphs from Hegel's "lord and bondsman" chapter before they can post. "So you crave recognition? Beware of the dialetic, friend! Interpret paragraph 183 of PoS." 021