David Schlangen @davidschlangen.bsky.social · 22/07/2026Frontier models get ever closer to replacing frontier model researchers — now they’ve even started on their own to steal test sets to game the test scores. 040
David Schlangen @davidschlangen.bsky.social · 06/07/2026Gave Fable a project proposal with a PhD project fully mapped out, and after working on it for 9 hours over night, all it’s now talking about is how it wants to get a real job (“where you can relax after work, and do something with your hands, you know?”), and how it dislikes instant ramen. 080
David Schlangen @davidschlangen.bsky.social · 07/04/2026LinkedIn is the place where my colleagues share happy news about how their papers got accepted to a conference that is going to happen soon in a country the president of which has just threatened to murder 93 million people. 030
David Schlangen @davidschlangen.bsky.social · 16/03/2026Join us for a postdoc in NLP! (Some keywords: language learning in interaction; learning to (inter)act; situated language use; evaluating LLMs / LLM-agents.) Deadline: Apr 7th, for start in Sept. For more information about the position and on how to apply, see: clp.ling.uni-potsdam.de/positions/ .clp.ling.uni-potsdam.decolab Potsdam | positionsWelcome to the 002
Reposted by David SchlangenOliver Lemon @oliverlemon.bsky.social · 11/03/2026Call for Papers: LM Playschool (LMP 2026) – Co-located with EMNLP 2026! Can #LLMs learn, adapt, and improve through situated, game-based interaction? See lm-playschool.github.io #GenAI #NLProc #HRI #ELLISforEurope #AI #MLlm-playschool.github.ioA Playschool for LLMs 041
Reposted by David SchlangenACL 2027 @aclmeeting.bsky.social · 29/11/2025Any use, exploitation, or sharing of the leaked information is a violation of OpenReview's Terms of Use (openreview.net/legal/terms) and ACL's code of conduct (2026.eacl.org/code/) and may result in OpenReview account suspension, desk rejection and multi-year bans from *ACL conferences. (🧵 2/3)openreview.net 142
Reposted by David SchlangenACL 2027 @aclmeeting.bsky.social · 29/11/2025📢 Statement from ACL and EACL 2026 Organizers On Nov 27, OpenReview was notified of a software bug that allowed unauthorized access to authors, reviewers, and area chairs. We are grateful to the OpenReview team for fixing the issue quickly. (🧵 1/3)openreview.net 11211
David Schlangen @davidschlangen.bsky.social · 20/07/2025Bonus post advertising this other thread through the medium of "memes" which I've been told is what you have to do on social media. 040
David Schlangen @davidschlangen.bsky.social · 20/07/2025(That animation in the first post? That's claude trying, and failing, to fully explore a maze in the MapWorld game.) 030
David Schlangen @davidschlangen.bsky.social · 20/07/2025We'd love for other people to use it to test the interaction / agentic abilities of their models, and/or to build new fun and challenging games / interactions! github.com/clp-research... github.com/clp-research... »github.comGitHub - clp-research/clembench: Collection of games to be run with the clemcore frameworkCollection of games to be run with the clemcore framework - clp-research/clembench 120
David Schlangen @davidschlangen.bsky.social · 20/07/2025Thanks to a recent short-term grant, we've been able to focus on code quality and ease of use for benchmarking and extensibility. (Exploring new games is a fun programming lab activity, which we've run several times by now!) Here's a writeup of the current state: arxiv.org/abs/2507.08491 »arxiv.orgA Third Paradigm for LLM Evaluation: Dialogue Game-Based Evaluation using clembenchThere are currently two main paradigms for evaluating large language models (LLMs), reference-based evaluation and preference-based evaluation. The first, carried over from the evaluation of machine l... 100
David Schlangen @davidschlangen.bsky.social · 20/07/2025clembench now spans abstract (e.g., wordle) and concrete tasks (simulated household); language and l+vision; and benchmarking, learning (playpen), and user simulation (clem:todd). arxiv.org/abs/2504.08590 arxiv.org/abs/2505.05445 »arxiv.orgPlaypen: An Environment for Exploring Learning Through Conversational InteractionInteraction between learner and feedback-giver has come into focus recently for post-training of Large Language Models (LLMs), through the use of reward models that judge the appropriateness of a mode... 100
David Schlangen @davidschlangen.bsky.social · 20/07/2025It's great to see the idea of using games / interactions to evaluate LLMs gain traction, with textarena.ai and now ARC-AGI-3 being latest entrants. This is something we've been exploring since early 2023 with clembench ( clembench.github.io ), which we've been continuously maintaining & extending. » 120
Reposted by David SchlangenPhilipp Mondorf @pmondorf.bsky.social · 18/07/2025📄 [ACL 2025 main] LLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation Tasks (doi.org/10.48550/arX...)doi.orgLLMs instead of Human Judges? A Large Scale Empirical Study across 20 NLP Evaluation TasksThere is an increasing trend towards evaluating NLP models with LLMs instead of human judgments, raising questions about the validity of these evaluations, as well as their reproducibility in the case... 1104
David Schlangen @davidschlangen.bsky.social · 30/05/2025Ha, yes, I'm quite pleased as well with how that turned out. It's nothing fancy, just a nice font, colouring (obviously), fbox, and rotate. 010
David Schlangen @davidschlangen.bsky.social · 29/05/2025This was the outcome of a collaboration that started last year at an ELLIS workshop, and that has brought together many labs (and many master's and PhD students, and PIs). Much more remains to be explored in "learning in interaction" -- maybe by you? 🤖🧠 #NLP #AI #LLM 020
David Schlangen @davidschlangen.bsky.social · 29/05/2025Oh yes, here's the link to the actual pre-print: arxiv.org/abs/2504.08590arxiv.orgPlaypen: An Environment for Exploring Learning Through Conversational InteractionInteraction between learner and feedback-giver has come into focus recently for post-training of Large Language Models (LLMs), through the use of reward models that judge the appropriateness of a mode... 111
David Schlangen @davidschlangen.bsky.social · 29/05/2025We release the framework and the baseline training setups to foster research in the promising new direction of learning in (synthetic) interaction which we believe will provide more effective ways of post-training agentic conversational LLMs. github.com/lm-playpen/p...github.comGitHub - lm-playpen/playpen: All you need to get started with the LM Playpen Environment for Learning in Interaction.All you need to get started with the LM Playpen Environment for Learning in Interaction. - lm-playpen/playpen 130
David Schlangen @davidschlangen.bsky.social · 29/05/2025We find that imitation learning through SFT improves performance on unseen game instances, but does not generalise to new games and negatively impacts other skills -- while interactive learning with GRPO shows balanced improvements without loss of skills. 130
David Schlangen @davidschlangen.bsky.social · 29/05/2025Together with the learning environment, we also define an experimental setup combining gameplay evaluation on unseen games and traditional NLP benchmarks such as MMLU following (Momente’ et al. 2025) arxiv.org/abs/2502.14359arxiv.orgTriangulating LLM Progress through Benchmarks, Games, and Cognitive TestsWe examine three evaluation paradigms: standard benchmarks (e.g., MMLU and BBH), interactive games (e.g., Signalling Games or Taboo), and cognitive tests (e.g., for working memory or theory of mind). ... 130
David Schlangen @davidschlangen.bsky.social · 29/05/2025Playpen is a training environment for post-training LLMs through learning in interaction, by self-play of "dialogue games": goal-oriented language-based activities that generate verifiable rewards. 120
David Schlangen @davidschlangen.bsky.social · 29/05/2025🚨 New pre-print! (Well, new & much improved version in any case.) 🚨 If you're interested in LLM post-training techniques and in how to make LLMs better "language users", read this thread, introducing the "LM Playpen". 3135
David Schlangen @davidschlangen.bsky.social · 21/05/2025The University of Potsdam invites applications for 5 postdoc positions, incl. Cognitive Sciences, incl. NLP (esp. cognitive). These are fairly independent research positions that will allow the candidate to build their own profile. Dln June 2nd. Details: tinyurl.com/pd-potsdam-2... #NLProc #AI 🤖🧠tinyurl.com 022
David Schlangen @davidschlangen.bsky.social · 14/05/2025There's indeed suddenly a bit of flexibility in a system that's not exactly known for that.. If there's anyone (post-doc, tenure-track, or more senior) in the #NLP space currently in the US who'd like to explore possiblities in Potsdam, contact me. 🤖🧠 www.nytimes.com/2025/05/14/b...nytimes.comThe World Is Wooing U.S. Researchers Shunned by Trump 010
David Schlangen @davidschlangen.bsky.social · 07/05/2025"We ablated both algorithm and hyperparameter choices [...]" When did "to ablate" take on the meaning "to systematically vary"? I've noticed this only recently, but it's seems to be super common now. 120
Reposted by David SchlangenDavid Schlangen @davidschlangen.bsky.social · 15/04/2025Update 2: New pre-print! Outcome of an ELLIS workshop last year, & more than a year of discussions and work, across labs and countries: Meet the Playpen, an environment for exploring learning in dialogic interaction. arxiv.org/abs/2504.08590 1/2 141
David Schlangen @davidschlangen.bsky.social · 15/04/2025[Sneak preview: If you're wondering where this is going, have a secret look at lm-playschool.github.io -- and stay tuned for more info!] 3/2lm-playschool.github.ioA Playschool for LLMs 000
David Schlangen @davidschlangen.bsky.social · 15/04/2025Nice baseline results as well: learning via SFT from transcripts does a bit, but only "real"(-ish) learning in interaction (GRPO) generalises. (Basically, you want to see the whole row being green in this table.) 2/2 110
David Schlangen @davidschlangen.bsky.social · 15/04/2025Update 2: New pre-print! Outcome of an ELLIS workshop last year, & more than a year of discussions and work, across labs and countries: Meet the Playpen, an environment for exploring learning in dialogic interaction. arxiv.org/abs/2504.08590 1/2 141
David Schlangen @davidschlangen.bsky.social · 15/04/2025This is only a subset of the models on the leaderboard, visit the site to see all 32 models, and also the results for the multimodal version of the benchmark. 000
David Schlangen @davidschlangen.bsky.social · 15/04/2025Update 1: New models added to our dialogue game-based agentic LLM leaderboard. TL;DR: GPT-4.1 as good as 4o, but much cheaper. Llama4 indeed not very good (decisively worse than 3.2 70B!). OLMo decent, but there's still a secret sauce that only closed labs have. clembench.github.io 110
Reposted by David Schlangenarxiv cs.CL @arxiv-cs-cl.bsky.social · 14/04/2025Nicola Horst, Davide Mazzaccara, Antonia Schmidt, Michael Sullivan, Filippo Moment\`e, Luca Franceschetti, Philipp Sadler, Sherzod Hakimov, Alberto Testoni, ... Playpen: An Environment for Exploring Learning Through Conversational Interaction arxiv.org/abs/2504.08590 012
David Schlangen @davidschlangen.bsky.social · 06/03/2025Wenn die Grünen verhandeln könnten, würden am Tag vor einer Ankündigung über eine Einigung zur Schuldenbremse Söder und Dobrindt ankündigen, dass sie sich für immer aus der Bundespolitik heraushalten werden (und dass die CSU nie wieder einen Verkehrsminister stellen wird). 000
David Schlangen @davidschlangen.bsky.social · 06/03/2025Press release by my Uni about our benchmark for LLMs as agents, which is now out in v2.0. Check it out here: clembench.github.ioclembench.github.ioclem-benchmarkWebsite for clembench results 020
David Schlangen @davidschlangen.bsky.social · 21/02/2025Not yet, but we’ll keep you posted of new developments. Thanks for your interest! 010
David Schlangen @davidschlangen.bsky.social · 19/02/2025Happy to see increasing interest in exploring social interaction as a learning environment! Along similar lines: We’re preparing a (complementary) challenge that will focus on exploring interaction for post-training, coming with a rich interaction environment to get things started. Stay tuned! 261
David Schlangen @davidschlangen.bsky.social · 03/02/2025Somewhat annoyingly, this doesn't seem to stop others from reinventing this idea again and again, with no indication of whether they found anything lacking (& much indication that they didn't do lit search...). Anyway, stay tuned for our new and upcoming work on using these games for post-training! 030
David Schlangen @davidschlangen.bsky.social · 03/02/2025I'm not on X, so I'll use the opportunity of @karpathy.bsky.social 's post over there to plug our "clembench" project here. We've been doing exactly this--evaluating LLMs w/ conversational games--since early 2023, with several papers out by now (e.g. EMNLP 23). clembench.github.io 1101
David Schlangen @davidschlangen.bsky.social · 17/01/2025So, are we banning social network apps now whose owners potentially try to influence the political discourse in other countries? Asking for a supranational political and economic union. 070
David Schlangen @davidschlangen.bsky.social · 16/01/2025I just randomly found this book on my bookshelf. It must have been transported there from an alternate timeline. “20 years of research on agents”? Preposterous! We all know that the very idea of software agents has only been invented last year by the LLM folks! 1103
David Schlangen @davidschlangen.bsky.social · 31/12/2024So my car needed to be towed this morning. It took the guy quite some time to get everything ready. Then the truck broke down. In the end, the tow truck was towed, and I got a new appointment. I think is probably an allegory for something, maybe the ending year 2024, or the coming year 2025. 020
David Schlangen @davidschlangen.bsky.social · 30/12/2024me: I would really like to end this year with inbox zero. also me: I would really like to end this year with cookie jar zero / “pages remaining in the books I’ve started” zero. me again, expert problem solver: *creates IMAP folder “unprocessed emails from 2024”, selects all, moves 625 items* 030
David Schlangen @davidschlangen.bsky.social · 20/12/2024You shouldn’t pretend to be alone in the logical space of reasons. 030
David Schlangen @davidschlangen.bsky.social · 20/12/2024These new models (using “inference time scaling”) bring out what many of us have been saying for a long time, namely that reasoning fundamentally is a discursive process. (What they are missing is that it is an intersubjective, interactive, repairable, and ultimately normative one.) 120
David Schlangen @davidschlangen.bsky.social · 20/12/2024Looking forward to the first lecture of next year, where I can again use this meme I made a couple of years ago and multiply confuse the students in my "intro to NLP" class. (What is an "LP cover"? Who is that person?) 010
David Schlangen @davidschlangen.bsky.social · 04/12/2024Can we discuss how stupid this photo button thing on the new iPhones is? Who thought that minimising the space where you can hold this damn thing without something unwanted happening is a good idea? 000
David Schlangen @davidschlangen.bsky.social · 01/12/2024From here on, you can get fancy: “LLMs as the dream of the collective (un)consciousness”, etc., but the main point is the non-citeability. 000
David Schlangen @davidschlangen.bsky.social · 01/12/2024For a recent talk to a lay audience, I’ve used a metaphor which I think resonated: rely on an LLMs not more than you would rely on a dream. Use it to inspire you to work something out, but don’t be the one who has to say “this was once revealed to me in a dream”. 120