Sign in

bastian bunzeck

@bbunzeck.bsky.social
582 followers 1.2K following 210 posts

wondering how humans and computers learn and use language 👶🧠🗣️🖥️💬 the work is mysterious and important, see bbunzeck.github.io mentally still at @ucsandiego.bsky.social, phd @clausebielefeld.bsky.social

PostsRepliesMedia
Reposted by bastian bunzeck
Supper Mario Broth @mariobrothblog.bsky.social · 02/10/2026
The Switch version of Paper Mario: The Thousand-Year Door includes its own fictional language that is a cipher for English. Amusingly, the misspelled "Game Cuve" graffiti from the original version is translated into the new language... but still reads "Game Cuve" when deciphered.
The Nintendo Switch version of Paper Mario: The Thousand-Year Door contains an in-universe language used for signs and labels that is a cipher (i.e. a one-to-one symbol conversion) for English.
Textures that were in English or unreadable in the original (left) are translated into the cipher in the remake (right). The remake's poster reads "Wanted" and the signboard reads "Signboard".
A famously misspelled "Game Cuve" graffiti appears in Rogueport in the original.
In the remake, it can be seen through this arch from the back.
Analyzing it reveals that it is still misspelled! The designers deliberately kept the misspelling in the translation.
Here is the key to the cipher used.
Source: research/images by Mario Wiki users "LadySophie17" and "Reese Rivers", mariowiki.com/List_of_fictional_languages
151648338
Reposted by bastian bunzeck
Dan Cohen @dancohen.org · 06/10/2026
On the 20th anniversary of Zotero's 1.0 release on October 5, 2006, I wrote a history of the collaboration and years it took to create the popular research tool now used by over 20 million people: newsletter.dancohen.org/archive/the-...
newsletter.dancohen.org
The Slow Formation of Durable Software
How a group of historians took years to imagine Zotero, the research application that would eventually be used by millions
213256
Reposted by bastian bunzeck
Lukas Edman @lukasnlp.bsky.social · 03/10/2026
Great news! Our workshop what-llms-can-not-do.org is going to COLING 2027 in Macau! Hope to see you there!
what-llms-can-not-do.org
What LLMs Can(not) Do
A living survey of benchmarks that compare large language models with humans.
0185
Reposted by bastian bunzeck
Leonie Weissweiler @weissweiler.bsky.social · 01/10/2026
📢I'm hiring a PhD student to work on interpretability for Protein Language Models! 🗓️There are now two weeks left to apply, find out more on the LION Lab website: lionlabnlp.github.io/jobs/ai4pf/
lionlabnlp.github.io
PhD Student Position in Interpretability for Protein Language Models — LION Lab
Fully funded PhD student position at LION Lab, Leipzig University, on interpretability for protein language models.
186
Reposted by bastian bunzeck
A. Feder Cooper @afedercooper.bsky.social · 30/09/2026
Our large-scale study of memorization of books in open-weight LLMs (e.g., Llama, Qwen) will appear at the 2026 Conference on Language Modeling as an oral. We made a website for exploring our results on 200 books and 14 models: books-memorization.github.io
books-memorization.github.io
How much do open-weight LLMs memorize specific books?
Open-weight LLMs memorize books far more than previously believed. Memorization varies by model family, model size, and book. In extreme cases, entire books are memorized, and we can generate them eff...
26336
Reposted by bastian bunzeck
Ethan Mollick @emollick.bsky.social · 30/09/2026
And its a 3-way race again...
1419115
Reposted by bastian bunzeck
Marco @mcognetta.bsky.social · 30/09/2026
Check out the paper on @alphaxiv.org ! www.alphaxiv.org/abs/2609.tok...
alphaxiv.org
Tokenization: A Survey for Modern NLP
While modern language models take raw text as their input and produce raw text as output, they do not operate over text directly. Hidden in the very first step of language model pipelines is...
2193
Reposted by bastian bunzeck
ACL 2027 @aclmeeting.bsky.social · 30/09/2026
🔔CFP for the #ACL2027NLP main conference is up! 2027.aclweb.org/calls/main/ 📝Submission deadline to ARR is Jan 4, 2027.
2027.aclweb.org
Main Conference Papers
ACL 2027 Call for Papers.
076
Reposted by bastian bunzeck
Lilian Weber @lilweb.bsky.social · 28/09/2026
✨New PhD Position!✨ Come and work with us on understanding cognition with a mix of cognitive modelling and ultrasound neuromodulation! More info on the lab & position: mic-lab.de/positions/ph... Official job ad: tinyurl.com/miclab-phd 🙏Please RT - deadline is Oct 14 ‼️
mic-lab.de
Research Associate (m/f/d): Models, Interventions, and Cognition — MIC-Lab
Models and interventions of cognition.
04239
Reposted by bastian bunzeck
Mike Frank @mcxfrank.bsky.social · 25/09/2026
We've just released a new version of childes-db, my lab's interface to the CHILDES database of child language transcripts. It lets you work with CHILDES data from R through a versioned, reproducible API. A few updates 🧵 childes-db.stanford.edu/
Screenshot of the childes-db website: 'A flexible and reproducible interface to CHILDES.' childes-db 2026.1 contains 56,579 transcripts from 9,151 children across 437 corpora, with panels for an R API tutorial and interactive visualizations of mean length of utterance by child age.
12913
Reposted by bastian bunzeck
Kanishka Misra @kanishka.bsky.social · 25/09/2026
Come be my colleague and teach meaning at UT Austin Linguistics!
078
Reposted by bastian bunzeck
DM Howlcroft 👻 @davehowcroft.com · 23/09/2026
I think one of the issues we're seeing is a disconnect between busyness/output and productivity--I think productivity assumes there is value in the product and it will be a while until the majority of "paper shaped objects" are in fact papers.
033
Reposted by bastian bunzeck
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 20/09/2026
Rising paper numbers won’t collapse the scientific system, it’ll just return it to an older state where you only read papers by people you know and trust. Not saying this is a good thing.
3503
Reposted by bastian bunzeck
bergelsonlab @bergelsonlab.bsky.social · 20/09/2026
reviewer 2: so uh, tell me about your filtering process bc that is a truly unexpected result me: ikr??
041
Reposted by bastian bunzeck
LION Lab @lionlab.bsky.social · 17/09/2026
🦁LION Lab is hiring! 🧑‍🔬One fully funded PhD student (TVL-13 100%) for 3 years 💉Topic: Interpretability for Protein Language Models 👥Advised by @weissweiler.bsky.social together with Clara Schoeder 🌍Leipzig, Germany 🔗Apply by Oct 15: lionlabnlp.github.io/jobs/ai4pf/ Please share! #NLProc #NLP
0109
Reposted by bastian bunzeck
Transactions on Machine Learning Research @tmlrorg.bsky.social · 16/09/2026
TMLR has faced a deluge of submissions, necessitating stricter desk rejection policies due to limited reviewer capacity Co-EiC Nihar Shah reached out to authors of 10 papers slated for desk reject. Could they answer questions about their *own* submission? medium.com/@TmlrOrg/ask...
medium.com
Asking Authors About Their Own Papers
By Nihar B. Shah
116965
Reposted by bastian bunzeck
Pranav A @pranav-nlp.bsky.social · 15/09/2026
Now at KONVENS,Barbara Plank starts off with a keynote on how variability and disagreement is useful for NLP tasks.
Barbara plank giving a keynote talk on human label variation
182
Reposted by bastian bunzeck
Leonie Weissweiler @weissweiler.bsky.social · 15/09/2026
#COLING2027 has joined forces with #EACL2027 and #NAACL2027 for the call for tutorials! 📢Joint Call for Tutorial Proposals (EACL/NAACL/COLING) 2027 🗓️Submission deadline: Friday, 2 October 2026 📝Submit via softconf: softconf.com/p/eacl-tutor... 🔗Full call here: 2027.coling-iccl.org/calls/tutori...
2027.coling-iccl.org
Joint Call for Tutorial Proposals 2027
COLING 2027 Tutorials
083
Reposted by bastian bunzeck
Tomás Ryan @tjryan.bsky.social · 13/09/2026
"MIT wants to rebuild education around the things AI can't replace. That means oral exams, semester portfolios, in-person project work, and a required social component in every subject. The report even floats the idea of rethinking grades entirely..." This is the way forward @tcddublin.bsky.social
515449
Reposted by bastian bunzeck
neutral @neutral.zone · 12/09/2026
Frontier Psychiatrist sounds like an Anthropic job posting
623032
Reposted by bastian bunzeck
Processor of Natural Languages @processorofnl.bsky.social · 12/09/2026
hot take: You should not need to be a cryptic crossword solving genius to comprehend the writing to a paper's abstract.
011
Reposted by bastian bunzeck
Open Encyclopedia of Cognitive Science bot @oecs-bot.bsky.social · 11/09/2026
Distinguishing between our species and other animals requires balancing subtle quantitative differences with the unique, complex interplay of our… Uniquely Human Cognition by Daniel B. M. Haun doi.org/10.21428/e2759450.2760bd31 #CognitiveScience
Open Encyclopedia of Cognitive Science: "Uniquely Human Cognition" by Daniel B. M. Haun. Uniquely human cognition refers to aspects of cognition that are specific to our evolutionary lineage and distinguish us from other animals, including our closest living relatives, the nonhuman primat
086
Reposted by bastian bunzeck
Lisa Bylinina @bylinina.bsky.social · 11/09/2026
look at my new babylm paper! arxiv.org/abs/2609.11870 basically, i initialize token embeddings with representations from an image encoder rather than randomly and then train text-only as usual. kind of a visual demonstration to start off the word learning process
1155
Reposted by bastian bunzeck
Jake Quilty-Dunn @quiltydunn.bsky.social · 11/09/2026
New paper, co-authored with Justin Wood, in @cp-trendscognsci.bsky.social: The Nativist-Empiricist Debate Is Broken We argue that radical nativism and radical empiricism are compatible and, in some cases, both plausibly true. So something is wrong. 🧵 Author share link here, valid til Oct 31:
authors.elsevier.com
Please wait whilst we redirect you
All content on this site: Copyright © 2026 Elsevier B.V., its licensors, and contributors. All rights are reserved, including those for text and data mining, AI training, and similar technologies. For all open access content, the relevant licensing terms apply.
17630
Reposted by bastian bunzeck
ACL 2027 @aclmeeting.bsky.social · 11/09/2026
ACL Sustainable Reviewing Policy: We are introducing changes in the @ReviewAcl reviewing and submissions. The changes will cap authors and introduce changes in the reviewing to keep our community sustainable. #NLProc #NLP
12712
bastian bunzeck @bbunzeck.bsky.social · 09/09/2026
Criticised the over-use of Claude-isms in 3/4 ARR papers this cycle 🥲
060
Reposted by bastian bunzeck
Kanishka Misra @kanishka.bsky.social · 09/09/2026
Understanding how learners conclude “X laughed Y” is incorrect is an age-old question, with several hypotheses, some of which have been ~impossible to disentangle! @tomyxw.bsky.social, @fredashi.bsky.social, and I use controlled rearing to shed light on this in our new EMNLP paper: 1/n
Title slide for “Disentangling Statistical Preemption from Entrenchment in Language Models’ Avoidance of Overgeneralization,” by Yixuan Wang, Freda Shi, and Kanishka Misra. Includes the main results plot and experimental design.
2179
Reposted by bastian bunzeck
Eugene Vinitsky 🍒 @eugenevinitsky.bsky.social · 08/09/2026
Cannot believe there’s a tremendous mathematical result and instead of being exciting it’s annoying
622310
Reposted by bastian bunzeck
Kanishka Misra @kanishka.bsky.social · 02/09/2026
My current opinion on semantic cognition x LMs will appear in Current Opinion in Behavioral Sciences' Special issue on AI and Psychology, edited by the incredible @neuranna.bsky.social and @stepalminteri.bsky.social Catch the updated version: osf.io/preprints/ps... Grateful to both reviewers!
Title page of “Semantic Cognition for and from Language Models” by me and an arrow pointing downwards to the fact that it will appear in Current Opinion in Behavioral Sciences!
0171
Reposted by bastian bunzeck
Daniel Duran @simphon.bsky.social · 02/09/2026
Thanks to @stefanhartmann.bsky.social and @bbunzeck.bsky.social for organizing the session “Computational approaches to language dynamics” at #DGKL2026 ! It was great to attend those fascinating presentations, have interesting discussions and also see that computational modelling is well and alive!
uni-bielefeld.de
DGKL2026 - Universität Bielefeld
This is the webpage of the DGKL conference 2026
042
Reposted by bastian bunzeck
Tom McCoy @rtommccoy.bsky.social · 01/09/2026
🤖🧠NEW PAPER🧠🤖 (The result of an 8-year project!) LLMs seem very different from symbolic systems. Yet LLMs excel in symbolic domains (e.g., language/code/math). How do they do it? Our finding: LLM representations have implicit symbolic structure! Link in thread ⬇️ 1/n
Overview of the paper. 
Title: The Emergent Symbolic Structure of Artificial Neural Networks
Authors: Tom McCoy, Paul Soulos, Tal Linzen, Paul Smolensky
Left: Neural networks encode information in vectors (there is then an image of a vector), yet they excel at tasks long thought to require symbolic structure (there is then an image of a symbolic representation, specifically a syntax tree). How do LLMs do it?
Right: We find that LLM representations can be closely approximated with symbolic structures. This approximation lets us edit the structure of an LLM’s output by editing the structure of its internal representations, as shown. There is then an image of two edits to LLMs. In the first one, the original input is 3 + 6 * 8, with an answer of 51. But if we swap the positions of the 3 and the 6, the output becomes 30. In the second one, the original input is a Python command repeating the list [Z, U] three times, producing [Z, U, Z, U, Z, U]. But if we edit the input in a way that adds a Q at the end of the input, the output becomes [Z, U, Q, Z, U, Q, Z, U, Q].
431988
Reposted by bastian bunzeck
Sean Trott @seantrott.bsky.social · 31/08/2026
New paper out in @openmindjournal.bsky.social! "Large Language Models as Distributional Baselines for Language Tasks". With @jamichaelov.bsky.social , @camrobjones.bsky.social , Tyler Chang, and Ben Bergen. direct.mit.edu/opmi/article...
direct.mit.edu
Large Language Models as Distributional Baselines for Language Tasks
Abstract. In order to ask questions about the mechanisms underpinning human cognition, researchers must control for properties of stimuli that could confound detected effects. In experiments involving...
1126
Reposted by bastian bunzeck
DGKL-GCLA @dgkl-gcla.bsky.social · 31/08/2026
Hello Bielefeld! 👋 #DGKL2026 is underway! That also means it's time for a general assembly and an election of a new DGKL/GLCA board. 🗳️ We are still looking for new board members. 🤓 Join us tonight at 6.00 pm and participate!
053
Reposted by bastian bunzeck
Daniel Duran @simphon.bsky.social · 31/08/2026
Bastian Bunzeck presenting Universal Dependencies annotations of constructions in child-caretaker dialogues at the 11th International Conference of the German Cognitive Linguistics Association #DGKL in Bielefeld. @bbunzeck.bsky.social‬ #LINCC #CRC1646
061
bastian bunzeck @bbunzeck.bsky.social · 30/08/2026
We did stop in Hamm in the end, but only after passing Altenbeken and Paderborn (both of which are closer to Bielefeld) without stopping 🥲 the DB works in mysterious ways…
020
bastian bunzeck @bbunzeck.bsky.social · 30/08/2026
What the picture is not showing: our stop in Bielefeld was cancelled, so nobody knows if and when I’ll get there tonight. Gracias Deutsche Bahn 🥲
240
bastian bunzeck @bbunzeck.bsky.social · 30/08/2026
Excited to finally be back in Bielefeld for @dgkl2026.bsky.social! Me and @stefanhartmann.bsky.social are organising a theme session on »Computational approaches« featuring @anthe.sevenants.net, @simphon.bsky.social and many others (that I did not find on here!). 🤖💬
1112
Reposted by bastian bunzeck
Morten H. Christiansen @mh-christiansen.bsky.social · 29/08/2026
I think that your article does a great job at mentioning much of the additional input kids get and how researchers are trying to incorporate that into AI models. My point was simply that given that, current models get rather impoverished input and do not learn from actual interaction (like children)
031
Reposted by bastian bunzeck
Anthony Moser @anthonymoser.com · 16/04/2026
are you there claude? it's me, margaret
210818
Reposted by bastian bunzeck
Xiulin Yang @xiulinyang.bsky.social · 28/08/2026
🍎🍊 How would you know if a language model is better at one language than another? Our #EMNLP2026 paper argues that only one metric can actually lead to fair crosslingual evaluation. This work is a collaboration with @wegotlieb.bsky.social & @catherinearnett.bsky.social! (1/5)
13312
bastian bunzeck @bbunzeck.bsky.social · 27/08/2026
When I lived in SD every aisle in my local Ralph‘s had 10x the variety of a well-stocked German supermarket, but the vegan section was half as big as in any small German discount store. I’m not vegan but I was very surprised. Especially in California of all places
010
Reposted by bastian bunzeck
mr. TIM @timkellogg.me · 27/08/2026
Apparently in the next version of Fable, Anthropic changed their tokenizer based on real world usage and made “Let me check rather than answer from memory” into a single token, reportedly saving 43% of all token generation very interesting tinyurl.com/58zf9d3c
1428422
Reposted by bastian bunzeck
Tal Linzen @tallinzen.bsky.social · 26/08/2026
Thanks to the National Science Foundation for featuring our new @pnas.org paper, and for the continued support for basic science projects like this one!
1183
Reposted by bastian bunzeck
Matt Henderson @matthen.com · 26/08/2026
To tell if a maze is solvable, just hang it by its corners If it tears into pieces, you’ve found a solution
694154868
Reposted by bastian bunzeck
Rachel Courtland @rachelcourtland.com · 24/08/2026
I've been waiting for this feature to post online! For our Kids issue, @elisecutts.bsky.social explores a big open question: why children seem to acquire language so much more efficiently than LLMs. ter.li/0lqpwc1w
ter.li
Kids outlearn AI—and we still don’t know why
LLMs need vastly more data than children to learn language. Understanding why could help us create more efficient models—and reveal more about developing minds.
294
bastian bunzeck @bbunzeck.bsky.social · 24/08/2026
Great slide by @kanishka.bsky.social on how he envisions the link between machine cognition and human cognition. 🤖👶
171
Reposted by bastian bunzeck
Joshua Raclaw @joshuaraclaw.com · 22/08/2026
Optimist: The cup is half full Pessimist: The cup is half empty Linguist: What even is a cup, exactly
image from labov's 1973ish study of cups and bowls and semantic boundaries, a bunch of weird cuppy thing are illustrated including a mug with a loooong fluted stem and a triangle mugneat bar graph of cuppiness and bowliness
617940
Reposted by bastian bunzeck
Nora Newcombe @noranewcombe.bsky.social · 22/08/2026
Optimist: The cup is half full Pessimist: The cup is half empty Developmentalist: Curious if the cup started out full and is emptying or started out empty and is filling.
2526
bastian bunzeck @bbunzeck.bsky.social · 21/08/2026
I’ve called so many people on accident while trying to copy their number 🥲
010
bastian bunzeck @bbunzeck.bsky.social · 20/08/2026
Me and some other participants were wondering if this incredible cross-table can also be found in one of your publications, or if it is summer school exclusive? 😇
110