Sign in

Siyuan Song

@siyuansong.bsky.social
192 followers 363 following 46 posts

Grad student@Princeton Psychology siyuansong.site Language, Learning, Intelligence Prev: Undergrad@UTexas, SJTU Summer Research Visit @MIT BCS, Harvard Psych Opinions are my own.

PostsRepliesMedia
Reposted by Siyuan Song
Simon Schug @smonsays.bsky.social · 17/09/2026
Human thought is thought to be systematic: Do reasoning models have systematicity of thought? With @brendenlake.bsky.social we study this question in our new preprint arxiv.org/abs/2609.13948 Thread 🧵
14014
Reposted by Siyuan Song
Mike Frank @mcxfrank.bsky.social · 27/04/2026
For a year and a half, @carorowland.bsky.social, @lehersingh.bsky.social, Marisa Casillas, Shanley Allen, and I have been meeting to discuss whether innateness is still a useful concept to think about in studying language acquisition. Here's our take: osf.io/preprints/ps...
Title pageThe origins of language have been a persistent object of philosophical and scientific inquiry, in part
because they offer a window into the origins of thought. Classic innatist theories of language have argued
that languages share universal elements that arise because language acquisition is guided by rich,
biologically specified structures in the form of a universal grammar. However, different versions of this
hypothesis that posit substantial amounts of innate content are not consistent with recent evidence. We
review new evidence from language diversity, human interaction, and large language models (LLMs), all
of which challenge classical innatist views. Yet we believe there is still value in asking what elements of
language development are innate. In this Perspective, we integrate evidence from these three areas
towards a broader conception of innateness, one that seeks to explain both consistency and variation in
language acquisition. On this view, innateness remains a key orienting principle for understanding the
origins of language.
66128
Reposted by Siyuan Song
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 24/04/2026
Chinese babyLM is on folks:
061
Siyuan Song @siyuansong.bsky.social · 24/04/2026
🚀 Announcing the Chinese BabyLM Challenge: the first shared task on data-efficient pretraining for Chinese. 📍 Co-located with NLPCC 2026 (Nov 3–5, Macau🇨🇳🇲🇴) Can you train a strong Chinese LM on just ~100M words? chinese-babylm.github.io 🧵 👇(1/6)
chinese-babylm.github.io
Chinese BabyLM Challenge
1114
Siyuan Song @siyuansong.bsky.social · 25/03/2026
Just arrived in Boston for #HSP2026! I'll be presenting my work with @thomashikaru.bsky.social on error sensitivity in next-word predictions of humans and LMs — Friday 12:10–2:00pm poster session. Come say hi!
181
Reposted by Siyuan Song
Martin Zettersten @mzettersten.bsky.social · 12/03/2026
I'm hiring a new lab manager for my lab @ UCSD! For more info on the lab, check out our website: lillab.ucsd.edu Target start date is June 1 (flexible) and application deadline is March 26. Please share with anyone you think might be a good fit! Apply here: employment.ucsd.edu/laboratory-c...
employment.ucsd.edu
Laboratory Coordinator - 138788
Laboratory Coordinator - 138788 | Careers at UC San Diego
03732
Reposted by Siyuan Song
Kanishka Misra @kanishka.bsky.social · 10/03/2026
What is the interplay between representations learned from (language) surface forms alone, and those learned from more grounded evidence (e.g.,vision)? Excited to share new work understanding “Cross-modal taxonomic generalization” in (V)LMs arxiv.org/abs/2603.07474 1/
title section of the paper: “Cross-Modal Taxonomic Generalization in (Vision) Language Models” by Tianyang Xu, Marcelo Sandoval-Castañeda, Karen Livescu, Greg Shakhnarovich, Kanishka Misra.
13311
Reposted by Siyuan Song
Harvey Lederman @harveylederman.bsky.social · 06/03/2026
Can large language models *introspect*? In a new paper, @kmahowald.bsky.social and I study the MECHANISM of introspection in big open-source models. tldr: Models detect internal anomalies through DIRECT ACCESS, but don't know what the anomalies are. And they love to guess “apple” 🍎
27115
Reposted by Siyuan Song
Kanishka Misra @kanishka.bsky.social · 27/01/2026
“All bears have a property”, “Some bears have a property”, “Bears have a property” are different in terms of how the property is generalized to a specific bear – a great example of how language constrains thought! This holds for kids, adults, and according to our new work, (V)LMs! 🧵
Title page of our paper: "Bears, all bears, and some bears. Language Constraints on Language Models' Inductive Inferences"
1269
Reposted by Siyuan Song
Juan Diego Rodriguez @juand-r.bsky.social · 22/01/2026
Our first South by Semantics lecture of the semester at UT Austin is happening next week on January 30th! I'm excited to hear Dr. Amir Zeldes (Associate Professor at Georgetown University) talk about saliency in discourse and the memorability of salient information for both humans and LLMs.
031
Reposted by Siyuan Song
Leonie Weissweiler @weissweiler.bsky.social · 11/12/2025
🧑‍🔬I’m recruiting PhD students in Natural Language Processing @unileipzig.bsky.social Computer Science, together with @scadsai.bsky.social! Topics include, but aren’t limited to: 🔎Linguistic Interpretability 🌍Multilingual Evaluation 📖Computational Typology Please share! #NLProc #NLP
14225
Reposted by Siyuan Song
Mike Frank @mcxfrank.bsky.social · 01/12/2025
Can we use VLMs to quantify multimodal alignment in children's experiences? We analyze a large corpus of headcam videos to find out! New preprint from our BabyView project, led by @alvinwmtan.bsky.social and Jane Yang: arxiv.org/abs/2511.18824
Figure 1 showing alignment pipeline using CLIP models on BabyView data.Figure 2: human judgments are correlated with CLIP scores.
1295
Reposted by Siyuan Song
James Michaelov @jamichaelov.bsky.social · 01/12/2025
Looking forward to #NeurIPS25 this week 🏝️! I'll be presenting at Poster Session 3 (11-2 on Thursday). Feel free to reach out!
0103
Reposted by Siyuan Song
Juan Diego Rodriguez @juand-r.bsky.social · 01/12/2025
I’m excited to present SimpleStories at EurIPS! Also if anyone at #EurIPS is interested in chatting about LLM data efficiency, interpretability, model inconsistency or other topics feel free to DM me. Dataset and models: lnkd.in/e_VGWqhP Code: lnkd.in/eEidmv74 Paper: lnkd.in/eH6jS9uY
2184
Siyuan Song @siyuansong.bsky.social · 10/11/2025
String probability might be the best tool for assessing LMs' grammatical knowledge, yet it does not directly tell you 'how grammatical' a string is. Here's why and how we should use string probability and minimal pairs: Excited to see this out - it's my great honor to be part of this amazing team!
030
Reposted by Siyuan Song
Kyle Mahowald @kmahowald.bsky.social · 10/11/2025
Oh cool! Excited this LM + construction paper was SAC-Highlighted! Check it out to see how LM-derived measures of statistical affinity separate out constructions with similar words like "I was so happy I saw you" vs "It was so big it fell over".
0174
Reposted by Siyuan Song
Kyle Mahowald @kmahowald.bsky.social · 07/11/2025
Delighted Sasha's (first year PhD!) work using mech interp to study complex syntax constructions won an Outstanding Paper Award at EMNLP! Also delighted the ACL community continues to recognize unabashedly linguistic topics like filler-gaps... and the huge potential for LMs to inform such topics!
aclanthology.org
1338
Reposted by Siyuan Song
Jennifer Hu @jennhu.bsky.social · 04/11/2025
Interested in doing a PhD at the intersection of human and machine cognition? ✨ I'm recruiting students for Fall 2026! ✨ Topics of interest include pragmatics, metacognition, reasoning, & interpretability (in humans and AI). Check out JHU's mentoring program (due 11/15) for help with your SoP 👇
02815
Reposted by Siyuan Song
Linyang He @linyanghe.bsky.social · 30/10/2025
🧠 New at #NeurIPS2025! 🎵 We're far from the shallow now🎵 TL;DR: We introduce the first "reasoning embedding" and uncover its unique spatio-temporal pattern in the brain. 🔗 arxiv.org/abs/2510.228...
184
Reposted by Siyuan Song
Multilingual Representation Workshop @ EMNLP 2026 @mrl-workshop.bsky.social · 29/10/2025
Introducing Global PIQA, a new multilingual benchmark for 100+ languages. This benchmark is the outcome of this year’s MRL shared task, in collaboration with 300+ researchers from 65 countries. This dataset evaluates physical commonsense reasoning in culturally relevant contexts.
12210
Reposted by Siyuan Song
Harvey Lederman @harveylederman.bsky.social · 24/10/2025
Very excited to be going to Chicago for @agnescallard.bsky.social's famous Night Owls next week! I'll be discussing my essay "ChatGPT and the Meaning of Life". Hope to see you there if you're local!
141
Reposted by Siyuan Song
Kanishka Misra @kanishka.bsky.social · 21/10/2025
If I spill the tea—“Did you know Sue, Max’s gf, was a tennis champ?”—but then if you reply “They’re dating?!” I’d be a bit puzzled, since that’s not the main point! Humans can track what’s ‘at issue’ in conversation. How sensitive are LMs to this distinction? New paper w/ @sangheekim.bsky.social!
Title of our paper: “Hey, wait a minute: on at-issue sensitivity in Language Models” by Sanghee Kim and Kanishka Misra.

Below: A person says “Sue, Max’s girlfriend, was a tennis champ!”; a second person responds with “What racket does she use?” (which targets at-issue content); a third person replies with “They’re dating?” (which targets not at-issue content)
3354
Reposted by Siyuan Song
Ethan Gotlieb Wilcox @wegotlieb.bsky.social · 21/10/2025
I will be recruiting PhD students via Georgetown Linguistics this application cycle! Come join us in the PICoL (pronounced “pickle”) lab. We focus on psycholinguistics and cognitive modeling using LLMs. See the linked flyer for more details: bit.ly/3L3vcyA
22914
Reposted by Siyuan Song
Kanishka Misra @kanishka.bsky.social · 16/10/2025
"Although I hate leafy vegetables, I prefer daxes to blickets." Can you tell if daxes are leafy vegetables? LM's can't seem to! 📷 We investigate if LMs capture these inferences from connectives when they cannot rely on world knowledge. New paper w/ Daniel, Will, @jessyjli.bsky.social
Title page of the paper: WUGNECTIVES: Novel Entity Inferences of Language Models from Discourse Connectives, with two figures at the bottom

Left: Our figure 1 -- comparing previous work, which usually predicted the connective given the arguments (grounded in the world); our work flips this premise by getting models to use their knowledge of connectives to predict something about the world.

Right: Our main results across 7 types of connective senses. Models are especially bad at Concession connectives.
2329
Siyuan Song @siyuansong.bsky.social · 15/10/2025
Honored to get the chance to contribute to the Chinese dataset! And had a great time working with all the awesome collaborators!
100
Reposted by Siyuan Song
Juan Diego Rodriguez @juand-r.bsky.social · 06/10/2025
Excited to present this at COLM tomorrow! (Tuesday, 11:00 AM poster session)
032
Reposted by Siyuan Song
Sasha Boguraev @sashaboguraev.bsky.social · 06/10/2025
I will be giving a short talk on this work at the COLM Interplay workshop on Friday (also to appear at EMNLP)! Will be in Montreal all week and excited to chat about LM interpretability + its interaction with human cognition and ling theory.
084
Reposted by Siyuan Song
Kanishka Misra @kanishka.bsky.social · 06/10/2025
Traveling to my first @colmweb.org🍁 Not presenting anything but here are two posters you should visit: 1. @qyao.bsky.social on Controlled rearing for direct and indirect evidence for datives (w/ me, @weissweiler.bsky.social and @kmahowald.bsky.social), W morning Paper: arxiv.org/abs/2503.20850
arxiv.org
Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models
Language models (LMs) tend to show human-like preferences on a number of syntactic phenomena, but the extent to which these are attributable to direct exposure to the phenomena or more general propert...
1135
Reposted by Siyuan Song
Jessy Li @jessyjli.bsky.social · 06/10/2025
On my way to #COLM2025 🍁 Check out jessyli.com/colm2025 QUDsim: Discourse templates in LLM stories arxiv.org/abs/2504.09373 EvalAgent: retrieval-based eval targeting implicit criteria arxiv.org/abs/2504.15219 RoboInstruct: code generation for robotics with simulators arxiv.org/abs/2405.20179
0124
Reposted by Siyuan Song
Kyle Mahowald @kmahowald.bsky.social · 06/10/2025
I’m at #COLM2025 from Wed with: @siyuansong.bsky.social Tue am introspection arxiv.org/abs/2503.07513 @qyao.bsky.social Wed am controlled rearing: arxiv.org/abs/2503.20850 @sashaboguraev.bsky.social INTERPLAY ling interp: arxiv.org/abs/2505.16002 I’ll talk at INTERPLAY too. Come say hi!
arxiv.org
Language Models Fail to Introspect About Their Knowledge of Language
There has been recent interest in whether large language models (LLMs) can introspect about their own internal states. Such abilities would make LLMs more interpretable, and also validate the use of s...
1206
Siyuan Song @siyuansong.bsky.social · 06/10/2025
Heading to #COLM2025 to present my first paper w/ @jennhu.bsky.social @kmahowald.bsky.social ! When: Tuesday, 11 AM – 1 PM Where: Poster #75 Happy to chat about my work and topics in computational linguistics & cogsci! Also, I'm on the PhD application journey this cycle! Paper info 👇:
073
Reposted by Siyuan Song
Tom McCoy @rtommccoy.bsky.social · 30/09/2025
🤖 🧠 NEW BLOG POST 🧠 🤖 What skills do you need to be a successful researcher? The list seems long: collaborating, writing, presenting, reviewing, etc But I argue that many of these skills can be unified under a single overarching ability: theory of mind rtmccoy.com/posts/theory...
Illustration of the blog post's main argument, summarized as: "Theory of Mind as a Central Skill for Researchers: Research involves many skills.If each skill is viewed separately, each one takes a long time to learn. These skills can instead be connected via theory of mind – the ability to reason about the mental states of others. This allows you to transfer your abilities across areas, making it easier to gain new skills."
2202
Reposted by Siyuan Song
Kanishka Misra @kanishka.bsky.social · 30/09/2025
The compling group at UT Austin (sites.utexas.edu/compling/) is looking for PhD students! Come join me, @kmahowald.bsky.social, and @jessyjli.bsky.social as we tackle interesting research questions at the intersection of ling, cogsci, and ai! Some topics I am particularly interested in:
Picture of the UT Tower with "UT Austin Computational Linguistics" written in bigger font, and "Humans processing computers processing human processing language" in smaller font
3189
Reposted by Siyuan Song
Jessy Li @jessyjli.bsky.social · 25/09/2025
Can AI aid scientists amidst their own workflows, when they do not know step-by-step workflows and may not know, in advance, the kinds of scientific utility a visualization would bring? Check out @sebajoe.bsky.social’s feature on ✨AstroVisBench:
083
Reposted by Siyuan Song
Harvey Lederman @harveylederman.bsky.social · 24/09/2025
Simon Goldstein and I have a new paper, “What does ChatGPT want? An interpretationist guide”. The paper argues for three main claims. philpapers.org/rec/GOLWDC-2 1/7
philpapers.org
Simon Goldstein & Harvey Lederman, What Does ChatGPT Want? An Interpretationist Guide - PhilPapers
This paper investigates LLMs from the perspective of interpretationism, a theory of belief and desire in the philosophy of mind. We argue for three conclusions. First, the right object of study ...
2236
Reposted by Siyuan Song
Naomi Saphra @nsaphra.bsky.social · 24/09/2025
I did a QA with Quanta about interpretability and training dynamics! I got to talk about a bunch of research hobby horses and how I got into them.
26512
Reposted by Siyuan Song
Andrew Lampinen @lampinen.bsky.social · 22/09/2025
Why does AI sometimes fail to generalize, and what might help? In a new paper (arxiv.org/abs/2509.16189), we highlight the latent learning gap — which unifies findings from language modeling to agent navigation — and suggest that episodic memory complements parametric learning to bridge it. Thread:
arxiv.org
Latent learning: episodic memory complements parametric learning by enabling flexible reuse of experiences
When do machine learning systems fail to generalize, and what mechanisms could improve their generalization? Here, we draw inspiration from cognitive science to argue that one weakness of machine lear...
15710
Reposted by Siyuan Song
Stefan Frank @stefanfrank.bsky.social · 19/09/2025
Announcing the first (and perhaps only) Multilingual Minds and Machines Meeting! Come join us in Nijmegen, June 22-23, 2026, if you are interested in computational models of human multilingualism: mmmm2026.github.io
0117
Reposted by Siyuan Song
Catherine Arnett @catherinearnett.bsky.social · 19/09/2025
Did you know? ❌77% of language models on @hf.co are not tagged for any language 📈For 95% of languages, most models are multilingual 🚨88% of models with tags are trained on English In a new blog post, @tylerachang.bsky.social and I dig into these trends and why they matter! 👇
1132
Reposted by Siyuan Song
Brenden Lake @brendenlake.bsky.social · 08/09/2025
Our new lab for Human & Machine Intelligence is officially open at Princeton University! Consider applying for a PhD or Postdoc position, either through Computer Science or Psychology. You can register interest on our new website lake-lab.github.io (1/2)
15414
Reposted by Siyuan Song
Conference on Language Modeling @colmweb.org · 26/08/2025
COLM 2025 accepted submissions are now public: openreview.net/group?id=col... Congratulations to all the authors, and see you all in Montreal!
openreview.net
COLM 2025 Conference
Welcome to the OpenReview homepage for COLM 2025 Conference
061
Reposted by Siyuan Song
Kyle Mahowald @kmahowald.bsky.social · 26/08/2025
Can AI introspect? Surprisingly tricky to define what that means! And also interesting to test. New work from @siyuansong.bsky.social, @harveylederman.bsky.social, @jennhu.bsky.social and me on introspection in LLMs. See paper and thread for a definition and some experiments!
0111
Reposted by Siyuan Song
Jennifer Hu @jennhu.bsky.social · 26/08/2025
Can AI models introspect? What does introspection even mean for AI? We revisit a recent proposal by Comșa & Shanahan, and provide new experiments + an alternate definition of introspection. Check out this new work w/ @siyuansong.bsky.social, @harveylederman.bsky.social, & @kmahowald.bsky.social 👇
1215
Reposted by Siyuan Song
Harvey Lederman @harveylederman.bsky.social · 26/08/2025
exciting new paper from Siyuan! I really enjoyed working with him on this, inspired by important work by Murray Shanahan and Julia Comsa. Hard questions about how to operationalize the notion of “introspection” that’s relevant for practical applications in AI today. Hope you’ll check it out!
062
Siyuan Song @siyuansong.bsky.social · 26/08/2025
How reliable is what an AI says about itself? The answer depends on whether models can introspect. But, if an LLM says its temperature parameter is high (and it is!)….does that mean it’s introspecting? Surprisingly tricky to pin down. Our paper: arxiv.org/abs/2508.14802 (1/n)
1162
Reposted by Siyuan Song
Leshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 25/08/2025
A dataset of ancient Chinese writings to study with LLMs, including 170K sentences for pretraining With 10K words, mapping to modern word (when applicable) There are so many fascinating questions out there www.arxiv.org/abs/2508.15791
092
Reposted by Siyuan Song
David Bau @davidbau.bsky.social · 18/08/2025
This Friday NEMI 2025 is at Northeastern in Boston, 8 talks, 24 roundtables, 90 posters; 200+ attendees. Thanks to goodfire.ai/ for sponsoring! nemiconf.github.io/summer25/ If you can't make it in person, the livestream will be here: www.youtube.com/live/4BJBis...
youtube.com
New England Mechanistic Interpretability Workshop
About:The New England Mechanistic Interpretability (NEMI) workshop aims to bring together academic and industry researchers from the New England and surround...
1167
Reposted by Siyuan Song
evolangconf.bsky.social @evolangconf.bsky.social · 07/08/2025
Call for papers for Evolang 2026 is now posted! evolang2026.org. Workshop proposals due Sept 22nd. Paper submissions October 26th. Evolang is interdisciplinarity at its best. We hope to see you in Plovdiv!!
02214
Reposted by Siyuan Song
Andrew Lampinen @lampinen.bsky.social · 05/08/2025
In neuroscience, we often try to understand systems by analyzing their representations — using tools like regression or RSA. But are these analyses biased towards discovering a subset of what a system represents? If you're interested in this question, check out our new commentary! Thread:
What do representations tell us about a system? Image of a mouse with a scope showing a vector of activity patterns, and a neural network with a vector of unit activity patterns
Common analyses of neural representations: Encoding models (relating activity to task features) drawing of an arrow from a trace saying [on_____on____] to a neuron and spike train. Comparing models via neural predictivity: comparing two neural networks by their R^2 to mouse brain activity. RSA: assessing brain-brain or model-brain correspondence using representational dissimilarity matrices
617453
Reposted by Siyuan Song
Hope Kean @hopekean.bsky.social · 03/08/2025
Is the Language of Thought == Language? A Thread 🧵 New Preprint (link: tinyurl.com/LangLOT) with @alexanderfung.bsky.social, Paris Jaggers, Jason Chen, Josh Rule, Yael Benn, @joshtenenbaum.bsky.social, ‪@spiantado.bsky.social‬, Rosemary Varley, @evfedorenko.bsky.social 1/8
tinyurl.com
Evidence from Formal Logical Reasoning Reveals that the Language of Thought is not Natural Language
Humans are endowed with a powerful capacity for both inductive and deductive logical thought: we easily form generalizations based on a few examples and draw conclusions from known premises. Humans al...
68133