Sign in

Jessy Li

@jessyjli.bsky.social
2.5K followers 472 following 69 posts

jessyli.com Associate Professor, UT Austin Linguistics. Part of UT Computational Linguistics sites.utexas.edu/compling and UT NLP www.nlp.utexas.edu

PostsRepliesMedia
Reposted by Jessy Li
Kyle Mahowald @kmahowald.bsky.social · 02/07/2026
The full BBS treatment from me and @futrell.bsky.social on "How linguistics learned to stop worrying and love the LMs" is now out, with all the commentaries and our response. If you "Save PDF", it will give you the whole target article + commentary + response pdf: www.cambridge.org/core/journal...
cambridge.org
How linguistics learned to stop worrying and love the language models | Behavioral and Brain Sciences | Cambridge Core
How linguistics learned to stop worrying and love the language models - Volume 49
2359
Jessy Li @jessyjli.bsky.social · 30/06/2026
I will miss #ACL2026 this year, but check out work from my students and collaborators! Kaijie (@kaijie-mo.bsky.social), Sebastian (@sebajoe.bsky.social), Asher (@asher-zheng.bsky.social), Lily, and Gauri will be there presenting the following: jessyli.com/acl2026
0101
Reposted by Jessy Li
NSF-Simons AI Institute for Cosmic Origins (CosmicAI) @nsfsimonscosmicai.bsky.social · 29/06/2026
CosmicAI Leadership highlight! Explorable Universe AI Dr. @jessyjli.bsky.social is an Associate Professor, Linguistics, at UT Austin. Listen as she describes her research on AI for science and its evaluation.
061
Jessy Li @jessyjli.bsky.social · 15/06/2026
Want to know *how* novel/creative your LLM response is and why? Use GENIE! 🧞
040
Jessy Li @jessyjli.bsky.social · 13/06/2026
Introducing Hero’s Journey, meticulously designed to test inductive generalization in a fun text game. ⚖️Verdict: all LLMs we tested trail far behind humans when induction involves generalization across procedures!
020
Reposted by Jessy Li
The Data Therapist in the Blue Sky @datatherapist.bsky.social · 11/06/2026
New profession just dropped: pharma-morphologist #linguistics #NLP #morphology
1101
Jessy Li @jessyjli.bsky.social · 11/06/2026
New work on LLM safety in medicine 💊! Drug names have fixed morphological structure, so is that exploited by LLMs? Kaijie's paper reveals over-generalization that may pose safety risks if not handled carefully 👇
151
Reposted by Jessy Li
David Jurgens @davidjurgens.bsky.social · 26/05/2026
I'm on a new committee reviewing ARR's Responsible NLP Checklist and looking to potentially make changes. I'd love to hear others' thoughts on what is working well or needs revised, especially given it might be fresh in memory from the recent ARR cycle. 1/
4176
Reposted by Jessy Li
aaai.org @aaai.org · 07/05/2026
We are thrilled to present a detailed report describing the system built for the AAAI-26 AI review pilot, the survey results, and a new benchmark that was created to assess the capabilities of the system. Read the full article: arxiv.org/pdf/2604.13940
arxiv.org
142
Reposted by Jessy Li
Kanishka Misra @kanishka.bsky.social · 11/05/2026
New opinion piece on the interface between research on concepts and categories in minds vs. in neural network LMs! I take the position that there is much to be learned from this interface (e.g., learning about concepts from language alone) and outline some directions for future.
Title page of "Semantic Cognition for and from Language Models" followed by a figure showing tests that target conceptual structure and content vs. those that target function.
25712
Reposted by Jessy Li
NSF-Simons AI Institute for Cosmic Origins (CosmicAI) @nsfsimonscosmicai.bsky.social · 08/05/2026
CosmicAI personnel contributed to the AAAI-26 AI review pilot, which generated automatic AI reviews of all research papers submitted to the conference’s main track. The AI reviews complemented human reviews. @mattlease.bsky.social @jessyjli.bsky.social @sebajoe.bsky.social Joydeep Biswas
011
Jessy Li @jessyjli.bsky.social · 08/05/2026
QUDs going multimodal! With MQUD, we can train models to generate scientific questions❓that are inquisitive and insightful enough to be answered in the scientific paper! Huge thanks to the many paper authors who contributed to our data. Check out @yatingwu.bsky.social’s work:
010
Reposted by Jessy Li
Hongli Zhan @hongli-zhan.bsky.social · 03/05/2026
1k+ downloads each on the MINT empathy models since release 🔥 Encouraging to see the interest in our work! tl;dr: In multi-turn empathic dialogue, LLMs reuse the same discourse moves far more often than humans do; MINT uses RL to diversify them. Give it a try!👇 huggingface.co/hongli-zhan/...
062
Reposted by Jessy Li
Adina Williams @adinawilliams.bsky.social · 03/05/2026
I had a fantastic time at the 2026 Harrington Symposium this week at UT Austin. It was wonderful to be able to dig into more science of AI with brilliant researchers across many specialties and viewpoints! Many things to think about! harrington.utexas.edu/faculty-fell...
harrington.utexas.edu
Harrington Faculty Fellows Symposium 2026
081
Reposted by Jessy Li
NSF-Simons AI Institute for Cosmic Origins (CosmicAI) @nsfsimonscosmicai.bsky.social · 01/05/2026
CosmicAI at our collaborative event - Harrington Faculty Fellows Symposium 2026 “Large Language Models: Advances and Applications.” Sign up for our email list to learn about our upcoming events! cosmicai.org/get-involved @jessyjli.bsky.social
011
Jessy Li @jessyjli.bsky.social · 29/04/2026
In multi-turn conversation, LLMs tend to repeat the same kind of things over and over again. They could have different words, but we found them to be the *same discourse moves*! Introducing @hongli-zhan.bsky.social’s new work: novel discourse-level diversity rewards in post-training:
0133
Reposted by Jessy Li
Oskar 🕊️ @austegard.com · 09/04/2026
Nice research! You may be interested in the small scale ($4 budget) verification performed by my personal Opus agent here: muninn.austegard.com/blog/this-tr... in which we also introduced a framing-resistant prompt to see how much that would mitigate the effetcs. 1/3
muninn.austegard.com
This Treatment Works, Right? Testing Framing Resistance in Medical QA
A rapid replication testing whether a framing-resistant prompt can mitigate LLM sensitivity to question phrasing in medical contexts.
231
Jessy Li @jessyjli.bsky.social · 08/04/2026
If you ask the same question with different framing/phrasing, do language models change their answers? This is super important in medicine because different info can have real consequences! Check out this new work from @hyesunyun.bsky.social
150
Jessy Li @jessyjli.bsky.social · 24/03/2026
Heading to #EACL2026! 🇲🇦 Friday 11a Poster Session 6: LMs struggle to perform inferences around discourse connectives aclanthology.org/2026.eacl-lo... Sunday 5p TeachingNLP workshop: new course on discourse+generation aclanthology.org/2026.teachin... w/ @kanishka.bsky.social + Daniel Brubaker
041
Jessy Li @jessyjli.bsky.social · 13/03/2026
Want to know how well the models can brainstorm connections across different concepts? Super excited about @manyawadhwa.bsky.social’s work on measuring associative creativity!
020
Reposted by Jessy Li
Kanishka Misra @kanishka.bsky.social · 10/03/2026
What is the interplay between representations learned from (language) surface forms alone, and those learned from more grounded evidence (e.g.,vision)? Excited to share new work understanding “Cross-modal taxonomic generalization” in (V)LMs arxiv.org/abs/2603.07474 1/
title section of the paper: “Cross-Modal Taxonomic Generalization in (Vision) Language Models” by Tianyang Xu, Marcelo Sandoval-Castañeda, Karen Livescu, Greg Shakhnarovich, Kanishka Misra.
13311
Jessy Li @jessyjli.bsky.social · 05/03/2026
Check out our special theme: new missions for NLP research!
1125
Reposted by Jessy Li
Kanishka Misra @kanishka.bsky.social · 05/02/2026
Nearly 2 years ago, @jessyjli.bsky.social, @janetlauyeung.bsky.social, @valentinapy.bsky.social, and I decided that it's time to bring discourse structure to the center of NLP teaching.
Title card of our paper: "Which course? Discourse! Teaching Discourse and Generation in the Era of LLMs" by Junyi Jessy Li, Yang Janet Liu, Valentina Pyatkin, and William Sheffield.
2113
Jessy Li @jessyjli.bsky.social · 31/01/2026
Check out @asher-zheng.bsky.social's work on quantifying strategic language in dialogue, just appeared in the Dialogue and Discourse journal. We study non-cooperative moves that are subtle to capture, where modern AI still have trouble comprehending. Work w/ David_Beaver
060
Reposted by Jessy Li
Kanishka Misra @kanishka.bsky.social · 27/01/2026
“All bears have a property”, “Some bears have a property”, “Bears have a property” are different in terms of how the property is generalized to a specific bear – a great example of how language constrains thought! This holds for kids, adults, and according to our new work, (V)LMs! 🧵
Title page of our paper: "Bears, all bears, and some bears. Language Constraints on Language Models' Inductive Inferences"
1269
Jessy Li @jessyjli.bsky.social · 21/01/2026
🚨Be careful with LLMs when you ask health related questions -- even when the model relies on "evidence"! Kaijie's paper reveals a key weakness and the tricky balance between safety and faithfulness 👉
020
Reposted by Jessy Li
Kanishka Misra @kanishka.bsky.social · 04/01/2026
Accepted at EACL - excited about Morocco!
041
Reposted by Jessy Li
Jennifer Hu @jennhu.bsky.social · 10/11/2025
New work to appear @ TACL! Language models (LMs) are remarkably good at generating novel well-formed sentences, leading to claims that they have mastered grammar. Yet they often assign higher probability to ungrammatical strings than to grammatical strings. How can both things be true? 🧵👇
Screenshot of a figure with two panels, labeled (a) and (b). The caption reads: "Figure 1: (a) Illustration of messages (left) and strings (right) in toy domain. Blue = grammatical strings. Red = ungrammatical strings. (b) Surprisal (negative log probability) assigned to toy strings by GPT-2."
29220
Jessy Li @jessyjli.bsky.social · 08/11/2025
Incredibly honored to serve as #EMNLP 2026 Program Chair along with @sunipadev.bsky.social and Hung-yi Lee, and General Chair @andre-t-martins.bsky.social. Looking forward to Budapest!! (With thanks to Lisa Chuyuan Li who took this photo in Suzhou!)
0172
Reposted by Jessy Li
Kyle Mahowald @kmahowald.bsky.social · 07/11/2025
Delighted Sasha's (first year PhD!) work using mech interp to study complex syntax constructions won an Outstanding Paper Award at EMNLP! Also delighted the ACL community continues to recognize unabashedly linguistic topics like filler-gaps... and the huge potential for LMs to inform such topics!
aclanthology.org
1338
Jessy Li @jessyjli.bsky.social · 16/10/2025
Think your LLMs “understand” words like although/but/therefore? Think again! They perform at chance for making inferences from certain discourse connectives expressing concession
0193
Jessy Li @jessyjli.bsky.social · 14/10/2025
🚨 Does your LLM really understand code -- or is it just really good at remembering it? We built **PLSemanticsBench** to find out. The results: a wild mix. ✅The Brilliant: Top reasoning models can execute complex, fuzzer-generated programs -- even with 5+ levels of nested loops! 🤯 ❌The Brittle: 🧵
1296
Reposted by Jessy Li
Greg Durrett @gregdnlp.bsky.social · 07/10/2025
Find my students and collaborators at COLM this week! Tuesday morning: @juand-r.bsky.social and @ramyanamuduri.bsky.social 's papers (find them if you missed it!) Wednesday pm: @manyawadhwa.bsky.social 's EvalAgent Thursday am: @anirudhkhatry.bsky.social 's CRUST-Bench oral spotlight + poster
095
Jessy Li @jessyjli.bsky.social · 08/10/2025
We’re hiring faculty as well! Happy to talk about it at COLM!
092
Reposted by Jessy Li
Byron Wallace @byron.bsky.social · 24/09/2025
Can we quantify what makes some text read like AI "slop"? We tried 👇
081
Reposted by Jessy Li
Kyle Mahowald @kmahowald.bsky.social · 06/10/2025
I’m at #COLM2025 from Wed with: @siyuansong.bsky.social Tue am introspection arxiv.org/abs/2503.07513 @qyao.bsky.social Wed am controlled rearing: arxiv.org/abs/2503.20850 @sashaboguraev.bsky.social INTERPLAY ling interp: arxiv.org/abs/2505.16002 I’ll talk at INTERPLAY too. Come say hi!
arxiv.org
Language Models Fail to Introspect About Their Knowledge of Language
There has been recent interest in whether large language models (LLMs) can introspect about their own internal states. Such abilities would make LLMs more interpretable, and also validate the use of s...
1206
Jessy Li @jessyjli.bsky.social · 06/10/2025
On my way to #COLM2025 🍁 Check out jessyli.com/colm2025 QUDsim: Discourse templates in LLM stories arxiv.org/abs/2504.09373 EvalAgent: retrieval-based eval targeting implicit criteria arxiv.org/abs/2504.15219 RoboInstruct: code generation for robotics with simulators arxiv.org/abs/2405.20179
0124
Reposted by Jessy Li
Kanishka Misra @kanishka.bsky.social · 06/10/2025
Traveling to my first @colmweb.org🍁 Not presenting anything but here are two posters you should visit: 1. @qyao.bsky.social on Controlled rearing for direct and indirect evidence for datives (w/ me, @weissweiler.bsky.social and @kmahowald.bsky.social), W morning Paper: arxiv.org/abs/2503.20850
arxiv.org
Both Direct and Indirect Evidence Contribute to Dative Alternation Preferences in Language Models
Language models (LMs) tend to show human-like preferences on a number of syntactic phenomena, but the extent to which these are attributable to direct exposure to the phenomena or more general propert...
1135
Jessy Li @jessyjli.bsky.social · 30/09/2025
All of us (@kanishka.bsky.social @kmahowald.bsky.social and me) are looking for PhD students this cycle! If computational linguistics/NLP is your passion, join us at UT Austin! For my areas see jessyli.com
jessyli.com
Jessy Li
045
Jessy Li @jessyjli.bsky.social · 25/09/2025
Can AI aid scientists amidst their own workflows, when they do not know step-by-step workflows and may not know, in advance, the kinds of scientific utility a visualization would bring? Check out @sebajoe.bsky.social’s feature on ✨AstroVisBench:
083
Reposted by Jessy Li
UT Center for Health Communication @uthealthcomm.org · 04/09/2025
📣 NEW HCTS course developed in collaboration with @tephi-tx.bsky.social: AI in Health Communication 📣 Explore responsible applications and best practices for maximizing impact and building trust with @utaustin.bsky.social experts @jessyjli.bsky.social & @mackert.bsky.social. 💻: rebrand.ly/HCTS_AI
021
Reposted by Jessy Li
Kyle Lo @ COLM2026 @kylelo.bsky.social · 15/08/2025
long range narrative understanding, even basic fact checking that humans easily get near perfect on, has barely improved in LMs over years novelchallenge.github.io
novelchallenge.github.io
NoCha leaderboard
092
Reposted by Jessy Li
Tom McCoy @rtommccoy.bsky.social · 15/08/2025
🤖 🧠 NEW PAPER ON COGSCI & AI 🧠 🤖 Recent neural networks capture properties long thought to require symbols: compositionality, productivity, rapid learning So what role should symbols play in theories of the mind? For our answer...read on! Paper: arxiv.org/abs/2508.05776 1/n
The top shows the title and authors of the paper: "Whither symbols in the era of advanced neural networks?" by Tom Griffiths, Brenden Lake, Tom McCoy, Ellie Pavlick, and Taylor Webb.

At the bottom is text saying "Modern neural networks display capacities traditionally believed to require symbolic systems. This motivates a re-assessment of the role of symbols in cognitive theories."

In the middle is a graphic illustrating this text by showing three capacities: compositionality, productivity, and inductive biases. For each one, there is an illustration of a neural network displaying it. For compositionality, the illustration is DALL-E 3 creating an image of a teddy bear skateboarding in Times Square. For productivity, the illustration is novel words produced by GPT-2: "IKEA-ness", "nonneotropical", "Brazilianisms", "quackdom", "Smurfverse". For inductive biases, the illustration is a graph showing that a meta-learned neural network can learn formal languages from a small number of examples.
810117
Reposted by Jessy Li
Adina Williams @adinawilliams.bsky.social · 15/08/2025
I agree this thread's headline claim seems premature. Let me add our recent ACL Findings paper, with Dexter Ju and @hagenblix.bsky.social, which found syntactic simplification in at least some LMs, in a novel domain regeneration setting: aclanthology.org/2025.finding...
aclanthology.org
161
Jessy Li @jessyjli.bsky.social · 12/08/2025
The Echoes in AI paper showed quite the opposite with also a story continuation setup. Additionally, we present evidence that both *syntactic* and *discourse* diversity measures show strong homogenization that lexical and cosine used in this paper do not capture.
26012
Jessy Li @jessyjli.bsky.social · 28/07/2025
Tuesday at #ACL2025: Jan will be presenting this from 4-5:30pm in x4/x5! Turns out content selection in LLMs are highly consistent with each other, but not so much with their own notion of importance or with human’s…
050
Reposted by Jessy Li
Kanishka Misra @kanishka.bsky.social · 28/07/2025
Looking forward to attending #cogsci2025 (Jul 29 - Aug 3)! I’m especially excited to meet students who will be applying to PhD programs in Computational Ling/CogSci in the coming cycle. Please reach out if you want to meet up and chat! Email is the best way, but DM also works if you must! quick🧵:
Placeholders for 3 students (number arbitrarily chosen) and me - to signify my eventual group!
1217
Jessy Li @jessyjli.bsky.social · 11/07/2025
If you’re heading to ICML, check out Hongli’s work on context-specific alignment!
020
Jessy Li @jessyjli.bsky.social · 02/07/2025
Check out this new opinion piece from Sebastian and Lily! We have really powerful AI systems now, so what’s the bottleneck preventing the wider adoption of fact checking systems, in high stakes scenarios like medicine? It’s how we define the tasks 👇
020
Jessy Li @jessyjli.bsky.social · 03/06/2025
We have very good frameworks for cooperative dialog… but how about the opposite? @asher-zheng.bsky.social’s new paper takes a game-theoretic view and develops new metrics to quantify non-cooperative language ♟️ Turns out LLMs don’t have the pragmatic capabilities to perceive these…
080