Sign in

Xiaoyan Bai

@elenal3ai.bsky.social
453 followers 185 following 71 posts

PhD @UChicagoCS / BE in CS @Umich / ✨AI/NLP transparency and interpretability/📷🎨photography painting

PostsRepliesMedia
Reposted by Xiaoyan Bai
chenhaotan.bsky.social @chenhaotan.bsky.social · 09/03/2026
Peer review is facing a death spiral, and AI production tools are speeding it up. AI-assisted reviewing is necessary and should be open. We built OpenAIReview: open AI reviewing for everyone, for the cost of a coffee. openaireview.github.io/blog.html 🧵
openaireview.github.io
AI-assisted Reviewing is Necessary and Should be Open
Peer review is facing a death spiral. AI production tools are speeding it up. AI-assisted reviewing is necessary and should be open.
12610
Reposted by Xiaoyan Bai
Mourad Heddaya @mheddaya.bsky.social · 19/02/2026
Democracy depends on an informed electorate. But political issues and ballot measures can be confusing, obscuring the effects of one outcome versus another. And politics is personal. Once we make an initial decision, it can be hard to see things from "the other side."
242
Xiaoyan Bai @elenal3ai.bsky.social · 10/02/2026
📖 ≠ 🧪 The Story is Not the Science. Code is submitted but rarely executed during peer review—an issue likely to worsen with research agents. 🧑‍🔬 We introduce 𝐌𝐞𝐜𝐡𝐄𝐯𝐚𝐥𝐀𝐠𝐞𝐧𝐭, an execution-grounded evaluation of narrative + execution. 𝐕𝐞𝐫𝐢𝐟𝐲 𝐭𝐡𝐞 𝐬𝐜𝐢𝐞𝐧𝐜𝐞, 𝐧𝐨𝐭 𝐣𝐮𝐬𝐭 𝐭𝐡𝐞 𝐬𝐭𝐨𝐫𝐲. 1/n
294
Reposted by Xiaoyan Bai
Data Science Institute @dsi-uchicago.bsky.social · 13/01/2026
Featured in UChicago News: @elenal3ai.bsky.social & @chenhaotan.bsky.social's research into why AI can write complex code but fails at 4-digit multiplication: tinyurl.com/5ukvm7p7
tinyurl.com
Why can’t powerful AIs learn basic multiplication?
New research reveals why even state-of-the-art large language models stumble on seemingly easy tasks—and what it takes to fix it
021
Xiaoyan Bai @elenal3ai.bsky.social · 24/11/2025
Will be at #NeurIPS2025 presenting “Concept Incongruence”! 🦄🦆 Curious about a unicorn duck? Stop by, get one, and chat with us! We made a new demo for detecting hidden conflicts in system prompts to spot “concept incongruence” for safer prompts. 🔗: github.com/ChicagoHAI/d... 🗓️ Dec 3 11AM - 2PM
161
Xiaoyan Bai @elenal3ai.bsky.social · 20/11/2025
Research agents are getting smarter. They can write convincing PhD-level reports 🧑‍🔬 But has anyone checked if the way they find their results makes any sense? Our framework, MechEvalAgents, verifies the science, not just the story 🤖 1/n🧵
130
Reposted by Xiaoyan Bai
Haokun Liu @haokunliu.bsky.social · 10/11/2025
We're launching a weekly competition where the community decides which research ideas get implemented. Every week, we'll take the top 3 ideas from IdeaHub, run experiments with AI agents, and share everything: code, successes, and failures. It's completely free and we'll try out ideas for you!
174
Reposted by Xiaoyan Bai
Lexing Xie @lexingxie.bsky.social · 30/10/2025
Identifying human morals and values in language is crucial for analysing lots of human- and AI-generated text. We introduce "MoVa: Towards Generalizable Classification of Human Morals and Values" - to be presented at @emnlpmeeting.bsky.social oral session next Thu #CompSocialScience #LLMs 🧵 (1/n)
885
Xiaoyan Bai @elenal3ai.bsky.social · 28/10/2025
🕸️ Here’s a network showing how much different models predict each other as the author of some text!
082
Xiaoyan Bai @elenal3ai.bsky.social · 27/10/2025
❓ Does an LLM know thyself? 🪞 Humans pass the mirror test at ~18 months 👶 But what about LLMs? Can they recognize their own writing—or even admit authorship at all? In our new paper, we put 10 state-of-the-art models to the test. Read on 👇 1/n 🧵
1124
Xiaoyan Bai @elenal3ai.bsky.social · 24/10/2025
In our new work, we reverse-engineer two models: a standard fine-tuned (SFT), and an implicit chain-of-thought (ICoT) model to see why models struggle with multi-digit multiplication. 👉Check out the paper here: arxiv.org/abs/2510.00184 🎉Big thanks to all my amazing collaborators!
arxiv.org
Why Can't Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls
Language models are increasingly capable, yet still fail at a seemingly simple task of multi-digit multiplication. In this work, we study why, by reverse-engineering a model that successfully learns m...
071
Reposted by Xiaoyan Bai
chenhaotan.bsky.social @chenhaotan.bsky.social · 23/10/2025
AI can accelerate scientific discovery, but only if we get the scientist–AI interaction right. The dream of “autonomous AI scientists” is tempting: machines that generate hypotheses, run experiments, and write papers. But science isn’t just automation. cichicago.substack.com/p/the-mirage... 🧵
cichicago.substack.com
The Mirage of Autonomous AI Scientists
Science as AI’s killer application cannot succeed without scientist-AI interaction: Introducing Hypogenic.ai.
2236
Reposted by Xiaoyan Bai
Dang Nguyen @divingwithorcas.bsky.social · 26/09/2025
HR Simulator™: a game where you gaslight, deflect, and “let’s circle back” your way to victory. Every email a boss fight, every “per my last message” a critical hit… or maybe you just overplayed your hand 🫠 Can you earn Enlightened Bureaucrat status? (link below!)
245
Reposted by Xiaoyan Bai
chenhaotan.bsky.social @chenhaotan.bsky.social · 25/09/2025
🚀 We’re thrilled to announce the upcoming AI & Scientific Discovery online seminar! We have an amazing lineup of speakers. This series will dive into how AI is accelerating research, enabling breakthroughs, and shaping the future of research across disciplines. ai-scientific-discovery.github.io
12215
Reposted by Xiaoyan Bai
chenhaotan.bsky.social @chenhaotan.bsky.social · 16/09/2025
As AI becomes increasingly capable of conducting analyses and following instructions, my prediction is that the role of scientists will increasingly focus on identifying and selecting important problems to work on ("selector"), and effectively evaluating analyses performed by AI ("evaluator").
2108
Reposted by Xiaoyan Bai
chenhaotan.bsky.social @chenhaotan.bsky.social · 29/08/2025
We are proposing the second workshop on AI & Scientific Discovery at EACL/ACL. The workshop will explore how AI can advance scientific discovery. Please use this Google form to indicate your interest (corrected link): forms.gle/MFcdKYnckNno... More in the 🧵! Please share! #MLSky 🧠
forms.gle
Program Committee Interest for the Second Workshop on AI & Scientific Discovery
We are proposing the second workshop on AI & Scientific Discovery at EACL/ACL (Annual meetings of The Association for Computational Linguistics, the European Language Resource Association and Internat...
1148
Xiaoyan Bai @elenal3ai.bsky.social · 31/07/2025
⚡️Ever asked an LLM-as-Marilyn Monroe about the 2020 election? Our paper calls this concept incongruence, common in both AI and how humans create and reason. 🧠Read my blog to learn what we found, why it matters for AI safety and creativity, and what's next: cichicago.substack.com/p/concept-in...
195
Reposted by Xiaoyan Bai
chenhaotan.bsky.social @chenhaotan.bsky.social · 09/07/2025
Prompting is our most successful tool for exploring LLMs, but the term evokes eye-rolls and grimaces from scientists. Why? Because prompting as scientific inquiry has become conflated with prompt engineering. This is holding us back. 🧵and new paper with @ari-holtzman.bsky.social .
23715
Reposted by Xiaoyan Bai
chenhaotan.bsky.social @chenhaotan.bsky.social · 02/07/2025
When you walk into the ER, you could get a doc: 1. Fresh from a week of not working 2. Tired from working too many shifts @oziadias.bsky.social has been both and thinks that they're different! But can you tell from their notes? Yes we can! Paper @natcomms.nature.com www.nature.com/articles/s41...
12611
Xiaoyan Bai @elenal3ai.bsky.social · 25/06/2025
Humbled to receive an honorable mention🌟
010
Reposted by Xiaoyan Bai
chenhaotan.bsky.social @chenhaotan.bsky.social · 23/06/2025
Since @elenal3ai.bsky.social cannot make it, I presented the poster on concept incongruence: arxiv.org/abs/2505.14905
072
Xiaoyan Bai @elenal3ai.bsky.social · 28/05/2025
I am glad that you found our paper entertaining! This is a great point for my follow-up thread on the implications of concept incongruence. Our main goal is to raise awareness and provide clarity around concept incongruence.
134
Xiaoyan Bai @elenal3ai.bsky.social · 27/05/2025
🚨 New paper alert 🚨 Ever asked an LLM-as-Marilyn Monroe who the US president was in 2000? 🤔 Should the LLM answer at all? We call these clashes Concept Incongruence. Read on! ⬇️ 1/n 🧵
13017
Reposted by Xiaoyan Bai
Mourad Heddaya @mheddaya.bsky.social · 01/05/2025
🧑‍⚖️How well can LLMs summarize complex legal documents? And can we use LLMs to evaluate? Excited to be in Albuquerque presenting our paper this afternoon at @naaclmeeting 2025!
22313
Reposted by Xiaoyan Bai
Haokun Liu @haokunliu.bsky.social · 28/04/2025
🚀🚀🚀Excited to share our latest work: HypoBench, a systematic benchmark for evaluating LLM-based hypothesis generation methods! There is much excitement about leveraging LLMs for scientific hypothesis generation, but principled evaluations are missing - let’s dive into HypoBench together.
1119
Reposted by Xiaoyan Bai
chenhaotan.bsky.social @chenhaotan.bsky.social · 21/04/2025
Encourage your students to submit posters and register! Limited free housing is provided for student participants only, on a first-come (i.e., request)-first-serve basis. We are also actively looking for sponsors. Reach out if you are interested! Please repost! Help spread the words!
21010
Reposted by Xiaoyan Bai
Dang Nguyen @divingwithorcas.bsky.social · 14/04/2025
1/n You may know that large language models (LLMs) can be biased in their decision-making, but ever wondered how those biases are encoded internally and whether we can surgically remove them?
11712
Reposted by Xiaoyan Bai
Julia Mendelsohn @jmendelsohn2.bsky.social · 20/02/2025
New preprint! Metaphors shape how people understand politics, but measuring them (& their real-world effects) is hard. We develop a new method to measure metaphor & use it to study dehumanizing metaphor in 400K immigration tweets Link: bit.ly/4i3PGm3 #NLP #NLProc #polisky #polcom #compsocialsci 🐦🐦
Screenshot of top half of first page of paper. The paper is titled: "When People are Floods: Analyzing Dehumanizing Metaphors in Immigration Discourse with Large Language Models". The authors are Julia Mendelsohn (University of Chicago) and Ceren Budak (University of Michigan). The top right corner contains a visual showing the sentence "They want immigrants to pour into and infest this country". The caption says: Figure 1: Dehumanizing sentence likening immigrants to the source domain concepts of Water and Vermin via the words "pour" and "infest". 

The abstract text on the left reads: Metaphor, discussing one concept in terms of another, is abundant in politics and can shape how people understand important issues. We develop a computational approach to measure metaphorical language, focusing on immigration discourse on social media. Grounded in qualitative social science research, we identify seven concepts evoked in immigration discourse (e.g. "water" or "vermin"). We propose and evaluate a novel technique that leverages both word-level and document-level signals to measure metaphor with respect to these concepts. We then study the relationship between metaphor, political ideology, and user engagement in 400K US tweets about immigration. While conservatives tend to use dehumanizing metaphors more than liberals, this effect varies widely across concepts. Moreover, creature-related metaphor is associated with more retweets, especially for liberal authors. Our work highlights the potential for computational methods to complement qualitative approaches in understanding subtle and implicit language in political discourse.
618264
Reposted by Xiaoyan Bai
Guillaume Lajoie @glajoie.bsky.social · 23/10/2024
Compositional representations are a key attributes of intelligent systems that generalize well. An issue is that there is no robust way to quantify compositionality. Below is our attempt at such a quantifiable measurement. arxiv.org/abs/2410.148... w/ E Elmoznino & T Jiralerspong & Y Bengio
arxiv.org
A Complexity-Based Theory of Compositionality
Compositionality is believed to be fundamental to intelligence. In humans, it underlies the structure of thought, language, and higher-level reasoning. In AI, compositional representations can enable ...
0164
Reposted by Xiaoyan Bai
Mourad Heddaya @mheddaya.bsky.social · 15/11/2024
How do everyday narratives reveal hidden cause-and-effect patterns that shape our beliefs and behaviors? In our paper, we propose Causal Micro-Narratives to uncover narratives from real-world data. As a case study, we characterize the narratives about inflation in news.
1347