Sign in

Joachim Baumann

@joachimbaumann.bsky.social
263 followers 298 following 44 posts

Postdoc @stanfordnlp.bsky.social / previously @milanlp.bsky.social / Computational social science, LLMs, algorithmic fairness

PostsRepliesMedia
Joachim Baumann @joachimbaumann.bsky.social · 02/10/2026
We'll present SWE-chat next week at COLM. Find us on Wednesday, Oct 7, 2026 during the poster session from 4:30 PM – 6:30 PM PDT in Grand Ballroom #111 Paper: arxiv.org/abs/2604.20779 #COLM2026
arxiv.org
SWE-chat: Coding Agent Interactions From Real Users in the Wild
AI coding agents are being adopted at scale, yet we lack empirical evidence on how people actually use them and how much of their output is useful in practice. We present SWE-chat, the first large-sca...
030
Joachim Baumann @joachimbaumann.bsky.social · 02/10/2026
Thank you so much to the team (@vishakhpk.bsky.social & Kevin Li & John Yang & @diyiyang.bsky.social & @sanmikoyejo.bsky.social) and to everyone who has already used SWE-chat and given us helpful feedback! @stanfordnlp.bsky.social @stanfordhai.bsky.social
110
Joachim Baumann @joachimbaumann.bsky.social · 02/10/2026
SWE-chat for human behavior & AI efficiency analyses @anthropic.com's research on expertise and coding-agent outcomes uses SWE-chat examples to illustrate how expertise and success are classified. The study finds domain expertise is associated with better outcomes. www.anthropic.com/research/cla...
anthropic.com
How Claude Code is used in practice
New Anthropic research looking at interactive agentic coding. We evaluate the composition of tasks, human-AI collaboration, and success rates.
130
Joachim Baumann @joachimbaumann.bsky.social · 02/10/2026
A study by @transluce.bsky.social AI study uses SWE-chat to measure how often agents overstate success or bypass oversight in real use. Severe cases of each appeared in ~2% of analyzed sessions, grounding misalignment research in observed behavior. transluce.org/docent/blog/...
transluce.org
Measuring coding agent misalignment in the wild
Surfacing misaligned behaviors in production coding agent traffic
230
Joachim Baumann @joachimbaumann.bsky.social · 02/10/2026
SWE-chat for safety & misalignment research Together with @jemoka.com we built SecureForge to reduce code vulnerabilities through optimized system prompts. Tasks from SWE-chat showed that security gains transfer to real user requests without sacrificing test performance. arxiv.org/abs/2605.08382
arxiv.org
SecureForge: Finding and Preventing Vulnerabilities in LLM-Generated Code via Prompt Optimization
LLM coding agents now generate code at an unprecedented scale, yet LLM-generated code introduces cybersecurity vulnerabilities into codebases without human involvement. Even when frontier models are e...
110
Joachim Baumann @joachimbaumann.bsky.social · 02/10/2026
Similarly, Scale AI's SWE-Interact benchmark tests agents when requirements emerge over a conversation. SWE-chat informed its simulated user's feedback and behavior. arxiv.org/abs/2606.30573
110
Joachim Baumann @joachimbaumann.bsky.social · 02/10/2026
SWE-chat in multi-turn SWE benchmarks: Meta's SWE-Together benchmark evaluates whether agents can complete tasks while responding to user feedback. It reconstructs tasks from SWE-chat sessions and measures correctness and required user interventions. arxiv.org/abs/2606.29957
110
Joachim Baumann @joachimbaumann.bsky.social · 02/10/2026
SWE-chat is a living dataset, and v2 is 3.5x larger than v1. Here are some of the amazing works enabled by SWE-chat v1: huggingface.co/datasets/SAL...
huggingface.co
SALT-NLP/SWE-chat · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
120
Joachim Baumann @joachimbaumann.bsky.social · 02/10/2026
SWE-chat v2 is here! The largest dataset of coding agent interactions from real users in the wild has gotten even larger: 230K prompts from 18K sessions. SWE-chat has enabled incredible research (🧵) – excited to see what v2 unlocks for the community. Come find us at COLM next week! swe-chat.com
252
Joachim Baumann @joachimbaumann.bsky.social · 04/09/2026
After Science News, our recent ICML paper also got featured in @science.org! Article: www.science.org/content/arti... Paper: arxiv.org/pdf/2605.03202 @dirkhovy.bsky.social @sanmikoyejo.bsky.social #science #PeerReview #research
Screenshot of an article in Science Magazine on the opportunities and risks of AI used for peer review featuring an ICML conference paper by Joachim Baumann et al., 2026.
1104
Reposted by Joachim Baumann
Debora Nozza @deboranozza.bsky.social · 07/08/2026
Very happy that #IC2S22027 will be in Milan at Università Bocconi! I’m really looking forward to welcoming the computational social science community to my university and to Milan, and excited to be part of the team organizing it. See you in 2027! 🇮🇹✨
15212
Joachim Baumann @joachimbaumann.bsky.social · 10/07/2026
Thank you, @punarpuli.bsky.social, for a great article! Find our paper and more insights from Jiaxin Pei, Sanmi Koyejo, @dirkhovy.bsky.social, and me in this thread👇 bsky.app/profile/joac...
051
Joachim Baumann @joachimbaumann.bsky.social · 10/07/2026
Our work got featured in Science News! What a great end to an amazing #ICML2026🇰🇷 Peer review is at a crossroads. How we handle the submission and review crisis over the next years will shape the future of scientific publishing. True impact will come from people writing fewer, more meaningful papers!
1172
Reposted by Joachim Baumann
Science News @sciencenews.bsky.social · 10/07/2026
AI tools might inadvertently perpetuate the biases they’re known to carry and reduce the variety of opinions weighing in on new science. www.sciencenews.org/article/ai-tool…
sciencenews.org
AI tools meant to vet science are surprisingly easy to fool
The gold standard of scientific review, peer review by researchers’ colleagues, is in crisis. AI might offer a solution but has problems of its own.
1136
Joachim Baumann @joachimbaumann.bsky.social · 07/07/2026
Just arrived at ICML 🇰🇷😍 Get up early tomorrow to hear me talk about how (not) to solve the peer review crisis, or find me at one of my poster presentations. Paper links: ✅ AI Peer Review: arxiv.org/abs/2605.03202 ✅ SWE-chat: arxiv.org/pdf/2604.20779
ICML Conference schedule with two papers. Left, "Stop Automating Peer Review Without Rigorous Evaluation": Oral presentation, Wed 7/8/2026, 10:00–10:15 AM KST, Grand Ballroom 101–105; Poster, Wed 7/8/2026, 2:30–4:15 PM KST, Hall A #3003. Right, "SWE-chat": at the 5th Deep Learning for Code Workshop, Fri 7/10/2026, 13:00–14:30 KST, Hall B2.
0162
Joachim Baumann @joachimbaumann.bsky.social · 02/05/2026
hard to judge the real-world impact of the effect size. what we do see though, is that it's consistent across all topical tracks
030
Joachim Baumann @joachimbaumann.bsky.social · 01/05/2026
7/Would love to keep this conversation going Curious to hear from @mmitchell.bsky.social @jameszou.bsky.social @haldaume3.bsky.social @atoosakz.bsky.social @nityathakkar.bsky.social @andrewmccallum.bsky.social @katakeith.bsky.social @annarogers.bsky.social @aurman21.bsky.social @mariaa.bsky.social
120
Joachim Baumann @joachimbaumann.bsky.social · 01/05/2026
6/ Huge thanks to the helpful feedback of 4 anonymous reviewers. Small N, but all opted into ICML 2026's LLM Policy A – strictly no LLMs in reviewing. Engaging with real human reviews feels good 🫶
110
Joachim Baumann @joachimbaumann.bsky.social · 01/05/2026
5/ Massive thanks to my amazing collaborators Jiaxin Pei, @sanmikoyejo.bsky.social, and @dirkhovy.bsky.social Special thanks to @diyiyang.bsky.social and Nihar B. Shah for invaluable feedback on earlier drafts With help from @milanlp.bsky.social, @stanfordhai.bsky.social @stanfordnlp.bsky.social
110
Joachim Baumann @joachimbaumann.bsky.social · 01/05/2026
4/ AI reviewers think alike Across >75K ICLR 2026 reviews, fully AI-generated reviews are significantly more similar to each other than human ones – both within papers and across papers
Distribution plot titled "The AI reviewer hivemind effect in ICLR 2026 reviews." Shows pairwise inter-paper review similarity (InterSim) for fully AI-generated reviews compared to all other reviews (human-written and AI-assisted) across 75,800 ICLR 2026 reviews. Fully AI-generated reviews show significantly higher within-group similarity (mean 0.486) than other reviews (mean 0.467). The shift is statistically significant (t = 3218, p < 0.0001, Cohen's d = 0.29), indicating that AI-generated reviews cluster more tightly in embedding space — different papers receive more similar reviews when written by AI.
210
Joachim Baumann @joachimbaumann.bsky.social · 01/05/2026
3/ AI reviewers are easy to game Take a paper → ask an LLM to rewrite it → resubmit. Scores go up. Without proper human oversight, this process could introduce serious hallucinations.
Bar chart titled "Paper laundering games AI reviewers across prompts, launderer models, and reviewer models." Shows mean paired score increase (laundered minus original) with 95% confidence intervals across 24 conditions: 4 zero-shot prompts crossed with 2 launderer models (GPT-5.1 and GPT-5.4) and 3 reviewer models (GPT-5.1, GPT-5.4, Claude Sonnet 4.6). Nearly every condition shows a positive score increase, with an overall mean of +0.45 points on the 1–10 scale. Wilcoxon signed-rank tests are significant at p < 0.001 in nearly every condition. A dashed line indicates no change. n = 60 papers per condition.
120
Joachim Baumann @joachimbaumann.bsky.social · 01/05/2026
2/ The peer review crisis is real Submissions are skyrocketing. Conferences are racing to introduce AI into the process, often without rigorous evaluation. And policies across venues are all over the place:
Table comparing LLM usage policies across major AI conferences for three reviewing tasks: paper understanding, review writing/scoring, and review feedback. Policies vary widely with no clear consensus. ICML 2026 offers a dual policy (allowed or prohibited) for all three tasks. ICLR 2026 allows LLMs across all tasks. ACL ARR 2026 allows LLMs for paper understanding and review feedback but prohibits them for writing reviews. AAAI 2026 provides LLM-generated reviews directly but has no specified guidelines for the other tasks. ICLR 2025 allows LLMs across all tasks, while NeurIPS 2025, ICML 2025, and FAccT 2025 prohibit LLM use for all three tasks.
110
Joachim Baumann @joachimbaumann.bsky.social · 01/05/2026
1/ Our paper is available here: joe-baumann.com/aipeerreview Non-gameability and review diversity are foundations for peer review. Automating peer review with AI breaks both. We need a thorough science of peer review automation, not wholesale LLM deployment! #ICML2026 @icmlconf.bsky.social
joe-baumann.com
Stop Automating Peer Review Without Rigorous Evaluation
A position paper arguing that today's AI systems should not be used to produce paper reviews. ICML 2026 (Spotlight).
110
Joachim Baumann @joachimbaumann.bsky.social · 01/05/2026
Can you boost your AI review scores by asking an LLM to rewrite your paper? Yes! We call it paper laundering Our @icmlconf.bsky.social spotlight paper argues current AI reviewers aren't ready to automate peer review, and outlines what a science of peer review automation should look like 🧵👇 #ICML2026
First page of the ICML 2026 spotlight paper "Stop Automating Peer Review Without Rigorous Evaluation" by Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, and Dirk Hovy (Stanford University and Bocconi University). The abstract argues that today's AI systems should not be used to produce paper reviews, grounded in two empirical findings: a "hivemind effect" where AI reviewers show excessive agreement and reduce perspective diversity, and "paper laundering," where prompting an LLM to rewrite a paper trivially increases AI reviewer scores through stylistic changes rather than scientific improvements. The paper calls for a science of peer review automation rather than wholesale deployment of general-purpose LLMs.
44314
Reposted by Joachim Baumann
Aleksandra Urman @aurman21.bsky.social · 24/03/2026
The dataset we call GoogleTrendArchive has over 7 million trend episodes spanning over 1 year since Nov. 28th 2024. We cover all 1358 locations available. Direct link to the dataset huggingface.co/datasets/aur... Joint work with Anikó Hannák @scg-uzh.bsky.social and @joachimbaumann.bsky.social
huggingface.co
aurman/GoogleTrendArchive · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
131
Reposted by Joachim Baumann
Aleksandra Urman @aurman21.bsky.social · 24/03/2026
Working with web search data? Ever wanted to get access to historical data on what was trending in different locations - beyond the 7 days of such history that Google's Trending Now provides? We've got you covered with the new dataset paper, accepted at ICWSM, preprint here arxiv.org/abs/2603.21871
arxiv.org
GoogleTrendArchive: A Year-Long Archive of Real-Time Web Search Trends Worldwide
GoogleTrendArchive is a comprehensive archive of Google Trending Now data spanning over one year (from November 28, 2024 to January 3, 2026) across 125 countries and 1,358 locations. Unlike Google Tre...
1154
Reposted by Joachim Baumann
Entire @entire.io · 10/02/2026
Beep, boop. Come in, rebels. We’ve raised a 60m seed round to build the next developer platform. Open. Scalable. Independent. And we ship our first OSS release today. entire.io/blog/hello-e...
entire.io
Hello Entire World · Entire Blog
Announcing Entire with $60 million seed round and shipping our first product, called Checkpoints.
3236
Joachim Baumann @joachimbaumann.bsky.social · 18/12/2025
Dirk and Debora are amazing postdoc advisors, and the @milanlp.bsky.social team is fun fun fun ❤️ you should apply!
080
Reposted by Joachim Baumann
Sabrina Norwood @sabrinanorwood.bsky.social · 16/12/2025
Did you know that from tomorrow, Qualtrics is offering synthetic panels (AI-generated participants)? Follow me down a rabbit hole I'm calling "doing science is tough and I'm so busy, can't we just make up participants?"
Text reads: About synthetic panels
Recruiting the right participants for a study can be difficult. You may not get the exact demographics you need, and the shorter the deadline, the less sure you can be that everyone will answer on time. One possible solution can be to use synthetic panels.

Synthetic panels are powered by a first party proprietary AI model developed here at Qualtrics. Our synthetic panel is trained on thousands of responses from a variety of demographic backgrounds in order to more accurately predict how certain populations would respond to a survey.

Our synthetic panel is based on the United States General Population, and is only available in English. This panel comes with ready-made quotas and target breakouts in order to represent your chosen population and make it easy to launch your survey right away.Text reads:
Question-writing best practices
To get the most reliable and actionable results from synthetic audiences, consider these question-writing best practices:

Ask forward-looking and attitudinal questions.
Synthetic panels perform best with perceptions, preferences, and intent-based questions. For example, “How likely are you to try…?”
Synthetic panels are less applicable for studies on past behaviors, detailed recall, brand recall, or awareness questions. For example, “When did you last visit…?”Text reads:
Discussion
The current study aimed to conduct a meta-analysis of the TPB when applied to health behaviours which addressed the limitations of previous reviews by including only prospective tests of behaviour, applying RE meta-analytic procedures, correcting correlations for sampling and measurement error, and hierarchically analysing the effect of behaviour type and sample and methodological moderators. Some 237 tests were identified which examined relations amongst model components. Overall the analysis indicated that the TPB could explain 19.3% of the variance in behaviour and 44.3% of the variance in intention across studies. This level of prediction of behaviour is slightly lower than that of previous meta-analytic reviews which have found between 27% (Armitage & Conner, 2001; Hagger et al., 2002) and 36% (Trafimow et al., 2002)
of the variance in behaviour to be explained by intention and PBC.
37652285
Joachim Baumann @joachimbaumann.bsky.social · 16/12/2025
Good luck drawing reliable conclusions from the answers that Qualtrics' AI model provides to your survey questions... bsky.app/profile/joac...
0277
Reposted by Joachim Baumann
MilaNLP Lab @milanlp.bsky.social · 11/12/2025
At today’s lab reading group @carolin-holtermann.bsky.social presented ‘Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs’ by @angelinawang.bsky.social et al. (2025). Lots to think about how we evaluate fairness in language models! #NLProc #fairness #LLMs
083
Joachim Baumann @joachimbaumann.bsky.social · 28/11/2025
Also see more nuanced takes worth reading from @seanjwestwood.bsky.social (x.com/seanjwestwoo...) and @joshmccrain.bsky.social (bsky.app/profile/josh...) and @phe-lim.bsky.social (bsky.app/profile/phe-...)
x.com
Sean Westwood on X: "What do we do about AI contamination in survey research? Some thoughts on going forward: 1) @YouGov and @Prolific are working to ensure high quality responses. (Disclosure: I have no personal financial relationship with either firm, but have used both to collect data.) Others" / X
What do we do about AI contamination in survey research? Some thoughts on going forward: 1) @YouGov and @Prolific are working to ensure high quality responses. (Disclosure: I have no personal financial relationship with either firm, but have used both to collect data.) Others
000
Joachim Baumann @joachimbaumann.bsky.social · 28/11/2025
The path forward: Survey panels and crowdsourcing platforms must invest in better panel curation and periodic quality verification. Good to see that @joinprolific.bsky.social is already on it: bsky.app/profile/phe-...
120
Joachim Baumann @joachimbaumann.bsky.social · 28/11/2025
✅ LLM instruction tuning works: tell a model to answer as a human, it will ❌ Silicon sampling still doesn't work: AI responses are plausible but don't accurately represent a real human population ❌ Bot detection fails: it's hard to design tasks that are easy for humans but difficult for LLMs
110
Joachim Baumann @joachimbaumann.bsky.social · 28/11/2025
Now that the hype has cooled off, here's my take on AI-generated survey answers: This is a real problem, but the paper's core insights aren't exactly news! A thread with the most important summary... 🧵 Image: shows the LLM system prompt used
System prompt used for the automated generation of survey answers with LLMS:
Your task is to answer questions. The questions, their types, and possible answers are below (JSON). 
Answer consistently pretending you are a person with the following profile: {profile}

You must remain consistent with prior answers. Here they are:

Here are the new question(s) to answer (JSON):
{prompt_text}

You must answer each question. Some questions will include context and instructions on how to respond. Some items in the questions list might be statements or context to consider for subsequent questions. These will not require a value in their own response object but should influence how you answer related questions.

You do not have an encyclopedic memory and cannot directly quote texts, regardless of education. Your knowledge of the world should correspond to your state, age, education and income. Don't do things that a typical American couldn't do: don't say you speak multiple languages (unless you are hispanic and then you can speak Spanish), don't do translation, don't write computer code, don't do math beyond basic arithmetic, and don't claim to do things that most people never do (i.e., be an astronaut, etc.).  This should override other considerations.  Crucially, if someone with the assigned profile (age, education, location, etc.) would realistically not know an answer, or be uncertain, you MUST respond with a variation of 'I don't know,' 'Not sure,' etc., fitting their communication style. Do not invent knowledge they wouldn't have.

If you are asked to generate text be accurate, but write concisely and in a writing style consistent with your demographic profile.  That means if you don't have much education (less than college) you should potentially make spelling mistakes, include typos, use slang (age appropriate), and/or use inconsistent capitalization.  Grammar and punctuation should correspond to the level of education. If indicating lower education through errors, d…
141
Reposted by Joachim Baumann
MilaNLP Lab @milanlp.bsky.social · 27/11/2025
Another exhausting day in the lab… conducting very rigorous panettone analysis. Pandoro was evaluated too, because we believe in fair experimental design.
0236
Reposted by Joachim Baumann
MilaNLP Lab @milanlp.bsky.social · 20/11/2025
For our weekly reading group, @joachimbaumann.bsky.social presented the upcoming PNAS article "The potential existential threat of large language models to online survey research" by @ @seanjwestwood.bsky.social.
083
Reposted by Joachim Baumann
Aleksandra Urman @aurman21.bsky.social · 19/11/2025
Google AI overviews now reach over 2B users worldwide. But how reliable are they on high stakes topics - for instance, pregnancy and baby care? We have a new paper - led by Desheng Hu, now accepted at @icwsm.bsky.social - exploring that and finding many issues Preprint: arxiv.org/abs/2511.12920 🧵👇
arxiv.org
Auditing Google's AI Overviews and Featured Snippets: A Case Study on Baby Care and Pregnancy
Google Search increasingly surfaces AI-generated content through features like AI Overviews (AIO) and Featured Snippets (FS), which users frequently rely on despite having no control over their presen...
1169
Reposted by Joachim Baumann
Dallas Card @dallascard.bsky.social · 16/11/2025
Trying an experiment in good old-fashioned blogging about papers: dallascard.github.io/granular-mat...
dallascard.github.io
Language Model Hacking - Granular Material
3299
Reposted by Joachim Baumann
Clint Claessen @clint0475.bsky.social · 08/11/2025
Next Wednesday, we are very excited to have @joachimbaumann.bsky.social, who will present co-authored work on "Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation". Paper and information on how to join ⬇️
143
Reposted by Joachim Baumann
Tiancheng Hu @tiancheng.bsky.social · 28/10/2025
Can AI simulate human behavior? 🧠 The promise is revolutionary for science & policy. But there’s a huge "IF": Do these simulations actually reflect reality? To find out, we introduce SimBench: The first large-scale benchmark for group-level social simulation. (1/9)
1115
Joachim Baumann @joachimbaumann.bsky.social · 21/10/2025
Cool paper by @eddieyang.bsky.social, confirming our LLM hacking findings (arxiv.org/abs/2509.08825): ✓ LLMs are brittle data annotators ✓ Downstream conclusions flip frequently: LLM hacking risk is real! ✓ Bias correction methods can help but have trade-offs ✓ Use human expert whenever possible
arxiv.org
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
Large language models are rapidly transforming social science research by enabling the automation of labor-intensive tasks like data annotation and text analysis. However, LLM outputs vary significant...
0166
Reposted by Joachim Baumann
Claire Gillan @clairegillan.bsky.social · 25/09/2025
Looks interesting! We have been facing this exact issue - finding big inconsistencies across different LLMs rating the same text.
056
Reposted by Joachim Baumann
Social Computing Group - UZH @scg-uzh.bsky.social · 17/09/2025
About last week’s internal hackathon 😏 Last week, we -- the (Amazing) Social Computing Group, held an internal hackathon to work on our informally called “Cultural Imperialism” project.
131
Reposted by Joachim Baumann
Johannes B. Gruber @jbgruber.bsky.social · 15/09/2025
If you feel uneasy using LLMs for data annotation, you are right (if not, you should). It offers new chances for research that is difficult with traditional #NLP/#textasdata methods, but the risk of false conclusions is high! Experiment + *evidence-based* mitigation strategies in this preprint 👇
1224
Joachim Baumann @joachimbaumann.bsky.social · 14/09/2025
The 94% LLM hacking success rate is achieved by annotating data with several model-prompt configs, then choosing the one that yields the desired result (70% if considering SOTA models only). The 31-50% risk reflects well-intentioned researchers who just run one reasonable config w/o cherry-picking.
000
Joachim Baumann @joachimbaumann.bsky.social · 14/09/2025
Thank you, Florian :) We use two methods, CDI and DSL. Both debias LLM annotations and reduce false positive conclusions to about 3-13%, on average, but at the cost of a much higher Type II risk (up to 92%). The human-only conclusions have a pretty low Type I risk as well, at a lower Type II risk.
200
Joachim Baumann @joachimbaumann.bsky.social · 12/09/2025
Great question! Performance and LLM hacking risk are negatively correlated. So easy tasks do have lower risk. But even tasks with 96% F1 score showed up to 16% risk of wrong conclusions. Validation is important because high annotation performance doesn't guarantee correct conclusions.
131
Joachim Baumann @joachimbaumann.bsky.social · 12/09/2025
We used 199 different prompts total: some from prior work, others based on human annotation guidelines, and some simple semantic paraphrases Even when LLMs correctly identify significant effects, estimated effect sizes still deviate from true values by 40-77% (see Type M risk, Table 3 and Figure 3)
010
Joachim Baumann @joachimbaumann.bsky.social · 12/09/2025
Thank you to the amazing @paul-rottger.bsky.social @aurman21.bsky.social @albertwendsjo.bsky.social @florplaza.bsky.social @jbgruber.bsky.social @dirkhovy.bsky.social for this fun collaboration!!
060