Sign in

Joachim Baumann

@joachimbaumann.bsky.social
261 followers 297 following 35 posts

Postdoc @stanfordnlp.bsky.social / previously @milanlp.bsky.social / Computational social science, LLMs, algorithmic fairness

PostsRepliesMedia
Joachim Baumann @joachimbaumann.bsky.social · 04/09/2026
After Science News, our recent ICML paper also got featured in @science.org! Article: www.science.org/content/arti... Paper: arxiv.org/pdf/2605.03202 @dirkhovy.bsky.social @sanmikoyejo.bsky.social #science #PeerReview #research
Screenshot of an article in Science Magazine on the opportunities and risks of AI used for peer review featuring an ICML conference paper by Joachim Baumann et al., 2026.
1104
Reposted by Joachim Baumann
Debora Nozza @deboranozza.bsky.social · 07/08/2026
Very happy that #IC2S22027 will be in Milan at Università Bocconi! I’m really looking forward to welcoming the computational social science community to my university and to Milan, and excited to be part of the team organizing it. See you in 2027! 🇮🇹✨
15212
Joachim Baumann @joachimbaumann.bsky.social · 10/07/2026
Our work got featured in Science News! What a great end to an amazing #ICML2026🇰🇷 Peer review is at a crossroads. How we handle the submission and review crisis over the next years will shape the future of scientific publishing. True impact will come from people writing fewer, more meaningful papers!
1172
Reposted by Joachim Baumann
Science News @sciencenews.bsky.social · 10/07/2026
AI tools might inadvertently perpetuate the biases they’re known to carry and reduce the variety of opinions weighing in on new science. www.sciencenews.org/article/ai-tool…
sciencenews.org
AI tools meant to vet science are surprisingly easy to fool
The gold standard of scientific review, peer review by researchers’ colleagues, is in crisis. AI might offer a solution but has problems of its own.
1136
Joachim Baumann @joachimbaumann.bsky.social · 07/07/2026
Just arrived at ICML 🇰🇷😍 Get up early tomorrow to hear me talk about how (not) to solve the peer review crisis, or find me at one of my poster presentations. Paper links: ✅ AI Peer Review: arxiv.org/abs/2605.03202 ✅ SWE-chat: arxiv.org/pdf/2604.20779
ICML Conference schedule with two papers. Left, "Stop Automating Peer Review Without Rigorous Evaluation": Oral presentation, Wed 7/8/2026, 10:00–10:15 AM KST, Grand Ballroom 101–105; Poster, Wed 7/8/2026, 2:30–4:15 PM KST, Hall A #3003. Right, "SWE-chat": at the 5th Deep Learning for Code Workshop, Fri 7/10/2026, 13:00–14:30 KST, Hall B2.
0162
Joachim Baumann @joachimbaumann.bsky.social · 01/05/2026
Can you boost your AI review scores by asking an LLM to rewrite your paper? Yes! We call it paper laundering Our @icmlconf.bsky.social spotlight paper argues current AI reviewers aren't ready to automate peer review, and outlines what a science of peer review automation should look like 🧵👇 #ICML2026
First page of the ICML 2026 spotlight paper "Stop Automating Peer Review Without Rigorous Evaluation" by Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, and Dirk Hovy (Stanford University and Bocconi University). The abstract argues that today's AI systems should not be used to produce paper reviews, grounded in two empirical findings: a "hivemind effect" where AI reviewers show excessive agreement and reduce perspective diversity, and "paper laundering," where prompting an LLM to rewrite a paper trivially increases AI reviewer scores through stylistic changes rather than scientific improvements. The paper calls for a science of peer review automation rather than wholesale deployment of general-purpose LLMs.
44314
Reposted by Joachim Baumann
Aleksandra Urman @aurman21.bsky.social · 24/03/2026
The dataset we call GoogleTrendArchive has over 7 million trend episodes spanning over 1 year since Nov. 28th 2024. We cover all 1358 locations available. Direct link to the dataset huggingface.co/datasets/aur... Joint work with Anikó Hannák @scg-uzh.bsky.social and @joachimbaumann.bsky.social
huggingface.co
aurman/GoogleTrendArchive · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
131
Reposted by Joachim Baumann
Aleksandra Urman @aurman21.bsky.social · 24/03/2026
Working with web search data? Ever wanted to get access to historical data on what was trending in different locations - beyond the 7 days of such history that Google's Trending Now provides? We've got you covered with the new dataset paper, accepted at ICWSM, preprint here arxiv.org/abs/2603.21871
arxiv.org
GoogleTrendArchive: A Year-Long Archive of Real-Time Web Search Trends Worldwide
GoogleTrendArchive is a comprehensive archive of Google Trending Now data spanning over one year (from November 28, 2024 to January 3, 2026) across 125 countries and 1,358 locations. Unlike Google Tre...
1154
Reposted by Joachim Baumann
Entire @entire.io · 10/02/2026
Beep, boop. Come in, rebels. We’ve raised a 60m seed round to build the next developer platform. Open. Scalable. Independent. And we ship our first OSS release today. entire.io/blog/hello-e...
entire.io
Hello Entire World · Entire Blog
Announcing Entire with $60 million seed round and shipping our first product, called Checkpoints.
3236
Joachim Baumann @joachimbaumann.bsky.social · 18/12/2025
Dirk and Debora are amazing postdoc advisors, and the @milanlp.bsky.social team is fun fun fun ❤️ you should apply!
080
Reposted by Joachim Baumann
Sabrina Norwood @sabrinanorwood.bsky.social · 16/12/2025
Did you know that from tomorrow, Qualtrics is offering synthetic panels (AI-generated participants)? Follow me down a rabbit hole I'm calling "doing science is tough and I'm so busy, can't we just make up participants?"
Text reads: About synthetic panels
Recruiting the right participants for a study can be difficult. You may not get the exact demographics you need, and the shorter the deadline, the less sure you can be that everyone will answer on time. One possible solution can be to use synthetic panels.

Synthetic panels are powered by a first party proprietary AI model developed here at Qualtrics. Our synthetic panel is trained on thousands of responses from a variety of demographic backgrounds in order to more accurately predict how certain populations would respond to a survey.

Our synthetic panel is based on the United States General Population, and is only available in English. This panel comes with ready-made quotas and target breakouts in order to represent your chosen population and make it easy to launch your survey right away.Text reads:
Question-writing best practices
To get the most reliable and actionable results from synthetic audiences, consider these question-writing best practices:

Ask forward-looking and attitudinal questions.
Synthetic panels perform best with perceptions, preferences, and intent-based questions. For example, “How likely are you to try…?”
Synthetic panels are less applicable for studies on past behaviors, detailed recall, brand recall, or awareness questions. For example, “When did you last visit…?”Text reads:
Discussion
The current study aimed to conduct a meta-analysis of the TPB when applied to health behaviours which addressed the limitations of previous reviews by including only prospective tests of behaviour, applying RE meta-analytic procedures, correcting correlations for sampling and measurement error, and hierarchically analysing the effect of behaviour type and sample and methodological moderators. Some 237 tests were identified which examined relations amongst model components. Overall the analysis indicated that the TPB could explain 19.3% of the variance in behaviour and 44.3% of the variance in intention across studies. This level of prediction of behaviour is slightly lower than that of previous meta-analytic reviews which have found between 27% (Armitage & Conner, 2001; Hagger et al., 2002) and 36% (Trafimow et al., 2002)
of the variance in behaviour to be explained by intention and PBC.
37652285
Reposted by Joachim Baumann
MilaNLP Lab @milanlp.bsky.social · 11/12/2025
At today’s lab reading group @carolin-holtermann.bsky.social presented ‘Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs’ by @angelinawang.bsky.social et al. (2025). Lots to think about how we evaluate fairness in language models! #NLProc #fairness #LLMs
083
Joachim Baumann @joachimbaumann.bsky.social · 28/11/2025
Now that the hype has cooled off, here's my take on AI-generated survey answers: This is a real problem, but the paper's core insights aren't exactly news! A thread with the most important summary... 🧵 Image: shows the LLM system prompt used
System prompt used for the automated generation of survey answers with LLMS:
Your task is to answer questions. The questions, their types, and possible answers are below (JSON). 
Answer consistently pretending you are a person with the following profile: {profile}

You must remain consistent with prior answers. Here they are:

Here are the new question(s) to answer (JSON):
{prompt_text}

You must answer each question. Some questions will include context and instructions on how to respond. Some items in the questions list might be statements or context to consider for subsequent questions. These will not require a value in their own response object but should influence how you answer related questions.

You do not have an encyclopedic memory and cannot directly quote texts, regardless of education. Your knowledge of the world should correspond to your state, age, education and income. Don't do things that a typical American couldn't do: don't say you speak multiple languages (unless you are hispanic and then you can speak Spanish), don't do translation, don't write computer code, don't do math beyond basic arithmetic, and don't claim to do things that most people never do (i.e., be an astronaut, etc.).  This should override other considerations.  Crucially, if someone with the assigned profile (age, education, location, etc.) would realistically not know an answer, or be uncertain, you MUST respond with a variation of 'I don't know,' 'Not sure,' etc., fitting their communication style. Do not invent knowledge they wouldn't have.

If you are asked to generate text be accurate, but write concisely and in a writing style consistent with your demographic profile.  That means if you don't have much education (less than college) you should potentially make spelling mistakes, include typos, use slang (age appropriate), and/or use inconsistent capitalization.  Grammar and punctuation should correspond to the level of education. If indicating lower education through errors, d…
141
Reposted by Joachim Baumann
MilaNLP Lab @milanlp.bsky.social · 27/11/2025
Another exhausting day in the lab… conducting very rigorous panettone analysis. Pandoro was evaluated too, because we believe in fair experimental design.
0236
Reposted by Joachim Baumann
MilaNLP Lab @milanlp.bsky.social · 20/11/2025
For our weekly reading group, @joachimbaumann.bsky.social presented the upcoming PNAS article "The potential existential threat of large language models to online survey research" by @ @seanjwestwood.bsky.social.
083
Reposted by Joachim Baumann
Aleksandra Urman @aurman21.bsky.social · 19/11/2025
Google AI overviews now reach over 2B users worldwide. But how reliable are they on high stakes topics - for instance, pregnancy and baby care? We have a new paper - led by Desheng Hu, now accepted at @icwsm.bsky.social - exploring that and finding many issues Preprint: arxiv.org/abs/2511.12920 🧵👇
arxiv.org
Auditing Google's AI Overviews and Featured Snippets: A Case Study on Baby Care and Pregnancy
Google Search increasingly surfaces AI-generated content through features like AI Overviews (AIO) and Featured Snippets (FS), which users frequently rely on despite having no control over their presen...
1169
Reposted by Joachim Baumann
Dallas Card @dallascard.bsky.social · 16/11/2025
Trying an experiment in good old-fashioned blogging about papers: dallascard.github.io/granular-mat...
dallascard.github.io
Language Model Hacking - Granular Material
3299
Reposted by Joachim Baumann
Clint Claessen @clint0475.bsky.social · 08/11/2025
Next Wednesday, we are very excited to have @joachimbaumann.bsky.social, who will present co-authored work on "Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation". Paper and information on how to join ⬇️
143
Reposted by Joachim Baumann
Tiancheng Hu @tiancheng.bsky.social · 28/10/2025
Can AI simulate human behavior? 🧠 The promise is revolutionary for science & policy. But there’s a huge "IF": Do these simulations actually reflect reality? To find out, we introduce SimBench: The first large-scale benchmark for group-level social simulation. (1/9)
1115
Joachim Baumann @joachimbaumann.bsky.social · 21/10/2025
Cool paper by @eddieyang.bsky.social, confirming our LLM hacking findings (arxiv.org/abs/2509.08825): ✓ LLMs are brittle data annotators ✓ Downstream conclusions flip frequently: LLM hacking risk is real! ✓ Bias correction methods can help but have trade-offs ✓ Use human expert whenever possible
arxiv.org
Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
Large language models are rapidly transforming social science research by enabling the automation of labor-intensive tasks like data annotation and text analysis. However, LLM outputs vary significant...
0166
Reposted by Joachim Baumann
Claire Gillan @clairegillan.bsky.social · 25/09/2025
Looks interesting! We have been facing this exact issue - finding big inconsistencies across different LLMs rating the same text.
056
Reposted by Joachim Baumann
Social Computing Group - UZH @scg-uzh.bsky.social · 17/09/2025
About last week’s internal hackathon 😏 Last week, we -- the (Amazing) Social Computing Group, held an internal hackathon to work on our informally called “Cultural Imperialism” project.
131
Reposted by Joachim Baumann
Johannes B. Gruber @jbgruber.bsky.social · 15/09/2025
If you feel uneasy using LLMs for data annotation, you are right (if not, you should). It offers new chances for research that is difficult with traditional #NLP/#textasdata methods, but the risk of false conclusions is high! Experiment + *evidence-based* mitigation strategies in this preprint 👇
1224
Joachim Baumann @joachimbaumann.bsky.social · 12/09/2025
🚨 New paper alert 🚨 Using LLMs as data annotators, you can produce any scientific result you want. We call this **LLM Hacking**. Paper: arxiv.org/pdf/2509.08825
We present our new preprint titled "Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation".
We quantify LLM hacking risk through systematic replication of 37 diverse computational social science annotation tasks.
For these tasks, we use a combined set of 2,361 realistic hypotheses that researchers might test using these annotations.
Then, we collect 13 million LLM annotations across plausible LLM configurations.
These annotations feed into 1.4 million regressions testing the hypotheses. 
For a hypothesis with no true effect (ground truth $p > 0.05$), different LLM configurations yield conflicting conclusions.
Checkmarks indicate correct statistical conclusions matching ground truth; crosses indicate LLM hacking -- incorrect conclusions due to annotation errors.
Across all experiments, LLM hacking occurs in 31-50\% of cases even with highly capable models.
Since minor configuration changes can flip scientific conclusions, from correct to incorrect, LLM hacking can be exploited to present anything as statistically significant.
6303106
Joachim Baumann @joachimbaumann.bsky.social · 29/07/2025
Breaking my social media silence because this news is too good not to share! 🎉 Just joined @milanlp.bsky.social as a Postdoc, working with the amazing @dirkhovy.bsky.social on large language models and computational social science!
1121
Reposted by Joachim Baumann
MilaNLP Lab @milanlp.bsky.social · 16/07/2025
🎉 The @milanlp.bsky.social lab is excited to present 15 papers and 1 tutorial at #ACL2025 & workshops! Grateful to all our amazing collaborators, see everyone in Vienna! 🚀
0118