Sign in

Tanise Ceron

@taniseceron.bsky.social
163 followers 173 following 38 posts

Postdoc @milanlp.bsky.social | Interested in language models and how they shape the information environment

PostsRepliesMedia
Reposted by Tanise Ceron
Data Science Lab - Hertie School @hertiedatascience.bsky.social · 18/09/2026
Join us for the first Brown Bag event of the fall semester, featuring @taniseceron.bsky.social Tanise Ceron, a Postdoctoral Research Fellow at the MilaNLP group at Università Bocconi. 📅 22 September ⏰ 12:00-13:00, CEST 📍Hertie School, Maker space Register 🔗 www.hertie-school.org/en/datascien...
041
Reposted by Tanise Ceron
MilaNLP Lab @milanlp.bsky.social · 12/06/2026
@taniseceron.bsky.social is presenting her work about political content in pre-training and post-training data at the AI & Society conference. #AIandSociety #NLProc
0175
Tanise Ceron @taniseceron.bsky.social · 30/01/2026
Come join our group! Still one day left for applying. 😊
000
Tanise Ceron @taniseceron.bsky.social · 30/01/2026
Some findings that I find particularly impactful for the area of political biases in LLMs: 1) Aligning LLMs with DPO on left-leaning opinions does not have a significant impact on the stance of the models given that vanilla LLMs already reflect a more left-leaning alignment.
110
Reposted by Tanise Ceron
MilaNLP Lab @milanlp.bsky.social · 18/12/2025
🚀 We’re opening 2 fully funded postdoc positions in #NLP! Join the MilaNLP team and contribute to our upcoming research projects. 🔗 More details: milanlproc.github.io/open_positio... ⏰ Deadline: Jan 31, 2026
01913
Tanise Ceron @taniseceron.bsky.social · 02/12/2025
I will be @euripsconf.bsky.social this week to present our paper as non-archival at the PAIG workshop (Beyong Regulation: Private Governance & Oversight Mechanisms for AI). Very much looking forward to the discussions! If you are at #EurIPS and want to chat about LLM's training data. Reach out!
094
Tanise Ceron @taniseceron.bsky.social · 27/11/2025
We go out of the routine every now and then at the lab. :)
140
Tanise Ceron @taniseceron.bsky.social · 24/11/2025
@agnesedaff.bsky.social presented our work on "Generalizability of Media Frames: Corpus creation and analysis across countries" at *SEM co-located with EMNLP 2025 in China.
291
Tanise Ceron @taniseceron.bsky.social · 18/11/2025
Does anyone know any good resource that systematically documents information about the training data of different LLMs (e.g. name of datasets, language proportion, etc whenever available)?
220
Reposted by Tanise Ceron
MilaNLP Lab @milanlp.bsky.social · 31/10/2025
Proud to present our #EMNLP2025 papers! Catch our team across Main, Findings, Workshops & Demos 👇
12114
Tanise Ceron @taniseceron.bsky.social · 29/09/2025
📣 New Preprint! Have you ever wondered what the political content in LLM's training data is? What are the political opinions expressed? What is the proportion of left- vs right-leaning documents in the pre- and post-training data? Do they correlate with the political biases reflected in models?
24613
Reposted by Tanise Ceron
arXiv cs.CL Computation and Language @cscl-bot.bsky.social · 29/09/2025
Tanise Ceron, Dmitry Nikolaev, Dominik Stammbach, Debora Nozza: What Is The Political Content in LLMs' Pre- and Post-Training Data? arxiv.org/abs/2509.22367 arxiv.org/pdf/2509.22367 arxiv.org/html/2509.22367
013
Tanise Ceron @taniseceron.bsky.social · 26/09/2025
Today Sourabh Dattawad presented our work "Leveraging Media Frames to Improve Normative Diversity in News Recommendations" at INRA (International Workshop on News Recommendation and Analytics) co-located with RecSys 2025 in Prague. arxiv.org/pdf/2509.02266
121
Reposted by Tanise Ceron
Joachim Baumann @joachimbaumann.bsky.social · 12/09/2025
🚨 New paper alert 🚨 Using LLMs as data annotators, you can produce any scientific result you want. We call this **LLM Hacking**. Paper: arxiv.org/pdf/2509.08825
We present our new preprint titled "Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation".
We quantify LLM hacking risk through systematic replication of 37 diverse computational social science annotation tasks.
For these tasks, we use a combined set of 2,361 realistic hypotheses that researchers might test using these annotations.
Then, we collect 13 million LLM annotations across plausible LLM configurations.
These annotations feed into 1.4 million regressions testing the hypotheses. 
For a hypothesis with no true effect (ground truth $p > 0.05$), different LLM configurations yield conflicting conclusions.
Checkmarks indicate correct statistical conclusions matching ground truth; crosses indicate LLM hacking -- incorrect conclusions due to annotation errors.
Across all experiments, LLM hacking occurs in 31-50\% of cases even with highly capable models.
Since minor configuration changes can flip scientific conclusions, from correct to incorrect, LLM hacking can be exploited to present anything as statistically significant.
6303106
Reposted by Tanise Ceron
MilaNLP Lab @milanlp.bsky.social · 07/07/2025
Last week we held our 1st MilaNLP retreat by beautiful Lago Maggiore! ⛰️🌊 We shared research ideas, stories (academic & beyond), and amazing food. It was a great time to connect outside of the usual lab working days, and most importantly, strengthen our bonds as a team. #ResearchLife #NLProc
0225
Reposted by Tanise Ceron
Beatrice Savoldi @bsavoldi.bsky.social · 03/06/2025
🔍 Stiamo studiando come l'AI viene usata in Italia e per farlo abbiamo costruito un sondaggio! 👉 bit.ly/sondaggio_ai... (è anonimo, richiede ~10 minuti, e se partecipi o lo fai girare ci aiuti un sacco🙏) Ci interessa anche raggiungere persone che non si occupano e non sono esperte di AI!
bit.ly
Qualtrics Survey | Qualtrics Experience Management
The most powerful, simple and trusted way to gather experience data. Start your journey to experience management and try a free account today.
11618
Tanise Ceron @taniseceron.bsky.social · 15/05/2025
Reminder for the importance of evaluating political biases robustly. :)
010
Reposted by Tanise Ceron
Dirk Hovy @dirkhovy.bsky.social · 03/05/2025
We (w/ @diyiyang.bsky.social, @zhuhao.me, & Bodhisattwa Prasad Majumder) are excited to present our #NAACL25 tutorial on Social Intelligence in the Age of LLMs! It will highlight long-standing and emerging challenges of AI interacting w humans, society & the world. ⏰ May 3, 2:00pm-5:30pm Room Pecos
0156
Reposted by Tanise Ceron
Verena Kunz @verenakunz.bsky.social · 30/04/2025
Join us in an hour at 17:00 (CEST) for @taniseceron.bsky.social's talk on "Evaluating Political Bias: Insights into Robustness and Multilinguality“. Access to Zoom at join.slack.com/t/tadapolisc... or send me a ✉️
011
Reposted by Tanise Ceron
Christopher Klamm 🏔️ @cklamm.bsky.social · 21/04/2025
🥁 It's the second half of our 🌱 speaker series (tada.cool) this term, and we couldn't be more excited! Next week (Wednesday, April 30 at 5pm CET), we have the pleasure of welcoming @taniseceron.bsky.social to share insights on "Facilitating Information Access Through Language Models". More details ⬇️
075
Reposted by Tanise Ceron
Dirk Hovy @dirkhovy.bsky.social · 05/03/2025
Wanna keep up with our @milanlp.bsky.social lab? Here is a starter pack of current and former members: bsky.app/starter-pack...
0137
Tanise Ceron @taniseceron.bsky.social · 05/03/2025
Happy to be presenting at #TaDa and looking forward to watching the great talks coming up. :)
032
Tanise Ceron @taniseceron.bsky.social · 17/02/2025
🤯 Perhaps it's time to start challenging the culture in our area of "submitting to get some feedback". Isn't it more productive to submit only when we feel the paper is really ready to be submitted?
3121