Sign in

Mila Gorecki

@milago.bsky.social
809 followers 232 following 4 posts

PhD student in Machine Learning @ MPI-IS Tübingen, Tübingen AI Center, IMPRS-IS

PostsRepliesMedia
Reposted by Mila Gorecki
Yatong Chen @yatongchen.bsky.social · 03/07/2026
Takeaway: leaderboard design is mechanism design! If you care about evals, post-training, or strategic behavior in ML, come talk to us! 📍Poster: HALL A #4413 📅 Thursday, July 9th, 2:30pm 📄 Paper link: arxiv.org/abs/2603.08371 Joint work w/ Guanhua Zhang & Moritz Hardt
arxiv.org
Leaderboard Incentives: Model Rankings under Strategic Post-Training
Influential benchmarks incentivize competing model developers to strategically allocate post-training resources toward improvements on the leaderboard, a phenomenon dubbed benchmaxxing or training on ...
021
Reposted by Mila Gorecki
Yatong Chen @yatongchen.bsky.social · 03/07/2026
LLM leaderboards aren't passive measurements — they're mechanisms that create incentives! Why does the current evaluation protocol reward "benchmaxxing": post-training for the benchmark without improving model capability? Our #ICML2026 paper gives a theoretical answer! 🧵
2144
Reposted by Mila Gorecki
minaremeli.bsky.social @minaremeli.bsky.social · 08/07/2026
Come meet me at the last poster session at ICML if you are interested in chatting about LLM evaluation and pairwise comparisons! I will be presenting joint work with my advisor (Moritz Hardt). Thu Jul 9, 5:00 PM – 6:45 PM KST, Hall A (#4411) arxiv.org/abs/2606.09409
arxiv.org
Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings
Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge bi...
1132
Reposted by Mila Gorecki
Florian Dorner @flodorner.bsky.social · 05/12/2025
Meet me at the Benchmarking workshop (sites.google.com/view/benchma...) at EurIPS on Saturday: We’ll present two works on errors in LLM-as-Judge and their impacts on benchmarking and test-time-scaling:
173
Reposted by Mila Gorecki
Gunnar König @gunnark.bsky.social · 02/12/2025
At #NeurIPS in San Diego this week? Interested in XAI, causality, or performative prediction? Come visit our poster! 💬 Performative Validity of Recourse Explanations 📆 Wednesday, 4.30 pm, Poster Session 2 w/ Hidde Fokkema, Timo Freiesleben, Celestine Mendler-Dünner, Ulrike von Luxburg
0123
Reposted by Mila Gorecki
Andreas Geiger @andreasgeiger.bsky.social · 02/12/2025
Attending #Neurips2025? Get your personalized Scholar Inbox conference program now to easily navigate the poster sessions and find what you are looking for: www.scholar-inbox.com/conference/n...
03612
Reposted by Mila Gorecki
Yatong Chen @yatongchen.bsky.social · 01/12/2025
I'll be @neuripsconf.bsky.social presenting Strategic Hypothesis Testing (spotlight!) tldr: Many high-stakes decisions (e.g., drug approval) rely on p-values, but people submitting evidence respond strategically even w/o p-hacking. Can we characterize this behavior & how policy shapes it? 1/n
1173
Mila Gorecki @milago.bsky.social · 02/12/2025
Excited to be at #Neurips2025 this week to present our paper "Monoculture or Multiplicity: Which is it?", joint work with Moritz Hardt. 📄 Paper #1000: openreview.net/pdf?id=DO5Lt... 📍 Wed, Dec 3, 2025 • 4:30 PM – 7:30 PM Feel free to come by and reach out! A short 🧵.
1164