Sign in

Mila Gorecki

@milago.bsky.social
810 followers 232 following 4 posts

PhD student in Machine Learning @ MPI-IS Tübingen, Tübingen AI Center, IMPRS-IS

PostsRepliesMedia
Reposted by Mila Gorecki
Yatong Chen @yatongchen.bsky.social · 03/07/2026
Takeaway: leaderboard design is mechanism design! If you care about evals, post-training, or strategic behavior in ML, come talk to us! 📍Poster: HALL A #4413 📅 Thursday, July 9th, 2:30pm 📄 Paper link: arxiv.org/abs/2603.08371 Joint work w/ Guanhua Zhang & Moritz Hardt
arxiv.org
Leaderboard Incentives: Model Rankings under Strategic Post-Training
Influential benchmarks incentivize competing model developers to strategically allocate post-training resources toward improvements on the leaderboard, a phenomenon dubbed benchmaxxing or training on ...
031
Reposted by Mila Gorecki
Yatong Chen @yatongchen.bsky.social · 03/07/2026
LLM leaderboards aren't passive measurements — they're mechanisms that create incentives! Why does the current evaluation protocol reward "benchmaxxing": post-training for the benchmark without improving model capability? Our #ICML2026 paper gives a theoretical answer! 🧵
2144
Reposted by Mila Gorecki
minaremeli.bsky.social @minaremeli.bsky.social · 08/07/2026
Come meet me at the last poster session at ICML if you are interested in chatting about LLM evaluation and pairwise comparisons! I will be presenting joint work with my advisor (Moritz Hardt). Thu Jul 9, 5:00 PM – 6:45 PM KST, Hall A (#4411) arxiv.org/abs/2606.09409
arxiv.org
Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings
Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues or display judge bi...
1132
Reposted by Mila Gorecki
Florian Dorner @flodorner.bsky.social · 05/12/2025
Meet me at the Benchmarking workshop (sites.google.com/view/benchma...) at EurIPS on Saturday: We’ll present two works on errors in LLM-as-Judge and their impacts on benchmarking and test-time-scaling:
173
Reposted by Mila Gorecki
Gunnar König @gunnark.bsky.social · 02/12/2025
At #NeurIPS in San Diego this week? Interested in XAI, causality, or performative prediction? Come visit our poster! 💬 Performative Validity of Recourse Explanations 📆 Wednesday, 4.30 pm, Poster Session 2 w/ Hidde Fokkema, Timo Freiesleben, Celestine Mendler-Dünner, Ulrike von Luxburg
0123
Reposted by Mila Gorecki
Andreas Geiger @andreasgeiger.bsky.social · 02/12/2025
Attending #Neurips2025? Get your personalized Scholar Inbox conference program now to easily navigate the poster sessions and find what you are looking for: www.scholar-inbox.com/conference/n...
03612
Reposted by Mila Gorecki
Yatong Chen @yatongchen.bsky.social · 01/12/2025
I'll be @neuripsconf.bsky.social presenting Strategic Hypothesis Testing (spotlight!) tldr: Many high-stakes decisions (e.g., drug approval) rely on p-values, but people submitting evidence respond strategically even w/o p-hacking. Can we characterize this behavior & how policy shapes it? 1/n
1173
Mila Gorecki @milago.bsky.social · 02/12/2025
The empirical landscape sits between the two extremes.  - Model similarity is high, yet disagreements let individuals find recourse by switching models.  - Systemic exclusion is rare, yet more likely than under strong multiplicity.  - Even in a single model, prompt variations induce multiplicity.
030
Mila Gorecki @milago.bsky.social · 02/12/2025
We evaluate 50 LLMs (various sizes & providers) across 6 tasks to assess how well each narrative fits the current LLM landscape, assuming that decision makers will increasingly rely on these models for consequential predictions.
110
Mila Gorecki @milago.bsky.social · 02/12/2025
There are two narratives about model ecosystems that grew out of the algorithmic fairness debate: 1. Monoculture: models converge toward homogeneity. 2. Multiplicity: many models solve tasks similarly but disagree on individual predictions, creating outcome variation.
110
Mila Gorecki @milago.bsky.social · 02/12/2025
Excited to be at #Neurips2025 this week to present our paper "Monoculture or Multiplicity: Which is it?", joint work with Moritz Hardt. 📄 Paper #1000: openreview.net/pdf?id=DO5Lt... 📍 Wed, Dec 3, 2025 • 4:30 PM – 7:30 PM Feel free to come by and reach out! A short 🧵.
1164