Sign in

Florian Dorner

@flodorner.bsky.social
85 followers 280 following 43 posts

PhD student in CS @ ETHZ / MPI-IS Theory of ML evaluation flodorner.github.io

PostsRepliesMedia
Florian Dorner @flodorner.bsky.social · 23/04/2026
📅 Thursday, April 23, 2026 ⏰ 10:30 AM – 1:00 PM 📍 Pavilion 4, P4-#4413 📄 Paper: arxiv.org/pdf/2507.12399 💻 GitHub: github.com/socialfounda... Joint work with @yatongchen.bsky.social @andcrz.bsky.social and Fanny Yang
arxiv.org
020
Florian Dorner @flodorner.bsky.social · 23/04/2026
This is not just a theoretical phenomenon: In our experiments with differently sized Qwen verifiers, we see similar performance for all sizes at small N, but larger verifiers yield noticeably better performance when N is increased.
120
Florian Dorner @flodorner.bsky.social · 23/04/2026
The top-right region of the ROC determines early scaling, while the bottom-left determines behavior at large N. Thus we cannot extrapolate scaling laws from small-N observations: For any observed early scaling, there are multiple consistent ROCs, each associated with different large-N performance.
220
Florian Dorner @flodorner.bsky.social · 23/04/2026
We show that for any query, the performance of resampling methods like Best-of-N is fully determined by initial model accuracy and the verifier ROC curve. In particular, concave ROC curves imply monotonic scaling.
110
Florian Dorner @flodorner.bsky.social · 23/04/2026
At ICLR and interested in theory for LLMs? Join us at our poster to learn more about the (im)possibility of scaling laws for test-time scaling methods like Best-of-N when verification is imperfect!
132
Florian Dorner @flodorner.bsky.social · 12/12/2025
In light of the discussions about LLM-generated ICLR reviews, I recently wondered whether a similar dynamic might play out for LLMs: While pre-training objectives promote approximate indistinguishability of generated text, more and more heavy post-training might make detection a lot easier...
010
Florian Dorner @flodorner.bsky.social · 05/12/2025
In the second paper (arxiv.org/abs/2410.13341), we show that LLM judges weaker than the models they evaluate are of limited use for benchmarking, even if their judgments are processed in a statistically optimal way. Correspondingly, we cannot rely on LLM judges for evaluating frontier models.
arxiv.org
Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefore become an importan...
130
Florian Dorner @flodorner.bsky.social · 05/12/2025
In the first paper (arxiv.org/abs/2507.12399), we characterize how LLM judge errors affect test-time-scaling via Best-of-N based on the verifier ROC curve. Our results point towards more efficient alternatives to Best-of-N, and explain why scaling laws for test-time-scaling are unreliable.
arxiv.org
ROC-n-reroll: How verifier imperfection affects test-time scaling
Test-time scaling aims to improve language model performance by leveraging additional compute during inference. Many works have empirically studied techniques such as Best-of-N (BoN) and Rejection Sam...
120
Florian Dorner @flodorner.bsky.social · 05/12/2025
Meet me at the Benchmarking workshop (sites.google.com/view/benchma...) at EurIPS on Saturday: We’ll present two works on errors in LLM-as-Judge and their impacts on benchmarking and test-time-scaling:
173
Reposted by Florian Dorner
Yatong Chen @yatongchen.bsky.social · 01/12/2025
I'll be @neuripsconf.bsky.social presenting Strategic Hypothesis Testing (spotlight!) tldr: Many high-stakes decisions (e.g., drug approval) rely on p-values, but people submitting evidence respond strategically even w/o p-hacking. Can we characterize this behavior & how policy shapes it? 1/n
1173
Florian Dorner @flodorner.bsky.social · 25/10/2025
Also, from time to time, the wrong proofs it suggests for more complicated things seem to contain non-trivial insights and are "fixable".
010
Florian Dorner @flodorner.bsky.social · 25/10/2025
Not much of a step up compared to the o1/o3 "thinking" versions of GPT-4. But quite a big step compared to base GPT-4. It still makes a lot of mistakes, but often produces correct proofs for simple Lemmata (not so much for more complicated stuff).
111
Reposted by Florian Dorner
Tübingen AI Center @tuebingen-ai.bsky.social · 24/10/2025
Congratulations also to Vivian Nastl (supervised by Moritz Hardt) and Ricardo Dominguez-Olmedo (Moritz Hardt and Bernhard Schölkopf) for winning 2025 Global Google PhD fellowships. Find out more about their work here: is.mpg.de/en/news/vivi... @maxplanckcampus.bsky.social @unituebingen.bsky.social
is.mpg.de
Vivian Nastl and Ricardo Dominguez-Olmedo receive 2025 Google Ph.D. Fellowship
Program supports exceptional graduate students working on innovative research in computer science and related fields
052
Reposted by Florian Dorner
Michael Saxon @saxon.me · 18/10/2025
The viral "Definition of AGI" paper tells you to read fake references which do not exist! Proof: different articles present at the specified journal/volume/page number, and their titles exist nowhere on any searchable repository. Take this as a warning to not use LMs to generate your references!
615536
Florian Dorner @flodorner.bsky.social · 17/10/2025
Assuming all problems are actually solvable...
000
Florian Dorner @flodorner.bsky.social · 17/10/2025
Is that not trivially true, since LLMs assign nonzero probability to any possible string?
100
Reposted by Florian Dorner
Yatong Chen @yatongchen.bsky.social · 22/09/2025
We (w/ Moritz Hardt, Olawale Salaudeen and @joavanschoren.bsky.social) are organizing the Workshop on the Science of Benchmarking & Evaluating AI @euripsconf.bsky.social 2025 in Copenhagen! 📢 Call for Posters: rb.gy/kyid4f 📅 Deadline: Oct 10, 2025 (AoE) 🔗 More info: rebrand.ly/bg931sf
1217
Florian Dorner @flodorner.bsky.social · 21/09/2025
Do you have a list of the best ones? I vaguely recall reading things in this direction, but cannot really remember specific titles.
010
Reposted by Florian Dorner
Millicent Li @millicentli.bsky.social · 17/09/2025
Wouldn’t it be great to have questions about LM internals answered in plain English? That’s the promise of verbalization interpretability. Unfortunately, our new paper shows that evaluating these methods is nuanced—and verbalizers might not tell us what we hope they do. 🧵👇1/8
1268
Florian Dorner @flodorner.bsky.social · 17/09/2025
The focus on evaluating checkpoints during a training run rather than different trained models is super interesting!
110
Florian Dorner @flodorner.bsky.social · 17/09/2025
Interesting work! Can you comment a bit on what you do different compared to previous IRT-based LLM evaluation methods? We recently did some work confirming IRTs efficacy for in-distribution models, but also found it to be quite brittle when it comes to novel models arxiv.org/abs/2506.07673
arxiv.org
How Benchmark Prediction from Fewer Data Misses the Mark
Large language model (LLM) evaluation is increasingly costly, prompting interest in methods that speed up evaluation by shrinking benchmark datasets. Benchmark prediction (also called efficient LLM ev...
210
Florian Dorner @flodorner.bsky.social · 14/09/2025
I guess in terms of the notation from section 4 in the paper, does this plot Type X risk, or Type X Error Feasibility rate?
000
Florian Dorner @flodorner.bsky.social · 14/09/2025
, at least for large n. So I am trying to understand whether the asymptotics kick in a lot slower than I would have thought, or whether I am missing something else about the setup., at least for large n.
000
Florian Dorner @flodorner.bsky.social · 14/09/2025
Thank you! Do I understand correctly that these results are independent/orthogonal from the success hacking ones? I guess my confusion stems from asymptotic theory for PPI (and by extension seemingly for DSL) suggesting that both type 1 and type 2 errors should be lower/at most very similar
100
Florian Dorner @flodorner.bsky.social · 12/09/2025
Are the reported errors for the case of selecting the model with the most significant results, post-hoc?
100
Florian Dorner @flodorner.bsky.social · 12/09/2025
Interesting work! Can you comment a bit more on the setup for the regression correction methods? As far as I understand, PPI++ (which should be quite similar to DSL) relatively reliably reduces variance compared to ground truth only, while remaining quite close to unbiased.
200
Florian Dorner @flodorner.bsky.social · 08/08/2025
Does anyone have background on this plot, compared to the 32% performance for o3-mini-high with tool use claimed by OpenAI in January? #GPT5 #GPT-5 openai.com/index/introd... openai.com/index/openai...
010
Florian Dorner @flodorner.bsky.social · 23/07/2025
Super interesting field, but worth keeping in mind that this usually only buys you a relatively small fraction of "extra ground truth labels" (this does not cover active sampling strategies, but I haven not seen them yielding much larger improvements in practice, either) arxiv.org/abs/2410.13341
arxiv.org
Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefore become an importan...
020
Florian Dorner @flodorner.bsky.social · 17/07/2025
Do you have a source re: attendance requirement? 👀
100
Florian Dorner @flodorner.bsky.social · 10/05/2025
Not sure this can ethically be done retroactively (due to participant consent). But given that 20% of data is shared with model providers, privacy concerns with instead sharing this data publically in the future seem surmountable.
000
Florian Dorner @flodorner.bsky.social · 10/05/2025
New blogpost by my colleague Ricardo, arguing that instead of limiting data collection from big labs, LMArena should publicly release all data for everyone. ricardodominguez.github.io/blogs/arena....
ricardodominguez.github.io
How to Fix the Chatbot Arena? Release All Data
110
Florian Dorner @flodorner.bsky.social · 30/04/2025
Is this just the prompts, or do model providers get information about whether or not they won (and the competing response)?
100
Florian Dorner @flodorner.bsky.social · 24/04/2025
Shout out to my colleagues Ricardo Dominguez-Olmedo, Vivian Nastl and Moritz Hardt! If you’d like to chat at the conference, send me a message, or visit us at one of the poster sessions!
000
Florian Dorner @flodorner.bsky.social · 24/04/2025
100
Florian Dorner @flodorner.bsky.social · 24/04/2025
Tomorrow, I will speak about our work on the limitations of LLM-as-a-Judge 🤖 when applied to evaluating frontier models. (Session 3D) arxiv.org/abs/2410.13341
arxiv.org
Limits to scalable evaluation at the frontier: LLM as Judge won't beat twice the data
High quality annotations are increasingly a bottleneck in the explosively growing machine learning ecosystem. Scalable evaluation methods that avoid costly annotation have therefore become an importan...
110
Florian Dorner @flodorner.bsky.social · 24/04/2025
In two hours, Ricardo is giving a talk about our paper on training on the test task, and its confounding impacts on LLM benchmarking 📉📈. (Session 1B) arxiv.org/abs/2407.07890
110
Florian Dorner @flodorner.bsky.social · 24/04/2025
In Singapore for #ICLR2025 and excited for two oral presentations on work I have contributed to! 🎉
100
Florian Dorner @flodorner.bsky.social · 30/03/2025
Wouldn't the ratio of the tax burden for the average income compared to the 5x income say more about how progressive taxation is? With that, Germany seems to be the "least progressive" according to the graphic (which honestly seems a bit surprising).
110
Florian Dorner @flodorner.bsky.social · 07/03/2025
Seems worth keeping in mind, that while uncertainty can improve LLM as judge, reliable results still require debiasing using a sample ground truth data.
000
Florian Dorner @flodorner.bsky.social · 07/03/2025
We had some (very preliminary and specific) results on this in our ICLR paper (arxiv.org/abs/2410.13341), glad to see this investigated in more detail!
100
Florian Dorner @flodorner.bsky.social · 04/03/2025
Also kinda wild how the 2018 paper does not even seem to be cited in the new work
020
Florian Dorner @flodorner.bsky.social · 25/02/2025
I would say greens and SPD are closer to each other than FDP along almost any metric I can imagine. Also would not say that Linke is more "anti-system/populist" than AFD, but that one is somewhat contingent on interpreting axis labels.
010
Florian Dorner @flodorner.bsky.social · 25/02/2025
As a german, this looks hella inaccurate...
100
Florian Dorner @flodorner.bsky.social · 21/01/2025
Starting to believe @natolambert.bsky.social's take that the o1 plots are misleading [1] (in the sense that OpenAI cannot fully control test compute at inference time). In particular, it seems like scaling up test compute might require extensive retraining. [1] www.interconnects.ai/p/openais-o1...
020
Florian Dorner @flodorner.bsky.social · 20/01/2025
I meant Figure 2 in the R1 report looks like the left o1 plot if you squint hard enough (and consider the x-axis is linear rather than logarithmic)
120
Florian Dorner @flodorner.bsky.social · 20/01/2025
What exactly do you mean by that? I thought figure 2 would show increasing performance with more training steps (i.e. RL compute)?
100
Florian Dorner @flodorner.bsky.social · 19/01/2025
Appears to still be available on github github.com/hendrycks/ma...
github.com
GitHub - hendrycks/math: The MATH Dataset (NeurIPS 2021)
The MATH Dataset (NeurIPS 2021). Contribute to hendrycks/math development by creating an account on GitHub.
000
Florian Dorner @flodorner.bsky.social · 20/12/2024
Part on FrontierMath seems misleading (via reddit user named like one of the benchmark authors): www.reddit.com/r/OpenAI/com... "Tao's comments were based on a sample of T3 problems. He could almost certainly do all the T1 problems [25€] and a good number of the T2 problems. [50%]"
reddit.com
elliotglazer's comment on "OpenAI's new model, o3, shows a huge leap in the world's hardest math benchmark"
Explore this conversation and more from the OpenAI community
130