Sign in

Florian Dorner

@flodorner.bsky.social
85 followers 280 following 43 posts

PhD student in CS @ ETHZ / MPI-IS Theory of ML evaluation flodorner.github.io

PostsRepliesMedia
Florian Dorner @flodorner.bsky.social · 23/04/2026
This is not just a theoretical phenomenon: In our experiments with differently sized Qwen verifiers, we see similar performance for all sizes at small N, but larger verifiers yield noticeably better performance when N is increased.
120
Florian Dorner @flodorner.bsky.social · 23/04/2026
The top-right region of the ROC determines early scaling, while the bottom-left determines behavior at large N. Thus we cannot extrapolate scaling laws from small-N observations: For any observed early scaling, there are multiple consistent ROCs, each associated with different large-N performance.
220
Florian Dorner @flodorner.bsky.social · 08/08/2025
Does anyone have background on this plot, compared to the 32% performance for o3-mini-high with tool use claimed by OpenAI in January? #GPT5 #GPT-5 openai.com/index/introd... openai.com/index/openai...
010
Florian Dorner @flodorner.bsky.social · 24/04/2025
100
Florian Dorner @flodorner.bsky.social · 24/04/2025
In two hours, Ricardo is giving a talk about our paper on training on the test task, and its confounding impacts on LLM benchmarking 📉📈. (Session 1B) arxiv.org/abs/2407.07890
110
Florian Dorner @flodorner.bsky.social · 21/01/2025
Starting to believe @natolambert.bsky.social's take that the o1 plots are misleading [1] (in the sense that OpenAI cannot fully control test compute at inference time). In particular, it seems like scaling up test compute might require extensive retraining. [1] www.interconnects.ai/p/openais-o1...
020
Florian Dorner @flodorner.bsky.social · 20/01/2025
I meant Figure 2 in the R1 report looks like the left o1 plot if you squint hard enough (and consider the x-axis is linear rather than logarithmic)
120