Sign in

David Heineman

@davidheineman.com
31 followers 176 following 6 posts

Pre-doc @ai2.bsky.social davidheineman.com

PostsRepliesMedia
David Heineman @davidheineman.com · 19/08/2025
Evaluating language models is tricky, how do we know if our results are real, or due to random chance? We find an answer with two simple metrics: signal, a benchmark’s ability to separate models, and noise, a benchmark’s random variability between training steps 🧵
2154
Reposted by David Heineman
Ai2 @ai2.bsky.social · 02/06/2025
RewardBench 2 is here! We took a long time to learn from our first reward model evaluation tool to make one that is substantially harder and more correlated with both downstream RLHF and inference-time scaling.
The RewardBench 2 Leaderboard on HuggingFace.
1208