Serina Chang @serinachang5.bsky.social · 25/04/2025Blog post on exciting research happening at MSR, including our recent work ChatBench on human-AI vs AI-alone evaluation! 000
Serina Chang @serinachang5.bsky.social · 22/04/2025Looking forward to taking part in this CHI'25 panel organized by @angelhwang.bsky.social !! 050
Serina Chang @serinachang5.bsky.social · 11/04/20251st post on bsky! What happens when a static benchmark comes to life? ✨ Introducing ChatBench, a large-scale user study where we *converted* MMLU questions into thousands of user-AI conversations. Then, we trained a user simulator on ChatBench to generate user-AI outcomes on unseen questions. 1/ 🧵 151
Reposted by Serina Changjake hofman @jakehofman.bsky.social · 09/04/2025Check out ChatBench, our new paper+dataset. We turned AI benchmarks into user-AI chats and show that AI-alone evals often fail to predict how real humans perform with AI. @serinachang5.bsky.social @ashtonanderson.bsky.social serinachang5.github.io/assets/files... huggingface.co/datasets/mic...serinachang5.github.io 0102