Arduin Findeis @arduin.io · 20/08/2026New blog post: Everything in AI seems to constantly change but two evaluation challenges have remained surprisingly persistent ... overfitting and saturation Like a stubborn tree in a storm of AI progress. Read the post for my intuition why I don't expect this to change: arduin.io/blog/persist... 000
Reposted by Arduin FindeisTiancheng Hu @tiancheng.bsky.social · 28/10/2025Can AI simulate human behavior? 🧠 The promise is revolutionary for science & policy. But there’s a huge "IF": Do these simulations actually reflect reality? To find out, we introduce SimBench: The first large-scale benchmark for group-level social simulation. (1/9) 1115
Arduin Findeis @arduin.io · 27/07/2025👋 I'll be at #ACL2025 presenting research from my Apple internship! Our poster is titled: "Can External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?" ☞ Let's meet: come by our poster on Tuesday (29/7), 10:30 - 12:00, Hall 4/5, or DM me to set up a meeting! ✍︎ Paper link below ↓machinelearning.apple.comCan External Validation Tools Improve Annotation Quality for LLM-as-a-Judge?Pairwise preferences over model responses are widely collected to evaluate and provide feedback to large language models (LLMs). Given two… 140
Arduin Findeis @arduin.io · 24/04/2025Excited to be in Singapore for ICLR! Keen to chat about interpreting feedback data and detecting model characteristics ⚖️ Reach out or come by our poster on Inverse Constitutional AI on Friday 25 April from 10am-12.30pm (#520 in Hall 2B) - @timokauf.bsky.social and I will be there! 000
Arduin Findeis @arduin.io · 17/04/2025How exactly was the initial Chatbot Arena version of Llama 4 Maverick different from the public HuggingFace version?🕵️ I used our Feedback Forensics app to quantitatively analyse how exactly these two models differ. An overview…👇🧵 100
Arduin Findeis @arduin.io · 17/03/2025🕵🏻💬 Introducing Feedback Forensics: a new tool to investigate pairwise preference data. Feedback data is notoriously difficult to interpret and has many known issues – our app aims to help! Try it at app.feedbackforensics.com Three example use-cases 👇🧵 172