Sign in

EvalEval Coalition

@eval-eval.bsky.social
129 followers 8 following 48 posts

We are a researcher community developing scientifically grounded research outputs and robust deployment infrastructure for broader impact evaluations. evalevalai.com

PostsRepliesMedia
EvalEval Coalition @eval-eval.bsky.social · 17/03/2026
3 days left! 📃 Writing, wrote, or just submitted a paper? Commit it to the EvalEval workshop at ACL 2026 in San Diego! evalevalai.com/events/2026-... (including ARR Submissions, non-archival, positions, and extended abstracts!) Submission Deadline: March 19th, 2026 AoE
evalevalai.com
2026 ACL Workshop on Evaluating AI in Practice
This workshop focuses on AI evaluation in practice, centering the tensions and collaborations between model developers and evaluation researchers and aims to surface practical insights from across the...
041
EvalEval Coalition @eval-eval.bsky.social · 11/03/2026
⏳ 9 more days! We extended the submission deadline for the EvalEval Workshop @ ACL 2026. If your work touches AI evaluation, submit! We welcome: ✅ Regular papers ✅ ARR submissions ✅ Non-archival work ✅ Position papers ✅ Extended abstracts 📅 Deadline: March 19 🌐 evalevalai.com/events/2026-...
openreview.net
ACL 2026 Workshop EvalEval
Welcome to the OpenReview homepage for ACL 2026 Workshop EvalEval
083
EvalEval Coalition @eval-eval.bsky.social · 17/02/2026
Read the full announcement: evalevalai.com/infrastructu... Shared Task: evalevalai.com/events/share... Project Webpage: evalevalai.com/projects/eve... #AIEvaluation #EvalEval
evalevalai.com
Every Eval Ever: Toward a Common Language for AI Eval Reporting
The multistakeholder coalition EvalEval launches Every Eval Ever, a shared format and central eval repository. We’re working to resolve AI evaluation fragmentation, improving formatting, settings, and...
000
EvalEval Coalition @eval-eval.bsky.social · 17/02/2026
Thankful to our partners for the feedback: CAISI, AIEleuther, Huggingface, NomaSecurity, TrustibleAI, InspectAI, Meridian, AVERI, CIP, Stanford HELM, Weizenbaum, Evidence Prime, MIT, TUM, IBM Research 🤝
120
EvalEval Coalition @eval-eval.bsky.social · 17/02/2026
How can you help? We are launching a shared task alongside our workshop at @aclmeeting.bsky.social → Two tracks: public + proprietary eval data → Co-authorship for qualifying contributors → Workshop at ACL 2026 (San Diego) → Deadline: May 1, 2026 📅
110
EvalEval Coalition @eval-eval.bsky.social · 17/02/2026
What we built: 📋 Metadata schema for cross-framework comparison 🔧 Validation via Hugging Face Jobs 🔌 Converters (Inspect AI, HELM, lm-eval-harness) 📊 Community repo organized by benchmark/model/run ✨ Captures scores AND context: settings, prompts, example-level data
110
EvalEval Coalition @eval-eval.bsky.social · 17/02/2026
This has real costs! 🔬 Signal buried in noise, can't tell if differences reflect model capability or just setup 📦 Evaluation debt piles up silently across the ecosystem 🔎Redundant re-runs of expensive evaluations 🌟That's where Every Eval Ever comes
110
EvalEval Coalition @eval-eval.bsky.social · 17/02/2026
🤔Consider the scenario LLaMA 65B scored 0.637 on HELM's MMLU LLaMA 65B scored 0.488 on lm-eval-harness's MMLU Same model. Same benchmark name. Different prompts, settings, extraction methods. 💡Which score is right? Both? Neither? We can't compare. 🤷
110
EvalEval Coalition @eval-eval.bsky.social · 17/02/2026
🚀 Launching Every Eval Ever: Toward a Common Language for AI Eval Reporting 🚀 A shared schema + crowdsourced repository so we can finally compare evals across frameworks and stop rerunning everything from scratch 🔧 A tale of broken AI evals 🧵👇 evalevalai.com/projects/eve...
evalevalai.com
Every Eval Ever | EvalEval Coalition
1134
EvalEval Coalition @eval-eval.bsky.social · 17/02/2026
We're seeking submissions on: 🔍 Evaluation validity & reliability 🌍 Sociotechnical impacts ⚙️ Infrastructure & costs 🤝 Community-centered approaches Full papers (6-8 pages), short papers (4 pages) or tiny papers (2 pages) welcome. Check out the full CFP: t.co/JRSr50V7Y6
t.co
https://evalevalai.com/events/2026-acl-workshop/
010
EvalEval Coalition @eval-eval.bsky.social · 17/02/2026
🚨 The next edition of EvalEval Workshop is coming to @aclmeeting.bsky.social 2026! 🧠 Workshop on "AI Evaluation in Practice: Bridging Research, Development, and Real-World Impact" 🎇 📢 CFP is now open!!! More details ⏬ 📍 San Diego 📝 Submission deadline: Mar 12, 2026
163
EvalEval Coalition @eval-eval.bsky.social · 10/12/2025
Thank you to everyone who attended, presented at, spoke at, or helped organize this workshop. You rock! Special thanks to the UK AI Security Institute for cohosting and their support.
000
EvalEval Coalition @eval-eval.bsky.social · 10/12/2025
It's a wrap on EvalEval in San Diego! A jam packed day of learning, making new friends, critically examining the field of evals, and walking away with renewed energy and new collaborations! We have a lot of announcements coming, but first: EvalEval will be back for #ACL2026!
151
EvalEval Coalition @eval-eval.bsky.social · 13/11/2025
📜Paper: arxiv.org/pdf/2511.056... 📝Blog: tinyurl.com/blogAI1 🤝At EvalEval, we are a coalition of researchers working towards better AI evals. Interested in joining us? Check out: evalevalai.com 7/7 🧵
arxiv.org
000
EvalEval Coalition @eval-eval.bsky.social · 13/11/2025
Continued.. 📉 Reporting on social impact dimensions has steadily declined, both in frequency and detail, across major providers 🧑‍💻 Sensitive content gets the most attention, as it’s easier to define and measure 🛡️Solution? Standardized reporting & safety policies (6/7)
100
EvalEval Coalition @eval-eval.bsky.social · 13/11/2025
Key Takeaways: ⛔️ First-party reporting is often sparse & superficial, with many reporting NO social impact evals 📉 On average, first-party scores are far lower than third-party evals (0.72 vs 2.62/3) 🎯 Third parties provide some complementary coverage (GPT-4 and LLaMA) (5/7)
110
EvalEval Coalition @eval-eval.bsky.social · 13/11/2025
💡 We also interviewed developers from for-profit and non-profit orgs to understand why some disclosures happen and why others don’t. 💬 TLDR: Incentives and constraints shape reporting (4/7)
100
EvalEval Coalition @eval-eval.bsky.social · 13/11/2025
📊 What we did: 🔎 Analyzed 186 first-party release reports from model developers & 183 post-release evaluations (third-party) 📏 Scored 7 social impact dimensions: bias, harmful content, performance disparities, environmental costs, privacy, financial costs, & labor (3/7)
100
EvalEval Coalition @eval-eval.bsky.social · 13/11/2025
While general capability evaluations are common, social impact assessments, covering bias, fairness, and privacy, etc., are often fragmented or missing. 🧠 🎯Our goal: Explore the AI Eval landscape to answer who evaluates what and identify gaps in social impact evals!! (2/7)
110
EvalEval Coalition @eval-eval.bsky.social · 13/11/2025
🚨 AI keeps scaling, but social impact evaluations aren’t–and the data proves it 🚨 Our new paper, 📎“Who Evaluates AI’s Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations,” analyzes hundreds of evaluation reports and reveals major blind spots ‼️🧵 (1/7)
1113
EvalEval Coalition @eval-eval.bsky.social · 06/11/2025
Note: General registration is constrained by space capacity! Please note that attendance will be confirmed by the organizers based on space availability. Accepted posters will be invited to register for free and attend the workshop in person!
000
EvalEval Coalition @eval-eval.bsky.social · 06/11/2025
📮 We are inviting students and early-stage researchers to submit an Abstract (Max 500 words) to be presented as posters during interactive session. Submit here: tinyurl.com/AbsEval We have a rock-star lineup of AI researchers and an amazing program. Please RSVP at the earliest! Stay tuned!
100
EvalEval Coalition @eval-eval.bsky.social · 06/11/2025
🚨 EvalEval is back - now in San Diego!🚨 🧠 Join us for the 2025 Workshop on "Evaluating AI in Practice Bridging Statistical Rigor, Sociotechnical Insights, and Ethical Boundaries" (Co-hosted with UKAISI) 📅 Dec 8, 2025 📝 Abstract due: Nov 20, 2025 Details below! ⬇️ evalevalai.com/events/works...
evalevalai.com
131
EvalEval Coalition @eval-eval.bsky.social · 31/10/2025
💡This paper was brought to you as part of our spotlight series featuring papers on evaluation methods & datasets, the science of evaluation, and many more. 📸Interested in working on better AI evals? Join us: evalevalai.com
020
EvalEval Coalition @eval-eval.bsky.social · 31/10/2025
🚫 The approach also avoids mislabeled data and delays benchmark saturation, continuing to distinguish model improvements even at high performance levels. 📑Read more: arxiv.org/abs/2509.11106
120
EvalEval Coalition @eval-eval.bsky.social · 31/10/2025
📊Results & Findings 🧪 Experiments across 6 LLMs and 6 major benchmarks: 🏃Fluid Benchmarking outperforms all baselines across all four evaluation dimensions: efficiency, validity, variance, and saturation. ⚡️It achieves lower variance with up to 50× fewer items needed!!
110
EvalEval Coalition @eval-eval.bsky.social · 31/10/2025
It combines two key ideas: ✍️Item Response Theory: Models LLM performance in a latent ability space based on item difficulty and discrimination across models 🧨Dynamic Item Selection: Adaptive benchmarking-weaker models get easier items, while stronger models face harder ones
110
EvalEval Coalition @eval-eval.bsky.social · 31/10/2025
🔍How to address this? 🤔 🧩Fluid Benchmarking: This work proposes a framework inspired by psychometrics that uses Item Response Theory (IRT) and adaptive item selection to dynamically tailor benchmark evaluations to each model’s capability level. Continued...👇
110
EvalEval Coalition @eval-eval.bsky.social · 31/10/2025
⚠️ Evaluation results can be noisy and prone to variance & labeling errors. 🧱As models advance, benchmarks tend to saturate quickly, reducing their longterm usefulness. 🪃Existing approaches typically tackle just one of these problems (e.g., efficiency or validity) What now⁉️
110
EvalEval Coalition @eval-eval.bsky.social · 31/10/2025
💣Current SOTA benchmarking setups face several systematic issues: 📉It’s often unclear which benchmark(s) to choose, while evaluating on all available ones is too expensive, inefficient, and not always aligned with the intended capabilities we want to measure. More 👇👇
110
EvalEval Coalition @eval-eval.bsky.social · 31/10/2025
✨ Weekly AI Evaluation Paper Spotlight ✨ 🤔Is it time to move beyond static tests and toward more dynamic, adaptive, and model-aware evaluation? 🖇️ "Fluid Language Model Benchmarking" by @valentinhofmann.bsky.social et. al introduces a dynamic benchmarking method for evaluating language models
130
EvalEval Coalition @eval-eval.bsky.social · 24/10/2025
💡This is part of our new weekly spotlight series that will feature papers on evaluation methods & datasets, the science of evaluation, and many more. 📷 Interested in working on better AI evals? Check out: evalevalai.com
evalevalai.com
EvalEval Coalition
We are a researcher community developing scientifically grounded research outputs and robust deployment infrastructure for broader impact evaluations.
010
EvalEval Coalition @eval-eval.bsky.social · 24/10/2025
🏗️Therefore, fixing leaderboard design, e.g., private eval sets, provenance checks, randomized human tests, etc., is critical for AI ecosystem security and safety Read more: arxiv.org/pdf/2507.08983
arxiv.org
110
EvalEval Coalition @eval-eval.bsky.social · 24/10/2025
📊Key insights 🗳️Popular leaderboards (e.g., ChatArena, MTEB) can be exploited to distribute poisoned LLMs at scale 🔐Derivative models (finetuned, quantized, “abliterated”) are easy backdoor vectors. For instance, unsafe LLM variants often get downloaded as much as originals! Continued...
100
EvalEval Coalition @eval-eval.bsky.social · 24/10/2025
🔍 Method: 🧮Introduces TrojanClimb, a framework showing how attackers can: ⌨️ Simulate leaderboard attacks where malicious models achieve high test scores while embedding harmful pay loads (4 modalities) 🔒 Leverage stylistic watermarks/tags to game voting-based leaderboards
100
EvalEval Coalition @eval-eval.bsky.social · 24/10/2025
🌟 Weekly AI Evaluation Spotlight 🌟 🤖 Did you know malicious actors can exploit trust in AI leaderboards to promote poisoned models in the community? This week's paper 📜"Exploiting Leaderboards for Large-Scale Distribution of Malicious Models" by @iamgroot42.bsky.social explores this!
152
EvalEval Coalition @eval-eval.bsky.social · 17/10/2025
💡This spotlight series will feature papers on evaluation methods & datasets, the science of evaluation, and many more. Stay tuned! 🤝 Interested in working on better AI evals? We are a coalition of researchers working towards better AI evals. Check out: evalevalai.com
evalevalai.com
EvalEval Coalition
We are a researcher community developing scientifically grounded research outputs and robust deployment infrastructure for broader impact evaluations.
030
EvalEval Coalition @eval-eval.bsky.social · 17/10/2025
🧮 Benchmark Saturation != Reliability. Models achieve near-perfect scores without demonstrating true reliability. 📢 Highlights the gap between apparent competence & dependable reliability - therefore systematic reliability testing is needed. Read more at: arxiv.org/pdf/2502.03461
arxiv.org
130
EvalEval Coalition @eval-eval.bsky.social · 17/10/2025
📊Key insights: ‼️Noise in benchmarks is substantial! For some datasets, up to 90% of reported “model errors” actually stem from *bad data* instead of model failures. 🧠 After benchmark cleaning, even top LLMs fail on simple, unambiguous platinum benchmark tasks. Continued...
143
EvalEval Coalition @eval-eval.bsky.social · 17/10/2025
🔍 Method: 🧹 Revise & clean 15 popular LLM benchmarks across 6 domains to create *platinum* benchmarks. 🤖 Use multiple LLMs to flag inconsistent samples via disagreement. ⚠️ Bad” questions fall into 4 types: mislabeled, contradictory, ambiguous, or ill-posed. Example 👇
130
EvalEval Coalition @eval-eval.bsky.social · 17/10/2025
✨Weekly AI Evaluation Paper Spotlight✨ 🕵️ Is benchmark noise and label errors masking the true fragility of LLMs? 🖇️"Do Large Language Model Benchmarks Test Reliability?" - This paper by @joshvendrow.bsky.social provides insights!
171
EvalEval Coalition @eval-eval.bsky.social · 11/08/2025
🚨New blog: The AI Evaluation Chart Crisis 📝 From misleading bar heights to missing error bars, recent model launches have sparked debate on AI evals. In our new blogpost, we dig into what’s broken, why it matters and how they should be presented 👇 evalevalai.com/documentatio...
evalevalai.com
The AI Evaluation Chart Crisis
Charts used to showcase performance demonstrate broader issues in the AI evaluation ecosystem: a lack of balance between competitive benchmarking and statistical rigor.
062
EvalEval Coalition @eval-eval.bsky.social · 16/07/2025
This kickoff post lays out: 1) 🔍 Why we need a science of evaluation; 2) 🤝 Our goals for the community; 3) 🛠️ How you can get involved (2/2) Interested in joining? Check out evalevalai.com
evalevalai.com
EvalEval Coalition
We are a researcher community developing scientifically grounded research outputs and robust deployment infrastructure for broader impact evaluations.
000
EvalEval Coalition @eval-eval.bsky.social · 16/07/2025
🚨 AI Evals Crisis: Officially kicking off the Eval Science Workstream 🚨 We’re building a shared scientific foundation for evaluating AI systems, one that’s rigorous, open, and grounded in real-world & cross-disciplinary best practices👇 (1/2) Read our new blog post: tinyurl.com/evalevalai
tinyurl.com
The Science of Evaluations: Workstream Kickoff Post
Announcing the launch of a research-driven initiative among a community of researchers to strengthen the science of AI evaluations.
121
EvalEval Coalition @eval-eval.bsky.social · 23/06/2025
Join us for the Eval Eval Coalition Social at @facct.bsky.social tomorrow Tuesday June 24th from 4-4:30 pm during the coffee break! We would love to have you join us and we look forward to seeing you there!! #FAccT2025 #EvalEval
042
EvalEval Coalition @eval-eval.bsky.social · 22/06/2025
We would love to have you join us!! Check out evaleval.github.io for more info and stay tuned for future updates!! #EvalEval #AIEvaluations (3/3)
evaleval.github.io
EvalEval Coalition
We are a researcher community developing scientifically grounded research outputs and robust deployment infrastructure for broader impact evaluations.
010
EvalEval Coalition @eval-eval.bsky.social · 22/06/2025
Our coalition is focused on producing scientifically grounded research outputs, robust deployment infrastructure for broader impact evaluations, and fostering a community of researchers passionate about developing better evaluations 🌎🌍🌏 (2/3)
110
EvalEval Coalition @eval-eval.bsky.social · 22/06/2025
Introducing the Eval Eval Coalition! ✨ We are a community of researchers dedicated to designing, developing, and deploying better evaluations (1/3)
131