Sign in

EvalEval Coalition

@eval-eval.bsky.social
128 followers 8 following 48 posts

We are a researcher community developing scientifically grounded research outputs and robust deployment infrastructure for broader impact evaluations. evalevalai.com

PostsRepliesMedia
EvalEval Coalition @eval-eval.bsky.social · 17/03/2026
3 days left! 📃 Writing, wrote, or just submitted a paper? Commit it to the EvalEval workshop at ACL 2026 in San Diego! evalevalai.com/events/2026-... (including ARR Submissions, non-archival, positions, and extended abstracts!) Submission Deadline: March 19th, 2026 AoE
evalevalai.com
2026 ACL Workshop on Evaluating AI in Practice
This workshop focuses on AI evaluation in practice, centering the tensions and collaborations between model developers and evaluation researchers and aims to surface practical insights from across the...
041
EvalEval Coalition @eval-eval.bsky.social · 11/03/2026
⏳ 9 more days! We extended the submission deadline for the EvalEval Workshop @ ACL 2026. If your work touches AI evaluation, submit! We welcome: ✅ Regular papers ✅ ARR submissions ✅ Non-archival work ✅ Position papers ✅ Extended abstracts 📅 Deadline: March 19 🌐 evalevalai.com/events/2026-...
openreview.net
ACL 2026 Workshop EvalEval
Welcome to the OpenReview homepage for ACL 2026 Workshop EvalEval
083
EvalEval Coalition @eval-eval.bsky.social · 17/02/2026
🚀 Launching Every Eval Ever: Toward a Common Language for AI Eval Reporting 🚀 A shared schema + crowdsourced repository so we can finally compare evals across frameworks and stop rerunning everything from scratch 🔧 A tale of broken AI evals 🧵👇 evalevalai.com/projects/eve...
evalevalai.com
Every Eval Ever | EvalEval Coalition
1134
EvalEval Coalition @eval-eval.bsky.social · 17/02/2026
🚨 The next edition of EvalEval Workshop is coming to @aclmeeting.bsky.social 2026! 🧠 Workshop on "AI Evaluation in Practice: Bridging Research, Development, and Real-World Impact" 🎇 📢 CFP is now open!!! More details ⏬ 📍 San Diego 📝 Submission deadline: Mar 12, 2026
163
EvalEval Coalition @eval-eval.bsky.social · 10/12/2025
It's a wrap on EvalEval in San Diego! A jam packed day of learning, making new friends, critically examining the field of evals, and walking away with renewed energy and new collaborations! We have a lot of announcements coming, but first: EvalEval will be back for #ACL2026!
151
EvalEval Coalition @eval-eval.bsky.social · 13/11/2025
🚨 AI keeps scaling, but social impact evaluations aren’t–and the data proves it 🚨 Our new paper, 📎“Who Evaluates AI’s Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations,” analyzes hundreds of evaluation reports and reveals major blind spots ‼️🧵 (1/7)
1113
EvalEval Coalition @eval-eval.bsky.social · 06/11/2025
🚨 EvalEval is back - now in San Diego!🚨 🧠 Join us for the 2025 Workshop on "Evaluating AI in Practice Bridging Statistical Rigor, Sociotechnical Insights, and Ethical Boundaries" (Co-hosted with UKAISI) 📅 Dec 8, 2025 📝 Abstract due: Nov 20, 2025 Details below! ⬇️ evalevalai.com/events/works...
evalevalai.com
131
EvalEval Coalition @eval-eval.bsky.social · 31/10/2025
✨ Weekly AI Evaluation Paper Spotlight ✨ 🤔Is it time to move beyond static tests and toward more dynamic, adaptive, and model-aware evaluation? 🖇️ "Fluid Language Model Benchmarking" by @valentinhofmann.bsky.social et. al introduces a dynamic benchmarking method for evaluating language models
130
EvalEval Coalition @eval-eval.bsky.social · 24/10/2025
🌟 Weekly AI Evaluation Spotlight 🌟 🤖 Did you know malicious actors can exploit trust in AI leaderboards to promote poisoned models in the community? This week's paper 📜"Exploiting Leaderboards for Large-Scale Distribution of Malicious Models" by @iamgroot42.bsky.social explores this!
152
EvalEval Coalition @eval-eval.bsky.social · 17/10/2025
✨Weekly AI Evaluation Paper Spotlight✨ 🕵️ Is benchmark noise and label errors masking the true fragility of LLMs? 🖇️"Do Large Language Model Benchmarks Test Reliability?" - This paper by @joshvendrow.bsky.social provides insights!
171
EvalEval Coalition @eval-eval.bsky.social · 11/08/2025
🚨New blog: The AI Evaluation Chart Crisis 📝 From misleading bar heights to missing error bars, recent model launches have sparked debate on AI evals. In our new blogpost, we dig into what’s broken, why it matters and how they should be presented 👇 evalevalai.com/documentatio...
evalevalai.com
The AI Evaluation Chart Crisis
Charts used to showcase performance demonstrate broader issues in the AI evaluation ecosystem: a lack of balance between competitive benchmarking and statistical rigor.
062
EvalEval Coalition @eval-eval.bsky.social · 16/07/2025
🚨 AI Evals Crisis: Officially kicking off the Eval Science Workstream 🚨 We’re building a shared scientific foundation for evaluating AI systems, one that’s rigorous, open, and grounded in real-world & cross-disciplinary best practices👇 (1/2) Read our new blog post: tinyurl.com/evalevalai
tinyurl.com
The Science of Evaluations: Workstream Kickoff Post
Announcing the launch of a research-driven initiative among a community of researchers to strengthen the science of AI evaluations.
121
EvalEval Coalition @eval-eval.bsky.social · 23/06/2025
Join us for the Eval Eval Coalition Social at @facct.bsky.social tomorrow Tuesday June 24th from 4-4:30 pm during the coffee break! We would love to have you join us and we look forward to seeing you there!! #FAccT2025 #EvalEval
042
EvalEval Coalition @eval-eval.bsky.social · 22/06/2025
Introducing the Eval Eval Coalition! ✨ We are a community of researchers dedicated to designing, developing, and deploying better evaluations (1/3)
131