Sign in

Chandler Smith

@chansmi.bsky.social
64 followers 230 following 17 posts

Multi-Agent Researcher at CAIF | applied research at IQT | Thinking about making MA systems go well

PostsRepliesMedia
Reposted by Chandler Smith
Anka Reuel ➡️ NeurIPS @ankareuel.bsky.social · 27/01/2025
Submitting a benchmark to ICML? Check out our NeurIPS Spotlight paper BetterBench! We outline best practices for benchmark design, implementation & reporting to help shift community norms. Be part of the change! 🙌 + Add your benchmark to our database for visibility: betterbench.stanford.edu
1133
Reposted by Chandler Smith
Vincent Conitzer @conitzer.bsky.social · 09/01/2025
The 2025 Cooperative AI summer school (9-13 July 2025 near London) is now accepting applications, due March 7th! www.cooperativeai.com/summer-schoo...
cooperativeai.com
Cooperative AI
1155
Chandler Smith @chansmi.bsky.social · 07/01/2025
Very excited to read this!
030
Chandler Smith @chansmi.bsky.social · 09/12/2024
On my way to NeurIPS ‘24 ✈️ to present our Spotlight paper Betterbench and the Concordia Contest! Would love to connect with folks and chat anything multi-agent, agentic AI, benchmarking, etc. I am applying for fall ‘25 PhDs. Ping me if you have advice or there may be a fit!
120
Chandler Smith @chansmi.bsky.social · 06/12/2024
🚀🚨 Excited to announce our work on Multi-Agent LLM Training! MALT is a multi-agent configuration that leverages synthetic data generation and credit assignment strategies for post-training specialized models solving problems together
110
Chandler Smith @chansmi.bsky.social · 26/11/2024
🚀 Check out our @neuripsconf.bsky.social Spotlight paper Betterbench, which outlines new standards in benchmarking AI! Delighted to have it featured in @techreviewjp.bsky.social
120
Reposted by Chandler Smith
Anka Reuel ➡️ NeurIPS @ankareuel.bsky.social · 25/11/2024
🚨 NeurIPS 2024 Spotlight Did you know we lack standards for AI benchmarks, despite their role in tracking progress, comparing models, and shaping policy? 🤯 Enter BetterBench–our framework with 46 criteria to assess benchmark quality: betterbench.stanford.edu 1/x
413925