Reposted by Sara HookerPrinceton Center for Information Technology Policy @princetoncitp.bsky.social · 05/05/2025⚠️ Leaderboard Illusion: "We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release & retract scores if desired..the ability of these providers to choose the best score leads to biased Arena scores" Paper out now!🔻 292
Sara Hooker @sarahooker.bsky.social · 30/04/2025It is critical for scientific integrity that we trust our measure of progress. The @lmarena.bsky.social has become the go-to evaluation for AI progress. Our release today demonstrates the difficulty in maintaining fair evaluations on the Arena, despite best intentions. 19569
Reposted by Sara HookerMarzieh Fadaee @mziizm.bsky.social · 30/04/20251/ Science is only as strong as the benchmarks it relies on. So how fair—and scientifically rigorous—is today’s most widely used evaluation benchmark? We took a deep dive into Chatbot Arena to find out. 🧵 1286
Reposted by Sara HookerJonathan Wenger @jwenger.bsky.social · 16/04/2025This has been a topic close to my heart for a long time. We have an awesome lineup of speakers who have made deep contributions to open-source in ML, e.g. @sarahooker.bsky.social , @chrisrackauckas.bsky.social, Matt Johnson, Tri Dao, @stellaathena.bsky.social, Evan Shelhamer. 0102
Reposted by Sara HookerIsra Salazar @israsalazar.bsky.social · 10/04/2025Today we are releasing Kaleidoscope 🎉 A comprehensive multimodal & multilingual benchmark for VLMs! It contains real questions from exams in different languages. 🌍 20,911 questions and 18 languages 📚 14 subjects (STEM → Humanities) 📸 55% multimodal questions 1266
Sara Hooker @sarahooker.bsky.social · 19/03/2025It is rare I get to completely disconnect. Very grateful for this week in Patagonia. 2330
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 05/03/2025We're particularly proud to release Aya Vision 8B - it's compact 🐭 and efficient 🐎, outperforming models up to 11x its size 📈. Releasing open weights helps to make breakthroughs in VLMs accessible to the research community. 1144
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 06/03/2025Just 2 days after launch, Aya Vision is trending on @hf.co 🔥🔥 We launched open-weights with the goal of making VLM breakthroughs accessible to the research community - so exciting to see such a positive response. huggingface.co/CohereForAI/... 072
Reposted by Sara Hooker(((Steve Chapman))) @stevechapman.bsky.social · 03/03/2025Love this post by @sarahooker.bsky.social on that other platform: "The first step of any meaningful pursuit is to severely underestimate its difficulty." 051
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 04/03/2025Introducing ✨ Aya Vision ✨ - an open-weights model to connect our world through language and vision Aya Vision adds breakthrough multimodal capabilities to our state-of-the-art multilingual 8B and 32B models. 🌿 183
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 25/02/2025An important topic in AI is the climate impacts of the energy-intensive computing hardware needed to train and deploy AI models ⚡ Our policy primer explores ways to move towards more sustainable AI. 🌱 📜 cohere.com/research/pap... 021
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 11/02/2025Does more compute equate with greater risk?⚡️What is our track record predicting what risks emerge with scale? 📈 In this work led by Sara Hooker, we seek to understand the viability of compute thresholds ⚖️ as a way to mitigate risk. 🦺 arxiv.org/abs/2407.05694 011
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 18/02/2025In this work, we ask "How does model merging stack up when optimizing language models for diverse multitask learning?" 📚🧩 📜https://arxiv.org/abs/2410.10801 051
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 22/01/2025Aya Expanse, our open-weight 32B model, outperforms drastically larger models including Claude, Mistral Large 2, & Llama 405B on Scale's Private Multilingual Protocol. We are proud to work on global AI that is efficient and accessible 🔥 163
Reposted by Sara HookerJekaterina Novikova @j-novikova-nlp.bsky.social · 23/01/2025Our paper is accepted to ICLR! INCLUDE: Evaluating Multilingual LLMs with Regional Knowledge (arxiv.org/abs/2411.19799) A benchmark of ~200k QA pairs across 44 languages, capturing real-world cultural nuances. A collaborative effort led by @cohereforai.bsky.social, with contributors worldwide. /1 1114
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 12/02/2025In this cross-institutional work, we introduce technical governance for AI and 100+ 🔢 open technical problems 🔧. We provide a taxonomy of open problem areas in TAIG organized by governance capacities and governance targets. 📜https://arxiv.org/pdf/2407.14981 022
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 13/02/2025The C4AI Research Grant program is proud to have supported a project focused on building LLM tools for teachers 🧑🏫 This project focused on adapting educational materials to students’ skill levels, ensuring more effective and responsible AI integration in classrooms. 141
Sara Hooker @sarahooker.bsky.social · 13/02/2025Many people have asked me about the France Action Summit. I think a summit is typically most valuable as a catalyst, not as a solution in itself. But, will share some observations. 24210
Reposted by Sara HookerKathy Baxter @baxterkb.bsky.social · 11/02/2025Boris Gamazaychikov, @salesforce.com Head of #AI #Sustainability announced the AI Energy Score we launched at the AI Action Summit in Paris. 🌍 This offers a standardized way to measure & compare the energy efficiency of AI models. 🫶 www.linkedin.com/posts/bgamaz...linkedin.com 1174
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 28/01/2025"Anyone who is serious about what the next generation of models is knows it can't be the current" Thanks to @baratunde.com for hosting Head of Cohere For AI, @sarahooker.bsky.social on the latest episode of Life with Machines. Check out their full conversation on YouTube: youtu.be/-BsobAoOJvkyoutu.beIs AI on the Verge of a Meltdown? | Sara Hooker (Ep. 8)YouTube video by Baratunde Thurston 1131
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 21/01/2025On Scale AI's private multilingual protocol, Aya Expanse is indexed as the best open-weights model. Additionally, in some languages we're outperforming: 🔒proprietary models 🐘larger models ⛰️models built by more researchers with more infrastructure Lots to be proud of today. 181
Sara Hooker @sarahooker.bsky.social · 18/01/2025As the @cohereforai.bsky.social joins the Bluesky family — we will be sharing paper gems from when we first started as a lab. This paper is part of a larger research agenda where we have focused on how to better represent the long tail = making AI work for almost all real world distributions. 0253
Sara Hooker @sarahooker.bsky.social · 17/01/2025Last year we published a fantastic cross-institutional survey on efficiency techniques for language models. Comprehensive and a good starting pointing for researchers working on efficiency. 091
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 16/01/2025How do we do more 🐘 with less 🐁? In an era of ever larger models, work on efficiency is ever more important. This cross-institutional collaboration provides a survey of the field for practitioners and researchers alike ⚙️. 📜Learn more: arxiv.org/pdf/2209.000... 031
Reposted by Sara HookerCohere Labs @cohereforai.bsky.social · 15/01/2025We are committed to making meaningful progress in machine learning research through open collaboration. Follow this 🧵to stay on top of our research contributions. 58222
Sara Hooker @sarahooker.bsky.social · 14/01/2025Changing the spaces where AI breakthroughs happen. ✨ Join us. 🔥 0173
Reposted by Sara HookerAlice Oh @aliceatkaist.bsky.social · 09/01/2025Bye Dagstuhl! Huge thanks to @841io.bsky.social et al for organizing, and @kanarinka.bsky.social @anitachan.bsky.social @sarahooker.bsky.social @haldaume3.bsky.social and many others for eye opening discussions 🙏 2101
Reposted by Sara HookerBob E Hayes @bobehayes.bsky.social · 20/12/2024This is where the data to build #AI comes from | @melissahei.bsky.social @splendidsteph.bsky.social “We are using these models all over the world, and there’s a massive discrepancy between the world we’re seeing and what’s invisible to these models." @sarahooker.bsky.social www.technologyreview... 0103
Reposted by Sara HookerBob E Hayes @bobehayes.bsky.social · 10/12/2024The Reality of #AI and Biorisk "We find that existing studies around AI-related biorisk are nascent, often speculative in nature, or limited in terms of their methodological maturity and transparency." ~ @sarahooker.bsky.social et al. arxiv.org/abs/2412.0... #GenerativeAIarxiv.orgThe Reality of AI and BioriskTo accurately and confidently answer the question 'could an AI model or system increase biorisk', it is necessary to have both a sound theoretical threat model for how AI models or systems could... 052
Reposted by Sara HookerBob E Hayes @bobehayes.bsky.social · 06/12/2024The Reality of #AI and Biorisk "We find that existing studies around AI-related biorisk are nascent, often speculative in nature, or limited in terms of their methodological maturity and transparency." ~ @sarahooker.bsky.social et al. arxiv.org/abs/2412.0... #GenerativeAIarxiv.orgThe Reality of AI and BioriskTo accurately and confidently answer the question 'could an AI model or system increase biorisk', it is necessary to have both a sound theoretical threat model for how AI models or systems could... 092
Reposted by Sara HookerMarzieh Fadaee @mziizm.bsky.social · 06/12/2024🚀 Our mission to strengthen the multilingual open-source ecosystem continues!👇 192
Reposted by Sara HookerAngelika Romanou @agromanou.bsky.social · 05/12/2024Introducing Global-MMLU🌍: A multilingual benchmark featuring MMLU translations in 42 languages crafted with: ✅ Human curation ✅ Extensive metadata ✅ Insights into cultural sensitivity Proud to have collaborated with Shivalika Singh, @sarahooker.bsky.social and Cohere For AI! 0134
Reposted by Sara HookerLeshem (Legend) Choshen @EMNLP @lchoshen.bsky.social · 05/12/2024You would think moral questions are universal, MMLU only asks about US morals... better translations and separation by sensitivity 👇👇 0132
Reposted by Sara HookerShayne Longpre @shaynelongpre.bsky.social · 05/12/2024Interested in how LLMs are really used? We are starting a research project to find out! In collaboration w/ @sarahooker.bsky.social @ankareuel.bsky.social and others. We are looking for two junior researchers to join us. Apply by Dec 15th! forms.gle/H2o3cNCPdG8e...forms.gleGoogle Forms: Sign-inAccess Google Forms with a personal Google account or Google Workspace account (for business use). 0151
Sara Hooker @sarahooker.bsky.social · 05/12/2024Is MMLU Western-centric? 🤔 As part of a massive cross-institutional collaboration: 🗽Find MMLU is heavily overfit to western culture 🔍 Professional annotation of cultural sensitivity data 🌍 Release improved Global-MMLU 42 languages 📜 Paper: arxiv.org/pdf/2412.03304 📂 Data: hf.co/datasets/Coh... 75912
Reposted by Sara HookerAidan P @aidanpeppin.bsky.social · 04/12/2024AI amplifying biorisk has been a major topic in policy & governance work. But does the available evidence match this level of attention? 🦠 ⚠️ Our new paper looks at the science underpinning ideas that AI could increase biorisks. arxiv.org/abs/2412.01946arxiv.orgThe Reality of AI and BioriskTo accurately and confidently answer the question 'could an AI model or system increase biorisk', it is necessary to have both a sound theoretical threat model for how AI models or systems could incre... 1162
Reposted by Sara HookerHal Daumé III @haldaume3.bsky.social · 04/12/2024« no » definitely this was strongly my prior but it’s good to see this worked out to hopefully shape where investments go 0142
Reposted by Sara HookerAntoine Bosselut @abosselut.bsky.social · 02/12/2024Translating MMLU is great, but global users of multilingual #LLMs don't care all that much about an LLM's understanding of US Law! Our new #NLProc work centers multilingual #LLM evaluations toward regional knowledge in 44 languages. 1283
Reposted by Sara HookerMarzieh Fadaee @mziizm.bsky.social · 03/12/2024INCLUDE is a massive benchmark across 44 languages curated from 52 countries and includes both regional and cultural knowledge. 193
Reposted by Sara HookerMarzieh Fadaee @mziizm.bsky.social · 03/12/2024Good performance shouldn’t mean 'just in English' anymore 🪩 We provide a robust way to assess models with a new benchmark that captures in-language nuances and cultural contexts. 1182
Sara Hooker @sarahooker.bsky.social · 04/12/2024AI amplifying biorisk has been a major focus in AI policy & governance work. Is the spotlight merited? Our recent cross-institutional work asks: Does the available evidence match the current level of attention? 📜 arxiv.org/abs/2412.01946 25912
Sara Hooker @sarahooker.bsky.social · 04/12/2024I'll be at NeurIPS next week -- looking forward to seeing many of you there! Let the Vancouver cantonese and sushi food tour begin. 2461
Reposted by Sara HookerAngelika Romanou @agromanou.bsky.social · 02/12/2024To build INCLUDE, we collected ~200K MCQ data from 44 languages and 58 knowledge domains, collected from local sources in 52 countries, representing a rich array of cultural and regional knowledge. 161
Reposted by Sara HookerAngelika Romanou @agromanou.bsky.social · 02/12/2024🤔 Why is regional knowledge so important? Users expect #LLMs to know information relevant to their environments— customs, culture, etc. To be relevant & relatable, LLMs need to know these nuances. It's not just global knowledge; it's about meeting user needs where they are. 141