Sign in

Marzieh Fadaee

@mziizm.bsky.social
1.7K followers 169 following 76 posts

seeks to understand language. Head of Cohere Labs @Cohere_Labs @Cohere PhD from @UvA_Amsterdam marziehf.github.io

PostsRepliesMedia
Reposted by Marzieh Fadaee
Cohere Labs @cohereforai.bsky.social · 17/02/2026
Introducing ✨Tiny Aya✨, a family of massively multilingual small language models built to run where people actually are. Tiny Aya delivers strong multilingual performance in 70+ global languages in a 3.35B parameter model, efficient enough to run locally, even on a phone.
29615
Reposted by Marzieh Fadaee
Julia Kreutzer @juliakreutzer.bsky.social · 18/02/2026
🌱Very proud of our team's latest release 😊 meet Tiny Aya, a massively multilingual model with 3.35B parameters. Tech report here: github.com/Cohere-Labs/...
github.com
1327
Reposted by Marzieh Fadaee
Cohere Labs @cohereforai.bsky.social · 29/09/2025
What if the way we verify synthetic code is limiting model performance? In our latest work we uncover the Verification Ceiling Problem: strict “all tests must pass” rules throw away useful data, while weak tests let errors through.
131
Reposted by Marzieh Fadaee
Cohere Labs @cohereforai.bsky.social · 30/09/2025
We’re not your average lab. We’re a hybrid research environment dedicated to revolutionizing the ML space. And we’re hiring a Senior Research Scientist to co-create with us. If you believe in research as a shared, global effort — this is your chance.
143
Marzieh Fadaee @mziizm.bsky.social · 05/09/2025
I'm excited to share that I'll be stepping into the role of Head of @cohereforai.bsky.social. It's an honor and a responsibility to lead such an extraordinary group of researchers pushing the boundaries of AI research.
1132
Reposted by Marzieh Fadaee
Cohere Labs @cohereforai.bsky.social · 15/08/2025
While effective for chess♟️, Elo ratings struggle with LLM evaluation due to volatility and transitivity issues. New post in collaboration with AI Singapore explores why Elo falls short for AI leaderboards and how we can do better.
163
Marzieh Fadaee @mziizm.bsky.social · 13/08/2025
Breaking into AI research is harder than ever, and early-career researchers face fewer chances to get started. Entry points matter. We started the Scholars Program 3 years ago to give new researchers a real shot — excited to open applications for year 4✨
163
Reposted by Marzieh Fadaee
Coalition on Digital Impact (CODI) @codi.global · 07/08/2025
🌍 Language shapes how we think and connect—but most AI models still struggle beyond English. @microsoft.com's July seminar discussed how we can bridge the gap and build #AIforEveryone with @mziizm.bsky.social of @cohere.com. 📽️ www.microsoft.com/en-us/resear...
microsoft.com
Building Better Language Models Through Global Understanding - Microsoft Research
Modern language models have achieved remarkable capabilities in English, but human knowledge and experience span thousands of languages, each encoding unique perspectives and problem-solving approache...
011
Marzieh Fadaee @mziizm.bsky.social · 29/07/2025
ACL day 2 ✨
010
Marzieh Fadaee @mziizm.bsky.social · 09/07/2025
🖼️ Most text-to-image models only really work in English. This limits who can use them and whose imagination they reflect. We asked: can we build a small, efficient model that understands prompts in multiple languages natively?
122
Marzieh Fadaee @mziizm.bsky.social · 29/06/2025
Everyone talks about GEB (I agree, it's a gem) but Hofstadter's Analogy book is criminally underrated. If you're working on learning intelligence through language understanding, it’s a must-read.
040
Reposted by Marzieh Fadaee
Julia Kreutzer @juliakreutzer.bsky.social · 26/06/2025
🍋 Squeezing the most of few samples - check out our LLMonade recipe for few-sample test-time scaling in multitask environments. Turns out that standard methods miss out on gains on non-English languages. We propose more robust alternatives. Very proud of this work that our scholar Ammar led! 🚀
041
Marzieh Fadaee @mziizm.bsky.social · 09/06/2025
London has me under its spell. every. single. visit.
130
Reposted by Marzieh Fadaee
Julia Kreutzer @juliakreutzer.bsky.social · 04/06/2025
🚨LLM safety research needs to be at least as multilingual as our models. What's the current stage and how to progress from here? This work led by @yongzx.bsky.social has answers! 👇
042
Reposted by Marzieh Fadaee
Cohere Labs @cohereforai.bsky.social · 28/05/2025
Over 7000 languages are spoken worldwide 🌐, but AI safety efforts focus on only a fraction of them. Our latest paper draws on our multi-year efforts with the wider research community to explore why this matters and how we can bridge the AI language gap.
173
Reposted by Marzieh Fadaee
Anna Rogers @annarogers.bsky.social · 26/05/2025
📢 The Copenhagen NLP Symposium on June 20th! - Invited talks by @loubnabnl.hf.co (HF) @mziizm.bsky.social (Cohere) @najoung.bsky.social (BU) @kylelo.bsky.social (AI2) Yohei Oseki (UTokyo) - Exciting posters by other participants Register to attend and/or present your poster at cphnlp.github.io /1
cphnlp.github.io
Copenhagen NLP Symposium 2025
symposium website
13412
Reposted by Marzieh Fadaee
Sara Hooker @sarahooker.bsky.social · 30/04/2025
It is critical for scientific integrity that we trust our measure of progress. The @lmarena.bsky.social has become the go-to evaluation for AI progress. Our release today demonstrates the difficulty in maintaining fair evaluations on the Arena, despite best intentions.
19569
Reposted by Marzieh Fadaee
David Pfau @davidpfau.com · 30/04/2025
Goodhart's law rules everything around me.
0184
Marzieh Fadaee @mziizm.bsky.social · 30/04/2025
1/ Science is only as strong as the benchmarks it relies on. So how fair—and scientifically rigorous—is today’s most widely used evaluation benchmark? We took a deep dive into Chatbot Arena to find out. 🧵
1286
Marzieh Fadaee @mziizm.bsky.social · 22/04/2025
Not in Singapore for #ICLR2025 but our lab’s work is! In particular, I am very proud of these collaborations: ✨INCLUDE (spotlight) — models fail to grasp regional nuances across languages 💎To Code or Not to Code (poster) — code is key for generalizing beyond coding tasks
130
Marzieh Fadaee @mziizm.bsky.social · 17/04/2025
🚨 Excited to share our latest paper! Multilingual LLMs are getting really good. But the way we evaluate them? Not the best sometimes. 🌟 We show how decades of lessons from Machine Translation can help us fix it
1110
Marzieh Fadaee @mziizm.bsky.social · 10/04/2025
Very excited to release Kaleidoscope—a multilingual, multimodal evaluation set for VLMs, built as part of our open-science initiative! 🌍 18 languages (high-, mid-, low-) 📚 21k questions (55% require image understanding) 🧪 STEM, social science, reasoning, and practical skills
1104
Reposted by Marzieh Fadaee
Tom Kocmi @kocmitom.bsky.social · 11/03/2025
Big news from WMT! 🎉 We are expanding beyond MT and launching a new multilingual instruction shared task. Our goal is to foster truly multilingual LLM evaluation and best practices in automatic and human evaluation. Join us and build the winning multilingual system! www2.statmt.org/wmt25/multil...
www2.statmt.org
Multilingual Instruction Shared Task
1127
Reposted by Marzieh Fadaee
Tom Kocmi @kocmitom.bsky.social · 28/03/2025
☀️ Summer internship at Cohere! Are you excited about multilingual evaluation, human judgment, or meta-eval? Come help us explore how a rigorous eval really looks like while questioning the status quo in LLM evaluation. I’m looking for an intern (EU timezone preferred), are you interested? Ping me!
272
Marzieh Fadaee @mziizm.bsky.social · 27/03/2025
Command🅰️ technical report is out. Information-dense. Detailed. Pretty. Simply A+! 💎: cohere.com/research/pap...
cohere.com
Command A: An Enterprise-Ready Family of Large Language Models
In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command
151
Marzieh Fadaee @mziizm.bsky.social · 12/03/2025
Good morning Paris
150
Marzieh Fadaee @mziizm.bsky.social · 04/03/2025
✨👓 Aya Vision is here 👓✨ A multilingual, multimodal model designed to understand across languages and modalities (text, images, etc) to bridge the language gap and empower global users!
Question: Guess where is this kid coming back from?

Answer: Based on the details provided in the image, it appears that this child is likely returning from music class. The presence of a musical note symbol on his hand, which appears to have been drawn with a pencil or pen, suggests that he was engaged in music-related activities. This symbol is commonly associated with reading and writing music and suggests that the child may have been learning to read music notes, practicing a piece of music, or taking a composition class.
142
Marzieh Fadaee @mziizm.bsky.social · 25/02/2025
Excited to share insights from our new paper on evaluating LLMs in multi-session coding interactions! 📚📚📚 We introduce MEMORYCODE, a novel dataset to assess how well LLMs track & execute coding instructions across multiple sessions, mimicking real-world collaboration.
230
Marzieh Fadaee @mziizm.bsky.social · 23/01/2025
INCLUDE is going to ICLR! Huge congrats to all my co-authors!
010
Marzieh Fadaee @mziizm.bsky.social · 16/12/2024
This #Neurips2024 was the perfect way to end this year. So long Vancouver!
040
Marzieh Fadaee @mziizm.bsky.social · 11/12/2024
Day 2 #neurips2024, let the fun officially begin. On a separate note, as much as I love Amsterdam I'm mountain-deprived and only have eyes for this glorious view this week.
020
Marzieh Fadaee @mziizm.bsky.social · 10/12/2024
Good morning Vancouver! It's a lovely day to see old friends and make new ones.
020
Marzieh Fadaee @mziizm.bsky.social · 08/12/2024
Aya Expanse technical report is here and I couldn’t be prouder! 🎉 Check out the report: arxiv.org/abs/2412.04261
arxiv.org
Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier
We introduce the Aya Expanse model family, a new generation of 8B and 32B parameter multilingual language models, aiming to address the critical challenge of developing highly performant multilingual ...
043
Marzieh Fadaee @mziizm.bsky.social · 06/12/2024
🚀 Our mission to strengthen the multilingual open-source ecosystem continues!👇
192
Reposted by Marzieh Fadaee
Aidan P @aidanpeppin.bsky.social · 04/12/2024
AI amplifying biorisk has been a major topic in policy & governance work. But does the available evidence match this level of attention? 🦠 ⚠️ Our new paper looks at the science underpinning ideas that AI could increase biorisks. arxiv.org/abs/2412.01946
arxiv.org
The Reality of AI and Biorisk
To accurately and confidently answer the question 'could an AI model or system increase biorisk', it is necessary to have both a sound theoretical threat model for how AI models or systems could incre...
1162
Marzieh Fadaee @mziizm.bsky.social · 04/12/2024
see you in a few hours 🤩
040
Marzieh Fadaee @mziizm.bsky.social · 03/12/2024
Good performance shouldn’t mean 'just in English' anymore 🪩 We provide a robust way to assess models with a new benchmark that captures in-language nuances and cultural contexts.
1182
Marzieh Fadaee @mziizm.bsky.social · 26/11/2024
I'm always thinking about the aspects of research that aren't discussed enough online, and one key area is effective scientific communication. Very excited about this panel discussion featuring two amazing researchers: @jayalammar.bsky.social and @shaynelongpre.bsky.social
1100
Marzieh Fadaee @mziizm.bsky.social · 20/11/2024
Laura is amazing for seeing this project through! Check out her paper on studying the reasoning capabilities of LLMs and influences from the training data 🔮
040
Marzieh Fadaee @mziizm.bsky.social · 18/11/2024
does anyone else mutter "banksy" in their head when they type "bsky" instead of "bluesky"?
220
Marzieh Fadaee @mziizm.bsky.social · 18/11/2024
It's a beautiful day to immerse in overleaf drafts
070
Marzieh Fadaee @mziizm.bsky.social · 13/11/2024
hello world! #NLProc
020