Sign in

Julia Kreutzer

@juliakreutzer.bsky.social
259 followers 190 following 31 posts

NLP & ML research @cohereforai.bsky.social 🇨🇦

PostsRepliesMedia
Reposted by Julia Kreutzer
Cohere Labs @cohereforai.bsky.social · 19/08/2026
Announcing new research on the state of the art of linguistic reasoning. 🧩🧠🏆 📜Read the paper: arxiv.org/abs/2608.18011
1102
Julia Kreutzer @juliakreutzer.bsky.social · 30/06/2026
🤨"But why linguistics" is the most common question when talking about linguistic reasoning benchmarks. Last year we organized a shared task at WMT...and no one participated 🤣 🤯Let me change your mind why this is one of the most challenging, focused and best reasoning benchmarks right now.
1142
Julia Kreutzer @juliakreutzer.bsky.social · 30/06/2026
🔥Possibly the most fun and underrated AI challenge this summer: Compete on unseen linguistic reasoning problems and present your solutions to the expert jury in a month!
011
Julia Kreutzer @juliakreutzer.bsky.social · 12/03/2026
Join us tomorrow for a discussion around culture in AI! 🌱
070
Julia Kreutzer @juliakreutzer.bsky.social · 03/03/2026
💭We need more research that focuses on aspects beyond accuracy, especially in multilingual AI. 👉Help us explore the importance of culture in building and testing AI, with a few minutes of your time. Happy to have a chat as well with anyone who's interested in that space!
163
Julia Kreutzer @juliakreutzer.bsky.social · 18/02/2026
🌱Very proud of our team's latest release 😊 meet Tiny Aya, a massively multilingual model with 3.35B parameters. Tech report here: github.com/Cohere-Labs/...
github.com
1327
Julia Kreutzer @juliakreutzer.bsky.social · 02/12/2025
At #Neurips2025 this week with @cohereforai.bsky.social 🤩 This is what brought me into research: the ✨Network Effect ✨ Let's build the next breakthrough together! #LabLegends
021
Reposted by Julia Kreutzer
Cohere Labs @cohereforai.bsky.social · 29/10/2025
We’re thrilled to announce that some of our research will be presented at @emnlpmeeting.bsky.social next week! 🥳 If you’re attending the conference, don’t miss the chance to explore our work and connect with our team.
131
Reposted by Julia Kreutzer
Cohere Labs @cohereforai.bsky.social · 30/10/2025
How well do LLMs handle multilinguality? 🌍🤖 🔬We brought the rigor from Machine Translation evaluation to multilingual LLM benchmarking and organized the WMT25 Multilingual Instruction Shared Task spanning 30 languages and 5 subtasks.
132
Reposted by Julia Kreutzer
Cohere Labs @cohereforai.bsky.social · 23/10/2025
🌍Most multilingual instruction data starts as English and translation can’t capture cultural nuance or linguistic richness What if we optimized prompts instead of completions? That’s the focus of our most recent work on prompt space optimization for multilingual synthetic data🗣️
111
Reposted by Julia Kreutzer
Antoine Bosselut @abosselut.bsky.social · 03/09/2025
The next generation of open LLMs should be inclusive, compliant, and multilingual by design. That’s why we @icepfl.bsky.social @ethz.ch @cscsch.bsky.social ) built Apertus.
2248
Julia Kreutzer @juliakreutzer.bsky.social · 10/10/2025
Let's do the venue justice. Very excited for today's multilingual workshops at #COLM2025 💙
In Montreal 140 languages are spoken
0101
Julia Kreutzer @juliakreutzer.bsky.social · 10/10/2025
Looking forward to tomorrow's #COLM2025 workshop on multilingual data quality! 🤩
063
Julia Kreutzer @juliakreutzer.bsky.social · 08/10/2025
Ready for our poster today at #COLM2025! 💭This paper has had an interesting journey, come find out and discuss with us! @swetaagrawal.bsky.social @kocmitom.bsky.social Side note: being a parent in research does have its perks, poster transportation solved ✅
A poster attached to a kid's bike seat
0121
Reposted by Julia Kreutzer
Cohere Labs @cohereforai.bsky.social · 30/09/2025
We’re not your average lab. We’re a hybrid research environment dedicated to revolutionizing the ML space. And we’re hiring a Senior Research Scientist to co-create with us. If you believe in research as a shared, global effort — this is your chance.
143
Julia Kreutzer @juliakreutzer.bsky.social · 02/10/2025
💡A collaborative➕diverse team is key. In real life as in the LLM world 💪🦾 Check out our latest work that builds on this insight. 👇
131
Reposted by Julia Kreutzer
Marzieh Fadaee @mziizm.bsky.social · 13/08/2025
Breaking into AI research is harder than ever, and early-career researchers face fewer chances to get started. Entry points matter. We started the Scholars Program 3 years ago to give new researchers a real shot — excited to open applications for year 4✨
163
Reposted by Julia Kreutzer
Cohere Labs @cohereforai.bsky.social · 15/08/2025
While effective for chess♟️, Elo ratings struggle with LLM evaluation due to volatility and transitivity issues. New post in collaboration with AI Singapore explores why Elo falls short for AI leaderboards and how we can do better.
163
Reposted by Julia Kreutzer
Conference on Language Modeling @colmweb.org · 14/07/2025
COLM 2025 is now accepting applications for: Financial Assistance Application -- docs.google.com/forms/d/e/1F... Volunteer Application -- docs.google.com/forms/d/e/1F... Childcare Financial Assistance Application -- docs.google.com/forms/d/e/1F... All due by July 31
docs.google.com
COLM 2025 Financial Assistance Application
Goal of the Financial Assistance Program. We at COLM believe our community should be diverse and inclusive. We recognize that some might be less likely to attend because of financial burden of travel ...
064
Julia Kreutzer @juliakreutzer.bsky.social · 26/06/2025
🍋 Squeezing the most of few samples - check out our LLMonade recipe for few-sample test-time scaling in multitask environments. Turns out that standard methods miss out on gains on non-English languages. We propose more robust alternatives. Very proud of this work that our scholar Ammar led! 🚀
041
Julia Kreutzer @juliakreutzer.bsky.social · 04/06/2025
🚨LLM safety research needs to be at least as multilingual as our models. What's the current stage and how to progress from here? This work led by @yongzx.bsky.social has answers! 👇
042
Julia Kreutzer @juliakreutzer.bsky.social · 28/05/2025
🚧No LLM safety without multilingual safety - what is missing to closing the language gap? And where does this gap actually originate from? Answers 👇
011
Julia Kreutzer @juliakreutzer.bsky.social · 09/05/2025
Multilingual 🤝reasoning 🤝 test-time scaling 🔥🔥🔥 New preprint! @yongzx.bsky.social has all the details 👇
051
Reposted by Julia Kreutzer
Marzieh Fadaee @mziizm.bsky.social · 30/04/2025
1/ Science is only as strong as the benchmarks it relies on. So how fair—and scientifically rigorous—is today’s most widely used evaluation benchmark? We took a deep dive into Chatbot Arena to find out. 🧵
1286
Julia Kreutzer @juliakreutzer.bsky.social · 17/04/2025
🤓MT eyes on multilingual LLM benchmarks 👉 Here's a bunch of simple techniques that we could adopt easily, and in total get a much richer understanding of where we are with multilingual LLMs. 🍬Bonus question: how can we spur research on evaluation of evaluations?
030
Reposted by Julia Kreutzer
Tom Kocmi @kocmitom.bsky.social · 17/04/2025
Tired of messy non-replicable multilingual LLM evaluation? So were we. In our new paper, we experimentally illustrate common eval. issues and present how structured evaluation design, transparent reporting, and meta-evaluation can help us to build stronger models.
071
Julia Kreutzer @juliakreutzer.bsky.social · 17/04/2025
📖New preprint with Eleftheria Briakou @swetaagrawal.bsky.social @mziizm.bsky.social @kocmitom.bsky.social! arxiv.org/abs/2504.11829 🌍It reflects experiences from my personal research journey: coming from MT into multilingual LLM research I missed reliable evaluations and evaluation research…
Screenshot of the paper header with title and author list and affiliations
1121
Reposted by Julia Kreutzer
Cohere Labs @cohereforai.bsky.social · 10/04/2025
🚀 We are excited to introduce Kaleidoscope, the largest culturally-authentic exam benchmark. 📌 Most VLM benchmarks are English-centric or rely on translations—missing linguistic & cultural nuance. Kaleidoscope expands in-language multilingual 🌎 & multimodal 👀 VLMs evaluation
1197
Reposted by Julia Kreutzer
Tom Kocmi @kocmitom.bsky.social · 28/03/2025
☀️ Summer internship at Cohere! Are you excited about multilingual evaluation, human judgment, or meta-eval? Come help us explore how a rigorous eval really looks like while questioning the status quo in LLM evaluation. I’m looking for an intern (EU timezone preferred), are you interested? Ping me!
272
Reposted by Julia Kreutzer
Marzieh Fadaee @mziizm.bsky.social · 27/03/2025
Command🅰️ technical report is out. Information-dense. Detailed. Pretty. Simply A+! 💎: cohere.com/research/pap...
cohere.com
Command A: An Enterprise-Ready Family of Large Language Models
In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command
151
Reposted by Julia Kreutzer
Conference on Language Modeling @colmweb.org · 20/03/2025
A bit of a mess around the conflict of COLM with the ARR (and to lesser degree ICML) reviews release. We feel this is creating a lot of pressure and uncertainty. So, we are pushing our deadlines: Abstracts due March 22 AoE (+48hr) Full papers due March 28 AoE (+24hr) Plz RT 🙏
33631
Julia Kreutzer @juliakreutzer.bsky.social · 12/03/2025
💬The first Q&A starts in a few hours. 🔔Also, a reminder to create your Open review profile if you haven't already. Non-institutional accounts require a verification process that can take time. One week till the abstract deadline!
002
Reposted by Julia Kreutzer
Cohere Labs @cohereforai.bsky.social · 11/03/2025
We’re excited to bring back Expedition Aya 🌍 A 6-week open build challenge to accelerate ML research progress in multilingual, multimodal and efficiency. Join us to expand the world that AI sees.
131
Julia Kreutzer @juliakreutzer.bsky.social · 11/03/2025
✨ Multilingual language modeling meets WMT✨ very exciting opportunity to get WMT-style evaluations for MLLMs: unseen tests, human evaluation, meta-evaluation, and that for multiple languages and tasks. Almost too good to be true! 🤩
020
Reposted by Julia Kreutzer
Conference on Language Modeling @colmweb.org · 09/03/2025
COLM's @juliakreutzer.bsky.social and @abosselut.bsky.social will hold two paper submission Q&A sessions. We run a simple process, but figured this can help authors, especially first-time authors. March 12: dateful.com/eventlink/14... March 13: dateful.com/eventlink/83... Plz RT 🙏
053
Reposted by Julia Kreutzer
Tom Kocmi @kocmitom.bsky.social · 20/02/2025
Guess what? The jubilee 🎉 20th iteration of WMT General MT 🎉 is here, and we want you to participate - as the entry barrier to make an impact is so low! This isn’t just any repeat. We’ve kept what worked, removed what was outdated, and introduced many exciting new twists! Among the key changes are:
1185
Reposted by Julia Kreutzer
Cohere Labs @cohereforai.bsky.social · 04/03/2025
Who will triumph - Aya Vision or its creators? 🤼 We challenged the research team behind Aya Vision to mystery object trivia. 💎🌽🧈 www.youtube.com/watch?v=iQEd...
youtube.com
Aya Vision Challenge
YouTube video by Cohere
141
Reposted by Julia Kreutzer
Marzieh Fadaee @mziizm.bsky.social · 04/03/2025
✨👓 Aya Vision is here 👓✨ A multilingual, multimodal model designed to understand across languages and modalities (text, images, etc) to bridge the language gap and empower global users!
Question: Guess where is this kid coming back from?

Answer: Based on the details provided in the image, it appears that this child is likely returning from music class. The presence of a musical note symbol on his hand, which appears to have been drawn with a pencil or pen, suggests that he was engaged in music-related activities. This symbol is commonly associated with reading and writing music and suggests that the child may have been learning to read music notes, practicing a piece of music, or taking a composition class.
142