Sign in

Kyutai

@kyutai-labs.bsky.social
587 followers 6 following 47 posts

kyutai.org Open-Science AI Research Lab based in Paris

PostsRepliesMedia
Kyutai @kyutai-labs.bsky.social · 02/10/2026
Earlier this year we released MuScriptor, the best open model for multi-instrument transcription to date, created in collaboration with MireloAI Give it a recording in any genre: pop, classical, metal, jazz, whatever, and it transcribes the individual instruments into MIDI. Link in 🧵
140
Reposted by Kyutai
David Picard @davidpicard.eurosky.social · 19/06/2026
I'll need to track how many time I refer to this paper. It's probably going to be my new language filler.
1124
Kyutai @kyutai-labs.bsky.social · 19/06/2026
🎰 Welcome to the FID Lottery. We pulled the lever 25 times on the same machine. Identical diffusion model, identical ImageNet class-cond recipe, only the seed changed. The house paid out anywhere from 33.59 to 35.69 FID. A 2.1-point spread, pure luck. Step onto the floor 👇🧵
14710
Kyutai @kyutai-labs.bsky.social · 16/06/2026
Hypnotizing to watch. Great work, co-authored by our very own @nicolasdufour.bsky.social
091
Reposted by Kyutai
Antoine Guédon @antoine-guedon.bsky.social · 16/06/2026
What if you could turn any number of photos (3, 8, 15, or even 60) into one clean 3D surface (pts & mesh) with Flow Matching? Check out our new work, Surflo: Consistent 3D Surface Flow Model with Global State. 🧵 1/N 🔗https://anttwo.github.io/surflo/
13512
Reposted by Kyutai
KE:SAI - Kyutai ELLIS Scalable Autonomous Intelligence @kesai.eu · 20/05/2026
Today @kyutai-labs.bsky.social and @ellisinsttue.bsky.social launch @kesai.eu! Robot learning is bottlenecked by the cost of physical interaction. Our mission is to advance the efficiency frontier of robust & safe physical AI through fully open and reproducible research. kesai.eu/blog/2026-05...
kesai.eu
Kyutai and ELLIS Tübingen launch KE:SAI: A Premier Franco-German Partnership for Open Science in Physical AI
Kyutai and ELLIS Tübingen announce the official launch of KE:SAI – Kyutai ELLIS Scalable Autonomous Intelligence, a non-profit open science lab for physical AI.
11710
Kyutai @kyutai-labs.bsky.social · 16/04/2026
We're releasing OVIE, a novel view generation model trained entirely on single images. No multi-view datasets needed. Given a single image, it generates novel views of any scene in real time, running orders of magnitude faster than competing approaches.
1152
Kyutai @kyutai-labs.bsky.social · 27/06/2025
Our latest open-source speech-to-text model just claimed 1st place among streaming models and 5th place overall on the OpenASR leaderboard 🥇🎙️ While all other models need the whole audio, ours delivers top-tier accuracy on streaming content. Open, fast, and ready for production!
143
Kyutai @kyutai-labs.bsky.social · 23/05/2025
Talk to unmute.sh 🔊, the most modular voice AI around. Empower any text LLM with voice, instantly, by wrapping it with our new speech-to-text and text-to-speech. Any personality, any voice. Interruptible, smart turn-taking. We’ll open-source everything within the next few weeks.
281
Kyutai @kyutai-labs.bsky.social · 05/05/2025
🚀 Thrilled to announce Helium 1, our new 2B-parameter LLM, now available alongside dactory, an open-source pipeline to reproduce its training dataset covering all 24 EU official languages. Helium sets new standards within its size class on European languages!
130
Kyutai @kyutai-labs.bsky.social · 01/04/2025
Have you enjoyed talking to 🟢Moshi and dreamt of making your own speech to speech chat experience🧑‍🔬🤖? It's now possible with the moshi-finetune codebase! Plug your own dataset and change the voice/tone/personality of Moshi 💚🔌💿. An example after finetuning w/ only 20 hours of the DailyTalk dataset. 🧵
161
Kyutai @kyutai-labs.bsky.social · 21/03/2025
Meet MoshiVis🎙️🖼️, the first open-source real-time speech model that can talk about images! It sees, understands, and talks about images — naturally, and out loud. This opens up new applications, from audio description for the visual impaired to visual access to information.
162
Kyutai @kyutai-labs.bsky.social · 11/02/2025
Even Kavinsky 🎧🪩 can't break Hibiki! Just like Moshi, Hibiki is robust to extreme background conditions 💥🔊.
084
Kyutai @kyutai-labs.bsky.social · 07/02/2025
Meet Hibiki, our simultaneous speech-to-speech translation model, currently supporting 🇫🇷➡️🇬🇧. Hibiki produces spoken and text translations of the input speech in real-time, while preserving the speaker’s voice and optimally adapting its pace based on the semantic content of the source speech. 🧵
1112
Kyutai @kyutai-labs.bsky.social · 14/01/2025
Helium 2B running locally on an iPhone 16 Pro at ~28 tok/s, faster than you can read your loga lessons in French 🚀 All that thanks to mlx-swift with q4 quantization!
011
Kyutai @kyutai-labs.bsky.social · 13/01/2025
Meet Helium-1 preview, our 2B multi-lingual LLM, targeting edge and mobile devices, released under a CC-BY license. Start building with it today! huggingface.co/kyutai/heliu...
huggingface.co
kyutai/helium-1-preview-2b · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1165