Sign in

Michael Noukhovitch

@mnoukhov.bsky.social
304 followers 210 following 38 posts

PhD in AI @mila-quebec.bsky.social RLHF and language grounding, whatever that means. Whitespace aficianado. mnoukhov.github.io

PostsRepliesMedia
Michael Noukhovitch @mnoukhov.bsky.social · 15/09/2026
Is RL actually making your LLM better? Gains from RL are mostly on easy questions🤯 We're calling this the Matthew Effect for RL on LLMs. We then leverage async RL to solve harder problems by Never Giving Up! arxiv.org/abs/2609.13443 and mnoukhov.github.io/posts/ngu/ and check out thread below 🧵👇
1267
Michael Noukhovitch @mnoukhov.bsky.social · 12/12/2025
Olmo 3.1: even more RL = even more RL-Zero! @saurabhshah2.bsky.social and I tweaked some hyperparams and prompts, @hamishivi.bsky.social and @finbarr.bsky.social improved the code and boom! New Olmo 3.1 RL-Zero 👾 An updated, solid baseline for your RL and reasoning research
162
Michael Noukhovitch @mnoukhov.bsky.social · 20/11/2025
Check out Olmo 3 RL-Zero: a clean and scientific setup to benchmark RLVR Everyone is finetuning with Qwen but its hard to know whether your eval is contaminated and skewing your RLVR results. Olmo 3 has a solution.
120
Reposted by Michael Noukhovitch
Nathan Lambert @natolambert.bsky.social · 20/11/2025
We present Olmo 3, our next family of fully open, leading language models. This family of 7B and 32B models represents: 1. The best 32B base model. 2. The best 7B Western thinking & instruct models. 3. The first 32B (or larger) fully open reasoning model.
310525
Reposted by Michael Noukhovitch
Dane Carnegie Malenfant @dvnxmvlhdf5.bsky.social · 05/06/2025
Preprint Alert 🚀 Multi-agent reinforcement learning (MARL) often assumes that agents know when other agents cooperate with them. But for humans, this isn’t always the case. For example, plains indigenous groups used to leave resources for others to use at effigies called Manitokan. 1/8
Manitokan are images set up where one can bring a gift or receive a gift. 1930s Rocky Boy Reservation, Montana, Montana State University photograph. Colourized with AI
13513
Michael Noukhovitch @mnoukhov.bsky.social · 24/04/2025
@dnllvy.bsky.social @oumarkaba.bsky.social presenting cool work at #ICLR2025 on generative models for crystals leveraging symmetry ❄️🪞, repping @mila-quebec.bsky.social
051
Reposted by Michael Noukhovitch
Sara Vera Marjanovic @saravera.bsky.social · 01/04/2025
Models like DeepSeek-R1 🐋 mark a fundamental shift in how LLMs approach complex problems. In our preprint on R1 Thoughtology, we study R1’s reasoning chains across a variety of tasks; investigating its capabilities, limitations, and behaviour. 🔗: mcgill-nlp.github.io/thoughtology/
A circular diagram with a blue whale icon at the center. The diagram shows 8 interconnected research areas around LLM reasoning represented as colored rectangular boxes arranged in a circular pattern. The areas include: §3 Analysis of Reasoning Chains (central cloud), §4 Scaling of Thoughts (discussing thought length and performance metrics), §5 Long Context Evaluation (focusing on information recall), §6 Faithfulness to Context (examining question answering accuracy), §7 Safety Evaluation (assessing harmful content generation and jailbreak resistance), §8 Language & Culture (exploring moral reasoning and language effects), §9 Relation to Human Processing (comparing cognitive processes), §10 Visual Reasoning (covering ASCII generation capabilities), and §11 Following Token Budget (investigating direct prompting techniques). Arrows connect the sections in a clockwise flow, suggesting an iterative research methodology.
15116
Michael Noukhovitch @mnoukhov.bsky.social · 07/04/2025
Llama 4 uses async RLHF and I would just like to announce that I called it t.co/w9qJxr944C
160
Michael Noukhovitch @mnoukhov.bsky.social · 18/03/2025
Our work on Asynchronous RLHF was accepted to #ICLR2025 ! (I was so excited to announce it, I forgot to say I was excited) Used by @ai2.bsky.social for OLMo-2 32B 🔥 New results show ~70% speedups for LLM + RL math and reasoning 🧠 🧵below or hear my DLCT talk online on March 28!
1133
Michael Noukhovitch @mnoukhov.bsky.social · 11/02/2025
Programming using an AI assistant in order to improve AI assistants is giving me strong sci-fi vibes. Specifically Isaac Asimov, who clearly invented vibe coding in 1956 users.ece.cmu.edu/~gamvrosi/th...
020
Michael Noukhovitch @mnoukhov.bsky.social · 11/12/2024
I'm at #NeurIPS2024 this week if anyone wants to talk about RLHF while drinking an overpriced (but excellent) pourover coffee or tea!
150