Sign in

Germán Kruszewski

@germank.bsky.social
109 followers 5 following 12 posts

Senior Scientist @Naver Labs Europe. MSCA Postdoctoral Research @UPF (COLT). Just trying out bsky for now...

PostsRepliesMedia
Germán Kruszewski @germank.bsky.social · 18/03/2026
Or dig directly into our paper: openreview.net/forum?id=pPW...
openreview.net
Whatever Remains Must Be True: Filtering Drives Reasoning in LLMs...
Reinforcement Learning (RL) has become the _de facto_ standard for tuning LLMs to solve tasks involving reasoning. However, growing evidence shows that models trained in such way often suffer from...
000
Germán Kruszewski @germank.bsky.social · 18/03/2026
For more details, check out our blog post: huggingface.co/blog/germank...
huggingface.co
Distribution Matching Prevents Mode Collapse in Training Reasoning Models
A Blog post by Germán Kruszewski on Hugging Face
111
Germán Kruszewski @germank.bsky.social · 18/03/2026
We’re excited about α-DPG as a tool for training next-gen reasoning models. One broader takeaway: split “what do you want?” from “how do you get there?” That separation makes design choices clearer—and failures easier to diagnose.
110
Germán Kruszewski @germank.bsky.social · 18/03/2026
Low-α α-DPG (and GRPO with strong KL regularization) shows a more conservative pattern: fewer medium problems converted to easy, but almost no hard problems lost.
110
Germán Kruszewski @germank.bsky.social · 18/03/2026
Categorizing problems by difficulty (in terms of the proportion of samples that are correct), we see that after training with GRPO or high-α α-DPG, many medium problems become easy — but a notable number of previously hard problems become unsolvable.
110
Germán Kruszewski @germank.bsky.social · 18/03/2026
On the LEAN theorem-proving benchmark: low α → best coverage of any method we tested. High α → matches state-of-the-art RL precision, usually with better coverage too.
Our trained models form a Pareto frontier between coverage and precision
120
Germán Kruszewski @germank.bsky.social · 18/03/2026
Once you fix the target explicitly, you can pick a better divergence. We use α-divergences, which smoothly interpolate between mode-seeking and mass-covering behavior. One parameter, α, moves you along that spectrum.
110
Germán Kruszewski @germank.bsky.social · 18/03/2026
The diversity loss comes from the "how". Standard RL minimizes a mode-seeking divergence (Reverse KL), which makes the model pile mass onto a few high-probability solutions and ignore the rest.
110
Germán Kruszewski @germank.bsky.social · 18/03/2026
The "what" has a clean answer: filter the base model's outputs through a verifier. Keep correct solutions, remove wrong ones, preserve their relative probabilities. Simple, well-motivated, and diverse by construction.
110
Germán Kruszewski @germank.bsky.social · 18/03/2026
We think the problem is that RL bundles two decisions that should be kept separate: → WHAT distribution do you want the model to represent? → HOW do you train it to get there?
110
Germán Kruszewski @germank.bsky.social · 18/03/2026
After RL training, models get more accurate but more repetitive. They converge on the same solutions over and over, quietly forgetting rare-but-valid paths they used to occasionally find.
110
Germán Kruszewski @germank.bsky.social · 18/03/2026
Inference-time scaling is great, but it only works if your model actually explores. RL training — by default — kills that exploration. If you are curious to learn why and what we did about it in about our recent ICLR 2026 paper, check out this 🧵!
Our trained models form a Pareto frontier between coverage and precision
141
Reposted by Germán Kruszewski
Thibaut Thonet @tthonet.bsky.social · 31/01/2025
🚨 Excited about ML/NLP and looking for a research internship on controlled text generation? Come work with us on advanced constraint processing in large language models at NAVER LABS Europe! ✨ Learn more and apply here: europe.naverlabs.com/job/2-7/
europe.naverlabs.com
Internship: Advanced Constraint Processing in LLMs
The ability to control the outputs of Large Language Models is a central topic in NLP and Machine Learning, with applications to safety, trustworthiness, reasoning, etc. Following […]
084