Sign in

Costa Huang

@vwxyzjn.bsky.social
549 followers 129 following 49 posts

RL + LLM @ai2.bsky.social; main dev of cleanrl.dev

PostsRepliesMedia
Costa Huang @vwxyzjn.bsky.social · 02/10/2025
Congrats on the launch!
110
Costa Huang @vwxyzjn.bsky.social · 01/05/2025
🥘 Excited to share our latest OLMo 1 B models! Almost summer RL time. We did another two-stage RL: * The first RLVR run uses allenai/RLVR-GSM-MATH-IF-Mixed-Constraints * The final RLVR run uses allenai/RLVR-MATH for targeted MATH improvement Short 🧵
130
Costa Huang @vwxyzjn.bsky.social · 13/03/2025
Introducing OLMo-2-0325-32B-Instruct! It's the spring RL curve time. This time, we used GRPO for RLVR and trained a pretty nice fully open source model!
1121
Costa Huang @vwxyzjn.bsky.social · 12/02/2025
🔥 allenai/Llama-3.1-Tulu-3-8B (trained with PPO) -> allenai/Llama-3.1-Tulu-3.1-8B (trained with GRPO) We are happy to "quietly" release our latest GRPO-trained Tulu 3.1 model, which is considerably better in MATH and GSM8K!
1225
Costa Huang @vwxyzjn.bsky.social · 11/02/2025
🤯 Check out our new iOS OLMoE app that runs the model on-device! We also trained new OLMoE-1B-7B-0125 this time using the Tulu 3 recipe. Very exciting that RLVR improved gsm8k by almost 10 points for OLMoE 🔥 A quick 🧵
151
Costa Huang @vwxyzjn.bsky.social · 31/01/2025
I nerd-snipped myself over the @deepseek.bsky.social GRPO's usage of John Schulman's kl3 estimator. I can now see why: When directly minimizing the KL loss, kl3 just appears much more numerically stable. And the >0 guarantee here is also really nice (kl1 could go negative).
150
Costa Huang @vwxyzjn.bsky.social · 06/01/2025
We released the OLMo 2 report! Ready for some more RL curves? 😏 This time, we applied RLVR iteratively! Our initial RLVR checkpoint on the RLVR dataset mix shows a low GSM8K score, so we did another RLVR on GSM8K only and another on MATH only 😆. And it works! A thread 🧵 1/N
1125
Reposted by Costa Huang
Nathan Lambert @natolambert.bsky.social · 16/12/2024
Maybe my favorite unexpected crossover model from the community. SmolLM from HuggingFace trained with part of the Tülu 3 recipe we release a month ago :D. Cool numerical explorations of post-training stuff. Nothing crazy. SultanR/SmolTulu-1.7b-Instruct buff.ly/41Epauv
0373
Reposted by Costa Huang
Valentina Pyatkin @valentinapy.bsky.social · 01/12/2024
I'll be attending NeurIPS in Vancouver next week! I would love to chat about LLM post-training research, the Faculty job market and anything in between - ping me if you'd like to meet up! You can also find me at the following:
4472
Costa Huang @vwxyzjn.bsky.social · 26/11/2024
So happy OLMo 2 is out! We applied the same Tülu 3 RLVR recipe and it worked very nicely for our final 13B instruct model. Here are the gains/losses of allenai/OLMo-2-1124-13B-Instruct (RLVR's checkpoint) over allenai/OLMo-2-1124-13B-DPO. More to share soon!
1140
Reposted by Costa Huang
Ai2 @ai2.bsky.social · 26/11/2024
Meet OLMo 2, the best fully open language model to date, including a family of 7B and 13B models trained up to 5T tokens. OLMo 2 outperforms other fully open models and competes with open-weight models like Llama 3.1 8B — As always, we released our data, code, recipes and more 🎁
The OLMo 2 models sit at the Pareto frontier of training FLOPs vs model average performance.
515235
Reposted by Costa Huang
vmoens @vmoens.bsky.social · 22/11/2024
One of my fav projects: LeanRL, a simple RL library that provides recipes for fast RL training using torch.compile and cudagraphs. Using these, we got >6x speed-ups compared to the original CleanRL implementations. github.com/pytorch-labs...
2325
Costa Huang @vwxyzjn.bsky.social · 23/11/2024
👀 @araffin.bsky.social says he can’t @ me. What’s going on? @bsky.app
100
Reposted by Costa Huang
Nathan Lambert @natolambert.bsky.social · 21/11/2024
I've spent the last two years scouring all available resources on RLHF specifically and post training broadly. Today, with the help of a totally cracked team, we bring you the fruits of that labor — Tülu 3, an entirely open frontier model post training recipe. We beat Llama 3.1 Instruct. Thread.
821343