Sign in

Lewis Tunstall

@lewtun.bsky.social
593 followers 2 following 13 posts

🤗 LLM whisperer @huggingface 📖 Co-author of "NLP with Transformers" book 💥 Ex-particle physicist 🤘 Occasional guitarist 🇦🇺 in 🇨🇭

PostsRepliesMedia
Lewis Tunstall @lewtun.bsky.social · 21/05/2026
We are releasing Carbon: a crazy fast DNA model Carbon is 275x faster than the next best model. So fast you can process the whole human genome on a single GPU in <2 days. We built a demo so you can explore how the model can generate DNA sequences and a lot more: huggingface.co/spaces/Huggi...
110
Lewis Tunstall @lewtun.bsky.social · 10/02/2025
Introducing OpenR1-Math-220k! huggingface.co/datasets/ope... The community has been busy distilling DeepSeek-R1 from inference providers, but we decided to have a go at doing it ourselves from scratch 💪 More details in 🧵
huggingface.co
open-r1/OpenR1-Math-220k · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1163
Lewis Tunstall @lewtun.bsky.social · 25/01/2025
We are reproducing the full DeepSeek R1 data and training pipeline so everybody can use their recipe. Instead of doing it in secret we can do it together in the open! Follow along: github.com/huggingface/...
github.com
GitHub - huggingface/open-r1: Fully open reproduction of DeepSeek-R1
Fully open reproduction of DeepSeek-R1. Contribute to huggingface/open-r1 development by creating an account on GitHub.
620036
Lewis Tunstall @lewtun.bsky.social · 16/12/2024
We outperform Llama 70B with Llama 3B on hard math by scaling test-time compute 🔥 How? By combining step-wise reward models with tree search algorithms :) We're open sourcing the full recipe and sharing a detailed blog post 👇
410921
Lewis Tunstall @lewtun.bsky.social · 12/12/2024
Hey ML peeps, we found a nice extension to beam search at Hugging Face that is far more scalable and produces more diverse candidates The basic idea is to split your N beams into N/M subtrees and then run greedy node selection in parallel Does anyone know what this algorithm is called?
060