Sign in

Anton

@anton-l.bsky.social
1.4K followers 116 following 16 posts

Feeding LLMs @ Hugging Face

PostsRepliesMedia
Anton @anton-l.bsky.social · 12/02/2025
LLM Reasoning labs will be eating good today🍔 We commandeered the HF cluster for a few days and generated 1.2M reasoning-filled solutions to 500k NuminaMath problems with DeepSeek-R1 🐳 Have fun!
2223
Reposted by Anton
Anton @anton-l.bsky.social · 19/12/2024
Introducing 📐FineMath: the best open math pre-training dataset with 50B+ tokens! Math remains challenging for LLMs and by training on FineMath we see considerable gains over other math datasets, especially on GSM8K and MATH. 🤗 huggingface.co/datasets/Hug... Here’s a breakdown 🧵
A plot showing increased performance of Llama-3.2-3B when pretrained on FineMath
24615
Anton @anton-l.bsky.social · 19/12/2024
Introducing 📐FineMath: the best open math pre-training dataset with 50B+ tokens! Math remains challenging for LLMs and by training on FineMath we see considerable gains over other math datasets, especially on GSM8K and MATH. 🤗 huggingface.co/datasets/Hug... Here’s a breakdown 🧵
A plot showing increased performance of Llama-3.2-3B when pretrained on FineMath
24615
Reposted by Anton
Thomas Wolf @thomwolf.bsky.social · 11/12/2024
The Open LLM Leaderboard got a new front page for Christmas Check it out at huggingface.co/spaces/open-...
26612
Reposted by Anton
Guilherme Penedo @guilherme.hf.co · 08/12/2024
Announcing 🥂 FineWeb2: A sparkling update with 1000s of 🗣️languages. We applied the same data-driven approach that led to SOTA English performance in🍷 FineWeb to thousands of languages. 🥂 FineWeb2 has 8TB of compressed text data and outperforms other datasets.
17619
Reposted by Anton
Andi @andimara.bsky.social · 26/11/2024
Let's go! We are releasing SmolVLM, a smol 2B VLM built for on-device inference that outperforms all models at similar GPU RAM usage and tokens throughputs. SmolVLM can be fine-tuned on a Google collab and be run on a laptop! Or process millions of documents with a consumer GPU!
410422
Reposted by Anton
merve @merve.bsky.social · 26/11/2024
Small yet mighty! 💫 We are releasing SmolVLM: a new 2B small vision language made for on-device use, fine-tunable on consumer GPU, immensely memory efficient 🤠 We release three checkpoints under Apache 2.0: SmolVLM-Instruct, SmolVLM-Synthetic and SmolVLM-Base huggingface.co/collections/...
1115927
Anton @anton-l.bsky.social · 25/11/2024
Check out how easy it is to do LLM evals with LightEval! * any dataset on the 🤗 Hub can become an eval task in a few lines of code: customize the prompt, metrics, parsing, few-shots, everything! * model- and data-parallel inference * auto batching with the new vLLM backend
A screenshot of LightEval benchmarking results in a terminal
27610
Reposted by Anton
Loubna Ben Allal @loubnabnl.hf.co · 24/11/2024
Making SmolLM2 more reproducible: open-sourcing our training & evaluation toolkit 🛠️ github.com/huggingface/... Pre-training & evaluation code, synthetic data generation pipelines, post-training scripts, on-device tools & demos Apache 2.0. V2 data mix coming soon! Which tools should we add next?
github.com
GitHub - huggingface/smollm: Everything about the SmolLM & SmolLM2 family of models
Everything about the SmolLM & SmolLM2 family of models - GitHub - huggingface/smollm: Everything about the SmolLM & SmolLM2 family of models
25910
Reposted by Anton
Gabriel Martín Blázquez @gabrielmb.com · 21/11/2024
Excited to announce the SFT dataset used for @huggingface.bsky.social SmolLM2! The dataset for SmolLM2 was created by combining multiple existing datasets and generating new synthetic datasets, including MagPie Ultra v1.0, using distilabel. Check out the dataset: huggingface.co/datasets/Hug...
huggingface.co
HuggingFaceTB/smoltalk · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1248
Anton @anton-l.bsky.social · 15/11/2024
10x followers in the past week, I guess it's happening!
030