Sign in

Loubna Ben Allal

@loubnabnl.hf.co
1.5K followers 142 following 8 posts

SmolLMs & Data @huggingface Training SmolLMs and curating high quality web and synthetic datasets ✨ loubnabnl.github.io

PostsRepliesMedia
Loubna Ben Allal @loubnabnl.hf.co · 19/12/2024
We built code datasets, English datasets, and now it’s time for math! 🚀 Check out Anton’s thread to learn how we curated the best public math pre-training dataset.
040
Loubna Ben Allal @loubnabnl.hf.co · 12/12/2024
Sharing my slides on "Synthetic data and smol models in 2024" from yesterday's Latent Space event at NeurIPS: docs.google.com/presentation... - Synthetic data is everywhere - Model collapse, is the web polluted? - 3B+ models running on your iPhone - When and why use smol models?
docs.google.com
Synthetic data & Smol models in 2024
Loubna Ben Allal Hugging Face Synthetic data and smol models in 2024 loubnabnl LoubnaBenAllal1
1245
Reposted by Loubna Ben Allal
Tavis Rudd @tavis.damnsimple.com · 11/12/2024
Another great talk at @latentspacepod.bsky.social NeurIPS: @loubnabnl.hf.co on Synthetic Data & Smol Models
042
Reposted by Loubna Ben Allal
Ben Burtenshaw @benburtenshaw.bsky.social · 03/12/2024
For anyone interested in fine-tuning or aligning LLMs, I’m running this free and open course called smol course. It’s not a big deal, it’s just smol. 🧵>>
932363
Reposted by Loubna Ben Allal
Caleb Fahlgren @calebfahlgren.hf.co · 02/12/2024
The amazing, new Qwen2.5-Coder 32B model can now write SQL for any @hf.co dataset ✨
1183
Loubna Ben Allal @loubnabnl.hf.co · 01/12/2024
We hit 1K ⭐ on our SmolLM repo—thank you! 🎉 New updates: • SmolLM2 nanotron checkpoints (with optimizer states) for easier continual pre-training • Local inference demos (MLC, Transformers.js, MLX, llama.cpp) • SmolVLM: Vision-language model built on SmolLM2 github.com/huggingface/...
0191
Loubna Ben Allal @loubnabnl.hf.co · 30/11/2024
📬 Summarize and rewrite your text/emails faster, and offline! Check @andimara.bsky.social's Smol Tools for summarization and rewriting. It uses SmolLM2 to summarize text and make it more friendly or professional, all running locally thanks to llama.cpp github.com/huggingface/...
github.com
smollm/smol_tools at main · huggingface/smollm
Everything about the SmolLM & SmolLM2 family of models - huggingface/smollm
1112
Reposted by Loubna Ben Allal
Xenova @xenova.bsky.social · 27/11/2024
WOW! 🤯 Language models are becoming smaller and more capable than ever! Here's SmolLM2 running 100% locally in-browser w/ WebGPU on a 6-year-old GPU. Just look at that speed! ⚡️😍 Powered by 🤗 Transformers.js and ONNX Runtime Web! How many tokens/second do you get? Let me know! 👇
24610
Reposted by Loubna Ben Allal
Simon Willison @simonwillison.net · 29/11/2024
This demo of structured data extraction running on an LLM that executes entirely in the browser (Chrome only for the moment since it uses WebGPU) is amazing My notes here: simonwillison.net/2024/Nov/29/...
simonwillison.net
Structured Generation w/ SmolLM2 running in browser & WebGPU
Extraordinary demo by Vaibhav Srivastav. Here's Hugging Face's [SmolLM2-1.7B-Instruct](https://huggingface.co/HuggingFaceTB/SmolLM2-1.7B-Instruct) running directly in a web browser (using WebGPU, so r...
518523
Reposted by Loubna Ben Allal
vb @reach-vb.hf.co · 28/11/2024
Fuck it! Structured Generation w/ SmolLM2 running in browser & WebGPU 🔥 Powered by MLC Web-LLM & XGrammar ⚡ Define a JSON schema, Input free text, get structured data right in your browser - profit!!
410613
Reposted by Loubna Ben Allal
Elie @eliebak.hf.co · 27/11/2024
We’re looking for an intern to join our SmolLM team! If you’re excited about training LLMs and building high-quality datasets, we’d love to hear from you. 🤗 US: apply.workable.com/huggingface/... EMEA: apply.workable.com/huggingface/...
apply.workable.com
ML Research Engineer Internship, SmolLMs pretraining and datasets - EMEA Remote - Hugging Face
Here at Hugging Face, we’re on a journey to advance good Machine Learning and make it more accessible. Along the way, we contribute to the development of technology for the better.We have built the fa...
76312
Reposted by Loubna Ben Allal
Andi @andimara.bsky.social · 26/11/2024
Let's go! We are releasing SmolVLM, a smol 2B VLM built for on-device inference that outperforms all models at similar GPU RAM usage and tokens throughputs. SmolVLM can be fine-tuned on a Google collab and be run on a laptop! Or process millions of documents with a consumer GPU!
410422
Reposted by Loubna Ben Allal
Anton @anton-l.bsky.social · 25/11/2024
Check out how easy it is to do LLM evals with LightEval! * any dataset on the 🤗 Hub can become an eval task in a few lines of code: customize the prompt, metrics, parsing, few-shots, everything! * model- and data-parallel inference * auto batching with the new vLLM backend
A screenshot of LightEval benchmarking results in a terminal
27610
Reposted by Loubna Ben Allal
Thomas Wolf @thomwolf.bsky.social · 24/11/2024
It's Sunday morning so taking a minute for a nerdy thread (on math, tokenizers and LLMs) of the work of our intern Garreth By adding a few lines of code to the base Llama 3 tokenizer, he got a free boost in arithmetic performance 😮 [thread]
527134
Loubna Ben Allal @loubnabnl.hf.co · 24/11/2024
Making SmolLM2 more reproducible: open-sourcing our training & evaluation toolkit 🛠️ github.com/huggingface/... Pre-training & evaluation code, synthetic data generation pipelines, post-training scripts, on-device tools & demos Apache 2.0. V2 data mix coming soon! Which tools should we add next?
github.com
GitHub - huggingface/smollm: Everything about the SmolLM & SmolLM2 family of models
Everything about the SmolLM & SmolLM2 family of models - GitHub - huggingface/smollm: Everything about the SmolLM & SmolLM2 family of models
25910
Reposted by Loubna Ben Allal
Gabriel Martín Blázquez @gabrielmb.com · 21/11/2024
Excited to announce the SFT dataset used for @huggingface.bsky.social SmolLM2! The dataset for SmolLM2 was created by combining multiple existing datasets and generating new synthetic datasets, including MagPie Ultra v1.0, using distilabel. Check out the dataset: huggingface.co/datasets/Hug...
huggingface.co
HuggingFaceTB/smoltalk · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1248
Reposted by Loubna Ben Allal
Leandro von Werra @lvwerra.bsky.social · 21/11/2024
What's the secret sauce of SmolLM2 to beat LLM titans like Llama3.2 and Qwen2.5? Unsurprisingly: data, data, data! The SmolTalk is open and available here: huggingface.co/datasets/Hug...
2627