Reposted by Guilherme Penedogarreth @garrethlee.bsky.social · 16/12/2024🚀 With Meta's recent paper replacing tokenization in LLMs with patches 🩹, I figured that it's a great time to revisit how tokenization has evolved over the years using everyone's favourite medium - memes! Let's take a trip down memory lane! [1/N] 4339
Guilherme Penedo @guilherme.hf.co · 08/12/2024We will very soon announce a big community project, and are working on a 📝 blogpost walking you through the entire dataset creation process. Stay tuned! 160
Guilherme Penedo @guilherme.hf.co · 08/12/2024The dataset is released under the permissive 📜 ODC-By 1.0 license, and the 💻 code to reproduce it and our evaluations is public. Find out all about 🥂 FineWeb2 on the 🤗 model page: huggingface.co/datasets/Hug...huggingface.coHuggingFaceFW/fineweb-2 · Datasets at Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 140
Guilherme Penedo @guilherme.hf.co · 08/12/2024Announcing 🥂 FineWeb2: A sparkling update with 1000s of 🗣️languages. We applied the same data-driven approach that led to SOTA English performance in🍷 FineWeb to thousands of languages. 🥂 FineWeb2 has 8TB of compressed text data and outperforms other datasets. 17619