Sara Han @sdiazlor.hf.co · 20/01/2025Start synthesizing 🚀: huggingface.co/spaces/argil... ✍ Blog post: huggingface.co/blog/sdiazlo...huggingface.coSynthetic Data Generator - a Hugging Face Space by argillaBuild datasets using natural language 030
Sara Han @sdiazlor.hf.co · 20/01/2025💫 Generate RAG data with the Synthetic Data Generator to improve your RAG system! 1️⃣ Generate from your documents, dataset, or dataset description. 2️⃣ Configure it. 3️⃣ Generate the synthetic dataset. 4️⃣ Fine-tune the retrieval and reranking models. 5️⃣ Build a RAG pipeline. 1143
Reposted by Sara HanJosé Francisco Calvo @jfcalvo.hf.co · 19/12/2024🚀 Argilla v2.6.0 is here! 🎉 Let me show you how EASY it is to export your annotated datasets from Argilla to the Hugging Face Hub. 🤩 Take a look to this quick demo 👇 💁♂️ More info about the release at github.com/argilla-io/a... #AI #MachineLearning #OpenSource #DataScience #HuggingFace #Argilla 0115
Sara Han @sdiazlor.hf.co · 18/12/2024🙅♀️ No-code end-to-end example to train your model 1️⃣ Use the Synthetic Data Generator to create your custom dataset 2️⃣ Use AutoTrain to use the generated dataset and train your model Check it here: huggingface.co/blog/synthet... 0113
Sara Han @sdiazlor.hf.co · 16/12/2024- No code required—everything can be handled through the interface. - 100% free to use. - Designed to create text classification and chat datasets. - Review in Argilla and push to the Hub. 000
Sara Han @sdiazlor.hf.co · 16/12/2024Where do I get quality data from? We often need to fine-tune models for very specific scenarios. And that’s where the Synthetic Data Generator comes in! Want to see how it works? Watch this quick video (www.youtube.com/watch?v=nXjV...) and get started here: t.co/hJ1b2TsMq0youtube.comSynthetic Data Generator - Build Datasets Using Natural LanguageYouTube video by Argilla 130
Sara Han @sdiazlor.hf.co · 12/12/2024Pouco a pouco avanzamos! 🚀 Anímovos a contribuir, tan só tedes que entrar na ligazón, ler as instrucións e comezar a anotar ✍ data-is-better-together-fineweb-c.hf.space/share-your-p...data-is-better-together-fineweb-c.hf.spaceglg - galego - GalicianJoin and contribute to the dataset glg - galego - Galician 010
Sara Han @sdiazlor.hf.co · 10/12/2024It only takes 2 steps: - Coordinate with your Language Lead: huggingface.co/spaces/Huggi.... Or become one if it is missing: huggingface.co/spaces/natal... - Read the guidelines and start annotating according to the educational value: huggingface.co/spaces/data-...huggingface.coDiscussion Forum - a Hugging Face Space by HuggingFaceFWDiscover amazing ML apps made by the community 110
Sara Han @sdiazlor.hf.co · 10/12/2024Spanish, Filipino, Amharic, French, German, Basque, Catalan, Galician, Guarani, Telugu, Italian, Pashto, Romanian, Tamil, Urdu, Danish... and many more! All included in the FineWeb2 Community Annotation Sprint! 🔥 💫 Join to build an impactful dataset for your language! 1132
Sara Han @sdiazlor.hf.co · 09/12/2024 Binarized dataset: huggingface.co/datasets/dat... Blog post: huggingface.co/blog/image-p...huggingface.codata-is-better-together/open-image-preferences-v1-binarized · Datasets at Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 000
Sara Han @sdiazlor.hf.co · 09/12/2024Open Image Preferences released! 🚀 - Open-source dataset for text2image - 10K samples manually evaluated by the HF community. - Binarized format for SFT, DPO, or ORPO. It comes with a nice blog post explaining the steps to pre-process and generate the data, along with the results. 151
Sara Han @sdiazlor.hf.co · 07/12/2024I'd say that more small models and focus on agents and on-device 010
Sara Han @sdiazlor.hf.co · 05/12/2024huggingface.co/spaces/huggi...huggingface.coOpen Source Ai Year In Review 2024 - a Hugging Face Space by huggingfaceWhat happened in open-source AI this year, and what’s next? 000
Sara Han @sdiazlor.hf.co · 03/12/2024Language is power! A multilingual annotation sprint for hundreds of languages is starting soon! Step up as a Language Lead and help drive this effort for your language. If there's already a Language Lead, stay tuned! Is this the start of a nice community? docs.google.com/forms/d/e/1F...docs.google.comLanguage Lead sign-upAt Hugging Face 🤗, we're launching a big community initiative to improve LLM training for many languages. We're looking for Language Leads to help us cultivate specific languages during this initiativ... 020
Sara Han @sdiazlor.hf.co · 29/11/2024Docs: docs.zenml.io/stack-compon...docs.zenml.ioArgilla | ZenML - Bridging the gap between ML & OpsAnnotating data using Argilla. 000
Sara Han @sdiazlor.hf.co · 29/11/2024Want to improve your model quality? Implement the data annotation stage in your MLOps effortlessly thanks to the enhanced integration of Argilla with ZenML. ✨Use the latest Argilla features ✨Improve human-in-the-loop workflows ✨Manage datasets, track progress, and coordinate your annotation team 170
Sara Han @sdiazlor.hf.co · 28/11/2024Model: huggingface.co/Qwen/QwQ-32B... Demo: huggingface.co/spaces/Qwen/...huggingface.coQwen/QwQ-32B-Preview · Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 020
Sara Han @sdiazlor.hf.co · 28/11/2024🚀 QwQ-32B-Preview is available on the Hub! > The results are very promising, beating o1-mini. > However, they also have several limitations you might notice even in the demo (I found endless reasoning trying to find out the number of 'r' in 🍓). So, let's see how they deal with them. 161
Sara Han @sdiazlor.hf.co · 28/11/2024I do. Big AI companies stealing our data have put us on guard, but good intentions also exist. So, let's learn together from this and find ways to continue building with consent and transparency for everyone, not just those in power. 100
Reposted by Sara Hanmerve @merve.bsky.social · 27/11/2024It's pretty sad to see the negative sentiment towards Hugging Face on this platform due to a dataset put by one of the employees. I want to write a small piece. 🧵 Hugging Face empowers everyone to use AI to create value and is against monopolization of AI it's a hosting platform above all. 2945570
Reposted by Sara HanXuan Son Nguyen @ngxson.hf.co · 27/11/2024Hugging Face inference endpoints now support CPU deployment for llama.cpp 🚀 🚀 Why this is a huge deal? Llama.cpp is well-known for running very well on CPU. If you're running small models like Llama 1B or embedding models, this will definitely save tons of money 💰 💰 3266
Sara Han @sdiazlor.hf.co · 26/11/2024Read the blog post: huggingface.co/blog/burtens...huggingface.coLet’s make a generation of amazing image generation modelsA Blog post by ben burtenshaw on Hugging Face 000
Sara Han @sdiazlor.hf.co · 26/11/2024Steps: 1️⃣ Log in to the Argilla Space with your HF account: huggingface.co/spaces/data-... 2️⃣ Check the guidelines. 3️⃣ Time to start annotating! Can you climb to the top of the leaderboard? huggingface.co/spaces/data-...huggingface.coImage Preferences - Argilla annotation space - a Hugging Face Space by data-is-better-togetherA community project to create an image preferences dataset. 100
Sara Han @sdiazlor.hf.co · 26/11/2024🎨 Help to build an image preference dataset! > Goal: Release an open-source image dataset, enabling the entire community to benefit from it. > Requirements: All you need is a Hugging Face account and a willingness to contribute. More in 🧵 140
Reposted by Sara HanDaniel Vila @dvilasuero.hf.co · 26/11/2024Let's make AI more inclusive. At @huggingface.bsky.social we'll launch a huge community sprint soon to build high-quality training datasets for many languages. We're looking for Language Leads to help with outreach. Find your language and nominate yourself: forms.gle/iAJVauUQ3FN8... 85321
Reposted by Sara HanDr Sasha Luccioni @sashamtl.bsky.social · 25/11/2024"Naftali was assigned to train AI to recognize and weed out pornography, hate speech and excessive violence, which meant sifting through the worst of the worst content online for hours on end." So much of AI is based on exploiting workers in precarious conditions 😔 www.cbsnews.com/news/labeler...cbsnews.comLabelers training AI say they're overworked, underpaid and exploited by big American tech companiesDigital workers in Kenya had to sift through horrific online content to train AI, but say they were underpaid, overworked, and got inadequate mental health support. So they're fighting back. 34522
Sara Han @sdiazlor.hf.co · 25/11/2024Reach out to us at: - GitHub: github.com/argilla-io/a... - Discord ( #argilla-distilabel-general or #argilla-distilabel-help): hf.co/join/discordgithub.comGitHub - argilla-io/argilla: Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasetsArgilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets - argilla-io/argilla 010
Sara Han @sdiazlor.hf.co · 25/11/2024Argilla has reached the 4K stars on GitHub! ✨ Thanks to everyone for the support! We'll continue shipping new updates for you to curate your data easily ✍️ And, for sure, your feedback is more than welcome 🙌 1200
Reposted by Sara HanLoubna Ben Allal @loubnabnl.hf.co · 24/11/2024Making SmolLM2 more reproducible: open-sourcing our training & evaluation toolkit 🛠️ github.com/huggingface/... Pre-training & evaluation code, synthetic data generation pipelines, post-training scripts, on-device tools & demos Apache 2.0. V2 data mix coming soon! Which tools should we add next?github.comGitHub - huggingface/smollm: Everything about the SmolLM & SmolLM2 family of modelsEverything about the SmolLM & SmolLM2 family of models - GitHub - huggingface/smollm: Everything about the SmolLM & SmolLM2 family of models 25910
Sara Han @sdiazlor.hf.co · 23/11/2024Go to your HF profile and spot the differences. Could you find them? . . . Yes, you're right! You can now add your Bluesky account and also check your recent activity 🙌 0150
Reposted by Sara HanEmily Witko @witko.bsky.social · 22/11/2024We're big on transparency and love all things open at @huggingface.bsky.social so we thought, why not share our internal mission, vision, and values with everyone? Take a read and let us know what you think! It's such a special place! 15414
Reposted by Sara HanNathan Lambert @natolambert.bsky.social · 22/11/2024A cool space from the community to explore the many open datasets we released from building Tulu 3 (SOTA open post training recipe). 0132
Sara Han @sdiazlor.hf.co · 22/11/2024It enhances instruction following and reasoning but also includes rewriting, summarization or function calling. Evaluations have proven it outperforms the recently released Orca-AgenInstruct 1M on several tasks! 030
Sara Han @sdiazlor.hf.co · 22/11/2024The SmolLM2 recipe isn't more of a secret! SmolTalk has been released on the Hub. It uses a mix of public and synthetic data, including Magpie Ultra using distilabel 🚀 huggingface.co/datasets/Hug...huggingface.coHuggingFaceTB/smoltalk · Datasets at Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 1271
Reposted by Sara HanGabriel Martín Blázquez @gabrielmb.com · 21/11/2024Excited to announce the SFT dataset used for @huggingface.bsky.social SmolLM2! The dataset for SmolLM2 was created by combining multiple existing datasets and generating new synthetic datasets, including MagPie Ultra v1.0, using distilabel. Check out the dataset: huggingface.co/datasets/Hug...huggingface.coHuggingFaceTB/smoltalk · Datasets at Hugging FaceWe’re on a journey to advance and democratize artificial intelligence through open source and open science. 1248
Sara Han @sdiazlor.hf.co · 21/11/2024✨ Time to introduce myself to the new followers! I'm working on ML and advocacy. Till now, I have focused on building Argilla and distilabel. But I'll also share updates and cool stuff from the AI community, tools, or notebooks. P.S. Maybe a bit of my 🐕 and 🎮 too! 1110
Reposted by Sara HanBen Burtenshaw @benburtenshaw.bsky.social · 21/11/2024synthetic data on easy_mode We noticed that folk were building synthetic datasets in a common way. Basically going from prompt to dataset, and iterating. So we implemented a friendly abstraction to help you learn and improve this flow. 4255
Sara Han @sdiazlor.hf.co · 19/11/2024You might not know me yet, and that’s perfect—no bias here ✨. So, I just have a question for you. > What three key aspects would make an annotation tool ideal for you? Feel free to share your thoughts! I'm reading you 🤗 040