Sign in

Daniel Vila

@dvilasuero.hf.co
3.7K followers 573 following 56 posts

Everything datasets and human feedback for AI at Hugging Face. Prev: co-founder and CEO of Argilla (acquired by Hugging Face)

PostsRepliesMedia
Reposted by Daniel Vila
Florent Daudens @fdaudens.bsky.social · 28/01/2025
🚀 The open source community is unstoppable: 4M total downloads for DeepSeek models on @hf.co , with 3.2M coming from the +600 models created by the community. That's 30% more than yesterday!
0101
Reposted by Daniel Vila
Sara Han @sdiazlor.hf.co · 20/01/2025
💫 Generate RAG data with the Synthetic Data Generator to improve your RAG system! 1️⃣ Generate from your documents, dataset, or dataset description. 2️⃣ Configure it. 3️⃣ Generate the synthetic dataset. 4️⃣ Fine-tune the retrieval and reranking models. 5️⃣ Build a RAG pipeline.
1143
Reposted by Daniel Vila
Natalia @nataliaelv.hf.co · 17/01/2025
New chapter in the Hugging Face NLP course! 🤗 🚀 We've added a new chapter about the very basics of Argilla to the Hugging Face NLP course. Learn how to set up an Argilla instance, load & annotate datasets, and export them to the Hub.  Any feedback for improvements welcome!
Screenshot of the Introduction to Argilla in Chapter 10 of the Hugging Face NLP course
1151
Reposted by Daniel Vila
Daniel van Strien @danielvanstrien.bsky.social · 16/01/2025
🎉 50,000+ annotations reached! The FineWeb2-C community is helping build better language models on annotation at a time. 📊 Current stats: - 115 languages represented - 419 amazing contributors - 24 languages with complete datasets But we're not done yet! 🧵
Screenshot of this text:   Total annotations submitted: 50,035  Languages with annotations: 115  Total contributors: 419
1196
Reposted by Daniel Vila
David Berenstein @davidberenstein.bsky.social · 07/01/2025
High-quality data for fine-tuning language models for free and at the click of a button! Prompt and wait for your dataset to push to Argilla or the Hub Evaluate, review and fine-tune a model. Blog:
buff.ly
Fine-tune a SmolLM on domain-specific synthetic data from a LLM
A Blog post by David Berenstein on Hugging Face
1102
Reposted by Daniel Vila
Daniel van Strien @danielvanstrien.bsky.social · 03/01/2025
Was 2024 the year of datasets? Is 2025 the year for community-built datasets? It's exciting to see the progress of many languages in FineWeb-C: - Total annotations submitted: 41,577 - Languages with annotations: 106 - Total contributors: 363
0294
Reposted by Daniel Vila
Daniel van Strien @danielvanstrien.bsky.social · 06/01/2025
The finish line is near! We're building FineWeb-Edu for many languages and need your help 🤗 Many FineWeb-C languages are close to 1,000 annotations! Assamese is 99.4% done, French needs 64 more annotations, Tamil: 216. Please help us reach the goal: huggingface.co/spaces/data-...
Progress bars showing remaining annotations needed for 15 languages in FineWeb-C dataset, ranging from 6 to 593 annotations needed
1205
Daniel Vila @dvilasuero.hf.co · 20/12/2024
💥 Ending 2024: A full data annotation journey on the Hugging Face Hub—from raw data to training-ready datasets! With Argilla 2.6.0, push your data to the Hub from the UI Let’s make 2025 the year anyone can build more transparent and accountable AI—no coding or model skills needed.
1213
Reposted by Daniel Vila
José Francisco Calvo @jfcalvo.hf.co · 19/12/2024
🚀 Argilla v2.6.0 is here! 🎉 Let me show you how EASY it is to export your annotated datasets from Argilla to the Hugging Face Hub. 🤩 Take a look to this quick demo 👇 💁‍♂️ More info about the release at github.com/argilla-io/a... #AI #MachineLearning #OpenSource #DataScience #HuggingFace #Argilla
0115
Reposted by Daniel Vila
David Berenstein @davidberenstein.bsky.social · 17/12/2024
🔥 We got great feedback on this: "Synthetic Data Generator" A no-code tool to create datasets with LLMs, making it a breeze, allowing ANYONE to create datasets and models in minutes and without any code. Blog: buff.ly/4gybyoT GitHub: buff.ly/49IDSmd Space: buff.ly/3Y1S99z
buff.ly
Introducing the Synthetic Data Generator - Build Datasets with Natural Language
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1142
Reposted by Daniel Vila
Ashvanth.S @ashvanths.bsky.social · 14/12/2024
Well, around 10 percent of the initial goal is complete, and so far, it's been quite a one-man army effort. We're still in the hunt for more people to join and contribute to this open-source initiative. @hf.co data-is-better-together-fineweb-c.hf.space/share-your-p...
data-is-better-together-fineweb-c.hf.space
tam - தமிழ் - Tamil
Join and contribute to the dataset tam - தமிழ் - Tamil
141
Reposted by Daniel Vila
Johannes @johko.bsky.social · 13/12/2024
The sprint for crowd sourced annotations with argilla is in full swing over at data-is-better-together-fineweb-c.hf.space I've just contributed 100 examples to this dataset: data-is-better-together-fineweb-c.hf.space/share-your-p... Big thanks to @dvilasuero.hf.co, @nataliaelv.hf.co and team 🙌
data-is-better-together-fineweb-c.hf.space
nds - Neddersass’sch - Low German
Join and contribute to the dataset nds - Neddersass’sch - Low German
0121
Reposted by Daniel Vila
Moritz Laurer @moritzlaurer.bsky.social · 12/12/2024
I've been building a small library for working with prompt templates on the @huggingface.bsky.social Hub: `pip install prompt-templates`. Motivation: The community currently shares prompt templates in a wide variety of formats: in datasets, in model cards, as strings in .py files, as .txt/... 🧵
1164
Reposted by Daniel Vila
Ben Burtenshaw @benburtenshaw.bsky.social · 12/12/2024
Desperate to contribute to the development of Scots language AI. I've just contributed 16 examples to this dataset: data-is-better-together-fineweb-c.hf.space/share-your-p...
data-is-better-together-fineweb-c.hf.space
sco - Scots - Scots
Join and contribute to the dataset sco - Scots - Scots
1103
Daniel Vila @dvilasuero.hf.co · 12/12/2024
I've just contributed 156 examples to the FineWeb 2 Spanish dataset: data-is-better-together-fineweb-c.hf.space/share-your-p... If you want to contribute, sign in with @hf.co and find your language
data-is-better-together-fineweb-c.hf.space
spa - español - Spanish
Join and contribute to the dataset spa - español - Spanish
1225
Daniel Vila @dvilasuero.hf.co · 10/12/2024
Help shape the future of multilingual Open Source AI! Join the FineWeb 2 Community Annotation Sprint to create an open training dataset with full transparency and human validation in many languages. Review datasets in your language and help identify the best sources for training.
1203
Reposted by Daniel Vila
frascuchon.bsky.social @frascuchon.bsky.social · 03/12/2024
✨ Argilla 2.5.0 is live and it comes with webhook listener support to supercharge your workflows! 🚀 #AI #MachineLearning #Webhooks #TechUpdate
182
Reposted by Daniel Vila
David Berenstein @davidberenstein.bsky.social · 09/12/2024
👐 Open Image Preferences is an Apache 2.0 licensed dataset for text-to-image generation by the @hf.co community. This dataset contains 10K text-to-image preference pairs across image generation categories, using different model families and prompt complexities. Blog: huggingface.co/blog/image-p...
huggingface.co
Open Preference Dataset for Text-to-Image Generation by the 🤗 Community
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1174
Reposted by Daniel Vila
Sara Han @sdiazlor.hf.co · 09/12/2024
Open Image Preferences released! 🚀 - Open-source dataset for text2image - 10K samples manually evaluated by the HF community. - Binarized format for SFT, DPO, or ORPO. It comes with a nice blog post explaining the steps to pre-process and generate the data, along with the results.
151
Daniel Vila @dvilasuero.hf.co · 06/12/2024
Announcing Global-MMLU - an improved MMLU Open dataset with evaluation coverage across 42 languages. The result of months of work with the goal of advancing Multilingual LLM evaluation. Built together with the community and amazing collaborators at Cohere4AI, MILA, MIT, and many more.
56611
Daniel Vila @dvilasuero.hf.co · 03/12/2024
We're about to launch the biggest collaboration effort since the Open Assistant. Let's get the highest quality data for open foundation models with all the nuances & diversity of each language, all with data provenance and transparency Join us as language lead: docs.google.com/forms/d/10XI...
docs.google.com
Language Lead sign-up
At Hugging Face 🤗, we're launching a big community initiative to improve LLM training for many languages. We're looking for Language Leads to help us cultivate specific languages during this initiativ...
073
Reposted by Daniel Vila
Natalia @nataliaelv.hf.co · 03/12/2024
Next week we're launching a collaborative annotation effort to build a big multilingual dataset, so you can have high-quality data in your language. We are really close to getting leads for 100 languages! Can you help us cover the remaining 200?
Screenshot of a dashboard showing the number of languages with a lead and languages without a lead
4154
Reposted by Daniel Vila
Ben Burtenshaw @benburtenshaw.bsky.social · 03/12/2024
For anyone interested in fine-tuning or aligning LLMs, I’m running this free and open course called smol course. It’s not a big deal, it’s just smol. 🧵>>
932363
Reposted by Daniel Vila
José Francisco Calvo @jfcalvo.hf.co · 02/12/2024
🙌 I just wanted to share a few thoughts about the latest Argilla release, 2.5.0, as it's a pretty big one! Argilla now has full support for webhooks, which means you can do some pretty cool stuff, like model training on the fly as annotations are created. 🤯 #MachineLearning #NLP #DataLabeling
153
Reposted by Daniel Vila
Ben Burtenshaw @benburtenshaw.bsky.social · 30/11/2024
[SATURDAY THREAD] ☕️ 🧑‍🎓 In case you spent the week reading GDPR legislation and missed everything. It’s all about vision language models and image preference datasets. >> 🧵 Here are the models and datasets you can use in your projects.
4395
Reposted by Daniel Vila
Damián Pumar @damianpumar.hf.co · 26/11/2024
Recently, I added a feature to #Argilla to optimize plugin loading 🎉. It removes unnecessary code, improves readability, and lets future plugins load automatically. 🚀 Check out the PR 👇 and make your first contribution to our repo. github.com/argilla-io/a... #dev_experience #clean_code
github.com
🔥 Improve plugins loaders by damianpumar · Pull Request #5697 · argilla-io/argilla
Remove duplicated names Improve the way to load plugins and extensions Auto loading new plugins Convert to typescript Delete unused directive
091
Reposted by Daniel Vila
Damián Pumar @damianpumar.hf.co · 29/11/2024
🚀 We’re excited to announce Argilla v2.5.0, which includes: * Argilla webhooks, * A new design for the datasets home page. * Python 3.13 and Pydantic v2 support. 📙 Read here 👇 the full release notes github.com/argilla-io/a...
github.com
Release v2.5.0 · argilla-io/argilla
🔆 Release highlights Webhooks You can now create and manage webhooks to support your workflows! Webhooks allow you to submit real-time information to other applications whenever a specific event oc...
0172
Reposted by Daniel Vila
Stella Biderman @stellaathena.bsky.social · 28/11/2024
A dataset of 1 million or 2 million Bluesky posts is completely irrelevant to training large language models. The primary usecase for the datasets that people are losing their shit over isn't ChatGPT, it's social science research and developing systems that improve Bluesky.
825139
Reposted by Daniel Vila
Margaret Mitchell @mmitchell.bsky.social · 27/11/2024
The best path forward in AI requires technologists to be reflective/self-critical about how their work impacts society. Transparency helps this. Appreciate Bsky for flagging AI ethics &my colleague’s response. Let’s make informed consent a real thing. More later; Recommend: bsky.app/profile/cfie...
611723
Reposted by Daniel Vila
Daniel van Strien @danielvanstrien.bsky.social · 27/11/2024
I've removed the Bluesky data from the repo. While I wanted to support tool development for the platform, I recognize this approach violated principles of transparency and consent in data collection. I apologize for this mistake.
15264493
Reposted by Daniel Vila
Ben Burtenshaw @benburtenshaw.bsky.social · 26/11/2024
The community has labelled over 3000 image preferences in a few hours. One open source image preferences dataset coming right up!
6385
Reposted by Daniel Vila
Natalia @nataliaelv.hf.co · 26/11/2024
At @huggingface.bsky.social 🤗 we're preparing a collaborative annotation effort to build an open-source multilingual dataset. If you'd like to get high-quality open data for your language, check if yours is listed in this form and sign up! forms.gle/DHJdtvoSNxAA...
forms.gle
Language Lead sign-up
At Hugging Face 🤗, we're launching a big community initiative to improve LLM training for many languages. We're looking for Language Leads to help us cultivate specific languages during this initiativ...
4309
Reposted by Daniel Vila
Florent Daudens @fdaudens.bsky.social · 26/11/2024
🎨 Want better open-source AI art models? We need your help! Most top image generators are trained on human preferences—but those datasets are closed. Let's build our own! Rate images in pairs and help make AI art accessible to everyone 🔓 👉 huggingface.co/blog/burtens...
8103
Reposted by Daniel Vila
Daniel van Strien @danielvanstrien.bsky.social · 26/11/2024
First dataset for the new @huggingface.bsky.social @bsky.app community organisation: one-million-bluesky-posts 🦋 📊 1M public posts from Bluesky's firehose API 🔍 Includes text, metadata, and language predictions 🔬 Perfect to experiment with using ML for Bluesky 🤗 huggingface.co/datasets/blu...
huggingface.co
bluesky-community/one-million-bluesky-posts · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
69852772
Reposted by Daniel Vila
José Francisco Calvo @jfcalvo.hf.co · 26/11/2024
Did you know that on Argilla, we’re adding a new feature to export labeled datasets directly to the Hugging Face Hub? 🤔 We’re leveraging the Hugging Face datasets library for seamless integration, including defining span labeling Stay tuned for the release!🧠✨ #MachineLearning #NLP #DataLabeling
A snippet of code with the following content:

Sequence({
    "label": ClassLabel(names=question.values),
    "start": features.Value(dtype="int64"),
    "end": features.Value(dtype="int64"),
})
0222
Reposted by Daniel Vila
Amélie Viallet @ameeelie.bsky.social · 26/11/2024
👀 Who said the Argilla tool was only for text? I am proud of my brilliant teammates for setting up this significant initiative 🤗 @benburtenshaw.bsky.social @davidberenstein.bsky.social @danielvanstrien.bsky.social @dvilasuero.hf.co
021
Daniel Vila @dvilasuero.hf.co · 26/11/2024
Super excited to launch the Open Images Preferences @huggingface.bsky.social community sprint Have fun browsing images generated with the latest OSS models while contributing to the future of Open Source AI 🧵
2162
Reposted by Daniel Vila
Mark Collier @markcollier.me · 24/11/2024
Added some more folks to the Open Source AI Starter Pack: go.bsky.app/N8yVZdW
217922
Reposted by Daniel Vila
Ashvanth.S @ashvanths.bsky.social · 26/11/2024
Glad to see HF do this initiative for training models
031
Daniel Vila @dvilasuero.hf.co · 26/11/2024
Let's make AI more inclusive. At @huggingface.bsky.social we'll launch a huge community sprint soon to build high-quality training datasets for many languages. We're looking for Language Leads to help with outreach. Find your language and nominate yourself: forms.gle/iAJVauUQ3FN8...
85321
Reposted by Daniel Vila
Emily Witko @witko.bsky.social · 25/11/2024
ICYMI, my colleague @dvilasuero.hf.co created a feed you can use to keep up to date on 🤗 datasets. Pin it!
031
Reposted by Daniel Vila
Daniel van Strien @danielvanstrien.bsky.social · 25/11/2024
The AT Protocol unlocks exciting possibilities: - Building custom feeds using ML - Creating dashboards for data exploration - Developing custom models for Bluesky To gather @bsky.app resources on @huggingface.bsky.social. I've established a community org 🤗 huggingface.co/bluesky-comm...
huggingface.co
bluesky-community (Bluesky Community)
Tools for Bluesky 🦋
1015933
Reposted by Daniel Vila
Anthony A. Gatti @aagatti.bsky.social · 25/11/2024
I'm excited to share that our work on #deeplearning based shape modeling is finally out in #ieee. 🧵 for info about Data, Benchmarks, and Models! @akshay-chaudhari.bsky.social @stanfordmedicine.bsky.social www.medrxiv.org/content/10.1... ieeexplore.ieee.org/document/107...
1228
Daniel Vila @dvilasuero.hf.co · 25/11/2024
Interested in open datasets for ML and AI? I've just created this feed with posts about @huggingface.bsky.social datasets! Don't miss the latest news and conversations about the secret sauce behind every AI model. bsky.app/profile/dvil...
34512
Daniel Vila @dvilasuero.hf.co · 25/11/2024
This is super useful for NLP lovers like myself
070
Reposted by Daniel Vila
Philipp Schmid @philschmid.bsky.social · 25/11/2024
Created a visual for how function calling works. Wdyt? 🤔
6242
Reposted by Daniel Vila
Ben Burtenshaw @benburtenshaw.bsky.social · 25/11/2024
TRL is a cornerstone of LLM post training and imo it's the default to learn. There are great alternatives like Unsloth, Axolotl, and AutoTrain. But if you want a daily drive that does experimentation to production, it's TRL. 🧵 these community notebooks guide you through TRL's core:
3568
Daniel Vila @dvilasuero.hf.co · 24/11/2024
I am very excited to launch a new community initiative next week. Let's build the largest open community dataset to evaluate and improve image generation models. Follow: huggingface.co/data-is-bett... And stay tuned here
huggingface.co
data-is-better-together (Data Is Better Together)
Building better datasets together
18512
Reposted by Daniel Vila
Hugging Face @hf.co · 22/11/2024
We created these values as a team, and we live them as a team 💪
7486
Reposted by Daniel Vila
Ben Burtenshaw @benburtenshaw.bsky.social · 23/11/2024
In case you passed out and woke up on saturday lunch. Small models and high quality data are back! ... if they ever left 🤔 - SmolTalk dataset from @huggingface.bsky.social - Tulu 3 models and datasets from @ai2.bsky.social - Nvidia Nymba model from @nvidiastudio.bsky.social
18610