Sign in

Hugging Face

@hf.co
17K followers 53 following 3 posts

The AI community building the future!

PostsRepliesMedia
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 21/11/2025
Building datasets to train smaller, task-focused models used to be incredibly time-consuming. Very excited to see SAM3 massively lower that barrier. Describe the class you want to detect and get annotated datasets automatically! Try it yourself: huggingface.co/datasets/uv-...!
Screenshot of a simple app showing bounding boxes for photographs detected in historic newspaper images. hf jobs uv run \
  --flavor a100-large \
  -s HF_TOKEN=HF_TOKEN \
  https://huggingface.co/datasets/uv-scripts/sam3/raw/main/detect-objects.py \
  -- davanstrien/newspapers-with-images-after-photography-big \
  davanstrien/newspapers-photo-predictions \
  --class-name "photograph" \
  --confidence-threshold 0.4
15712
Reposted by Hugging Face
Julien Chaumond @julien-c.hf.co · 02/11/2025
Training LLMs end to end is hard. But way more people should, and will, be doing it in the future. The @hf.co Research team is excited to share their new e-book that covers the full pipeline: · pre-training, · post-training, · infra. 200+ pages of what worked and what didn’t. ⤵️
415227
Reposted by Hugging Face
EvE Bio @evebio.bsky.social · 18/11/2025
💻 Our pharmome mapping data is now accessible to ML developers on @hf.co, making our purpose-built drug-target interaction data easily accessible for model development. huggingface.co/blog/hugging...
huggingface.co
The Pharmome Map: a comprehensive public dataset for drug-target interaction modeling
A Blog post by Hugging Science on Hugging Face
192
Reposted by Hugging Face
Jacob S. Zelko @thecedarprince.bsky.social · 20/10/2025
JuliaHealth is on @hf.co! 🤗 If you are interested in #julialang, #llm or #agentic workflows, and how #GenAI can be used within public health, medical informatics, and survey-based research, drop us a line or a follow! 🤓 How are you using GenAI in your medical research? #opensource #medsky #data
huggingface.co
JuliaHealthOrg (The JuliaHealth Organization)
Org profile for The JuliaHealth Organization on Hugging Face, the AI community building the future.
0155
Reposted by Hugging Face
Towards Data Science @towardsdatascience.com · 15/10/2025
When we build an app, it’s only natural to want to share it. Ivo Bernardo walks you through a short tutorial on how to deploy your own @hf.co Space. If you want to highlight your work and applications, this is a strong option.
towardsdatascience.com
Showcasing Your Work on HuggingFace Spaces | Towards Data Science
Building an app is exciting - but sharing it is where the real value kicks in. Back when Heroku offered a free tier, deploying demos was effortless. Those days are gone, and finding a simple, free…
041
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 16/10/2025
Small models work great for GLAM but there aren't enough examples! With @wjbmattingly.bsky.social I'm launching small-models-for-glam on @hf.co to create/curate models that run on modest hardware and address GLAM use cases. Follow the org to keep up-to-date! huggingface.co/small-models...
0127
Reposted by Hugging Face
Thibault Clérice @ponteineptique.bsky.social · 15/10/2025
(10/🧵) The corpus isn’t just readable 👁️ — it’s also fully downloadable! Now hosted on @hf.co : 🧾 JSONL dataset → huggingface.co/datasets/com... 📂 More formats (ALTO, TEI, etc.) coming soon — we’re uploading the GBs as we speak.
huggingface.co
comma-project/comma-jsonl · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
141
Reposted by Hugging Face
allegra @sequelbox.bsky.social · 06/10/2025
Esper 3.1 is here on @hf.co - our DevOps, coding, and architecture specialist is back, trained on higher difficulty data! For everyone to use: huggingface.co/ValiantLabs/...
huggingface.co
ValiantLabs/Qwen3-4B-Thinking-2507-Esper3.1 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1101
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 02/10/2025
New @hf.co BigLAM dataset: 9,363 OA books with page images + rich MARC metadata for evaluating (and training) VLMs on metadata extraction. Libraries are starting to explore AI-assisted cataloguing, but we lack public evaluation data. Hoping this helps fill that gap. huggingface.co/datasets/big...
Screenshot of the dataset viewer showing a column of marc data + the first few pages of an open access monograph
2329
Reposted by Hugging Face
Erik @erikkaum.bsky.social · 06/10/2025
We have Nvidia B200s ready to go for you in Hugging Face Inference Endpoints 🔥 I tried them out myself and the performance is amazing. On top of that we just got a fresh batch of H100s as well. At $4.5/hour it's a clear winner in terms of price/perf compared to the A100.
061
Reposted by Hugging Face
Martin Mundt @martinmundt.bsky.social · 06/10/2025
Wonder how LLMs learn over long time horizons & how hate-checks deal with time? Look at our new work "Chronoberg": an open-source dataset spanning 250 years of books with analysis of shifts in meaning & continual learning of LLMs: arxiv.org/pdf/2509.22360 huggingface.co/datasets/spa...
051
Reposted by Hugging Face
jsulz @jsulz.com · 03/10/2025
The Hub is on 100% on Xet. 🚀 A little over a year ago, @hf.co acquired XetHub to unlock the next phase of growth in models and datasets. huggingface.co/blog/xethub-... In April, there were 1,000 Hugging Face repos on Xet. Now every repo (over 6M) on the Hub is on Xet.
Graph showing the conversion of Hugging Face repositories from LFS storage to Xet storage.
2125
Reposted by Hugging Face
Giada Pistilli @giadapistilli.com · 29/09/2025
One of the hardest challenges in AI safety is finding the right balance: how do we protect people from harm without undermining their agency? This tension is especially visible in conversational systems, where safeguards can sometimes feel more paternalistic than supportive.
1111
Reposted by Hugging Face
Ben Trent @benwtrent.bsky.social · 01/10/2025
The @hf.co community is awesome. Real work that moves everyone forward: huggingface.co/blog/rteb
huggingface.co
Introducing RTEB: A New Standard for Retrieval Evaluation
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
021
Reposted by Hugging Face
Earth Species Project (ESP) @earthspecies.bsky.social · 04/09/2025
🐋🐦🐸 We just launched an interactive demo of NatureLM-audio on @hf.co! 👉 Try the demo with your audio or ours, share your feedback, and help us shape the future of decoding animal communication: huggingface.co/blog/EarthSp...
0113
Reposted by Hugging Face
Gradio @gradio-hf.bsky.social · 05/08/2025
You only need one line of code to start exploring the new @OpenAI models! gr.load("models/openai/gpt-oss-120b", provider="fireworks-ai").launch()
1112
Reposted by Hugging Face
Florent Daudens @fdaudens.bsky.social · 05/08/2025
Well, it took just 2 hours for GPT-OSS to hit #1 on @hf.co. Don’t remember seeing anything rise that fast!
0122
Reposted by Hugging Face
pngwn @pngwn.at · 06/08/2025
OpenAI have released their new open source models! One thing I really like about this release is that while they are only open weight, the model is not gated in any way (anyone can download it) and it has a permissive OSS license (apache 2). Very refreshing. huggingface.co/openai/gpt-o...
huggingface.co
openai/gpt-oss-120b · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
2135
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 06/08/2025
You can now generate synthetic data using OpenAIs GPT OSS models on @hf.co Jobs! One command, no setup: hf jobs uv run --flavor l4x4 [script-url] \ --input-dataset your/dataset \ --output-dataset your/output Works on L4 GPUs ⚡ huggingface.co/datasets/uv-...
huggingface.co
uv-scripts/openai-oss · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
0111
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 01/08/2025
Many VLM-based OCR models have been released recently. Are they useful for libraries and archives? I made a quick Space to compare VLM OCR with "traditional" OCR using 11k Scottish exam papers from @natlibscot.bsky.social huggingface.co/spaces/davanstrien/ocr-time-capsule
Screenshot of the app showing a page from a book + different views of existing and new ocr.
44715
Reposted by Hugging Face
jsulz @jsulz.com · 30/07/2025
We just crossed 1 million repositories backed by Xet storage on @hf.co I celebrated by reviving the early 2000s web design aesthetics that I love so much. Here's our dashboard showing our progress converting the Hub from Git LFS to Xet (and demonstrating my questionable design sensibilities).
huggingface.co
Ready Xet Go - a Hugging Face Space by jsulz
This app helps you monitor the progress of migrating repositories to Xet, showing you stats and charts on migration status and file types.
171
Reposted by Hugging Face
Rajat Arya @rajatarya.com · 30/07/2025
Built my 1st app exclusively using @hf.co Hub features! It helps me keep track/summarize the latest HF News. Uses Datasets, Inference Endpoints, the newly announced `hf jobs`, and Spaces to visualize results. Check it out here: huggingface.co/spaces/rajat...
Screenshot of running HF News Aggregator.
061
Reposted by Hugging Face
Rasmus Aagaard @rasgaard.com · 28/07/2025
`hf jobs` looks super interesting. Send off workloads easily to remote infra. And with inline metadata uv scripts `hf jobs uv run my_script.py` all dependencies can be defined in a single file. So simple and so useful. huggingface.co/docs/hugging...
huggingface.co
Run and manage Jobs
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
0143
Reposted by Hugging Face
William J.B. Mattingly @wjbmattingly.bsky.social · 28/07/2025
Working to port VLaMy to an entirely free mode where you can just cache all your data in the browser for a project. Slowly adding all the features from the full version to this user-free version. Available now on @hf.co @danielvanstrien.bsky.social Link: huggingface.co/spaces/wjbma...
2206
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 08/07/2025
465 people. 122 languages. 58,185 annotations! FineWeb-C v1 is complete! Communities worldwide have built their own educational quality datasets, proving that we don't need to wait for big tech to support languages. Huge thanks to all who contributed! huggingface.co/blog/davanst...
23311
Reposted by Hugging Face
Tuana @tuana.dev · 26/06/2025
Last week, we concluded the @gradio-hf.bsky.social‬ MCP hackathon with @hf.co‬. The project that one the @llamaindex.bsky.social prize was the "Nasa Space Explorer" 🔭🪐 3 servers that provide live data on: ☄️ Asteroids 🤖 the Mars Rover 🌌 Astronomy Here's the space: huggingface.co/spaces/Agen...
062
Reposted by Hugging Face
Tim @nerddis.co · 05/07/2025
i love the simplicity of LeRobot from @hf.co to interact with robots, especially for beginners like me there is one very huge problem: it's written in python, but i love js introducing: LeRobot.js interact with your robot directly in the browser www.youtube.com/watch?v=H1iU...
youtube.com
introducing LeRobot.js - interact with your robot in the browser
YouTube video by Tim Pietrusky
1123
Reposted by Hugging Face
Georgia Channing @cgeorgiaw.bsky.social · 06/07/2025
🧬 super psyched to announce a new collaboration between @hf.co and Ginkgo Datapoints to open up high-quality biological datasets for the machine learning community! Just dropped the GDPx and GDPa dataset series on the Hub (x1000 boost to AI for drug development) 🔗 huggingface.co/ginkgo-datap...
071
Reposted by Hugging Face
Reihaneh Rabbany @reirab.com · 19/06/2025
If you are interested in a unified collection of common misinformation detection benchmarks, check out our recent repo @hf.co
0166
Reposted by Hugging Face
William J.B. Mattingly @wjbmattingly.bsky.social · 25/06/2025
Have high quality data sitting on Transkribus? Want to make it available on @hf.co with a single command line? Introducing Transkribus-HF which allows you to take a Transkribus export zip and make it into a HF dataset! It can parse pages, regions, lines, or windows! github.com/wjbmattingly...
2125
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 24/06/2025
Everyone’s dropping VLM-based OCR models lately… But are they actually better than traditional OCR engines, which output XML for historical docs? I built OCR Time Machine to test it! 📄 Upload image + ALTO/PAGE XML ⚖️ Compare outputs side by side 🔗 huggingface.co/spaces/davan...
Screenshot showing a document page image on the left with corresponding OCR output on the right of the page.
2309
Reposted by Hugging Face
Adina Yakup @adinayakup.bsky.social · 16/06/2025
MiniMax-M1 🔥 The first reasoning model by MiniMax AI is now live on @hf.co huggingface.co/collections/... ✨ 40k/80k thinking budget ✨ Powered by Hybrid MoE + Lightning Attention 👀 ✨ 1M context length 🤯 ✨ Apache 2.0 ✨ RL-trained for math, coding & real-world software
huggingface.co
MiniMax-M1 - a MiniMaxAI Collection
MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model.
1165
Reposted by Hugging Face
Daniël de Kok @danieldk.eu · 17/06/2025
Over the past few months, we have worked on the @hf.co Kernel Hub. Kernel Hub allows you to get cutting-edge compute kernels directly from the hub in a few lines of code. David Holz made a great writeup of how you can use kernels in your projects: huggingface.co/blog/hello-h...
huggingface.co
Learn the Hugging Face Kernel Hub in 5 Minutes
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
092
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 17/06/2025
“AI Scraping Bots Are Breaking Open Libraries, Archives, and Museums” – interesting piece via @404media.co Not a perfect fix, but making ML-ready datasets from collections can help. If you want help getting your data on @hf.co, I'd be happy to help.
Screenshot of the header of the article with text:

AI Scraping Bots Are Breaking Open Libraries, Archives, and Museums
0144
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 16/06/2025
Institutional Books: Massive Historical Text Corpus - 983K books, 242B tokens, 386M pages - 19th-20th century texts in 254 languages - Refined OCR with quality scores & metadata - Noncommercial early-access release huggingface.co/datasets/ins...
huggingface.co
institutional/institutional-books-1.0 · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
03815
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 13/06/2025
How to Make Gallery, Library, Archive & Museum Collections Ready for AI 📚✨ Join me this Tuesday – I'm hosting a hands-on session for the #ai4lam community, where I'll show cultural institutions how to share data on Hub. 🗓️ June 17, 16:00 UK 📍 Details: docs.google.com/document/d/1...
docs.google.com
2025-06-17 ai4lam Community Call
ai4lam Community Call Tuesday, June 17, 2025 8:00 AM California | 11:00 AM Washington DC | 16:00 UK | 17:00 Oslo & Paris | 01:00 Sydney Find your local time here Connection Information: Join Zoom Me...
0112
Reposted by Hugging Face
Kevin Slote @kevinslote.bsky.social · 11/06/2025
huggingface.co/Belykh-Lab I’m trying to start a trend in the applied math community of releasing data sets on @hf.co
huggingface.co
Belykh-Lab (Biological and Engineering Networks Lab)
Org profile for Biological and Engineering Networks Lab on Hugging Face, the AI community building the future.
0123
Reposted by Hugging Face
IDI @institutional.org · 12/06/2025
We look forward to growing Institutional Books through community. We welcome collaboration from researchers and model makers as we: - Evaluate the dataset’s impact on model outputs - Continuing to refine our OCR pipelines View the dataset on Hugging Face: huggingface.co/datasets/ins...
huggingface.co
institutional/institutional-books-1.0 · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1162
Reposted by Hugging Face
Marian Veteanu @mveteanu.bsky.social · 06/06/2025
Hugging Face released their MCP server. You can see it here configured and running in Claude Desktop. @hf.co
1153
Reposted by Hugging Face
Matt Biddulph @biddul.ph · 07/06/2025
first steps with the @hf.co SO-101 AI robot arm
0183
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 09/06/2025
Inspired by @hf.co's official MCP server, I built my own to expose my semantic search API for the HF ecosystem! Features AI-powered search, parameter analysis via safetensors, and tools to find similar models/datasets. Try: "Find non maths reasoning datasets from 2025"!
194
Reposted by Hugging Face
Karen Hao @karenhao.bsky.social · 18/05/2025
Such an important project: @hf.co put up an interactive site to see the real time energy costs of chatting with genAI. "Calculate how much water it would take to cool the world's largest supercomputer" took 13% of a smartphone battery. Complete with hallucinations. 😆 huggingface.co/spaces/jdela...
huggingface.co
Chat UI Energy Score - a Hugging Face Space by jdelavande
Chat with an AI assistant and see how much energy your conversation uses. Get real-time energy estimates compared to everyday activities like phone charging or driving.
34612
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 20/05/2025
🗞️ Just released a Parquet version of the Newspaper Navigator dataset on @hf.co! - 3M+ visual elements from historic US newspapers — photos, maps, cartoons, OCR + metadata. - Parquet = fast filters, easier analysis. - Great for ML + cultural research. 👉 huggingface.co/datasets/big...
Screenshot of the dataset viewer on the Hugging Face Hub. Shows a set of metadata for the newspaper navigator dataset. It also has previews of a few rows showing images alongside metadata columns.
1137
Reposted by Hugging Face
Daniel van Strien @danielvanstrien.bsky.social · 08/05/2025
Finally documented the Beyond Words dataset from the @librarycongress.bsky.social labs / @bcgl.bsky.social for the BigLAM @hf.co org! - 3.5K annotated historical newspaper pages - Bounding boxes + category labels - Photos, ads, headlines, cartoons & more
Image of a historic newspaper with bounding box predictions for "photographs" "headline" "illustration" etc.
13310
Reposted by Hugging Face
Nicolas Barradeau @nicoptere.bsky.social · 04/05/2025
Not sure how long it has been around but it’s a really great ML resource huggingface.co/learn by @hf.co
huggingface.co
Hugging Face - Learn
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
0225
Reposted by Hugging Face
Jack (in SF) Langerman @jacklangerman.bsky.social · 07/05/2025
🚨 Just one month left to submit your solutions for The Structured Semantic 3D Reconstruction (S23DR-2025) Challenge!!! It is not too late to join! Comp on @hf.co, part of the Workshop on Urban Scene Modeling at @cvprconference.bsky.social 2025 🔥$25,000 prize pool. Deadline: June 5, 2025. 🧵 (1/7)
154
Reposted by Hugging Face
Giada Pistilli @giadapistilli.com · 07/05/2025
Ever notice how some AI assistants feel like tools while others feel like companions? Turns out, it's not always about fancy tech upgrades, because sometimes it's just clever design. huggingface.co/blog/giadap/...
huggingface.co
AI Personas: The Impact of Design Choices
A Blog post by Giada Pistilli on Hugging Face
1105
Reposted by Hugging Face
TNG Technology Consulting GmbH @tngtech.com · 02/05/2025
Introducing DeepSeek-R1T-Chimera: our new open weights model adds #R1 reasoning to #DeepSeek V3-0324. In benchmarks, the hybrid child model appears to be as smart as R1 but uses 40% fewer output tokens. Available on @hf.co & Open Router.
2112
Reposted by Hugging Face
William J.B. Mattingly @wjbmattingly.bsky.social · 02/05/2025
New free HTR app nearly ready to share! This is a simple local app that creates a local database that lets you create projects which have document images that you can then send off to Qwen 2.5 VL and GliNER on my Caracal app on @hf.co entirely for free. You just need a free HF account/token.
051