jsulz @jsulz.com · 03/10/2025The Hub is on 100% on Xet. 🚀 A little over a year ago, @hf.co acquired XetHub to unlock the next phase of growth in models and datasets. huggingface.co/blog/xethub-... In April, there were 1,000 Hugging Face repos on Xet. Now every repo (over 6M) on the Hub is on Xet. 2125
jsulz @jsulz.com · 15/07/2025A sneaky part of making this all work is our backward compatibility with Git LFS. This allows us to roll out a significant protocol change without forcing workflow changes We call this the Git LFS Bridge internally, and like our migration process, it's power is in its simplicity. 000
jsulz @jsulz.com · 15/07/2025You can see over the past few months some of the biggest migrations show up in our cluster throughput. Each spike corresponds to a significant migration (where we download from LFS and upload to Xet) with the baseline steadily increasing to just shy of 100 Gb/s 100
jsulz @jsulz.com · 15/07/2025The engine behind moving from Git LFS to Xet is our migration process. It's simple, powerful, and has moved well over a dozen PB just by itself. Here's a high level view of how it works. 100
jsulz @jsulz.com · 26/06/2025Meanwhile, our migrations have pushed throughput to numbers that are bonkers. In June, we hit upload speeds of 577Gb/s (crossing 500Gb/s for the first time). 100
jsulz @jsulz.com · 26/06/2025It's been a bit since I took a step back and looked at our progress to migrate @hf.co from Git LFS to Xet, but every time I do it's mind boggling. A month ago there were 5,500 users/orgs on Xet with 150K repos and 4PB. Today? 🤗 700,000 users/orgs 📈 350,000 repos 🚀 15PB 121
jsulz @jsulz.com · 12/06/2025Along with the rest of the crucial services of the internet, it is back up. But I'll never know how often she was chasing squirrels for those few hours. 010
jsulz @jsulz.com · 12/06/2025Learned GCP was out by seeing my dog door monitor was broken (backed by a cloud SQL instance). How am I going to replay/track the events of her going in and out the dog door during this outage? These are the important questions. 110
jsulz @jsulz.com · 31/05/2025I misspelled avocado once while grocery shopping. Now I only buy avacardos. 010
jsulz @jsulz.com · 21/05/2025Continuing to move all the LFS bytes into Xet storage on Hugging Face! Currently up to: 🤗 5,500 users and orgs with Xet access 🚀 150,000 Xet-backed models and datasets 🤯 4+ PB managed by Xet How much more to go? If the Hub's top storage users are any indication: many bytes 010
jsulz @jsulz.com · 30/04/2025And moving all these bytes is no joke. Our content-addressed-store (CAS) is doing a lot of hard work, hitting up to 150 Gb/s as we migrate repos from LFS to Xet. 000
jsulz @jsulz.com · 30/04/2025We've also updated our repo graph which shows how Xet-backed repos share bytes with each other. Here you can see how different versions of the Qwen, Llama, and Phi models are grouped together. Interactive graph here: huggingface.co/spaces/xet-t... 100
jsulz @jsulz.com · 09/04/2025Come find BERT island. Or see how datasets relate in practice, and how model libraries or tasks can tie repos together. It's a byte-level map of the Hub. The result is a beautiful visualization from Saba Noorassa and @reverius42.bsky.social that I’ve already lost way too much time to. 011
jsulz @jsulz.com · 07/04/2025This graph shows requests per second (rps) to our content-addressed store (CAS) right as the release went live (h/t to @rajatarya.com for the screenshot) yellow = GETs; dashed line = launch time. I think it's pretty easy to spot when Xet started to send the first bytes to excited downloaders 👀 000
jsulz @jsulz.com · 07/04/2025If you go to any model in the collection, you'll see the Xet logo supporting the many TB of tensor files. Every request to download these files comes to our infrastructure. 100
jsulz @jsulz.com · 05/04/2025With the models on our infrastructure, we can peer in and see how well our dedupe performs across the Llama 4 family. On average, we're seeing ~25% dedupe, providing huge savings to the community who iterate on these state-of-the-art models. Here's a few selected models and how they perform on Xet. 131
jsulz @jsulz.com · 26/03/2025You know you've made it when you're in the @hf.co docs. 🤗🤓 Check it out! huggingface.co/docs/hub/xet 061
jsulz @jsulz.com · 18/03/2025These are the kinds of challenges you only see when you move from theory to practice. There's nothing more satisfying than working on infrastructure for months and seeing requests funnel through and take off like a rocket 🚀 110
jsulz @jsulz.com · 18/03/2025The first migrations routed ~6% of @hf.co download traffic through Xet infrastructure. Real requests gave us an opportunity to see pods experiencing load imbalances and overhead from streaming entire blocks for partial requests. We fixed these issues on the fly without any major disruption. 110
jsulz @jsulz.com · 12/03/2025You can apply for yourself, or your entire organization. Head over to your account settings for more information or join anywhere you see the Xet logo on a repository you know. 000
jsulz @jsulz.com · 26/02/2025All good. Just communicate through GitHub comments. (somewhere, an email angel 📧 😇 lost its wings) 000
jsulz @jsulz.com · 21/02/2025Here's what you can expect with this step: ✅ First off, no action needed - this migration is helping us test and scale the infrastructure before a broader rollout 👀🔎 But you can play "spot the Xet logo" - if you see our logo on a file in a repo, that's a file we're serving now! Download away 🌐 100
jsulz @jsulz.com · 12/02/2025Bringing the infrastructure that supports this to life has amazing benefits. We can take a real repository on the Hub and see impressive savings in upload speeds as we deduplicate and aggregate. This visualization shows the dedupe across the files in a repository: Darker blocks == more dedupe. 120
jsulz @jsulz.com · 03/02/2025Cleaning out my saved Reddit posts and found this gem for interpreting sourdough crumb structure as it relates to fermentation time. Plenty of other factors that go into crumb structure (shaping/folding, starter, flour mix, etc), but fermentation time is crucial. This reference is the 🐐 #breadsky 1164
jsulz @jsulz.com · 10/01/2025I spent five weeks almost totally logged off. On Dec 9, we welcomed a new family member; Daphne, a beautiful baby girl. Now I'm gearing up to ease the transition back to work at @hf.co by spending lots of time in the kitchen. Here's 10 weeks of frozen dinners (plus a few frozen cookie balls 😋) 130
jsulz @jsulz.com · 06/12/2024This is nguha/legalbench used for evaluating legal reasoning in LLMs where (if you squint) you can see the types of reasoning being tested. huggingface.co/datasets/ngu... (this is one where it might be best to head over to huggingface.co/spaces/jsulz... and give it a spin to see for yourself. 010
jsulz @jsulz.com · 06/12/2024And here is mozilla-foundation/common_voice_17_0, the "common voice dataset" with over 30k hours of MP3 files and corresponding text files huggingface.co/datasets/moz... 110
jsulz @jsulz.com · 06/12/2024More interesting is when the directory/file naming convention lets you see the inequity in the bytes; most apparent with multilingual NLP datasets. Yellow sections for directories/files == more bytes devoted to that language. This is facebook/multilingual_librispeech huggingface.co/datasets/fac... 100
jsulz @jsulz.com · 06/12/2024I thought that large datasets would be the most interesting, and they can be - here's blanchon/RESISC45, a dataset of 31k images from Google Earth bucketed into 45 taxonomies with 700 photos per taxonomy huggingface.co/datasets/bla... 100
jsulz @jsulz.com · 06/12/2024And here is wikimedia/wikipedia, the Wikipedia dataset containing cleaned articles of all languages. Each directory contains the language abbreviation at the end. Fascinating to see which languages are more and less represented. huggingface.co/datasets/wik... 141
jsulz @jsulz.com · 06/12/2024I would not suggest trying to search for HuggingFaceFW/fineweb - huggingface.co/datasets/Hug... With over 41TB and nearly 25,000 Parquet files, this viz almost always crashes my browser. 130
jsulz @jsulz.com · 06/12/2024Always a hoot when you run across a repository with nearly 10,000 .tar files like laion/laion-audio-preview - huggingface.co/datasets/lai... 120
jsulz @jsulz.com · 06/12/2024We're hunting for interesting repos on Hugging Face: Unique file formats, big files, small files, no directories, many directories, etc. I built a tool to visualize repos and the treemaps are fun. Each one preserves the directory structure with colors for size. Here's huggingface.co/black-forest... 172
jsulz @jsulz.com · 04/12/2024No one asked me to capture the current state of my life in three words and a series of pictures, but that didn't stop me. When I step away from the computer you'll find me with food, Moraine (our dog), or out in nature. More at: huggingface.co/datasets/jsu... 050
jsulz @jsulz.com · 02/12/2024Of course we also decorate our house. Every year I jump on the roof and risk my life to string up lights. This year, we got full coverage on both floors. 110
jsulz @jsulz.com · 02/12/2024And this year we added a non-Santa to the mix. Our dog (Moraine) is now represented front and center. Yes, that's Mr. Potato Head Santa in the background. 110
jsulz @jsulz.com · 02/12/2024It's the holiday season, which means it's time to deck out our house and cover a tree with Santas. We have a collection of Santas, handed down to us by my mother-in-law who had been collecting them for decades. Every year I find several new Santas. Here's Santa with an old-school lawnmower. 230
jsulz @jsulz.com · 26/11/2024The result? A control plane with a custom protocol routing requests to content addressed stores in 3 AWS regions: us-east-1, eu-west-3, and ap-southeast-1. 🌍 130
jsulz @jsulz.com · 26/11/2024Our goal? Replace Git LFS with Xet-backed storage to make uploads smarter and downloads faster. The analysis? A 24 hour snapshot of upload requests in October - 88 countries, 8.2M requests, and 130.8TB of data - to design a scalable infrastructure. 160
jsulz @jsulz.com · 23/11/2024All this tech culminates in a novel and flexible ecosystem where I can be a more active participant, knowing my content is mine and the experience is mine to choose. It's a very new community here, but for now it seems like a nice place to post pictures of my dog and random musings like these. 010
jsulz @jsulz.com · 23/11/2024These components are open. Anyone is free to host their own PDS, read from the relay, or run their own labelers/feed generators. This is a design choice and goes back to the core philosophy of the system. 110
jsulz @jsulz.com · 23/11/2024The protocol is supported by a few key pieces of infrastructure to make the underlying data model useful/consumable - the personal data servers (PDS), relay, app view, labelers, feed generators. The figure below illustrates the interplay between each of these pieces. 110
jsulz @jsulz.com · 23/11/2024Many decentralized social media platform choose one focus, but combining both is tough. Building on that philosophy, however, unlocks exciting possibilities, like separating the social media platform from its underlying technology. 110
jsulz @jsulz.com · 23/11/2024Distributed systems are magical. Built for communication, their protocols reflect the values and philosophies of their clients. Bluesky is built with an ethos of decentralization and user experience. Check out the abstract of the paper covering the AT Protocol arxiv.org/pdf/2402.03239 110