Sign in

Daniel van Strien

@danielvanstrien.bsky.social
4.8K followers 520 following 417 posts

Machine Learning Librarian at @hf.co

PostsRepliesMedia
Daniel van Strien @danielvanstrien.bsky.social · 30/09/2026
Uploaded a dataset of 98,877 historical newspaper pages (1700s–1940s) to the Hub, each with its original OCR text, word boxes and confidence scores. huggingface.co/datasets/big...
Sample grid of newspapers in the dataset
24917
Daniel van Strien @danielvanstrien.bsky.social · 16/09/2026
Made some improvements to the OCR scripts onboarding in my uv-scripts collection. First run is now a small example: seven scanned NASA pages in, Markdown out, one HF jobs command. It also works on your own images or a folder of PDFs. You can pick from 28 OCR models. huggingface.co/datasets/uv-...
Image of a old scanned typed page on the left and on the right is markdown output
1141
Daniel van Strien @danielvanstrien.bsky.social · 11/09/2026
Need a historical illustration? Search 1.49 million images from British Library books and Britannica (1500s–1920s). Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources. huggingface.co/spaces/davan...
47620
Daniel van Strien @danielvanstrien.bsky.social · 03/09/2026
A 15.9M-parameter specialist OCR model outperforms many VLMs hundreds of times larger. On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved. huggingface.co/spaces/fineb...
Plot showing parameters vs performance. Kraken is top left.
2346
Daniel van Strien @danielvanstrien.bsky.social · 26/08/2026
Every illustration in the Encyclopaedia Britannica dataset now has an instance mask: 411,385 cut-outs from 115,293 pages, 1768–1929, each linked to its full-resolution scan. Public domain, no image generation involved. Work in progress.
The left image shows an illustration surrounded by text. right image shows image cropped via mask prediction
03710
Daniel van Strien @danielvanstrien.bsky.social · 25/08/2026
Uploaded a dataset of 115,293 illustrated pages from the Encyclopaedia Britannica, 1st edition (1768) to 14th (1929) to the Hub huggingface.co/datasets/big...
grid showing examples from the dataset
0335
Daniel van Strien @danielvanstrien.bsky.social · 20/08/2026
Every illustration in the British Library's 19th-century book images dataset now has an instance mask: 1.02M clean cutouts of public-domain engravings, no image generation involved! One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.
111425
Daniel van Strien @danielvanstrien.bsky.social · 10/08/2026
FineBooks wants to help re-OCR the public domain to unlock a rich source of data for AI, for researchers, and for enjoyment. Step one: work out which modern OCR models are actually good enough.
Leaderboard preview
2255
Daniel van Strien @danielvanstrien.bsky.social · 07/08/2026
Uploaded a dataset of 1,080,814 public domain images, mostly from 19th-century books, to the Hugging Face Hub: huggingface.co/datasets/big... You can also do semantic search against the images here: huggingface.co/spaces/davan...
15412
Daniel van Strien @danielvanstrien.bsky.social · 03/08/2026
Public data on which coding agents actually use the Hugging Face Hub, refreshed monthly. Claude Code now sends a majority of everything the Hub can attribute to a named agent. In May, it led on 4 days out of 30. In July, all 30. huggingface.co/datasets/hug...
plot showing growth of different coding agents using the hub
4132
Daniel van Strien @danielvanstrien.bsky.social · 31/07/2026
A film catalogue tells you what a film is about, not what happens inside it. So: 1,864 public-domain Prelinger films, described every ~60 seconds by an open 2B video model. 23,148 searchable moments. Search "typing on a computer keyboard", land on the second it happens.
2144
Daniel van Strien @danielvanstrien.bsky.social · 30/07/2026
New recipe: timestamped video captions on @hf.co Jobs. Point it at a bucket of videos → parquet dataset out: scene descriptions + second-precise <start – end> events. ~$0.05 per hour of footage on a single A10G (Marlin-2B on vLLM). huggingface.co/datasets/uv-...
Image of a hand pointing at a thermometer scale. Caption below shows timestampe from the video the image is captured from with "A hand points to the thermometer scale" as the caption
091
Daniel van Strien @danielvanstrien.bsky.social · 14/07/2026
Small open models are getting genuinely good at document parsing: OvisOCR2 (0.9B, Apache 2.0) is claiming SOTA on OmniDocBench v1.6. Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command. huggingface.co/datasets/uv-scripts/ocr
input image and ocr output side by side
0363
Daniel van Strien @danielvanstrien.bsky.social · 09/07/2026
Open ASR models with speaker diarization are now fast and cheap: I diarized 174 hours of Apollo 11 mission audio (the real July 1969 NASA tapes) for $9.46 with a 0.9B open model. Search it, hear any moment on the original tape: huggingface.co/spaces/davan...
demo screenshot showing search results and play buttons next to transcripts
3251
Daniel van Strien @danielvanstrien.bsky.social · 07/07/2026
I ran 10 newer OCR models on @ai2.bsky.social's olmOCR-bench "old scans" subset. The ranking flips depending on what you actually want.
slopegraph showing flips in ratings
3193
Daniel van Strien @danielvanstrien.bsky.social · 02/07/2026
Coding agents are real users of the @hf.co Hub! They're searching for models, building and pushing datasets, training models on Jobs, spinning up Spaces... Now there's public data: each agent's share of Hub traffic, updated monthly 👇
Horizontal bar chart titled "Agents calling the Hugging Face Hub", showing each coding agent's share of agent-attributed huggingface_hub requests for June 2026. claude-code leads at 23.9%, followed by codex 19.8%, cursor-cli 10.2%, antigravity 2.7%, openclaw 2.3%, hermes-agent 1.4%, then pi, opencode, github-copilot and cursor below 1%. A gray bar at the bottom shows 37.7% of traffic as "unknown" — unregistered tools. Source: hf.co/datasets/huggingface/agent-usage.
2201
Daniel van Strien @danielvanstrien.bsky.social · 02/07/2026
The latest @commoncrawl.bsky.social crawl indexes way more than HTML pages. 20.9M PDFs. Plus calendars, BibTeX, markdown... The whole index now lives in a @hf.co Bucket, so I pulled this with one SQL query straight over it (new S3 API). 2.1B rows, nothing downloaded, $0 to read.
Screenshot of duckdb query and an output table ranking content types in common crawl as page counts and pcts
2246
Daniel van Strien @danielvanstrien.bsky.social · 30/06/2026
You can now use 100s of tools with @hf.co Buckets, thanks to the new S3 API! Usually just one or two lines to change. huggingface.co/docs/hub/sto...
Diff between boto3 client using buckets and S3 showing two new lines for enpoint_url and config
060
Daniel van Strien @danielvanstrien.bsky.social · 29/06/2026
Fairly benchmarking OCR models is hard! Ran a few newer OCR models on @ai2.bsky.social's olmOCR-bench "old scans" subset. Worth knowing: the score "punishes" models for extracting too much (letterheads, stamps, etc.)
1130
Daniel van Strien @danielvanstrien.bsky.social · 24/06/2026
If libraries, archives and museums pooled their (labelled) data, they could build state-of-the-art open models for the things they actually care about! I tried a small version: one open model (NuExtract-3, 4B) fine-tuned to read archival index cards across several collections.
1295
Daniel van Strien @danielvanstrien.bsky.social · 23/06/2026
Ran the same 650M model across 7 languages: German, Polish, Finnish, Russian, Serbian, Estonian, Swedish via @europeana.bsky.social newspapers. Still some challenging layouts, but pretty mind-blowing to me that a 650M model is doing this when a year ago 72B VLMs failed very often.
3306
Daniel van Strien @danielvanstrien.bsky.social · 22/06/2026
The text it pulls out is really clean i.e. reads the columns in the right order, gets the headers, even caught a handwritten note in the margin!
Side by side: on the left, a scanned 1901 newspaper page (The Commoner); on the right, the clean text Surya OCR 2 read from it — the masthead and opening editorial, in reading order.
160
Daniel van Strien @danielvanstrien.bsky.social · 22/06/2026
I think VLM-based OCR might finally be close to working on historic newspapers! Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow. Surya OCR 2 (a 650M model!) does a very good job!
A 1901 newspaper page overlaid with Surya OCR 2's detected layout: colored boxes around each block (text, section-header, picture) and numbered dots joined by a path showing the model's reading order across the columns.
18814
Daniel van Strien @danielvanstrien.bsky.social · 11/06/2026
Can the new DiffusionGemma model help fix broken OCR? In theory, denoising tokens in parallel could work better for OCR correction since context is seen upfront? Pointed it at 19th-century newspaper OCR. It corrected better than the autoregressive baseline — at ~8x the speed.
47015
Daniel van Strien @danielvanstrien.bsky.social · 02/06/2026
What could a rich ecosystem of small GLAM AI models enable? IMO: cheaper, better-fitted, more robust models. Example: I extended an existing @natlibscot.bsky.social archival card detector to 4 collections to make a more generic index card detector. Took an hour or two and minimal $
A scanned archival page showing nine typewritten index cards arranged in a 3×3 grid. Each card has been detected by the model and outlined with a green bounding box, with confidence scores between 0.93 and 0.99.
1141
Daniel van Strien @danielvanstrien.bsky.social · 27/05/2026
Derived datasets are bigger on Hugging Face Hub than people realise. ~73% of analysed datasets on the Hub are derivatives of something else, i.e. cleaned, translated, extended, etc. Built an explorer that infers the missing lineage from content: huggingface.co/spaces/davan...
Screenshot of the overview of the lineage appScreenshot showing the lineage of one dataset with two steps of children.
1203
Daniel van Strien @danielvanstrien.bsky.social · 26/05/2026
Built a useful GLAM model this week: flags blank index cards before expensive OCR. Cheap, scales, adapts. IMO the technical barrier is now very low, the main one is knowing how. So I wrote this up as a chapter: danielvanstrien.xyz/ai-patterns-for-glam/patterns/index-card-classifier.html
Two archival index cards side by side. Left, labelled "blank": a bare card with only a punch hole near the bottom centre. Right, labelled "content (not-blank)": a typed catalogue card with shelf number, book title, date, and a library form stamp.
1175
Daniel van Strien @danielvanstrien.bsky.social · 22/05/2026
You can now run SQL over 2.19 BILLION web pages — zero download. @commoncrawl.bsky.social April 2026 crawl + URL index are on Hugging Face Storage Buckets. DuckDB reads it straight over hf:// — I counted all 2.19B in ~35s. Or point your own agent at it 👇 huggingface.co/spaces/davan...
Language breakdown prompt for AI agentCopy and paste for DuckDB syntax
1529
Daniel van Strien @danielvanstrien.bsky.social · 20/05/2026
NuExtract3 (4B, Apache-2.0) does OCR *and* structured extraction. Point it at a dataset of scanned index cards + a JSON schema → clean catalog JSON One command on huggingface Jobs. (or skip the schema for plain Markdown OCR) Script + dataset 👇
A worked example of NuExtract3 turning a document image into structured data. Top: a terminal command — hf jobs uv run --image vllm/vllm-openai nuextract3.py index-cards cards-json --template schema.json. Middle (input): a scanned typewritten library index card reading "ABAD (Joseph), Captain, Spanish Army, letter of (1783), 5538, f.11." Bottom (output): the JSON extracted from the card — image_type "index_card", heading "Abad J.", heading_type "person", epithet "Captain, Spanish Army", and one entry with ms_no "5538", folios ["f.11"], description "letter of (1783)".
24112
Daniel van Strien @danielvanstrien.bsky.social · 02/04/2026
2 million book titles visualised as an interactive Embedding Atlas. Fiction, History, Science — each forms its own cluster. The trick: mount the same @hf.co Bucket to a GPU Job AND a Space. Job writes, Space reads. No uploads!
Screenshot of an Embedding Atlas visualization showing 2 million Open Library book titles as a scatter plot. Points are colored by subject category — History (orange), Fiction (green), Religion (purple), Law & Politics (red), and others. Clusters are labeled with topic keywords like "Church-Christian-Bible-God", "Shakespeare", "surgery-anatomy-physiology". A tooltip shows a selected book titled "Rosabelle" categorized as Society, published 1995. The right panel shows a category distribution bar chart.
2192
Daniel van Strien @danielvanstrien.bsky.social · 31/03/2026
Segment any object in an image dataset with a text prompt — one command. uv run segment-objects.py data output --class-name deer Pixel-level masks via SAM3. Perfect for agents building their own training data. Runs on @hf.co Jobs. huggingface.co/datasets/uv-...
Night-vision camera trap image of a white-tailed deer in a forest. The deer is highlighted with a red segmentation mask overlay showing pixel-level object boundaries, with a green bounding box and "deer 96%" confidence label. Generated automatically by SAM3 from a text prompt with zero training data.
171
Daniel van Strien @danielvanstrien.bsky.social · 19/03/2026
Is olmOCR-bench getting close to saturation? Top score is now 85.9%. Yesterday, Datalab took #1 with chandra-ocr-2. A year ago, the best was 79. Visualised the race to get there using @hf.co leaderboard data
0101
Daniel van Strien @danielvanstrien.bsky.social · 16/03/2026
One of the nicest things about Nvidia model releases is that they ship the training data. What does it look like? I sampled 250k examples from 24 datasets in the Nemotron post-training v3 collection and built an interactive Embedding Atlas to explore it.
Interactive Embedding Atlas visualization of 250,000 training examples from NVIDIA's Nemotron post-training v3 collection, colored by category. Distinct clusters are visible for Math (blue, 56k), Code (orange, 45k), Agentic (green, 41k), Instruction Following (red, 25k), Finance (purple, 18k), Multilingual (brown, 18k), Science (pink, 18k), Safety (grey, 16k), and Identity (yellow, 8k). A selected point shows an example from the Agentic function-calling dataset: "I need to know the weight of the creature that can evolve into a Fire Dragon."
3377
Daniel van Strien @danielvanstrien.bsky.social · 13/03/2026
The new @hf.co storage Buckets open up the Hub beyond models and datasets. Example: IIIF image hosting. With Buckets, just upload static tiles and any IIIF viewer zooms straight from CDN!
Screenshot of Mirador IIIF viewer showing a panoramic map of Amesville, Athens County Ohio from 1875, served from a Hugging Face Storage Bucket. The map shows detailed bird's-eye view illustrations of buildings, streets, and surrounding farmland.
2143
Daniel van Strien @danielvanstrien.bsky.social · 11/03/2026
74GB of Dutch PDFs, filtered and written back to the Hub - without touching local disk! Hub is your disk! I built a PoC adding sink_parquet for @pola.rs to stream writes to @hf.co's new Storage Buckets via Xet. Constant memory ~18 min on a 2-vCPU machine.
Code screenshot show loading from the hub filtering and pushing to a bucket
1101
Daniel van Strien @danielvanstrien.bsky.social · 05/03/2026
There is no best VLM OCR model - rankings can flip completely by document type. I built ocr-bench: run open OCR models on YOUR documents, get a per-collection leaderboard. VLM-as-judge with Bradley-Terry ELO, all running on @hf.co. No local GPU needed.
Screenshot of plot showing ELO vs paramter count for different OCR models
15211
Daniel van Strien @danielvanstrien.bsky.social · 27/02/2026
Is it worth re-OCR'ing old library index cards? Re-OCR'd 453,000 from @bpl.boston.gov's rare books catalogue. ~$50 compute using @huggingface Jobs BPL's own guide calls their search "extremely unreliable." Does better OCR + semantic search help fix it? Demo space link below
Screenshot of a search UI showing a text box with search results showing index cards next to the ocr for the card
1418
Daniel van Strien @danielvanstrien.bsky.social · 23/02/2026
Ran the same OCR models on 68 pages of historic newspaper. Every model hallucinated or looped. DeepSeek-OCR-2, LightOnOCR-2, GLM-OCR – all melt down on dense newspaper columns. You can try yourself using this @hf.co dataset: huggingface.co/datasets/dav...
Image with historic newspaper on the left and output from OCR models sampled on the right. The output on the right shows okay starting text and then a lot of repetition.
4203
Daniel van Strien @danielvanstrien.bsky.social · 20/02/2026
Llama.cpp joins Hugging Face github.com/ggml-org/lla...
llama.cpp logo + Hugging Face logo
2547
Daniel van Strien @danielvanstrien.bsky.social · 19/02/2026
Re-OCR'd the complete 1771 Encyclopaedia Britannica (2,724 pages) with a single command on @hf.co Jobs. - 0.9B model (GLM-OCR) ~$0.002/page ~$5 total on an L4 GPU Before (old Tesseract ocr) → After
Screenshot of old vs new ocr. 

old ocr text is garbled. New ocr much cleaner.
59616
Daniel van Strien @danielvanstrien.bsky.social · 17/02/2026
The uv-scripts/ocr collection now includes 13 models, including GLM-OCR, a 0.9B model that scores 94.6% on OmniDocBench. One command to run any of them on your dataset via @hf.co Jobs. huggingface.co/datasets/uv-...
table of contents showing ocr models supported in the repo
0101
Daniel van Strien @danielvanstrien.bsky.social · 09/02/2026
Datasets and benchmarks drive AI progress, but finding papers that introduce new ones means digging through thousands of arXiv abstracts. Updated the Dataset Papers on ArXiv app to surface them: 52K+ papers classified as introducing new datasets from 212K CS papers.
191
Daniel van Strien @danielvanstrien.bsky.social · 02/02/2026
Built an object detector from zero-labelled data in one afternoon with help from Claude Code (it can do more than vibe code, TODO apps...) SAM3 on HF Jobs → correct the errors → train YOLO → repeat. Three rounds: 31% → 99% accuracy on historical index cards from @natlibscot.bsky.social
image of a index card with a green bounding box prediction around the card contents
3103
Daniel van Strien @danielvanstrien.bsky.social · 19/12/2025
Built a 2.5MB image classifier that runs in the browser in an evening with Claude Code. I used a dataset I labelled in 2022 and left on @hf.co for 3 years 😬. It finds illustrated pages in historical books. No server. No GPU.
28618
Daniel van Strien @danielvanstrien.bsky.social · 09/12/2025
Just posted my slides from the AI4LAM #FF2025 workshop on open source AI for GLAMs. Probably slides on their own aren't that useful, but they do feature one of my growing collection of libraries-and-AI memes, so there's that danielvanstrien.xyz/slides.html
Here's alt text for the meme:

Alt text: "Flex Tape meme format. Top panel: Phil Swift (labeled 'Library Systems Vendor') aggressively spraying water representing 'Outdated systems, metadata issues, disjointed search and complex user needs.' Bottom panel: A hand slapping Flex Tape underwater, with the tape labeled 'AI-powered chat interface.'"

This captures the joke that vendors are positioning AI chat as a quick fix for deep-seated library infrastructure problems—a bit like slapping tape on a leak rather than fixing the plumbing.
273
Daniel van Strien @danielvanstrien.bsky.social · 21/11/2025
Building datasets to train smaller, task-focused models used to be incredibly time-consuming. Very excited to see SAM3 massively lower that barrier. Describe the class you want to detect and get annotated datasets automatically! Try it yourself: huggingface.co/datasets/uv-...!
Screenshot of a simple app showing bounding boxes for photographs detected in historic newspaper images. hf jobs uv run \
  --flavor a100-large \
  -s HF_TOKEN=HF_TOKEN \
  https://huggingface.co/datasets/uv-scripts/sam3/raw/main/detect-objects.py \
  -- davanstrien/newspapers-with-images-after-photography-big \
  davanstrien/newspapers-photo-predictions \
  --class-name "photograph" \
  --confidence-threshold 0.4
15712
Daniel van Strien @danielvanstrien.bsky.social · 05/11/2025
Very much looking forward to presenting at this tomorrow. I will be making my usual pitch that datasets are the foundational infrastructure for cultural heritage to benefit from and create useful AI models and tools. Be warned, I did fire up the meme generator for my slides...
Meme showing two versions of the doge dog. 
Top: muscular buff doge labeled 'Public Domain circa 2000' saying 'Open access to knowledge and culture is a universal good'.
Bottom: small weak doge labeled 'Public Domain post LLMs?' saying 'Someone might use public domain content for training an LLM' with worried emoji.
0368
Daniel van Strien @danielvanstrien.bsky.social · 23/10/2025
huggingface.co/nanonets/Nan... might be worth a try for this. Can extract formulas into LaTeX
screenshot of latex output
110
Daniel van Strien @danielvanstrien.bsky.social · 22/10/2025
The command (using @hf.co Jobs - serverless GPU compute) Full script at huggingface.co/datasets/uv-...
hf jobs uv run --flavor a100-large --timeout 2h \
    -s HF_TOKEN \
    https://huggingface.co/datasets/uv-scripts/ocr/raw/main/deepseek-ocr-vllm.py \
    NationalLibraryOfScotland/Britain-and-UK-Handbooks-Dataset \
    davanstrien/handbooks-deep-ocr \
    --resolution-mode base \
    --batch-size 2048 \
    --prompt-mode free
020
Daniel van Strien @danielvanstrien.bsky.social · 22/10/2025
DeepSeek-OCR just got vLLM support 🚀 Currently processing @natlibscot.bsky.social's 27,915-page handbook collection with one command. Processing at ~350 images/sec on A100 Using @hf.co Jobs + uv - zero setup batch OCR! Will share final time + cost when done!
Logs showing ocr progress
1182