Sign in

Daniel van Strien

@danielvanstrien.bsky.social
4.8K followers 520 following 417 posts

Machine Learning Librarian at @hf.co

PostsRepliesMedia
Daniel van Strien @danielvanstrien.bsky.social · 30/09/2026
Newspapers from @europeana.bsky.social
070
Daniel van Strien @danielvanstrien.bsky.social · 30/09/2026
Uploaded a dataset of 98,877 historical newspaper pages (1700s–1940s) to the Hub, each with its original OCR text, word boxes and confidence scores. huggingface.co/datasets/big...
Sample grid of newspapers in the dataset
24716
Daniel van Strien @danielvanstrien.bsky.social · 23/09/2026
The Europeana Newspapers dataset on @hf.co now has an `alto` config: the raw ALTO XML for all 5.9M pages. The coordinates for every word, line and block, per-word OCR confidence and font info that the flattened text dropped are back. huggingface.co/datasets/biglam/europeana_newspapers
huggingface.co
biglam/europeana_newspapers · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1164
Daniel van Strien @danielvanstrien.bsky.social · 16/09/2026
Made some improvements to the OCR scripts onboarding in my uv-scripts collection. First run is now a small example: seven scanned NASA pages in, Markdown out, one HF jobs command. It also works on your own images or a folder of PDFs. You can pick from 28 OCR models. huggingface.co/datasets/uv-...
Image of a old scanned typed page on the left and on the right is markdown output
1141
Reposted by Daniel van Strien
Daniel van Strien @danielvanstrien.bsky.social · 11/09/2026
Need a historical illustration? Search 1.49 million images from British Library books and Britannica (1500s–1920s). Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources. huggingface.co/spaces/davan...
47620
Daniel van Strien @danielvanstrien.bsky.social · 11/09/2026
Need a historical illustration? Search 1.49 million images from British Library books and Britannica (1500s–1920s). Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources. huggingface.co/spaces/davan...
47620
Daniel van Strien @danielvanstrien.bsky.social · 10/09/2026
This includes Kraken’s BLLA segmentation, with raw text output and no table/LaTeX reconstruction. Full results + scripts to reproduce the run: github.com/davanstrien/...
github.com
ocr-bench/experiments/kraken-olmocr-bench at main · davanstrien/ocr-bench
Per-collection OCR leaderboards using VLM-as-judge - davanstrien/ocr-bench
040
Daniel van Strien @danielvanstrien.bsky.social · 10/09/2026
Used Astra + Jobs to see how well this model performs on the full olmOCR-bench Unsurprisingly, it doesn’t do brilliantly overall: 36.8% But for a ~16M-parameter recogniser, I think 74.4% on long/tiny text and 57.9% on multi-column pages are pretty interesting.
190
Reposted by Daniel van Strien
Adina Yakup @adinayakup.bsky.social · 10/09/2026
DeepSeek v4.1 Flash is just another level 🤯 huggingface.co/deepseek-ai/... - Asymmetric Causal-Encoder-Decoder: 550B MoE, input 8B / output 16B - Native vision merged into one endpoint - KV cache crushed: ~1/4 the HBM vs last one, 437× smaller than their first model
huggingface.co
deepseek-ai/DeepSeek-V4.1-Flash · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
2754
Daniel van Strien @danielvanstrien.bsky.social · 07/09/2026
OCR for Japanese manga, Swedish handwriting or Arabic print? There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines. I’ve gathered 41 models into four collections, with short notes to help you choose: huggingface.co/collections/...
huggingface.co
OCR on the Hub - a davanstrien Collection
Curated OCR models for documents, languages, handwriting and text in images. Browse four collections with short practical notes.
0288
Daniel van Strien @danielvanstrien.bsky.social · 03/09/2026
Fun fact: I had to adjust the plotting logic to get it to show properly because the model does so well! Kraken OCR library: kraken.re/main/index.h... Model: huggingface.co/small-models...
kraken.re
kraken — kraken documentation
030
Daniel van Strien @danielvanstrien.bsky.social · 03/09/2026
A 15.9M-parameter specialist OCR model outperforms many VLMs hundreds of times larger. On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved. huggingface.co/spaces/fineb...
Plot showing parameters vs performance. Kraken is top left.
2346
Reposted by Daniel van Strien
Tom Aarsen @tomaarsen.com · 26/08/2026
📈 New blog post: Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers. As a practical example, I finetuned a ColBERT-style model for medical retrieval. 14.5 hours on one RTX 3090, and it beats every general-purpose retriever I could find. Thread 🧵
2214
Daniel van Strien @danielvanstrien.bsky.social · 26/08/2026
Every illustration in the Encyclopaedia Britannica dataset now has an instance mask: 411,385 cut-outs from 115,293 pages, 1768–1929, each linked to its full-resolution scan. Public domain, no image generation involved. Work in progress.
The left image shows an illustration surrounded by text. right image shows image cropped via mask prediction
03710
Daniel van Strien @danielvanstrien.bsky.social · 25/08/2026
Uploaded a dataset of 115,293 illustrated pages from the Encyclopaedia Britannica, 1st edition (1768) to 14th (1929) to the Hub huggingface.co/datasets/big...
grid showing examples from the dataset
0335
Daniel van Strien @danielvanstrien.bsky.social · 20/08/2026
Dataset (crop_masks config): huggingface.co/datasets/big... Search + cutouts: huggingface.co/spaces/davan... Model: huggingface.co/small-models...
huggingface.co
biglam/british-library-book-images · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
0131
Daniel van Strien @danielvanstrien.bsky.social · 20/08/2026
Every illustration in the British Library's 19th-century book images dataset now has an instance mask: 1.02M clean cutouts of public-domain engravings, no image generation involved! One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.
111425
Daniel van Strien @danielvanstrien.bsky.social · 18/08/2026
Synthetic data at scale without owning a GPU: datatrove's new Jobs backend + Qwen3.8-27B → 35,837 length-controllable TL;DRs of Hugging Face cards, $0.43 per 1,000. Full guide: danielvanstrien.xyz/posts/2026/d...
danielvanstrien.xyz
Distilling Qwen3.8 with datatrove on Hugging Face Jobs – Daniel van Strien
datatrove’s new Jobs backend plus a days-old 27B teacher: regenerating a 35,837-summary training dataset in one afternoon for $15.58, with a calibration-first workflow and length-controllable outputs.
051
Daniel van Strien @danielvanstrien.bsky.social · 14/08/2026
Hopefully lessons have been learned!
000
Daniel van Strien @danielvanstrien.bsky.social · 10/08/2026
For this project, we want to potentially scale to a very large number of pages and also want the approach to be somewhat reproducible, so we're not really considering closed models. My guess is Gemini would be fairly highly ranking on this leaderboard.
010
Daniel van Strien @danielvanstrien.bsky.social · 10/08/2026
A new @hf.co and @eleutherai.bsky.social project, with @storytracer.com! Blog post: huggingface.co/blog/fineboo...
huggingface.co
FineBooks: are open OCR models good enough to unlock historical knowledge?
A Blog post by FineBooks on Hugging Face
142
Daniel van Strien @danielvanstrien.bsky.social · 10/08/2026
So we built a leaderboard. 14 open models, scored on 2,165 pages of expert human transcription. The best read historical print at ~97.6% character accuracy, for under $2 per 1,000 pages.
150
Daniel van Strien @danielvanstrien.bsky.social · 10/08/2026
FineBooks wants to help re-OCR the public domain to unlock a rich source of data for AI, for researchers, and for enjoyment. Step one: work out which modern OCR models are actually good enough.
Leaderboard preview
2255
Reposted by Daniel van Strien
Sebastian Majstorovic @storytracer.com · 10/08/2026
Are open OCR models good enough to unlock historical knowledge? @danielvanstrien.bsky.social and I built a leaderboard using six expert-transcribed volumes from the @biodivlibrary.bsky.social sky.social. Meet FineBooks, from @hf.co and @eleutherai.bsky.social! huggingface.co/blog/fineboo...
huggingface.co
FineBooks: are open OCR models good enough to unlock historical knowledge?
A Blog post by FineBooks on Hugging Face
06322
Daniel van Strien @danielvanstrien.bsky.social · 07/08/2026
Uploaded a dataset of 1,080,814 public domain images, mostly from 19th-century books, to the Hugging Face Hub: huggingface.co/datasets/big... You can also do semantic search against the images here: huggingface.co/spaces/davan...
15412
Daniel van Strien @danielvanstrien.bsky.social · 03/08/2026
Shares are relative, so a falling share doesn't mean falling usage; the attributed base is growing. 23% of agent traffic still comes from tools that haven't registered, down from 37% in June. Curious whether the open harnesses close the gap in August!
040
Daniel van Strien @danielvanstrien.bsky.social · 03/08/2026
Public data on which coding agents actually use the Hugging Face Hub, refreshed monthly. Claude Code now sends a majority of everything the Hub can attribute to a named agent. In May, it led on 4 days out of 30. In July, all 30. huggingface.co/datasets/hug...
plot showing growth of different coding agents using the hub
4132
Daniel van Strien @danielvanstrien.bsky.social · 31/07/2026
The model describes what's on screen; it can't tell you who anyone is or why a film was made. That part stays human. ~$10 in compute on @hf.co Jobs Unreviewed output: good for finding, not citing. huggingface.co/spaces/davan...
huggingface.co
Prelinger moments - a Hugging Face Space by davanstrien
Timestamped search over VLM-captioned Prelinger films
191
Daniel van Strien @danielvanstrien.bsky.social · 31/07/2026
A film catalogue tells you what a film is about, not what happens inside it. So: 1,864 public-domain Prelinger films, described every ~60 seconds by an open 2B video model. 23,148 searchable moments. Search "typing on a computer keyboard", land on the second it happens.
2144
Daniel van Strien @danielvanstrien.bsky.social · 30/07/2026
New recipe: timestamped video captions on @hf.co Jobs. Point it at a bucket of videos → parquet dataset out: scene descriptions + second-precise <start – end> events. ~$0.05 per hour of footage on a single A10G (Marlin-2B on vLLM). huggingface.co/datasets/uv-...
Image of a hand pointing at a thermometer scale. Caption below shows timestampe from the video the image is captured from with "A hand points to the thermometer scale" as the caption
091
Reposted by Daniel van Strien
Daniel van Strien @danielvanstrien.bsky.social · 14/07/2026
Small open models are getting genuinely good at document parsing: OvisOCR2 (0.9B, Apache 2.0) is claiming SOTA on OmniDocBench v1.6. Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command. huggingface.co/datasets/uv-scripts/ocr
input image and ocr output side by side
0363
Daniel van Strien @danielvanstrien.bsky.social · 14/07/2026
Small open models are getting genuinely good at document parsing: OvisOCR2 (0.9B, Apache 2.0) is claiming SOTA on OmniDocBench v1.6. Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command. huggingface.co/datasets/uv-scripts/ocr
input image and ocr output side by side
0363
Daniel van Strien @danielvanstrien.bsky.social · 11/07/2026
This model does it e2e so nothing for me to choose!
040
Reposted by Daniel van Strien
Jim Clifford @jimclifford.bsky.social · 10/07/2026
Reading the Archive by Machine: An OCR Benchmark for Historians, 1612–1921 Here is version 1 of a working paper on the new OCR tools that are transforming digital history. working-papers-in-critical-search.github.io/paper-004-oc...
working-papers-in-critical-search.github.io
Reading the Archive by Machine – Working Papers in Critical Search
A benchmark of six OCR systems (Tesseract, olmOCR 2, Chandra 2, Infinity Parser 2, GLM-OCR, and Gemini 3.5 Flash) on human-transcribed archival documents spanning 1612–1921: early-modern print, ninete...
24114
Reposted by Daniel van Strien
James Feigenbaum @jamesfeigenbaum.bsky.social · 09/07/2026
I don't have it yet, but there's an econ history paper in these transcripts or transcripts like them (that we can now turn into textual data at the cost of pennies per hour)
1133
Daniel van Strien @danielvanstrien.bsky.social · 09/07/2026
It's unedited machine output over scratchy radio (it keeps hearing "Follow eleven, this is Houston"). The recipe runs on any audio collection in one command: huggingface.co/datasets/uv-...
huggingface.co
uv-scripts/transcription · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
020
Daniel van Strien @danielvanstrien.bsky.social · 09/07/2026
The transcript is a finding aid, not a replacement — every segment links back to the Internet Archive originals. And the pipeline flagged which reels are blank carrier hiss: useful collection metadata in its own right. Dataset (CC0): huggingface.co/datasets/dav...
huggingface.co
davanstrien/apollo-11-diarized · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
110
Daniel van Strien @danielvanstrien.bsky.social · 09/07/2026
Open ASR models with speaker diarization are now fast and cheap: I diarized 174 hours of Apollo 11 mission audio (the real July 1969 NASA tapes) for $9.46 with a 0.9B open model. Search it, hear any moment on the original tape: huggingface.co/spaces/davan...
demo screenshot showing search results and play buttons next to transcripts
3251
Daniel van Strien @danielvanstrien.bsky.social · 07/07/2026
Write-up + plus how I ran the whole thing on HF Jobs, no local GPU: danielvanstrien.xyz/posts/2026/o...
danielvanstrien.xyz
Benchmarking OCR fairly is harder than it looks – Daniel van Strien
I ran ten current OCR models on olmOCR-bench’s ‘old scans’ subset. The ranking flips depending on one thing: whether you want a model that keeps what’s on the page, or one that cleans it up. A note on...
072
Daniel van Strien @danielvanstrien.bsky.social · 07/07/2026
Two other things: a 1B model (LightOnOCR-2) has the best raw transcription in the field, and PaddleOCR-VL 1.6 sometimes hallucinates Chinese characters on English scans.
130
Daniel van Strien @danielvanstrien.bsky.social · 07/07/2026
IMO this is because a lot of VLM-based OCR models were made to provide tokens for training. It's less useful if you want faithful OCR of the whole page, like an archive where the letterhead is part of the record.
220
Daniel van Strien @danielvanstrien.bsky.social · 07/07/2026
The score rewards dropping boilerplate, i.e. letterheads, stamps, page numbers, so a model that reads the page more faithfully can rank lower.
120
Daniel van Strien @danielvanstrien.bsky.social · 07/07/2026
On the headline score, PaddleOCR-VL beats NuExtract3 (38.6 vs 37.8). But rank by how much of the page each model actually reads, and NuExtract3 is well ahead (41.6 vs 31.2). Same two models, opposite order.
110
Daniel van Strien @danielvanstrien.bsky.social · 07/07/2026
I ran 10 newer OCR models on @ai2.bsky.social's olmOCR-bench "old scans" subset. The ranking flips depending on what you actually want.
slopegraph showing flips in ratings
3193
Daniel van Strien @danielvanstrien.bsky.social · 02/07/2026
Explore it in the Dataset Viewer, no code needed: - the monthly leaderboard + how it shifts as new tools launch - request share vs user share (a few heavy pipelines vs many light users) - daily data: launch spikes, weekday vs weekend patterns... huggingface.co/datasets/hug...
huggingface.co
huggingface/agent-usage · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
010
Daniel van Strien @danielvanstrien.bsky.social · 02/07/2026
Coding agents are real users of the @hf.co Hub! They're searching for models, building and pushing datasets, training models on Jobs, spinning up Spaces... Now there's public data: each agent's share of Hub traffic, updated monthly 👇
Horizontal bar chart titled "Agents calling the Hugging Face Hub", showing each coding agent's share of agent-attributed huggingface_hub requests for June 2026. claude-code leads at 23.9%, followed by codex 19.8%, cursor-cli 10.2%, antigravity 2.7%, openclaw 2.3%, hermes-agent 1.4%, then pi, opencode, github-copilot and cursor below 1%. A gray bar at the bottom shows 37.7% of traffic as "unknown" — unregistered tools. Source: hf.co/datasets/huggingface/agent-usage.
2201
Daniel van Strien @danielvanstrien.bsky.social · 02/07/2026
Can run it locally. Just streams the data from HF so you don't need the space to store common crawl locally!
110
Daniel van Strien @danielvanstrien.bsky.social · 02/07/2026
The latest @commoncrawl.bsky.social crawl indexes way more than HTML pages. 20.9M PDFs. Plus calendars, BibTeX, markdown... The whole index now lives in a @hf.co Bucket, so I pulled this with one SQL query straight over it (new S3 API). 2.1B rows, nothing downloaded, $0 to read.
Screenshot of duckdb query and an output table ranking content types in common crawl as page counts and pcts
2246
Daniel van Strien @danielvanstrien.bsky.social · 30/06/2026
You can now use 100s of tools with @hf.co Buckets, thanks to the new S3 API! Usually just one or two lines to change. huggingface.co/docs/hub/sto...
Diff between boto3 client using buckets and S3 showing two new lines for enpoint_url and config
060
Daniel van Strien @danielvanstrien.bsky.social · 29/06/2026
A model that reads the page more faithfully can rank lower. Works well as a score if you want clean "reading-order" tokens for LLM training, but less useful as a metric if you care about is "faithful" OCR.
010