Sign in

Daniel van Strien

@danielvanstrien.bsky.social
4.7K followers 520 following 415 posts

Machine Learning Librarian at @hf.co

PostsRepliesMedia
Daniel van Strien @danielvanstrien.bsky.social · 23/09/2026
The Europeana Newspapers dataset on @hf.co now has an `alto` config: the raw ALTO XML for all 5.9M pages. The coordinates for every word, line and block, per-word OCR confidence and font info that the flattened text dropped are back. huggingface.co/datasets/biglam/europeana_newspapers
huggingface.co
biglam/europeana_newspapers · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
1164
Daniel van Strien @danielvanstrien.bsky.social · 16/09/2026
Made some improvements to the OCR scripts onboarding in my uv-scripts collection. First run is now a small example: seven scanned NASA pages in, Markdown out, one HF jobs command. It also works on your own images or a folder of PDFs. You can pick from 28 OCR models. huggingface.co/datasets/uv-...
Image of a old scanned typed page on the left and on the right is markdown output
1141
Reposted by Daniel van Strien
Daniel van Strien @danielvanstrien.bsky.social · 11/09/2026
Need a historical illustration? Search 1.49 million images from British Library books and Britannica (1500s–1920s). Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources. huggingface.co/spaces/davan...
47620
Daniel van Strien @danielvanstrien.bsky.social · 11/09/2026
Need a historical illustration? Search 1.49 million images from British Library books and Britannica (1500s–1920s). Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources. huggingface.co/spaces/davan...
47620
Daniel van Strien @danielvanstrien.bsky.social · 10/09/2026
Used Astra + Jobs to see how well this model performs on the full olmOCR-bench Unsurprisingly, it doesn’t do brilliantly overall: 36.8% But for a ~16M-parameter recogniser, I think 74.4% on long/tiny text and 57.9% on multi-column pages are pretty interesting.
190
Reposted by Daniel van Strien
Adina Yakup @adinayakup.bsky.social · 10/09/2026
DeepSeek v4.1 Flash is just another level 🤯 huggingface.co/deepseek-ai/... - Asymmetric Causal-Encoder-Decoder: 550B MoE, input 8B / output 16B - Native vision merged into one endpoint - KV cache crushed: ~1/4 the HBM vs last one, 437× smaller than their first model
huggingface.co
deepseek-ai/DeepSeek-V4.1-Flash · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
2754
Daniel van Strien @danielvanstrien.bsky.social · 07/09/2026
OCR for Japanese manga, Swedish handwriting or Arabic print? There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines. I’ve gathered 41 models into four collections, with short notes to help you choose: huggingface.co/collections/...
huggingface.co
OCR on the Hub - a davanstrien Collection
Curated OCR models for documents, languages, handwriting and text in images. Browse four collections with short practical notes.
0288
Daniel van Strien @danielvanstrien.bsky.social · 03/09/2026
A 15.9M-parameter specialist OCR model outperforms many VLMs hundreds of times larger. On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved. huggingface.co/spaces/fineb...
Plot showing parameters vs performance. Kraken is top left.
2346
Reposted by Daniel van Strien
Tom Aarsen @tomaarsen.com · 26/08/2026
📈 New blog post: Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers. As a practical example, I finetuned a ColBERT-style model for medical retrieval. 14.5 hours on one RTX 3090, and it beats every general-purpose retriever I could find. Thread 🧵
2214
Daniel van Strien @danielvanstrien.bsky.social · 26/08/2026
Every illustration in the Encyclopaedia Britannica dataset now has an instance mask: 411,385 cut-outs from 115,293 pages, 1768–1929, each linked to its full-resolution scan. Public domain, no image generation involved. Work in progress.
The left image shows an illustration surrounded by text. right image shows image cropped via mask prediction
03710
Daniel van Strien @danielvanstrien.bsky.social · 25/08/2026
Uploaded a dataset of 115,293 illustrated pages from the Encyclopaedia Britannica, 1st edition (1768) to 14th (1929) to the Hub huggingface.co/datasets/big...
grid showing examples from the dataset
0335
Daniel van Strien @danielvanstrien.bsky.social · 20/08/2026
Every illustration in the British Library's 19th-century book images dataset now has an instance mask: 1.02M clean cutouts of public-domain engravings, no image generation involved! One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.
111425
Daniel van Strien @danielvanstrien.bsky.social · 18/08/2026
Synthetic data at scale without owning a GPU: datatrove's new Jobs backend + Qwen3.8-27B → 35,837 length-controllable TL;DRs of Hugging Face cards, $0.43 per 1,000. Full guide: danielvanstrien.xyz/posts/2026/d...
danielvanstrien.xyz
Distilling Qwen3.8 with datatrove on Hugging Face Jobs – Daniel van Strien
datatrove’s new Jobs backend plus a days-old 27B teacher: regenerating a 35,837-summary training dataset in one afternoon for $15.58, with a calibration-first workflow and length-controllable outputs.
051
Daniel van Strien @danielvanstrien.bsky.social · 10/08/2026
FineBooks wants to help re-OCR the public domain to unlock a rich source of data for AI, for researchers, and for enjoyment. Step one: work out which modern OCR models are actually good enough.
Leaderboard preview
2255
Reposted by Daniel van Strien
Sebastian Majstorovic @storytracer.com · 10/08/2026
Are open OCR models good enough to unlock historical knowledge? @danielvanstrien.bsky.social and I built a leaderboard using six expert-transcribed volumes from the @biodivlibrary.bsky.social sky.social. Meet FineBooks, from @hf.co and @eleutherai.bsky.social! huggingface.co/blog/fineboo...
huggingface.co
FineBooks: are open OCR models good enough to unlock historical knowledge?
A Blog post by FineBooks on Hugging Face
06322
Daniel van Strien @danielvanstrien.bsky.social · 07/08/2026
Uploaded a dataset of 1,080,814 public domain images, mostly from 19th-century books, to the Hugging Face Hub: huggingface.co/datasets/big... You can also do semantic search against the images here: huggingface.co/spaces/davan...
15412
Daniel van Strien @danielvanstrien.bsky.social · 03/08/2026
Public data on which coding agents actually use the Hugging Face Hub, refreshed monthly. Claude Code now sends a majority of everything the Hub can attribute to a named agent. In May, it led on 4 days out of 30. In July, all 30. huggingface.co/datasets/hug...
plot showing growth of different coding agents using the hub
4132
Daniel van Strien @danielvanstrien.bsky.social · 31/07/2026
A film catalogue tells you what a film is about, not what happens inside it. So: 1,864 public-domain Prelinger films, described every ~60 seconds by an open 2B video model. 23,148 searchable moments. Search "typing on a computer keyboard", land on the second it happens.
2144
Daniel van Strien @danielvanstrien.bsky.social · 30/07/2026
New recipe: timestamped video captions on @hf.co Jobs. Point it at a bucket of videos → parquet dataset out: scene descriptions + second-precise <start – end> events. ~$0.05 per hour of footage on a single A10G (Marlin-2B on vLLM). huggingface.co/datasets/uv-...
Image of a hand pointing at a thermometer scale. Caption below shows timestampe from the video the image is captured from with "A hand points to the thermometer scale" as the caption
091
Reposted by Daniel van Strien
Daniel van Strien @danielvanstrien.bsky.social · 14/07/2026
Small open models are getting genuinely good at document parsing: OvisOCR2 (0.9B, Apache 2.0) is claiming SOTA on OmniDocBench v1.6. Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command. huggingface.co/datasets/uv-scripts/ocr
input image and ocr output side by side
0363
Daniel van Strien @danielvanstrien.bsky.social · 14/07/2026
Small open models are getting genuinely good at document parsing: OvisOCR2 (0.9B, Apache 2.0) is claiming SOTA on OmniDocBench v1.6. Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command. huggingface.co/datasets/uv-scripts/ocr
input image and ocr output side by side
0363
Reposted by Daniel van Strien
Jim Clifford @jimclifford.bsky.social · 10/07/2026
Reading the Archive by Machine: An OCR Benchmark for Historians, 1612–1921 Here is version 1 of a working paper on the new OCR tools that are transforming digital history. working-papers-in-critical-search.github.io/paper-004-oc...
working-papers-in-critical-search.github.io
Reading the Archive by Machine – Working Papers in Critical Search
A benchmark of six OCR systems (Tesseract, olmOCR 2, Chandra 2, Infinity Parser 2, GLM-OCR, and Gemini 3.5 Flash) on human-transcribed archival documents spanning 1612–1921: early-modern print, ninete...
24114
Reposted by Daniel van Strien
James Feigenbaum @jamesfeigenbaum.bsky.social · 09/07/2026
I don't have it yet, but there's an econ history paper in these transcripts or transcripts like them (that we can now turn into textual data at the cost of pennies per hour)
1133
Daniel van Strien @danielvanstrien.bsky.social · 09/07/2026
Open ASR models with speaker diarization are now fast and cheap: I diarized 174 hours of Apollo 11 mission audio (the real July 1969 NASA tapes) for $9.46 with a 0.9B open model. Search it, hear any moment on the original tape: huggingface.co/spaces/davan...
demo screenshot showing search results and play buttons next to transcripts
3251
Daniel van Strien @danielvanstrien.bsky.social · 07/07/2026
I ran 10 newer OCR models on @ai2.bsky.social's olmOCR-bench "old scans" subset. The ranking flips depending on what you actually want.
slopegraph showing flips in ratings
3193
Daniel van Strien @danielvanstrien.bsky.social · 02/07/2026
Coding agents are real users of the @hf.co Hub! They're searching for models, building and pushing datasets, training models on Jobs, spinning up Spaces... Now there's public data: each agent's share of Hub traffic, updated monthly 👇
Horizontal bar chart titled "Agents calling the Hugging Face Hub", showing each coding agent's share of agent-attributed huggingface_hub requests for June 2026. claude-code leads at 23.9%, followed by codex 19.8%, cursor-cli 10.2%, antigravity 2.7%, openclaw 2.3%, hermes-agent 1.4%, then pi, opencode, github-copilot and cursor below 1%. A gray bar at the bottom shows 37.7% of traffic as "unknown" — unregistered tools. Source: hf.co/datasets/huggingface/agent-usage.
2201
Daniel van Strien @danielvanstrien.bsky.social · 02/07/2026
The latest @commoncrawl.bsky.social crawl indexes way more than HTML pages. 20.9M PDFs. Plus calendars, BibTeX, markdown... The whole index now lives in a @hf.co Bucket, so I pulled this with one SQL query straight over it (new S3 API). 2.1B rows, nothing downloaded, $0 to read.
Screenshot of duckdb query and an output table ranking content types in common crawl as page counts and pcts
2246
Daniel van Strien @danielvanstrien.bsky.social · 30/06/2026
You can now use 100s of tools with @hf.co Buckets, thanks to the new S3 API! Usually just one or two lines to change. huggingface.co/docs/hub/sto...
Diff between boto3 client using buckets and S3 showing two new lines for enpoint_url and config
060
Daniel van Strien @danielvanstrien.bsky.social · 29/06/2026
Fairly benchmarking OCR models is hard! Ran a few newer OCR models on @ai2.bsky.social's olmOCR-bench "old scans" subset. Worth knowing: the score "punishes" models for extracting too much (letterheads, stamps, etc.)
1130
Reposted by Daniel van Strien
Daniel van Strien @danielvanstrien.bsky.social · 24/06/2026
If libraries, archives and museums pooled their (labelled) data, they could build state-of-the-art open models for the things they actually care about! I tried a small version: one open model (NuExtract-3, 4B) fine-tuned to read archival index cards across several collections.
1295
Reposted by Daniel van Strien
Melanie Walsh @mellymeldubs.bsky.social · 24/06/2026
Excited to share this. @neel2112.bsky.social, @mariaa.bsky.social, and I analyzed 500K anonymous ChatGPT convos (shared w/ consent from WildChat) to see if people were generating fiction. We found tons of stories, fanfiction & erotica. Many users iterated on the same stories for days and weeks.
Screenshot of paper abstract that reads: 

AI FICTION IN THE WILD Neel Gupta  Maria Antoniak  Melanie Walsh

Some professional authors are beginning to use AI tools to help produce their fiction writing. Are readers using AI to generate fiction, too? Drawing on over 500,000 anonymized, English-language ChatGPT-user conversations (Zhao et al.), we find that more than one third of the conversations involve some form of fiction generation—including original stories, roleplay, fanfiction, and erotica. This AI-generated fiction is notably dominated by power users. We identify common fiction generation patterns and profiles among these users, including what we call infinite story demanders, who repeatedly request and revise variations of the same or similar narratives over extended periods of time. We show that users especially gravitate toward fanfiction and erotica, and that they are broadly drawn to generic forms, repetition, immediacy, and niche combinations of story elements. Our findings motivate two theoretical provocations. First, we argue that AI technologies may lead to a shift in the conventional relationship between the author and reader, potentially producing what we call a solipsistic reader-writer, who both generates and consumes fiction within a closed conversational loop, interacting with a machine rather than a human other. Second, we note that LLMs enable interactivity, play, and permutation in ways that are seemingly pleasurable for users, raising questions about where AI will fit into contemporary storytelling and entertainment ecosystems. We situate these developments within broader transformations in literature and media, including self-publishing, fanfiction, and pornography, and suggest that AI-generated fiction shares structural affinities with on-demand, personalized, and repetitive cultural forms.
515645
Reposted by Daniel van Strien
Josh Hadro @hadro.bsky.social · 24/06/2026
Say, for example, if we had 3 examples of labeled city directory data from every state in the US, we could easily train a generic city directory parser! (I’m actually working on this, and I’ll have it done relatively soon, but @danielvanstrien.bsky.social’s point holds in so many other cases!)
191
Daniel van Strien @danielvanstrien.bsky.social · 24/06/2026
If libraries, archives and museums pooled their (labelled) data, they could build state-of-the-art open models for the things they actually care about! I tried a small version: one open model (NuExtract-3, 4B) fine-tuned to read archival index cards across several collections.
1295
Daniel van Strien @danielvanstrien.bsky.social · 23/06/2026
Ran the same 650M model across 7 languages: German, Polish, Finnish, Russian, Serbian, Estonian, Swedish via @europeana.bsky.social newspapers. Still some challenging layouts, but pretty mind-blowing to me that a 650M model is doing this when a year ago 72B VLMs failed very often.
3306
Daniel van Strien @danielvanstrien.bsky.social · 22/06/2026
I think VLM-based OCR might finally be close to working on historic newspapers! Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow. Surya OCR 2 (a 650M model!) does a very good job!
A 1901 newspaper page overlaid with Surya OCR 2's detected layout: colored boxes around each block (text, section-header, picture) and numbered dots joined by a path showing the model's reading order across the columns.
18814
Reposted by Daniel van Strien
Maria Antoniak @mariaa.bsky.social · 17/06/2026
I've been offline more than usual because of paper and grant deadlines plus work travel, but I'm back with my first real newsletter today! Click for rambling about infrastructure, benchmarks, Babylonian tablets, and more.
field-notes.leaflet.pub
Dispatch: Humanistic AI, OCR, and Hugging Face
In which I reflect on the a recent trip to Chicago, overview the state of OCR for humanities data, and fangirl over Hugging Face infrastructure.
1389
Reposted by Daniel van Strien
Daniel van Strien @danielvanstrien.bsky.social · 11/06/2026
Can the new DiffusionGemma model help fix broken OCR? In theory, denoising tokens in parallel could work better for OCR correction since context is seen upfront? Pointed it at 19th-century newspaper OCR. It corrected better than the autoregressive baseline — at ~8x the speed.
47015
Daniel van Strien @danielvanstrien.bsky.social · 11/06/2026
Can the new DiffusionGemma model help fix broken OCR? In theory, denoising tokens in parallel could work better for OCR correction since context is seen upfront? Pointed it at 19th-century newspaper OCR. It corrected better than the autoregressive baseline — at ~8x the speed.
47015
Reposted by Daniel van Strien
Daniel van Strien @danielvanstrien.bsky.social · 10/06/2026
Got a digitised collection that needs OCR? uv-scripts is a set of single-file Python scripts that OCR a whole image dataset to markdown in one command — 20+ open VLMs to pick from, nothing to install but uv. github.com/davanstrien/...
13813
Daniel van Strien @danielvanstrien.bsky.social · 10/06/2026
Got a digitised collection that needs OCR? uv-scripts is a set of single-file Python scripts that OCR a whole image dataset to markdown in one command — 20+ open VLMs to pick from, nothing to install but uv. github.com/davanstrien/...
13813
Daniel van Strien @danielvanstrien.bsky.social · 02/06/2026
What could a rich ecosystem of small GLAM AI models enable? IMO: cheaper, better-fitted, more robust models. Example: I extended an existing @natlibscot.bsky.social archival card detector to 4 collections to make a more generic index card detector. Took an hour or two and minimal $
A scanned archival page showing nine typewritten index cards arranged in a 3×3 grid. Each card has been detected by the model and outlined with a green bounding box, with confidence scores between 0.93 and 0.99.
1141
Daniel van Strien @danielvanstrien.bsky.social · 27/05/2026
Derived datasets are bigger on Hugging Face Hub than people realise. ~73% of analysed datasets on the Hub are derivatives of something else, i.e. cleaned, translated, extended, etc. Built an explorer that infers the missing lineage from content: huggingface.co/spaces/davan...
Screenshot of the overview of the lineage appScreenshot showing the lineage of one dataset with two steps of children.
1203
Daniel van Strien @danielvanstrien.bsky.social · 26/05/2026
Built a useful GLAM model this week: flags blank index cards before expensive OCR. Cheap, scales, adapts. IMO the technical barrier is now very low, the main one is knowing how. So I wrote this up as a chapter: danielvanstrien.xyz/ai-patterns-for-glam/patterns/index-card-classifier.html
Two archival index cards side by side. Left, labelled "blank": a bare card with only a punch hole near the bottom centre. Right, labelled "content (not-blank)": a typed catalogue card with shelf number, book title, date, and a library form stamp.
1175
Daniel van Strien @danielvanstrien.bsky.social · 22/05/2026
You can now run SQL over 2.19 BILLION web pages — zero download. @commoncrawl.bsky.social April 2026 crawl + URL index are on Hugging Face Storage Buckets. DuckDB reads it straight over hf:// — I counted all 2.19B in ~35s. Or point your own agent at it 👇 huggingface.co/spaces/davan...
Language breakdown prompt for AI agentCopy and paste for DuckDB syntax
1529
Daniel van Strien @danielvanstrien.bsky.social · 21/05/2026
The 'why' behind yesterday's demo: libraries don't (just) want searchable text from catalogue cards. They want structured records they can ingest into existing systems. Short blog + a live demo for this library use case: danielvanstrien.xyz/posts/2026/s...
danielvanstrien.xyz
How to turn catalogue card images into structured JSON with a 4B open model – Daniel van Strien
Re-OCR made digitised catalogue cards searchable as text. But libraries need structured records they can ingest into their catalogues. NuExtract3 (4B, Apache-2.0) extracts schema-shaped JSON from card...
092
Daniel van Strien @danielvanstrien.bsky.social · 20/05/2026
NuExtract3 (4B, Apache-2.0) does OCR *and* structured extraction. Point it at a dataset of scanned index cards + a JSON schema → clean catalog JSON One command on huggingface Jobs. (or skip the schema for plain Markdown OCR) Script + dataset 👇
A worked example of NuExtract3 turning a document image into structured data. Top: a terminal command — hf jobs uv run --image vllm/vllm-openai nuextract3.py index-cards cards-json --template schema.json. Middle (input): a scanned typewritten library index card reading "ABAD (Joseph), Captain, Spanish Army, letter of (1783), 5538, f.11." Bottom (output): the JSON extracted from the card — image_type "index_card", heading "Abad J.", heading_type "person", epithet "Captain, Spanish Army", and one entry with ms_no "5538", folios ["f.11"], description "letter of (1783)".
24112
Daniel van Strien @danielvanstrien.bsky.social · 18/05/2026
What is "AI for libraries" beyond a catalogue chatbot? IMO: design patterns (OCR, extraction, classification, search), and agents that both run them and develop the small models behind them. As part of work with @natlibscot.bsky.social started a book on this: danielvanstrien.xyz/ai-patterns-...
danielvanstrien.xyz
AI Design Patterns for Information Professionals
65114
Reposted by Daniel van Strien
Programming Historian @proghist.bsky.social · 13/05/2026
Nouvelle leçon ! 💫 Une nouvelle traduction de la leçon de @danielvanstrien.bsky.social, Kaspar Beelen, @melvinwevers.bsky.social, @thomassmits.bsky.social, et @kmcdono.bsky.social. doi.org/10.46430/phfr0041
doi.org
La vision par ordinateur pour les sciences humaines : introduction à la classification d'images par apprentissage profond (partie 1)
Ceci est le premier volet d'une leçon en deux parties qui présente des méthodes de vision par ordinateur basées sur l'apprentissage profond pour la recherche en sciences humaines. À partir d'un ...
165
Reposted by Daniel van Strien
Daniel van Strien @danielvanstrien.bsky.social · 01/05/2026
Can an open-weight coding agent match Claude Code at training a domain-specific model? Same prompt. ~13 min each. Pushed to Hugging Face. Pi + Kimi K2.6 vs Claude Code + Opus 4.7. Task: classify NC session laws (1866-1967) as Jim Crow or not. Full write-up: danielvanstrien.xyz/posts/2026/a...
danielvanstrien.xyz
Two agents, one prompt – Daniel van Strien
I gave two coding agents the same one-line prompt and watched what each decided to do. Both shipped working classifiers via hf jobs. The headline metrics were within 1.5pp — the interesting difference...
2183
Daniel van Strien @danielvanstrien.bsky.social · 01/05/2026
Can an open-weight coding agent match Claude Code at training a domain-specific model? Same prompt. ~13 min each. Pushed to Hugging Face. Pi + Kimi K2.6 vs Claude Code + Opus 4.7. Task: classify NC session laws (1866-1967) as Jim Crow or not. Full write-up: danielvanstrien.xyz/posts/2026/a...
danielvanstrien.xyz
Two agents, one prompt – Daniel van Strien
I gave two coding agents the same one-line prompt and watched what each decided to do. Both shipped working classifiers via hf jobs. The headline metrics were within 1.5pp — the interesting difference...
2183