Sign in

Sebastian Majstorovic

@storytracer.com
853 followers 241 following 64 posts

Managing Director of datovis.com. Open Data Consultant for @eleutherai.bsky.social. Technical Director of @datarescueproject.org. Personal website: storytracer.com.

PostsRepliesMedia
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 16/09/2026
Made some improvements to the OCR scripts onboarding in my uv-scripts collection. First run is now a small example: seven scanned NASA pages in, Markdown out, one HF jobs command. It also works on your own images or a folder of PDFs. You can pick from 28 OCR models. huggingface.co/datasets/uv-...
Image of a old scanned typed page on the left and on the right is markdown output
1141
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 11/09/2026
Need a historical illustration? Search 1.49 million images from British Library books and Britannica (1500s–1920s). Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources. huggingface.co/spaces/davan...
47620
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 07/09/2026
OCR for Japanese manga, Swedish handwriting or Arabic print? There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines. I’ve gathered 41 models into four collections, with short notes to help you choose: huggingface.co/collections/...
huggingface.co
OCR on the Hub - a davanstrien Collection
Curated OCR models for documents, languages, handwriting and text in images. Browse four collections with short practical notes.
0288
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 03/09/2026
A 15.9M-parameter specialist OCR model outperforms many VLMs hundreds of times larger. On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved. huggingface.co/spaces/fineb...
Plot showing parameters vs performance. Kraken is top left.
2346
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 20/08/2026
Every illustration in the British Library's 19th-century book images dataset now has an instance mask: 1.02M clean cutouts of public-domain engravings, no image generation involved! One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open.
111425
Reposted by Sebastian Majstorovic
Dominique Reill @dominiquereill.bsky.social · 12/08/2026
Interested in participating in the 2027 Central European History Convention July 15-17, 2027? Submit a paper proposal! Deadline September 13, 2026. Check here for all the different options of how to apply to be part of the program! ceh-c.univie.ac.at/news/call-fo... Please share far and wide!
ceh-c.univie.ac.at
Call for 2027
Call for 2027
086
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 10/08/2026
FineBooks wants to help re-OCR the public domain to unlock a rich source of data for AI, for researchers, and for enjoyment. Step one: work out which modern OCR models are actually good enough.
Leaderboard preview
2255
Sebastian Majstorovic @storytracer.com · 10/08/2026
Are open OCR models good enough to unlock historical knowledge? @danielvanstrien.bsky.social and I built a leaderboard using six expert-transcribed volumes from the @biodivlibrary.bsky.social sky.social. Meet FineBooks, from @hf.co and @eleutherai.bsky.social! huggingface.co/blog/fineboo...
huggingface.co
FineBooks: are open OCR models good enough to unlock historical knowledge?
A Blog post by FineBooks on Hugging Face
06322
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 07/08/2026
Uploaded a dataset of 1,080,814 public domain images, mostly from 19th-century books, to the Hugging Face Hub: huggingface.co/datasets/big... You can also do semantic search against the images here: huggingface.co/spaces/davan...
15412
Sebastian Majstorovic @storytracer.com · 26/07/2026
Heon Cha Haus was a great tip for Seoul @tedunderwood.com. I had their signature iced tea and an extremely tasty Red Bean & Cream Rice Cake. Not to mention the wonderful service! #DH2026
161
Sebastian Majstorovic @storytracer.com · 25/07/2026
Enjoying the weekend in Seoul before #DH2026 kicks off next week in Daejeon @dh2026daejeon.bsky.social.
091
Reposted by Sebastian Majstorovic
ReMona Lisa - News Reader - @thebookwormturns.co.uk · 11/07/2026
I saw this and had to share.
32855704
Reposted by Sebastian Majstorovic
Jim Clifford @jimclifford.bsky.social · 10/07/2026
Reading the Archive by Machine: An OCR Benchmark for Historians, 1612–1921 Here is version 1 of a working paper on the new OCR tools that are transforming digital history. working-papers-in-critical-search.github.io/paper-004-oc...
working-papers-in-critical-search.github.io
Reading the Archive by Machine – Working Papers in Critical Search
A benchmark of six OCR systems (Tesseract, olmOCR 2, Chandra 2, Infinity Parser 2, GLM-OCR, and Gemini 3.5 Flash) on human-transcribed archival documents spanning 1612–1921: early-modern print, ninete...
24114
Reposted by Sebastian Majstorovic
Data Rescue Project #DataRescue @datarescueproject.org · 08/07/2026
We are so very grateful for this honor! Thank you to our supporters and everyone who voted. Congratulations to everyone who was nominated -- we feel very privileged to be counted among your incredible work.
0155
Reposted by Sebastian Majstorovic
APDU @apduorg.bsky.social · 08/07/2026
Congratulations to the 2026 Data Integrity Award winners - @datarescueproject.org - Gina Plata-Nino of @fracposts.bsky.social - @mapresearch.bsky.social & Williams Institute - @hudgov.bsky.social & FEMA - Disaster Recovery Data Portal Details in 🧵
static.klipy.com
Congratulations Fireworks Display
ALT: Congratulations Fireworks Display
283
Reposted by Sebastian Majstorovic
Europeana @europeana.bsky.social · 08/07/2026
Read the feasibility study for the European Books Data Commons and discover the next steps! Explore how European libraries can make the full text of millions of digitised books available for research, innovation and contribute to a European AI infrastructure built on public values: bit.ly/4gs09tY
dataspace-culturalheritage.eu
Feasibility study paves the way for a European Books Data Commons
The National Library of the Netherlands (KB) and the Europeana Foundation have just published a feasibility study for the European Books Data Commons (EBDC).
01211
Reposted by Sebastian Majstorovic
𝐁𝐫𝐞𝐭 𝐁𝐚𝐭𝐬𝐨𝐧 @bretbatson.bsky.social · 26/06/2026
#DulaPeepShow #DuaLipa
consequence.net
Dua Lipa Opening Physical Library for Banned and Censored Books
Dua Lipa, founder of the Service95 Book Club, has partnered with Livraria Lello to open a physical library in Porto, Portugal.
11277041610
Reposted by Sebastian Majstorovic
samdonvil.bsky.social @samdonvil.bsky.social · 25/06/2026
@storytracer.com shows how volunteers or #GLAMS workers can be responsible actors in the current political climate by protecting datasets from war, political censorship and defunding! #glamlabsfutures
022
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 22/06/2026
I think VLM-based OCR might finally be close to working on historic newspapers! Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow. Surya OCR 2 (a 650M model!) does a very good job!
A 1901 newspaper page overlaid with Surya OCR 2's detected layout: colored boxes around each block (text, section-header, picture) and numbered dots joined by a path showing the model's reading order across the columns.
18814
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 23/06/2026
Ran the same 650M model across 7 languages: German, Polish, Finnish, Russian, Serbian, Estonian, Swedish via @europeana.bsky.social newspapers. Still some challenging layouts, but pretty mind-blowing to me that a 650M model is doing this when a year ago 72B VLMs failed very often.
3306
Reposted by Sebastian Majstorovic
Europeana @europeana.bsky.social · 23/06/2026
Read our newly published paper on AI 'The case for Public AI: making it happen with cultural heritage.' 400+ professionals, one shared position - cultural heritage should shape AI, not just feed it. 👉 bit.ly/4xLdyU3 #PublicAI #AI #ArtificialIntelligence #CulturalHeritage #DataSpace
dataspace-culturalheritage.eu
Making Public AI reality: How cultural heritage can lead the way
After months of collaborative development and iteration with the data space community and the Europeana Initiative, we are excited to share the newly published paper, ‘The case for Public AI: making i...
065
Reposted by Sebastian Majstorovic
Data Rescue Project #DataRescue @datarescueproject.org · 19/06/2026
Unfortunate to see this outcome but the fight for the panels at Washington’s house will continue in the community! This is our history and we want to #SaveOurSigns
095
Reposted by Sebastian Majstorovic
Biodiversity Heritage Library @biodivlibrary.bsky.social · 18/06/2026
In the 90 minutes after @theguardian.com article went live, BHL received more than US$2,500 in donations. Thank you to everyone helping keep biodiversity knowledge free and open. Please read, share, and help us build momentum.
theguardian.com
A bonanza for fans of the natural world: the digital library sharing 64m pages of scientific knowledge with everyone
The Biodiversity Heritage Library is an invaluable online archive of historic texts on species living and lost supplied by the world’s leading museums and universities. Now its future is in doubt
19849
Reposted by Sebastian Majstorovic
CNN @cnn.com · 13/06/2026
A federal judge in Massachusetts has ordered the Trump administration to restore all signs that were changed or removed at national parks across the country as part of President Donald Trump’s controversial directive last year.
cnn.it
Judge orders Trump administration to restore signs changed at national parks | CNN Politics
A federal judge in Massachusetts has ordered the Trump administration to restore all signs that were changed or removed at national parks across the country as part of President Donald Trump’s controv...
1431676
Reposted by Sebastian Majstorovic
Thibault Clérice @ponteineptique.bsky.social · 29/05/2026
OCR for Ancient Greek is constrained by a lack of open training data. But the deeper challenge tackled by 2 papers from Inria is doing OCR and document structure recovery together: section hierarchies, milestone numbering, marginal references. Paper 1: arxiv.org/html/2603.02...
1257
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 22/05/2026
You can now run SQL over 2.19 BILLION web pages — zero download. @commoncrawl.bsky.social April 2026 crawl + URL index are on Hugging Face Storage Buckets. DuckDB reads it straight over hf:// — I counted all 2.19B in ~35s. Or point your own agent at it 👇 huggingface.co/spaces/davan...
Language breakdown prompt for AI agentCopy and paste for DuckDB syntax
1529
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 20/05/2026
NuExtract3 (4B, Apache-2.0) does OCR *and* structured extraction. Point it at a dataset of scanned index cards + a JSON schema → clean catalog JSON One command on huggingface Jobs. (or skip the schema for plain Markdown OCR) Script + dataset 👇
A worked example of NuExtract3 turning a document image into structured data. Top: a terminal command — hf jobs uv run --image vllm/vllm-openai nuextract3.py index-cards cards-json --template schema.json. Middle (input): a scanned typewritten library index card reading "ABAD (Joseph), Captain, Spanish Army, letter of (1783), 5538, f.11." Bottom (output): the JSON extracted from the card — image_type "index_card", heading "Abad J.", heading_type "person", epithet "Captain, Spanish Army", and one entry with ms_no "5538", folios ["f.11"], description "letter of (1783)".
24112
Reposted by Sebastian Majstorovic
Free Law Project ⚖ @free.law · 07/05/2026
We're announcing two changes CourtListener API access: 1. Full API access is now open to everyone, including the PACER APIs that previously required a conversation with us. 2. Higher tiers are available through FLP memberships (including edu!) or commercial agreements. 👇 free.law/2026/05/07/a...
free.law
Full CourtListener Data Access via API Now Included with Membership
Researchers, journalists, developers, and vibe coders can now access the full CourtListener API, including PACER data, with a membership. No contact form. No waiting for approval.
12713
Reposted by Sebastian Majstorovic
infoDOCKET @infodocket.bsky.social · 07/05/2026
The Guardian: “‘Things Were Going Dark Left and Right’: The Race to Save US Government #Datasets Before They’re Deleted” www.infodocket.com/2026/05/07/t... @datarescueproject.org @envirodgi.bsky.social @archive.org #data #libraries #librarians
174
Reposted by Sebastian Majstorovic
Data Rescue Project #DataRescue @datarescueproject.org · 07/05/2026
Thanks to @theguardian.com for highlighting the people involved in the effort to save and demonstrate the importance of federal public data. ❤️🛟 www.theguardian.com/us-news/2026...
theguardian.com
‘Things were going dark left and right’: the race to save US government datasets before they’re deleted
Group has banded together to rescue data as Trump administration has removed or altered data on climate change, reproductive health, LGBTQ people and more
22413
Reposted by Sebastian Majstorovic
Data Rescue Project #DataRescue @datarescueproject.org · 24/04/2026
As we head into the weekend, and near the end of our #BobsBurgers month of memes, I wanted to share this joyous depiction of our truly amazing DRP volunteers. They give me hope without intending to; they're creating the good they want to see in the world; they are all gold stars.
A still from Bob's Burgers. There is a big yellow star shooting upwards in front of a cloudy blue sky. The star has the facial features of Tina Belcher.
063
Reposted by Sebastian Majstorovic
Ana Stevenson @dranastevenson.bsky.social · 24/04/2026
“[T]he popularity of AskHistorians demonstrates that there can and should be more history-related jobs in both academic and public-facing settings. People want history to understand the world, and I can also tell you that those people do not want the AI-generated version of their answers.”
07323
Reposted by Sebastian Majstorovic
Europeana @europeana.bsky.social · 16/04/2026
We’ve signed the #OpenHeritageStatement to support a global call for equitable access to #PublicDomain heritage in the digital environment! Discover more about the Statement, how it links to our advocacy for open #CulturalHeritage and how you can sign ➡️ bit.ly/4sB1Ur0 @creativecommons.bsky.social
bit.ly
The Europeana Initiative is proud to sign the Open Heritage Statement
Read on to discover more about the Statement’s unified vision for the public domain, and how it links to our work and advocacy for open cultural heritage!
0105
Reposted by Sebastian Majstorovic
Peter Baker @peterbakernyt.bsky.social · 05/04/2026
After Trump took office last year, the US Holocaust Museum quietly removed educational material about American racism from its website and canceled a workshop on the "fragility of democracy," @iriesentner.bsky.social reports. www.politico.com/news/2026/04...
politico.com
‘Proactively fall in line:’ Holocaust Memorial Museum quietly changed content after Trump returned to office
Two former employees said they believed the museum was altering its content preemptively to avoid unwanted negative attention from the Trump administration.
47725352
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 30/03/2026
You never know what data will be used for! I uploaded a @britishlibrary.bsky.social dataset to Hugging Face in 2022. IIRC one of my first PR to a HF repo! 4 years later, someone trains a Victorian chatbot on it More libraries should be sharing their public domain collections for AI to build on!
7837
Reposted by Sebastian Majstorovic
Cameron Blevins @cblevins.bsky.social · 27/03/2026
📣 New historical data visualization! "How Fast Was the Mail?" is an interactive map showing how long information took to travel across the US between 1882-1908: cblevins.github.io/mail-time/ +
cblevins.github.io
How Fast Was the Mail?
Explore mail transit times via railway between major U.S. cities, 1882–1908.
922191
Reposted by Sebastian Majstorovic
Catherine Arnett @catherinearnett.bsky.social · 26/03/2026
I’m at #HSP2026 at MIT this week! I’ll be giving a talk Friday at 5:25pm entitled “Structural Priming Effects in Language Models are Less Human-like in Languages Other Than English”. Looking forward to chatting to everyone!
0121
Reposted by Sebastian Majstorovic
Anne Schechinger @anneschech.bsky.social · 18/03/2026
Interesting new piece from @iatp.bsky.social showing how confidence in USDA data is eroding under the Trump Administration, with serious consequences. www.iatp.org/farming-risk...
iatp.org
Farming in the Dark: Unreliable USDA data jeopardizes a sustainable farming transition
As farmers struggle to adjust to a changing climate that has led to an increased frequency of droughts, floods, and other extreme weather events, incomplete or inaccurate public data has become a grow...
12811
Reposted by Sebastian Majstorovic
Data Rescue Project #DataRescue @datarescueproject.org · 27/02/2026
New DRP post: #Philly interpretive panels returned to the President's House. We congratulate the local community that fought HARD for their return. We hope this inspires other communities to #SaveOurSigns www.datarescueproject.org/independence...
datarescueproject.org
Independence National Historical Park - A Hopeful Update
We recently posted about the takedown of signs at the President’s House site at Independence National Historical Park. We also posted a call for more photos for Save Our Signs, both before and after…
0155
Reposted by Sebastian Majstorovic
Rainer Simon @aboutgeo.bsky.social · 20/02/2026
Small update to #tinyIIIF, my no-nonsense #IIIF server! You can now choose your image server during setup: • IIPImage • Cantaloupe Running small- to mid-sized collections? Teaching with IIIF materials? Building IIIF-enabled tools? Check out tiny.iiif! github.com/rsimon/tiny-... #DigitalFriday
Decorative images with a screenshot of tiny.iiif in the background, and image server choices (as text labels) in the foreground: IIPImage, Cantaloupe.
174
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 19/02/2026
Re-OCR'd the complete 1771 Encyclopaedia Britannica (2,724 pages) with a single command on @hf.co Jobs. - 0.9B model (GLM-OCR) ~$0.002/page ~$5 total on an L4 GPU Before (old Tesseract ocr) → After
Screenshot of old vs new ocr. 

old ocr text is garbled. New ocr much cleaner.
59616
Reposted by Sebastian Majstorovic
Data Rescue Project #DataRescue @datarescueproject.org · 18/02/2026
Get ready for live updates and quotes from @lyndamk.bsky.social and @mikalarae.bsky.social's #IDCC26 Keynote!
7103
Reposted by Sebastian Majstorovic
Rainer Simon @aboutgeo.bsky.social · 17/02/2026
Yay! The first image served from my new #tinyIIIF test instance. I'm running it on a 2 CPU/4GB RAM VM. Seems to be the absolute minimum & performance is pretty slow. If anyone has advice on which specs I'd need to get more #IIIF speed out of Cantaloupe – let me know!
The Waldseemüller map in liiive (a IIIF annotation tool).
021
Reposted by Sebastian Majstorovic
eleutherai.bsky.social @eleutherai.bsky.social · 13/02/2026
Announcing our latest paper: CommonLID In collaboration with @commoncrawl.bsky.social @mlcommons.org @jhu.edu we built a LID benchmark on actual Common Crawl text covering 109 languages. Existing evaluations overestimate how well LangID works on web data. arxiv.org/abs/2601.18026
arxiv.org
CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data
Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data of...
12212
Reposted by Sebastian Majstorovic
Data Rescue Project #DataRescue @datarescueproject.org · 05/02/2026
It's official. One year ago today we formalized as the Data Rescue Project. We know this work is exhausting and are so grateful to everyone who has showed up, day in and day out.
Meme from the Simpsons, reading "0 days without needing to save data from the US federal government."
1293
Reposted by Sebastian Majstorovic
rdapassociation.bsky.social @rdapassociation.bsky.social · 03/02/2026
🎉Congrats to the winners of the 2025 RDAP Work of the Year Award, The Data Rescue Project! (@datarescueproject.org)🛟🏆 This award acknowledges the work’s impact on the wider research and scholarly communication ecosystem in support of RDAP’s mission and values. rdapassociation.org/news/13593532
rdapassociation.org
Research Data Access and Preservation Association - 2025 RDAP Work of the Year Award
0207
Reposted by Sebastian Majstorovic
Lynda M. Kellam #DataRescue @lyndamk.bsky.social · 30/01/2026
Say hi! to the wonderful people doing the @datarescueproject.org AMA tonight! @quetzal1234.bsky.social @nurnberger.bsky.social @storytracer.com Tess Just missing for now: @mikalarae.bsky.social and @katscade.bsky.social
2122
Reposted by Sebastian Majstorovic
Daniel van Strien @danielvanstrien.bsky.social · 19/12/2025
Built a 2.5MB image classifier that runs in the browser in an evening with Claude Code. I used a dataset I labelled in 2022 and left on @hf.co for 3 years 😬. It finds illustrated pages in historical books. No server. No GPU.
28618
Reposted by Sebastian Majstorovic
Melanie Walsh @mellymeldubs.bsky.social · 17/12/2025
Very happy to introduce a new tool, BookReconciler! You can take spreadsheets with book data and add subject headings, descriptions, ISBNs, HathiTrust IDs, & more. You can also cluster editions & variations of the same "Work." Led by @thisismattmiller.com and supported by @post45data.bsky.social.
Diagram illustrating the BookReconciler workflow. On the left, a book cover of The Book of Salt by Monique Truong appears alongside “Minimal Metadata,” listing Author: Truong, Monique and Title: The Book of Salt. An arrow points to a box labeled “BookReconciler” with book and diamond icons. A downward arrow leads to “Enriched + Clustered Metadata,” showing multiple editions of the book cover and expanded metadata, including several ISBNs, subject headings (e.g., Vietnamese–France fiction, women authors, household employees, gay men, cooking), and an author VIAF identifier.
712256
Reposted by Sebastian Majstorovic
Anna Kijas @akijas.bsky.social · 12/12/2025
SUCHO has received the 2025 Karl Preusker medal from Bibliothek & Information Deutschland (BID)! The jury selected us as an example of “courage, solidarity, professional excellence, and the central role of libraries, archives, and digital infrastructures in the resilience of democratic societies.”
bsb-muenchen.de
Auszeichnungen
Der Dachverband Bibliothek & Information Deutschland (BID) e. V. hat die Karl-Preusker-Medaille 2025 dem Ukrainischen Bibliotheksverband verliehen. „Ausgezeichnet wird damit der außergewöhnliche Einsa...
1232