Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 16/09/2026Made some improvements to the OCR scripts onboarding in my uv-scripts collection. First run is now a small example: seven scanned NASA pages in, Markdown out, one HF jobs command. It also works on your own images or a folder of PDFs. You can pick from 28 OCR models. huggingface.co/datasets/uv-... 1141
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 11/09/2026Need a historical illustration? Search 1.49 million images from British Library books and Britannica (1500s–1920s). Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources. huggingface.co/spaces/davan... 47620
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 07/09/2026OCR for Japanese manga, Swedish handwriting or Arabic print? There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines. I’ve gathered 41 models into four collections, with short notes to help you choose: huggingface.co/collections/...huggingface.coOCR on the Hub - a davanstrien CollectionCurated OCR models for documents, languages, handwriting and text in images. Browse four collections with short practical notes. 0288
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 03/09/2026A 15.9M-parameter specialist OCR model outperforms many VLMs hundreds of times larger. On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved. huggingface.co/spaces/fineb... 2346
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 20/08/2026Every illustration in the British Library's 19th-century book images dataset now has an instance mask: 1.02M clean cutouts of public-domain engravings, no image generation involved! One small segmentation model, $3.24 of compute. Model, masks and a search Space with cutout view all open. 111425
Reposted by Sebastian MajstorovicDominique Reill @dominiquereill.bsky.social · 12/08/2026Interested in participating in the 2027 Central European History Convention July 15-17, 2027? Submit a paper proposal! Deadline September 13, 2026. Check here for all the different options of how to apply to be part of the program! ceh-c.univie.ac.at/news/call-fo... Please share far and wide!ceh-c.univie.ac.atCall for 2027Call for 2027 086
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 10/08/2026FineBooks wants to help re-OCR the public domain to unlock a rich source of data for AI, for researchers, and for enjoyment. Step one: work out which modern OCR models are actually good enough. 2255
Sebastian Majstorovic @storytracer.com · 10/08/2026Are open OCR models good enough to unlock historical knowledge? @danielvanstrien.bsky.social and I built a leaderboard using six expert-transcribed volumes from the @biodivlibrary.bsky.social sky.social. Meet FineBooks, from @hf.co and @eleutherai.bsky.social! huggingface.co/blog/fineboo...huggingface.coFineBooks: are open OCR models good enough to unlock historical knowledge?A Blog post by FineBooks on Hugging Face 06322
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 07/08/2026Uploaded a dataset of 1,080,814 public domain images, mostly from 19th-century books, to the Hugging Face Hub: huggingface.co/datasets/big... You can also do semantic search against the images here: huggingface.co/spaces/davan... 15412
Sebastian Majstorovic @storytracer.com · 26/07/2026Heon Cha Haus was a great tip for Seoul @tedunderwood.com. I had their signature iced tea and an extremely tasty Red Bean & Cream Rice Cake. Not to mention the wonderful service! #DH2026 161
Sebastian Majstorovic @storytracer.com · 25/07/2026Enjoying the weekend in Seoul before #DH2026 kicks off next week in Daejeon @dh2026daejeon.bsky.social. 091
Reposted by Sebastian MajstorovicReMona Lisa - News Reader - @thebookwormturns.co.uk · 11/07/2026I saw this and had to share. 32855704
Reposted by Sebastian MajstorovicJim Clifford @jimclifford.bsky.social · 10/07/2026Reading the Archive by Machine: An OCR Benchmark for Historians, 1612–1921 Here is version 1 of a working paper on the new OCR tools that are transforming digital history. working-papers-in-critical-search.github.io/paper-004-oc...working-papers-in-critical-search.github.ioReading the Archive by Machine – Working Papers in Critical SearchA benchmark of six OCR systems (Tesseract, olmOCR 2, Chandra 2, Infinity Parser 2, GLM-OCR, and Gemini 3.5 Flash) on human-transcribed archival documents spanning 1612–1921: early-modern print, ninete... 24114
Reposted by Sebastian MajstorovicData Rescue Project #DataRescue @datarescueproject.org · 08/07/2026We are so very grateful for this honor! Thank you to our supporters and everyone who voted. Congratulations to everyone who was nominated -- we feel very privileged to be counted among your incredible work. 0155
Reposted by Sebastian MajstorovicAPDU @apduorg.bsky.social · 08/07/2026Congratulations to the 2026 Data Integrity Award winners - @datarescueproject.org - Gina Plata-Nino of @fracposts.bsky.social - @mapresearch.bsky.social & Williams Institute - @hudgov.bsky.social & FEMA - Disaster Recovery Data Portal Details in 🧵static.klipy.comCongratulations Fireworks DisplayALT: Congratulations Fireworks Display 283
Reposted by Sebastian MajstorovicEuropeana @europeana.bsky.social · 08/07/2026Read the feasibility study for the European Books Data Commons and discover the next steps! Explore how European libraries can make the full text of millions of digitised books available for research, innovation and contribute to a European AI infrastructure built on public values: bit.ly/4gs09tYdataspace-culturalheritage.euFeasibility study paves the way for a European Books Data CommonsThe National Library of the Netherlands (KB) and the Europeana Foundation have just published a feasibility study for the European Books Data Commons (EBDC). 01211
Reposted by Sebastian Majstorovic𝐁𝐫𝐞𝐭 𝐁𝐚𝐭𝐬𝐨𝐧 @bretbatson.bsky.social · 26/06/2026#DulaPeepShow #DuaLipaconsequence.netDua Lipa Opening Physical Library for Banned and Censored BooksDua Lipa, founder of the Service95 Book Club, has partnered with Livraria Lello to open a physical library in Porto, Portugal. 11277041610
Reposted by Sebastian Majstorovicsamdonvil.bsky.social @samdonvil.bsky.social · 25/06/2026 @storytracer.com shows how volunteers or #GLAMS workers can be responsible actors in the current political climate by protecting datasets from war, political censorship and defunding! #glamlabsfutures 022
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 22/06/2026I think VLM-based OCR might finally be close to working on historic newspapers! Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow. Surya OCR 2 (a 650M model!) does a very good job! 18814
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 23/06/2026Ran the same 650M model across 7 languages: German, Polish, Finnish, Russian, Serbian, Estonian, Swedish via @europeana.bsky.social newspapers. Still some challenging layouts, but pretty mind-blowing to me that a 650M model is doing this when a year ago 72B VLMs failed very often. 3306
Reposted by Sebastian MajstorovicEuropeana @europeana.bsky.social · 23/06/2026Read our newly published paper on AI 'The case for Public AI: making it happen with cultural heritage.' 400+ professionals, one shared position - cultural heritage should shape AI, not just feed it. 👉 bit.ly/4xLdyU3 #PublicAI #AI #ArtificialIntelligence #CulturalHeritage #DataSpacedataspace-culturalheritage.euMaking Public AI reality: How cultural heritage can lead the wayAfter months of collaborative development and iteration with the data space community and the Europeana Initiative, we are excited to share the newly published paper, ‘The case for Public AI: making i... 065
Reposted by Sebastian MajstorovicData Rescue Project #DataRescue @datarescueproject.org · 19/06/2026Unfortunate to see this outcome but the fight for the panels at Washington’s house will continue in the community! This is our history and we want to #SaveOurSigns 095
Reposted by Sebastian MajstorovicBiodiversity Heritage Library @biodivlibrary.bsky.social · 18/06/2026In the 90 minutes after @theguardian.com article went live, BHL received more than US$2,500 in donations. Thank you to everyone helping keep biodiversity knowledge free and open. Please read, share, and help us build momentum.theguardian.comA bonanza for fans of the natural world: the digital library sharing 64m pages of scientific knowledge with everyoneThe Biodiversity Heritage Library is an invaluable online archive of historic texts on species living and lost supplied by the world’s leading museums and universities. Now its future is in doubt 19849
Reposted by Sebastian MajstorovicCNN @cnn.com · 13/06/2026A federal judge in Massachusetts has ordered the Trump administration to restore all signs that were changed or removed at national parks across the country as part of President Donald Trump’s controversial directive last year.cnn.itJudge orders Trump administration to restore signs changed at national parks | CNN PoliticsA federal judge in Massachusetts has ordered the Trump administration to restore all signs that were changed or removed at national parks across the country as part of President Donald Trump’s controv... 1431676
Reposted by Sebastian MajstorovicThibault Clérice @ponteineptique.bsky.social · 29/05/2026OCR for Ancient Greek is constrained by a lack of open training data. But the deeper challenge tackled by 2 papers from Inria is doing OCR and document structure recovery together: section hierarchies, milestone numbering, marginal references. Paper 1: arxiv.org/html/2603.02... 1257
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 22/05/2026You can now run SQL over 2.19 BILLION web pages — zero download. @commoncrawl.bsky.social April 2026 crawl + URL index are on Hugging Face Storage Buckets. DuckDB reads it straight over hf:// — I counted all 2.19B in ~35s. Or point your own agent at it 👇 huggingface.co/spaces/davan... 1529
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 20/05/2026NuExtract3 (4B, Apache-2.0) does OCR *and* structured extraction. Point it at a dataset of scanned index cards + a JSON schema → clean catalog JSON One command on huggingface Jobs. (or skip the schema for plain Markdown OCR) Script + dataset 👇 24112
Reposted by Sebastian MajstorovicFree Law Project ⚖ @free.law · 07/05/2026We're announcing two changes CourtListener API access: 1. Full API access is now open to everyone, including the PACER APIs that previously required a conversation with us. 2. Higher tiers are available through FLP memberships (including edu!) or commercial agreements. 👇 free.law/2026/05/07/a...free.lawFull CourtListener Data Access via API Now Included with MembershipResearchers, journalists, developers, and vibe coders can now access the full CourtListener API, including PACER data, with a membership. No contact form. No waiting for approval. 12713
Reposted by Sebastian MajstorovicinfoDOCKET @infodocket.bsky.social · 07/05/2026The Guardian: “‘Things Were Going Dark Left and Right’: The Race to Save US Government #Datasets Before They’re Deleted” www.infodocket.com/2026/05/07/t... @datarescueproject.org @envirodgi.bsky.social @archive.org #data #libraries #librarians 174
Reposted by Sebastian MajstorovicData Rescue Project #DataRescue @datarescueproject.org · 07/05/2026Thanks to @theguardian.com for highlighting the people involved in the effort to save and demonstrate the importance of federal public data. ❤️🛟 www.theguardian.com/us-news/2026...theguardian.com‘Things were going dark left and right’: the race to save US government datasets before they’re deletedGroup has banded together to rescue data as Trump administration has removed or altered data on climate change, reproductive health, LGBTQ people and more 22413
Reposted by Sebastian MajstorovicData Rescue Project #DataRescue @datarescueproject.org · 24/04/2026As we head into the weekend, and near the end of our #BobsBurgers month of memes, I wanted to share this joyous depiction of our truly amazing DRP volunteers. They give me hope without intending to; they're creating the good they want to see in the world; they are all gold stars. 063
Reposted by Sebastian MajstorovicAna Stevenson @dranastevenson.bsky.social · 24/04/2026“[T]he popularity of AskHistorians demonstrates that there can and should be more history-related jobs in both academic and public-facing settings. People want history to understand the world, and I can also tell you that those people do not want the AI-generated version of their answers.” 07323
Reposted by Sebastian MajstorovicEuropeana @europeana.bsky.social · 16/04/2026We’ve signed the #OpenHeritageStatement to support a global call for equitable access to #PublicDomain heritage in the digital environment! Discover more about the Statement, how it links to our advocacy for open #CulturalHeritage and how you can sign ➡️ bit.ly/4sB1Ur0 @creativecommons.bsky.socialbit.lyThe Europeana Initiative is proud to sign the Open Heritage StatementRead on to discover more about the Statement’s unified vision for the public domain, and how it links to our work and advocacy for open cultural heritage! 0105
Reposted by Sebastian MajstorovicPeter Baker @peterbakernyt.bsky.social · 05/04/2026After Trump took office last year, the US Holocaust Museum quietly removed educational material about American racism from its website and canceled a workshop on the "fragility of democracy," @iriesentner.bsky.social reports. www.politico.com/news/2026/04...politico.com‘Proactively fall in line:’ Holocaust Memorial Museum quietly changed content after Trump returned to officeTwo former employees said they believed the museum was altering its content preemptively to avoid unwanted negative attention from the Trump administration. 47725352
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 30/03/2026You never know what data will be used for! I uploaded a @britishlibrary.bsky.social dataset to Hugging Face in 2022. IIRC one of my first PR to a HF repo! 4 years later, someone trains a Victorian chatbot on it More libraries should be sharing their public domain collections for AI to build on! 7837
Reposted by Sebastian MajstorovicCameron Blevins @cblevins.bsky.social · 27/03/2026📣 New historical data visualization! "How Fast Was the Mail?" is an interactive map showing how long information took to travel across the US between 1882-1908: cblevins.github.io/mail-time/ +cblevins.github.ioHow Fast Was the Mail?Explore mail transit times via railway between major U.S. cities, 1882–1908. 922191
Reposted by Sebastian MajstorovicCatherine Arnett @catherinearnett.bsky.social · 26/03/2026I’m at #HSP2026 at MIT this week! I’ll be giving a talk Friday at 5:25pm entitled “Structural Priming Effects in Language Models are Less Human-like in Languages Other Than English”. Looking forward to chatting to everyone! 0121
Reposted by Sebastian MajstorovicAnne Schechinger @anneschech.bsky.social · 18/03/2026Interesting new piece from @iatp.bsky.social showing how confidence in USDA data is eroding under the Trump Administration, with serious consequences. www.iatp.org/farming-risk...iatp.orgFarming in the Dark: Unreliable USDA data jeopardizes a sustainable farming transitionAs farmers struggle to adjust to a changing climate that has led to an increased frequency of droughts, floods, and other extreme weather events, incomplete or inaccurate public data has become a grow... 12811
Reposted by Sebastian MajstorovicData Rescue Project #DataRescue @datarescueproject.org · 27/02/2026New DRP post: #Philly interpretive panels returned to the President's House. We congratulate the local community that fought HARD for their return. We hope this inspires other communities to #SaveOurSigns www.datarescueproject.org/independence...datarescueproject.orgIndependence National Historical Park - A Hopeful UpdateWe recently posted about the takedown of signs at the President’s House site at Independence National Historical Park. We also posted a call for more photos for Save Our Signs, both before and after… 0155
Reposted by Sebastian MajstorovicRainer Simon @aboutgeo.bsky.social · 20/02/2026Small update to #tinyIIIF, my no-nonsense #IIIF server! You can now choose your image server during setup: • IIPImage • Cantaloupe Running small- to mid-sized collections? Teaching with IIIF materials? Building IIIF-enabled tools? Check out tiny.iiif! github.com/rsimon/tiny-... #DigitalFriday 174
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 19/02/2026Re-OCR'd the complete 1771 Encyclopaedia Britannica (2,724 pages) with a single command on @hf.co Jobs. - 0.9B model (GLM-OCR) ~$0.002/page ~$5 total on an L4 GPU Before (old Tesseract ocr) → After 59616
Reposted by Sebastian MajstorovicData Rescue Project #DataRescue @datarescueproject.org · 18/02/2026Get ready for live updates and quotes from @lyndamk.bsky.social and @mikalarae.bsky.social's #IDCC26 Keynote! 7103
Reposted by Sebastian MajstorovicRainer Simon @aboutgeo.bsky.social · 17/02/2026Yay! The first image served from my new #tinyIIIF test instance. I'm running it on a 2 CPU/4GB RAM VM. Seems to be the absolute minimum & performance is pretty slow. If anyone has advice on which specs I'd need to get more #IIIF speed out of Cantaloupe – let me know! 021
Reposted by Sebastian Majstoroviceleutherai.bsky.social @eleutherai.bsky.social · 13/02/2026Announcing our latest paper: CommonLID In collaboration with @commoncrawl.bsky.social @mlcommons.org @jhu.edu we built a LID benchmark on actual Common Crawl text covering 109 languages. Existing evaluations overestimate how well LangID works on web data. arxiv.org/abs/2601.18026arxiv.orgCommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web DataLanguage identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data of... 12212
Reposted by Sebastian MajstorovicData Rescue Project #DataRescue @datarescueproject.org · 05/02/2026It's official. One year ago today we formalized as the Data Rescue Project. We know this work is exhausting and are so grateful to everyone who has showed up, day in and day out. 1293
Reposted by Sebastian Majstorovicrdapassociation.bsky.social @rdapassociation.bsky.social · 03/02/2026🎉Congrats to the winners of the 2025 RDAP Work of the Year Award, The Data Rescue Project! (@datarescueproject.org)🛟🏆 This award acknowledges the work’s impact on the wider research and scholarly communication ecosystem in support of RDAP’s mission and values. rdapassociation.org/news/13593532rdapassociation.orgResearch Data Access and Preservation Association - 2025 RDAP Work of the Year Award 0207
Reposted by Sebastian MajstorovicLynda M. Kellam #DataRescue @lyndamk.bsky.social · 30/01/2026Say hi! to the wonderful people doing the @datarescueproject.org AMA tonight! @quetzal1234.bsky.social @nurnberger.bsky.social @storytracer.com Tess Just missing for now: @mikalarae.bsky.social and @katscade.bsky.social 2122
Reposted by Sebastian MajstorovicDaniel van Strien @danielvanstrien.bsky.social · 19/12/2025Built a 2.5MB image classifier that runs in the browser in an evening with Claude Code. I used a dataset I labelled in 2022 and left on @hf.co for 3 years 😬. It finds illustrated pages in historical books. No server. No GPU. 28618
Reposted by Sebastian MajstorovicMelanie Walsh @mellymeldubs.bsky.social · 17/12/2025Very happy to introduce a new tool, BookReconciler! You can take spreadsheets with book data and add subject headings, descriptions, ISBNs, HathiTrust IDs, & more. You can also cluster editions & variations of the same "Work." Led by @thisismattmiller.com and supported by @post45data.bsky.social. 712256
Reposted by Sebastian MajstorovicAnna Kijas @akijas.bsky.social · 12/12/2025SUCHO has received the 2025 Karl Preusker medal from Bibliothek & Information Deutschland (BID)! The jury selected us as an example of “courage, solidarity, professional excellence, and the central role of libraries, archives, and digital infrastructures in the resilience of democratic societies.”bsb-muenchen.deAuszeichnungenDer Dachverband Bibliothek & Information Deutschland (BID) e. V. hat die Karl-Preusker-Medaille 2025 dem Ukrainischen Bibliotheksverband verliehen. „Ausgezeichnet wird damit der außergewöhnliche Einsa... 1232