Sign in

William J.B. Mattingly

@wjbmattingly.bsky.social
555 followers 116 following 385 posts

Digital Nomad · Historian · Data Scientist · NLP · Machine Learning Cultural Heritage Data Scientist at Yale Former Postdoc in the Smithsonian Maintainer of Python Tutorials for Digital Humanities linktr.ee/wjbmattingly

PostsRepliesMedia
William J.B. Mattingly @wjbmattingly.bsky.social · 01/10/2026
Finetuned GliNER Decision is now beating Jev in entity linking. In our tests, we start to see it enter parity with around 600 examples, but we pushed it a bit further with around 3,000 examples. We'll be sharing the model and the data soon!
191
William J.B. Mattingly @wjbmattingly.bsky.social · 25/09/2026
woot woot! First finetune of GliNER-Decide!
071
William J.B. Mattingly @wjbmattingly.bsky.social · 24/09/2026
Here's my blog on how to leverage the stochastic nature of VLMs to identify areas of potential divergence (most likely areas for errors). Read more here! wjbmattingly.com/blog/where-t...
1144
William J.B. Mattingly @wjbmattingly.bsky.social · 16/09/2026
One of the fun things about LLMs is we can use their greatest weakness in OCR/HTR (their non-determinism) to be their greatest strength in identifying potential errors. The same model running over the same image multiple times shows where divergence (errors) appear. Blog and app coming soon!
2323
William J.B. Mattingly @wjbmattingly.bsky.social · 09/09/2026
Qwen 3.5 0.8B finetune comma base model now has LoRA adapters that can do bbox alignment in the same pass as HTR.
161
William J.B. Mattingly @wjbmattingly.bsky.social · 08/09/2026
Still finetuning these models. Page-level VLM Qwen 3.5 0.8B finetunes for medieval Latin manuscripts. All models were trained on 10k images. I am now training further on an additional 20k images from the comma dataset. You can test the models on @hf.co : huggingface.co/spaces/wjbma...
262
William J.B. Mattingly @wjbmattingly.bsky.social · 22/07/2026
I've added 210 articles to this local database and the knowledge graph is starting to get extensive. KGs get more useful as they get larger. Of these 210 articles, there are 30 references to Alcuin corresponding with (only type of many relationships in the KG) someone with citations and excerpts.
030
William J.B. Mattingly @wjbmattingly.bsky.social · 21/07/2026
Prototype in place for going from a collection of scholarly articles to knowledge graph. Also, there's a component in place that automatically links to linked open data, like Wikidata. All of this knowledge is from only 5 articles.
1100
William J.B. Mattingly @wjbmattingly.bsky.social · 10/07/2026
Daniel does a lot of cool stuff, but this is awesome!
050
William J.B. Mattingly @wjbmattingly.bsky.social · 10/07/2026
Sneak peak at what I've been working on =). Zero-shot with Gemini 3.5 Flash. Object detection for lines and images and transcription at the line level all in a single prompt. I'll be talking about this in Vienna in September. Fable 5 made the viewer.
060
William J.B. Mattingly @wjbmattingly.bsky.social · 15/06/2026
We found a 398-node geographic cycle in Yale's LUX database that poisoned 458,214 downstream records. By using Tarjan's SCC algorithm and cleaning the data with Gemini Flash Lite, we fixed the issue for about $6.50.
1103
Reposted by William J.B. Mattingly
Rebecca Sutton Koeser @suttonkoeser.bsky.social · 18/05/2026
Recently published: article on the research that went into the 2.0 version of the Shakespeare and Company Project datasets, and the potential of the updates and additional data on member addresses and authors. doi.org/10.22148/jca... One of the first few in the newly launched JCA!
doi.org
<em>Shakespeare and Company Project</em> Data Sets, Version 2.0
The Shakespeare and Company Project data sets provide a detailed portrait of Shakespeare and Company, Sylvia Beach’s bookshop and lending library in interwar Paris. This article outlines the research,...
2115
Reposted by William J.B. Mattingly
Peter T. Evans @petertevans.bsky.social · 18/05/2026
New #DigitalHumanities journal
041
William J.B. Mattingly @wjbmattingly.bsky.social · 12/05/2026
We tested OpenAI’s Whisper on 1,847 Holocaust testimonies from the Fortunoff at Yale. It achieved 85% accuracy, but the errors reveal how AI struggles in this domain. Full findings: wjbmattingly.com/blog/transcribing-…
1182
William J.B. Mattingly @wjbmattingly.bsky.social · 10/05/2026
Mallorca is so photogenic. It’s hard to take a bad photo.
050
Reposted by William J.B. Mattingly
Jane Winters @jfwinters.bsky.social · 07/05/2026
‘Exploring Digital Cultural Heritage’ is now out in the wild! You can download an open-access copy (PDF) or read the free digital edition on Manifold uolpress.co.uk/book/explori.... Loved working on this with Eirini Goudarouli, @amsichani.bsky.social & the @uolpress.bsky.social team.
Book cover for Exploring Digital Cultural Heritage in purple, pink and yellow with University of London Press logo and Digital Cultural Heritage series strapline
39661
William J.B. Mattingly @wjbmattingly.bsky.social · 07/05/2026
Off to Mallorca to present on our work at Fortunoff! More coming soon!
040
William J.B. Mattingly @wjbmattingly.bsky.social · 04/05/2026
We parsed 3.6 million historical names with 96% accuracy by moving away from expensive frontier models to fine-tuned Qwen 3.5 models. We also found that switching our output format from JSON to YAML dramatically increased stability.
64911
William J.B. Mattingly @wjbmattingly.bsky.social · 30/04/2026
Haven't gotten to use graph theory in a long time. The last few days have been a lot of fun! I just finished writing a blog about how we are using graph theory to improve how places are represented in LUX at Yale. More coming soon!
051
William J.B. Mattingly @wjbmattingly.bsky.social · 27/04/2026
I’ve replaced Conda with uv by @crmarsh.com for my Python projects. The slow dependency resolution and cross-platform friction between Windows and Linux was the catalyst. uv is faster, leaner, and more reliable for team workflows. Full breakdown here: wjbmattingly.com/blog/bye-conda-hel…
070
William J.B. Mattingly @wjbmattingly.bsky.social · 27/04/2026
I spent the weekend migrating my old PythonHumanities website from WordPress to a custom design (thanks, Claude) that allows for more flexibility, including live coding exercises in each page. I also updated things to make getting started with Python easier by using uv. More soon!
1395
Reposted by William J.B. Mattingly
Patrick J. Burns @diyclassics.bsky.social · 26/03/2026
✨ LatinCy v3.9 sm/md/lg/trf pipelines for SpaCy available ✨ - Improved tokenization and u/v norm - New custom Latin-specific XPOS tags - Better, more consistent lemma/morph coverage huggingface.co/latincy/la_c... #digiclass #nlproc
"Album" cover for the LatinCy v3.9 pipelines with "catus" from Gesner's 1551 Historia Animalium
0137
Reposted by William J.B. Mattingly
Ben Lee @bcgl.bsky.social · 18/03/2026
Last May, @toddpresner.bsky.social, Robert Ehrenreich, and I convened a symposium at @ischool.uw.edu to explore "AI and the Future of Holocaust Research & Memory." Our resulting white paper is now officially online and contains provocations, reflections, and refusals: doi.org/10.6069/cx72...
doi.org
AI and The Future of Holocaust Research & Memory
How will the advent of AI impact the future of Holocaust studies? Will it provide new methods for analyzing data and displaying information for research and education that will benefit the field, or w...
191
William J.B. Mattingly @wjbmattingly.bsky.social · 17/03/2026
Want an easy way to find Scripture references in texts? Now you can. I modified our AI pipeline to perform scripture detection. It leverages Gemini 3.1 Pro Preview (requires API key). It's fully open-source and found here: github.com/yale-ch/scri...
2157
William J.B. Mattingly @wjbmattingly.bsky.social · 12/03/2026
Open-sourcing this soon. Leverage Gemini to do vulgate scriptural identification in texts (of most languages).
071
William J.B. Mattingly @wjbmattingly.bsky.social · 19/02/2026
Opportunity: we have two post-doc positions at Yale that are about to be opened! If you have finished your Ph.D. within the last 3 years in one of these areas, let's talk! First: Using AI and knowledge graphs to accelerate provenance research. Second: Knowledge graphs for research dataset discovery.
yale.edu
Yale University
Since its founding in 1701, Yale University has been dedicated to expanding and sharing knowledge, inspiring innovation, and preserving cultural and scientific information for future generations.
141
Reposted by William J.B. Mattingly
Patrick J. Burns @diyclassics.bsky.social · 02/02/2026
Announcing—LatinCy Readers v.1.0.2, i.e. LatinCy-powered corpus readers for Latin text collections. Quickly get sentences, lines, words annotated for lemma, POS, morphology, NER, etc. Supporting .txt, .xml, .tess, .conllu, and more. github.com/diyclassics/... #digiclass #nlproc
LatinCy Readers logo
22916
William J.B. Mattingly @wjbmattingly.bsky.social · 23/01/2026
Finally getting back to this app that lets you edit/transcribe/annotate the Voynich manuscript.
070
William J.B. Mattingly @wjbmattingly.bsky.social · 30/10/2025
You can now process Hebrew archival documents with Qwen 3 VL =) --- Will be using this to finetune further on handwritten Hebrew. Metrics are on the test set that is fairly close in style and structure to the training data. I tested on out-of-training edge cases and it worked (Link to model below)
150
William J.B. Mattingly @wjbmattingly.bsky.social · 27/10/2025
Does anyone have a dataset of 1,000 + pages of handwritten text on Transkribus that they want to use for finetuning a VLM? If so, please let me know. This would be for any language and any script.
035
William J.B. Mattingly @wjbmattingly.bsky.social · 24/10/2025
More coming soon but finetuned Qwen 3 VL-8B on 150k lines of synthetic Yiddish typed and handwritten data. Results are pretty amazing. Even on the harder heldout set it gets a CER of 1% and a WER of 2%. Preparing page-level dataset and finetunes now, thanks to the John Locke Jr.
091
William J.B. Mattingly @wjbmattingly.bsky.social · 24/10/2025
Over the last 24 hours, I have finetuned three Qwen3-VL models (2B, 4B, and 8B) on the CATmuS dataset on @hf.co . The first version of the models are now available on the Small Models for GLAM organization with @danielvanstrien.bsky.social (Links below) Working on improving them further.
1112
Reposted by William J.B. Mattingly
Daniel van Strien @danielvanstrien.bsky.social · 22/10/2025
DeepSeek-OCR just got vLLM support 🚀 Currently processing @natlibscot.bsky.social's 27,915-page handbook collection with one command. Processing at ~350 images/sec on A100 Using @hf.co Jobs + uv - zero setup batch OCR! Will share final time + cost when done!
Logs showing ocr progress
1182
William J.B. Mattingly @wjbmattingly.bsky.social · 21/10/2025
Want an easy way to edit the output from Dots.OCR? Introducing Dots.OCR editor, an easy way to edit outputs from the model. Features: 1) Edit bounding boxes 2) Edit OCR 3) Edit reading order 4) Group sections (good for newspapers) Vibe coded with Claude 4.5 github.com/wjbmattingly...
github.com
GitHub - wjbmattingly/dots-ocr-editor
Contribute to wjbmattingly/dots-ocr-editor development by creating an account on GitHub.
062
Reposted by William J.B. Mattingly
Daniel van Strien @danielvanstrien.bsky.social · 16/10/2025
Small models work great for GLAM but there aren't enough examples! With @wjbmattingly.bsky.social I'm launching small-models-for-glam on @hf.co to create/curate models that run on modest hardware and address GLAM use cases. Follow the org to keep up-to-date! huggingface.co/small-models...
0127
William J.B. Mattingly @wjbmattingly.bsky.social · 24/09/2025
🚨Job ALERT🚨! My old postdoc is available! I cannot emphasize enough how much a life-altering position this was for me. It gave me the experience that I needed for my current role. As a postdoc, I was able to define my projects and acquire a lot of new skills as well as refine some I already had.
067
Reposted by William J.B. Mattingly
Ben Lee @bcgl.bsky.social · 09/09/2025
Excited to be co-editing a special issue of @dhquarterly.bsky.social on Artificial Intelligence for Digital Humanities: Research problems and critical approaches dhq.digitalhumanities.org/news/news.html We're inviting abstracts now - please feel free to reach out with any questions!
dhq.digitalhumanities.org
DHQ: Digital Humanities Quarterly: News
02110
William J.B. Mattingly @wjbmattingly.bsky.social · 14/08/2025
Something I've realized over the last couple weeks with finetuning various VLMs is that we just need more data. Unfortunately, that takes a lot of time. That's why I'm returning to my synthetic HTR workflow. This will be packaged now and expanded to work with other low-resource languages. Stay tuned
0100
William J.B. Mattingly @wjbmattingly.bsky.social · 13/08/2025
I've been getting asked training scripts when a new VLM drops. Instead of scripts, I'm going to start updating this new Python package. It's not fancy. It's for full finetunes. This was how I first trained Qwen 2 VL last year.
1142
William J.B. Mattingly @wjbmattingly.bsky.social · 13/08/2025
Let's go! Training LFM2-VL 1.6B on Catmus dataset on @hf.co now. Will start posting some benchmarks on this model soon.
041
William J.B. Mattingly @wjbmattingly.bsky.social · 13/08/2025
Training on full catmus now and the results after first checkpoint are very promising. Character and massive word-level improvement.
131
William J.B. Mattingly @wjbmattingly.bsky.social · 13/08/2025
LiquidAI cooked with LFM2-VL. At the risk of sounding like an X AI influencer, don't sleep on this model. I'm finetuning right now on Catmus. A small test over night on only 3k examples is showing remarkable improvement. Training now on 150k samples. I see this as potentially replacing TrOCR.
192
William J.B. Mattingly @wjbmattingly.bsky.social · 12/08/2025
New super lightweight VLM just dropped from Liquid AI in two flavors: 450M and 1.6B. Both models can work out-of-the-box with medieval Latin at the line level. I'm fine-tuning on Catmus/medieval right now on an h200.
193
Reposted by William J.B. Mattingly
Rainer Simon @aboutgeo.bsky.social · 12/08/2025
With #IMMARKUS, you can already use popular AI services for image transcription. Now, you can also use them for translation! Transcribe a historic source, select the annotation—and translate it with a click.
172
William J.B. Mattingly @wjbmattingly.bsky.social · 12/08/2025
GLM-4.5V with line-level transcription of medieval Latin in Caroline Miniscule. Inference was run through @hf.co Inferencevia Novita.
041
William J.B. Mattingly @wjbmattingly.bsky.social · 11/08/2025
Qwen 3-4B Thinking finetune nearly ready to share. It can convert unstructured natural language, non-linkedart JSON, and HTML into LinkedArt JSON.
050
Reposted by William J.B. Mattingly
Claire @chirila.bsky.social · 08/08/2025
Discover Magazine did a nice feature on the Voynich Manuscript. I had a delightful conversation with Sam Waters, and here's the result. (there's also a print version in the current issue) #linguistics #Yale #YaleLibrary #BeineckeLibrary #conlangs www.discovermagazine.com/was-the-worl...
discovermagazine.com
Was the World’s Most Mysterious Manuscript from the Middle Ages A Hoax?
Its indecipherable text and illustrations have stumped scholars throughout history. Here’s what we know about the Voynich Manuscript’s content and creation, and whether the text is truly as mystifying...
282
William J.B. Mattingly @wjbmattingly.bsky.social · 07/08/2025
I spent the last 24 hours finetuning Dots.OCR with different datasets. Here are some of the things I learned... TLDR Don't sleep on this model! If you do HTR or OCR try this.
283
William J.B. Mattingly @wjbmattingly.bsky.social · 06/08/2025
Vibe coding is great for quick visualizers for projects. Single Claude 4 Sonnet prompt and viola.
010