Sign in

Patrick J. Burns

@diyclassics.bsky.social
701 followers 285 following 125 posts

Associate Research Scholar, Digital Projects @ ISAW Library | prev. Quantitative Criticism Lab (UT-Austin) & Culture, Cognition, Coevolution Lab (Harvard) | Fordham PhD, #Classics | #DigiClass, #Latin, #Greek | LatinCy dev, CLTK contrib | #Python

PostsRepliesMedia
Patrick J. Burns @diyclassics.bsky.social · 05/10/2026
Will be at #CAAS2026 this week—hosting a workshop Friday 10/9 @ 10:30am on the computational methods behind generating Latin vocabulary lists from plaintext passages #digiclass #nlproc diyclassics.github.io/automating/
Program text: "Workshop 1: Automating Vocabulary List Production King Sullivan Room. Organizer: Patrick Burns (Institute for the Study of the Ancient World, NYU). Hands-on introduction to state-of-the-art computational text analysis for Latin pedagogy using the example of
generating publication-quality vocabulary lists from plain text."
131
Patrick J. Burns @diyclassics.bsky.social · 04/08/2026
✨Version bump✨ latincy-preprocess v0.4.0 Added Ancient Greek preprocessing features (based on @jtauber.com excellent greek-normalisation work!) github.com/latincy/lati...
Header text for LatinCy Preprocess (v0.4.0) GitHub page, including badges. Text reads: Latin text preprocessing: U/V normalization, long-s OCR correction, diacritics stripping, macron removal, and Beta Code → Unicode Greek conversion — plus Ancient Greek elision/accent normalization — with optional Rust acceleration and spaCy integration.
093
Patrick J. Burns @diyclassics.bsky.social · 31/07/2026
📝 New post, new series 📝 "Homeric Hapax Tables in Kumpf 1984" Quantified Homer post on a data/code reconstruction of the four summary tables in Kumpf’s 1984 book »Four Indices of the Homeric Hapax Legomena« exploratoryphilology.org/posts/kumpf-1984/ #digiclass
First page of Kumpf's Index I: An Alphabetical Listing of All the Homeric Hapax Legomena with some alpha entries
031
Patrick J. Burns @diyclassics.bsky.social · 26/07/2026
Here is a screenshot with one example of possible output from the repo's demo notebook: github.com/latincy/lati....
Code to produce a formatted vocabulary list for a Latin passage using latincy-vocab
110
Patrick J. Burns @diyclassics.bsky.social · 24/07/2026
🚀 New release 🚀 latincy-vocab v0.1.0 Build well-formatted Latin vocabulary lists from plaintext github.com/latincy/lati... #digiclass #teachlatin
Logo and badge information for LatinCy Vocab v0.1.0
1116
Patrick J. Burns @diyclassics.bsky.social · 10/07/2026
✨Version bump✨ latincy-lexicon v0.9.0 Whitaker's Words as LatinCy component Improved handling of archaic & rare forms in both lookups and paradigms, fixes to irregular verbs, better future tense handling, beta Lewis & Short integration, etc. github.com/latincy/lati...
LatinCy Lexicon logo, badge, description for v0.9.0
040
Patrick J. Burns @diyclassics.bsky.social · 09/07/2026
✨Version bump✨ txtdown v0.3.1 Minimal markup for Latin text collections Improved document validation; improved quotation handling; hierarchy-aware citations github.com/diyclassics/...
Release details for txtdown v0.3.1
0104
Patrick J. Burns @diyclassics.bsky.social · 14/04/2026
We can also "reinflect" Latin forms based on existing contextual morph annotations, e.g.
Code to generate "reinflected" forms on Latin words in context with LatinCy Lexicon
010
Patrick J. Burns @diyclassics.bsky.social · 14/04/2026
Not only can we start using LatinCy annotations to disambiguate words, we can also use the WW word formation logic to generate paradigms for spaCy tokens...
Code to generate all forms of the Latin verb `amo` in the present indicative active from a single form.
120
Patrick J. Burns @diyclassics.bsky.social · 14/04/2026
Announcing—LatinCy Lexicon v0.1, a refactored version of Whitaker's Words that uses LatinCy annotations to disambiguate words/meanings. Can be added as a custom component to any LatinCy pipeline. github.com/latincy/lati... #digiclass #nlproc
LatinCy Lexicon logo
3218
Patrick J. Burns @diyclassics.bsky.social · 26/03/2026
✨ LatinCy v3.9 sm/md/lg/trf pipelines for SpaCy available ✨ - Improved tokenization and u/v norm - New custom Latin-specific XPOS tags - Better, more consistent lemma/morph coverage huggingface.co/latincy/la_c... #digiclass #nlproc
"Album" cover for the LatinCy v3.9 pipelines with "catus" from Gesner's 1551 Historia Animalium
0137
Patrick J. Burns @diyclassics.bsky.social · 12/03/2026
Attending the »AI & the Study of Antiquity« conference @ Rutgers University today and tomorrow...
Program for AI & the Study of Antiquity at Rutgers March 12 & 13, 2026
010
Patrick J. Burns @diyclassics.bsky.social · 06/03/2026
Wanted an easier way to preview CONLLU files in vscode, couldn't find one, worked up my own...
Side-by-side conllu file in plaintext vs. formatted preview on an example sentence
152
Patrick J. Burns @diyclassics.bsky.social · 26/02/2026
Model drop! Some (beta!) LatinCy releases ahead of Friday's dev meeting/"birthday" party, trained on same data as spaCy models… - LatinCy Stanza huggingface.co/latincy/la_s... - LatinCy UDPipe huggingface.co/latincy/la_u... - LatinCy Flair huggingface.co/latincy/la_f... #nlproc #digiclass
Comparative metrics for different LatinCy models—spaCy but also Stanza, Flair, UDPipeLatinCy @ 3 graphic with birthday cake
032
Patrick J. Burns @diyclassics.bsky.social · 25/02/2026
Super-experimental at this point—but I am now embedding a local LLM inside Prodigy to assist with NER annotations... 1. Using `correct` recipe, LatinCy model tags likely entities 2. Optional "Ask LLM" button queries Mistral based on text/tags 3. Add'l RAG runs over the project's NER guidelines
Prodigy interface for LatinCy NER annotations... here Ov. Met. 2.5 and a PERSON label assigned to MulciberProdigy interface for LatinCy NER annotations... here Ov. Met. 2.5 and a PERSON label assigned to Mulciber; "Ask LLM" button queries an "NER Assistant" and provides more context
3153
Patrick J. Burns @diyclassics.bsky.social · 23/02/2026
Ever have a lot of all-u Latin text, ever need a superfast way to replace only the consonants, i.e. uerbum → verbum... new feature in latincy-preprocess v0.1... github.com/diyclassics/... #nlproc #digiclass
Python code for using latincy-preprocess to change u → v in Latin text, e.g. uerbum → verbum
072
Patrick J. Burns @diyclassics.bsky.social · 23/02/2026
Ever have like a million long-s errors in your Latin OCR, ever need a superfast way to correct them against a Latin character ngram model... new feature in latincy-preprocess v0.1... github.com/diyclassics/... #nlproc #digiclass #digiclafs
Python code demonstrating long-s correction using latincy-preprocess, e.g. "funt in fundamento reipublicae ftatua" → "sunt in fundamento reipublicae statua"
143
Patrick J. Burns @diyclassics.bsky.social · 13/02/2026
Trying to figure out how to offer new ways to work with the collections... Here is an example of combined Latin/Greek search in a single call. Uses regex for now—working on memory managment/speeding up the annotations for true lemma search etc.... github.com/diyclassics/...
LatinCy readers using both Latin and Greek corpus readers showing regex search for Latin and Greek versions of "Caesar"
010
Patrick J. Burns @diyclassics.bsky.social · 10/02/2026
I can add that LatinCy Reader pattern matching is (as shown here) flexible enough to take agreement into account—of course, assuming accuracy of the tagger & morpher; updated notebook here github.com/diyclassics/...
LatinCy annotations used to for noun-adjective agreement in Caesar (e.g. magnis itineribus, magnam partem).
010
Patrick J. Burns @diyclassics.bsky.social · 09/02/2026
Such an important point and really working on it these days—between higher visibility for different parts of the project, better documentation, more demos, etc. One example I can point to now is trying to make small single-use dashboards to show certain features... latincy.streamlit.app/~/+/
Streamlit dashboard demo of LatinCy Text Analyzer with sentence annotations from Seneca's Epistulae Morales
120
Patrick J. Burns @diyclassics.bsky.social · 09/02/2026
Searching for any form of "magnus" followed by any noun—just one example of flexible pattern matching possible with LatinCy Readers `find_sents` call... from this demo notebook: github.com/diyclassics/...
Demo of LatinCy Readers `find_sents` method; returns five examples of 'magnus' + NOUN in a Latin text collection.
2114
Patrick J. Burns @diyclassics.bsky.social · 05/02/2026
Program announced & RSVPs open for the LatinCy Developers/Users Meeting on Feb. 27, 2026, a remote event via @isawlibrary.bsky.social Registration link at diyclassics.github.io/latincy2026/ #digiclass #nlproc
043
Patrick J. Burns @diyclassics.bsky.social · 02/02/2026
Announcing—LatinCy Readers v.1.0.2, i.e. LatinCy-powered corpus readers for Latin text collections. Quickly get sentences, lines, words annotated for lemma, POS, morphology, NER, etc. Supporting .txt, .xml, .tess, .conllu, and more. github.com/diyclassics/... #digiclass #nlproc
LatinCy Readers logo
22916
Patrick J. Burns @diyclassics.bsky.social · 08/01/2026
Stop by the ISAW Monographs / NYU Press table in the #AIASCS Book Exhibit—let’s talk about texts, archaeology, art history, and material culture inter alia isaw.nyu.edu/publications...
Patrick Burns standing with the ISAW Monographs book tableRecent publications from the ISAW Monographs series
041
Patrick J. Burns @diyclassics.bsky.social · 07/01/2026
Headed to SF for #AIASCS — presenting Saturday morning (10 Jan. @ 8am, DCA panel) on the Latin content of LLM training data repos, massive-scale philology, and classics-comp sci collaboration w. D. Bamman, C. Brooks, M. Hudspeth & B. O’Connor #digiclass #nlproc
Title slide from SCS 2026 talk called “Recovering 34 Billion Latin Words from AI Training Data”
0103
Patrick J. Burns @diyclassics.bsky.social · 15/12/2025
Reminder that we are hosting the first LatinCy Developers/Users Meeting this February (and it's also a third "birthday" for the pipelines!)...
Graphic that read "la [birthday cake emoji] 3"
011
Patrick J. Burns @diyclassics.bsky.social · 15/12/2025
Have a chapter on "quantitative dialogism" in this open-access Brill volume on »Direct Speech in Greek and Latin Epic« brill.com/display/book...
012
Patrick J. Burns @diyclassics.bsky.social · 16/07/2025
Presenting at #ianls25 on… “Neo-Latin as a Pragmatic Source of Language Model Training Data” This morning (16 Jul) @ 10:30am, Egger 011
Program contents for IANLS DIGITAL TECHNOLOGY IV Special Session Chair: Carolin A. Giere
10.00 am - 10.30 am GRABSKA-GRADZINSKA Iwona and WIATER Mateusz "Neolatina Sarmatica - How Programming Languages can Revive Interest in Latin Texts"
10.30 am - 11.00 am BURNS Patrick "Neo-Latin as a Pragmatic Source of Language Model Training Data"
11.00 am - 11.30 am CECCHINI Flavio Massimiliano "The Tongueprint: Assessing and Quantifying Bilingualism in Renaissance Neo-Latin Texts"
042
Patrick J. Burns @diyclassics.bsky.social · 19/02/2025
Giving a talk tomorrow (Th. 2/20 @ 4:30pm Eastern) at U. Cincinnati’s Taft Center called “The Digital Afterlife of a Dead Language: Or Recovering 34 Billion(!) Latin Words from AI Training Data”. Talk will be over Zoom as well—link to follow. There will be unicorns! diyclassics.github.io/afterlife/
Woodcut of a unicorn from Gesner’s Historiae Animalium
010
Patrick J. Burns @diyclassics.bsky.social · 04/01/2025
Talking about LLM prompts and Latin teaching/research on this morning’s DCA panel at 8am (Sat. 1/4) in Salon K/Hybrid #aiascs #scs40
8:00am-
10:30am, Salon K
(Hybrid)
SCS-40: HYBRID: Opening
Up Classics with Al (organized by the Digital
Classics Association)
Neil Coffee, University at Buffalo, SUNY, Organizer
1. Neil Coffee, University at
Buffalo, SUNY
Introduction
2. Samuel Huskey, University of Oklahoma Opening Up Bottlenecks in Digital Classics
Workflows with Human-in-the-Loop Al
3. Patrick Burns, New York
University
Prompt Engineering for
Latin Teachers4. Edward Ross, University of Reading, and Jackie Baines, University of Reading
Generative Image Al and Teaching Classics: A Case of Exaggeration
5. Gregory Crane, Tufts
University
Al, Machine Actionable
Publication and
Assigning Credit
6. Joseph Dexter, Harvard University, and Pramit Chaudhuri, University of Texas at Austin Benchmarking
Generative Al Models for
Classical Literary
Criticism
062
Patrick J. Burns @diyclassics.bsky.social · 03/06/2024
New article in NECJ on the relationship between computational chat and active Latin: crossworks.holycross.edu/necj/vol51/i...
Abstract for "(Re)active Latin" article in new issue of NECJ
022
Patrick J. Burns @diyclassics.bsky.social · 09/08/2023
Introducing LatinCy... Trained spaCy pipelines for Latin NLP... 📦Models: huggingface.co/latincy 🌌Universe: spacy.io/universe/project/latincy 📝Preprint: arxiv.org/abs/2305.04365 #NLProc #DigiClass
Card for LatinCy at spaCy Universe website, inc. code example for how to install and run the pipeline on a sample Latin sentence.
091