Sign in

Craig Messner

@cmessner.bsky.social
65 followers 122 following 20 posts

Humanities machine learning/NLP/language modeling @ JHU | cmessner.me

PostsRepliesMedia
Craig Messner @cmessner.bsky.social · 06/10/2026
\sbox0{\csname @firstofone\endcsname{\lowercase{\romannumeral-`0\noindent What ever do you mean? }}}% \kvset[mode=text,validate=true]\usebox0
050
Craig Messner @cmessner.bsky.social · 04/10/2026
No preprint at the moment (due to new father status) but a model here: No preprint (yet) due to new baby, but a model here: huggingface.co/Hplm/edgar-b... Self elicited IT/RL posttrained version to come, with an emphasis on binding to user context. Say hi at TADA if you happen to be going!
huggingface.co
Hplm/edgar-base · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
000
Craig Messner @cmessner.bsky.social · 04/10/2026
I'm currently located in the bay area, so I'm excited to be at #tada2026 tomorrow presenting a poster on synthetic data augmentation for domain-restricted humanities language modeling and its implications for lexical leakage. I really enjoyed my last Text as Data experience, happy to see it back
131
Reposted by Craig Messner
Alexander Doria @dorialexander.bsky.social · 01/10/2026
After a long wait, releasing the SYNTH paper! It’s not pretraining, mid-training or post-training, it’s just training: a fully synthetic single-stage pipeline to train workable reasoning models with unprecedented data efficiency. arxiv.org/abs/2609.378...
14115
Craig Messner @cmessner.bsky.social · 29/09/2026
Unforeseen but somewhat useful consequence of having a newborn: I'm back on eastern time
000
Reposted by Craig Messner
Ted Underwood @tedunderwood.com · 22/09/2026
Models like Talkie-1930 sound like voices from the past. If they could reliably speak from specified historical vantage points, researchers might also use them to simulate the past. But how reliable are they? Today we release a benchmark answering that question for English contexts 1831-1930.
Free-text evaluation of answers to character modeling and constrained generation questions. This is just one of several scores Chronologic-EN can produce; we focus on it here because it's both the hardest test and the one most relevant to simulation of the past. Frontier models reach 72%; Talkie-1930 is stronger than several larger competitors, but not at the frontier by this measure. Note that this score has improved ~45% in the last two years, but still falls perceptibly short of ground truth.
516344
Craig Messner @cmessner.bsky.social · 14/07/2026
Last year I was strapping together Wikidata, book repos, and library circulation pdfs to do this. Made some small (quantified) concessions due to time. Now I can reach the same confidence level again quickly -- and with functionally, the same methods. Great if you know what you need to do.
020
Craig Messner @cmessner.bsky.social · 08/07/2026
Insert Charli xcx "brat" edit here
110
Reposted by Craig Messner
Sil Hamilton @srhm.ca · 11/06/2026
@404media.co wrote about our new preprint on tell-tale signs of AI-generated stories! Cc @dmimno.bsky.social Paper: arxiv.org/abs/2605.26492 Article: www.404media.co/elias-thorne...
404media.co
Chatbots Keep Telling Stories About Lighthouse Keeper 'Elias Thorne'. We Might Know Why
LLMs including ChatGPT, Gemini and Claude are obsessed with telling stories about lighthouse keepers and clockmakers, and one character named 'Elias Thorne' has made his way from chatbots to Amazon bo...
11710
Reposted by Craig Messner
Ted Underwood @tedunderwood.com · 11/06/2026
was reading this and enjoying it — and then enjoyed it more when it turned out to be a paper by Sil Hamilton and @dmimno.bsky.social
2336
Reposted by Craig Messner
Ted Underwood @tedunderwood.com · 10/06/2026
For fans of Talkie-1930 and all the methodological questions raised by historical models, here's a new entrant into the field, TypewriterLM, trained up to 1913. Corpus, instruction-tuning datasets, and event dataset are released. arxiv.org/abs/2606.02991
arxiv.org
Pretraining Language Models on Historical Text
We introduce TypewriterLM, a 7.24B History language model (LM) trained exclusively on English text predating 1913. Developing History LMs requires addressing challenges in data quality and availabilit...
611320
Reposted by Craig Messner
philpax @philpax.me · 02/05/2026
excellent analysis of Talkie resobscura.substack.com/p/are-vintag...
resobscura.substack.com
Are "Vintage LLMs" the start of a new humanistic field?
Thoughts on Historical Language Models and Talkie-1930
2348
Reposted by Craig Messner
David Duvenaud @davidduvenaud.bsky.social · 28/04/2026
Announcing Talkie: a new, open-weight historical LLM! We trained and finetuned a 13B model on a newly-curated dataset of only pre-1930 data. Try it below! with @alecrad.bsky.social and @nicklevin01.bsky.social
25615
Craig Messner @cmessner.bsky.social · 28/04/2026
At this point tack a sign to my forehead that reads "no, LLMs aren't complete black boxes"
000
Craig Messner @cmessner.bsky.social · 25/04/2026
Mostly MatSci : www.google.com/maps/place/D...
google.com
Google Maps
Find local businesses, view maps and get driving directions in Google Maps.
000
Craig Messner @cmessner.bsky.social · 25/04/2026
It's possibly the worst building on campus. I spent all of last spring teaching in a classroom unceasingly beset by an HVAC system in terminal, noisy, decline.
100
Reposted by Craig Messner
JHU Computer Science @jhucompsci.bsky.social · 23/03/2026
“Pretraining Language Models for Diachronic Linguistic Change Discovery” by @tom-lippincott.bsky.social, @cmessner.bsky.social, & more shows that efficient pretraining techniques produce useful models over corpora too large for easy manual inspection and too small for “typical” LLM approaches: (3/5)
arxiv.org
Pretraining Language Models for Diachronic Linguistic Change Discovery
Large language models (LLMs) have shown potential as tools for scientific discovery. This has engendered growing interest in their use in humanistic disciplines, such as historical linguistics and lit...
121
Craig Messner @cmessner.bsky.social · 10/03/2026
Maybe a bit afield, but I always like telling students about Hugh Kenner's experiments with ngram modeling for literary style in the early 1980s (so roughly contemporaneous with Jelenik's adoption of them for ASR). As found in Byte magazine in 1984: archive.org/details/byte...
archive.org
Byte Magazine Volume 09 Number 12 - New Chips : Free Download, Borrow, and Streaming : Internet Archive
010
Craig Messner @cmessner.bsky.social · 07/03/2026
this is both predictable and informative
000
Craig Messner @cmessner.bsky.social · 07/03/2026
I regret to inform you that while LLMs may well automate numerous rote linguistic tasks they cannot as yet replace your teammates in CS2
110
Craig Messner @cmessner.bsky.social · 20/02/2026
New to me as well, will be using this in class this semester!
010
Craig Messner @cmessner.bsky.social · 16/02/2026
From my own local extrapolations, it feels like methods tied to classic quant dh/cultural analytics are getting broader play, and that "humanities machine learning" (which extends ML fields like interpretability, data-efficient training, evaluation and etc.) is emerging underneath
110
Craig Messner @cmessner.bsky.social · 13/01/2026
The real question is: who is winning?
000
Craig Messner @cmessner.bsky.social · 13/10/2025
Any leads on research that examines what disjoint exists between human-recoverable and SAE recoverable features out there? (Perhaps after: arxiv.org/html/2506.15..., which features are "naturally" distinguishable by humans but represented densely by SAEs, which end up in "noisy" dense reps.)
arxiv.org
Dense SAE Latents Are Features, Not Bugs
000
Craig Messner @cmessner.bsky.social · 09/10/2025
Anyone have on instruction tuning for models trained solely on historical data? Turns out texts from 1750 have very few "reddit-like" constructs.
030
Craig Messner @cmessner.bsky.social · 07/10/2025
Interestingly this reads to me a lot like a description of how close reading works in practice, especially post new-historicism -- "why this word here, knowing what we know about its contemporaneous use"
010
Craig Messner @cmessner.bsky.social · 09/09/2025
More importantly, its frustrations will hopefully serve as useful critical lessons. I also hereby claim the use of the name "Poetaster" for any further such systems!
000
Craig Messner @cmessner.bsky.social · 09/09/2025
I teach a machine learning class for students in traditionally less computational fields (cdh.jhu.edu/teaching/4/). Recent students had questions about RAG, so I used EmbeddingGemma's release as an excuse to put together an example in the form of a poetry criticism game github.com/messner1/poe...
github.com
GitHub - messner1/poetaster: Demo of a RAG-based educational game on an EmbeddingGemma/quantized Llama backbone
Demo of a RAG-based educational game on an EmbeddingGemma/quantized Llama backbone - messner1/poetaster
140