Sign in

Craig Messner

@cmessner.bsky.social
59 followers 115 following 19 posts

Humanities machine learning/NLP/language modeling @ JHU | cmessner.me

PostsRepliesMedia
Craig Messner @cmessner.bsky.social · 23h
I'm currently located in the bay area, so I'm excited to be at #tada2026 tomorrow presenting a poster on synthetic data augmentation for domain-restricted humanities language modeling and its implications for lexical leakage. I really enjoyed my last Text as Data experience, happy to see it back
121
Reposted by Craig Messner
Alexander Doria @dorialexander.bsky.social · 01/10/2026
After a long wait, releasing the SYNTH paper! It’s not pretraining, mid-training or post-training, it’s just training: a fully synthetic single-stage pipeline to train workable reasoning models with unprecedented data efficiency. arxiv.org/abs/2609.378...
14014
Craig Messner @cmessner.bsky.social · 29/09/2026
Unforeseen but somewhat useful consequence of having a newborn: I'm back on eastern time
000
Reposted by Craig Messner
Ted Underwood @tedunderwood.com · 22/09/2026
Models like Talkie-1930 sound like voices from the past. If they could reliably speak from specified historical vantage points, researchers might also use them to simulate the past. But how reliable are they? Today we release a benchmark answering that question for English contexts 1831-1930.
Free-text evaluation of answers to character modeling and constrained generation questions. This is just one of several scores Chronologic-EN can produce; we focus on it here because it's both the hardest test and the one most relevant to simulation of the past. Frontier models reach 72%; Talkie-1930 is stronger than several larger competitors, but not at the frontier by this measure. Note that this score has improved ~45% in the last two years, but still falls perceptibly short of ground truth.
515942
Reposted by Craig Messner
Sil Hamilton @srhm.ca · 11/06/2026
@404media.co wrote about our new preprint on tell-tale signs of AI-generated stories! Cc @dmimno.bsky.social Paper: arxiv.org/abs/2605.26492 Article: www.404media.co/elias-thorne...
404media.co
Chatbots Keep Telling Stories About Lighthouse Keeper 'Elias Thorne'. We Might Know Why
LLMs including ChatGPT, Gemini and Claude are obsessed with telling stories about lighthouse keepers and clockmakers, and one character named 'Elias Thorne' has made his way from chatbots to Amazon bo...
11710
Reposted by Craig Messner
Ted Underwood @tedunderwood.com · 11/06/2026
was reading this and enjoying it — and then enjoyed it more when it turned out to be a paper by Sil Hamilton and @dmimno.bsky.social
2336
Reposted by Craig Messner
Ted Underwood @tedunderwood.com · 10/06/2026
For fans of Talkie-1930 and all the methodological questions raised by historical models, here's a new entrant into the field, TypewriterLM, trained up to 1913. Corpus, instruction-tuning datasets, and event dataset are released. arxiv.org/abs/2606.02991
arxiv.org
Pretraining Language Models on Historical Text
We introduce TypewriterLM, a 7.24B History language model (LM) trained exclusively on English text predating 1913. Developing History LMs requires addressing challenges in data quality and availabilit...
611320
Reposted by Craig Messner
philpax @philpax.me · 02/05/2026
excellent analysis of Talkie resobscura.substack.com/p/are-vintag...
resobscura.substack.com
Are "Vintage LLMs" the start of a new humanistic field?
Thoughts on Historical Language Models and Talkie-1930
2348
Reposted by Craig Messner
David Duvenaud @davidduvenaud.bsky.social · 28/04/2026
Announcing Talkie: a new, open-weight historical LLM! We trained and finetuned a 13B model on a newly-curated dataset of only pre-1930 data. Try it below! with @alecrad.bsky.social and @nicklevin01.bsky.social
25615
Craig Messner @cmessner.bsky.social · 28/04/2026
At this point tack a sign to my forehead that reads "no, LLMs aren't complete black boxes"
000
Reposted by Craig Messner
JHU Computer Science @jhucompsci.bsky.social · 23/03/2026
“Pretraining Language Models for Diachronic Linguistic Change Discovery” by @tom-lippincott.bsky.social, @cmessner.bsky.social, & more shows that efficient pretraining techniques produce useful models over corpora too large for easy manual inspection and too small for “typical” LLM approaches: (3/5)
arxiv.org
Pretraining Language Models for Diachronic Linguistic Change Discovery
Large language models (LLMs) have shown potential as tools for scientific discovery. This has engendered growing interest in their use in humanistic disciplines, such as historical linguistics and lit...
121
Craig Messner @cmessner.bsky.social · 07/03/2026
I regret to inform you that while LLMs may well automate numerous rote linguistic tasks they cannot as yet replace your teammates in CS2
110
Craig Messner @cmessner.bsky.social · 13/10/2025
Any leads on research that examines what disjoint exists between human-recoverable and SAE recoverable features out there? (Perhaps after: arxiv.org/html/2506.15..., which features are "naturally" distinguishable by humans but represented densely by SAEs, which end up in "noisy" dense reps.)
arxiv.org
Dense SAE Latents Are Features, Not Bugs
000
Craig Messner @cmessner.bsky.social · 09/10/2025
Anyone have on instruction tuning for models trained solely on historical data? Turns out texts from 1750 have very few "reddit-like" constructs.
030
Craig Messner @cmessner.bsky.social · 09/09/2025
I teach a machine learning class for students in traditionally less computational fields (cdh.jhu.edu/teaching/4/). Recent students had questions about RAG, so I used EmbeddingGemma's release as an excuse to put together an example in the form of a poetry criticism game github.com/messner1/poe...
github.com
GitHub - messner1/poetaster: Demo of a RAG-based educational game on an EmbeddingGemma/quantized Llama backbone
Demo of a RAG-based educational game on an EmbeddingGemma/quantized Llama backbone - messner1/poetaster
140