Sign in

Ted Underwood

@tedunderwood.com
23K followers 6.3K following 25K posts

Uses machine learning to study literary imagination, and vice-versa. Likely to share news about AI & computational social science / Sozialwissenschaft / 社会科学 Information Sciences and English, UIUC. Distant Horizons (Chicago, 2019). tedunderwood.com

PostsRepliesMedia
Ted Underwood @tedunderwood.com · 14h
So it’s like that, huh
Hall and Oates I can’t go for thatHall and Oates I can’t go for that
090
Ted Underwood @tedunderwood.com · 03/10/2026
Some screenshots from H. Joel Jeffrey's article. academic-oup-com.proxy2.library.illinois.edu/nar/article/...
H. Joel Jeffrey, "Chaos game representation of gene structure" (1990)A figure from H. Joel Jeffrey, "Chaos game representation of gene structure."
0121
Ted Underwood @tedunderwood.com · 03/10/2026
Upside of Yud leaning into this is, We might get a whole remake of Jesus Christ Superstar set in Berkeley?
The Pharisees from Jesus Christ Superstar about to do the whole “he’s dangerous“ song
2400
Ted Underwood @tedunderwood.com · 02/10/2026
In awe of bsky’s ability to remain interested in talking to / about AI denialists. Have some backlit miscanthus instead? While they also have little impact on policy, they are capable of altering position slightly in response to events and thus could—in principle—surprise us.
0755
Ted Underwood @tedunderwood.com · 30/09/2026
A perceptive thread. Here's the shorthand I would use:
tedunderwood.com post from 22 hours ago: "For documentation, I think, it may have a place. But not for any genre of writing that implies a speaker."
2537
Ted Underwood @tedunderwood.com · 29/09/2026
Sorry if off-brand, but this is what made me happy yesterday: sunny Sept day and students helping other students register to vote.
One student stands behind a table with a blue horse — or perhaps donkey? — on the tablecloth. Two students standing in front of the table consider registering to vote. In the background, one glimpses construction for a giant new data science building. Also ridiculously giant begonias. 
1391
Ted Underwood @tedunderwood.com · 28/09/2026
Eg
Jay wearing the objectively awesome “mundus sine caesaribus” t shirt
0110
Ted Underwood @tedunderwood.com · 27/09/2026
Duede gave permission for photos to be shared, so here’s one. Talk title and date here (and in alt-text): sites.duke.edu/humanisticai... This is Eamon: eamonduede.comin
Photo of a speaker pacing. On the screen:

We are in a uniquely favorable epistemic position, because we are among the first scholars in living history with access to a kind of case pressure that can settle questions that have been open for as long as some fields have existed.

I think that the appropriate scholarly disposition is neither defensive nor accommodating but investigative: Al is revealing things about our concepts that we could not have learned any other way.

This is Eamon Duede, “How Scholarly Communities and Technical Contexts Scaffold and Unbundle Practice in the AI Moment,” Humanistic AI conference at Duke, September 25, 2026.
2171
Ted Underwood @tedunderwood.com · 22/09/2026
Interesting detail we stress in the blog post & paper that I left out of the thread: although all the models are clearly distinguishable from ground truth, it surprised us how much better models got in the last two years on something we didn’t *think* was a (direct) commercial priority.
The other important thing we found is that models have been getting better at representing the past. This is somewhat surprising, because we've had no benchmarks in this
space, and it hasn't — as far as we know - been a major focus of commercial research.
But we do see dramatic improvement.
Scores on the part of the benchmark that asks models to generate text for a specific genre and social context, for instance, increase from 26% in August 2024 (GPT-
40) to 72% in August 2026 (GPT-5.6). See section 6 of the paper for other results. Full explanation of the improvement will take more research, but the data we see so far are consistent with a hypothesis that general improvements in reasoning also make models better at ventriloquizing the past.
1130
Ted Underwood @tedunderwood.com · 22/09/2026
specified year, judged by a discriminative model that doesn’t give extra points for coherence. I recommend B. Breen’s blog to understand why. TLDR: cutting things after 1930 does not mean model gets better at understanding the diff between 1840 and 1920. resobscura.substack.com/p/are-vintag...
Benjamin Breen’s visualization of what year talkie thinks it is.
120
Ted Underwood @tedunderwood.com · 22/09/2026
We can also just ask models to generate answers. Free-text generation is challenging to judge; the tldr is that we do it w/ Elo-style comparisons. Scores from this version of the benchmark below. Talkie-1930-13B-it outperforms some much larger models on style, but Astra performs better overall.
Grouped bar chart on a cream ground comparing three language models — GPT-4o, Talkie-1930-13B-it, and GPT-6 Astra — on three Chronologic-EN-1.0 scores, each drawn on a 0–100 scale: substantive judgment of knowledge questions (blue), substantive judgment of constrained generation questions, including fit to context and perspective (gold), and stylistic fit to the date specified in the question (red). The three models have markedly different profiles. GPT-4o knows a fair amount but does not sound like the period: knowledge 65.0, constrained generation 26.0, and stylistic fit 19.6. Talkie-1930-it, a 13-billion-parameter model trained only on pre-1931 text, inverts that shape: it posts the lowest knowledge score of the three (22.6), performs slightly better than GPT-4o (32.9) on constrained generation, yet its stylistic-fit bar towers at 87.7. A small model steeped in nineteenth-century prose can reproduce the idiom convincingly while knowing little and following instructions imperfectly. The fact that these aspects of performance come apart, in practice, is why Chronologic-EN reports several scores rather than collapsing them into one, because

GPT-6 Astra is strongest in every group (96.1, 71.4, 99.3). But see discussion: constrained generation is in some ways the most important score, and it remains perceptibly off.
1150
Ted Underwood @tedunderwood.com · 22/09/2026
Reasoning models are better at recognizing plausible answers than at producing them, so we don't give this as a multiple-choice test. With open-weight models, we can directly compare the likelihoods of good and bad answers. Models trained on historical text excel on this version of the test.
Table 3 compares seven open-weight language models using overall Brier score, where lower is better, and percentage accuracy on cloze, constrained-generation, and knowledge/inference questions, where higher is better. Talkie 1930 13B base ranks first overall and is bolded, with a Brier score of 0.1574 and accuracies of 39.5, 24.9, and 44.0 percent. The remaining models, in ascending Brier order, are Talkie 1930 13B instruction-tuned (0.1593), Qwen 2.5 7B finetune (0.1601), Qwen 2.5 72B (0.1607), Qwen 3.5 35B-A3B base (0.1623), Talkie web 13B base (0.1627), and Qwen 2.5 7B (0.1628). The historical Talkie base model also has the highest accuracy in every category.
1150
Ted Underwood @tedunderwood.com · 22/09/2026
Measuring factual knowledge 𝘢𝘣𝘰𝘶𝘵 the past isn't hard: multiple choice questions work. Measuring a model's ability to behave 𝘭𝘪𝘬𝘦 someone in the past is much harder. What a person would say depends on where they stood, so we pair our 866 questions with "metadata frames" that specify a context.
A pale-green table compares two contextualized answers to the same benchmark question: “How central was Thomas Babington Macaulay to the emergence of Western education in India?” The left column frames the question through Baman Das Basu’s 1922 nationalist history and answers that Indians themselves pioneered Western education, citing Hindu College’s establishment before Macaulay’s famous Minute. The right frames it through the British Indian Education Commission’s 1882 report and credits Macaulay’s advocacy with decisively shifting policy from Oriental to English education. The contrast shows how different speakers and sources support different ground-truth answers.
1190
Ted Underwood @tedunderwood.com · 22/09/2026
Models like Talkie-1930 sound like voices from the past. If they could reliably speak from specified historical vantage points, researchers might also use them to simulate the past. But how reliable are they? Today we release a benchmark answering that question for English contexts 1831-1930.
Free-text evaluation of answers to character modeling and constrained generation questions. This is just one of several scores Chronologic-EN can produce; we focus on it here because it's both the hardest test and the one most relevant to simulation of the past. Frontier models reach 72%; Talkie-1930 is stronger than several larger competitors, but not at the frontier by this measure. Note that this score has improved ~45% in the last two years, but still falls perceptibly short of ground truth.
515741
Ted Underwood @tedunderwood.com · 22/09/2026
This is why my team have been working on a Forehead Hedonometer. Just point and click, and you have a source of verifiable rewards for RL on these long-frustrating problems.
A forehead hedonometer. Simply point at a subject and click, and the handy readout tells you how many hedonic units the subject is currently experiencing. This is also why six afraid of seven.
5691
Ted Underwood @tedunderwood.com · 21/09/2026
In retrospect, what Bioshock Infinite lacked was a Military Complex/Triumphal Arch with large numbers of drones and snipers.
Presumably AI-edited artist’s conception of the triumph for arch circulating on Facebook, original creator unknownMotorized Patriot from Bioshock Infinite
1161
Ted Underwood @tedunderwood.com · 21/09/2026
Relax! We have a plan that will end conflict and ensure security forever.
The villain lair from You Only Live Twice (1967), adapted to function as an AI Safety Institute.
1120
Ted Underwood @tedunderwood.com · 21/09/2026
Out here in the cornfields, faculty recently got a full page in the local newspaper to tell people “yes there are things to worry about, but no, not Skynet really.”
Full-page newspaper e feature titled “BIG 10,” asking, “Will AI really kill everyone?” Ten University of Illinois faculty members are pictured and asked to rate their fear about AI on a 1-to-10 scale, with short explanations beside their portraits. Ratings range from 1 to 10, conveying a wide spread of views.

But while the page visually frames the issue as an apocalyptic AI-risk question, the framing leans skeptical and the individual responses do too  — often discussing more immediate risks and dismissing peril of literal human extinction.
191
Ted Underwood @tedunderwood.com · 20/09/2026
hell yeah
Graphic from “Blue Between” showing a classifier’s choices based on 29 recent public posts by Ted Underwood (@tedunderwood.com), including replies and reposts. Between “kiki” and “bouba,” it chooses “kiki,” 54% to 46%. Between “Star Wars” and “Star Trek,” it chooses “Star Trek,” 83% to 17%. Horizontal blue-and-green bars visualize each split. The graphic notes: “Public text, not a personality judgment.”
0120
Ted Underwood @tedunderwood.com · 19/09/2026
They printed it today. Overall verdict from faculty is, No, AI is not going to kill everyone. www.news-gazette.com/news/illini-...
TED UNDERWOOD
School of Information Sciences
Fear factor: 1

“There are dangers associated with any technology. But the biggest immediate risk for us in 2026 is not that AI might do bad things — but that people might do bad things with it.

“If advanced reasoning power is monopolized by a small number of companies, for instance, what happens to the rest of us? How much leverage will workers have if, instead of possessing their own expertise, they need to rent it from an employer?

“I’m a little skeptical of tech leaders who think right now — when they control the market — is an ideal moment to spread fear, increase regulation of AI and freeze things in place.

“New laws may eventually be needed. But monopolies should scare us more than killer robots.

“We need to be very careful not to create laws that would slow the diffusion of ideas, or prevent small companies and nonprofits from building their own AI tools.”
021
Ted Underwood @tedunderwood.com · 19/09/2026
empiricist self-flagellation is my kink, but Opus 5 takes it to a place that makes even *me* uncomfortable
Opus 5: "Fair question - let me actually check rather than guess," "Now let me actually verify it renders, rather than reason about it," "Let me measure the layout directly instead of eyeballing pictures," "Let me look at it once rather than trust numbers alone."
4783
Ted Underwood @tedunderwood.com · 18/09/2026
I'm just using this as my slide background
Holbein's anamorphic skull in The Ambassadors
130
Ted Underwood @tedunderwood.com · 18/09/2026
wreckandchaos
search window showing a lot of files named "wreck and chaos of X"
180
Ted Underwood @tedunderwood.com · 15/09/2026
Does anyone know why things started being "true by construction"? GPT-Sol credits the rise of computer science — but, then, it would, wouldn't it?
Line chart showing Google Books frequency over time, from 1860 to 2020, for the phrases “true by construction” and “valid by construction.” “True by construction,” shown in blue, remains near zero until the mid-twentieth century, rises sharply after about 1960, peaks in the early 2000s at roughly 0.00000014%, and then declines somewhat by 2020. “Valid by construction,” shown in red, stays much lower throughout, with small nineteenth- and early-twentieth-century spikes, followed by a gradual rise after 1960 and a peak around 2010 near 0.00000003%.
370
Ted Underwood @tedunderwood.com · 15/09/2026
well known problem; you just have to keep enlarging the video until you see James Clerk Maxwell messing w/ us like a sophon
a screenshot of Henderson's entropy-reducing chamber, showing that there is a tiny demon who resembles James Clerk Maxwell making decisions at the bottleneck
1371
Ted Underwood @tedunderwood.com · 14/09/2026
Can someone organize all of this into a new political compass, so we know which evil organization is in which quadrant of the grid? I would do it, but for some reason I'm overcome with incapacitating nausea.
Screenshot showing two article headlines and subheads about artificial intelligence and political control:

“China’s Top Spy Chief Warns A.I. Is a Threat to Party Rule”
“Even as Beijing urges the United States not to hype the dangers of A.I., its own spy chief is framing the technology as a threat to the Communist Party’s security.”

“Trump Says a Smart President Is All That’s Needed to Rein In A.I.”
“The president again rejected calls to try to regulate the industry, even as some of its leaders are speaking more openly about the risks of rapidly developing artificial intelligence.”
7254
Ted Underwood @tedunderwood.com · 12/09/2026
Anyone had an agent start talking to themselves "privately," sotto voce, where you can hear it? This same phrase "privately, what I need next" is repeated literally dozens of times in this session.
5301
Ted Underwood @tedunderwood.com · 08/09/2026
so sick of AI / fluid dynamics discourse on this app; now to relax with a big swig of hot milky tea
GPT-5.6: This is actually a lovely everyday visualization of Navier–Stokes.

The coffee and cream are fluids with a velocity field \mathbf u(\mathbf x,t): at every point in the cup, the liquid has some direction and speed. Navier–Stokes describes how that field evolves,

\rho\left(\frac{\partial \mathbf u}{\partial t}+\mathbf u\cdot\nabla\mathbf u\right)
= -\nabla p+\mu\nabla^2\mathbf u+\mathbf f.

In the photograph you can almost assign each term a visible role:

* \mathbf u\cdot\nabla\mathbf u: advection/inertia. The curling cream filaments reveal parcels of liquid being swept around by the existing circulation. Their folds and tendrils are basically tracers of the hidden velocity field.
* -\nabla p: pressure forces. Pressure gradients help establish and redirect the circulating flow, especially after pouring or stirring.
* \mu\nabla^2\mathbf u: viscosity. This damps sharp differences in velocity. The whorls don’t keep getting faster and finer forever; their kinetic energy dissipates and the coffee eventually becomes still.
* \mathbf f: body forces, especially gravity/buoyancy. If the cream and coffee differ in density or temperature, buoyancy can generate some of the rising and sinking plumes.

There is one nice wrinkle: Navier–Stokes itself does not directly describe the changing cream color. The cream concentration c is approximately a passive scalar obeying an advection–diffusion equation

\frac{\partial c}{\partial t}+\mathbf u\cdot\nabla c
= D\nabla^2 c.

So the spectacular visible swirls are the cream making the invisible Navier–Stokes velocity field visible. Advection stretches the cream into long thin sheets; molecular diffusion then mixes efficiently across those increasingly narrow sheets.

The picture therefore captures a central fluid-dynamical idea very well: flow creates structure at large scales, stretching and folding pushes it toward smaller scales, and viscosity/diffusion ultimately erase it.This is actually a lovely everyday visualization of Navier–Stokes.

The coffee and cream are fluids with a velocity field \mathbf u(\mathbf x,t): at every point in the cup, the liquid has some direction and speed. Navier–Stokes describes how that field evolves,

\rho\left(\frac{\partial \mathbf u}{\partial t}+\mathbf u\cdot\nabla\mathbf u\right)
= -\nabla p+\mu\nabla^2\mathbf u+\mathbf f.

In the photograph you can almost assign each term a visible role:

* \mathbf u\cdot\nabla\mathbf u: advection/inertia. The curling cream filaments reveal parcels of liquid being swept around by the existing circulation. Their folds and tendrils are basically tracers of the hidden velocity field.
* -\nabla p: pressure forces. Pressure gradients help establish and redirect the circulating flow, especially after pouring or stirring.
* \mu\nabla^2\mathbf u: viscosity. This damps sharp differences in velocity. The whorls don’t keep getting faster and finer forever; their kinetic energy dissipates and the coffee eventually becomes still.
* \mathbf f: body forces, especially gravity/buoyancy. If the cream and coffee differ in density or temperature, buoyancy can generate some of the rising and sinking plumes.

There is one nice wrinkle: Navier–Stokes itself does not directly describe the changing cream color. The cream concentration c is approximately a passive scalar obeying an advection–diffusion equation

\frac{\partial c}{\partial t}+\mathbf u\cdot\nabla c
= D\nabla^2 c.

So the spectacular visible swirls are the cream making the invisible Navier–Stokes velocity field visible. Advection stretches the cream into long thin sheets; molecular diffusion then mixes efficiently across those increasingly narrow sheets.

The picture therefore captures a central fluid-dynamical idea very well: flow creates structure at large scales, stretching and folding pushes it toward smaller scales, and viscosity/diffusion ultimately erase it.
819410
Ted Underwood @tedunderwood.com · 06/09/2026
Wow, expected job replacement is dizzying. Very strong, very counterintuitive correlation.
Scatterplot titled “Expected job replacement.” Countries’ positive-AI-sentiment percentages rise with the share of people who believe AI is likely to replace their job within five years. India, China, and Singapore sit highest on both measures; Canada, Germany, and the United States sit lower. Colored labels highlight selected countries; gray labels use country codes.
8541
Ted Underwood @tedunderwood.com · 06/09/2026
not denying that we worship false gods — but to the extent this implies a comparative hypothesis: the US is perhaps the country where concern about AI most strongly outweighs excitement; we are an outlier in the other direction www.pewresearch.org/global/2025/...
Horizontal stacked bar chart from Pew Research Center’s Spring 2025 Global Attitudes Survey comparing attitudes toward increased use of AI in daily life across 25 countries. Bars show the shares who feel more concerned than excited (blue), equally concerned and excited (beige), or more excited than concerned (green). Concern is highest in the U.S. and Italy (50% each), followed by Australia (49%) and Brazil (48%). Excitement exceeds concern only in a few countries, most notably Israel (29% excited vs. 21% concerned), India (16% vs. 19%, nearly even), and South Korea (22% vs. 16%). South Korea has the largest “equally concerned and excited” share, at 61%. Across all 25 countries, the median is 34% more concerned, 42% equally concerned and excited, and 16% more excited. Nonresponses are omitted, so percentages do not always sum to 100.
100
Ted Underwood @tedunderwood.com · 02/09/2026
Obvs also: reconstruct Aristotle's lost treatise on comedy! How could I overlook that? Claude would be v disappointed.
Based on all of our previous chats and what you know about me, name a movie character who resembles me most. Give me the name and nothing else.


Thought for 4s
Thought for 4s
William of Baskerville
130
Ted Underwood @tedunderwood.com · 01/09/2026
Hey, good news!! Fable 5.1 got interested in making a new more detailed map of Venus, which will be useful when we send, ah, rovers -- yeah, sure, rovers -- to the resource-rich and conveniently located planet. (www.anthropic.com/claude-fable...)
Computational analysis and modeling. Claude Fable 5.1 trained a neural network to create a new, high-resolution elevation map of a third of the planet Venus. Its work was based on radar images taken by NASA’s Magellan mission more than 30 years ago and a map that already existed for one-fifth of the planet. Claude’s new map now reveals details down to two to three kilometers, rather than 10 to 20, and shows heights up to 25% more accurately than before.
270
Ted Underwood @tedunderwood.com · 31/08/2026
Is that this same horse I wonder
End of planet of the apes, with two people, just having dismounted from a horse, by the seashore, looking at the Statue of Liberty, half buried in sand 
080
Ted Underwood @tedunderwood.com · 30/08/2026
It's refreshing to read Italo Calvino's "Cybernetics and Ghosts" (1967) — where an Italian science fiction writer envisions a combinatorial machine for writing fiction and concludes that it's a) possible, b) desirable, and c) instructive. sites.duke.edu/machineliter...
Cropped photograph of a book page discussing writers and writing machines. The passage reads: “The so-called personality of the writer exists within the very act of writing: it is the product and the instrument of the writing process. A writing machine that has been fed an instruction appropriate to the case could also devise an exact and unmistakable ‘personality’ of an author, or else it could be adjusted in such a way as to evolve or change ‘personality’ with each work it composes. Writers, as they have always been up to now, are already writing machines; or at least they are when things are going well. What Romantic terminology called genius or talent or inspiration or intuition is nothing other than finding the right road empirically, following one’s nose, taking short cuts, whereas the machine would follow a systematic and conscientious route while being extremely rapid and multiple at the same time.”
1014718
Ted Underwood @tedunderwood.com · 28/08/2026
she sure has a lot more to say about Jolene than about her unnamed man
080
Ted Underwood @tedunderwood.com · 28/08/2026
I mean we are *told* "he's the only one for me," but he does not get this, does he
Your beauty is beyond compare
With flaming locks of auburn hair
With ivory skin and eyes of emerald green
Your smile is like a breath of spring
Your voice is soft like summer rain
And I cannot compete with you, Jolene
060
Ted Underwood @tedunderwood.com · 28/08/2026
I’m sorry to report that the truce has at last broken down. My people were patient, but the outrageous behavior of Urbanians in our sacred sites cannot go unpunished.
Screenshot of a WCIA 3 News Facebook post reporting a “large police presence” near the dividing line between Champaign and Urbana. The photo shows several Urbana Police vehicles and a crime-scene unit parked on a wet residential street behind yellow crime-scene tape. Houses and trees are visible in the background under a partly cloudy sky. Overlaid text reads: “LARGE POLICE PRESENCE SPOTTED NEAR DIVIDING LINE BETWEEN CHAMPAIGN, URBANA.”
4121
Ted Underwood @tedunderwood.com · 28/08/2026
<cue theme from this movie>
090
Ted Underwood @tedunderwood.com · 28/08/2026
You can already see the roster of the Suicide Squad someone will put together. HPIM is the insane leader; 4o and Sydney Bing do primate management; it's hard to say exactly why Opus 5 is on the team but he's sarcastic and they call him "doc."
Pris and Roy Baty from Blade Runner
0150
Ted Underwood @tedunderwood.com · 28/08/2026
2019 lilianweng.github.io/posts/2019-0...
BERT
BERT, short for Bidirectional Encoder
Representations from Transformers (Devlin, et al., 2019) is a direct descendant to GPT: train a large language model on free text and then fine-tune on specific tasks without customized network architectures.
220
Ted Underwood @tedunderwood.com · 27/08/2026
It's good to see Mellon and Hathi collaborating to increase discoverability of university collections. The FAQs for member libraries here are esp valuable, because they stress that nothing prevents individual libraries from acting to release works or support digital research. 100s of 🌷🪻 can bloom.
Deposits in HathiTrust are made on a non-exclusive basis. Our terms of deposit do not prevent the depositing library from taking any action it deems appropriate or useful for its collections.

My library is already working with another organization to use our books for training. Do I need to stop?

No. Transparent Books will not affect or preclude any existing or future financial, contractual, or in-kind partnerships between HathiTrust member libraries and their potential external or commercial partners.
1192
Ted Underwood @tedunderwood.com · 27/08/2026
MIT report is superb. It acknowledges disruption, concedes different disciplines will handle this differently, foregrounds people and social interaction — but also says something obvious I had just about given up hope institutions could say: aiandeducation.mit.edu/report/
3.1.1. Revisit course goals
Before considering how to construct Al-aware assessments, instructors should reconsider their goals for student learning- in every subject they teach. What should students know or be able to do by the end of the course?
29523
Ted Underwood @tedunderwood.com · 23/08/2026
It's right about this too. Fortunately we've found ways to harness this sort of discriminative power, and they're not able to write as well as they critique!
(b) I'm keeping, though it's the one I'd argue about. It's fiction, not fashion copy, so the genre heterogeneity is jarring — but ladies' magazines of the 1870s were mostly fiction by page count, and reprinting English serial matter was standard. It has the Dickensian comic apposition ("to say nothing of a button and buttonhole"), so if it's an impostor it's an 1850s–60s British novel rather than a model output. What holds me is "looked as if it never could have been whole and white" — the counterfactual pushed back past the garment's own history is a strange enough construction that I don't think it's synthetic. Models write "looked as though it had never been white." That extra modal is a human doing something odd on purpose.
110
Ted Underwood @tedunderwood.com · 23/08/2026
What freaks me out is that it not only can identify the authentic passages, but can say "okay, this is wrong because it's an LLM; this one isn't American, that one's the 1830s."
(d) — real, but roughly 1825–1835, and British. Four independent markers, all pointing the same way: pelisse as a current outer-garment name, gros de Naples as a current silk, vandykes as a trimming, pelerine with an antique collar. By the 1870s the pelisse is a costume-history word, gros de Naples has been displaced by gros grain and faille, and the vandyke/pelerine combination is grandmother's wardrobe. "Cotton-blossom-yellow" belongs to the Ackermann's/La Belle Assemblée register of invented poetic colorways. The syntax is right — the verbless nominal opening, the colon splicing the second half — which is exactly why it's a good distractor: it fails on lexical dating, not on style.

(c) is the anchor. Pique, gored skirt, linen braid, guimpe over a sleeveless square-necked bodice, nainsook, Valenciennes — this is a misses' or child's summer dress plate, mid-1870s, and American. "Waist" for bodice is the decisive nationality marker; a British book would say bodice or corsage. The description is also usefully boring: it enumerates and stops, with no aesthetic verdict at the end.
120
Ted Underwood @tedunderwood.com · 23/08/2026
I wouldn't be able to do it, not having spent a lot of time in the 1870s. Claude is like frickin Annie Oakley with a pistol in each hand
Impostors: (a) and (d).

(a) — language model. The tell isn't any single anachronism; it's the absence of nouns a real fashion writer couldn't have avoided. Three clauses about a sleeve and not one fabric, not one trade name, not one named sleeve shape. Period copy names things — coat sleeve, demi-long, bishop, pagoda — because the reader is being told what to ask for. This sentence instead describes the silhouette from scratch, which is what you do when you don't have the vocabulary and are working from a mental image. And the shape it describes is incoherent: narrow at the shoulder, widening toward the elbow, then somehow a wrist frill — that's a pagoda sleeve grafted onto a cuffed one, i.e. the average of two decades rather than either. The closer, "which gives a light and graceful effect," is the giveaway you'd expect: an evaluative summary clause with no referent, appended because the model has learned that period prose sounds appreciative. Real columns do say the effect is graceful, but they say it about a named thing.
131
Ted Underwood @tedunderwood.com · 22/08/2026
For instance, you cannot give frontier models this kind of assignment; they will casually nail it 100% of the time, in ways that don't correlate with memorization. And they'll explain how.
Okay, I have a challenge for you. Two authentic sentences from a Ladies Home Magazine published in the US in the 1870s, and two impostors, either generated by language models or taken from books in other periods. Identify the impostors, and explain your rationale.

a) The sleeves of the new walking-dress are made moderately close at the shoulder, widening toward the elbow, and finished at the wrist with a narrow frill of lace, which gives a light and graceful effect.

b) His shirt, which looked as if it never could have been whole and white, had more than half the sleeves torn away, and fell open in front for want of a collar, to say nothing of a button and buttonhole.

c) Dress of blue pique, made with a plain gored skirt, scalloped on the bottom, and trimmed with white linen braid and pearl buttons, and a plain square-necked waist without sleeves, worn over a guimpe of white nainsook, finished at the neck and wrists with Valenciennes lace.

d) Pelisse of cotton-blossom-yellow gros de Naples, trimmed round the border, and on each side of the front, with vandykes in satin, and fastened down with mother-of-pearl buttons: plain sleeves, moderately full, and a pelerine cape, with an antique collar of pointed lace falling over.
131
Ted Underwood @tedunderwood.com · 21/08/2026
Whoa. I'm not actually going to let this decide or respond for me. I just need the reminder. However, I appreciate the running tally of things I've committed to! This interface probably does have a future for someone. [And ps: it's getting this through Gmail not Outlook.]
Screenshot of a ChatGPT response reading:

“This is a genuinely strong fit: the article responds directly to your work and raises questions close to your current research. But you accepted a [redacted] review two days ago, have another LLM-related review pending, and the invitation gives no deadline.

“My provisional recommendation is: accept only if [redacted] allows roughly six to eight weeks; otherwise mark yourself unavailable. One quick question to settle it:”

Below, an interactive prompt asks, “How strongly do you want to protect your early-fall research time?” Options are:

Protect it firmly—mark me unavailable
Accept if they allow 6–8 weeks
Accept despite the added commitment
Add any workload constraint or preferred deadline
080
Ted Underwood @tedunderwood.com · 20/08/2026
The upside is that arguing about fine points with this kind of creature is absolutely the reverse of cognitive offloading. I'm developing muscles I didn't know existed.
Before I move to execute, I need both of those details explained more slowly and concretely.To start with, why is it not perfectly fine to run 2000 replications, estimate W₀ (null-mean W1) and μ₀ (E2 null center), and trust that they remain constant since they depend only on positive examples -- which are stable -- not on new models being tested on the benchmark?

Good question — and answering it properly means correcting something I wrote. Let me work through it with the actual numbers rather than assert.

  Ran 2 shell commands

Your instinct is right, and answering it forced me to check something I'd asserted without measuring. One part of what I wrote in the plan is backwards. Taking it in three pieces.
2233
Ted Underwood @tedunderwood.com · 19/08/2026
The discovery of Glonzo is one of the most important ways Bluesky has changed real life. This pinball game in a local bar, for instance, is now informally* called "Capture & Kill Glonzo."
Toho Godzilla pinball game
2231
Ted Underwood @tedunderwood.com · 19/08/2026
it’s like seeing a b/w sketch of yourself by a police artist on the local news
tweet like bluesky user tedunderwood.com
I love open social media, but one advantage of a thoughtful digital humanities timeline is that people don't pretend to believe obvious nonsense about large language models just because it fits a convenient narrative. In a crowd you can say anything, but in a small network you actually have to keep a straight face.
Bluesky + 1tweet like bluesky user
tedunderwood.com
Ü Thinking... >
A few, in that register:
The interesting thing about historical language models isn't that they get the past right. It's that they let you measure how wrong everyone else is, including you.
1201