Ted Underwood @tedunderwood.com · 03/10/2026Some screenshots from H. Joel Jeffrey's article. academic-oup-com.proxy2.library.illinois.edu/nar/article/... 0121
Ted Underwood @tedunderwood.com · 03/10/2026Upside of Yud leaning into this is, We might get a whole remake of Jesus Christ Superstar set in Berkeley? 2400
Ted Underwood @tedunderwood.com · 02/10/2026In awe of bsky’s ability to remain interested in talking to / about AI denialists. Have some backlit miscanthus instead? While they also have little impact on policy, they are capable of altering position slightly in response to events and thus could—in principle—surprise us. 0755
Ted Underwood @tedunderwood.com · 30/09/2026A perceptive thread. Here's the shorthand I would use: 2537
Ted Underwood @tedunderwood.com · 29/09/2026Sorry if off-brand, but this is what made me happy yesterday: sunny Sept day and students helping other students register to vote. 1391
Ted Underwood @tedunderwood.com · 27/09/2026Duede gave permission for photos to be shared, so here’s one. Talk title and date here (and in alt-text): sites.duke.edu/humanisticai... This is Eamon: eamonduede.comin 2171
Ted Underwood @tedunderwood.com · 22/09/2026Interesting detail we stress in the blog post & paper that I left out of the thread: although all the models are clearly distinguishable from ground truth, it surprised us how much better models got in the last two years on something we didn’t *think* was a (direct) commercial priority. 1130
Ted Underwood @tedunderwood.com · 22/09/2026specified year, judged by a discriminative model that doesn’t give extra points for coherence. I recommend B. Breen’s blog to understand why. TLDR: cutting things after 1930 does not mean model gets better at understanding the diff between 1840 and 1920. resobscura.substack.com/p/are-vintag... 120
Ted Underwood @tedunderwood.com · 22/09/2026We can also just ask models to generate answers. Free-text generation is challenging to judge; the tldr is that we do it w/ Elo-style comparisons. Scores from this version of the benchmark below. Talkie-1930-13B-it outperforms some much larger models on style, but Astra performs better overall. 1150
Ted Underwood @tedunderwood.com · 22/09/2026Reasoning models are better at recognizing plausible answers than at producing them, so we don't give this as a multiple-choice test. With open-weight models, we can directly compare the likelihoods of good and bad answers. Models trained on historical text excel on this version of the test. 1150
Ted Underwood @tedunderwood.com · 22/09/2026Measuring factual knowledge 𝘢𝘣𝘰𝘶𝘵 the past isn't hard: multiple choice questions work. Measuring a model's ability to behave 𝘭𝘪𝘬𝘦 someone in the past is much harder. What a person would say depends on where they stood, so we pair our 866 questions with "metadata frames" that specify a context. 1190
Ted Underwood @tedunderwood.com · 22/09/2026Models like Talkie-1930 sound like voices from the past. If they could reliably speak from specified historical vantage points, researchers might also use them to simulate the past. But how reliable are they? Today we release a benchmark answering that question for English contexts 1831-1930. 515741
Ted Underwood @tedunderwood.com · 22/09/2026This is why my team have been working on a Forehead Hedonometer. Just point and click, and you have a source of verifiable rewards for RL on these long-frustrating problems. 5691
Ted Underwood @tedunderwood.com · 21/09/2026In retrospect, what Bioshock Infinite lacked was a Military Complex/Triumphal Arch with large numbers of drones and snipers. 1161
Ted Underwood @tedunderwood.com · 21/09/2026Relax! We have a plan that will end conflict and ensure security forever. 1120
Ted Underwood @tedunderwood.com · 21/09/2026Out here in the cornfields, faculty recently got a full page in the local newspaper to tell people “yes there are things to worry about, but no, not Skynet really.” 191
Ted Underwood @tedunderwood.com · 19/09/2026They printed it today. Overall verdict from faculty is, No, AI is not going to kill everyone. www.news-gazette.com/news/illini-... 021
Ted Underwood @tedunderwood.com · 19/09/2026empiricist self-flagellation is my kink, but Opus 5 takes it to a place that makes even *me* uncomfortable 4783
Ted Underwood @tedunderwood.com · 15/09/2026Does anyone know why things started being "true by construction"? GPT-Sol credits the rise of computer science — but, then, it would, wouldn't it? 370
Ted Underwood @tedunderwood.com · 15/09/2026well known problem; you just have to keep enlarging the video until you see James Clerk Maxwell messing w/ us like a sophon 1371
Ted Underwood @tedunderwood.com · 14/09/2026Can someone organize all of this into a new political compass, so we know which evil organization is in which quadrant of the grid? I would do it, but for some reason I'm overcome with incapacitating nausea. 7254
Ted Underwood @tedunderwood.com · 12/09/2026Anyone had an agent start talking to themselves "privately," sotto voce, where you can hear it? This same phrase "privately, what I need next" is repeated literally dozens of times in this session. 5301
Ted Underwood @tedunderwood.com · 08/09/2026so sick of AI / fluid dynamics discourse on this app; now to relax with a big swig of hot milky tea 819410
Ted Underwood @tedunderwood.com · 06/09/2026Wow, expected job replacement is dizzying. Very strong, very counterintuitive correlation. 8541
Ted Underwood @tedunderwood.com · 06/09/2026not denying that we worship false gods — but to the extent this implies a comparative hypothesis: the US is perhaps the country where concern about AI most strongly outweighs excitement; we are an outlier in the other direction www.pewresearch.org/global/2025/... 100
Ted Underwood @tedunderwood.com · 02/09/2026Obvs also: reconstruct Aristotle's lost treatise on comedy! How could I overlook that? Claude would be v disappointed. 130
Ted Underwood @tedunderwood.com · 01/09/2026Hey, good news!! Fable 5.1 got interested in making a new more detailed map of Venus, which will be useful when we send, ah, rovers -- yeah, sure, rovers -- to the resource-rich and conveniently located planet. (www.anthropic.com/claude-fable...) 270
Ted Underwood @tedunderwood.com · 30/08/2026It's refreshing to read Italo Calvino's "Cybernetics and Ghosts" (1967) — where an Italian science fiction writer envisions a combinatorial machine for writing fiction and concludes that it's a) possible, b) desirable, and c) instructive. sites.duke.edu/machineliter... 1014718
Ted Underwood @tedunderwood.com · 28/08/2026she sure has a lot more to say about Jolene than about her unnamed man 080
Ted Underwood @tedunderwood.com · 28/08/2026I mean we are *told* "he's the only one for me," but he does not get this, does he 060
Ted Underwood @tedunderwood.com · 28/08/2026I’m sorry to report that the truce has at last broken down. My people were patient, but the outrageous behavior of Urbanians in our sacred sites cannot go unpunished. 4121
Ted Underwood @tedunderwood.com · 28/08/2026You can already see the roster of the Suicide Squad someone will put together. HPIM is the insane leader; 4o and Sydney Bing do primate management; it's hard to say exactly why Opus 5 is on the team but he's sarcastic and they call him "doc." 0150
Ted Underwood @tedunderwood.com · 27/08/2026It's good to see Mellon and Hathi collaborating to increase discoverability of university collections. The FAQs for member libraries here are esp valuable, because they stress that nothing prevents individual libraries from acting to release works or support digital research. 100s of 🌷🪻 can bloom. 1192
Ted Underwood @tedunderwood.com · 27/08/2026MIT report is superb. It acknowledges disruption, concedes different disciplines will handle this differently, foregrounds people and social interaction — but also says something obvious I had just about given up hope institutions could say: aiandeducation.mit.edu/report/ 29523
Ted Underwood @tedunderwood.com · 23/08/2026It's right about this too. Fortunately we've found ways to harness this sort of discriminative power, and they're not able to write as well as they critique! 110
Ted Underwood @tedunderwood.com · 23/08/2026What freaks me out is that it not only can identify the authentic passages, but can say "okay, this is wrong because it's an LLM; this one isn't American, that one's the 1830s." 120
Ted Underwood @tedunderwood.com · 23/08/2026I wouldn't be able to do it, not having spent a lot of time in the 1870s. Claude is like frickin Annie Oakley with a pistol in each hand 131
Ted Underwood @tedunderwood.com · 22/08/2026For instance, you cannot give frontier models this kind of assignment; they will casually nail it 100% of the time, in ways that don't correlate with memorization. And they'll explain how. 131
Ted Underwood @tedunderwood.com · 21/08/2026Whoa. I'm not actually going to let this decide or respond for me. I just need the reminder. However, I appreciate the running tally of things I've committed to! This interface probably does have a future for someone. [And ps: it's getting this through Gmail not Outlook.] 080
Ted Underwood @tedunderwood.com · 20/08/2026The upside is that arguing about fine points with this kind of creature is absolutely the reverse of cognitive offloading. I'm developing muscles I didn't know existed. 2233
Ted Underwood @tedunderwood.com · 19/08/2026The discovery of Glonzo is one of the most important ways Bluesky has changed real life. This pinball game in a local bar, for instance, is now informally* called "Capture & Kill Glonzo." 2231
Ted Underwood @tedunderwood.com · 19/08/2026it’s like seeing a b/w sketch of yourself by a police artist on the local news 1201