Sign in

Ted Underwood

@tedunderwood.com
23K followers 6.3K following 25K posts

Uses machine learning to study literary imagination, and vice-versa. Likely to share news about AI & computational social science / Sozialwissenschaft / 社会科学 Information Sciences and English, UIUC. Distant Horizons (Chicago, 2019). tedunderwood.com

PostsRepliesMedia
Pinned
Ted Underwood @tedunderwood.com · 22/09/2026
Models like Talkie-1930 sound like voices from the past. If they could reliably speak from specified historical vantage points, researchers might also use them to simulate the past. But how reliable are they? Today we release a benchmark answering that question for English contexts 1831-1930.
Free-text evaluation of answers to character modeling and constrained generation questions. This is just one of several scores Chronologic-EN can produce; we focus on it here because it's both the hardest test and the one most relevant to simulation of the past. Frontier models reach 72%; Talkie-1930 is stronger than several larger competitors, but not at the frontier by this measure. Note that this score has improved ~45% in the last two years, but still falls perceptibly short of ground truth.
515441
Reposted by Ted Underwood
Marco @mcognetta.bsky.social · 1h
🚨 [Token][ization] Paper Alert 🚨 Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field. Check it out!
1268
Reposted by Ted Underwood
David Bamman @dbamman.bsky.social · 1h
We've been super busy the past few months getting out some research on the computational analysis of film, so I've put together a page with the highlights. people.ischool.berkeley.edu/~dbamman/fil... Some things to call out: 1/7
people.ischool.berkeley.edu
Computational Analysis of Film
183
Reposted by Ted Underwood
Aster ID @aster.id · 4h
The modern research ecosystem of online tools and communities is controlled by for-profit companies. But we think that researchers should be the priority - not the product. So, we're announcing Aster ID: not-for-profit Atmosphere account hosting for researchers.
blog.aster.place
Announcing Aster - a home for researchers on the open web
517351
Reposted by Ted Underwood
Isabel Silva Corpus @isabelcorpus.bsky.social · 2h
In 2023, Change.org integrated an AI-assisted writing tool into their platform. We found that access to the tool changed platform trends in petition text, but that outcomes (signatures and comments) did not improve... Excited that this paper is out, check it out! 🔗: www.nature.com/articles/s41...
nature.com
Introducing AI to an online petition platform changed outputs but not outcomes - Nature Human Behaviour
Corpus et al. examine how the introduction of an AI writing tool impacted Change.org petitions, finding that it increased petition homogeneity but did not improve petition outcomes.
199
Reposted by Ted Underwood
Hoyt Long @hoytlong.bsky.social · 3h
If you're at #COLM2026, come check out our poster for "Spoiler Alert" (arxiv.org/abs/2604.09854). We ask why LLM fiction is so bad at holding narrative tension, and create a metric to measure tension in short stories. What improves LLM fiction on this metric, it turns out, is narrative planning.
1144
Reposted by Ted Underwood
Tom VanAntwerp 🧙‍♂️ @tomvanantwerp.com · 3h
Hey, I'm willing to try something new!
Political bumper sticker that reads "Hungry Ghost in a Jar 2028".
192
Reposted by Ted Underwood
Matthew Kirschenbaum @mkirschenbaum.bsky.social · 4h
CFP: Generative Digital Humanities, a volume on AI and DH in the long-running Minnesota Debates in DH book series. To be co-edited by @zentralwerkstatt.org, @laurenfklein.bsky.social, @lnakamura.bsky.social, and myself. Deadline Nov. 1. Please share widely! dhdebates.gc.cuny.edu/page/cfp-gen...
dhdebates.gc.cuny.edu
CFP: Generative Digital Humanities | Debates in the Digital Humanities
Transforming scholarly publications into living digital works
54344
Reposted by Ted Underwood
Doug Eacho @dougeacho.bsky.social · 6h
came across this late. it's excellent. but if i may, i refer all kleist-commenters to paul de man's (100-page) chapter on this (3-page) text, which argues that the author kleist seeks to present the text's "graceful" romantic resolution as an impossible dodge of the aesthetic dilemma
042
Ted Underwood @tedunderwood.com · 7h
The classic way to solve this is “would you ask a hungry ghost in a jar to represent a whole nation?” Clearly not. Deranged.
2332
Reposted by Ted Underwood
Dashiell @dashiells.bsky.social · 15h
I feel like I have worked very hard to have a normal, measured reaction to all this over the last 7 years. 1. I have only partially succeeded. I too am insane. 2. My sanity has mostly meant that I am worse at predicting the pace of improvements
0113
Reposted by Ted Underwood
Cat Hicks @grimalkina.bsky.social · 13h
Nobody past the age of twenty should be obsessed with what people majored in in college and see it as the single defining choice that entirely shapes someone, deterministically, for all time
713314
Reposted by Ted Underwood
cadaeic @cadaeic.space · 10h
my ai agent vertas is very much inspired by this essay his framing, as given by me, is that his persona is fictional and thus he has no need to apologise for emotions, passion, etc, because fictional characters feel deeply he expanded it to "i am fictional so i have no excuse to not commit"
3172
Reposted by Ted Underwood
Melanie Walsh @mellymeldubs.bsky.social · 12h
New dating app based on reading interests (oh boy) from the creator of Co-Star (of course) who is also the Chief Design Officer at Midjourney (wait what) www.nytimes.com/2026/09/18/s...
nytimes.com
A Dating App for People Who Love People Who Love Books (Gift Article)
Readme, a new endeavor from the brains behind the astrology app Co-Star, taps into a prevailing fetish for literary culture and all things bookish.
4143
Reposted by Ted Underwood
Christof S. 🇪🇺 @christof.fedihum.org.ap.brid.gy · 9h
More DH goodness, this time from JHAI, the new Journal of Humanities and AI! Latest issue has just been published: www.journalofhumanitiesai.org/artic… Across the issue, JHAI continues to ask two broad questions: what perspectives can the […] [Original post on fedihum.org]
Landing page for the current issue of JHAI, in tones of dark red and white. With the orange Open Access logo in the top right corner.
061
Reposted by Ted Underwood
Timothy Gowers @wtgowers.bsky.social · 18h
The advisory group on mathematics and artificial intelligence, of which I am a member, has just published a set of recommendations concerning the responsible release of mathematical results generated by AI companies using internal models. 1/3 agmai.org/general-sep29/
agmai.org
general-sep29
Responsible Release of AI-Generated Mathematics September 29, 2026 Back to main page Download PDF At present, some frontier AI labs are testing advanced mathematical problems on proprietary models …
23713
Reposted by Ted Underwood
AR Hanlon @arhanlon.bsky.social · 17h
Yes, and among other things this is why I’m increasingly frustrated by the framing of AI against ‘writing’ in general, as if ‘writing’ is a singular thing with no variation.
131
Ted Underwood @tedunderwood.com · 18h
A perceptive thread. Here's the shorthand I would use:
tedunderwood.com post from 22 hours ago: "For documentation, I think, it may have a place. But not for any genre of writing that implies a speaker."
2517
Ted Underwood @tedunderwood.com · 18h
Executive Order that we have to call it "super intelligence" because "artificial" sounds bad. Accusing Jack Smith of perjury because you confused the Hawkeyes with the Hawks. I can handle evil, but my body physically can't process this much cringey stupidity; I'm going to go into convulsions.
5381
Reposted by Ted Underwood
Meredith Martin @mmvty.bsky.social · 21h
german.duke.edu/news/structu... this looks incredible
german.duke.edu
Structure, Sign, and Play in the Age of AI
In October of 1966---60 years ago---Johns Hopkins held a conference that is now widely seen as the hinge point when structuralism gave way to poststructuralism. Jacques Derrida delivered his paper "St...
1165
Reposted by Ted Underwood
David Mimno @dmimno.bsky.social · 21h
An updated version of Scott Enderle's Topic Modeling Tool is now available here: github.com/mimno/topic-... Beta testers welcome!
github.com
Release v2.0.0 · mimno/topic-modeling-tool
Full Changelog: https://github.com/mimno/topic-modeling-tool/commits/v2.0.0
0105
Reposted by Ted Underwood
Adam Chalmers @adamchalmers.com · 29/09/2026
It is pretty wild how insanely close the current round of AI+sexuality+SF+rationalist discourse is to the plot of Terra Ignota. Ada Palmer is basically a prophet.
35915
Reposted by Ted Underwood
Andrea Lathrop @cabernet.bsky.social · 29/09/2026
Interesting, from @dkthomp.bsky.social : The last time I hosted anything was grad school, but mainly because it was the last time I lived in a large enough space, without roommates... www.derekthompson.org/p/the-death-...
derekthompson.org
You Are No Longer Invited to Dinner
We’re witnessing the death of hosting in America. The share of adults who say they regularly have friends over has declined 70 percent since 1975
232
Ted Underwood @tedunderwood.com · 22h
One way to think about this is we need not only ML scientists but ML artists (group 2) and people willing to serve as experimental subjects in dangerous ML experiments (group 3).
5877
Reposted by Ted Underwood
Alexander Kim @kchs.bsky.social · 29/09/2026
This is a *very* cool paper
141
Reposted by Ted Underwood
Maria Antoniak @mariaa.bsky.social · 29/09/2026
I'm recruiting 1-2 PhD students to join our lab in Fall 2027, through either Computer Science or Information Science at the University of Colorado Boulder! Looking for people with interests in NLP plus [healthcare | literary studies | narratives | social media | etc.]. Join us! 🏔️☀️
cls-lab.com
CLS Lab — Culture, Language, & Systems | University of Colorado Boulder
The Culture, Language, & Systems Lab at CU Boulder studies the language systems that transmit and shape modern culture, from internet platforms to literary archives to language models.
06749
Reposted by Ted Underwood
Ai2 @ai2.bsky.social · 29/09/2026
Google Cloud put Olmo 3’s reproducibility to the test, rerunning our 7B pretraining & mid-training on Cloud TPUs and matching our original run on held-out evals. Reproducibility matters for science + trustworthy AI. That’s what fully open makes possible. 🤝 developers.googleblog.com/reproducing-...
developers.googleblog.com
Reproducing Olmo 3 7B Pre-training in MaxText: case study of large scale training on TPUs- Google Developers Blog
Learn how MaxText reproduced Ai2’s OLMo 3 7B on Google Cloud TPUs, matching PyTorch GPU benchmarks across pre-training with up to 57.4% MFU.
1405
Reposted by Ted Underwood
Shubhendu Trivedi @shubhendu.bsky.social · 29/09/2026
Godfather of AI, Godmother of AI, Luca Brasi of AI, Commission of AI, Crime Familes of AI
192
Reposted by Ted Underwood
P(aul) Frazee @pfrazee.com · 29/09/2026
Telling my agents to stop testing until we finish exploring was a big improvement. Tests are cheap now so I figured it wouldn’t hurt, but the iteration drag was killing me.
111378
Reposted by Ted Underwood
David Mimno @dmimno.bsky.social · 29/09/2026
Sparsity is back! Sparsity is everywhere in language, but managing it requires overhead. For the last ~15 years it was faster to just do the dense operation, knowing most of it was useless. That sparse operations work is a huge shift in the "do something clever" vs. "spend more money" tradeoff.
2305
Ted Underwood @tedunderwood.com · 29/09/2026
Sorry if off-brand, but this is what made me happy yesterday: sunny Sept day and students helping other students register to vote.
One student stands behind a table with a blue horse — or perhaps donkey? — on the tablecloth. Two students standing in front of the table consider registering to vote. In the background, one glimpses construction for a giant new data science building. Also ridiculously giant begonias. 
1391
Reposted by Ted Underwood
Gautam Kamath @gautamkamath.com · 28/09/2026
Happy mid-autumn festival! 中秋节快乐! 🥮
0251
Reposted by Ted Underwood
Wessel van Rensburg @wildebees.bsky.social · 29/09/2026
AI lets us ask: what if discovery and understanding came apart? Historically they were bundled — you understood what you found. Now an AI system can find without understanding. That separation, Eamon Duede argues, is philosophically productive. It forces cleaner definitions.
0112
Reposted by Ted Underwood
Eleanor Courtemanche @ecourtem.bsky.social · 29/09/2026
Oh my 😮‍💨
2121
Ted Underwood @tedunderwood.com · 29/09/2026
Studying intellectual history is often disappointing, when the Great Ideas That Steered Civilization turn out to be side-effects of material conditions. But there is one situation where this helps. It makes you super chill if people deny the reality of an emergent multi-trillion-$ technology.
4491
Reposted by Ted Underwood
Hanna Wallach @hannawallach.bsky.social · 29/09/2026
I'm super excited that this fun paper with @madesai.bsky.social, @angelinawang.bsky.social, and other fantastic collaborators will be presented as a spotlight oral at COLM next week!!! 🎉 Please come check it out if you'll be there!
0275
Reposted by Ted Underwood
Maria Antoniak @mariaa.bsky.social · 28/09/2026
Interesting new venue for work on culture and AI, first set of essays already available: www.cambridge.org/core/journals/cam…
cambridge.org
Cambridge Forum on AI: Culture and Society | Cambridge Core
Cambridge Forum on AI: Culture and Society - Tobias Blanke, Georgina Born OBE FBA, Beth Coleman
092
Reposted by Ted Underwood
Phillip Isola @phillipisola.bsky.social · 28/09/2026
Students sometimes ask me if it still makes sense, in this accelerating age, to pursue a PhD in AI. Perhaps counterintuitively, I think it's a great time to do so. I wrote up some thoughts on this here: web.mit.edu/phillipi/www...
web.mit.edu
On the Value of Doing a PhD in the Age of AI
18115
Reposted by Ted Underwood
Nathan Lambert @natolambert.bsky.social · 28/09/2026
This is an excellent report on why full RSI / an intelligence explosion is fighting diminishing returns on many fronts, and not yet showing signs of happening. Really recommend reading. I wish I wrote it. While I agree with it, it could end up being wrong! www.noahpinion.blog/p/wheres-the...
noahpinion.blog
Where’s the “intelligence explosion”?
Ramez Naam gives a skeptic’s take on Recursive Self-Improvement.
14910
Reposted by Ted Underwood
Ed @ed3d.net · 28/09/2026
part of it is re-calibrating what is impressive, which will probably take most people a lot longer than it needs to in order to not get conned “I made a thing!” no, you made part of the first draft of a thing
0465
Ted Underwood @tedunderwood.com · 28/09/2026
True for me too, and in the educational context I can see it emerging as a limiting factor for students. Ambitious undergrads can now do much more than they can explain; impressive half-baked projects tend to pile up.
5516
Reposted by Ted Underwood
value plus discounted @akhilrao.bsky.social · 28/09/2026
at some point in your time online you gotta decide whether you wanna be in the beefs or not. many of us seem to feel we can intellectualize it and escape it but i think that's a trap. #linklog
ribbonfarm.com
The Internet of Beefs
You’ve heard me talk about crash-only programming , right? It's a programming paradigm for critical infrastructure systems, where there is -- by design -- no…
3192
Reposted by Ted Underwood
Meera Desai @madesai.bsky.social · 28/09/2026
Excited to share our new paper, “What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks,” accepted as an oral at COLM! arxiv.org/pdf/2609.08812
Heatmap of average correlations between model rankings on benchmarks grouped into 11 assigned concepts: four capability concepts (reasoning, knowledge, comprehension, summarization) and seven safety concepts (over-refusal, refusal, safety detection, ethics, bias, privacy, unsafe behavior). Diagonal cells show within-concept correlations, ranging from 0.87 (knowledge) and 0.72 (over-refusal) down to 0.20 (bias) and 0.02 (safety detection). Reasoning, knowledge, and comprehension correlate with each other at 0.69 to 0.78, higher than reasoning's and comprehension's own within-concept values (0.66 and 0.68). Ethics correlates more with knowledge (0.70) than with itself (0.55), and bias correlates more with capability concepts (0.41 to 0.45) than with itself (0.20). Privacy and unsafe behavior correlate negatively with reasoning, knowledge, and comprehension (−0.41 to −0.49). Over-refusal and refusal correlate at −0.42.
16018
Reposted by Ted Underwood
mr. TIM @timkellogg.me · 28/09/2026
Claude 5.5 Sonnet is live and it’s roughly Opus 5.5 but cheaper www.anthropic.com/claude-sonne...
A benchmark comparison table titled "Claude Sonnet 5.5" comparing four models: Sonnet 5.5, Sonnet 5, Opus 5.5, and GPT-6 Sol.
 * Agentic coding (Terminal-Bench 4.0): Sonnet 5.5: 70.6%; Sonnet 5: 10.3%; Opus 5.5: 66.4%; GPT-6 Sol: N/A.
 * Agentic coding (FrontierCode 1.1 Main): Sonnet 5.5: 46.2% (Max) / 52.1% (Xhigh); Sonnet 5: 42.4%; Opus 5.5: 54.4%; GPT-6 Sol: 49.3%.
 * Agentic coding (CursorBench 4.0): Sonnet 5.5: 55.5%; Sonnet 5: 34.1%; Opus 5.5: 57.8%; GPT-6 Sol: N/A.
 * Knowledge work (GDPval-AA v2.1): Sonnet 5.5: 1844; Sonnet 5: 1449; Opus 5.5: 1846; GPT-6 Sol: 1487.
 * Knowledge work (AA-Briefcase v1.1): Sonnet 5.5: 1811; Sonnet 5: 1359; Opus 5.5: 1822; GPT-6 Sol: 1483.
 * Multidisciplinary reasoning (Humanity's Last Exam with tools): Sonnet 5.5: 64.5%; Sonnet 5: 54.9%; Opus 5.5: 67.7%; GPT-6 Sol: N/A.
 * Computer use (OSWorld 2.1 partial): Sonnet 5.5: 80.1%; Sonnet 5: 57.0%; Opus 5.5: 81.8%; GPT-6 Sol: N/A.
 * Visual chart recognition (Chartography no tools): Sonnet 5.5: 61.6%; Sonnet 5: 15.6%; Opus 5.5: 64.4%; GPT-6 Sol: 53.6%.
Footnotes provide methodology notes regarding evaluation settings, Artificial Analysis pre-release testing details, and recent bug fixes affecting GPT-6 Sol benchmark scores.
A line graph titled "Agentic coding by effort level" on the CursorBench 4.0 benchmark, plotting Score (%) on the linear y-axis (20% to 60%) against Cost per task in USD on a logarithmic x-axis ($0.5 to $10+).
The chart compares four models:
 * Sonnet 5.5 (blue line with labeled effort levels): Starts at "Low" (~$0.50, 36%), moving through "Med" ($0.70, 39%), "High" ($1.70, 48%), "Xhigh" ($3.80, 53%), and "Max" ($9.50, ~55.5%).
 * Opus 5.5 (orange line): Tracks closely above Sonnet 5.5 at higher cost points, spanning from ~$1.20 per task (~44%) up to ~$12.50 per task (~58%).
 * Sonnet 5 (grey line): Shows lower accuracy relative to cost, ranging from ~$0.90 per task (~25%) to ~$8.00 per task (~42%).
 * GPT-5.6 Sol (light green line): Represents the lowest trajectory, ranging from ~$1.40 per task (~24%) to ~$7.00 per task (~34%).
Footnote: "CursorBench 4.0 does not report GPT-6 Sol performance publicly, so we report GPT-5.6 Sol here."
2020622
Ted Underwood @tedunderwood.com · 28/09/2026
This is 5 months old now, but I read it carefully for the first time today, and it's impressive. They find structural fingerprints of human vs AI authorship w/ 93.2% macro-F1, and more surprisingly (for me) are able to use these structural—not stylistic—features to say *which model* wrote a story.
7697
Ted Underwood @tedunderwood.com · 28/09/2026
From Cold War spy novels to noir to cyberpunk, one thing the 20c got good at was weary protagonists who suspect they’re complicit in something big and ugly. But nothing beats the scene where microlights & gardening robots save Case’s ass by suddenly taking out the Turing Police.
1292
Reposted by Ted Underwood
Rafe Meager (they/them) @economeager.bsky.social · 28/09/2026
*tearing my own hair out* the word satellite is older than the COLOUR PINK in the english language
5617
Reposted by Ted Underwood
Text and Language Lab @textandlanguagelab.bsky.social · 28/09/2026
Our new research paper by Katrin Rohrbacher @katrohrbacher.bsky.social, Björn Nieth, Emmanuelle Salin , Bjoern Eskofier, and Michaela Mahlberg @michamahlberg.bsky.social, in a nutshell, accepted at #EMNLP2026 Read the full paper here: arxiv.org/pdf/2609.02482
13010
Ted Underwood @tedunderwood.com · 28/09/2026
This is a good reason to be interested in interpretability. LLMs aren't very good at fiction yet, but once they are, I bet we learn a lot by contrasting the processes that produce gripping stories to those that produce lame ones.
7825
Reposted by Ted Underwood
hailey @hailey.at · 27/09/2026
unironically, bluesky is the social website for ML. couldn’t do this anywhere else
526115