Sign in

David Mimno

@dmimno.bsky.social
7K followers 4.7K following 860 posts

He teaches information science at Cornell. mimno.infosci.cornell.edu

PostsRepliesMedia
Reposted by David Mimno
David Bamman @dbamman.bsky.social · 30/09/2026
We've been super busy the past few months getting out some research on the computational analysis of film, so I've put together a page with the highlights. people.ischool.berkeley.edu/~dbamman/fil... Some things to call out: 1/7
people.ischool.berkeley.edu
Computational Analysis of Film
1339
Reposted by David Mimno
Marco @mcognetta.bsky.social · 30/09/2026
🚨 [Token][ization] Paper Alert 🚨 Tokenization is a wildly understudied area of language modeling despite it having effects across all of NLP. Over the past ~8 months, 32 (!) tokenizer researchers put together the most comprehensive survey of the field. Check it out!
112935
Reposted by David Mimno
Isabel Silva Corpus @isabelcorpus.bsky.social · 30/09/2026
In 2023, Change.org integrated an AI-assisted writing tool into their platform. We found that access to the tool changed platform trends in petition text, but that outcomes (signatures and comments) did not improve... Excited that this paper is out, check it out! 🔗: www.nature.com/articles/s41...
nature.com
Introducing AI to an online petition platform changed outputs but not outcomes - Nature Human Behaviour
Corpus et al. examine how the introduction of an AI writing tool impacted Change.org petitions, finding that it increased petition homogeneity but did not improve petition outcomes.
22211
Reposted by David Mimno
Hoyt Long @hoytlong.bsky.social · 30/09/2026
If you're at #COLM2026, come check out our poster for "Spoiler Alert" (arxiv.org/abs/2604.09854). We ask why LLM fiction is so bad at holding narrative tension, and create a metric to measure tension in short stories. What improves LLM fiction on this metric, it turns out, is narrative planning.
2237
David Mimno @dmimno.bsky.social · 29/09/2026
An updated version of Scott Enderle's Topic Modeling Tool is now available here: github.com/mimno/topic-... Beta testers welcome!
github.com
Release v2.0.0 · mimno/topic-modeling-tool
Full Changelog: https://github.com/mimno/topic-modeling-tool/commits/v2.0.0
0115
David Mimno @dmimno.bsky.social · 29/09/2026
Has anyone made Yes / Yes to all / No USB foot pedals?
081
Reposted by David Mimno
Alison Babeu @alicatlibrarian.bsky.social · 29/09/2026
The Perseus Digital Library has been hard at work on Minimum Viable Perseus (Perseus MVP). Designed as a static site and heavily inspired by the beloved but faltering P4 (and its awesome users), please check out our efforts here (sites.tufts.edu/perseusupdates) and try it out: beta.perseus.tufts.edu
14119
David Mimno @dmimno.bsky.social · 29/09/2026
Sparsity is back! Sparsity is everywhere in language, but managing it requires overhead. For the last ~15 years it was faster to just do the dense operation, knowing most of it was useless. That sparse operations work is a huge shift in the "do something clever" vs. "spend more money" tradeoff.
2315
David Mimno @dmimno.bsky.social · 25/09/2026
I'm at the combination classifier and language model
010
David Mimno @dmimno.bsky.social · 24/09/2026
"Will no agents rid me of this turbulent priest?"
060
David Mimno @dmimno.bsky.social · 24/09/2026
This was happening back when it was just mail-merge templates, but I wonder with agents if prospectives are sometimes now even aware of what's being sent
010
Reposted by David Mimno
naitian @naitian.org · 24/09/2026
If you haven't yet registered for TADA 2026 on October 5th (day before COLM), registration is open to the public! Detailed schedule forthcoming, but lots of fantastic posters and talks. We even have a second half-day of lightning talks scheduled! tada2026.eventbrite.com
tada2026.eventbrite.com
Text as Data 2026
A leading forum for interdisciplinary research on the study of politics, society, and culture through computational analysis of documents.
0177
David Mimno @dmimno.bsky.social · 23/09/2026
Laya seems similar and fully open. It’s a modernbert fine tune.
010
Reposted by David Mimno
Ted Underwood @tedunderwood.com · 22/09/2026
Models like Talkie-1930 sound like voices from the past. If they could reliably speak from specified historical vantage points, researchers might also use them to simulate the past. But how reliable are they? Today we release a benchmark answering that question for English contexts 1831-1930.
Free-text evaluation of answers to character modeling and constrained generation questions. This is just one of several scores Chronologic-EN can produce; we focus on it here because it's both the hardest test and the one most relevant to simulation of the past. Frontier models reach 72%; Talkie-1930 is stronger than several larger competitors, but not at the frontier by this measure. Note that this score has improved ~45% in the last two years, but still falls perceptibly short of ground truth.
516344
David Mimno @dmimno.bsky.social · 22/09/2026
Are they groundbreaking technically? Probably not. I'm sure there's training tricks, but everyone will figure them out. But as a new *design pattern* for how models are used, this feels totally different and BERT-level exciting.
4190
David Mimno @dmimno.bsky.social · 22/09/2026
And I'm going to bet that a lot of what AI companies imagined as the potential exponential-growth big-model use cases aren't writing code or emails, but automating business processes. And Jev/Laya/whatever-comes-next are going to make that vastly simpler and cheaper.
1153
David Mimno @dmimno.bsky.social · 22/09/2026
That's where a model optimized for making simple decisions from complex inputs is so exciting. Clear answers, fast cheap inference, and compatible with the kind of sequence-of-decisions logic we were already moving towards. Every project we're working on can use this.
2130
David Mimno @dmimno.bsky.social · 22/09/2026
Sometimes you need text output, especially for code. But a lot of the time we're just trying to get a simple 👍👎 for some fairly complicated question from a lot of documents. The hard part is often specifying the problem, which usually ends up looking like a flowchart.
1110
David Mimno @dmimno.bsky.social · 22/09/2026
This is great! And it's so flexible and effective that we've started to assume that everything is just a text generation problem. But besides being ruinously slow and expensive, just figuring out what the answer was in the generated free text becomes its own hard problem.
170
David Mimno @dmimno.bsky.social · 22/09/2026
In the past few years, text generation models have become good enough that you can do structured prediction purely as a text gen process: the text that the model generates just happens to be a JSON object that happens to be the structure. No cleverness beyond clear specification, no custom code.
160
David Mimno @dmimno.bsky.social · 22/09/2026
After 2013, word2vec embeddings started making the classifiers work better. Contextualized embeddings (ELMo, BERT) were even better, but still just fed vectors into stacks of classifiers.
140
David Mimno @dmimno.bsky.social · 22/09/2026
This was a big leap, but still required a lot of cleverness in constructing the mapping, and custom code to train and run. Moving to a new task meant starting mostly from scratch.
150
David Mimno @dmimno.bsky.social · 22/09/2026
Around 2010, the most exciting approaches were clever ways to map complex structured prediction tasks into huge numbers of tiny decisions, usually performed with single-layer linear classifiers. For example, shift-reduce parsers phrased parsing as a four-way choice made at each token.
170
David Mimno @dmimno.bsky.social · 22/09/2026
It's possible for Jev/Laya/Decision Models to be not that big a deal as tech and massive as a new paradigm. Here's why I'm really excited from an NLP history perspective (thread)
17519
David Mimno @dmimno.bsky.social · 22/09/2026
The python package is well documented and feels modern, and most of the obvious python LDAs use stochastic variational. Both BERTopic and SV LDA are good enough that they seem to work, but not really that good. The weakness in BERTopic is that embedding data doesn't really work with HDBSCAN.
010
Reposted by David Mimno
Sung Kim @sungkim.bsky.social · 22/09/2026
This seems to be the most popular alternative to jev. Laya: Multilingual, non-autoregressive System 1 decision model. huggingface.co/convaiinnova... Laya Repo: github.com/NandhaKishor... Laya MLX: huggingface.co/aac6fef/laya... Laya MLX Repo: github.com/mizorewww/la...
51057
Reposted by David Mimno
Gordon Pennycook @gordpennycook.bsky.social · 21/09/2026
We have a new paper on AI debunking that includes a bunch of interesting data on various "epistemically suspect beliefs". Check it out!
1176
Reposted by David Mimno
Melanie Walsh @mellymeldubs.bsky.social · 21/09/2026
We are hiring for several positions in the UW Information School, including a tenure-track position in Library and Information Science. Please share with anyone interested in these areas! Deadline October 30. Job: apply.interfolio.com/192127
Screenshot of job ad that reads: "The successful candidate will be expected to contribute to our Library and Information Science (LIS) research, with areas of relevance including, but not limited to:

Libraries - Academic, Digital, or Public 

Librarianship - Leadership, Innovation, or Education  

Information Technology - Ethics, Public Interest, and Justice-Oriented Approaches  

Data - Archives, Public / Open, and Humanistic or Scientific Applications

Artificial Intelligence - Public Interest, Library or Cultural Heritage-Centered, Open-Source Applications in Information Institutions  "
24745
David Mimno @dmimno.bsky.social · 21/09/2026
Thanks for this! The problem of document import is tough, almost a complete app in itself.
010
Reposted by David Mimno
Richard Jean So @richardjeanso.bsky.social · 21/09/2026
Prospective PhD students: if you work on computational humanities, cultural analytics and/or cultural AI, pls apply to Duke English! We're recruiting in these areas & you'd join a vibrant community. Candidates with training in both lit studies + CS/sciences esp welcome. Email me if any questions!
12824
Reposted by David Mimno
Quinn Daedal @quinnanya.me · 21/09/2026
The #DataSittersClub is back... but the news is Terminal 💀. You'll survive, though, with @alyssavi.bsky.social new DSC little tl;dr #4: Navigating the Terminal for the First Time. It's the honest, practical, funny guide to the Mac Terminal that you need before you get down to work with other tools.
datasittersclub.github.io
Navigating the Terminal for the First Time
Yikes!
1177
Reposted by David Mimno
Quinn Daedal @quinnanya.me · 21/09/2026
Part 2 of the #DataSittersClub back-to-school spree is @alyssavi.bsky.social DSC little tl;dr #5: What's the Deal with Topic Modeling? Topic modeling is one of the first methods people hear about with digital humanities, but really... what's up with it? And how are you supposed to do it? And when?
datasittersclub.github.io
What's the Deal with Topic Modeling?
Are Topics Real?
2118
David Mimno @dmimno.bsky.social · 20/09/2026
emmaxxing
050
David Mimno @dmimno.bsky.social · 20/09/2026
It seems like what matters is the amount of computation, not whether the model can create a human-interpretable narration of the computation?
100
David Mimno @dmimno.bsky.social · 20/09/2026
Will Jev itself be significant? Almost certainly no. But is it pointing to a significantly new paradigm of AI interaction? Absolutely.
1131
David Mimno @dmimno.bsky.social · 18/09/2026
A university founded in the Middle Ages in a city that centers itself around a Medieval castle does not see a need to study Medieval History
35122
David Mimno @dmimno.bsky.social · 16/09/2026
whoa, the claude code update is suddenly much more aggressive about going off and doing things without asking
460
David Mimno @dmimno.bsky.social · 16/09/2026
I’d check kalshi
010
Reposted by David Mimno
Maria Antoniak @mariaa.bsky.social · 14/09/2026
We're hiring in Computer Science at the University of Colorado Boulder! ☀️⛰️ Machine learning and NLP people, please apply!
05531
Reposted by David Mimno
danah boyd @zephoria.bsky.social · 14/09/2026
PhD students (& new PhDs): I'm hiring a postdoc at Cornell (Ithaca) to conduct a novel study at the intersection of political economy and tech. Applications are due Oct 16. There are a LOT more details in the job ad so make sure to read it thoroughly: academicjobsonline.org/ajo/jobs/32502
lnkd.in
LinkedIn
This link will take you to a page that’s not on LinkedIn
22221
Reposted by David Mimno
David Mimno @dmimno.bsky.social · 13/09/2026
The silence from industry in the face of concerted attacks on US competitiveness has been disgraceful. Will this finally wake them up? Or will they just offshore all skilled work? And yes, this is a constitutional crisis if the executive can apply a “ridiculous fee” veto to any legislation.
0349
David Mimno @dmimno.bsky.social · 13/09/2026
The silence from industry in the face of concerted attacks on US competitiveness has been disgraceful. Will this finally wake them up? Or will they just offshore all skilled work? And yes, this is a constitutional crisis if the executive can apply a “ridiculous fee” veto to any legislation.
0349
David Mimno @dmimno.bsky.social · 11/09/2026
Killer AI is sufficient, but not necessary
171
Reposted by David Mimno
Advait @ COLM @advaitdeshmukh.com · 10/09/2026
1/7 In 1941, Borges imagined an impossible novel that follows all narrative branches at once. Today, we can observe chatbot users exploring narrative possibilities as they repeatedly edit story prompts. In our COLM 2026 paper, we studied this branching exploration via the WildChat dataset. 🧵
Paraphrased edit trees from the dataset
24512
David Mimno @dmimno.bsky.social · 11/09/2026
This is the exact use case it was designed for! ☺️ I could try updating / adopting Enderle's tool. Can you say more about what people were excited about, and what they found frustrating? I did a lot over the summer in python/rust and connections to embeddings/UMAP.
220
David Mimno @dmimno.bsky.social · 11/09/2026
Very cool! Is there an easy way to grab just the OCR text, say for one newspaper or date range? I'd like to use this for a class.
010
Reposted by David Mimno
Schloss Dagstuhl – Leibniz-Zentrum für Informatik (LZI) @dagstuhl.de · 10/09/2026
It is time for possible Dagstuhl proposals again! Proposals can be submitted between October 15 and November 1, 2026. For links to guidelines, more details on the process, and important dates see www.dagstuhl.de/en/institute...
dagstuhl.de
Call for Proposals (Deadline November 1, 2026)
Call for Proposals (Deadline November 1, 2026)
0108
David Mimno @dmimno.bsky.social · 10/09/2026
Reposting since a lot has happened this week
0141