Sign in

Andy Halterman

@ahalterman.bsky.social
542 followers 60 following 7 posts

Assistant professor of political science at MSU. NLP, text, and conflict.

PostsRepliesMedia
Reposted by Andy Halterman
David Mimno @dmimno.bsky.social · 15/07/2026
Text as Data is happening at Berkeley right before COLM. Please share!
03014
Andy Halterman @ahalterman.bsky.social · 13/07/2026
If you're interested in learning how to create custom event data from text, my workshop at ICPSR is coming up soon (Aug 3-7). We'll be using my brand new textbook (!) that's meant to make it much easier to define your what you're looking for and implement a pipeline for doing it. myumi.ch/w9nbN
myumi.ch
Topical Workshops - 2026 ICPSR Summer Program Registration
The Summer Program offers excellent opportunities for training in research design, analytic strategies, statistical methods, and data management.
000
Reposted by Andy Halterman
Dallas Card @dallascard.bsky.social · 10/07/2026
As some may have heard me talk about at #ACL2026, I'm excited to share a new preprint on approaches to validation when using LLMs to measure concepts in social science, led by @madesai.bsky.social and @azjacobs.bsky.social !! Paper: arxiv.org/abs/2607.07915
Title page from "Validating LLMs in social science: Epistemic threats and emerging norms" by Meera Desai, Dallas Card, and Abigail Z. Jacobs
56819
Reposted by Andy Halterman
Lucy Li @lucy3.bsky.social · 31/03/2026
Another one of @ahalterman.bsky.social and @katakeith.bsky.social's papers that I think should be cited more by CSS researchers: What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification arxiv.org/abs/2510.03541
arxiv.org
What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification
Generative large language models (LLMs) are now used extensively for text classification in computational social science (CSS). In this work, focus on the steps before and after LLM prompting -- conce...
1315
Reposted by Andy Halterman
Political Analysis @polanalysis.bsky.social · 27/11/2025
Currently in FirstView: In “Codebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts,” @ahalterman.bsky.social and @katakeith.bsky.social show how “off-the-shelf” LLMs have limitations in faithfully following real-world codebook operationalizations.
131
Reposted by Andy Halterman
Aidan Milliff @aidanmilliff.com · 19/09/2025
Very short summary of this paper:
Dr Strangelove "it could easily be accomplished with a computer" meme rewritten to say "it could be accomplished with a computer if you are careful."
071
Andy Halterman @ahalterman.bsky.social · 19/09/2025
Very excited that my paper with @katakeith.bsky.social is now out in @polanalysis.bsky.social. We investigate whether LLMs actually follow the instructions/definitions provided in codebooks, propose some diagnostics, and release a new evaluation dataset. www.cambridge.org/core/journal...
cambridge.org
Codebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts | Political Analysis | Cambridge Core
Codebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts
04717
Reposted by Andy Halterman
Political Analysis @polanalysis.bsky.social · 28/04/2025
Currently in FirstView: “Synthetically generated text for supervised text analysis.” @ahalterman.bsky.social proposes using LLMs to generate synthetic training data for training smaller, traditional supervised text models.
134
Andy Halterman @ahalterman.bsky.social · 31/01/2025
New paper in Political Analysis on synthetic text data for training classifiers. Main idea: generate training examples with LLMs, then fit classifiers on synthetic (+real) text. Paper has validations and guidance. Blog: andrewhalterman.com/post/synthet... Paper: www.cambridge.org/core/journal...
cambridge.org
Synthetically generated text for supervised text analysis | Political Analysis | Cambridge Core
Synthetically generated text for supervised text analysis
1144
Andy Halterman @ahalterman.bsky.social · 03/12/2024
A McSweeney's style parody.

TREADSTONE RECRUITMENT OR YOUR PHD PROGRAM?

Can you tell whether each quote is discussing a top-secret CIA black ops program or your department's graduate program?

1. “We'd hoped it might build into a good training platform, but quite honestly, for a strictly theoretical exercise, we thought it was far too expensive."

2. “You came to us. You volunteered. You said you'd do anything it takes."

3. "You haven't slept for a long time, have you? Have you made a decision? This can't go on, you know. You have to decide."

4. "The details? No. I mean, I was told it was voluntary. I don't know if that's true or not, but that's what I was told."

5. "Stop running from the truth. You chose to come here! You chose to stay! And no matter how much you want to forget it... eventually you're going to have to face how you chose to become [who you are]."

6.  "You could have left at any time. And you knew exactly what it meant for you if you chose to stay."

7. "You're not a liar are you? Or too weak to see this through?"

8. “Look, they took vulnerable subjects, okay? You mix that with the right pharmacology and some serious behavior modification..."

9.  “You made yourself into who you are."
042
Andy Halterman @ahalterman.bsky.social · 08/11/2023
I couldn't find a tutorial I liked to get students who know R up and running with Python, so I wrote my own! Part 1 is here: andrewhalterman.com/post/python_... (And if you have a tutorial you like, please let me know!)
31710
Andy Halterman @ahalterman.bsky.social · 17/10/2023
The optimal amount of dad joke cringe when teaching undergrads is > 0.
The "focus group" meme from I Think You Should Leave. The original text, "A great steering wheel that doesn't whiff out the window while I driving" is replaced with "A great steering wheel that whiffs out the window while I driving (Schelling 1966)"
030