Reposted by Andy HaltermanDavid Mimno @dmimno.bsky.social · 15/07/2026Text as Data is happening at Berkeley right before COLM. Please share! 03014
Andy Halterman @ahalterman.bsky.social · 13/07/2026If you're interested in learning how to create custom event data from text, my workshop at ICPSR is coming up soon (Aug 3-7). We'll be using my brand new textbook (!) that's meant to make it much easier to define your what you're looking for and implement a pipeline for doing it. myumi.ch/w9nbNmyumi.chTopical Workshops - 2026 ICPSR Summer Program RegistrationThe Summer Program offers excellent opportunities for training in research design, analytic strategies, statistical methods, and data management. 000
Reposted by Andy HaltermanDallas Card @dallascard.bsky.social · 10/07/2026As some may have heard me talk about at #ACL2026, I'm excited to share a new preprint on approaches to validation when using LLMs to measure concepts in social science, led by @madesai.bsky.social and @azjacobs.bsky.social !! Paper: arxiv.org/abs/2607.07915 56819
Reposted by Andy HaltermanLucy Li @lucy3.bsky.social · 31/03/2026Another one of @ahalterman.bsky.social and @katakeith.bsky.social's papers that I think should be cited more by CSS researchers: What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification arxiv.org/abs/2510.03541arxiv.orgWhat is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classificationGenerative large language models (LLMs) are now used extensively for text classification in computational social science (CSS). In this work, focus on the steps before and after LLM prompting -- conce... 1315
Reposted by Andy HaltermanPolitical Analysis @polanalysis.bsky.social · 27/11/2025Currently in FirstView: In “Codebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts,” @ahalterman.bsky.social and @katakeith.bsky.social show how “off-the-shelf” LLMs have limitations in faithfully following real-world codebook operationalizations. 131
Reposted by Andy HaltermanAidan Milliff @aidanmilliff.com · 19/09/2025Very short summary of this paper: 071
Andy Halterman @ahalterman.bsky.social · 19/09/2025Very excited that my paper with @katakeith.bsky.social is now out in @polanalysis.bsky.social. We investigate whether LLMs actually follow the instructions/definitions provided in codebooks, propose some diagnostics, and release a new evaluation dataset. www.cambridge.org/core/journal...cambridge.orgCodebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts | Political Analysis | Cambridge CoreCodebook LLMs: Evaluating LLMs as Measurement Tools for Political Science Concepts 04717
Reposted by Andy HaltermanPolitical Analysis @polanalysis.bsky.social · 28/04/2025Currently in FirstView: “Synthetically generated text for supervised text analysis.” @ahalterman.bsky.social proposes using LLMs to generate synthetic training data for training smaller, traditional supervised text models. 134
Andy Halterman @ahalterman.bsky.social · 31/01/2025New paper in Political Analysis on synthetic text data for training classifiers. Main idea: generate training examples with LLMs, then fit classifiers on synthetic (+real) text. Paper has validations and guidance. Blog: andrewhalterman.com/post/synthet... Paper: www.cambridge.org/core/journal...cambridge.orgSynthetically generated text for supervised text analysis | Political Analysis | Cambridge CoreSynthetically generated text for supervised text analysis 1144
Andy Halterman @ahalterman.bsky.social · 08/11/2023I couldn't find a tutorial I liked to get students who know R up and running with Python, so I wrote my own! Part 1 is here: andrewhalterman.com/post/python_... (And if you have a tutorial you like, please let me know!) 31710
Andy Halterman @ahalterman.bsky.social · 17/10/2023The optimal amount of dad joke cringe when teaching undergrads is > 0. 030