Sign in

Dallas Card

@dallascard.bsky.social
2.8K followers 431 following 166 posts

Assistant professor at si.umich.edu working in computational social science, machine learning, and NLP | dallascard.github.io

PostsRepliesMedia
Reposted by Dallas Card
brendan o’connor @brenocon.bsky.social · 05/10/2026
This year UMass Amherst CICS will be hiring tenure-track faculty in natural language processing! Job ad to be posted soon. For anyone at #TADA2026 or #COLM2026 this week, I or my colleague Hamed Zamani would love to chat or answer questions about it - just say hi or send me an email to meet!
01413
Reposted by Dallas Card
Meera Desai @madesai.bsky.social · 28/09/2026
Excited to share our new paper, “What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks,” accepted as an oral at COLM! arxiv.org/pdf/2609.08812
Heatmap of average correlations between model rankings on benchmarks grouped into 11 assigned concepts: four capability concepts (reasoning, knowledge, comprehension, summarization) and seven safety concepts (over-refusal, refusal, safety detection, ethics, bias, privacy, unsafe behavior). Diagonal cells show within-concept correlations, ranging from 0.87 (knowledge) and 0.72 (over-refusal) down to 0.20 (bias) and 0.02 (safety detection). Reasoning, knowledge, and comprehension correlate with each other at 0.69 to 0.78, higher than reasoning's and comprehension's own within-concept values (0.66 and 0.68). Ethics correlates more with knowledge (0.70) than with itself (0.55), and bias correlates more with capability concepts (0.41 to 0.45) than with itself (0.20). Privacy and unsafe behavior correlate negatively with reasoning, knowledge, and comprehension (−0.41 to −0.49). Over-refusal and refusal correlate at −0.42.
16619
Dallas Card @dallascard.bsky.social · 28/09/2026
Amazing postdoc opportunity!
002
Dallas Card @dallascard.bsky.social · 11/09/2026
What's the best way to engage publicly with a high-profile, polarizing, but mostly speculative debate that's been going on ad nauseum for years on end? Asking for a friend.
040
Dallas Card @dallascard.bsky.social · 08/09/2026
Note that the initial deadline is very soon! Applicants to UMSI must submit their application materials to Interfolio by September 20, 2026 and to the PPFP Office by November 1, 2026.
000
Dallas Card @dallascard.bsky.social · 08/09/2026
We define privacy broadly, and welcome applications from researchers approaching privacy through technical, policy, design, ethical, and/or social lenses.
100
Dallas Card @dallascard.bsky.social · 08/09/2026
Job alert! We are hiring in the area of privacy and technology, with a particular interest in how privacy intersects with AI. This position is via the Presidential Postdoctoral Fellowship Program, which we would expect to lead to a tenure-track position within UMSI. www.si.umich.edu/people/facul...
si.umich.edu
Presidential Postdoctoral Fellow/Assistant Professor - Privacy and Technology | umsi
Job posting for the Presidential Postdoctoral Fellow/Assistant Professor - Privacy and Technology position
175
Dallas Card @dallascard.bsky.social · 04/09/2026
One thing I love about the ACL anthology: when you click on "PDF", it just gives you the pdf, rather than opening some janky, bespoke pdf viewer in your browser.
0140
Reposted by Dallas Card
brendan o’connor @brenocon.bsky.social · 01/08/2026
The deadline has been extended to Monday, August 3 --
0126
Reposted by Dallas Card
David Jurgens @davidjurgens.bsky.social · 28/07/2026
Podcasts are an important part of the media ecosystem but woefully understudied. A year ago @blitt.bsky.social, @dallascard.bsky.social and I put out SPoRC, a new dataset of 1.1M podcast episodes. Today at #IC2S2, Dallas and I led a tutorial for how you can work with SPoRC and collect your own data!
1193
Reposted by Dallas Card
Sarah Fox @perhaxis.bsky.social · 24/07/2026
The Tech Solidarity Lab is on the move! This fall, I'm joining the University of Michigan School of Information as the John Derby Evans Associate Professor. I'm grateful to the students and colleagues at CMU for what we built together. The work continues, just a little deeper in the Rust Belt. 💙
2201
Dallas Card @dallascard.bsky.social · 24/07/2026
We're so excited to have you join us here!! 🎉🎉🎉
010
Dallas Card @dallascard.bsky.social · 16/07/2026
That also makes sense : ) There is definitely more startup and maintenance cost to making your own thing, although it's easier now than ever. I used to post stuff on medium but ultimately switched to a static site on github pages, and I've been happy with that solution (for my purposes).
030
Dallas Card @dallascard.bsky.social · 15/07/2026
Personally I find both substack and medium to be terrible reading experiences (too many pop-ups and annoying web design) and much prefer following independently-hosted blogs via RSS!
110
Dallas Card @dallascard.bsky.social · 15/07/2026
Substack seems to have been enormously successful, partly by their support for managing subscribers and sending posts out as email newsletters. But with both it and medium you're locked in, limited by their affordances, and embedded within a universe of slop (medium arguably more so than substack).
130
Dallas Card @dallascard.bsky.social · 15/07/2026
For more context, you can get a sense of the type of work that typically appears by looking at the proceedings from previous events, many of which are available here: textasdata.github.io/events/
textasdata.github.io
Events
Website of the Text as Data Society
020
Dallas Card @dallascard.bsky.social · 15/07/2026
Text as Data has been a wonderful, long-running, non-archival workshop for empirical research at the intersection of AI and social science (especially work involving text). After a few years off, it will be happening again this year as a one-day event in early October!
2268
Dallas Card @dallascard.bsky.social · 10/07/2026
Thanks so much for these pointers! We were also surprised to not find any examples of papers using prompting for measurement in the sociology journals we selected, even though people are clearly experimenting with it. I'll see if we can at least determine how many we would find in those you suggest!
120
Dallas Card @dallascard.bsky.social · 10/07/2026
Yes! Would love to hear your thoughts on it!
010
Dallas Card @dallascard.bsky.social · 10/07/2026
There's lots more in our paper (arxiv.org/abs/2607.07915), including our complete annotation scheme, and the data itself is available on OSF: osf.io/ab5zc/overvi...
arxiv.org
Validating LLMs in social science: Epistemic threats and emerging norms
Large language models (LLMs) are reshaping social science methodology. Researchers increasingly prompt language models to generate quantitative measurements of social concepts, for example labeling da...
000
Dallas Card @dallascard.bsky.social · 10/07/2026
Also relevant is work by @ahalterman.bsky.social and @katakeith.bsky.social on codebooks and conceptualization (arxiv.org/abs/2510.03541), and the paper by @dashapruss.bsky.social and Jessie Allen on AI jurisprudence (ojs.aaai.org/index.php/AI...).
arxiv.org
What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification
Generative large language models (LLMs) are now used extensively for text classification in computational social science (CSS). In this work, focus on the steps before and after LLM prompting -- conce...
130
Dallas Card @dallascard.bsky.social · 10/07/2026
We also suggest looking at recent examples that have made more thorough attempts at validation, and use these as inspiration, such as work by @patrickwu.bsky.social (arxiv.org/abs/2303.12057), @aaronschein.bsky.social (arxiv.org/abs/2312.09203), and James Bisbee (osf.io/preprints/so...)
arxiv.org
Large Language Models Can Be Used to Estimate the Latent Positions of Politicians
Existing approaches to estimating politicians' latent positions along specific dimensions often fail when relevant data is limited. We leverage the embedded knowledge in generative large language mode...
120
Dallas Card @dallascard.bsky.social · 10/07/2026
For rigorous validation, we recommend that researchers assess multiple aspects of construct validity, and we include examples from this set of papers, as well as hypothetical examples, mapping them to parts of the construct validity framework.
110
Dallas Card @dallascard.bsky.social · 10/07/2026
Overall, we recommend that researchers and experts should not simply cede interpretation of concepts to models. Despite the temptation, we should not assume that models will necessarily interpret terms in a way that is consistent, objective, or match the researcher's intended meaning.
120
Dallas Card @dallascard.bsky.social · 10/07/2026
Less commonly, we found cases of authors describing more varied approaches to validation, including face validity, discriminant validity, etc., with the use of multiple coordinated approaches shown in this plot.
110
Dallas Card @dallascard.bsky.social · 10/07/2026
Most commonly the reference for convergent validity was human annotated data, but papers also compared against measurements from other models. Both the quality of the reference and the rigor of the comparison varied greatly (e.g., amount of data, intercoder reliability, chance correction, etc.)
120
Dallas Card @dallascard.bsky.social · 10/07/2026
For validation, there is more consistency, but limited scope. In 8 tasks (across 6 papers), authors reported no attempt to validate their measurements. Most commonly, researchers assessed convergent validity (e.g., comparing individual model predictions to a reference), but details varied greatly.
111
Dallas Card @dallascard.bsky.social · 10/07/2026
For operationalization, we find reasonable reporting, but with a wide variety of approaches, including some use of few-shot examples, model roles (e.g., "You are a medical expert"), attempts to constrain the output (e.g., "Answer only with a number."), prompt engineering, and so forth.
111
Dallas Card @dallascard.bsky.social · 10/07/2026
Overall, we find a strong lack of consistency or clear norms. Surprisingly, researchers often do not define their concept of interest (29/50 tasks). Rather, they might ask a model "Is the following post offensive?", relying on the model to correctly interpret the meaning of "offensive".
131
Dallas Card @dallascard.bsky.social · 10/07/2026
In this paper, we survey what has so far been published in eight top social science journals (2022-2025), and assess how researchers are using LLM as measurement instruments, how they are defining and operationalizing their constructs, and how they are approaching validation.
100
Dallas Card @dallascard.bsky.social · 10/07/2026
Social scientists have long used a variety of approaches to get quantitative measurements of fuzzy concepts like "ideology", including human coding and supervised learning. Unsurprisingly, there has been great interest in using LLMs for this, but LLMs present unique challenges for valid measurement.
130
Dallas Card @dallascard.bsky.social · 10/07/2026
As some may have heard me talk about at #ACL2026, I'm excited to share a new preprint on approaches to validation when using LLMs to measure concepts in social science, led by @madesai.bsky.social and @azjacobs.bsky.social !! Paper: arxiv.org/abs/2607.07915
Title page from "Validating LLMs in social science: Epistemic threats and emerging norms" by Meera Desai, Dallas Card, and Abigail Z. Jacobs
56819
Dallas Card @dallascard.bsky.social · 27/05/2026
Congratulations! That's so exciting!!
110
Reposted by Dallas Card
michael veale @michae.lv · 05/05/2026
2 year research postdoc at UCL @laws.ucl.ac.uk on data and democracy, public law focus — could focus on surveillance and/or AI issues www.ucl.ac.uk/work-at-ucl/...
ucl.ac.uk
UCL – University College London
UCL is consistently ranked as one of the top ten universities in the world (QS World University Rankings 2010-2022) and is No.2 in the UK for research power (Research Excellence Framework 2021).
22036
Reposted by Dallas Card
Melanie Walsh @mellymeldubs.bsky.social · 26/04/2026
Postdoc in Digital Humanities at WashU in St. Louis. Deadline is May 1. Also, St. Louis is wonderful. apply.interfolio.com/185361
Screenshot of Interfolio position that reads: "Postdoctoral Fellowship, Digital Humanities - Department of Comparative Literature & Thought, Washington University in St. Louis
Washington University in St. Louis - Danforth Campus: School of Arts & Sciences: Comparative Literature
Location
St. Louis, MO
Open Date
Apr 24, 2026

Deadline
May 01, 2026 at 11:59 PM Eastern Time
Description
The Department of Comparative Literature and Thought at Washington University in St. Louis invites applications for a two-year postdoctoral teaching fellowship in digital humanities. The fellowship period will run from July 1, 2026 to June 30, 2028, with the possibility of a one-year extension.

In addition to pursuing their own research agenda, the postdoctoral fellow will teach the equivalent of four courses per year and is expected to participate fully in the thriving DH community on campus. Teaching responsibilities will include at least one regular course each Fall and Spring semester. Additional teaching responsibilities may include additional semester-long lab courses or a series of several short-term workshops. Fellows are expected to be in residence during the entire fellowship period, apart from research-related travel. Fellows will receive a salary of $64,380 per year, plus Washington University postdoctoral benefits and a $2,500 annual research/travel stipend."
13935
Dallas Card @dallascard.bsky.social · 17/04/2026
It is wild how quickly Claude Code lets you progress from rapid prototyping to feature creep.
0252
Dallas Card @dallascard.bsky.social · 09/04/2026
@lucy3.bsky.social Yes, don't leave us in suspense!
120
Reposted by Dallas Card
nlpandcss.bsky.social @nlpandcss.bsky.social · 06/04/2026
To accommodate ACL decisions, we are further extending the commitment deadline for pre-reviewed ARR submissions to April 7!
044
Dallas Card @dallascard.bsky.social · 01/04/2026
Thank you to everyone who has completed their reviews for the NLP+CSS workshop! If anyone has the bandwidth to help with emergency reviewing over the next couple of days, please send me a DM!
012
Dallas Card @dallascard.bsky.social · 21/03/2026
I guess the appropriate baseline comparison would be a map with each state's population visualized using the dot encoding.
100
Dallas Card @dallascard.bsky.social · 21/03/2026
Yes, nice observation here! It looks to me like they've just computed a value per state, and then randomly distributed that many dots in the corresponding area. But as you say, it looks like that's only for the US, whereas Canada and other countries just have one value for the entire country.
100
Reposted by Dallas Card
nlpandcss.bsky.social @nlpandcss.bsky.social · 03/03/2026
Deadline extended! Get your direct submissions in by March 8 AoE
074
Dallas Card @dallascard.bsky.social · 04/03/2026
If any of them are interested in podcasts, they could try looking at some subset of SPORC (e.g., Arts, Leisure, Music, Fiction, etc.): huggingface.co/datasets/bli...
huggingface.co
blitt/SPoRC · Datasets at Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
150
Dallas Card @dallascard.bsky.social · 04/03/2026
Looks great! Can't wait to read!
010
Dallas Card @dallascard.bsky.social · 22/02/2026
Finally, many thanks to Stuart Soroka (@snsoroka.bsky.social) for championing this work!
030
Dallas Card @dallascard.bsky.social · 22/02/2026
If you've made it this far, you might also want to check out Amber's earlier work on media storms: www.tandfonline.com/doi/abs/10.1..., or my student Ben Litterer's (@blitt.bsky.social) ACL paper on the same topic: aclanthology.org/2023.finding...
aclanthology.org
When it Rains, it Pours: Modeling Media Storms and the News Ecosystem
Benjamin Litterer, David Jurgens, Dallas Card. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023.
141
Dallas Card @dallascard.bsky.social · 22/02/2026
For additional details, including coding protocols, teaching resources, and side-by-side case comparisons, you can refer to the accompanying website: www.amber-boydstun.com/catching-fir...
amber-boydstun.com
Catching Fire Appendix
Cambridge Political Communication Element APPENDIX FOR   CATCHING FIRE IN THE NEWS
141
Dallas Card @dallascard.bsky.social · 22/02/2026
We also discuss additional factors that can influence the course of a storm, such as journalistic gatekeeping, attention fatigue, political activism, and strategic communication online. For a more in-depth summary, please take a look at Jill's thread here: bsky.app/profile/jill... or read the book!
121
Dallas Card @dallascard.bsky.social · 22/02/2026
The book is build around a series of paired case studies -- similar events, where one became a full-fledged media storm, and the other did not -- such as the Titan Submersible Implosion vs. the Messenia Migrant Boat Disaster, occurring just days apart in 2023.
132
Dallas Card @dallascard.bsky.social · 22/02/2026
The heart of this work uses the fire triangle model (heat, fuel, and oxygen) as a metaphor to characterize the necessary conditions for an event to become into a media storm -- those stories that are so pervasive in the news that they are practically inescapable.
151