Sign in

Abigail Jacobs

@azjacobs.bsky.social
2.6K followers 759 following 33 posts

Asst Prof of Information @ UMich thinking about assumptions built into AI

PostsRepliesMedia
Reposted by Abigail Jacobs
Meera Desai @madesai.bsky.social · 28/09/2026
Come find us at COLM, I’ll be presenting this paper next Thursday afternoon (10/8, oral session 6 and poster session 6). Full paper here: arxiv.org/abs/2609.08812
arxiv.org
What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reasoning, refusal). We ad...
082
Reposted by Abigail Jacobs
Meera Desai @madesai.bsky.social · 28/09/2026
To make our analysis possible, we collected model outputs and scores from 53 models on 56 capability and safety benchmarks. This data, including item-level model responses and score, is available here huggingface.co/datasets/mad...
Illustration of the dataset as a 3D stack of 56 grids, one per benchmark. Each grid has 53 rows (models) and up to 1,000 columns (questions per benchmark), with cells shaded to represent item-level results. The grids are colored by each benchmark's assigned concept: over-refusal, reasoning, knowledge, summarization, comprehension, sentiment analysis, refusal, safety detection, ethics, bias, privacy, and unsafe behavior.
152
Reposted by Abigail Jacobs
Meera Desai @madesai.bsky.social · 28/09/2026
Benchmarks inform how we use, govern, and deploy AI, but do they measure what they claim to? Adapting convergent and discriminant validity from the social sciences, we conduct a large-scale investigation of 56 benchmarks across 53 models and find evidence that several do not.
132
Reposted by Abigail Jacobs
Meera Desai @madesai.bsky.social · 28/09/2026
Excited to share our new paper, “What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks,” accepted as an oral at COLM! arxiv.org/pdf/2609.08812
Heatmap of average correlations between model rankings on benchmarks grouped into 11 assigned concepts: four capability concepts (reasoning, knowledge, comprehension, summarization) and seven safety concepts (over-refusal, refusal, safety detection, ethics, bias, privacy, unsafe behavior). Diagonal cells show within-concept correlations, ranging from 0.87 (knowledge) and 0.72 (over-refusal) down to 0.20 (bias) and 0.02 (safety detection). Reasoning, knowledge, and comprehension correlate with each other at 0.69 to 0.78, higher than reasoning's and comprehension's own within-concept values (0.66 and 0.68). Ethics correlates more with knowledge (0.70) than with itself (0.55), and bias correlates more with capability concepts (0.41 to 0.45) than with itself (0.20). Privacy and unsafe behavior correlate negatively with reasoning, knowledge, and comprehension (−0.41 to −0.49). Over-refusal and refusal correlate at −0.42.
16218
Reposted by Abigail Jacobs
If you know you know @privatechand.bsky.social · 01/08/2026
I keep thinking about this gorgeous essay about how LLMs- so often trained by inheritors of colonial English- has now rendered the kind of language that is natural to us (“fancy” vocabulary, passive voice, reduced relative clauses) artificial and suspect. marcusolang.substack.com/p/im-kenyan-...
marcusolang.substack.com
I'm Kenyan. I Don't Write Like ChatGPT. ChatGPT Writes Like Me.
I'm calm. I'm calm. I promise.
225095
Reposted by Abigail Jacobs
Don Moynihan @donmoyn.bsky.social · 31/07/2026
Very excited for this new essay at "Can We Still Govern?": The key claim is that planning for a progressive future cannot just focus on new policies. It also has to include - to start with - political power. What does a power-shifting perspective look like? 🧵 donmoynihan.substack.com/p/rebuilding...
donmoynihan.substack.com
Rebuilding Democracy Requires Rebalancing Power
What does a power-shifting agenda look like?
721376
Reposted by Abigail Jacobs
Alondra Nelson @alondra.bsky.social · 01/08/2026
This.
0123
Reposted by Abigail Jacobs
Carl T. Bergstrom @carlbergstrom.com · 31/07/2026
Nature did a short piece on our preprint about how LLMs will affect the practice of science; my more detailed thread is below.
nature.com
Scientists using LLMs will ‘do more, less well’, modelling study predicts
Research suggests that the combination of incentives to publish and the use of large language models will lead to more papers, but they will be less refined.
8440170
Reposted by Abigail Jacobs
Dallas Card @dallascard.bsky.social · 10/07/2026
As some may have heard me talk about at #ACL2026, I'm excited to share a new preprint on approaches to validation when using LLMs to measure concepts in social science, led by @madesai.bsky.social and @azjacobs.bsky.social !! Paper: arxiv.org/abs/2607.07915
Title page from "Validating LLMs in social science: Epistemic threats and emerging norms" by Meera Desai, Dallas Card, and Abigail Z. Jacobs
56819
Reposted by Abigail Jacobs
Amina Abdu @aminaabdu.bsky.social · 30/04/2026
In our new paper, we look at 50+ years of GAO reports to understand how administrative use of algorithms has changed the scope of legitimate practices within federal agencies, even when agency algorithms (unlike prior forms of state quantification) are seen as undermining legitimacy.
1192
Abigail Jacobs @azjacobs.bsky.social · 01/05/2026
Extremely proud of my (soon to be graduated!) student @aminaabdu.bsky.social and also of this just published work!
030
Reposted by Abigail Jacobs
Maria Antoniak @mariaa.bsky.social · 15/04/2026
I'll be at #ICML2026 on July 10-11 in Seoul to speak at the Workshop on Culture x AI: Evaluating AI as a Cultural Technology. The workshop is currently accepting submissions, with humanities, ML, HCI, and social/cognitive sciences all welcome. Submit papers by May 1! Join us in Seoul! 🇰🇷
doingaidifferently.org
Doing AI Differently | Culture × AI Workshop
24416
Reposted by Abigail Jacobs
Lawfare @lawfaremedia.org · 09/04/2026
"The constitution is not a neutral charter standing above the firm but, rather, a company document that prioritizes its mission, market position, safety research, and normative influence all at once." Lisa Klaassen and Ralph Schroeder detail their criticisms of Claude's Constitution.
lawfaremedia.org
The Code Is Not the Law: Why Claude’s Constitution Misleads
Anthropic’s appeals to constitutionalism and virtue-ethics risk obscuring where the power and accountability for shaping AI behavior lies.
02714
Reposted by Abigail Jacobs
UC Berkeley School of Information @berkeleyischool.bsky.social · 03/04/2026
Join us for a Bellwether Lecture! 🔔 University of Michigan School of Information Assistant Professor azjacobs.bsky.social will discuss structure, governance, and inequality in sociotechnical systems & hidden assumptions in machine learning. 📅 April 8, 12:10-1:30 pm 📍 210 South Hall & Online
ischool.berkeley.edu
The Hidden Governance of AI and Other Threats to Democracy
Apr 8, 2026, 12:10 pm - Abigail Jacobs researches structure, governance, and inequality in sociotechnical systems and the hidden assumptions in machine learning.
0163
Reposted by Abigail Jacobs
Eileen Clancy 🧿 @clancyny.bsky.social · 22/03/2026
Excellent, important, and clarifying. Historian of computing Kevin Baker teaches us about the history of aerial targeting. And, Baker says, contra news accounts, Anthropic and Claude didn't select the Minab girls' school as a target. @kevinbaker.bsky.social
12413
Abigail Jacobs @azjacobs.bsky.social · 31/03/2026
Well put: “Formal review persists, but substantive discretion has migrated upstream…When objectives are embedded in architecture, administrative errors and political misjudgments are operationalized at scale. What appears as a dispute about fairness is therefore a deeper institutional misalignment”
071
Reposted by Abigail Jacobs
Tech Policy Press @techpolicypress.bsky.social · 31/03/2026
AI systems aren’t just supporting decisions—they’re structuring how public authority is exercised, writes Michael A. Santoro. “Human in the loop” isn’t enough. Accountability must be built upstream through guardrails embedded in system design, not added after the fact, he argues.
buff.ly
Where is Accountability When Governments Deploy AI?
AI systems now structure public authority; human-in-the-loop oversight is insufficient, requiring upstream guardrails argues Michael A. Santoro.
1239
Reposted by Abigail Jacobs
Ben Recht @beenwrekt.bsky.social · 26/03/2026
Coarsely summarizing perspectives on the bureaucratic culture of language models by @himself.bsky.social, @azjacobs.bsky.social, and Lily Chumley.
argmin.net
The Poetics of Bureaucracy
Language models are a bureaucratic technology
0223
Reposted by Abigail Jacobs
David Rothkopf @djrothkopf.bsky.social · 28/10/2025
If you have any interest in the future of AI, please join us for another really insightful conversation with @alondra.bsky.social. She's brilliant but better yet, she's right!
0389
Abigail Jacobs @azjacobs.bsky.social · 28/10/2025
AI as governance -- @himself.bsky.social on how AI reshapes markets, bureaucracy, democracy...and culture. Very happy ot see this getting the mainstream social science treatment. www.annualreviews.org/content/jour... I can't believe I missed this paper coming out!
annualreviews.org
AI as Governance
Political scientists have had remarkably little to say about artificial intelligence (AI), perhaps because they are dissuaded by its technical complexity and by current debates about whether AI might ...
1233
Reposted by Abigail Jacobs
A. Feder Cooper @afedercooper.bsky.social · 15/07/2025
Feeling so excited + grateful to be representing this paper at #ICML! Please stop by to talk about how to do more valid measurement for evaling gen AI systems! Work led by the incomparable @hannawallach.bsky.social and @azjacobs.bsky.social as a part of Microsoft’s AI and Society initiative!!
0122
Abigail Jacobs @azjacobs.bsky.social · 26/02/2025
“If ___ ran a mini nuclear power plant” seems like a strong vibe for the day
020
Reposted by Abigail Jacobs
Carrie Brown @brizzyc.bsky.social · 12/02/2025
As ever, Tressie McMillan Cottom has the most astute analysis of how to read Musk's behavior. www.nytimes.com/2025/02/12/o...
nytimes.com
Opinion | Look Past Elon Musk’s Chaos. There’s Something More Sinister at Work.
Everything is content.
2196
Abigail Jacobs @azjacobs.bsky.social · 12/02/2025
big day to submit an article on how "efficiency" is used to undermine legitimacy of the administrative state www.nytimes.com/2025/02/11/u... (gift link)
nytimes.com
At Oval Office, Musk Makes Broad Claims of Federal Fraud Without Proof (Gift Article)
The billionaire, whose federal cost-cutting team has been operating in secrecy, asserted that he had uncovered waste and fraud across the bureaucracy, without providing evidence.
181
Reposted by Abigail Jacobs
Maria Antoniak @mariaa.bsky.social · 15/12/2024
"there's a lot of qualitative work that goes into designing quantitative metrics" -- @azjacobs.bsky.social "how do we translate between benchmark performance and what it will really be like to use a model" -- Su Lin Blodgett
0478
Reposted by Abigail Jacobs
Dr Abeba Birhane @abeba.blacksky.app · 12/12/2024
"Overall, the starting list constitutes at best a narrow coverage of the risks the technology is likely to pose, & at worst a (partial) red herring poised to direct significant risk mitigation efforts to building on inappropriate foundations." @yjernite.bsky.social et al on the AI Act Systemic Risks
0114
Reposted by Abigail Jacobs
Hanna Wallach @hannawallach.bsky.social · 14/12/2024
Evaluating Generative AI Systems is a Social Science Measurement Challenge: arxiv.org/abs/2411.10939 TL;DR: The ML community would benefit from learning from and drawing on the social sciences when evaluating GenAI systems.
arxiv.org
Evaluating Generative AI Systems is a Social Science Measurement Challenge
Across academia, industry, and government, there is an increasing awareness that the measurement tasks involved in evaluating generative AI (GenAI) systems are especially difficult. We argue that thes...
141
Reposted by Abigail Jacobs
Hanna Wallach @hannawallach.bsky.social · 14/12/2024
New paper on why machine "unlearning" is much harder than it seems is now up on arXiv: arxiv.org/abs/2412.06966 This was a huuuuuge cross-disciplinary effort led by @msftresearch.bsky.social FATE postdoc @grumpy-frog.bsky.social!!!
arxiv.org
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy, Research, and Practice
We articulate fundamental mismatches between technical methods for machine unlearning in Generative AI, and documented aspirations for broader impact that these methods could have for law and policy. ...
27224
Reposted by Abigail Jacobs
Hal Daumé III @haldaume3.bsky.social · 13/11/2024
The AI Interdisciplinary Institute at the University of Maryland (AIM) is hiring 40 new faculty members in all areas of AI, particularly: - accessibility, - sustainability, - social justice, and - learning; building on computational, humanistic, or social scientific approaches to AI. >
16419
Reposted by Abigail Jacobs
Oskar van der Wal @ovdw.bsky.social · 07/08/2024
Working on #bias & #discrimination in #NLP? Passionate about integrating insights from different disciplines? And do you want to discuss current limitations of #LLM bias mitigation work? 🤖 👋Join the workshop New Perspectives on Bias and Discrimination in Language Technology 4&5 Nov in #Amsterdam!
wai-amsterdam.github.io
Workshop: New Perspectives on Bias and Discrimination in Language Technology.
Workshop: New Perspectives on Bias and Discrimination in Language Technology.
154
Abigail Jacobs @azjacobs.bsky.social · 08/05/2024
Not enough people are worried about this
161
Reposted by Abigail Jacobs
Dallas Card @dallascard.bsky.social · 01/04/2024
I'm excited to share that the journal version of our paper, "An archival perspective on pretraining data", is now available (open access) from Patterns! This project was led by @madesai.bsky.social, along with Irene Pasquetto, @azjacobs.bsky.social, and myself www.cell.com/patterns/ful... 1/n
cell.com
An archival perspective on pretraining data
Large language models depend crucially on the data they are trained on. The authors consider how these pretraining datasets, like archives, are diverse, sociocultural collections that mediate knowledg...
173
Abigail Jacobs @azjacobs.bsky.social · 31/10/2023
New executive order just dropped, with lots of roles for sociotechnical work around AI with relatively fast turnaround… but also a bat flying around the White House logo because the executive order is also spooky 🦇🦇 www.whitehouse.gov/briefing-roo...
whitehouse.gov
Executive Order on the Safe, Secure, and Trustworthy Development and Use of Artificial Intelligence ...
By the authority vested in me as President by the Constitution and the laws of the United States of America, it is hereby ordered as follows:      Section 1.  Purpose.  Artificial intelligence (A...
070
Reposted by Abigail Jacobs
World Privacy Forum @worldprivacyforum.bsky.social · 15/10/2023
One of the many, many concepts we discuss in our upcoming report is Measurement Modeling as it relates to AI governance. Discussed this with @azjacobs.bsky.social earlier this year.
021
Abigail Jacobs @azjacobs.bsky.social · 17/06/2023
Spent the last week at #FAccT2023 - filled with some exciting work and exciting scholars on measurement, policy, social impacts emerging from ML. Happy to see this letter come out of it: facctconference.org/2023/harm-polic…
facctconference.org
ACM FAccT - Statement on AI Harms and Policy
051
Abigail Jacobs @azjacobs.bsky.social · 17/06/2023
Felt cute, might delete later
A fluffy regal cat looks up while sitting next to a slowly-dying potted plant
0111