Sign in

Ryan Steed

@rbsteed.com
570 followers 167 following 24 posts

AI Policy Fellow @ Princeton | PhD Carnegie Mellon | privacy, accountability, & algorithmic systems

PostsRepliesMedia
Reposted by Ryan Steed
Suresh Venkatasubramanian @geomblog.bsky.social · 01/04/2026
It's about what's hidden, and what new deficiencies the tech carries with it. @victorojewale.bsky.social opines on the evolution of deployed AI and its limits. victorojewale.substack.com/p/from-exper...
victorojewale.substack.com
Every Generation of AI Solves One Problem and Hides Another
What brittleness looked like then, what it looks like now, and why “adding a layer” never fully solves it
053
Reposted by Ryan Steed
Suresh Venkatasubramanian @geomblog.bsky.social · 02/03/2026
It's been a journey of nearly 3 years, but I'm very excited to announce the CNTR AISLE Portal! 🚀 cntr-aisle.org It’s a new way to review and evaluate the 1,000+ AI bills introduced in the U.S. over the last three years. Check out the Bill Library and our Profiles#AIPolicy #OpenData
cntr-aisle.org
CNTR AISLE
CNTR AISLE Portal
12815
Ryan Steed @rbsteed.com · 19/02/2026
GLMMs are just one approach — we’re looking forward to more work on statistical frameworks for AI evaluation. Send questions/comments to caisi-metrology@nist.gov. Paper (w/ the talented Drew Keller, Kweku Kwegyir-Aggrey, Anita Rao, Julia Sharp, and Stevie Bergman): nvlpubs.nist.gov/nistpubs/ai/...
010
Ryan Steed @rbsteed.com · 19/02/2026
GLMMs have other benefits, too: - We can estimate question difficulties to identify problematic questions and other patterns in benchmarks. - Variance decomposition (between- and within-questions) can highlight nuances in performance between tasks, languages, and other subsets of a benchmark.
Fig. 6a from the paper: Distribution of estimated question difficulties by domain and labeled difficulty (GPQA-Diamond). Distribution of random effects in each domain. Chemistry questions were particularly difficult
for the LLMs we tested. Each dot indicates a GPQA-Diamond question’s GLMM-estimated difficulty (i.e., random effect value). Box plots display quartiles and violin plots display estimated density. These estimates show that GPQA-Diamond’s chemistry questions were particularly difficult for the 22 tested LLMs. On the other hand, question difficulty for LLMs has a weak relationship with question-writer-labeled difficulty. This may suggest that humans and the tested LLMs find different questions difficult, and/or could call into question whether writer annotations are accurate even for human difficulty.Fig. 6b from the paper: Distribution of estimated question difficulties by domain and labeled difficulty (GPQA-Diamond). Distribution of random effects for the 191 questions at the three most common writer-annotated difficulty levels. Question difficulty for LLMs has a weak relationship with
human-labeled difficulty. Each dot indicates a GPQA-Diamond question’s GLMM-estimated difficulty (i.e., random effect value). Box plots display quartiles and violin plots display estimated density. These estimates show that GPQA-Diamond’s chemistry questions were particularly difficult for the 22 tested LLMs. On the other hand, question difficulty for LLMs has a weak relationship with question-writer-labeled difficulty. This may suggest that humans and the tested LLMs find different questions difficult, and/or could call into question whether writer annotations are accurate even for human difficulty.
100
Ryan Steed @rbsteed.com · 19/02/2026
Ideally, evaluators should use a statistical model to explicitly define the estimand & other statistical assumptions. We propose one approach using generalized linear mixed models. GLMMs can often estimate uncertainty more precisely than typical “regression-free” approaches.
100
Ryan Steed @rbsteed.com · 19/02/2026
AI evals rarely specify which question is being answered — but the choice matters, especially when it comes to computing error bars. (Assuming error bars are included at all…) In particular, error bars for generalized accuracy tend to be larger and may yield different rankings.
Fig. 1 from the paper: Comparing accuracy estimates (GPQA-Diamond). Lower plots show the estimated accuracy of a selection of tested LLMs* with 95% confidence intervals. Upper plots show corresponding confidence interval (CI) widths. Generalized accuracy CIs are larger than benchmark accuracy CIs because they account for the selection of benchmark items from a superpopulation. Notably, some pairs of LLMs may have significantly different benchmark accuracy but not generalized accuracy. The simple average (pink) estimates reflect the average across all n benchmark questions with standard error calculated as standard deviation of results divided by √n. For estimates of benchmark accuracy, the simple average method results in under-confident CIs compared to a valid regression-free method (blue). For estimates of generalized accuracy, the simple average method provides valid CIs, but precision can be increased by running more trials per item (as in the regression-free method). Generalized linear mixed model (GLMM, orange) estimates require additional assumptions but further increase precision.
100
Ryan Steed @rbsteed.com · 19/02/2026
We identify two distinct questions about accuracy: - Benchmark accuracy: How well does the LLM perform on this specific, fixed benchmark? - Generalized accuracy: How well would the LLM perform across the larger population of questions similar to those in this benchmark?
100
Ryan Steed @rbsteed.com · 19/02/2026
AI benchmark evals commonly report "accuracy" metrics — but what’s really being measured? And how should we compute the error bars? New NIST report from my team at CAISI outlines a better statistical framework for eval analysis: www.nist.gov/news-events/...
nist.gov
New Report: Expanding the AI Evaluation Toolbox with Statistical Models
We're Hiring!Please visit the CAISI Careers Page to learn more abou
110
Reposted by Ryan Steed
Avijit Ghosh @evijit.io · 17/02/2026
This has been a massive community project, and we need you all to participate! See more: evalevalai.com/projects/eve...
073
Ryan Steed @rbsteed.com · 17/02/2026
Had the chance to give feedback on this project on CAISI’s behalf. I’m very excited to see this develop!
030
Ryan Steed @rbsteed.com · 11/02/2026
If this kind of work speaks to you, come work with us! My team at CAISI is hiring an Applied Systems AI Research Scientist, among many other roles. www.nist.gov/caisi/career...
nist.gov
Careers at CAISI
About CAISICAISI, within NIST at the Department of Commerce, acts as a startup within g
022
Ryan Steed @rbsteed.com · 11/02/2026
CAISI invites input on any aspect of this draft, including from orgs that conduct AI evals and from users of eval reports (for decision-making, procurement, integration, etc.) Public comment closes March 31 — details here: www.nist.gov/news-events/...
nist.gov
Towards Best Practices for Automated Benchmark Evaluations
Comments Sought on Initial Public Draft of NIST AI 800-2 through March 31
010
Ryan Steed @rbsteed.com · 11/02/2026
Automated benchmarks are not all you need, but they are popular tools in AI development. Hoping this doc is a foundation for future guidelines on field testing and other kinds of evals.
Table I.1 from https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-2.ipd.pdf
100
Ryan Steed @rbsteed.com · 11/02/2026
Section 3 covers critical practices related to responsible and transparent reporting — including uncertainty quantification, reproducibility, and properly qualified claims.
Table 3.1 from https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-2.ipd.pdf
100
Ryan Steed @rbsteed.com · 11/02/2026
Section 2 dives into the nitty-gritty operational details of setting up and running a benchmark — including helpful lists of relevant settings and design principles.
100
Ryan Steed @rbsteed.com · 11/02/2026
I’m especially excited about the focus on practical measurement validity. Sections 1 describes ways to assess the relationship between the contents of a benchmark and what evaluators really want to measure.
Table 1.1 from https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.800-2.ipd.pdf
100
Ryan Steed @rbsteed.com · 11/02/2026
Excited to co-author a new public draft from NIST CAISI on best practices for automated benchmark evals. We want your feedback! Public comment open til March 31. Highlights below 🧵 www.nist.gov/news-events/...
nist.gov
Towards Best Practices for Automated Benchmark Evaluations
Comments Sought on Initial Public Draft of NIST AI 800-2 through March 31
120
Ryan Steed @rbsteed.com · 14/01/2026
NIST CAISI is also hiring post-docs — applications due Feb. 1. Come work with our team on AI evaluations and metrology! Apply here: ra.nas.edu/RAPLab10/Opp... More detail: www.linkedin.com/posts/astevi...
010
Ryan Steed @rbsteed.com · 09/01/2026
Late notice, but: NIST has a multi-year, funded graduate fellowship + summer intern program (potentially with our team at CAISI!). Due January 15. stemfellowships.org/applicants/
stemfellowships.org
Applicants
Fellowship Information *** Application is currently OPEN *** Number of Awards: Varies annually  Schedule: Online applications open in late August and close on January 15, 2026. Type: Fellowshi…
122
Reposted by Ryan Steed
AI Accountability Lab @aial.ie · 17/12/2025
Are you passionate about exploring what a conceptually cogent, methodologically sound, and well-founded AI evaluation and safety research might look like? Come do a PhD with us. Closing Date: 10 February 2026 Apply here aial.ie/hiring/phd-a...
About the PhD

Audits and evaluation of AI systems — and the broader context that AI systems operate in — have become central to conceptualising, quantifying, measuring and understanding the operations, failures, limitations, underlying assumptions, and downstream societal implications of AI systems. Existing AI audit and evaluation efforts are fractured, done in a siloed and ad-hoc manner, and with little deliberation and reflection around conceptual rigour and methodological validity.

This PhD is for a candidate that is passionate about exploring what a conceptually cogent, methodologically sound, and well-founded AI evaluation and safety research might look like. This requires grappling with questions such as:

    What does it mean to represent “ground truth” in proxies, synthetic data, or computational simulation?
    How do we reliably measure abstract and complex phenomena?
    What are the epistemological or methodological implications of quantification and measurement approaches we choose to employ? Particularly, what underlying presuppositions, values, or perspectives do they entail?
    How do we ensure the lived experiences of impacted communities play a critical role in the development and justification of measurement metrics and proxies?
    Through exploration of these questions, the candidate is expected to engage with core concepts in the philosophy of science, history of science, Black feminist epistemologies, and similar schools of thought to develop an in-depth understanding of existing practices with the aim of applying it to advance shared standards and best practice in AI evaluation.

The candidate is expected to integrate empirical (for example, through analysis or evaluation of existing benchmarks) or practical (for example, by executing evaluation of AI systems) components into the overall work.
01610
Reposted by Ryan Steed
Deb Raji @rajiinio.bsky.social · 11/12/2025
US CAISI is hiring -- the internal govt name for the role is "IT Specialist" but it is effectively a research scientist role! Salary is $120,579 to - $195,200 per year, and you get to work on AI evaluation within government agencies! Job posting (**closes EOD 12/28/2025**): lnkd.in/exJgkqr5
12410
Ryan Steed @rbsteed.com · 08/12/2025
Note that this position requires a specially formatted resume, 2 pages max: help.usajobs.gov/faq/applicat...
help.usajobs.gov
USAJOBS Help Center - How do I write a resume for a federal job?
USAJOBS Help Center
010
Ryan Steed @rbsteed.com · 08/12/2025
Also, our team is hiring an AI Research Scientist! www.usajobs.gov/job/851528400
usajobs.gov
USAJOBS connects job seekers with federal jobs across the United States and around the world as the official employment site for the federal government
<p>NIST works with industry and science to advance innovation and improve quality of life. We're looking for an IT Specialist (AI) to join our team!<br> <br> <a href="https://www.nist.gov/caisi">CAISI...
1107
Ryan Steed @rbsteed.com · 04/12/2025
Also, belated announcement that I joined @steviebergman.bsky.social’s wonderful Applied Systems team at CAISI — with @anitakrao.bsky.social, Drew Keller, & (formerly) @kwekuka.bsky.social. 

More to come!
020
Ryan Steed @rbsteed.com · 04/12/2025
"Building gold-standard AI systems requires gold-standard AI measurement science... Today, many evaluations of AI systems do not precisely articulate what has been measured, much less whether the measurements are valid."

 We highlight open q's about construct validity, field studies, and more.
120
Ryan Steed @rbsteed.com · 04/12/2025
Our team at NIST's Center for AI Standards and Innovation (CAISI) just released a blog post with open questions for AI measurement science: www.nist.gov/blogs/caisi-...
nist.gov
Accelerating AI Innovation Through Measurement Science
Building gold-standard AI systems requires gold-standard AI measurement science – the scientific study of methods used to assess AI systems’ properties and impacts. NIST works to improve measurements ...
151
Reposted by Ryan Steed
Deb Raji @rajiinio.bsky.social · 04/12/2025
US CAISI (the equivalent of the US "AI Safety Institute") just put out their approach to AI measurement & there's such a significant portion on construct validity (nist.gov/blogs/caisi-...). Great to see this after ongoing advocating about this issue (arxiv.org/abs/2511.04703)!
1258
Reposted by Ryan Steed
Emma Harvey @emmharv.bsky.social · 14/07/2025
After having such a great time at #CHI2025 and #FAccT2025, I wanted to share some of my favorite recent papers here! I'll aim to post new ones throughout the summer and will tag all the authors I can find on Bsky. Please feel welcome to chime in with thoughts / paper recs / etc.!! 🧵⬇️:
25510
Ryan Steed @rbsteed.com · 24/07/2025
Thank you :) Love this thread
020
Reposted by Ryan Steed
Nari Johnson @narijohnson.bsky.social · 14/05/2025
Q: What do school buses, desktop computers, and AI have in common? A: The same decades-old laws and processes apply when governments go to purchase them. Our new #FAccT2025 paper examines how these legacy public procurement norms apply to AI. 🧵
A title slide with the paper title: "Legacy Procurement Practices Shape How U.S. Cities Govern AI". The title includes a small illustration that is a simple chart: A government provides an AI vendor with money, and in exchange the vendor provides the government with an AI system.
281
Reposted by Ryan Steed
Victor Ojewale @victorojewale.bsky.social · 28/04/2025
Excited to present "Towards AI Accountability Infrastructure: Gaps and Opportunities in AI Audit Tooling" at #CHI2025 tomorrow(today)! 🗓 Tue, 29 Apr | 9:48–10:00 AM JST (Mon, 28 Apr | 8:48–9:00 PM ET) 📍 G401 (Pacifico North 4F) 📄 dl.acm.org/doi/10.1145/...
2218
Reposted by Ryan Steed
Deb Raji @rajiinio.bsky.social · 22/04/2025
Our lecture topic in our ML eval class today. ✨ Evals have always played an important socio technical role in policy decision-making and surprisingly little changes when you add "AI" to the mix
0112
Reposted by Ryan Steed
Garrett Johnson @garjoh.bsky.social · 19/03/2025
Great webinar featuring Alessandro Acquisti on "The Economics of Privacy at a Crossroads." Alessandro has such a wide view of the privacy literature: both its history and inside & outside of economics. www.youtube.com/watch?v=ecc9...
youtube.com
The Economics of Privacy at a Crossroads | Luohan Webinar
YouTube video by Luohan Academy
082
Ryan Steed @rbsteed.com · 11/03/2025
TIL Minecraft has responsible AI lessons for kids education.minecraft.net//content/dam...
education.minecraft.net
000
Reposted by Ryan Steed
Dr Abeba Birhane @abeba.blacksky.app · 18/02/2025
I was in Paris last week for the #AIActionSummit & was honoured to participate in the “AI in public interest” panel. 👇🏾 my thoughts & reflections on what AI in public interest is/isn’t & some concrete steps/initiatives for `bending the arc of AI towards the public interest' aial.ie/pages/aiparis/
aial.ie
Bending the arc of AI towards the public interest
By Abeba Birhane, 18/02/2025 Following the first in Bletchley Park in 2023 and the second in Seoul in 2024, the third AI Action Summit took place in February 2025 in Paris. In the context of previous ...
921779
Reposted by Ryan Steed
Todd Feathers @toddfeathers.bsky.social · 11/02/2025
When Toledo police rolled out their Fusus system, allowing officers to tap into the live feeds of privately owned cameras, they promised to only use the power in emergency situations. We obtained data that tells a very different story about when, and who, TPD watches. gizmodo.com/clearly-disc...
gizmodo.com
‘Clearly Discrimination’: How a City Uses Fusus to Spy on Its Poorest Residents
Fusus’s technology allows police to tap into live feeds from public and privately owned surveillance cameras. In Toledo, Ohio, cops use the power to watch one particular type of location.
34131
Reposted by Ryan Steed
Karen Hao @karenhao.bsky.social · 27/01/2025
As someone who has reported on AI for 7 years and covered China tech as well, I think the biggest lesson to be drawn from DeepSeek is the huge cracks it illustrates with the current dominant paradigm of AI development. A long thread. 1/
21161252340
Reposted by Ryan Steed
Ryan Calo @rcalo.bsky.social · 11/01/2025
How often does the FTC base its enforcement proceedings on third party research? Often! ftcreverse.engineering
ftcreverse.engineering
FTC RE
1112
Reposted by Ryan Steed
Dr Abeba Birhane @abeba.blacksky.app · 10/12/2024
oh, this is out and they make me sound like an angel but a really good read “We are not evaluating systems for some hypothetical, potential risks in the future,” Birhane says www.fastcompany.com/91238006/how...
fastcompany.com
How Abeba Birhane is cleaning up AI’s dirty data
“We are not evaluating systems for some hypothetical, potential risks in the future. These audits are uncovering actual real issues, real problems, whether it's racism, sexism, or encoding of stereoty...
717655
Reposted by Ryan Steed
Ken Holstein @kenholstein.bsky.social · 10/12/2024
Come work with @jessicahullman.bsky.social and I as a postdoc! We'll apply approaches from HCI, AI/ML, statistics, and decision theory to design new tools & methods for evaluating human-AI decision-making (e.g., tools to elicit, represent, & validate specifications of real-world decision problems).
1183
Reposted by Ryan Steed
Ben Winters @benwinters.bsky.social · 03/12/2024
tough morning for data brokers, great morning for people in addition to the CFPB FCRA rule (bsky.app/profile/benw...), the FTC announced a case against one who sold location data of visits to healthcare and places of worship, securing a ban on sale of this type of data www.ftc.gov/news-events/...
media.tenor.com
a man with a beard is sitting in a golf cart with his hand on his head .
Alt: dj khaled "another one" meme
042
Reposted by Ryan Steed
Hypervisible @hypervisible.blacksky.app · 03/12/2024
The proposed consent order with IntelliVision Technologies would prevent the company “from making misleading claims about its software. That would include any misleading statements about its performance identifying people of different genders, ethnicities, and skin tones.”
yahoo.com
The FTC is cracking down on a firm that claims its AI-powered face recognition tech has 'zero' bias
A software provider won’t be allowed to lie about the accuracy of its artificial intelligence-powered facial recognition technology under a new proposed order from the Biden administration.
2338
Ryan Steed @rbsteed.com · 28/11/2024
Something (someone) to be thankful for today — very excited to see this!!
1141
Reposted by Ryan Steed
Dr Abeba Birhane @abeba.blacksky.app · 28/11/2024
Friends, this week marks a monumental step forward in my path towards pushing for a greater ecology of accountability in the age of AI. Thrilled to officially announce the AI Accountability Lab (AIAL) @theaial.bsky.social at Trinity College Dublin is launching today, Thursday Nov 28, 2024 aial.ie
aial.ie
AI Accountability Lab
Stewarding a greater ecology of accountability in the age of AI
911253310
Reposted by Ryan Steed
Leon Yin @leonyin.org · 20/11/2024
Professors and researchers: are you booking speakers for the Spring semester? I have talks and workshops on data-driven investigations using crowdsourcing with WhatsApp, auditing black box AI systems and advanced web scraping. inspectelement.org/apis.html gist.github.com/yinleon/4e39...
inspectelement.org
Finding Undocumented APIs
Introduction, case studies, and exercises for finding and using undocumented APIs hidden in plain sight.
1226
Reposted by Ryan Steed
Jessica Hullman @jessicahullman.bsky.social · 19/11/2024
Differential privacy is a SOTA approach for protecting individual privacy in data releases. But how it interacts w/common expectations & workflows can be hard to judge. New article for policy-makers w/P. Nanayakkara summarizes some gaps between theory & practice journals.sagepub.com/doi/abs/10.1...
journals.sagepub.com
Sage Journals: Discover world-class research
Subscription and open access journals from Sage, the world's leading independent academic publisher.
2236