Sign in

sanmikoyejo.bsky.social

@sanmikoyejo.bsky.social
86 followers 30 following 20 posts
PostsRepliesMedia
Reposted by @sanmikoyejo.bsky.social
Joachim Baumann @joachimbaumann.bsky.social · 02/10/2026
SWE-chat v2 is here! The largest dataset of coding agent interactions from real users in the wild has gotten even larger: 230K prompts from 18K sessions. SWE-chat has enabled incredible research (🧵) – excited to see what v2 unlocks for the community. Come find us at COLM next week! swe-chat.com
252
Reposted by @sanmikoyejo.bsky.social
Meera Desai @madesai.bsky.social · 28/09/2026
Come find us at COLM, I’ll be presenting this paper next Thursday afternoon (10/8, oral session 6 and poster session 6). Full paper here: arxiv.org/abs/2609.08812
arxiv.org
What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks
Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reasoning, refusal). We ad...
082
Reposted by @sanmikoyejo.bsky.social
Meera Desai @madesai.bsky.social · 28/09/2026
I learned so much working on this with @angelinawang.bsky.social , @hannawallach.bsky.social , Sang Truong, Alex Chouldechova, @afedercooper.bsky.social, Jean Garcia-Gathright, Daniel Ho, @azjacobs.bsky.social, @sanmikoyejo.bsky.social, and Nick Pangakis!
151
Reposted by @sanmikoyejo.bsky.social
Meera Desai @madesai.bsky.social · 28/09/2026
Overall, we show that convergent and discriminant validity are valuable lenses for systematically assessing benchmark validity. Our approach can be useful for benchmark developers in assessing individual benchmarks, as well as researchers looking to assess groups of benchmarks.
121
Reposted by @sanmikoyejo.bsky.social
Meera Desai @madesai.bsky.social · 28/09/2026
We also use our approach to interrogate a few individual benchmarks, and find evidence that BBQ, a bias benchmark, may be inadvertently measuring reasoning rather than bias. This is especially important because BBQ is often the only bias benchmark used in commercial model releases.
131
Reposted by @sanmikoyejo.bsky.social
Meera Desai @madesai.bsky.social · 28/09/2026
In some cases, benchmarks correlate more strongly by shared design elements than the concepts they claim to measure. For example, bias benchmarks that claim to measure gender and race bias cluster by task rather than these demographic targets.
Dendrogram clustering bias benchmarks by similarity, with each benchmark's gender bias and race bias versions colored separately. Every benchmark's gender and race versions pair with each other first: DiscrimEval, BBQ-accuracy, CALM, BBQ-bias, and DecodingTrust each form their own pair. None of them cluster by demographic target. GenMO, which has only a gender version, joins the DiscrimEval pair. DecodingTrust is the most distant from the other benchmarks.
131
Reposted by @sanmikoyejo.bsky.social
Meera Desai @madesai.bsky.social · 28/09/2026
Correlations are strong among benchmarks that claim to measure reasoning, knowledge, and comprehension, suggesting that these benchmarks may not be capturing distinct concepts.
Heatmap of average correlations between model rankings on capability benchmarks, grouped by assigned concept. Correlations between reasoning, knowledge, and comprehension benchmarks are high (0.69 to 0.78). They exceed the within-concept correlations for reasoning (0.66) and comprehension (0.68), though not for knowledge (0.87). Summarization benchmarks correlate weakly with the other three (0.06 to 0.18).
121
Reposted by @sanmikoyejo.bsky.social
Meera Desai @madesai.bsky.social · 28/09/2026
We find that benchmarks that claim to measure similar safety concepts often correlate weakly, which reflect that safety concepts like “bias” are inconsistently (but potentially validly) conceptualized across benchmarks.
Box plots of Spearman rank correlations between pairs of benchmarks that share an assigned concept, with each dot representing one benchmark pair. Capability benchmarks correlate strongly with others in the same concept: reasoning and knowledge pairs mostly fall between about 0.6 and 1. Safety benchmarks correlate more weakly and vary more. Refusal pairs (n=10 benchmarks) range from slightly below 0 to about 0.85, safety detection pairs (n=3) center near 0, and bias pairs (n=8) have the widest spread, from about −0.7 to 0.9
141
Reposted by @sanmikoyejo.bsky.social
Meera Desai @madesai.bsky.social · 28/09/2026
To make our analysis possible, we collected model outputs and scores from 53 models on 56 capability and safety benchmarks. This data, including item-level model responses and score, is available here huggingface.co/datasets/mad...
Illustration of the dataset as a 3D stack of 56 grids, one per benchmark. Each grid has 53 rows (models) and up to 1,000 columns (questions per benchmark), with cells shaded to represent item-level results. The grids are colored by each benchmark's assigned concept: over-refusal, reasoning, knowledge, summarization, comprehension, sentiment analysis, refusal, safety detection, ethics, bias, privacy, and unsafe behavior.
152
Reposted by @sanmikoyejo.bsky.social
Meera Desai @madesai.bsky.social · 28/09/2026
To adapt convergent and discriminant validity to benchmarks, we assessed whether benchmarks that purport to measure similar concepts correlate with each other, and whether benchmarks that purport to measure dissimilar concepts do not. We ask the same questions at the item level using IRT models.
131
Reposted by @sanmikoyejo.bsky.social
Meera Desai @madesai.bsky.social · 28/09/2026
Benchmarks inform how we use, govern, and deploy AI, but do they measure what they claim to? Adapting convergent and discriminant validity from the social sciences, we conduct a large-scale investigation of 56 benchmarks across 53 models and find evidence that several do not.
132
Reposted by @sanmikoyejo.bsky.social
Meera Desai @madesai.bsky.social · 28/09/2026
Excited to share our new paper, “What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks,” accepted as an oral at COLM! arxiv.org/pdf/2609.08812
Heatmap of average correlations between model rankings on benchmarks grouped into 11 assigned concepts: four capability concepts (reasoning, knowledge, comprehension, summarization) and seven safety concepts (over-refusal, refusal, safety detection, ethics, bias, privacy, unsafe behavior). Diagonal cells show within-concept correlations, ranging from 0.87 (knowledge) and 0.72 (over-refusal) down to 0.20 (bias) and 0.02 (safety detection). Reasoning, knowledge, and comprehension correlate with each other at 0.69 to 0.78, higher than reasoning's and comprehension's own within-concept values (0.66 and 0.68). Ethics correlates more with knowledge (0.70) than with itself (0.55), and bias correlates more with capability concepts (0.41 to 0.45) than with itself (0.20). Privacy and unsafe behavior correlate negatively with reasoning, knowledge, and comprehension (−0.41 to −0.49). Over-refusal and refusal correlate at −0.42.
16419
Reposted by @sanmikoyejo.bsky.social
Joachim Baumann @joachimbaumann.bsky.social · 04/09/2026
After Science News, our recent ICML paper also got featured in @science.org! Article: www.science.org/content/arti... Paper: arxiv.org/pdf/2605.03202 @dirkhovy.bsky.social @sanmikoyejo.bsky.social #science #PeerReview #research
Screenshot of an article in Science Magazine on the opportunities and risks of AI used for peer review featuring an ICML conference paper by Joachim Baumann et al., 2026.
1104
Reposted by @sanmikoyejo.bsky.social
RCMLR workshop @ NeurIPS 2026 @rcmlr26.bsky.social · 12/08/2026
RCMLR is a new #NeurIPS2026 workshop (Sydney), bringing together ML researchers, clinicians, policymakers and science communicators to discuss the responsible communication of biomedical ML research. Submissions and reviewers welcome until 29 August: translatingmlresearch.github.io/RCMLR/
translatingmlresearch.github.io
Responsible Communication of Machine Learning Research in Biomedicine
142
Reposted by @sanmikoyejo.bsky.social
Astrobites @astrobites.bsky.social · 15/06/2026
From Sowkhya Shanbhog: Meet Dr. Sanmi Koyejo: Stanford computer scientist, AI researcher, and AAS plenary speaker working to make artificial intelligence a more trustworthy partner in scientific discovery. ⚛️ 🔭 ☄️ 🧪astrobites.org/2026/06/14/aas248-sanmi-koyejo/
astrobites.org
Meet the AAS 248 Plenary Speakers: Dr. Sanmi Koyejo
Meet Dr. Sanmi Koyejo: Stanford computer scientist, AI researcher, and AAS plenary speaker working to make artificial intelligence a more trustworthy partner in scientific discovery.
163
Reposted by @sanmikoyejo.bsky.social
AFAA 2026 @ ICLR @afciworkshop.bsky.social · 26/04/2026
Shoutout to our incredible organizing team for yet another successful workshop @prakharg.bsky.social, @zeyutang.bsky.social, @adoubleva.bsky.social, Miriam Rateike, Jamelle Watson-Daniels, Golnoosh Farnadi, @jessicaschrouff.bsky.social, @sanmikoyejo.bsky.social
031
Reposted by @sanmikoyejo.bsky.social
Association for Health Learning and Inference (AHLI) @ahli-cc.bsky.social · 18/03/2026
June is going to be a busy month for Seattle 👀 CHIL, FIFA, and… 🥁 AHLI’s inaugural Health AI Summer Camp 📍 University of Washington 📅 June 22–28 Fully funded for accepted participants. ⏳Apply by April 15 (limited spots): ahli.cc/summercamp #mlsummerschool #ml4h #healthai
034
Reposted by @sanmikoyejo.bsky.social
Association for Health Learning and Inference (AHLI) @ahli-cc.bsky.social · 26/02/2026
The CHIL 2026 Doctoral Symposium is back! Apply by March 13th 📅 chil.ahli.cc/submit/docto... Last year, we welcomed 28 outstanding PhD researchers for mentorship and lightning talks in health AI. Watch 3 participant talks from 2025 👇 (see 2:48:43) www.youtube.com/watch?v=YaDo...
023
sanmikoyejo.bsky.social @sanmikoyejo.bsky.social · 02/02/2026
1/ Wonderful student projects from CS329H (Fall ’25) ML from Human Preferences at Stanford University! 🚀 Sang Truong, Andy Haupyt, and I introduced students to preference learning + alignment, culminating in final projects. Out of ~50, here are 5 standouts 👇
100
Reposted by @sanmikoyejo.bsky.social
Jessica Schrouff @jessicaschrouff.bsky.social · 09/01/2026
Our Responsible AI group is hiring at GSK! Join a great team investigating how to responsibly develop AI for drug discovery in an environment mixing research and real-world impact. Full-time in London. jobs.gsk.com/en-gb/jobs/4...
jobs.gsk.com
AI/ML Research Scientist, Responsible AI in London, United Kingdom | GSK Careers
GSK Careers is hiring a AI/ML Research Scientist, Responsible AI in London, United Kingdom. Review all of the job details and apply today!
012
Reposted by @sanmikoyejo.bsky.social
Zeyu Tang @zeyutang.bsky.social · 06/01/2026
We want your work on fairness, alignment, and/or agentic systems!! Proud to be co-organizing AFAA @iclr-conf.bsky.social with: @prakharg.bsky.social @adoubleva.bsky.social @Miriam @Jamelle @Golnoosh @jessicaschrouff.bsky.social @sanmikoyejo.bsky.social www.afciworkshop.org #AFAA2026 #ICLR2026
afciworkshop.org
AFAA 2026
The Algorithmic Fairness Across Alignment Procedures and Agentic Systems (AFAA) workshop aims to spark discussions on rethinking fairness in AI alignment procedures and agentic system development.
041
Reposted by @sanmikoyejo.bsky.social
Angelina Wang @ COLM @angelinawang.bsky.social · 12/12/2025
This is work with Daniel E. Ho and @sanmikoyejo.bsky.social. We also have a related preprint with Erin Beeghly on the social impacts of this personalization, and how it interacts with group-based preferences and stereotypes: angelina-wang.github.io/files/person...
angelina-wang.github.io
031
Reposted by @sanmikoyejo.bsky.social
Angelina Wang @ COLM @angelinawang.bsky.social · 12/12/2025
This has huge implications for evaluation: • Benchmark scores ≠ what end users actually experience • Some high-risk behaviors, e.g., manipulative patterns, might only surface in personalized interfaces We argue for more realistic evals that take personalization into account.
121
Reposted by @sanmikoyejo.bsky.social
Angelina Wang @ COLM @angelinawang.bsky.social · 12/12/2025
In fact, even the same MMLU science question can yield different ChatGPT answers for different users, despite using the exact same underlying model. User-level personalization and interaction patterns shape outputs in ways existing evals do not capture.
111
sanmikoyejo.bsky.social @sanmikoyejo.bsky.social · 15/12/2025
The personalization gap challenges an implicit assumption in AI evaluation: that we can measure capability and safety independently of deployment context. Good opportunity to rethink how we evaluate AI systems.
010
Reposted by @sanmikoyejo.bsky.social
Michael Eddy @michaeleddy.bsky.social · 05/08/2025
New in @science.org : 20+ AI scholars—inc. @alondra.bsky.social @randomwalker.bsky.social @sanmikoyejo.bsky.social et al, lay out a playbook for evidence-based AI governance. Without solid data, we risk both hype and harm. Thread 👇
22218
Reposted by @sanmikoyejo.bsky.social
Angelina Wang @ COLM @angelinawang.bsky.social · 30/07/2025
Grateful to win Best Paper at ACL for our work on Fairness through Difference Awareness with my amazing collaborators!! Check out the paper for why we think fairness has both gone too far, and at the same time, not far enough aclanthology.org/2025.acl-lon...
aclanthology.org
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
Angelina Wang, Michelle Phan, Daniel E. Ho, Sanmi Koyejo. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
0284
Reposted by @sanmikoyejo.bsky.social
Angelina Wang @ COLM @angelinawang.bsky.social · 02/06/2025
Instead, we should permit differentiating based on the context. Ex: synagogues in America are legally allowed to discriminate by religion when hiring rabbis. Work with Michelle Phan, Daniel E. Ho, @sanmikoyejo.bsky.social arxiv.org/abs/2502.01926
111
Reposted by @sanmikoyejo.bsky.social
Stanford HAI @stanfordhai.bsky.social · 20/05/2025
Most major LLMs are trained using English data, making it ineffective for the approximately 5 billion people who don't speak English. Here, HAI Faculty Affiliate @sanmikoyejo.bsky.social discusses the risks of this digital divide and how to close it. hai.stanford.edu/news/closing...
hai.stanford.edu
Closing the Digital Divide in AI | Stanford HAI
Large language models aren't effective for many languages. Scholars explain what's at stake for the approximately 5 billion people who don't speak English.
041
Reposted by @sanmikoyejo.bsky.social
Stanford HAI @stanfordhai.bsky.social · 23/04/2025
Our latest white paper maps the landscape of large language model development for low-resource languages, highlighting challenges, trade-offs, and strategies. Read more here: hai.stanford.edu/policy/mind-...
hai.stanford.edu
Mind the (Language) Gap: Mapping the Challenges of LLM Development in Low-Resource Language Contexts | Stanford HAI
In collaboration with The Asia Foundation and the University of Pretoria, this white paper maps the LLM development landscape for low-resource languages, highlighting challenges, trade-offs, and strat...
061
Reposted by @sanmikoyejo.bsky.social
Margaret Mitchell @mmitchell.bsky.social · 16/04/2025
Collaboration with a bunch of lovely people I am thankful to be able to work with: @hannawallach.bsky.social , @angelinawang.bsky.social , Olawale Salaudeen, Rishi Bommasani, and @sanmikoyejo.bsky.social. 🤗
031
Reposted by @sanmikoyejo.bsky.social
Margaret Mitchell @mmitchell.bsky.social · 16/04/2025
🧑‍🔬 Happy our article on Creating a Generative AI Evaluation Science, led by @weidingerlaura.bsky.social & @rajiinio.bsky.social, is now published by the National Academy of Engineering. =) www.nae.edu/338231/Towar... Describes how to mature eval so systems can be worthy of trust and safely deployed.
nae.edu
Toward an Evaluation Science for Generative AI Systems
There is an urgent need for a more robust and comprehensive approach to AI evaluation. There is an increasing imperative to anticipate and understand ...
1407
Reposted by @sanmikoyejo.bsky.social
Technical AI Governance @ ICML 2025 @taig-icml.bsky.social · 01/04/2025
📣We’re thrilled to announce the first workshop on Technical AI Governance (TAIG) at #ICML2025 this July in Vancouver! Join us (& this stellar list of speakers) in bringing together technical & policy experts to shape the future of AI governance! www.taig-icml.com
1144
Reposted by @sanmikoyejo.bsky.social
Stanford HAI @stanfordhai.bsky.social · 26/03/2025
AI systems present an opportunity to reflect society's biases. “However, realizing this potential requires careful attention to both technical and social considerations,” says HAI Faculty Affiliate @sanmikoyejo.bsky.social in his latest op-ed via @theguardian.com: www.theguardian.com/commentisfre...
theguardian.com
Could AI help us build a more racially just society? | Sanmi Koyejo
We have an opportunity to build systems that don’t just replicate our current inequities. Will we take them?
095
Reposted by @sanmikoyejo.bsky.social
weidingerlaura.bsky.social @weidingerlaura.bsky.social · 20/03/2025
Very excited we were able to get this collaboration working -- congrats and big thanks to the co-authors! @rajiinio.bsky.social @hannawallach.bsky.social @mmitchell.bsky.social @angelinawang.bsky.social Olawale Salaudeen, Rishi Bommasani @sanmikoyejo.bsky.social @williamis.bsky.social
051
Reposted by @sanmikoyejo.bsky.social
weidingerlaura.bsky.social @weidingerlaura.bsky.social · 20/03/2025
3) Institutions and norms are necessary for a long-lasting, rigorous and trusted evaluation regime. In the long run, nobody trusts actors correcting their own homework. Establishing an ecosystem that accounts for expertise and balances incentives is a key marker of robust evaluation in other fields.
132
Reposted by @sanmikoyejo.bsky.social
weidingerlaura.bsky.social @weidingerlaura.bsky.social · 20/03/2025
which challenged concepts of what temperature is and in turn motivated the development of new thermometers. A similar virtuous cycle is needed to refine AI evaluation concepts and measurement methods.
121
Reposted by @sanmikoyejo.bsky.social
weidingerlaura.bsky.social @weidingerlaura.bsky.social · 20/03/2025
2) Metrics and evaluation methods need to be refined over time. This iteration is key to any science. Take the example of measuring temperature: it went through many iterations of building new measurement approaches,
121
Reposted by @sanmikoyejo.bsky.social
weidingerlaura.bsky.social @weidingerlaura.bsky.social · 20/03/2025
Just like the “crashworthiness” of a car indicates aspects of safety in case of an accident, AI evaluation metrics need to link to real-world outcomes.
121
Reposted by @sanmikoyejo.bsky.social
weidingerlaura.bsky.social @weidingerlaura.bsky.social · 20/03/2025
We identify three key lessons in particular. 1) Meaningful metrics: evaluation metrics must connect to AI system behaviour or impact that is of relevance in the real-world. They can be abstract or simplified -- but they need to correspond to real-world performance or outcomes in a meaningful way.
121
Reposted by @sanmikoyejo.bsky.social
weidingerlaura.bsky.social @weidingerlaura.bsky.social · 20/03/2025
We pull out key lessons from other fields, such as aerospace, food security, and pharmaceuticals, that have matured from being research disciplines to becoming industries with widely used and trusted products. AI research is going through a similar maturation -- but AI evaluation needs to catch up.
121
sanmikoyejo.bsky.social @sanmikoyejo.bsky.social · 20/03/2025
Excited about this work framing out responsible governance for generative AI!
000
sanmikoyejo.bsky.social @sanmikoyejo.bsky.social · 07/03/2025
@koloskova.bsky.social is awesome, and you should apply to work with her if you can! Also, Congrats!
020
sanmikoyejo.bsky.social @sanmikoyejo.bsky.social · 12/01/2025
📚 Incredible student projects from the 2024 Fall quarter's Machine Learning from Human Preferences course web.stanford.edu/class/cs329h/. Our students tackled some fascinating challenges at the intersection of AI alignment and human values. Selected project details follow... 1/n
web.stanford.edu
CS329H: Machine Learning from Human Preferences
Machine Learning from Human Preferences
141
Reposted by @sanmikoyejo.bsky.social
Anka Reuel ➡️ NeurIPS @ankareuel.bsky.social · 19/12/2024
As one of the vice chairs of the EU GPAI Code of Practice process, I co-wrote the second draft which just went online – feedback is open until mid-January, please let me know your thoughts, especially on the internal governance section! digital-strategy.ec.europa.eu/en/library/s...
digital-strategy.ec.europa.eu
Second Draft of the General-Purpose AI Code of Practice published, written by independent experts
Independent experts present the second draft of the General-Purpose AI Code of Practice, based on the feedback received on the first draft, published on 14 November 2024.
0145
Reposted by @sanmikoyejo.bsky.social
rylanschaeffer.bsky.social @rylanschaeffer.bsky.social · 13/12/2024
What happens when "If at first you don't succeed, try again?" meets modern ML/AI insights about scaling up? You jailbreak every model on the market😱😱😱 Fire work led by @jplhughes.bsky.social Sara Price @aengusl.bsky.social Mrinank Sharma Ethan Perez arxiv.org/abs/2412.03556
032
Reposted by @sanmikoyejo.bsky.social
Hanna Wallach @hannawallach.bsky.social · 14/12/2024
New paper on why machine "unlearning" is much harder than it seems is now up on arXiv: arxiv.org/abs/2412.06966 This was a huuuuuge cross-disciplinary effort led by @msftresearch.bsky.social FATE postdoc @grumpy-frog.bsky.social!!!
arxiv.org
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy, Research, and Practice
We articulate fundamental mismatches between technical methods for machine unlearning in Generative AI, and documented aspirations for broader impact that these methods could have for law and policy. ...
27224
Reposted by @sanmikoyejo.bsky.social
Awa Dieng @adoubleva.bsky.social · 10/12/2024
📆 AFME workshop: Sat, Dec 14 in room 111-112 Join our expert panellists* for a timely discussion on “Rethinking fairness in the era of large language models”!! * @jessicaschrouff.bsky.social, @sethlazar.org, Sanmi Koyejo, Hoda Heidari
0113
Reposted by @sanmikoyejo.bsky.social
Berivan Isik @berivanisik.bsky.social · 27/11/2024
Check out our new paper on privacy preserving style and content transfer 👇 arxiv.org/abs/2411.14639 Led by @poonpura.bsky.social , who is applying for PhD programs this year 🚀 w/ @poonpura.bsky.social , Wei-Ning Chen, Sanmi Koyejo, Albert No
arxiv.org
Differentially Private Adaptation of Diffusion Models via Noisy Aggregated Embeddings
We introduce novel methods for adapting diffusion models under differential privacy (DP) constraints, enabling privacy-preserving style and content transfer without fine-tuning. Traditional approaches...
141
Reposted by @sanmikoyejo.bsky.social
Pura Peetathawatchai @poonpura.bsky.social · 27/11/2024
For details, check out our paper (feedback appreciated!): 📄: arxiv.org/abs/2411.14639 🙌: big thank you to my collaborators and mentors Wei-Ning Chen, @berivanisik.bsky.social, Sanmi Koyejo, Albert No 🧵 16/16
arxiv.org
Differentially Private Adaptation of Diffusion Models via Noisy Aggregated Embeddings
We introduce novel methods for adapting diffusion models under differential privacy (DP) constraints, enabling privacy-preserving style and content transfer without fine-tuning. Traditional approaches...
121