Sign in

Meera Desai

@madesai.bsky.social
1.3K followers 249 following 14 posts

PhD student at University of Michigan School of Information. meera-desai.com

PostsRepliesMedia
Meera Desai @madesai.bsky.social · 28/09/2026
Excited to share our new paper, “What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks,” accepted as an oral at COLM! arxiv.org/pdf/2609.08812
Heatmap of average correlations between model rankings on benchmarks grouped into 11 assigned concepts: four capability concepts (reasoning, knowledge, comprehension, summarization) and seven safety concepts (over-refusal, refusal, safety detection, ethics, bias, privacy, unsafe behavior). Diagonal cells show within-concept correlations, ranging from 0.87 (knowledge) and 0.72 (over-refusal) down to 0.20 (bias) and 0.02 (safety detection). Reasoning, knowledge, and comprehension correlate with each other at 0.69 to 0.78, higher than reasoning's and comprehension's own within-concept values (0.66 and 0.68). Ethics correlates more with knowledge (0.70) than with itself (0.55), and bias correlates more with capability concepts (0.41 to 0.45) than with itself (0.20). Privacy and unsafe behavior correlate negatively with reasoning, knowledge, and comprehension (−0.41 to −0.49). Over-refusal and refusal correlate at −0.42.
16419
Meera Desai @madesai.bsky.social · 15/07/2026
rejoice! text as data is back!
050
Reposted by Meera Desai
Dallas Card @dallascard.bsky.social · 10/07/2026
As some may have heard me talk about at #ACL2026, I'm excited to share a new preprint on approaches to validation when using LLMs to measure concepts in social science, led by @madesai.bsky.social and @azjacobs.bsky.social !! Paper: arxiv.org/abs/2607.07915
Title page from "Validating LLMs in social science: Epistemic threats and emerging norms" by Meera Desai, Dallas Card, and Abigail Z. Jacobs
56819
Reposted by Meera Desai
David Mimno @handle.invalid · 11/12/2025
Lavinia Dunagan and @dallascard.bsky.social find implicit references to bible verses using a combination of neural embeddings and text similarity—neither is enough on its own #CHR2025
Slide showing significant differences in implicit references to bible verses by US political parties
1285
Reposted by Meera Desai
Ted Underwood @tedunderwood.com · 03/05/2025
Have we talked about this paper yet on bsky? It's really good, and ought to be cited widely by people working at the humanities-AI interface. TLDR: If AI is going to be applied to fuzzy human problems, social science measurement theory becomes essential. #MLSky
128211
Reposted by Meera Desai
Victor Ray @victorerikray.bsky.social · 12/02/2025
So-called “institutional neutrality” was pushed by folks in coalition with the ghouls currently gutting science and higher ed. Their goal was always for universities to neutrally (and silently) watch as they were systematically dismantled by people who hate verifiable knowledge.
1022562
Reposted by Meera Desai
Hanna Wallach @hannawallach.bsky.social · 04/02/2025
Remember this @neuripsconf.bsky.social workshop paper? We spent the past month writing a newer, better, longer version!!! You can find it online here: arxiv.org/abs/2502.00561
arxiv.org
Position: Evaluating Generative AI Systems is a Social Science Measurement Challenge
The measurement tasks involved in evaluating generative AI (GenAI) systems are especially difficult, leading to what has been described as "a tangle of sloppy tests [and] apples-to-oranges comparisons...
28614
Meera Desai @madesai.bsky.social · 20/01/2025
Submissions are open for our workshop on data and benchmarking practices in ML! Accepting both full-length submissions (up to 10 pages) and tiny papers submissions (3-5 pages)
0122
Reposted by Meera Desai
rlongjohn.bsky.social @rlongjohn.bsky.social · 16/01/2025
📣Announcing MLDPR 2025, an @iclr-conf.bsky.social workshop on data and benchmarking practices in ML! 🔗 mldpr2025.com 📭accepting submissions related to: ▫️data curation ▫️FAIR datasets ▫️ML data repositories ▫️reproducibility ▫️holistic benchmarking
Flyer for the ICLR 2025 workshop entitled The Future of Machine Learning Data Practices and Repositories. The date of the workshop if to be determined but will be one of the two workshop dates (April 27 or 28) for ICLR 2025 in Singapore. The submission deadline is February 3, 2025 anywhere on earth, and the review deadline is March 5, 2025 anywhere on earth. Call for papers and more information is available at the website https://mldpr2025.com/. Reach out to mldpr2025@gmail.com for questions! The workshop is organized by Markelle Kelly, Rachel Longjohn, Meera Desai, Shivani Kapania, Maria Antoniak, Padhraic Smyth, Joaquin Vanschoren, Sameer Singh, Daniel S. Katz, and Amy Winecoff.
0157
Reposted by Meera Desai
Hanna Wallach @hannawallach.bsky.social · 02/12/2024
Evaluating Generative AI Systems is a Social Science Measurement Challenge: arxiv.org/abs/2411.10939 TL;DR: The ML community would benefit from learning from and drawing on the social sciences when evaluating GenAI systems.
23911
Meera Desai @madesai.bsky.social · 22/10/2024
Applying to the PhD program at U Mich School of Information? Submit your application materials to our feedback program—led by current Ph.D. students who want to support applicants from underrepresented backgrounds in higher ed 📅 Materials accepted until Nov 11 🔗 Link: bit.ly/UMSI-Review-...
010