Sign in

Isabel Silva Corpus

@isabelcorpus.bsky.social
179 followers 240 following 21 posts

PhD student in Info Sci at Cornell (Tech) isabelsilvacorpus.github.io

PostsRepliesMedia
Reposted by Isabel Silva Corpus
Meera Desai @madesai.bsky.social · 28/09/2026
Excited to share our new paper, “What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks,” accepted as an oral at COLM! arxiv.org/pdf/2609.08812
Heatmap of average correlations between model rankings on benchmarks grouped into 11 assigned concepts: four capability concepts (reasoning, knowledge, comprehension, summarization) and seven safety concepts (over-refusal, refusal, safety detection, ethics, bias, privacy, unsafe behavior). Diagonal cells show within-concept correlations, ranging from 0.87 (knowledge) and 0.72 (over-refusal) down to 0.20 (bias) and 0.02 (safety detection). Reasoning, knowledge, and comprehension correlate with each other at 0.69 to 0.78, higher than reasoning's and comprehension's own within-concept values (0.66 and 0.68). Ethics correlates more with knowledge (0.70) than with itself (0.55), and bias correlates more with capability concepts (0.41 to 0.45) than with itself (0.20). Privacy and unsafe behavior correlate negatively with reasoning, knowledge, and comprehension (−0.41 to −0.49). Over-refusal and refusal correlate at −0.42.
16018
Reposted by Isabel Silva Corpus
Kate Donahue @kpaxdonahue.bsky.social · 01/09/2026
I’m recruiting students this upcoming cycle at UIUC CS. I’m excited about Qs on societal impact of AI, especially human-AI collaboration, agentic teams, and benchmark design. Also, check out my grad seminar on these topics! tinyurl.com/cs598aisinth...
1184
Reposted by Isabel Silva Corpus
Emma Harvey @emmharv.bsky.social · 12/08/2026
✨New Work✨ forthcoming at #AIES2026: 1️⃣ "Data Annotation as Measurement" by me, @allisonkoe.bsky.social, and @kizilcec.bsky.social explores how data annotation can go wrong, and how measurement theory can help improve it. 🔗: arxiv.org/pdf/2608.07297
A screenshot of our paper title, authors, and abstract.
Title: Data Annotation as Measurement
Authors: Emma Harvey, Allison Koenecke, Rene Kizilcec
Abstract: Modern AI systems depend on annotated data, but annotation is rarely treated as the act of measurement that it is. Instead, annotation quality is commonly reduced to agreement: if multiple annotators assign the same annotation to a data instance, the annotations are taken to be high-quality. Yet agreement does not establish whether annotations validly capture the underlying concept they are meant to represent. In this paper, we argue that data annotation should be understood as a measurement problem. Like other forms of measurement, annotation requires defining a concept, operationalizing it through an instrument, applying that instrument, and evaluating the reliability and validity of the resulting measurements. Drawing on a literature review of annotation quality research (N=132) and semi-structured interviews with annotation team members (N=10), we develop a framework for diagnosing and correcting annotation issues. First, we map key decision points across annotation processes—including task design, annotator management, quality assessment, quality improvement, and adjudication—that shape annotation outcomes. Second, we identify five distinct sources of annotation issues: error, ambiguity, impossibility, subjectivity, and annotator identity. Annotation problems that appear similar at the level of outcomes often require different process-level interventions based on their sources. Finally, we translate measurement theory into practical guidance for annotation teams, showing how assessments of reliability and validity can move beyond agreement alone. By reframing annotation as measurement, we offer a conceptual foundation for improving the quality of annotated data used in AI research and practice.
3266
Reposted by Isabel Silva Corpus
hal @harold.bsky.social · 11/08/2026
just now getting around to watching the openai presentation about the HF hack— my lab and I put out a paper about what we call "agent meltdowns" (which I think is a useful conceptual framework) in May will post some thoughts about the hack + our research as I watch www.youtube.com/watch?v=87Dy...
abstract of our paper, at https://arxiv.org/abs/2605.19149
131
Isabel Silva Corpus @isabelcorpus.bsky.social · 30/07/2026
Tomorrow I’ll be presenting at #ic2s2 during the 10:45am Human-AI Interaction session! Have you ever wondered… what happened to petitions when change.org integrated an AI-assisted writing tool to the platform ?? If so, come find out! pre-print 🔗: arxiv.org/abs/2511.13949
1171
Isabel Silva Corpus @isabelcorpus.bsky.social · 29/07/2026
It was great to have the chance to share about this work led by @mariannealq.bsky.social on AI overviews in local news environments… keep an eye out for more on this soon! #ic2s2
0141
Reposted by Isabel Silva Corpus
Kenny Peng @kennypeng.bsky.social · 06/07/2026
🧵Can we reconcile excitement for SAEs with negative results? Our #ICML2026 position paper argues that even if SAEs underperform baselines when acting on knowns (e.g. probing, steering), they're a powerful tool for ~discovering unknowns~ Poster: Tue 2pm, HALL A #1715
2125
Reposted by Isabel Silva Corpus
Nikhil Garg @nkgarg.bsky.social · 03/07/2026
1/7 Excited to share EconCSLib: a Lean library/workflow. Our vision is to enable researchers who don't know Lean to formalize their Applied Modeling papers. Why? I, and most theorists, are terrified of fatal mistakes in our papers Paper: arxiv.org/abs/2606.13306 Project: gargnikhil.com/EconCSLib/
gargnikhil.com
EconCSLib
EconCSLib is a Lean 4 library and workflow for formalizing Economics and Computation research papers with language-model assistance.
1283
Reposted by Isabel Silva Corpus
Ben Green @benzevgreen.bsky.social · 26/06/2026
Super excited to share a new #FAccT2026 paper with Nel Escher and Nikola Banovic! We show that algorithm auditing policies in the US are wildly insufficient, as they don't account for how public sector algorithms actually function in practice. dl.acm.org/doi/abs/10.1...
dl.acm.org
Algorithm Auditing Policies Rest on Flawed Assumptions About Public Sector Systems | Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency
You will be notified whenever a record that you have chosen has been cited.
1193
Reposted by Isabel Silva Corpus
Lauren Chambers @laurenmarietta.bsky.social · 28/06/2026
it's the last day of #FAccT2026 in Montreal (bonjour hiii), and I'm presenting my paper with Diag Davenport on the promise of #PublicInterestTech clinics for training the next gen of critical sociotechnical thinkers. 🏆 plus we got an honorable mention!? come thru @ 10:45! 📄 tiny.cc/pit-clinics-26
screenshot of a title slide, navy background with light blue icons and white and gold text: "a decision-making pedagogy for the public interest technology clinic: putting the ‘practice’ in critical technical practice." lauren m. chambers & diag davenport, uc berkeley. facct @ montreal, june 28, 2026
1217
Reposted by Isabel Silva Corpus
Jennah Gosciak @jennahgosciak.bsky.social · 25/06/2026
I am excited to be at @facct.bsky.social this year presenting a new 📝 "Scrutinizing Index-Based Risk Assessments: A Case Study in NYC Decision-making for Heat Emergency Management" (work with Luke Boyce, @angelinawang.bsky.social , and @allisonkoe.bsky.social ). 🔗: dl.acm.org/doi/10.1145/... (1/10)
1196
Isabel Silva Corpus @isabelcorpus.bsky.social · 24/06/2026
Excited to attend FAccT 2026 in Montreal this week! Let me know if you'll be there and want to catch up :) I'll be presenting a paper with @allisonkoe.bsky.social about ad delivery skew in the context of government advertising, come by and check it out!
2257
Reposted by Isabel Silva Corpus
Emma Harvey @emmharv.bsky.social · 23/06/2026
I'm so excited to attend #FAccT2026 in 🇨🇦Montreal🇨🇦 to present "Tradeoffs are Domain Dependent: Improving Accuracy and Fairness in Property Tax Assessments" by Evelyn Smith, me, Chris Berry, @jacobsgoldin.bsky.social, and Dan Ho!! 🔗: arxiv.org/pdf/2605.15020
A screenshot of our paper: 

Title: Tradeoffs are Domain Dependent: Improving Accuracy and Fairness in Property Tax Assessments
Authors: EVELYN SMITH and EMMA HARVEY (co-first authors), CHRISTOPHER BERRY, JACOB GOLDIN and DANIEL E. HO (co-senior authors)
Abstract: Algorithmic fairness research often assumes a tradeoff between fairness and accuracy. Yet this tradeoff may not be universal. We test this assumption in the context of U.S. property tax assessment - a setting in which the output of predictive algorithms directly determines the distribution of tax obligations among homeowners. Currently, systematic assessment errors cause owners of lower-valued properties to face disproportionately high tax burdens, creating regressivity in the property tax system. Using data on 26 million property sales spanning 95% of U.S. counties, we conduct three complementary analyses. First, we find that assessment accuracy and fairness - measured using domain-relevant metrics - are strongly correlated across counties under status quo practices. Second, in simulated assessment models, we show that adding property features improves accuracy in most cases, and that when accuracy improves, fairness almost always improves as well. Third, we show that incorporating publicly available Census data into assessment models - a feasible reform in most counties - would significantly improve both accuracy and fairness relative to status quo assessments. Together, these results challenge the presumed universality of the fairness-accuracy tradeoff and demonstrate that well-designed modeling improvements can advance both fairness and accuracy in large-scale public sector systems.
1113
Reposted by Isabel Silva Corpus
Sil Hamilton @srhm.ca · 11/06/2026
@404media.co wrote about our new preprint on tell-tale signs of AI-generated stories! Cc @dmimno.bsky.social Paper: arxiv.org/abs/2605.26492 Article: www.404media.co/elias-thorne...
404media.co
Chatbots Keep Telling Stories About Lighthouse Keeper 'Elias Thorne'. We Might Know Why
LLMs including ChatGPT, Gemini and Claude are obsessed with telling stories about lighthouse keepers and clockmakers, and one character named 'Elias Thorne' has made his way from chatbots to Amazon bo...
11710
Reposted by Isabel Silva Corpus
Sophie Greenwood @sjgreenwood.bsky.social · 26/03/2026
Introducing skytrails by @kennypeng.bsky.social, a way to browse posts across Bluesky! Follow trails to navigate the space of content/conversations here, and discover new interests beyond your usual feeds 🕸️ I'll talk more about it tomorrow in my talk at #atscience #atmosphereconf!
1239
Reposted by Isabel Silva Corpus
Kenny Peng @kennypeng.bsky.social · 17/02/2026
New paper! The Linear Representation Hypothesis is a powerful intuition for how language models work, but lacks formalization. We give a mathematical framework in which we can ask and answer a basic question: how many features can be stored under the hypothesis? 🧵 arxiv.org/abs/2602.11246
14514
Reposted by Isabel Silva Corpus
travis lloyd (træve) @travislloydphd.bsky.social · 09/02/2026
"Community Notes" are reshaping how millions encounter information on social media--but what makes them work (or not)? We term these "Crowdsourced Context Systems" (CCS) and introduce a framework for designing and evaluating them in a new #CHI26 paper 🧵
2296
Reposted by Isabel Silva Corpus
Sophie Greenwood @sjgreenwood.bsky.social · 14/01/2026
Excited to present a new preprint with @nkgarg.bsky.social: presenting usage statistics and observational findings from Paper Skygest in the first six months of deployment! 🎉📜 arxiv.org/abs/2601.04253
Title + abstract of the preprint
417150
Reposted by Isabel Silva Corpus
travis lloyd (træve) @travislloydphd.bsky.social · 06/12/2025
I spoke with @kattenbarge.bsky.social for this @wired.com piece about my research into reddit moderators' experiences moderating AI-generated content. Moderators are working hard to keep Reddit "one of the most human spaces left on the internet," but it's a trying and often thankless task.
1114
Isabel Silva Corpus @isabelcorpus.bsky.social · 01/12/2025
Excited to share a new working paper! What happened when Change.org integrated an AI writing tool into their platform? We provide causal evidence that petition text changed significantly while outcomes did not improve. 1/ arxiv.org/abs/2511.13949
arxiv.org
Introducing AI to an Online Petition Platform Changed Outputs but not Outcomes
The rapid integration of AI writing tools into online platforms raises critical questions about their impact on content production and outcomes. We leverage a unique natural experiment on Change$.$org...
45418
Reposted by Isabel Silva Corpus
hal @harold.bsky.social · 20/11/2025
@davidingram.bsky.social covered @mantzarlis.com and my work on Grokipedia citations for NBC! much more to analyze here www.nbcnews.com/news/amp/rcn...
nbcnews.com
Elon Musk’s Grokipedia cites neo-Nazi website 42 times: study
An analysis by researchers at Cornell University is the first comprehensive look at Grokipedia since Musk launched his project last month.
185
Isabel Silva Corpus @isabelcorpus.bsky.social · 18/11/2025
Had a great time at CODE@MIT this weekend, and wanted to highlight a few (of the many) cool talks!
1175
Reposted by Isabel Silva Corpus
John Holbein @johnholbein1.bsky.social · 01/05/2025
If you run conjoint experiments, you need to read this. Most conjoints estimate average effects for each attribute. But what if the effect of one attribute depends on the others? This paper has got you covered!
45011
Reposted by Isabel Silva Corpus
Emma Harvey @emmharv.bsky.social · 23/06/2025
I am so excited to be in 🇬🇷Athens🇬🇷 to present "A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms" by me, @kizilcec.bsky.social, and @allisonkoe.bsky.social, at #FAccT2025!! 🔗: arxiv.org/pdf/2506.04419
A screenshot of our paper's:

Title: A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
Authors: Emma Harvey, Rene Kizilcec, Allison Koenecke
Abstract: Increasingly, individuals who engage in online activities are expected to interact with large language model (LLM)-based chatbots. Prior work has shown that LLMs can display dialect bias, which occurs when they produce harmful responses when prompted with text written in minoritized dialects. However, whether and how this bias propagates to systems built on top of LLMs, such as chatbots, is still unclear. We conduct a review of existing approaches for auditing LLMs for dialect bias and show that they cannot be straightforwardly adapted to audit LLM-based chatbots due to issues of substantive and ecological validity. To address this, we present a framework for auditing LLM-based chatbots for dialect bias by measuring the extent to which they produce quality-of-service harms, which occur when systems do not work equally well for different people. Our framework has three key characteristics that make it useful in practice. First, by leveraging dynamically generated instead of pre-existing text, our framework enables testing over any dialect, facilitates multi-turn conversations, and represents how users are likely to interact with chatbots in the real world. Second, by measuring quality-of-service harms, our framework aligns audit results with the real-world outcomes of chatbot use. Third, our framework requires only query access to an LLM-based chatbot, meaning that it can be leveraged equally effectively by internal auditors, external auditors, and even individual users in order to promote accountability. To demonstrate the efficacy of our framework, we conduct a case study audit of Amazon Rufus, a widely-used LLM-based chatbot in the customer service domain. Our results reveal that Rufus produces lower-quality responses to prompts written in minoritized English dialects.
13110
Reposted by Isabel Silva Corpus
Jennah Gosciak @jennahgosciak.bsky.social · 24/06/2025
I am presenting a new 📝 “Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments” at @facct.bsky.social on Thursday, with @aparnabee.bsky.social, Derek Ouyang, @allisonkoe.bsky.social, @marzyehghassemi.bsky.social, and Dan Ho. 🔗: arxiv.org/abs/2506.13735 (1/n)
"Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments"

Conducting disparity assessments at regular time intervals is critical for surfacing potential biases in decision-making and improving outcomes across demographic groups. Because disparity assessments fundamentally depend on the availability of demographic information, their efficacy is limited by the availability and consistency of available demographic identifiers. While prior work has considered the impact of missing data on fairness, little attention has been paid to the role of delayed demographic data. Delayed data, while eventually observed, might be missing at the critical point of monitoring and action -- and delays may be unequally distributed across groups in ways that distort disparity assessments. We characterize such impacts in healthcare, using electronic health records of over 5M patients across primary care practices in all 50 states. Our contributions are threefold. First, we document the high rate of race and ethnicity reporting delays in a healthcare setting and demonstrate widespread variation in rates at which demographics are reported across different groups. Second, through a set of retrospective analyses using real data, we find that such delays impact disparity assessments and hence conclusions made across a range of consequential healthcare outcomes, particularly at more granular levels of state-level and practice-level assessments. Third, we find limited ability of conventional methods that impute missing race in mitigating the effects of reporting delays on the accuracy of timely disparity assessments. Our insights and methods generalize to many domains of algorithmic fairness where delays in the availability of sensitive information may confound audits, thus deserving closer attention within a pipeline-aware machine learning framework.Figure contrasting a conventional approach to conducting disparity assessments, which is static, to the analysis we conduct in this paper. Our analysis (1) uses comprehensive health data from over 1,000 primary care practices and 5 million patients across the U.S., (2) timestamped information on the reporting of race to measure delay, and (3) retrospective analyses of disparity assessments under varying levels of delay.
1134
Reposted by Isabel Silva Corpus
Emma Harvey @emmharv.bsky.social · 09/06/2025
📣 "Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems" is forthcoming at #ACL2025NLP - and you can read it now on arXiv! 🔗: arxiv.org/pdf/2506.04482 🧵: ⬇️
A screenshot of our paper: 

Title: Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems

Authors: Emma Harvey, Emily Sheng, Su Lin Blodgett, Alexandra Chouldechova, Jean Garcia-Gathright, Alexandra Olteanu, Hanna Wallach

Abstract: The NLP research community has made publicly available numerous instruments for measuring representational harms caused by large language model (LLM)-based systems. These instruments have taken the form of datasets, metrics, tools, and more. In this paper, we examine the extent to which such instruments meet the needs of practitioners tasked with evaluating LLM-based systems. Via semi-structured interviews with 12 such practitioners, we find that practitioners are often unable to use publicly available instruments for measuring representational harms. We identify two types of challenges. In some cases, instruments are not useful because they do not meaningfully measure what practitioners seek to measure or are otherwise misaligned with practitioner needs. In other cases, instruments---even useful instruments---are not used by practitioners due to practical and institutional barriers impeding their uptake. Drawing on measurement theory and pragmatic measurement, we provide recommendations for addressing these challenges to better meet practitioner needs.
1174
Reposted by Isabel Silva Corpus
Emma Harvey @emmharv.bsky.social · 27/03/2025
🎉 So excited to share that "Don't Forget the Teachers" has received a Best Paper Award at #CHI2025!! @allisonkoe.bsky.social @kizilcec.bsky.social
A screenshot of our paper on the CHI program website, with a "Best Paper" badge visible. The screenshot includes the title: "Don't Forget the Teachers": Towards an Educator-Centered Understanding of Harms from Large Language Models in Education. The screenshot also includes the author list: Emma Harvey, Allison Koenecke, Rene F. Kizilcec.
7353
Reposted by Isabel Silva Corpus
travis lloyd (træve) @travislloydphd.bsky.social · 26/03/2025
*NEW DATASET AND PAPER* (CHI2025): How are online communities responding to AI-generated content (AIGC)? We study this by collecting and analyzing the public rules of 300,000+ subreddits in 2023 and 2024. 1/
1165
Reposted by Isabel Silva Corpus
Sophie Greenwood @sjgreenwood.bsky.social · 10/03/2025
Please repost to get the word out! @nkgarg.bsky.social and I are excited to present a personalized feed for academics! It shows posts about papers from accounts you’re following bsky.app/profile/pape...
8171118
Reposted by Isabel Silva Corpus
Emma Harvey @emmharv.bsky.social · 13/03/2025
✨New Work✨ by me, @allisonkoe.bsky.social, and @kizilcec.bsky.social forthcoming at #CHI2025: "Don't Forget the Teachers": Towards an Educator-Centered Understanding of Harms from Large Language Models in Education 🔗: arxiv.org/pdf/2502.14592
A screenshot of our paper:

Title: “Don’t Forget the Teachers”: Towards an Educator-Centered Understanding of Harms from Large Language Models in Education

Authors: Emma Harvey, Allison Koenecke, Rene Kizilcec

Abstract: Education technologies (edtech) are increasingly incorporating new features built on LLMs, with the goals of enriching the processes of teaching and learning and ultimately improving learning outcomes. However, it is still too early to understand the potential downstream impacts of LLM-based edtech. Prior attempts to map the risks of LLMs have not been tailored to education specifically, even though it is a unique domain in many respects: from its population (students are often children, who can be especially impacted by technology) to its goals (providing the ‘correct’ answer may be less important than understanding how to arrive at an answer) to its implications for higher-order skills that generalize across contexts (e.g. critical thinking and collaboration). We conducted semi-structured interviews with six edtech providers representing leaders in the K-12 space, as well as a diverse group of 23 educators with varying levels of experience with LLM-based edtech. Through a thematic analysis, we explored how each group is anticipating, observing, and accounting for potential harms from LLMs in education. We find that, while edtech providers focus primarily on mitigating technical harms, i.e. those that can be measured based solely on LLM outputs themselves, educators are more concerned about harms that result from the broader impacts of LLMs, i.e. those that require observation of interactions between students, educators, school systems, and edtech to measure. Overall, we (1) develop an education-specific overview of potential harms from LLMs, (2) highlight gaps between conceptions of harm by edtech providers and those by educators, and (3) make recommendations to facilitate the centering of educators in the design and development of edtech tools.
1548