Sign in

andrea wen-yi wang

@andreawwenyi.bsky.social
1.9K followers 96 following 19 posts

phd candidate @ cornell infosci andreawwenyi.github.io

PostsRepliesMedia
Reposted by andrea wen-yi wang
Melanie Walsh @mellymeldubs.bsky.social · 02/10/2026
In this powerful piece on libraries, this comment stuck out. I went to ALA for the first time this year, and I was struck by the prominent presence of corporate vendors, many of which are sucking libraries dry… www.nytimes.com/interactive/...
Photo of librarian Crystal Gates with quote: I would increase access to electronic materials. Our budget can’t keep pace with demand. The book that I can personally get on Kindle or Audible for $3.99 costs my library $80 to $120, and then we have to return it in three years. I’m limited in the number of checkouts or the number of people. I don’t mind one-copy, one-user restrictions. But telling me I’m going to spend $80 on a book that my patrons no longer have access to in 12 months? It’s not getting damaged. It’s not getting wet. It’s not getting torn up.
Crystal Gates, North Little Rock, Ark.
North Little Rock Public Libraries
06022
Reposted by andrea wen-yi wang
David Mimno @dmimno.bsky.social · 15/07/2026
Text as Data is happening at Berkeley right before COLM. Please share!
03014
andrea wen-yi wang @andreawwenyi.bsky.social · 14/07/2026
Love this!!
001
andrea wen-yi wang @andreawwenyi.bsky.social · 03/07/2026
ABS works in baseball coz MLB recognized that if u only care about optimizing call accuracy, the tech would destroy the sport. They spent 7 years finding a solution that balance accuracy, entertainment, speed, and the art of baseball. (we wrote a FAccT paper on this: arxiv.org/abs/2605.16237)
arxiv.org
Inside Baseball: The Automated Ball-Strike System as an Object Lesson in Technological Rule Enforcement
Clearly-defined rules are often assumed to be straightforward to automate and evaluate. We challenge this assumption through an in-depth study of Major League Baseball's (MLB) seven-year experimentati...
0209
Reposted by andrea wen-yi wang
404 Media @404media.co · 11/06/2026
When you ask ChatGPT or any popular LLM to tell you a story, one name keeps coming up: "Elias Thorne." Depending which chatbot you ask, he's a lighthousekeeper, clockmaker or explorer. His stories are also flooding Amazon's AI-generated book market, YouTube slop, and fake news. Who is he?
404media.co
Chatbots Keep Telling Stories About Lighthouse Keeper 'Elias Thorne'. We Might Know Why
LLMs including ChatGPT, Gemini and Claude are obsessed with telling stories about lighthouse keepers and clockmakers, and one character named 'Elias Thorne' has made his way from chatbots to Amazon bo...
25389140
Reposted by andrea wen-yi wang
Ted Underwood @tedunderwood.com · 26/02/2026
We're in a strange situation rn where Google can train freely on books from university libraries—but researchers *at* universities have limited access. I'm optimistic this can be fixed, but if you're in admin or working at a foundation, please know: univs are failing here & resources are needed.
Building benchmarks is only one way scholars can help steer AI development. We can also measure the effects of AI on students, build better datasets, or tune new open models. Openness itself could be our most important contribution. Universities have huge libraries, and the legal doctrine of fair use should protect models trained on those collections for a nonprofit educational purpose. At the moment, we are not pressing this advantage. Higher education has been so cautious about fair use that the private sector can now train more freely on our libraries (via Google Books) than is possible for academic AI researchers. We need to be bolder: It is our duty to ensure library collections remain open to the public in a form that empowers 21st-century readers. If our intellectual heritage gets enclosed in proprietary tools, we will find ourselves making the same bad bargain we made with scientific publishers, who sell our own research back to us at a steep markup.
921150
Reposted by andrea wen-yi wang
travis lloyd (træve) @travislloydphd.bsky.social · 09/02/2026
"Community Notes" are reshaping how millions encounter information on social media--but what makes them work (or not)? We term these "Crowdsourced Context Systems" (CCS) and introduce a framework for designing and evaluating them in a new #CHI26 paper 🧵
2296
Reposted by andrea wen-yi wang
Karen Levy @karenlevy.bsky.social · 03/02/2026
Interesting new statistical analysis, again confirming that mandatory electronic monitoring in trucking—ostensibly a safety measure—led to *increased* accident and fatality rates after implementation. www.sciencedirect.com/science/arti...
sciencedirect.com
The unintended consequences of monitoring technologies: Evidence from the Electronic Logging Device mandate
The electronic logging device (ELD) final rule was passed in 2015 and implemented in phases over the course of four years. The mandate required that m…
0103
Reposted by andrea wen-yi wang
Alvaro M. Bedoya @bedoyausa.bsky.social · 15/10/2025
I used to focus on left versus right. Now I’m much more worried about the money at the top. But while it might seem strange to say it, I think this is a hopeful way of looking at the world that opens the door to coalitions that seemed impossible before. My first for the @newrepublic.com:
914660
Reposted by andrea wen-yi wang
Jennah Gosciak @jennahgosciak.bsky.social · 24/06/2025
I am presenting a new 📝 “Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments” at @facct.bsky.social on Thursday, with @aparnabee.bsky.social, Derek Ouyang, @allisonkoe.bsky.social, @marzyehghassemi.bsky.social, and Dan Ho. 🔗: arxiv.org/abs/2506.13735 (1/n)
"Bias Delayed is Bias Denied? Assessing the Effect of Reporting Delays on Disparity Assessments"

Conducting disparity assessments at regular time intervals is critical for surfacing potential biases in decision-making and improving outcomes across demographic groups. Because disparity assessments fundamentally depend on the availability of demographic information, their efficacy is limited by the availability and consistency of available demographic identifiers. While prior work has considered the impact of missing data on fairness, little attention has been paid to the role of delayed demographic data. Delayed data, while eventually observed, might be missing at the critical point of monitoring and action -- and delays may be unequally distributed across groups in ways that distort disparity assessments. We characterize such impacts in healthcare, using electronic health records of over 5M patients across primary care practices in all 50 states. Our contributions are threefold. First, we document the high rate of race and ethnicity reporting delays in a healthcare setting and demonstrate widespread variation in rates at which demographics are reported across different groups. Second, through a set of retrospective analyses using real data, we find that such delays impact disparity assessments and hence conclusions made across a range of consequential healthcare outcomes, particularly at more granular levels of state-level and practice-level assessments. Third, we find limited ability of conventional methods that impute missing race in mitigating the effects of reporting delays on the accuracy of timely disparity assessments. Our insights and methods generalize to many domains of algorithmic fairness where delays in the availability of sensitive information may confound audits, thus deserving closer attention within a pipeline-aware machine learning framework.Figure contrasting a conventional approach to conducting disparity assessments, which is static, to the analysis we conduct in this paper. Our analysis (1) uses comprehensive health data from over 1,000 primary care practices and 5 million patients across the U.S., (2) timestamped information on the reporting of race to measure delay, and (3) retrospective analyses of disparity assessments under varying levels of delay.
1134
Reposted by andrea wen-yi wang
Emma Harvey @emmharv.bsky.social · 23/06/2025
I am so excited to be in 🇬🇷Athens🇬🇷 to present "A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms" by me, @kizilcec.bsky.social, and @allisonkoe.bsky.social, at #FAccT2025!! 🔗: arxiv.org/pdf/2506.04419
A screenshot of our paper's:

Title: A Framework for Auditing Chatbots for Dialect-Based Quality-of-Service Harms
Authors: Emma Harvey, Rene Kizilcec, Allison Koenecke
Abstract: Increasingly, individuals who engage in online activities are expected to interact with large language model (LLM)-based chatbots. Prior work has shown that LLMs can display dialect bias, which occurs when they produce harmful responses when prompted with text written in minoritized dialects. However, whether and how this bias propagates to systems built on top of LLMs, such as chatbots, is still unclear. We conduct a review of existing approaches for auditing LLMs for dialect bias and show that they cannot be straightforwardly adapted to audit LLM-based chatbots due to issues of substantive and ecological validity. To address this, we present a framework for auditing LLM-based chatbots for dialect bias by measuring the extent to which they produce quality-of-service harms, which occur when systems do not work equally well for different people. Our framework has three key characteristics that make it useful in practice. First, by leveraging dynamically generated instead of pre-existing text, our framework enables testing over any dialect, facilitates multi-turn conversations, and represents how users are likely to interact with chatbots in the real world. Second, by measuring quality-of-service harms, our framework aligns audit results with the real-world outcomes of chatbot use. Third, our framework requires only query access to an LLM-based chatbot, meaning that it can be leveraged equally effectively by internal auditors, external auditors, and even individual users in order to promote accountability. To demonstrate the efficacy of our framework, we conduct a case study audit of Amazon Rufus, a widely-used LLM-based chatbot in the customer service domain. Our results reveal that Rufus produces lower-quality responses to prompts written in minoritized English dialects.
13110
Reposted by andrea wen-yi wang
John Garrison Marks @johngmarks.com · 10/06/2025
Worth noting today that the entire budget of the NEH is about $200M.
6415225
Reposted by andrea wen-yi wang
David Mimno @dmimno.bsky.social · 10/06/2025
New NEH-supported tutorial on running LLMs locally with ollama! Your laptop is more powerful than you think. Save money, privacy, and energy. aiforhumanists.com/tutorials/
aiforhumanists.com
Code Tutorials
The AI for Humanists project is developing resources to enable DH scholars to explore how large language models and AI technologies can be used in their research and teaching. Find an annotated biblio...
36224
Reposted by andrea wen-yi wang
Lucy Li @lucy3.bsky.social · 05/05/2025
I'm joining Wisconsin CS as an assistant professor in fall 2026!! There, I'll continue working on language models, computational social science, & responsible AI. 🌲🧀🚣🏻‍♀️ Apply to be my PhD student! Before then, I'll postdoc for a year in the NLP group at another UW 🏔️ in the Pacific Northwest
Wisconsin-Madison's tree-filled campus, next to a big shiny lake A computer render of the interior of the new computer science, information science, and statistics building. A staircase crosses an open atrium with visibility across multiple floors
1614514
Reposted by andrea wen-yi wang
Alex Gil @elotroalex.bsky.social · 27/04/2025
For the HTR and OCR crew: New paper by Jonathan Bourne. He's been working to help DLOC handle OCR for a whole bunch of Caribbean historical newspapers. "Scrambled text: fine-tuning language models for OCR error correction using synthetic data" link.springer.com/article/10.1...
link.springer.com
Scrambled text: fine-tuning language models for OCR error correction using synthetic data - International Journal on Document Analysis and Recognition (IJDAR)
OCR errors are common in digitised historical archives significantly affecting their usability and value. Generative Language Models (LMs) have shown potential for correcting these errors using the co...
44414
Reposted by andrea wen-yi wang
Maria Antoniak @mariaa.bsky.social · 25/04/2025
Slightly paraphrasing @oms279.bsky.social during his talk at #COMPTEXT2025: "The single most important use case for LLMs in sociology is turning unstructured data into structured data." Discussing his recent work on codebooks, prompts, and information extraction: osf.io/preprints/so...
2295
Reposted by andrea wen-yi wang
Simona Liao @simonaliao.bsky.social · 02/12/2024
Hi everyone, I am excited to share our large-scale survey study with 800+ researchers, which reveals researchers’ usage and perceptions of LLMs as research tools, and how the usage and perceptions differ based on demographics. See results in comments! 🔗 Arxiv link: arxiv.org/abs/2411.05025
arxiv.org
LLMs as Research Tools: A Large Scale Survey of Researchers' Usage and Perceptions
The rise of large language models (LLMs) has led many researchers to consider their usage for scientific work. Some have found benefits using LLMs to augment or automate aspects of their research pipe...
910332
andrea wen-yi wang @andreawwenyi.bsky.social · 09/04/2025
[New preprint!] Do Chinese AI Models Speak Chinese Languages? Not really. Chinese LLMs like DeepSeek are better at French than Cantonese. Joint work with Unso Jo and @dmimno.bsky.social . Link to paper: arxiv.org/pdf/2504.00289 🧵
1256
Reposted by andrea wen-yi wang
Sung Kim @sungkim.bsky.social · 31/03/2025
You’ve probably heard about how AI/LLMs can solve Math Olympiad problems ( deepmind.google/discover/blo... ). So naturally, some people put it to the test — hours after the 2025 US Math Olympiad problems were released. The result: They all sucked!
917350
Reposted by andrea wen-yi wang
travis lloyd (træve) @travislloydphd.bsky.social · 26/03/2025
*NEW DATASET AND PAPER* (CHI2025): How are online communities responding to AI-generated content (AIGC)? We study this by collecting and analyzing the public rules of 300,000+ subreddits in 2023 and 2024. 1/
1165
Reposted by andrea wen-yi wang
Dr. Casey Fiesler @cfiesler.bsky.social · 19/03/2025
hey it's that time of year again, when people start to wonder whether AIES is actually happening and when this year’s paper deadline might be if so! anyone know anything about the ACM/AAAI conference on AI Ethics & Society for 2025? (I used to ask about this every year on Twitter haha.)
1216
Reposted by andrea wen-yi wang
David Mimno @dmimno.bsky.social · 11/11/2024
Best Student Paper at #AIES 2024 went to @andreawwenyi.bsky.social! Annotating gender-biased narratives in the courtroom is a complex, nuanced task with frequent subjective decision-making by legal experts. We asked: What do experts desire from a language model in this annotation process?
1194
andrea wen-yi wang @andreawwenyi.bsky.social · 21/02/2024
How do LLMs represent relationships between languages? By studying the embedding layers of XLM-R and mT5, we find they are highly interpretable. LLMs can find semantic alignment as an emergent property! Joint work with @dmimno.bsky.social. 🧵
2264