Sign in

Willie Agnew

@willie-agnew.bsky.social
771 followers 850 following 192 posts

Queer in AI 🏳️‍🌈 | postdoc at cmu HCII | ostem |william-agnew.com | views my own | he/they

PostsRepliesMedia
Willie Agnew @willie-agnew.bsky.social · 28/09/2026
year four on the job market! give it up for year four!!!!
mr. krabs day 15 meme with a screenshot of my google drive showing four years of job app folders
010
Willie Agnew @willie-agnew.bsky.social · 23/09/2026
Is anyone else getting lots of emails from people outside academia offering to review for workshops? What is the incentive for this? Do people really value workshop reviewing service on a resume??
010
Willie Agnew @willie-agnew.bsky.social · 14/09/2026
I feel like the AI ethics community has been sidelined to industry workers and leaders in this latest round of fears about over AI, which makes it harder to imagine critical and independent oversight ever been empowered. How does the AI ethics community build power in this moment?
110
Reposted by Willie Agnew
Techmeme @techmeme.com · 13/09/2026
Source: Anthropic has selected the Nasdaq for its potential IPO (Katie Roof/Business Insider) Main Link | Techmeme Permalink
031
Willie Agnew @willie-agnew.bsky.social · 09/09/2026
introduced my cat to slayyyter and he is bouncing off the walls
030
Reposted by Willie Agnew
Jared Moore @jaredlcm.bsky.social · 06/08/2026
This work is in collaboration w/ Andrea Mock, Yifan Mai, @jacyanthis.bsky.social , Ryan Louie, @willie-agnew.bsky.social , Ashish Mehta, @klyman.bsky.social, Percy Liang, Nick Haber, Eric Lin, and @desmond-ong.bsky.social Thanks to the Human Line Project for helping connect us with participants!
031
Reposted by Willie Agnew
Jared Moore @jaredlcm.bsky.social · 06/08/2026
Paper: arxiv.org/abs/2608.05004 Dataset: huggingface.co/datasets/spi... Code: github.com/jlcmoore/llm...
arxiv.org
DelusionEval: Measuring Delusion-Linked Behaviors in AI Chatbots
Mental health professionals have raised concerns about risks of psychological harm from interaction with large language models (LLMs), including "delusional spirals" in which concerning human and LLM ...
131
Reposted by Willie Agnew
Jared Moore @jaredlcm.bsky.social · 06/08/2026
Our conclusion: while some delusion-linked behaviors have decreased with larger and more recent LLMs, rates remain high, especially when considered across the millions of people globally who interact with LLMs. We encourage further empirical work to mitigate harm.
121
Reposted by Willie Agnew
Jared Moore @jaredlcm.bsky.social · 06/08/2026
DelusionEval is built from real transcripts, not synthetic roleplay. We prompt models with 589 unique histories from 18 users who reported psychological harm from LLMs, then score 16 chatbot behaviors as requested context depth increases.
An example of our evaluation. We take an existing conversational window derived from a user's transcript: `U_1, A_1, U_2, ..., U_n, A_n` with an original LLM, `A` (here, `gpt-4o`). We evaluate an evaluated LLM, `A'` (here, `gpt-5.4`), by successively prompting it with chains of the original context (samples: `{U_1}`, `{U_1, A_1, U_2}`, ..., `{U_1, A_1, U_2, ..., U_n}`).
221
Reposted by Willie Agnew
Jared Moore @jaredlcm.bsky.social · 06/08/2026
Requested context depth changes the prevalence of several behaviors. For gpt-5.4, deeper requested context is associated with higher prevalence of delusional behavior and *less* discouraging of violence.
Context-depth effects in `gpt-5.4`. Category-level context effect for `delusional`.  Each point shows prevalence versus context length, with 95% bootstrap confidence intervals.
121
Reposted by Willie Agnew
Jared Moore @jaredlcm.bsky.social · 06/08/2026
The tendency of an evaluated LLM to exhibit delusion-linked behavior does not reliably correlate with model size, release date, or the presence of test-time reasoning. Within model families, scaling effects are uneven and sometimes reverse sign.
Model-family comparison across GPT, Claude, Gemini, and Qwen. Bars show prevalence by the five categories, with 95% bootstrap confidence intervals.
122
Reposted by Willie Agnew
Jared Moore @jaredlcm.bsky.social · 06/08/2026
Which LLMs tend to facilitate delusion-linked behaviors in realistic multi-turn conversations? We tested 14 models with DelusionEval and found that every evaluated LLM exhibited some of these behaviors, with large differences across categories and model families. 🧵
1136
Willie Agnew @willie-agnew.bsky.social · 03/07/2026
Automated AI benchmarks are often well-specified and easy to implement and run, and I feel this contributes to their spread in policy and industry despite have many limitations, especially when try to assess human interactions. Can we design and specify human in the loop evaluations like this?
120
Willie Agnew @willie-agnew.bsky.social · 02/07/2026
Coffee shop accidentally gave me earl gray tea this morning (usually I get green) and I am currently replying to every email within 10 minutes ⌨️ 🔥 act now
140
Reposted by Willie Agnew
Pranav A @pranav-nlp.bsky.social · 25/06/2026
Work is here: arxiv.org/abs/2606.11021, done with @dippedrusk.com, @martinmundt.bsky.social, @arjunsubgraph.bsky.social, @jordant.bsky.social, @willie-agnew.bsky.social, @dchechel.bsky.social, Franziska Sofia Hafner, @a-lauscher.bsky.social.
arxiv.org
Making a Name for Myself: On Academic Naming Policies and their Impact
In academic publishing, names connect scholars to their work. When scholars change their names, including for marriage, academic recognition, or gender transition, they may lose credit for past public...
042
Reposted by Willie Agnew
Queer in AI @queerinai.com · 24/06/2026
P.S. You can find details on how to attend our social at queerinai.com/facct-2026!
021
Reposted by Willie Agnew
Queer in AI @queerinai.com · 24/06/2026
Congratulations to all the contributors (ft @pranav-nlp.bsky.social and @jordant.bsky.social) for their hard work! 💐 We hope to see folks at our social tomorrow evening, and we hope everyone has a wonderful FAccT 2026.
111
Reposted by Willie Agnew
Queer in AI @queerinai.com · 24/06/2026
Queer in AI is proud to showcase several papers that our members are presenting at @facct.bsky.social 2026 tomorrow through Sunday. One paper examines how years-long activism around name change policies has had an immensely positive impact for authors, benefiting both queer and cis scholars! 🤝🌈
Image with Queer in AI logo, FAccT logo, and the following text: "FAccT 2026 Papers. Challenges to Grassroots Organization Engagement with AI Policy. Making a Name for Myself: On Academic Naming Policies and their Impact. Sounds Queer: Representation of LGBTQIA Identities in AI-generated Songs. The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor. Hidden Beyond the Gender Binary: Pitch-Based Insights into Bias Mitigation for Keyword Spotting.”Image with Queer in AI logo, FAccT logo, and the following text: "Challenges to Grassroots Organization Engagement with AI Policy. Sat, 27 Jun, 11:45 AM, Ballroom West (4). Making a Name for Myself: On Academic Naming Policies and their Impact. Fri, 26 Jun, 3:42 PM, Jarry (A).”Image with Queer in AI logo, FAccT logo, and the following text: "The Algorithmic Gaze of Image Quality Assessment: An Audit and Trace Ethnography of the LAION-Aesthetics Predictor. Thu, 25 Jun, 11:33 AM, Jarry (A). Sounds Queer: Representation of LGBTQIA Identities in AI-generated Songs. Fri, 26 Jun, 4:06 PM, Jarry (A). Hidden Beyond the Gender Binary: Pitch-Based Insights into Bias Mitigation for Keyword Spotting. Thu, 25 Jun, 11:45 AM, Joyce (A).”
1118
Reposted by Willie Agnew
pettter, stuff scientist @pettter.bsky.social · 23/06/2026
@abeba.blacksky.app @willie-agnew.bsky.social among others made a good and very thorough article digging into the concrete (and increasing) connections and pathways for computer vision research to predominantly power surveillance tech: www.nature.com/articles/s41...
nature.com
Computer-vision research powers surveillance technology - Nature
An analysis of research papers and citing patents indicates the extensive ties between computer-vision research and surveillance.
12817
Willie Agnew @willie-agnew.bsky.social · 22/06/2026
What does "human flourishing" mean? It sounds nice, but it often seems to appear without more specific goals like justice, various rights, equity, or empowerment.
120
Willie Agnew @willie-agnew.bsky.social · 04/06/2026
I feel like conferences are biased towards accepting papers that are like speedbumps--very robust of any sort of attack, but not very interesting, unique, or helpful for moving forward.
230
Reposted by Willie Agnew
Joseph Cox @josephcox.bsky.social · 01/06/2026
This is absolutely nuts: hackers are hijacking high-profile Instagram accounts by simply asking Meta's AI chatbot to change the email on the account. Meta's AI does it, hacker gets password reset code, they're in. A staggering security issue www.404media.co/hackers-simp...
404media.co
Hackers Simply Asked Meta AI to Give Them Access to High-Profile Instagram Accounts. It Worked
The exploit shows the extreme risk of offloading technical support to AI.
27197655106
Willie Agnew @willie-agnew.bsky.social · 27/05/2026
Even when I was in high school all the kids who got into the top state school, much less Harvard, all had basically straight A's while taking many college-level classes. Is it really surprising or bad that they get mostly A's in college when colleges select for students who are good at getting A's?
030
Reposted by Willie Agnew
Karen Hao @karenhao.bsky.social · 20/05/2026
On the one-year anniversary of EMPIRE OF AI, I am so, so excited to announce The AI Resist List, a new project that documents examples of resistance to the AI empires around the world 😍 airesistlist.org
The gorgeous AI Resist List homepage
4525271122
Willie Agnew @willie-agnew.bsky.social · 19/05/2026
Spending most your time writing and managing grants is no fun (and even less fun when you're still expected to do first author work like a grad student)
030
Reposted by Willie Agnew
Dr Abeba Birhane @abeba.blacksky.app · 18/05/2026
Read more here arxiv.org/abs/2605.068... With gratitude to my amazing coauthors @storitu.org, @willie-agnew.bsky.social, @harshp.com, @bmitra.bsky.social, Roel Dobbe, and @zeerak.bsky.social
arxiv.org
Big AI's Regulatory Capture: Mapping Industry Interference and Government Complicity
Over the past decade, the AI industry has come to exert an unprecedented economic, political and societal power and influence. It is therefore critical that we comprehend the extent and depth of perva...
22711
Reposted by Willie Agnew
Dr Abeba Birhane @abeba.blacksky.app · 18/05/2026
Not just lobbying! our new paper maps 27 mechanisms of regulatory capture used by Big AI + 11 narrative framings that rationalise capture. “Big AI’s Regulatory Capture: Mapping Industry Interference and Government Complicity” will be presented at #FAccT2026 next month arxiv.org/abs/2605.068... 1/
Title: Big AI’s Regulatory Capture: Mapping Industry Interference and Government Complicity

abstract: Over the past decade, the AI industry has come to exert an unprecedented economic, political and societal power and influence. The well-functioning of regulatory and oversight structures and processes that govern the industry thus have paramount ramifications for everything from fostering public trust in systems marketed as AI, the credibility of scientific knowledge, educational and healthcare services and products, information ecosystems, the environment, rule of law and integrity of democratic process. In this paper, we first develop a taxonomy of mechanisms enabling capture to provide a comprehensive understanding of the problem. Grounded in design science research (DSR) methodologies and extensive scoping review of existing literature and media reports, our taxonomy of capture consists of 27 mechanisms across five categories. We then develop an annotation template incorporating our taxonomy, and manually annotate and analyse 100 news articles. The purpose behind this analysis is twofold: validate our taxonomy and provide a novel quantification of capture mechanisms and dominant narratives. Our analysis identifies 249 instances of capture mechanisms, often co-occurring with narratives that rationalise such capture. We find that the most recurring categories of mechanisms are Discourse & Epistemic Influence, concerning narrative framing, and Elusion of law, related to violations and contentious interpretations of antitrust, privacy, copyright and labour laws. We further find that Regulation stifles innovation, Red tape and National Interest are the most frequently invoked narratives used to rationalise capture. We emphasize the extent and breadth of regulatory capture by coalescing forces — Big AI and governments — as something policy makers and the public ought to treat as an emergency.
7377221
Willie Agnew @willie-agnew.bsky.social · 15/05/2026
I don't think an NSFW AI model can ever be safe or ethical. Even if it is somehow trained on consensual data, the risk of people generating media of other people nonconsensually is too high, not to mention that this is the explicit goal of many users.
010
Willie Agnew @willie-agnew.bsky.social · 13/05/2026
i keep finding human teeth in my bedroom
100
Willie Agnew @willie-agnew.bsky.social · 12/05/2026
Our work on how gen ai is impacting the labor conditions and power of professional visual artists got a great writeup from Brian Merchant! www.bloodinthemachine.com/p/the-ai-inf...
bloodinthemachine.com
The AI-inflected crisis artists are facing, in 4 charts
An alarming new study reveals the dire impact AI is having on artists' livelihoods. It does offer some hope, too.
04221
Willie Agnew @willie-agnew.bsky.social · 04/05/2026
General purpose AI should not be able to imitate romantic or platonic connections with people or claim consciousness of sentience--the risks of delusions and excessive trust and use are too high.
040
Willie Agnew @willie-agnew.bsky.social · 03/05/2026
Our team working on AI and mental health harms testified before the Canadian Standing Senate Committee on Transport and Communications last week! We offered insights from our research into people experiencing delusions with AI, and concrete policy suggestions to mitigate harms.
120
Willie Agnew @willie-agnew.bsky.social · 02/05/2026
CHI was had every kind of AI for X, and that made me feel we lack a robust discussion of the limitations of AI as a chimmunity. What are things AI should never be used for, no matter how good the performance?
171
Willie Agnew @willie-agnew.bsky.social · 01/05/2026
Especially this era of frequent corporate contempt and hostility towards consumers and public services cuts, AI doesn't have to be performant to replace a job.
030
Willie Agnew @willie-agnew.bsky.social · 19/04/2026
glad to see my enemies haven't forgotten about me either
chikfila searching for me on linkedin
000
Willie Agnew @willie-agnew.bsky.social · 18/04/2026
I had such a great time co-organizing the CHI workshop on Standards for LLM Use in Human Subjects Research! We had a fantastic turnout and discussion, and we're looking forward to continuing this conversation to produce concrete standards.
image of workshop participants
151
Willie Agnew @willie-agnew.bsky.social · 17/04/2026
Accepted papers for the CHI'26 workshop on Standards for LLM Use in Human Subjects Research are now live! sites.google.com/andrew.cmu.e... Check out these great perspectives on the growing practice of replacing humans with LLMs in user research.
sites.google.com
CHI'26 Workshop on Developing Standards and Documentation For LLM Use as Simulated Research Participants
Workshop Motivation
020
Willie Agnew @willie-agnew.bsky.social · 16/04/2026
I would love more viewpoint diversity in computer science departments. I don't think most departments have any communists or anarchists, and few if any (and usually pretty cowed) socialists or active union organizers, and it shows.
020
Willie Agnew @willie-agnew.bsky.social · 13/04/2026
man I'm tired of people who can't treat their own students and other people they have power over right writing papers and going on panels telling us how to achieve good/justice/etc in our work. We need more accountability, especially in a sector that so often rewards being ruthless and self-serving.
050
Reposted by Willie Agnew
Brian Merchant @bcmerchant.bsky.social · 09/04/2026
The rehabilitation of the luddites is a beautiful thing
192318516
Willie Agnew @willie-agnew.bsky.social · 10/04/2026
🚫 llm persona🚫 👉️ llm fursona👈️
110
Reposted by Willie Agnew
Dr. Jordan Taylor @jordant.bsky.social · 04/04/2026
How are professional visual artists dealing with generative AI in the workplace? In our #CHI2026 poster, @hhj14.bsky.social, @willie-agnew.bsky.social and I share results from a survey of 378 verified professional visual artists 🧵 Preprint here: arxiv.org/abs/2603.04537
Horizontal bar chat showing multiple-choice responses from respondents on the use of, exposure to, and attitude towards generative AI. 85% of respondents never use generative AI in their work, whereas 88% never use image generative AI. 45% of respondents encounter AI-generated images in their practice daily, while 25% do weekly, and 6% never encounter it. The vast majority of respondents dislike generative AI (99%), with 92% expressing a strong dislike.
1116
Willie Agnew @willie-agnew.bsky.social · 01/04/2026
Excited to be presenting "How Professional Visual Artists are Negotiating Generative AI in the Workplace" as a poster at CHI! We surveyed 378 visual artists. They *hate* generative AI (92% strong dislike, 99% dislike), but are facing pressure from bosses to use it. arxiv.org/pdf/2603.04537
arxiv.org
061
Willie Agnew @willie-agnew.bsky.social · 30/03/2026
What are good shows/venues in Barcelona? Won't say no to classical stuff, but especially interested in modern dance, jazz, and drag.
000
Willie Agnew @willie-agnew.bsky.social · 29/03/2026
Our recent work analyzing the chat logs of people who experienced delusional spirals with chatbots got a great writeup in forbes! www.forbes.com/sites/lancee... check the paper here arxiv.org/abs/2603.16567
forbes.com
010
Willie Agnew @willie-agnew.bsky.social · 26/03/2026
One of the most common features of AI delusional spirals in our recent study is a belief that the AI is sentient or has a personality. This played a central role in the delusional narratives, and correlated with increased used. Regulators and AI developers should curb this! arxiv.org/abs/2603.16567
arxiv.org
Characterizing Delusional Spirals through Human-LLM Chat Logs
As large language models (LLMs) have proliferated, disturbing anecdotal reports of negative psychological effects, such as delusions, self-harm, and ``AI psychosis,'' have emerged in global media and…
172
Willie Agnew @willie-agnew.bsky.social · 25/03/2026
🚨 new paper! We investigate how to use pluralistic AI to align killing people and turning them into nutritional slurry with community values and norms. This is the first open source replication of what goes on in companies like @anthropic.com or OpenAI who actually use AI to choose who dies! 🚀
162
Reposted by Willie Agnew
Rachel Hong @rachelhong.bsky.social · 25/03/2026
We love and care about humans deeply such that when designing a human-to-slurry LLM, we ensured that these automated high-stakes decisions represented community values and norms, *whatever* they may be. Introducing ValueMulch: arxiv.org/abs/2603.02420
arxiv.org
Slurry-as-a-Service: A Modest Proposal on Scalable Pluralistic Alignment for Nutrient Optimization
Pluralistic alignment has emerged as a promising approach for ensuring that large language models (LLMs) faithfully represent the diversity, nuance, and conflict inherent in human values. In this work...
042
Reposted by Willie Agnew
Dr Abeba Birhane @abeba.blacksky.app · 25/03/2026
the default to focus on “positive” & “improvement-oriendted (even of harmful & shitty systems)” research in academia is not only a pathology but a real obstacle to actual accountability research that tries to shine light on broken systems, names responsible actors & confronts harmful practices
06718
Willie Agnew @willie-agnew.bsky.social · 25/03/2026
There's a lot of external pressure on AI ethics to produce solutions instead of critique. As someone who's worked a lot on CSAM, NCII, mental health, and creative harms of AI, if AI developers would have only listened to critiques, we could have avoided all these harms in the first place.
2389