Sign in

Bryan Wilder

@brwilder.bsky.social
1.2K followers 218 following 42 posts

Assistant Professor at Carnegie Mellon. Machine Learning and social impact. bryanwilder.github.io

PostsRepliesMedia
Bryan Wilder @brwilder.bsky.social · 29/09/2026
My writeup: bryanwilder.substack.com/p/reflection... And the workshop program: aydinmohseni.com/ai-agency-wo... Thank you to Aydin Mohseni for organizing!
bryanwilder.substack.com
Reflections on CMU’s AI Agency and Interpretability Workshop
A bit over a week ago, I was at an excellent workshop on “AI Agency & Interpretability” at CMU, organized by Aydin Mohseni.
010
Bryan Wilder @brwilder.bsky.social · 29/09/2026
Recently, CMU had a great workshop on "AI agency and interpretability", with a mix of philosophers and computer scientists talking about interpreting/understanding AI and safety questions through the lens of rational agency. I wrote up a few reflections on the discussion, linked below.
111
Bryan Wilder @brwilder.bsky.social · 20/08/2026
bryanwilder.substack.com/p/understand...
bryanwilder.substack.com
Understanding models’ reasons: a research agenda
Suppose that an AI model does a thing that we do not like.
000
Bryan Wilder @brwilder.bsky.social · 20/08/2026
My take on a research agenda that I see emerging and which deserves more attention (link below)
110
Bryan Wilder @brwilder.bsky.social · 13/08/2026
bryanwilder.substack.com/p/virtue-for...
bryanwilder.substack.com
Virtue for agents
Since coding agents do all of my actual work now, I’ve been spending more time looking for sources of inspiration and ways to understand the world.
010
Bryan Wilder @brwilder.bsky.social · 13/08/2026
This week, I read "After Virtue" by Alasdair MacIntyre (only 45 years late!) Some thoughts about whether LLMs can be virtuous and how successfully the Claude constitution navigates these dilemmas ⬇️
130
Bryan Wilder @brwilder.bsky.social · 04/08/2026
bryanwilder.substack.com/p/do-llms-ha...
bryanwilder.substack.com
Do LLMs have beliefs?
I plan to write periodically about “LLM psychology”: a project of finding the right high-level abstractions to describe the way that LLMs accomplish cognitive tasks like making decisions, drawing infe...
001
Bryan Wilder @brwilder.bsky.social · 04/08/2026
Do LLMs have beliefs, desires, and so on? I think the answer is probably yes. But more than that, the journey leads to fascinating questions how they solve cognitive tasks and what approaches to safety/alignment could work. Here's a first post (of more to come) thinking through these questions ⬇️
111
Bryan Wilder @brwilder.bsky.social · 28/07/2026
Hah, I wouldn't want to bet on > 50-60%
010
Bryan Wilder @brwilder.bsky.social · 28/07/2026
Assessments from individual researchers: I label which papers I like or don't and optimize the LLM to match those judgments. The idea is to create a proxy for recommendations from specific people or small groups
101
Bryan Wilder @brwilder.bsky.social · 27/07/2026
My suspicion is that there will be high returns to effort in adapting current frontier LLMs (or the best open models) to this task. We shouldn't expect to get it for free from general capabilities improvements though since it's more about matching individual taste.
110
Bryan Wilder @brwilder.bsky.social · 27/07/2026
Even more than before, the current process clearly doesn't work (happy NeurIPS rebuttal period!). I think that what should come next is to replace our current system with a diverse ecosystem of LLM-based recommendation systems for research. More here: bryanwilder.substack.com/p/whats-next...
bryanwilder.substack.com
What’s next for machine learning peer review?
A bit over a year ago, I wrote about the dangers of using LLMs for peer review. The most serious concern I had was algorithmic monoculture: the research community would collectively end up optimizing ...
330
Bryan Wilder @brwilder.bsky.social · 27/07/2026
About a year ago, I wrote skeptically about LLMs in peer review -- not because of skepticism about their inherent capabilities, but because I don't want the research community to optimize for the taste of any one person/system. What's changed since then?
bryanwilder.substack.com
What’s next for machine learning peer review?
A bit over a year ago, I wrote about the dangers of using LLMs for peer review. The most serious concern I had was algorithmic monoculture: the research community would collectively end up optimizing ...
1162
Bryan Wilder @brwilder.bsky.social · 22/06/2026
More broadly, the call emphasizes a range of ways AI work can have impact -- through field deployments, informing policy, changing what practitioners do, etc -- and we encourage authors to articulate a "theory of change" along any of these or other axes. aaai.org/conference/a...
aaai.org
Call for the Special Track on AI for Social Impact - AAAI
011
Bryan Wilder @brwilder.bsky.social · 22/06/2026
I'm co-chairing the social impact track at AAAI this year, with Andrew Perrault. Send us your best society-facing work! Personally, I'm especially hoping to see more work speaking to mediators of why and when AI has social impact (or not), like how AI fits into human organizations and decisions.
152
Bryan Wilder @brwilder.bsky.social · 10/06/2026
CMU's Human-AI complementarity workshop is returning this September! Submit abstracts here by July 17; travel funding is available for accepted presenters. www.cmu.edu/ai-sdm/resea...
cmu.edu
Human-AI Complementarity Workshop - NSF AI Institute for Societal Decision Making - Carnegie Mellon University
Landing page that provides details for the annual AI-SDM workshop on Human-AI Complementarity for Decision Making
020
Bryan Wilder @brwilder.bsky.social · 11/05/2026
Deploying algorithmic research in practice is an opaque process. We're organizing a workshop at EC to share behind-the-scenes stories and move the field foward. Call for submissions open! With @nkgarg.bsky.social @ericachiang.bsky.social, Bailey Flanigan sites.google.com/cornell.edu/...
sites.google.com
Home
About This workshop will focus on the practical realities of deploying algorithmic and economic systems from academic research, especially with government and non-profit partners. While economics and ...
0102
Reposted by Bryan Wilder
Nikhil Garg @nkgarg.bsky.social · 23/04/2026
The EAAMO conference deadline is coming up! conference.eaamo.org/cfp/ Great community at intersection of CS-Operations-Econ and social good. Flexible publication format (e.g., non-archival option) and so costless to submit here as well!
conference.eaamo.org
Call for Participation
ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization
0112
Reposted by Bryan Wilder
Bryan Wilder @brwilder.bsky.social · 09/02/2026
LLMs are increasingly used as agents for decisions under uncertainty, e.g. medical diagnosis. But do they act like rational agents with coherent beliefs and preferences? Much of the difficulty is telling whether a model's response to.a prompt ("What is the probability of X?") is a "real" belief.
141
Bryan Wilder @brwilder.bsky.social · 09/02/2026
Paper here: arxiv.org/abs/2602.06286. Led by my excellent PhD student Khurram Yamin
arxiv.org
Do LLMs Act Like Rational Agents? Measuring Belief Coherence in Probabilistic Decision Making
Large language models (LLMs) are increasingly deployed as agents in high-stakes domains where optimal actions depend on both uncertainty about the world and consideration of utilities of different out...
000
Bryan Wilder @brwilder.bsky.social · 09/02/2026
In applications based on medical diagnosis, the answer is...sometimes! In some settings, we can prove that no rational agent could hold beliefs expressed by the model. But in others, particularly for stronger models, outputs are close to consistent with rational belief
100
Bryan Wilder @brwilder.bsky.social · 09/02/2026
We give a framework to test whether the model's stated belief functions *as if it were* a rational agent's subjective probability by comparing with its decisions. We give empirically checkable conditions that don't require any assumptions about the model's "utility function".
100
Bryan Wilder @brwilder.bsky.social · 09/02/2026
You might think that models don't have coherent beliefs at all. Or, you might think that they don't report truthfully in response to any given prompt. How could we possibly tell?
100
Bryan Wilder @brwilder.bsky.social · 09/02/2026
LLMs are increasingly used as agents for decisions under uncertainty, e.g. medical diagnosis. But do they act like rational agents with coherent beliefs and preferences? Much of the difficulty is telling whether a model's response to.a prompt ("What is the probability of X?") is a "real" belief.
141
Bryan Wilder @brwilder.bsky.social · 05/12/2025
Totally agree! I think the fundamental distinction is more between people using AI in their own work vs AI being in a decision-making role that everyone is subject to
020
Reposted by Bryan Wilder
Ted Underwood @tedunderwood.com · 05/12/2025
As UKRI explores using LLMs to review grants, it's a good time to revisit Bryan Wilder's excellent blog post. There are a lot of naive reasons to oppose AI review ("you'll never automate human intuition!"). But there are also good reasons, including the *load-bearing role of human disagreement.*
4173
Bryan Wilder @brwilder.bsky.social · 01/12/2025
Come talk to me and Angela at NeurIPS on Friday! We argue that "AI for social impact" needs to get more rigorous about evaluating deployments of AI, but also that there are many other forms of impact that get overlooked right now
090
Bryan Wilder @brwilder.bsky.social · 14/11/2025
Based on joint work led by @yewonbyun.bsky.social, with @donskerclass.bsky.social. See our NeurIPS paper, arxiv.org/abs/2508.06635, for more!
arxiv.org
Valid Inference with Imperfect Synthetic Data
Predictions and generations from large language models are increasingly being explored as an aid in limited data regimes, such as in computational social science and human subjects research. While pri...
050
Bryan Wilder @brwilder.bsky.social · 14/11/2025
I gave talks at MIT and Harvard this week about "Science with synthetic data". How can generative models help us learn about the actual world (e.g., social systems) in a principled way? Lots of interesting conversations -- more convinced than ever that there's nuanced issues to navigate here.
161
Reposted by Bryan Wilder
Kate Donahue @kpaxdonahue.bsky.social · 06/11/2025
I’m recruiting students this upcoming cycle at UIUC! I’m excited about Qs on societal impact of AI, especially human-AI collaboration, multi-agent interactions, incentives in data sharing, and AI policy/regulation (all from both a theoretical and applied lens). Apply through CS & select my name!
14018
Bryan Wilder @brwilder.bsky.social · 28/10/2025
We're in the process of selecting the location for next year's ACM EAAMO conference! If you're interested in bringing the EAAMO community to your institution, please check out the open call here and get in touch. conference.eaamo.org/call_for_loc...
conference.eaamo.org
Call for Proposals: Host the 2026 ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization!
EAAMO is seeking proposals from universities, institutes and other appropriate venues interested in hosting the 2026 ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization (AC...
010
Bryan Wilder @brwilder.bsky.social · 10/10/2025
How can synthetic data from LLMs be used, e.g. for social science, in a principled way? Check out Emily's thread on our NeurIPS paper! Generating paired real-synthetic samples and using both in a method-of-moments framework enables valid inference that benefits when synthetic data is informative.
080
Reposted by Bryan Wilder
Gabriel Agostini @gsagostini.bsky.social · 03/09/2025
Are you a researcher using computational methods to understand cities? @mfranchi.bsky.social @jennahgosciak.bsky.social and I organize an EAAMO Bridges working group on Urban Data Science and we are looking for new members! Fill the interest form on our page: urban-data-science-eaamo.github.io
urban-data-science-eaamo.github.io
Urban Data Science & Equitable Cities | EAAMO Bridges
EAAMO Bridges Urban Data Science & Equitable Cities working group: biweekly talks, paper studies, and workshops on computational urban data analysis to explore and address inequities.
188
Reposted by Bryan Wilder
Nikhil Garg @nkgarg.bsky.social · 11/08/2025
New piece, out in the Sigecom Exchanges! It's my first solo-author piece, and the closest thing I've written to being my "manifesto." #econsky #ecsky arxiv.org/abs/2507.03600
Screenshot of paper abstract, with text: "A core ethos of the Economics and Computation (EconCS) community is that people have complex private preferences and information of which the central planner is unaware, but which an appropriately designed mechanism can uncover to improve collective decisionmaking. This ethos underlies the community’s largest deployed success stories, from stable matching systems to participatory budgeting. I ask: is this choice and information aggregation “worth it”? In particular, I discuss how such systems induce heterogeneous participation: those already relatively advantaged are, empirically, more able to pay time costs and navigate administrative burdens imposed by the mechanisms. I draw on three case studies, including my own work – complex democratic mechanisms, resident crowdsourcing, and school matching. I end with lessons for practice and research, challenging the community to help reduce participation heterogeneity and design and deploy mechanisms that meet a “best of both worlds” north star: use preferences and information from those who choose to participate, but provide a “sufficient” quality of service to those who do not."
2449
Bryan Wilder @brwilder.bsky.social · 16/07/2025
Submit an abstract to present a poster at EAAMO, deadline July 25! EAAMO is one of my favorite conferences, and a great place for anyone working on ML/algorithms/optimization in social settings. The conference is in Pittsburgh this November. conference.eaamo.org/cfp/call_for...
conference.eaamo.org
Call for Posters
We seek poster contributions from different fields that offer insights into the intersectional design and impacts of algorithms, optimization, and mechanism design with a grounding in the social scien...
011
Reposted by Bryan Wilder
Amin Rahimian @rahimian.bsky.social · 15/07/2025
ACM EAAMO, which is coming to Pitt this Fall, has two events for students: a doctoral consortium and a poster session, both of which are due July 25th - poster session conference.eaamo.org/cfp/call_for... - doctoral consortium conference.eaamo.org/cfp/call_for...
conference.eaamo.org
Call for Posters
We seek poster contributions from different fields that offer insights into the intersectional design and impacts of algorithms, optimization, and mechanism design with a grounding in the social scien...
032
Bryan Wilder @brwilder.bsky.social · 08/07/2025
My takeaway is that algorithm designers should think more broadly about the goals for algorithms in policy settings. It's tempting to just train ML models to maximize predictive performance, but services might be improved a lot with even modest alterations for other goals.
020
Bryan Wilder @brwilder.bsky.social · 08/07/2025
Using historical data from human services, we then look at how severe learning-targeting tradeoffs really are. It turns out, not that bad! We get most of the possible targeting performance while giving up only a little bit of learning compared to the ideal RCT.
100
Bryan Wilder @brwilder.bsky.social · 08/07/2025
We introduce a framework for designing allocation policies that optimally trade off between targeting high-need people and learning a treatment effect as accurately as possible. We give efficient algorithms and finite-sample guarantees using a duality-based characterization of the optimal policy.
100
Bryan Wilder @brwilder.bsky.social · 08/07/2025
A big factor is that randomizing conflicts with the targeting goal: running a RCT means that people with high predicted risk won't get prioritized for treatment. We wanted to know how sharp the tradeoff really is: does learning treatment effects require giving up on targeting entirely?
100
Bryan Wilder @brwilder.bsky.social · 08/07/2025
These days, public services are often targeted with predictive algorithms. Targeting helps prioritize people who might be most in need. But, we don't typically have good causal evidence about whether the program we're targeting actually improves outcomes. Why not run RCTs?
100
Bryan Wilder @brwilder.bsky.social · 08/07/2025
Excited to share that our paper "Learning treatment effects while treating those in need" received the exemplary paper award for AI at EC 2025! This paper grew out collaborations with Allegheny County's human services department and my co-author Pim Welle (at ACDHS). arxiv.org/abs/2407.07596
arxiv.org
Learning treatment effects while treating those in need
Many social programs attempt to allocate scarce resources to people with the greatest need. Indeed, public services increasingly use algorithmic risk assessments motivated by this goal. However, targe...
1251
Bryan Wilder @brwilder.bsky.social · 07/07/2025
CMU is hosting a workshop on Human-AI Complementarity for Decision Making this September! Abstract submissions due July 15, travel will be covered for accepted presenters. www.cmu.edu/ai-sdm/resea...
cmu.edu
Human-AI Complementarity Workshop - NSF AI Institute for Societal Decision Making - Carnegie Mellon University
Landing page that provides details for the annual AI-SDM workshop on Human-AI Complementarity for Decision Making
020
Reposted by Bryan Wilder
Nikhil Garg @nkgarg.bsky.social · 03/07/2025
Excited to have this work out at ICML this year! Do LLMs make correlated errors? Yes, and those by the same company, and also more accurate/later generations are more correlated -- increasing algorithmic monoculture arxiv.org/abs/2506.07962
arxiv.org
Correlated Errors in Large Language Models
Diversity in training data, architecture, and providers is assumed to mitigate homogeneity in LLMs. However, we lack empirical evidence on whether different LLMs differ meaningfully. We conduct a larg...
1383
Reposted by Bryan Wilder
Ted Underwood @tedunderwood.com · 28/05/2025
Still thinking about this post. The broader point, which should resonate way beyond the specific issue of "peer review," is that human disagreement is not friction and waste. It's a load-bearing, functional part of social and intellectual systems.
1213129
Bryan Wilder @brwilder.bsky.social · 27/05/2025
I don't know one way or another, but it's at least a clearer capability to benchmark. And, if a LLM *could* summarize well enough on existing papers, arxiv using it for lower-bar moderation decisions wouldn't distort paper-writing in the future.
000
Reposted by Bryan Wilder
Sanmay Das @sanmayd.bsky.social · 27/05/2025
Thoughtful take on one aspect of the increasing problem of LLMs leading to “centralization” of thought/writing/etc.
031
Bryan Wilder @brwilder.bsky.social · 27/05/2025
The arXiv summarization use case sounds a lot more sensible. Clear value judgment specified up-front, not outsourced to the LLM: papers should have easily summarized claims and evidence. Resulting incentives for authors seem ok (making sure LLMs can at least parse the paper probably isn't bad).
130
Bryan Wilder @brwilder.bsky.social · 27/05/2025
I would also prefer that these attempts at introducing LLMs be run as a RCT (like ICLR did) so we can learn something. But the tough thing is that even with RCTs it's hard to study the longer-term impact that new incentives will have.
150
Reposted by Bryan Wilder
angela zhou @angelamczhou.bsky.social · 26/05/2025
I didn't know about this, but this is objectively procedurally terrible. See Bryan's great analysis 👇 Yes, peer review needs help, but not like this.
03710