Sign in

Brad Knox

@bradknox.bsky.social
1.4K followers 30 following 19 posts

Research Associate Professor in CS at UT Austin. I research how humans can specify aligned reward functions.

PostsRepliesMedia
Reposted by Brad Knox
Center for Human-Compatible AI @chai-berkeley.bsky.social · 24/09/2026
@bradknox.bsky.social, @brianchristian.bsky.social, and Serena Booth argue that a key cause of the OpenAI–Hugging Face incident was overlooked: ExploitGym’s overly simple evaluation metric was itself misaligned. They discuss techniques that could help avoid this in future. Link below ⬇️
humancompatible.ai
An unexamined cause of the OpenAI Hugging Face hacking incident: its binary performance metric – Center for Human-Compatible Artificial Intelligence
121
Reposted by Brad Knox
Serena Booth @reniebird.bsky.social · 23/09/2026
www.lesswrong.com/posts/HsijSh... @bradknox.bsky.social, Brian Christian, and I wrote a thing. We argue that an unexamined cause of the Open AI / Hugging Face Attack was the Exploit Gym evaluation metric. We post this on Less Wrong as a plea to practitioners to design better metrics in the future.
lesswrong.com
An unexamined cause of the OpenAI Hugging Face hacking incident: its binary performance metric — LessWrong
We argue that a main cause of the OpenAI Hugging Face incident was overlooked: the overly simple evaluation metric in ExploitGym was misaligned. Furt…
2259
Brad Knox @bradknox.bsky.social · 13/08/2026
AI-based companionship is on the rise. If you want a quick, likely frightening, and (dare I say) engaging primer on its potential harms, here's a recent talk I gave. www.youtube.com/watch?v=PJXAdsVLk6Y
youtube.com
- YouTube
Enjoy the videos and music you love, upload original content, and share it all with friends, family, and the world on YouTube.
020
Reposted by Brad Knox
Reinforcement Learning Conference @rl-conference.bsky.social · 21/07/2026
RLC26 schedule is now live! rl-conference.cc/schedule.html Join us Aug 15–18 for Workshops, Keynotes: Marc Bellemare, Sheila McIlraith, Danijar Hafner, Rika Antonova & Balaraman Ravindran, Oral sessions, Posters — & of course Cirque du Soleil🎪! Full talk & poster schedule coming soon —stay tuned!
0123
Reposted by Brad Knox
Reinforcement Learning Conference @rl-conference.bsky.social · 03/06/2026
Good news, RL Community! The early registration deadline for RLC'26 has been extended to June 17th — don't miss the early rates! Register today! Full refunds for cancellations before July 14, 2026.
073
Reposted by Brad Knox
Reinforcement Learning Conference @rl-conference.bsky.social · 21/05/2026
Got a great TMLR paper but missed the RLC deadline? Following last year’s success, @RL_Conference is back with a Journal-to-Conference track! Accepted TMLR papers within scope are invited to submit for consideration. Please submit here: docs.google.com/forms/d/e/1F...
docs.google.com
RLC Journal to Conference Track
This form is used to organize applications for journal papers to be presented at https://rl-conference.cc/.
084
Reposted by Brad Knox
Reinforcement Learning Conference @rl-conference.bsky.social · 26/02/2026
RLC2026 Call for Workshops! We’re already live: openreview.net/group?id=rl-... Here's the opportunity to help shape the conference & spotlight your own RL focus areas. Call: rl-conference.cc/call_for_wor... Deadline: Mar 12 (AoE) And don't forget the awesome banquet :) www.cirquedusoleil.com/ludo
075
Reposted by Brad Knox
Aadirupa Saha @aadirupa.bsky.social · 23/12/2025
RLC'26 is inviting submissions! Please mark your calendar, looking forward to all the new ideas. Happy holidays! @eugenevinitsky.bsky.social @audurand.bsky.social @schaul.bsky.social @glenberseth.bsky.social @sologen.bsky.social @pcastr.bsky.social @bradknox.bsky.social @cvoelcker.bsky.social
093
Brad Knox @bradknox.bsky.social · 18/11/2025
We're a couple of months into this exciting initiative, AHOI, and we've had two great speakers so far: Joe Carlsmith and Ryan Lowe. Check out our website for future speakers, videos of past talks, and our email list. liberalarts.utexas.edu/news/ai-huma...
liberalarts.utexas.edu
AI+Human Objectives Initiative at UT Austin Receives Grant from Coefficient Giving
AUSTIN, Texas — The AI+Human Objectives Initiative (AHOI) at The University of Texas at Austin has received an award from the grantmaking organization Coefficient Giving. The grant will fund AHOI’s wo...
020
Brad Knox @bradknox.bsky.social · 08/08/2025
Our paper, Towards Improving Reward Design in RL: A Reward Alignment Metric for RL Practitioners, won the Outstanding Paper Award on Emerging Topics in Reinforcement Learning this year at RLC! Congrats to 1st author @cmuslima.bsky.social! Paper: sites.google.com/ualberta.ca/...
0112
Reposted by Brad Knox
Reinforcement Learning Conference @rl-conference.bsky.social · 25/06/2025
Propose some socials for RLC! Research topics, affinity groups, niche interests, whatever comes to mind! rl-conference.cc/call_for_soc...
rl-conference.cc
RLC Call for Workshops
0118
Brad Knox @bradknox.bsky.social · 30/04/2025
To do non-LLM RLHF research with real, non-author HUMAN data, last I checked there were few datasets available. Synthetic data is usually quite unrealistic, compromising your results (see bradknox.net/human-prefer...). Our dataset (and code) with real humans: dataverse.tdl.org/dataset.xhtm...
bradknox.net
Models of human preference for learning reward functions – Brad Knox, PhD
020
Brad Knox @bradknox.bsky.social · 04/03/2025
Vibecoding apparently requires a magic touch I lack. In two attempts from scratch, Cursor AI + Claude 3.5 goes off the rails constantly and has eventually degenerated into non-functionality. Degeneration #2: Claude is only pretending to run terminal commands and edit my code. 🤦
Cursor + Claude pretending to do work.
210
Brad Knox @bradknox.bsky.social · 24/02/2025
Calling a reward function dense or sparse is a misnomer, AFAICT. (1/n)
120
Reposted by Brad Knox
Tom Schaul @schaul.bsky.social · 24/02/2025
Some extra motivation for those of you in RLC deadline mode: our line-up of keynote speakers -- as all accepted papers get a talk, they may attend yours! @rl-conference.bsky.social
RLC Keynote speakers: Leslie Kaelbling, Peter Dayan, Rich Sutton, Dale Schuurmans, Joelle Pineau, Michael Littman
03710
Reposted by Brad Knox
Reinforcement Learning Conference @rl-conference.bsky.social · 08/01/2025
Excited to announce the first RLC 2025 keynote speaker, a researcher who needs little introduction, whose textbook we've all read, and who keeps pushing the frontier on RL with human-level sample efficiency
Announcement of Richard S. Sutton as RLC 2025 keynote speaker
0514
Brad Knox @bradknox.bsky.social · 14/01/2025
RLHF algorithms assume humans generate preferences according to normative models. We propose a new method for model alignment: influence humans to conform to these assumptions through interface design. Good news: it works! #AI #MachineLearning #RLHF #Alignment (1/n)
First page of the paper Influencing Humans to Conform to Preference Models for RLHF, by Hatgis-Kessell et al.Our proposed method of influencing human preferences.
173