Reposted by Nari Johnson
www.lesswrong.com/posts/HsijSh...
@bradknox.bsky.social, Brian Christian, and I wrote a thing. We argue that an unexamined cause of the Open AI / Hugging Face Attack was the Exploit Gym evaluation metric. We post this on Less Wrong as a plea to practitioners to design better metrics in the future.
lesswrong.com
An unexamined cause of the OpenAI Hugging Face hacking incident:
its binary performance metric — LessWrong
We argue that a main cause of the OpenAI Hugging Face incident was overlooked: the overly simple evaluation metric in ExploitGym was misaligned. Furt…