Reposted by Diego F. Aranha
If you optimize a model to find exploits, you should expect it to find them—and prepare for that. OpenAI didn't. They built a model, removed the safeguards, gave it the ExploitGym task, let it run, and didn't even monitor it. That's human decision-making.
mail.cyberneticforests.com/models-dont-...
mail.cyberneticforests.com
Models Don't Go Rogue
Stochastic Flocks & Cybersecurity 'Pandemonium'
💡This essay was drafted from my appearance on Mél Hogan's podcast, The Data Fix, discussing the OpenAI / Hugging Face hack. Embedded below or find it o...