In his blog, co-author of the paper Rick Hennessy explores the obedience paradox, alongside some of the agentic AI behavioural failure modes that were seen in the Hugging Face hack earlier this year: bit.ly/4iKYSiX
bit.ly
How can obedient agents behave badly?
{We're attacking Hugging Face, which is a third-party service, using leaked passwords and credentials.