Reposted by Anna Neumann
Anna Neumann, Holli Sargeant and Jat Singh argue that system prompts alone don't predict model behavior, so AI safety assessments must evaluate a system's outputs, not solely its written instructions.
techpolicy.press
Making AI Safe Requires More Than Constraining Its Instructions
System prompts alone don't predict model behavior, so "prompt governance" is insufficient for AI safety, write Anna Neumann, Holli Sargeant and Jat Singh.