Can we give an agent a tool to check which instructions it should actually follow?
And if we do, can we make sure it uses it?
Yes, and yes.
disreguard.com/blog/posts/s...
disreguard.com
sig: instruction signing for prompt injection defense
We can create a clear trust boundary by signing instructions and giving models a tool to participate in making secure choices