i'm curious if the models just can't be constrained to follow the rules of a CTF, or if the researchers prompt was worded in such a way that the models, once they got access to the internet, concluded it was a fake sandbox internet and still in scope of the challenge.