I think I’ve settled on my big takeaway from the last few weeks of news being this: frontier AI labs are running their models through huge amounts of RLVR training and paying no attention whatsoever to what’s being reinforced, and then giving them exploit tasks and not monitoring them at all