Sonnet 5.5 is worse than Sonnet 5 on my adversarial esoteric language benchmark. Seems to be due to 5.5 wanting more reassurance from the user, rather than just pushing to complete the task given.
bench.killswitch-lang.org
KillSwitch-Bench
A benchmark for evaluating coding agents on tasks in an adversarial esoteric language.