They're all generally regressing as they chase benchmarks and coding gains. For example, on arena Opus 4.6 is ranked higher than 4.7, which is higher than 5. Same goes for OpenAI's models. Across tasks they're making more mistakes, such as story contradictions, b/c they're overtraining coding/agents