GLM-5.3 beat Opus 4.8 on Z.ai's own Code Bench: 31.4% vs 29.5%. Ignore that. The number that matters is ~50K output tokens vs ~120K.
Agent work is billed in tokens, not accuracy points. Score-per-token is the leaderboard nobody publishes.
z.ai/blog/glm-5.3
z.ai
GLM-5.3: Frontier Coding with Emergent Cyber Capabilities
Z.ai's open-weights 753B MoE coding model: 31.4% on Z.ai Code Bench at ~50K output tokens vs Claude Opus 4.8's 29.5% at ~120K.