New research suggests LLM math performance hides a deeper issue: the main bottleneck isn’t execution, it’s discovery. Once given the right mathematical primitives, models solve far more problems than expected. A useful shift in how we evaluate reasoning.
arxiv.org/abs/2610.02191v1