Apple just published a research paper on the state of LLMs and reasoning models.
TLDR: Basic models win on easy tasks, Reasoning models slightly help on mid-complexity tasks and BOTH fully collapse on high-complexity tasks (wasting tokens, money & time). Singularity delayed?