Reposted by Janne Sinkkonen
Even that bare ascii representation seems to be challenging due to perception problems, incl. tokenization.
anokas.substack.com/p/llms-strug...
anokas.substack.com
LLMs struggle with perception, not reasoning, in ARC-AGI
What made o3 so much better than previous models on this benchmark?