Reposted by Keshav Ramji
> the token usage is 2.3x GPT Luna Max and almost 2x Kimi K3!
Imagine this is not benchmaxxed and test-time compute trade-off can be stretched to this level ... and text can definitely be compressed 🙂 cc @keshavramji.bsky.social
arxiv.org/abs/2604.22709
arxiv.org
Thinking Without Words: Efficient Latent Reasoning with Abstract Chain-of-Thought
While long, explicit chains-of-thought (CoT) have proven effective on complex reasoning tasks, they are costly to generate during inference. Non-verbal reasoning methods have emerged with shorter gene...