PS: Huang et al.'s concurrent Looped Models Done Right, Part II studies terminal KV reuse (one cache per physical layer), learned training-depth priors and orthogonal injection. We share across core blocks (less memory) and study donor consistency
www.alphaxiv.org/abs/2610.lo...
alphaxiv.org
Towards Looped Models Done Right Part II: Rethinking at Fixed Points
The work studies how approximate fixed-point behavior in looped language models can reduce the costs of training, decoding, prompt processing, and reinforcement learning. A learned recurrence-depth...