“Language Models Can Control Their Own Attention”
So this paper introduces Declarative Attention, where the model explicitly switches between global, focused, and local attention during reasoning, letting the inference engine skip irrelevant KV cache regions.
alphaxiv.org/abs/2609.02737