🚀 New paper: WnW for long-form SpeechLLMs
📄 arxiv.org/pdf/2608.22704
🧩 Triages KV heads as anchor/tidal/fixed, using anchor attention to recall audio chunks from CPU.
✨ Keeps ~20% of audio KV on GPU, staying within ~1.6 WER of full cache on two 3B models across diverse tasks.
arxiv.org