🔮At #ICML2026 we presented Attentive Multi-Layer Fusion for Vision Transformers, an efficient method for decoding information from multiple layers of ViTs for downstream task predictions. It’s almost as performant as full fine-tuning but only an inch more expensive than linear probing!