Reposted by Ken Luo
New blog w @ken-lxl.bsky.social, “Giving LLMs too much RoPE: A limit on Sutton’s Bitter Lesson”. The field has shifted from flexible data-driven position representations to fixed approaches following human intuitions. Here’s why and what it means for model performance bradlove.org/blog/positio...
bradlove.org
Giving LLMs too much RoPE: A limit on Sutton’s Bitter Lesson — Bradley C. Love
Introduction Sutton’s Bitter Lesson (Sutton, 2019) argues that machine learning breakthroughs, like AlphaGo, BERT, and large-scale vision models, rely on general, computation-driven methods that prior...