A new, highly recommended article from @rasbt.bsky.social:
"Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention"
-> magazine.sebastianraschka.com/p/recent-dev...
#MLsky
magazine.sebastianraschka.com
Recent Developments in LLM Architectures: KV Sharing, mHC, and Compressed Attention
From Gemma 4 to DeepSeek V4, How New Open-Weight LLMs Are Reducing Long-Context Costs