Long-context models are coming to our phones and laptops.
Sevren (the new kid on the European block) ported FlashPrefill V2 to MLX, bringing block-sparse attention to Apple silicon giving a 1.78X speedup on 32K tokens.
OS code: github.com/sevren-ai/Fl...
Blog post: sevren.ai/blog/flashpr...
github.com
GitHub - sevren-ai/FlashPrefillv2-MLX: FlashPrefill V2 block-sparse long-context prefill attention for MLX on Apple silicon.
FlashPrefill V2 block-sparse long-context prefill attention for MLX on Apple silicon. - sevren-ai/FlashPrefillv2-MLX