Uzu fast #inference engine 4 #macOS & #iOS. Native #Rust. Bindings: #Python, #TypeScript, #Swift or #CLI. ~ x4 faster vs. llama.cpp & #MLX with #Qwen 3.5 9B Q4, on #M5 #MacBook Pro:
- Uzu 92 tokens/s
- llama.cpp 22 t/s
- MLX 25 t/s
Needs lalamo model format converter
#AI #LLM
#OpenSource MIT Lic
github.com
GitHub - trymirai/uzu: A high-performance inference engine for AI models
A high-performance inference engine for AI models. Contribute to trymirai/uzu development by creating an account on GitHub.