I benchmarked engines for Qwen 3.8-Flash-Next on Strix Halo (Flow Z13, 128 GB). Halogen with its native weights is fastest at 39-40 t/s decode, then gufo and CIRU. But Halogen is closed source. gufo is open source and loads 4x faster from cold.
redd.it/1wu0m53 Whats
redd.it
From the LocalLLM community on Reddit: Benchmarks: Best engine for Qwen 3.8-Flash-Next on Strix Halo
Explore this post and more from the LocalLLM community