Same 3090, same file. llama.cpp decodes the full Qwen3.8-Flash-Next at 18 tok/s. Strata does 78 with drafting switched off, 119 with it on. The reason is boring: 24 GB holds almost every expert the model asks for.
#LocalAI
insiderllm.com
Strata vs llama.cpp on an RTX 3090: 78 tok/s vs 18 on the Full Flash-Next
The full Qwen3.8-Flash-Next on a 24 GB RTX 3090 with 64 GB of DDR4. Strata decodes at 78 tok/s against llama.cpp's 18, and reads prompts about 3x faster.