Reposted by Ben Berman
16000 tokens per second on a decent model. This type of speed is the future. Opens up an entirely new class of user experiences
reddit.com
From the LocalLLaMA community on Reddit: Free ASIC Llama 3.1 8B inference at 16,000 tok/s - no, not a joke
Explore this post and more from the LocalLLaMA community