github.com
k-quants by ikawrakow · Pull Request #1684 · ggml-org/llama.cpp
What
This PR adds a series of 2-6 bit quantization methods, along with quantization mixes, as proposed in #1240 and #1256. Scalar, AVX2, ARM_NEON, and CUDA implementations are provided.
Why
This is...