So cool! transformers can now run GGUF models efficiently.
To bring the performance close to llama.cpp, it uses the underlying ggml kernels through our kernels library.
Awesome work by @marcsun.bsky.social et al.!
huggingface.co/blog/transfo...
huggingface.co
Transformers now runs llama.cpp quants
We’re on a journey to advance and democratize artificial intelligence through open source and open science.