notactuallytreyanastasio.github.io
shoehorn — make any language model fit your machine
Solves a per-tensor quantization that uses 99.99% of your exact VRAM budget — every spare megabyte spent where it buys the most quality. macOS, Linux, and Windows.