The paper is entitled "SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression" and https://twitter.com/Tim_Dettmers/status/1666076553665744896 is a nice summary
"SpQR allows lossless LLM inference at 4.75 bits with a 15% speedup. You can run a 33B LLM on a single 24GB GPU fully lossless. SpQR works by isolating sensitive weights with higher precision and roughly doubles improvements from GPTQ"
Code here: https://github.com/Vahe1994/SpQR (https://news.ycombinator.com/item?id=36219128 but no traction )
The paper is entitled "SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression" and https://twitter.com/Tim_Dettmers/status/1666076553665744896 is a nice summary
"SpQR allows lossless LLM inference at 4.75 bits with a 15% speedup. You can run a 33B LLM on a single 24GB GPU fully lossless. SpQR works by isolating sensitive weights with higher precision and roughly doubles improvements from GPTQ"
Code here: https://github.com/Vahe1994/SpQR (https://news.ycombinator.com/item?id=36219128 but no traction )