How GPTQ reduces 175B-parameter models to 3–4 bits: a practical guide to post-training quantization
GPTQ can quantize 175B-parameter GPT-class models to 3–4 bits in about four GPU-hours using approximate second-order information — enough to run a 175B model on a single GPU — but accuracy and speed gains depend on the calibration data and kernel stack.