AI engineering · 26 of 42
Fewer bits per number, same shape
Scroll
Fewer bits per number, same shape
A model's weights are numbers, and they do not all need full precision. Quantization stores them in fewer bits — eight, or four — shrinking the model without changing its shape.
The gain is not only disk. A smaller model fits cheaper hardware and leaves more room for the KV cache, which is usually what actually limits how many people you can serve.
Quality loss is small on average and not evenly distributed. It tends to show up on the hardest and rarest inputs — exactly the ones a benchmark average hides. Run your own difficult cases before and after, not just an aggregate score.
Inference