Skip to content
Ishaan Reddy

Cookbook · Inference

Quantization

3 minquantizationinferencedeploymentgptq

The idea, in one analogy

A photograph stored at full resolution captures every detail, but most of that detail is redundant: you can compress the file substantially before a viewer notices the difference. Quantization does the same thing to a model's weights: instead of storing each number in high precision (16 or 32 bits), it rounds them to a much smaller set of representable values (8 bits, 4 bits, sometimes fewer), shrinking the model's memory footprint with a small, often barely-noticeable cost in accuracy.

What's being rounded

A model's weights are ordinarily stored as 16-bit or 32-bit floating-point numbers. Quantization maps that continuous range of values down onto a small fixed set of integers, then stores a scale factor so the integers can be converted back to approximate real values on the fly.

Two different times to quantize

Post-training quantization (PTQ): take an already-trained model and quantize its weights afterward, usually using a small calibration dataset to decide good scale factors per layer. Fast and simple, the standard choice when you just want a smaller model to deploy.

Quantization-aware training (QAT): simulate the rounding error during training itself (or fine-tuning), so the model's weights adjust to compensate for the precision loss before it ever gets deployed in quantized form. More expensive (it's a training run, not a one-time conversion), but recovers more accuracy at very low bit-widths than PTQ alone.

QLoRA is a related but distinct idea: it quantizes the frozen base model during fine-tuning specifically to reduce memory while training, not primarily to shrink the model for deployment afterward, though the resulting quantized weights can serve both purposes.

The tradeoff, concretely

Going from 16-bit to 8-bit roughly halves memory and often costs little to no measurable accuracy on standard benchmarks. Going to 4-bit roughly quarters memory but starts trading off real accuracy, especially on models that were already small to begin with (there's less redundancy to sacrifice). Going below 4-bit (2-3 bit schemes) is an active research area with much larger accuracy costs unless paired with more careful techniques (per-channel scales, outlier handling for the small number of unusually large weight values that don't compress well).

Where to look further

  • Dettmers et al., "LLM.int8()": 8-bit quantization that handles outlier weight values separately to avoid their accuracy cost.
  • Frantar et al., "GPTQ" and Lin et al., "AWQ": two widely-used post-training quantization methods for 4-bit weights.
  • GGUF: the quantized model format used by llama.cpp, with several bit-width variants (Q4, Q5, Q8, etc).
  • bitsandbytes: the library behind LLM.int8() and the NF4 format QLoRA uses.