Model Quantization definition
Model quantization is a technique that reduces the numerical precision of a neural network's weights, and sometimes its activations, for example from 16-bit floating point to 8-bit or 4-bit integers. Quantized models need far less memory and often run faster and cheaper, which makes it possible to serve large language models on smaller GPUs, laptops or phones.
How quantization works
Neural network weights are stored as floating-point numbers. A 7-billion-parameter model in 16-bit precision needs roughly 14 GB just for its weights; at 4 bits it needs roughly 3.5 to 4 GB. Quantization maps the range of values in each block of weights onto a small set of integer levels, storing a scale factor so the values can be approximately reconstructed during computation.
The trade-off is rounding error. Good methods keep that error small: they quantize in small groups, protect the most important weights, or calibrate on sample data to see which weights matter most. Done well, 8-bit models are usually close to the original quality, and good 4-bit models lose surprisingly little on many tasks, although results vary by model and task. Current hardware also supports low-precision formats natively, such as FP8 and, on the latest data center GPUs, FP4, which narrows the gap further.
Types of quantization
Several approaches exist, differing in when quantization happens and what gets quantized. The right choice depends on whether you are serving a model, training it, or running it on a device with limited memory, and on the hardware and serving software available:
- Post-training quantization (PTQ): quantize a trained model directly, optionally using a small calibration dataset
- Quantization-aware training (QAT): simulate low precision during training so the model learns to tolerate it
- Weight-only quantization: compress weights while computing in higher precision, common for LLMs
- Weight and activation quantization, such as INT8 or FP8, for faster matrix math on supporting hardware
- Popular LLM formats and methods: GGUF for llama.cpp and Ollama, GPTQ, AWQ and bitsandbytes
Why quantization matters for AI products
Inference cost is dominated by memory and memory bandwidth. A smaller model fits on cheaper GPUs, serves more concurrent users per GPU and often generates tokens faster, which lowers both cost and latency. Quantization is how teams run capable open models on a single GPU in their own cloud account, keeping sensitive data inside their environment.
It also brings AI to the edge. Quantized models run on laptops, phones and embedded devices, enabling offline assistants, on-device transcription and privacy-preserving features. Combined with small language models, quantization makes useful AI possible where a data center model would be too slow, too expensive or simply unreachable.
Accuracy trade-offs and how to evaluate
Do not assume a quantized model is as good as the original for your use case. Quality loss tends to show first on tasks that need precise reasoning, long contexts, code or less common languages. Benchmark original and quantized versions on the same task-specific test set, compare latency and cost, and choose the smallest precision that still meets your quality bar.
Nexzem evaluates quantized open models alongside hosted APIs when clients need lower cost, data residency or on-device AI, and deploys them with serving stacks such as vLLM, SGLang or Ollama depending on scale. Our AI inference explainer covers how serving fits into the picture.