0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how much memory does quantization save

How Much Memory Does Quantization Save? A Practical Guide

  1. aigi

    Quantization reduces the numerical precision used to store and compute model weights, activations, or both. The headline arithmetic is simple: moving from 32-bit floating point (FP32) to 8-bit integers reduces the raw storage for each value by 75%; moving to 4-bit reduces it by 87.5%. The practical answer to how much memory quantization saves is more nuanced because runtime memory also includes temporary activations, the KV cache for language models, metadata, working buffers, and sometimes unquantized layers.

    For Indian builders deploying models on laptops, phones, shared GPUs, or cost-sensitive cloud instances, that distinction matters. A model that fits in storage may still fail during inference if its KV cache or workspace exceeds available RAM or VRAM.

    The basic memory calculation

    For a model with N parameters, the raw weight memory is approximately:

    memory in bytes = N × bits per parameter ÷ 8

    Ignoring metadata and framework overhead, a 7-billion-parameter model requires roughly:

    • FP32: 28 GB
    • FP16 or BF16: 14 GB
    • INT8: 7 GB
    • INT4: 3.5 GB

    Compared with FP32, the theoretical savings are:

    • FP16/BF16: 50%
    • INT8: 75%
    • INT4: 87.5%

    Compared with FP16, INT8 saves 50%, while INT4 saves 75%. These figures describe weights only, not the complete inference footprint. For the underlying concepts and deployment trade-offs, see what model quantization is.

    What the headline percentages leave out

    Quantization overhead

    Most quantized formats store scale factors, zero points, group metadata, or lookup tables. A nominal 4-bit model may therefore use more than 4 bits per parameter. Group-wise formats often deliver an effective rate closer to 4.1–4.8 bits per parameter, depending on the implementation.

    Some layers are deliberately kept at 8-bit or 16-bit precision because they are sensitive to rounding. Embeddings, output heads, normalisation layers, and attention components may remain higher precision. As a result, the complete model rarely achieves the exact theoretical percentage.

    Activations and working memory

    Weights are only one part of a deployed model’s memory use. In vision and speech systems, large intermediate activations can dominate peak usage. Quantising activations may reduce this cost, but it requires compatible kernels and calibration. A weight-only 4-bit model can still need substantial FP16 workspace during inference.

    KV cache in LLMs

    For autoregressive language models, the key-value cache grows with the conversation or generated sequence. Long prompts, multiple concurrent users, and large context windows can make the KV cache larger than the quantised weights. Weight quantisation does not automatically quantise this cache.

    This is why a 7B model with 4-bit weights may fit comfortably on a consumer GPU for a short prompt but run out of memory during a long-context, multi-user workload. If you are building an agent, separate model memory from application memory; an overview of AI systems with persistent memory can help clarify that architecture.

    A more realistic example

    Suppose a 7B model is converted from FP16 to a 4-bit format. Its FP16 weights require about 14 GB. The theoretical 4-bit footprint is 3.5 GB, but a production file might occupy approximately 4–5 GB after scales, metadata, and unquantised components.

    At runtime, you may also need:

    • 0.5–2 GB for the KV cache, depending on context length and batch size
    • additional memory for temporary buffers and kernels
    • memory for the tokenizer, runtime, and operating system
    • more capacity if serving several requests concurrently

    The result could be 5–8 GB or more, rather than exactly 3.5 GB. Always measure peak resident memory or VRAM, not just the model file size.

    Which quantisation level should you choose?

    8-bit quantisation

    INT8 is usually the conservative choice when accuracy and predictable behaviour matter. It often preserves quality well, works with mature CPU and accelerator kernels, and cuts weight memory by roughly 75% versus FP32. It is suitable for production classifiers, speech pipelines, and LLM serving where hardware supports efficient integer arithmetic.

    4-bit quantisation

    INT4 is attractive when GPU memory is the main constraint. It can reduce FP32 weight storage by about 87.5% and FP16 storage by about 75%. Quality depends heavily on the model, calibration data, group size, and quantisation method. Evaluate on your actual Indian-language, domain-specific, and safety-critical test sets rather than relying on a generic benchmark.

    Mixed precision

    Mixed-precision quantisation keeps sensitive layers at 8-bit or 16-bit while compressing the rest more aggressively. This frequently offers a better accuracy-to-memory balance than applying one format uniformly. For LLMs, compare formats and runtimes using the same prompt set, context lengths, and throughput target. See which quantization format is best for Transformers for the relevant format-level decisions.

    PTQ versus QAT

    Post-training quantisation (PTQ) converts an already-trained model. It is fast, inexpensive, and often the right first experiment. Calibration with representative data can materially improve activation quantisation, especially for multilingual or specialised workloads. A practical guide to post-training quantization covers the workflow.

    Quantisation-aware training (QAT) simulates low-precision effects during training so the model can adapt. It costs more compute and engineering time but may preserve accuracy when PTQ causes unacceptable degradation. QAT is more compelling for high-volume deployments where a small quality improvement justifies retraining.

    How to measure savings before deployment

    Use a repeatable benchmark rather than quoting theoretical ratios:

    • record the original model’s file size and peak RAM or VRAM
    • quantise weights, then measure the new file size
    • test representative prompt lengths, image sizes, or audio durations
    • measure peak memory, latency, throughput, and energy where relevant
    • evaluate accuracy, hallucination rate, language coverage, and task-specific failures
    • test concurrency and failure behaviour under realistic load

    For low-memory phones, edge devices, and modest Indian startup infrastructure, start with the smallest format that meets your quality threshold. The low-memory device optimisation guide provides a broader deployment checklist.

    Bottom line

    Quantisation can save 50% of weight memory with FP16, about 75% with INT8, and up to 87.5% with INT4 when compared with FP32. Real end-to-end savings are lower when metadata, unquantised layers, activations, KV cache, and runtime buffers are included. Treat the percentages as a planning estimate, then benchmark the complete workload on the hardware you will actually use.

    FAQ

    Does 4-bit quantisation always reduce total memory by 87.5%?
    No. That is the ideal weight-only reduction from FP32. Metadata, higher-precision layers, KV cache, activations, and runtime overhead reduce the realised saving.

    Does quantisation make inference faster?
    It can, but not automatically. Speed depends on kernels, memory bandwidth, batch size, and hardware support. A smaller model may reduce memory traffic while an unsupported format may add conversion overhead.

    Is quantisation only useful for LLMs?
    No. Computer vision, speech, recommendation, and multimodal models can benefit, particularly on edge hardware. The best precision depends on the model architecture and available operators.

    Should I quantise the KV cache?
    For long-context or high-concurrency LLM serving, KV-cache quantisation can be valuable, but it may affect quality and requires runtime support. Test it separately from weight quantisation.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.