0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · which quantized models can run on cpu

Which Quantized Models Can Run on CPU? A Practical Guide

  1. aigi

    Quantization makes AI models smaller and cheaper to run by representing weights, activations, or both with fewer bits. The important question is not simply which quantized models can run on CPU, but which model, quantization scheme, runtime, and processor combination delivers acceptable latency and accuracy.

    For Indian startups, research teams, and public-sector deployments, CPU inference can be a sensible default. It avoids GPU procurement, works on ordinary cloud instances and on-premise servers, and can keep sensitive workloads close to users. A quantized model will not automatically be fast, however. Kernel support, memory bandwidth, instruction sets, context length, and preprocessing often matter as much as the nominal model size.

    The short answer

    Most modern model families have a CPU-compatible quantized version. Common examples include:

    • MobileNet, EfficientNet, ResNet, and SqueezeNet for image classification and edge vision.
    • DistilBERT, MiniLM, ALBERT, and other encoder models for classification, search, and extraction.
    • Whisper variants for speech recognition, particularly small and medium checkpoints with optimised runtimes.
    • Llama, Mistral, Gemma, Qwen, Phi, and similar small language models in GGUF or comparable 4-bit and 8-bit formats.
    • Embedding and reranking models quantized with ONNX Runtime, OpenVINO, or specialised vector-inference libraries.

    The best choice depends on the task. A 4-bit language model may be practical for a local assistant, while INT8 is often the safer option for production classification or vision workloads where predictable accuracy matters.

    Quantization formats that matter on CPU

    FP16 halves storage compared with FP32, but it is not always faster on CPUs. Many processors lack strong native FP16 arithmetic, so FP16 may primarily reduce memory use rather than improve latency.

    INT8 is the most established choice for CPU inference. It is widely supported by TensorFlow Lite, ONNX Runtime, OpenVINO, and vendor-optimised libraries. Post-training INT8 quantization is often sufficient for image classifiers and encoder-style NLP models. Quantization-aware training can preserve accuracy better when the model is sensitive to reduced precision.

    INT4 and other low-bit formats are especially popular for large language models. They reduce memory pressure enough to run models on desktops, laptops, and modest servers. GGUF files used with llama.cpp are a common local-deployment route. Low-bit inference can be very effective, but quality varies by calibration method, group size, CPU architecture, and runtime implementation.

    Dynamic quantization quantizes weights ahead of time while handling some activations at runtime. It is convenient for transformer encoders and CPU-based PyTorch workflows, although deployment support varies by operator and framework version.

    Model families that work well on CPU

    Computer vision models

    MobileNet and EfficientNet are strong starting points for camera, document, and retail applications. ResNet remains useful when compatibility and predictable accuracy matter, while SqueezeNet is suitable for very tight memory budgets. INT8 TensorFlow Lite or ONNX models can run efficiently on x86 servers, ARM boards, and many Android devices.

    For teams building a vision pipeline rather than selecting a single checkpoint, this guide to building computer vision models on GitHub is a useful companion. Measure the complete pipeline, including image decoding and resizing, not just neural-network execution.

    NLP and classical transformer models

    DistilBERT, MiniLM, ALBERT, and compact multilingual encoders are well suited to CPU classification, semantic search, named-entity recognition, and reranking. Dynamic or static INT8 quantization can substantially reduce latency and memory use. These models are generally easier to operate on CPU than generative language models because their output length is fixed and inference does not require token-by-token decoding.

    For Indian-language applications, benchmark the target languages separately. A model that performs well on English may degrade on Hindi, Marathi, Telugu, or code-mixed text. Teams working specifically with Hindi can compare compact options in this guide to open-source small language models for Hindi.

    Speech models

    Whisper and other speech-recognition models can run on CPU, especially small checkpoints with CTranslate2, faster-whisper, or other optimised backends. Real-time performance depends heavily on audio duration, beam size, sampling rate, and whether transcription is streamed. For Indian deployments, test accents, background noise, code-switching, and names of local places rather than relying on English-language benchmarks.

    Small and medium language models

    CPU inference is most practical with small language models, typically in the 1B–8B range, quantized to 4-bit or 8-bit formats. Llama, Mistral, Gemma, Qwen, and Phi families are commonly available through runtimes such as llama.cpp, Ollama, MLX on supported Apple hardware, and vendor-specific engines. A 7B model in 4-bit form may fit in roughly 4–6 GB of memory before runtime overhead, but actual requirements increase with context length and concurrent requests.

    If your objective is a private local assistant, review the architecture and operational considerations in how to deploy large language models locally. For a voice product, also compare generation latency with the requirements of a voice agent versus chatbot, since CPU text generation may be acceptable for asynchronous tasks but frustrating in live conversations.

    CPU runtimes and hardware considerations

    Choose the runtime before finalising the model. ONNX Runtime and OpenVINO are strong options for quantized vision and encoder models. TensorFlow Lite is practical for mobile and embedded deployments. llama.cpp and GGUF are widely used for local generative models. PyTorch remains valuable for experimentation, but a production CPU service may benefit from exporting to ONNX or another specialised backend.

    Check whether the target processor supports AVX2, AVX-512, VNNI, or ARM NEON. These instruction sets can materially change throughput. Intel and AMD server CPUs, Apple Silicon, and ARM devices may require different builds and tuning. Thread count is not the only variable: excessive threads can increase contention, power use, and tail latency.

    Memory is often the first bottleneck. Keep room for the operating system, tokenizer, key-value cache, application code, and multiple concurrent requests. For language models, reduce context length where the product allows it; the key-value cache can become larger than expected during long conversations.

    A practical benchmarking method

    Use representative workloads and record more than average tokens per second. A useful test includes:

    • Cold-start time and model-load time.
    • Peak resident memory and model file size.
    • Single-request latency at p50, p95, and p99.
    • Throughput at the expected concurrency level.
    • Prompt-processing speed and generation speed separately.
    • Accuracy, word error rate, F1, retrieval quality, or task-specific success.
    • Power draw and cost per 1,000 requests where relevant.

    Compare FP32, INT8, and 4-bit versions on the same machine and runtime. Keep tokenizers, prompts, input sizes, and thread settings consistent. A model that wins a synthetic benchmark may lose in production because network calls, document parsing, or database retrieval dominate total latency.

    Choosing the right quantized model

    Use INT8 when you need stable accuracy, broad framework support, and predictable production behaviour. Use 4-bit when memory is the main constraint and you are deploying a generative model locally. Prefer a smaller full-precision model over an aggressively quantized larger model if quality is sensitive or the CPU lacks optimised low-bit kernels.

    For an India-focused deployment, include offline operation, intermittent connectivity, multilingual evaluation, data residency, and device availability in the decision. A modest CPU model that works reliably in a district office or field device can be more valuable than a larger model that requires a GPU and continuous cloud access.

    Final checklist

    Before shipping, confirm that the model licence permits your use, the runtime supports every required operator, and the quantized checkpoint has been evaluated on your real data. Pin versions, record the quantization method, and keep a higher-precision fallback for regression testing. Quantization is a deployment strategy—not a substitute for profiling, quality evaluation, or sensible model selection.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.