AI inference is often limited by memory before it is limited by raw compute. Model weights, activations, attention key-value (KV) caches, temporary buffers, and framework overhead can exhaust GPU or system RAM even when the processor has spare capacity. Memory optimized inference is the disciplined process of reducing that footprint while preserving the accuracy, latency, throughput, and reliability a product requires.
For Indian startups and research teams, this matters at several levels: cloud GPU prices can make an otherwise viable product uneconomical; edge deployments may have only a few gigabytes of RAM; and data-residency or offline requirements can rule out sending every request to a large hosted model. The objective is not simply to make a model smaller. It is to deliver the required result on a defined hardware target at a sustainable cost.
Start with a memory budget
Before changing the model, establish what must fit in memory and when. A useful budget separates:
- Weights: Parameters stored in FP32, FP16, BF16, INT8, INT4, or another format.
- Activations: Intermediate tensors created as layers execute.
- KV cache: Per-request attention state in autoregressive language models; this often dominates long-context serving.
- Runtime overhead: Allocators, graph metadata, CUDA or accelerator libraries, token buffers, and fragmentation.
- Concurrency headroom: Additional memory required when multiple requests run together.
Profile a representative workload rather than a single short prompt. Record peak allocated memory, reserved memory, tokens per second, time to first token, latency at the target concurrency, and out-of-memory failures. Include long prompts and generated outputs if the application supports them. A model that fits one request but fails at eight concurrent users is not production-ready.
The main techniques
Quantization
Quantization stores weights and, where supported, activations at lower numerical precision. Moving from FP16 to INT8 can roughly halve weight memory; INT4 can reduce it further, although actual savings depend on scales, metadata, kernels, and runtime implementation.
Choose the method based on your operating constraints:
- Post-training quantization is fast and inexpensive to test.
- Calibration-based quantization uses representative data to reduce accuracy loss.
- Quantization-aware training adapts the model during training and can protect quality for sensitive tasks.
- Weight-only quantization is useful when activations remain difficult to quantize, especially for large language models.
Do not judge quantization only by average benchmark accuracy. Test Indian languages, code-mixed prompts, domain terminology, numerals, and safety-sensitive outputs. For a Hindi voice product, for example, evaluate recognition quality on the target accents and noise conditions; the Hindi ASR low WER guide offers a relevant evaluation lens.
Pruning and structured sparsity
Pruning removes parameters or channels that contribute less to the output. Unstructured sparsity can reduce the number of stored non-zero weights, but it helps latency only when the hardware and kernels exploit sparse computation. Structured pruning—removing complete channels, heads, or layers—is usually easier to deploy because it produces a smaller dense model.
Apply pruning gradually, fine-tune after each major change, and compare quality against the original model. If your inference stack does not support sparse kernels, pruning may lower storage without improving runtime memory or latency.
Distillation and smaller architectures
Knowledge distillation trains a student model to reproduce a larger teacher’s useful behaviour. This can be more reliable than aggressively compressing a production model, particularly for classification, ranking, extraction, and narrow support workflows. Distil on examples that reflect real traffic, including difficult and low-frequency cases.
Architecture selection also matters. Smaller attention dimensions, grouped-query attention, efficient convolution blocks, and reduced sequence lengths can deliver structural savings. For edge deployments, pair model design with suitable hardware planning; the guide to custom silicon for edge AI inference explains why memory bandwidth and on-chip memory can matter as much as accelerator TOPS.
KV-cache and context management
For generative models, the KV cache grows with context length, layers, heads, and concurrent sequences. Practical controls include:
- Limiting maximum input and output tokens by product requirement.
- Using grouped-query or multi-query attention when supported.
- Evicting or compressing old context for long-running sessions.
- Sharing prefixes for repeated system prompts.
- Quantizing the KV cache where quality and kernels permit it.
- Batching requests with similar sequence lengths to reduce padding waste.
Long-term conversation memory should not be confused with the active KV cache. Store durable facts selectively outside the model and retrieve only what is relevant. For architecture patterns, see AI system memory for personalised LLMs and the guide to persistent AI memory loops.
Runtime and serving optimizations
Model compression is only one part of the result. Use an inference engine that supports the target accelerator, fused operators, memory-aware allocation, continuous batching, and paged KV-cache management where applicable. Exporting through ONNX, TensorRT, OpenVINO, or an accelerator-specific compiler can improve execution, but validate numerical equivalence and unsupported operations before committing to a stack.
Avoid unnecessary copies between CPU and GPU memory. Reuse buffers, preallocate predictable tensor shapes, and keep tokenization and post-processing from becoming hidden bottlenecks. When serving multiple models, measure the cost of loading, unloading, and fragmentation; a smaller model can still cause failures if the process retains stale allocations.
For startups, memory optimization should be tied to unit economics. Compare cost per 1,000 requests or per million tokens—not only hourly GPU price. The low-cost AI inference playbook for Indian startups provides a useful framework for comparing cloud, colocated, and edge options, while optimizing LLM inference costs across regions helps account for regional pricing and traffic placement.
A practical optimization workflow
1. Define acceptance thresholds: Set quality, peak memory, p95 latency, throughput, availability, and cost targets.
2. Profile the baseline: Test realistic prompts, sequence lengths, concurrency, and failure cases.
3. Apply the least risky change: Start with buffer reuse, batching, operator fusion, and context limits before changing model precision.
4. Quantize and validate: Compare task quality, calibration subsets, languages, and edge cases.
5. Test combinations: Quantization plus distillation or structured pruning may outperform any single technique.
6. Load-test the deployment: Include cold starts, autoscaling, sustained traffic, and memory fragmentation.
7. Canary release: Monitor quality drift, out-of-memory events, latency, and cost in production.
Keep an uncompressed reference model and a reproducible evaluation set. Optimization is an ongoing release activity, not a one-time conversion.
Common mistakes
- Treating parameter count as total runtime memory.
- Measuring only single-request latency.
- Assuming lower precision is automatically faster on every accelerator.
- Ignoring tokenizer, padding, and data-transfer overhead.
- Optimizing average accuracy while missing failures in Indian languages or domain-specific inputs.
- Setting context limits so high that concurrency becomes uneconomical.
- Choosing a framework before confirming kernel and hardware support.
Choosing the right target
Use quantization first when you need a quick reduction in weight memory and have a compatible runtime. Prefer distillation or a smaller architecture when predictable latency and high concurrency matter. Prioritize KV-cache controls for chat, agents, and long-context applications. For low-volume products, a modest model on a CPU or shared accelerator may beat a heavily optimized large model on total cost.
The best design is the smallest system that meets the product’s quality and service-level requirements. In 2026, that often means combining lower-precision weights, disciplined context management, efficient serving, and hardware-aware profiling rather than relying on one compression trick.