Why inference speed matters
Inference speed optimization is the work of reducing the time and resources required for a trained model to produce a result. It is not simply a race for the lowest benchmark latency. A production system must meet a defined service-level objective while preserving accuracy, reliability, privacy, and unit economics.
For an Indian AI startup, these trade-offs are especially important. A consumer application may serve users on variable mobile networks, while a healthcare or industrial system may need predictable response times on local infrastructure. Every unnecessary millisecond can increase abandonment, queue length, and compute cost. Before changing the model, define the target: for example, p95 latency below 300 ms, 99.5% request success, and a maximum cost per 1,000 predictions.
Teams building the surrounding serving layer can also use the principles in How to Build High-Performance AI Pipelines, particularly for data movement, queuing, and parallel execution.
Measure the right latency
Start with a production-like workload rather than a single local prediction. Record:
- End-to-end latency: network, authentication, preprocessing, inference, postprocessing, and response transfer.
- Model latency: time spent inside the runtime or accelerator.
- Queueing delay: time a request waits for a worker or batch.
- Throughput: requests or tokens processed per second.
- Tail latency: p95, p99, and p99.9 values, not only the average.
- Resource utilisation: CPU, GPU, memory, accelerator occupancy, and power.
Create a baseline for representative inputs. Image dimensions, prompt length, sequence length, language, and batch size can change performance substantially. For generative AI, measure time to first token, tokens per second, and total completion time separately. Add tracing around tokenisation, retrieval, model execution, and postprocessing so an apparently slow model does not hide an inefficient application layer. LLM Application Performance Monitoring in India offers a useful framework for observability across these components.
Optimise the model without losing quality
Choose an architecture that fits the job
The fastest model is often a smaller model selected for the actual task. Distillation can transfer behaviour from a large teacher to a compact student. Structured pruning removes channels, layers, or attention heads in a way that hardware can execute efficiently; unstructured sparsity may reduce parameter count without delivering a real-world speed-up unless the runtime supports it.
For mobile and edge deployments, evaluate compact architectures, input resolution, context length, and maximum output length together. A model with fewer parameters can still be slow if preprocessing or memory transfers dominate. See the AI model optimization guide for mobile devices when deploying on Android, iOS, or low-power edge hardware.
Quantise with a validation plan
Quantisation converts weights and, in some cases, activations from FP32 to FP16, BF16, INT8, or lower precision. FP16 and BF16 are common on modern GPUs; INT8 can deliver strong gains on CPUs and supported accelerators. Post-training quantisation is quick, while quantisation-aware training usually protects accuracy better for sensitive models.
Build a calibration set that reflects Indian languages, accents, image conditions, and real customer inputs. Compare not only aggregate accuracy but also class-level recall, safety checks, hallucination rates, and performance on low-frequency cases. Keep a full-precision fallback for inputs where quality drops beyond an agreed threshold.
Reduce unnecessary work
Cache tokenisation, embeddings, feature extraction, and repeated results where inputs are stable. For retrieval-augmented generation, cache retrieval results carefully and invalidate them when source content changes. Dynamic batching can improve accelerator utilisation, but excessive batch waiting harms interactive latency. Use separate policies for real-time, conversational, and offline jobs.
For streaming generation, stop early when the answer is complete, cap output length, and avoid sending unused context. Speculative decoding, prefix caching, continuous batching, and paged attention can materially improve LLM serving, but they should be tested against actual prompt distributions rather than copied from a benchmark.
Select the runtime and hardware deliberately
Export models to a deployment format supported by the target environment, then benchmark the complete graph. ONNX Runtime, TensorRT, OpenVINO, Core ML, TensorFlow Lite, and vendor-specific SDKs can fuse operations, select kernels, and manage memory more effectively than a training framework used unchanged in production. Building high-performance AI applications with open-source tools can help teams compare practical software stacks.
Hardware decisions should follow workload characteristics:
- CPU: suitable for small models, low traffic, preprocessing, and deployments where simplicity matters.
- GPU: effective for parallel workloads, large language models, and high throughput; account for memory and cold-start costs.
- NPU or mobile accelerator: useful for on-device workloads with strict power and privacy requirements.
- FPGA or custom silicon: justified when volume, latency, and power targets support the engineering investment.
Benchmark the same model across realistic concurrency levels. A GPU that wins at batch 32 may lose for single-request latency. Include model loading, autoscaling, thermal throttling, and cloud-region network time. For specialised edge products, Custom silicon for edge AI inference explains when dedicated hardware becomes commercially sensible.
Build a production optimisation loop
Optimisation should be a controlled engineering cycle:
1. Define an SLO and quality floor. Specify latency percentiles, throughput, error rate, accuracy, and cost.
2. Profile before changing code. Locate time spent in preprocessing, memory copies, kernels, queues, and postprocessing.
3. Change one variable at a time. Compare architecture, precision, runtime, and hardware using the same test set.
4. Run shadow traffic. Replay anonymised production requests without affecting users.
5. Canary the release. Monitor tail latency, quality, saturation, and cost by device, geography, and model version.
6. Automate regression checks. Block deployments that breach latency or accuracy thresholds.
In India, also account for data-residency requirements, intermittent connectivity, regional language coverage, and the cost difference between hosted inference and self-managed infrastructure. On-device inference can reduce network dependency and protect sensitive data, while a hybrid design can route difficult cases to a larger cloud model.
Common mistakes to avoid
- Optimising average latency while ignoring p99 queueing and cold starts.
- Reporting accelerator kernel time without measuring network and preprocessing overhead.
- Quantising on a narrow calibration set that excludes production languages or edge cases.
- Increasing batch size until interactive requests wait too long.
- Choosing a framework because of a benchmark without checking operator support and maintenance.
- Treating lower latency as automatically better when it increases error rates or cost.
- Forgetting observability after deployment, making regressions difficult to explain.
A practical checklist for Indian AI builders
Before launch, confirm that you can answer these questions:
- What are the p50, p95, and p99 latency targets for each request type?
- Which stage is the current bottleneck?
- What accuracy or safety metric must not decline?
- Does the selected precision work on the target hardware?
- How do concurrency, batch size, and input length affect cost?
- Can the service degrade gracefully when an accelerator or network link is unavailable?
- Are model versions, hardware, runtime, and benchmark data recorded for reproducibility?
Inference speed is a product, architecture, and operations concern—not a final tuning step. Measure the complete path, select a model and runtime for the workload, validate quality on representative Indian data, and roll out changes with observability. Done well, optimisation delivers faster experiences while improving the economics of serving AI at scale.