What inference speed really measures
Inference speed is the time and capacity required to turn an input into a model output. Treating it as one number leads to poor engineering decisions. A production system should track:
- Time to first token (TTFT): important for conversational AI and streaming interfaces.
- Inter-token latency: how quickly a generative model continues responding.
- End-to-end latency: includes networking, authentication, retrieval, preprocessing, inference, postprocessing, and response delivery.
- Throughput: requests, images, tokens, or records processed per second.
- Tail latency: p95 and p99 response times, which reveal the experience of users during traffic spikes.
- Cost per request: useful when comparing a cloud GPU, CPU deployment, API provider, or an Indian edge installation.
A model can have excellent benchmark latency and still feel slow because tokenisation, database calls, retrieval, or serial Python code dominates the request path. Measure the complete service before optimising the model.
Start with code understanding, not guesswork
Code understanding means being able to trace an input from the API boundary to the final prediction and identify which operations consume time, memory, and accelerator capacity. Build a simple request map:
1. Parse and validate the request.
2. Load or retrieve documents, images, features, or prompts.
3. Transform data into tensors or model inputs.
4. Execute the model.
5. Decode, rank, filter, or postprocess the output.
6. Record metrics and return the response.
Instrument every stage with a request ID and monotonic timestamps. For language models, record prompt tokens, generated tokens, context length, batching behaviour, and cache hits. For vision or tabular systems, record input dimensions and preprocessing time. This makes it easier to distinguish a slow model from a slow data path.
Teams building AI-assisted development workflows should also keep inference code small and documented. Guidance on documenting open-source AI codebases is directly relevant: clear module boundaries make profiling safer and reduce regressions during optimisation.
The highest-impact optimisation levers
Reduce unnecessary work
The fastest operation is the one you do not perform. Remove duplicate embeddings, repeated database queries, oversized prompts, unused output fields, and redundant conversions between Python, NumPy, and tensor formats. Cache stable results such as embeddings, classification outputs, and retrieved context where correctness permits.
For retrieval-augmented generation, limit the number of retrieved chunks, deduplicate near-identical passages, and set a maximum context budget. A longer prompt increases both prefill time and memory use. Cache tokenised system prompts when the serving stack supports prefix caching.
Optimise the data pipeline
Keep preprocessing close to the execution device and avoid per-item Python loops. Prefer vectorised operations, pinned memory for transfers, and parallel data loading when profiling shows input starvation. Convert data once, use consistent tensor layouts, and avoid sending small batches through a network boundary one at a time.
Asynchronous request handling can improve throughput, but it does not automatically reduce single-request latency. Use bounded queues, timeouts, cancellation, and back-pressure. An unbounded queue may turn a short overload into minutes of waiting.
Choose the right model and precision
A smaller model that meets the quality target often beats a larger model with aggressive optimisation. Evaluate distillation, structured pruning, smaller architectures, and task-specific fine-tuning. Quantisation from FP32 to FP16, BF16, INT8, or lower precision can reduce memory traffic and improve throughput, but test accuracy on representative Indian languages, accents, scripts, and domain data.
Do not assume quantisation is harmless. Compare quality, tail latency, cold-start time, and cost—not only tokens per second. Keep a fixed evaluation set and a rollback path for production.
Use a suitable runtime
Export and serve models with runtimes that match the hardware and workload, such as ONNX Runtime, TensorRT, OpenVINO, vendor SDKs, or a framework’s compiled execution path. For generative models, use serving systems that support continuous batching, paged attention, prefix caching, and streaming where appropriate.
Compilation can introduce startup cost and unsupported operators. Benchmark warm and cold paths separately. If you are deploying near users or devices, the custom silicon for edge AI inference guide provides useful context on when specialised hardware is justified.
Hardware and deployment decisions in India
Hardware selection should follow a measured workload. CPUs can be cost-effective for small models, classical ML, and low-volume services. GPUs are valuable for parallel tensor workloads, while edge accelerators help when connectivity, privacy, or predictable latency matters. Consider memory capacity, interconnect bandwidth, availability, power, cooling, and procurement lead times—not just peak TOPS or FLOPS.
For Indian deployments, also account for intermittent connectivity, regional traffic patterns, data-residency requirements, and the cost of moving data between zones. A hybrid design may keep sensitive preprocessing or retrieval in-country while routing selected workloads to an external accelerator. Autoscaling should include warm capacity if cold starts would breach your service-level objective.
A practical benchmarking method
Create a repeatable benchmark before changing code:
- Use production-like inputs and a fixed representative dataset.
- Test cold start, warm single-request latency, and concurrent traffic.
- Report median, p95, and p99 latency, plus throughput and cost.
- Separate prefill, decoding, preprocessing, network, and postprocessing time.
- Test realistic prompt lengths, output lengths, batch sizes, and failure cases.
- Compare quality and safety metrics alongside performance.
Profile with framework profilers, system metrics, distributed traces, and GPU or CPU utilisation tools. Low accelerator utilisation may indicate small batches, synchronisation points, data-transfer overhead, or unsupported operations falling back to the CPU. High utilisation does not prove the service is healthy if memory pressure causes tail-latency spikes.
Production checklist
Before release, verify that:
- Model weights and tokenisers are loaded once, not per request.
- Inference runs in evaluation mode and disables unnecessary gradients.
- Batching has a maximum wait time as well as a maximum batch size.
- Inputs have explicit limits for tokens, dimensions, and payload size.
- Timeouts, retries, circuit breakers, and graceful degradation are implemented.
- Metrics cover latency, throughput, errors, queue depth, cache hit rate, and cost.
- Model, runtime, hardware, and configuration versions are recorded.
- A canary or shadow deployment protects against quality regressions.
For teams using AI to improve their own engineering process, automated production-grade code reviews can help catch accidental synchronous calls, missing limits, and inefficient data handling—but performance findings still need real profiling.
What to prioritise first
Start with the largest measured contributor to end-to-end latency. If retrieval takes 60% of the request, quantising the model will not solve the problem. If decoding dominates, reduce output length, improve batching, or select a faster serving runtime. If India-wide traffic creates tail spikes, redesign capacity and routing before changing model architecture.
A useful optimisation cycle is: measure, change one variable, validate quality, benchmark under load, and document the result. This discipline turns inference speed from a vague aspiration into an engineering target that builders can manage, forecast, and improve.