Fast LLM applications are designed around latency budgets, not just powerful GPUs. A chat assistant, coding tool, voice agent, or agentic workflow must start responding quickly and maintain a predictable token rate under load. For Indian teams, the challenge also includes regional availability, bandwidth, data residency, and the economics of serving users across India and global markets.
This guide explains how to build a low latency LLM API as a production system. It focuses on the decisions that materially affect response time: model and hardware selection, prefill and decode optimization, request scheduling, caching, network placement, streaming, and measurement.
Start with a latency budget
Do not begin by choosing an inference engine. First define service-level targets for the actual user experience:
- Time to First Token (TTFT): Request arrival to the first streamed token. This captures queueing, prompt processing, and server overhead.
- Time Per Output Token (TPOT): Time between generated tokens after the first token. This determines reading and typing speed.
- End-to-end latency: Time until the complete response arrives.
- Requests per second and concurrency: The load at which your latency targets must hold.
- Tail latency: p95 and p99 matter more than averages for production reliability.
A useful budget might be p95 TTFT below 500 milliseconds for short chat prompts, with a stable 30–60 tokens per second during generation. Voice and interactive agent applications need tighter targets; teams building real-time voice agents with fast barge-in must account for speech recognition, orchestration, inference, and text-to-speech—not just model decoding.
Instrument each stage separately: DNS and TLS, gateway queueing, tokenization, prompt prefill, first-token generation, decode, serialization, and network delivery. Without this breakdown, a team may optimise GPU kernels while the real bottleneck is a slow retrieval call or an overloaded API gateway.
Choose the smallest model that meets the quality bar
Model choice is the highest-leverage latency decision. A smaller model generally reduces memory traffic, prompt processing time, and per-token decode cost. Establish an evaluation set from real Indian product traffic, including code-mixed English, Indian names, regional terminology, and the languages your users actually speak.
Compare models on quality, TTFT, TPOT, context length, memory footprint, and cost per generated token. A 7B or 14B model with strong prompting and retrieval can outperform a much larger model for a narrow support workflow. Route difficult requests to a larger model only when an evaluator or classifier identifies a genuine need. This model cascade usually produces better economics than serving every request with the largest available checkpoint.
For agentic systems, reduce unnecessary turns as well. A workflow that calls an LLM five times will often feel slower than one carefully structured call, even if each individual request is fast. Apply the same discipline when designing generative AI agents: constrain tool calls, parallelise independent work, and return intermediate progress to the client.
Quantize carefully and validate quality
Quantization lowers memory use and often improves throughput by reducing memory bandwidth. Common deployment choices include FP8, INT8, AWQ, GPTQ, and other weight-only formats. The right option depends on the GPU, serving engine, model architecture, and whether your workload is prefill- or decode-heavy.
Use a representative evaluation set before shipping. Measure factual accuracy, tool-call correctness, structured-output validity, multilingual performance, and refusal behaviour—not only perplexity. Four-bit quantization can be an excellent trade-off for many workloads, but sensitive legal, financial, or medical applications may need higher precision. Teams building a private AI chatbot for lawyers should also test whether quantization changes citation fidelity and long-context retrieval behaviour.
Use a production inference engine
A raw PyTorch loop is useful for experiments but rarely provides the scheduling and memory management required in production. Evaluate engines such as:
- vLLM: A strong default for open models, with PagedAttention, continuous batching, streaming, prefix caching, and an OpenAI-compatible API.
- TensorRT-LLM: Worth testing when you run NVIDIA hardware and need aggressively optimised kernels and compiled execution.
- Hugging Face TGI: A mature option with batching, streaming, and operational integrations.
- Specialised vendor runtimes: These may win for a particular GPU, model family, or latency target, but benchmark them against an open baseline.
Paged KV-cache management prevents the memory waste caused by reserving large contiguous blocks for every sequence. Continuous batching admits and removes requests at iteration boundaries, keeping the GPU busy without waiting for an entire batch to finish. These features improve throughput and can reduce queueing latency, but they must be tuned against your workload rather than enabled blindly.
Optimise prefill, decode, and caching separately
LLM serving has two distinct phases. Prefill processes the input prompt and is usually compute-intensive; decode generates one token at a time and is often constrained by memory bandwidth and KV-cache capacity.
For long prompts, use FlashAttention-compatible kernels, avoid sending redundant instructions, and trim retrieved documents before they reach the model. Prefix caching is valuable when many requests share a system prompt, policy text, or stable document prefix. KV caching avoids recomputing prior conversation tokens, but it consumes GPU memory, so configure eviction and maximum sequence lengths deliberately.
Speculative decoding can improve TPOT by using a smaller draft model to propose several tokens, which the target model verifies in parallel. It works best when the draft model closely matches the target distribution and requests generate sufficiently long outputs. Benchmark acceptance rates; a poorly matched draft model can add overhead rather than remove it.
Make the network path part of the design
Streaming is essential for perceived responsiveness. Use Server-Sent Events or an equivalent protocol to flush tokens as they are produced, send heartbeats during slow tool calls, and define clear cancellation behaviour. Streaming does not reduce total generation time, but it reduces the time users wait before seeing progress.
Place inference close to your users and dependent services. An Indian user routed to a US region may incur substantial round-trip delay before inference begins. Compare Mumbai, Hyderabad, Bengaluru, Singapore, and other viable regions using measured p50 and p95 network latency—not provider marketing claims. Keep retrieval databases, tool services, and the model regionally close where privacy and availability permit.
Put a lightweight gateway in front of the engine for authentication, quotas, request validation, retries, circuit breaking, and cancellation. Keep orchestration off the critical path where possible. Rust or Go can help at very high concurrency, but a carefully configured asynchronous Python gateway is sufficient for many products; moving languages will not fix poor batching or a slow downstream dependency.
Plan GPU capacity around concurrency
GPU memory must hold model weights, runtime workspace, activations, and KV cache. Estimate capacity using the maximum concurrent sequences, prompt length, output length, precision, and batching policy. Avoid sizing only for a single benchmark request.
Benchmark on the exact model and quantization format you intend to deploy. Record throughput and p95 TTFT/TPOT at increasing concurrency, then identify the point where queueing rises sharply. A smaller GPU fleet with predictable headroom is usually better than running every card near saturation. Add admission control and per-tenant limits so one long-context workload cannot consume the entire KV cache.
Build observability before optimisation
Track metrics by model, route, region, prompt-length bucket, output-length bucket, and status code:
- p50, p95, and p99 TTFT and end-to-end latency
- TPOT and output tokens per second
- queue time, prefill time, decode time, and cancellation rate
- cache hit rate, GPU memory utilisation, and KV-cache occupancy
- tokens per request, cost per successful response, and error rate
Maintain a replayable benchmark suite and run it after every model, driver, engine, or prompt change. Include load tests with realistic concurrency and long-tail prompts. For distributed agent systems, trace every model and tool call; the observability approach used for distributed systems with AI agents is directly relevant to finding hidden serial dependencies.
A practical rollout sequence
1. Define TTFT, TPOT, tail-latency, quality, and cost targets.
2. Establish a representative multilingual and domain-specific evaluation set.
3. Benchmark two or three model sizes in FP16/BF16 and suitable quantized formats.
4. Deploy vLLM or another production engine with streaming and continuous batching.
5. Add prefix/KV caching, prompt trimming, and speculative decoding only after baseline measurement.
6. Test regional placement and downstream service latency from Indian user locations.
7. Introduce load shedding, quotas, autoscaling, and rollback paths.
8. Optimise kernels or move runtimes only when traces show that software overhead is material.
A low-latency LLM API is not defined by one trick. It is the result of matching model quality to the task, keeping prompts and workflows compact, scheduling requests efficiently, streaming early, and measuring the complete path. For Indian builders, regional deployment and disciplined capacity planning can be as important as quantization or GPU choice.