Why inference speed matters
For an AI product, a model is only as useful as the experience and economics around it. Slow generation increases user drop-off, makes voice and agent workflows feel broken, and raises the cost of every request. For Indian startups, the challenge is sharper: products may need to serve multiple Indic languages, operate across uneven network conditions, and reach customers at price points that leave little room for inefficient compute.
Optimizing large language model inference speed starts with defining the target rather than chasing a single benchmark. Track time to first token (TTFT), time per output token, end-to-end latency, throughput, error rate, and cost per request. Measure p50 and p95 values separately; a system that is fast on average but unreliable for the slowest users is not production-ready.
Set a useful performance budget
Before changing the model, establish budgets for each product path. A retrieval chatbot may prioritise TTFT, while document summarisation may tolerate a longer wait if it processes large batches. Voice applications need consistent streaming and short gaps between tokens. Agentic systems must also account for multiple model calls, tool execution, retrieval, and network overhead.
Create a representative test set containing short and long prompts, different output lengths, concurrent users, and the Indian languages your product supports. Include Devanagari, Tamil, Bengali, and code-mixed inputs where relevant. Tokenisation can vary significantly across languages, so comparing requests by character count alone can produce misleading conclusions. Keep quality checks alongside speed tests: factual accuracy, instruction following, refusal behaviour, and Indic-language fluency should not regress unnoticed.
Optimise the model first
Choose the smallest model that meets the quality bar
A larger model is not automatically better for every task. Test smaller instruction-tuned or distilled models for classification, routing, extraction, FAQ answering, and constrained generation. A two-stage design—small model for routine requests and a larger model for difficult cases—often improves both latency and cost.
If your use case is Hindi or another low-resource language, inspect the model’s actual performance instead of relying on English benchmarks. Guidance on open-source small language models for Hindi can help teams shortlist candidates, while fine-tuning Llama for Indian regional languages is useful when a general model misses domain terminology or local language patterns.
Quantise with a measured quality test
Quantisation reduces weight precision, commonly from 16-bit floating point to 8-bit or 4-bit formats. It can lower memory use and improve throughput, particularly when the model otherwise fails to fit comfortably on available GPU memory. Compare formats such as AWQ, GPTQ, and GPTQ-like or native low-bit implementations using your own evaluation set.
Do not treat quantisation as a free speed upgrade. Some kernels are faster only on specific GPUs, and aggressive quantisation can damage reasoning, long-context accuracy, or language quality. Record memory consumption, TTFT, decode speed, and quality for every candidate before selecting a production configuration.
Reduce unnecessary context and output
Prompt length is a major driver of prefill latency and memory use. Remove duplicated instructions, trim conversation history, retrieve fewer but better passages, and summarise old turns. Put stable instructions in a reusable prefix where the serving stack supports prefix caching. Set practical output limits and use structured responses when an application needs fields rather than prose.
For repetitive or low-variance requests, semantic caching can return a validated previous answer. Cache only where freshness, privacy, and user-specific context permit it; never allow one customer’s sensitive response to be reused for another.
Improve the serving runtime
Use continuous batching and efficient attention
Static batching waits for a group of requests and can leave GPU capacity idle. Continuous or in-flight batching admits new requests as others finish, improving utilisation for chat workloads with variable prompt and output lengths. Modern serving engines such as vLLM, SGLang, and TensorRT-LLM can provide paged attention, scheduling, and kernel optimisations, but the best choice depends on the model architecture and hardware.
KV-cache management is equally important. The cache stores attention states for previous tokens, speeding generation but consuming substantial memory. Configure cache limits, eviction behaviour, prefix reuse, and maximum context length deliberately. A runaway context window can reduce concurrency even when average requests are short.
Consider speculative decoding
Speculative decoding uses a smaller draft model to propose tokens that a larger model verifies in groups. When the draft model is well matched to the target model and outputs are predictable, this can increase decode speed without changing final quality. Benchmark it on real prompts: verification overhead may outweigh gains for short answers or difficult reasoning tasks.
Stream responses and isolate slow work
Streaming does not reduce total computation, but it improves perceived latency by showing useful output as soon as generation begins. Send tokens through a stable API path, handle client disconnects, and cap queues so one traffic spike does not exhaust the service. Move retrieval, document parsing, and tool calls into separately measured stages. This makes it easier to find whether the bottleneck is the model, the database, or the application layer.
Match hardware to the workload
GPU selection should consider memory capacity, memory bandwidth, supported kernels, power costs, and availability—not just theoretical FLOPS. A smaller model on a cost-efficient GPU may outperform a larger model with poor utilisation. For on-premise or edge deployments, AI model optimisation for mobile devices offers relevant principles around quantisation, memory limits, and offline execution.
Use multi-GPU serving only when the model or traffic justifies its communication overhead. Tensor parallelism can make a model fit, but interconnect speed matters. For high-throughput workloads, replicate a model across devices and route requests between replicas. For regional products, place inference near users where practical, while accounting for data-residency requirements and the availability of Indian cloud regions.
Build a repeatable benchmark
A credible benchmark should report:
- TTFT and p50/p95 end-to-end latency
- Output tokens per second and requests per second
- Prompt and completion token counts
- GPU memory use, utilisation, and power draw
- Cost per million input and output tokens
- Quality scores by language, task, and prompt length
- Performance under realistic concurrency and failure conditions
Warm up the runtime before measuring, separate prefill from decode, and test both steady-state and burst traffic. Compare the same prompts, sampling settings, context limits, and hardware. Store benchmark results in version control so a model, driver, CUDA, or serving-engine upgrade can be reviewed like any other production change.
A practical optimisation sequence
Start with profiling, not guesswork. First remove unnecessary context and output, then test a smaller or distilled model. Next evaluate quantisation, continuous batching, KV-cache settings, and an appropriate serving engine. Only after those changes should you scale hardware or introduce speculative decoding. Re-run quality and safety evaluations at each step.
For teams deploying in constrained environments, how to deploy large language models locally covers the operational trade-offs around privacy, hardware, updates, and offline access. If your application handles Indic text, also account for the data and evaluation layer described in low-resource language datasets for AI training in India. Better data and shorter, cleaner prompts often deliver more dependable gains than infrastructure changes alone.
Common mistakes to avoid
- Optimising average latency while ignoring p95 latency and queueing
- Comparing models with different prompt lengths or output limits
- Assuming 4-bit quantisation is faster on every GPU
- Allowing unbounded conversation history or retrieval context
- Measuring generated text quality only in English
- Scaling hardware before fixing batching and memory pressure
- Treating streaming as a substitute for lower compute latency
- Sharing caches without strict tenant and privacy controls
Conclusion
Fast LLM inference is a systems problem spanning model choice, tokenisation, prompts, kernels, scheduling, hardware, networking, and product design. Indian builders should optimise for the complete user journey: reliable first-token latency, consistent Indic-language quality, predictable cost, and graceful behaviour during traffic spikes. Establish a representative benchmark, make one change at a time, and keep quality and safety gates in the deployment pipeline.