What LLM inference speed actually measures
LLM inference speed is not one number. A production team should track several measures because each describes a different user or infrastructure outcome:
- Time to first token (TTFT): how long a user waits before streaming begins.
- Inter-token latency (ITL): the gap between generated tokens after the first one.
- Time to last token (TTLT): total time from request arrival to completion.
- Tokens per second: useful for comparing generation performance, but incomplete without latency context.
- Requests per second and tokens per minute: measures of system throughput.
- p50, p95, and p99 latency: the experience of typical and worst-served users.
For a customer-support assistant, TTFT and p95 latency may matter most. For document extraction, batch throughput and cost per million tokens may be more important. Define the target before changing the model or buying hardware.
A useful service-level objective might be: *p95 TTFT below 800 ms, p95 completion below six seconds, and a fixed cost per resolved conversation*. This is more actionable than promising that a model will be “fast.”
Start with the workload, not the model
Profile traffic before optimising. Record prompt length, output length, concurrency, arrival patterns, context-window usage, and the percentage of requests that use tools or retrieval. Indian products often need to handle uneven traffic: predictable daytime peaks, campaign-driven bursts, and users on high-latency mobile networks. Separate model latency from network, retrieval, queueing, and application overhead.
Keep a representative benchmark set containing short chat requests, long retrieved contexts, multilingual prompts, structured JSON outputs, and failure or retry cases. Hindi, Tamil, Bengali, and code-switched inputs should be included when they reflect your users. A benchmark made only of short English prompts will produce misleading capacity estimates.
When the bottleneck is application orchestration rather than generation, full-stack AI engineering best practices can help teams optimise the complete request path instead of focusing narrowly on GPU utilisation.
The highest-impact model optimisations
Choose the smallest model that meets the quality bar
A larger model is not automatically better for production. Compare candidate models on task accuracy, refusal behaviour, structured-output reliability, multilingual performance, and latency. A smaller instruction model with targeted retrieval or fine-tuning can outperform a general model for a narrow workflow.
Use routing when the workload supports it: send simple classification, rewriting, and FAQ requests to a smaller model, while reserving a larger model for ambiguous or high-risk cases. Monitor quality by route so cost and speed improvements do not hide regressions.
Quantise carefully
Quantisation reduces weight precision, often lowering memory use and improving throughput. Common choices include FP16 or BF16 for broad compatibility, INT8 for a conservative reduction, and INT4 for more aggressive compression. The best option depends on the model, hardware, kernel support, and workload.
Do not evaluate quantisation only on a generic language benchmark. Test long-context retrieval, exact figures, tool calls, Indian-language prompts, and your most commercially important outputs. Measure quality loss alongside TTFT, tokens per second, memory consumption, and cost.
Use distillation or task-specific fine-tuning
Distillation can transfer the behaviour of a capable teacher into a smaller student model. It is particularly useful for stable tasks such as intent classification, ticket routing, extraction, and response drafting. Fine-tuning can also reduce the amount of prompt context needed for every request, improving both latency and input-token cost. Follow best practices for fine-tuning LLMs on custom data before treating fine-tuning as a shortcut.
Pruning is less universally beneficial than quantisation or distillation. Removing parameters only improves real-world speed when the serving runtime supports the resulting sparsity efficiently. Otherwise, a smaller dense model may be faster and easier to operate.
Improve serving efficiency
Separate prefill from decode
LLM serving has two distinct phases. Prefill processes the input prompt and is sensitive to prompt length and compute. Decode generates tokens one at a time and is often constrained by memory bandwidth and key-value cache access. A long retrieved context can therefore damage TTFT even when the final answer is short.
Measure both phases and optimise them differently. Reduce unnecessary context, deduplicate retrieved passages, cap conversation history, and summarise old turns. For decode-heavy workloads, use efficient kernels, suitable batching, and enough memory bandwidth.
Use continuous batching
Static batching waits for a group of requests to form. Continuous or in-flight batching admits new requests as others finish, keeping accelerators busier while limiting unnecessary waiting. This is usually more suitable for conversational traffic with variable prompt and output lengths.
Batching is a trade-off: larger batches can improve throughput but increase queueing and memory pressure. Set a maximum queue delay and benchmark under realistic concurrency. A system that reports excellent tokens per second but poor p95 TTFT is not optimised for interactive use.
Manage the KV cache
The key-value cache stores attention data from prior tokens and can consume substantial accelerator memory. Paged or block-based KV-cache management reduces fragmentation and allows more concurrent sequences. Prefix caching can also avoid recomputing repeated system prompts, policy text, or common document prefixes.
Track cache hit rate, eviction rate, memory usage, and the effect on tail latency. Cache only content that is stable and safe to reuse; tenant-specific or sensitive context requires strict isolation.
Stream responses and control output length
Streaming does not reduce total generation time, but it improves perceived responsiveness by exposing the first token sooner. Pair it with output-token limits, concise system prompts, and structured schemas. Unbounded generation is a common source of avoidable latency and cost.
For agentic systems, cap tool retries and parallelise independent calls where safe. The principles in best practices for developing agentic workflows are directly relevant because orchestration delays can dominate model generation time.
Select hardware and inference software together
GPU choice should follow the model and workload. Compare accelerator memory, memory bandwidth, supported precision, interconnects, power limits, and hourly cost—not just theoretical FLOPS. Multi-GPU deployment can enable larger models, but communication overhead may reduce gains for smaller or decode-bound workloads.
For edge or constrained deployments, purpose-built silicon can be attractive when volume justifies engineering effort. Review custom silicon for edge AI inference when latency, power, privacy, or offline operation are central requirements.
Evaluate serving engines using the exact model and quantisation format you plan to deploy. Features worth testing include continuous batching, paged KV caches, prefix caching, speculative decoding, tensor parallelism, quantised kernels, and OpenAI-compatible APIs. A mature runtime with strong observability can outperform a theoretically faster stack that is difficult to tune.
Benchmark like a production team
Run load tests with realistic prompt and output distributions, not a single fixed prompt. Vary concurrency and report p50, p95, and p99 TTFT, total latency, throughput, error rate, GPU memory, power, and cost per million input and output tokens.
Use a controlled comparison:
- Keep model weights, prompt templates, and hardware constant when testing one change.
- Warm up the runtime before measuring.
- Include cold starts and autoscaling behaviour separately.
- Test failure, timeout, retry, and queue saturation paths.
- Record quality metrics beside performance metrics.
- Repeat tests across peak and off-peak traffic patterns.
A simple optimisation loop is: measure, identify the dominant constraint, change one variable, rerun the benchmark, and validate quality. This prevents teams from stacking several changes without knowing which one helped.
Keep latency and cost aligned
The fastest configuration is not always the most economical. Compare reserved capacity, on-demand GPUs, managed APIs, and regional deployments using actual utilisation. For Indian startups, a lower-cost accelerator with slightly higher latency may be the right choice for asynchronous workloads, while interactive applications may justify premium capacity.
Model routing, caching, prompt compression, and output limits often deliver better economics than hardware upgrades alone. For a structured approach, see this 2026 playbook for low-cost LLM inference for startups. Also account for data residency, support, observability, egress, and operational complexity when comparing providers.
A practical production checklist
Before launch, confirm that you can answer these questions:
- What are the TTFT, p95 completion latency, and throughput targets?
- Which prompt and output distributions were used in testing?
- Is the model quantised, and has quality been validated on Indian-language and domain-specific data?
- Are continuous batching, KV-cache management, and prefix caching enabled where appropriate?
- What happens when the accelerator is full or a provider times out?
- Can the system route simple requests to a smaller model?
- Are token usage, queue time, cache behaviour, and cost visible per tenant or product feature?
- Is there a fallback model or degraded mode for peak traffic?
LLM inference speed is ultimately a systems problem. The strongest deployments combine an appropriately sized model, efficient serving, realistic benchmarks, disciplined prompt and context management, and cost-aware capacity planning. Improve the slowest part of the request path first, then verify that the gain survives real traffic and real users.