Why GPU capacity planning matters
GPU capacity for LLM inference is not simply the number of GPUs in a server. It is the usable combination of memory, memory bandwidth, compute, interconnect, and serving software required to meet a workload’s latency, throughput, availability, and cost targets.
That distinction matters for Indian startups and engineering teams. A GPU that can load a model may still fail in production when users send long prompts, requests arrive concurrently, or the application runs tool calls and retries. Conversely, buying the largest available accelerator can leave a low-volume product with an unnecessarily high cost base. A sound plan starts with workload measurements, then chooses the smallest configuration that meets the service-level objective.
Start with the workload, not the GPU
Define these variables before comparing hardware:
- Model: parameter count, architecture, context window, and whether it uses dense or mixture-of-experts layers.
- Precision: FP16, BF16, FP8, INT8, or 4-bit quantisation.
- Traffic: average and peak requests per second, daily volume, and concurrency.
- Tokens: average and p95 input and output tokens per request.
- Service targets: time to first token (TTFT), time per output token, end-to-end latency, and availability.
- Serving pattern: interactive chat, batch jobs, retrieval-augmented generation, agents, or embeddings.
A customer-support bot with short Indian-language queries has a very different profile from an agent that sends a 100,000-token context to a model and performs multiple calls. For multi-model applications, routing requests according to quality, latency, and price is often more valuable than adding GPUs; see this guide to multi-model inference orchestration for Indian startups.
Estimate VRAM requirements
VRAM must hold more than model weights. A practical estimate is:
Required VRAM ≈ weights + KV cache + runtime overhead + safety margin
For dense models, unquantised weights are approximately:
- FP16 or BF16: 2 bytes per parameter
- INT8: 1 byte per parameter
- 4-bit formats: roughly 0.5 bytes per parameter, plus scales and metadata
A 7-billion-parameter model therefore needs about 14 GB for FP16 weights before runtime overhead. A 70-billion-parameter model needs about 140 GB, so it generally requires quantisation, multiple GPUs, or both. These are planning estimates, not guaranteed footprints: implementation, tensor parallelism, kernels, and checkpoint format change the result.
The KV cache is frequently the hidden constraint. During autoregressive generation, the server stores attention keys and values for every active sequence. Its size grows with context length, concurrent sequences, layers, and hidden dimensions. Long-context workloads can exhaust VRAM even when the model weights fit comfortably. Measure memory at the concurrency and token lengths you actually expect rather than testing one short prompt.
Leave headroom for CUDA graphs, framework allocations, token buffers, batching, and traffic spikes. Running at 98% VRAM utilisation may look efficient but leaves little room for a longer request or a rolling deployment.
What GPU specifications actually affect inference?
Memory capacity and bandwidth
VRAM capacity determines whether the model and active KV cache fit. Memory bandwidth affects how quickly weights and cache data can be read, making it especially important for small-batch, memory-bound decoding. A newer accelerator with less total memory may be a poor choice if it forces aggressive quantisation or model sharding.
Tensor performance and precision support
Tensor cores and support for BF16, FP8, INT8, and 4-bit kernels can materially improve throughput. Compare benchmarks using the same model, quantisation, prompt length, output length, and serving engine. Peak teraFLOPS alone is not a useful production forecast.
Interconnect and multi-GPU scaling
When a model spans GPUs, tensor or pipeline parallelism introduces communication. High-bandwidth links can preserve performance; ordinary PCIe or network links may create a bottleneck. Multi-GPU inference is justified when the model cannot fit on one device, when throughput is high enough to amortise communication, or when availability requires redundancy. It is not automatically faster for every request.
Choose a deployment shape
- Single-GPU serving: simplest operations and often best latency for small and medium models.
- Replicated GPUs: run one model copy per device to increase concurrency and isolate failures.
- Sharded serving: split a large model across GPUs using tensor or pipeline parallelism.
- Reserved cloud capacity: predictable workloads can benefit from commitments, while bursty traffic may suit on-demand or serverless options.
- On-premises or colocation: attractive when utilisation is consistently high, but requires power, cooling, networking, and hardware support.
Indian teams should price the full system, not only hourly GPU rental. Include storage, egress, idle capacity, observability, engineering time, electricity where relevant, and the cost of keeping a second region or fallback provider. Regional latency and data-residency requirements can also make a slightly more expensive local deployment the better product decision. For a structured cost view, use optimizing LLM inference costs across regions.
Optimise before adding capacity
1. Quantise carefully. Test 8-bit and 4-bit variants against a representative evaluation set, including English and Indian languages where relevant. Lower precision can reduce memory and cost, but quality loss may affect safety, retrieval, or tool use.
2. Use continuous batching. Modern serving engines combine active requests dynamically, improving utilisation without waiting for a fixed batch to complete.
3. Separate prefill and decode. Prompt processing is compute-heavy, while token generation is often memory- and latency-sensitive. Disaggregated or specialised pools can help at scale.
4. Control context. Trim irrelevant history, cap retrieved documents, cache stable prefixes, and avoid sending duplicate tool output.
5. Use speculative decoding. A smaller draft model can accelerate generation when acceptance rates are high.
6. Keep a model ladder. Route simple requests to a smaller model and reserve the larger model for difficult cases. This approach often delivers better economics than running one large model for every query.
Teams comparing open models and local serving options may also benefit from the India open-source AI inference engines deployment guide. For a startup budget review, see low-cost LLM inference for startups.
Measure capacity with production metrics
Track metrics by model, GPU type, precision, prompt length, and traffic class:
- TTFT: responsiveness before the first generated token.
- Inter-token latency: smoothness of streaming output.
- p50, p95, and p99 latency: tail behaviour matters more than averages.
- Output tokens per second: generation capacity.
- Requests and tokens per second: useful throughput measures.
- GPU memory, utilisation, power, and temperature: hardware health and saturation.
- Queue time and rejection rate: evidence of insufficient capacity or poor admission control.
- Cost per million input and output tokens: the metric finance and product teams can act on.
Benchmark at peak-like concurrency with realistic token distributions. A single-user benchmark can hide queueing, KV-cache pressure, and tail latency. Run load tests after every model, quantisation, kernel, or serving-engine change.
A practical sizing workflow
1. Collect at least several days of request, token, and latency data—or create a realistic forecast for a new product.
2. Set p95 TTFT, p95 completion latency, availability, and cost targets.
3. Calculate weight memory and estimate KV-cache growth at peak concurrency.
4. Select two or three candidate GPU configurations, including a fallback option.
5. Benchmark with the exact model, quantisation, context distribution, and serving stack.
6. Add capacity for failures, deployments, and forecast error rather than targeting continuous saturation.
7. Reassess monthly as prompts, models, traffic, and cloud prices change.
Bottom line
The right GPU capacity for LLM inference is the configuration that meets real latency and throughput targets at an acceptable cost—not the one with the highest advertised compute. Start with tokens and concurrency, budget VRAM for the KV cache, validate quantisation, benchmark multi-GPU communication, and monitor cost per token. For teams under pressure to reduce spend, practical ways to reduce LLM inference costs for developers provide a useful next step.
FAQ
How much VRAM does an LLM need?
Estimate model parameters multiplied by bytes per parameter, then add KV cache, runtime overhead, and safety headroom. A 7B model may fit on a 16 GB GPU in a compressed format, while a 70B model usually needs substantial quantisation or multiple GPUs.
Is more GPU memory always better?
No. More memory improves model fit and concurrency, but performance also depends on bandwidth, precision kernels, serving software, and interconnects. Benchmark the complete stack.
Should I use one large GPU or several smaller GPUs?
Use one GPU when the model and expected cache fit and latency is the priority. Use several when the model cannot fit, throughput justifies parallelism, or redundancy is required. Account for communication overhead and operational complexity.
How often should capacity be reviewed?
Review after any major model or context-window change and at least monthly for a growing production service. Traffic mix and token counts can change capacity requirements faster than request volume alone suggests.