Start with the workload, not the GPU
GPU capacity for LLMs means more than the number of accelerators in a server. A useful capacity plan accounts for GPU memory (VRAM), memory bandwidth, compute throughput, interconnect speed, storage, concurrency, latency targets, and operating cost. The right configuration for fine-tuning a 7B model is very different from the infrastructure needed to serve a multilingual model to thousands of users.
Before selecting hardware, write down four requirements:
- Model size and precision: parameter count, context length, and whether weights use FP16, BF16, FP8, INT8, or INT4.
- Workload type: pre-training, continued pre-training, fine-tuning, evaluation, batch inference, or interactive serving.
- Service target: requests per second, time to first token, tokens per second, uptime, and peak concurrency.
- Data constraints: residency, privacy, network access, and whether the workload must run in an Indian region or on-premises.
Teams training on Indian-language or domain-specific corpora should also plan data cleaning, tokenisation, and evaluation infrastructure. The guide to training LLMs on Indian datasets covers these choices in greater depth.
VRAM: the first sizing constraint
GPU memory must hold more than model weights. During training it may also hold gradients, optimiser states, activations, temporary buffers, and communication overhead. A rough weight-only estimate is:
VRAM for weights ≈ parameter count × bytes per parameter
For example, a 7B-parameter model requires approximately 14 GB for FP16 weights, 7 GB for INT8 weights, or 3.5 GB for INT4 weights. These figures are only a starting point. Inference requires room for the KV cache, which grows with context length and the number of concurrent sequences. Training usually requires several times the weight memory, depending on the optimiser, batch size, activation checkpointing, and precision.
Practical implications:
- A 24 GB GPU can be suitable for quantised inference, small-model fine-tuning, and parameter-efficient methods such as LoRA.
- A 48–80 GB accelerator provides more headroom for longer contexts, larger batches, full-parameter fine-tuning, and higher concurrency.
- Models that do not fit on one GPU can use tensor parallelism, pipeline parallelism, CPU offload, or quantisation—but each adds complexity or latency.
Do not treat quantisation as free. Lower precision reduces memory use and often improves throughput, but it can affect accuracy, calibration, tool use, and long-context behaviour. Validate the quantised model on representative Indian languages, accents, code, and domain terminology before production.
Training capacity versus inference capacity
Training and fine-tuning
Pre-training frontier-scale models is a distributed-systems project. It requires many accelerators, fast GPU-to-GPU communication, high-throughput storage, checkpoint management, fault recovery, and experienced operations staff. For most Indian startups, universities, and public-interest teams, training from scratch is not the economical first step.
More practical options include:
- Supervised fine-tuning on a strong open model.
- LoRA or QLoRA when full-parameter updates do not fit in memory.
- Continued pre-training on a carefully filtered domain corpus.
- Distillation to create a smaller model for local or edge deployment.
- Batch evaluation using rented GPUs instead of maintaining a permanent cluster.
For custom datasets, use the recommendations in best practices for fine-tuning LLMs, especially around validation splits, leakage checks, and reproducible experiments.
Inference and serving
Inference capacity is driven by token volume and user behaviour rather than parameter count alone. A model may fit comfortably in VRAM but still fail its service-level objective because generation is too slow or the KV cache exhausts memory.
Estimate capacity using:
- Input tokens per request.
- Output tokens per request.
- Requests per second at peak, not average demand.
- Maximum context length and concurrent sessions.
- Time-to-first-token and tokens-per-second targets.
- Whether requests can be batched or must be served interactively.
Continuous batching, paged attention, prefix caching, speculative decoding, and prompt caching can increase utilisation. However, benchmark the complete serving stack—model server, tokenizer, networking, logging, and safety checks—not just raw GPU throughput.
For teams constrained by hardware, deploying lightweight LLMs locally and deploying open-source LLMs for mobile apps offer useful paths to lower memory and network requirements.
Choosing hardware in 2026
GPU model names change quickly, so compare measurable characteristics rather than buying by brand or headline FLOPS. Prioritise:
- VRAM capacity: determines model fit, context length, and concurrency.
- Memory bandwidth: affects how quickly weights and KV-cache data move.
- Compute support: BF16, FP16, FP8, INT8, and INT4 support can materially change throughput.
- Interconnect: NVLink or equivalent high-speed links matter for multi-GPU workloads.
- Power and cooling: electricity, rack density, and thermal limits affect total cost.
- Software compatibility: drivers, CUDA or alternative runtimes, kernels, and serving frameworks.
Cloud GPUs are useful for bursty experiments and access to newer accelerators without capital expenditure. On-premises hardware can be cheaper for predictable, sustained utilisation, but only after accounting for procurement, networking, storage, maintenance, power, and staff. For sensitive faculty, healthcare, or government data, private LLM deployment for faculty research data provides a relevant architecture reference.
A practical capacity-planning workflow
1. Define the model and precision. Record parameters, layers, context window, quantisation method, and expected adapter size.
2. Calculate a memory budget. Include weights, KV cache, activations, runtime buffers, and at least 10–20% operational headroom.
3. Run a representative benchmark. Use real prompts, language mix, output lengths, concurrency, and safety tooling.
4. Measure the service objective. Track p50 and p95 latency, throughput, GPU utilisation, VRAM use, queue time, and errors.
5. Model peak demand. Include exam periods, campaign launches, batch jobs, and traffic spikes common in Indian deployments.
6. Choose scaling strategy. Scale up for simplicity, scale out for resilience and throughput, or route requests to smaller models when possible.
7. Add redundancy and observability. A single GPU may be adequate for a prototype but is not a production availability plan.
A cost model should use cost per million input and output tokens, not hourly GPU price alone. Include idle time, retries, storage, egress, orchestration, monitoring, and engineering effort. If external model APIs are part of the design, review common AI API cost blockers before committing to a provider.
Common mistakes to avoid
- Sizing by parameter count only: context length and concurrency can dominate memory use.
- Using average traffic: peak demand determines production capacity.
- Ignoring interconnects: multi-GPU communication can erase theoretical compute gains.
- Benchmarking synthetic prompts: real workloads expose token, language, and latency differences.
- Overtraining when fine-tuning is enough: smaller interventions reduce both GPU time and data risk.
- Skipping evaluation after quantisation: lower cost is not useful if quality or safety falls.
- Running one GPU without failover: plan maintenance and hardware failure from the start.
Bottom line
The best GPU capacity for LLMs is the smallest, reliable configuration that meets quality, latency, throughput, privacy, and cost targets. Start with a measured workload, reserve memory headroom, use parameter-efficient training where appropriate, and benchmark production-like traffic before purchasing hardware. For Indian builders, a hybrid approach—local or private infrastructure for sensitive workloads and cloud capacity for bursts—often delivers the strongest balance between control and speed.
FAQ
How much VRAM does a 7B LLM need?
Weight-only memory ranges from roughly 14 GB in FP16 to 3.5 GB in INT4. Inference also needs KV-cache and runtime headroom; fine-tuning needs substantially more.
Is one 24 GB GPU enough for LLM work?
It can support quantised inference, experimentation, and LoRA-style fine-tuning for smaller models. It is unlikely to support large full-parameter training or high-concurrency serving.
Should I buy GPUs or use the cloud?
Use cloud GPUs for uncertain or bursty demand. Consider purchasing hardware when utilisation is predictable and you can operate power, cooling, networking, security, and maintenance reliably.
How do I size GPUs for a chatbot?
Measure prompt and completion tokens, peak requests per second, concurrency, context length, latency targets, and batching efficiency. Then benchmark the exact model and serving stack under peak-like load.
Can quantisation replace more GPUs?
Often for inference, yes—but it may change quality and does not remove all compute or KV-cache requirements. Validate accuracy and safety on your actual workload.