0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · gpu capacity for llm

GPU Capacity for LLMs: Memory, Sizing and Cost Guide

  1. aigi

    Large language models are often limited by memory capacity and data movement, not raw GPU core count. A model may fit on paper yet fail at runtime because weights, activations, the key-value cache and framework overhead exceed available VRAM. For Indian builders, the right decision also depends on cloud availability, electricity, data residency, bandwidth and the cost of keeping a GPU busy.

    This guide explains how to estimate GPU capacity for LLM workloads, choose between local and cloud infrastructure, and improve utilisation without buying more hardware than the product needs.

    What GPU capacity means for an LLM

    GPU capacity combines several resources:

    • VRAM: The fastest constraint for model weights, activations, gradients and the KV cache.
    • Memory bandwidth: Determines how quickly the GPU can move model data during inference and training.
    • Compute throughput: Usually measured in FP16, BF16 or FP8 teraFLOPS; it affects how quickly matrix operations run.
    • Interconnect bandwidth: NVLink, PCIe and networking determine how efficiently multiple GPUs cooperate.
    • Host resources: CPU cores, system RAM, local SSD storage and network capacity can bottleneck data loading and serving.

    CUDA core count alone is a poor purchasing metric. A GPU with more cores but insufficient VRAM cannot load the model, while a high-memory accelerator with weak interconnects may struggle with distributed training.

    Estimate VRAM before choosing a GPU

    Start with the model’s parameter count and numerical precision. A useful first estimate for weights only is:

    • FP32: about 4 bytes per parameter
    • FP16 or BF16: about 2 bytes per parameter
    • INT8: about 1 byte per parameter
    • 4-bit quantisation: about 0.5 bytes per parameter, plus quantisation metadata

    For example, a 7-billion-parameter model requires roughly 14 GB for FP16 weights before runtime overhead. In practice, you also need room for the tokenizer, framework allocations, temporary tensors and the KV cache. A 16 GB GPU may therefore be uncomfortable for serving a 7B model at useful context lengths; 24 GB provides a safer starting point.

    Training requires much more memory. Full-parameter training stores weights, gradients and optimiser states, often reaching several times the size of the model. Adam-style training can require roughly 12–16 bytes per parameter or more, before activations. Activation memory also rises with sequence length and batch size.

    Use these figures as planning ranges, not guarantees. Measure the actual model with the intended framework, context window and batch size.

    Capacity needs by workload

    Inference

    Inference memory is driven by model weights and the KV cache. The cache grows with:

    • Number of concurrent requests
    • Context length and generated output length
    • Number of layers and attention heads
    • Precision used for cached keys and values

    A chatbot serving many users may need more VRAM for concurrency than for the model itself. Reduce pressure with quantisation, paged attention, continuous batching, shorter default context windows and request-level limits.

    Fine-tuning

    Parameter-efficient methods such as LoRA and QLoRA substantially reduce memory requirements because the base model remains frozen. They still need memory for activations and adapter training, so sequence length and micro-batch size matter. Gradient checkpointing, gradient accumulation and 4-bit loading can make a single high-memory workstation viable for smaller models.

    Pre-training or full fine-tuning

    Large-scale training typically requires multiple data-centre GPUs, fast storage and a reliable interconnect. The challenge is not simply adding cards: communication during gradient synchronisation can dominate runtime. Select a platform with strong intra-node bandwidth and benchmark the exact training stack before committing to a long reservation.

    A practical GPU-sizing workflow

    1. Define the workload: training, fine-tuning, batch inference or interactive serving.
    2. Record model requirements: parameters, precision, context length and architecture.
    3. Set service targets: latency, tokens per second, concurrent users and uptime.
    4. Calculate a memory budget: weights plus KV cache, activations, buffers and 15–25% headroom.
    5. Benchmark representative traffic: synthetic tests often hide long prompts and concurrency spikes.
    6. Compare total cost: include storage, egress, idle time, orchestration and engineering effort.
    7. Design a fallback: queue batch jobs, route overflow to another provider or use a smaller model.

    This workflow is particularly important when selecting among Indian cloud providers, regional data centres and international GPU clouds. Availability and hourly pricing can change quickly in 2026, so verify current quotes rather than relying on old hardware comparisons.

    Choosing hardware for Indian AI products

    For prototyping, a 16–24 GB consumer GPU can support quantised small models and selected fine-tuning workloads. A 48–80 GB data-centre GPU is more suitable for larger models, higher concurrency and production reliability. Multi-GPU systems make sense when the model or throughput target cannot fit on one card, but they add networking, cooling and operational complexity.

    Prioritise:

    • VRAM headroom over headline core count
    • BF16 or FP16 support for modern training stacks
    • Strong software support for PyTorch, inference servers and quantisation
    • Proven driver and container compatibility
    • Power draw and cooling requirements for on-premise deployments
    • Data-location, compliance and support requirements

    Startups building a product should also estimate utilisation. A dedicated GPU running at 10% utilisation may cost more than an API or shared cloud endpoint. Teams comparing build-versus-buy decisions can review affordable AI development tools for Indian startups before locking in infrastructure.

    Techniques that stretch GPU capacity

    • Quantisation: Use 8-bit or 4-bit weights where quality and tooling permit.
    • Mixed precision: Use BF16 or FP16 to reduce memory use and increase throughput.
    • Gradient checkpointing: Recompute selected activations instead of storing them.
    • LoRA or QLoRA: Train adapters rather than every model parameter.
    • Continuous batching: Keep inference hardware busy as requests arrive.
    • Speculative decoding: Use a smaller draft model to accelerate generation.
    • Model routing: Send simple requests to smaller models and complex tasks to larger ones.
    • Profiling: Track VRAM, GPU utilisation, tokens per second, latency and queue time separately.

    If your application includes document or image inputs, GPU sizing must include preprocessing and vision workloads. The practical trade-offs are similar to those discussed in evaluating OpenRouter vision models for video understanding.

    Common mistakes to avoid

    • Sizing only for model weights: Runtime memory and KV cache can be substantial.
    • Ignoring context length: Doubling context can significantly increase cache requirements.
    • Assuming multiple GPUs combine automatically: Sharding requires compatible software and communication paths.
    • Benchmarking with one request: Production concurrency changes latency and memory behaviour.
    • Buying the newest accelerator without a workload test: Software maturity and availability may matter more.
    • Forgetting storage and networking: Slow data pipelines leave expensive GPUs idle.

    For teams integrating an LLM into a broader product, infrastructure decisions should follow the application architecture. An enterprise AI app development platform in India may reduce custom operational work, while a specialised stack is preferable when you need tight control over latency, models or data.

    Frequently asked questions

    How much GPU memory does a 7B model need?
    A 7B model in FP16 needs about 14 GB for weights alone. Plan for additional memory, especially with long contexts or concurrent inference; 24 GB is a more practical starting point than 16 GB.

    Is one large GPU better than several smaller GPUs?
    Often, yes, for simpler deployment and lower communication overhead. Multiple GPUs are useful when memory or throughput requirements exceed one card, provided the interconnect and software stack are suitable.

    Can Indian startups use consumer GPUs?
    Yes, for prototyping, quantised inference and some parameter-efficient fine-tuning. Production teams should also evaluate reliability, cooling, warranty, power costs and the opportunity cost of maintaining hardware.

    Should I rent or buy GPUs?
    Rent when demand is uncertain, workloads are bursty or the team needs rapid experimentation. Buying can make sense with sustained utilisation, predictable workloads and suitable power and support infrastructure.

    What should I monitor?
    Track VRAM allocation, GPU utilisation, tokens per second, first-token latency, end-to-end latency, queue time, error rates and cost per million tokens. These metrics reveal whether the bottleneck is memory, compute, networking or application design.

    Build with a measured capacity plan

    GPU capacity for LLM workloads should be treated as an engineering budget, not a shopping-list specification. Begin with model size, precision, context and service targets; validate them under realistic traffic; then scale through quantisation, routing and batching before adding hardware. For early-stage teams, grants and non-dilutive support can help fund infrastructure experiments—explore AI Grants India for relevant opportunities.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.