0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · gpu capacity for inference

GPU Capacity for Inference: A Practical Sizing Guide

  1. aigi

    GPU capacity for inference determines whether an AI product feels responsive, scales predictably, and remains affordable. The right choice depends less on a GPU’s headline specifications than on the model, request pattern, latency target, context length, and deployment constraints.

    For Indian startups and engineering teams, this matters at every stage: a developer laptop, a rented cloud GPU, a shared inference server, and an edge device have different economics and operational trade-offs. This guide explains how to size capacity, benchmark it properly, and improve utilisation without compromising user experience.

    What GPU capacity means for inference

    Inference capacity is the amount of model-serving work a system can complete within a defined time and service-level target. It combines several resources:

    • VRAM: Holds model weights, activations, KV cache, runtime buffers, and sometimes multiple model replicas. Memory capacity is often the first hard limit.
    • Memory bandwidth: Controls how quickly weights and intermediate data move. Many language-model workloads are memory-bandwidth-bound, especially at low batch sizes.
    • Compute throughput: Tensor cores, CUDA cores, or equivalent accelerators determine how quickly matrix operations run, particularly for large batches and vision models.
    • Interconnect: PCIe, NVLink, and network links affect multi-GPU and disaggregated serving. A fast GPU cannot compensate for a slow path to CPU memory or another accelerator.
    • Host resources: CPU preprocessing, tokenisation, networking, storage, and system RAM can bottleneck a GPU.
    • Software stack: Kernels, drivers, quantisation support, compilation, and serving systems materially change real-world performance.

    A useful capacity statement is specific: “This deployment sustains 40 requests per second at p95 latency below 300 ms,” not “it uses a high-end GPU.”

    Start with the workload, not the hardware

    Before comparing GPU models, record the workload you need to serve:

    • Model architecture and parameter count
    • Precision: FP32, FP16, BF16, INT8, or INT4
    • Input and output token lengths for language models
    • Image, audio, or video resolution and sequence length
    • Target time to first token, time per output token, and end-to-end latency
    • Average and peak requests per second
    • Concurrent users and burst behaviour
    • Availability target, geographic location, and data-residency requirements

    For LLMs, weight memory is only the starting point. The serving system also needs space for the KV cache, which grows with concurrent sequences and context length. A model that fits in VRAM for one request may fail under production concurrency. Leave headroom for runtime allocations, batching, observability, and rolling updates rather than sizing to the exact theoretical minimum.

    Teams building multilingual products should test representative Indian-language traffic. Tokenisation can produce different sequence lengths across English, Hindi, Tamil, Bengali, and other languages, changing both memory use and latency. For local-language workloads, compare the economics of local LLM inference for Indian languages instead of assuming that an English benchmark predicts production behaviour.

    Estimating GPU memory requirements

    A practical memory estimate includes four major components:

    1. Model weights: Parameter count multiplied by bytes per parameter. Quantisation reduces this footprint, but scales, metadata, and dequantisation buffers add overhead.
    2. KV cache: Driven by context length, number of layers, attention dimensions, precision, and active sequences. This is often the dominant variable for conversational applications.
    3. Activations and temporary buffers: Depend on architecture, batch size, kernel implementation, and input dimensions.
    4. Serving overhead: Includes CUDA graphs, communication buffers, allocator fragmentation, and framework runtime requirements.

    Do not treat quantisation as a free speed upgrade. INT4 or INT8 can increase effective capacity and reduce cost, but accuracy, kernel availability, and hardware support must be validated on your own evaluation set. For a cost-sensitive deployment, compare quantised quality and latency with the economics discussed in how to reduce LLM inference costs for developers.

    Capacity planning: latency, throughput, and utilisation

    Inference has competing objectives. Increasing batch size usually improves throughput, but can increase queueing delay. A low-latency chatbot may need continuous batching with strict queue limits, while an offline document pipeline can use larger batches and accept slower individual responses.

    Track at least these metrics:

    • Time to first token (TTFT): Important for interactive LLM applications.
    • Time per output token: Determines how quickly a response streams.
    • End-to-end p50 and p95 latency: Exposes the user experience and tail behaviour.
    • Throughput: Requests per second or tokens per second at a defined concurrency.
    • GPU utilisation: Useful, but not sufficient; high utilisation can coexist with poor latency.
    • VRAM usage and KV-cache occupancy: Shows whether concurrency will cause failures.
    • Queue depth and rejection rate: Reveals saturation before users see outages.
    • Cost per request or per million tokens: Connects infrastructure to unit economics.

    Benchmark at expected and peak concurrency using production-shaped prompts. A single-request benchmark is not a capacity plan. Include cold starts, model loading, autoscaling delays, network transfer, and observability overhead. Run tests long enough to expose thermal throttling and allocator fragmentation.

    Selecting a deployment shape

    A single large GPU is simple and can offer strong performance for a model that needs shared memory. Multiple smaller GPUs may provide better availability or price flexibility, but model sharding introduces communication overhead and operational complexity. Replicated single-GPU workers are often easier to scale horizontally when the model fits comfortably on one device.

    Cloud GPUs are useful for variable demand, rapid experiments, and teams without hardware operations expertise. On-premise or colocated infrastructure can win for steady workloads, sensitive data, or predictable utilisation, but requires power, cooling, spares, and monitoring. For rural, industrial, or mobile use cases, edge deployment may be preferable; custom silicon for edge AI inference explains when specialised hardware can beat general-purpose GPUs.

    In India, include regional availability, egress charges, GST treatment, support response times, and data-location requirements in the comparison. The cheapest hourly GPU can become expensive if traffic crosses regions or if low utilisation persists. Use optimising LLM inference costs across regions when choosing between Indian and overseas capacity.

    Improving capacity without buying more GPUs

    Apply optimisation in this order:

    • Use an efficient serving engine: Select one with continuous batching, paged KV cache, fused kernels, streaming, and appropriate quantisation support.
    • Compile and profile: Remove CPU bottlenecks, pin data transfers, use asynchronous execution, and profile kernels rather than guessing.
    • Control context: Set practical input limits, summarise old conversation history, and avoid sending repeated system prompts where architecture permits.
    • Batch intelligently: Dynamic batching improves throughput, but enforce maximum queueing time for interactive requests.
    • Route by task: Send simple requests to smaller models and reserve larger GPUs for complex cases. Multi-model routing is covered in multi-model inference orchestration for Indian startups.
    • Cache safely: Cache embeddings, repeated retrieval results, or deterministic responses where privacy and freshness rules allow.
    • Separate workloads: Keep latency-sensitive traffic away from batch jobs and background evaluations.
    • Autoscale with useful signals: Combine queue depth, active sequences, tokens per second, and latency—not GPU utilisation alone.

    A practical sizing workflow

    1. Define the service-level objective and peak traffic.
    2. Measure real prompt, output, resolution, and language distributions.
    3. Calculate a conservative memory budget, including concurrency headroom.
    4. Benchmark two or three GPU and precision configurations.
    5. Test failure modes: traffic bursts, long contexts, model reloads, and GPU loss.
    6. Calculate cost per successful request at normal and peak utilisation.
    7. Deploy with dashboards, alerts, admission control, and a fallback model.
    8. Re-test after model, tokenizer, driver, or serving-engine changes.

    This approach prevents overbuying based on model size alone and prevents underprovisioning based on an optimistic demo. For startups comparing hosted APIs with self-hosting, low-cost AI inference for Indian startups provides a useful decision frame.

    FAQ

    Is more VRAM always better?
    No. More VRAM helps fit larger models and higher concurrency, but memory bandwidth, kernels, latency, and price may matter more for a model that already fits.

    How much GPU utilisation should I target?
    There is no universal number. Sustainable throughput with acceptable p95 latency and failure headroom matters more than a utilisation percentage.

    Can CPU inference replace GPUs?
    For small, quantised models and low-volume workloads, yes. GPUs generally become attractive when concurrency, latency, or model size increases.

    Should I buy or rent GPUs?
    Rent first while the workload is uncertain. Consider owned or colocated capacity when demand is steady enough to keep hardware productive and you can operate it reliably.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.