0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · gpu for llms

GPU for LLMs: Choosing Hardware for Training and Inference

  1. aigi

    Start with the workload, not the GPU

    The best GPU for LLMs depends on what you are building. Training a foundation model, fine-tuning an open model, serving a multilingual chatbot, and running batch embeddings have very different hardware requirements. Buying the most powerful card available can leave an Indian startup with unnecessary capital expenditure, electricity costs, and idle capacity.

    Define four variables first:

    • Model size and precision: 7B, 14B, 32B, or larger; FP16, BF16, FP8, or 4-bit quantisation.
    • Workload type: pre-training, supervised fine-tuning, preference optimisation, inference, or embeddings.
    • Latency and throughput: response time for one user versus tokens per second across many users.
    • Deployment model: local workstation, colocated server, Indian cloud region, or a global managed service.

    For teams working with Indian-language data, hardware planning should follow the data and model strategy. Read how to train LLMs on Indian datasets before estimating a training cluster: dataset cleaning, deduplication, and evaluation can materially change compute requirements.

    How much GPU memory do LLMs need?

    VRAM is usually the first constraint. A rough inference estimate for model weights is:

    Parameters × bytes per parameter

    A 7B model needs approximately 14 GB for FP16 weights, 7 GB for 8-bit weights, or about 4 GB for 4-bit weights. The real requirement is higher because the runtime also needs space for the KV cache, activations, CUDA graphs, temporary buffers, and batching.

    Typical starting points are:

    • 8–12 GB: small models, embeddings, experimentation, and heavily quantised 7B models with modest context windows.
    • 16–24 GB: a strong developer setup for 7B–14B inference, fine-tuning with parameter-efficient methods, and computer-vision or NLP pipelines.
    • 40–48 GB: more comfortable fine-tuning, longer contexts, larger batches, and 30B-class quantised inference.
    • 80 GB or more: serious training and fine-tuning, large-context serving, and multi-GPU workloads.

    Quantisation reduces memory use but can affect quality and kernel compatibility. Test the exact model, context length, batch size, and serving engine rather than relying on a theoretical capacity figure.

    GPU options worth considering in 2026

    NVIDIA data-centre GPUs

    NVIDIA remains the practical default for many LLM teams because CUDA, cuBLAS, NCCL, PyTorch integrations, and inference engines have broad support. The A100 80GB remains useful on the cloud market for training and fine-tuning, while newer H100 and H200 systems offer substantially stronger tensor performance, memory bandwidth, and support for modern low-precision workloads. B100/B200-class systems can be attractive for large-scale production, but availability, pricing, and access in India vary by provider.

    Choose these when you need mature distributed training, predictable framework support, or high utilisation across multiple workloads. For a single developer workstation, their price and power draw are usually difficult to justify.

    NVIDIA RTX workstations

    Consumer and workstation RTX cards can deliver excellent value for prototyping and local inference. Cards with 24 GB or more of VRAM are particularly useful for 7B–14B models, quantised larger models, LoRA fine-tuning, and multimodal experiments. They generally lack the data-centre features, ECC memory, and NVLink configurations of enterprise systems, but the lower acquisition cost can be decisive for an early-stage Indian team.

    Plan for cooling, a reliable power supply, and physical space. A workstation that throttles under sustained load is not cheaper in practice than a well-configured cloud instance.

    AMD Instinct GPUs

    AMD Instinct accelerators can be compelling where ROCm support matches the stack. MI250, MI300, and newer generations provide substantial memory capacity and bandwidth, especially for teams able to validate kernels and dependencies. However, compatibility is workload-specific. Check support for PyTorch, vLLM, FlashAttention, quantisation libraries, and distributed communication before committing.

    Cloud GPUs and specialised accelerators

    Cloud access is often the best starting point when usage is irregular. Compare cost per useful output token, not hourly price alone. Include storage, data transfer, attached CPU and RAM, checkpoint persistence, idle time, and regional availability. Google TPUs and other specialised accelerators can work well for teams already invested in their software ecosystem, but they may require more adaptation than a CUDA-based deployment.

    For local deployment, how to deploy lightweight LLMs locally in 2026 covers practical model and runtime choices that can reduce dependence on expensive accelerators.

    Training, fine-tuning, and inference require different choices

    Pre-training is dominated by total compute, data throughput, inter-GPU networking, and checkpoint reliability. Look for high-bandwidth memory, fast GPU-to-GPU communication, and a cluster with strong networking rather than simply counting GPUs.

    Fine-tuning is more flexible. LoRA, QLoRA, gradient checkpointing, mixed precision, and smaller sequence lengths can make a 24–48 GB card sufficient for many domain adaptations. Teams should benchmark the quality trade-off before scaling hardware. The guide to fine-tuning LLMs on custom data is useful for designing that experiment.

    Inference is governed by latency, concurrent requests, context length, and utilisation. A quantised model on a smaller GPU may beat a larger model on an expensive accelerator when traffic is low. At higher concurrency, continuous batching and efficient KV-cache management become as important as raw tensor performance.

    A practical selection framework for Indian teams

    Use this sequence before purchasing or reserving capacity:

    1. Create a representative benchmark. Measure tokens per second, time to first token, peak VRAM, error rate, and quality on real Indian-language or domain prompts.
    2. Test three precisions. Compare BF16 or FP16 with 8-bit and 4-bit variants, including long-context behaviour.
    3. Calculate monthly utilisation. A rented GPU used 20 hours a month is usually preferable to owning a server; steady, high utilisation can reverse that decision.
    4. Account for operations. Include data residency, backups, observability, patching, cooling, power, and support.
    5. Validate the software path. Confirm drivers, CUDA or ROCm versions, serving engine support, and multi-GPU communication before procurement.
    6. Plan for the next model size. Leave headroom for longer contexts, retrieval, reranking, or a second model in the same service.

    If your product is already serving users, pair GPU benchmarking with LLM application performance monitoring in India. Monitoring actual queue time, cache usage, GPU utilisation, and cost per request prevents overprovisioning.

    Recommended starting points

    • Solo developer: a 16–24 GB RTX workstation or an hourly cloud GPU; use quantised open models.
    • Startup fine-tuning a 7B–14B model: 24–48 GB VRAM with QLoRA or LoRA, then validate on a larger cloud instance.
    • Production inference: select a GPU based on concurrency and context length; benchmark vLLM or an equivalent engine with real traffic patterns.
    • Enterprise training: use 80 GB-class or newer data-centre GPUs, fast networking, robust checkpointing, and reserved capacity where utilisation is predictable.
    • Budget-sensitive deployment: optimise the model and request path first. Smaller models, caching, batching, retrieval quality, and quantisation often save more than a hardware upgrade.

    The right GPU for LLMs is the one that meets quality and latency targets at an acceptable cost per useful result. Start with a reproducible benchmark, choose the smallest configuration that clears the target, and scale only when measured demand requires it.

    Last updated 24 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.