0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · h200 gpu scaling

H200 GPU Scaling: Architecture, Performance and Deployment

  1. aigi

    The NVIDIA H200 is built for memory-intensive AI workloads, but buying more GPUs does not automatically produce proportionally more performance. H200 GPU scaling is the discipline of increasing throughput, reducing time to train or serve a model, and keeping utilisation high as workloads move from one accelerator to a multi-GPU server or cluster.

    For Indian AI teams, the decision is both technical and financial. Power, cooling, cloud availability, interconnects, data movement and engineering time can matter as much as GPU count. A sound scaling plan starts with the workload and its bottleneck—not with a target number of cards.

    What H200 GPU scaling actually means

    Scaling has two dimensions:

    • Scale-up: Use multiple H200 GPUs inside one server, connected through high-bandwidth GPU interconnects and a suitable host architecture.
    • Scale-out: Add servers and distribute work across a cluster using fast networking, collective communication and a reliable scheduler.
    • Serve-out: Replicate an inference service across GPUs or nodes to handle more requests while meeting latency targets.

    The H200’s large HBM3e memory capacity is especially useful for large language models, long-context inference, recommendation systems and scientific workloads. It can reduce paging and offloading compared with smaller-memory accelerators. However, memory capacity is not the same as compute throughput. A model may fit comfortably while remaining limited by data loading, kernel efficiency, synchronisation or CPU preprocessing.

    Choose the right parallelism strategy

    The best strategy depends on model size, batch size, sequence length and whether the objective is training speed or serving throughput.

    • Data parallelism: Each GPU holds a copy of the model and processes different samples. Gradients are synchronised after each step. It is usually the simplest approach when the model fits on one H200.
    • Fully sharded data parallelism: Parameters, gradients and optimiser states are partitioned across GPUs. This reduces per-GPU memory pressure, but adds communication and implementation complexity.
    • Tensor parallelism: Individual layers are split across GPUs. It helps large models fit and can support low-latency inference, provided GPU-to-GPU communication is fast enough.
    • Pipeline parallelism: Layers are divided into stages across GPUs. Micro-batching is essential; otherwise, pipeline bubbles leave accelerators idle.
    • Expert parallelism: Mixture-of-experts models route tokens to selected experts. Network traffic and load imbalance become central scaling concerns.

    In practice, hybrid parallelism is common. Start with the least complex approach that meets memory and latency requirements. Teams already designing high-performance AI pipelines should treat data layout, checkpointing and communication as first-class pipeline stages rather than afterthoughts.

    Measure scaling efficiency before expanding

    Do not evaluate a cluster using GPU utilisation alone. Capture a baseline on one GPU and compare it with two, four, eight and—where relevant—multiple nodes.

    Track:

    • Throughput: tokens per second, samples per second or requests per second.
    • Step time: Separate compute, input, checkpoint and synchronisation time.
    • Scaling efficiency: achieved speed-up divided by ideal linear speed-up.
    • Memory headroom: allocated and reserved HBM, activation memory and fragmentation.
    • Communication time: all-reduce, all-gather, all-to-all and point-to-point traffic.
    • End-to-end latency: Include tokenisation, retrieval, queueing and network time for inference.
    • Cost per useful unit: For example, cost per million training tokens or per million served tokens.

    A four-GPU job that delivers 3.6 times the single-GPU throughput has 90% scaling efficiency. If it delivers only 2.1 times, adding more GPUs may worsen economics unless the larger configuration enables a necessary model size or service-level target.

    Use repeatable datasets, fixed batch sizes and warm runs. Profile both the fast path and failure path, including retries, checkpoint recovery and node replacement. For production services, pair accelerator metrics with the practices described in LLM application performance monitoring in India.

    Build the platform around the bottleneck

    A reliable H200 deployment requires more than GPU servers.

    • Interconnect: Confirm the topology and bandwidth between GPUs, hosts and switches. Poor placement can turn tensor or expert parallelism into a network-bound workload.
    • CPU and PCIe capacity: Host-side tokenisation, data decompression and input pipelines must feed the GPUs fast enough.
    • Storage: Use local NVMe caches or high-throughput shared storage for datasets and checkpoints. Avoid repeatedly reading large shards over a congested network.
    • Software stack: Pin compatible versions of drivers, CUDA, NCCL, PyTorch or another framework, container runtime and orchestration layer.
    • Scheduling: Enforce GPU-aware placement, quotas, pre-emption policy and fair sharing. Kubernetes, Slurm or managed cloud schedulers can work; operational consistency matters more than branding.
    • Thermal and power planning: Validate rack power, cooling and redundancy. In India, data-centre location and cloud region can affect both availability and energy cost.

    For application teams, scaling backend infrastructure for AI applications provides the broader service architecture around queues, APIs, databases and autoscaling. The GPU layer should fit that architecture instead of becoming an isolated cluster.

    Optimise the workload before buying more GPUs

    Several changes can improve performance without increasing hardware:

    • Use mixed precision such as BF16 where model quality and numerical stability permit it.
    • Apply fused kernels, FlashAttention-style implementations and framework compilation where supported.
    • Tune batch size, sequence packing, gradient accumulation and activation checkpointing.
    • Keep data loaders asynchronous and use pinned host memory where appropriate.
    • Cache tokenised data and avoid repeated CPU transformations.
    • Use quantisation for inference after validating accuracy and tail latency.
    • Save checkpoints asynchronously and test restore time on the target storage system.

    Open-source kernels and frameworks can materially lower cost, but compatibility testing is essential. A review of open-source tools for high-performance AI applications is useful when assembling a production stack rather than relying on default framework settings.

    Cost and capacity planning for Indian teams

    Compare configurations using the workload’s useful output, not hourly GPU price alone. Include:

    • GPU rental or depreciation
    • storage, data transfer and egress
    • orchestration and observability
    • engineering and platform operations
    • idle capacity between jobs
    • failed runs and checkpoint recovery
    • electricity and cooling for on-premises systems

    Cloud bursting can be sensible for short training campaigns, while a reserved or owned fleet may win for predictable inference. Keep datasets and model artefacts close to the compute region when possible, and test availability before committing to a launch schedule. If API-based components are part of the system, account for AI API cost blockers alongside GPU expenditure.

    A practical rollout plan

    1. Define the target: training time, throughput, p95 latency or cost per token.
    2. Establish a single-H200 baseline with a production-like workload.
    3. Profile compute, memory, input, storage and communication separately.
    4. Test one-server scaling before multi-node scaling.
    5. Validate fault recovery, checkpointing and job resumption.
    6. Add autoscaling and quotas only after measuring queue and service behaviour.
    7. Review quality, security and unit economics at every scale point.

    The result should be a capacity model that states when to add GPUs, when to optimise software and when to redesign the model or serving path. That is more useful than a headline benchmark and more defensible when applying for infrastructure funding or planning a production launch.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.