0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · compute for large models

Compute for Large Models: A Practical Guide for 2026

  1. aigi

    Large models are not defined only by parameter count. Compute for large models is a systems problem spanning GPUs, memory, data pipelines, networking, software, and the economics of serving each request. A model that trains successfully can still be too expensive to fine-tune, too slow for production, or impossible to run on the hardware available to an Indian startup or research team.

    The right approach is to measure the workload first, then select the smallest reliable system that meets quality and latency targets. As of 2026, that often means combining rented accelerators for training with quantised, smaller models for inference rather than attempting to run every workload on the largest available GPU.

    What compute does a large model actually need?

    Compute requirements differ across three stages:

    • Pre-training: Usually the most demanding stage, involving massive datasets, long accelerator runs, high-speed interconnects, and fault-tolerant checkpointing.
    • Fine-tuning: Less expensive than pre-training, but memory-intensive methods such as full-parameter fine-tuning can still require multiple GPUs.
    • Inference: Repeated serving may become the largest long-term cost. Memory capacity, tokens per second, concurrency, and response latency matter more than raw training throughput.

    Estimate requirements using more than parameter count. Track sequence length, batch size, precision, number of training tokens, expected users, requests per second, and maximum response length. A 7B-parameter model with long context and high concurrency can be harder to serve than a larger model handling short prompts.

    For builders working on regional-language or multimodal systems, data preparation is also part of the compute budget. Projects involving Indian languages, OCR, speech, or vision may spend substantial time on preprocessing, deduplication, tokenisation, evaluation, and transfer learning.

    Start with a workload and budget specification

    Before selecting a cloud instance or buying hardware, write a short compute specification:

    • Target model and task
    • Training, fine-tuning, or inference objective
    • Dataset size and storage format
    • Context length and expected batch size
    • Target quality metrics
    • Maximum acceptable training time
    • Inference latency and concurrency targets
    • Monthly budget and data-residency requirements

    Separate experimentation compute from production compute. Researchers may benefit from flexible, interruptible instances, while a production API needs predictable availability and capacity. For Indian teams, compare regional cloud availability, egress charges, GST-inclusive pricing, support, and whether data can remain within the required jurisdiction.

    A simple benchmark is more useful than a provider’s headline specification. Run a representative workload for several minutes and record throughput, peak memory, power or instance cost, failure recovery, and output quality. Repeat the test at the context lengths your application will actually use.

    Hardware choices: GPU, accelerator, or CPU?

    GPUs remain the default for most large-model training and inference because they combine high memory bandwidth with parallel computation. The important specifications are:

    • VRAM or accelerator memory: Determines whether weights, activations, optimiser states, and the KV cache fit.
    • Memory bandwidth: Strongly affects inference and many training operations.
    • Interconnect: NVLink, InfiniBand, or equivalent high-speed networking can determine multi-GPU efficiency.
    • Precision support: BF16, FP16, FP8, and INT8 support can improve speed and reduce memory use.
    • Availability and reliability: A cheaper accelerator is not useful if capacity is difficult to obtain or failures lose long training runs.

    CPUs remain valuable for data loading, retrieval, orchestration, preprocessing, and smaller models. Dedicated inference accelerators can be cost-effective at scale, but only after the model and serving stack are compatible and benchmarked.

    Do not assume that adding GPUs produces linear speedups. Communication overhead, uneven batches, storage bottlenecks, and synchronisation can reduce scaling efficiency. Measure throughput per rupee, not just throughput per GPU.

    Reduce memory before adding hardware

    Memory optimisation is often the fastest route to lower compute costs.

    • Quantisation: Use FP16 or BF16 where appropriate, then evaluate INT8 or INT4 for inference. Quantisation-aware testing is essential because quality may vary by language, domain, and task.
    • Parameter-efficient fine-tuning: LoRA and related adapter methods train a small set of parameters instead of updating the full model.
    • Gradient checkpointing: Recomputes selected activations to reduce memory at the cost of extra computation.
    • Gradient accumulation: Simulates a larger batch when the device cannot hold it directly.
    • Pruning and distillation: Remove low-value capacity or train a smaller student model for a defined task.
    • Paged or efficient attention: Reduces memory pressure for long contexts and concurrent requests.

    For production, test quality on representative Indian names, scripts, code-switching, accents, and domain terminology. A lower-bit model that fails on these cases can create more operational cost than it saves in hardware.

    Teams building open-source small language models for Hindi can often achieve better economics by fine-tuning a compact base model with adapters, quantising it, and serving it on a single accelerator or CPU-compatible runtime.

    Distributed training without wasting accelerators

    Distributed training is useful when a model or dataset exceeds one machine’s capacity. Data parallelism replicates the model across devices, while tensor and pipeline parallelism divide model computation. Each approach introduces communication and operational complexity.

    Use a staged plan:

    1. Validate the data pipeline and training loop on one GPU.
    2. Profile memory, dataloader throughput, and kernel utilisation.
    3. Scale to two or more devices and measure parallel efficiency.
    4. Add checkpointing, automatic recovery, and experiment tracking.
    5. Only then commit to a longer multi-node run.

    Keep datasets close to the compute environment, use fast local caching where possible, and ensure checkpoints are written to durable storage. A failed run that cannot resume wastes more compute than an imperfectly optimised kernel.

    Kubernetes can help teams share accelerators, but it is not automatically the best first step. For smaller teams, managed jobs or a simple queue may be easier to operate. When deployment grows complex, the guidance on deploying deep learning models on GKE is relevant to cluster design, packaging, and serving decisions.

    Inference: optimise for tokens, latency, and utilisation

    Inference economics depend on how efficiently the system converts accelerator time into useful tokens. Use continuous batching, prefix caching, KV-cache management, streaming responses, and model-appropriate runtimes. Set limits on context length and output tokens; unrestricted prompts can overwhelm capacity.

    Benchmark at several concurrency levels. A system that is fast for one request may collapse under simultaneous traffic. Record:

    • Time to first token
    • Inter-token latency
    • End-to-end latency
    • Tokens per second
    • Requests per second
    • Peak memory
    • Cost per million input and output tokens
    • Error and timeout rates

    Route requests intelligently. A small model can handle classification, extraction, summarisation, and routine support, while a larger model handles difficult reasoning. Retrieval can reduce the need to enlarge the model or context window. For repetitive enterprise workloads, caching and deduplication may produce greater savings than hardware changes.

    For complex applications, building high-performance AI applications with open-source tools offers a useful direction: keep the serving layer modular so that models, runtimes, and hardware can be changed without rewriting the product.

    Build an India-ready cost and operations plan

    Model compute is only one line item. Include storage, data transfer, observability, evaluation, orchestration, failed jobs, support, and engineering time. Compare reserved capacity with on-demand instances, and use spot or preemptible capacity only when checkpointing and retry logic are robust.

    Security and governance matter when processing health, financial, education, or public-sector data. Apply access controls, encryption, audit logs, retention limits, and clear separation between training data and production prompts. Track model versions and dataset versions so that quality regressions can be reproduced.

    A practical launch sequence is:

    • Benchmark a baseline model on real workloads.
    • Establish quality, latency, and cost thresholds.
    • Optimise precision, batching, caching, and adapters.
    • Load-test at expected and peak traffic.
    • Add fallback models and rate limits.
    • Monitor cost per request and quality after deployment.

    Builders exploring student or early-stage products can also study best machine learning projects for computer science students for project scopes that are ambitious but feasible on limited compute.

    FAQ

    How much GPU memory does a large model need?
    It depends on parameter count, precision, context length, batch size, optimiser states, and KV-cache requirements. Inference may fit through quantisation, while full fine-tuning usually needs substantially more memory.

    Should a startup buy GPUs or use the cloud?
    Use the cloud while workloads are uncertain or bursty. Buying hardware can make sense when utilisation is consistently high, workloads are stable, and the team can manage cooling, networking, security, and maintenance.

    Is the biggest model always the best choice?
    No. A smaller, well-evaluated model with retrieval, task-specific fine-tuning, and good serving can outperform a larger model on cost, latency, and domain accuracy.

    How do I reduce compute costs first?
    Measure the baseline, then reduce context and output lengths, use batching and caching, apply parameter-efficient fine-tuning, test quantisation, and route simple requests to smaller models before adding hardware.

    Last updated 24 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.