0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · scaling deep learning models on low budget infrastructure

Scaling Deep Learning Models on Low-Budget Infrastructure

  1. aigi

    Start with a workload and budget, not a GPU

    Scaling deep learning models on low budget infrastructure is an engineering optimisation problem. The objective is not to recreate a hyperscaler’s cluster; it is to meet a defined quality, latency, and delivery target at a sustainable cost.

    First classify the workload:

    • Inference only: use an existing model, quantise it, and optimise serving.
    • Supervised fine-tuning: use LoRA or QLoRA before considering full fine-tuning.
    • Continued pretraining: budget for data quality, checkpoint storage, and repeated experiments.
    • Pretraining from scratch: attempt this only with a narrow model, a clear data advantage, and a credible compute plan.

    Set a maximum monthly spend, an experiment budget, and a completion deadline. Track cost per training run, cost per million tokens or images, GPU utilisation, and validation quality. A cheap run that produces unusable output is not economical.

    Teams still building their fundamentals should document these experiments in a reproducible repository. A strong machine learning portfolio project for beginners in India can demonstrate the same discipline: fixed datasets, tracked metrics, and an honest infrastructure bill.

    Reduce the model’s memory footprint first

    Hardware upgrades should be the last response to a memory problem. Apply these techniques before adding nodes:

    • Parameter-efficient fine-tuning: LoRA updates small low-rank matrices rather than every model weight. QLoRA combines this approach with low-bit base-model loading and is often the most practical route for adapting 7B–14B models.
    • Mixed precision: Use BF16 where supported for stability, and FP16 where hardware and software stacks are compatible. Keep sensitive operations in higher precision when required.
    • Gradient accumulation: Several small micro-batches can simulate a larger batch without requiring equivalent VRAM.
    • Activation checkpointing: Recompute activations during backpropagation to reduce memory, accepting a moderate increase in compute time.
    • Efficient attention kernels: FlashAttention-class implementations reduce memory traffic and can improve throughput, particularly for long contexts.
    • Sequence packing: Combine shorter examples into full training sequences so padding does not consume paid compute.

    Quantisation is not automatically safe. Test quality on a task-specific evaluation set, including Indian languages, code-mixed text, and production edge cases if those are part of the product.

    Choose hardware by workload

    A used RTX 3090 can remain attractive for experimentation because its 24GB of VRAM supports many fine-tuning workloads. RTX 4090 systems provide more compute but require careful attention to power, cooling, and total system cost. Consumer cards usually lack data-centre features such as ECC memory, strong interconnects, and predictable multi-GPU networking.

    For a local machine, calculate the full cost:

    • GPU, motherboard, CPU, RAM, NVMe storage, and power supply
    • Electricity, cooling, maintenance, and replacement risk
    • Engineer time spent managing drivers and failed jobs
    • The residual value of the equipment after the project

    Cloud GPUs are preferable when demand is irregular or the project needs a different accelerator for each phase. Compare providers using delivered cost per useful training step, not the hourly rate alone. Include storage, egress, idle time, startup delays, and regional taxes or currency exposure. Indian teams should also evaluate domestic providers and data-centre locations when latency, procurement, or data residency matter.

    Use distributed training selectively

    More GPUs do not always reduce the time or cost to a useful checkpoint. On ordinary Ethernet, communication can dominate when gradients are synchronised frequently.

    Use PyTorch DistributedDataParallel for replicas of a model that fits on each GPU. Use FSDP or DeepSpeed ZeRO when parameters, gradients, and optimiser states must be sharded. Start with the smallest configuration that works and measure scaling efficiency after adding each node.

    A sensible progression is:

    1. Single-GPU baseline with a fixed data order and evaluation set.
    2. Single machine with multiple GPUs and local storage.
    3. Multi-node training only after profiling communication and input throughput.
    4. Larger or faster instances if they deliver a lower cost per completed step.

    For many startup workloads, one well-utilised GPU beats a loosely connected cluster. Read the broader principles in this guide to scaling backend infrastructure for AI applications, especially around queues, observability, and capacity planning.

    Make spot and preemptible compute reliable

    Spot instances can reduce compute prices substantially, but interruption is part of the design. Never run a long job without resumability.

    Implement the following:

    • Save model, optimiser, scheduler, random-state, and data-sampler checkpoints to durable object storage.
    • Checkpoint by time and by completed samples, not only by epoch.
    • Keep the last few checkpoints and verify that they can be restored in a clean environment.
    • Separate immutable training data from temporary local caches.
    • Automatically requeue failed jobs with an upper limit on retries.
    • Record the exact container image, commit, configuration, and dependency lockfile.

    A 15-minute checkpoint interval may be appropriate for volatile capacity, but benchmark the storage overhead. For short jobs, on-demand capacity can be cheaper once interruption, restart, and engineer time are included.

    Fix the data pipeline before buying compute

    Low GPU utilisation often comes from slow storage, tokenisation, image decoding, or network access. Pre-tokenise stable text datasets, store shards in formats suited to sequential reads, and keep hot data near the workers. WebDataset, Parquet-based pipelines, memory mapping, and asynchronous prefetching can help, but measure rather than adding tools by default.

    Run a short profiling job and inspect:

    • GPU utilisation and memory use
    • Data-loader wait time
    • CPU saturation and RAM pressure
    • Storage throughput and network traffic
    • Tokens or samples processed per second

    Data quality also determines how much compute is wasted. Deduplicate, remove corrupted examples, filter sensitive data appropriately, and maintain a validation split that is not used for tuning. For high-stakes products, data veracity infrastructure for high-stakes AI is directly relevant: provenance and quality checks can prevent expensive training on unreliable data.

    Optimise inference separately

    Training economics do not predict serving economics. A model that is affordable to fine-tune can still be expensive to serve at low traffic.

    Use dynamic batching, continuous batching, and prefix or KV-cache reuse where appropriate. Evaluate vLLM, TensorRT-LLM, ONNX Runtime, or other serving stacks against your latency target and model architecture. Quantise to INT8, INT4, or another supported format only after measuring accuracy and tail latency. Route simple requests to a smaller model and reserve the larger model for cases that need it.

    CPU inference is viable for smaller classifiers, embeddings, rerankers, and asynchronous jobs. For voice products, GPU spend is only one part of the bill; telephony, transcription, storage, and concurrency planning matter too. The principles in telephony infrastructure for scalable voice agents provide a useful comparison.

    A practical 2026 operating plan

    For an Indian startup, a staged plan reduces risk:

    • Week 1: establish a single-GPU baseline, evaluation suite, and cost dashboard.
    • Weeks 2–3: add PEFT, mixed precision, checkpointing, data profiling, and quantised inference.
    • Weeks 4–6: compare local, domestic cloud, and spot pricing using identical workloads.
    • After validation: scale only the bottleneck—training throughput, memory, storage, or serving concurrency.

    Keep production and experimentation isolated. Apply quotas, automatic shutdowns, budget alerts, and lifecycle rules for object storage. Treat every experiment as a decision: continue, change the method, or stop.

    For founders moving from a research prototype to a product, transitioning from research to a deep tech startup covers the commercial and operational decisions that infrastructure planning alone cannot solve.

    FAQ

    Can a consumer GPU fine-tune a large language model?

    Often, yes. QLoRA, gradient accumulation, checkpointing, and short context windows can make 7B-class fine-tuning practical on a 24GB card. Larger models may require sharding, offloading, or a cloud instance.

    Is a multi-GPU cluster always faster?

    No. Synchronisation and data-loading overhead can erase the benefit of additional GPUs. Benchmark throughput and cost per useful checkpoint before expanding.

    Should a startup buy GPUs or rent them?

    Buy when utilisation is predictable and the team can operate the hardware. Rent when workloads are bursty, capital is limited, or different projects need different accelerators.

    What should be tracked first?

    Track cost per experiment, useful samples or tokens per hour, GPU utilisation, checkpoint recovery time, validation quality, and inference cost per request. These metrics connect infrastructure choices to product outcomes.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.