Compute is a constraint, but it does not have to be a ceiling. For most Indian startups, student teams, and independent developers, the winning strategy is not training the largest model. It is building a system that delivers the required quality at the lowest repeatable cost.
How to build AI at scale with limited compute resources starts with product and workload discipline: define the task, measure quality, select the smallest capable model, and spend compute only where it improves user outcomes. A specialised 3B–8B model with good retrieval and evaluation can be more useful than a much larger general-purpose model.
Start with the workload, not the model
Before choosing a GPU or foundation model, write down the workload’s actual requirements:
- Inputs: text, images, audio, video, or mixed modalities
- Latency: interactive, near-real-time, or batch
- Traffic: requests per minute, peak concurrency, and expected growth
- Quality: accuracy, groundedness, safety, and language coverage
- Context: average and maximum prompt length
- Data rules: residency, retention, encryption, and customer isolation
Separate offline work from online serving. Dataset cleaning, embedding generation, evaluation, and batch inference can run on interruptible machines. User-facing inference needs predictable availability and latency. This separation often reduces costs more than a model change.
For Indic products, evaluate each target language rather than relying on an English benchmark. The low-resource Indic NLP guide covers data selection, language-specific evaluation, and practical approaches for building systems that work beyond English.
Use retrieval before fine-tuning
Fine-tuning is useful for changing behaviour, output format, tone, or task execution. It is usually the wrong first tool for adding frequently changing facts. For that, use retrieval-augmented generation (RAG): store trusted documents as chunks and embeddings, retrieve relevant passages, and ask the model to answer from that evidence.
A lean RAG system should include:
- A curated document-ingestion pipeline with deduplication and access controls
- Chunking that preserves headings, tables, citations, and document structure
- Hybrid search combining semantic embeddings with keyword or metadata filters
- A reranker for difficult queries, used only when it improves measured accuracy
- Citations, refusal rules, and a test set of real user questions
RAG avoids repeatedly retraining a model whenever policies, prices, schemes, or internal documents change. It also makes errors easier to inspect. Fine-tune only after you can show that retrieval and prompt design cannot solve the problem.
Choose the smallest capable model
Begin with a strong small model and establish a baseline. Compare models on your own evaluation set, not just public leaderboards. Measure answer quality alongside tokens per second, time to first token, memory use, and cost per successful task.
Useful options include:
- Small language models: suitable for classification, extraction, routing, and constrained generation
- Distillation: train a smaller student model on high-quality teacher outputs, then validate carefully for factuality and safety
- Mixture-of-experts models: potentially efficient at inference, but only when the serving stack supports their routing and memory pattern
- Task-specific models: often better for embeddings, reranking, speech recognition, OCR, or moderation than a general LLM
For voice products, do not use an LLM for every stage. A smaller speech-to-text model, a compact intent classifier, and a lightweight response model may deliver lower latency and cost. Compare the architecture against practical deployment patterns in this voice agent architecture and deployment guide.
Reduce memory with quantization and parameter-efficient tuning
Quantization stores model weights at lower precision, commonly INT8 or 4-bit formats. Post-training quantization (PTQ) is the fastest starting point: quantize an existing model, then test it against your evaluation suite. If accuracy drops on sensitive tasks, try a higher precision format or quantization-aware training.
Quantization is not automatically safe for every workload. Test long contexts, code, numbers, multilingual prompts, tool calls, and refusal behaviour. Keep a full-precision or higher-precision fallback for difficult requests.
When adaptation is necessary, use LoRA or QLoRA instead of updating every parameter. These methods freeze the base model and train small adapter weights, reducing VRAM and storage requirements. Practical safeguards include:
- Keep training and validation data strictly separate
- Use gradient accumulation when the GPU cannot fit the desired batch
- Enable gradient checkpointing to trade compute for memory
- Monitor adapter overfitting with held-out, production-like prompts
- Store adapters separately so one base model can serve multiple customers
A single consumer GPU can be enough for a focused adaptation project, but the total budget must include data preparation, evaluation, storage, failed experiments, and serving.
Make training data do more work
Compute efficiency begins with data efficiency. Remove duplicates, corrupt samples, boilerplate, personally identifiable information, and contradictory labels before training. A smaller, cleaner dataset often beats a large noisy crawl.
Use active learning to send uncertain or high-impact examples for review. Generate synthetic data only for clearly defined gaps, and mix it with verified human or real-world examples. For India-focused systems, include spelling variation, code-switching, transliteration, accents, regional formats, and domain terminology in evaluation—not merely in training.
Create a compact regression set before every model or prompt change. Track failures by category: hallucination, missed retrieval, language error, unsafe completion, formatting failure, latency, and tool-use error. This prevents teams from spending GPU hours optimising a metric that users do not care about.
Serve efficiently under real traffic
Inference cost is driven by model size, input and output tokens, concurrency, and repeated context. Optimise each one:
- Use KV caching for autoregressive generation and repeated conversation context
- Apply prompt compression and summarisation to control long histories
- Stream responses only when it improves perceived latency
- Use continuous batching with a production inference engine such as vLLM or an equivalent stack
- Route simple requests to a small model and escalate only difficult cases
- Cache deterministic embeddings, classifications, and safe repeat queries
- Use speculative decoding when a compatible draft model improves measured throughput
Set budgets at the API boundary: maximum input tokens, output tokens, retries, tool calls, and per-user concurrency. Without limits, one long-running agent can consume the capacity intended for hundreds of ordinary requests.
If your product uses multiple specialised agents, model orchestration becomes a systems problem. The principles in building distributed systems with AI agents are useful for queues, retries, state, observability, and failure isolation.
Plan cloud capacity for India and cost volatility
Use a mixed compute strategy rather than committing every workload to premium GPUs. Run experiments and batch jobs on spot or preemptible instances with checkpointing. Keep a stable serving pool for production, and compare regional availability, egress charges, storage, and support—not just hourly GPU prices.
Design jobs to resume safely:
- Save checkpoints and dataset manifests frequently
- Make training deterministic enough to reproduce failures
- Store model versions, tokenizer versions, prompts, and evaluation results
- Use queues so interrupted batch jobs return automatically
- Separate persistent data from disposable machines
Measure cost per successful task, not cost per GPU hour. A cheaper model that causes more retries, human review, or customer abandonment may be the expensive option.
A practical 30-day execution plan
Week 1: define quality, latency, privacy, and cost targets; collect a representative evaluation set.
Week 2: benchmark two or three small models with RAG and strict token limits; profile memory and throughput.
Week 3: add quantization, routing, caching, and continuous batching; test failure cases and peak concurrency.
Week 4: fine-tune only the remaining gaps with LoRA or QLoRA; deploy behind monitoring and a rollback path.
Track quality, p50 and p95 latency, GPU utilisation, tokens per second, cache hit rate, cost per request, and cost per successful outcome. Revisit the architecture when traffic or task mix changes.
Frequently asked questions
Can a small team build a competitive AI product?
Yes. Compete on a narrow workflow, proprietary data, reliability, language coverage, or distribution rather than raw parameter count. A focused system can outperform a general model on a defined task.
Is 4-bit quantization safe?
It can be, but validate it. Test multilingual output, long context, arithmetic, structured generation, tool calls, and safety behaviour. Keep a higher-precision fallback for sensitive requests.
When should I fine-tune instead of using RAG?
Use RAG for changing knowledge and source-grounded answers. Fine-tune when you need consistent behaviour, formatting, classification, style, or task-specific reasoning that prompting and retrieval cannot provide.
What should I optimise first?
Start with data quality, model selection, token budgets, and routing. Then optimise kernels and infrastructure. Hardware tuning cannot rescue an undefined task or a poor evaluation set.
Build lean, then scale deliberately
Limited compute rewards teams that measure before they expand. Start with a narrow product loop, a small capable model, clean data, retrieval where appropriate, and observable serving. Add GPUs only when evidence shows that more capacity—not better routing, caching, or data—is the constraint.
Indian builders looking for structured support can explore AI Grants India for funding and ecosystem opportunities. The strongest applications will show a clear user problem, a defensible data or distribution advantage, and a credible plan for delivering quality within compute limits.