Generic model APIs are useful for validating a product, but they rarely create a durable advantage. For an Indian startup, building custom fine-tuned LLMs for Indian startups makes sense when the product depends on a specialised workflow, local-language performance, controlled tone, predictable latency, or lower unit economics.
Fine-tuning is not automatically the right answer. It changes a model’s behaviour; it does not reliably turn a model into a current knowledge base. A strong production system usually combines a capable base model, carefully curated examples, retrieval-augmented generation (RAG), deterministic business logic, and an evaluation pipeline.
Decide what the model must learn
Start with the failure that is costing the business money or trust. Fine-tuning is most useful when the model must repeatedly learn:
- A consistent output format, such as structured underwriting notes, claim summaries, or support dispositions.
- A specialised tone and workflow for a narrow customer segment.
- Indian language and code-switching patterns, including Hindi-English, Tamil-English, or Telugu-English interactions.
- Domain-specific classification, extraction, routing, or tool-calling behaviour.
- Short, predictable answers that reduce latency and inference cost.
Do not fine-tune merely to add changing facts such as interest rates, government schemes, inventory, or regulations. Put those facts in an authoritative retrieval layer and show citations where the use case demands it. Teams building best practices for fine-tuning LLMs on custom data should separate behavioural knowledge from business knowledge before collecting a single example.
Choose a base model for the actual workload
Benchmark candidate models on your own prompts rather than relying on leaderboard scores. Compare quality in the languages, scripts, accents, and document formats your customers use. Review licensing, commercial restrictions, context length, tool calling, quantisation support, and availability of local deployment options.
A practical shortlist may include open-weight Llama, Mistral, Gemma, Qwen, and Indic-focused models from Indian research and commercial teams. The best choice depends on the task:
- Small models, roughly 3B–8B parameters: suitable for classification, extraction, routing, and high-volume support when fine-tuned well.
- Mid-sized models: useful for complex reasoning, multilingual assistance, and multi-step tool use, but more expensive to serve.
- Large models: valuable as teachers, evaluators, or fallback systems; often uneconomical as the default model for every request.
Test tokenisation explicitly. Indian scripts can consume more tokens than English, affecting context limits and cost. Measure latency and accuracy separately for Romanised and native-script inputs; a model that performs well on Hindi may still struggle with Hinglish typed in Latin characters.
Build a trustworthy training dataset
Data preparation is usually the largest source of quality improvement. Begin with real, consented product interactions, expert-written examples, support transcripts, public documents with clear usage rights, and synthetic examples reviewed by people who understand the domain.
A useful instruction record should include:
- The user request and relevant context.
- The desired answer or structured output.
- An abstention or escalation response when the request is unsafe, ambiguous, or outside scope.
- Metadata such as language, domain, difficulty, and source.
- A quality label or reviewer rationale where possible.
Remove duplicates, boilerplate, prompt injection attempts, and low-quality conversations. Split data by customer, document, or time period—not randomly—so near-duplicates do not leak from training into testing. Keep a locked evaluation set that the training team cannot repeatedly tune against.
For fintech, health, education, and employment products, document the source and permission for every dataset. Apply data minimisation and remove or mask personal information before training. DPDP compliance is not achieved simply by hosting a GPU in India: establish a lawful purpose, retention policy, access controls, vendor agreements, deletion process, and incident response plan with qualified legal advice.
Start with PEFT, not full-model training
Most startups should begin with supervised fine-tuning using LoRA or QLoRA. These parameter-efficient methods train small adapter layers instead of updating every model parameter. They reduce memory requirements, speed up experiments, and let teams keep multiple task-specific adapters over one base model.
A sensible first experiment is:
1. Select one narrow task and 500–2,000 high-quality examples.
2. Establish a no-fine-tuning baseline using prompting and RAG.
3. Train a LoRA adapter with conservative learning rates and several checkpoints.
4. Evaluate on held-out languages, edge cases, and adversarial inputs.
5. Compare quality, latency, GPU memory, and cost per successful task.
Full fine-tuning is justified only when you have substantial, clean data and a clear reason adapters cannot meet the requirement. Continued pre-training may help when you possess a large corpus in a specialised language or domain, but it requires stronger data governance and evaluation than ordinary instruction tuning.
Pair fine-tuning with RAG and tools
A fine-tuned model should not be expected to memorise every policy or document. Use RAG for changing information, with access controls, document versioning, chunking tests, metadata filters, and citation checks. Use APIs or deterministic code for calculations, eligibility rules, payments, and other high-risk operations.
For voice products, fine-tune the language model only where it improves the conversation policy or task completion. Speech recognition, translation, and text-to-speech need separate evaluation for accents, noisy calls, and regional languages. This is especially relevant to teams exploring voice agent vs IVR for customer support or deploying fintech customer onboarding with voice agents.
Evaluate what customers actually experience
Accuracy alone is insufficient. Create a scorecard covering:
- Task completion and exact-match or structured-output accuracy.
- Factuality, citation correctness, and refusal quality.
- Performance by language, script, accent, gendered forms, and code-switching pattern.
- Safety, privacy leakage, prompt injection resistance, and jailbreak resilience.
- P95 latency, throughput, context usage, GPU utilisation, and cost per successful interaction.
- Human reviewer preference and escalation rate.
Use production shadow tests before replacing an API model. Log prompts, retrieved passages, outputs, tool calls, and user corrections with appropriate redaction. Monitor drift after launch: customer language changes, policies are updated, and a model that passed last quarter’s tests can quietly degrade.
Plan infrastructure and unit economics
Training can happen on Indian GPU providers, hyperscalers, or approved international infrastructure depending on data restrictions, price, and availability. Compare total experiment cost rather than hourly GPU price: storage, data transfer, failed runs, engineering time, observability, and idle capacity matter.
For serving, quantise only after measuring quality. A 4-bit model may deliver excellent economics for a narrow task, but aggressive quantisation can damage multilingual output or tool calling. Consider batching, continuous serving, response caching for safe requests, smaller routing models, and CPU inference for lightweight classifiers. Keep adapter and base-model versions in a registry, with reproducible training configurations and rollback procedures.
A practical 90-day delivery plan
Weeks 1–2: define the task, risk tier, success metrics, data permissions, and baseline model.
Weeks 3–5: clean and label data, build the evaluation harness, and implement RAG or tools where required.
Weeks 6–8: run LoRA/QLoRA experiments, test multilingual and adversarial cases, and compare economics.
Weeks 9–10: deploy in shadow mode, add monitoring, redaction, rate limits, and human escalation.
Weeks 11–12: launch to a small cohort, review failures weekly, and decide whether to expand, retrain, or keep the API baseline.
The goal is not to own the largest model. It is to own a reliable, measurable system that solves an Indian customer problem better and more affordably than a generic endpoint. Founders working on Indian open-source AI developer projects can also reuse open tooling, evaluation assets, and community benchmarks instead of rebuilding every component.
Frequently asked questions
How much data is needed? A few hundred excellent examples can improve a narrow workflow; broader multilingual behaviour usually needs much more. Validate with held-out data instead of using a fixed sample-count rule.
How much does fine-tuning cost? A small adapter experiment may cost a few thousand rupees to tens of thousands, depending on GPU, sequence length, retries, and dataset size. Data curation and engineering frequently cost more than compute.
Should startups train from scratch? Almost never. Start with an open-weight or hosted base model, then invest in data, evaluation, retrieval, and workflow integration.
Is Indian hosting mandatory? Not universally. Requirements depend on the data, contracts, sector, and applicable law. Treat location as one part of a broader governance and security assessment.
AI Grants India supports founders building practical, India-relevant AI systems. If your startup is developing a multilingual model, domain adapter, or efficient inference stack, apply for AI Grants India with evidence of the problem, dataset readiness, evaluation plan, and path to production.