Supervised fine-tuning (SFT) teaches a base model to follow examples. Preference optimisation and reinforcement learning then push it towards responses that are more useful, safe, truthful, and consistent. The engineering challenge is not simply adding GPUs: it is building a repeatable system in which data, training, evaluation, and deployment improve together.
For teams building in India, the right design must also handle uneven GPU availability, multilingual data, privacy requirements, and strict cost limits. Treat alignment as a production pipeline rather than a one-off training run.
Start with a measurable alignment target
Before choosing FSDP, DPO, or PPO, define what “better” means. Write a short model specification covering:
- Supported languages, domains, and user personas
- Desired behaviours, such as citation, refusal, brevity, or tool use
- Failure modes that are unacceptable
- Offline metrics and human review criteria
- Latency, memory, and inference-cost limits
Create a small, locked evaluation set before training. Keep it separate from SFT and preference data, and include difficult cases: ambiguous prompts, code-switching, unsafe requests, factual questions, long context, and domain-specific terminology. For Indian deployments, include Hindi and other relevant Indic languages rather than assuming English benchmarks transfer.
A reliable data veracity infrastructure approach is valuable here: record provenance, licence, language, annotator instructions, transformations, and known limitations for every dataset version.
Build the data pipeline before increasing model size
Scaling low-quality data only makes mistakes cheaper to reproduce. Begin with a canonical schema for prompts, responses, preference pairs, metadata, and evaluation labels. Store immutable raw data separately from processed training shards so that every checkpoint can be traced back to its inputs.
For SFT, prioritise diverse, instruction-complete examples over volume. Remove duplicates, near-duplicates, malformed conversations, leaked benchmark items, and answers with unsupported claims. Pack examples of similar lengths into training sequences to reduce padding waste, but preserve conversation boundaries and mask loss on user messages where appropriate.
Synthetic data can expand coverage, but it needs filtering. Generate candidate instructions with a stronger teacher, score them for difficulty and relevance, then sample human reviews from both high- and low-confidence groups. Use Python scripts for automating data preprocessing for deterministic cleaning, hashing, language identification, deduplication, and shard validation.
For preference data, ensure comparisons are genuinely informative. Pairs where both answers are poor, nearly identical, or judged solely on stylistic preference add little signal. Capture the reason for a preference—factuality, safety, instruction following, completeness, or tone—so evaluation can diagnose regressions.
Choose the least complex alignment method that works
SFT is usually the best first alignment stage. Train a strong instruction-following baseline, evaluate it, and only then add preference optimisation. This isolates data problems from reward-model or policy-training problems.
DPO and related offline methods are often the practical default for teams with limited infrastructure. They learn from preferred and rejected responses without maintaining a separate online rollout loop. Variants such as IPO, ORPO, and KTO may be useful when preference data is limited or differently structured, but compare them on the same held-out set rather than adopting them by reputation.
PPO and online RL become relevant when exploration, tool use, long-horizon behaviour, or continually refreshed feedback is central to the product. They require careful management of policy, reference, reward, and value components, as well as rollout quality and reward hacking. Start with small-scale experiments and explicit stop conditions.
AI-generated feedback can reduce annotation costs, but it should not replace human review for high-risk behaviour. Use RLAIF to expand coverage, calibrate the teacher against expert labels, and audit disagreement cases. For systems that can act autonomously, pair alignment training with secure autonomous AI workflows so model quality is not treated as the only safety control.
Scale training with a staged parallelism plan
Use the simplest distributed strategy that fits the model and batch. Data parallelism is easy to operate when each GPU can hold a full replica. When memory becomes the constraint, use FSDP or ZeRO to shard parameters, gradients, and optimizer states. Activation checkpointing, mixed precision, FlashAttention, and fused optimisers can further reduce memory pressure, but measure their effect on throughput rather than enabling every option blindly.
For very large models, combine approaches carefully:
- FSDP or ZeRO for sharded training state
- Tensor parallelism when individual layers exceed device memory
- Pipeline parallelism when the model is split across stages
- Sequence parallelism for long-context workloads
- Gradient accumulation when per-device batches are small
Benchmark tokens per second, step time, communication overhead, checkpoint time, and failure recovery. A cluster that is nominally larger may deliver less useful work if network bandwidth, data loading, or synchronisation is poor.
Separate training and serving environments where possible. Use reproducible containers, pinned dependencies, automatic checkpoint uploads, and resumable jobs. Indian teams can reduce exposure to GPU shortages by supporting more than one provider and using orchestration tools such as Kubernetes, KubeRay, or SkyPilot. Best AI developer tools for cloud automation can help standardise provisioning, secrets, logs, and teardown across providers.
Control cost and operational risk
Track cost per million training tokens, cost per accepted preference, and cost per evaluation improvement. These metrics reveal whether a larger run is producing meaningful gains. Use parameter-efficient adapters for iteration and ablations; reserve full fine-tuning for cases where adapters cannot meet quality or deployment requirements.
Use a three-tier checkpoint policy:
- Frequent lightweight checkpoints for recovery
- Periodic adapter or model snapshots for experiments
- Fewer full checkpoints retained for release candidates
Validate every dataset shard and checkpoint before training continues. Test restoration from a failed worker, not just successful completion. Spot instances can be useful for resumable jobs, but do not place the only copy of data, metadata, or checkpoints on ephemeral storage.
Evaluate continuously, not at the end
A single aggregate score can hide serious regressions. Run evaluations after each meaningful SFT or preference-training change across capability, safety, factuality, multilingual performance, refusal behaviour, and latency. Combine automated graders with human review, especially for culturally specific or high-stakes outputs.
Use a fixed regression suite, a rotating challenge set, and production-like samples with personal information removed. Monitor reward-model disagreement, refusal rates, verbosity, hallucination patterns, and performance by language. LLM judges are useful for scale, but calibrate them against expert ratings and inspect adversarial examples.
Before release, test the complete serving path: quantisation, batching, retrieval, tool calls, timeouts, and fallback behaviour. A model that wins offline but fails under real latency or context limits is not production-ready. For teams building internal engineering systems, building open source AI tools for Indian developers offers relevant patterns for transparent evaluation and reusable infrastructure.
A practical 2026 rollout plan
1. Week 1–2: define the model specification, build the evaluation set, and audit data rights and provenance.
2. Week 3–4: train an SFT baseline with packing, efficient attention, and reproducible checkpoints.
3. Week 5–6: create a small, high-quality preference set; compare DPO or another offline method against the SFT baseline.
4. Week 7–8: profile distributed throughput, test failure recovery, and tune data-loader and communication settings.
5. After validation: introduce synthetic or AI feedback selectively, expand multilingual coverage, and consider online RL only where offline methods plateau.
The core principle is straightforward: scale the feedback and measurement system before scaling the GPU bill. Clean data, explicit objectives, efficient preference methods, resilient orchestration, and continuous evaluation will usually outperform a larger but poorly controlled training run.