Foundation models are not built with one library. They require a coordinated stack for data collection and filtering, tokenisation, distributed training, evaluation, inference, deployment, and ongoing monitoring. The best open source tools for building foundation models let Indian research teams and startups control more of that stack without locking themselves into a single cloud or vendor.
The important distinction is between building a model from scratch and adapting an existing open-weight model. Pre-training a competitive language or multimodal model requires large, carefully governed datasets, substantial GPU capacity, and a team that can operate distributed systems. Fine-tuning, continued pre-training, distillation, or retrieval augmentation is far more achievable for most builders.
Start with the right level of ambition
Before selecting tools, define the model project:
- Fine-tuning: Adapt an existing model to a domain, language, style, or task.
- Continued pre-training: Extend training on specialised or underrepresented data.
- From-scratch pre-training: Train a new base model with a new tokenizer and corpus.
- Multimodal training: Combine text with images, audio, video, or structured data.
- Inference optimisation: Make an existing model cheaper and faster to serve.
For most Indian startups, a strong baseline is to begin with an open-weight model, evaluate it on representative Indian data, and fine-tune only when prompting or retrieval is insufficient. Teams working in Indic languages should also review this guide to low-resource Indic NLP, since data quality, script coverage, transliteration, and evaluation often matter more than adding parameters.
Core frameworks for model training
PyTorch
PyTorch remains the default foundation for much academic and commercial model development. Its eager execution model makes experiments relatively easy to inspect, while its ecosystem supports custom architectures, mixed precision, checkpointing, and GPU acceleration.
Use PyTorch when you need:
- Fast iteration on novel architectures
- Broad support for language, vision, audio, and multimodal models
- Integration with distributed training and optimisation libraries
- Access to current research implementations
JAX
JAX is valuable for teams that prioritise compilation, accelerator efficiency, and large-scale numerical workloads. It can deliver excellent performance on TPU and GPU infrastructure, but the functional programming model and debugging workflow require more specialised experience.
TensorFlow
TensorFlow remains relevant where teams already use its production tooling, mobile deployment stack, or established internal workflows. For a new large language model project, compare its ecosystem and team expertise against PyTorch and JAX rather than choosing it by default.
Model architectures, tokenisation, and data
The Hugging Face Transformers ecosystem provides reusable implementations, configuration formats, training utilities, and model distribution workflows. Its libraries are especially useful when adapting open-weight language, vision, speech, or multimodal models rather than implementing every component yourself.
For large datasets, pair model libraries with dedicated data tools:
- Hugging Face Datasets: Load, transform, stream, and version datasets.
- WebDataset: Read sharded samples efficiently from object storage.
- Apache Arrow and Parquet: Store columnar data for repeatable processing.
- SentencePiece and Hugging Face Tokenizers: Train and benchmark tokenizers.
- DataTrove and similar filtering pipelines: Deduplicate, filter, and prepare web-scale corpora.
Do not treat scraped text as training-ready data. Build a pipeline that records source, licence, language, quality score, personal-data handling, and transformation history. For Indian deployments, test whether the corpus represents regional scripts, code-switching, spelling variation, speech-derived text, and domain-specific terminology. A smaller, clean, legally usable corpus can be more valuable than a larger noisy crawl.
Distributed training and memory efficiency
Foundation model training is usually limited by memory movement, communication, and data throughput—not just raw FLOPS. Common building blocks include:
- DeepSpeed: ZeRO optimisations, pipeline parallelism, and memory-saving training features.
- torchtitan and native PyTorch distributed tools: Practical patterns for scaling modern transformer training.
- Megatron-LM: Tensor, pipeline, and data parallelism for large transformer workloads.
- FSDP: Shard parameters, gradients, and optimizer states across workers.
- Accelerate: Simplify multi-GPU and multi-node training configurations.
- FlashAttention: Reduce attention memory use and improve throughput where supported.
Track tokens per second, GPU utilisation, checkpoint size, failed-job recovery time, and cost per training run. A cluster that appears cheaper on hourly rates can become expensive if storage, networking, idle capacity, or repeated failed experiments are ignored. Indian teams should benchmark on the actual GPU types and cloud regions they can reliably access rather than relying on published performance figures.
Fine-tuning and alignment
Parameter-efficient methods make adaptation practical on smaller clusters. LoRA, QLoRA, adapters, and quantisation-aware workflows update a small portion of the model while preserving the base weights. Libraries such as PEFT, TRL, and bitsandbytes support supervised fine-tuning, preference optimisation, and low-memory experiments.
Define the target behaviour before training. A dataset should include realistic inputs, expected outputs, edge cases, refusal examples, and a clear split between training and evaluation data. For an assistant serving Indian users, test unsupported languages, mixed-language prompts, sensitive financial or health queries, and low-bandwidth usage—not only polished English benchmarks.
Evaluation, safety, and reproducibility
A model is not ready because its loss curve improves. Build an evaluation harness that measures:
- Task accuracy and calibration
- Hallucination and citation behaviour
- Robustness to spelling, scripts, code-switching, and dialect variation
- Toxicity, privacy leakage, and unsafe instruction following
- Latency, memory use, throughput, and cost per request
Use tools such as lm-evaluation-harness, EleutherAI evaluation suites, and task-specific test sets, but do not rely on public leaderboards alone. Maintain versioned datasets, prompts, model checkpoints, configuration files, and random seeds. Publish a model card covering intended use, limitations, training data categories, licences, known risks, and evaluation results.
Teams building public-interest or education products can also learn from open-source AI projects for student developers, particularly the discipline of documenting setup, constraints, and reproducible experiments.
Inference and deployment
Training tools and serving tools solve different problems. For production inference, evaluate:
- vLLM: High-throughput serving with continuous batching and broad LLM support.
- Text Generation Inference: A production-oriented serving stack for transformer models.
- llama.cpp: Efficient local and CPU-friendly inference for supported quantised models.
- SGLang: Structured generation and efficient serving for suitable workloads.
- ONNX Runtime, TensorRT-LLM, and OpenVINO: Hardware-specific optimisation options.
Choose based on traffic shape, context length, model size, hardware, streaming requirements, and reliability targets. Quantisation can reduce cost and memory, but always re-run quality and safety tests after quantising. For applications serving the next wave of Indian internet users, latency, intermittent connectivity, language coverage, and predictable operating costs deserve as much attention as benchmark scores; see this guide to building AI apps for the next billion users in India.
A practical build sequence
1. Establish a baseline with an open-weight model and a small, representative evaluation set.
2. Audit data provenance, licences, privacy risks, language coverage, and duplication.
3. Prototype retrieval or prompting before committing to fine-tuning.
4. Run a small parameter-efficient fine-tuning experiment with held-out tests.
5. Profile memory, throughput, and cost on the target hardware.
6. Add safety tests, monitoring, rollback procedures, and model documentation.
7. Scale training only after the data and evaluation pipeline are stable.
Open source reduces barriers, but it does not remove engineering responsibility. Licences may restrict commercial use, redistribution, or certain applications; “open weights” is not always equivalent to an OSI-approved open-source licence. Review every model, dataset, and dependency before release.
FAQ
Do I need to train a foundation model from scratch?
Usually not. Fine-tuning, continued pre-training, retrieval, or distillation can deliver a specialised system at a fraction of the data and compute cost.
What is the best framework to start with?
PyTorch is a practical default because of its research adoption and ecosystem. Choose JAX or TensorFlow when accelerator support, existing expertise, or deployment requirements justify them.
Can a small Indian startup build a useful model?
Yes, if “useful” is narrowly defined. A language-specific, domain-specific, or workflow-specific model can be valuable without competing with the largest general-purpose models. Start with a measurable user problem and a high-quality evaluation set.
Are open-weight models automatically safe to use commercially?
No. Check the model, training-data, and dependency licences, along with privacy, security, and sector-specific obligations. Record those decisions before deployment.
For more India-focused examples, explore Indian open-source AI developer projects and the open-source AI projects for beginners. If your team is building a credible AI product or research system, apply for AI grants from AI Grants India to explore funding support.