QLoRA makes domain adaptation practical when you do not have access to a large multi-GPU cluster. It combines 4-bit quantisation of a pretrained language model with trainable low-rank adapters, so most weights stay frozen while the model learns patterns from your data. For an Indian-language assistant, customer-support bot, education tool, or government-service workflow, this can reduce cost without requiring you to train a model from scratch.
This guide shows how to fine-tune with QLoRA on Hugging Face using Indian datasets. The examples target causal language models and instruction tuning. Exact model names, VRAM requirements, and library APIs change, so pin tested versions and check the model card before starting.
Choose the right task and base model
QLoRA is most useful when you have a clear adaptation objective. Decide whether you are training:
- An instruction-following assistant
- A Hindi, Tamil, Bengali, Marathi, Telugu, or multilingual response model
- A domain model for law, healthcare, finance, agriculture, or education
- A structured text generator that must return JSON or a fixed format
Start with a base model that already supports the target scripts and languages. A smaller multilingual or Indic-capable model is often a better choice than a larger English-first model. Review its license, commercial-use terms, tokenizer coverage, context length, and existing safety limitations before collecting data.
For broader implementation decisions, best practices for fine-tuning LLMs on custom data covers dataset design, validation, and deployment trade-offs that apply beyond QLoRA.
Prepare an Indian dataset that is actually trainable
Do not treat “Indian dataset” as a quality guarantee. Data may contain code-switching, spelling variation, transliteration, duplicated web text, machine translations, or personally identifiable information. Build a dataset that reflects the product’s real users and permissions.
Useful sources can include AI4Bharat resources, Indic language benchmarks, licensed public records, synthetic examples reviewed by native speakers, and your own support or workflow data. Before training:
- Check rights and consent: confirm that text can be used for model training and redistribution.
- Remove sensitive information: redact phone numbers, Aadhaar-like identifiers, addresses, account details, and private conversations.
- Preserve language labels: record language, script, region, and whether text is transliterated.
- Deduplicate: remove repeated prompts, near-duplicate answers, and train-validation overlap.
- Balance the data: avoid allowing Hindi or English to overwhelm smaller target languages.
- Keep a held-out test set: have native speakers review representative examples before training.
For instruction tuning, store examples in a consistent format such as JSONL:
{"instruction":"मौसम के आधार पर किसान को क्या सावधानी रखनी चाहिए?","input":"बारिश अगले दो दिन जारी रहेगी।","output":"खेत में जल निकासी की जाँच करें..."}If your dataset is conversational, use messages with role and content fields. Apply the base model’s chat template rather than manually inventing separators. Incorrect formatting is a common reason for apparently successful training that produces poor responses.
Install a reproducible Hugging Face environment
Use a CUDA-enabled machine with sufficient VRAM, or a rented GPU environment. A 4-bit setup can make a 7B-class model feasible on a single GPU, but memory still depends on sequence length, batch size, optimizer, and checkpoints. QLoRA is not a promise that every model will fit on a laptop GPU.
pip install -U transformers datasets accelerate peft bitsandbytes trl
accelerate configPin the working versions in requirements.txt, record the GPU and CUDA versions, and authenticate with Hugging Face only when the model or dataset requires gated access. Never place tokens in notebooks committed to a repository.
Load the model in 4-bit mode
The following pattern uses BitsAndBytesConfig and current PEFT conventions. Replace the model ID with one whose license and language coverage suit your project.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer, BitsAndBytesConfig
model_id = "your-compatible-model"
quant_config = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
bnb_4bit_use_double_quant=True,
bnb_4bit_compute_dtype=torch.bfloat16,
)
model = AutoModelForCausalLM.from_pretrained(
model_id,
quantization_config=quant_config,
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_tokenUse float16 instead of bfloat16 if your GPU does not support BF16. Quantisation lowers memory use, but it can affect quality and does not eliminate the need for evaluation.
Configure the LoRA adapter
QLoRA trains adapter weights attached to selected transformer modules. Module names differ between architectures, so inspect the model or consult its documentation rather than copying a target list blindly.
from peft import LoraConfig, prepare_model_for_kbit_training
model = prepare_model_for_kbit_training(model)
lora_config = LoraConfig(
r=16,
lora_alpha=32,
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
)Higher rank can improve capacity but increases trainable parameters. Start with r=8 or r=16, then compare against a held-out set. Print trainable parameters before training; only a small fraction should be unfrozen.
Format, tokenise, and train
With TRL’s supervised fine-tuning trainer, format each record using the model’s chat template. A simplified example is:
from datasets import load_dataset
from trl import SFTTrainer, SFTConfig
train = load_dataset("your-org/your-indic-dataset", split="train")
valid = load_dataset("your-org/your-indic-dataset", split="validation")
def format_example(example):
return tokenizer.apply_chat_template(
example["messages"], tokenize=False, add_generation_prompt=False
)
args = SFTConfig(
output_dir="./indic-qlora-adapter",
max_seq_length=2048,
per_device_train_batch_size=1,
gradient_accumulation_steps=16,
learning_rate=2e-4,
num_train_epochs=2,
logging_steps=10,
eval_strategy="steps",
eval_steps=100,
save_steps=100,
bf16=True,
gradient_checkpointing=True,
packing=False,
report_to="none",
)
trainer = SFTTrainer(
model=model,
processing_class=tokenizer,
train_dataset=train,
eval_dataset=valid,
peft_config=lora_config,
formatting_func=format_example,
args=args,
)
trainer.train()
trainer.save_model("./indic-qlora-adapter")Library arguments change across releases; if processing_class, eval_strategy, or max_seq_length is rejected, consult the installed TRL version’s documentation. Keep sequences within the model context window and use gradient accumulation to simulate a larger batch. Monitor loss, GPU memory, tokens per second, and checkpoint size.
Evaluate Indian-language quality, not just loss
Validation loss cannot tell you whether a model handles respectful address, regional vocabulary, code-switching, or numerals correctly. Build an evaluation set separated by language, script, task, and difficulty. Test:
- Factual accuracy and refusal behaviour
- Translation and transliteration consistency
- Named entities, dates, currency, and Indian numbering formats
- Toxic, biased, casteist, communal, or privacy-invasive outputs
- Robustness to spelling variation and mixed Hindi-English prompts
- Instruction following and structured output validity
Use native-speaker review alongside automated metrics. Compare the QLoRA model with the untouched base model and a retrieval-augmented baseline. For high-stakes uses, keep retrieval and human review in the system rather than expecting fine-tuning to memorise changing facts. If you are building language tooling, open-source vision-language models for Indian languages may also be relevant when inputs include documents, images, or scanned forms.
Save, merge, and deploy safely
Save the adapter and tokenizer separately from the base model. This produces a compact artifact and makes it easier to update or roll back. Merge the adapter only when your serving stack requires a standalone model, and test quality after merging. Publish a model card that documents training data provenance, languages, known limitations, evaluation results, quantisation settings, and the exact base-model license.
Before production deployment, add prompt and output logging with privacy controls, rate limits, abuse monitoring, and a rollback path. For an Indian startup, adapter-based deployment can lower storage and experimentation costs, but serving costs still depend on traffic, context length, concurrency, and GPU availability. Teams exploring local AI development can also review Indian open-source AI developer projects for reusable tooling and community practices.
Common mistakes to avoid
- Training on unlicensed scraped content
- Mixing raw completions and chat-formatted examples
- Evaluating only in English
- Letting duplicated Hindi or English data dominate the corpus
- Using deprecated
prepare_model_for_int8_trainingfor a 4-bit workflow - Selecting LoRA target modules without checking the architecture
- Setting a learning rate so high that the model forgets general capability
- Publishing adapter weights without documenting the base model and data limits
QLoRA is best viewed as a focused adaptation method, not a substitute for product engineering. Strong data governance, language-aware evaluation, and a carefully chosen base model will matter more than increasing rank or training for extra epochs.