JSONL is a practical interchange format for Hugging Face datasets: one valid JSON object per line, readable as a stream and easy to inspect in version control. For Indian-language fine-tuning, however, writing one JSON object per line is only the beginning. You also need a stable schema, Unicode-safe processing, language-aware splits, licensing records, and checks that prevent leakage or low-quality examples from reaching training.
This guide covers classification, instruction tuning, and conversational data. It assumes Python, the Hugging Face datasets library, and a dataset stored locally or in object storage.
Choose the training format before writing JSONL
Your JSONL schema should match the model objective. Do not mix unrelated structures merely because they are all valid JSON.
- Classification: use fields such as
textandlabel. - Instruction tuning: use
instruction, optionalinput, andoutput, or a singlemessagesfield for chat models. - Continued pretraining: use a
textfield containing clean, licensed text, with optional metadata kept outside the training text. - Preference tuning: use fields such as
prompt,chosen, andrejected.
For a multilingual Indian dataset, add explicit metadata such as language, script, source_id, and license. Metadata helps auditing and stratified evaluation, but avoid feeding sensitive or irrelevant metadata to the model. Teams building serious pipelines should also review best practices for fine-tuning LLMs on custom data before selecting the objective.
A classification record might look like this:
{"text":"यह सेवा बहुत उपयोगी है।","label":"positive","language":"hi","script":"Devanagari","source_id":"review-001"}An instruction-tuning record can use:
{"instruction":"ग्राहक को भुगतान की स्थिति बताइए।","input":"लेन-देन आईडी: TXN-4821","output":"आपका भुगतान सत्यापन में है।","language":"hi","source_id":"support-014"}For chat models, prefer the model's documented conversation format:
{"messages":[{"role":"system","content":"You are a helpful support assistant."},{"role":"user","content":"मेरा ऑर्डर कब आएगा?"},{"role":"assistant","content":"कृपया अपना ऑर्डर नंबर साझा करें।"}],"language":"hi"}Prepare Indian-language data carefully
Indian datasets often contain multiple scripts, transliterated text, code-switching, abbreviations, and regional spelling variation. Cleaning should improve consistency without erasing meaningful language patterns.
- Preserve Unicode characters; never force ASCII conversion.
- Normalize Unicode consistently, usually with NFC, while retaining the original record if auditability matters.
- Keep punctuation that affects meaning, including danda characters such as
।and Tamil or Bengali punctuation where used. - Decide whether emojis, hashtags, phone numbers, URLs, and currency symbols are useful for the task.
- Record language and script separately. Hindi written in Latin script is not the same as Hindi written in Devanagari.
- Detect and review code-mixed examples rather than silently discarding them.
- Remove duplicated, templated, or near-duplicated records before splitting the data.
Do not lowercase every language by default. Lowercasing can be harmless for some Latin-script classification tasks but is unnecessary or damaging for scripts where case is not relevant and for data containing names, product codes, or English terms. If your corpus includes medical, financial, or public-service content, establish a review process for privacy and factual accuracy. Data veracity infrastructure for high-stakes AI is a useful companion topic for this stage.
Convert records to UTF-8 JSONL
Use ensure_ascii=False so Indian scripts remain readable in the file. Use compact separators if storage matters, but readable output is often better during development.
import json
from pathlib import Path
records = [
{
"instruction": "ग्राहक को भुगतान की स्थिति बताइए।",
"input": "लेन-देन आईडी: TXN-4821",
"output": "आपका भुगतान सत्यापन में है।",
"language": "hi",
"script": "Devanagari",
"source_id": "support-014",
}
]
output_path = Path("indian_support_train.jsonl")
with output_path.open("w", encoding="utf-8", newline="\n") as file:
for record in records:
file.write(json.dumps(record, ensure_ascii=False) + "\n")Each record must occupy exactly one physical line. Escape embedded line breaks inside string values automatically through json.dumps; do not manually concatenate JSON strings. Keep secrets, raw phone numbers, Aadhaar numbers, email addresses, and other personal data out of training files unless there is a documented lawful basis and a strong de-identification process.
Validate syntax, schema, and content
A file can be valid JSONL and still be unusable for training. Validate every line, required fields, field types, empty values, language codes, and label vocabulary.
import json
required = {"instruction", "input", "output", "language", "source_id"}
allowed_languages = {"hi", "en", "bn", "ta", "te", "mr", "gu", "kn", "ml", "pa", "or"}
with open("indian_support_train.jsonl", encoding="utf-8") as file:
for line_number, line in enumerate(file, start=1):
if not line.strip():
raise ValueError(f"Blank line at {line_number}")
try:
row = json.loads(line)
except json.JSONDecodeError as error:
raise ValueError(f"Invalid JSON on line {line_number}: {error}")
missing = required - row.keys()
if missing:
raise ValueError(f"Missing {missing} on line {line_number}")
if row["language"] not in allowed_languages:
raise ValueError(f"Unsupported language on line {line_number}")
if not all(isinstance(row[field], str) for field in required):
raise ValueError(f"Non-string field on line {line_number}")Then load the result with Hugging Face:
from datasets import load_dataset
dataset = load_dataset("json", data_files="indian_support_train.jsonl", split="train")
print(dataset)
print(dataset.features)
print(dataset[0])Inspect random rows and aggregate counts by language, script, source, and label. Check for rows with unusually long inputs, empty outputs, repeated prompts, or outputs that merely copy the input. For production pipelines, add automated tests to CI and retain a rejected-records file with reasons.
Split and tokenize without leakage
Create train, validation, and test splits by source, user, document, or conversation—not only by random row. Random splitting can place near-identical translations, support tickets, or turns from the same conversation in both training and evaluation. Keep a balanced test set for each important language and task category, including code-mixed and transliterated examples when they matter to users.
Choose a tokenizer compatible with the base model. A tokenizer trained primarily on English may produce inefficient token sequences for Indic scripts, increasing memory use and truncation. Measure token lengths by language before setting max_length. Do not truncate instructions or answers blindly; report how many examples are affected and revise the data or sequence length.
For chat fine-tuning, apply the model's chat template during preprocessing rather than inventing a new delimiter. For supervised instruction tuning, calculate loss on assistant responses when the training framework supports it, instead of treating system and user text as target labels.
Hugging Face loading and training checks
For a classification dataset, map string labels to stable integer IDs and use a model with the correct number of labels. For causal language modelling or chat fine-tuning, confirm that the collator, tokenizer, padding side, and special tokens match the base model. A JSONL file does not determine these settings.
Before a full run, perform a small smoke test:
- Load ten to one hundred records.
- Tokenize and inspect decoded samples.
- Run one training step and confirm loss is finite.
- Check that labels are not entirely padding or masked.
- Verify that evaluation examples are never included in training.
- Generate outputs in every priority language, not only English.
Use a reproducible data manifest containing the file hash, source licences, language counts, preprocessing version, split seed, and model revision. This makes it possible to explain what was trained and reproduce a result months later.
Common JSONL mistakes
- Multiple JSON objects on one line: write one record per line.
- Python dictionaries instead of JSON: use
json.dumps; JSON requires double quotes. - Escaped Indian scripts: use
ensure_ascii=Falseand UTF-8. - Inconsistent columns: standardize the schema before loading.
- Null or blank target text: reject or repair it before training.
- Unsupported chat structure: follow the base model's documented template.
- PII and copyrighted text: document consent, licence, retention, and removal procedures.
- Unbalanced languages: report per-language metrics and sample intentionally.
Final checklist
Before uploading the dataset to a hub or starting a costly run, confirm that every line parses, required fields are present, Unicode is preserved, duplicates are removed, personal data is handled appropriately, and splits prevent leakage. Compare language and script distributions, inspect token lengths, and run a small multilingual evaluation set.
The best JSONL dataset is not merely valid—it is traceable, representative, legally usable, and aligned with the model's training objective. If your application serves customers through speech or chat, also consider how the fine-tuned model will connect to voice agent services for Indian businesses and how response quality will be monitored after deployment.