Gujarati evaluation needs more than a single accuracy number. A model may produce fluent Gujarati while ignoring constraints, answering in the wrong format, or relying on English reasoning that changes the result. This guide presents a reproducible workflow for benchmarking Gujarati instruction following on IndicEval with Hugging Face tools.
Use the process to compare checkpoints, validate a fine-tune, or establish a baseline before deploying a Gujarati assistant. For wider context on dataset selection, see this Indian language LLM benchmark datasets guide, and use the same discipline when comparing Gujarati results with other Indian languages through this multilingual LLM benchmarking framework.
What IndicEval should measure
Treat IndicEval as an evaluation protocol, not merely a dataset download. Before running a model, identify:
- The Gujarati instruction-following split and its version or commit.
- Task categories, such as classification, extraction, rewriting, question answering, and constrained generation.
- Whether each item has a fixed answer, reference answers, or rubric-based evaluation.
- The expected output language, script, format, and length.
- Any licensing or access restrictions on the data and model.
Instruction following has at least two dimensions: task correctness and instruction compliance. A response can be factually correct but still fail because it ignores a requested number of points, returns prose instead of JSON, or answers in Hindi or English rather than Gujarati. Preserve these dimensions separately in your report.
Prepare a reproducible Hugging Face environment
Pin the software and hardware configuration. Generation behavior can change with library versions, quantisation, kernels, or decoding defaults. A typical starting environment is:
python -m venv .venv
source .venv/bin/activate
pip install "transformers>=4.45" "datasets>=2.20" accelerate evaluate sentencepiece torchInstall the exact IndicEval package or clone the official repository specified by its documentation. Do not substitute a placeholder repository or silently modify evaluation scripts. Record the IndicEval commit, Python version, GPU type, CUDA version, model revision, and tokenizer revision in a README or run manifest.
For a production-grade comparison, keep one configuration file containing:
- Model ID and revision.
- Dataset ID, split, and revision.
- Prompt template.
- Maximum input and output tokens.
- Decoding parameters.
- Batch size and precision.
- Random seed and evaluation timestamp.
Load the model and tokenizer correctly
Choose a causal instruction-tuned model that supports Gujarati, then verify its actual language coverage. A model card mentioning multilingual training does not guarantee useful Gujarati instruction following. Inspect sample generations before investing in a full run.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "your-org/your-instruct-model"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision="main")
model = AutoModelForCausalLM.from_pretrained(
model_id,
revision="main",
torch_dtype=torch.bfloat16,
device_map="auto",
)
model.eval()Use the model's official chat template when one is available. Avoid manually adding system, user, and assistant markers unless the model card requires it. A mismatched template can make a capable model appear weak and makes comparisons unfair.
def format_prompt(example):
messages = [
{"role": "user", "content": example["instruction"]}
]
return tokenizer.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)If the benchmark specifies a prompt format, follow it exactly. Do not translate Gujarati prompts into English for convenience: that changes the task being measured.
Validate and prepare the Gujarati split
Load the documented split with the Hugging Face datasets library, then inspect it before generation. Check for missing instructions, duplicate IDs, malformed Unicode, unexpected Romanised Gujarati, and examples that contain answers in the prompt.
from datasets import load_dataset
data = load_dataset("your-org/indic-eval", name="gujarati", revision="main")
print(data)
print(data["test"][0])Gujarati text should normally use Gujarati Unicode code points, but do not assume every item is clean. Measure the proportion of Gujarati-script characters, identify mixed Gujarati-English prompts, and preserve punctuation and numerals. Normalise only when the official evaluator does so. Aggressive normalisation can erase meaningful differences in names, numbers, or formatting.
Keep the official test set untouched. If you need prompt or parser development, create a local validation slice and never tune generation settings against the final test answers.
Run controlled generation
Use deterministic decoding for a primary score so that model comparisons are stable. Typical settings are do_sample=False, a defined max_new_tokens, and an explicit stopping policy. Do not truncate Gujarati outputs solely because they contain more bytes than English outputs; tokenisation and script affect length.
inputs = tokenizer(
prompt,
return_tensors="pt",
truncation=True,
max_length=4096,
).to(model.device)
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=256,
do_sample=False,
pad_token_id=tokenizer.eos_token_id,
)
answer = tokenizer.decode(
output[0][inputs["input_ids"].shape[1]:],
skip_special_tokens=True,
).strip()Run a second, clearly labelled sensitivity test only if useful—for example, temperature sampling or multiple seeds. Never mix sampled and greedy results into one headline score. Record latency, tokens per second, memory use, and failures alongside quality metrics; these matter when selecting a model for an Indian-language product.
Score correctness and compliance separately
Use the evaluator supplied by IndicEval wherever possible. Match the metric to the task:
- Exact match or normalised match for short, fixed-format answers.
- Accuracy and macro-F1 for classification.
- Structured-field accuracy for JSON or key-value outputs.
- Reference-based metrics only when references are appropriate.
- Rubric or judge-based scoring for open-ended responses, with a fixed rubric and sampled human audit.
BLEU alone is a poor measure of instruction following. It can reward surface overlap while missing whether the response obeyed constraints. Add Gujarati-specific checks for script, required entities, number preservation, refusal format, and answer length. If an LLM judge is used, blind the model identity, keep the rubric in Gujarati where possible, and report judge model, prompt, temperature, and agreement with human ratings.
Analyse errors, not just rankings
Create an error taxonomy and label a representative sample of failures:
- Wrong answer despite following the requested format.
- Correct answer in the wrong language or script.
- Omitted constraints, fields, or steps.
- Hallucinated facts or unsupported citations.
- Prompt leakage or copied reference text.
- Unsafe or culturally inappropriate response.
- Truncation, empty output, or malformed JSON.
Break results down by task type, input length, script purity, and difficulty. Compare Gujarati performance with a related-language or multilingual baseline, but avoid treating language ranking as a proxy for model quality. For speech interfaces, use a separate speech-to-text accuracy benchmark workflow, because recognition errors and instruction-following errors should not be conflated.
Gujarati-English transfer can also expose failure modes in translation-heavy tasks. When such tasks are included, consult this practical guide to Gujarati-English neural machine translation models rather than interpreting translation scores as general instruction-following ability.
Report results so others can reproduce them
Publish a table containing model revision, IndicEval revision, Gujarati split, prompt template, decoding settings, aggregate scores, per-category scores, failure rate, and confidence intervals where applicable. Include at least 20–50 manually inspected examples, with sensitive content redacted responsibly.
State whether scores came from zero-shot prompting, few-shot prompting, supervised fine-tuning, or retrieval augmentation. If the model was fine-tuned, document the data source and contamination checks. A useful benchmark report should let another team rerun the exact command and explain why a model failed—not merely announce a leaderboard position.
FAQ
Is Gujarati fluency enough to pass IndicEval? No. The model must answer correctly and follow explicit constraints, formats, language requirements, and safety instructions.
Should I translate the benchmark into English first? No. Evaluate the Gujarati input as published. Translation creates a different benchmark and hides script- and language-specific failures.
Can I use a general LLM judge? Yes, but disclose the judge configuration, calibrate it against human ratings, and retain deterministic checks for objective fields.
What should I do after finding weaknesses? Build a clean Gujarati error set, fix the prompt or data issue, rerun the untouched test split, and report both the original and improved results. For a small deployment prototype, a small Gujarati chatbot can provide realistic product-level checks beyond the benchmark.