Why Indian-language maths needs a dedicated benchmark
A model can appear strong on English maths while failing on a Hindi word problem, a Tamil fraction question, or a Kannada prompt containing local number words. The difficulty is not limited to arithmetic. Models must parse language, preserve quantities, understand units, follow culturally familiar phrasing, and produce a verifiable answer.
A useful benchmark therefore separates language understanding, mathematical reasoning, and answer formatting. It should also expose performance differences between languages rather than hiding them inside one overall score. This matters for education products, assessment tools, public-service assistants, and research on low-resource Indic NLP. For background on data scarcity, tokenisation, and evaluation issues, see this builder’s guide to low-resource Indic NLP.
Define the benchmark before choosing a model
Start with a written evaluation specification. Record:
- Languages: Choose one or more languages and state the script, dialect, and spelling conventions. Do not treat “Hindi” or “Bengali” as a single uniform data source.
- Maths coverage: Include arithmetic, algebra, geometry, percentages, ratios, measurement, probability, and multi-step word problems according to the intended use case.
- User level: Label questions by school grade or difficulty. A benchmark for Class 5 learning support should not be compared directly with one for competitive-exam preparation.
- Output contract: Decide whether the model must return a final number, a multiple-choice option, a symbolic expression, or a worked solution.
- Allowed tools: Run separate tracks for no-tool reasoning, calculator use, and code execution. Mixing them makes results difficult to interpret.
Create a held-out test set that is never used for prompt tuning or fine-tuning. Keep a development set for iteration and a private test set for final reporting. Deduplicate by meaning, not only by exact text: translated versions of the same problem can otherwise leak between splits.
Build and audit the dataset
Hugging Face Datasets makes it practical to store, version, and load benchmark data, but the platform does not guarantee that a dataset is mathematically or linguistically sound. Every item needs review.
A useful record might contain:
{
"id": "hi_arithmetic_0042",
"language": "hi",
"script": "Devanagari",
"domain": "arithmetic",
"difficulty": "grade_6",
"question": "...",
"answer": "42",
"solution": "...",
"source": "author-created",
"license": "CC BY 4.0"
}For each question, verify the reference answer with an independent solver or symbolic calculation. Use native-language reviewers to check that translation has not changed the quantities, gender, tense, units, or implied relationships. Track whether numerals appear in Arabic digits, native digits, or words; this allows you to measure robustness instead of accidentally rewarding one formatting choice.
Avoid translating an English benchmark mechanically and calling it an Indic benchmark. Commission original items where possible, document translator qualifications, and include realistic Indian contexts without making culture a proxy for difficulty. Record licence, provenance, annotator agreement, and known limitations in the dataset card.
Select models and establish fair comparison
Compare models that are appropriate for the task: multilingual foundation models, Indic-focused open models, and strong English baselines. Keep the evaluation conditions consistent across models:
- Use the same prompt template and number of demonstrations.
- Fix decoding settings, or report temperature, top-p, maximum tokens, and random seeds.
- Keep model versions and quantisation settings in the experiment log.
- Distinguish base, instruction-tuned, and fine-tuned checkpoints.
- Report parameter count, context length, hardware, and inference cost.
For a builder, the best model is not always the one with the highest accuracy. Latency, memory, licence terms, privacy, and deployment cost matter. The Indian open-source AI developer projects guide is useful when deciding whether to build on a community checkpoint or a commercial API.
Run evaluation on Hugging Face
Install the core libraries and load a versioned dataset. For generative maths, use AutoModelForCausalLM rather than a sequence-classification head. A minimal evaluation pattern is:
from datasets import load_dataset
from transformers import pipeline
DATASET = "org/indic-math-benchmark"
MODEL = "your-org/your-model"
data = load_dataset(DATASET, split="test")
generator = pipeline("text-generation", model=MODEL, device_map="auto")
predictions = []
for item in data:
prompt = (
"Solve the problem. End with FINAL: followed by only the answer.\n"
f"Language: {item['language']}\n"
f"Problem: {item['question']}\n"
)
output = generator(prompt, max_new_tokens=256, do_sample=False)[0]["generated_text"]
predictions.append({"id": item["id"], "prediction": output})In production evaluation, add batching, timeouts, retry handling, and deterministic logging. Save the exact prompt, raw output, parsed answer, model revision, and runtime. Upload predictions and metrics as an evaluation artifact rather than relying on a manually copied spreadsheet.
Score more than exact match
Exact-match accuracy is a useful headline metric when answers are normalised, but it is insufficient for generated solutions. Add:
- Numerical accuracy: Parse equivalent forms such as
0.5,1/2, and50%where the task permits them. - Unit-aware scoring: Mark
5 kgand5differently when units are required. - Multiple-choice accuracy: Score the selected option separately from any explanation.
- Solution validity: Check intermediate steps with a verifier or symbolic tool; a correct final answer reached through invalid reasoning is a separate category.
- Format compliance: Measure whether the model follows the requested output structure.
- Calibration: Compare confidence with correctness if the system will decide when to defer to a teacher or human reviewer.
Report macro averages by language, maths domain, difficulty, and script. Include confidence intervals through bootstrap resampling. A single pooled score can conceal a serious failure in Marathi or Malayalam, especially when English or Hindi dominates the test set.
Analyse errors and publish reproducible results
Create an error taxonomy: translation error, number extraction error, operation selection error, arithmetic error, unsupported assumption, hallucinated explanation, and formatting failure. Review a balanced sample of failures with native speakers and maths educators. Plot a confusion matrix for multiple-choice tasks and compare performance on native-script, transliterated, and mixed-script prompts.
Publish the dataset card, evaluation script, prompt templates, model revisions, hardware details, and known contamination checks. If you fine-tune, keep the test set private or document exactly how it was protected. For education deployments, test with the same age groups and device constraints you expect in the field; platforms aimed at schools may also benefit from the lessons in this guide to interactive live learning platforms for Indian schools.
A practical release checklist
Before claiming that a model is strong at Indian-language maths, confirm that you have:
- Native-language review for every test language.
- Verified answers and documented licences.
- Separate development and test data with leakage checks.
- Results broken down by language, domain, difficulty, and script.
- Both final-answer and reasoning-quality measures.
- Reproducible code and pinned model and dataset revisions.
- Safety guidance stating that outputs require human review in high-stakes education.
The most credible benchmark is not the one with the largest dataset. It is the one that makes failures visible, comparisons fair, and improvements measurable. Hugging Face can provide the distribution and tooling; benchmark quality still depends on careful Indian-language data work, transparent evaluation, and disciplined reporting.