Indian-language AI systems can produce fluent answers that are factually wrong, culturally inappropriate, or unsupported by the supplied context. A benchmark should expose those failures—not hide them behind a single BLEU or perplexity score. This guide lays out a reproducible workflow for measuring hallucination on Hugging Face across Indic languages, scripts, tasks, and deployment conditions.
The approach is useful for chatbots, retrieval-augmented generation (RAG), translation, summarisation, voice agents, and public-service applications. For background on data constraints, tokenisation, and evaluation design, see this builder’s guide to low-resource Indic NLP.
Define hallucination before writing tests
Hallucination is task-dependent. Separate at least four failure types:
- Unsupported claim: The answer introduces information absent from a provided source.
- Contradiction: The answer conflicts with the source or a verified fact.
- Fabrication: The model invents citations, names, schemes, statistics, or events.
- Instruction failure: The response ignores uncertainty, changes the user’s meaning, or answers in the wrong language or script.
Create a benchmark card before running experiments. Record the languages, dialects, scripts, tasks, model versions, decoding settings, context sources, and pass criteria. Do not combine Hindi written in Devanagari with Hindi transliterated into Latin script unless the distinction is intentional; script changes can materially alter performance.
Build a representative Indic test set
Start with a small, high-quality seed set rather than a large noisy crawl. Include prompts from real product traffic, carefully reviewed public documents, and adversarial examples. For each item, store:
- The prompt and requested language or script
- A source passage or structured fact record
- One or more acceptable answers
- Claims that must be present, absent, or qualified
- Metadata such as language, domain, region, and difficulty
- Annotation guidance and an escalation label for ambiguity
Cover code-mixed queries, spelling variation, transliteration, honorifics, numerals, dates, names, and local entities. Test multiple registers: formal Hindi, conversational Hindi-English, Tamil with English product terms, and regional vocabulary. Include public-health, education, agriculture, government-scheme, financial, and legal examples where an apparently minor error can cause harm.
Use balanced labels. A benchmark containing only answerable questions cannot measure whether a model refuses or expresses uncertainty when evidence is missing. Add unanswerable prompts, conflicting sources, outdated facts, and near-miss entities. Maintain a hidden test split so model developers cannot tune directly against every example.
Set up Hugging Face evaluation infrastructure
Install the core libraries and pin their versions for reproducibility:
pip install -U transformers datasets evaluate accelerate sentencepieceLoad data with datasets, keep source fields separate from prompts, and avoid leaking reference answers into generation inputs. A simple dataset schema might include id, language, script, prompt, context, reference_answer, claims, and answerable.
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "your-org/your-indic-model"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
data = load_dataset("your-org/indic-hallucination-set")Generate under fixed settings first: temperature, top-p, maximum tokens, system prompt, and random seed. Then run a separate robustness pass with several seeds and decoding configurations. Report both the mean and the worst-case result; a model that succeeds only under one carefully selected setting is not production-ready.
Measure factuality instead of relying on overlap scores
BLEU and ROUGE can be useful for translation and summarisation, but they are weak hallucination measures. A correct paraphrase may have low lexical overlap, while a fluent fabricated answer may score well. Perplexity measures token prediction, not factual correctness.
Use a layered scorecard:
- Claim support rate: Percentage of answer claims entailed by the supplied context or verified knowledge base.
- Contradiction rate: Percentage containing a claim that conflicts with the evidence.
- Unsupported-claim rate: Percentage adding unverifiable details.
- Abstention quality: Whether the model declines appropriately when the answer is not supported.
- Answerability accuracy: Whether it distinguishes answerable from unanswerable prompts.
- Language and script compliance: Whether it responds in the requested Indic language, register, and script.
- Entity and numeric accuracy: Exact correctness of names, dates, quantities, units, and scheme eligibility conditions.
For RAG systems, evaluate retrieval separately from generation. A model cannot cite evidence it never receives, so log retrieved passages, rank, source date, and context length. Score citation precision and whether each important claim is actually supported by the cited passage.
Add automated checks carefully
Automated entailment or judge-model scoring can accelerate triage, but Indic-language quality varies considerably by model and domain. Translate-everything-to-English evaluation may erase politeness, ambiguity, morphology, and culturally specific meaning. Prefer multilingual evaluators tested on the target language, and validate them against a human-labelled sample.
Normalise Unicode before comparison, but preserve the original output for review. Use language identification, script detection, named-entity checks, numeric extraction, and citation validation as targeted tests—not as substitutes for factual review. Flag outputs for human assessment when evaluators disagree or confidence is low.
A practical scoring record can include binary labels for support, contradiction, fabrication, refusal, language compliance, and safety. Aggregate by language, script, task, domain, and prompt type. Never publish only an overall average: a strong aggregate score can conceal severe failure in a low-resource language.
Run human evaluation with Indic reviewers
Human review remains essential for semantic equivalence and cultural context. Recruit fluent speakers who understand the evaluation domain, not only general bilingual annotators. Give them the source text, prompt, model response, and a short rubric. Ask reviewers to label each claim as supported, contradicted, unsupported, or unclear, and to mark harmful implications separately.
Use double annotation for every test item and adjudicate disagreements. Measure agreement, but do not treat disagreement as annotator failure; ambiguity often reveals a weak prompt or an under-specified reference. Pay reviewers fairly, document dialect and proficiency, and remove personally identifiable information from production samples.
Compare models and report results responsibly
Benchmark at least one baseline, one current open model, and the production candidate. Keep prompts and decoding settings identical. Publish the model revision, dataset version, evaluation code, hardware constraints, latency, and cost per request where possible. Include confidence intervals or bootstrap estimates for headline rates.
A useful report table breaks results down by:
- Language and script
- Answerable versus unanswerable prompts
- RAG versus no-context generation
- Domain and risk level
- Short versus long context
- Standard versus code-mixed input
- Automated versus human-verified labels
Track regressions in continuous evaluation. Every model, prompt, tokenizer, retrieval index, or safety-policy change should trigger the benchmark. Store failed examples in a review queue, but keep a frozen test set to prevent overfitting.
Common pitfalls for Indian-language benchmarks
- Treating English translation as the ground truth for every language
- Mixing dialects and scripts without reporting the difference
- Using synthetic data without checking naturalness and factual validity
- Comparing models with different context windows or generation limits
- Counting refusal as success even when the question is answerable
- Ignoring code-mixing, transliteration, and speech-to-text errors
- Publishing scores without sample sizes or annotation details
For voice products, evaluate the complete pipeline: audio transcription, language identification, retrieval, generation, and text-to-speech. This matters for builders exploring voice agent services for Indian businesses, where a hallucination can be amplified by an apparently confident spoken response.
A practical release gate
Before deployment, set thresholds by risk tier rather than one universal number. For a low-risk content assistant, you may tolerate some unsupported phrasing with clear uncertainty. For health, finance, education, or government workflows, require near-zero contradiction on critical facts, verified citations, reliable abstention, and human escalation.
A release checklist should confirm that the model:
- Passes every critical factuality test in each supported language
- Does not invent sources, schemes, people, or statistics
- Handles missing and conflicting evidence safely
- Preserves names, numbers, dates, and units
- Responds in the requested language and script
- Has monitoring, feedback capture, rollback, and incident ownership
FAQ
Is perplexity enough to measure hallucination?
No. Perplexity measures language-model fit, not whether generated claims are true or supported. Pair it with claim-level factuality and human review.
Should I use BLEU or ROUGE?
Use them for suitable translation or summarisation comparisons, but do not use them as the primary hallucination metric. Semantic entailment and contradiction checks are more informative.
How many examples are needed?
Begin with a reviewed pilot of roughly 100–300 examples per important language and task, then expand based on observed failure modes. Report sample sizes and uncertainty.
Can an LLM judge Indian-language outputs?
It can assist with triage, but validate its decisions against fluent human reviewers and audit disagreements, especially for dialects, code-mixing, and culturally specific claims.
Reliable Indic evaluation is infrastructure for safer products, not a one-time leaderboard exercise. If your team is building open models, datasets, or evaluation tools for India, explore Indian open-source AI developer projects and consider support through AI Grants India.