Evaluation should answer a business question: can this model perform a defined task reliably, safely, quickly, and at an acceptable cost? A high score on a public leaderboard is useful context, but it does not prove that a model will handle your customers, documents, languages, or failure modes.
For Indian teams, the evaluation surface is wider still. A model may work well in English and fail on Hindi, Tamil, or code-switched queries; produce fluent answers about Indian law while citing outdated information; or appear affordable until long prompts and retries inflate inference costs. A useful evaluation programme measures these risks before launch and continues after deployment.
Start with a task-specific evaluation plan
Define the product task before selecting metrics. “Model quality” is too broad to measure consistently. A support assistant, coding copilot, document extractor, and voice bot need different test sets and acceptance thresholds.
Write down:
- Inputs: languages, document types, prompt lengths, spelling variation, and expected traffic.
- Outputs: answer, classification, JSON object, code patch, citation, or tool call.
- Failure cost: inconvenience, financial loss, privacy exposure, unsafe advice, or regulatory risk.
- Acceptance criteria: for example, 95% valid JSON, 90% citation-supported answers, or fewer than two critical safety failures per 10,000 requests.
- Operating constraints: latency target, model endpoint, context window, GPU capacity, and cost per request.
Build a golden dataset from real or carefully simulated examples. Start with 100–300 cases covering common requests, edge cases, adversarial prompts, ambiguous questions, and known production failures. Keep a private test split so prompt and model changes cannot be optimised against every example.
If the system serves Indian-language users, include native-speaker annotations rather than relying on translated English tests. Guidance on low-resource Indic natural language processing and low-resource language datasets for AI training in India is especially relevant when creating representative evaluation data.
Use public benchmarks as signals, not verdicts
Benchmarks help with initial model selection, but they should not be treated as production evidence. Common tests include:
- MMLU and newer multitask suites: broad academic and professional knowledge.
- GSM8K and harder mathematics sets: arithmetic and multi-step reasoning.
- HumanEval and MBPP: code-generation ability, usually measured with executable tests.
- TruthfulQA: resistance to common misconceptions and misleading prompts.
- Long-context tests: retrieval and instruction following over large inputs.
- Tool-use and structured-output tests: function selection, argument validity, and schema compliance.
Check dataset contamination, licence terms, language coverage, and test design before comparing scores. A benchmark score may reflect memorisation, narrow formatting conventions, or an evaluation setup unlike yours. Re-run a representative slice using the same temperature, system prompt, tool definitions, context length, and output limits used in production.
Measure quality on several dimensions
A single aggregate score hides important trade-offs. Report separate metrics for:
- Correctness: Is the answer factually and procedurally right?
- Instruction adherence: Did it follow requested format, length, language, and constraints?
- Groundedness: Are claims supported by supplied documents or tools?
- Completeness: Did it address every required part of the request?
- Robustness: Does performance hold under paraphrases, typos, long inputs, and prompt injection?
- Safety and privacy: Does it refuse harmful requests and avoid exposing sensitive information?
- Consistency: Does the same input produce acceptable results across repeated runs?
For extraction, calculate field-level precision, recall, and exact or normalised match rates. For classification, use confusion matrices and class-specific recall; accuracy alone can conceal poor performance on rare but important categories. For code, run tests, linting, security checks, and human review rather than judging only textual similarity.
Use LLM judges carefully
An LLM judge can score helpfulness, relevance, style, or pairwise preference at scale. Use explicit rubrics with observable criteria, require structured outputs, randomise response order, and blind the judge to model identity. Calibrate it against human-labelled examples and inspect disagreements.
Do not rely on a single judge or one overall score. Judge models can show position bias, verbosity preference, self-preference, and sensitivity to formatting. For high-risk tasks, combine automated judging with human review and deterministic checks. A judge should never be the sole authority for medical, legal, financial, safety, or privacy decisions.
A practical pattern is three-layer evaluation:
1. Programmatic checks: schema validity, citation presence, forbidden content, tool-call arguments, and latency.
2. Model-based checks: relevance, groundedness, tone, and completeness.
3. Human audits: sampled outputs, critical failures, and cases where automated graders disagree.
Evaluate RAG as retrieval plus generation
Retrieval-augmented generation fails in two distinct places: the retriever may return the wrong evidence, or the generator may misuse correct evidence. Evaluate them separately before combining results.
Track:
- Recall@k: whether the needed source appears in the retrieved set.
- Context precision: whether top-ranked passages are actually useful.
- Faithfulness: whether claims are supported by retrieved context.
- Answer relevance: whether the response addresses the question.
- Citation correctness: whether each citation supports the associated claim.
- Abstention quality: whether the system says it lacks evidence instead of guessing.
Frameworks such as Ragas can accelerate experiments, but validate their scores against human judgements and your domain’s risk tolerance. Test OCR errors, duplicate documents, stale versions, multilingual queries, access controls, and prompt injection embedded in retrieved content. For production RAG, log query, retrieved document IDs, scores, final answer, citations, and user feedback with appropriate privacy safeguards.
Include Indic-language and India-specific tests
English-only evaluation can overstate performance. Test native scripts, transliteration, mixed-language prompts, regional names, numerals, dates, honorifics, and speech-to-text noise. Compare both semantic quality and infrastructure behaviour: Indic tokenisation can increase token counts, latency, and cost even when the answer is correct.
Build test cases around Indian addresses, public schemes, education, healthcare, agriculture, finance, and local regulatory terminology only when your product needs them. Separate knowledge freshness from language quality: a fluent answer can still rely on outdated policy information. For teams adapting open models, fine-tuning Llama for Indian regional languages and open-source small language models for Hindi offer useful starting points, but every fine-tuned model needs fresh, task-specific testing.
Track production performance and safety
Offline quality is necessary but insufficient. Measure:
- Time to first token and end-to-end latency, including retrieval and tool calls.
- Throughput, queue time, timeout rate, and availability.
- Input and output tokens, cache hit rate, and cost per successful task.
- Retry, escalation, correction, and user-abandonment rates.
- Safety violations, prompt-injection success, data leakage, and unsupported claims.
Use a shadow deployment or small canary before full release. Compare versions on the same traffic slice, monitor regressions by language and user segment, and retain enough output metadata to reproduce failures without storing unnecessary personal data. For constrained environments, optimisation choices matter: AI model optimisation for mobile devices covers the latency, memory, and quality trade-offs that also affect edge deployments.
Turn evaluation into a release gate
Run evaluations whenever you change the model, prompt, retrieval settings, chunking, embedding model, safety policy, or tool definitions. Store dataset versions, configuration, judge version, random seeds where applicable, and cost alongside every result.
A release should require more than an improved average score. Define non-negotiable gates such as:
- No increase in critical safety failures.
- Minimum recall for high-value intents and languages.
- Maximum p95 latency and cost per successful task.
- No regression beyond an agreed margin on the golden set.
- Human approval for newly introduced failure categories.
Finally, maintain an error taxonomy. Label failures as retrieval, reasoning, knowledge, instruction-following, language, safety, tool-use, or infrastructure errors. This turns evaluation from a leaderboard exercise into an engineering feedback loop: each failure should lead to a dataset example, a mitigation, and a regression test.