0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to evaluate a small language model

How to Evaluate a Small Language Model: A Practical Guide

  1. aigi

    Small language models are attractive to Indian builders because they can run with lower latency, reduced inference costs, and tighter control over data. But a smaller model is not automatically better for a production use case. The right evaluation asks whether the model performs reliably on your users’ language, workflows, devices, and failure conditions.

    A useful evaluation plan combines offline metrics, human review, adversarial testing, and deployment measurements. This matters especially for applications handling Hindi and other Indic languages, where spelling variation, code-mixing, transliteration, and limited benchmark coverage can hide serious quality gaps. For background on the data challenges, see this guide to low-resource Indic natural language processing.

    Start with a clear evaluation target

    Define the job before choosing a metric. A model used for customer support should not be evaluated like one used for document extraction or on-device autocomplete.

    Write down:

    • Primary task: generation, classification, extraction, summarisation, translation, retrieval-augmented answering, or tool calling.
    • Users and languages: include the actual mix of English, Hindi, regional languages, transliterated text, abbreviations, and speech-to-text errors.
    • Failure tolerance: decide what is unacceptable. A wrong product suggestion may be inconvenient; an incorrect health or financial answer may be dangerous.
    • Deployment limits: specify latency, memory, throughput, battery, bandwidth, and inference-cost targets.
    • Success threshold: set minimum quality and maximum error rates before testing models.

    Create a fixed evaluation set before tuning the model. Keep training, validation, and test examples separate, and prevent near-duplicates from crossing those splits. A small but carefully labelled test set is more useful than a large, noisy benchmark.

    Build a representative test set

    Your test data should reflect production traffic rather than generic internet text. Sample queries by intent, language, difficulty, user segment, and expected answer length. Include both common cases and rare cases that could cause material harm.

    For Indian deployments, test:

    • Native-script Hindi and regional-language inputs.
    • Romanised Hindi and code-mixed prompts such as Hinglish.
    • Spelling variation, names, addresses, local institutions, and currency formats.
    • Low-bandwidth or noisy speech transcripts.
    • Domain terms used by Indian customers, such as GST, UPI, pin codes, and local government schemes.
    • Ambiguous prompts, incomplete requests, and adversarial instructions.

    Use a gold answer only where one correct answer exists. For open-ended generation, define a rubric with criteria such as factuality, relevance, completeness, tone, citation quality, and refusal behaviour. Have at least two reviewers score a meaningful sample, then resolve disagreements and record the reasons.

    If the model targets Hindi specifically, compare it with relevant open models and datasets using a guide to open-source small language models for Hindi. Do not assume that performance in English transfers to Hindi or other Indian languages.

    Choose metrics that match the task

    Classification and detection

    Use accuracy only when classes are balanced and errors have similar consequences. Otherwise report:

    • Precision: how many predicted positives are correct.
    • Recall: how many actual positives the model finds.
    • F1 score: a balance of precision and recall.
    • Macro-F1: useful when minority classes matter.
    • Confusion matrix: shows which categories the model confuses.

    For moderation, fraud, or triage, select an operating threshold based on business risk rather than defaulting to 0.5.

    Generation and summarisation

    Perplexity can compare language modelling quality under controlled conditions, but it is not a proxy for usefulness, factuality, or instruction following. BLEU and ROUGE can help with translation or reference-heavy summarisation, yet they often undervalue valid answers that use different wording.

    Add task-specific measures:

    • Exact-match or token-level F1 for structured answers.
    • Field-level accuracy for extraction.
    • Citation precision and unsupported-claim rate for retrieval systems.
    • JSON validity and schema compliance for tool-enabled applications.
    • Human ratings for relevance, clarity, completeness, and factuality.

    For chat systems, maintain a labelled set of expected refusals, safe completions, and escalation cases. Track hallucination rate separately from general answer quality.

    Measure production performance, not just quality

    A model that scores well offline may still fail on a mobile device or become too expensive at scale. Record:

    • Time to first token and total response latency.
    • Tokens per second and requests per second.
    • Peak RAM or VRAM use.
    • Model size after quantisation.
    • CPU, GPU, or accelerator utilisation.
    • Energy consumption for on-device use.
    • Cost per request and cost per successful task.
    • Failure rates, timeouts, and fallback frequency.

    Test under realistic concurrency and with the longest prompts users will send. If deployment is on phones, edge devices, or unreliable networks, pair evaluation with AI model optimisation for mobile devices. A small increase in accuracy may not justify a large increase in memory or latency.

    Test robustness, safety, and security

    Small models can be more brittle when prompts are long, instructions conflict, or inputs contain unfamiliar language. Create targeted test suites for:

    • Prompt injection and instruction hierarchy attacks.
    • Personally identifiable information leakage.
    • Unsafe advice and overconfident answers.
    • Toxic, discriminatory, or culturally inappropriate outputs.
    • Jailbreaks, repeated prompting, and multilingual attacks.
    • Long-context degradation and irrelevant-document distraction.
    • Misspellings, transliteration, emojis, and mixed scripts.

    Evaluate refusal quality, not merely refusal frequency. The model should decline unsafe requests clearly while still helping with legitimate adjacent tasks. For sensitive applications, route uncertain cases to a human or a stronger model and measure how often escalation is triggered.

    Run error analysis systematically

    After scoring, inspect failures by category instead of looking only at the aggregate number. Build an error table with the input, expected behaviour, actual output, severity, likely cause, and proposed fix.

    Useful categories include:

    • Missing knowledge or outdated information.
    • Poor instruction following.
    • Language or script confusion.
    • Retrieval failure.
    • Hallucination or unsupported inference.
    • Formatting and tool-use errors.
    • Latency, truncation, or context-window failure.

    Slice results by language, device, geography, user type, and task difficulty. A strong overall score can conceal poor performance for a smaller but important user group. When the model is adapted for regional languages, compare fine-tuning strategies and data quality; fine-tuning Llama for Indian regional languages provides useful implementation context.

    Use a repeatable evaluation workflow

    A practical 2026 workflow looks like this:

    1. Freeze a versioned test set and evaluation rubric.
    2. Establish a baseline with a simple rules system, current model, or larger reference model.
    3. Run automated task metrics and infrastructure benchmarks.
    4. Sample outputs for blind human review.
    5. Execute safety, robustness, and multilingual test suites.
    6. Analyse severe and recurring errors.
    7. Re-test after every data, prompt, quantisation, or model change.
    8. Log results in a dashboard with model version, dataset version, configuration, and hardware.

    Tools such as Hugging Face Evaluate, lm-evaluation-harness, custom Python tests, and experiment trackers can automate much of this work. Keep the evaluation code in version control and make the test command reproducible. For production, use shadow traffic or a carefully controlled A/B test before a broad rollout.

    The decision framework

    Choose the model that meets the complete requirement, not the one with the highest single benchmark score. A useful release gate might require:

    • Minimum quality for every priority language and task.
    • No critical safety failures in red-team testing.
    • Latency and memory within device limits.
    • Cost per successful task below the target.
    • Stable results across model updates.
    • A documented fallback and monitoring plan.

    The final evaluation report should state what was tested, what was excluded, where the model fails, and what safeguards are in place. That level of transparency helps teams make responsible trade-offs and gives funders, customers, and engineering partners confidence in the system.

    FAQ

    Is perplexity enough to evaluate a small language model?
    No. Perplexity measures next-token prediction under a particular dataset and setup. Combine it with task accuracy, human review, safety tests, and deployment benchmarks.

    How many examples are needed?
    There is no universal number. Start with a balanced, representative set, then add examples for every recurring or high-severity failure. More important than raw size are clear labels, clean splits, and coverage of real inputs.

    Should I compare a small model with a larger model?
    Yes. Use a stronger model as a reference where practical, but compare cost, latency, privacy, and reliability as well as quality. The best production choice may be a small model with a fallback path.

    How should Indic-language quality be reviewed?
    Use native or highly proficient reviewers, test native and Romanised scripts, include code-mixed prompts, and report results separately by language. Do not hide weak regional-language performance inside one aggregate score.

    If you are building an AI product in India, AI Grants India can help you identify funding and support opportunities for responsible model development.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.