0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark small language models

How to Benchmark Small Language Models in 2026

  1. aigi

    Small language models (SLMs) can be faster and cheaper to deploy than frontier models, but parameter count alone does not determine whether a model is fit for production. A 1–7 billion parameter model may perform well on English classification yet struggle with code-mixed Hindi, noisy user input, or long-context retrieval. The right benchmark measures the complete system: quality, latency, memory, cost, robustness, and safety on the workload you actually serve.

    This guide explains how to benchmark small language models in a repeatable way, with particular attention to Indian products, on-device inference, and constrained cloud deployments.

    1. Define the deployment decision first

    Start with a decision, not a leaderboard. Write down what the model must do and the constraints it must satisfy.

    Useful questions include:

    • Is the model classifying support tickets, extracting fields, answering questions, summarising documents, or generating text?
    • Must inference run on a CPU, Android device, edge GPU, or shared cloud instance?
    • What is the maximum acceptable p95 latency and monthly inference budget?
    • Which languages, scripts, dialects, and code-mixed formats are required?
    • What level of factual accuracy, refusal quality, and privacy protection is necessary?

    For Indian deployments, include Hindi-English code-mixing, transliterated text, regional spelling variation, and low-resource languages in the acceptance criteria. The practical challenges overlap with those described in this guide to low-resource Indic natural language processing.

    Create a model card for the evaluation itself: model version, quantisation, prompt template, tokenizer, hardware, software versions, dataset revisions, and decoding settings. Without this record, two apparently different results may not be comparable.

    2. Build a representative evaluation set

    Use public benchmarks for broad comparison, but do not treat them as a production test. Assemble a private, representative set from real or carefully redacted inputs.

    A useful test suite contains:

    • Core task examples: typical requests covering the main product workflow.
    • Hard cases: ambiguous wording, long inputs, misspellings, tables, names, and incomplete context.
    • Distribution slices: language, script, geography, customer segment, document type, and input length.
    • Failure probes: prompt injection, unsupported claims, personally identifiable information, and unsafe requests.
    • Regression cases: examples from incidents or errors observed in earlier releases.

    Keep a locked test set that is never used for tuning. If you repeatedly optimise against the same examples, the benchmark becomes a training target rather than an independent measurement. For generation tasks, use human-labelled rubrics alongside automated scores because lexical overlap often undervalues a correct paraphrase and rewards fluent but inaccurate output.

    3. Select metrics that match the task

    There is no single “best” metric for an SLM. Report task quality and operating cost together.

    Classification and extraction

    Use accuracy when classes are balanced and errors have similar consequences. For imbalanced problems, report precision, recall, macro-F1, and per-class results. For structured extraction, measure exact match, field-level precision and recall, and validity rate for the required JSON schema.

    Also inspect the confusion matrix. A high aggregate score can conceal poor performance on a minority language, a safety-critical class, or a commercially important customer segment.

    Generation and question answering

    For summarisation or translation, BLEU, ROUGE, and BERTScore can support comparison, but they should not be the only evidence. Evaluate factual consistency, completeness, relevance, terminology, and reading quality.

    For question answering, measure answer accuracy and abstention quality. A model that says “I don’t have enough information” when evidence is absent may be more useful than one that produces confident hallucinations. If retrieval is involved, separately test retrieval recall, grounded answer rate, citation correctness, and unsupported-claim rate.

    Language modelling

    Perplexity is useful for comparing models under the same tokenisation and corpus, but it is not a direct proxy for user value. Compare it across English, Indian languages, transliteration, and code-mixed text rather than reporting one pooled number. Tokenisation can significantly affect apparent efficiency and quality for Indic scripts.

    4. Measure latency, throughput, and memory correctly

    Performance claims are meaningless without hardware and measurement details. Warm up the model before timing it, exclude one-time loading unless cold-start performance matters, and run enough repetitions to report median, p95, and p99 latency.

    Track:

    • Time to first token for interactive generation.
    • End-to-end latency, including tokenisation and post-processing.
    • Decode speed in tokens per second.
    • Throughput at realistic concurrency levels.
    • Peak RAM or VRAM, model load time, and disk footprint.
    • Energy use where on-device or battery-powered deployment matters.

    Benchmark several sequence lengths and batch sizes. A model may look fast at 128 tokens but become impractical at 4,000 tokens. Test the exact quantisation used in production—FP16, INT8, INT4, or another format—because compression changes both speed and output quality. For production infrastructure, compare the result with the deployment considerations in how to deploy deep learning models on GKE, especially around autoscaling and resource limits.

    5. Include cost and quality trade-offs

    Calculate cost per successful task, not merely cost per token. A cheaper model that requires retries, human review, or frequent escalation may be more expensive overall.

    A practical cost model includes:

    • Compute price per hour or per device.
    • Average input and output tokens.
    • Batching and concurrency.
    • Storage, network, observability, and evaluation overhead.
    • Human review or fallback-model costs.

    Plot quality against p95 latency and cost. This Pareto view makes it easier to choose between a larger model, a quantised SLM, or a cascade in which the SLM handles routine cases and escalates uncertain requests.

    6. Test robustness, safety, and privacy

    A benchmark should reveal how a model fails, not just how often it succeeds. Perturb inputs with spelling errors, punctuation changes, mixed scripts, speech-transcription noise, irrelevant context, and adversarial instructions. Test whether outputs remain stable when the same request is phrased differently.

    For safety, measure refusal correctness rather than refusal rate. Penalise both unsafe compliance and unnecessary refusal of legitimate requests. Check for leakage of system prompts, memorised personal information, sensitive-data reproduction, and insecure tool calls. For a customer-facing assistant, include escalation behaviour and whether the model clearly communicates uncertainty.

    If your product processes Indian-language documents or voice transcripts, add language-specific red-team cases. SLMs may perform well on standard Hindi while failing on transliteration, regional vocabulary, or code-mixed instructions. Teams building multilingual products can also compare the model choices discussed in open-source small language models for Hindi and fine-tuning Llama for Indian regional languages.

    7. Use a reproducible benchmark workflow

    A reliable workflow can be implemented as follows:

    1. Freeze model weights, tokenizer, prompt, decoding parameters, and test data.
    2. Run a smoke test to catch formatting, dependency, and device errors.
    3. Evaluate quality by task and by important data slice.
    4. Run latency, throughput, memory, and cold-start tests on target hardware.
    5. Repeat tests across seeds or fixed runs where generation is stochastic.
    6. Conduct human review on a stratified sample of outputs.
    7. Record failures with input, output, expected behaviour, and severity.
    8. Compare against a baseline, such as a trusted API, previous release, or rules-based system.
    9. Publish confidence intervals or uncertainty ranges where sample sizes permit.
    10. Set release gates before deployment and rerun the suite after every model, prompt, or infrastructure change.

    Open-source evaluation harnesses and the Transformers ecosystem can automate much of the scoring, but custom business and Indic-language tests remain essential. Keep benchmark code versioned and store raw outputs, not only aggregate scores.

    8. Report results in a decision-ready table

    A useful report should allow a technical and non-technical reviewer to make the same decision. Include:

    • Model and quantisation configuration.
    • Dataset size, provenance, language mix, and contamination checks.
    • Quality metrics overall and by slice.
    • Median, p95, and p99 latency.
    • Throughput, peak memory, and cold-start time.
    • Cost per request or successful task.
    • Safety, refusal, privacy, and robustness results.
    • Known failure modes and recommended mitigations.

    Do not rank models by one score. State the selection rule—for example, “macro-F1 above 0.85, p95 latency below 500 ms, peak RAM below 4 GB, and no critical safety failures.” This makes trade-offs explicit and prevents impressive benchmark numbers from overriding deployment requirements.

    FAQ

    Is perplexity enough to compare small language models?
    No. Perplexity measures next-token prediction under a particular corpus and tokenizer. Pair it with task quality, latency, memory, cost, and safety tests.

    Should I use standard benchmarks?
    Yes, for broad comparison and regression tracking. Add a private, representative test set before making a production decision.

    How many examples are needed?
    Use enough examples to cover important slices and estimate uncertainty. More examples are needed for rare failure modes than for a simple smoke test; document the sample size and selection method.

    Should quantised models be benchmarked separately?
    Always. Quantisation can change accuracy, output format compliance, latency, and memory use. Benchmark the exact artifact you intend to deploy.

    What is a strong result for an Indian-language application?
    There is no universal threshold. Report results separately for each required language, script, transliteration pattern, and code-mixed format, then set thresholds based on the product’s risk and user expectations.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.