0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark sarvam model on hugging face

How to Benchmark a Sarvam Model on Hugging Face

  1. aigi

    Sarvam models are built for Indian-language use cases, so a useful benchmark must test more than a generic English score. You need to identify the exact checkpoint and task, use representative data, measure quality and serving performance, and record the conditions that produced each result. This guide explains how to benchmark a Sarvam model on Hugging Face in a way that another builder can reproduce in 2026.

    Start with the exact model and task

    “ Sarvam model” is not a sufficient benchmark specification. Hugging Face may contain checkpoints for translation, speech, embeddings, or text generation, and each requires a different evaluation method. Begin by recording:

    • The complete Hugging Face repository ID and revision or commit hash.
    • Model type, parameter count, context length, tokenizer, and required libraries.
    • Whether the checkpoint is base, instruction-tuned, quantised, or fine-tuned.
    • Target languages, scripts, domains, and expected input length.
    • Hardware, precision, batch size, and decoding settings.

    Read the model card before writing evaluation code. Check its licence, intended use, known limitations, language coverage, and any published scores. For Indian-language comparisons, pair this work with a task-specific reference such as benchmarking NLP models for Telugu and Sanskrit, rather than treating an English benchmark as a proxy for local-language capability.

    Build a reproducible Hugging Face environment

    Create an isolated environment and pin the versions used in the run. A minimal setup for text-generation or encoder models is:

    python -m venv .venv
    source .venv/bin/activate
    pip install -U "transformers" "datasets" "evaluate" "accelerate" \
      "torch" "sentencepiece" "sacrebleu" "jiwer" "pandas"

    Use a GPU when measuring throughput, but do not compare GPU and CPU results as though they were equivalent. Record the GPU model, CUDA version, available memory, operating-system details, and whether other workloads were running. Set deterministic seeds where supported, while recognising that some GPU operations and generation paths can still vary slightly.

    Load the checkpoint with the class appropriate to its architecture. For a causal language model:

    import torch
    from transformers import AutoTokenizer, AutoModelForCausalLM
    
    model_id = "sarvam/<exact-checkpoint>"
    tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
    model = AutoModelForCausalLM.from_pretrained(
        model_id,
        torch_dtype=torch.bfloat16,
        device_map="auto",
        trust_remote_code=True,
    ).eval()

    Do not copy trust_remote_code=True blindly into production. Inspect repository code first and pin a reviewed revision. If the model card specifies a custom class or processor, follow those instructions instead of forcing AutoModelForCausalLM.

    Choose data that reflects Indian usage

    A benchmark is only as credible as its test set. Use a held-out split and avoid examples that appeared in training or public demonstrations. Include the languages and conditions your application will face:

    • Native scripts and Romanised text where relevant.
    • Code-mixing, spelling variation, abbreviations, and regional vocabulary.
    • Short queries, long documents, conversational turns, and noisy inputs.
    • Formal and informal registers, including government, education, health, and commerce domains.
    • Dialect or transliteration cases that are important to your users.

    Keep a private test set for final validation. Public datasets are useful for comparison, but they can become targets for prompt tuning or contamination. For work involving Marathi dialects, for example, fine-tuning AI models for Marathi dialect provides useful context on why aggregate scores can hide variation between speech communities.

    Match metrics to the task

    Do not report accuracy for every workload. Select metrics that reflect the failure modes that matter:

    • Generation: exact match for constrained answers, ROUGE for summarisation, and human or model-assisted quality checks for open-ended responses.
    • Translation: SacreBLEU, chrF, COMET where appropriate, plus human review for terminology, adequacy, and fluency.
    • Speech recognition: word error rate and character error rate, separated by language, script, speaker, and audio condition.
    • Classification: accuracy, macro-F1, per-class recall, and confusion matrices. Macro-F1 is important when classes are imbalanced.
    • Embeddings: retrieval recall@k, mean reciprocal rank, and task-level accuracy rather than only cosine similarity.
    • Serving: time to first token, inter-token latency, tokens per second, peak VRAM, RAM, and cost per request.

    For Sanskrit or Telugu, report scores separately instead of presenting one pooled number. A model can look strong overall while failing on one language, script, or domain.

    Run a controlled quality benchmark

    The Hugging Face datasets and evaluate libraries can manage data and standard metrics, but the inference loop must respect the model’s input format. For generation:

    from datasets import load_dataset
    import torch
    
    data = load_dataset("your-org/your-eval-set", split="test")
    outputs = []
    
    for row in data:
        prompt = row["prompt"]
        inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
        with torch.inference_mode():
            generated = model.generate(
                **inputs,
                max_new_tokens=128,
                do_sample=False,
                pad_token_id=tokenizer.eos_token_id,
            )
        text = tokenizer.decode(generated[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True)
        outputs.append({"prediction": text, "reference": row.get("reference")})

    Use the same prompts, maximum output length, stopping rules, and decoding parameters for every model in a comparison. Run at least one warm-up pass, then evaluate the complete test set. Save raw predictions, errors, configuration, and per-example metadata—not only the final score. This makes regressions diagnosable and supports later human review.

    Measure latency and memory separately

    Quality and speed answer different questions. For serving tests, use fixed prompt-length buckets and batch sizes, then repeat each run enough times to report median and tail latency. Synchronise CUDA before reading timers:

    import time
    
    def timed_generate(prompt, repeats=20):
        inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
        for _ in range(3):
            with torch.inference_mode():
                model.generate(**inputs, max_new_tokens=64)
        times = []
        for _ in range(repeats):
            if torch.cuda.is_available(): torch.cuda.synchronize()
            start = time.perf_counter()
            with torch.inference_mode():
                result = model.generate(**inputs, max_new_tokens=64)
            if torch.cuda.is_available(): torch.cuda.synchronize()
            times.append(time.perf_counter() - start)
        return sorted(times), result

    Report p50 and p95 latency, tokens per second, peak memory, and the number of concurrent requests. Test full precision and the quantisation mode you actually intend to deploy; results from one cannot reliably predict the other. If the target is a phone or edge device, compare against a mobile-specific workflow such as AI model optimization for mobile devices.

    Compare fairly and inspect failures

    A strong benchmark includes a baseline model, not just Sarvam’s absolute score. Select a baseline with similar size, task, language coverage, and licensing constraints. Keep hardware and software constant. Report confidence intervals or bootstrap intervals when the test set is small, and separate aggregate results by language, length, domain, and input quality.

    Then inspect failures manually. Categorise hallucinations, untranslated spans, script errors, unsafe answers, formatting failures, and refusals. For applications using retrieval or agents, test the complete pipeline; a standalone model score may not predict production behaviour. If repetitive generation is a concern, track duplicate n-grams and review reducing repetitive responses in LLM applications.

    Publish a benchmark card

    Make the result useful to other builders by publishing:

    • Model ID, revision, tokenizer, and prompt template.
    • Dataset source, licence, split, size, and contamination safeguards.
    • Hardware, software versions, precision, batch size, and decoding settings.
    • Quality metrics with per-language and per-category breakdowns.
    • Latency methodology, memory readings, throughput, and cost assumptions.
    • Raw predictions or an evaluation repository where redistribution permits.
    • Known limitations and examples of representative failures.

    This level of reporting prevents inflated comparisons and gives teams enough information to reproduce the result or decide whether the model fits their product.

    FAQ

    Can I benchmark Sarvam with a generic Hugging Face pipeline?

    You can use a pipeline for a quick smoke test, but serious evaluation should use the model’s documented processor, prompt format, and generation settings. Pipelines can also obscure batching and device placement.

    Should I use public benchmarks only?

    No. Public benchmarks support comparability, while a private, representative test set protects against contamination and tests your actual users’ languages and workflows.

    What is the most important metric?

    There is no universal winner. For a translation product, quality by language and terminology may matter most; for an API, p95 latency, throughput, and cost may determine viability. Report both task quality and operational performance.

    How often should I rerun the benchmark?

    Rerun it whenever the model, tokenizer, prompt, quantisation, runtime, hardware, or dataset changes. Keep historical results so regressions are visible.

    For Indian AI builders seeking support for evaluation infrastructure or deployment, explore AI Grants India and review the eligibility requirements before applying.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.