0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to use hugging face to benchmark sanskrit on indicgenbench

How to Use Hugging Face to Benchmark Sanskrit on IndicGenBench

  1. aigi

    Sanskrit benchmarking needs more than running a model over a test file and reporting one score. Script variation, sandhi, rich morphology, domain differences, and limited high-quality data can all change the result. Hugging Face provides the model, tokenizer, dataset, and experiment infrastructure; IndicGenBench provides a task-specific basis for comparing generative performance across Indian languages.

    This guide shows a reproducible workflow for Sanskrit evaluation. It also highlights where generic example code can mislead: verify the actual IndicGenBench repository, dataset configuration, task format, and evaluation command before installing packages or writing a submission script. Package names and APIs may change, and IndicGenBench releases may not expose a single universal Python class.

    What you are measuring

    Start by defining the benchmark question. Sanskrit evaluation may cover translation, generation, question answering, summarisation, or another task. Each requires a different input-output format and metric. Record the following before running an experiment:

    • The exact IndicGenBench release, commit, or dataset revision.
    • The Sanskrit task and its language direction, if applicable.
    • The model checkpoint and tokenizer revision.
    • The decoding settings: greedy decoding, beam search, temperature, top-p, and maximum length.
    • The evaluation split and whether it is public, hidden, or held out.
    • The metrics and text-normalisation rules used by the benchmark.

    For broader context, compare your setup with this guide to Indian language LLM benchmark datasets and the more general framework for benchmarking multilingual LLMs in India. Those comparisons help separate Sanskrit-specific performance from general model quality.

    Set up a reproducible Hugging Face project

    Use an isolated environment and pin versions. A minimal starting point is:

    python -m venv .venv
    source .venv/bin/activate
    pip install -U pip
    pip install transformers datasets evaluate accelerate sentencepiece sacrebleu

    Install any IndicGenBench dependency exactly as its official documentation specifies. Do not assume that indic-gen-bench is the correct PyPI package or that IndicGenBench is an importable class. Many benchmark projects distribute evaluation scripts through a Git repository instead.

    Create a small configuration file or record the run in a table:

    benchmark_revision: <commit-or-release>
    task: <sanskrit-task>
    model: <org>/<checkpoint>
    tokenizer: <org>/<checkpoint>
    seed: 42
    max_new_tokens: 256

    Set a seed where supported, save the command used, and retain the raw predictions. A score without predictions cannot be audited or diagnosed.

    Load the model and tokenizer safely

    For a text-generation model, begin with the task-appropriate Hugging Face class rather than assuming every checkpoint is encoder-only:

    from transformers import AutoTokenizer, AutoModelForCausalLM
    
    model_id = "<model-repository>/<checkpoint>"
    tokenizer = AutoTokenizer.from_pretrained(model_id, revision="<revision>")
    model = AutoModelForCausalLM.from_pretrained(model_id, revision="<revision>")

    For encoder-decoder translation or sequence-to-sequence generation, use AutoModelForSeq2SeqLM. For classification or extractive tasks, use the corresponding task head. Check the model card for supported scripts, maximum context length, language tags, and licensing terms.

    Before benchmarking, run a smoke test on five examples. Confirm that the tokenizer preserves Devanagari correctly, special tokens are configured, padding is handled, and generated text is not accidentally truncated. If your Sanskrit corpus uses transliteration, such as IAST or Harvard-Kyoto, keep that representation consistent throughout the run.

    Load and validate the IndicGenBench data

    Use the dataset loader or repository instructions supplied by IndicGenBench. A typical Hugging Face pattern looks like this, but the dataset identifier and configuration must come from the benchmark release:

    from datasets import load_dataset
    
    data = load_dataset("<official-dataset-id>", "<sanskrit-config>")
    print(data)
    print(data["test"].column_names)
    print(data["test"][0])

    Do not create a new random test split from training data when the benchmark already defines an official test set. That breaks comparability and can introduce leakage. Instead, use the published split and preserve its example IDs.

    Validate the data before inference:

    • Check for empty prompts, missing references, duplicate IDs, and unexpected Unicode.
    • Inspect punctuation, danda characters (।, ॥), whitespace, and normalisation.
    • Confirm that source and target fields match the selected task.
    • Measure input and reference lengths to identify truncation risk.
    • Look for near-duplicate training and test examples if you fine-tuned the model.

    Keep the original fields untouched. If you normalise text for scoring, save both the original and normalised versions.

    Run generation and save predictions

    Use batched inference and disable gradients. The exact prompt template must match the benchmark specification; changing it can materially change results.

    import torch
    
    model.eval()
    predictions = []
    references = []
    
    for batch in batches:
        inputs = tokenizer(
            batch["source"],
            return_tensors="pt",
            padding=True,
            truncation=True,
            max_length=1024,
        ).to(model.device)
    
        with torch.no_grad():
            output_ids = model.generate(
                **inputs,
                max_new_tokens=256,
                do_sample=False,
            )
    
        predictions.extend(tokenizer.batch_decode(output_ids, skip_special_tokens=True))
        references.extend(batch["target"])

    For encoder-decoder models, the generated sequence normally contains only the target. For causal models, remove the prompt portion before scoring if the generation output includes it. Save a JSONL file containing the example ID, input, reference, prediction, model revision, and decoding parameters.

    Score the right metrics

    IndicGenBench may define task-specific metrics, so use its official evaluator wherever possible. Generic metrics can still help during debugging:

    • Exact match or accuracy is useful for constrained answers but is harsh for open-ended generation.
    • ChrF is often informative for morphologically rich languages because it rewards character-level overlap.
    • BLEU can support translation comparisons, but interpret it alongside other metrics and the benchmark’s tokenisation rules.
    • ROUGE may help for summarisation, though it should not be treated as a complete measure of Sanskrit quality.
    • Semantic or human evaluation is important when multiple valid Sanskrit formulations exist.

    Do not compare scores produced with different normalisation, tokenisation, reference counts, or decoding settings. Report confidence intervals or bootstrap estimates when the test set is small. For a useful baseline, include a multilingual checkpoint, a Sanskrit-focused checkpoint, and—where permitted—a fine-tuned model.

    Diagnose errors instead of chasing one score

    Create an error taxonomy and review a stratified sample of outputs. Useful categories include sandhi handling, inflection and agreement, compounds, named entities, punctuation, transliteration, hallucination, copying, and instruction-following failures. Break results down by source length, genre, and linguistic phenomenon where metadata allows.

    If the model struggles with domain-specific Sanskrit, fine-tuning may help, but data quality matters more than simply adding examples. This Sanskrit translation fine-tuning guide covers data preparation, validation, and training decisions. For language-pair comparisons, see benchmarking NLP models for Telugu and Sanskrit.

    Publish a benchmark report others can reproduce

    A credible report should include the dataset revision, model and tokenizer identifiers, hardware, software versions, prompt template, decoding parameters, preprocessing rules, metrics, raw scores, and known limitations. Link to the prediction file when licensing permits. State whether the model was fine-tuned on any Sanskrit or related benchmark data.

    As of 2026, the most useful Sanskrit benchmark result is not necessarily the highest single number. It is the result another team can rerun, inspect, and improve without guessing which hidden preprocessing or configuration produced it.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.