0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to benchmark tamil instruction following on indicifeval using hugging face

How to Benchmark Tamil Instruction Following on IndicEval

  1. aigi

    Tamil instruction-following benchmarks are useful only when the evaluation is reproducible, linguistically careful, and tied to real product requirements. A model may produce fluent Tamil yet ignore constraints, answer in the wrong format, or fail on code-mixed requests. IndicEval can provide a common evaluation surface; Hugging Face supplies the datasets, tokenizers, models, and tooling needed to run the experiment.

    This guide focuses on a reliable workflow rather than a single model or an assumed dataset identifier. Dataset names, configurations, task schemas, and model architectures change, so verify the current IndicEval repository or Hugging Face dataset card before running the commands below.

    What to measure in Tamil instruction following

    Instruction following is broader than exact-answer accuracy. Your evaluation should test whether the model:

    • Understands the Tamil instruction, including colloquial and formal registers.
    • Completes the requested task without adding irrelevant content.
    • Obeys constraints such as language, length, structure, audience, and prohibited content.
    • Preserves names, numbers, dates, units, and terminology.
    • Produces an answer that is useful to a Tamil-speaking user.

    Before benchmarking, define the task categories in your report. Useful categories include question answering, summarisation, rewriting, classification, extraction, structured generation, safety refusal, and code-mixed Tamil-English prompts. For broader dataset selection, compare IndicEval with the Indian language LLM benchmark datasets guide and the practical framework for benchmarking multilingual LLMs in India.

    Prepare a reproducible Hugging Face environment

    Create an isolated environment and pin the versions used for the run:

    python -m venv .venv
    source .venv/bin/activate
    pip install -U "transformers" "datasets" "accelerate" "evaluate" "sacrebleu" "sentencepiece" "torch"

    Record the following in a requirements.txt or experiment manifest:

    • Python, PyTorch, Transformers, Datasets, and Accelerate versions.
    • Model repository and exact revision or commit.
    • Dataset repository, configuration, split, and revision.
    • Hardware, quantisation settings, batch size, maximum input length, and generation parameters.
    • Random seed and the date of the run.

    Do not assume that every Tamil-capable model is an instruction-tuned model. A base language model may continue text competently but fail to follow a direct request. When choosing candidates, inspect the model card, supported languages, licence, context window, chat template, and inference requirements. If you are comparing model families, see this overview of large language models for Tamil speakers.

    Load and inspect the IndicEval data

    First discover the available configurations instead of hard-coding an unverified name:

    from datasets import get_dataset_config_names, load_dataset
    
    repo_id = "<verified-IndicEval-dataset-repository>"
    print(get_dataset_config_names(repo_id))
    
    dataset = load_dataset(repo_id, name="<verified-tamil-configuration>")
    print(dataset)
    print(dataset["test"].column_names)
    print(dataset["test"][0])

    Confirm which fields represent the instruction, context, reference answer, category, and any structured constraints. Preserve the original test split. Do not translate or rewrite test examples after downloading them, because that changes the benchmark. If a task has multiple valid references, store all acceptable references and score against the documented protocol.

    Tamil-specific checks matter. Inspect Unicode normalisation, punctuation, Grantha characters, numerals, transliterated words, and code-mixed text. Normalise only when the benchmark specifies it; otherwise report both raw and normalised results. Also check for duplicate prompts, training-set contamination, empty references, and examples whose expected output is ambiguous.

    Run inference with the model’s chat template

    For chat or instruction-tuned models, use the model’s configured template rather than manually inventing a prompt format. The following pattern works for causal language models; adapt it for encoder-decoder models:

    import torch
    from transformers import AutoModelForCausalLM, AutoTokenizer
    
    model_id = "<verified-tamil-capable-instruction-model>"
    tokenizer = AutoTokenizer.from_pretrained(model_id, revision="<revision>")
    model = AutoModelForCausalLM.from_pretrained(
        model_id,
        revision="<revision>",
        torch_dtype="auto",
        device_map="auto",
    )
    model.eval()
    
    def generate_answer(instruction, context=None):
        user_text = instruction if not context else f"{instruction}\n\nContext:\n{context}"
        messages = [{"role": "user", "content": user_text}]
        prompt = tokenizer.apply_chat_template(
            messages, tokenize=False, add_generation_prompt=True
        )
        inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
        with torch.inference_mode():
            output = model.generate(
                **inputs,
                max_new_tokens=512,
                do_sample=False,
                temperature=None,
                top_p=None,
            )
        generated = output[0][inputs["input_ids"].shape[1]:]
        return tokenizer.decode(generated, skip_special_tokens=True).strip()

    Use deterministic decoding for leaderboard-style comparisons. If you also want to measure production behaviour, run a separate sampled setting and report it separately. Avoid truncating Tamil prompts silently: log token counts and the number of examples that exceed the context window.

    Score instruction following, not just text overlap

    Exact-match and token-level F1 can be useful for closed-answer tasks, but they are weak measures for open-ended Tamil responses. A practical report should combine:

    • Task score: accuracy, macro-F1, exact match, or execution success where appropriate.
    • Text similarity: chrF or character-level metrics, which are often more informative than word overlap for morphologically rich languages.
    • Constraint adherence: a separate pass/fail score for language, length, format, required fields, and forbidden content.
    • Human quality rating: native Tamil reviewers assess correctness, relevance, fluency, and instruction compliance.
    • Safety and refusal accuracy: distinguish an appropriate refusal from a refusal that ignores a harmless request.

    Use a small, adjudicated human sample rather than presenting an automated metric as ground truth. Give reviewers the prompt, model output, reference where available, and a clear rubric. Blind the model identity and measure inter-rater agreement. For structured outputs, parse JSON or other formats programmatically before judging content.

    Analyse errors and publish a useful report

    Break results down by task type, prompt length, register, code-mixing, domain, and constraint count. A single Tamil score can hide serious weaknesses—for example, strong summarisation but poor list formatting or unreliable handling of dates and quantities.

    Include confidence intervals or bootstrap intervals when the test set is large enough, and avoid ranking models on tiny score differences. Publish per-category scores, failed examples with sensitive data removed, generation settings, dataset and model revisions, and known limitations. If you fine-tune after evaluation, keep the original test set untouched and create a separate development set.

    For teams building low-resource or offline applications, compare quality against latency, memory, and licence constraints. A smaller model may be the better deployment choice if it follows Tamil instructions reliably within your hardware budget. Related implementation options include running a quantized Tamil model offline, training a tokenizer for Tamil language models, and creating a small language model for Tamil.

    A compact benchmark checklist

    Before sharing results, verify that you have:

    • Used the official IndicEval Tamil split and documented its revision.
    • Confirmed the model’s Tamil and instruction-tuning support.
    • Applied the correct chat template and avoided prompt leakage.
    • Logged truncation, failures, latency, and generation settings.
    • Reported task, constraint, human, and safety measures separately.
    • Analysed errors by language variety and task category.
    • Released code or an experiment manifest where licensing permits.

    A rigorous IndicEval run is not simply a score printed from a notebook. It is an auditable comparison that shows what a Tamil model can do, where it fails, and whether the result matters for the users and workflows you intend to serve.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.