0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to use karpathy autoresearch to evaluate indicbert accuracy on punjabi legal documents

How to Use AutoResearch to Evaluate IndicBERT on Punjabi Legal Text

  1. aigi

    Start with the right evaluation question

    Evaluating IndicBERT on Punjabi legal documents is not a single accuracy test. It is a set of controlled experiments that should answer a specific product question: can the model classify a petition, identify a court or statute, retrieve a relevant passage, or label entities reliably enough for a legal workflow?

    That distinction matters because a model can score well on document classification while failing on legal named-entity recognition, clause extraction, or long-context retrieval. Before using Karpathy’s AutoResearch workflow, define the task, input format, labels, acceptable error rate, and whether the output is advisory or used to trigger an action. Teams building production systems should also review how Indian legal teams should evaluate and deploy AI law tools before treating benchmark results as deployment evidence.

    Verify the tools and model assumptions

    The original AutoResearch project is best understood as an experiment loop: a researcher defines a training or evaluation program, runs repeatable trials, and compares changes against a fixed validation setup. Do not assume that a package named karpathy-autoresearch, an autoresearch evaluate command, or a particular IndicBERT class exists in your environment. Check the project repository, commit, Python version, and supported framework first.

    Similarly, verify the exact IndicBERT checkpoint and tokenizer. IndicBERT variants differ in architecture, pretraining data, vocabulary, and language coverage. Confirm that the checkpoint supports Punjabi in Gurmukhi and record the model revision. Use a pinned environment, for example:

    python -m venv .venv
    source .venv/bin/activate
    pip install -U torch transformers datasets evaluate accelerate scikit-learn
    pip freeze > requirements.lock.txt

    If your task involves a small model or a compressed checkpoint, compare it using the same protocol described in how to evaluate a small language model. That prevents a favourable result caused by inconsistent preprocessing or test splits.

    Build a defensible Punjabi legal dataset

    Legal text is sensitive, repetitive, and highly vulnerable to leakage. Assemble documents from clearly documented sources, check usage rights, and remove personal identifiers where possible. Record the document type, court or authority, date, language, script, source, and any OCR or translation step.

    Use a split based on documents—not random sentences. If paragraphs from the same judgment appear in both training and test data, the score will be inflated. A practical starting structure is:

    • Training set: 70% of documents.
    • Validation set: 15% for model and threshold decisions.
    • Held-out test set: 15%, opened only for final reporting.
    • Time-based test: newer documents, if the system will process current filings.

    Include difficult variation: scanned PDFs, OCR noise, spelling variants, mixed Punjabi-English text, citations, section numbers, headings, and long judgments. Preserve the original text alongside normalized text so errors can be traced back to the source. For OCR-heavy collections, document character substitutions and confidence scores rather than silently correcting them.

    Create annotation guidelines before labelling. Define how annotators handle nested entities, abbreviations, incomplete clauses, quotations, and ambiguous legal references. Use at least two annotators for a sample and report agreement. A disagreement rate may reveal that the label scheme—not IndicBERT—is the main bottleneck.

    Choose tasks and metrics that reflect legal use

    For document or paragraph classification, report macro-F1, per-class precision and recall, a confusion matrix, and accuracy only as a secondary metric. Accuracy can conceal failure on rare but important classes. For named-entity recognition, use entity-level precision, recall, and F1 with exact-span matching, plus a relaxed overlap score if useful.

    For extractive question answering or clause spans, report exact match and token-level F1. For retrieval, use recall@k, mean reciprocal rank, and nDCG. If the model feeds a RAG pipeline evaluation, separately test retrieval quality, answer correctness, citation support, and refusal behaviour; do not call an answer accurate merely because the underlying classifier performed well.

    Add slices that matter in India:

    • Gurmukhi-only versus Punjabi-English mixed text.
    • Clean digital text versus OCR text.
    • Short notices versus long judgments.
    • Different courts, years, and document types.
    • Common versus rare legal entities.
    • Documents with and without citations or section references.

    Connect AutoResearch to a reproducible experiment

    Create an experiment configuration that fixes the dataset revision, tokenizer, maximum sequence length, random seeds, learning rate, batch size, number of epochs, and evaluation frequency. Each run should save the configuration, model revision, predictions, metrics, logs, and environment lockfile. AutoResearch should compare runs—not replace the evaluation design.

    A useful run manifest might contain:

    model: ai4bharat/indic-bert
    model_revision: <pinned-commit-or-tag>
    task: legal_ner
    language: pa-Guru
    dataset_revision: <dataset-hash>
    max_length: 512
    seed: 42
    metrics: [macro_f1, entity_f1, precision, recall]

    Your evaluation script should load the held-out test set only after the experiment plan is frozen. Produce a machine-readable JSON file and a human-readable report for every run. Store predictions with document IDs, but keep the document text in a controlled location. If AutoResearch’s interface differs from your setup, wrap the script in the repository’s documented runner rather than inventing unsupported commands.

    For long documents, test chunking strategies explicitly: fixed windows, heading-aware chunks, and overlapping windows. Compare how chunk length and overlap affect both score and compute cost. A multi-stage LLM pipeline can help separate OCR cleanup, retrieval, classification, and human review, but each stage needs its own test set and failure budget.

    Analyse errors, not just leaderboard scores

    After each run, inspect false positives and false negatives by slice. Look for script confusion, OCR substitutions, statute-number formatting, rare names, negation, quoted language, and boundary errors. Read examples with a Punjabi-speaking legal expert; automated metrics cannot explain whether a prediction is legally misleading.

    Report confidence calibration as well. A model that is 90% accurate but confidently wrong on rare legal categories can be more dangerous than a lower-scoring model that flags uncertainty. Consider calibration curves, expected calibration error, abstention thresholds, and a human-review queue. Never present IndicBERT output as legal advice or final interpretation without qualified review.

    Minimum report for a credible result

    Publish the checkpoint, dataset composition, annotation process, split policy, preprocessing, hardware, training budget, seeds, metrics, confidence intervals where feasible, and representative errors. State what the test does not cover—especially translation, OCR, unseen courts, and extremely long judgments. Compare against simple baselines such as majority class, TF-IDF with a linear classifier, and a multilingual alternative.

    The conclusion should be operational: which task is ready for assisted use, which slices require more data, and what review controls are mandatory. For broader benchmarking context, Punjabi small-language-model comparisons can inform model selection, but they cannot substitute for a Punjabi legal test set built around your actual workflow.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.