0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to use karpathy style research agents to evaluate sarvam ai model performance

How to Use Research Agents to Evaluate Sarvam AI Models

  1. aigi

    What this evaluation approach means

    A “Karpathy-style” research agent is not a specific product or official framework. It is a disciplined workflow inspired by Andrej Karpathy’s emphasis on small experiments, transparent code, reproducible runs, and tight feedback loops. The agent helps you generate test cases, run model calls, inspect failures, compare versions, and turn observations into the next experiment.

    That distinction matters. The agent should orchestrate evaluation, not act as an unquestioned judge of Sarvam AI model quality. Your benchmark design, reference answers, human review, and production telemetry remain the source of truth.

    For teams building Indian-language products, this approach is especially useful because aggregate scores can hide major differences between Hindi, Tamil, Telugu, Bengali, Marathi, Kannada, Malayalam, Gujarati, Punjabi, and English. It can also expose failures in code-switching, transliteration, noisy audio, regional phrasing, and low-resource language coverage.

    Start with a precise evaluation contract

    Before writing an agent, define what “better” means for your application. A voice assistant for restaurants has different requirements from a document extraction system or a hospital follow-up workflow. For example, teams designing multilingual voice agents for restaurants in India should test menu names, addresses, quantities, local pronunciation, interruptions, and order confirmation—not just generic language understanding.

    Write an evaluation contract containing:

    • Use cases: the user journeys Sarvam must support.
    • Languages and scripts: include native scripts, Romanised Indian languages, English, and code-switched inputs where relevant.
    • Expected behaviour: define acceptable answers, refusal conditions, escalation rules, and formatting.
    • Hard constraints: maximum latency, cost per request, context length, uptime, and data residency requirements.
    • Release thresholds: specify which failures block deployment and which are acceptable trade-offs.

    Keep this contract versioned. A score is meaningful only when the dataset, prompt, model version, decoding settings, and evaluator version are recorded alongside it.

    Build a representative Indian-language test set

    Avoid relying exclusively on public benchmarks. Create a private evaluation set from anonymised production examples, domain-written prompts, and deliberately difficult cases. Stratify it by language, task, difficulty, and risk.

    A useful test matrix includes:

    • Language: native script, Roman transliteration, mixed-language input, and regional variants.
    • Task: classification, extraction, summarisation, translation, question answering, tool use, and dialogue.
    • Input quality: clean text, spelling errors, speech-recognition errors, abbreviations, and incomplete requests.
    • Context: short prompts, long conversations, conflicting instructions, and missing information.
    • Risk: benign, sensitive, regulated, adversarial, and out-of-scope requests.

    For voice products, measure the complete pipeline rather than only the language model. Evaluate speech-to-text errors, Sarvam’s response, text-to-speech quality, barge-in handling, and end-to-end time to a useful answer. The practical considerations in how voice agents work are a useful checklist for separating model quality from orchestration and audio-stack problems.

    Do not put every example in one static test set. Maintain three partitions: a development set for iteration, a locked holdout set for release decisions, and a live canary set sampled from production with privacy safeguards.

    Design the research agent as an experiment runner

    A reliable agent can be implemented as a small set of explicit tools rather than a fully autonomous system. Give it functions to:

    1. Load a versioned dataset and sampling plan.
    2. Call the Sarvam API or deployment with fixed parameters.
    3. Record request IDs, model versions, prompts, outputs, latency, token or character usage, and errors.
    4. Run deterministic metrics and rule-based checks.
    5. Send only selected cases to an LLM judge or human reviewer.
    6. Cluster failures and propose the next experiment.
    7. Produce a report with raw evidence and confidence intervals.

    Use a structured schema for every trial. At minimum, store case_id, language, task, input, expected output or rubric, model configuration, output, timestamp, latency, cost estimate, evaluator result, and failure category. Never let the agent silently rewrite prompts or remove poor results.

    For larger workloads, place the runner behind a queue and isolate tenants, credentials, and sensitive datasets. Patterns from building distributed systems with AI agents are relevant when parallelising evaluations, but reproducibility should take priority over maximum throughput.

    Combine automated metrics with human review

    Use metrics that match the task. Exact match works for structured fields; character or word error rate helps evaluate transcription; precision, recall, and F1 suit classification; and schema-validity checks are valuable for tool calls and JSON output. For generation, automatic similarity scores can support analysis but should not be treated as a complete quality measure.

    For open-ended responses, create a rubric with observable criteria:

    • factual correctness and completeness;
    • language and script appropriateness;
    • instruction following;
    • groundedness in supplied context;
    • clarity and politeness;
    • safety and appropriate refusal;
    • successful task completion.

    LLM judges can scale review, but calibrate them against labelled human examples and inspect disagreements. Blind the judge to model identity where possible, randomise comparison order, and report agreement rates. Maintain a human-review queue for safety failures, high-impact decisions, and cases where judges disagree.

    If the product handles health information, do not infer compliance from a high benchmark score. Review access controls, retention, consent, audit logs, and escalation separately; the HIPAA-compliant voice agents for hospitals guide offers a useful risk-oriented frame, even where Indian regulations and contractual requirements differ.

    Measure quality, latency, cost, and reliability together

    A model that is slightly more accurate but twice as slow or expensive may be unsuitable for an Indian consumer product. Report results by language and task, not only as one average.

    Track:

    • success rate and task completion;
    • per-language accuracy, recall, and refusal quality;
    • p50, p95, and p99 latency;
    • timeout, retry, and malformed-output rates;
    • cost per request and per completed task;
    • context-window and rate-limit failures;
    • safety and privacy incident counts.

    Use bootstrap confidence intervals or repeated samples to distinguish real improvements from noise. For voice systems, include time to first audio and total turn duration. For production planning, compare the model against a baseline—such as the current Sarvam version, a smaller fallback, and a human workflow.

    Run controlled experiments and close the feedback loop

    Change one major variable at a time: model version, prompt, retrieval policy, decoding settings, tool schema, or audio component. Run the same locked cases across candidates, then use the agent to identify regressions by slice. A useful report should show both aggregate movement and examples of newly introduced failures.

    Do not automatically fine-tune or update production from the agent’s recommendations. Route proposals through an experiment owner, review data provenance, check for leakage, and rerun the holdout set. For deployment, use shadow traffic or a small canary with rollback criteria. Log consent and redact personally identifiable information before examples enter the research loop.

    Common failure modes

    Several shortcuts produce misleading conclusions:

    • English-heavy datasets: hide weak performance in Indian languages.
    • Synthetic-only prompts: miss real speech, slang, ambiguity, and incomplete context.
    • Single-score reporting: conceals safety, latency, and language-specific regressions.
    • Uncalibrated LLM judges: reward fluent but incorrect answers.
    • Prompt leakage: lets the evaluator see reference answers or test labels.
    • No cost accounting: makes an impractical system look competitive.
    • Unversioned runs: prevent teams from reproducing a result.
    • Blind automation: allows the agent to discard inconvenient failures.

    Treat every benchmark as a decision instrument with a defined owner, review date, and release consequence.

    A practical release checklist

    Before approving a Sarvam model or configuration, confirm that:

    • the test set covers target Indian languages, scripts, domains, and edge cases;
    • the holdout set has not been used for prompt or model tuning;
    • automated metrics and human rubric scores are reported separately;
    • latency, cost, reliability, and safety meet written thresholds;
    • regressions are reviewed by language and use case;
    • prompts, code, datasets, model identifiers, and evaluator versions are archived;
    • sensitive data is redacted and access is auditable;
    • a canary, rollback path, and post-release monitoring plan exist.

    The strongest outcome is not a flattering benchmark. It is a repeatable evaluation system that tells an Indian product team where Sarvam performs well, where it needs guardrails or routing, and whether each new release improves the user experience at an acceptable operational cost.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.