0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · post training experiments

Post-Training Experiments: A Practical Evaluation Guide

  1. aigi

    What post-training experiments should answer

    Post training experiments begin after a model has completed its initial training or fine-tuning. Their purpose is not simply to produce a better score. They should answer practical questions: Does the model generalise? Where does it fail? Is it useful at the required cost and latency? Does it remain safe and fair for the people who will use it?

    For Indian AI products, evaluation often needs to cover multilingual inputs, code-mixed language, varied accents, low-connectivity environments, and uneven data quality. A model that performs well on a clean English benchmark may behave very differently on Hindi-English queries, regional spellings, scanned documents, or noisy mobile audio. Treat these conditions as core evaluation requirements rather than optional edge cases.

    Teams building language systems can start by defining representative test slices using low-resource language datasets for AI training in India. For vision-language products, the same principle applies to Indian scripts, lighting conditions, document layouts, and local objects.

    Build an evaluation plan before changing the model

    Write an evaluation plan before running experiments. This prevents teams from selecting metrics after seeing favourable results and makes comparisons between model versions defensible.

    Specify:

    • The task and users: Define the exact workflow, user groups, acceptable errors, and human fallback.
    • The baseline: Record the current model, a simple heuristic, and—where possible—a human or production benchmark.
    • The evaluation splits: Keep training, development, test, and production-monitoring data separate. Prevent duplicates, near-duplicates, and user-level leakage across splits.
    • The success thresholds: Set minimum quality, latency, cost, safety, and fairness thresholds before testing.
    • The slices: Break results down by language, dialect, geography, device, document type, class frequency, and other factors that can change outcomes.
    • The decision rule: Decide whether a result leads to deployment, more data collection, rollback, or another training run.

    A fixed, versioned test set is valuable, but it should not be your only test. Maintain a hidden holdout set and a continuously refreshed set of production-like examples so that optimisation does not overfit to a public benchmark.

    Choose metrics that match the real risk

    Accuracy is often inadequate, especially when classes are imbalanced or errors have unequal consequences. Select metrics based on the product decision.

    • Classification: Report precision, recall, F1, balanced accuracy, confusion matrices, and calibration. For fraud, safety, or medical triage, examine false negatives separately.
    • Ranking and retrieval: Use recall at k, mean reciprocal rank, nDCG, and task completion rate.
    • Regression: Track MAE, RMSE, percentile errors, and error by value range. Averages can hide severe failures at the high or low end.
    • Generative AI: Measure factuality, instruction adherence, refusal quality, citation accuracy, toxicity, and human task success. Use structured rubrics and independent reviewers rather than one aggregate score.
    • Speech and vision: Report word error rate or character error rate, detection metrics, localisation quality, and performance across noise, resolution, lighting, and script conditions.
    • Operations: Measure p50 and p95 latency, throughput, memory, uptime, token or inference cost, and energy or hardware requirements.

    For language work, compare performance across scripts and code-mixed inputs. Benchmarking NLP models for Telugu and Sanskrit, for example, requires more than one headline score; it should expose where morphology, transliteration, and limited labelled data affect reliability.

    High-value post-training experiments

    1. Holdout and slice evaluation

    Run the candidate model on untouched test data, then repeat the analysis across meaningful slices. A model may improve overall while regressing badly for a smaller language group. Report confidence intervals or bootstrap ranges where sample sizes permit, and flag slices with too few examples instead of presenting unstable conclusions.

    2. Ablation and model comparison

    Change one factor at a time: fine-tuning data, prompt format, retrieval source, quantisation level, reward model, or decoding settings. Compare against the baseline using the same test set and infrastructure. Record quality gains alongside latency and cost; a small score improvement may not justify a large increase in serving expense.

    3. Robustness and stress testing

    Create targeted perturbations: spelling errors, transliteration, missing fields, long inputs, adversarial instructions, noisy audio, low-resolution images, and out-of-distribution examples. For Indian deployments, include intermittent connectivity, mobile-sized screens, regional terminology, and mixed English-language usage.

    Test graceful failure. A reliable system should ask for clarification, abstain, or route to a human when confidence is low—not produce a plausible but unsafe answer.

    4. Calibration and threshold testing

    If the model outputs probabilities or confidence scores, check whether they reflect actual correctness. Reliability diagrams, expected calibration error, and threshold curves help set operating points. Evaluate several thresholds against business costs, such as manual review load, missed cases, and false alerts.

    5. Fairness and subgroup analysis

    Measure quality and error rates for relevant groups without exposing sensitive data unnecessarily. Compare false-positive and false-negative rates, abstention rates, and access or completion outcomes. Fairness analysis should involve domain experts and affected users; statistical parity alone cannot establish that a system is appropriate.

    6. Online, shadow, and longitudinal tests

    Before a full launch, run the new model in shadow mode so it receives realistic traffic without influencing decisions. Then use a carefully designed A/B or canary test with rollback controls. Monitor drift over time: input distributions, label rates, user behaviour, latency, and error reports can all change after release.

    Make experiments reproducible

    Every result should be traceable to a model artifact, dataset version, code commit, configuration, random seed, hardware profile, and evaluator version. Tools such as MLflow or Weights & Biases can help, but a lightweight system of versioned files and documented run IDs is better than undocumented notebooks.

    Maintain an evaluation report containing:

    • The hypothesis and baseline
    • Dataset provenance, consent, licensing, and known limitations
    • Metrics, slices, sample sizes, and uncertainty
    • Qualitative failure examples with sensitive information removed
    • Cost, latency, and infrastructure measurements
    • Safety, fairness, and privacy findings
    • The deployment decision, owner, and follow-up date

    For models that must run within Indian data-governance or connectivity constraints, include a local or private-serving experiment. Guidance on deploying large language models locally is especially relevant when data cannot be sent to an external API. Deployment experiments should also test quantisation, batching, fallback behaviour, and recovery after service interruption.

    A practical 2026 workflow

    Use this sequence for a disciplined post-training cycle:

    1. Define the user outcome, risk levels, and release thresholds.
    2. Freeze a representative holdout set and create a separate challenge set.
    3. Establish baseline quality, cost, latency, and human-review measures.
    4. Run automated metrics and slice analysis.
    5. Inspect failures with domain experts and label recurring error categories.
    6. Test robustness, calibration, fairness, privacy, and operational limits.
    7. Compare candidate changes using the same protocol.
    8. Run shadow or canary deployment with monitoring and rollback.
    9. Document the decision and schedule a drift review.

    For production systems, connect monitoring to an incident process. Define who investigates a regression, how users can report harmful outputs, how data is retained, and when a model is withdrawn. If you are evaluating a vision model, related implementation choices are covered in evaluating OpenRouter vision models for video understanding, while teams serving models at scale can review how to deploy deep learning models on GKE.

    Common mistakes to avoid

    • Optimising a public benchmark while ignoring production-like data
    • Reporting only averages and hiding subgroup regressions
    • Changing test data or prompts between model comparisons
    • Treating automated judges as ground truth for generative outputs
    • Ignoring latency, cost, privacy, and human-review capacity
    • Running A/B tests without guardrails, statistical planning, or rollback
    • Collecting user data for evaluation without clear governance and consent

    FAQ

    How many examples are needed?

    There is no universal number. Use enough examples to estimate the metrics and slices that matter, and report uncertainty. Rare but high-impact failure modes may need targeted challenge sets rather than large random samples.

    Should post-training experiments use human evaluation?

    Yes, when quality involves relevance, safety, factuality, tone, or usefulness. Use a written rubric, trained reviewers, blinded comparisons, and agreement checks. Automated metrics are useful for scale but should not be the sole release criterion.

    What is the difference between offline and online evaluation?

    Offline evaluation uses fixed datasets and is safer for controlled comparisons. Online evaluation measures behaviour under real traffic, including latency, user completion, distribution shift, and operational failures. Mature teams use both.

    Apply for AI Grants India

    If your Indian AI venture is building a reliable evaluation pipeline, multilingual model, or production-ready application, apply for AI Grants India with a clear problem statement, evidence from post-training experiments, and a plan for responsible deployment.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.