Fine-tuning is only useful when you can demonstrate that the adapted model is better for its intended job. A lower training loss is not enough: the model may overfit, fail on rare inputs, regress on general capability, or behave poorly in the languages and formats your users actually provide.
This guide explains how to run evaluation after fine tuning with Hugging Face MCP—using the Hugging Face ecosystem and Model Card Playground (MCP) as part of a repeatable evaluation workflow. The exact MCP interface can change, so treat it as a presentation and testing layer around a properly defined dataset, evaluation script, and model artefact rather than as a replacement for measurement.
What to evaluate after fine-tuning
Start by writing down the model’s job and the failure that matters most. A classifier, instruction-following model, translation system, and retrieval-augmented assistant need different tests.
Evaluate at four levels:
- Task quality: accuracy, F1, exact match, ROUGE, BLEU, word error rate, or another task-appropriate measure.
- Reliability: performance across lengths, topics, writing styles, languages, and difficult examples.
- Safety and policy behaviour: refusal quality, privacy leakage, unsafe completion rates, and resistance to prompt injection where relevant.
- Operational fitness: latency, memory use, throughput, context-window behaviour, and inference cost.
For Indian deployments, do not assume an English benchmark predicts performance in Hindi, Marathi, Tamil, Bengali, or code-mixed inputs. Use representative regional data and document the script, dialect, spelling variation, and transliteration conventions. The Indian language LLM benchmark datasets guide is useful when designing this coverage.
1. Freeze a clean evaluation split
Keep three distinct datasets:
- Training data: used to update model weights.
- Validation data: used during model and hyperparameter selection.
- Test data: held back until the final comparison.
Avoid evaluating on examples copied from training conversations, synthetic prompts generated from the same templates, or near-duplicates of training records. Deduplicate by prompt, document, user, and source where possible. If records come from the same customer, patient, school, or conversation, split by entity or conversation—not randomly—so related examples cannot leak across partitions.
Store the test set in a versioned format such as JSONL or a Hugging Face Dataset. Include stable IDs, the input, the reference answer or label, the task type, and useful slices such as language, difficulty, domain, and source. Never place secrets, personal data, Aadhaar numbers, phone numbers, or confidential customer content in a public repository or Model Card.
Before running MCP, define a baseline: the original base model, an existing production model, or a simple heuristic. A fine-tuned model should beat a meaningful baseline on the target task without unacceptable regressions elsewhere. For guidance on dataset design and training decisions, see best practices for fine-tuning LLMs on custom data.
2. Load the model and dataset reproducibly
Install the libraries needed for your task and pin versions in a requirements file or lockfile:
pip install -U transformers datasets evaluate accelerate scikit-learnLoad the exact model revision produced by training. If the model is hosted privately on the Hugging Face Hub, authenticate through an environment variable or secret manager rather than hard-coding a token.
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForSequenceClassification
model_id = "org-or-user/fine-tuned-model"
revision = "COMMIT_OR_TAG"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForSequenceClassification.from_pretrained(
model_id, revision=revision
)
test_set = load_dataset("json", data_files="data/test.jsonl", split="train")For causal language models, use AutoModelForCausalLM and generate outputs with fixed decoding settings. Record the model revision, tokenizer revision, hardware, precision, maximum input length, generation parameters, and random seed. These details make a future MCP comparison meaningful.
3. Run task-specific metrics
For classification, calculate macro F1 alongside accuracy. Macro F1 gives minority classes equal weight; this matters when a rare but costly class is underrepresented. Inspect a confusion matrix and per-class support rather than reporting one headline number.
For generation, combine automatic metrics with human or expert review. Exact match can be appropriate for structured extraction, while ROUGE or BLEU may help with summarisation and translation but should not be treated as a complete quality judgment. For open-ended answers, assess factuality, instruction adherence, citation correctness, format compliance, and refusal behaviour. LLM-as-judge scoring can accelerate triage, but calibrate it against human ratings and check for verbosity, language, and model-family bias.
A simple classification evaluation with the Transformers trainer looks like this:
import numpy as np
from sklearn.metrics import accuracy_score, f1_score
# predictions should come from your Trainer or inference loop
predicted = np.array([0, 1, 1, 0])
actual = np.array([0, 1, 0, 0])
results = {
"accuracy": accuracy_score(actual, predicted),
"macro_f1": f1_score(actual, predicted, average="macro"),
}
print(results)Do not use model.predict() on a raw Transformers model; that method belongs to specific trainer or pipeline APIs. Use a tested inference loop, Trainer.predict(), or a task pipeline, and make preprocessing identical between evaluation and production.
4. Use Hugging Face MCP for inspection and comparison
Once the evaluation artefact and model are available, use MCP to inspect the model card, run representative examples, and compare outputs across the base and fine-tuned versions. Upload or connect only the data permitted by your organisation’s privacy policy. A useful MCP review includes:
- The model revision and intended use.
- Evaluation-set provenance and split method.
- Overall and slice-level metrics.
- Several correct, borderline, and failed examples.
- Known limitations, excluded cases, and safety considerations.
- Hardware and inference settings used for the reported results.
MCP is most valuable when it makes failures visible to reviewers. Create fixed prompt suites for common user journeys instead of testing only a few impressive examples. For production-grade workflows, pair it with best tools for LLM evaluation and experiment tracking or an automated regression system.
5. Perform slice-based error analysis
A single score can hide serious gaps. Break results down by:
- Indian language, dialect, script, and transliteration style.
- Input length and formatting.
- Domain, label, or intent.
- New versus familiar entities.
- Ambiguous, adversarial, or low-quality inputs.
- Demographic or geographic categories, only where lawful and ethically justified.
Read failed examples and classify the cause: insufficient training coverage, incorrect labels, truncation, tokenisation problems, retrieval errors, prompt ambiguity, hallucination, or a genuine capability limit. Then decide whether to improve the data, change the prompt, alter the model, add retrieval, or reject the use case. Do not automatically fine-tune again; repeated training on the same failure pattern can amplify label noise and overfitting.
If the model serves regional-language users, test code-mixing and spelling variation explicitly. Projects involving Marathi, Sanskrit, or other lower-resource settings may need custom benchmarks and expert reviewers; compare your process with work on fine-tuning AI models for Marathi dialect and fine-tuning large language models for Sanskrit translation.
6. Set a release decision, not just a score
Define pass/fail thresholds before looking at the final results. A practical release gate might require:
- Macro F1 above the agreed target on the test set.
- No critical slice below its minimum threshold.
- No statistically meaningful regression against the base model on general prompts.
- Safety and privacy checks passed.
- Latency and memory within the deployment budget.
- Model card, dataset version, licence, and limitations documented.
Use confidence intervals or bootstrap resampling when the test set is small. For high-risk applications, conduct a pilot with human escalation and monitor real-world error rates. Re-evaluate after changes to the model, tokenizer, prompt, retrieval index, data, or serving stack—not only after another fine-tuning run.
Evaluation checklist
Before publishing or deploying the model, confirm that you have:
- A leakage-checked, versioned test set.
- A baseline and predeclared metrics.
- Overall, per-class, and slice-level results.
- Representative failure examples reviewed by domain experts.
- Safety, privacy, and licence checks.
- Reproducible code, model revision, and inference settings.
- A completed Model Card with intended use and limitations.
- A monitoring plan for drift and user feedback.
The core principle is simple: use MCP to make evaluation accessible and reviewable, but keep the measurement process independent, reproducible, and tied to the real Indian deployment context. A model is ready when it meets the agreed quality, safety, and operational gates—not merely when it produces plausible examples in a playground.