What the LM Evaluation Harness does
EleutherAI’s LM Evaluation Harness gives builders a consistent way to run language-model benchmarks across models, datasets, and hardware. It supports Hugging Face Transformers models, common-generation and multiple-choice tasks, batch evaluation, reproducible configuration, and result exports. For Indian language models, it is useful because a single headline score rarely captures performance across Hindi, Tamil, Telugu, Bengali, Marathi, Kannada, Malayalam, Gujarati, Punjabi, Odia, Assamese, or code-mixed text.
The harness is not a substitute for a carefully designed Indic benchmark. It is the execution layer: you still need to verify language coverage, script quality, contamination risk, prompt format, and whether the task reflects your product. This matters especially for low-resource Indic NLP, where spelling variation, transliteration, morphology, tokenisation, and uneven training data can distort results. The low-resource Indic natural language processing guide is a useful companion when designing that evaluation set.
Install a reproducible environment
Use a fresh virtual environment rather than installing into a system Python setup. GPU support depends on your PyTorch installation and CUDA version.
python -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
pip install lm-eval transformers accelerate datasets sentencepiece
lm_eval --versionOn Windows, activate the environment with .venv\\Scripts\\activate. For larger models, install the appropriate PyTorch build first, then add bitsandbytes only if you have verified that quantisation is supported by your hardware and model.
Pin package versions in requirements.txt or a lockfile. Record the model revision, tokenizer revision, dataset revision, GPU type, precision, batch size, and harness version with every run. A result that cannot be reproduced is a weak basis for a model or product decision.
Load a Hugging Face model correctly
The harness can load many causal and sequence-to-sequence models through its Hugging Face integration. For a causal language model, begin with a small smoke test:
lm_eval \\
--model hf \\
--model_args pretrained=org/model-name,revision=main \\
--tasks hellaswag \\
--device cuda:0 \\
--batch_size 4 \\
--limit 10The example task is only a connectivity check, not an Indic evaluation. Replace it with tasks that match your model’s architecture and intended use. Use --device cpu for a first test on a machine without a GPU. --batch_size auto can help identify a workable batch size, but save the final resolved configuration in your experiment log.
For a gated or private Hugging Face repository, authenticate before running the harness:
huggingface-cli loginSome repositories provide a base model but not a compatible tokenizer configuration. Check the model card, tokenizer files, chat template, context length, and supported languages before debugging the evaluation command. A model trained for masked-language modelling is not evaluated like a causal chat model; architecture and objective must match the harness adapter and task.
Choose Indic tasks instead of chasing one score
Start with established multilingual or language-specific tasks available in your installed harness version. List tasks locally:
lm_eval --tasks listThen inspect a task before running it:
lm_eval --include_path lm_eval/tasks --tasks listTask availability and names change across releases, so treat the local installation as authoritative. For an Indic model, create an evaluation matrix covering:
- Language understanding: classification, natural-language inference, sentiment, and topic recognition in each target language.
- Knowledge and reasoning: multiple-choice tasks, with translated or native-language prompts where possible.
- Generation: perplexity or exact-match tasks on held-out Indic text, if the model and dataset format support them.
- Cross-lingual behaviour: the same intent expressed in multiple Indian languages, including code-mixed and transliterated inputs.
- Safety and reliability: harmful-content refusal, hallucination checks, and culturally specific edge cases relevant to your application.
Do not average languages into one number without reporting the per-language scores. A model can look strong overall because it performs well in Hindi and English while failing on Assamese or Malayalam. For production decisions, examine macro-averages, worst-language performance, confidence intervals where feasible, and the gap between native script and Romanised input.
Run a standard evaluation
A typical run looks like this:
lm_eval \\
--model hf \\
--model_args pretrained=org/indic-model,trust_remote_code=False \\
--tasks task_name \\
--device cuda:0 \\
--batch_size auto \\
--output_path results/indic-model \\
--log_samplesUse trust_remote_code=True only after reviewing the repository and understanding the code being executed. Save --log_samples outputs for error analysis, but remove sensitive or personally identifiable data before sharing artifacts.
For generative tasks, prompt formatting is part of the experiment. Compare a plain prompt with the model’s documented instruction or chat template, and report exactly which format you used. Small changes in language labels, punctuation, answer options, or few-shot examples can materially change results in Indic languages.
Evaluate your own dataset
When no built-in task reflects your application, create a custom task configuration. A practical dataset should contain a stable split, explicit language metadata, a prompt or input field, and a reference answer or target. Keep test data private if it is intended to measure future product performance.
A custom task configuration normally specifies the dataset, split, document fields, prompt template, target, and metric. The exact YAML fields depend on the harness release, so copy the structure of a nearby built-in task and validate it with a small run. Store the configuration in version control and include examples for every language and script.
Before trusting the metric, test for:
- Unicode normalisation differences, especially combining marks and nukta characters.
- Multiple valid spellings and orthographic conventions.
- Transliteration, code-mixing, and punctuation patterns common in real Indian user input.
- Label imbalance and translation artefacts.
- Leakage from training data, public benchmarks, or repeated prompts.
- Human agreement, particularly for open-ended translation, summarisation, and safety labels.
Exact match is often too strict for free-form Indic generation. Depending on the task, use token-level F1, character or subword similarity, semantic similarity reviewed by native speakers, or a human rubric. Automated metrics should support—not replace—native-speaker quality review.
Make results fair and useful
Use the same decoding settings, number of examples, context limits, and evaluation splits when comparing models. Report whether the model was evaluated in zero-shot, few-shot, or instruction-tuned mode. Include tokenisation statistics: an Indic model that uses unusually many tokens per sentence may incur higher latency and cost even when its task score is competitive.
Benchmark both open-source baselines and the candidate model. Track latency, peak memory, throughput, and quantised-versus-full-precision differences alongside quality. If the target product is a voice assistant, text benchmarks alone are insufficient; pair them with speech recognition and response tests. Teams building multilingual customer experiences may also benefit from reviewing voice agent services for Indian businesses and the related implementation considerations.
Use a results table with at least these columns: model and revision, language, script, task, prompt format, shots, metric, score, hardware, precision, harness version, and evaluation-set version. Publish the command and configuration where licensing and privacy allow. Re-run a small fixed regression suite on every model or prompt change.
Common failure modes
- Task not found: update the harness or use the task name listed by your installed version.
- Tokenizer errors: verify
tokenizer.json, SentencePiece files, special tokens, and the model card’s loading instructions. - Out-of-memory failures: reduce batch size, use a smaller model, enable supported quantisation, or evaluate by language in separate jobs.
- Unexpectedly poor scores: inspect rendered prompts and logged samples before changing the model. Language tags, answer ordering, and chat templates are frequent causes.
- Misleading high scores: check contamination, duplicate examples, translation shortcuts, and whether the benchmark measures your real use case.
- Unsafe remote execution: avoid unreviewed custom code and pin trusted revisions.
A practical evaluation workflow
1. Define the product decision and target languages.
2. Run a ten-example smoke test to confirm model and tokenizer loading.
3. Select standard tasks plus a native-speaker-reviewed custom set.
4. Evaluate zero-shot and, where relevant, few-shot or instruction prompts.
5. Log samples, inspect errors by language and script, and measure operational cost.
6. Repeat the run with pinned versions and publish a per-language report.
7. Add the strongest failure cases to a regression suite.
The harness is most valuable when it turns evaluation into a repeatable engineering process. For Indian language models, that means resisting a single leaderboard number and building evidence across languages, scripts, user intents, safety cases, and deployment constraints. Teams also exploring Indian open-source AI developer projects can use the same discipline to make model comparisons easier to audit and extend.