Why embedding benchmarks need an Indic lens
Embedding models turn text into vectors used for semantic search, retrieval-augmented generation, clustering, recommendation, classification, and duplicate detection. A model can perform well on an English leaderboard yet fail on Indian-language queries because of script variation, code-mixing, spelling differences, morphology, or uneven training data.
A useful benchmark therefore measures more than one headline score. It should reveal how a model handles Hindi, Bengali, Tamil, Telugu, Marathi, Kannada, Malayalam, Gujarati, Punjabi, Odia, Urdu, and mixed-language text in the domains your product serves. The methodology in this guide complements the practical concerns covered in low-resource Indic natural language processing.
Define the task before choosing a metric
Start with the production task. “Best embedding model” has no meaning without a target use case.
- Semantic search: Does the correct document appear near the top of the results?
- Question-answer retrieval: Can the model retrieve the passage needed to answer a query?
- Classification: Do embeddings separate intent, topic, sentiment, or support categories?
- Clustering: Do similar texts form coherent groups without labels?
- Duplicate detection: Can the system identify paraphrases, transliteration variants, and repeated complaints?
- Cross-lingual retrieval: Can a Hindi query retrieve a relevant English or Tamil document?
Record the task, languages, domains, latency target, maximum input length, and acceptable model size before running tests. This prevents teams from selecting a model solely because it has a strong score on an unrelated benchmark.
Select models on Hugging Face carefully
Use the Hugging Face Model Hub to shortlist models, but treat model cards as evidence rather than guarantees. Check the supported languages, training data, licence, pooling recommendation, tokenizer behaviour, context length, and whether the model is intended for sentence embeddings or only masked-language modelling.
Create a comparison set that includes at least:
- A multilingual baseline such as multilingual BERT or XLM-R-derived models.
- An Indic-focused model trained or adapted for Indian languages.
- A strong multilingual sentence-embedding model suitable for retrieval.
- A smaller model that could meet an Indian startup’s latency and infrastructure budget.
Keep model revisions pinned by commit hash where possible. Log the exact tokenizer, pooling method, quantisation setting, device, batch size, and library versions. A model’s default pooler_output is not automatically the best sentence representation; mean pooling over the attention mask is common, but follow the model’s documentation and test alternatives.
Build a representative evaluation set
Do not translate an English test set and assume it represents Indian usage. Assemble examples from the channels and domains where the system will operate, using consented, licensed, or public data.
Stratify the set by:
- Language and script, including Devanagari, Bengali-Assamese, Tamil, Telugu, Kannada, Malayalam, Gujarati, Gurmukhi, Odia, and Urdu scripts where relevant.
- Native script, Romanised text, and code-mixed text such as Hinglish or Tanglish.
- Formal, conversational, misspelled, abbreviated, and speech-transcribed text.
- Short queries, long passages, named entities, numbers, dates, addresses, and product terms.
- Domains such as education, finance, government services, healthcare, commerce, and customer support.
For retrieval, create query-document relevance labels. A practical test set can contain a few hundred carefully reviewed queries per important language, with hard negatives that look lexically similar but are semantically wrong. Split development and test data by user, source, or document—not randomly by sentence—so near-duplicates do not leak across splits.
Measure the right outcomes
For semantic retrieval, prioritise ranking metrics:
- Recall@k: Whether a relevant item appears in the first *k* results.
- MRR: Rewards the position of the first relevant result.
- nDCG@k: Handles graded relevance and multiple relevant documents.
- Precision@k: Useful when the interface displays only a small number of results.
For pair and classification tasks, use Spearman or Pearson correlation for similarity labels, and accuracy, macro-F1, and per-language F1 for classification. Macro-F1 is especially important when one Indian language or intent has fewer examples. For clustering, report adjusted Rand index or normalized mutual information only when gold labels are available; otherwise combine quantitative scores with human review.
Perplexity and BLEU are generally not the primary metrics for a sentence-embedding benchmark. Perplexity evaluates language modelling, while BLEU evaluates generated translations. They can be relevant to other model assessments, but they do not tell you whether vectors rank the right document.
Run a reproducible Hugging Face benchmark
Install a fixed environment and use deterministic seeds where supported. The following pattern produces normalised embeddings and can be adapted for each candidate model:
import torch
import torch.nn.functional as F
from transformers import AutoTokenizer, AutoModel
model_id = "your-org/your-embedding-model"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id).eval()
texts = ["आपका आवेदन कहाँ तक पहुँचा है?", "Where is my application status?"]
inputs = tokenizer(
texts, padding=True, truncation=True,
max_length=512, return_tensors="pt"
)
with torch.inference_mode():
outputs = model(**inputs)
mask = inputs["attention_mask"].unsqueeze(-1)
vectors = (outputs.last_hidden_state * mask).sum(1) / mask.sum(1).clamp(min=1)
vectors = F.normalize(vectors, p=2, dim=1)
similarity = vectors @ vectors.T
print(similarity)For a production-grade run, batch inputs, warm up the device, measure peak memory, and report throughput in texts or tokens per second. Evaluate both CPU and GPU if your deployment may run on affordable Indian cloud instances or local servers. Test truncation explicitly: long documents may lose their decisive information when a model silently cuts the input.
Compare quality, cost, and robustness
Publish a table with overall and per-language scores rather than one average. Include confidence intervals or bootstrap estimates, especially when test sets are small. A two-point improvement may not be meaningful if it falls within sampling noise.
Track these operational measures alongside quality:
- Parameters, disk size, and peak RAM or VRAM.
- Median and p95 latency at realistic batch sizes.
- Embedding generation throughput and indexing cost.
- Performance under quantisation, if used.
- Failure rates for empty input, very long input, emojis, punctuation, and mixed scripts.
- Licence and restrictions on commercial use or redistribution.
Inspect the worst-ranked examples manually. Common failure patterns include transliteration mismatch, named-entity confusion, overly generic vectors, and dominance by high-resource languages. If the system will support voice interfaces, evaluate speech-recogniser errors as well; the top-rated voice agent services for Indian businesses illustrate why spoken, code-mixed language can differ sharply from clean text.
Avoid benchmark mistakes
Do not compare models with different preprocessing, pooling, truncation, or normalisation settings and call the result fair. Do not let translated test examples, training documents, or near-duplicate customer tickets enter the evaluation set. Do not report only an aggregate score that hides poor performance in a smaller language.
Human evaluation remains necessary. Ask native or highly proficient reviewers to judge the top retrieved results using a clear rubric: relevant, partially relevant, irrelevant, or unsafe. For sensitive domains, add checks for privacy leakage, harmful associations, and incorrect retrieval of medical, legal, or financial information.
Turn results into a deployment decision
Choose the model that meets the product threshold, not necessarily the leaderboard winner. A slightly weaker model may be preferable if it is faster, easier to host, commercially licensed, and more consistent across languages. Keep a frozen test set for release decisions, and add anonymised production failures to a separate regression suite after review.
Document the benchmark as a reproducible asset: dataset version, label guidelines, model commit, code, hardware, metrics, and known limitations. Teams building open tools can also learn from Indian open-source AI developer projects and contribute language-specific evaluation data back to the community.
As of 2026, the strongest Indic embedding workflow is not simply “download a model and calculate cosine similarity.” It is a controlled comparison across languages, scripts, domains, retrieval quality, infrastructure cost, and real user failure modes.