0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · retrieval models sota

Retrieval Models SOTA: Benchmarks, Methods and Trends

  1. aigi

    Retrieval models SOTA (state of the art) are the strongest-performing information retrieval systems reported for a defined benchmark, task, language, dataset and evaluation protocol. In practice, there is no single universal winner: a model that leads on English passage retrieval may lose on multilingual search, domain-specific documents, long-context evidence or latency-constrained production workloads.

    For engineers building semantic search, recommendation, enterprise search or retrieval-augmented generation (RAG), the useful question is not simply “which retrieval model is SOTA?” It is: which retriever provides the best recall, ranking quality, robustness, cost and latency for my corpus and query distribution? This guide maps the modern retrieval landscape, explains the main benchmarks and metrics, and gives a practical framework for selecting and evaluating a model.

    What “retrieval models SOTA” means

    Information retrieval has two major stages:

    • Candidate retrieval: Find a relatively small set of potentially relevant documents from a large corpus.
    • Reranking: Reorder those candidates using a more expressive model, often a cross-encoder or instruction-following language model.

    A retrieval model is usually evaluated by whether it returns relevant passages, documents, products or entities near the top of the result list. SOTA claims are meaningful only when the following details are specified:

    • Dataset and split, such as MS MARCO, BEIR, MIRACL or a private corpus
    • Retrieval unit: passage, document, page, chunk or answer-bearing span
    • Language and script
    • Indexing and chunking strategy
    • Number of retrieved candidates, commonly top-k
    • Whether relevance labels are binary, graded or generated
    • Whether the model uses external training data or benchmark-specific supervision
    • Hardware, quantisation, batch size and latency constraints

    A model may therefore be SOTA on a leaderboard while being unsuitable for a production system with Indian-language queries, rapidly changing documents or strict response-time requirements.

    The main classes of modern retrieval models

    Sparse lexical retrieval

    Sparse methods represent text using exact or weighted terms. BM25 remains a strong baseline because it is fast, interpretable and effective when query words overlap with document words. Modern learned sparse systems, including expansion-based approaches, predict additional terms or importance weights while retaining an inverted-index architecture.

    Sparse retrieval is particularly useful for:

    • Product names, legal clauses, IDs and error codes
    • Exact quotations and rare entities
    • Fresh content that cannot be repeatedly embedded
    • Transparent enterprise search
    • Low-cost, high-throughput candidate generation

    Its weakness is vocabulary mismatch. A query using “vehicle financing” may not match a document that only says “car loan” unless the index or query-expansion system bridges the terms.

    Dense bi-encoder retrieval

    Dense retrievers encode a query and a document independently into vectors. Similarity is commonly measured using cosine similarity or a dot product. Because document embeddings can be precomputed, dense retrieval scales efficiently with approximate nearest-neighbour indexes such as HNSW, IVF or product quantisation.

    The standard pipeline is:

    1. Split and normalise documents.
    2. Encode each document or chunk.
    3. Build a vector index.
    4. Encode the user query at request time.
    5. Retrieve the nearest vectors.
    6. Optionally rerank the candidates.

    Dense models are strong at semantic matching, paraphrases and natural-language questions. However, they can miss exact identifiers, numbers, code tokens and newly emerging terminology. Embedding dimension, pooling strategy, chunk length and similarity function also affect results.

    Late-interaction retrieval

    Late-interaction models retain token-level representations rather than compressing an entire document into one vector. Query and document token embeddings are compared at retrieval time using a compressed interaction function. This preserves more fine-grained matching information than a single-vector bi-encoder while remaining more scalable than a full cross-encoder.

    The trade-off is index size and serving complexity. Late-interaction systems can be attractive when passage-level relevance and difficult semantic matching justify higher storage and compute costs.

    Cross-encoder rerankers

    A cross-encoder reads the query and candidate document together, allowing full attention between both sequences. It generally provides better pairwise relevance judgements than a bi-encoder, but every query-document pair must be scored at request time.

    A common high-quality architecture is hybrid retrieval followed by reranking:

    • Retrieve 50–1,000 candidates using sparse and dense methods.
    • Fuse the candidate lists.
    • Score the top 50–200 candidates with a cross-encoder.
    • Return the top results or pass them to a RAG generator.

    This design often outperforms an attempt to use a very large model for first-stage retrieval across the entire corpus.

    Benchmarks used to compare retrieval models

    BEIR

    BEIR is a heterogeneous zero-shot evaluation suite covering multiple information-retrieval tasks and domains. It is valuable because it tests transfer beyond a single training distribution. Results can expose whether a model relies too heavily on benchmark-specific patterns.

    When reading BEIR results, check whether the reported score is averaged across datasets, whether a model was trained on overlapping data and whether the evaluation uses the original official protocol.

    MS MARCO

    MS MARCO is a widely used passage-retrieval benchmark built from real search queries. It has driven progress in dense retrieval, hard-negative mining and reranking. Strong MS MARCO performance is useful evidence, but it should not be treated as proof of general-domain or multilingual superiority.

    MIRACL and multilingual benchmarks

    MIRACL evaluates retrieval across many languages and is more relevant for systems serving multilingual users. For India-focused deployments, benchmark language coverage is only a starting point. Hindi, Bengali, Tamil, Telugu, Marathi and other languages have different morphology, spelling variation, code-mixing patterns and content availability.

    A retriever should be tested on real queries such as Hindi-English code-mixed questions, transliterated terms and regional-language variants rather than relying only on an aggregate multilingual score.

    Domain-specific evaluations

    Legal, medical, financial, scientific and technical search require their own test sets. A domain benchmark should include difficult cases:

    • Synonyms and abbreviations
    • Conflicting or outdated documents
    • Tables and structured fields
    • Long documents with localised evidence
    • Negation and numerical constraints
    • Queries with multiple requirements
    • Unanswerable questions

    For RAG, retrieval evaluation should also measure whether the returned evidence actually supports the generated answer, not merely whether it contains similar words.

    Metrics that define SOTA retrieval

    Recall@k

    Recall@k measures whether at least one relevant item appears in the top k results. It is critical for candidate generation. If the correct passage is absent from the candidate set, no reranker or language model can recover it.

    For RAG systems, track recall at the number of chunks passed into the context window. Recall@5, Recall@10 and Recall@20 often reveal different operating points.

    MRR and nDCG

    Mean Reciprocal Rank rewards placing the first relevant result high in the list. It is useful when users typically need one best answer. Normalised Discounted Cumulative Gain (nDCG) supports graded relevance and is better when several results may be useful with different importance levels.

    Precision@k

    Precision@k measures how many of the top k results are relevant. It matters when context windows, analyst attention or downstream processing are limited. A system can have strong recall but poor precision if it retrieves many loosely related chunks.

    Latency, throughput and cost

    A production definition of SOTA must include:

    • p50, p95 and p99 query latency
    • Queries per second under realistic concurrency
    • CPU, GPU and memory requirements
    • Vector-index size and update time
    • Embedding cost for new documents
    • Cost per thousand or million searches
    • Time required for reranking

    A small model that achieves slightly lower nDCG but cuts latency by 80% may be the better engineering choice.

    Training techniques behind leading retrievers

    Contrastive learning

    Dense retrievers commonly train a query encoder and passage encoder so positive pairs are close and negative pairs are distant. In-batch negatives make each other passage a negative for the current query, improving training efficiency.

    The quality of negatives is crucial. Random negatives are usually too easy. Hard negatives retrieved by BM25, an earlier dense model or a teacher reranker force the model to distinguish semantically similar but incorrect passages.

    Distillation

    Knowledge distillation transfers relevance signals from a stronger teacher, such as a cross-encoder, to a faster bi-encoder. Distillation can improve ranking quality without requiring the production model to perform expensive joint query-document attention.

    Instruction tuning

    Instruction-aware embedding models are trained to interpret task-specific prompts, including “represent this query for retrieving passages” or “represent this document for semantic search.” The instruction format must be used consistently at indexing and query time when the model documentation requires it.

    Synthetic data and self-training

    Large language models can create queries, positive passages, explanations and hard negatives from unlabeled documents. Synthetic data expands coverage, but it can also introduce artificial language patterns, label errors and confirmation bias. Validate synthetic examples against human-labelled queries before using them as the sole training source.

    Multilingual and domain adaptation

    Fine-tuning on local queries and documents can substantially improve performance. For India, include English, regional languages, transliteration and code-mixed examples when they occur in production. Preserve difficult entities, numerals, addresses, product SKUs and government scheme names during tokenisation and evaluation.

    Hybrid retrieval is often the practical SOTA

    Hybrid retrieval combines lexical and semantic signals. A simple implementation merges BM25 and dense results using Reciprocal Rank Fusion (RRF), which combines ranks without requiring scores to be calibrated. A learned fusion model can use features such as BM25 score, vector similarity, field matches, freshness and document authority.

    Hybrid retrieval is effective because the two methods fail differently:

    • Sparse search captures exact names, codes and rare terms.
    • Dense search captures paraphrases and conceptual similarity.
    • Metadata filters enforce permissions, date ranges, geography and product constraints.
    • Reranking resolves close semantic alternatives.

    For enterprise and RAG systems, this layered approach is usually more reliable than replacing every lexical component with embeddings.

    How to evaluate a retrieval model properly

    Start with a representative query set rather than a public leaderboard. Sample queries by language, department, intent, length, difficulty and traffic volume. Create graded relevance labels where possible, and record the exact document version and chunk that should support the answer.

    Use a repeatable offline evaluation:

    1. Establish BM25 and a simple dense baseline.
    2. Evaluate Recall@k, MRR or nDCG at realistic k values.
    3. Compare chunk sizes and overlap independently from model changes.
    4. Test hybrid fusion and reranking.
    5. Measure latency, memory and index update cost.
    6. Inspect failures manually by category.
    7. Run online A/B tests using click quality, reformulation rate, resolution rate and user feedback.

    Avoid data leakage. If queries or documents from the evaluation set were used during training, label the result accordingly. Also avoid tuning repeatedly on the test set; use a development set for iteration and reserve the test set for final comparison.

    Choosing a retrieval model for production

    A useful decision framework is:

    • Need exact matching: Start with BM25 or a learned sparse method.
    • Need semantic search: Add a strong dense bi-encoder.
    • Need high recall across varied queries: Use hybrid retrieval.
    • Need top-ranked precision: Add a cross-encoder reranker.
    • Need multilingual Indian-language support: Validate language-specific and code-mixed queries locally.
    • Need low latency at scale: Prefer compact embeddings, approximate indexes and bounded reranking.
    • Need strict freshness: Re-embed incrementally and retain lexical retrieval for new content.
    • Need access control: Apply security filters before exposing retrieved content to the generator.

    Model selection should follow corpus characteristics and service-level objectives, not leaderboard rank alone.

    Common mistakes when chasing retrieval SOTA

    • Comparing scores from incompatible benchmarks or protocols
    • Ignoring chunking, metadata and document parsing quality
    • Using only semantic retrieval for exact identifiers
    • Measuring generated answer quality without checking evidence recall
    • Reranking too many candidates and violating latency budgets
    • Failing to test permissions and document freshness
    • Treating multilingual averages as proof of Indian-language performance
    • Publishing a single aggregate score without confidence intervals or failure analysis

    A strong retrieval system is an end-to-end pipeline. Better OCR, table extraction, deduplication, metadata filters and query rewriting can produce larger gains than changing the embedding model.

    FAQ: Retrieval models SOTA

    What is the current SOTA retrieval model?

    There is no universal current winner. Results depend on benchmark, language, domain, retrieval stage, candidate count and hardware. Compare leading dense, sparse and hybrid systems on your own evaluation set.

    Are dense retrievers better than BM25?

    Not always. Dense retrievers are strong for semantic similarity, while BM25 excels at exact terms, rare entities and fresh content. Hybrid retrieval is often the safest default.

    Is a cross-encoder a retrieval model?

    It is usually used as a reranker rather than a first-stage retriever because it must jointly process each query-document pair and is more computationally expensive.

    What should I use for RAG?

    Begin with hybrid retrieval, measure evidence Recall@k, then add reranking if precision is insufficient. Evaluate the complete pipeline with real queries, languages, access rules and latency targets.

    How can Indian AI teams evaluate multilingual retrieval?

    Build a labelled test set covering English, regional languages, transliteration, code-mixing, local entities and domain terminology. Report results by language and intent instead of only one combined score.

    Apply for AI Grants India

    If you are an Indian AI founder building search, RAG, language technology or retrieval infrastructure, apply through AI Grants India for relevant grant opportunities and support. Share your technical approach, evaluation evidence and product vision to take the next step.

    Last updated 9 October 2026

AIGI may be inaccurate. Replies seeded from the guide above.