0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build scalable rag applications

How to Build Scalable RAG Applications in Production

  1. aigi

    Retrieval-Augmented Generation (RAG) is easy to demonstrate and difficult to operate. A prototype can load documents, run a similarity search, and pass the results to an LLM. A production system must also handle changing data, tenant isolation, traffic spikes, model failures, multilingual queries, predictable latency, and an answer that can be audited.

    For Indian startups, the design challenge is sharper: budgets are often constrained, users may mix English with Hindi or other Indian languages, and enterprise customers increasingly ask where their data is processed. The right approach is to treat RAG as a distributed application with an evaluation loop, not as a vector database attached to a chatbot.

    Start with a measurable production target

    Before selecting a framework or database, define the workload. Record:

    • Number of documents, pages, chunks, and daily updates
    • Concurrent users and expected peak requests per second
    • Target time to first token and complete-answer latency
    • Maximum acceptable retrieval and generation cost per request
    • Required answer-grounding rate and citation coverage
    • Tenant, role, geography, and retention constraints

    Separate the request path from the indexing path. User queries need predictable latency; document parsing, OCR, embedding, and re-indexing can run asynchronously. Teams already operating queues and services can apply the same principles described in this guide to scaling backend infrastructure for AI applications.

    Build an incremental ingestion pipeline

    A scalable RAG application begins with durable source data and repeatable processing. Store original files in object storage, keep document metadata in a transactional database, and treat parsed text and embeddings as derived data. Every chunk should carry a stable document ID, version, tenant ID, source location, language, timestamp, and access-control metadata.

    Use an ingestion workflow with these stages:

    1. Detect new, changed, and deleted files.
    2. Parse text, tables, images, and scanned pages with the appropriate extractor or OCR service.
    3. Normalise encoding, remove boilerplate, and preserve headings and page references.
    4. Split content according to structure rather than an arbitrary character count.
    5. Generate embeddings in batches and write them idempotently.
    6. Mark an index version as active only after validation completes.

    Semantic or structure-aware chunking is usually better than fixed windows, but it is not automatically correct. Keep headings with their content, avoid splitting tables mid-row, and include limited parent context when a passage depends on a larger section. Test chunk sizes against real questions rather than optimising for a benchmark alone.

    For large corpora, parallelise parsing and embedding while respecting provider rate limits. Use a dead-letter queue for malformed files and make retries safe. Incremental updates should replace only affected document versions; rebuilding an entire index for every change is expensive and creates avoidable downtime.

    Choose retrieval architecture deliberately

    A vector database is only one component. PostgreSQL with pgvector can be a strong starting point when relational filters, permissions, and transactional updates matter. Dedicated systems such as Qdrant, Milvus, Weaviate, or managed vector platforms become useful when index size, ingestion throughput, or operational requirements justify another service.

    Evaluate the following capabilities:

    • Filtering: tenant, user, product, language, date, and permission filters must be applied before or during retrieval.
    • Sharding and replication: understand how the system handles growth, failover, and uneven tenant sizes.
    • Index tuning: HNSW provides fast approximate search, while other indexes may reduce memory use or suit very large collections.
    • Backup and rebuilds: test recovery from a clean snapshot, not just a dashboard toggle.
    • Operational limits: measure memory, write throughput, compaction behaviour, and p95 latency under realistic load.

    Never rely on application-side filtering after retrieval. That pattern can leak restricted content and wastes search capacity. Enforce access control in the retrieval query and validate it again before context reaches the model.

    Use a multi-stage retrieval pipeline

    Single-pass top-k vector search is a useful baseline, not a production strategy. A robust pipeline commonly looks like this:

    1. Classify the query and apply tenant and security filters.
    2. Optionally rewrite ambiguous queries or generate focused sub-queries.
    3. Run hybrid retrieval using dense vectors and keyword search such as BM25.
    4. Merge and deduplicate candidates.
    5. Re-rank a manageable candidate set with a cross-encoder or specialised reranker.
    6. Select compact evidence with diversity and token limits.
    7. Generate an answer with citations and explicit uncertainty rules.

    Hybrid retrieval matters for exact identifiers, policy names, SKUs, legal terms, and Indian addresses that embeddings may represent poorly. Query expansion can improve recall, but it adds latency and may introduce drift; enable it selectively for difficult or underspecified questions.

    Reranking 50–200 candidates and passing the best five to the LLM often offers a better quality-cost balance than sending a large, noisy context. Add parent-child retrieval when small chunks match well but require surrounding section context to answer correctly.

    Design for latency, reliability, and cost

    Instrument each stage separately: queue wait, parsing, embedding, retrieval, reranking, prompt construction, first token, and completion. Set budgets for each stage so a slow dependency cannot consume the entire request timeout.

    Practical controls include:

    • Stream answers to reduce time to first token while preserving a complete response contract.
    • Cache embeddings, normalised queries, retrieval results, and safe final responses at different TTLs.
    • Use asynchronous jobs for bulk analysis, long PDFs, and report generation.
    • Apply circuit breakers, timeouts, retries with jitter, and fallback models.
    • Limit context by relevance, token budget, and duplicate-content detection.
    • Route simple questions to smaller models and reserve larger models for synthesis or difficult reasoning.
    • Batch embedding jobs and run them during lower-cost compute windows.

    Semantic caching needs care. Cache keys should include tenant, permissions, retrieval configuration, model version, and relevant data version. A cached answer that ignores document updates or access changes is a correctness and security failure.

    Model choice should reflect the workload. Hosted APIs may accelerate launch, while self-hosted open models can improve economics and data control at sustained volume. For Indian-language products, evaluate the complete pipeline—including retrieval and generation—on code-mixed Hindi, regional-language queries, transliterated text, and domain terminology rather than assuming English benchmark results transfer.

    Make quality and safety observable

    RAG quality has at least two separate dimensions: retrieval quality and answer quality. Build a small, versioned evaluation set from real user questions. Include unanswerable questions, permission-sensitive queries, conflicting documents, long-tail terminology, and multilingual examples.

    Track:

    • Recall and precision of relevant retrieved evidence
    • Context coverage and ranking quality
    • Faithfulness to supplied sources
    • Answer relevance and completeness
    • Citation correctness and refusal quality
    • p50, p95, and p99 latency
    • Token consumption, cache hit rate, and cost per successful answer
    • Error rates by tenant, language, document type, and model route

    Use tracing to inspect the full chain from query classification to final output. Tools such as OpenTelemetry-compatible collectors and RAG evaluation platforms can help, but a simple structured event log is better than unobserved automation. Run regression tests whenever you change chunking, embeddings, prompts, rerankers, or models.

    Treat prompt injection in retrieved documents as an expected threat. Delimit evidence, instruct the model to treat documents as untrusted data, restrict tools by policy, and scan outputs for accidental disclosure. For regulated workflows such as legal support, a dedicated private AI chatbot for lawyers architecture offers useful patterns for access controls, audit trails, and deployment boundaries.

    A practical rollout plan

    Start with one domain and a narrow answer contract: cite sources, refuse unsupported claims, and expose retrieval traces internally. Establish a baseline with a modest corpus and load test before adding query expansion or agents. Then:

    • Move ingestion to a durable, replayable queue.
    • Add hybrid retrieval, metadata filters, and reranking.
    • Introduce caching and model routing after measuring bottlenecks.
    • Test failure recovery, index rebuilds, tenant isolation, and data deletion.
    • Add dashboards and a golden evaluation set to every release.
    • Scale shards, workers, and inference capacity only when measurements justify it.

    Agentic workflows can sit on top of this foundation, but they should not replace deterministic retrieval controls. If your product requires planning or multiple tool calls, review the design principles in building distributed systems with AI agents before allowing agents to mutate data or access unrestricted sources.

    The most scalable RAG systems are not the ones with the most components. They are the ones with clear data ownership, bounded latency, enforceable permissions, measurable retrieval quality, and an operating model that makes failures diagnosable. Build that foundation first, then add models and agent workflows where they create measurable value.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.