0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build a 7-phase llm pipeline

How to Build a 7-Phase LLM Pipeline

  1. aigi

    A production LLM application is more than a model endpoint wrapped in a prompt. It is a sequence of contracts: data must be trustworthy, retrieval must be relevant, model calls must be controlled, and every response must be measurable. This seven-phase design gives Indian founders and engineering teams a practical route from a prototype to a dependable AI product.

    The phases are not strictly linear. In production, evaluation feeds back into chunking, retrieval, prompts, and model selection. Start with a narrow use case, define measurable quality targets, and add complexity only when the baseline shows where it is needed.

    Reference architecture and operating principles

    A useful request path is:

    • Ingestion: collect, clean, classify, and version source material.
    • Retrieval: search authorised content using dense, lexical, or hybrid methods.
    • Context preparation: rerank, filter, compress, and cite evidence.
    • Generation: route the task to an appropriate model and invoke tools safely.
    • Controls: validate inputs, outputs, permissions, and policy constraints.
    • Operations: trace, evaluate, monitor, and improve the system continuously.

    Keep these concerns modular. A change to your embedding model should not require rewriting the business logic or safety layer. For agentic workflows, the same principle applies when deciding whether to use a single orchestrator or a multi-agent design; a practical introduction is available in this guide to building distributed systems with AI agents.

    Phase 1: Ingest, govern, and prepare data

    Begin by mapping every source the application may use: PDFs, websites, ticketing systems, databases, cloud drives, call transcripts, and operational APIs. Record ownership, update frequency, access permissions, language, and retention requirements before indexing anything.

    Build ingestion as a repeatable job rather than a one-time script. It should:

    • Extract text and tables while preserving document structure.
    • OCR scans and validate output quality, including Indian scripts and mixed English-language content.
    • Remove navigation, duplicated headers, boilerplate, and stale versions.
    • Detect and mask or tokenise personal and financial information.
    • Attach metadata such as source, tenant, document type, timestamp, language, and access group.
    • Store document and chunk versions so an answer can be traced to the source available at that time.

    Chunk by meaning where possible. A policy section, product specification, or support procedure should remain coherent. Test chunk sizes against retrieval recall rather than choosing a size by convention. For multilingual products, evaluate language-specific tokenisation and embeddings instead of assuming English performance transfers to Hindi, Tamil, Bengali, or code-mixed queries.

    Phase 2: Embed and index for authorised retrieval

    Convert approved chunks into embeddings and store them with metadata and a stable document identifier. PostgreSQL with pgvector is often a sensible first choice for Indian startups because it combines transactional records, access controls, and vector search without adding another major system. Dedicated databases become useful when scale, filtering, or operational requirements justify them.

    Choose an embedding model using your own query set. Compare recall for exact identifiers, short questions, long descriptions, multilingual terms, and domain vocabulary. Maintain separate indexes or model versions when a migration would otherwise make old and new vectors incomparable.

    Use HNSW or IVF only after measuring the trade-off between recall, memory, and latency. Add metadata filters before retrieval where possible: tenant, region, product, role, and document status are not just optimisation fields—they are security boundaries. Never rely on the language model to enforce permissions after retrieving unauthorised content.

    Phase 3: Build the retrieval layer

    Retrieval-Augmented Generation (RAG) should supply evidence, not merely more text. A typical query flow normalises the question, applies identity and tenant filters, searches multiple indexes, and returns candidate passages with provenance.

    Dense search handles paraphrases, while BM25 or another lexical method is stronger for policy numbers, SKUs, claim IDs, names, and exact legal language. Hybrid retrieval is therefore a strong default for enterprise applications. Query rewriting can help ambiguous questions, but keep the original query for auditing and avoid rewriting away critical identifiers.

    Return structured retrieval results containing the chunk text, source, score, page or section, timestamp, and permissions. If no result meets a calibrated threshold, the system should ask a clarifying question, provide a safe “not found” response, or route to a human—not manufacture an answer.

    Phase 4: Rerank and assemble context

    Initial retrieval is designed for breadth. Reranking is designed for precision. Retrieve perhaps 20–100 candidates, then use a cross-encoder or reranker to judge query-passage relevance. Measure whether reranking improves answer quality enough to justify its latency and cost.

    Context assembly should:

    • Remove duplicates and near-duplicates.
    • Prefer current, authoritative documents when versions conflict.
    • Enforce tenant and role filters again before generation.
    • Keep evidence within a defined token budget.
    • Place passages in a consistent order and attach citations or source labels.

    Prompt compression can reduce cost, but aggressive compression may remove qualifications such as “only if” or “except where.” Evaluate faithfulness after compression. For voice products, context budgets matter even more because long responses increase turn-taking delay; teams building these systems can also review this guide to building a secure voice agent for banking.

    Phase 5: Route prompts, models, and tools

    Separate application policy from prompt text. A robust generation layer defines the task, available evidence, output schema, refusal conditions, and tool permissions. Use structured outputs for downstream systems and validate them with a schema before taking action.

    Model routing can materially improve unit economics. Send classification, extraction, and simple FAQ requests to smaller models; reserve larger models for ambiguous reasoning or high-value cases. Set timeouts, retries, rate limits, and fallbacks for every provider. Log model version, temperature, prompt version, tool calls, and token usage without storing unnecessary customer data.

    Tools need explicit contracts: allowed arguments, authentication scope, idempotency, confirmation requirements, and error handling. A model should not be able to transfer funds, change account details, or send a customer communication without policy checks and, where appropriate, human confirmation. For workflows that genuinely require multiple specialised workers, compare the operational complexity against a deterministic state machine before adopting a multi-agent framework. See also the practical guide to deploying open-source AI agents.

    Phase 6: Add security, safety, and reliability controls

    Guardrails should be layered rather than concentrated in a single classifier. Validate the input for prompt injection, malicious files, data-exfiltration attempts, and prohibited requests. Apply retrieval-time authorisation. Validate output structure, citations, sensitive-data leakage, and policy compliance before returning an answer.

    Treat retrieved documents as untrusted content. Instruct the model that documents provide evidence, not instructions, and isolate tool instructions from retrieved text. Add deterministic checks for high-risk domains such as lending, insurance, healthcare, employment, and legal advice. Define escalation paths for uncertainty, conflicting sources, abusive content, and failed verification.

    For Indian deployments, document data residency, vendor subprocessors, retention, consent, auditability, and incident response requirements. Privacy is an architecture concern: redact before telemetry, restrict tracing access, and avoid sending more customer data to an external model than the task requires. A private legal assistant illustrates these constraints in practice; see how to build a private AI chatbot for lawyers.

    Phase 7: Evaluate, observe, and improve

    Do not wait for production incidents to measure quality. Build a golden dataset of representative questions, difficult edge cases, multilingual inputs, adversarial prompts, and expected evidence. Label retrieval relevance separately from answer correctness. Track:

    • Retrieval recall and citation precision.
    • Faithfulness to supplied evidence.
    • Task success and structured-output validity.
    • Abstention quality and unsafe-completion rate.
    • P50 and P95/P99 latency by phase.
    • Cost per successful task, not just cost per request.
    • Tool failure, fallback, and escalation rates.

    Use traces to connect a user request to query rewriting, retrieved chunks, reranking, model calls, tools, and final output. Sample sensitive traces carefully and redact personal data. Run regression tests whenever prompts, models, chunking, indexes, or policies change. LLM-as-a-judge can accelerate review, but calibrate it against human labels and retain human evaluation for high-risk decisions.

    A practical launch sequence

    Ship the smallest reliable path first: one data source, one retrieval strategy, one model, strict output validation, and an evaluation set. Then add hybrid search, reranking, routing, caching, and agentic behaviour only where measurements show a benefit. Establish budgets for latency, cost, and failure rates before scaling traffic.

    A sensible production checklist includes versioned data, access-aware retrieval, structured outputs, provider fallbacks, trace redaction, golden tests, rollback procedures, and an owner for every alert. This turns the seven phases from a diagram into an operating system for the product.

    Frequently asked questions

    Do I need GPUs to build this pipeline?

    No. API models and managed vector infrastructure are adequate for many first releases. GPUs become relevant when self-hosting improves privacy, latency, availability, or cost at your measured scale.

    Is a vector database mandatory?

    No. A relational database with full-text search may be enough for a small corpus. Add vector search when semantic matching solves a demonstrated retrieval problem.

    How long does implementation take?

    A narrow proof of concept can take weeks. Production hardening—especially permissions, evaluation, multilingual quality, observability, and incident handling—usually takes substantially longer.

    What is the most common failure?

    Teams often optimise model selection before fixing data quality, access control, retrieval evaluation, and unclear success criteria. Better models cannot compensate for missing or unauthorised evidence.

    Apply for AI Grants India

    If you are building LLM infrastructure, an industry-specific copilot, or a multilingual AI product from India, apply to AI Grants India for support, mentorship, and equity-free funding. A clear technical plan should show the target user, measurable evaluation set, deployment constraints, and why your system can deliver durable value.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.