0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to integrate large language models into web applications

How to Integrate Large Language Models into Web Applications

  1. aigi

    Why LLM integration needs an application architecture

    The fastest way to add an LLM feature is to call a model API from a backend route. The reliable way is to treat the model as one component in a larger product system: your application owns authentication, business rules, data access, observability, and user experience; the model handles language tasks within clearly defined limits.

    Common web-app use cases include support assistants, document search, form completion, summarisation, translation, structured extraction, and workflow automation. Start with one measurable job rather than a general-purpose chatbot. Define the expected input, output format, latency target, acceptable error rate, and what happens when the model is uncertain.

    For Indian products, language coverage and data handling matter from the beginning. If your application serves Hindi or other Indic languages, review low-resource Indic natural language processing and the available low-resource language datasets for AI training in India before selecting a model.

    Choose a model and deployment pattern

    Compare models using your own representative test set, not only benchmark scores. Evaluate:

    • Task quality: accuracy, instruction following, extraction reliability, and multilingual performance.
    • Context window: the amount of conversation or source material the model can process.
    • Latency: time to first token and total response time.
    • Cost: input and output tokens, embedding charges, reranking, hosting, and retries.
    • Controls: data retention, regional availability, safety settings, structured outputs, and rate limits.
    • Operational fit: API stability, SDK quality, monitoring, and fallback options.

    A hosted API usually provides the quickest route to production and avoids GPU operations. Self-hosting can improve control, predictable high-volume economics, or data residency, but introduces model serving, capacity planning, patching, and hardware costs. Smaller open models are often suitable for classification, routing, extraction, and simple rewriting. For demanding Indian-language applications, compare general models with open-source small language models for Hindi and models designed for regional-language adaptation.

    Use a model gateway or a thin provider abstraction so application code does not depend on one vendor’s request format. Keep prompts, model names, token limits, and temperature-like settings in configuration rather than scattering them through route handlers.

    Recommended web architecture

    A practical request flow looks like this:

    1. The browser sends an authenticated request to your application backend.
    2. The backend validates the input, applies permissions and quotas, and creates a request ID.
    3. An orchestration layer retrieves approved context, selects tools, and builds the model request.
    4. The provider or self-hosted inference service returns a response.
    5. The backend validates the output, records safe telemetry, and sends the result to the browser.

    Do not expose provider API keys in browser JavaScript. The backend should enforce tenant isolation, redact sensitive fields, and decide which database records or tools the model may access. For long-running jobs, place requests on a queue and notify the client through polling, server-sent events, or WebSockets. For interactive chat, stream tokens only after access checks have completed.

    If your expected traffic is substantial, plan capacity, queues, retries, and graceful degradation alongside the feature. The guide to scaling backend infrastructure for AI applications is useful when moving from a prototype to a multi-tenant service.

    Minimal backend integration

    A provider-neutral Python pattern keeps the model call on the server and validates the result before returning it:

    import os
    from fastapi import FastAPI, HTTPException
    from pydantic import BaseModel
    from openai import OpenAI
    
    app = FastAPI()
    client = OpenAI(api_key=os.environ["LLM_API_KEY"])
    
    class ChatRequest(BaseModel):
        message: str
    
    @app.post("/api/assistant")
    def assistant(request: ChatRequest):
        if not request.message.strip() or len(request.message) > 4000:
            raise HTTPException(status_code=400, detail="Invalid message")
    
        response = client.responses.create(
            model=os.environ.get("LLM_MODEL", "your-model"),
            instructions="Answer using approved product information. Say when unsure.",
            input=request.message,
        )
        return {"answer": response.output_text}

    In production, add authentication, per-user rate limits, request timeouts, retries with backoff, circuit breakers, structured logging, and output moderation. Use environment secrets or a managed secret store; never commit keys or place them in client-side bundles. Provider SDKs change, so pin versions and test upgrades in staging.

    Add retrieval when answers depend on your data

    A model’s training data is not a substitute for your product database, current policies, or private documents. Retrieval-augmented generation (RAG) fetches relevant content at request time and places it in the prompt. A basic pipeline is:

    • Extract and clean source documents.
    • Split them into meaningful chunks with titles and metadata.
    • Generate embeddings and store them in a vector-capable database.
    • Retrieve candidates using semantic, keyword, or hybrid search.
    • Filter by tenant, access rights, language, and freshness.
    • Ask the model to answer only from the supplied context and cite sources where possible.

    RAG reduces stale-answer risk but does not guarantee truth. Test retrieval separately from generation: measure whether the correct passage is found, whether the answer is supported, and whether the system admits missing information. For multilingual products, preserve the original text, language metadata, transliteration where useful, and regional terminology. Open-source multilingual tooling can be especially relevant; see building high-performance AI applications with open-source tools.

    Make outputs dependable

    Free-form text is difficult for downstream code to trust. When the result drives an action, request a strict schema and validate it server-side. Reject malformed JSON, unknown enum values, missing fields, and unsafe URLs. A useful pattern is to have the model propose an action while your application independently checks permissions, pricing, inventory, and policy before execution.

    Use tool calling narrowly. Define each tool’s name, arguments, validation rules, and authorization scope. Never allow a model to execute arbitrary SQL, shell commands, payments, or account changes without deterministic checks and explicit user confirmation.

    Security, privacy, and responsible deployment

    Treat every user message and retrieved document as untrusted input. Prompt injection can attempt to override instructions, expose secrets, or manipulate tool use. Mitigations include:

    • Separating system instructions, user content, and retrieved content.
    • Allowlisting tools and validating every argument.
    • Filtering sensitive data before sending it to an external provider.
    • Enforcing row-level access controls outside the model.
    • Limiting context size and output length.
    • Logging security events without storing unnecessary personal data.
    • Providing a human escalation path for high-impact decisions.

    For India-focused services, map data flows, retention, consent, and deletion requirements before launch. Do not assume that a provider’s “enterprise” label automatically satisfies your obligations. Document which data is sent to which service and whether it is used for training.

    Evaluate before and after launch

    Create a versioned evaluation set from real, consented examples and difficult edge cases. Include regional spelling, code-mixed language, noisy speech transcripts, ambiguous requests, and attempts to extract private information. Track task correctness, groundedness, refusal quality, latency, cost per request, and user feedback.

    Run offline tests for every prompt or model change, then use staged rollout and shadow traffic where practical. Monitor sudden token growth, repeated retries, provider errors, unsafe outputs, and language-specific quality drops. Store prompt and response traces only under a documented retention policy, with redaction applied before analytics.

    Control cost and latency

    Set token budgets and context limits, trim conversation history, cache stable results, and route simple tasks to smaller models. Batch offline jobs rather than running them synchronously in user requests. Stream responses to improve perceived latency, but show a clear loading state and allow cancellation. Keep deterministic fallbacks for outages: search results, templates, human support, or a queued response.

    A production LLM feature is not finished when the first answer appears. It is ready when users can understand its limits, your backend can contain failures, and your team can measure quality over time.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.