0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build ai chatbots with flask and react

How to Build AI Chatbots with Flask and React

  1. aigi

    What you are building

    A Flask–React chatbot has three layers:

    • React renders the conversation, handles input, and consumes streamed responses.
    • Flask authenticates requests, validates payloads, manages conversation state, and calls the model provider.
    • The AI layer generates answers, optionally retrieves documents, and applies safety and business rules.

    This separation keeps the browser free of secret API keys and makes it easier to replace a model, add retrieval, or expose the same chatbot through a mobile app. For Indic-language products, plan language detection, transliteration, and evaluation early; low-resource Indic NLP often requires different data and testing choices than English-first systems.

    Choose the chatbot architecture

    Start with the smallest architecture that supports your product requirement. A basic assistant sends the user message and recent conversation to an LLM. A knowledge assistant adds retrieval-augmented generation (RAG), where Flask searches approved documents and places relevant excerpts in the model prompt. A workflow assistant adds tools such as ticket creation, database lookups, or payment-status checks.

    Keep these concerns separate in code:

    • routes/: HTTP endpoints and authentication
    • services/llm.py: model-provider calls and streaming
    • services/retrieval.py: embeddings and document search
    • prompts/: versioned system instructions
    • schemas/: request and response validation
    • tests/: API, prompt, retrieval, and safety tests

    If the assistant must take multi-step actions, study patterns for generative AI agents, but do not turn a simple FAQ bot into an agent unnecessarily.

    Build the Flask backend

    Create an isolated environment and install the API dependencies:

    python -m venv .venv
    source .venv/bin/activate       # Windows: .venv\\Scripts\\activate
    pip install flask flask-cors python-dotenv pydantic gunicorn

    A minimal endpoint should validate input, enforce a size limit, and return a stable response shape:

    from flask import Flask, request, jsonify
    from pydantic import BaseModel, ValidationError
    
    app = Flask(__name__)
    
    class ChatRequest(BaseModel):
        message: str
        conversation_id: str | None = None
    
    @app.post("/api/chat")
    def chat():
        try:
            body = ChatRequest.model_validate(request.get_json(silent=True) or {})
        except ValidationError as exc:
            return jsonify({"error": "Invalid request", "details": exc.errors()}), 400
    
        if not body.message.strip() or len(body.message) > 4000:
            return jsonify({"error": "Message must contain 1–4000 characters"}), 400
    
        # Replace this with an authenticated model-service call.
        return jsonify({"message": "Connect this endpoint to your model service."})

    In production, load credentials from environment variables or a secret manager, never from React source code. Add authentication before exposing the endpoint, apply per-user rate limits, and log request IDs rather than raw sensitive conversations. Configure CORS for the exact frontend origin, not *, when credentials are involved.

    Connect a model safely

    Your Flask service should construct a message list containing a system instruction, a controlled amount of conversation history, and the current user message. Set a maximum output token limit and a timeout. Catch provider errors and return a generic, retryable error to the client; do not expose stack traces or provider keys.

    For a production chatbot, also implement:

    • Conversation limits: trim or summarise old turns to control cost and context size.
    • Prompt-injection resistance: treat retrieved documents and user text as untrusted data.
    • Grounding rules: instruct the model to say when an answer is unsupported.
    • PII controls: redact or avoid storing Aadhaar numbers, phone numbers, financial data, and legal or health records unless the product has a clear compliance basis.
    • Observability: capture latency, token usage, model version, retrieval hits, refusal rate, and user feedback.

    If you need private deployments for sensitive organisations, the design considerations in building a private AI chatbot for lawyers are useful beyond the legal sector.

    Add streaming for a responsive interface

    Non-streaming responses are easiest to ship, but users perceive a long wait when the model takes several seconds. Use Server-Sent Events (SSE) or WebSockets to send partial output. SSE is usually sufficient for one-way token delivery:

    from flask import Response, stream_with_context
    
    @app.post("/api/chat/stream")
    def stream_chat():
        data = request.get_json() or {}
        message = data.get("message", "").strip()
    
        @stream_with_context
        def events():
            for chunk in generate_answer(message):
                yield f"data: {chunk}\\n\\n"
            yield "data: [DONE]\\n\\n"
    
        return Response(events(), mimetype="text/event-stream")

    Use a production WSGI setup such as Gunicorn and confirm that your reverse proxy does not buffer SSE output. Set connection and model timeouts, and provide a cancel action so users can stop expensive generations.

    Build the React chat interface

    For a new application, use a maintained React setup such as Vite rather than starting a new project with the older Create React App workflow:

    npm create vite@latest chatbot-ui -- --template react
    cd chatbot-ui
    npm install
    npm run dev

    Keep messages as structured data rather than plain strings:

    const [messages, setMessages] = useState([]);
    
    async function sendMessage(text) {
      setMessages(old => [...old, { role: "user", content: text }]);
      const response = await fetch("/api/chat", {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify({ message: text, conversation_id: conversationId })
      });
      const data = await response.json();
      setMessages(old => [...old, { role: "assistant", content: data.message }]);
    }

    Add disabled and loading states, retry handling, keyboard submission, accessible labels, markdown sanitisation, and a visible error state. Never render model output as unsanitised HTML. For streaming, use fetch() with a readable stream and append decoded chunks to the current assistant message.

    Add retrieval and conversation memory

    For company or government information, do not rely on model training data. Ingest approved files, split them into meaningful sections, create embeddings, and store vectors with document IDs, access controls, and source metadata. At query time, retrieve a small set of relevant passages, apply permission filters, and show citations in the React response.

    Store conversation records with a tenant or user ID, retention policy, model metadata, and consent status. PostgreSQL is a practical starting point; Redis can handle short-lived sessions and rate limits. Do not use browser local storage for confidential chats unless you have assessed the risk.

    Test before deployment

    Create a test set from real user intents, including Hindi, English, code-mixed queries, misspellings, abusive inputs, prompt injection, and questions outside the product scope. Measure:

    • factual accuracy and citation correctness;
    • refusal and escalation behaviour;
    • first-token and full-response latency;
    • cost per conversation;
    • retrieval precision and failure cases;
    • accessibility and mobile usability.

    Use deterministic mock model responses for API tests, then run a smaller set of live evaluations on every prompt or model change. Human review remains essential for high-impact domains.

    Deploy the application

    Build the React app into static assets and serve it through a CDN or the same domain as Flask. Run Flask behind Nginx or a managed load balancer with HTTPS, health checks, structured logs, and environment-specific configuration. Containerise the API if your team needs repeatable deployments, and scale workers according to model-call latency rather than CPU alone.

    For voice or multimodal extensions, the same API boundary can support additional clients; review this voice-agent architecture guide before adding real-time audio complexity.

    Start with a narrow use case, instrument every model call, and expand only when evaluation shows that the chatbot is reliable, secure, and affordable for its intended Indian users.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.