0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · build custom chatbot using openai api

Build a Custom Chatbot Using the OpenAI API

  1. aigi

    A useful chatbot is not a chat window wrapped around an API call. It is a product system that combines a reliable interface, backend orchestration, model calls, business tools, data retrieval, safety controls, and observability. For Indian startups, the design must also account for multilingual users, WhatsApp-led workflows, variable network quality, rupee-denominated budgets, and obligations under India’s Digital Personal Data Protection framework.

    This guide explains how to build a custom chatbot using the OpenAI API in 2026, with a focus on architecture that can move from prototype to production.

    Start with a narrow, measurable job

    Define the chatbot’s first job before choosing a model. “Answer anything” is difficult to evaluate and expensive to operate. Strong initial use cases include:

    • Answering questions from approved product or policy documents
    • Qualifying leads and collecting structured information
    • Helping support agents draft replies
    • Checking application status through authenticated tools
    • Guiding users through onboarding or troubleshooting

    Set success metrics early: answer accuracy, grounded-answer rate, escalation rate, first-response latency, resolution rate, cost per conversation, and user satisfaction. A bot that sounds impressive but invents policy details is not production-ready.

    If your product may eventually support calls, compare the experience carefully using this guide to voice agent vs chatbot trade-offs. Text is usually simpler to launch; voice introduces speech recognition, interruption handling, telephony, and stricter latency requirements.

    Use a production-ready architecture

    A typical architecture has six layers:

    • Client: Web, mobile, WhatsApp, or an agent console
    • API backend: Authentication, session management, rate limits, streaming, and business rules
    • Conversation orchestrator: Prompt assembly, routing, tool selection, retries, and fallback logic
    • OpenAI model layer: Response generation, structured outputs, embeddings, or other supported capabilities
    • Data and tools: Your database, search index, CRM, ticketing system, payment status service, or internal APIs
    • Operations: Logging, tracing, moderation, evaluation, alerts, and cost reporting

    Keep the OpenAI key on your server. The browser or mobile application should call your backend, never the model provider directly. Store user and conversation identifiers separately from sensitive message content where possible, and define retention periods before launch.

    For multi-agent workflows, do not add agents simply because the framework supports them. A single orchestrator with well-defined tools is easier to secure and debug. Use a distributed design only when separate agents have genuinely different responsibilities, permissions, or workloads; the principles in building distributed systems with AI agents are useful at that stage.

    Set up the API safely

    Create an API project, generate a key, and load it through environment variables or a managed secrets service. A minimal Python setup is:

    pip install openai fastapi uvicorn python-dotenv
    import os
    from openai import OpenAI
    
    client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])
    
    response = client.responses.create(
        model=os.environ.get("OPENAI_MODEL", "gpt-4o-mini"),
        instructions=(
            "You are a support assistant for an Indian SaaS company. "
            "Answer only from supplied context or approved tools. "
            "If information is missing, say so and offer escalation."
        ),
        input="How can I reset my account password?"
    )
    
    print(response.output_text)

    OpenAI’s API surface changes over time, so pin SDK versions, review current model and endpoint documentation, and keep model names configurable. In production, handle timeouts, rate-limit responses, transient failures, malformed outputs, and provider unavailability. Return a useful fallback instead of exposing stack traces or silently fabricating an answer.

    Design prompts as policies, not personalities

    Your system instructions should define:

    • The assistant’s scope and audience
    • What sources it may rely on
    • When it must ask a clarifying question
    • When it must refuse or escalate
    • Required language, tone, and output format
    • Rules for handling personal, financial, medical, or legal information

    Do not put secrets, private credentials, or irreversible business rules only in a prompt. Enforce permissions in application code. Treat retrieved documents and tool results as untrusted data; they can contain instructions that conflict with your policy.

    For workflows consumed by software, request a strict schema such as intent, answer, citations, and needs_human. Validate the result before acting on it. A model should propose a refund, account change, or loan-status update; your backend must verify identity, permissions, limits, and required confirmations before executing it.

    Add conversation memory deliberately

    The API does not automatically know a user’s previous messages. Your application must manage state. Three practical patterns are:

    1. Recent-message window: Send the latest turns for short conversations.
    2. Rolling summary: Periodically compress older turns into a structured summary.
    3. Persistent user state: Store durable facts—such as plan type or language preference—separately from chat history, with consent and deletion controls.

    Never assume every message should be retained. Redact unnecessary personal data, set expiry rules, and give users a way to delete or correct stored information.

    Implement RAG for private or changing knowledge

    Retrieval-augmented generation is generally preferable to fine-tuning when the chatbot needs current company information. A dependable RAG pipeline has four stages:

    • Ingest: Parse source files, remove duplicates, preserve headings and metadata, and record document versions.
    • Chunk: Split content by meaning rather than an arbitrary character count. Keep policy sections and tables intact where possible.
    • Index: Create embeddings and store vectors alongside source, product, language, access level, and effective-date metadata.
    • Retrieve and answer: Search for relevant passages, apply permission and freshness filters, then instruct the model to answer only from the supplied evidence.

    Return citations or document references when users need to verify an answer. Test retrieval separately from generation: a fluent answer cannot compensate for missing or irrelevant evidence. For Indic content, preserve the original script and language metadata, and evaluate transliterated queries such as Hinglish. Teams working on Hindi, Tamil, Telugu, or other lower-resource languages can also learn from this low-resource Indic NLP builder’s guide.

    Support Indian users without guessing

    Ask users for their preferred language when it matters, or infer it conservatively and allow them to switch. Define whether names, dates, currency, addresses, and phone numbers should follow Indian conventions. Test code-switching, regional spelling, speech-like text, and short mobile messages.

    Do not translate regulated or contractual content casually. For lending, insurance, healthcare, or government workflows, preserve approved wording and route ambiguity to a trained human. A Hindi response that changes the meaning of an eligibility condition is a product failure, not a localisation issue.

    Control cost and latency

    Build a simple cost model before launch. Track input tokens, output tokens, embedding volume, retries, tool calls, storage, and channel fees such as WhatsApp or telephony. Practical controls include:

    • Route classification, intent detection, and simple FAQs to a smaller model.
    • Use a stronger model only for complex synthesis or difficult cases.
    • Limit output length and conversation history.
    • Cache stable answers and repeated retrieval results where safe.
    • Stream responses to improve perceived latency.
    • Summarise long sessions instead of resending every turn.
    • Set per-user, per-tenant, and global budgets.
    • Log token usage by feature, not only by application.

    Never optimise purely for the lowest token bill. Measure cost against successful resolution and escalation. A cheap model that causes repeat questions can be more expensive than a reliable first response.

    Secure, evaluate, and launch in stages

    Protect the system against prompt injection, data leakage, abusive usage, and unsafe tool execution. Apply authentication before exposing account-specific information, use least-privilege service credentials, validate tool arguments, and require confirmation for irreversible actions. Minimise personal data sent to the API and document your retention, deletion, consent, and grievance processes under applicable Indian law and contracts.

    Create an evaluation set from real but sanitised conversations. Include normal requests, ambiguous questions, adversarial prompts, multilingual queries, stale documents, and tool failures. Score factuality, groundedness, refusal quality, language fidelity, latency, and cost. Run it automatically whenever prompts, models, retrieval settings, or source documents change.

    Launch with a narrow cohort, human escalation, and visible feedback controls. Review failed conversations weekly. Expand the bot’s scope only when evidence shows that accuracy and operational handling are strong.

    When to choose a different approach

    RAG is best for changing knowledge; fine-tuning is better suited to consistent style, classification, or structured behaviour when you have high-quality examples. A rules engine may be safer for deterministic eligibility or compliance calculations. If the core interaction is a phone call, evaluate a purpose-built voice architecture rather than forcing a text chatbot into a telephony workflow. Teams building that path can use this voice agent architecture and deployment guide.

    The strongest custom chatbot is not the one with the longest prompt. It is the one that knows its limits, uses verified data, protects user information, and gives your team enough evidence to improve it.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.