0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to deploy generative ai agents

How to Deploy Generative AI Agents in Production

  1. aigi

    Generative AI agents are software systems that use models to decide what to do next, call tools, update state, and continue until a task is complete. That loop makes them useful for customer support, research, operations, developer tooling, and workflow automation—but it also makes deployment more demanding than exposing a single model API.

    A production agent needs bounded autonomy, durable state, predictable failure handling, secure tool access, and enough telemetry to explain every important decision. For Indian startups, the design must also account for data residency, intermittent integrations, multilingual users, variable traffic, and strict unit economics.

    This guide explains how to deploy generative AI agents from a working prototype to a reliable production service.

    Start with a bounded job, not a general-purpose agent

    Before choosing a framework or cloud service, define the task the agent is allowed to complete. A narrow agent with explicit tools is easier to evaluate and safer to operate than an unrestricted assistant that can browse, execute code, and modify records.

    Document four things:

    • Goal: the business outcome, such as resolving a support ticket or reconciling an invoice.
    • Inputs and outputs: the data the agent receives and the exact result your application expects.
    • Permitted tools: APIs, databases, search systems, browsers, or internal services the agent may call.
    • Stop conditions: success, refusal, human approval, timeout, iteration limit, or failed dependency.

    If you are still designing the workflow, compare the deployment concerns here with the architecture patterns in Build Generative AI Agents. Voice products have additional streaming and interruption requirements; the voice agent architecture guide covers those separately.

    Use a durable agent architecture

    A production deployment commonly has six layers:

    1. Client layer: web, mobile, WhatsApp, voice, or an internal application.
    2. API layer: authentication, request validation, rate limits, and routing.
    3. Orchestrator: the agent graph or state machine that selects steps and tools.
    4. Model gateway: provider routing, retries, fallbacks, token budgets, and caching.
    5. Tool services: isolated functions for search, CRM actions, payments, databases, and code execution.
    6. State and telemetry: durable checkpoints, business records, traces, metrics, and audit logs.

    Frameworks such as LangGraph are useful when an agent must pause, resume, branch, retry, or request human approval. A graph-based design is generally easier to reason about than an opaque loop because each node can define its inputs, outputs, permissions, and failure policy. CrewAI and similar frameworks can work for role-based prototypes, but production teams should still make state transitions and tool contracts explicit.

    Treat the model as a probabilistic component inside a deterministic system. The application—not the model—should enforce permissions, schemas, spending limits, and irreversible actions.

    Choose the right runtime

    Short synchronous requests

    A containerised HTTP service is suitable for tasks that complete within a predictable time. Use FastAPI, Node.js, or an equivalent runtime, and stream partial status where helpful. Do not keep critical state only in process memory: deployments, crashes, and autoscaling will lose it.

    Long-running or interruptible jobs

    For research, document processing, browser tasks, and multi-step operations, place work on a queue. A request should create a job and return an ID; workers then execute the agent and write progress to durable storage. Redis-backed queues, Celery, BullMQ, managed cloud queues, or workflow engines can all work, provided they support retries and idempotency.

    Kubernetes and managed containers

    Docker provides reproducible dependencies, while Kubernetes or a managed container platform supports independent scaling of API servers, workers, tool services, and model gateways. Start with a managed container service when traffic is modest. Adopt Kubernetes when you genuinely need workload isolation, custom scheduling, GPU management, or multi-service operational control.

    Serverless functions are useful for stateless tool calls and lightweight orchestration, but hard timeouts, cold starts, and limited local storage make them a poor default for long-running agents.

    Make state durable and privacy-aware

    Separate agent state from business data. A useful state model includes:

    • Thread state: the current messages, tool results, and pending step.
    • Checkpoint state: the last safe point from which execution can resume.
    • Retrieved context: documents or records selected for this task.
    • Business state: orders, tickets, approvals, and other authoritative records.
    • Audit state: who initiated an action, which tool was called, and what was returned.

    Redis can support short-lived coordination, but do not treat it as the only source of truth. Use PostgreSQL or another durable database for checkpoints and transactional records. Add a vector store only when semantic retrieval solves a real problem; embedding every conversation creates cost, retention, and deletion obligations.

    For Indian deployments, classify personal and confidential data before sending it to a model provider. Apply retention limits, tenant isolation, encryption, access controls, and deletion workflows. If your product serves healthcare, study the operational implications in HIPAA-compliant voice agents for hospitals, while remembering that your legal and compliance obligations depend on the actual jurisdiction and data flows.

    Secure tools and enforce human control

    Prompt instructions are not a security boundary. Every tool should have a server-side permission check, typed input schema, timeout, and narrowly scoped credentials. Prefer read-only tools by default and require explicit approval for payments, deletion, external messages, account changes, or production deployments.

    Implement these controls:

    • Tenant-aware authorisation: verify the user and organisation on every tool request.
    • Input and output validation: reject malformed arguments and constrain returned data.
    • Network isolation: keep databases and internal services behind private networking.
    • Sandboxing: run code execution and browser automation in disposable, restricted environments.
    • Prompt-injection resistance: treat retrieved documents and web pages as untrusted content.
    • Budget limits: cap tokens, tool calls, wall-clock time, and maximum iterations.
    • Kill switches: disable a tool, workflow, tenant, or model route without redeploying.

    For distributed workflows, the principles in Building Distributed Systems with AI Agents are especially relevant: retries, duplicate messages, partial failure, and eventual consistency must be designed rather than discovered in production.

    Design for failure and observability

    An agent should fail clearly, not continue improvising after a dependency or permission error. Use exponential backoff for transient model and API failures, circuit breakers for unhealthy services, and idempotency keys for actions that may be retried. Store a correlation ID across the client request, agent run, model calls, tool calls, and database writes.

    Track at least:

    • completion and escalation rates;
    • latency by model, node, and tool;
    • input and output tokens and cost per successful task;
    • retry, timeout, refusal, and validation-failure rates;
    • retrieval quality and citation or grounding failures;
    • unsafe-action attempts and policy violations.

    Trace complete runs with tools such as LangSmith, OpenTelemetry-compatible platforms, or Arize Phoenix. Redact secrets and unnecessary personal data before logging. Build a replayable evaluation set from real, consented, anonymised tasks, then run it against every prompt, model, tool, or graph change.

    Control cost and capacity

    Agent cost is driven by the number of model calls, context size, tool traffic, and failed retries—not just the headline price of a model. Set a cost budget per task and expose it in operations dashboards.

    Practical controls include:

    • route classification and extraction to smaller models;
    • reserve stronger models for ambiguous or high-value decisions;
    • summarise or retrieve history instead of sending full transcripts;
    • cache stable retrieval and deterministic tool results;
    • batch offline jobs and schedule them during lower-cost periods;
    • use open models through vLLM when volume and hardware utilisation justify the operational burden;
    • measure cost per completed business outcome, not cost per API call.

    For teams considering self-hosted inference, How to Deploy Llama 3 Agents offers a useful model-serving comparison. Keep provider abstraction at the model gateway, but test each model independently: tool-calling quality, multilingual performance, latency, refusal behaviour, and structured-output reliability vary substantially.

    A production deployment checklist

    Before exposing an agent to customers, confirm that:

    • the workflow has a defined scope and explicit stop conditions;
    • every tool has authentication, authorisation, validation, timeout, and audit logging;
    • state survives process restarts and can resume safely;
    • retries are idempotent and duplicate actions are prevented;
    • prompts, models, tools, and datasets are versioned;
    • evaluation tests cover normal, adversarial, multilingual, and failure cases;
    • dashboards show latency, quality, safety, and cost;
    • human escalation is available for uncertain or high-impact cases;
    • secrets are stored in a managed vault and never committed to source control;
    • rollback and emergency-disable procedures have been tested.

    Deploy in stages: internal users first, then a small production cohort, followed by gradual expansion. Review traces and business outcomes at each stage. The objective is not maximum autonomy; it is reliable completion of a valuable task within controlled operational boundaries.

    FAQ

    Can I deploy a generative AI agent on a single server?
    Yes, for an early pilot with low traffic and non-sensitive workloads. Use a durable database, process supervisor, queue for long jobs, strict limits, and backups. Move to managed containers or Kubernetes when isolation and scaling needs justify the complexity.

    How do I prevent infinite loops?
    Enforce maximum steps, wall-clock time, model calls, and spend per run. Add state-based loop detection and require a human decision when the agent repeats a failed action.

    Should every agent use a vector database?
    No. Use one when semantic retrieval improves measured task performance. For structured records, normal database queries are often more accurate, cheaper, and easier to secure.

    What should Indian startups prioritise first?
    Start with a narrow workflow, strong authorisation, durable state, evaluation data, and cost telemetry. Scaling infrastructure before proving task quality usually increases expense without improving the product.

    AI Grants India supports Indian founders building practical AI infrastructure and applications. Explore AI Grants India for grant and ecosystem opportunities as you take an agent from prototype to production.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.