0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to deploy open source ai agents

How to Deploy Open-Source AI Agents in Production

  1. aigi

    Open-source AI agents are easier to prototype than to operate. A notebook demo can call tools and return useful answers; a production system must also manage concurrency, failures, permissions, model updates, private data, and operating cost. The right deployment plan treats an agent as a distributed application—not simply a model behind an API.

    This guide explains how to deploy open source AI agents for internal workflows, customer-facing products, and domain-specific automation. It focuses on practical decisions for Indian engineering teams: data residency, GPU economics, multilingual workloads, and the operational discipline needed beyond a proof of concept.

    Define the production boundary first

    An agent normally combines five layers:

    • Model: Llama, Qwen, Mistral, Gemma, or another open-weight model.
    • Agent runtime: A graph or orchestration layer that manages state, tool calls, retries, and approvals.
    • Inference server: vLLM, SGLang, Hugging Face TGI, Ollama, or llama.cpp.
    • Application services: APIs, authentication, databases, queues, and business integrations.
    • Execution environment: Containers, virtual machines, Kubernetes, or managed GPU infrastructure.

    Decide what the agent is allowed to do before selecting technology. A read-only research assistant has very different requirements from an agent that can issue refunds, modify records, or execute code. Define approved tools, user permissions, maximum steps, timeout limits, and human approval points in writing.

    For teams building complex workflows, patterns from building distributed systems with AI agents are useful: separate deterministic services from model-driven decisions, make every operation observable, and design for partial failure.

    Choose the model and serving strategy

    Start with the smallest model that meets your quality and latency target. Larger models may improve reasoning, but they increase VRAM requirements, queueing delays, and cost. Evaluate representative Indian-language and domain-specific prompts rather than relying only on public benchmarks.

    A practical selection process is:

    1. Create a test set covering normal requests, ambiguous instructions, tool errors, prompt injection, and sensitive data.
    2. Compare answer quality, structured-output reliability, tool-call accuracy, and tokens per second.
    3. Measure performance at expected concurrency, not just with one request.
    4. Test quantized and full-precision variants before purchasing hardware.

    vLLM is a strong default for GPU-hosted, OpenAI-compatible serving and high-throughput workloads. SGLang can be effective for structured generation and complex serving patterns. Ollama is convenient for development and small private deployments, while llama.cpp is valuable for CPU, edge, and low-cost environments. For a model-specific implementation, see how to deploy Llama 3 agents.

    Do not confuse an OpenAI-compatible endpoint with OpenAI-equivalent behaviour. Tool-calling formats, JSON enforcement, context limits, streaming, and safety characteristics vary by model and server. Pin model versions and test the exact serving configuration you deploy.

    Size GPUs using workload data

    VRAM planning should account for more than model weights. You also need memory for the KV cache, runtime overhead, batching, and the agent’s context window. A quantized model may fit on a single GPU but still perform poorly when several long conversations run concurrently.

    Use these as starting points, not guarantees:

    • 7B–8B models: often suitable for development and modest workloads on 8–16 GB GPUs, depending on precision and context length.
    • 13B–14B models: commonly require 16–24 GB or more for comfortable serving.
    • 30B– ned models: usually need multi-GPU or carefully chosen quantization.
    • 70B-class models: generally require substantial multi-GPU capacity, even when quantized.

    Benchmark cold-start time, steady-state throughput, p95 latency, and maximum concurrent sessions. In India, compare domestic GPU providers, reserved capacity, and on-premise options with global clouds. The cheapest hourly GPU is not always the cheapest deployment once egress, idle capacity, support, and engineering time are included.

    Containerise the agent and inference server

    Keep the agent application and model server in separate containers. This lets you update orchestration code without rebuilding model images and scale the two layers independently.

    A minimal application image might look like this:

    FROM python:3.12-slim
    WORKDIR /app
    COPY requirements.txt .
    RUN pip install --no-cache-dir -r requirements.txt
    COPY src/ ./src/
    CMD ["python", "-m", "src.main"]

    Use multi-stage builds where appropriate, run as a non-root user, pin dependencies, and scan images before release. Store secrets in a vault or cloud secret manager rather than in Docker images or environment files committed to source control.

    For a first production release, Docker Compose can be enough for a single host. Move to Kubernetes when you genuinely need rolling deployments, multiple GPU nodes, workload isolation, or automated scaling. Kubernetes adds operational cost; it is not a substitute for a clear service boundary.

    Design state, memory, and tools deliberately

    Agents need state, but indiscriminately storing every conversation creates privacy, cost, and retrieval problems.

    • Session state: Keep active graph state in Redis or PostgreSQL with expiry and tenant isolation.
    • Long-term knowledge: Use PostgreSQL with pgvector or a dedicated vector database when semantic retrieval is justified.
    • Document retrieval: Track source, permissions, version, and timestamp for every chunk.
    • Tool execution: Place code execution and file handling in isolated sandboxes with CPU, memory, network, and time limits.
    • External APIs: Use allowlists, scoped credentials, idempotency keys, and explicit confirmation for irreversible actions.

    For multilingual products, evaluate tokenisation and retrieval in the languages your users actually speak. Indic language support can be uneven across embedding models, OCR systems, and speech pipelines; the low-resource Indic NLP guide offers relevant design considerations. If the product is voice-first, deployment also needs streaming audio, interruption handling, and telephony reliability, as outlined in how to build a voice agent.

    Optimise latency and cost

    Quantisation is often the first optimisation, but it should follow evaluation. GGUF is convenient for llama.cpp-based deployments; AWQ and GPTQ are widely used for GPU inference; FP8 can deliver strong throughput on compatible hardware. Compare quality on tool selection and structured outputs—not only on conversational answers.

    Additional levers include:

    • Limit context growth with summarisation and state compaction.
    • Cache stable retrieval results and deterministic tool responses.
    • Stream tokens while the agent is working, but never stream unvalidated actions as completed.
    • Route simple requests to a smaller model and escalate difficult cases.
    • Set per-request budgets for tokens, tool calls, wall-clock time, and spend.
    • Queue background jobs instead of holding interactive GPU capacity for long tasks.

    Track cost per successful task, not just cost per token. A cheap model that requires repeated retries or human correction may be more expensive overall.

    Secure the deployment

    Agent security is broader than prompt filtering. Treat model output as untrusted input. Validate tool arguments against schemas, enforce authorisation in the tool service, and prevent the model from choosing credentials or network destinations.

    Minimum controls should include:

    • Authentication, tenant isolation, rate limits, and request quotas.
    • Network egress restrictions for sandboxes and model containers.
    • PII redaction before logs, traces, and evaluation datasets.
    • Prompt-injection tests for retrieved documents, web pages, emails, and tool responses.
    • Human approval for financial, legal, medical, or irreversible operations.
    • Audit logs recording user, model version, tools invoked, arguments, outcome, and policy decision.

    If you are building for healthcare, deployment must cover access controls, retention, consent, and auditability; compare the operational requirements in HIPAA-compliant voice agents for hospitals, while adapting them to Indian law and contracts. The DPDP Act and sector-specific obligations should be reviewed with qualified legal counsel rather than treated as a simple hosting-location checklist.

    Monitor the agent as a system

    Instrument every run with a trace ID and capture model latency, queue time, input and output tokens, tool duration, retries, error type, retrieval quality, and final task success. OpenTelemetry-compatible tracing and self-hosted tools such as Langfuse can help teams retain sensitive traces within their controlled environment.

    Create alerts for rising p95 latency, GPU memory pressure, tool failure rates, loop counts, empty retrieval results, and unexpected spend. Maintain a regression suite that runs before changing the model, prompt, retrieval index, or inference engine. A deployment is not production-ready until you can roll back the model and application independently.

    A practical rollout plan

    1. Prototype: Run the smallest viable model locally and define tool contracts.
    2. Evaluate: Build a domain test set with adversarial and multilingual cases.
    3. Containerise: Separate the agent API, inference server, data stores, and workers.
    4. Harden: Add authentication, schemas, sandboxing, budgets, audit logs, and secrets management.
    5. Pilot: Serve a limited user group with capped concurrency and human review.
    6. Scale: Add batching, queues, replicas, autoscaling, and capacity alerts only after measuring demand.
    7. Operate: Review quality, incidents, cost per successful task, and model updates continuously.

    Open-source deployment delivers control, but it also transfers responsibility to your team. The strongest Indian implementations begin with narrow, measurable workflows, keep sensitive actions behind policy gates, and expand only after reliability and unit economics are proven.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.