0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · developing multi agent systems with python

Developing Multi-Agent Systems with Python: 2026 Guide

  1. aigi

    Multi-agent systems are useful when a workflow needs multiple specialised capabilities, not simply more prompts. One agent may retrieve information, another may validate it, and a third may decide whether an action is safe to execute. Python is a strong foundation for these systems because it combines mature API libraries, asynchronous programming, data validation, observability, and access to both hosted and open-weight models.

    The important shift is architectural: treat agents as software components with explicit responsibilities, tools, permissions, state, and failure handling. Do not start by adding agents because a demo sounds more impressive. Start with a measurable business workflow, then introduce collaboration only where it improves quality, latency, cost, or control.

    What a multi-agent system should contain

    An agent is more than an LLM call. A production agent normally has:

    • A role and objective: what it is allowed to accomplish and what it must refuse.
    • Inputs and outputs: preferably validated with typed schemas rather than free-form text.
    • Tools: APIs, databases, search, calculators, code execution, or internal business systems.
    • State: the current task, previous decisions, tool results, and approvals.
    • Policies: authentication, access controls, rate limits, and escalation rules.
    • Evaluation criteria: how you determine whether the result is correct, complete, safe, and affordable.

    Use multiple agents when responsibilities are genuinely different, when independent review reduces risk, or when parallel work improves turnaround time. A single well-designed agent is usually better for a short, deterministic workflow.

    Choose the right orchestration pattern

    The orchestration pattern determines how agents communicate and who controls the workflow.

    Sequential pipeline

    Agent A produces a structured result for Agent B, which passes a refined result to Agent C. This works well for research, document processing, and content quality checks. The benefit is predictability; the weakness is that one poor intermediate result can affect every downstream step.

    Parallel workers with a synthesiser

    Several agents perform independent tasks at the same time, and a final agent combines their outputs. Python's asyncio.gather() can reduce waiting time when calls are independent. Add deduplication, timeouts, and a clear rule for handling conflicting findings.

    Supervisor and workers

    A supervisor assigns tasks to specialist agents and decides whether the work is complete. This is flexible but expensive and harder to test. Constrain delegation with an allowlist of workers, a maximum task budget, and structured completion states.

    Graph-based workflow

    A graph is appropriate when the process contains branching, retries, human approval, or loops. Tools such as LangGraph model nodes, transitions, and persistent state explicitly rather than hiding control flow inside a long prompt.

    For voice-first customer operations, the same patterns can support call routing, qualification, and escalation. Before building one, understand what a voice agent is and how voice AI works in 2026, especially the distinction between conversation handling and backend task execution.

    Python frameworks and a practical stack

    Framework choice should follow your workflow, not the other way around.

    • LangGraph: strong for explicit state machines, cycles, checkpoints, and human approval.
    • CrewAI: approachable role-and-task abstractions for teams that want fast prototyping.
    • AutoGen: useful for conversational agent collaboration and configurable human involvement.
    • Plain Python with `asyncio`: often the best option for narrow workflows where you need maximum control.

    A sensible baseline includes Python 3.11 or newer, pydantic for schemas, an async HTTP client, a relational database for business state, and a tracing system for model and tool calls. Add a vector store only when semantic retrieval is demonstrably useful; do not use embeddings as a substitute for a well-designed database query.

    Define contracts before prompts

    The most reliable multi-agent systems communicate through typed contracts. For example:

    from pydantic import BaseModel, Field
    
    class ResearchFinding(BaseModel):
        claim: str
        source_url: str
        confidence: float = Field(ge=0, le=1)
        needs_review: bool

    A research agent should return this object, not an unstructured paragraph. The next agent can then reject missing URLs, route low-confidence findings to review, and measure quality consistently. Keep prompts focused on decisions and evidence; keep validation and business rules in Python.

    Tools should also be narrow and permissioned:

    async def get_customer_order(order_id: str, user_id: str) -> dict:
        # Verify the caller can access this order before querying it.
        ...

    Never expose unrestricted database access or arbitrary shell execution to an agent in production. Put sensitive actions behind deterministic functions, approval gates, and audit logs.

    Manage state, memory, and context

    Separate three kinds of information:

    • Run state: current inputs, intermediate outputs, retries, and status.
    • Conversation memory: relevant user preferences or previous interactions.
    • Business records: authoritative data such as orders, accounts, invoices, or claims.

    The database should remain the source of truth for business records. A vector store can help retrieve unstructured policy documents, but it should not decide whether a payment was made. Summarise long histories, cap context size, and pass only the evidence each agent needs. This reduces cost and limits accidental exposure of personal data.

    For Indian deployments, design for multilingual input, code-switching, intermittent connectivity, and regional data requirements from the beginning. If the system handles calls, link the agent to CRM and ticketing systems rather than treating the transcript as the complete record. Restaurant operators, for example, may benefit from multilingual voice agents for restaurants in India, while sales teams need different confirmation and escalation rules.

    Reliability, security, and cost controls

    Agentic failures are often operational rather than purely model-related. Build these controls into the first version:

    • Set maximum steps, wall-clock time, tokens, and tool calls per run.
    • Use timeouts, retries with exponential backoff, and circuit breakers for external services.
    • Make tools idempotent so a retry cannot create duplicate payments or tickets.
    • Require human approval for irreversible, regulated, or high-value actions.
    • Redact secrets and personal information from prompts and traces.
    • Log prompt versions, model versions, tool arguments, outputs, latency, and cost.
    • Maintain an allowlist of domains, APIs, and actions available to each agent.

    Security also includes prompt-injection defence. Treat retrieved documents and user messages as untrusted data. An instruction found in a webpage should never override the system's permissions or approval policy.

    Testing and evaluation

    Unit-test every tool independently. Then test the workflow with a fixed dataset containing normal cases, ambiguous requests, missing data, malicious inputs, and tool failures. Evaluate more than final text:

    • Did the system choose the correct route?
    • Were tool arguments valid and authorised?
    • Did it cite or preserve required evidence?
    • Did it stop when confidence was low?
    • Was latency and cost within budget?

    Use traces to inspect each transition, not only the final response. Run regression evaluations whenever you change prompts, models, tools, or routing logic. For customer-facing automation, compare automation success with escalation quality; a system that handles fewer calls but escalates the right ones may be the better product. Teams assessing business impact can also review the benefits of using a voice agent for Indian businesses.

    A build plan for Indian teams

    Start with one workflow and one success metric. In week one, map the human process, identify systems of record, and define approval boundaries. Next, implement a single agent with typed tools. Add a second agent only when you can show that specialist review or parallel execution improves the metric. Pilot with synthetic and historical data, then run a limited production rollout with manual oversight.

    For founders seeking support, document the problem, baseline process, data access, evaluation set, expected unit economics, and safeguards. AI Grants India supports builders working on practical systems across logistics, finance, healthcare, agriculture, education, and public-service use cases. Learn more about applying for AI grants in India and present the workflow as an accountable product, not just a collection of agents.

    FAQ

    Do I need a GPU?

    No. Hosted models are sufficient for most prototypes and many production workloads. A GPU becomes relevant when latency, privacy, volume, or model customisation justifies operating an open-weight model.

    Can agents use different models?

    Yes. Use a stronger model for planning or high-risk review and smaller models for classification, extraction, or routing. Measure quality and total workflow cost rather than comparing model prices in isolation.

    Is a framework mandatory?

    No. Plain Python is often easier to understand for a small, deterministic workflow. Adopt a framework when you need durable execution, graph state, visual tracing, checkpointing, or standard integrations.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.