Python is a strong starting point for building LLM agents because it combines mature web tooling, data libraries, model SDKs, and deployment options. But a production agent is not simply a chatbot with a longer prompt. It is a controlled software system that uses a language model to interpret goals, select tools, retrieve information, and produce an outcome.
For Indian founders and engineering teams, the opportunity is practical: automate support operations, reconcile documents, assist field staff, search internal policies, or coordinate workflows across fragmented systems. The winning design is usually not the most autonomous one. It is the agent with the clearest scope, safest permissions, measurable quality, and a sensible operating cost.
What a custom LLM agent actually contains
A useful agent has six layers:
- Model: A hosted or self-hosted LLM that interprets requests and generates structured decisions.
- Instructions: System policies, task rules, examples, and constraints.
- Tools: Typed Python functions for APIs, databases, search, calculations, and business actions.
- State: Conversation history, task status, user identity, and workflow checkpoints.
- Knowledge: Retrieved content from documents, databases, or approved external sources.
- Runtime controls: Timeouts, retries, budgets, approvals, logging, and evaluation hooks.
Separate these layers in code. Keep business rules outside the prompt where possible, validate every model-produced argument, and make side effects explicit. If your use case involves several services or long-running jobs, study the design principles in building distributed systems with AI agents before adding more agents.
Start with a narrow workflow
Define one job in operational terms before choosing a framework. For example: “Classify an inbound support ticket, retrieve the relevant policy, draft a response, and route refunds above ₹5,000 for approval.” This is more useful than “build an autonomous customer-service agent.”
Write down:
- The user and the measurable business outcome.
- Inputs the agent may read.
- Tools it may call and actions it may perform.
- Actions requiring human approval.
- Failure states and escalation routes.
- Accuracy, latency, and cost targets.
Use a conventional application when the workflow is deterministic. An agent is justified when inputs are ambiguous, the correct sequence varies, or natural-language interpretation materially reduces manual work.
Set up a maintainable Python project
Use a virtual environment and pin dependencies. A lightweight starting point might include a model SDK, Pydantic for schemas, an HTTP client, an observability package, and a test runner.
python -m venv .venv
source .venv/bin/activate
pip install pydantic httpx pytest python-dotenvYou can add an orchestration library when the workflow needs graph execution, tool routing, persistence, or human hand-offs. LangGraph is suited to explicit stateful workflows; other frameworks can accelerate prototyping. Do not let a framework hide the control flow you need to audit.
Keep secrets in environment variables or a managed secret store. Never place API keys in prompts, source control, notebooks, or client-side code.
Define tools as typed, limited capabilities
A tool should do one thing, accept validated input, and return a predictable result. Avoid exposing raw database queries, unrestricted shell commands, or Python exec() to a model.
from pydantic import BaseModel, Field
class OrderLookup(BaseModel):
order_id: str = Field(pattern=r"^[A-Z0-9-]{6,30}$")
def get_order_status(request: OrderLookup) -> dict:
# Enforce tenant access and query through a safe repository layer.
return {"order_id": request.order_id, "status": "in_transit"}Add authentication, tenant checks, rate limits, timeouts, and idempotency. A tool that sends a message or changes a record should support a dry-run mode and produce an audit event. For payments, cancellations, medical advice, legal submissions, or other high-impact actions, require confirmation or human approval.
Choose the right control loop
The common patterns are straightforward:
- Router: Classifies a request and sends it to a specialist workflow.
- Single tool-using agent: Selects tools for a bounded task.
- Graph workflow: Moves through explicit states with conditional branches and retries.
- Planner-worker pattern: Creates subtasks, runs workers, then validates their results.
- Multi-agent workflow: Assigns distinct roles, but only when separation improves quality or ownership.
Do not introduce a multi-agent system because it sounds advanced. Multiple model calls increase latency, cost, coordination failures, and debugging effort. A state machine with one model and well-defined tools is often the better first release. For code-focused teams exploring collaborative agents, swarm-based IDE agents offers a useful comparison of role-based coordination.
Add retrieval without turning RAG into a dumping ground
Retrieval-augmented generation (RAG) is useful when the model needs current, private, or domain-specific information. Build the pipeline deliberately:
1. Ingest approved documents and record source ownership.
2. Parse, clean, and split content by meaningful sections.
3. Store embeddings with metadata such as tenant, language, date, and permissions.
4. Filter before semantic search, not after generation.
5. Retrieve a small set of relevant passages.
6. Require citations or source identifiers in the final response.
7. Return “not found” when evidence is insufficient.
For Indian deployments, plan for English plus regional-language content, scanned PDFs, inconsistent document formats, and changing regulations. Test retrieval separately from answer generation. If your corpus is highly specialised, compare RAG with fine-tuning LLMs on custom data; fine-tuning changes behaviour, while RAG supplies changing facts.
Make outputs structured and testable
Use schemas for classifications, tool arguments, routing decisions, and API responses. Pydantic models can reject missing fields, invalid enums, and unsafe values before they reach application code. Prefer a response such as {"decision": "escalate", "reason": "...", "source_ids": [...]} over free-form text when downstream software must act.
Build an evaluation set before launch. Include common requests, ambiguous phrasing, multilingual queries, adversarial inputs, stale documents, missing data, and permission violations. Track:
- Task success and factual correctness.
- Retrieval precision and citation coverage.
- Tool-selection and argument errors.
- Escalation quality.
- Latency, token usage, and cost per completed task.
- Unsafe or unauthorised actions.
Run regression tests whenever prompts, models, tools, or retrieval settings change. Production traces should capture model version, prompt version, tool calls, retrieved source IDs, timings, and redacted inputs.
Security and privacy for Indian products
Treat model output as untrusted input. Defend against prompt injection in user messages and retrieved documents. Enforce permissions in the application and tool layer, not through instructions alone. Redact personal data from logs, isolate tenants, encrypt sensitive records, and define retention periods.
Where data residency, confidentiality, or connectivity matters, evaluate self-hosted models through options such as vLLM or Ollama. Local inference can improve control, but it shifts responsibility for GPU capacity, patching, model updates, monitoring, and quality. For healthcare workflows, compare your architecture with guidance on private AI chatbots for lawyers and privacy-sensitive professional applications, while obtaining sector-specific legal review.
Control latency and cost
Use the smallest model that passes your evaluation set. Reserve stronger models for ambiguous cases, planning, or review. Cache safe, repeatable retrieval and computation results; do not cache responses containing tenant-specific or sensitive information without a clear policy. Limit context size, cap iterations, set per-request budgets, and terminate stalled tools.
A practical production policy is: one model call for routing, one or two calls for the main task, and a deterministic fallback when confidence or evidence is low. Measure cost per successful resolution rather than cost per API call.
A production checklist
Before exposing the agent to customers, confirm that:
- Every tool has typed inputs, authentication, timeouts, and audit logs.
- Side-effecting actions support approval, idempotency, and rollback where possible.
- Retrieval enforces tenant permissions and returns source references.
- The agent stops after a defined step, token, and time budget.
- Human escalation is visible and operationally staffed.
- Evaluation covers English, relevant Indian languages, edge cases, and abuse.
- Prompts, models, tools, and datasets are versioned.
- Monitoring detects failures, cost spikes, drift, and unsafe behaviour.
Voice is another interface rather than a different safety model. If your Python agent will handle calls, review how to build a voice agent and design interruption handling, consent, transcripts, and human transfer from the beginning.
FAQ
Should I use LangChain, LangGraph, or a custom loop? Use a custom loop for a small workflow, a graph for explicit state and branching, and a framework when it removes real integration work. Keep the underlying state and policies understandable.
Can I run a custom agent on an open model? Yes. Hosted APIs, self-hosted Llama-family models, and Indian or regional providers can all work. Select based on quality, language coverage, latency, privacy, and total operating cost—not model size alone.
How do I reduce hallucinations? Restrict the task, retrieve authoritative evidence, require structured outputs and citations, validate tool results, and escalate when evidence is missing. A second model call is not a substitute for good data and permissions.
When should an agent take action automatically? Only when the action is reversible or low-risk, the inputs are validated, and the measured error rate is acceptable. Start in draft or approval mode, then expand autonomy using production evidence.
AI Grants India supports Indian builders working on applied AI products with funding, mentorship, and ecosystem access. Explore AI Grants India if you are turning an agent prototype into a defensible product.