What makes an AI agent different from a chatbot?
A chatbot mainly generates replies. An AI agent combines a model with instructions, tools, context, and a control loop so it can decide what to do next. It may retrieve a document, call an API, update a database, ask for approval, or hand a task to a human.
For Indian builders, useful applications include multilingual customer support, internal knowledge assistants, field-service workflows, collections, healthcare administration, and small-business automation. The best starting point is not a fully autonomous system. It is a narrow workflow with clear permissions and a measurable outcome. If your product serves users across languages, this guide to building AI apps for the next billion users in India is a useful companion.
Start with a bounded problem
Write a one-sentence specification before choosing a model or framework:
> Given a customer’s request and account context, classify the intent, retrieve the relevant policy, and either answer or create a support ticket.
Then define:
- Inputs: text, voice transcript, uploaded file, API event, or structured form.
- Allowed actions: search, calculate, send, create, update, or escalate.
- Success metric: resolution rate, factual accuracy, task completion, response time, or cost per task.
- Failure behaviour: ask a clarifying question, refuse, retry, or route to a person.
- Risk boundary: actions requiring explicit user confirmation or staff approval.
Avoid vague goals such as “make a smart assistant”. A narrow agent is easier to test, cheaper to run, and safer to deploy.
A practical Python architecture
A maintainable agent usually has six layers:
1. Interface: API, web app, WhatsApp integration, voice pipeline, or command line.
2. Orchestrator: decides whether to answer, retrieve information, call a tool, or escalate.
3. Model adapter: provides one stable interface across model providers and local models.
4. Tools: typed Python functions for search, business APIs, calculators, and databases.
5. State and memory: stores the current task, conversation summary, user preferences, and durable records.
6. Observability: captures traces, tool calls, latency, token usage, errors, and outcomes.
Keep business logic outside prompts. A prompt can guide decisions, but it should not be the only place where pricing rules, permissions, or compliance controls live.
A minimal project structure might look like this:
agent_app/
api.py
agent.py
models.py
tools/
search.py
tickets.py
policies.py
prompts/
tests/
evals/
observability.pyUse environment variables or a secrets manager for credentials. Never place API keys in source code, notebooks, logs, or client-side JavaScript.
Choose the simplest useful control loop
Most first agents need a bounded loop rather than unrestricted autonomy:
for step in range(MAX_STEPS):
decision = model.decide(state, available_tools)
if decision.type == "final":
return decision.answer
if decision.type == "tool_call":
result = execute_validated_tool(decision.name, decision.arguments)
state = state.with_result(result)
continue
return escalate_to_human(state)Add hard limits for steps, time, tokens, tool calls, and spending. Validate every tool argument with typed schemas such as Pydantic. Treat model output as untrusted input; the model may propose an action, but application code must authorise and execute it.
For complex workflows, represent the process as a state machine or directed graph. This makes retries, approvals, and recovery explicit. Multi-agent designs can help when specialist roles are genuinely independent, but they also multiply latency, cost, and failure modes. Use them only after a single-agent workflow has reached its limits. Distributed designs require additional coordination; see building distributed systems with AI agents before introducing multiple workers.
Tools, retrieval, and memory
Tools
A good tool has one purpose, a strict schema, a clear description, and predictable errors. Prefer get_order_status(order_id) over a generic tool such as run_database_query. Apply least-privilege permissions: a support agent may read an order and create a ticket, but not issue a refund without approval.
Retrieval
Use retrieval-augmented generation when the agent must answer from changing or private information. A typical pipeline is:
- Ingest approved documents and attach source metadata.
- Split content by meaningful sections, not arbitrary character counts alone.
- Create embeddings and store them in a vector or hybrid search index.
- Retrieve a small set of relevant passages.
- Instruct the model to answer only from supplied evidence.
- Return citations or document references where users need verification.
Evaluate retrieval separately from answer quality. A fluent answer based on the wrong document is still a failure. For short customer messages, a dedicated intent extraction workflow can be more reliable than sending every request through a large agent.
Memory
Separate working memory from long-term records. Working memory is the current task context; long-term memory should contain only information with a defined purpose, retention period, and deletion process. Do not automatically save sensitive conversation content as “memory”. Obtain consent where required and redact personal data in logs.
India-specific product considerations
Support English, Hindi, and the languages your users actually need, but test code-switching, names, addresses, numbers, and regional accents independently. Translation can change intent, especially for payments, healthcare, and legal requests. Voice systems need careful handling of latency, interruptions, background noise, and fallback to keypad or human support. Compare the workflow against conventional telephony using a voice agent vs IVR guide.
Design for intermittent connectivity and low-end devices. Keep payloads small, provide concise responses, and avoid assuming that every user can read long text. For healthcare, finance, education, and government-facing products, map data flows, consent, retention, access controls, and incident response before launch. Human review is essential for high-impact decisions.
Evaluation before production
Create a test set from real or realistically anonymised tasks. Include normal requests, ambiguous wording, adversarial prompts, missing data, tool failures, and attempts to exceed permissions. Track:
- Task completion and correct escalation.
- Factual accuracy and citation correctness.
- Tool-selection and argument-validation errors.
- Prompt-injection and data-exfiltration resistance.
- Latency, token consumption, and cost per successful task.
- Performance by language, device, customer segment, and network condition.
Use deterministic unit tests for tools and policies, scenario-based tests for workflows, and human review for quality and safety. Run evaluations in CI when prompts, models, retrieval indexes, or tool definitions change. Keep production traces privacy-safe and sample them for review.
Deployment and operations
Expose the agent through a small FastAPI service or an equivalent application layer. Use asynchronous jobs for long-running tasks, queues for retries, and idempotency keys for actions such as payments, messages, and ticket creation. Add timeouts, exponential backoff, circuit breakers, rate limits, and a kill switch.
Start with a low-risk pilot and a restricted user group. Monitor real outcomes rather than only model scores. Review failed traces weekly, update tools and policies, and maintain a rollback path for model or prompt changes. Cost control usually comes from shorter context, caching, smaller models for classification, and routing difficult cases to stronger models—not from removing safeguards.
A sensible build sequence
1. Implement the deterministic workflow without an agent.
2. Add one model-powered decision, such as intent classification.
3. Add one read-only tool and strict validation.
4. Add retrieval with citations.
5. Add approval gates for write actions.
6. Build an evaluation set and observability before broad rollout.
7. Introduce memory, voice, or multiple agents only when evidence justifies it.
This sequence keeps the system understandable and gives you a baseline for measuring whether each new capability improves the product.
FAQ
Do I need to train a model from scratch?
No. Most teams should begin with a capable hosted or open-weight model, strong tool schemas, retrieval, and evaluation. Fine-tuning is useful when you have a stable task, sufficient examples, and a demonstrated gap that prompting cannot close.
Which Python libraries should I learn?
Start with Python’s standard library, FastAPI, Pydantic, an HTTP client, a database driver, and your chosen model SDK. Add an orchestration framework only when it reduces real complexity. Frameworks should not replace understanding of state, permissions, retries, and testing.
How autonomous should an agent be?
As autonomous as the risk allows. Let agents draft, search, classify, and prepare actions freely; require confirmation or human approval for irreversible, financial, privacy-sensitive, or high-impact actions.
What is the most common implementation mistake?
Building a broad agent before defining success, permissions, and failure handling. A small, observable workflow with one reliable tool is a better foundation than a general assistant that can call everything.