Start with a workflow, not an autonomous persona
Building custom AI agents from scratch in India is most effective when you begin with a measurable business workflow rather than a broad “general-purpose agent”. Choose a task with clear inputs, bounded tools, and an observable outcome: reconciling invoices, triaging support tickets, preparing GST research, qualifying leads, or following up with patients.
Write down the agent’s job, authority, constraints, escalation path, and success metric before selecting a model. A useful first version should answer four questions:
- What information may the agent read?
- Which actions may it take without approval?
- Which actions require a human confirmation?
- What evidence must it retain for audit and debugging?
This approach also makes distributed execution easier later; teams designing larger systems can study patterns in building distributed systems with AI agents before introducing queues, workers, and multiple agents.
A production architecture for Indian deployments
An agent is not just an LLM with a prompt. A dependable system normally contains six layers:
1. Interface: Web, mobile, WhatsApp, email, voice, or an internal dashboard.
2. Orchestrator: A state machine that manages steps, retries, approvals, and termination.
3. Model gateway: A consistent interface for hosted and self-hosted models, with routing, budgets, and fallbacks.
4. Knowledge layer: Retrieval over approved documents, databases, and APIs.
5. Tool layer: Typed functions for search, SQL, payments, CRM updates, or case management.
6. Control plane: Authentication, permissions, audit logs, monitoring, evaluation, and incident response.
Use explicit state rather than relying on an unstructured conversation history. Store the user request, retrieved sources, tool arguments, tool results, approval events, and final response as separate fields. This makes failures reproducible and prevents sensitive data from being copied unnecessarily into every prompt.
LangGraph is a practical choice for stateful workflows with loops and human approval nodes. A lightweight Python state machine may be better for a single narrow workflow. Multi-agent frameworks should be introduced only when separate roles genuinely reduce complexity; otherwise, they add coordination failures, latency, and cost.
Select models by task and risk
Do not choose one model for every operation. Route work according to capability and sensitivity:
- Use a stronger reasoning model for ambiguous requests, planning, and exception handling.
- Use a smaller model for classification, extraction, routing, and summarisation.
- Use deterministic code for calculations, validation, permissions, and policy checks.
- Consider open-weight models such as Llama-family models when self-hosting, predictable pricing, or network isolation matters.
A model gateway should record model version, prompt version, token usage, latency, tool calls, and outcome. In India’s price-sensitive market, cost per completed task is more meaningful than cost per million tokens. Set budgets at both request and workflow level, and stop execution when a task exceeds its maximum steps or spend.
For teams deploying open models, how to deploy Llama 3 agents offers a useful starting point. Fine-tuning is appropriate when the model repeatedly fails at a stable behaviour or format; it is not a substitute for current business knowledge. Use retrieval for changing policies, catalogues, schemes, and operational documents. Review best practices for fine-tuning LLMs on custom data before preparing a training set.
Build tools as controlled APIs
Tools are where an agent can create real value—and real damage. Every tool should have a narrow schema, strict authentication, validation, timeout, retry policy, and permission boundary. Never expose a raw shell, unrestricted SQL connection, or broad production API to an LLM.
For each tool, specify:
- Accepted arguments and allowed values
- Data classification and access scope
- Whether the operation is read-only or mutating
- Idempotency behaviour for retries
- Human approval requirements
- A safe, structured error response
Separate planning from execution. The model can propose a payment, refund, account change, or message, but the application should verify policy, user authorisation, limits, and destination before execution. High-impact actions should produce a preview and require explicit approval. This is especially important for fintech, healthcare, employment, education, and government-facing workflows.
Design retrieval for Indian data
RAG quality depends more on document preparation and retrieval policy than on the vector database brand. Ingest authoritative documents with titles, dates, issuing authority, jurisdiction, language, and version metadata. Preserve tables where they carry meaning, and remove duplicated or superseded material.
Use hybrid retrieval—keyword plus semantic search—when users ask for legal clauses, product codes, circular numbers, or names. Apply metadata filters for state, language, customer, date, and access role. Return citations or source references in the final answer, and instruct the agent to say when evidence is missing rather than inventing a conclusion.
Indian businesses often have mixed data: English contracts, Hindi or regional-language conversations, PDFs, spreadsheets, and code-switched queries. Test retrieval separately in English, Hindi, Tamil, Telugu, Bengali, Marathi, and the languages relevant to your users. Do not assume that translation preserves legal or financial meaning.
Indic language and voice considerations
Indic-language performance involves more than translation. Evaluate spelling variation, transliteration, code-switching, numerals, names, addresses, honourifics, and local terminology. Tokenisation can increase latency and cost for some languages, so measure actual tokens and response quality on representative queries instead of relying on English benchmarks.
For voice workflows, evaluate speech recognition on accents, noisy environments, names, amounts, and domain vocabulary. Design confirmation steps for phone numbers, rupee values, dates, and addresses. A restaurant or field-service workflow may benefit from the practical patterns in multilingual voice agents for restaurants in India, while healthcare teams should separately assess consent, escalation, and clinical risk.
DPDP-ready data governance
Treat privacy and security as architecture requirements, not a final compliance checklist. Map the personal data your agent receives, stores, retrieves, sends to model providers, and writes back to business systems. Define retention periods and delete data that is not needed for the stated purpose.
At minimum, implement:
- Consent and notice flows appropriate to the use case
- Purpose limitation and role-based access controls
- Encryption in transit and at rest
- PII detection, masking, and redaction before external model calls
- Provider contracts covering data use, retention, and security
- Audit logs for prompts, tools, approvals, and administrative actions
- A process for correction, deletion, grievance handling, and incident response
The Digital Personal Data Protection framework is only one part of the risk picture. Sector-specific rules, contractual duties, RBI or health-sector expectations, and cross-border transfer requirements may also apply. Obtain qualified legal and security advice for high-risk deployments; an agent prototype is not evidence of compliance.
Evaluation, observability, and release gates
Create a test set from real but anonymised cases before launch. Include normal requests, ambiguous inputs, adversarial prompts, outdated documents, tool failures, multilingual queries, and attempts to access another user’s data. Track:
- Task completion and factual accuracy
- Groundedness and citation correctness
- Tool-selection and argument errors
- Escalation and approval rates
- Latency, token use, and cost per successful task
- Harmful-action blocks and privacy violations
Run offline evaluations on every prompt, model, retrieval, or tool change. In production, use traces and sampled human review. Start with shadow mode or read-only actions, then expand permissions gradually. A maximum step count, deadlines, circuit breakers, schema validation, and idempotency keys are essential protections against loops and duplicate actions.
A practical 90-day build plan
Days 1–15: Select one workflow, document the baseline process, define permissions, collect evaluation cases, and establish the data map.
Days 16–35: Build a read-only prototype with retrieval, typed tools, structured outputs, logging, and a human review queue.
Days 36–60: Add multilingual tests, model routing, redaction, failure handling, cost controls, and offline evaluation gates.
Days 61–90: Pilot with a small user group, compare against the human baseline, expand only the safest permissions, and document incidents and rollback procedures.
The strongest Indian agent products will not be the most autonomous. They will be the ones that complete a valuable workflow reliably, show their evidence, protect user data, and hand control to people at the right moment.