Agentic applications are not simply chatbots with longer prompts. They are software systems in which Claude can decide which operation to perform, receive results from external tools, revise its plan, and stop when the task is complete. The application—not the model—must control permissions, state, retries, budgets, and user approvals.
For Indian startups and engineering teams, this distinction matters. A useful agent might reconcile invoices, classify support requests, query internal policies, prepare a GST workpaper, or coordinate actions across business systems. In each case, Claude provides reasoning and language capability, while your backend remains responsible for execution and safety.
This guide explains how to design that boundary with the Claude API in 2026.
Start with a workflow, not an autonomous agent
Before writing an agent loop, define the business process as a sequence of observable steps. Ask:
- What information enters the workflow?
- Which decisions require model judgment?
- Which operations can be deterministic code?
- Which actions change data, spend money, contact a customer, or create legal exposure?
- Where must a person approve or correct the result?
Many “agent” projects are better implemented as a controlled workflow with one or two model calls. Use Claude for classification, extraction, planning, or drafting; use ordinary code for validation, calculations, authentication, and writes. This approach is cheaper, easier to test, and more predictable than granting a model broad freedom.
If the system must coordinate several services, establish clear contracts between components. Patterns discussed in building distributed systems with AI agents are useful when queues, workers, retries, and service ownership become part of the design.
The Claude API tool-use loop
Claude’s tool-use interface lets you describe callable operations with JSON Schema. A tool should represent one narrow capability, not an entire backend. For example:
{
"name": "find_invoice",
"description": "Find an invoice by supplier name and invoice number.",
"input_schema": {
"type": "object",
"properties": {
"supplier_name": { "type": "string" },
"invoice_number": { "type": "string" }
},
"required": ["supplier_name", "invoice_number"],
"additionalProperties": false
}
}A typical cycle is:
1. Send the user request, system instructions, conversation state, and available tools.
2. Inspect the response and its stop_reason.
3. If Claude returns a tool_use block, validate its input against your schema.
4. Authorise the operation for the current user and tenant.
5. Execute the tool in your application.
6. Return a tool_result containing a structured success or error payload.
7. Repeat within a strict step and token budget.
8. Present the final answer only after validating citations, totals, and side effects.
Treat tool arguments as untrusted input. JSON validity does not prove that a request is safe or commercially acceptable. A tool such as issue_refund should enforce amount limits, identity checks, idempotency keys, and approval status independently of Claude.
Design tools that are easy to control
Good tools are narrow, typed, idempotent where possible, and explicit about failure. Prefer get_customer_balance and create_draft_invoice over a generic run_sql or execute_admin_action tool.
Include these properties in tool design:
- Clear input schemas: Mark required fields and reject unknown fields.
- Stable outputs: Return structured fields such as
status,data,error_code, andretryable. - Permission boundaries: Check tenant, role, resource ownership, and environment in the backend.
- Idempotency: Ensure retries do not create duplicate payments, tickets, or messages.
- Timeouts: Return a controlled failure rather than leaving the loop hanging.
- Audit metadata: Record actor, tool name, arguments, result, and approval context.
For customer-facing products, combine tool use with multilingual response design. A support agent serving Indian users may need to preserve product names and numbers while responding in English, Hindi, Tamil, or another requested language; the principles in building multilingual chatbots for Indian startups provide a useful product lens.
Choose the right orchestration pattern
Use the simplest pattern that meets the requirement.
Single agent with tools: Best for bounded tasks such as retrieving account information, summarising documents, or creating a draft after validation.
Router: A lightweight classifier sends a request to a specialist workflow. Keep routing categories mutually understandable and include an “uncertain” path that asks a human or the user for clarification.
Planner and workers: A planner creates a task list, while workers execute independent subtasks. Require each worker to return evidence and a defined output schema; do not pass unstructured prose between every component.
Generator and evaluator: One call creates an output and another checks it against explicit tests. For code, run the code or test suite rather than relying only on a model critique.
Multi-agent designs add latency, cost, and more failure points. Use them for genuine separation of expertise or parallelism—not because “agent” sounds more advanced.
State, memory, and durable execution
Keep workflow state outside the prompt. Store a run identifier, current step, tool calls, tool results, user approvals, and failure status in a durable database or queue. Persist after every meaningful transition so a process can resume after an API timeout or worker restart.
Separate three kinds of context:
- Task state: Facts and outputs needed to complete the current run.
- Conversation context: User-visible history relevant to the current interaction.
- Knowledge retrieval: Documents or records fetched for this task, with source identifiers.
Do not treat a vector database as memory by default. Retrieve only relevant, permission-filtered records and attach provenance. For regulated or financial use cases in India, retain the source document, retrieval timestamp, tenant identity, and transformation history.
A durable workflow should support replay. Given the same recorded inputs and tool results, engineers should be able to reconstruct why the system made a decision. This is more valuable than exposing hidden chain-of-thought; log concise decisions, tool intents, evidence, and outcomes instead.
Guardrails, approvals, and security
An agent must have less authority than the human or service account behind it. Apply defence in depth:
- Use scoped credentials and separate read and write tools.
- Put sensitive actions behind approval gates.
- Redact secrets and personal data from prompts and logs.
- Defend against prompt injection in retrieved documents and web content.
- Enforce network egress, file access, and tenant isolation at infrastructure level.
- Add maximum steps, wall-clock time, token budget, and tool-call limits.
- Require confirmation for payments, deletion, external messages, and policy exceptions.
The system should fail closed. If a tool returns ambiguous data, stop or request clarification rather than guessing. For a deeper implementation checklist, see secure autonomous AI workflows.
Evaluation and production observability
A successful demo is not evidence of a reliable agent. Build an evaluation set from real, anonymised tasks and include ordinary, ambiguous, adversarial, and failure cases. Measure:
- Task completion and factual accuracy
- Correct tool selection and argument validity
- Unnecessary tool calls and loop length
- Approval bypass attempts
- Latency, token consumption, and cost per completed task
- Recovery from timeouts, malformed data, and partial outages
Trace every run with a correlation ID. Capture model version, prompt version, tool schema version, latency, stop reason, errors, and final outcome. Avoid storing raw sensitive content where it is not required. Sample traces for human review and turn recurring failures into regression tests.
For engineering teams working with newer Claude coding capabilities, Claude Opus coding: a deep dive can complement this workflow guidance, but code-generation agents still require sandboxing, tests, review, and restricted credentials.
India-specific deployment considerations
Plan for intermittent networks, regional-language inputs, and integrations with Indian business systems. Queue long-running tasks instead of holding an HTTP request open. Use idempotent workers and clear user-visible statuses such as “awaiting approval” or “retrying a bank-data fetch.”
For GST, lending, healthcare, education, or Account Aggregator-related products, keep the model away from unreviewed final decisions unless the applicable compliance and risk process explicitly permits it. Minimise personal data, document retention rules, and establish where data is processed and who can access traces. UPI, banking, and government-system actions should be mediated by verified APIs and deterministic controls, never by free-form browser automation alone.
A practical launch checklist
Before exposing an agent to customers, confirm that you can answer yes to these questions:
- Does every tool have a narrow purpose, schema, timeout, and permission check?
- Can the workflow resume safely after any tool failure?
- Are writes idempotent and approval-protected?
- Is there a hard limit on steps, cost, and execution time?
- Can an engineer replay and explain a failed run?
- Do evaluations cover Indian names, addresses, currencies, languages, and business formats where relevant?
- Is there a human escalation path?
Agentic workflows with the Claude API become dependable when they are engineered as constrained software systems rather than handed over to model autonomy. Start with a narrow workflow, instrument every transition, and expand permissions only after measured reliability. For founders building this infrastructure in India, building high-performance AI applications with open-source tools offers complementary options for retrieval, orchestration, and cost control.