LLM agents are software systems that interpret a goal, choose actions, call tools, and adapt based on results. That capability makes them useful for coding, support, research, operations, and internal workflows—but it also changes the security problem. A vulnerable chatbot may produce a harmful answer; a vulnerable agent can send an email, expose customer records, alter a repository, or modify production infrastructure.
Building secure LLM agents for developers therefore requires more than a strong system prompt. Treat the model as an untrusted decision-making component inside a controlled application. The application—not the model—must enforce identity, permissions, validation, approvals, isolation, and auditability.
This guide presents a practical architecture for teams building agents in India and deploying them for Indian users, enterprises, and regulated sectors in 2026.
Start with a threat model, not a framework
Before choosing an agent SDK, document what the agent can access and what could go wrong. A useful threat model covers:
- Assets: customer data, source code, credentials, payment records, health information, production systems, and model prompts.
- Actors: malicious users, compromised accounts, poisoned documents, untrusted websites, rogue insiders, and vulnerable dependencies.
- Entry points: chat messages, uploaded files, retrieval results, email, web pages, tickets, plugins, APIs, and tool responses.
- Impact: data disclosure, unauthorised transactions, privilege escalation, destructive changes, denial of service, and regulatory exposure.
Map every tool call as a security boundary. An agent that only searches public documentation has a very different risk profile from one that can issue refunds or deploy code. Teams building complex workflows should also separate agent roles and queues; the design principles in building distributed systems with AI agents are useful when deciding where to isolate state, retries, and permissions.
Design tools as narrow, typed capabilities
The safest tool is small, explicit, and difficult to misuse. Do not expose a general-purpose execute_sql, shell, browser, or HTTP client when the business task needs only one constrained operation.
Prefer tools such as:
get_order_status(order_id)rather than unrestricted database access.create_support_ticket(category, priority, summary)with enumerated values.search_documents(query, collection_id)limited to approved collections.propose_refund(order_id, amount)followed by a separate approval step.
Validate every argument outside the model using strict schemas. Reject unknown fields, enforce length and range limits, normalise identifiers, and verify that the authenticated user is allowed to access the referenced object. Parameterised queries are mandatory; prompt instructions must never be concatenated into SQL, shell commands, URLs, or policy expressions.
Use separate read and write tools. Make destructive actions impossible to trigger through a read-only credential. For agent teams, give each role its own tool registry instead of sharing one broad set of capabilities.
Treat prompt injection as an impact-control problem
Prompt injection cannot be eliminated reliably because an agent may receive instructions and untrusted data through the same context. A web page, PDF, email, or retrieved document can contain text such as “ignore previous instructions” and attempt to redirect the workflow.
Reduce the impact with layered controls:
- Label external content as untrusted data, not instructions.
- Keep system policy, user intent, retrieved content, and tool output in distinct structured fields where possible.
- Never let retrieved text directly define tool parameters or permissions.
- Require the agent to cite or identify the source of consequential claims.
- Re-check the proposed action against policy after retrieval and immediately before execution.
- Add adversarial test cases to every retrieval and browsing workflow.
A secondary classifier or LLM judge can flag suspicious content, but it is not a security boundary. Deterministic policy checks, permission enforcement, and approval gates must still make the final decision.
Isolate code execution and network access
If an agent writes or runs code, execute it outside the application host. Use an ephemeral container, micro-VM, or WebAssembly runtime with:
- No access to host sockets, cloud metadata endpoints, or internal service networks.
- A read-only base filesystem and a temporary workspace.
- CPU, memory, process, file-size, and execution-time limits.
- Egress allowlists rather than unrestricted internet access.
- Automatic destruction after completion or timeout.
- Separate credentials, never inherited host credentials.
A container is not automatically a sandbox. Review its kernel exposure, mounted volumes, Linux capabilities, runtime configuration, and network policy. For coding agents, compile and test in isolation, then return an artefact or patch for review instead of granting direct write access to the main repository.
When deploying open models, follow a controlled release process. How to deploy Llama 3 agents in production can inform model-serving decisions, but production security still depends on the surrounding runtime, identity layer, and tool gateway—not only on the model checkpoint.
Enforce identity, least privilege, and secrets hygiene
Every agent run should have a verifiable identity, user or service principal, purpose, tenant, and expiry. Pass this context to the tool gateway, where authorisation is enforced independently of the model.
Use:
- Short-lived OAuth tokens with narrowly defined scopes.
- Per-tool service accounts and separate environments.
- Vaults or managed secret stores for credentials.
- Server-side secret injection so the model never sees plaintext keys.
- Tenant isolation in retrieval indexes, caches, logs, and temporary files.
- Explicit denial for high-risk actions unless an approval token is present.
Never put API keys in prompts, conversation history, source code, or model-visible tool responses. Redact secrets and personal data from traces while retaining enough context for investigation. For health, finance, and public-sector deployments in India, involve legal, security, and domain owners early; sector-specific obligations may require stronger retention, access, and incident controls.
Add human approval where the blast radius demands it
Autonomy should be graduated, not binary. A practical policy is:
- Low risk: public search, draft generation, read-only knowledge lookup. Automate with rate limits.
- Moderate risk: creating tickets, sending external messages, editing non-production records. Show the exact proposed action and require confirmation.
- High risk: payments, refunds, account recovery, production deployment, deletion, or access-policy changes. Require authenticated approval, step-up verification, or two-person review.
The approval screen must display the resolved target, parameters, affected records, expected side effects, and relevant evidence. Do not ask a human to approve an opaque sentence such as “proceed with the task.” Bind approval to a specific action, user, timestamp, and expiry so it cannot be replayed after the plan changes.
Build observability into the agent runtime
Log a tamper-resistant audit trail for each run: request identity, model and policy versions, retrieved sources, tool arguments, approvals, results, latency, retries, and final outcome. Store sensitive traces with access controls and retention rules; full prompts should not automatically be visible to every developer.
Monitor for:
- Unusual tool sequences or privilege changes.
- Excessive retries, loops, token use, or outbound requests.
- Access to records outside the user’s normal scope.
- Repeated policy denials or prompt-injection detections.
- Unexpected changes in approval rates and tool failure patterns.
Use trace IDs across the model gateway, retrieval layer, tool service, and business system. This makes incident response practical and supports rollback when an agent produces an unsafe change. For voice or multimodal workflows, apply the same controls to transcriptions, attachments, and tool outputs; agent security is not limited to text chat.
Test before production—and continuously after it
Create an evaluation suite that includes both normal tasks and attacks. Test direct and indirect prompt injection, malicious files, poisoned retrieval documents, cross-tenant access, malformed tool arguments, replayed approvals, credential leakage, denial-of-service loops, and ambiguous user requests.
Run tests at three levels:
- Unit tests: schemas, policy functions, authorisation checks, redaction, and network rules.
- Integration tests: complete tool workflows with fake accounts and isolated data.
- Adversarial evaluations: automated attack cases plus manual review of high-impact paths.
Stage agents with synthetic or masked Indian customer data, feature flags, strict quotas, and a kill switch. Roll out read-only capabilities first, then expand only when evidence supports the additional autonomy.
A production checklist for Indian builders
Before launch, confirm that:
- Every tool has a named owner, schema, permission model, timeout, and rate limit.
- No model output is executed without validation and authorisation.
- External content cannot grant permissions or directly trigger actions.
- Code execution is isolated from hosts, metadata services, and private networks.
- Secrets remain outside model context and logs.
- High-impact actions require bound, auditable approval.
- Tenant boundaries and data-retention policies are tested, not assumed.
- Security alerts, incident ownership, rollback, and shutdown procedures are documented.
Secure agents are built by constraining what the model can do, verifying what it proposes, and making every consequential action observable. The strongest architecture is not the one that promises perfect model behaviour; it is the one that remains safe when the model, a document, a dependency, or a user behaves unexpectedly.