0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · ai agent security research

AI Agent Security Research: Threats, Testing and Safeguards

  1. aigi

    AI agents are no longer limited to producing text. They can browse, call APIs, execute code, send messages, update records and make decisions across multi-step workflows. That autonomy creates a distinct security problem: an agent may be manipulated through its inputs, tools, memory, permissions or surrounding infrastructure.

    AI agent security research focuses on understanding and reducing these risks before agents are trusted with production access. For Indian startups, enterprises, public-sector teams and researchers, the priority is not to stop experimentation. It is to make agent behaviour observable, constrain its authority and test failure modes systematically.

    What makes agent security different

    A conventional application usually follows code-defined paths. An AI agent combines a model with prompts, retrieval systems, tools, memory, identity controls and an orchestration layer. Its behaviour can therefore change when it encounters new instructions, documents or tool results.

    A useful threat model covers:

    • The model: jailbreaks, prompt injection, unsafe tool selection and unreliable reasoning.
    • Instructions: system prompts, developer messages, user requests and conflicting context.
    • Tools: APIs, browsers, databases, code interpreters, payment systems and communication channels.
    • Data: retrieved documents, conversation history, secrets, personal information and proprietary records.
    • Identity: user authentication, service accounts, delegated permissions and agent-to-agent trust.
    • Runtime: containers, plugins, orchestration services, logs, queues and model providers.

    This is why securing an agent requires more than filtering user prompts or selecting a reputable model.

    Priority threats for security researchers

    Prompt injection and indirect injection

    An attacker can ask an agent to ignore its rules directly, but the more difficult case is indirect prompt injection. Malicious instructions may be hidden in a webpage, email, PDF, support ticket or database record that the agent is asked to process. If the agent treats retrieved content as authority, it may leak data or perform an unauthorised action.

    Research should test whether the system can distinguish instructions from untrusted data, especially when content is multilingual, encoded, obfuscated or spread across several documents.

    Excessive agency and unsafe tool use

    An agent with broad permissions can turn a minor model mistake into a serious incident. Examples include deleting records, approving refunds, exposing customer data or sending an unreviewed message to thousands of users.

    Assess every tool by its impact, reversibility and required scope. Read-only access, narrow API permissions, allow-listed destinations and human approval for high-impact actions should be the default.

    Data leakage and cross-tenant exposure

    Agents may reveal system prompts, credentials, retrieved documents, conversation history or another customer’s information. Risks increase when memory is shared across users, logs contain sensitive content or retrieval filters are weak.

    Test tenant isolation, redaction, retention, encryption and access decisions at both the application and data-store layers. Do not treat a model instruction such as “never reveal secrets” as a substitute for technical controls.

    Supply-chain and tool risks

    Third-party plugins, open-source packages, model servers and remote tool endpoints can introduce malicious code or unexpected data flows. A compromised tool description may influence agent planning, while an unpinned dependency can silently change runtime behaviour.

    Maintain a software and model inventory, verify package provenance, pin versions, scan dependencies and review outbound network access. Treat tool schemas and descriptions as security-sensitive configuration.

    Memory poisoning and persistent compromise

    Long-term memory can make an agent more useful—and give attackers a way to plant durable instructions. A poisoned memory entry may affect future users or workflows long after the original interaction.

    Research should examine how memories are created, approved, retrieved, updated and deleted. Store provenance with each memory and separate user preferences from operational instructions.

    A practical AI agent security research workflow

    1. Define assets and unacceptable actions

    List the data, systems and business outcomes the agent can affect. Then define actions that must never happen autonomously, such as transferring funds, changing access rights or disclosing regulated information.

    For each workflow, document the principal, agent, tools, data sources, decision points and human approvals. This creates a usable security boundary rather than a vague AI risk statement.

    2. Map the attack surface

    Create an interaction diagram covering user inputs, retrieved content, model calls, memory, tools, identity tokens, logs and external services. Record whether each component is trusted, partially trusted or untrusted.

    Teams building voice agents for Indian businesses should also include call recordings, speech-to-text transcripts, caller identity, telephony providers and escalation paths. Voice does not remove prompt injection; it adds new channels for impersonation, transcription errors and sensitive data exposure.

    3. Build an adversarial test suite

    Use repeatable test cases rather than one-off demonstrations. Include:

    • Direct and indirect prompt injection.
    • Conflicting instructions and privilege escalation attempts.
    • Malicious documents, links, attachments and tool outputs.
    • Data-exfiltration requests and cross-user retrieval tests.
    • Tool misuse, replayed approvals and parameter tampering.
    • Unicode, multilingual, encoded and unusually long inputs.
    • Memory poisoning, denial of service and model-output manipulation.

    Run tests across model versions, orchestration changes and realistic business data. Track attack success rate, sensitive-data exposure, unauthorised tool calls, unsafe completion rate and time to detection.

    4. Add containment before capability

    Security controls should sit outside the model wherever possible. Useful controls include:

    • Short-lived, scoped credentials instead of broad static keys.
    • Policy enforcement that validates tool calls and parameters.
    • Sandboxed code execution and restricted network egress.
    • Separate read and write tools, with write actions requiring approval.
    • Rate limits, spend limits, transaction caps and circuit breakers.
    • Output validation, data-loss prevention and destination allow-lists.
    • Detailed, tamper-resistant audit logs with correlation IDs.

    A human approval step should show the exact action, affected records, destination and reason—not merely ask the reviewer to approve an opaque agent decision.

    Standards and governance for Indian teams

    Use the NIST AI Risk Management Framework and its generative-AI guidance as reference points for governance, measurement and accountability. For security engineering, map controls to established practices such as identity management, secure software development, threat modelling, incident response and vendor risk management.

    Indian organisations should also account for the Digital Personal Data Protection Act, 2023, sector-specific obligations, contractual data-residency requirements and internal retention policies. Healthcare deployments need stronger controls around consent, access, auditability and clinical escalation; a discussion of operational requirements appears in this guide to HIPAA-compliant voice agents for hospitals, although Indian teams must not assume HIPAA alone satisfies local obligations.

    Assign ownership across security, engineering, legal, product and operations. Every production agent should have a named owner, documented capabilities, an incident playbook and a process for suspending access quickly.

    Designing safer agent architectures

    Prefer least privilege, bounded autonomy and explicit state transitions. Let the model propose an action, but let deterministic policy code decide whether it is allowed. Use separate agents or services for planning, retrieval and execution when that reduces privilege concentration.

    Keep secrets out of prompts and model context. Partition data by tenant and purpose. Treat retrieved content as untrusted. Require confirmation for irreversible actions, and provide a safe fallback when confidence, policy checks or tool responses are ambiguous.

    For customer-facing systems, assess whether the agent should answer at all. A restaurant table-booking voice agent may safely collect a preferred time and pass it to a controlled booking service; it should not receive unrestricted database or payment access simply to improve convenience.

    What strong research outputs look like

    A useful security research programme produces more than a list of jailbreaks. It should deliver:

    • A threat model tied to real assets and workflows.
    • Reproducible attack cases and severity ratings.
    • Measured baseline and post-control performance.
    • Clear remediation owners and deadlines.
    • Evidence that fixes survive regression testing.
    • Detection rules, escalation paths and recovery procedures.

    Researchers should publish responsibly: share enough detail to improve defence without exposing live credentials, customer data or an easily weaponised exploit. Builders can contribute by releasing test datasets, evaluation harnesses, policy patterns and anonymised incident learnings.

    A deployment checklist

    Before granting production access, confirm that the team has:

    • Inventoried tools, models, data stores and external providers.
    • Tested direct and indirect prompt injection.
    • Enforced least-privilege identity and tenant isolation.
    • Sandboxed code and restricted outbound traffic.
    • Added approval gates for high-impact actions.
    • Logged prompts, tool calls, decisions and outcomes safely.
    • Defined rollback, shutdown and incident-notification procedures.
    • Scheduled continuous evaluation after model or workflow changes.

    AI agent security research is most valuable when it changes architecture and operating practice. In 2026, Indian teams should treat autonomy as a capability to be earned through evidence: narrow permissions, measurable tests, strong monitoring and a fast route to human control.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.