AI agents rarely operate alone. A production agent may call a model provider, search service, payment gateway, CRM, database, messaging platform, or internal API during one task. Each connection introduces credentials, permissions, cost, and failure modes. API keys for AI agents therefore need to be treated as production security controls—not configuration strings copied into a prompt or code repository.
This guide explains how to design, store, restrict, rotate, monitor, and revoke API keys for agentic applications. The principles apply whether you are building a customer-support bot, a multilingual voice system, or a distributed workflow with several specialised agents. For broader architecture decisions, see this guide to building distributed systems with AI agents.
What API keys do—and what they do not do
An API key is a credential that identifies a client to an API provider. The provider uses it to decide whether a request is allowed and often to associate the request with a project, account, quota, or billing profile.
API keys commonly support:
- Authentication: identifying the calling application or project.
- Authorisation: limiting access to selected endpoints, models, tools, or operations.
- Usage accounting: attributing requests, tokens, storage, or transactions to a team or project.
- Rate control: applying quotas and request limits.
- Incident investigation: tracing activity to a specific service or workload.
An API key is not automatically a complete identity system. Many keys are bearer credentials: anyone who possesses one may use it until the provider rejects, expires, or revokes it. A key also does not prove which end user authorised an action. For high-risk operations, combine provider credentials with application authentication, user-level authorisation, audit logs, and approval workflows.
API keys in an AI agent workflow
A typical agent receives a task, selects a tool, constructs an API request, and processes the response. The key should be used by a trusted backend or tool gateway—not exposed to the model, browser, mobile app, or end user.
A safer flow looks like this:
1. The user authenticates with your application.
2. Your backend evaluates the user’s permissions and the agent’s allowed tools.
3. The agent requests a tool action through a controlled interface.
4. A tool service retrieves the relevant secret from a secrets manager.
5. The service validates inputs, calls the external API, and filters the response.
6. The system records the action, outcome, latency, and cost without logging the secret.
This separation matters for voice and customer-facing deployments. For example, an agent handling fintech customer onboarding with voice agents may need access to identity, notification, and case-management services, but should not receive unrestricted credentials for any of them.
Design credentials around least privilege
Create credentials for a specific workload, environment, and purpose. Avoid one shared key for development, staging, production, and every agent in the company.
Use these controls where the provider supports them:
- Restrict keys to required APIs and endpoints.
- Separate read, write, delete, and administrative operations.
- Set project, tenant, account, or resource boundaries.
- Apply spending limits, quotas, and request-rate limits.
- Bind credentials to a service identity, workload, or network location.
- Use short-lived tokens or workload identity instead of long-lived keys where possible.
- Create separate keys for local development, CI/CD, staging, and production.
A customer-support agent may need to read a ticket and draft a response, while a billing agent may create a refund only after human approval. Model these as different tools and credentials. Do not rely on the language model to enforce the distinction.
Store secrets outside code and prompts
Never place a production key in source code, a frontend bundle, a notebook shared with a team, an agent prompt, or a system message. Prompt text is not a secure vault, and an agent can reveal sensitive context through logs, tool calls, or manipulated output.
Use a secrets manager or cloud key-management service, with access granted to the runtime identity. Common implementation practices include:
- Inject secrets at runtime rather than committing them to repositories.
- Keep
.envfiles local and excluded from version control. - Scan commits, container images, logs, and artifacts for leaked credentials.
- Mask secrets in application, tracing, and CI/CD logs.
- Restrict who can read or update secrets, and log those access events.
- Use separate secret versions during rotation so old and new credentials can overlap briefly.
For Indian teams, account for the data flows created by third-party model, messaging, and analytics providers. Credential security is only one part of the design; review where prompts, personal data, call recordings, and tool responses are processed and retained. This is especially important when building patient follow-up with voice agents in India.
Build rotation and revocation into the product
Key rotation should be a routine deployment operation, not an emergency-only procedure. Maintain an inventory showing each key’s owner, provider, environment, permissions, creation date, expiry date, dependent services, and last observed use.
A practical rotation process is:
1. Create a new credential with the same or narrower permissions.
2. Deploy it through the secrets manager without exposing its value.
3. Verify successful requests and monitor error rates.
4. Disable the previous credential after the overlap window.
5. Revoke it and record the change in the incident or change log.
Define a rapid revocation path for suspected compromise. Automate alerts for unusual geography, sudden request spikes, new models or endpoints, quota exhaustion, and unexpected spend. If a key appears in a public repository, assume it is compromised: revoke it first, investigate usage, review logs, and then issue a replacement.
Monitor agent activity, not just key activity
Provider dashboards are useful, but an agent platform should also record context around each tool call. Capture:
- Agent, service, environment, and tenant identifiers.
- Tool name, endpoint category, status code, latency, and retry count.
- Request and response sizes, token usage, and estimated cost.
- Approval state for sensitive actions.
- A correlation ID linking the user request to downstream calls.
Do not log raw API keys, access tokens, full payment details, health information, or unnecessary personal data. Store only what is needed for debugging, security review, billing, and compliance. Add budgets and circuit breakers so a prompt-injection attack or software loop cannot generate uncontrolled calls.
Common mistakes to avoid
- One key for every agent: compromise spreads across the entire platform.
- Keys in client-side code: users can extract them from browsers and apps.
- Unrestricted tool wrappers: the agent can call more operations than its task requires.
- Logging complete requests: secrets and personal data often leak through observability systems.
- No expiry or owner: abandoned credentials remain active indefinitely.
- Retries without limits: transient failures can become expensive request storms.
- Treating prompt instructions as access control: a model can be manipulated; enforce policy in code.
A production checklist
Before launching an agent, confirm that you can answer yes to the following:
- Does every environment and major workload have distinct credentials?
- Are secrets stored in a managed vault and absent from code, prompts, and client bundles?
- Are permissions, endpoints, quotas, and budgets narrowly scoped?
- Can the team rotate or revoke a key without rebuilding the entire application?
- Are tool calls authenticated, authorised, rate-limited, and auditable?
- Are sensitive operations protected by validation and human approval?
- Are alerts configured for anomalous usage and unexpected spend?
- Has the team rehearsed credential-compromise response?
API keys remain useful for straightforward service integrations, but mature agent systems should progressively adopt short-lived credentials, workload identity, gateway-based access, and policy enforcement. Secure the tool layer first: the model may decide what to ask for, but trusted application code must decide what is permitted.
Apply for AI Grants India
Building a secure AI product in India? Apply for AI Grants India to explore funding support for experimentation, infrastructure, and responsible deployment.