Artificial intelligence becomes a security engineering problem the moment it reaches production. A model may expose sensitive data through prompts, follow malicious instructions hidden in retrieved documents, generate unsafe outputs, or become an attack surface through its APIs and surrounding infrastructure. These risks are amplified when AI systems connect to customer records, internal tools, payment workflows, or autonomous actions.
AI security production therefore requires more than protecting a model file. It means securing the complete AI application: data pipelines, training and inference infrastructure, prompts, retrieval systems, agents, APIs, identities, logs, and human approval processes. For Indian startups and enterprises, the programme should also account for the Digital Personal Data Protection Act, sectoral obligations, cloud-residency expectations, and the realities of operating with constrained security teams.
What AI security in production covers
A production AI system usually contains several interacting layers:
- User and application layer: Web, mobile, chat, and API interfaces that accept untrusted input.
- Orchestration layer: Prompt templates, tool calling, agents, workflow logic, and business rules.
- Model layer: Hosted foundation models, fine-tuned models, embedding models, classifiers, and guard models.
- Data layer: Training data, evaluation sets, vector databases, document stores, telemetry, and feedback.
- Infrastructure layer: Containers, Kubernetes, GPUs, cloud accounts, CI/CD pipelines, secrets, and networks.
- Governance layer: Access policies, human review, incident response, audit evidence, and vendor management.
The key design principle is to treat every model output as untrusted. A model can be highly accurate and still be manipulated, leak information, make an unsafe tool call, or produce an answer that violates policy. Security controls must exist outside the model, with deterministic enforcement at the application and infrastructure layers.
Threat model before deployment
Start with a documented threat model rather than a generic checklist. Identify assets, trust boundaries, attackers, abuse cases, and acceptable impact. Relevant assets can include personally identifiable information, health or financial records, proprietary prompts, retrieval indexes, model weights, API credentials, customer conversations, and generated decisions.
Common production threats include:
1. Prompt injection: An attacker causes the model to ignore system instructions or disclose protected information.
2. Indirect prompt injection: Malicious instructions are embedded in web pages, PDFs, emails, or knowledge-base documents consumed by retrieval or agents.
3. Sensitive information disclosure: Prompts, outputs, logs, embeddings, or model responses reveal personal or confidential data.
4. Insecure tool use: An agent invokes email, database, payment, shell, or ticketing tools without sufficient authorization.
5. Data poisoning: Training, fine-tuning, feedback, or retrieval data is altered to change model behaviour.
6. Model extraction and abuse: Attackers replicate capabilities, harvest outputs, or consume quota through automated queries.
7. Supply-chain compromise: A vulnerable package, container, model, dataset, plugin, or third-party API enters the stack.
8. Denial of service: Expensive prompts, long contexts, repeated requests, or GPU exhaustion increase cost or reduce availability.
9. Excessive agency: The system performs consequential actions without confirmation, rate limits, or rollback.
For each threat, record the attack path, business impact, likelihood, preventive controls, detection signals, owner, and residual risk. The result should drive architecture and testing—not sit as a compliance document.
Secure AI architecture for production
Enforce least privilege for models and agents
Give each workflow only the permissions it requires. A customer-support assistant should not possess unrestricted database access merely because the underlying language model can generate SQL. Use separate service identities, narrowly scoped tokens, read-only access by default, and explicit allowlists for tools and destinations.
For agentic systems, place a policy enforcement point between the model and every tool. Validate arguments against schemas, check the user’s authorization, apply transaction limits, and require approval for irreversible actions. Do not allow model-generated URLs, SQL, shell commands, or API destinations to pass directly into execution.
Separate tenants and data domains
Multi-tenant applications must enforce tenant boundaries in the application, database, retrieval, and caching layers. Apply tenant filters before retrieval, not after a model has received the documents. Verify that embeddings, conversation memory, response caches, and evaluation logs cannot cross tenants.
Use data classification labels such as public, internal, confidential, and restricted. Route restricted data only to approved models and regions. Where possible, redact or tokenize sensitive fields before inference and restore them only within a controlled application boundary.
Protect secrets and model endpoints
Store API keys, database credentials, signing keys, and cloud tokens in a managed secrets vault. Never place secrets in prompts, source code, notebooks, model weights, or client-side applications. Rotate credentials regularly and after suspected exposure.
Expose inference services through authenticated APIs with TLS, request-size limits, quotas, timeouts, and abuse controls. Keep administrative model endpoints on private networks. Restrict egress from inference workloads so a compromised process cannot freely exfiltrate data to the internet.
Data security and privacy controls
AI security production depends heavily on data discipline. Maintain an inventory showing where data originates, how it is transformed, which model processes it, where outputs are stored, and how long records are retained.
Practical controls include:
- Use approved datasets with provenance, licensing, and integrity checks.
- Scan documents for malware and active content before ingestion.
- Remove unnecessary personal data from prompts and training corpora.
- Apply deterministic PII detection and redaction before model calls.
- Encrypt data in transit and at rest, with managed key rotation.
- Separate production data from development and evaluation environments.
- Disable provider training on customer data unless explicitly approved.
- Define retention and deletion workflows for prompts, outputs, embeddings, and backups.
- Log access to sensitive data without copying the sensitive content into every log record.
For Indian organisations, map processing activities to the Digital Personal Data Protection Act, contractual commitments, and sector-specific requirements. Banking, insurance, healthcare, education, and government deployments may require additional controls for auditability, localisation, access, and data sharing. Obtain legal review for the actual use case rather than assuming that a generic AI policy is sufficient.
Model and application testing
Traditional software security testing is necessary but not enough. Combine application penetration testing with AI-specific evaluation.
Security testing categories
- Prompt-injection testing: Test direct and indirect attacks across languages, obfuscation, role-play, encoded instructions, and conflicting documents.
- Data-leakage testing: Attempt to retrieve secrets, system prompts, other users’ records, memorised training examples, and restricted documents.
- Tool-abuse testing: Submit malformed or adversarial arguments and verify authorization, schema validation, rate limits, and approval gates.
- Robustness testing: Measure behaviour under long contexts, contradictory evidence, unavailable tools, malformed files, and model timeouts.
- Supply-chain testing: Scan packages, images, dependencies, model files, datasets, and plugins for vulnerabilities and provenance gaps.
- Adversarial evaluation: Use a maintained test suite of harmful, biased, privacy-sensitive, and policy-sensitive prompts.
Run tests at pull request, pre-release, and scheduled intervals. Keep attack prompts versioned and treat failures as regression bugs. A single benchmark score cannot establish production security: evaluate the exact prompts, tools, policies, retrieval configuration, and model version used in the deployed system.
Runtime monitoring and detection
Security controls must continue after launch. Collect enough telemetry to detect abuse without creating a second privacy problem. Useful signals include authentication events, policy decisions, tool calls, retrieval metadata, token volume, latency, error rates, refusal rates, unusual geographic patterns, and model or prompt-version changes.
Build alerts for:
- Sudden increases in requests, token usage, or context length.
- Repeated attempts to override system instructions.
- Access to unusual tools, records, tenants, or data classifications.
- High-volume extraction-like querying.
- Responses containing secrets, personal data, or prohibited content.
- New outbound destinations or unexpected egress.
- Model drift, retrieval-quality degradation, and rising human overrides.
Use correlation IDs to connect a user request, model call, retrieval event, tool action, and final response. Protect logs with access controls, integrity mechanisms, retention limits, and redaction. For high-risk actions, retain an immutable audit record showing who initiated the request, what the model proposed, which policy approved it, and what actually executed.
Incident response for AI systems
An AI incident may involve both cyber compromise and unsafe model behaviour. Prepare playbooks before the first production failure. Define severity levels and owners across security, engineering, product, legal, privacy, and communications teams.
A practical response sequence is:
1. Contain: Disable a tool, revoke a key, block an IP range, pause a workflow, or switch to a safer model.
2. Preserve evidence: Secure relevant prompts, outputs, policy decisions, traces, configuration versions, and access logs.
3. Assess impact: Identify affected users, tenants, data types, actions, and time window.
4. Remediate: Patch code, tighten permissions, remove poisoned data, update prompts or policies, and rotate credentials.
5. Validate recovery: Re-run security tests and confirm that containment did not create new failures.
6. Notify where required: Follow contractual, regulatory, and internal breach-reporting procedures.
7. Learn and improve: Add the incident to regression tests and update the threat model.
Do not rely on deleting a conversation as an incident response plan. If an agent can send payments, modify records, or contact customers, implement transaction reversibility and emergency kill switches.
Governance, ownership, and assurance
Assign clear accountability. The product owner should own business risk; engineering should own implementation; security should define controls and test them; privacy and legal teams should assess data processing; and operations should own availability and incident response.
Create an AI system record containing:
- Intended use and prohibited use cases.
- Model, prompt, retrieval, and tool versions.
- Data sources, classifications, retention, and vendors.
- Known limitations and evaluation results.
- Access-control design and human approval points.
- Monitoring metrics, escalation thresholds, and rollback procedures.
- Security and privacy review dates.
Use recognised references such as the NIST AI Risk Management Framework, NIST Secure Software Development Framework, OWASP guidance for large language model applications, ISO/IEC 27001, and ISO/IEC 42001 where appropriate. These frameworks are most valuable when translated into engineering tickets, acceptance criteria, evidence requirements, and operational ownership.
A production readiness checklist
Before enabling a high-impact AI feature, verify that:
- The threat model covers prompt injection, data leakage, tool abuse, extraction, poisoning, and denial of service.
- The model and tools use least-privilege identities.
- Tenant isolation is tested across retrieval, memory, caches, and logs.
- Sensitive data is minimised, redacted, encrypted, and governed by retention rules.
- Model providers, subprocessors, data-use terms, and regions are documented.
- Input, output, and tool-call validation is implemented outside the model.
- Red-team and regression tests run against the production configuration.
- Rate limits, quotas, timeouts, cost controls, and circuit breakers are active.
- Runtime monitoring and privacy-safe audit logs are available.
- Human approval exists for irreversible or high-impact actions.
- Kill switches, rollback paths, and incident playbooks have been tested.
- Security ownership and review dates are explicit.
Building an India-ready AI security programme
Indian AI companies can make security a product advantage by designing it into the first production architecture. Start with a narrow, well-defined use case; classify data before selecting a model; prefer managed controls that a small team can operate; and document every external data and model dependency.
For startups, the highest-value early investments are identity and access management, secrets management, tenant isolation, secure logging, provider due diligence, automated prompt and output tests, and a tested incident response plan. As the product gains customers, add independent penetration testing, formal risk assessments, continuous supply-chain scanning, disaster recovery exercises, and evidence aligned to enterprise procurement requirements.
The goal is not to eliminate every uncertain model response. It is to ensure that uncertainty cannot silently become a data breach, unauthorised transaction, regulatory violation, or uncontrolled operational event. Secure AI in production is a layered system of constrained permissions, verified data flows, measurable behaviour, and rapid human intervention.
FAQ: AI security production
What is AI security production?
It is the practice of securing an AI system throughout live operation, including its models, data, prompts, APIs, agents, tools, infrastructure, monitoring, and governance.
Is prompt engineering enough to secure an AI application?
No. Prompts can guide behaviour but cannot enforce authorization or guarantee confidentiality. Use application-level policy checks, access controls, validation, isolation, and monitoring.
How do I protect an AI agent that uses business tools?
Use least-privilege service identities, strict tool schemas, argument validation, user authorization checks, rate limits, transaction limits, approval gates, audit logs, and a kill switch.
What should Indian startups prioritise first?
Prioritise data classification, tenant isolation, secrets management, secure APIs, provider due diligence, prompt-injection testing, privacy-safe logging, and incident response. Expand controls as risk and customer requirements grow.
Which standards can guide an AI security programme?
NIST AI RMF, NIST SSDF, OWASP LLM application guidance, ISO/IEC 27001, and ISO/IEC 42001 are useful references. Select controls based on your actual risk, sector, customers, and regulatory obligations.
Apply for AI Grants India
If you are an Indian AI founder building a secure, scalable product, apply for support and funding through AI Grants India. Get help turning responsible AI security into a production-ready advantage.