0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · ai codebase security

AI Codebase Security: A Practical Guide for Indian Teams

  1. aigi

    AI codebases need more than conventional application security. They combine source code, training data, model weights, prompts, evaluation sets, infrastructure and third-party packages—often across notebooks, cloud services and production APIs. A compromised dependency, leaked secret or poisoned dataset can affect both the software and the decisions it produces.

    For Indian startups, research teams and enterprises, AI codebase security should be treated as an engineering process that begins before training and continues through production monitoring. The goal is not to eliminate every risk; it is to make assets discoverable, access controlled, testable and recoverable.

    Map the AI attack surface first

    Create an inventory before choosing tools. Record where each asset lives, who owns it and how it moves through the system:

    • Source repositories, notebooks, CI/CD workflows and deployment manifests
    • Training, validation and evaluation datasets, including personal or regulated information
    • Model checkpoints, adapters, embeddings, tokenisers and prompt templates
    • Secrets such as API keys, cloud credentials, database passwords and signing keys
    • External models, packages, containers, plugins, agents and data providers
    • Production endpoints, vector databases, queues, GPUs and observability systems

    Classify assets by sensitivity and business impact. A public demo model and a model trained on health, financial or government data should not receive the same controls. Maintain a data-flow diagram showing collection, preprocessing, training, storage, inference and deletion. This makes it easier to identify unnecessary copies and exposed interfaces.

    Teams building reliable repositories should also apply the practices in machine learning GitHub repositories, particularly around repository structure, contribution rules and reproducible experiments.

    Secure source code, credentials and collaboration

    Use a central Git provider with organisation-level security settings. Require MFA, protect the default branch, enforce pull-request review and prevent direct pushes to production code. Apply least privilege to repository, cloud and registry access; a developer who can submit code does not necessarily need permission to read production data or publish model artefacts.

    At minimum:

    • Scan every commit and pull request for secrets, including deleted-file history.
    • Store credentials in a managed secret vault, never in notebooks, .env files or configuration committed to Git.
    • Rotate keys after exposure and use short-lived, workload-specific credentials where possible.
    • Sign releases, containers and model artefacts; retain provenance for each production version.
    • Separate development, staging and production accounts, networks and datasets.
    • Log repository, registry, cloud-console and model-access events, and review high-risk activity.

    Collaborative teams should establish CODEOWNERS, security escalation paths and a responsible disclosure process. The guidance on collaborative AI development is useful when several teams share datasets, evaluation code or deployment infrastructure.

    Make the software supply chain verifiable

    AI projects inherit risk from Python packages, JavaScript dependencies, CUDA libraries, container images, pre-trained models and hosted APIs. Pin dependencies and generate a software bill of materials (SBOM) for release builds. Run software composition analysis, static analysis and container scans in CI, but define a clear policy for exceptions so teams do not simply ignore noisy findings.

    Use private package and model registries where appropriate. Verify checksums and signatures for downloaded artefacts, review licences, and quarantine untrusted models before they enter a training or inference environment. Do not load arbitrary serialised objects or plugins: unsafe deserialisation can enable code execution.

    For open-source projects, generative AI for open-source security offers complementary approaches for triage and vulnerability discovery, but AI-generated fixes still require human review, tests and provenance checks.

    Protect data throughout the ML lifecycle

    Data security is inseparable from AI codebase security. Define collection purpose, retention periods, lawful basis and deletion procedures before training. Minimise fields, redact identifiers where feasible, and separate raw data from curated training data. Restrict access by project and environment rather than granting broad storage-bucket permissions.

    Use encryption in transit and at rest, with managed key rotation and tightly controlled decryption rights. Mask sensitive values in logs, traces, prompts and error reports. Test whether models memorise or reproduce sensitive records through extraction, membership-inference and prompt-based evaluations.

    For Indian deployments, map controls to applicable obligations under the Digital Personal Data Protection Act, 2023, sectoral rules and contractual requirements. A DPIA-style review is especially valuable for systems processing children’s data, health information, financial data or large-scale personal information. Keep evidence: consent or purpose records, access logs, retention decisions, incident reports and deletion outcomes.

    Harden training and model artefacts

    Training pipelines should verify dataset origin, integrity and licensing. Use checksums, immutable snapshots and documented transformations so a later change can be investigated. Test for poisoning, label manipulation, anomalous samples and train-test contamination. Keep evaluation data access separate from training access to preserve trustworthy measurements.

    Protect model checkpoints and adapters as sensitive intellectual property. Restrict download permissions, encrypt storage and record who promoted an artefact to production. Require reproducible build metadata: code commit, dataset version, base model, hyperparameters, package lockfile and evaluation results.

    When custom data is involved, secure experimentation practices matter as much as model quality. The guide to fine-tuning LLMs on custom data can help teams define safer preprocessing, evaluation and release gates.

    Secure inference, agents and APIs

    Treat model output as untrusted input. Validate schemas, enforce length and format limits, and apply business rules outside the model. Protect APIs with authentication, authorisation, rate limits, quotas and abuse monitoring. Never expose system prompts, internal tools or raw retrieval results by default.

    For retrieval-augmented systems, enforce document-level permissions at retrieval time—not only in the user interface. Defend against prompt injection by isolating instructions from retrieved content, limiting tool capabilities and requiring confirmation for irreversible actions. Agent workflows deserve additional controls: narrow tool scopes, sandbox execution, network egress restrictions and human approval for payments, account changes or data deletion. Teams implementing these systems can cross-check agentic workflow best practices for approval and failure-handling patterns.

    Build security into CI/CD and operations

    Create security gates that match risk rather than blocking every build. A practical pipeline includes secret scanning, dependency and licence checks, SAST, container scanning, infrastructure-as-code checks, unit tests, adversarial input tests and model evaluation. Block releases for exposed secrets, critical exploitable vulnerabilities, unsigned artefacts or failed privacy tests; document accepted lower-risk findings with owners and deadlines.

    In production, monitor for unusual token usage, prompt-injection attempts, data exfiltration, excessive tool calls, model drift, suspicious downloads and unexpected infrastructure changes. Maintain tested rollback paths for code, models and prompts. Keep backups isolated from the primary environment and regularly rehearse restoration.

    Incident response and governance

    Prepare playbooks for leaked credentials, poisoned data, vulnerable dependencies, model theft, unsafe output and personal-data exposure. Define who can disable an endpoint, revoke keys, quarantine a model, notify customers and preserve evidence. Run tabletop exercises at least annually and after major architecture changes.

    Use a risk register with an owner, severity, affected asset, mitigation, deadline and residual risk. Review it during product and model-release meetings. Security is stronger when product, engineering, legal, privacy and operations share the same release criteria.

    A practical 30-day starting plan

    • Week 1: inventory repositories, datasets, models, secrets and production endpoints.
    • Week 2: enforce MFA, branch protection, secret scanning, dependency pinning and environment separation.
    • Week 3: add SBOM generation, artefact signing, data-access reviews and adversarial model tests.
    • Week 4: run an incident exercise, fix the highest-risk findings and publish owners and recovery targets.

    AI codebase security is an ongoing operating discipline. Start with visibility and access control, then add provenance, testing, monitoring and recovery. This approach gives Indian builders a defensible foundation without slowing responsible experimentation.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.