0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · python sandboxes for agents

Python Sandboxes for Agents: Secure Code Execution

  1. aigi

    Agents that write or execute Python can calculate, transform files, query data, call internal tools, and automate workflows. That flexibility is also a security boundary: a model may generate unsafe code, a user may deliberately submit it, or a dependency may behave unexpectedly. Running such code directly inside the agent service is an avoidable risk.

    Python sandboxes for agents provide a separate execution environment with explicit limits on CPU, memory, files, network access, runtime, and permissions. They are useful for coding agents, data-analysis assistants, workflow automation, and enterprise copilots—but only when treated as infrastructure isolation, not as a Python feature.

    For teams building production systems, sandboxing belongs alongside identity, observability, approval workflows, and data governance. This matters especially for Indian startups handling payments, healthcare records, customer conversations, or business documents. For context on agent architecture at scale, see building distributed systems with AI agents.

    What a Python sandbox should protect

    A sandbox should assume that executed code is untrusted. The code may be malicious, accidentally destructive, or simply too expensive to run. Define the assets and actions it must not reach:

    • Host access: no access to the host filesystem, kernel interfaces, process table, sockets, or container runtime.
    • Credentials: no cloud keys, database passwords, service-account tokens, environment secrets, or mounted credential files.
    • Network destinations: block the internet by default and allow only specific APIs through a controlled gateway.
    • Resources: enforce CPU, memory, disk, process-count, output-size, and wall-clock limits.
    • Tenant data: prevent one user, customer, or agent run from reading another's files or results.
    • Availability: stop infinite loops, fork bombs, oversized outputs, and repeated expensive jobs.

    A sandbox cannot compensate for excessive privileges in the surrounding application. If the orchestrator gives the sandbox a production database token, isolation has already failed at the design level.

    Choose the isolation layer carefully

    Not all approaches provide the same protection. Use the weakest mechanism only when the threat model supports it.

    Language-level restrictions

    Tools such as RestrictedPython can limit syntax and selected built-ins. They are useful for tightly controlled expressions or educational environments, but they are not a complete security boundary for hostile code. Python's introspection, native extensions, serialization, and dependency ecosystem make in-process restrictions difficult to audit.

    Do not rely on exec(), a filtered globals dictionary, or removed built-ins as your only defence. These controls can support policy enforcement, but untrusted code should not share the agent server's process.

    Browser and WebAssembly execution

    Pyodide and similar WebAssembly-based runtimes can execute Python in a browser or isolated worker. They are attractive for interactive notebooks and low-risk transformations because the runtime has a narrower operating-system surface. They still need limits for memory, execution time, package availability, and data transfer.

    Containers

    A short-lived container is a practical baseline for many development teams. Run one job per container, use a read-only root filesystem, drop Linux capabilities, apply seccomp and AppArmor or SELinux profiles, set resource quotas, and run as a non-root user. Keep the container image minimal and pin dependencies.

    A container is not the same as a virtual machine. Container escape vulnerabilities, kernel exposure, misconfigured mounts, and access to the container socket can create serious risk. For hostile multi-tenant workloads, consider microVMs or a managed isolated execution service.

    MicroVMs and dedicated workers

    MicroVMs provide stronger workload separation with faster startup than traditional virtual machines. They are appropriate when agents execute arbitrary code for multiple customers or when regulatory and contractual requirements demand a stronger boundary. The operational cost is higher, so use them for the workloads that need them rather than every deterministic function.

    A production execution flow

    A robust agent should not send model-generated Python straight to a worker. Use an explicit pipeline:

    1. Classify the request. Decide whether code execution is necessary and whether the task is low, medium, or high risk.
    2. Create an isolated job. Generate a unique run ID and provision a fresh container, microVM, or worker.
    3. Pass minimal inputs. Copy only the required files or structured data; never mount broad application directories.
    4. Apply policy. Set package, filesystem, network, resource, and time limits before starting execution.
    5. Execute with a deadline. Stream bounded logs and capture structured results rather than unrestricted stdout.
    6. Validate outputs. Check type, size, schema, provenance, and whether the result is safe for the next tool.
    7. Destroy the environment. Delete temporary files, revoke short-lived credentials, and remove the worker after completion.
    8. Record an audit event. Store the requester, model version, code hash, policy decision, resource use, and outcome.

    For sensitive operations—such as sending money, changing a medical record, or contacting a customer—require human approval outside the sandbox. Sandboxing limits execution; it does not decide whether an action should happen.

    Practical controls for Indian deployments

    Design for the data flows your product will actually support. If a Bengaluru-based support agent processes recordings, transcripts, or CRM exports, document where each artifact is stored and whether the sandbox can reach it. Use region-appropriate cloud controls and confirm provider commitments for retention, logging, and cross-border processing.

    Recommended controls include:

    • Use short-lived, scoped credentials issued per job.
    • Keep network egress disabled unless a documented allowlist is required.
    • Proxy approved API calls so the sandbox never receives raw long-lived keys.
    • Encrypt inputs, outputs, and temporary object storage.
    • Redact personal, financial, and health information from logs.
    • Apply per-user and per-tenant quotas to control cost and abuse.
    • Scan dependencies and pin package versions; avoid allowing arbitrary pip install during a run.
    • Separate development, staging, and production data completely.
    • Test escape attempts, prompt injection, denial of service, and data exfiltration.

    These measures are particularly relevant to voice and customer-service systems. A voice agent that invokes code to look up orders or calculate eligibility should expose narrow tools, not a path into the entire backend. Teams working on fintech customer onboarding with voice agents or patient follow-up with voice agents should model sensitive fields and approval paths before enabling execution.

    A safer minimal pattern

    The following conceptual configuration is more appropriate than calling exec() in the web process:

    job = {
        "image": "agent-python:2026-01",
        "command": ["python", "/workspace/main.py"],
        "network": "none",
        "read_only_root": True,
        "user": "10001:10001",
        "cpu_limit": "500m",
        "memory_limit": "256Mi",
        "timeout_seconds": 10,
        "pids_limit": 32,
        "workspace_bytes": 10_000_000,
    }
    result = sandbox_runner.run(job, input_files=approved_files)

    This is not a security guarantee by itself. The runner must enforce these settings at the container, microVM, or managed-service layer. The application should also validate result, cap output size, and treat all returned text as untrusted.

    Measure security and reliability

    Track more than successful executions. Useful metrics include startup latency, p95 runtime, timeout rate, memory peaks, denied network attempts, rejected packages, output truncations, cleanup failures, and cost per run. Alert on unusual behaviour such as a sudden increase in outbound attempts or repeated resource exhaustion from one tenant.

    Run adversarial tests before launch and after every change to images, kernels, runtimes, policies, or model providers. Maintain a threat model and an incident procedure: disable execution, revoke worker credentials, preserve audit records, and notify affected customers where required.

    When not to use a sandbox

    Do not add arbitrary Python merely because an agent can generate it. Prefer deterministic, typed tools for routine operations such as currency conversion, database lookups, document retrieval, or booking workflows. A narrow function with validation is easier to secure, test, and explain than a general-purpose interpreter.

    Use a sandbox when computation or code flexibility provides clear value—data analysis, simulation, code repair, spreadsheet transformation, or user-authored scripts—and when the expected latency and infrastructure cost fit the product. For agent teams exploring richer development environments, how to build swarm-based IDE agents offers a useful adjacent architecture, but every worker still needs independent permission boundaries.

    FAQ

    Are Python sandboxes completely secure?
    No. They reduce blast radius. Strong isolation, least privilege, network controls, patching, monitoring, and careful application design are still required.

    Is RestrictedPython enough for production?
    Usually not for hostile or multi-tenant code. Treat it as a language-policy layer, not a replacement for process, container, microVM, or managed isolation.

    Should sandboxed agents have internet access?
    Default to no network access. If access is necessary, route requests through an allowlisted proxy with authentication, rate limits, content controls, and detailed logs.

    What is the best starting point for a startup?
    Begin with typed tools where possible. For genuine code execution, use ephemeral non-root workers, read-only images, strict quotas, no default egress, short-lived credentials, and automated cleanup. Move to microVMs or a specialist service as threat exposure and tenancy requirements grow.

    Last updated 24 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.