0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to deploy ml models on aws lambda india

How to Deploy ML Models on AWS Lambda in India

  1. aigi

    AWS Lambda is a good fit for intermittent, CPU-based ML inference: document classification, fraud scoring, recommendations, OCR post-processing, and lightweight computer vision. It is less suitable for continuously busy endpoints, GPU-dependent models, or large language models that must remain loaded in memory. For Indian teams, the decision should be based on traffic shape, model size, latency targets, and data-handling requirements—not simply on whether Lambda is serverless.

    This guide explains how to deploy ML models on AWS Lambda India, with Mumbai (ap-south-1) as the primary regional example. AWS availability, quotas, pricing, and service features can change, so verify current limits and prices before production launch.

    When Lambda is the right deployment choice

    Lambda works particularly well when:

    • Requests arrive in bursts and the function can scale down between peaks.
    • The model runs on CPU and completes within Lambda’s timeout and memory limits.
    • You want a managed endpoint without operating EC2 instances, Kubernetes, or autoscaling groups.
    • Your application already uses S3, API Gateway, EventBridge, SQS, or DynamoDB.
    • You can tolerate occasional cold starts or budget for Provisioned Concurrency.

    For steady, high-volume inference, compare Lambda with SageMaker real-time inference, ECS, or an EC2 service. For GPU inference, Lambda is not the right target. Teams serving generative or multimodal workloads should first evaluate model size and hardware needs; related deployment patterns are covered in how to deploy Llama 3 agents and how to deploy large language models locally.

    Choose a deployment architecture

    There are three practical patterns.

    Container image: Package the runtime, dependencies, handler, and model inside an image stored in Amazon ECR. Lambda supports images up to 10 GB uncompressed, making this the default for NumPy, scikit-learn, PyTorch CPU, and ONNX Runtime deployments.

    S3 model download: Keep the model in S3 and download it during initialization or on the first invocation. This keeps images smaller, but increases cold-start time and requires careful caching. Use versioned object keys and checksum validation rather than downloading an unpinned latest file.

    Amazon EFS: Mount EFS when several functions need access to large shared model files. EFS does not remove model-loading cost: the function still has to read weights into memory, and VPC networking plus file-system throughput can add latency. Use it selectively for multi-gigabyte assets, not as a default replacement for a compact image.

    For most small and medium models, start with an immutable ECR image. For larger language or vision systems, a managed endpoint or container service may provide a better operational profile. If your workload includes Indian-language vision or text, benchmark the actual model rather than assuming a framework’s package size predicts inference speed; open source vision language models for Indian languages is a useful model-selection reference.

    Build a Lambda-compatible inference image

    Export the model before packaging it. Prefer a portable, versioned format such as ONNX where supported. For scikit-learn, joblib can work, but pin the exact library versions used during training. Avoid loading untrusted pickle files: deserialisation can execute arbitrary code.

    A minimal Dockerfile looks like this:

    FROM public.ecr.aws/lambda/python:3.12
    
    COPY requirements.txt .
    RUN pip install --no-cache-dir -r requirements.txt --target ${LAMBDA_TASK_ROOT}
    
    COPY model.onnx ${LAMBDA_TASK_ROOT}/
    COPY app.py ${LAMBDA_TASK_ROOT}/
    
    CMD ["app.handler"]

    Keep imports and model loading outside the handler so warm invocations reuse the loaded object:

    import json
    import os
    import onnxruntime as ort
    
    MODEL_PATH = os.path.join(os.environ["LAMBDA_TASK_ROOT"], "model.onnx")
    session = ort.InferenceSession(MODEL_PATH, providers=["CPUExecutionProvider"])
    
    def handler(event, context):
        body = json.loads(event.get("body", "{}"))
        prediction = run_inference(session, body)
        return {
            "statusCode": 200,
            "headers": {"content-type": "application/json"},
            "body": json.dumps(prediction),
        }

    Build for the architecture you will run. arm64 can offer strong price-performance for compatible Python wheels, while x86_64 may have broader support for older scientific libraries. Do not switch architectures without rebuilding and testing native dependencies.

    Deploy in the Mumbai region

    Create an ECR repository in ap-south-1, build and push the image, then create a Lambda function from that image:

    aws ecr create-repository \
      --repository-name ml-inference \
      --region ap-south-1
    
    aws ecr get-login-password --region ap-south-1 | \
      docker login --username AWS --password-stdin \
      ACCOUNT_ID.dkr.ecr.ap-south-1.amazonaws.com
    
    docker build --platform linux/amd64 -t ml-inference:1.0.0 .
    docker tag ml-inference:1.0.0 \
      ACCOUNT_ID.dkr.ecr.ap-south-1.amazonaws.com/ml-inference:1.0.0
    docker push \
      ACCOUNT_ID.dkr.ecr.ap-south-1.amazonaws.com/ml-inference:1.0.0

    Use linux/arm64 instead if the function is configured for Arm. Deploy with a specific image tag or digest, not a mutable latest tag. Set memory and timeout through infrastructure as code, using AWS SAM, CDK, or Terraform so the configuration is reviewable and reproducible.

    Expose synchronous predictions through API Gateway or a Lambda Function URL. API Gateway is usually preferable for authentication, throttling, request validation, and consistent API policies. For asynchronous workloads such as batch document scoring, place requests on SQS and let Lambda process them with a dead-letter queue.

    Control cold starts and latency

    Measure three separate timings: container initialisation, model loading, and actual inference. CloudWatch duration alone does not explain every bottleneck.

    • Provisioned Concurrency: Keep a chosen number of execution environments initialized for predictable latency. Test the cost against your peak-hour SLA.
    • Memory sizing: More memory also provides more CPU. Benchmark several settings; the fastest configuration can be cheaper overall if it reduces duration substantially.
    • Model optimisation: Quantise where accuracy permits, remove unused operators, and use ONNX Runtime or another suitable CPU backend.
    • Lazy loading: Load optional assets only when needed, but do not introduce unpredictable first-request latency for a synchronous API.
    • Payload discipline: Pass references to large files in S3 rather than embedding them in API requests. Validate payload size and content type before inference.
    • Concurrency limits: Set reserved concurrency to protect downstream databases and control spend. Do not assume unlimited parallelism is safe for a model endpoint.

    Benchmark with realistic Indian traffic patterns, including mobile-network latency, regional-language text, and peak loads around your product’s actual usage. A model used for computer vision models on GitHub may need a different image preprocessing path from a tabular classifier, so include decoding and preprocessing in the measurement.

    Security, privacy, and operations

    Give the Lambda execution role only the permissions it needs: read access to a specific S3 prefix, decrypt access to the required KMS key, and logging permissions. Store credentials and API keys in Secrets Manager or Systems Manager Parameter Store, not in the image or environment variables committed to source control.

    Use encryption in transit and at rest. If the function accesses private databases, place it in a VPC with appropriate subnets and security groups, then verify that outbound access to ECR, S3, and other services works through VPC endpoints or controlled networking. VPC attachment can affect startup time, so test it rather than treating it as cost-free.

    For personal data, document what enters the endpoint, where logs are retained, who can access predictions, and how deletion requests are handled. Align the design with the Digital Personal Data Protection Act and your organisation’s contractual and sector-specific obligations. Never log raw Aadhaar-like identifiers, health records, or full user prompts by default.

    Add structured logs, request IDs, latency metrics, error rates, throttling alarms, and a model version field to every response or trace. Store model artefacts with immutable versioning and maintain a rollback path. Canary or weighted deployments are safer than replacing the production function in place.

    Cost and architecture checklist

    Before launch, calculate Lambda requests, GB-seconds, Provisioned Concurrency, API Gateway, ECR, S3, CloudWatch, NAT Gateway, and EFS charges. In India, NAT Gateway costs can dominate a small inference service if every invocation reaches the public internet; VPC endpoints and private service access may materially change the bill.

    Choose Lambda when burst economics and operational simplicity matter. Choose a persistent container or managed inference endpoint when the model is always hot, traffic is sustained, startup is expensive, or GPU acceleration is required. Then validate the choice with a load test and a 30-day cost projection—not a single successful demo.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.