0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to deploy deep learning models on cloud platforms

How to Deploy Deep Learning Models on Cloud Platforms

  1. aigi

    A notebook that produces accurate predictions is not yet a product. Production deployment adds constraints that training experiments often hide: predictable latency, repeatable environments, secure data flows, GPU availability, model versioning, and a bill your team can sustain.

    This guide explains how to deploy deep learning models on cloud platforms for real-time APIs, asynchronous jobs, and streaming workloads. The principles apply to PyTorch, TensorFlow, Hugging Face models, computer-vision systems, speech models, and smaller multimodal or language models. For teams building more advanced agent systems, the deployment patterns also complement this guide to deploying Llama 3 agents.

    Start with the workload, not the cloud provider

    Define the service-level requirements before selecting an instance or managed platform. Record:

    • Latency target: p50 and p95 response time, including preprocessing and network overhead.
    • Throughput: requests per second, images or tokens per minute, and expected peak traffic.
    • Payload size: large images, audio files, video frames, or long prompts can dominate costs.
    • Availability: whether occasional interruption is acceptable.
    • Data requirements: residency, encryption, retention, and access controls.
    • Scaling pattern: steady traffic, sharp peaks, scheduled batches, or continuous streams.

    Use online inference for interactive applications such as search, fraud scoring, or a voice agent. Use batch inference for catalogue enrichment, document processing, and recurring analytics; it allows you to schedule cheaper capacity and retry failed jobs. Streaming inference suits camera feeds and sensor data, but requires back-pressure, buffering, and careful GPU utilisation.

    Do not begin with a GPU by default. A quantised vision classifier or small language model may perform well on a CPU. Benchmark the complete request path, then choose hardware based on cost per prediction—not merely raw accelerator speed.

    Prepare and optimise the model

    Before packaging, make the artefact reproducible. Store the model weights, architecture, tokenizer or label map, preprocessing rules, and compatible library versions together. A model that cannot reconstruct the exact training-time preprocessing is not production-ready.

    Common optimisation routes include:

    • Quantisation: Convert FP32 weights to FP16, BF16, or INT8 where accuracy permits. Validate representative Indian-language, regional, and low-quality inputs rather than relying only on aggregate accuracy.
    • Pruning and distillation: Remove redundant computation or train a smaller student model for lower latency.
    • Graph optimisation: Export compatible models to ONNX and test ONNX Runtime or TensorRT for NVIDIA hardware.
    • Dynamic batching: Combine compatible requests to improve accelerator utilisation, while enforcing a maximum queue delay.
    • Caching: Cache embeddings, repeated lookups, and immutable preprocessing outputs when privacy rules allow it.

    Measure cold-start time, model-load time, warm inference latency, memory use, and accuracy after optimisation. For generative systems, track input and output tokens, time to first token, and total generation time. Model quality and operational performance must be evaluated together.

    Package a reproducible inference service

    A container should include the inference code, dependency lockfile, model-serving runtime, health checks, and a clearly defined request schema. Keep large weights outside the image where possible—typically in object storage or a model registry—so application changes do not require rebuilding multi-gigabyte images.

    A robust service should expose:

    • Readiness endpoint: Reports whether weights and accelerator memory are loaded.
    • Liveness endpoint: Confirms the process is responsive without running an expensive prediction.
    • Prediction endpoint: Validates inputs, applies preprocessing, runs inference, and returns a versioned response.
    • Metrics endpoint: Exposes request count, latency, errors, queue depth, memory, and accelerator utilisation.

    Use a CUDA-compatible base image only when the workload needs NVIDIA acceleration. Pin Python and framework versions, scan images for vulnerabilities, run as a non-root user, and keep secrets out of Dockerfiles. Build once and promote the same image through staging and production.

    For teams still strengthening fundamentals, a small machine learning portfolio project can be a useful place to practise this packaging workflow before operating a high-cost endpoint.

    Choose a managed deployment route

    The major clouds offer similar building blocks, but their operational trade-offs differ.

    AWS

    Amazon SageMaker endpoints support managed real-time, asynchronous, serverless, and batch inference. Use a custom container when you need a specialised runtime; use built-in serving options when speed of implementation matters more than fine-grained control. AWS also offers Inferentia and Trainium families for workloads that can be compiled and tested against those accelerators.

    For a broader platform stack, combine ECR for images, S3 for artefacts, IAM for least-privilege access, CloudWatch for logs and metrics, and Application Load Balancer or API Gateway for controlled ingress.

    Google Cloud

    Vertex AI provides model registries, endpoints, batch prediction, evaluation, and monitoring in one managed workflow. Artifact Registry stores containers, while Cloud Storage holds model artefacts. Google Cloud is attractive for teams already using TensorFlow or TPU-compatible workloads, but benchmark TPU migration effort rather than assuming it is economical for every model.

    Microsoft Azure

    Azure Machine Learning supports registries, managed online endpoints, batch endpoints, deployment slots, and monitoring. It fits organisations using Microsoft identity, networking, and data platforms. ONNX Runtime can be especially useful where cross-framework portability is a priority.

    Kubernetes and self-managed serving

    Kubernetes with KServe, NVIDIA Triton, or similar runtimes offers control over networking, scheduling, and multi-model serving. It also transfers more responsibility to your team: cluster upgrades, GPU drivers, autoscaling, node failures, security, and observability. Choose it when workload scale or platform standardisation justifies that operational burden—not because it is fashionable.

    A production deployment workflow

    1. Freeze the artefact: Register the model, code commit, dataset version, and evaluation report.
    2. Create a test contract: Include valid, malformed, boundary, oversized, and adversarial inputs.
    3. Build and scan the image: Pin dependencies and produce a versioned image digest.
    4. Deploy to staging: Use production-like hardware and realistic request distributions.
    5. Load-test the full path: Test concurrency, queueing, model loading, failure recovery, and scale-out.
    6. Release gradually: Use shadow traffic, canary percentages, or blue-green deployment.
    7. Set autoscaling rules: Combine request rate, queue depth, latency, and accelerator memory; GPU utilisation alone can be misleading.
    8. Record rollback conditions: Define accuracy, latency, error-rate, and cost thresholds before launch.

    Keep preprocessing close to inference when network transfer is expensive. For sensitive workloads, use private endpoints, customer-managed encryption where required, network policies, and strict retention controls. Do not send personal or regulated data to a third-party service merely to simplify deployment.

    Control GPU cost and availability

    GPU pricing varies by region, capacity, reservation, and accelerator family. Compare cost per successful prediction across CPU, GPU, inference-specific chips, and quantised versions. Include storage, egress, idle endpoint time, logging, and managed-service fees.

    Use scale-to-zero or serverless options for intermittent traffic only after measuring cold starts. Large models may take longer to download and initialise than your product can tolerate. For batch jobs, interruptible or spot capacity can reduce costs, but design checkpointing, retries, and idempotent jobs around interruption. Keep a small on-demand fallback if deadlines matter.

    In India, benchmark Mumbai, Hyderabad, and Delhi NCR options where available against other regions. Local placement can reduce user latency and simplify data-governance decisions, while overseas regions may offer better accelerator availability or pricing. Put a CDN in front of static assets, not blindly in front of private model weights, and avoid cross-region data movement that quietly increases both latency and cost.

    Monitor quality, reliability, and drift

    Infrastructure dashboards are necessary but insufficient. Track:

    • Request volume, error rate, p50/p95/p99 latency, queue time, and cold starts.
    • GPU memory, utilisation, power, host CPU, and container restarts.
    • Input validation failures, payload sizes, and timeouts.
    • Prediction confidence, abstention rate, and human-review rate.
    • Data drift, segment-level accuracy, and changes in language, geography, or device mix.
    • Cost per request, cost per batch, and idle accelerator hours.

    Log correlation IDs and model versions, but redact prompts, documents, faces, voices, and other sensitive payloads by default. Establish retention and access policies before production traffic arrives. Keep a rollback-ready registry and test rollback regularly.

    For applications that combine models with tools or workflows, deployment must also cover prompt and policy versions, not only weights. The operational principles in this guide to deploying open-source AI agents are relevant when a model becomes one component in a larger system.

    A practical launch checklist

    Before exposing an endpoint to customers, confirm that:

    • The model and preprocessing pipeline are versioned and reproducible.
    • Accuracy has been tested on representative Indian languages, accents, devices, and network conditions where relevant.
    • Authentication, rate limits, encryption, and private networking are configured.
    • Health checks distinguish startup, readiness, and failure states.
    • Autoscaling, timeout, retry, and circuit-breaker behaviour has been load-tested.
    • Dashboards and alerts cover latency, errors, drift, quality, and spend.
    • A canary release and documented rollback have been completed.
    • The team knows who owns incidents and how to pause an unsafe model.

    Cloud deployment is successful when users receive reliable predictions at an acceptable cost—not when a model merely responds from a public URL. Start with a measurable service contract, optimise the artefact, automate promotion, and keep the first production architecture simpler than the research stack.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.