0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to deploy scalable ai models

How to Deploy Scalable AI Models in Production

  1. aigi

    A model that performs well in a notebook is not yet a product. Production deployment introduces variable traffic, slow or unreliable networks, GPU scarcity, model updates, privacy obligations, and an infrastructure bill that can grow faster than revenue. For Indian AI teams, the challenge is amplified by multilingual workloads, uneven connectivity, price-sensitive customers, and traffic spikes around campaigns, examinations, public services, or seasonal commerce.

    Knowing how to deploy scalable AI models means designing the complete serving system—not simply placing a model behind a REST endpoint. The goal is predictable latency, controlled cost, safe releases, and a clear path from one customer to many.

    Start with a production contract

    Before choosing Kubernetes or a GPU provider, define what “scalable” means for your product. Record the requirements in a short deployment contract:

    • Latency: Set targets for p50, p95, and p99 latency. A batch document classifier can tolerate seconds; a voice agent cannot.
    • Throughput: Measure requests per second, tokens per second, images per minute, or jobs per hour.
    • Availability: Decide whether the service needs 99.9% uptime or can process work asynchronously.
    • Quality: Track accuracy, groundedness, refusal behaviour, or business outcomes—not just infrastructure metrics.
    • Data constraints: Identify whether prompts, images, audio, or outputs contain personal or regulated information.
    • Budget: Establish a cost ceiling per request, conversation, document, or active user.

    Separate online inference from offline workloads early. Real-time APIs need predictable latency and reserved capacity. Embedding generation, fine-tuning, evaluation, and bulk transcription can run through queues on cheaper spot capacity.

    If you are deploying a voice product, the model endpoint is only one component. Review the architecture alongside telephony infrastructure for scalable voice agents, particularly when media streaming and regional latency affect the user experience.

    Package the model as a repeatable artifact

    A production model should be reproducible from source control, not copied manually from a notebook. Your build pipeline should pin the model version, tokenizer, runtime, system libraries, and configuration. Store large weights in an object store or model registry; build immutable container images that reference a specific artifact checksum.

    A robust request path usually contains:

    1. An API gateway for authentication, rate limits, and request validation.
    2. A preprocessing layer for tokenisation, resizing, normalisation, or language detection.
    3. A model server responsible for batching and inference.
    4. A post-processing layer for formatting, filtering, and policy checks.
    5. A queue or event stream for work that does not need an immediate response.
    6. Metrics, logs, traces, and feedback capture at every stage.

    Keep application logic separate from model serving. This lets you upgrade a model without rewriting billing, permissions, or business workflows. It also makes it easier to route different languages, customers, or workload classes to different models.

    Choose the serving stack by workload

    There is no universal “best” inference server. Match the tool to the model and traffic pattern:

    • GPU vision and multimodel workloads: NVIDIA Triton supports multiple frameworks, dynamic batching, and concurrent model execution.
    • PyTorch services: TorchServe can be suitable for straightforward deployments, although teams should assess its maintenance and ecosystem fit before standardising on it.
    • TensorFlow models: TensorFlow Serving provides model versioning and a mature serving interface.
    • LLM generation: vLLM and Hugging Face TGI support continuous batching and efficient key-value cache management. Evaluate context length, quantisation support, structured output, and tool-calling requirements—not headline throughput alone.
    • CPU-friendly models: ONNX Runtime, OpenVINO, or a carefully optimised native service can be cheaper and simpler for classifiers, ranking models, and small language models.
    • Mobile and edge use cases: ONNX Runtime, TensorFlow Lite, and platform-native runtimes reduce server load and improve resilience when connectivity is poor. See the guide to AI model optimisation for mobile devices before sending every prediction to the cloud.

    For Llama-based applications, serving the model is only part of the problem. Retrieval, tool permissions, prompt versioning, and fallback behaviour matter equally; the Llama 3 agents deployment guide covers these additional layers.

    Optimise before adding machines

    Scaling an inefficient model multiplies waste. Establish a baseline using representative Indian-language text, real image resolutions, realistic prompt lengths, and production concurrency. Then optimise systematically:

    • Use FP16 or BF16 where the hardware and quality profile support it.
    • Test INT8 or INT4 quantisation against a fixed evaluation set; lower memory use is not useful if accuracy or safety declines.
    • Export compatible models to ONNX or TensorRT when kernel and hardware support justify the engineering effort.
    • Apply dynamic batching for traffic that can tolerate a small queueing delay.
    • Use streaming responses for generative interfaces, while enforcing output and token limits.
    • Cache deterministic results, embeddings, and repeated system prompts where privacy permits.
    • Keep payloads small and avoid sending full conversation history when a compact state representation is sufficient.

    Measure end-to-end latency. A fast GPU can still produce a slow product if tokenisation, network transfer, queueing, database reads, or serialisation dominate the request.

    Build autoscaling around the real bottleneck

    CPU utilisation alone is a poor autoscaling signal for AI. A service may show moderate CPU usage while GPU memory is full, the KV cache is exhausted, or requests are waiting in an inference queue. Track and scale using signals such as:

    • Queue depth and queue wait time
    • In-flight requests and tokens per second
    • GPU memory utilisation and compute utilisation
    • Batch size and cache hit rate
    • p95/p99 latency and timeout rate
    • Error rate by model version and customer tier

    Kubernetes is useful when you need multi-service orchestration, node pools, rollout controls, and workload isolation. Use separate pools for CPU services, general GPUs, and specialised accelerators. Apply resource requests and limits, pod disruption budgets, readiness probes, and topology rules so a node failure does not remove every replica.

    For smaller teams, managed container platforms or a direct VM deployment may be more economical. Adopt Kubernetes when its operational benefits exceed the cost of running it—not because it is the default answer.

    Release models without breaking users

    Treat model changes like software releases. Maintain a versioned evaluation set covering accuracy, latency, safety, and representative languages or dialects. Release with one of these patterns:

    • Shadow traffic: Send a copy of production requests to the new model without returning its output.
    • Canary: Route a small percentage of real traffic to the candidate model.
    • Blue-green: Keep the previous and new environments available and switch traffic after validation.
    • Rollback: Make reverting to the last known-good model a one-command or one-click operation.

    Log the model version, prompt or preprocessing version, hardware profile, and relevant request metadata for every prediction. Redact sensitive content and define retention limits. For products involving open-source AI agents, also log tool calls, permissions, retries, and external side effects.

    Monitor quality, safety, and cost

    Infrastructure observability is necessary but insufficient. Create dashboards for availability, latency, throughput, GPU utilisation, queueing, and cost per successful request. Pair them with quality indicators such as human review scores, user corrections, retrieval relevance, hallucination reports, language-specific performance, and drift.

    Set alerts that lead to action: rising p99 latency, falling cache hit rate, abnormal token consumption, a spike in refusals, or quality degradation for one language. Sample prompts and outputs securely for investigation, and never treat raw user data as unrestricted logs.

    For Indian deployments, compare regions and providers using delivered cost and reliability, not advertised hourly price alone. Include egress, storage, idle capacity, support, GPU availability, and data residency requirements. Use spot instances for retryable batch work, but keep critical online inference on capacity you can reliably obtain. Quantisation, batching, and smaller specialist models often save more than switching cloud vendors.

    A practical launch checklist

    Before opening the endpoint to customers, verify that you can:

    • Rebuild and deploy the exact model from source and artifact references.
    • Load-test at expected peak concurrency with realistic inputs.
    • Enforce authentication, quotas, payload limits, and tenant isolation.
    • Scale on queueing and inference metrics, not CPU alone.
    • Roll back a bad model without rebuilding the application.
    • Detect quality, data drift, cost anomalies, and sensitive-data leakage.
    • Run a degraded mode: cached output, smaller model, asynchronous queue, or human review.
    • Estimate cost per transaction at current and projected traffic.

    The right architecture is usually incremental: begin with a reproducible container and a simple managed service, add a dedicated inference server when profiling shows the need, and introduce Kubernetes or multi-region failover when traffic and uptime requirements justify the operational burden. That approach lets Indian teams ship sooner while preserving a credible path to scale.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.