Google Kubernetes Engine (GKE) is a strong choice for teams that need Kubernetes-level control over deep learning inference without building a cluster platform from scratch. It supports NVIDIA GPU nodes, Google Cloud networking and observability, managed upgrades, and integrations with storage and identity services.
For an Indian startup, the deployment decision is rarely just technical. GPU availability, quota limits, latency to users, model-loading time, data residency, and per-request cost all affect whether a model is commercially viable. This guide explains how to deploy deep learning models on GKE in a way that is reproducible, observable, and ready for production workloads.
Choose the right GKE architecture
Start by separating workloads instead of placing every pod on one node pool:
- CPU system pool: Runs Kubernetes add-ons, ingress, monitoring, and lightweight APIs.
- GPU inference pool: Runs latency-sensitive model servers with taints and labels that prevent unrelated workloads from consuming GPUs.
- Spot GPU pool: Handles batch jobs, evaluation, embedding generation, and other interruptible work.
- Optional training pool: Uses separate node types, storage, and permissions from online inference.
Choose a regional cluster when availability and resilience matter. For users in India, compare asia-south1 (Mumbai) and asia-south2 (Delhi) against the location of your database, object storage, and customers. Do not select a region based on latency alone: confirm that the GPU type you need is available and that your project has sufficient regional quota.
Autopilot can simplify some operational work, but teams with specialised GPU scheduling, custom DaemonSets, or strict node-pool controls may prefer GKE Standard. Validate current GPU and accelerator support before committing to the architecture; hardware availability and GKE capabilities change over time.
If you are moving from a research prototype into a company, document the model, data, latency target, and operating cost first. The guidance in transitioning from research to a deep tech startup is useful for turning that technical baseline into a product plan.
Create GPU capacity safely
You can create a regional Standard cluster and a dedicated GPU node pool. Keep system services away from expensive accelerators, and apply a taint to the GPU pool so only inference workloads can schedule there.
gcloud container clusters create ai-prod \\
--region asia-south1 \\
--machine-type e2-standard-4 \\
--num-nodes 1 \\
--release-channel regular
gcloud container node-pools create gpu-inference \\
--cluster ai-prod \\
--region asia-south1 \\
--machine-type g2-standard-8 \\
--accelerator type=nvidia-l4,count=1,gpu-driver-version=latest \\
--enable-autoscaling --min-nodes=0 --max-nodes=10 \\
--node-taints workload=gpu:NoSchedule \\
--node-labels workload=gpuThe exact machine and accelerator flags depend on the region and GKE release. Prefer a supported accelerator configuration and let GKE manage drivers where that option is available. Avoid copying an old driver DaemonSet into production without checking compatibility with the node image and Kubernetes version.
Before deployment, verify:
- GPU quota in the selected region and project.
- Availability of the chosen accelerator and machine family.
- IAM permissions for cluster, Artifact Registry, and model storage.
- Private networking requirements and outbound access for model downloads.
- A budget alert and a maximum node-pool size.
A zero GPU quota is a common reason an otherwise correct deployment fails. Request quota early, then test with a small node pool before reserving capacity for a launch.
Package the model and serving runtime
Use a purpose-built inference server rather than exposing a raw notebook or development API. The right choice depends on the model:
- NVIDIA Triton: A flexible option for TensorFlow, PyTorch, ONNX, TensorRT, and multi-model serving.
- vLLM: Well suited to transformer and LLM serving, especially when continuous batching and OpenAI-compatible APIs are useful.
- TorchServe or TensorFlow Serving: Appropriate for teams already invested in those ecosystems, though evaluate their maintenance status and feature fit.
- Custom FastAPI or gRPC service: Useful for pre/post-processing, provided the model server remains isolated from business logic where possible.
For LLM workloads, compare vLLM with other runtimes using your real prompt-length distribution, concurrency, quantisation format, and latency target. Teams deploying agentic systems can also review how to deploy Llama 3 agents and how to deploy open-source AI agents in production.
Build a pinned image and keep model weights outside the image when they are large or updated frequently. Store versioned artefacts in Cloud Storage or an equivalent registry-backed workflow, download them during an explicit startup phase, and expose readiness only after the model is loaded.
FROM vllm/vllm-openai:<pinned-version>
COPY start-server.sh /usr/local/bin/start-server.sh
ENTRYPOINT ["/usr/local/bin/start-server.sh"]Do not use latest for production images. Scan dependencies, generate a software bill of materials where practical, and test the exact image on the target GPU before release.
Deploy with GPU requests and scheduling rules
A pod must request the GPU resource explicitly. Add node affinity, tolerations, probes, and a controlled rollout strategy:
apiVersion: apps/v1
kind: Deployment
metadata:
name: model-server
spec:
replicas: 2
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
selector:
matchLabels:
app: model-server
template:
metadata:
labels:
app: model-server
spec:
nodeSelector:
workload: gpu
tolerations:
- key: workload
operator: Equal
value: gpu
effect: NoSchedule
containers:
- name: server
image: asia-south1-docker.pkg.dev/PROJECT/ai/model:v1.2.0
ports:
- containerPort: 8000
resources:
requests:
cpu: "4"
memory: 16Gi
nvidia.com/gpu: "1"
limits:
cpu: "8"
memory: 32Gi
nvidia.com/gpu: "1"
readinessProbe:
httpGet:
path: /health/ready
port: 8000
initialDelaySeconds: 30
periodSeconds: 10Use a Service and an ingress or gateway for traffic. Add a PodDisruptionBudget, spread replicas across failure domains where possible, and ensure the node pool can actually scale to the number of replicas requested. A second replica does not provide meaningful availability if both pods depend on one node or one unavailable GPU type.
Autoscale on demand, not CPU alone
CPU-based Horizontal Pod Autoscaling is often a poor signal for inference. A model can saturate GPU memory or build a request queue while CPU remains mostly idle. Track request rate, queue depth, tokens per second, GPU duty cycle, GPU memory, and P95/P99 latency.
Use Managed Service for Prometheus or another metrics pipeline, then expose a carefully chosen metric to HPA or use a queue-driven system such as KEDA. Scale the deployment when pods are overloaded and the node pool when pending pods cannot be scheduled. Cluster autoscaling will not help if the project has no quota or if the selected accelerator is unavailable.
Set scale-up thresholds below the point where latency becomes unacceptable. Scale-down conservatively because loading a large model onto a fresh GPU can take minutes. For bursty workloads, keep a small warm capacity; for batch work, use Spot GPUs and tolerate interruptions with checkpoints and retries.
Optimise latency and cost
Benchmark before and after every optimisation. Useful levers include:
- FP16, BF16, or INT8 quantisation where quality permits.
- TensorRT or compiled kernels for supported architectures.
- Dynamic or continuous batching for compatible workloads.
- Request limits and maximum sequence lengths to prevent noisy-neighbour failures.
- Model warm-up during startup to avoid first-request latency.
- Separate deployments for model sizes, languages, or latency tiers.
- Cached model artefacts and regional storage close to the cluster.
Do not assume the largest GPU is the cheapest option. Measure cost per successful request at the concurrency your product actually needs. For computer vision products, serving and preprocessing may dominate the bill; for LLMs, output tokens and context length can dominate. Teams working with visual models may also find how to build computer vision models on GitHub useful when designing a reproducible model pipeline.
Secure and monitor production traffic
Use Workload Identity Federation for GKE instead of embedding service-account keys in images or Kubernetes Secrets. Keep the cluster private where practical, restrict ingress, encrypt sensitive data, and apply least-privilege IAM to model storage. Do not log prompts, images, or personal information by default. Define retention and redaction rules before collecting traces.
Monitor:
- Request success rate and error classes.
- P50, P95, and P99 latency.
- Queue depth, throughput, and batch size.
- GPU utilisation and memory pressure.
- Pod restarts, OOM kills, image-pull failures, and model-load duration.
- Node provisioning time and unschedulable pods.
- Cost by namespace, model, environment, and customer where possible.
Test failure modes deliberately: kill a pod during traffic, revoke model-storage access in staging, exhaust GPU memory, and simulate regional capacity loss. Keep a previous model image available for rollback, and separate model version from application version so either can be reverted independently.
A production checklist
Before exposing the endpoint to customers, confirm that you have:
- A pinned, scanned container image.
- Versioned model weights and a repeatable loading process.
- GPU quota and tested capacity in the selected Indian region.
- Readiness and liveness probes that understand model startup.
- Queue- or latency-based autoscaling, not CPU alone.
- Budget alerts, node limits, and a Spot strategy for interruptible jobs.
- Authentication, rate limits, request-size limits, and audit logging.
- Dashboards, alerts, rollback procedures, and a load-test report.
GKE is most valuable when it becomes a repeatable platform rather than a one-off cluster. Start with one model and one well-defined SLO, automate the image-to-deployment path, and only add specialised GPU pools or multi-model serving after measurements justify the complexity.