Kubernetes can add replicas quickly, but scaling microservices on Kubernetes in India requires more than enabling an autoscaler. Your application must absorb traffic spikes, recover from pod and node failures, protect databases and queues, and keep latency predictable across Indian users and cloud regions.
This guide focuses on the decisions that matter in production: what to scale, which signal to use, how to provision capacity, and how to prove that the system remains reliable as demand grows. If your architecture is still evolving, pair this guide with how to build scalable microservices for AI systems or the more language-specific guide to scalable microservices with Go.
Start with a scaling model
Before changing Kubernetes settings, map each service’s workload and bottleneck. A request-driven API, an asynchronous worker, a GPU inference service, and a stateful database need different scaling strategies.
For every service, document:
- Traffic pattern: steady, seasonal, bursty, or event-driven.
- Scaling unit: requests per second, queue messages, concurrent sessions, tokens, or batch jobs.
- Latency target: especially p95 and p99 latency, not only average response time.
- Dependencies: databases, caches, third-party APIs, Kafka, SQS, or internal services.
- Failure behaviour: what happens when a dependency is slow or unavailable.
- Cost boundary: the maximum sustainable spend per request, tenant, or transaction.
This prevents a common failure mode: scaling an API tier while its database, connection pool, or downstream payment provider remains the real limit.
Configure HPA around useful signals
The Horizontal Pod Autoscaler (HPA) changes replica count according to observed metrics. CPU utilisation is a reasonable starting point for CPU-bound services, but it is often a poor proxy for user demand. A service can have low CPU usage while request latency rises because it is waiting on a database or external API.
A production HPA should usually combine:
- CPU and memory as safety signals.
- Requests per second per pod for stateless HTTP services.
- In-flight requests or concurrency for slow APIs.
- Queue depth and message age for workers.
- GPU utilisation, tokens per second, or inference queue time for AI workloads.
Define realistic resource requests and limits. HPA’s utilisation calculations depend on requests, and inaccurate requests distort scaling decisions. Use the Metrics Server for basic metrics and Prometheus with the Prometheus Adapter for custom metrics.
Set conservative scale-down behaviour to avoid flapping during short traffic dips. Set a suitable behavior policy, stabilisation window, and maximum rate of scale-up. Test these settings with a load generator rather than relying on a quiet staging environment.
Scale event-driven services from the queue
For asynchronous systems, replicas should follow work waiting to be processed rather than CPU alone. KEDA can scale deployments from Kafka lag, RabbitMQ messages, Redis lists, cloud queues, and other event sources.
Track oldest message age as well as queue length. A queue with many small messages may be healthy, while a smaller queue containing long-running jobs may already violate its service-level objective. Set worker concurrency carefully: adding pods does not help if every worker opens too many database connections or calls a rate-limited API.
For AI pipelines, separate ingestion, preprocessing, inference, and post-processing workers. This lets you scale the expensive stage independently and makes it easier to use GPU nodes only where required. Teams building high-volume data workflows may also benefit from optimising Python scripts for large-scale AI data.
Provision nodes before pods become a problem
Pod autoscaling cannot work when the cluster has no room for new replicas. Use the Cluster Autoscaler with managed node groups, or a provider-native provisioner such as Karpenter on AWS where its operational model fits your environment. Choose node pools by workload rather than creating one oversized pool.
A practical Indian production layout may include:
- General-purpose nodes for APIs and control-plane-adjacent workloads.
- Memory-optimised nodes for caches and memory-heavy services.
- Compute- or GPU-optimised nodes for inference and batch processing.
- Spot or preemptible capacity for interruption-tolerant workers.
- On-demand capacity for critical APIs and quorum-based systems.
Use taints, tolerations, node affinity, and topology spread constraints to keep incompatible workloads apart. Define pod disruption budgets, but do not make them so strict that node upgrades or autoscaling become impossible.
For AWS deployments, Mumbai is a common primary region; Hyderabad or another region may support disaster recovery depending on data residency, service availability, and latency requirements. Multi-zone placement within a region improves resilience, but cross-zone traffic and replicated storage can increase costs. Validate the trade-off with measured traffic and failure tests.
Protect traffic during scale-up
New pods need time to start, load configuration, warm caches, and establish connections. Use startup probes for slow initialisation and readiness probes that reflect actual ability to serve requests. A liveness probe should detect a deadlock, not restart a pod merely because a downstream dependency is temporarily slow.
At the ingress and service layers:
- Apply timeouts, retries, and circuit breakers deliberately.
- Avoid retry storms by using bounded retries and jitter.
- Use graceful shutdown and a
preStophook so terminating pods stop receiving work first. - Set connection and keep-alive limits appropriate to the application server.
- Use canary or weighted traffic during releases that alter capacity or latency.
For multi-region products, route users to the nearest healthy region where practical, while keeping writes consistent with the system’s data model. A service mesh can help with traffic policy and telemetry, but it also adds operational complexity. Introduce it for a clear requirement, not as a default.
Treat the database as part of the scaling plan
The application tier is often the easiest layer to scale. Databases and external dependencies are harder.
Use connection pooling, such as PgBouncer for PostgreSQL, and cap connections per pod. Otherwise, an HPA event can multiply connections faster than the database can handle them. Add indexes based on query plans, cache hot reads, and use read replicas only when the application can tolerate replica lag.
For high-volume systems, define ownership boundaries between services and avoid distributed transactions where possible. Use idempotency keys, outbox patterns, and asynchronous workflows for operations that cross service boundaries. Sharding or distributed SQL may eventually be appropriate, but only after query, schema, and workload analysis show that simpler approaches are exhausted.
Build observability around service objectives
Autoscaling needs trustworthy feedback. Instrument every service with metrics, logs, and traces using OpenTelemetry where possible. Monitor:
- Request rate, error rate, and p50/p95/p99 latency.
- HPA recommendations, desired replicas, and scaling events.
- Pending pods, node utilisation, scheduling failures, and evictions.
- Queue depth, message age, worker throughput, and retry counts.
- Database connections, locks, cache hit rate, and replica lag.
- Cost by namespace, workload, environment, and customer where possible.
Create alerts around user impact and saturation, not every transient CPU spike. Distributed tracing is especially valuable when a slow request crosses several Indian payment, identity, or data services.
Control Kubernetes costs in India
Start with right-sized requests based on observed usage, then review them after meaningful traffic changes. Apply namespace quotas and limit ranges so an experimental workload cannot consume the production cluster’s capacity. Use separate node pools for expensive GPUs and schedule batch work during cheaper or less congested periods when your provider’s pricing model allows it.
Spot capacity can reduce costs for stateless workers, but design for interruption: checkpoint jobs, make consumers idempotent, and maintain enough on-demand capacity for recovery. Track egress, cross-zone traffic, logs, storage snapshots, and managed database charges; compute is not the whole bill.
If Kubernetes operational overhead is exceeding the product’s needs, compare it with managed containers or serverless platforms. The right question is not whether Kubernetes is powerful, but whether your team can operate it safely. For broader deployment trade-offs, see how to deploy and scale web applications in India.
A production readiness checklist
Before a major launch or traffic campaign, verify:
- Load tests cover normal traffic, burst traffic, dependency failure, and recovery.
- HPA or event-driven scaling uses metrics tied to user impact.
- Node autoscaling can add capacity before pods time out.
- Readiness, startup, and shutdown behaviour have been tested.
- Database connections remain safe at maximum replica count.
- PDBs, topology spread, backups, and restore procedures are documented.
- Dashboards show latency, errors, saturation, queue age, and cost.
- Runbooks identify who responds to regional, database, and capacity incidents.
Scale gradually, measure the result, and keep a tested rollback path. For AI products with strict throughput or GPU requirements, how to scale enterprise AI applications efficiently provides a useful companion perspective.