Microservices are not simply a way to split a monolith into smaller repositories. They are a set of operational and organisational commitments: independently deployable services, explicit ownership, reliable APIs, isolated data, and automation strong enough to manage distributed failure. Containers make those commitments easier to package and reproduce, but they do not solve poor service boundaries or weak operations.
This guide explains how to build a containerized microservices architecture that can start locally, run predictably in CI, and scale on Kubernetes or a managed Indian cloud environment. It is suitable for product teams building SaaS, fintech, commerce, logistics, or AI-enabled systems.
Start with boundaries, not containers
The first design decision is the service boundary. Begin with business capabilities rather than technical layers. For an e-commerce product, identity, catalogue, orders, payments, inventory, and notifications may be reasonable domains. Avoid creating a service for every database table or CRUD endpoint.
Use these tests before extracting a service:
- Business ownership: Can one team own the service and its roadmap?
- Data ownership: Does it control its own tables and business rules?
- Change cadence: Does it change independently from neighbouring capabilities?
- Failure tolerance: Can the wider product handle it being slow or temporarily unavailable?
- Deployment value: Is independent deployment worth the network and operational overhead?
Use domain events for meaningful state changes, such as OrderPlaced or PaymentCaptured, rather than allowing every service to query another service’s database. If the system is small, a modular monolith may be the better starting point. Extract services when scale, team autonomy, or fault isolation creates a measurable benefit.
Systems that include agents, model inference, or asynchronous workflows need an additional distinction between request-serving services and long-running workers. The principles in building distributed systems with AI agents are useful when queues, retries, tool calls, and partial results become part of the architecture.
Define contracts and communication patterns
Choose communication based on the job, not fashion:
- REST or gRPC: Use synchronous calls for short, user-facing operations with a clear response.
- Events and queues: Use asynchronous messaging for workflows, notifications, retries, and workload smoothing.
- Batch jobs: Use scheduled workers for reports, reconciliation, and data processing.
Publish an API contract using OpenAPI or protobuf and validate it in CI. Apply timeouts to every outbound call. Add idempotency keys to payment, booking, and other state-changing operations so retries do not create duplicate effects.
For events, define ownership, schema versions, retention, and replay behaviour. Consumers should tolerate fields being added and should send failed messages to a dead-letter queue. A simple event envelope can include an event ID, type, version, timestamp, producer, correlation ID, and payload.
An API gateway can handle authentication, routing, request limits, and external API versioning. It should not become a second monolith containing business logic. Keep service-to-service authorisation and validation inside the services as well.
Build secure, reproducible container images
A production Dockerfile should use a supported runtime, a multi-stage build, a non-root user, a minimal final image, and a pinned dependency lockfile. Do not copy secrets into an image or use development servers in production.
FROM node:22-alpine AS build
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
RUN npm run build && npm prune --omit=dev
FROM node:22-alpine
WORKDIR /app
ENV NODE_ENV=production
RUN addgroup -S app && adduser -S app -G app
COPY --from=build --chown=app:app /app ./
USER app
EXPOSE 3000
CMD ["node", "dist/server.js"]Build and test the image locally:
docker build --pull -t registry.example.in/orders:1.0.0 .
docker run --rm -p 3000:3000 registry.example.in/orders:1.0.0Use .dockerignore, scan images for vulnerabilities, generate a software bill of materials, and sign images before deployment. Push immutable tags such as a Git commit SHA rather than repeatedly overwriting latest. For Indian deployments, select a registry and cluster region that meet your latency, residency, and regulatory requirements; the right choice depends on the data classification of each service.
Create a local development environment
Use Docker Compose or an equivalent local tool for the smallest useful slice: the API, database, queue, and dependent mock services. Do not reproduce an entire production cluster on every laptop. Provide seed data, health checks, documented environment variables, and one command to run migrations and tests.
Keep local and production configuration structurally similar while separating secrets. Use environment-specific configuration for endpoints, resource limits, and feature flags. A service should fail clearly at startup when a required configuration value is missing.
Deploy with Kubernetes deliberately
Kubernetes is valuable when you need automated rollouts, service discovery, workload scheduling, and horizontal scaling. It also adds complexity. Start with managed Kubernetes or a simpler container platform unless your team can operate clusters, networking, upgrades, and security.
A basic deployment should specify an immutable image, resource requests and limits, probes, and a dedicated service account:
apiVersion: apps/v1
kind: Deployment
metadata:
name: orders
spec:
replicas: 2
selector:
matchLabels:
app: orders
template:
metadata:
labels:
app: orders
spec:
serviceAccountName: orders
containers:
- name: orders
image: registry.example.in/orders:abc123
ports:
- containerPort: 3000
readinessProbe:
httpGet:
path: /ready
port: 3000
livenessProbe:
httpGet:
path: /health
port: 3000
resources:
requests:
cpu: 100m
memory: 256Mi
limits:
cpu: 500m
memory: 512MiAdd a Service for internal discovery, horizontal pod autoscaling for suitable stateless workloads, and a PodDisruptionBudget for critical services. Use namespaces, network policies, admission controls, and least-privilege RBAC. Store secrets in a managed secret system or sealed-secret workflow, not plain manifests.
Use rolling deployments with automatic rollback conditions. Canary releases are worthwhile for high-risk changes, especially model-serving or payment services. Database migrations must be backwards compatible: deploy additive schema changes first, then switch application behaviour, and remove old fields only after all versions have moved on.
Design data and resilience explicitly
Give each service ownership of its data. Cross-service transactions should use sagas, compensating actions, or an outbox pattern rather than distributed two-phase commits. The outbox pattern writes a business change and its event record in one local transaction; a publisher then delivers the event reliably.
Every dependency needs a failure policy:
- Set bounded timeouts and retry only transient failures.
- Use exponential backoff with jitter and a maximum retry count.
- Add circuit breaking or load shedding for overloaded dependencies.
- Make consumers and commands idempotent.
- Use queues to absorb bursts, but monitor queue age and dead letters.
- Return useful degraded responses where the product can tolerate them.
Test failure with dependency outages, slow databases, expired credentials, full disks, and dropped messages. Resilience is an application property, not something Kubernetes provides automatically.
Build observability into every service
A production service should emit three linked signals:
- Metrics: Request rate, error rate, latency percentiles, saturation, queue depth, and business outcomes.
- Logs: Structured JSON with severity, service, version, request ID, trace ID, and safe business context.
- Traces: Distributed spans across gateways, services, databases, and queues.
Use OpenTelemetry for instrumentation and send telemetry to a controlled backend. Create alerts around user impact and service-level objectives rather than CPU alone. Track cost, especially for clusters, databases, egress, and GPU-backed workloads. Never log access tokens, full payment details, Aadhaar numbers, or other sensitive personal data.
Secure the software supply chain
Protect the path from source code to running workload:
- Require code review and automated tests before merging.
- Scan dependencies, source, container images, and infrastructure manifests.
- Generate SBOMs and sign release artifacts.
- Restrict CI credentials and use short-lived workload identity where possible.
- Apply Kubernetes Pod Security standards and minimal Linux capabilities.
- Rotate secrets and define an incident response process.
- Record audit events for privileged actions and sensitive data access.
For products serving Indian users, map data flows and retention before choosing databases or observability vendors. Security, privacy, and residency requirements should shape architecture at the design stage, not become a release blocker later.
Operate through an automated delivery path
A practical CI/CD pipeline is: lint and unit tests, contract tests, build an immutable image, scan and sign it, run integration tests, deploy to a staging namespace, execute smoke tests, and promote with an approval or policy gate. Infrastructure should be version-controlled and reviewed. GitOps can help teams audit what is running and roll back configuration changes.
Keep a service catalogue with owners, dependencies, runbooks, SLOs, dashboards, data classification, and recovery expectations. Define recovery point and recovery time objectives, then test backups and restoration. A service without an owner or runbook is an operational liability regardless of its code quality.
Common mistakes to avoid
- Splitting a stable monolith before measuring bottlenecks.
- Sharing one database schema across all services.
- Treating containers as virtual machines and running as root.
- Using retries without timeouts, idempotency, or backoff.
- Adding a service mesh before basic networking and observability work.
- Scaling replicas without checking database capacity and downstream limits.
- Relying on logs alone to debug distributed requests.
- Deploying mutable image tags or unreviewed Kubernetes YAML.
A practical adoption sequence
Start with one bounded capability and a thin vertical slice. Containerize it, add health endpoints and telemetry, automate its build, and deploy it to a non-production environment. Next, introduce contract testing, independent data ownership, and a rollback strategy. Only then add more services, asynchronous workflows, autoscaling, or a service mesh.
If your platform will power conversational or voice products, separate latency-sensitive inference from background work and benchmark the full path—not only model latency. Guides such as how to build a voice agent and natural-sounding TTS for voice agents illustrate why streaming, cancellation, and provider failure handling need dedicated boundaries.
The goal is not the maximum number of containers. It is a system in which teams can change one capability safely, understand its behaviour, recover from failure, and control infrastructure cost. Build those properties first; Kubernetes and the rest of the platform should follow the evidence.