FastAPI is a strong choice for Python teams building APIs that need speed, clear contracts, and a short path from prototype to production. GitHub adds the engineering workflow around the framework: code review, automated testing, dependency checks, release management, and deployment through GitHub Actions.
The important distinction is that FastAPI alone does not make a system scalable. Scalability depends on service boundaries, data ownership, failure handling, infrastructure, and operational discipline. This guide presents a practical architecture for teams building microservices in India and elsewhere, with patterns that remain useful as traffic, contributors, and operational complexity increase.
Decide whether you need microservices
Microservices are useful when parts of a product need to scale, deploy, or evolve independently. They are not automatically better than a modular monolith. Splitting a small application too early creates network calls, duplicated infrastructure, and difficult debugging.
Start with a modular monolith when:
- The product has a small team or limited traffic.
- Domain boundaries are still changing.
- A single database and deployment are operationally simpler.
Consider separate services when a domain has clear ownership, distinct scaling needs, different release cycles, or strong isolation requirements. A payments service, notification worker, and document-processing service often have different performance profiles and are reasonable candidates for separation.
Teams working on more complex workflows should also understand the principles behind building distributed systems with AI agents, especially around retries, state, idempotency, and partial failure.
Design service boundaries and contracts
Define services around business capabilities, not technical layers such as “the database service” or “the utility service.” Each service should own its data and expose a deliberate API. Avoid allowing one service to query another service’s database directly; that creates hidden coupling and makes independent deployment impossible.
For each service, document:
- The responsibility it owns.
- REST or event-based interfaces it provides.
- Request and response schemas.
- Authentication and authorisation rules.
- Timeouts, retry behaviour, and error codes.
- Data it owns and events it publishes.
FastAPI’s type annotations and Pydantic models make API contracts explicit. Use separate models for input, output, and internal database records so implementation details do not leak into public responses. Add an API prefix such as /api/v1 when you need a managed versioning strategy.
Create a production-ready FastAPI service
Use a layout that keeps routing, business logic, configuration, and integrations separate:
service/
├── app/
│ ├── main.py
│ ├── api/routes/
│ ├── core/config.py
│ ├── models/
│ ├── services/
│ └── db/
├── tests/
├── Dockerfile
├── pyproject.toml
└── .github/workflows/ci.ymlA minimal application can begin like this:
from fastapi import FastAPI
app = FastAPI(title="orders-service", version="1.0.0")
@app.get("/healthz", tags=["operations"])
async def healthcheck():
return {"status": "ok"}Use lifespan handlers for controlled startup and shutdown tasks, such as opening a connection pool. Keep request handlers thin: validate input, call a service-layer function, and return a response. Long-running work should move to a queue and worker rather than blocking an HTTP request.
FastAPI’s asynchronous endpoints are valuable for I/O-bound work, but async does not make CPU-heavy code faster. Run image processing, model inference, and large file transformations in workers or separate services. For AI products, this separation is particularly important when combining API traffic with GPU or CPU-intensive workloads; building high-performance AI applications with open-source tools offers useful architectural context.
Handle communication and failure correctly
Use synchronous HTTP for short operations that require an immediate response. Use a broker such as RabbitMQ, Kafka, or a managed queue for events and jobs that can be processed asynchronously. Define event schemas and include a unique event ID, creation time, producer, and schema version.
Production safeguards should include:
- Explicit connection and request timeouts.
- Bounded retries with exponential backoff.
- Idempotency keys for payments, orders, and other repeatable commands.
- Circuit breakers or failure limits for unstable dependencies.
- Dead-letter queues for messages that repeatedly fail.
- Correlation IDs propagated across service calls.
Do not retry every error. A validation failure should return a client error; a temporary database or network failure may be retried. Set deadlines at the gateway and service layers so one slow dependency does not consume every worker.
Organise the GitHub repository
For a small platform, a monorepo is often the most efficient starting point. Keep each service in its own directory, share only carefully maintained libraries, and make every service independently testable and deployable. Separate repositories can work better when teams, permissions, or release schedules differ significantly.
Your repository should include:
- A clear README with local setup and architecture notes.
.env.examplecontaining names, not secret values.- Pinned or locked dependencies.
- Database migration instructions.
- OpenAPI documentation and example requests.
- CODEOWNERS for sensitive directories.
- Issue templates and pull-request checks.
GitHub is also a useful entry point for Indian developers building public portfolios. Contributors can learn practical workflows through how to contribute to AI GitHub repositories in India, including issue selection, focused pull requests, and maintainer communication.
Test the service and its boundaries
Use multiple test levels rather than relying only on endpoint tests:
- Unit tests for business rules and pure functions.
- API tests with FastAPI’s
TestClientor an async HTTP client. - Integration tests against real or containerised databases and brokers.
- Contract tests to detect incompatible changes between consumers and providers.
- Load tests to measure latency, throughput, and resource limits.
A simple endpoint test looks like this:
from fastapi.testclient import TestClient
from app.main import app
client = TestClient(app)
def test_healthcheck():
response = client.get("/healthz")
assert response.status_code == 200
assert response.json() == {"status": "ok"}Run formatting, linting, type checks, security scans, and tests on every pull request. Treat failing checks as merge blockers for production branches.
Build CI/CD with GitHub Actions
A practical workflow should install a supported Python version, cache dependencies, run quality checks, build the container, and publish an immutable image. Deploy only after the image is verified. Tag images with a commit SHA rather than relying only on latest, making rollback deterministic.
Use environment-specific configuration outside the repository. Store secrets in GitHub Actions secrets or, preferably, a cloud secret manager. Use protected environments and required approvals for production. A safe release sequence is:
1. Open a pull request and run all automated checks.
2. Build and scan the container image.
3. Deploy to a staging environment.
4. Run smoke and migration checks.
5. Perform a canary or rolling production deployment.
6. Monitor the release and roll back by image version if required.
Deploy with security and observability built in
Package the service with Docker using a small, non-root base image and a multi-stage build. Run multiple replicas behind a load balancer or Kubernetes service when traffic requires it. Configure readiness and liveness probes separately: readiness determines whether the instance receives traffic, while liveness identifies a process that must be restarted.
Secure the API with TLS, short-lived credentials, least-privilege service identities, input validation, rate limits, and carefully managed CORS. Never log tokens, passwords, personal data, or raw payment details. For Indian products, account for data residency, sector-specific obligations, and retention requirements relevant to your users and industry.
Expose structured logs, metrics, and traces. Track request rate, error rate, latency percentiles, saturation, queue depth, database pool usage, and dependency failures. Prometheus and Grafana are common choices, while OpenTelemetry can standardise traces across services. Define service-level objectives before adding autoscaling; scaling on CPU alone may miss queue backlogs or latency spikes.
A practical production checklist
Before launch, confirm that:
- Every service has an owner, health checks, and documented dependencies.
- APIs have versioned schemas and backward-compatibility rules.
- Timeouts, retries, idempotency, and dead-letter handling are tested.
- CI blocks unsafe merges and produces reproducible artifacts.
- Secrets are externalised and access is audited.
- Database migrations are reviewed and reversible where possible.
- Dashboards and alerts cover user-facing symptoms, not just machine metrics.
- A rollback and incident-response procedure has been rehearsed.
The best FastAPI microservice is not the one with the most endpoints or infrastructure. It is a small, well-owned component with a stable contract, predictable failure behaviour, automated delivery, and enough telemetry for its team to operate it confidently.