Rust is a strong fit for small cloud services: a compiled binary starts quickly, uses little memory, and can handle concurrent requests without a garbage collector. That does not make cloud deployment automatically free. Egress, public IPv4 addresses, managed databases, storage, build minutes, logs, and idle resources can all create charges.
The practical goal is to design a small, stateless service that stays inside documented free allowances, then add billing alerts and a shutdown plan. Free tiers change, quotas vary by region, and some offers require a payment card. Confirm current pricing before deploying production traffic.
Choose the right free deployment model
Your best option depends on traffic, persistence, latency, and how much infrastructure you want to operate.
- Cloud Run: Best for HTTP services that can scale to zero. It reduces server maintenance and works well with Rust containers, but requests, CPU allocation, networking, and outbound traffic still have limits.
- Fly.io: Convenient for containerised services and regional placement. Check the current promotional terms rather than assuming an old “free tier” is permanent. Mumbai availability, included resources, and billing rules can change.
- Oracle Cloud Infrastructure: OCI’s Always Free Ampere A1 resources can be useful for a small self-managed cluster, including in an Indian region when capacity is available. You manage patching, firewalls, backups, and reliability yourself.
- A small VM: A free eligible VM can host several services with Docker Compose or a lightweight reverse proxy. This is inexpensive but creates an operations burden and a single point of failure.
For an API with irregular traffic, start with Cloud Run or an equivalent scale-to-zero platform. For several always-on internal services, an ARM64 VM may offer better value. If the workload includes AI inference, separate the API from the model-serving process; deployment patterns for open-source AI agents in production have different memory and GPU requirements.
Design the service before writing the Dockerfile
Keep the first deployment deliberately boring:
- One stateless HTTP service
- One health endpoint, such as
/healthz - Configuration supplied through environment variables
- No files written to the container filesystem except temporary data
- A managed or separately hosted database
- Structured logs sent to standard output
- A clear timeout for every outbound request
Use Axum for a straightforward Tokio-based service, Actix Web when you already know its ecosystem, or another framework only when it solves a specific requirement. Avoid putting a queue, scheduler, database, and API into one container merely to reduce the number of deployments. That makes failures and scaling harder to diagnose.
For an India-focused application, select the nearest supported region—often Mumbai, Hyderabad, or Delhi NCR where available—but measure actual latency from your users. Regional placement also affects data residency, database round trips, and egress charges. If your service is part of an AI product, review low-latency AI model deployment before choosing a region around the web tier alone.
Build a small, reproducible container
Never compile Rust on a 256 MB or similarly constrained production VM. Build in CI or a sufficiently large local environment, then copy only the release binary into the runtime image.
FROM rust:1.85-bookworm AS builder
WORKDIR /app
COPY Cargo.toml Cargo.lock ./
COPY src ./src
RUN cargo build --release
FROM debian:bookworm-slim
RUN useradd --system --uid 10001 app
WORKDIR /app
COPY --from=builder /app/target/release/orders-api /app/orders-api
USER 10001
EXPOSE 8080
ENTRYPOINT ["/app/orders-api"]Use a pinned Rust toolchain and commit Cargo.lock for an application. A dependency-only build layer improves CI cache hits. For smaller images, a musl build or a distroless-compatible runtime can work, but test TLS, DNS, certificates, and native dependencies before switching. Image size is useful; peak memory, startup time, and outbound traffic matter more to the bill.
A sensible release profile is:
[profile.release]
lto = "thin"
codegen-units = 1
strip = "symbols"
panic = "abort"Measure before enabling aggressive settings. Full LTO can make CI slow, while allocator changes such as mimalloc are not automatically improvements. Use a load test to compare resident memory and latency with your real dependency set.
Deploy and configure safely
For a container platform, make the application listen on 0.0.0.0, read the platform-provided port, and terminate gracefully on SIGTERM. Set conservative limits:
- Request timeout: 15–60 seconds, depending on the endpoint
- Maximum request body size
- Connection pool size based on the database quota
- One or two workers initially
- CPU and memory limits with headroom for bursts
- Minimum instances at zero unless latency requirements justify otherwise
Store secrets in the provider’s secret manager or encrypted CI variables—not in fly.toml, Terraform state committed to Git, Docker layers, or logs. Add authentication and rate limits before exposing an administrative endpoint. A free deployment still needs HTTPS, dependency updates, least-privilege IAM, and a firewall policy.
If you use OCI, restrict ingress to required ports, disable password SSH, use key-based access, enable unattended security updates where appropriate, and keep backups outside the VM. Treat the Always Free machine as a useful development or low-risk production host, not as a replacement for redundancy.
Add a database without losing control of cost
Do not run a production database on ephemeral container storage. For prototypes, hosted Postgres, serverless SQLite, or a small external database can be practical, but assess sleeping, storage caps, connection limits, backups, and deletion policies. Use SQLx or another driver with a small pool and migrations that run as an explicit release step.
Keep object files out of the application container. Use an object-storage service only after checking request, storage, and egress pricing; “free” storage can still become billable through downloads. For bookkeeping or other small-business workloads, separate transaction data from uploaded documents, as the architecture in cloud-based bookkeeping for small shops in India illustrates.
Automate checks and deployment
A minimal GitHub Actions pipeline should run:
1. cargo fmt --check
2. cargo clippy --all-targets --all-features -- -D warnings
3. Unit and integration tests
4. A release build and container vulnerability scan
5. Deployment only from a protected branch or signed release
Cache Cargo’s registry and target directories, but invalidate caches when the toolchain or lockfile changes. Build for the target architecture of the host. ARM64 is often attractive on OCI, but an x86 build will not run there; use Docker Buildx or a native ARM runner and test the resulting image.
For AI-heavy systems, keep model downloads out of every deployment. Package only the API, fetch approved artefacts during a controlled startup process, or use a dedicated inference service. See how to deploy ML models on AWS Lambda in India for a useful contrast between compact request handlers and model-serving workloads.
Monitor the free-tier boundary
Before launch, record every allowance: compute hours, memory, requests, build minutes, storage, database capacity, logs, IP addresses, and egress. Set budget alerts at low thresholds and review usage weekly. Disable preview environments after testing, delete unattached volumes, cap log retention, and avoid debug logging in production.
Track p50 and p95 latency, error rate, restart count, resident memory, CPU time, database connections, and outbound bytes. Add a synthetic health check, but keep its frequency modest. A service that scales to zero may show cold starts; decide whether that is acceptable instead of paying for an always-on instance by default.
A practical launch checklist
- Confirm the provider’s 2026 terms, region support, and payment requirements.
- Build and scan an image outside the production VM.
- Test startup, shutdown, health checks, timeouts, and oversized requests.
- Use non-root containers and least-privilege credentials.
- Externalise database and object storage with backups.
- Add CI checks, deployment rollback, budget alerts, and log retention limits.
- Run a small load test and document the point at which the free tier ends.
Free cloud is best treated as a constrained engineering environment, not a permanent promise. A lean Rust service can run comfortably within those constraints, but reliability comes from measurement, security, and an exit plan—not from the word “free” in a pricing table.