Serverless deployment is useful when an AI endpoint has uneven traffic, a small model, or a clear request-response pattern. GitHub should not be treated as the place that runs your model: it stores source code, configuration, tests, and deployment workflows, while a cloud platform provides the runtime. That separation makes releases auditable and repeatable.
For Indian teams, serverless can be a sensible way to launch an MVP without committing early to a dedicated GPU or Kubernetes cluster. It is less suitable for large language models, long-running inference, streaming generation, or workloads that require consistently warm accelerators. Start with the workload, not the hosting trend.
Choose the right serverless shape
A typical architecture has five parts:
- GitHub repository: application code, infrastructure configuration, tests, and documentation.
- GitHub Actions: automated tests, packaging, security checks, and deployment.
- Serverless compute: AWS Lambda, Google Cloud Run functions, Azure Functions, or an equivalent managed runtime.
- Object storage or a model registry: versioned model artefacts, rather than large binary files committed to Git.
- API gateway and observability: authentication, rate limits, logs, metrics, and traces.
Use a function when inference completes within the platform’s execution and memory limits. For larger containers or Python dependencies, a serverless container service is often easier than forcing everything into a ZIP package. If the model is a voice or multimodal system, review the architecture used in this voice agent deployment guide before choosing a request-response function.
Prepare the model for inference
Training code and inference code should be separate. The production package should load a tested model and expose a small, deterministic prediction path.
Before deployment:
- Export the model to a portable format such as ONNX, TorchScript, TensorFlow SavedModel, or a framework-native format supported by the runtime.
- Pin dependency versions in
requirements.txt,package-lock.json, or an equivalent lockfile. - Keep model weights in object storage or a registry when they are too large for the deployment package.
- Add input validation, a model version, and a health-check response.
- Measure package size, memory use, cold-start time, and inference latency with production-like inputs.
- Remove training libraries, notebooks, datasets, and unused development dependencies.
For image or edge workloads, quantisation and smaller architectures can reduce both latency and cost. The AI model optimisation guide for mobile devices covers techniques that are also relevant to serverless inference, including reduced precision and model-size trade-offs.
A practical repository layout is:
ai-service/
├── src/
│ ├── handler.py
│ ├── inference.py
│ └── schemas.py
├── tests/
├── models/
│ └── README.md
├── infrastructure/
├── requirements.txt
├── Dockerfile
└── .github/workflows/deploy.ymlDo not place API keys, cloud credentials, or private model weights in GitHub. Use repository or environment secrets, and prefer short-lived cloud identity federation from GitHub Actions over long-lived access keys.
Build a small, production-ready handler
The handler should parse the request, validate its schema, load the model efficiently, run inference, and return a bounded response. Load the model outside the request function where the runtime permits it; warm invocations can then reuse the in-memory object.
import json
from inference import predict
def handler(event, context):
try:
body = json.loads(event.get("body") or "{}")
values = body.get("values")
if not isinstance(values, list) or not values:
return {"statusCode": 400, "body": json.dumps({"error": "values is required"})}
result = predict(values)
return {
"statusCode": 200,
"headers": {"content-type": "application/json"},
"body": json.dumps({"model_version": "2026-01", "prediction": result}),
}
except Exception:
return {"statusCode": 500, "body": json.dumps({"error": "inference failed"})}Use structured logs without recording sensitive prompts, images, health information, or personally identifiable information. For Indian deployments, confirm where request data and logs are stored, especially when serving financial, health, education, or government users. Data residency, retention, and consent requirements should be part of the design review rather than an afterthought.
Configure the cloud deployment
You can use the Serverless Framework, Terraform, AWS SAM, native cloud templates, or a managed container configuration. The exact syntax varies, but every deployment should define:
- Runtime and memory allocation
- Timeout and concurrency limits
- Environment variables
- API route and authentication
- Model location and immutable model version
- Network permissions and least-privilege access
- Log retention and alarms
Avoid unpinned runtimes and outdated examples. For example, do not copy a Node.js 12 configuration into a new 2026 project. Select a currently supported Python or Node.js runtime from your provider, then test upgrades in a staging environment.
If the model is too large for a function, package the inference server as a container and deploy it to a serverless container platform. If you are already operating Google Cloud infrastructure, compare this design with deploying deep learning models on GKE; Kubernetes gives more control, but it also creates more operational work.
Add GitHub Actions safely
A deployment workflow should test before it publishes. A minimal pattern is:
name: Deploy inference service
on:
push:
branches: [main]
workflow_dispatch:
permissions:
contents: read
id-token: write
jobs:
deploy:
runs-on: ubuntu-latest
environment: production
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'
- run: pip install -r requirements.txt
- run: pytest
- run: python scripts/check_model.py
- name: Deploy
run: ./scripts/deploy.shConfigure workload identity federation with your cloud provider so the workflow receives temporary credentials. Protect the production environment with reviewers, restrict deployments to main, and separate staging from production. Add dependency scanning, secret scanning, an infrastructure plan or preview, and an explicit rollback procedure.
GitHub is also a useful place to learn deployment patterns. Contributors exploring reusable inference services can start with AI GitHub repositories in India, while teams deploying open models should account for model licences, attribution, and acceptable-use restrictions. Those checks matter particularly when packaging open-source AI agents for production.
Test and operate the endpoint
Before exposing the URL publicly, test:
- Valid, missing, oversized, and malformed inputs
- Authentication failures and rate limits
- Cold and warm latency at expected concurrency
- Model output against a fixed evaluation set
- Timeouts, retries, and duplicate requests
- Rollback to the previous model version
- Logs and alerts without sensitive data leakage
Track p50, p95, and p99 latency, error rate, throttling, invocation count, memory use, and estimated cost. Set budgets and alerts from the first production release. A serverless bill can rise quickly if an endpoint is public, unthrottled, or repeatedly downloads model files.
For generative systems, also monitor token usage, response length, refusal behaviour, prompt-injection attempts, and factual quality. Large models such as Llama 3 may need dedicated inference infrastructure; use this Llama 3 production deployment guide to compare alternatives before placing one behind a function.
When serverless is the wrong choice
Choose a managed GPU endpoint, long-running container, or Kubernetes when you need large model weights, GPU acceleration, streaming responses, high sustained throughput, websockets, or predictable low latency. Serverless remains strong for lightweight classifiers, embeddings, OCR post-processing, moderation, scheduled batch jobs, and bursty APIs.
The best deployment is the smallest reliable system that meets your latency, privacy, reliability, and cost requirements. Use GitHub to make that system reproducible; use the cloud runtime that matches the model rather than forcing every AI workload into a function.