0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to deploy scalable llm apps on github

How to Deploy Scalable LLM Apps on GitHub

  1. aigi

    GitHub is the control plane for an LLM application, not the place where production inference runs. Use the repository to manage source code, tests, infrastructure definitions, container images, and release approvals; run the application on a cloud or specialised inference platform that can supply CPU, GPU, storage, networking, and observability.

    That distinction matters. A production LLM system may combine a FastAPI or Node.js API, a hosted model provider or vLLM server, a relational database, a vector store, queues, object storage, and a frontend. GitHub Actions connects these components into a repeatable delivery process. This guide shows a practical path from repository to production, with choices suited to Indian startups and teams serving users in India.

    Choose the deployment model first

    Your model strategy determines almost every infrastructure decision:

    • Hosted APIs: Use OpenAI, Anthropic, Google, or another provider when speed, reliability, and a small operations team matter most. Your app usually needs CPU hosting only.
    • Managed open models: Use a provider that hosts an open-weight model when you need model choice without operating GPUs.
    • Self-hosted inference: Run vLLM, Hugging Face TGI, or another serving stack when traffic volume, data residency, custom fine-tuning, or predictable unit economics justify GPU operations.
    • Hybrid routing: Send routine requests to a lower-cost model and escalate complex tasks to a larger model. Keep routing logic in the application rather than hard-coding it into deployment scripts.

    Teams deploying Llama-family models should also review how to deploy Llama 3 agents, especially when tool use, memory, and multi-step workflows are involved.

    Define a service-level target before choosing infrastructure. Record expected requests per second, maximum prompt and output tokens, acceptable time to first token, p95 total latency, availability, and regional requirements. Without these numbers, “scalable” usually becomes an expensive collection of guesses.

    Structure the GitHub repository for release safety

    A useful repository separates product code from deployment configuration:

    app/
      api/                 # HTTP routes and authentication
      llm/                 # providers, prompts, routing, fallbacks
      retrieval/           # embedding and vector-search logic
    tests/
      unit/
      integration/
    infrastructure/
      terraform/
      kubernetes/
    .github/workflows/
      test.yml
      build.yml
      deploy-staging.yml
      deploy-production.yml
    Dockerfile
    pyproject.toml

    Keep prompts, model identifiers, retrieval settings, and safety policies versioned. Test structured outputs, tool-call validation, prompt-injection handling, and provider fallback behaviour. For Python web services that call external models, integrating LLM APIs in Python web apps provides a useful application-layer reference.

    Do not commit API keys, model weights, .env files, customer documents, or production logs. Use GitHub secret scanning and push protection, but treat them as safeguards rather than a substitute for disciplined access control.

    Build a production-ready container

    Use a multi-stage Docker build where possible. Pin dependency versions, run as a non-root user, add a health endpoint, and avoid downloading model weights during every image build. For self-hosted inference, create a separate inference image from the API image so application releases do not force a multi-gigabyte model layer to rebuild.

    FROM python:3.12-slim AS runtime
    
    ENV PYTHONDONTWRITEBYTECODE=1 \
        PYTHONUNBUFFERED=1
    WORKDIR /app
    
    RUN useradd --create-home appuser
    COPY pyproject.toml uv.lock ./
    RUN pip install --no-cache-dir uv && uv sync --frozen --no-dev
    
    COPY app ./app
    USER appuser
    EXPOSE 8000
    
    CMD ["uv", "run", "uvicorn", "app.api:app", "--host", "0.0.0.0", "--port", "8000"]

    Publish immutable image tags to GitHub Container Registry (GHCR), such as the commit SHA. A latest tag is convenient for experiments but unsafe for production because it makes rollbacks ambiguous. Sign images where your platform supports it and scan them for vulnerabilities before deployment.

    Create a GitHub Actions pipeline

    A dependable pipeline should test first, build once, and promote the same image across environments. Use least-privilege permissions and modern, pinned action versions.

    name: release
    
    on:
      push:
        branches: [main]
      pull_request:
    
    permissions:
      contents: read
    
    jobs:
      test:
        runs-on: ubuntu-latest
        steps:
          - uses: actions/checkout@v4
          - uses: actions/setup-python@v5
            with:
              python-version: "3.12"
          - run: pip install -r requirements-dev.txt
          - run: pytest -q
    
      image:
        needs: test
        if: github.event_name == 'push'
        runs-on: ubuntu-latest
        permissions:
          contents: read
          packages: write
        steps:
          - uses: actions/checkout@v4
          - uses: docker/login-action@v3
            with:
              registry: ghcr.io
              username: ${{ github.actor }}
              password: ${{ secrets.GITHUB_TOKEN }}
          - uses: docker/build-push-action@v6
            with:
              push: true
              tags: ghcr.io/${{ github.repository }}:${{ github.sha }}

    Add integration tests with mocked model responses, then run a small number of live-provider tests separately. Live tests can be slow, costly, and nondeterministic. Store deployment credentials in GitHub Environments with required reviewers, environment-specific secrets, and branch restrictions. Prefer OIDC federation to long-lived cloud access keys.

    Deploy to the right runtime

    For an API that calls hosted models, Cloud Run, Azure Container Apps, ECS, or a similar managed container service can provide autoscaling with limited operations. Configure minimum instances if cold-start latency affects your user experience, and set maximum instances to prevent an accidental cost surge.

    For self-hosted GPU inference, use Kubernetes, a managed GPU service, or a specialist platform. Autoscaling must account for GPU memory and model loading time, not just CPU utilisation. Keep a warm replica for latency-sensitive workloads, use continuous batching where supported, and separate interactive requests from bulk jobs.

    A queue-based design is often better for document ingestion, evaluation, report generation, and other long-running tasks. Return a job ID from the API, process work through workers, and expose status through polling or webhooks. For a broader infrastructure view, see scalable machine learning infrastructure for developers.

    Use Terraform or another infrastructure-as-code tool to define networks, registries, databases, IAM, queues, and services. GitHub Actions can plan changes on pull requests and apply them only after approval. For Kubernetes, GitOps tools such as Argo CD or Flux can reconcile a reviewed manifest rather than granting every CI job direct cluster-admin access.

    Design for reliability and cost

    LLM traffic is bursty and token-heavy. Apply controls at several layers:

    • Set request, token, concurrency, and timeout limits.
    • Retry only transient failures, with exponential backoff and jitter.
    • Use idempotency keys for jobs that may be retried.
    • Stream responses where it improves perceived latency.
    • Cache embeddings and safe, repeatable completions; never cache private responses across users.
    • Route simple requests to smaller models and cap maximum output tokens.
    • Record cost per request, model, tenant, and workflow.

    For GPU deployments, monitor utilisation, VRAM, queue depth, time to first token, tokens per second, and model-load time. For API-based deployments, monitor provider errors, rate-limit responses, token usage, and fallback frequency. A scalable system is one whose cost and performance remain explainable as traffic grows.

    Secure data and regional traffic

    Treat prompts, retrieved documents, and model outputs as potentially sensitive. Redact personal data before sending it to external providers where required, encrypt data in transit and at rest, and enforce tenant-level access checks before retrieval. Keep audit logs free of raw secrets and unnecessary personal information.

    For users in India, placing application and data services in Mumbai or Hyderabad can reduce network latency and simplify regional requirements, but test actual provider availability and egress costs. Co-locate the vector database with the retrieval service, and use a CDN for static frontend assets. If you need a serverless GPU workflow, compare specialist options in building serverless AI apps with Modal.

    Release, observe, and roll back

    Deploy to staging on every protected-branch merge. Run smoke tests against the deployed endpoint, including authentication, retrieval, structured output validation, and streaming. Promote to production through an approval gate, then use canary or blue-green deployment for model and prompt changes.

    Track application metrics alongside quality metrics. Useful signals include groundedness, citation accuracy, refusal correctness, tool-call success, user feedback, latency percentiles, and cost per successful task. Store a small, privacy-reviewed evaluation set in the repository or a controlled evaluation service, and run it before changing models or prompts.

    Every release should have a one-command rollback to the previous image digest and configuration. Roll back application code independently from model configuration when possible. This separation makes it safer to test a new model without changing the API surface.

    Pre-launch checklist

    • Repository contains reproducible tests and pinned dependencies.
    • Images are immutable, scanned, and stored in GHCR.
    • Secrets use GitHub Environments or cloud secret management, not source files.
    • CI builds once and promotes the tested image.
    • Autoscaling, concurrency, timeouts, and maximum spend are configured.
    • Logs exclude sensitive prompts and credentials.
    • Dashboards cover latency, errors, tokens, GPU health, quality, and cost.
    • Staging smoke tests and production rollback have been rehearsed.

    GitHub can provide an excellent delivery foundation for LLM products, but scalability comes from the complete system: model serving, queues, data locality, security, observability, and disciplined release management. Start with a managed model and container platform when uncertainty is high; move to dedicated GPUs only when measured traffic, latency, or data requirements make the trade-off worthwhile.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.