0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to deploy ml models on github

How to Deploy ML Models on GitHub in 2026

  1. aigi

    GitHub is excellent for storing the code, configuration, tests, documentation, and deployment workflow behind a machine-learning application. It is not normally the runtime that serves predictions. GitHub Pages can host static front ends, while a Flask or FastAPI model API needs a separate service such as a cloud VM, container platform, Kubernetes cluster, or managed inference endpoint.

    This distinction prevents a common beginner mistake: pushing model.pkl to a repository and assuming the model is deployed. A practical workflow is to use GitHub as the source of truth, validate every change with GitHub Actions, and deploy the application to infrastructure suited to its latency, traffic, privacy, and GPU requirements.

    Choose the right deployment pattern

    Start by deciding what users need to access:

    • Demo or portfolio project: A GitHub repository plus a lightweight hosted app is usually enough.
    • Prediction API: Package the model behind Flask or FastAPI and deploy the service as a container or web service.
    • Batch inference: Run scheduled jobs that read data, generate predictions, and store results.
    • Static interface: Host HTML, JavaScript, and precomputed outputs on GitHub Pages; do not expect Pages to run Python inference.
    • GPU inference: Use a GPU-backed VM, Kubernetes workload, or managed inference platform. GitHub only coordinates the code and release process.

    For a model involving Indian-language text, speech, or vision, also document the supported languages, scripts, input formats, and known failure cases. Teams exploring larger open models can compare their deployment needs with guides on deploying Llama 3 agents or deploying Mistral-7B on consumer hardware.

    1. Create a reproducible repository

    Create a repository and clone it locally:

    git clone https://github.com/your-user/ml-prediction-api.git
    cd ml-prediction-api
    python -m venv .venv

    Activate the environment and install only the packages your application needs:

    # macOS/Linux
    source .venv/bin/activate
    
    # Windows PowerShell
    .venv\Scripts\Activate.ps1
    
    pip install fastapi uvicorn scikit-learn joblib
    pip freeze > requirements.txt

    Use a clear structure rather than placing everything in the root:

    ml-prediction-api/
    ├── app/
    │   ├── __init__.py
    │   └── main.py
    ├── model/
    │   └── model.joblib
    ├── tests/
    │   └── test_api.py
    ├── .github/workflows/ci.yml
    ├── .gitignore
    ├── Dockerfile
    ├── requirements.txt
    └── README.md

    Add virtual environments, caches, local secrets, datasets, and generated files to .gitignore. Never commit API keys, database passwords, Aadhaar-related data, customer records, or other personal information. For large weights, use Git Large File Storage, an object store, or a model registry instead of ordinary Git history. Private or regulated workloads should use a private repository and a controlled deployment environment.

    2. Package the model and preprocessing together

    A production prediction is only reliable when it applies the same preprocessing used during training. With scikit-learn, save a complete pipeline rather than saving a classifier separately from its encoder or scaler:

    # training/export_model.py
    import joblib
    from sklearn.pipeline import Pipeline
    from sklearn.preprocessing import StandardScaler
    from sklearn.linear_model import LogisticRegression
    
    pipeline = Pipeline([
        ("scale", StandardScaler()),
        ("classifier", LogisticRegression())
    ])
    pipeline.fit(X_train, y_train)
    joblib.dump(pipeline, "model/model.joblib")

    Record the Python version, framework versions, training-data snapshot, feature order, target definition, and evaluation metrics in the README or a model card. Pin dependencies where reproducibility matters, and confirm that the serving environment can load the artifact before pushing it.

    3. Expose a small, validated API

    FastAPI provides typed request validation and useful local documentation. Keep model loading outside the request handler so the artifact is loaded once when the process starts:

    # app/main.py
    import joblib
    from fastapi import FastAPI
    from pydantic import BaseModel
    
    app = FastAPI(title="Prediction API")
    model = joblib.load("model/model.joblib")
    
    class PredictionRequest(BaseModel):
        features: list[float]
    
    @app.get("/health")
    def health():
        return {"status": "ok"}
    
    @app.post("/predict")
    def predict(request: PredictionRequest):
        result = model.predict([request.features])
        return {"prediction": result.tolist()}

    Run it locally:

    uvicorn app.main:app --reload

    Test both valid and invalid inputs. Add authentication, rate limiting, structured logs, request IDs, and HTTPS before exposing an endpoint publicly. Do not return raw stack traces to users. For an image model, validate file size and MIME type; for language models, limit prompt length and log neither sensitive prompts nor personally identifiable information by default.

    4. Test before every deployment

    A repository should prove that the service still works, not merely contain source code. Add tests for:

    • Model loading and artifact compatibility.
    • Request schema validation and useful error responses.
    • A known input with an expected output range.
    • Empty, extreme, malformed, or oversized inputs.
    • The /health endpoint.

    A minimal test might call the API with a fixed fixture and check the response shape. Keep evaluation data separate from training data, and track metrics relevant to the application—not only accuracy. For Indian deployments, check performance across languages, accents, devices, regions, and data quality levels where those differences affect users. If your project is based on computer vision, review the workflow in how to build computer vision models on GitHub.

    5. Automate checks with GitHub Actions

    Create .github/workflows/ci.yml:

    name: test
    on: [push, pull_request]
    
    jobs:
      test:
        runs-on: ubuntu-latest
        steps:
          - uses: actions/checkout@v4
          - uses: actions/setup-python@v5
            with:
              python-version: "3.11"
          - run: pip install -r requirements.txt
          - run: pip install pytest
          - run: pytest -q

    Use pull requests for changes to model code, preprocessing, and infrastructure. A stronger pipeline can lint code, scan dependencies, build a container, run smoke tests, and publish an image only after tests pass. Store deployment credentials in GitHub Actions secrets or, preferably, use short-lived cloud identity federation. Never place secrets in YAML, notebooks, README files, or model metadata.

    6. Deploy the application, not just the repository

    Containerise the service so local, staging, and production environments use the same process:

    FROM python:3.11-slim
    WORKDIR /service
    COPY requirements.txt .
    RUN pip install --no-cache-dir -r requirements.txt
    COPY app ./app
    COPY model ./model
    CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8080"]

    Connect the repository to a container host, VM, serverless service, or Kubernetes platform. Configure the platform’s start command, port, CPU or GPU limits, environment variables, autoscaling, and health checks. For larger workloads, a deep-learning deployment on GKE offers more control, but it also requires operational expertise.

    Use separate staging and production environments. Deploy an immutable version—such as a Git commit SHA or container digest—rather than pulling an unpinned main branch at runtime. Run a smoke test after deployment, monitor latency and error rates, and keep a rollback path to the previous model version.

    Common mistakes to avoid

    • Treating GitHub Pages as a Python inference server.
    • Committing multi-gigabyte weights into normal Git history.
    • Saving preprocessing separately from the model.
    • Using pickle files from untrusted sources.
    • Running Flask or Uvicorn in debug mode in production.
    • Shipping an endpoint without authentication, validation, or rate limits.
    • Reporting only offline accuracy while ignoring drift, latency, and subgroup performance.
    • Releasing a model without documenting its licence, training data, limitations, and intended use.

    Deployment checklist

    Before sharing the URL, confirm that you have:

    • A README with setup, API examples, licence, model card, and limitations.
    • A pinned or deliberately managed dependency set.
    • Tests running on every pull request.
    • No secrets or sensitive data in Git history.
    • A versioned model artifact and reproducible release.
    • Health checks, logs, monitoring, and rollback instructions.
    • A privacy and security review appropriate to the data.

    GitHub makes ML deployment repeatable when it is treated as an engineering control plane rather than a hosting shortcut. Keep the repository reproducible, automate quality gates, and run inference on infrastructure matched to the model and its users. Builders working on open source can also learn how to contribute to AI GitHub repositories in India to improve documentation, testing, and deployment practices across the ecosystem.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.