0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build custom machine learning models github

How to Build Custom Machine Learning Models on GitHub

  1. aigi

    GitHub is more than a place to store model code. Used well, it becomes the operating layer for a machine learning project: requirements are documented, experiments are reproducible, data handling is auditable, tests run automatically, and deployment changes can be reviewed before they reach users.

    This guide shows how to build a custom machine learning model on GitHub in a way that is practical for Indian builders, student teams, startups, and research groups. The examples use Python and scikit-learn, but the workflow also applies to PyTorch, TensorFlow, XGBoost, and domain projects such as Indic language applications.

    Start with a narrow, testable problem

    Do not begin by choosing a library. Begin with a decision the model must improve.

    Define:

    • Input: What data will be available at prediction time?
    • Output: A class, score, ranking, forecast, or generated response.
    • User and action: Who uses the prediction, and what will they do with it?
    • Success metric: Accuracy, F1, recall, mean absolute error, latency, cost, or a business metric.
    • Constraints: Data residency, language coverage, device limits, privacy, and inference budget.

    For example, “predict whether a customer support ticket needs escalation within two hours” is more useful than “build an NLP model.” If the project involves Indian languages, specify the scripts, dialects, code-mixed inputs, and spelling variation you expect. A project working with low-data languages can learn from the practices in this guide to low-resource Indic NLP.

    Create a repository that another builder can run

    A clean repository reduces the gap between a promising notebook and a usable system. Start with a virtual environment and pin dependencies:

    mkdir custom-ml-model
    cd custom-ml-model
    git init
    python -m venv .venv
    source .venv/bin/activate       # Windows: .venv\\Scripts\\activate
    pip install pandas scikit-learn joblib pytest
    pip freeze > requirements.txt

    A practical structure is:

    custom-ml-model/
    ├── data/                  # keep raw data out of Git when sensitive or large
    ├── notebooks/             # exploration only
    ├── src/
    │   ├── data.py
    │   ├── features.py
    │   ├── train.py
    │   └── predict.py
    ├── tests/
    ├── models/                # generated artefacts, usually stored elsewhere
    ├── .gitignore
    ├── README.md
    └── requirements.txt

    Your README should state the problem, dataset source and licence, setup steps, training command, evaluation results, known limitations, and example inference request. Never commit API keys, personal information, customer records, or unlicensed datasets. Use .env files locally and GitHub Actions secrets for CI/CD.

    If you are building a portfolio project, document the trade-offs rather than only displaying a final score. A well-scoped repository can support machine learning portfolio projects for beginners in India and gives reviewers evidence that you understand engineering, not just model fitting.

    Build a reproducible training pipeline

    Separate data preparation, feature creation, training, and evaluation. This prevents training-serving skew, where the model sees one transformation during training and another in production.

    A compact scikit-learn pipeline might look like this:

    from pathlib import Path
    import joblib
    import pandas as pd
    from sklearn.compose import ColumnTransformer
    from sklearn.impute import SimpleImputer
    from sklearn.linear_model import LogisticRegression
    from sklearn.pipeline import Pipeline
    from sklearn.preprocessing import OneHotEncoder, StandardScaler
    
    DATA = Path("data/train.csv")
    TARGET = "approved"
    
    df = pd.read_csv(DATA)
    X = df.drop(columns=[TARGET])
    y = df[TARGET]
    
    numeric = ["income", "age"]
    categorical = ["city", "employment_type"]
    
    preprocess = ColumnTransformer([
        ("num", Pipeline([
            ("impute", SimpleImputer(strategy="median")),
            ("scale", StandardScaler()),
        ]), numeric),
        ("cat", Pipeline([
            ("impute", SimpleImputer(strategy="most_frequent")),
            ("encode", OneHotEncoder(handle_unknown="ignore")),
        ]), categorical),
    ])
    
    model = Pipeline([
        ("preprocess", preprocess),
        ("classifier", LogisticRegression(max_iter=1000)),
    ])
    model.fit(X, y)
    joblib.dump(model, "models/model.joblib")

    Use a fixed random seed for experiments where appropriate, record the dataset version, and log parameters and metrics. Git tracks source code well, but large datasets and model binaries belong in object storage or a data-versioning system. In India, also check whether personal data, consent, retention, and cross-border processing requirements affect your design; consult your organisation’s legal and security teams before using production data.

    Evaluate beyond a single accuracy number

    Create separate training, validation, and test sets. For time-dependent data, split chronologically rather than randomly. For users, patients, merchants, or households, group records so the same entity does not leak into multiple splits.

    Choose metrics that reflect harm and cost:

    • Classification: precision, recall, F1, ROC-AUC, and a confusion matrix.
    • Imbalanced outcomes: precision-recall AUC and recall for the critical class.
    • Regression: MAE, RMSE, and error by important segments.
    • Ranking or recommendations: precision@k, recall@k, and conversion impact.
    • Production: latency, memory, throughput, failure rate, and drift.

    Evaluate by language, geography, device type, gender where lawful and appropriate, and other relevant cohorts. A model that performs well overall may fail on rural users, code-mixed text, or low-bandwidth devices. Record limitations in the README and create a model card describing intended use, out-of-scope use, training data, metrics, and risks.

    Add tests and GitHub Actions

    Machine learning needs tests just like any other software. Add checks for schema columns, missing-value handling, feature dimensions, deterministic preprocessing, prediction shape, and unacceptable metric regressions. Keep a small, versioned fixture dataset for fast tests rather than downloading the full training set on every pull request.

    A basic GitHub Actions workflow can install dependencies and run tests:

    name: tests
    on: [push, pull_request]
    jobs:
      test:
        runs-on: ubuntu-latest
        steps:
          - uses: actions/checkout@v4
          - uses: actions/setup-python@v5
            with:
              python-version: "3.11"
          - run: pip install -r requirements.txt
          - run: pytest

    For larger projects, add linting, data validation, security scanning, and a workflow that trains only when the dataset or training code changes. Do not expose private training data in logs or pull-request artefacts. Use branch protection and require review for changes to inference code, access controls, and evaluation logic.

    Package the model for inference

    A simple API is useful for demos, but production deployment needs validation, authentication, rate limits, structured logging, timeouts, and a rollback plan. FastAPI is a strong option for typed request validation; Flask is adequate for small services.

    from fastapi import FastAPI
    from pydantic import BaseModel
    import joblib
    
    app = FastAPI()
    model = joblib.load("models/model.joblib")
    
    class Request(BaseModel):
        income: float
        age: int
        city: str
        employment_type: str
    
    @app.post("/predict")
    def predict(request: Request):
        result = model.predict([request.model_dump()])[0]
        return {"prediction": int(result)}

    Containerise the service, pin the runtime, and expose a /health endpoint. For a public-facing product, monitor model quality after launch using delayed labels, not just HTTP success rates. Store model versions and make predictions traceable to the code, feature, and data versions that produced them.

    If the model powers a conversational product, treat it as one component in a larger system: retrieval, business rules, safety checks, and observability matter as much as inference. The architecture lessons in how to build a voice agent are relevant when ML predictions are used inside real-time customer interactions.

    A practical GitHub checklist

    Before calling the project complete, verify that:

    • A new contributor can install and run it from the README.
    • No secrets, personal data, or restricted datasets are committed.
    • Training and test data are clearly separated and versioned.
    • Metrics include relevant cohorts and a baseline.
    • The model can be rebuilt from documented commands.
    • Tests run on every pull request.
    • Model artefacts have version labels and rollback instructions.
    • Licence, attribution, security, and known limitations are documented.
    • Monitoring covers drift, latency, errors, and outcome quality.

    GitHub will not make an unreliable model reliable by itself. Its value is disciplined collaboration: clear assumptions, repeatable experiments, reviewable changes, and an honest record of what the model can and cannot do. Build that foundation first, then optimise the algorithm.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.