0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · best practices for machine learning github repositories

Best Practices for Machine Learning GitHub Repositories

  1. aigi

    A machine learning repository should do more than store notebooks. It should explain where data comes from, identify the code and environment used to train a model, make results reproducible, and give another developer a safe path from clone to deployment. That standard matters whether you are building a student portfolio, an open-source project, or an AI product for Indian users.

    The core principle is simple: version the decisions that affect model behaviour. Git tracks source code well, but an ML system also depends on datasets, feature logic, hyperparameters, model artefacts, hardware, secrets, and evaluation procedures. The practices below turn a fragile experiment into a repository that can be audited and maintained.

    Start with a repository structure that reflects the system

    Keep exploratory work separate from reusable application code. A practical structure is:

    project/
    ├── README.md
    ├── pyproject.toml
    ├── uv.lock or poetry.lock
    ├── Dockerfile
    ├── Makefile
    ├── .env.example
    ├── data/
    │   ├── raw/
    │   ├── interim/
    │   └── processed/
    ├── notebooks/
    ├── src/project_name/
    │   ├── data/
    │   ├── features/
    │   ├── training/
    │   ├── inference/
    │   └── evaluation/
    ├── tests/
    ├── configs/
    ├── models/
    └── .github/workflows/

    Do not commit raw datasets, credentials, generated outputs, or large model binaries by default. Add them to .gitignore, then document how an authorised contributor can obtain them. Use predictable notebook names such as 01_data_audit.ipynb and 02_baseline.ipynb; move stable logic into src/ as soon as it is validated.

    If the repository is intended to demonstrate capability, connect it to a wider machine learning portfolio project strategy. A clear README and repeatable setup often communicate more than a long notebook.

    Make the README an operating manual

    A strong README should let a new contributor answer five questions quickly:

    • What problem does the project solve, and for whom?
    • What data and licence does it use?
    • How can I install dependencies and run a baseline?
    • How are models evaluated, and what are the limitations?
    • Where can I find the training, inference, and deployment entry points?

    Include a short architecture diagram, expected hardware, dataset preparation commands, sample input and output, test commands, and known failure modes. State whether metrics are reproducible or approximate. For projects using Indian languages, public-sector data, education records, or health information, document language coverage, geography, consent assumptions, and demographic limitations rather than presenting one aggregate score.

    Add CONTRIBUTING.md, a code of conduct, a licence, and an issue template. A pull request template should ask for the dataset or feature change, evaluation impact, tests added, and any privacy or cost implications.

    Pin environments and separate configuration from code

    Use pyproject.toml with a lockfile where possible. Pin direct dependencies and record the Python version; for GPU workloads, document the CUDA, driver, and framework compatibility matrix. A Dockerfile should reproduce the runtime, but it does not replace local developer instructions.

    Keep configuration in typed YAML, TOML, or structured Python settings rather than scattering constants through notebooks. Store non-secret deployment values in version control, such as image size or batch size. Store tokens, database passwords, and cloud credentials outside Git. Commit .env.example, never .env, and use GitHub Actions secrets or an approved Indian cloud secret manager in CI.

    Run a fresh-clone test before calling a release complete. The minimum path should resemble:

    make setup
    make test
    make train-baseline
    make evaluate

    If the project needs specialised infrastructure, provide a CPU-compatible smoke test so reviewers can validate the pipeline without renting an expensive GPU.

    Version data, models, and lineage

    Git commits alone cannot reproduce a model if the dataset changed between runs. Use DVC, lakeFS, or a managed dataset registry for large files and pipeline stages. Git LFS can work for a small number of fixed model weights, but it is not a substitute for dataset lineage or experiment metadata.

    Every training run should record:

    • Git commit or release tag
    • Dataset version and preprocessing configuration
    • Dependency and hardware details
    • Random seeds and split strategy
    • Hyperparameters and feature set
    • Metrics by relevant segment
    • Model artefact location and checksum

    Avoid relying on mutable paths such as models/latest.pkl. Give releases immutable identifiers and record the promotion decision. For LLM or fine-tuning projects, capture base-model revision, tokenizer, prompt or template version, training examples, safety filters, and evaluation set. The same discipline applies to fine-tuning LLMs on custom data.

    Turn notebooks into tested pipelines

    Notebooks are useful for exploration but conceal execution order, state, and unrecorded manual edits. Keep them short, restartable, and parameterised. Strip outputs before committing with nbstripout, and make notebook execution part of CI only when it is deterministic and reasonably fast.

    Move preprocessing, feature engineering, training, and inference into importable functions. Test them independently:

    • Unit tests: validate transformations, tokenisation, schema handling, and edge cases.
    • Data tests: check columns, types, ranges, null rates, duplicates, and label leakage.
    • Pipeline tests: run a tiny fixture dataset end to end.
    • Model tests: verify output shape, class labels, probability ranges, and a minimum baseline.
    • Regression tests: compare approved metrics or predictions against a fixed evaluation set.

    Use Pandera, Great Expectations, or equivalent checks at ingestion boundaries. Include examples for empty inputs, malformed Unicode, missing Indian-language text, extreme values, and duplicate records. These cases frequently expose production failures earlier than aggregate accuracy does.

    Build CI/CD for ML, not just Python

    A useful GitHub Actions workflow should run on pull requests and protected branches. Begin with formatting, linting, type checks, dependency audits, unit tests, and a small data validation job. Add a CPU smoke-training job that produces a metric and verifies that the model can be loaded for inference.

    Do not retrain a full model on every pull request. Separate workflows into:

    • Pull request checks: fast tests, schema validation, and smoke inference.
    • Scheduled jobs: data refresh, drift checks, and dependency updates.
    • Release workflows: approved training, model evaluation, container build, and deployment.
    • Promotion gates: human review plus thresholds for quality, fairness, latency, and cost.

    Use protected branches, required reviews, signed releases where appropriate, and least-privilege GitHub tokens. Pin third-party Actions to trusted commit SHAs. For teams planning production growth, document the path to scalable machine learning infrastructure, including queues, registries, observability, and rollback.

    Document evaluation, risk, and responsible use

    A model card should cover intended use, out-of-scope use, training data, evaluation methodology, subgroup performance, known limitations, and monitoring requirements. Add a data card describing provenance, consent or licence, retention, transformations, and access controls.

    For Indian deployments, consider language, script, dialect, region, connectivity, accessibility, and representation. Do not publish personal data merely because it is technically accessible. Apply the Digital Personal Data Protection Act obligations relevant to your organisation, minimise retained data, mask identifiers in fixtures, and scan commits with tools such as Gitleaks or GitHub secret scanning.

    Track operational metrics after deployment: latency, error rate, input drift, output distribution, abstention rate, and cost per request. For models used in education, finance, health, hiring, or public services, retain an escalation path to a human reviewer.

    A release checklist for maintainers

    Before merging or publishing a model, confirm:

    • A fresh clone can install and run a baseline.
    • Data and model versions are identified and retrievable.
    • No secrets, personal data, or unlicensed assets are committed.
    • Tests cover preprocessing, schema, inference, and a known baseline.
    • Metrics include relevant segments and a documented evaluation set.
    • The README, model card, licence, and limitations are current.
    • CI passes and the deployment artefact has a rollback path.

    A repository that meets these checks is easier to review, fund, hand over, and operate. It also becomes a credible public record of engineering judgement—especially valuable when open-source contributors or grant reviewers assess an Indian AI project. For practical collaboration guidance, see how to contribute to AI GitHub repositories.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.