Deep learning frameworks make it easy to define a model, call loss.backward(), and start training. That convenience is valuable for production, but it can hide the mechanics that determine whether a model learns, fails silently, or wastes scarce compute. Building deep learning models from scratch on GitHub gives you a controlled way to understand tensors, computational graphs, gradients, optimizers, and training systems before relying on high-level abstractions.
This guide focuses on a useful definition of “from scratch”: implement the learning mechanics yourself with Python and NumPy, while using ordinary tools for file handling, plotting, testing, and experiment tracking. You do not need to write a GPU driver or recreate an entire framework. The goal is a repository that explains the mathematics, produces reproducible results, and can be extended into a serious research or engineering project.
What a strong from-scratch repository should prove
A credible repository should answer four questions clearly:
- Can the model learn? It should fit a tiny synthetic dataset before tackling MNIST, Indic language data, or an image corpus.
- Are the gradients correct? Analytical gradients should agree with numerical finite-difference checks.
- Are experiments reproducible? Seeds, configuration, dataset versions, and environment details should be recorded.
- Is the code understandable? A reader should be able to move from an equation in the README to the corresponding implementation.
This makes the project more valuable than a copied notebook. If you are building a wider portfolio, pair it with other machine learning portfolio projects for beginners in India, but keep this repository narrowly focused on learning mechanics and engineering discipline.
Recommended project structure
Start with a small, modular layout rather than one large notebook:
from-scratch-dl/
├── src/
│ ├── tensor.py
│ ├── layers.py
│ ├── losses.py
│ ├── optimizers.py
│ └── data.py
├── tests/
│ ├── test_gradients.py
│ └── test_layers.py
├── examples/
│ ├── xor.py
│ └── mnist_mlp.py
├── experiments/
├── README.md
├── requirements.txt
└── pyproject.tomlBuild in stages. First implement a scalar automatic differentiation engine, then generalise operations to arrays. A scalar engine makes the chain rule visible; a tensor implementation exposes shape management, broadcasting, memory use, and vectorisation. Andrej Karpathy’s micrograd is a useful reference for the first stage, but your README should explain your own design decisions rather than simply reproducing it.
A practical implementation roadmap
1. Establish the mathematical baseline
Implement a linear layer, activation functions, a loss function, and a training loop before adding convolution or attention. For a dense layer, define the shapes explicitly:
- Input:
Xwith shape(batch_size, input_features) - Weights:
Wwith shape(input_features, output_features) - Bias:
bwith shape(output_features,) - Output:
Y = XW + b
For classification, use softmax cross-entropy. Compute softmax stably by subtracting the maximum logit in each row before exponentiation. Prefer a combined log-softmax and cross-entropy implementation so you do not create unnecessary intermediate values or invite overflow.
Implement ReLU, sigmoid, and tanh with both forward and backward methods. Record the cache required by the backward pass, but clear or replace it after each iteration where appropriate. This teaches an important distinction between parameters, gradients, and temporary activations.
2. Validate every gradient
Gradient checking is not optional. For a parameter θ, estimate its numerical derivative as:
(f(θ + ε) - f(θ - ε)) / (2ε)Compare this with the analytical gradient from your backward pass. Use a small network and a small batch, calculate relative error, and test each layer independently. A mismatch often comes from a transpose error, incorrect broadcasting, an omitted reduction factor, or an activation derivative evaluated at the wrong value.
Write automated tests for:
- Linear-layer forward shapes
- ReLU behaviour at positive and negative inputs
- Loss values on hand-calculated examples
- Gradient agreement within a chosen tolerance
- Parameter updates after one optimizer step
Use continuous integration to run these checks on every pull request. Developers exploring how to contribute to AI GitHub repositories in India will recognise this as a stronger contribution standard than adding an untested demo.
3. Add optimizers deliberately
Begin with batch gradient descent and stochastic gradient descent. Then add momentum and Adam only after the basic update is correct. Document the state each optimizer stores: momentum vectors, second-moment estimates, timestep, learning rate, and numerical stabiliser epsilon.
Expose hyperparameters through a configuration file or command-line arguments. Save the training and validation loss after every epoch, and report learning rate, batch size, parameter count, and runtime. A result without these details is difficult to reproduce or compare.
4. Build a dependable data pipeline
Keep data loading separate from model code. Your pipeline should handle deterministic train-validation-test splits, shuffling, batching, normalisation, and missing or malformed records. For Indian use cases, document language, script, licence, preprocessing, and demographic limitations rather than describing a dataset vaguely as “local data.”
For a first project, use XOR, a two-moons dataset, or MNIST. Move to Devanagari characters, speech features, or Indic text only when the core implementation is stable. If your goal is education technology, a carefully scoped experiment can connect to work on an interactive live learning platform for Indian schools without turning the repository into an unfocused product prototype.
Benchmarking: compare the right things
A NumPy implementation is primarily a learning and reference system. It is not expected to beat PyTorch on a GPU. Benchmark it for correctness, throughput, memory, and convergence—not just wall-clock time.
Compare the same architecture, dataset, batch size, precision, and stopping condition against a framework implementation. Track:
- Examples processed per second
- Peak memory consumption
- Final validation loss and accuracy
- Number of updates to reach a target loss
- Time spent loading data versus computing gradients
Profile before optimising. Replace avoidable Python loops with vectorised operations, but retain a readable reference version when it helps explain the algorithm. If you later target edge devices, measure model size and inference latency separately from training performance.
Extending the project beyond an MLP
Once dense networks are reliable, add one capability at a time:
- Convolution: implement cross-correlation, padding, stride, and backward propagation; verify against a tiny hand-worked example.
- Batch normalisation: test training and inference behaviour separately.
- Embeddings and attention: begin with a small self-attention block and explicit tensor shapes.
- Transformer language modelling: use a tiny corpus and inspect causal masking carefully.
- GPU execution: try CuPy after the NumPy version is correct; treat CUDA or Triton kernels as a later optimisation project.
For computer vision builders, a related guide to building computer vision models on GitHub can help with dataset and repository practices, while this project remains the reference implementation for the underlying training loop.
GitHub practices that improve technical credibility
Your README should include the project scope, equations, setup commands, dataset licence, experiment commands, known limitations, and a results table. Add plots for loss and accuracy, but include the raw metrics in CSV or JSON so others can inspect them. Pin dependencies and provide a small CPU-friendly command that completes in minutes.
Use issues for bugs and design questions, pull requests for changes, and releases for stable milestones. Avoid claiming that a scratch implementation is “production ready” unless it has profiling, error handling, tests, and deployment evidence. A clean commit history showing the progression from scalar autodiff to tensors is itself a useful portfolio signal, especially for Indian student developers building open-source AI.
Common failure modes
- The loss becomes NaN: inspect softmax stability, learning rate, input scale, and gradient magnitudes.
- Accuracy stays at chance: verify labels, output dimensions, gradient accumulation, and parameter updates.
- Training works but validation collapses: check leakage, normalisation, model capacity, and split quality.
- Gradients are always zero: inspect ReLU dead units, accidental integer arrays, detached values, and overwritten caches.
- Results change on every run: seed all relevant random generators and record the execution environment.
The most valuable outcome is not a large model. It is a compact, tested repository that lets another engineer understand why the model learns and where it fails. That foundation transfers directly to efficient inference, custom kernels, research prototypes, and deep-tech products built under Indian compute and data constraints.