0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · building deep learning models from scratch in python

Building Deep Learning Models from Scratch in Python

  1. aigi

    Deep learning frameworks can hide the mathematics that makes a model learn. Implementing a small network yourself is one of the fastest ways to understand tensors, gradients, optimisation, and failure modes before you build production systems. This guide uses Python and NumPy to create a binary classifier, then shows how to validate the implementation and transition to modern frameworks.

    The goal is not to replace PyTorch or TensorFlow. It is to build a mental model that helps you debug them. If you are assembling a broader beginner portfolio, pair this project with machine learning portfolio projects for beginners in India and publish the code, experiments, and results rather than only a notebook screenshot.

    What “from scratch” should mean

    There are two useful interpretations:

    • Educational from scratch: use Python and NumPy, but rely on array operations for matrix multiplication and numerical work.
    • Production from scratch: implement kernels, GPU support, distributed training, and deployment infrastructure yourself. This is rarely sensible for an individual project.

    This tutorial follows the first approach. You will implement the model parameters, forward pass, loss, analytical gradients, and optimiser. You will not implement a BLAS library or a GPU runtime.

    Set up a reproducible environment

    Create an isolated environment and install only the dependencies needed for the first experiment:

    python -m venv .venv
    source .venv/bin/activate        # Windows: .venv\\Scripts\\activate
    pip install numpy matplotlib scikit-learn

    Use a fixed random seed, keep training and test data separate, and record Python and package versions. These habits matter when you later compare experiments on Indian-language, vision, or speech datasets. For larger projects, a clear repository structure helps:

    project/
      data.py
      model.py
      train.py
      evaluate.py
      tests/
      README.md

    Build a small classification dataset

    The following dataset labels points according to whether their two coordinates sum to more than one. It is simple enough to visualise and non-linear enough to require a hidden layer.

    import numpy as np
    
    rng = np.random.default_rng(7)
    X = rng.random((1200, 2)).astype(np.float32)
    y = (X[:, 0] + X[:, 1] > 1.0).astype(np.int64)
    
    indices = rng.permutation(len(X))
    split = int(0.8 * len(X))
    train_idx, test_idx = indices[:split], indices[split:]
    X_train, y_train = X[train_idx], y[train_idx]
    X_test, y_test = X[test_idx], y[test_idx]

    For real data, fit normalisation statistics on the training set only. Applying test-set statistics during preprocessing is a form of leakage and can make evaluation look better than the model really is.

    Implement the forward pass

    A two-layer network computes:

    1. z1 = XW1 + b1
    2. a1 = ReLU(z1)
    3. z2 = a1W2 + b2
    4. p = softmax(z2)

    Use small, variance-aware initial weights. The softmax implementation must subtract the row maximum to avoid overflow.

    class MLP:
        def __init__(self, input_dim=2, hidden_dim=16, classes=2, seed=7):
            rng = np.random.default_rng(seed)
            self.W1 = (rng.standard_normal((input_dim, hidden_dim)) *
                       np.sqrt(2 / input_dim)).astype(np.float32)
            self.b1 = np.zeros((1, hidden_dim), dtype=np.float32)
            self.W2 = (rng.standard_normal((hidden_dim, classes)) *
                       np.sqrt(2 / hidden_dim)).astype(np.float32)
            self.b2 = np.zeros((1, classes), dtype=np.float32)
    
        def forward(self, X):
            self.z1 = X @ self.W1 + self.b1
            self.a1 = np.maximum(self.z1, 0)
            self.z2 = self.a1 @ self.W2 + self.b2
            shifted = self.z2 - self.z2.max(axis=1, keepdims=True)
            exp_scores = np.exp(shifted)
            self.probs = exp_scores / exp_scores.sum(axis=1, keepdims=True)
            return self.probs

    Use a stable cross-entropy loss

    For integer class labels, select the predicted probability for each correct class:

    def cross_entropy(probs, y):
        chosen = np.clip(probs[np.arange(len(y)), y], 1e-12, 1.0)
        return -np.log(chosen).mean()

    Clipping prevents log(0) from producing infinite values. It does not fix exploding gradients or a bad learning rate; those require diagnosis rather than masking.

    Derive and implement backpropagation

    For softmax followed by cross-entropy, the output gradient simplifies to (probs - one_hot_labels) / batch_size. Then apply the chain rule through the second linear layer, ReLU, and first linear layer.

    def gradients(model, X, y):
        n = len(y)
        dz2 = model.probs.copy()
        dz2[np.arange(n), y] -= 1
        dz2 /= n
    
        dW2 = model.a1.T @ dz2
        db2 = dz2.sum(axis=0, keepdims=True)
        da1 = dz2 @ model.W2.T
        dz1 = da1 * (model.z1 > 0)
        dW1 = X.T @ dz1
        db1 = dz1.sum(axis=0, keepdims=True)
        return dW1, db1, dW2, db2

    A common mistake is using the wrong array shape for biases or forgetting to average gradients across the batch. NumPy broadcasting can hide these errors, so inspect every tensor shape while debugging.

    Train with mini-batch gradient descent

    def accuracy(probs, y):
        return (probs.argmax(axis=1) == y).mean()
    
    def train(model, X, y, epochs=300, batch_size=32, lr=0.05):
        rng = np.random.default_rng(11)
        for epoch in range(epochs):
            order = rng.permutation(len(X))
            for start in range(0, len(X), batch_size):
                idx = order[start:start + batch_size]
                model.forward(X[idx])
                dW1, db1, dW2, db2 = gradients(model, X[idx], y[idx])
                model.W1 -= lr * dW1
                model.b1 -= lr * db1
                model.W2 -= lr * dW2
                model.b2 -= lr * db2
    
            if epoch % 50 == 0 or epoch == epochs - 1:
                train_probs = model.forward(X)
                print(epoch, cross_entropy(train_probs, y), accuracy(train_probs, y))
    
    model = MLP()
    train(model, X_train, y_train)
    print("test accuracy:", accuracy(model.forward(X_test), y_test))

    Verify gradients before trusting training

    Gradient checking compares analytical gradients with finite differences. Select a few parameters, perturb each by a small value, and compare the numerical derivative with the backpropagated value. The relative error should be very small—often around 1e-5 for a carefully implemented toy model.

    Also test that:

    • Loss decreases on a tiny dataset.
    • The model can overfit 10–20 examples.
    • Shuffling labels destroys meaningful performance.
    • Removing the hidden layer changes the decision boundary as expected.
    • Training and test metrics are reported separately.

    These tests are more informative than adding layers blindly. For image work, the next practical step is understanding how to build computer vision models on GitHub, including dataset documentation and reproducible evaluation.

    Move from NumPy to PyTorch responsibly

    Once the mechanics are clear, use a framework for real work. PyTorch provides automatic differentiation, GPU support, data loaders, mixed precision, checkpointing, and tested optimisers. The equivalent model is short:

    import torch
    from torch import nn
    
    model = nn.Sequential(
        nn.Linear(2, 16), nn.ReLU(), nn.Linear(16, 2)
    )
    loss_fn = nn.CrossEntropyLoss()
    optimiser = torch.optim.Adam(model.parameters(), lr=1e-3)

    Frameworks reduce implementation risk, but they do not remove the need to understand data leakage, class imbalance, calibration, or deployment constraints. For Indian-language applications, document language coverage, script variation, annotation quality, and performance across regions rather than reporting one aggregate score. Open-source work on vision-language models for Indian languages illustrates why dataset and evaluation choices are central to useful AI.

    A practical project checklist

    Before calling the project complete, include:

    • A README explaining the equations, assumptions, and setup.
    • Unit tests for softmax, loss, shapes, and gradient checks.
    • A learning curve and a held-out test result.
    • A comparison against a PyTorch implementation.
    • Seeded experiments and a requirements file.
    • Error analysis, not only accuracy.
    • A note on compute, data licensing, and limitations.

    If the work becomes a research prototype or product, map the technical milestone to a funding and validation plan. The transition from a working notebook to a defensible company is covered in transitioning from research to a deep tech startup in India.

    Final perspective

    Building a neural network with NumPy teaches the complete learning loop: represent data, compute predictions, measure error, differentiate, update parameters, and evaluate honestly. That foundation makes high-level tools easier to use and harder to misuse. Start with a tiny, testable model; prove that every component works; then scale through established frameworks, better data, and disciplined experimentation.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.