0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · implementing neural networks from scratch in python

Implementing Neural Networks from Scratch in Python

  1. aigi

    A neural network is easier to use well when you understand what happens between the input array and the final prediction. Implementing neural networks from scratch in Python gives you that understanding: matrix shapes become explicit, gradients are no longer mysterious, and training failures become easier to diagnose.

    This guide builds a small feedforward network with NumPy for binary classification. It is intended as a learning implementation, not a replacement for PyTorch, JAX, or TensorFlow in production. Once the fundamentals are clear, you can move to customizable neural network architectures for beginners, where the same ideas are extended into reusable designs.

    What you will build

    The example network has:

    • An input layer with two features
    • One hidden layer using ReLU activation
    • One output neuron using sigmoid activation
    • Binary cross-entropy loss
    • Full-batch gradient descent
    • Accuracy tracking during training

    The model will learn the XOR-like decision boundary in a small synthetic dataset. This is deliberately compact: a small dataset makes it possible to inspect values, shapes, and gradients without hiding the logic behind a framework.

    Prerequisites and setup

    You should be comfortable with Python functions, array indexing, and basic matrix multiplication. You do not need advanced calculus, but you should understand that backpropagation applies the chain rule to calculate how each parameter affects the loss.

    Install the only required package:

    python -m pip install numpy

    Use a current Python release and set a random seed so that your results are reproducible. In a real project, keep preprocessing separate from model code. Scripts for automating data preprocessing in Python can help when you move beyond toy data.

    The mathematics behind each layer

    For a batch of inputs X, a dense layer computes:

    Z = XW + b
    A = activation(Z)

    If X has shape (samples, input_features), W has shape (input_features, units), and b has shape (1, units), then Z and A have shape (samples, units).

    For binary classification, sigmoid converts the final logit into a probability:

    sigmoid(z) = 1 / (1 + exp(-z))

    The binary cross-entropy loss is:

    L = -mean(y log(p) + (1-y) log(1-p))

    Using sigmoid with binary cross-entropy produces a particularly useful output gradient: dL/dz = p - y. That simplifies the final layer's backpropagation.

    A complete NumPy implementation

    import numpy as np
    
    
    def sigmoid(z):
        z = np.clip(z, -50, 50)  # avoids overflow in exp
        return 1.0 / (1.0 + np.exp(-z))
    
    
    def relu(z):
        return np.maximum(0, z)
    
    
    def relu_derivative(z):
        return (z > 0).astype(float)
    
    
    class NeuralNetwork:
        def __init__(self, input_size=2, hidden_size=4, learning_rate=0.1, seed=7):
            rng = np.random.default_rng(seed)
            # Small random weights break symmetry between hidden units.
            self.W1 = rng.normal(0, np.sqrt(2 / input_size), (input_size, hidden_size))
            self.b1 = np.zeros((1, hidden_size))
            self.W2 = rng.normal(0, np.sqrt(2 / hidden_size), (hidden_size, 1))
            self.b2 = np.zeros((1, 1))
            self.learning_rate = learning_rate
    
        def forward(self, X):
            self.Z1 = X @ self.W1 + self.b1
            self.A1 = relu(self.Z1)
            self.Z2 = self.A1 @ self.W2 + self.b2
            self.y_hat = sigmoid(self.Z2)
            return self.y_hat
    
        def loss(self, y, y_hat):
            eps = 1e-8
            y_hat = np.clip(y_hat, eps, 1 - eps)
            return -np.mean(y * np.log(y_hat) + (1 - y) * np.log(1 - y_hat))
    
        def backward(self, X, y):
            samples = X.shape[0]
    
            # Sigmoid + binary cross-entropy derivative
            dZ2 = self.y_hat - y
            dW2 = (self.A1.T @ dZ2) / samples
            db2 = np.sum(dZ2, axis=0, keepdims=True) / samples
    
            dA1 = dZ2 @ self.W2.T
            dZ1 = dA1 * relu_derivative(self.Z1)
            dW1 = (X.T @ dZ1) / samples
            db1 = np.sum(dZ1, axis=0, keepdims=True) / samples
    
            self.W2 -= self.learning_rate * dW2
            self.b2 -= self.learning_rate * db2
            self.W1 -= self.learning_rate * dW1
            self.b1 -= self.learning_rate * db1
    
        def fit(self, X, y, epochs=5000, log_every=500):
            for epoch in range(1, epochs + 1):
                y_hat = self.forward(X)
                current_loss = self.loss(y, y_hat)
                self.backward(X, y)
    
                if epoch == 1 or epoch % log_every == 0:
                    predictions = (y_hat >= 0.5).astype(int)
                    accuracy = np.mean(predictions == y)
                    print(f"epoch={epoch:4d} loss={current_loss:.4f} accuracy={accuracy:.2f}")
    
        def predict(self, X):
            probabilities = self.forward(X)
            return (probabilities >= 0.5).astype(int)
    
    
    X = np.array([
        [0, 0], [0, 1], [1, 0], [1, 1]
    ], dtype=float)
    y = np.array([[0], [1], [1], [0]], dtype=float)
    
    model = NeuralNetwork()
    model.fit(X, y)
    print(model.predict(X).ravel())

    The forward method caches intermediate values because backpropagation needs them. The backward method computes derivatives in reverse order and updates parameters. The cache is not a shortcut: retaining these values is how most neural-network implementations avoid recomputing the forward pass.

    Why the implementation works

    Weight initialisation

    If every weight starts at zero, hidden neurons receive identical gradients and learn identical features. Random initialisation breaks that symmetry. He-style scaling, used above, is a sensible starting point for ReLU networks.

    Activation functions

    ReLU is computationally simple and generally easier to optimise than sigmoid in hidden layers. Sigmoid remains appropriate for a single binary-classification output. For multiclass classification, use one output per class and a numerically stable softmax with categorical cross-entropy.

    Learning rate

    A learning rate that is too high can make the loss oscillate or become nan; one that is too low can make training appear frozen. Try values such as 0.01, 0.05, and 0.1, while comparing loss curves rather than relying on a single final prediction.

    Debugging and validation checklist

    Scratch implementations usually fail because of shape, scaling, or gradient errors—not because the architecture is too small. Check these systematically:

    • Print the shape of every matrix in the forward and backward passes.
    • Standardise features when their scales differ substantially.
    • Confirm labels use the expected shape, such as (samples, 1).
    • Verify that loss decreases on a dataset small enough to memorise.
    • Compare analytical gradients with numerical finite-difference estimates.
    • Check for nan values after every major operation.
    • Keep validation and test data separate from training data.

    A useful gradient check perturbs one parameter by a small value epsilon and estimates its derivative as (L_plus - L_minus) / (2 * epsilon). Compare that estimate with the derivative from backpropagation. This catches transposed matrices and missing averaging factors quickly.

    From toy code to a usable training pipeline

    The example uses full-batch training and no validation split. A practical workflow needs mini-batches, shuffling, early stopping, checkpointing, and experiment tracking. It should also calculate metrics that match the business problem: accuracy can be misleading for fraud, medical screening, or rare-event classification, where precision, recall, and calibration matter more.

    For repeatable deployments, separate data ingestion, preprocessing, training, evaluation, and serving. A guide to building end-to-end ML pipelines in Python covers this structure, while implementing scalable ML pipelines for predictive analytics is useful when datasets and inference workloads grow.

    Do not implement optimisers, GPU kernels, distributed training, or automatic differentiation from scratch unless that is the learning objective. Use a production framework after you understand the mechanics. Frameworks provide tested tensor operations, automatic differentiation, hardware acceleration, mixed precision, and deployment integrations that a small educational class cannot safely reproduce.

    Recommended next experiments

    Extend the class in small, testable steps:

    • Replace XOR with a real tabular dataset and add standardisation.
    • Add mini-batch training and compare it with full-batch updates.
    • Implement L2 regularisation and measure its effect on validation loss.
    • Add a multiclass output layer with softmax.
    • Record loss and metrics for plotting after each epoch.
    • Write unit tests for activation functions and gradient shapes.
    • Rebuild the same model in PyTorch and compare predictions and gradients.

    The goal of implementing neural networks from scratch in Python is not to avoid mature tools. It is to develop enough mechanical understanding to choose architectures sensibly, inspect training behaviour, and recognise when a model or data pipeline is failing. That foundation transfers directly to larger systems, including Python-based AI applications and production ML workflows.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.