0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · build custom neural network from scratch

Build a Custom Neural Network from Scratch with NumPy

  1. aigi

    A custom neural network is worth building when you need to understand the mechanics of learning, validate a research idea, or create a deliberately small model for a constrained device. It is not automatically the right route for production: mature frameworks provide faster kernels, automatic differentiation, distributed training, and deployment tooling. The practical goal is to learn the complete training loop and know where custom code creates value.

    This guide builds a binary-classification multilayer perceptron (MLP) with Python and NumPy. The same structure extends to regression, multiclass classification, and domain-specific experiments, including models for Indian-language or edge settings. If your project is specifically focused on Indic data, pair this tutorial with the guide to low-resource Indic natural language processing.

    What you will build

    The network accepts a batch of examples, applies two hidden layers with ReLU activations, and produces one probability with a sigmoid output:

    • Input: X with shape (features, examples)
    • Hidden layers: affine transformation followed by ReLU
    • Output: sigmoid probability for binary classification
    • Loss: binary cross-entropy
    • Optimiser: mini-batch gradient descent

    Keeping examples in columns makes the matrix operations explicit. For a layer, the forward equation is:

    Z[l] = W[l] A[l-1] + b[l]

    A[l] = activation(Z[l])

    Here, W has shape (units in current layer, units in previous layer), while b broadcasts across the batch.

    Step 1: Prepare and validate the data

    Before changing the architecture, establish reliable data contracts. Scale continuous features using statistics from the training split only, encode labels as zero or one, and keep validation and test data untouched until evaluation. Leakage in preprocessing can make a weak model appear accurate.

    For an Indian deployment, also inspect language, geography, device, and income-related slices where relevant. Overall accuracy can hide poor performance for a regional language, a low-connectivity user group, or a minority class.

    import numpy as np
    
    rng = np.random.default_rng(42)
    
    # X: (n_features, n_examples), y: (1, n_examples)
    # Use training-set mean and standard deviation for real data.
    def standardise(X, mean=None, std=None):
        if mean is None:
            mean = X.mean(axis=1, keepdims=True)
        if std is None:
            std = X.std(axis=1, keepdims=True)
        std = np.where(std < 1e-8, 1.0, std)
        return (X - mean) / std, mean, std

    Check shapes at every boundary. Silent broadcasting errors are among the most expensive bugs in handwritten neural-network code.

    Step 2: Initialise parameters correctly

    Zero biases are fine, but hidden-layer weights must break symmetry. He initialisation is a strong default for ReLU layers because it scales the variance by the number of incoming units. For sigmoid or tanh layers, Xavier initialisation is usually more suitable.

    def initialise(layer_dims, seed=42):
        rng = np.random.default_rng(seed)
        params = {}
        for l in range(1, len(layer_dims)):
            fan_in = layer_dims[l - 1]
            params[f"W{l}"] = (rng.standard_normal((layer_dims[l], fan_in))
                                 * np.sqrt(2.0 / fan_in))
            params[f"b{l}"] = np.zeros((layer_dims[l], 1))
        return params

    Use a seed while debugging so that a change to the optimiser or derivative can be reproduced. Remove dependence on a fixed seed only when you are measuring robustness across multiple runs.

    Step 3: Implement activations and the forward pass

    ReLU is inexpensive and works well in hidden layers, but it can produce permanently inactive neurons if learning rates are excessive or initialisation is poor. The sigmoid output should be numerically stable when logits become large.

    def relu(z):
        return np.maximum(0.0, z)
    
    def relu_grad(z):
        return (z > 0).astype(z.dtype)
    
    def sigmoid(z):
        out = np.empty_like(z, dtype=float)
        positive = z >= 0
        out[positive] = 1.0 / (1.0 + np.exp(-z[positive]))
        exp_z = np.exp(z[~positive])
        out[~positive] = exp_z / (1.0 + exp_z)
        return out
    
    def forward(X, params):
        caches = [(X, None)]
        L = len(params) // 2
        A = X
        for l in range(1, L):
            Z = params[f"W{l}"] @ A + params[f"b{l}"]
            A = relu(Z)
            caches.append((A, Z))
    
        Z = params[f"W{L}"] @ A + params[f"b{L}"]
        AL = sigmoid(Z)
        caches.append((AL, Z))
        return AL, caches

    Store the pre-activation Z and activation A values needed by backpropagation. Recomputing them later is slower and makes debugging harder.

    Step 4: Compute loss and gradients

    For binary classification, binary cross-entropy is:

    J = -1/m * sum(y log(a) + (1-y) log(1-a))

    Clip probabilities before taking logarithms. In the final layer, sigmoid plus cross-entropy simplifies the derivative to dZ = AL - Y, which reduces numerical work.

    def binary_cross_entropy(AL, Y):
        eps = 1e-12
        AL = np.clip(AL, eps, 1.0 - eps)
        return float(-np.mean(Y * np.log(AL) + (1 - Y) * np.log(1 - AL)))
    
    def backward(Y, params, caches):
        grads = {}
        m = Y.shape[1]
        L = len(params) // 2
        AL = caches[-1][0]
        A_prev = caches[-2][0] if L > 1 else caches[0][0]
    
        dZ = AL - Y
        grads[f"dW{L}"] = (dZ @ A_prev.T) / m
        grads[f"db{L}"] = np.sum(dZ, axis=1, keepdims=True) / m
        dA = params[f"W{L}"].T @ dZ
    
        for l in range(L - 1, 0, -1):
            A_prev, Z = caches[l - 1]
            dZ = dA * relu_grad(Z)
            grads[f"dW{l}"] = (dZ @ A_prev.T) / m
            grads[f"db{l}"] = np.sum(dZ, axis=1, keepdims=True) / m
            dA = params[f"W{l}"].T @ dZ
        return grads

    The most common implementation errors are an incorrect transpose, averaging over the wrong axis, and using the wrong cached activation. Add assertions for every gradient shape before updating parameters.

    Step 5: Train with mini-batches

    Full-batch updates are easy to understand but can be slow and memory-heavy. Mini-batches usually offer a better balance. Shuffle examples at every epoch, track validation loss, and stop when validation performance no longer improves.

    def update(params, grads, learning_rate):
        for key in params:
            params[key] -= learning_rate * grads["d" + key]
    
    def train(X, Y, layer_dims, epochs=1000, batch_size=64, lr=0.01):
        params = initialise(layer_dims)
        n = X.shape[1]
        rng = np.random.default_rng(42)
    
        for epoch in range(epochs):
            order = rng.permutation(n)
            for start in range(0, n, batch_size):
                idx = order[start:start + batch_size]
                AL, caches = forward(X[:, idx], params)
                grads = backward(Y[:, idx], params, caches)
                update(params, grads, lr)
    
            if epoch % 100 == 0:
                AL, _ = forward(X, params)
                print(epoch, binary_cross_entropy(AL, Y))
        return params

    For larger experiments, Adam, learning-rate schedules, weight decay, and gradient clipping can improve convergence. Add one change at a time; otherwise you will not know which intervention fixed or caused a training issue.

    Step 6: Verify before trusting results

    A low training loss is not proof that the implementation is correct. Use a repeatable verification checklist:

    • Gradient checking: Compare analytical gradients with finite differences on a tiny network. The relative error should be very small.
    • Overfit a tiny batch: A correct model should memorise a handful of examples.
    • Baseline comparison: Compare against logistic regression or a shallow tree.
    • Held-out evaluation: Report precision, recall, F1, calibration, and confusion matrices where class imbalance matters.
    • Slice tests: Evaluate by language, region, device, and other meaningful cohorts.
    • Stability tests: Run several seeds and record mean and variance, not only the best result.

    For computer vision, a handwritten MLP is rarely competitive with convolutional or pretrained architectures. Use this exercise to understand gradients, then move to a framework when the experiment requires serious scale. A similar boundary applies when building computer vision models on GitHub: repository quality, data licensing, reproducibility, and deployment constraints matter as much as the model code.

    When custom code is the right choice

    Use NumPy-first code for education, algorithm prototyping, small tabular models, numerical research, or highly constrained inference experiments. Move to PyTorch, JAX, or another framework when you need automatic differentiation, GPU kernels, mixed precision, distributed training, checkpointing, or production observability.

    For product teams, the valuable output is often not a permanent handwritten framework. It is a tested understanding of the architecture, a benchmark, and a clear decision about latency, memory, accuracy, and maintenance. If the model will become one component in a larger autonomous workflow, review the architectural trade-offs in building distributed systems with AI agents before adding complexity.

    Indian deployment considerations

    A model that performs well on a laptop may fail in the field. Measure inference time on the actual Android handset, gateway, or server used by customers. Quantise only after establishing a quality baseline, and test offline behaviour, retries, model updates, and corrupted inputs.

    For multilingual products, evaluate scripts and code-mixed speech or text separately. A small custom model can be useful for routing, classification, fraud signals, or on-device ranking, while larger models may remain behind an API. Protect user data, document consent and retention, and keep an auditable model version for every prediction.

    Frequently asked questions

    Should I build a neural network from scratch instead of using PyTorch?

    Build one from scratch to learn and validate the mathematics. Use a mature framework for serious training and deployment unless your research specifically requires a custom numerical stack.

    What is the best hidden-layer activation?

    Start with ReLU and He initialisation. Consider alternatives such as GELU or leaky ReLU when dead neurons or optimisation instability appear, and verify the change with controlled experiments.

    How much data is required?

    There is no universal threshold. A small tabular classifier may work with hundreds or thousands of labelled examples, while high-dimensional tasks need substantially more data or transfer learning. Always compare against a simple baseline and inspect validation curves.

    What should I build next?

    Add multiclass softmax, regularisation, an Adam optimiser, gradient checks, and model checkpointing. Then reproduce the same experiment in PyTorch and compare correctness, speed, memory, and maintainability.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.