A neural network is easier to use well when you understand what happens between the input array and the final prediction. Implementing neural networks from scratch in Python gives you that understanding: matrix shapes become explicit, gradients are no longer mysterious, and training failures become easier to diagnose.
This guide builds a small feedforward network with NumPy for binary classification. It is intended as a learning implementation, not a replacement for PyTorch, JAX, or TensorFlow in production. Once the fundamentals are clear, you can move to customizable neural network architectures for beginners, where the same ideas are extended into reusable designs.
What you will build
The example network has:
- An input layer with two features
- One hidden layer using ReLU activation
- One output neuron using sigmoid activation
- Binary cross-entropy loss
- Full-batch gradient descent
- Accuracy tracking during training
The model will learn the XOR-like decision boundary in a small synthetic dataset. This is deliberately compact: a small dataset makes it possible to inspect values, shapes, and gradients without hiding the logic behind a framework.
Prerequisites and setup
You should be comfortable with Python functions, array indexing, and basic matrix multiplication. You do not need advanced calculus, but you should understand that backpropagation applies the chain rule to calculate how each parameter affects the loss.
Install the only required package:
python -m pip install numpyUse a current Python release and set a random seed so that your results are reproducible. In a real project, keep preprocessing separate from model code. Scripts for automating data preprocessing in Python can help when you move beyond toy data.
The mathematics behind each layer
For a batch of inputs X, a dense layer computes:
Z = XW + b
A = activation(Z)If X has shape (samples, input_features), W has shape (input_features, units), and b has shape (1, units), then Z and A have shape (samples, units).
For binary classification, sigmoid converts the final logit into a probability:
sigmoid(z) = 1 / (1 + exp(-z))The binary cross-entropy loss is:
L = -mean(y log(p) + (1-y) log(1-p))Using sigmoid with binary cross-entropy produces a particularly useful output gradient: dL/dz = p - y. That simplifies the final layer's backpropagation.
A complete NumPy implementation
import numpy as np
def sigmoid(z):
z = np.clip(z, -50, 50) # avoids overflow in exp
return 1.0 / (1.0 + np.exp(-z))
def relu(z):
return np.maximum(0, z)
def relu_derivative(z):
return (z > 0).astype(float)
class NeuralNetwork:
def __init__(self, input_size=2, hidden_size=4, learning_rate=0.1, seed=7):
rng = np.random.default_rng(seed)
# Small random weights break symmetry between hidden units.
self.W1 = rng.normal(0, np.sqrt(2 / input_size), (input_size, hidden_size))
self.b1 = np.zeros((1, hidden_size))
self.W2 = rng.normal(0, np.sqrt(2 / hidden_size), (hidden_size, 1))
self.b2 = np.zeros((1, 1))
self.learning_rate = learning_rate
def forward(self, X):
self.Z1 = X @ self.W1 + self.b1
self.A1 = relu(self.Z1)
self.Z2 = self.A1 @ self.W2 + self.b2
self.y_hat = sigmoid(self.Z2)
return self.y_hat
def loss(self, y, y_hat):
eps = 1e-8
y_hat = np.clip(y_hat, eps, 1 - eps)
return -np.mean(y * np.log(y_hat) + (1 - y) * np.log(1 - y_hat))
def backward(self, X, y):
samples = X.shape[0]
# Sigmoid + binary cross-entropy derivative
dZ2 = self.y_hat - y
dW2 = (self.A1.T @ dZ2) / samples
db2 = np.sum(dZ2, axis=0, keepdims=True) / samples
dA1 = dZ2 @ self.W2.T
dZ1 = dA1 * relu_derivative(self.Z1)
dW1 = (X.T @ dZ1) / samples
db1 = np.sum(dZ1, axis=0, keepdims=True) / samples
self.W2 -= self.learning_rate * dW2
self.b2 -= self.learning_rate * db2
self.W1 -= self.learning_rate * dW1
self.b1 -= self.learning_rate * db1
def fit(self, X, y, epochs=5000, log_every=500):
for epoch in range(1, epochs + 1):
y_hat = self.forward(X)
current_loss = self.loss(y, y_hat)
self.backward(X, y)
if epoch == 1 or epoch % log_every == 0:
predictions = (y_hat >= 0.5).astype(int)
accuracy = np.mean(predictions == y)
print(f"epoch={epoch:4d} loss={current_loss:.4f} accuracy={accuracy:.2f}")
def predict(self, X):
probabilities = self.forward(X)
return (probabilities >= 0.5).astype(int)
X = np.array([
[0, 0], [0, 1], [1, 0], [1, 1]
], dtype=float)
y = np.array([[0], [1], [1], [0]], dtype=float)
model = NeuralNetwork()
model.fit(X, y)
print(model.predict(X).ravel())The forward method caches intermediate values because backpropagation needs them. The backward method computes derivatives in reverse order and updates parameters. The cache is not a shortcut: retaining these values is how most neural-network implementations avoid recomputing the forward pass.
Why the implementation works
Weight initialisation
If every weight starts at zero, hidden neurons receive identical gradients and learn identical features. Random initialisation breaks that symmetry. He-style scaling, used above, is a sensible starting point for ReLU networks.
Activation functions
ReLU is computationally simple and generally easier to optimise than sigmoid in hidden layers. Sigmoid remains appropriate for a single binary-classification output. For multiclass classification, use one output per class and a numerically stable softmax with categorical cross-entropy.
Learning rate
A learning rate that is too high can make the loss oscillate or become nan; one that is too low can make training appear frozen. Try values such as 0.01, 0.05, and 0.1, while comparing loss curves rather than relying on a single final prediction.
Debugging and validation checklist
Scratch implementations usually fail because of shape, scaling, or gradient errors—not because the architecture is too small. Check these systematically:
- Print the shape of every matrix in the forward and backward passes.
- Standardise features when their scales differ substantially.
- Confirm labels use the expected shape, such as
(samples, 1). - Verify that loss decreases on a dataset small enough to memorise.
- Compare analytical gradients with numerical finite-difference estimates.
- Check for
nanvalues after every major operation. - Keep validation and test data separate from training data.
A useful gradient check perturbs one parameter by a small value epsilon and estimates its derivative as (L_plus - L_minus) / (2 * epsilon). Compare that estimate with the derivative from backpropagation. This catches transposed matrices and missing averaging factors quickly.
From toy code to a usable training pipeline
The example uses full-batch training and no validation split. A practical workflow needs mini-batches, shuffling, early stopping, checkpointing, and experiment tracking. It should also calculate metrics that match the business problem: accuracy can be misleading for fraud, medical screening, or rare-event classification, where precision, recall, and calibration matter more.
For repeatable deployments, separate data ingestion, preprocessing, training, evaluation, and serving. A guide to building end-to-end ML pipelines in Python covers this structure, while implementing scalable ML pipelines for predictive analytics is useful when datasets and inference workloads grow.
Do not implement optimisers, GPU kernels, distributed training, or automatic differentiation from scratch unless that is the learning objective. Use a production framework after you understand the mechanics. Frameworks provide tested tensor operations, automatic differentiation, hardware acceleration, mixed precision, and deployment integrations that a small educational class cannot safely reproduce.
Recommended next experiments
Extend the class in small, testable steps:
- Replace XOR with a real tabular dataset and add standardisation.
- Add mini-batch training and compare it with full-batch updates.
- Implement L2 regularisation and measure its effect on validation loss.
- Add a multiclass output layer with softmax.
- Record loss and metrics for plotting after each epoch.
- Write unit tests for activation functions and gradient shapes.
- Rebuild the same model in PyTorch and compare predictions and gradients.
The goal of implementing neural networks from scratch in Python is not to avoid mature tools. It is to develop enough mechanical understanding to choose architectures sensibly, inspect training behaviour, and recognise when a model or data pipeline is failing. That foundation transfers directly to larger systems, including Python-based AI applications and production ML workflows.