0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · open source handwritten digit recognition python

Open-Source Handwritten Digit Recognition with Python

  1. aigi

    Handwritten digit recognition is a compact but valuable computer-vision project: it teaches the complete machine-learning workflow from image preprocessing to model deployment. A useful system must do more than achieve a high MNIST score. It should handle images captured on phones or scanners, expose uncertainty, and be tested on writing styles that differ from the training data.

    This guide shows how to build open source handwritten digit recognition in Python with TensorFlow/Keras, NumPy, and optional OpenCV preprocessing. The same approach can support form digitisation, examination workflows, inventory labels, postal processing, and offline education tools. If you are new to the ecosystem, compare this project with other open-source AI projects for student developers before choosing your scope.

    Choose the right dataset

    MNIST is the standard starting point: 70,000 grayscale images of handwritten digits, each sized 28×28 pixels. It is excellent for learning and benchmarking, but it is not representative of every real deployment. Scanned Indian forms, classroom worksheets, and mobile-camera images may contain different stroke widths, backgrounds, scripts, and digit conventions.

    Use the dataset according to your objective:

    • MNIST: fast experiments and baseline models.
    • EMNIST: additional handwritten characters and digits.
    • Your own images: realistic evaluation for a specific form, region, device, or user group.
    • Synthetic augmentation: controlled variation in rotation, scale, translation, and stroke thickness.

    CIFAR-10 is a colour object dataset, not a handwritten-digit dataset, so it should not be presented as a direct alternative. For multilingual document systems, digit recognition may eventually sit alongside low-resource Indic natural language processing, especially when numbers appear next to names, addresses, or local-language text.

    Set up the Python environment

    Create an isolated environment and install only the packages required for the baseline:

    python -m venv .venv
    source .venv/bin/activate        # Windows: .venv\\Scripts\\activate
    pip install tensorflow opencv-python scikit-learn matplotlib

    TensorFlow provides the training and inference stack, while OpenCV is useful for resizing, thresholding, deskewing, and extracting a digit from a larger photograph. Pin package versions in requirements.txt for reproducibility, and record the Python version, random seed, dataset version, and hardware used for each experiment.

    Build a reliable baseline

    The following model is intentionally small. A convolutional neural network generally performs better than a fully connected network because it learns local stroke patterns and preserves spatial structure.

    import numpy as np
    import tensorflow as tf
    from tensorflow.keras import layers, models
    
    SEED = 42
    tf.keras.utils.set_random_seed(SEED)
    
    (x_train, y_train), (x_test, y_test) = tf.keras.datasets.mnist.load_data()
    x_train = x_train.astype("float32") / 255.0
    x_test = x_test.astype("float32") / 255.0
    
    # Add the channel dimension expected by Conv2D.
    x_train = x_train[..., np.newaxis]
    x_test = x_test[..., np.newaxis]
    
    model = models.Sequential([
        layers.Input(shape=(28, 28, 1)),
        layers.Conv2D(32, 3, activation="relu"),
        layers.MaxPooling2D(),
        layers.Conv2D(64, 3, activation="relu"),
        layers.MaxPooling2D(),
        layers.Flatten(),
        layers.Dropout(0.3),
        layers.Dense(128, activation="relu"),
        layers.Dense(10, activation="softmax")
    ])
    
    model.compile(
        optimizer="adam",
        loss="sparse_categorical_crossentropy",
        metrics=["accuracy"]
    )
    
    callbacks = [
        tf.keras.callbacks.EarlyStopping(
            monitor="val_accuracy", patience=2, restore_best_weights=True
        )
    ]
    
    model.fit(
        x_train, y_train,
        validation_split=0.1,
        epochs=10,
        batch_size=128,
        callbacks=callbacks
    )
    
    loss, accuracy = model.evaluate(x_test, y_test, verbose=0)
    print(f"Test accuracy: {accuracy:.4f}")
    model.save("digit_classifier.keras")

    A single accuracy number is not enough. Generate a confusion matrix, inspect incorrectly classified images, and calculate per-class precision and recall. Digits such as 4, 7, and 9 can be confused when writing styles overlap. Also evaluate confidence: a prediction with 0.51 probability should be routed differently from one with 0.99 probability.

    Preprocess images outside MNIST

    Real inputs rarely arrive as clean 28×28 arrays. A practical pipeline usually includes:

    • Convert the image to grayscale.
    • Correct uneven lighting and remove background noise.
    • Threshold foreground strokes from the page.
    • Find the digit’s bounding box and add padding.
    • Preserve the aspect ratio while resizing.
    • Centre the digit using its ink mass rather than blindly cropping.
    • Invert colours if the model expects white strokes on a black background.

    OpenCV can perform these operations, but preprocessing must match the training distribution. A model trained on centred MNIST images may fail on a tightly cropped phone image even when the digit is visually obvious. Build a small validation set from the target camera, scanner, paper, or form and test the complete pipeline—not just the classifier.

    Improve generalisation with augmentation

    Use augmentation to model realistic variation, not to create arbitrary distortions. Small rotations, translations, zoom changes, and mild contrast variation are usually useful. Excessive warping can teach the model patterns that never occur in production.

    augment = tf.keras.Sequential([
        layers.RandomRotation(0.08),
        layers.RandomTranslation(0.08, 0.08),
        layers.RandomZoom(0.1),
    ])

    Place augmentation inside the model or apply it only to training data. Keep validation and test images untouched. When collecting your own data, obtain consent, remove unnecessary personal information, and separate contributors across train, validation, and test splits to avoid overly optimistic results.

    Evaluate for a real Indian deployment

    MNIST performance is a baseline, not a launch criterion. Define the operational metric first: missed digits may be more costly than manual review in financial forms, while latency may matter more in an offline classroom application. Track:

    • Accuracy and macro F1 score.
    • Per-digit confusion and error rates.
    • Accuracy by device, lighting condition, writer, and paper type.
    • Abstention or human-review rate for low-confidence predictions.
    • Inference latency and model size on the target hardware.

    For Indian deployments, test variation in handwriting across regions and age groups rather than assuming a single writing style. If digits appear in multilingual documents, design the full document pipeline separately from the digit model. Open-source work benefits from clear documentation; projects can also learn from Indian student developers building open-source AI and from established Indian open-source AI developer projects.

    Deploy the classifier

    For a Python API, load the .keras model once at process startup and validate input shape, data type, and pixel range before inference. Never accept arbitrary image dimensions without a defined preprocessing path. Return the predicted digit, confidence, and model version so downstream systems can audit results.

    For edge or offline use, convert the model to TensorFlow Lite and benchmark it on the actual Android device, Raspberry Pi, or low-cost computer. Quantisation can reduce storage and latency, but measure its effect on difficult examples. If you need a broader production architecture, review guidance on building high-performance AI applications with open-source tools.

    Open-source project checklist

    A useful repository should include:

    • A clear README with setup and licence information.
    • Dataset sources, terms of use, and preprocessing assumptions.
    • Reproducible training and evaluation scripts.
    • Saved metrics, confusion matrices, and example failure cases.
    • A small inference command or API example.
    • Tests for image shape, normalisation, and expected output format.
    • Model and dependency versioning.

    As of 2026, lightweight open-source models are easy to run, but responsible engineering still depends on data quality and evaluation discipline. Start with MNIST, then earn confidence on representative data before connecting predictions to financial, educational, or administrative decisions.

    FAQs

    Is MNIST enough for production?

    Usually not. It is suitable for a baseline, but production systems should be tested on images collected from the intended device, form, language context, and user population.

    Should I use TensorFlow or PyTorch?

    Either can solve the task. TensorFlow/Keras offers a concise beginner workflow and straightforward TensorFlow Lite export; PyTorch is also a strong choice when your team already uses its training and deployment tools.

    How can I improve accuracy?

    Fix preprocessing mismatches first, then add representative data, targeted augmentation, regularisation, and a better CNN. Inspect errors before increasing model size.

    Can this recognise numbers written in Indian scripts?

    The MNIST model recognises Arabic numerals. Recognition of local numeral systems or mixed-script documents requires suitable labelled data and a separate evaluation plan.

    Apply for AI Grants India

    If you are an Indian builder turning a digit-recognition prototype into an education, document, accessibility, or public-service product, explore AI Grants India for funding and support opportunities.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.