0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build plant disease detection system using python

How to Build a Plant Disease Detection System Using Python

  1. aigi

    What you will build

    This guide shows how to build a plant disease image-classification system in Python. A farmer, field worker, or agronomist uploads a leaf photograph; the system returns a probable disease class and a confidence score. It is a decision-support tool—not a replacement for laboratory diagnosis or an agronomist.

    For an India-ready prototype, start with a narrow scope: one crop, a small set of diseases, and a healthy class. A model trained on clean laboratory images may perform well in a notebook and fail on a phone photograph taken in shade, dust, or a crowded field. Treat data quality and field validation as core engineering work.

    Define the problem before writing code

    Write down these decisions first:

    • Crop: tomato, cotton, rice, grape, chilli, or another clearly defined crop.
    • Classes: healthy plus diseases that look sufficiently different in photographs.
    • Input: a single leaf, whole plant, or a close-up of a suspected lesion.
    • Output: disease name, confidence, recommended next action, and an “uncertain” response.
    • Users: farmers, extension workers, nurseries, or researchers.
    • Connectivity: cloud inference, on-device inference, or an offline queue for low-connectivity areas.

    Do not begin with “all Indian plant diseases.” Similar symptoms, mixed infections, nutrient deficiencies, and pest damage make that target difficult to label reliably. If your project also needs regional-language guidance, plan a separate content layer; work on low-resource Indic natural language processing can inform multilingual labels and farmer-facing explanations.

    Set up the Python environment

    Use a virtual environment and pin dependencies so training and deployment remain reproducible. TensorFlow or PyTorch can work; the example below uses TensorFlow and Keras.

    python -m venv .venv
    source .venv/bin/activate        # Windows: .venv\\Scripts\\activate
    pip install tensorflow opencv-python numpy pandas scikit-learn matplotlib fastapi uvicorn pillow

    For a larger project, keep a requirements.txt, record the Python version, and store configuration separately from code. A useful repository layout is:

    data/raw/
    data/processed/
    src/train.py
    src/infer.py
    models/
    api/
    notebooks/

    A GPU is helpful for training but not essential for a small transfer-learning experiment. For deployment, measure latency and memory on the actual Android phone, edge device, or server you expect to use.

    Collect and label representative images

    Public datasets such as PlantVillage are useful for learning, but many images have uniform backgrounds and ideal lighting. Add locally collected photographs from farms, agricultural universities, Krishi Vigyan Kendras, nurseries, and extension programmes. Obtain permission, remove personally identifying information, and record metadata such as crop variety, location, date, growth stage, and capture conditions where appropriate.

    Create a labelling guide with example images and explicit rules for:

    • disease severity and mixed symptoms;
    • healthy leaves with natural blemishes;
    • nutrient deficiency, insect damage, and mechanical injury;
    • unusable images, including blur and severe occlusion;
    • images where the correct label requires expert inspection.

    Use two reviewers for difficult samples and retain an unknown or needs_expert_review category. Split data by plant, plot, or collection session, not by randomly copying near-identical images into train and test sets. Otherwise, leakage will inflate accuracy.

    Prepare images and datasets

    Resize images to a fixed shape such as 224 × 224 pixels, normalise according to the selected model, and apply realistic augmentation. Horizontal flips may be reasonable for leaves; aggressive colour changes can destroy disease cues. Useful transformations include modest rotation, crop, brightness variation, blur, and background variation.

    import tensorflow as tf
    
    IMG_SIZE = (224, 224)
    BATCH_SIZE = 32
    
    train_ds = tf.keras.utils.image_dataset_from_directory(
        "data/processed/train",
        image_size=IMG_SIZE,
        batch_size=BATCH_SIZE,
        label_mode="int",
    )
    
    val_ds = tf.keras.utils.image_dataset_from_directory(
        "data/processed/val",
        image_size=IMG_SIZE,
        batch_size=BATCH_SIZE,
        label_mode="int",
        shuffle=False,
    )

    Check class counts before training. If one disease dominates, use class weights, targeted collection, or carefully designed oversampling. Keep a completely untouched field-test set for the final evaluation.

    Train a transfer-learning model

    A lightweight pretrained model such as MobileNetV3 or EfficientNet can outperform a small CNN when labelled data is limited. Start with the backbone frozen, train a classification head, then unfreeze the final layers for low learning-rate fine-tuning.

    import tensorflow as tf
    from tensorflow.keras import layers
    
    num_classes = 5
    base = tf.keras.applications.MobileNetV3Small(
        input_shape=(224, 224, 3),
        include_top=False,
        weights="imagenet"
    )
    base.trainable = False
    
    inputs = tf.keras.Input(shape=(224, 224, 3))
    x = tf.keras.applications.mobilenet_v3.preprocess_input(inputs)
    x = base(x, training=False)
    x = layers.GlobalAveragePooling2D()(x)
    x = layers.Dropout(0.25)(x)
    outputs = layers.Dense(num_classes, activation="softmax")(x)
    model = tf.keras.Model(inputs, outputs)
    
    model.compile(
        optimizer=tf.keras.optimizers.Adam(1e-3),
        loss="sparse_categorical_crossentropy",
        metrics=["accuracy"]
    )
    
    callbacks = [
        tf.keras.callbacks.EarlyStopping(patience=5, restore_best_weights=True),
        tf.keras.callbacks.ReduceLROnPlateau(patience=2)
    ]
    model.fit(train_ds, validation_data=val_ds, epochs=30, callbacks=callbacks)

    Save the class-to-index mapping alongside the model. A prediction is unusable if the API cannot reliably translate index 2 into the correct disease name and language-specific display label.

    Evaluate beyond accuracy

    Report a confusion matrix, per-class precision, recall, and macro F1. Accuracy can hide poor performance on a rare but economically important disease. Also measure calibration: a model that says “98%” when it is frequently wrong is unsafe in the field.

    Test separately on:

    • images from farms not represented in training;
    • different phones, lighting, backgrounds, and distances;
    • early and severe disease stages;
    • healthy leaves with dust, holes, or natural colour variation;
    • crops and varieties outside the initial collection sites.

    Add an abstention rule. For example, return “retake image or consult an expert” when confidence is below a threshold or when image-quality checks fail. Thresholds must be selected on validation data, not chosen merely because they look reassuring.

    For deeper computer-vision implementation patterns, the guide to building computer vision models on GitHub is a useful companion, especially for experiment tracking, reproducible training, and model versioning.

    Expose predictions through an API

    A small FastAPI service can load the model once and accept an image upload. Validate file type and size, convert to RGB, resize consistently, and never trust a client-supplied filename.

    from fastapi import FastAPI, UploadFile, HTTPException
    from PIL import Image
    import io
    
    app = FastAPI()
    
    @app.post("/predict")
    async def predict(file: UploadFile):
        if file.content_type not in {"image/jpeg", "image/png"}:
            raise HTTPException(415, "Upload a JPEG or PNG image")
        image = Image.open(io.BytesIO(await file.read())).convert("RGB")
        # preprocess image, run model, apply confidence threshold
        return {"label": "tomato_leaf_mold", "confidence": 0.87}

    Return the model version, predicted class, confidence, and guidance such as “capture a closer image of the underside of the leaf.” Do not present pesticide dosage automatically unless recommendations have been reviewed by qualified agricultural experts and comply with local rules.

    Deploy for Indian field conditions

    Choose the deployment route based on connectivity, privacy, cost, and expected traffic:

    • Cloud API: simplest to update, but requires reliable connectivity and sends images off-device.
    • Android or edge inference: better for intermittent connectivity; export to TensorFlow Lite or ONNX and benchmark quantised models.
    • Hybrid: run quality checks and provisional inference offline, then sync images and feedback when online.

    Use HTTPS, authentication, rate limits, structured logs, and deletion policies for uploaded photographs. Monitor latency, failed uploads, confidence distributions, and disagreement with expert reviews. New varieties and seasonal conditions can create data drift, so plan periodic relabelling and retraining rather than treating the first model as finished.

    A practical launch checklist

    Before a pilot, confirm that you have:

    • a clearly bounded crop and disease taxonomy;
    • expert-reviewed labels and leakage-free splits;
    • field images from multiple regions and devices;
    • per-class evaluation and an abstention path;
    • a tested image-quality check;
    • model and label-map versioning;
    • consent, privacy, and data-retention rules;
    • feedback from agronomists and actual users.

    The strongest system is not the one with the highest notebook accuracy. It is the one that recognises its limits, works on imperfect photographs, communicates uncertainty clearly, and improves through verified field feedback.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.