0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · python scripts for automating image data labeling

Python Scripts for Automating Image Data Labeling

  1. aigi

    Automated labeling can turn a computer-vision dataset from a months-long annotation project into a controlled engineering workflow. But the objective is not to remove people from the process. It is to use models to generate reviewable drafts, direct human attention to difficult images, and preserve enough metadata to understand where labels came from.

    For Indian teams working with traffic cameras, retail shelves, crop disease, industrial inspection, or multilingual visual content, this distinction matters. A model trained on generic internet imagery may perform poorly on local road conditions, regional products, low-light footage, crowded scenes, or images captured on inexpensive phones. Python gives you the flexibility to combine open models, custom rules, annotation tools, and validation checks without locking the project into one platform.

    What an automated labeling pipeline should do

    A production-ready script normally handles six steps:

    • Ingest: discover images from local storage, object storage, or an annotation export.
    • Inspect: read dimensions, colour channels, file integrity, and duplicate status.
    • Predict: run an object detector, vision-language model, tracker, or segmentation model.
    • Normalise: map predictions to a consistent class taxonomy and coordinate system.
    • Review: send low-confidence or unusual examples to an annotator.
    • Validate and export: check geometry, class IDs, splits, and output format before training.

    Keep these stages separate. If inference, conversion, and review logic are embedded in one notebook, it becomes difficult to reproduce a dataset or diagnose a bad model. Store the image ID, model name, checkpoint version, prompt, confidence threshold, timestamp, and reviewer decision alongside each annotation. This provenance is a practical foundation for data veracity infrastructure for high-stakes AI.

    Start with a small, representative seed set

    Do not auto-label the entire archive immediately. First annotate a carefully selected seed set covering the conditions your model will encounter: camera angles, weather, lighting, object sizes, occlusion, regional variations, and empty scenes. For many early projects, a few hundred well-chosen images are more useful than thousands of easy examples.

    Use the seed set to test the label taxonomy. Decide whether visually similar categories should be merged, whether an object can have multiple labels, and how to mark truncated or ambiguous objects. These decisions should be written down before scaling. Inconsistent class definitions create more damage than imperfect model predictions.

    A lightweight Python inspection script can identify corrupt files and unusual dimensions before inference:

    from pathlib import Path
    from PIL import Image
    
    root = Path("raw_images")
    valid = []
    problems = []
    
    for path in root.rglob("*"):
        if path.suffix.lower() not in {".jpg", ".jpeg", ".png", ".webp"}:
            continue
        try:
            with Image.open(path) as im:
                im.verify()
            with Image.open(path) as im:
                valid.append({"file": str(path), "width": im.width, "height": im.height})
        except Exception as exc:
            problems.append({"file": str(path), "error": str(exc)})

    For larger collections, add perceptual hashes to detect near-duplicates and prevent the same scene from appearing in both training and validation sets.

    Generate draft boxes with detection models

    For fixed categories, a YOLO-family detector is often fast and economical once you have a seed model. For categories that change frequently, text-prompted detectors such as Grounding DINO can provide useful first drafts. Prompts should describe the object clearly, but test synonyms and context: “helmeted two-wheeler rider” may behave differently from “person wearing a helmet.”

    Treat confidence thresholds as a review policy, not a guarantee of correctness. A high threshold reduces false positives but can systematically miss small, occluded, or unusual objects. A better workflow commonly uses three bands:

    • High confidence: accept automatically only after validation on a held-out sample.
    • Middle confidence: route to quick human review.
    • Low confidence: reject, or send to detailed annotation.

    Save raw predictions before applying filtering. This lets you change thresholds later without rerunning expensive inference.

    Use SAM for masks, not as a complete labeling system

    Segment Anything Model (SAM) and newer variants can produce strong masks, but they still need a prompt such as a box, point, or coarse mask. A reliable pipeline pairs a detector with a segmentation model: detect the object, pass the box to SAM, then apply post-processing to remove tiny regions and fill obvious holes.

    import numpy as np
    from segment_anything import SamPredictor
    
    # predictor is loaded once and reused for a batch
    predictor.set_image(rgb_image)
    box = np.array([x_min, y_min, x_max, y_max], dtype=np.float32)
    masks, scores, _ = predictor.predict(
        box=box[None, :],
        multimask_output=True,
    )
    best = masks[int(np.argmax(scores))]

    Inspect masks visually. Hair, wires, reflective surfaces, crop leaves, and overlapping objects are common failure cases. For medical imagery, do not transfer ordinary computer-vision assumptions directly. DICOM handling, clinical review, privacy controls, and domain-specific validation are essential; teams should also consult ICMR-compliant medical AI data verification in India.

    Convert annotations safely

    Different training frameworks use different conventions. COCO stores absolute pixel coordinates and polygon or run-length masks. YOLO detection labels use normalised values in the order class_id centre_x centre_y width height. Conversion errors can silently produce unusable training data.

    def xyxy_to_yolo(box, image_width, image_height):
        x1, y1, x2, y2 = box
        x1 = max(0, min(x1, image_width))
        x2 = max(0, min(x2, image_width))
        y1 = max(0, min(y1, image_height))
        y2 = max(0, min(y2, image_height))
    
        cx = ((x1 + x2) / 2) / image_width
        cy = ((y1 + y2) / 2) / image_height
        w = (x2 - x1) / image_width
        h = (y2 - y1) / image_height
        return cx, cy, w, h

    Add automated assertions: coordinates must be finite, widths and heights must be positive, normalised values must fall between zero and one, and every class ID must exist in the taxonomy. Render a random sample of converted labels over the original images. A visual spot-check catches axis swaps and off-by-one errors quickly.

    Teams that already maintain ingestion and cleaning jobs can pair this workflow with Python scripts for automating data preprocessing, rather than duplicating file-handling logic.

    Build human review around uncertainty

    Human-in-the-loop review is where auto-labeling becomes dependable. Create queues for low confidence, disagreement between two models, unusually large or small boxes, crowded images, and samples selected from new locations or devices. Reviewers should be able to accept, edit, reject, and record a reason.

    Measure more than average accuracy. Track precision and recall by class, device, geography, lighting condition, and object size. In India, a model may work on urban daylight footage but fail on monsoon glare, dusty roads, regional vehicle designs, or low-bandwidth compressed images. Stratified metrics reveal these failures before deployment.

    Use active learning: after each review cycle, retrain or fine-tune the teacher model on corrected examples, then score the unlabelled pool again. Keep a fixed evaluation set that is never used for pseudo-label generation. For sensitive domains, use dual review and retain an audit trail.

    Scale without losing reproducibility

    For a small dataset, a Python command-line tool, a requirements lockfile, and a clear directory convention may be enough. Larger teams should add batch processing, resumable jobs, GPU monitoring, structured logs, and content-addressed image IDs. Separate raw data from generated labels so that a model update never overwrites the original source.

    Before training, produce a dataset report containing:

    • image and annotation counts by split and class;
    • rejected, missing, and duplicate files;
    • confidence distributions and review outcomes;
    • samples with empty labels or extreme geometry;
    • model and script versions;
    • known limitations and unresolved reviewer disagreements.

    If your startup is building broader automation around these workflows, Python data science automation for Indian startups offers a useful complementary direction. For tool selection, compare this custom approach with automated image labeling tools for developers, especially when you need collaboration, permissions, or hosted review queues.

    Practical recommendations for 2026

    Use open models for draft generation, but benchmark them on your own images before promising time savings. Keep a modest, high-quality human-reviewed set permanently available for regression testing. Version prompts, checkpoints, thresholds, taxonomies, and conversion code. Most importantly, treat every automatically created annotation as a hypothesis until it passes your quality gates.

    The best Python labeling system is not the one that produces the most boxes. It is the one that produces traceable, representative, reviewable labels at a cost and speed your team can sustain.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.