Object detection predicts both what appears in an image and where it appears. A useful system must return a class, confidence score, and bounding box for every detected object—and do so reliably across lighting, camera, scale, and background changes.
Building an object detector from scratch in Python is an excellent way to understand the mechanics behind modern vision systems. It is also a demanding engineering project. You must define the label schema, prevent annotation errors, implement batching for variable numbers of boxes, balance classification and localisation losses, and measure performance beyond a single demo image.
This guide uses PyTorch concepts because they expose the training loop clearly. The same design applies if you later export the model to ONNX, TensorRT, TFLite, or a mobile runtime.
Decide what “from scratch” means
The phrase can mean two different things:
- From-scratch pipeline: You implement the dataset, model head, losses, training loop, evaluation, and post-processing yourself, while using a standard backbone.
- From-scratch weights: Every model weight starts randomly initialised, with no ImageNet or detection pre-training.
For learning, implement the pipeline yourself. For a production system, use transfer learning unless you have a large, diverse dataset and enough compute. Random initialisation generally needs substantially more data and careful scheduling, while a pre-trained backbone converges faster and often performs better on small Indian datasets.
If you want to understand the wider computer-vision workflow, building computer vision models on GitHub is a useful companion to this implementation-focused guide.
1. Define the detection problem and dataset contract
Start with a precise specification before writing model code. Record:
- The object classes, including an explicit background convention.
- Whether objects can be truncated, heavily occluded, or extremely small.
- The minimum object size worth detecting.
- Required latency, image resolution, and hardware.
- The target evaluation metric and acceptable false-positive rate.
For Indian deployments, camera placement and conditions matter as much as architecture. A traffic camera, warehouse camera, mobile phone, and agricultural drone will produce very different data. Include regional variation such as monsoon glare, dust, low-light footage, crowded scenes, local vehicle types, clothing, signage, and compressed CCTV streams.
Use COCO JSON or a similarly structured format. Each annotation should include an image identifier, category identifier, and bounding box in a documented coordinate convention. Validate that boxes have positive width and height, lie within image boundaries, and reference existing images and categories.
Split by location, camera, person, or video sequence, not only by random image. Adjacent frames in both training and validation sets can make results look artificially strong.
2. Annotate, inspect, and augment
Tools such as CVAT and Label Studio can accelerate labelling, but human review remains essential. Sample annotations from every class and source, then inspect them for missing objects, inconsistent class names, loose boxes, and duplicate labels.
Useful augmentations include horizontal flips where semantically valid, modest scale changes, crops, colour shifts, blur, noise, and brightness variation. Apply geometric transforms to both the image and its boxes. An augmentation that changes pixels but leaves coordinates untouched silently corrupts training.
Avoid aggressive transformations that create unrealistic examples. If small objects are important, random crops may remove them or reduce them to unlearnable sizes. Track the percentage of images containing each class after augmentation.
3. Build a PyTorch dataset and collate function
Detection images rarely contain the same number of objects. Your dataset should return an image tensor and a target dictionary for each sample:
import torch
from torch.utils.data import Dataset
from PIL import Image
class DetectionDataset(Dataset):
def __init__(self, records, image_dir, transforms=None):
self.records = records
self.image_dir = image_dir
self.transforms = transforms
def __getitem__(self, index):
record = self.records[index]
image = Image.open(self.image_dir / record["file_name"]).convert("RGB")
boxes = torch.tensor(record["boxes"], dtype=torch.float32).reshape(-1, 4)
labels = torch.tensor(record["labels"], dtype=torch.int64)
target = {"boxes": boxes, "labels": labels,
"image_id": torch.tensor([record["image_id"]])}
if self.transforms:
image, target = self.transforms(image, target)
return image, target
def collate_fn(batch):
return tuple(zip(*batch))A production dataset should also check for empty annotations, invalid coordinates, unreadable files, and category IDs outside the configured range. Keep a small deterministic validation sample so that changes to preprocessing can be tested consistently.
4. Choose an architecture
Most detectors contain three parts:
- Backbone: extracts visual features, using a CNN or vision transformer.
- Neck: combines features at different resolutions, often with an FPN or PAN-style structure.
- Head: predicts class scores and box coordinates.
An anchor-based head predicts offsets from predefined boxes. This approach is established but requires choices about anchor sizes, aspect ratios, matching thresholds, and negative examples. An anchor-free head predicts object centres or distances to box sides and usually has fewer design parameters.
For a first implementation, use a small backbone and one or more feature scales. Larger models cannot compensate for faulty labels or a broken loss. A lightweight backbone is also a better starting point if the final target is a low-cost GPU, CPU, or mobile device.
5. Implement targets and losses
A detector normally optimises a combined objective:
Total loss = classification loss + box regression loss + optional objectness or centerness loss.
Classification may use cross-entropy or focal loss. Focal loss reduces the influence of abundant easy background examples. For box regression, Smooth L1 is stable during early experiments; IoU-based losses align more directly with detection quality. GIoU, DIoU, and CIoU can help when predicted and target boxes barely overlap.
The difficult part is target assignment. For each feature location or anchor, decide whether it is positive, negative, or ignored. Log positive counts per batch. A sudden zero-positive batch, extreme class imbalance, or exploding box loss usually indicates a coordinate-format or assignment bug rather than a model-capacity problem.
Test IoU independently before training:
def box_iou(a, b):
lt = torch.maximum(a[:, None, :2], b[None, :, :2])
rb = torch.minimum(a[:, None, 2:], b[None, :, 2:])
wh = (rb - lt).clamp(min=0)
inter = wh[..., 0] * wh[..., 1]
area_a = (a[:, 2] - a[:, 0]) * (a[:, 3] - a[:, 1])
area_b = (b[:, 2] - b[:, 0]) * (b[:, 3] - b[:, 1])
return inter / (area_a[:, None] + area_b[None, :] - inter + 1e-6)6. Train reproducibly
Use a warm-up period, an appropriate learning-rate schedule, mixed precision when supported, and gradient clipping if instability appears. Record the random seed, code version, dataset revision, image size, batch size, learning rate, and augmentation configuration.
Begin with a tiny overfit test: train on 10–20 images and confirm that the model can drive training loss down and recover the boxes. If it cannot, do not scale to thousands of images. Check image normalisation, label indexing, box conversion, target assignment, and loss signs first.
Monitor more than total loss. Save checkpoints based on validation performance, not merely the final epoch. For experiments, use TensorBoard, Weights & Biases, or a simple CSV log containing losses, learning rate, GPU memory, and validation metrics.
7. Evaluate with mAP and error analysis
Mean Average Precision depends on the IoU threshold. Report at least AP at 0.50 IoU and a stricter COCO-style average across thresholds where practical. Also report per-class precision, recall, and the number of false positives per image.
Review false negatives and false positives by category and operating condition. A low overall score may hide excellent performance on large objects and failure on distant ones. Create slices for night scenes, cameras, districts, weather, object size, and crowd density. These slices are often more useful to a product team than one headline number.
8. Apply non-maximum suppression correctly
Inference produces many candidate boxes. Class-aware NMS sorts candidates by confidence and removes boxes whose IoU with a higher-scoring box exceeds a threshold. Use score thresholds and NMS thresholds chosen on validation data; do not assume defaults are suitable for crowded scenes. Soft-NMS or weighted box fusion may help when nearby objects overlap, but they do not fix poor localisation.
Keep preprocessing and post-processing identical between validation and deployment. Differences in resizing, padding, RGB/BGR ordering, or coordinate scaling can erase the accuracy you measured in Python.
9. Optimise for deployment in India
Benchmark the complete pipeline—not only the neural network—on the intended device. Measure preprocessing, inference, NMS, memory, power, and sustained latency. Consider:
- Smaller backbones and lower input sizes for CPU or edge devices.
- FP16 or INT8 quantisation after checking accuracy on representative data.
- ONNX, TensorRT, OpenVINO, or TFLite export where supported.
- Frame skipping and tracking when processing fixed CCTV streams.
- Local buffering and privacy-preserving retention policies for sensitive footage.
Quantisation calibration data should reflect real deployment conditions, including low light and compression. If the detector supports safety, payments, healthcare, or public infrastructure, define an escalation path for uncertain predictions rather than forcing every output into a class.
Practical checklist
Before calling the model production-ready, confirm that you can answer yes to these questions:
- Are labels and coordinate conventions validated automatically?
- Is the split free from video or camera leakage?
- Can the model overfit a tiny sample?
- Are AP and per-condition errors tracked on a fixed validation set?
- Does exported inference match PyTorch inference?
- Has latency been measured on the actual target hardware?
- Are confidence thresholds calibrated for the business cost of errors?
Object detection from scratch is valuable because it makes every assumption visible. Start with a small, testable detector, establish trustworthy data and metrics, then add capacity only when error analysis shows that the model—not the pipeline—is the limiting factor.