Real-time object detection turns camera frames into structured events: a person enters a zone, a vehicle crosses a line, or a defect appears on a production surface. The difficult part is not loading a YOLO model. It is building a system that remains accurate under Indian lighting, camera, connectivity, and hardware conditions while meeting a defined response-time target.
This guide explains how to implement real-time object detection algorithms from prototype to deployment. It covers architecture selection, dataset design, video pipelines, benchmarking, optimization, and operational safeguards for applications such as traffic analytics, warehouse automation, retail monitoring, agriculture, and industrial inspection.
Start with a measurable inference target
Define the operating requirement before choosing a model. “Real time” can mean very different things:
- 30 FPS: approximately 33 ms per frame, including capture, preprocessing, inference, post-processing, and display or transmission.
- 15 FPS: approximately 67 ms per frame, often adequate for fixed-camera monitoring.
- Event latency: the time between an object appearing and your application raising an alert. This may matter more than raw FPS.
- Accuracy by class: a missed helmet, pothole, or product defect may be more costly than a false alert.
Measure end-to-end latency, not only the model’s inference time. A fast network can still produce a slow application if frames queue up, decoding blocks the inference thread, or results are sent to a distant cloud service. For bandwidth-sensitive deployments, process video at the edge and transmit metadata or event clips instead of every frame.
Select an architecture for the deployment environment
One-stage detectors are usually the practical choice for live video because they predict classes and bounding boxes in a single pass. YOLO-family models offer a strong speed-accuracy baseline, while SSD variants remain useful on constrained CPUs and mobile hardware. EfficientDet can be suitable when scalable accuracy is more important than implementation simplicity.
Choose the smallest model that meets the business accuracy requirement:
- Nano or small YOLO variants: good starting points for Jetson-class devices, industrial gateways, and edge GPUs.
- MobileNet-SSD or similarly compact detectors: useful for CPU-only, mobile, and low-power deployments.
- Medium or large detectors: appropriate when small objects, heavy occlusion, or strict recall requirements justify additional compute.
Do not select a model based only on published FPS. Benchmark the exact camera resolution, batch size, runtime, and hardware you will deploy. A model that performs well on a desktop GPU may be unsuitable for an Indian retail outlet with intermittent connectivity and a low-power edge computer.
For systems that combine detection with downstream automation, define the event contract early. A detector can publish object class, confidence, timestamp, camera ID, bounding box, and tracking ID to the application layer. This is the same discipline used in other real-time monitoring systems in India: separate model output from business rules, alerts, and audit logs.
Build a dataset that matches the camera scene
COCO-pretrained weights are valuable for transfer learning, but generic data rarely captures your production environment. Collect footage from the actual camera positions, lenses, mounting heights, and operating hours you expect to support.
Prioritise examples that expose failure modes:
- Strong sunlight, glare, shadows, rain, dust, and low-light scenes.
- Crowded roads, markets, stations, or factory floors with occluded objects.
- Regional vehicles, uniforms, packaging, signage, and safety equipment.
- Different camera compression levels, frame rates, and viewpoints.
- Empty scenes and difficult negatives that resemble target objects.
Use CVAT or another annotation tool and maintain a clear class policy. Decide whether partially visible objects count, how small objects are labelled, and how overlapping instances are separated. Store annotations in YOLO or COCO format, validate bounding boxes automatically, and split data by location or recording session—not random adjacent frames. A random frame split can make validation look excellent while hiding poor performance on a new camera.
Augment carefully. Mosaic, scale, crop, brightness, blur, and compression augmentation can improve robustness, but aggressive transformations may create scenes that do not exist in production. Keep a fixed, untouched evaluation set representing real Indian operating conditions.
Train for the target, not for a leaderboard
Start with pretrained weights and fine-tune on your labelled data. Track per-class precision, recall, confusion matrices, and performance across lighting and camera groups. Overall mAP can conceal a serious failure on a high-value class.
Use a practical training loop:
1. Establish a baseline at the intended input resolution.
2. Identify the classes and conditions causing the most errors.
3. Add representative samples rather than indiscriminately collecting more frames.
4. Retrain and compare against the same evaluation set.
5. Export the candidate model to the production runtime and benchmark it there.
Input resolution is a key trade-off. Reducing 640×640 to 320×320 can substantially improve throughput but may eliminate small objects. Consider tiling or region-of-interest inference when only part of a high-resolution frame matters. For fixed cameras, masking irrelevant areas can reduce false positives and compute cost.
Implement the video pipeline correctly
A basic Python prototype can use OpenCV with Ultralytics, ONNX Runtime, or another supported inference library. The production pipeline should have separate stages:
- Capture and decode: read from USB, RTSP, or recorded video without blocking inference.
- Preprocessing: resize and pad consistently, convert colour format, and normalise as required by the model.
- Inference: execute the model using the selected accelerator.
- Post-processing: apply confidence filtering and non-maximum suppression.
- Tracking: associate detections across frames when stable object IDs or dwell time are needed.
- Business logic: trigger line crossings, zone violations, counts, or defect workflows.
- Output: render selectively; store structured events and evidence clips according to policy.
Use bounded queues between stages. If inference cannot keep up, dropping old frames is often better than processing a growing backlog. For alerts, the newest frame usually has more operational value than a delayed frame. Multi-threaded capture and inference, hardware decoding, and asynchronous copies can improve throughput without changing the model.
Optimise and deploy at the edge
Export the trained model to an interoperable format such as ONNX, then benchmark the runtime on the target device. NVIDIA deployments can use TensorRT for kernel fusion and FP16 or INT8 execution. Intel-based systems may benefit from OpenVINO, while ONNX Runtime provides a flexible baseline across platforms.
Quantisation reduces memory use and can lower latency significantly. Prefer calibration data that reflects production scenes, especially when using INT8. After conversion, retest every important class; a small average accuracy change can hide a large recall loss for small or dark objects.
Also optimise the surrounding system:
- Use hardware-accelerated video decode where available.
- Avoid unnecessary copies between CPU and GPU memory.
- Match camera FPS to the actual processing budget.
- Batch only when latency permits; batching can increase throughput but delay individual frames.
- Run health checks for camera availability, queue depth, temperature, and inference time.
- Version models, labels, preprocessing settings, and runtime dependencies together.
Cloud inference can simplify central management, but connectivity and data costs matter. Edge inference is often preferable for sensitive footage, rapid alerts, and sites with unreliable connectivity. Retain only the video needed for investigation, and define access controls and retention periods before deployment.
Validate with operational metrics
A production acceptance test should include more than mAP. Record p50, p95, and p99 end-to-end latency; sustained FPS; dropped frames; CPU, GPU, and memory use; power draw; thermal throttling; and alert precision over a representative shift or week.
Test camera disconnection, time synchronisation, process restarts, storage exhaustion, model rollback, and network loss. Confirm that an alert is explainable: operators should be able to see the camera, timestamp, confidence, region, and evidence supporting the decision.
For safety-critical or high-consequence uses, keep a human review path. Detection should support decisions—not silently become the only source of truth. In defect and infrastructure applications, pair computer vision with a documented inspection process, similar to the controls required for automated railway track defect detection.
Common implementation mistakes
- Optimising FPS while ignoring alert latency and dropped frames.
- Training on near-duplicate frames from one camera.
- Reporting only aggregate mAP instead of per-class recall.
- Using confidence thresholds copied from another dataset.
- Treating tracking, counting, and detection as the same problem.
- Deploying an INT8 model without recalibrating and retesting it.
- Sending all raw video to the cloud when metadata would suffice.
- Failing to monitor drift after camera repositioning, seasonal changes, or new object types.
A reliable deployment is an iterative system: collect failure cases, label them, retrain, benchmark, and roll out gradually. If detections trigger customer communication or operational workflows, pair the vision service with a well-defined event and escalation layer; related design considerations appear in this practical guide to real-time voice agents.
FAQ
What language should I use? Python is ideal for experimentation and training. Use C++, Rust, or a vendor runtime integration when strict latency, memory control, or embedded deployment requirements justify it.
Can real-time detection run without a GPU? Yes. Compact models with OpenVINO or ONNX Runtime can work on modern CPUs, but benchmark the complete pipeline at the required resolution and camera count.
What is the best model? There is no universal winner. Select the smallest model that meets per-class accuracy and end-to-end latency targets on your actual hardware.
Should every frame be processed? Not always. Frame skipping, adaptive sampling, or region-of-interest inference can reduce cost, provided the application can tolerate missed short-lived events.
Build and fund the next deployment
Indian teams are applying computer vision to manufacturing, mobility, agriculture, logistics, public infrastructure, and safety. A clear target metric, representative dataset, and measured edge deployment make a stronger technical and funding case than a demo showing headline FPS. AI Grants India supports builders working to turn applied AI systems into deployable products.