0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build real time object detection systems

How to Build Real-Time Object Detection Systems

  1. aigi

    Real-time object detection is not simply a model that scores well on a benchmark. A usable system must capture frames, decode them, preprocess them, run inference, filter detections, track objects when necessary, and deliver decisions within a predictable latency budget. That changes how you choose data, hardware, model size, and deployment architecture.

    This guide explains how to build real time object detection systems for Indian startups, product teams, and applied-AI engineers. It focuses on production trade-offs: end-to-end latency rather than headline FPS, edge inference rather than unconditional cloud streaming, and reliability under crowded scenes, dust, glare, weak connectivity, and changing camera conditions.

    Start with the operating requirements

    Define the application before selecting a detector. A warehouse safety system, traffic camera, retail counter, and drone have different constraints.

    Write down:

    • Objects and actions: What must be detected, and what decision follows—alert, count, route, stop, or record?
    • Camera conditions: Resolution, field of view, frame rate, night performance, mounting height, and compression.
    • Latency budget: Separate capture-to-decode, preprocessing, inference, post-processing, tracking, and network or UI delay.
    • Throughput: One 30 FPS stream is a different problem from 32 cameras at 15 FPS.
    • Failure cost: A missed person, false alarm, or delayed industrial stop may have very different consequences.
    • Privacy and retention: Decide whether raw video leaves the site, how long it is stored, and who can access it.

    A 30 FPS camera produces a frame every 33 milliseconds, but processing every frame is not always necessary. Many applications can sample at 10–15 FPS while maintaining acceptable tracking. Conversely, a robotic control loop may need low glass-to-glass latency even if average throughput is high. Measure both p50 and p95 latency, not only average FPS.

    Build a representative dataset

    Model quality is usually constrained by data quality. Collect footage from the actual camera positions, not only convenient internet images. Include day and night scenes, seasonal variation, motion blur, partial occlusion, crowded layouts, empty scenes, and difficult negatives that resemble target objects.

    For an Indian deployment, test conditions such as harsh sunlight, monsoon reflections, dust, low-light streets, multilingual signage, dense pedestrian movement, and camera vibration. Split data by location and time rather than randomly by frame; adjacent video frames in both training and validation sets can produce misleadingly strong results.

    Label bounding boxes consistently and define edge cases in a short annotation policy. Decide how to label partially visible objects, truncated objects, reflections, mannequins, screens, and heavily blurred instances. Tools such as CVAT are useful for controlled annotation workflows. If the detector must support local-language signage or document-like visual content, pair the vision work with practices from low-resource Indic natural language processing.

    Track more than mAP. Report precision, recall, per-class performance, small-object recall, false alarms per camera-hour, and performance by lighting or location. A detector that performs well overall but misses helmets at night may be unsuitable for workplace safety.

    Select a model by workload, not popularity

    Single-stage detectors remain the practical default for low-latency applications. YOLO-family models offer a broad range of sizes, export options, and deployment tooling. Smaller variants are appropriate for mobile and low-power edge devices; larger variants improve accuracy when GPU capacity and latency allow.

    Consider the following alternatives:

    • YOLO-family detectors: Strong general-purpose choice with mature training and export workflows.
    • SSD and MobileNet variants: Useful for constrained mobile or embedded deployments where model simplicity matters.
    • RT-DETR-style detectors: Worth evaluating when accuracy and transformer-based architecture are priorities, provided the target runtime supports the required operators efficiently.
    • Segmentation or pose models: Prefer these when a box is insufficient—for example, measuring occupied floor area or detecting body posture.

    Benchmark the exact exported model on the target device. A model that is fast in PyTorch on a development workstation may be slower after conversion if unsupported operators trigger CPU fallbacks. Compare input sizes such as 416, 640, and 960 pixels, then select the smallest resolution that preserves critical recall. Upscaling a low-quality source cannot recover missing detail.

    Teams building reusable vision components can also review workflows for building computer vision models on GitHub, especially for dataset versioning, reproducible experiments, and export pipelines.

    Design the video pipeline as a system

    A reliable pipeline separates capture, decode, preprocessing, inference, post-processing, and output. Use bounded queues so a slow stage cannot consume unlimited memory. When the application values freshness over completeness, drop stale frames rather than allowing latency to grow indefinitely.

    A practical architecture is:

    1. Capture frames asynchronously from RTSP, USB, or WebRTC sources.
    2. Decode with hardware acceleration where available.
    3. Resize and letterbox using GPU or vectorized operations.
    4. Batch only when the latency budget allows it; batching improves throughput but can delay individual frames.
    5. Run inference in a dedicated worker.
    6. Apply confidence filtering and NMS, then track objects if temporal identity matters.
    7. Emit structured events and selected thumbnails rather than continuously uploading raw video.

    Keep timestamps from the camera and attach them to every detection. This makes it possible to distinguish model delay from network delay and to audit alerts later. For multi-camera deployments, use a stream framework such as NVIDIA DeepStream when its decoder, inference, and pipeline integrations match your hardware. Containerize the service and pin model, runtime, and driver versions before field deployment.

    Optimise inference and post-processing

    Export the trained model to the runtime used in production. NVIDIA deployments commonly use TensorRT with FP16 or INT8; Intel systems may benefit from OpenVINO; mobile and browser deployments often use ONNX Runtime, Core ML, NNAPI, or device-specific NPUs. Validate numerical output after conversion—speed gains are irrelevant if confidence scores or box coordinates change unexpectedly.

    Use calibration data that reflects production images for INT8 quantization. Inspect per-class recall after quantization, particularly for small or low-contrast objects. Fuse compatible layers, preallocate buffers, avoid unnecessary CPU-GPU copies, and profile memory transfers separately from kernel execution.

    Non-maximum suppression can become significant when scenes are crowded or input resolution is high. Use an optimized GPU implementation where available, limit candidate boxes before NMS, and tune confidence and IoU thresholds on validation footage. For dense scenes, class-agnostic NMS may remove valid overlapping objects, so treat it as a measured optimisation—not a universal shortcut.

    Add tracking only when the product needs it

    Detection answers “what is in this frame?” Tracking answers “is this the same object across frames?” If the product counts entries, measures dwell time, or prevents repeated alerts, add a tracker such as a Kalman-filter-based SORT variant or a stronger association method for crowded scenes.

    Tracking can reduce inference frequency: detect periodically, then track between detections. However, it can also propagate errors when an object is occluded or the camera moves. Evaluate identity switches, track loss, alert duplication, and end-to-end delay rather than assuming tracking automatically improves accuracy.

    Choose edge, cloud, or hybrid deployment

    Edge inference is usually preferable when video is sensitive, connectivity is unreliable, or response time matters. Jetson Orin devices, modern x86 systems with accelerators, mobile NPUs, and Coral-class TPUs can support different cost and power envelopes. Select hardware after measuring sustained performance, thermal throttling, storage, and field servicing—not just peak benchmark FPS.

    A hybrid design can send structured events, embeddings, or short evidence clips to the cloud while keeping continuous inference onsite. This reduces bandwidth and helps control operating costs across distributed Indian sites. Encrypt streams and event payloads, enforce device authentication, rotate credentials, and define retention policies before installation.

    If the system is part of a broader agent workflow, keep vision events separate from orchestration logic. Patterns from building distributed systems with AI agents can help with retries, queues, idempotent alerts, and service boundaries, but do not let an agent add unpredictable delay to a safety-critical control loop.

    Test field performance and monitor drift

    Create a replayable test set from recorded production streams. Test camera reconnects, packet loss, malformed frames, GPU memory pressure, model reloads, clock drift, and sudden lighting changes. Use load tests that match the intended number of cameras and concurrent users.

    Monitor:

    • End-to-end p50, p95, and p99 latency.
    • Effective FPS and dropped-frame rate.
    • Decoder, CPU, GPU, memory, temperature, and power usage.
    • Detections per class and false alerts per camera-hour.
    • Camera health, stream uptime, and reconnect time.
    • Drift in scene composition, lighting, camera angle, and object size.

    Add a human review path for uncertain detections and feed corrected examples into scheduled retraining. Version datasets, labels, model weights, calibration files, and runtime containers together so a field result can be reproduced.

    A practical build sequence

    Start with one camera and a narrow success criterion. Build a baseline model, measure the complete pipeline, and establish failure categories. Then optimise the bottleneck—often decoding, memory transfer, or stale queues rather than neural-network inference. Pilot across different sites, quantify false-alert cost, and only then scale camera count or model complexity.

    The strongest real-time systems are disciplined products: clear latency budgets, representative data, hardware-specific optimisation, safe failure behaviour, and observable operations. For Indian founders building such infrastructure, AI Grants India can be a starting point for exploring grant support and ecosystem resources.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.