0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build a quantized model for last mile delivery in india

How to Build a Quantized Model for Last-Mile Delivery in India

  1. aigi

    Last-mile delivery in India is a prediction and operations problem, not only a routing problem. A useful system must estimate delivery time, identify failed-delivery risk, prioritise orders, and support dispatchers while handling monsoon conditions, dense neighbourhoods, inconsistent addresses, traffic variability, and limited connectivity.

    Quantization makes these systems easier to run close to the operation. By converting floating-point weights and activations to lower-precision formats such as INT8, you can reduce model size, memory use, inference latency, and device costs. The goal is not to quantize blindly. The goal is to build a compact model that preserves the accuracy needed for a measurable business decision.

    Start with a narrow delivery decision

    Do not begin by asking for an “AI route optimiser”. Define one decision and one owner. Strong first use cases include:

    • ETA prediction: Estimate arrival time for each order using location, time, vehicle, route, and historical delivery signals.
    • Failed-delivery risk: Flag orders likely to fail because of unreachable customers, address ambiguity, cash-on-delivery issues, or restricted access.
    • Stop-time prediction: Estimate loading, parking, building-entry, and handover time at each stop.
    • Order batching: Decide which nearby orders can be grouped without violating delivery promises.
    • Demand forecasting: Predict neighbourhood-level order volume for rider and vehicle allocation.

    For a first deployment, ETA or stop-time prediction is usually more practical than end-to-end route optimisation. It has a clear target, can work alongside an existing routing engine, and produces metrics operations teams understand.

    Design an India-ready data pipeline

    A model is only as useful as the operational data behind it. Build a training table where each row represents a delivery event or route stop, with features available before the prediction is made. Avoid leakage, such as using actual arrival time or post-delivery status as an input.

    Useful fields include:

    • Pickup and drop coordinates, geohash, service zone, and approximate distance
    • Order creation time, promised slot, weekday, holiday, and festival period
    • Vehicle type, rider experience, delivery density, and route position
    • Historical travel time by road segment and time of day
    • Weather, rainfall, road closures, traffic, and local event signals
    • Building type, floor, lift availability, gated-community access, and parking difficulty
    • Address quality, landmark presence, phone reachability, and previous attempt history
    • Payment mode, cash-on-delivery status, and customer availability patterns

    India-specific address data deserves special treatment. Pin codes are too coarse for many decisions, while GPS points can be noisy around high-rises and informal settlements. Combine coordinates with geohashes, landmarks, locality names, and delivery-zone identifiers. If customer or rider interactions include regional languages, review the guidance on low-resource Indic natural language processing before adding text-derived features.

    Create a data contract covering schema, timestamps, location precision, consent, retention, and access controls. Remove unnecessary personal information, hash identifiers, and restrict raw phone numbers and addresses to systems that genuinely need them.

    Choose a model that fits the device

    Start with a baseline before selecting a sophisticated architecture. A gradient-boosted tree model, linear model, or small multilayer perceptron can perform well on structured delivery data and is often easier to explain than a large neural network.

    Match the model to the deployment target:

    • Cloud API: Useful when connectivity is reliable and central monitoring is the priority.
    • Rider smartphone: Requires low memory, low battery consumption, and graceful offline behaviour.
    • Warehouse or vehicle gateway: Supports local predictions with more compute than a phone.
    • Dispatcher workstation: Allows richer models but still benefits from low latency.

    Keep routing and prediction separate where possible. A route solver can generate candidate routes, while the quantized model estimates travel or stop time for each candidate. This modular design makes testing and replacement easier. If several services must coordinate decisions, document interfaces and failure handling as carefully as the model; systems built with distributed AI agents can otherwise become difficult to operate.

    Train and evaluate against operations metrics

    Split data by time, not randomly alone. Train on earlier weeks and test on later weeks so the evaluation reflects deployment and captures seasonal drift. Use separate evaluations for dense metros, tier-2 cities, rural edges, peak periods, rain, and new service zones.

    Track both model and business metrics:

    • Median absolute ETA error and the 80th or 90th percentile error
    • On-time delivery rate and promise-window violations
    • Failed-delivery rate and repeat-attempt rate
    • Rider kilometres, stops per hour, battery consumption, and API cost
    • Inference latency, memory use, model size, and offline success rate

    A model with a lower average error can still be worse if it performs poorly during evening peaks or systematically underestimates delivery time in certain neighbourhoods. Inspect errors by language, zone, vehicle type, weather, and address quality. Set minimum service thresholds before rollout rather than optimising only for aggregate accuracy.

    Quantize the model safely

    Quantization is a deployment step with accuracy and hardware trade-offs. Common approaches are:

    • Dynamic post-training quantization: Quantizes selected weights and calculates activation scales at runtime. It is a fast first experiment, especially for CPU inference.
    • Static post-training quantization: Uses a representative calibration dataset to quantize weights and activations. It often gives better latency and memory results.
    • Quantization-aware training: Simulates lower-precision arithmetic during training and is useful when post-training methods cause unacceptable accuracy loss.

    Use representative Indian delivery samples for calibration: peak and off-peak trips, different cities, rain, long-distance routes, failed attempts, and varied device conditions. Export a small set of candidate formats, such as INT8 CPU inference, and benchmark them on the exact target hardware rather than a developer laptop.

    TensorFlow Lite, PyTorch’s quantization tooling, and ONNX Runtime are practical options. Test operator support, preprocessing parity, threading, and fallback behaviour. A model may appear quantized while silently executing some operations in floating point, reducing the expected benefit.

    Compare the original and quantized models on the same frozen test set. Record accuracy change, p50 and p95 latency, peak memory, package size, battery impact, and cold-start time. Establish a release gate—for example, no material increase in high-risk ETA errors and no regression in on-time delivery—before promotion.

    Build for unreliable connectivity and human workflows

    For rider-facing applications, package the model and required preprocessing locally where privacy and size permit. Cache recent zone features, queue events offline, and synchronise when connectivity returns. Show confidence or an explanation that a dispatcher can act on; do not present uncertain predictions as facts.

    Operations teams should be able to override a recommendation, record why, and continue working when the model is unavailable. Treat dispatchers and riders as part of the feedback loop. Their observations about building access, temporary road closures, and address landmarks often expose data problems faster than dashboards do.

    Use clear interfaces with minimal text and support relevant Indian languages where the workflow needs it. If voice is useful for hands-free rider interaction, evaluate a focused voice agent architecture and deployment guide rather than adding a conversational layer without an operational purpose.

    Monitor drift, fairness, and security

    Deploy gradually: shadow mode first, then one zone or fleet, followed by a controlled expansion. Monitor:

    • Feature drift in distance, traffic, weather, order mix, and address quality
    • Accuracy by city, pin code, locality, vehicle type, and time window
    • Missing-data rates and preprocessing failures
    • Quantized-versus-reference prediction differences
    • Latency, crashes, battery use, and offline queue health
    • Changes in rider workload, customer complaints, and failed deliveries

    Retrain when drift or business thresholds require it, not only on a fixed calendar. Version datasets, calibration samples, model files, feature code, and rollback procedures together. Protect model endpoints, encrypt sensitive data, limit location retention, and audit access to customer information.

    A practical 90-day delivery plan

    • Weeks 1–2: Select one use case, define the KPI, map data sources, and establish privacy controls.
    • Weeks 3–5: Build the baseline, clean historical events, and create time-based evaluation sets.
    • Weeks 6–7: Train candidate models and analyse errors by zone, weather, vehicle, and delivery type.
    • Weeks 8–9: Quantize, calibrate, benchmark on target devices, and test offline behaviour.
    • Weeks 10–11: Run shadow mode with dispatcher review and measure operational impact.
    • Week 12: Launch a limited pilot with monitoring, rollback, and a retraining decision recorded in advance.

    The strongest result is not the smallest model. It is a model that delivers reliable predictions at the point of action, improves a defined logistics KPI, and remains maintainable as Indian delivery conditions change.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.