Small retailers do not need a large cloud model to make better inventory decisions. A compact model running on a billing computer, Android phone, or low-cost edge device can forecast demand, flag likely stock-outs, and work through unreliable connectivity. The engineering challenge is to build for messy retail data, mixed-language workflows, and strict cost limits—not simply to compress a model.
This guide explains how to build a quantized model for kirana stores in a way that is measurable and deployable in India.
Start with one operational decision
Choose a narrow use case before collecting data or selecting a model. Good first projects include:
- Next-day or next-week demand forecasting for fast-moving items.
- Stock-out alerts based on current inventory, sales velocity, and supplier lead time.
- Reorder recommendations that account for pack size, minimum order quantity, and cash constraints.
- Receipt or shelf recognition using a camera, if manual stock entry is the main bottleneck.
- Voice-assisted billing or stock lookup for shopkeepers who prefer Hindi, Tamil, Bengali, or another local language.
Avoid starting with a general-purpose “AI assistant”. A model that recommends replenishment for 100 high-volume SKUs can create more value, and produce clearer evaluation results, than an oversized system covering every store process. For broader product design, see this guide to building AI apps for the next billion users in India.
Define the decision, user, response time, and acceptable error in advance. For example: “Every morning, recommend quantities for the top 100 SKUs, with inference under two seconds and fewer than 10% missed stock-outs.”
Build a reliable retail dataset
Kirana data is often spread across billing software, notebooks, WhatsApp orders, distributor invoices, and handwritten stock counts. Begin with a simple, consistent schema:
- Store ID and product ID.
- Timestamp, quantity sold, selling price, discount, and return status.
- Current stock, wastage, expiry, and stock adjustments.
- Supplier, lead time, pack size, and minimum order quantity.
- Day of week, festival period, payday period, weather proxy, and local events where relevant.
Keep a product master table. Record aliases, spelling variants, brand, category, unit, and local-language names. “Amul milk 500 ml”, “amul 1/2 litre”, and a regional shorthand should resolve to one product ID. Do not train on customer phone numbers or unnecessary personal information. Hash identifiers, restrict access, and establish a deletion process.
For language-enabled interfaces, collect consented examples of real queries and code-switching. A retailer may say, “Kal Maggi ka stock kitna hai?” rather than use one language consistently. Low-resource language handling benefits from the methods described in this builder’s guide to Indic natural language processing.
Select the smallest model that solves the problem
For tabular demand forecasting, start with a strong baseline before trying a neural network:
- Seasonal naive forecasting for comparison.
- Regularised linear regression for simple trends.
- Gradient-boosted trees for nonlinear demand patterns.
- A small multilayer perceptron or temporal model when you have enough clean history.
Quantization is most valuable when the target model uses floating-point weights and activations, such as a small neural network, image model, or language model. If a gradient-boosted model already meets latency and memory requirements, compression may not be necessary.
Choose deployment hardware early. An Android phone, Raspberry Pi-class device, POS terminal, and cloud API have different supported runtimes and numeric kernels. Measure on the actual device, not only on a developer laptop.
Train and validate without leaking future information
Split sales data by time, not randomly. Train on earlier weeks, validate on a later period, and reserve the most recent period for a final test. A random split can place future promotions or the same demand pattern in both sets and produce misleading results.
Useful metrics depend on the decision:
- MAE or weighted MAE for average unit error.
- WAPE for store-level demand comparisons.
- Stock-out recall for detecting items likely to run out.
- Overstock rate and estimated working capital tied up.
- Service level: the share of demand fulfilled.
- Latency, RAM use, model size, and battery impact on the target device.
Evaluate separately for staples, perishables, slow-moving products, promotions, and festival periods. A model with a good average score can still fail badly on milk, bread, or regional bestsellers.
Quantize the model
There are three practical approaches:
1. Dynamic-range post-training quantization converts weights to lower precision while selecting activation ranges at runtime. It is a fast first experiment.
2. Full integer post-training quantization converts weights and activations, commonly to INT8. It usually offers stronger edge performance but requires representative calibration data.
3. Quantization-aware training (QAT) simulates low-precision arithmetic during training. Use it when post-training quantization causes unacceptable accuracy loss.
For a TensorFlow Lite workflow, export the trained model, provide a representative dataset covering typical inputs, convert to INT8, and verify the input and output tensor types. For PyTorch, compare an appropriate mobile or edge export path, then validate the resulting runtime. ONNX Runtime can be useful when you need portability across supported devices.
Do not assume a smaller file guarantees faster inference. Benchmark all of the following:
- Original FP32 model.
- Weight-only or dynamic quantized model.
- Fully INT8 model.
- QAT model if needed.
Check operator support, fallback to floating point, numerical range, and batch size one. A model that silently falls back to FP32 may save storage but deliver little latency or energy benefit.
Design for intermittent connectivity
A kirana deployment should continue basic operations when the network is unavailable. Keep prediction, product lookup, and the last known inventory state on-device where practical. Synchronise transactions later using an append-only event queue and resolve conflicts with explicit timestamps and business rules.
Use a small API only for tasks that genuinely need the cloud, such as fleet-wide retraining, analytics, or model distribution. Sign model files, verify versions before installation, encrypt sensitive data in transit and at rest, and maintain a rollback path. If voice is part of the product, keep commands short and provide a visible confirmation before changing stock or placing an order. See the voice agent architecture and deployment guide for the components involved in reliable speech workflows.
Run a field pilot
Pilot with a small group of stores representing different formats: neighbourhood grocery shops, semi-urban outlets, stores with high perishables, and shops using different billing systems. Run the current process and AI-assisted process in parallel before automating reorder decisions.
Track operational outcomes, not just model scores:
- Stock-outs per 100 key products.
- Wastage and expiry value.
- Time spent on stock counting.
- Recommendation acceptance and override rate.
- Sales uplift, gross margin, and cash tied up in inventory.
- App crashes, sync failures, and daily active usage.
A high override rate may indicate poor recommendations, unclear explanations, or an inconvenient workflow. Let shopkeepers record a reason for overrides; those labels become valuable training data.
Monitor drift and improve safely
Retail patterns change with seasons, local festivals, price changes, new competitors, and supplier disruptions. Monitor input distributions, missing fields, product mapping failures, forecast error, and outcome metrics by store and category. Retrain on a schedule only after checking data quality; frequent retraining on corrupted stock counts can make the model worse.
Use staged releases: deploy to a small percentage of stores, compare against the previous model, and keep the last known-good version available. Document model version, training window, quantization method, calibration set, hardware, and evaluation results. This makes debugging possible when a recommendation fails in the field.
A practical minimum viable stack
A lean first implementation could include:
- POS export or CSV ingestion with a product master.
- A seasonal baseline and a small forecasting model.
- Python for training, with TensorFlow Lite, ONNX Runtime, or a suitable mobile runtime for inference.
- INT8 post-training quantization with representative calibration data.
- Local SQLite storage, queued sync, and signed model updates.
- A simple Hindi-English or regional-language interface with human confirmation.
- A dashboard showing stock-outs, forecast error, latency, and recommendation overrides.
Start with measurable inventory outcomes, prove value on real devices, and only then expand into computer vision, voice, or multi-store optimisation. For image-based shelf counting, compare your approach with computer vision model building on GitHub, especially around dataset labelling and edge inference.
FAQ
Does every kirana AI model need quantization?
No. Quantize when memory, latency, bandwidth, or battery limits affect the product. Benchmark against a simpler unquantized baseline first.
Should I use INT8 or FP16?
INT8 often gives better memory and CPU efficiency, while FP16 may be easier on hardware with GPU or accelerator support. Test both on the actual device.
How much data is enough?
There is no universal threshold. Start with several months of clean, timestamped sales for fast-moving products, then measure performance by category and store. Sparse products may need category-level or hierarchical methods.
Can the model work without internet?
Yes, if the model, product data, and inference runtime are stored locally. Synchronisation, retraining, and fleet analytics can happen when connectivity returns.