Mandi price systems have a deceptively hard job: return a useful answer despite fragmented data, changing markets, multilingual queries, and uneven connectivity. A quantized model can reduce latency, memory use, and serving cost—but quantization is only one part of a reliable product. The larger system must distinguish a reported market price from a forecast, preserve units and dates, and show the source behind every answer.
This guide explains how to build such a system for Indian agricultural markets, with practical choices for a mobile app, IVR service, WhatsApp bot, or lightweight API. For language and query coverage, pair the pricing model with the principles in this guide to low-resource Indic natural language processing.
Define the product before choosing a model
Start by deciding what “price query” means. Common requests include:
- “What is today’s tomato price in Azadpur?”
- “What was the modal price for onion in Nashik yesterday?”
- “Which nearby mandi is paying more for paddy?”
- “What could I receive next week?”
The first three are retrieval, filtering, ranking, or simple aggregation tasks. They may not require a neural model at all. The last is forecasting and needs a separate model, uncertainty estimate, and clearer wording. Do not generate a confident forecast when the user asked for an official reported price.
Define an output contract containing commodity, variety, mandi, state, date, unit, price type, currency, source, and freshness. In India, “price” may mean minimum, maximum, or modal price; users may also mix quintals, kilograms, and local measures. Normalize these internally, but display the original unit and conversion when it matters.
Build a trustworthy data pipeline
Use a reproducible ingestion layer rather than placing raw data directly into the model. Potential sources include government market feeds, mandi boards, cooperative records, and validated partner data. Record the source, retrieval timestamp, reporting date, and any correction history for each observation.
A useful canonical schema might include:
commodityandvarietystate,district,market, and stable location IDsarrival_dateandpublished_atmin_price,max_price, andmodal_priceunit, currency, and conversion factorquantity_arrivedwhere availablesource, confidence, and data version
Clean duplicate records, impossible values, spelling variants, and sudden unit changes. Keep missing values distinct from zero. Map local names and transliterations—such as “pyaz”, “onion”, and regional equivalents—to a controlled vocabulary. Preserve the original query and the resolved interpretation so incorrect entity matches can be audited.
Split evaluation data by time, not randomly. A random split can leak future seasonal patterns into training and produce an unrealistic score. Hold out recent weeks or months, and include difficult slices: low-volume mandis, festival periods, monsoon disruptions, new varieties, and states absent from training.
Choose the smallest model that solves the task
For exact price retrieval, use a database or search index with deterministic filters. For ranking nearby mandis, a gradient-boosted model or lightweight learning-to-rank system may be sufficient. For forecasting, begin with strong baselines such as seasonal naive, moving average, and gradient boosting before trying an LSTM or transformer.
A practical architecture often separates responsibilities:
1. Query understanding extracts commodity, location, date, variety, and price type.
2. Data retrieval fetches authoritative observations.
3. Forecasting or ranking runs only when required.
4. Response generation formats the result, source, freshness, and caveat.
This hybrid approach is safer and cheaper than asking one language model to invent the entire answer. It also makes it easier to deploy an edge model for intent and entity extraction while keeping current prices on a server. If users interact by speech, the same architecture can sit behind a voice agent built for Indian deployments.
Establish a baseline and metrics
Before quantization, freeze a full-precision baseline. Track more than aggregate MAE. Recommended metrics include:
- MAE and RMSE for price forecasts, reported in rupees per stated unit.
- WAPE or sMAPE for comparisons across commodities with different price scales.
- Top-k retrieval accuracy for mandi recommendations.
- Entity resolution accuracy for commodity, variety, and location.
- Freshness and coverage for reported prices.
- Latency, memory, and energy on the target device or server.
Evaluate by commodity, language, state, mandi size, and date range. A small overall accuracy loss may be unacceptable for one crop or region. Set release thresholds before quantization—for example, no more than a defined MAE increase and no critical regression in low-resource languages.
Apply INT8 quantization carefully
For a neural model, post-training dynamic-range quantization is a fast first experiment. Full integer quantization generally provides better edge performance but requires a representative calibration set. Select calibration examples from real queries and feature distributions across commodities, seasons, states, and languages. Do not calibrate only on the largest mandi.
A TensorFlow Lite workflow can look like this:
import tensorflow as tf
model = tf.keras.models.load_model("mandi_model.keras")
converter = tf.lite.TFLiteConverter.from_keras_model(model)
converter.optimizations = [tf.lite.Optimize.DEFAULT]
# For full INT8 deployment, provide representative_data_gen and set:
# converter.representative_dataset = representative_data_gen
# converter.target_spec.supported_ops = [tf.lite.OpsSet.TFLITE_BUILTINS_INT8]
# converter.inference_input_type = tf.int8
# converter.inference_output_type = tf.int8
quantized_bytes = converter.convert()
with open("mandi_model_int8.tflite", "wb") as file:
file.write(quantized_bytes)Validate the converter’s input and output scales, zero points, tensor types, and unsupported operators. Quantization can affect embeddings, attention layers, rare categorical values, and outlier-heavy numeric features. If accuracy falls sharply, try dynamic quantization, selectively retain sensitive layers at higher precision, improve calibration, or use quantization-aware training (QAT). Always compare predictions on the exact same test set as the baseline.
Deploy for Indian connectivity and cost constraints
For a mobile or kiosk product, package the model with an on-device runtime such as TensorFlow Lite or ONNX Runtime Mobile. Keep current prices and updates server-side unless the product has a reliable synchronization strategy. For an API, expose a versioned endpoint such as GET /v1/prices, return structured JSON, and include observed_at, source, and model_version fields.
Design for practical conditions:
- Cache recent queries by commodity, mandi, and date.
- Return a useful fallback when a market feed is delayed.
- Support low-bandwidth JSON and text responses.
- Use Indian language labels and local numerals only when requested.
- Encrypt data in transit and restrict access to sensitive user or partner data.
- Log model version, retrieved records, and final response for auditability.
If the product serves the next billion users, test on entry-level Android devices and intermittent networks; the broader principles in building AI apps for India’s next billion users are directly relevant.
Monitor drift and prevent harmful answers
Market behavior changes with weather, policy, arrivals, transport costs, and crop cycles. Monitor missing-feed rates, distribution drift, stale responses, latency, and error by slice. Set alerts for impossible prices, abrupt unit changes, and confidence deterioration.
A response should say “reported modal price” or “model forecast”, never blur the two. Show the observation date and source. For forecasts, provide a range where possible and state the horizon. Add a human escalation path for disputed records. In a multilingual system, review translations and speech recognition errors separately; a correctly quantized model cannot fix a wrongly resolved mandi name.
A practical release checklist
Before production, confirm that you have:
- A time-based holdout and commodity/state/language slice evaluation.
- Full-precision versus INT8 accuracy, latency, memory, and energy comparisons.
- Calibration data representing rare markets and seasonal conditions.
- Versioned data, model, feature, and schema artifacts.
- Source and freshness fields in every price response.
- Tests for units, dates, missing values, transliterations, and adversarial inputs.
- Rollback and retraining procedures when a feed or model fails.
Quantization should be treated as an engineering optimization, not a claim of accuracy. Build the data and retrieval layer first, measure a credible baseline, then quantize against a fixed acceptance threshold. That sequence produces a mandi price service that is faster on constrained hardware without sacrificing traceability or farmer trust.