What you can predict—and at what scale
A useful model starts with a precise target. “Predict Rabi crops” can mean identifying which crop is planted, estimating yield, forecasting cropped area, or detecting crop stress. These are different machine-learning problems:
- Crop classification: wheat, mustard, barley, gram, fodder, fallow or other classes.
- Yield regression: tonnes per hectare or total production for a field, village, block or district.
- Area estimation: the number of hectares under each crop.
- Early-season forecasting: a prediction made before harvest using only observations available by a defined date.
For an initial Haryana pilot, crop classification at field or parcel level is often more achievable than yield prediction. Define the unit of prediction—field, 10-metre pixel, village or district—and the forecast cut-off date before collecting data. A model that uses March imagery cannot be presented as an early-season forecast if the intended decision must be made in January.
Build a Haryana-specific dataset
Use a consistent geography and season. Haryana’s Rabi landscape includes wheat, mustard, gram, barley and fodder, with meaningful differences across irrigation access, soil, sowing dates and district. Start with a small set of districts and expand only after the pipeline is stable.
Recommended data layers include:
- Sentinel-2 Level-2A: 10-metre optical bands and frequent revisit for crop signals.
- Landsat: useful for longer historical series and filling gaps in a multi-year archive.
- Weather: rainfall, minimum and maximum temperature, humidity, solar radiation and reference evapotranspiration.
- Soil and terrain: texture, organic carbon, pH, elevation and drainage where available.
- Irrigation and administrative data: canals, groundwater dependence, village boundaries and procurement or crop-survey records.
- Ground truth: geotagged field observations, crop-cutting records, farmer surveys or verified crop declarations.
Your labels are more important than model complexity. Record crop type, observation date, field geometry, source, confidence and season. Avoid treating an unverified crop label as fact. For high-stakes agricultural decisions, apply the same discipline described in data veracity infrastructure for high-stakes AI: preserve provenance, flag conflicts and maintain an auditable correction process.
Preprocess imagery without leaking future information
A practical workflow can be built with Google Earth Engine for collection and compositing, or with Python, Rasterio, GeoPandas and STAC-compatible catalogues for a reproducible local pipeline.
1. Filter imagery to the Rabi season and area of interest.
2. Apply the sensor’s cloud and cirrus mask, then remove poor-quality pixels.
3. Harmonise reflectance, projection, resolution and band names across dates.
4. Create weekly or fortnightly composites using median or quality-weighted mosaics.
5. Clip each composite to field boundaries and calculate summary statistics.
6. Keep acquisition dates and processing versions alongside every feature.
Do not fill missing observations with values from after the prediction date. That is temporal leakage. Likewise, do not randomly split neighbouring pixels from the same field into training and test sets; the model may simply learn field-specific patterns. For repeatable pipelines, combine geospatial processing with Python scripts for automating data preprocessing, including tests for missing bands, invalid geometries, duplicate labels and unexpected date ranges.
Engineer features that represent crop development
Raw bands are rarely the best input for XGBoost. Build features that capture both crop condition and its trajectory:
- Spectral indices: NDVI, EVI, NDWI, red-edge chlorophyll indices and bare-soil indices.
- Temporal statistics: median, maximum, minimum, slope, amplitude and date of peak vegetation signal.
- Phenology: cumulative growing-degree proxies, green-up timing and duration above an NDVI threshold.
- Spatial summaries: field-level median, percentiles, variance and the share of anomalous pixels.
- Weather windows: rainfall totals, temperature summaries and heat-stress days over defined periods.
- Context features: soil class, irrigation proxy, district, field area and previous-season land cover.
Calculate features only from data available at the forecast date. For classification, the target might be crop_type; for yield regression, use a measured yield value and document its measurement method. Missing values should be meaningful: retain a missingness indicator where cloud gaps or unavailable soil information carry operational significance rather than silently replacing everything with a global average.
Train an XGBoost model correctly
XGBoost is well suited to mixed, tabular features and nonlinear interactions. It does not require feature standardisation, although consistent units and sensible value ranges remain essential. A baseline implementation might look like this:
from xgboost import XGBClassifier
from sklearn.metrics import classification_report
model = XGBClassifier(
n_estimators=600,
max_depth=6,
learning_rate=0.05,
subsample=0.8,
colsample_bytree=0.8,
objective="multi:softprob",
eval_metric="mlogloss",
tree_method="hist",
random_state=42,
)
model.fit(
X_train, y_train,
eval_set=[(X_valid, y_valid)],
verbose=False,
)
predictions = model.predict(X_test)
print(classification_report(y_test, predictions))For yield, use XGBRegressor with an appropriate objective and inspect the target distribution before training. Tune a limited parameter grid rather than blindly searching hundreds of combinations. Useful controls include max_depth, min_child_weight, learning_rate, subsample, colsample_bytree, regularisation terms and early stopping.
Validate by season, geography and decision date
An 80/20 random split is not enough for satellite agriculture. Use grouped and time-aware validation:
- Spatial holdout: reserve entire fields, villages or districts for testing.
- Seasonal holdout: train on earlier Rabi seasons and test on a later one.
- Rolling forecast evaluation: replicate the information available at January, February and March cut-offs.
- Stratified reporting: break results down by crop, district, irrigation status and label confidence.
For classification, report macro-F1, per-class recall, balanced accuracy and a confusion matrix. Overall accuracy can hide poor performance on smaller crops such as gram or barley. For yield, report MAE, RMSE, bias and percentage error, alongside uncertainty intervals where possible. Compare XGBoost with simple baselines such as last-season crop maps, majority class and linear regression. A complex model is useful only when it beats a credible operational baseline.
Interpret, map and deploy the results
Use SHAP values or permutation importance to check whether the model relies on plausible signals. If a district code dominates predictions, the model may be memorising geography rather than learning transferable crop characteristics. Test performance after removing suspicious features.
Export predictions with field ID, crop probabilities or yield estimate, forecast date, model version and confidence flag. Display them as maps and tables for agriculture officers, researchers and farmer organisations. Tools for AI data visualization design can help communicate uncertainty, but every dashboard should show the source date and avoid presenting a probability as a guarantee.
A production system needs monitoring: track cloud coverage, feature drift, label delays, district-wise error and changes in crop patterns. Recalibrate or retrain when performance declines. Keep a human review route for low-confidence fields and disputed labels.
Haryana deployment checklist
Before using predictions for procurement, advisories, insurance or resource allocation, confirm that you have:
- A documented target, forecast date and decision use case.
- Multi-season labels with field-level provenance.
- Spatial and temporal holdout results, not only random-split accuracy.
- Crop-wise error metrics and confidence thresholds.
- A process for correcting labels and handling farmer disputes.
- Privacy controls for farmer, landholding and location data.
- An exportable model card covering limits, training period and known failure cases.
For smaller teams, begin with a district-level pilot and a transparent dashboard rather than promising nationwide field-level precision. A lightweight workflow can also use no-code data analytics platforms in India for exploratory reporting while the geospatial and modelling pipeline remains reproducible in code.
Frequently asked questions
Can Sentinel-2 alone predict Rabi crops? It can provide strong crop signals, but performance improves when imagery is combined with weather, field boundaries, soil and verified labels. Cloud gaps and mixed pixels remain important limitations.
Should I use NDVI only? No. NDVI is useful but can saturate in dense crops and cannot capture every crop distinction. Add red-edge, moisture, temporal and contextual features, then validate their contribution.
How much ground truth is needed? There is no universal number. Diverse, accurately geolocated samples across crops, districts, sowing dates and irrigation conditions are more valuable than many duplicate samples from one village.
Can the model predict yield from satellite data? Yes, but yield labels are noisier than crop labels. Report uncertainty, validate against independent crop-cutting or harvest records, and avoid making field-level claims from district-level training data.
If your project builds reliable agricultural AI for India, document the evidence, deployment safeguards and measurable farmer benefit clearly when exploring support through AI Grants India.