Basmati yield prediction is useful only when it is accurate at the right place and time. A model trained on one district, season or irrigation pattern may fail in another, so the goal is not simply to maximise an R² score. The goal is to produce a dependable estimate of yield—typically tonnes per hectare or kilograms per hectare—that agronomists, procurement teams, extension workers and farmers can act on.
This guide explains how to use random forest models to predict basmati rice yield in Haryana. It covers data design, leakage-resistant validation, model training, interpretation and the operational checks needed before using predictions in the field.
Define the prediction problem first
Start by fixing four decisions:
- Target: harvested basmati yield per hectare, with units recorded consistently.
- Prediction date: before sowing, during tillering, at flowering or several weeks before harvest.
- Geographic unit: plot, village, block or district. Do not mix units without recording the hierarchy.
- Forecast horizon: the number of days between the latest input and harvest.
A district-level model may support procurement planning, while a plot-level model can guide irrigation or nutrient decisions. These are different products and require different datasets. Record the crop variety, transplanting or direct-seeding method, sowing date, harvest date and whether yield was measured or self-reported.
For Haryana, preserve district and season information. Conditions in Karnal, Kurukshetra, Kaithal, Ambala and adjoining rice-growing areas can differ in soil, irrigation access, monsoon timing and management practices. These fields are not merely administrative labels; they help test whether a model generalises across production environments.
Assemble a defensible dataset
Useful features usually fall into five groups:
- Weather: rainfall totals and dry-spell counts, minimum and maximum temperature, humidity, solar radiation and heat-stress indicators.
- Soil: texture, pH, electrical conductivity, organic carbon, available nitrogen, phosphorus and potassium, plus drainage characteristics.
- Crop and management: variety, planting date, plant population, fertiliser applications, irrigation events, weed control and pest pressure.
- Remote sensing: vegetation indices such as NDVI or EVI, canopy temperature, radar backscatter and temporal summaries from satellite imagery.
- Location and history: district, soil zone, previous crop, previous yield and field elevation where available.
Use a consistent row definition. For example, one row can represent one field in one kharif season, with weather and satellite variables aggregated only up to the intended forecast date. This prevents information from after the prediction date leaking into training.
Potential sources include farm records, agricultural universities, state departments, weather stations, gridded weather products and satellite platforms. Create a data dictionary that states each variable’s unit, time window, source, missing-value code and collection method. A small, well-documented dataset is more valuable than a large table whose measurements cannot be reproduced.
If satellite imagery is part of the pipeline, an understanding of how to build computer vision models on GitHub can help with reproducible preprocessing, although Random Forest itself does not require deep learning.
Prepare features without leaking information
Random Forests do not generally require feature scaling, so normalisation is not essential. Cleaning and temporal alignment matter much more.
1. Remove duplicate field-season records and investigate implausible yields.
2. Standardise units, dates, district names and categorical labels.
3. Impute missing values using rules that would have been available at prediction time.
4. Add meaningful aggregates, such as cumulative rainfall from sowing to prediction date or average NDVI during tillering.
5. Encode categorical variables with one-hot encoding, ordinal encoding or native categorical handling in an appropriate library.
6. Keep a final untouched test set from later seasons or held-out districts.
Do not calculate a seasonal weather total that includes rainfall after the forecast date. Do not use final harvest biomass, post-harvest yield estimates or a revised management record as an input. Leakage can create an impressive score and a useless field tool.
Train a baseline Random Forest regressor
A Random Forest regressor averages predictions from many decision trees trained on varied bootstrap samples and feature subsets. It handles nonlinear relationships and interactions—such as heat stress becoming more damaging under limited irrigation—without requiring a precise equation.
pip install pandas scikit-learn numpy joblibA minimal training example is:
import pandas as pd
from sklearn.ensemble import RandomForestRegressor
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
from sklearn.model_selection import GroupShuffleSplit
# One row per field-season; yield_t_ha is the target
df = pd.read_csv("haryana_basmati_yield.csv")
features = ["rain_mm_to_date", "tmax_c_mean", "ndvi_tillering",
"soil_ph", "available_n", "irrigation_count"]
X = df[features]
y = df["yield_t_ha"]
groups = df["season"] # keep records from the same season together
splitter = GroupShuffleSplit(n_splits=1, test_size=0.2, random_state=42)
train_idx, test_idx = next(splitter.split(X, y, groups=groups))
model = RandomForestRegressor(
n_estimators=500,
max_features="sqrt",
min_samples_leaf=2,
random_state=42,
n_jobs=-1
)
model.fit(X.iloc[train_idx], y.iloc[train_idx])
pred = model.predict(X.iloc[test_idx])
print("MAE:", mean_absolute_error(y.iloc[test_idx], pred))
print("RMSE:", mean_squared_error(y.iloc[test_idx], pred) ** 0.5)
print("R2:", r2_score(y.iloc[test_idx], pred))Treat these settings as a starting point, not a universal recipe. Tune n_estimators, max_depth, max_features, min_samples_leaf and max_samples with cross-validation. Larger forests usually improve stability but increase training and inference cost.
Validate for Haryana conditions
A random row-wise 70:30 split can overstate performance when neighbouring fields or repeated seasons appear in both sets. Prefer validation that reflects deployment:
- Leave-one-season-out: tests performance in a new kharif season.
- Leave-one-district-out: tests geographic transfer.
- Group-based splits: keep all records from a field, village or farm in one partition.
- Time-based testing: train on earlier seasons and test on a later season.
Report MAE in the original yield unit, RMSE to expose large misses, and R² as a supplementary measure. Also report errors by district, variety, irrigation category and yield range. A model with good average accuracy may still systematically underpredict high-yield fields or perform poorly for rainfed plots.
Compare Random Forest with a simple historical mean, linear regression and boosted-tree baseline. If the forest barely beats a baseline, improve data quality before adding complexity. For operational decisions, prediction intervals or quantile estimates are often more useful than a single number.
Interpret and deploy the model responsibly
Feature importance can indicate which variables deserve investigation, but impurity-based importance can favour high-cardinality or continuous features. Use permutation importance on held-out data and, where appropriate, partial-dependence or SHAP analysis. Treat these outputs as associations, not proof that changing one variable will cause a yield increase.
Before deployment, save the complete feature-generation logic, model version, training period and data provenance. Monitor missingness, feature distributions and errors each season. Trigger review when satellite coverage drops, weather inputs are delayed, a new variety is introduced or predictions fall outside the training range.
A lightweight dashboard can show predicted yield, uncertainty, forecast date and the main contributing variables. Keep a human review step for extension and procurement decisions. This is especially important when predictions affect credit, input recommendations or farmer payments.
The same reproducibility discipline applies to other Indian AI projects, including AI predictive maintenance for railway infrastructure assets. If you later add local-language farmer interfaces, consider open-source small language models for Hindi for explanation and translation—but keep the yield model and its evidence separate from generated text.
Practical checklist
- Define the forecast date and target unit before collecting data.
- Keep post-forecast information out of every feature.
- Split by season, location or field rather than only by row.
- Compare against simple baselines.
- Report MAE and error ranges in tonnes per hectare.
- Test performance across districts, varieties and irrigation regimes.
- Store the data dictionary, preprocessing code and model version.
- Communicate uncertainty and retain agronomist oversight.
Random Forest is a strong first model for Haryana basmati yield forecasting because it works well with mixed, nonlinear agricultural data and is relatively straightforward to inspect. Its value, however, comes from disciplined measurement and validation—not from the algorithm alone.