Weather prediction for a stadium is not simply a matter of training a model on temperature and humidity. At Arun Jaitley Stadium in New Delhi, forecasts must account for monsoon rainfall, winter fog and pollution, intense pre-monsoon heat, short-duration thunderstorms, and changing conditions across a dense urban area. A useful model should support decisions such as pitch-cover readiness, player and spectator safety, broadcast planning, and event staffing—not just produce a single number.
This guide shows how to use XGBoost to predict weather in Arun Jaitley Stadium using a reproducible, time-aware workflow. The examples focus on short-term prediction, such as estimating temperature, rainfall probability, or humidity several hours ahead. For operational deployment, combine the model with official forecasts and on-site observations rather than treating it as a replacement for the India Meteorological Department (IMD).
Define the prediction task first
Choose one target and forecast horizon before collecting data. Different targets require different labels, features, and evaluation metrics:
- Temperature regression: predict air temperature one, three, or six hours ahead.
- Rainfall regression: predict accumulated rainfall in the next hour, although many observations will be zero.
- Rain classification: predict whether measurable rain will occur within a defined window.
- Event-risk classification: flag conditions such as heavy rain, extreme heat, poor visibility, or high wind.
A practical first project is binary rainfall prediction for the next three hours. Define the label clearly—for example, rain_next_3h = 1 if rainfall exceeds 0.1 mm during the following three hours. Avoid vague labels such as “bad weather”; they cannot be audited or improved reliably.
For broader guidance on production-grade feature stores, monitoring, and retraining, see implementing scalable ML pipelines for predictive analytics.
Assemble data for the stadium location
The stadium is in central Delhi, but the nearest weather station may not represent conditions inside the ground. Start with a stable location identifier and document the latitude, longitude, elevation, source, timestamp standard, and measurement interval for every dataset.
Useful inputs include:
- Historical observations: temperature, relative humidity, rainfall, wind speed, wind direction, pressure, cloud cover, visibility, and dew point.
- Forecast data: numerical weather prediction variables and hourly forecast updates available at prediction time.
- Radar or satellite signals: precipitation nowcasts and cloud movement, when legally and technically accessible.
- Air-quality and visibility data: particulate matter, haze, and visibility can be valuable for winter operations, but should be treated as separate environmental signals.
- Stadium observations: an automated weather station or calibrated sensors installed near the playing area provide the most relevant local measurements.
Potential sources include IMD datasets and services, a reputable commercial weather API, ERA5 or other reanalysis products, and weather-station networks. Check licensing before using data in a public or commercial system. Align all records to Indian Standard Time (IST), retain the original UTC timestamp where available, and record API retrieval time so that future leakage can be detected.
A weather model benefits from the same disciplined data lineage used in industrial forecasting and AI predictive maintenance for railway infrastructure assets: every feature should have a known source, timestamp, unit, and expected update delay.
Engineer features without leaking the future
XGBoost performs well on structured tabular data, but feature quality matters more than adding a large number of columns. Useful features include:
- Calendar variables: hour, month, day of year, weekend, and match or event status.
- Lagged observations: temperature, humidity, pressure, wind, and rainfall from one, three, six, and 24 hours earlier.
- Rolling statistics: three-hour and 24-hour rainfall totals, rolling humidity averages, pressure change, and recent temperature range.
- Cyclical time features:
sinandcostransformations for hour and day of year, which avoid treating 23:00 and 00:00 as far apart. - Forecast variables: predicted precipitation, cloud cover, wind gust, and temperature available at the moment the model makes its prediction.
- Interaction features: humidity combined with temperature, wind direction represented as sine and cosine, and month-by-hour combinations.
The most important rule is availability-time discipline. At 14:00, the model must use only measurements and forecasts available by 14:00. Do not calculate a rolling value using observations from 15:00, and do not randomly split rows when adjacent observations come from the same weather event.
Prepare a time-aware training dataset
Clean units and timestamps before modelling. Convert rainfall to millimetres, wind speed to a consistent unit, and categorical weather descriptions to controlled values. Investigate sensor outages separately from genuine zero rainfall. For missing values, use forward filling only where it reflects real operational availability; otherwise add a missingness indicator and let the model learn that the value was unavailable.
Split chronologically, for example:
- Training: the earliest 70% of dates.
- Validation: the next 15%.
- Test: the most recent 15%.
A stronger approach is walk-forward validation: train on an expanding historical window, validate on the next block, and repeat across seasons. This reveals whether the model works during Delhi's monsoon and winter regimes rather than only during the dominant season. Compare XGBoost with a persistence baseline (“the next hour resembles the current hour”) and an official forecast. A sophisticated model that cannot beat these baselines is not ready for deployment.
Train an XGBoost model in Python
Install the core packages:
pip install xgboost pandas scikit-learn matplotlibFor rainfall probability, use a classifier. The example below assumes df contains a timestamp, engineered features, and a target named rain_next_3h.
import pandas as pd
from xgboost import XGBClassifier
from sklearn.metrics import (
roc_auc_score, average_precision_score,
precision_recall_curve, classification_report
)
features = [
"temperature", "humidity", "pressure", "wind_speed",
"rain_1h", "rain_3h", "rain_24h", "pressure_change_3h",
"hour_sin", "hour_cos", "doy_sin", "doy_cos",
"forecast_precipitation", "forecast_cloud_cover"
]
df = df.sort_values("timestamp").dropna(subset=features + ["rain_next_3h"])
cutoff = df["timestamp"].quantile(0.85)
train = df[df["timestamp"] < cutoff]
test = df[df["timestamp"] >= cutoff]
model = XGBClassifier(
objective="binary:logistic",
n_estimators=500,
max_depth=6,
learning_rate=0.05,
subsample=0.8,
colsample_bytree=0.8,
min_child_weight=5,
reg_lambda=2.0,
eval_metric="auc",
tree_method="hist",
random_state=42
)
model.fit(
train[features], train["rain_next_3h"],
eval_set=[(test[features], test["rain_next_3h"])],
verbose=False
)
probability = model.predict_proba(test[features])[:, 1]
print("ROC-AUC:", roc_auc_score(test["rain_next_3h"], probability))
print("PR-AUC:", average_precision_score(test["rain_next_3h"], probability))For temperature, replace the classifier with XGBRegressor and evaluate MAE and RMSE in degrees Celsius. For rainfall amounts, consider a two-stage system: classify rain versus no rain, then train a regression model only on positive-rain cases. This handles the large number of zero observations better than applying ordinary regression to every row.
Evaluate for decisions, not just scores
Weather events are often imbalanced: most hours may have no heavy rain. Accuracy can therefore be misleading. Use ROC-AUC and PR-AUC for ranking, then select a probability threshold based on operational costs. If missing a rain event is more expensive than issuing an alert, choose a threshold that improves recall while measuring the resulting false-alarm rate.
Also check:
- Calibration: when the model says 70% rain probability, does rain occur about 70% of the time?
- Seasonal performance: compare summer, monsoon, post-monsoon, and winter results.
- Lead-time performance: test one-, three-, and six-hour horizons separately.
- Event-day performance: evaluate match days and high-attendance events independently.
- Error by intensity: distinguish drizzle, moderate rain, and heavy downpours.
Plot reliability curves, confusion matrices, predicted-versus-observed time series, and feature importance. SHAP explanations can help operators understand whether an alert was driven by rising humidity, pressure drops, radar precipitation, or recent rainfall. Explanations should support investigation, not imply that weather causation has been proven.
Deploy with safeguards
A dependable stadium service should ingest new observations on a fixed schedule, validate ranges, generate features using only available data, and write predictions with a model version and timestamp. Store the input snapshot alongside each forecast so an incorrect alert can be reconstructed later.
Set monitoring thresholds for missing data, sensor drift, changing API schemas, and performance deterioration. Retrain by season or when drift is demonstrated—not simply because a calendar reminder fires. For high-impact decisions, use a human review process and retain IMD warnings as an authoritative safety input. The model should recommend actions such as “inspect covers” or “review heat protocol,” not independently override venue safety procedures.
This operational mindset also applies to building predictive maintenance systems with AI, where alerts must be traceable, prioritised, and connected to a clear response.
Common mistakes to avoid
- Randomly splitting time-series rows and reporting inflated test performance.
- Using forecast revisions or future observations that were unavailable at prediction time.
- Treating a distant weather station as ground truth without measuring spatial error.
- Optimising accuracy while ignoring rare but operationally important heavy-rain events.
- Claiming precise forecasts from a small or poorly calibrated dataset.
- Deploying without fallbacks when the API, sensor, or model is unavailable.
Final checklist
Before using the model for Arun Jaitley Stadium operations, confirm that you have a defined target, documented data rights, IST-aligned timestamps, leakage-safe features, walk-forward validation, seasonal error analysis, calibrated probabilities, and an escalation procedure. Start with a baseline and one forecast horizon, then expand only when the evidence supports it.
For related city-level modelling approaches, compare this workflow with Bhubaneswar weather prediction using Hugging Face models and Guwahati weather prediction using Hugging Face models. If you are building a research or applied-AI project in India, explore AI Grants India for relevant funding opportunities.