Why feature selection matters in Eastern UP
Weather models for Eastern Uttar Pradesh must handle monsoon rainfall, heat, fog, cold waves, local convection, and strong variation between districts and observation sites. A dataset may contain dozens or hundreds of candidate predictors: lagged temperature, humidity, pressure, wind, rainfall, satellite indices, soil moisture, reanalysis variables, and calendar signals. More inputs do not automatically produce a better forecast. Redundant, noisy, or poorly timed variables can increase training cost and cause a model to learn patterns that fail outside the training period.
A genetic algorithm (GA) treats feature selection as a search problem. Each candidate solution represents a subset of variables, and the algorithm evolves subsets according to forecast performance and complexity. The method is especially useful when interactions matter—for example, when humidity and temperature jointly signal fog, or when rainfall depends on several lagged atmospheric conditions.
For a broader view of model families and deployment choices, see this guide to AI models, selection, applications, and deployment. Genetic algorithms are not a forecasting model by themselves; they are an optimisation layer around a forecasting pipeline.
Define the forecasting task first
Before writing the GA, specify exactly what the model must predict. A feature subset that works for next-day maximum temperature may be unsuitable for three-hour rainfall nowcasting.
Record these decisions:
- Target: rainfall amount, rain/no-rain, temperature, fog probability, heat index, or another variable.
- Forecast horizon: for example, 6 hours, 24 hours, or 3 days ahead.
- Geography: one station, a district, a grid cell, or all of Eastern UP.
- Update frequency: daily, hourly, or event-triggered.
- Operational constraint: latency, memory, missing-data tolerance, and acceptable false alarms.
If the target is binary, use metrics such as F1, precision-recall AUC, and the Brier score. For continuous rainfall or temperature, use MAE, RMSE, and calibration of prediction intervals. Do not use accuracy alone for rare events such as heavy rainfall.
Build a defensible Eastern UP dataset
Combine sources only after aligning timestamps, units, spatial references, and measurement quality. Potential inputs include IMD observations where access and licensing permit, automatic weather stations, satellite products, and gridded reanalysis. District-level modelling should account for station density and elevation rather than silently treating every location as equally observed.
Useful feature groups include:
- Current and lagged temperature, dew point, relative humidity, pressure, wind speed, and direction.
- Rainfall totals over rolling windows such as 3, 6, 12, 24, and 72 hours.
- Calendar variables, monsoon phase, day length, and hour-of-day encodings.
- Neighbouring-station or nearby-grid summaries.
- Satellite cloud, land-surface, vegetation, and soil-moisture indicators where temporally available.
- Reanalysis fields at multiple pressure levels, with careful attention to their forecast or analysis availability.
Create a data dictionary before modelling. It should state the source, unit, timestamp convention, spatial resolution, missing-value code, and the earliest time at which each feature becomes available. This prevents a common failure: using a final daily observation or a revised gridded product that would not have been available at prediction time.
For examples of location-specific weather modelling workflows, compare the approaches used for Bhubaneswar weather prediction with Hugging Face models and Guwahati weather prediction with Hugging Face models. The methods are transferable, but Eastern UP requires its own validation and climatology.
Prevent leakage before running the GA
Feature selection must happen inside the training process. If you select variables using the full dataset and then report test performance, information from the test period has influenced the result.
Use a chronological split, such as:
- Training period: earlier years and seasons.
- Validation period: a later block used to evolve and tune subsets.
- Final test period: the most recent unseen period, used once.
For stronger evidence, use rolling-origin validation: train on an initial period, validate on the next block, expand the training window, and repeat. Keep entire weather events together where practical. If neighbouring stations are highly correlated, use spatial holdouts as well as temporal holdouts to test generalisation across districts.
Imputation, scaling, feature engineering, and any dimensionality reduction must be fitted within each training fold. The GA should never see validation targets while creating a candidate subset.
Encode the genetic algorithm
Assume there are *p* candidate features. Represent each subset as a binary chromosome of length *p*: 1 means the feature is selected and 0 means it is excluded. A typical workflow is:
1. Generate an initial population of random or heuristic chromosomes.
2. Train the chosen forecasting model using each chromosome’s selected variables.
3. Score every chromosome through time-aware cross-validation.
4. Select stronger candidates, apply crossover, and mutate bits.
5. Repeat for a fixed number of generations or until improvement stalls.
6. Refit the selected subset on the complete training data and evaluate once on the test period.
A practical fitness function balances predictive skill and simplicity:
fitness = validation_error + λ × (selected_features / total_features) + penalties
Set λ according to operational priorities. Add penalties for excessive missingness, unstable performance across folds, or feature groups that violate availability rules. Use a minimum-feature constraint if the model must remain deployable on low-bandwidth or edge systems.
Start with a population of roughly 30–100 chromosomes and 20–60 generations, then tune these values experimentally. Keep elitism modest so the population does not converge too early. Mutation is essential when the search becomes repetitive; its rate can be reduced gradually, but should not be driven to zero.
Select and evaluate the forecasting model
The GA can wrap a random forest, gradient-boosted trees, linear model, support-vector regression, or neural network. Begin with a fast baseline such as logistic regression, linear regression, or a tree ensemble. A slow deep model inside every GA evaluation can make experimentation unnecessarily expensive.
Compare at least four systems:
- A climatology or persistence baseline.
- The full-feature model.
- The GA-selected model.
- A simple filter method, such as correlation screening or mutual information.
Report the mean and spread across temporal folds, not only the best score. Also measure inference time, number of selected variables, missing-data failures, and calibration. A subset that improves RMSE by a tiny amount but removes half the inputs may be valuable operationally; a subset that wins on one season and fails during the monsoon is not.
Use repeated GA runs with different random seeds. Track selection frequency: variables chosen in 80–100% of runs are more credible than variables appearing once in a lucky chromosome. SHAP values or permutation importance can explain the final model, but calculate them on held-out data and avoid presenting importance as causality.
A compact implementation pattern
Python libraries such as scikit-learn can provide the estimator and time-aware splits, while a GA library or a custom loop can manage chromosomes. Keep the fitness function deterministic where possible, cache evaluations for duplicate chromosomes, parallelise independent candidates, and log every configuration.
Store:
- Dataset version and feature dictionary.
- Random seeds and GA parameters.
- Fold boundaries and availability assumptions.
- Selected feature names for every run.
- Validation metrics, runtime, and missingness.
- The final model, preprocessing pipeline, and prediction timestamp.
For rapid experimentation, GenAI for feature prototyping can help generate pipeline scaffolding, but inspect every transformation manually. Generated code must not bypass time-aware splitting or introduce future variables.
Common mistakes and a deployment checklist
Avoid selecting features on the entire dataset, using random k-fold splits for time series, mixing station timestamps without timezone checks, and comparing models on different missing-value treatments. Do not assume that a larger chromosome or more generations guarantees a better forecast.
Before deployment, verify that:
- Every input is available at the stated forecast issue time.
- The pipeline handles missing stations and delayed feeds.
- Performance is reported separately for monsoon, winter fog, summer heat, and extreme events.
- Drift monitoring detects changes in station coverage and climate behaviour.
- Retraining and rollback procedures are documented.
- The final test set remains untouched until model selection is complete.
FAQ
Are genetic algorithms always better than simpler selection methods? No. They are useful for nonlinear, interacting features, but mutual information, regularisation, or tree-based importance may be faster and equally effective. Compare methods against the same time-aware baseline.
Should scaling be used? Scale inputs for linear models, SVMs, and neural networks. Tree models generally do not require scaling, but consistent preprocessing is still important when pipelines are compared.
How often should the feature subset be rerun? Reassess after major changes in station coverage, sensors, data products, or forecast horizon. Monitor drift rather than changing the subset after every small metric fluctuation.
Can this approach support edge deployment? Yes. Penalise feature count and latency in the fitness function, then export a fixed preprocessing-and-model pipeline. For resource-constrained systems, also review techniques for efficient image classification on edge devices, particularly its principles for measuring memory and inference cost.