India’s agriculture data is broad, uneven, and highly valuable for machine learning. It spans crop production, mandi arrivals, commodity prices, rainfall, soil health, irrigation, satellite imagery, pest and disease images, livestock, and farmer-facing advisories. The challenge is not simply finding a CSV. It is identifying data that is relevant to an Indian operating context, legally reusable, geographically consistent, and documented well enough to support a trustworthy model.
This guide explains where to look for open source Indian agriculture datasets for machine learning, how to evaluate them, and how to turn them into useful prototypes without overstating what the data can prove.
Start with the problem, not the dataset
Define the decision your model should support before downloading data. A yield-prediction project, for example, needs a target variable, a forecast horizon, and a geographic unit. A disease classifier needs labelled images and a clear definition of each disease class. A price-forecasting system must specify the market, commodity, unit, and prediction window.
Common project directions include:
- Yield estimation: Predict production or yield by district, season, crop, and year.
- Price forecasting: Estimate wholesale or mandi prices using arrivals, historical prices, weather, and market attributes.
- Crop classification: Use satellite or field images to identify crop types and planting cycles.
- Disease detection: Classify visible symptoms in crop images, while accounting for lighting and field conditions.
- Irrigation planning: Combine rainfall, soil, evapotranspiration, and crop-stage information.
- Resource mapping: Detect land use, water stress, or crop damage from geospatial imagery.
For a portfolio project, combine a well-scoped dataset with reproducible documentation. The best machine learning projects for beginners in India offer useful patterns for choosing a manageable target, baseline, and evaluation plan.
Reliable sources of Indian agriculture data
data.gov.in and government departments
The Open Government Data platform is one of the most useful starting points for official Indian statistics. Search by crop, district, state, season, production, area, yield, rainfall, soil, livestock, or agricultural marketing. Datasets may be available as CSV, APIs, catalogues, or downloadable reports.
Government records are valuable because they often cover multiple years and administrative units. They also require careful interpretation: definitions can change, missing values may be coded inconsistently, and figures reported at state, district, and market levels are not automatically comparable.
Agricultural marketing and price data
For price and market models, examine datasets associated with the Agmarknet system and the National Agriculture Market (eNAM). Useful fields may include commodity, variety, market, arrival date, arrivals, minimum price, maximum price, and modal price. Before modelling, standardise commodity names, units, market identifiers, and dates.
Do not treat modal price as a guaranteed farmer realisation. It is a market-level statistic, and the actual price depends on quality, timing, location, transaction costs, and buyer relationships.
ICAR, agricultural universities, and research repositories
The Indian Council of Agricultural Research (ICAR) institutes, state agricultural universities, and international research organisations publish crop trials, soil studies, pest observations, climate analyses, and agronomic datasets. ICRISAT is particularly relevant for semi-arid agriculture and smallholder farming research.
Research datasets can be more technically detailed than administrative data, but their coverage is often narrower. Read the accompanying paper and codebook to understand sampling, treatment conditions, measurement methods, and limitations before combining them with national statistics.
Satellite and geospatial data
Satellite imagery is useful for crop classification, vegetation monitoring, drought assessment, and land-use analysis. Platforms such as ISRO’s Bhuvan, Sentinel-2, Landsat, and Google Earth Engine provide access to imagery or derived geospatial layers, subject to their respective terms.
For an India-focused model, align imagery with crop calendars, cloud cover, resolution, and the exact geography in your labels. A district-level label paired with a small field image can create severe noise. Cloud masking, compositing, reprojection, and temporal sampling are often more important than choosing a complex neural network.
Kaggle, GitHub, and academic collections
Kaggle and GitHub are convenient for discovering crop-image datasets, rice and wheat disease collections, and cleaned versions of public records. They are useful for experimentation, but provenance varies. Check the original publisher, licence, collection date, duplicate images, class balance, and whether train and test images come from the same farms or photographers.
Open-source projects can help you build ingestion and evaluation pipelines; the Indian open-source AI developer projects guide is a useful companion for finding implementation ideas and contribution patterns.
Dataset categories and what they support
Tabular production and yield data
These datasets usually contain crop, season, area, production, yield, district, and state. They support forecasting, anomaly detection, and comparative analysis. Watch for data leakage: using final production figures or post-harvest information to predict a result that should be available before harvest will inflate performance.
Weather, soil, and water data
Rainfall, temperature, humidity, soil nutrients, groundwater, and irrigation records can improve agronomic models. Join them using consistent coordinates and time windows. A simple district average may hide major differences between rainfed and irrigated farms.
Image datasets
Crop and disease images support classification and detection models. Inspect the labels manually, separate near-duplicates, and split by farm, location, or collection session rather than randomly when possible. Random image splits can place almost identical images in both training and test sets.
Remote-sensing and time-series data
Vegetation indices such as NDVI can support crop monitoring and stress detection. Time-series models need regular timestamps, clear handling of missing observations, and a baseline such as seasonal averages or gradient-boosted trees. Deep learning is not automatically better when the labelled sample is small.
How to evaluate a dataset before using it
Use this checklist:
- Provenance: Who collected or published the data?
- Coverage: Which states, districts, crops, seasons, and years are represented?
- Granularity: Are observations field-, village-, district-, market-, or state-level?
- Schema: Are units, codes, categories, and missing-value conventions documented?
- Licence: Can you modify, redistribute, and use the data commercially?
- Freshness: When was it last updated, and is the collection process stable?
- Bias: Are smallholders, rainfed farms, marginal regions, or specific crops underrepresented?
- Reproducibility: Can another developer download the same version and recreate your features?
A dataset being downloadable does not mean it is open source. Treat the licence and terms of use as part of the technical specification.
A practical modelling workflow
1. Create a data card. Record source, licence, dates, geography, units, known gaps, and intended use.
2. Build a validation table. Check duplicates, impossible values, missingness, outliers, and inconsistent names.
3. Establish a baseline. Use a seasonal mean, last-value forecast, linear model, or decision tree before testing advanced methods.
4. Split by time or geography. For future prediction, use a temporal holdout. For generalisation, hold out districts, farms, or markets.
5. Measure the right outcome. Use MAE or RMSE for continuous targets, macro-F1 for imbalanced classes, and calibration when predictions drive decisions.
6. Test robustness. Compare performance across states, crops, seasons, farm types, and data availability levels.
7. Document uncertainty. A prediction should include its expected error and the conditions under which it may fail.
If you are still building core Python, data handling, and model evaluation skills, machine learning portfolio projects for beginners in India can help you structure the work without skipping fundamentals.
Responsible use in Indian agricultural settings
Agricultural data can expose sensitive information about land, income, production, or farmer behaviour. Avoid publishing personally identifiable information, precise farm locations, or inferred financial details without a legitimate basis and appropriate safeguards. Aggregate results when individual-level disclosure is unnecessary.
Models should support—not replace—agronomists, extension workers, traders, and farmers. Validate recommendations locally, communicate uncertainty in plain language, and test whether the system works for marginal farmers and low-connectivity settings. If the interface serves farmers in Indian languages, review the low-resource Indic natural language processing guide before deploying a text or voice layer.
A strong 2026 project blueprint
A credible starter project could combine district-level crop production from government data, rainfall from an official or reputable climate source, and a transparent model that forecasts yield one season ahead. Publish the data dictionary, cleaning script, licence notes, baseline, time-based evaluation, error analysis, and limitations. That package is more valuable than a headline accuracy score.
Open-source Indian agriculture datasets can accelerate useful research and products, but only when treated as evidence with context—not as ready-made ground truth. Start with a decision, verify the data-generating process, and build a model that remains understandable to the people expected to use it.
FAQ
Are all datasets on Kaggle open source?
No. Check the dataset page, original source, and licence. A public download may still restrict redistribution or commercial use.
Which dataset is best for a beginner?
A clean, documented tabular dataset with crop, year, geography, and production fields is usually easier than satellite imagery or disease images. Begin with a baseline and a time-based split.
Can I combine data from different government portals?
Yes, but first reconcile geographic codes, units, crop names, calendars, and reporting periods. Record every transformation in a reproducible pipeline.
What should I do if a dataset has missing values?
Measure missingness by feature, location, and time. Investigate why values are absent before imputing; missingness may itself reflect reporting or infrastructure differences.
Where can agriculture AI builders find support?
Researchers and founders can explore relevant public programmes, incubators, and funding opportunities through AI Grants India.