Data preprocessing is not a one-off cleaning exercise. It is the boundary between unreliable inputs and a model that can be tested, monitored, and trusted. For Indian teams, the challenge is amplified by mixed scripts, inconsistent address formats, rupee values, regional languages, duplicated records, changing vendor schemas, and data collected through low-bandwidth or offline workflows.
A good set of Python scripts for automating data preprocessing should do more than fill nulls. It should validate incoming data, preserve reproducibility, prevent training-serving skew, and make failures visible before they reach production. The examples below use Pandas and scikit-learn, but the design principles also apply to PySpark, Dask, and feature-store workflows.
Start with a data contract
Before writing transformations, define what the pipeline expects. A data contract should specify column names, data types, allowed ranges, units, missing-value rules, and whether a field is required. For example, amount_inr should be numeric and non-negative, while customer_id should be treated as an identifier rather than a model feature.
Use a schema library such as Pandera or Pydantic to reject malformed batches early. This is especially useful when data comes from multiple Indian states, vendors, languages, or legacy systems. Teams working on high-stakes applications should also review principles from data veracity infrastructure for high-stakes AI, including lineage, provenance, and human review.
import pandas as pd
REQUIRED_COLUMNS = {"customer_id", "age", "amount_inr", "state"}
def validate_input(df: pd.DataFrame) -> None:
missing = REQUIRED_COLUMNS - set(df.columns)
if missing:
raise ValueError(f"Missing required columns: {sorted(missing)}")
if (df["amount_inr"].dropna() < 0).any():
raise ValueError("amount_inr cannot be negative")Log the row count, schema version, source, and validation errors. Avoid logging personally identifiable information in application logs.
Handle missing values without leakage
Imputation statistics must be learned from the training split only. Fitting an imputer on the full dataset allows information from validation or test rows to influence training, producing over-optimistic results. Prefer median imputation for skewed numeric fields and an explicit category such as __missing__ for categorical fields.
from sklearn.impute import SimpleImputer
numeric_imputer = SimpleImputer(strategy="median")
categorical_imputer = SimpleImputer(
strategy="constant", fill_value="__missing__"
)Do not automatically treat every missing value as an error or replace it with zero. A missing salary, a zero salary, and a salary that was not collected represent different realities. Add indicators when absence itself may carry signal:
df["income_missing"] = df["income"].isna().astype("int8")Build one leakage-safe preprocessing pipeline
The most reusable approach is a ColumnTransformer inside a scikit-learn Pipeline. It keeps transformations attached to the estimator and ensures that training, validation, batch scoring, and API inference use the same fitted objects.
from sklearn.compose import ColumnTransformer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler
numeric_features = ["age", "amount_inr"]
categorical_features = ["state", "product_category"]
numeric_pipeline = Pipeline([
("imputer", SimpleImputer(strategy="median")),
("scaler", StandardScaler()),
])
categorical_pipeline = Pipeline([
("imputer", SimpleImputer(strategy="constant", fill_value="__missing__")),
("encoder", OneHotEncoder(handle_unknown="ignore")),
])
preprocessor = ColumnTransformer([
("numeric", numeric_pipeline, numeric_features),
("categorical", categorical_pipeline, categorical_features),
])handle_unknown="ignore" matters in production: a new district, product, or language code should not crash inference. For very high-cardinality fields, do not use arbitrary integer labels as though categories were ordered. Consider frequency encoding, hashing, or a model designed for categorical features, and measure the effect on performance and fairness.
Treat outliers as decisions, not automatic errors
An outlier may be a sensor fault, a duplicate transaction, a genuine high-value customer, or a fraud event—the exact observation a model needs. Blindly removing rows with the IQR rule can erase important cases and can disproportionately affect smaller regions or customer groups.
Start by flagging rather than deleting:
def add_iqr_flag(df, column):
q1 = df[column].quantile(0.25)
q3 = df[column].quantile(0.75)
iqr = q3 - q1
low, high = q1 - 1.5 * iqr, q3 + 1.5 * iqr
df[f"{column}_outlier"] = ((df[column] < low) | (df[column] > high)).astype("int8")
return dfFor monetary features, log transforms or domain-based caps may be more defensible than row deletion. Record every rule and compare model metrics before and after the change.
Normalise dates, categories, and Indian text data
Dates should be parsed with an explicit timezone and converted into features appropriate to the use case. For operational data, retain the original timestamp and derive hour, weekday, month, and elapsed-time features. Do not derive future information when predicting a past event.
Categorical values need canonicalisation: trim whitespace, normalise case where appropriate, and map known variants such as abbreviations. Preserve the raw value in a separate column when auditability matters.
Text requires more care than lowercasing and deleting punctuation. Unicode normalisation can help reconcile visually equivalent forms, but aggressive cleaning may damage Devanagari, Tamil, Bengali, or mixed-language text. Keep meaningful digits, punctuation, and code-mixed tokens when they carry intent. Teams preparing Indic-language corpora should pair preprocessing with low-resource language datasets for AI training in India and document script-specific choices.
import re
import unicodedata
def clean_text(value: str) -> str:
value = unicodedata.normalize("NFKC", str(value))
value = re.sub(r"\s+", " ", value).strip()
return valueAdd quality checks and reproducibility
A production script should fail loudly when assumptions break. Test null rates, duplicate identifiers, category growth, numeric ranges, and train-serving feature parity. Store the code version, dependency lockfile, schema version, input snapshot or hash, and configuration used for each run.
Useful checks include:
- Compare current distributions with the training baseline.
- Alert when a column's missingness changes materially.
- Check that labels are unavailable at inference time.
- Measure duplicate and near-duplicate records.
- Sample transformed rows for human review.
- Keep transformations deterministic with fixed random seeds where applicable.
For larger datasets, process CSV files in chunks or move to PySpark or Dask. Chunking reduces memory pressure but requires care: global statistics such as medians and category counts cannot always be computed correctly from one chunk. Use scalable aggregation methods or a distributed pipeline instead. Teams looking to automate the surrounding workflow can also learn from Python data science automation for Indian startups.
A practical project structure
Separate reusable transformations from orchestration and configuration:
preprocessing/
schema.py
transforms.py
pipeline.py
checks.py
config.yaml
tests/
test_transforms.py
test_schema.pyWrite unit tests for edge cases: empty strings, unknown categories, null-only columns, extreme values, Unicode text, duplicate IDs, and timezone boundaries. Run them in CI before a model or data pipeline is deployed. For a lightweight analytics team, a scheduled Python job may be enough; for a larger system, package the pipeline as a versioned container and expose the same transformation artifact to batch and online inference.
A production checklist
Before shipping, confirm that you can answer these questions:
- What schema and data version produced this training set?
- Which statistics were fitted, and on which split?
- What happens when a new category or language appears?
- Can an operator reproduce a failed batch?
- Are sensitive fields minimised, masked, and access-controlled?
- Are quality metrics monitored after deployment?
Reliable preprocessing is a product capability, not merely notebook code. Start with explicit contracts, use leakage-safe pipelines, preserve provenance, and make every transformation testable. For teams building open tooling around these practices, building open-source AI tools for Indian developers offers a useful path for sharing components while keeping data and credentials private.