0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · beginner guide to building machine learning models in python

Beginner Guide to Building Machine Learning Models in Python

  1. aigi

    Python makes machine learning accessible, but a working model requires more than importing a library and calling .fit(). You need a repeatable process for defining the problem, preparing trustworthy data, selecting a suitable baseline, measuring performance, and deciding whether the result is useful in practice.

    This beginner guide to building machine learning models in Python uses scikit-learn and a small tabular dataset as its foundation. The same workflow applies to projects such as predicting demand for an Indian retail business, classifying support tickets, estimating delivery times, or identifying unusual transactions. Start with classical machine learning before moving to deep learning: it is faster to iterate, easier to debug, and usually sufficient for structured data.

    What you need before writing a model

    You do not need an expensive GPU for the first stage. A laptop, Python 3.10 or newer, and a virtual environment are enough. Use Jupyter Notebook or Google Colab for exploration, then move stable code into scripts as the project grows.

    Create an isolated environment and install the core stack:

    python -m venv .venv
    source .venv/bin/activate        # macOS/Linux
    # .venv\\Scripts\\activate       # Windows
    pip install pandas numpy scikit-learn matplotlib seaborn joblib

    The main tools have distinct roles:

    • Pandas reads and transforms tabular data.
    • NumPy supports numerical operations.
    • Matplotlib and Seaborn help inspect distributions and relationships.
    • Scikit-learn provides preprocessing, models, validation, and metrics.
    • Joblib can save a trained pipeline for later inference.

    If you need project ideas rather than another toy example, compare these machine learning portfolio projects for beginners in India and choose one with a clear user, prediction deadline, and measurable outcome.

    Step 1: Define the prediction problem

    Write down four things before opening the dataset:

    1. Input: What information is available when a prediction is made?
    2. Target: What exactly should the model predict?
    3. Unit of prediction: Is one row a customer, order, image, or day?
    4. Success metric: What error is acceptable for the user or business?

    This prevents a common beginner mistake: optimising an impressive metric for a problem nobody can act on. For example, a delivery-delay model may need high recall for late orders, while a house-price estimator may be judged by mean absolute error in rupees.

    Also decide whether the task is supervised or unsupervised. Regression predicts a number, classification predicts a category, and clustering groups similar records without labelled outcomes.

    Step 2: Inspect and split the data correctly

    Load the data, inspect its shape and types, and check the target before performing transformations.

    import pandas as pd
    
    raw = pd.read_csv("orders.csv")
    print(raw.shape)
    print(raw.head())
    print(raw.dtypes)
    print(raw.isna().sum().sort_values(ascending=False).head())
    print(raw["late_delivery"].value_counts(normalize=True))

    Look for duplicate rows, impossible values, inconsistent categories, and target imbalance. Do not automatically delete unusual observations: an extreme value may be a genuine event or a data-entry error, and the correct treatment depends on the domain.

    Split the data before learning imputation values, encodings, or scaling parameters. For classification, use a stratified split so the class proportions remain similar:

    from sklearn.model_selection import train_test_split
    
    X = raw.drop(columns="late_delivery")
    y = raw["late_delivery"]
    
    X_train, X_test, y_train, y_test = train_test_split(
        X, y, test_size=0.2, random_state=42, stratify=y
    )

    Keep the test set untouched until the end. If your data is time-ordered, use a chronological split instead of randomly mixing past and future records.

    Step 3: Build preprocessing into a pipeline

    A pipeline keeps transformations consistent during training, cross-validation, and production inference. It also reduces data leakage, where information from the validation or test set accidentally influences training.

    Separate numeric and categorical columns, then combine their transformations with a ColumnTransformer:

    from sklearn.compose import ColumnTransformer
    from sklearn.impute import SimpleImputer
    from sklearn.pipeline import Pipeline
    from sklearn.preprocessing import OneHotEncoder, StandardScaler
    
    numeric_features = X_train.select_dtypes(include="number").columns
    categorical_features = X_train.select_dtypes(exclude="number").columns
    
    numeric_pipe = Pipeline([
        ("imputer", SimpleImputer(strategy="median")),
        ("scaler", StandardScaler()),
    ])
    
    categorical_pipe = Pipeline([
        ("imputer", SimpleImputer(strategy="most_frequent")),
        ("encoder", OneHotEncoder(handle_unknown="ignore")),
    ])
    
    preprocess = ColumnTransformer([
        ("numeric", numeric_pipe, numeric_features),
        ("categorical", categorical_pipe, categorical_features),
    ])

    Scaling matters for distance- and margin-based models such as K-nearest neighbours, logistic regression, and support vector machines. Tree-based models are generally less sensitive to feature scale, but a pipeline is still valuable for consistency.

    Step 4: Start with a baseline model

    A baseline gives you a reference point. For binary classification, begin with logistic regression or a small decision tree before trying complex ensembles.

    from sklearn.linear_model import LogisticRegression
    
    model = Pipeline([
        ("preprocess", preprocess),
        ("classifier", LogisticRegression(max_iter=1000, class_weight="balanced")),
    ])
    
    model.fit(X_train, y_train)

    class_weight="balanced" can help when late deliveries are uncommon, but it is not a substitute for checking the business cost of false positives and false negatives. Compare the model against a simple majority-class baseline. If it cannot beat that baseline, investigate the data and features before tuning parameters.

    For regression, replace the classifier with Ridge, RandomForestRegressor, or HistGradientBoostingRegressor. For clustering, remember that there is no target column; you must define how cluster usefulness will be assessed.

    Step 5: Evaluate with the right metrics

    Generate predictions only after fitting the pipeline:

    from sklearn.metrics import classification_report, confusion_matrix, roc_auc_score
    
    predictions = model.predict(X_test)
    probabilities = model.predict_proba(X_test)[:, 1]
    
    print(confusion_matrix(y_test, predictions))
    print(classification_report(y_test, predictions))
    print("ROC-AUC:", roc_auc_score(y_test, probabilities))

    Accuracy can mislead on imbalanced data. Precision measures how many predicted positives are correct; recall measures how many actual positives were found; F1-score balances both. Examine the confusion matrix and select a decision threshold based on real consequences, not a default value of 0.5.

    For regression, use mean absolute error, root mean squared error, and, where appropriate, R². Report errors in understandable units. An MAE of ₹800 is more useful to a product team than an unexplained decimal score.

    Use cross-validation on the training set for a more reliable estimate:

    from sklearn.model_selection import cross_validate
    
    scores = cross_validate(model, X_train, y_train, cv=5, scoring=["f1", "roc_auc"])
    print(scores["test_f1"].mean())

    Step 6: Tune only after the workflow is sound

    Once the baseline is stable, search a small, reasoned parameter space with RandomizedSearchCV or GridSearchCV. Tune model settings such as regularisation strength, tree depth, or number of estimators—not arbitrary combinations of every available option.

    Keep the final test set for one honest evaluation. Save the complete fitted pipeline, not just the estimator:

    import joblib
    joblib.dump(model, "delivery_model.joblib")

    Record the Python version, package versions, dataset snapshot, feature definitions, metric results, and known limitations. This is essential when a student project becomes a service used by an Indian business or public programme.

    Common mistakes and responsible practice

    • Leakage: Remove features that would only become available after the prediction event.
    • Overfitting: Prefer cross-validation and simpler models over chasing training accuracy.
    • Unstable categories: Use handle_unknown="ignore" and monitor new values.
    • Unrepresentative data: Check whether language, geography, device type, gender, income, or connectivity gaps affect performance.
    • Unclear consent: Avoid collecting personal data that the project does not need.
    • No monitoring: Track input drift, prediction rates, latency, and real-world outcomes after deployment.

    For computer vision, tabular pipelines are not enough; start with this guide to build computer vision models on GitHub. If you are ready to explore neural networks, first understand the trade-offs in customizable neural network architectures for beginners.

    A practical learning path

    Build three progressively harder projects: a regression model, a classification model with imbalanced data, and a time-aware forecasting project. Publish the problem statement, data card, baseline, evaluation method, error analysis, and limitations—not just a notebook with a score. These best machine learning projects for computer science students can help you choose a portfolio direction.

    The goal is not to memorise every algorithm. It is to develop the habit of asking whether the data reflects the intended users, whether the metric matches the decision, and whether the model remains reliable after it leaves the notebook.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.