0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build your first machine learning model from scratch

How to Build Your First Machine Learning Model from Scratch

  1. aigi

    Building your first machine learning model is less about choosing a sophisticated algorithm and more about creating a reliable workflow. You need a clearly defined prediction problem, representative data, a reproducible training process, and an evaluation method that reflects real-world costs. This approach works whether you are learning on a laptop, preparing a portfolio project, or prototyping an India-focused product.

    “From scratch” should mean understanding each stage rather than avoiding useful libraries. You will still use Python, pandas, NumPy, and scikit-learn; the goal is to know what those tools are doing, what can go wrong, and how to improve the result responsibly.

    1. Start with a measurable problem

    Choose one outcome and define it precisely before collecting data. A vague goal such as “predict customer behaviour” is difficult to train and evaluate. A stronger version might be: “Predict whether a customer will miss the next repayment within 30 days.”

    Most beginner projects use supervised learning:

    • Classification predicts a category, such as fraud/not fraud or churn/no churn.
    • Regression predicts a number, such as delivery time, electricity demand, or property price.

    Write down the target column, the prediction time, the available inputs, and the action that follows the prediction. For an Indian product, also ask whether the data covers the language, geography, device types, income ranges, and usage patterns you expect in production.

    For practice, select a dataset with a clear target and manageable size. If you want a project that is useful for a portfolio, compare your idea with these machine learning portfolio projects for beginners in India and choose one where you can explain the business value, not just the accuracy score.

    2. Set up a reproducible Python project

    A notebook is useful for exploration, but keep the final workflow reproducible. Use Python 3, create a virtual environment, and record dependencies in requirements.txt or pyproject.toml.

    A practical starter stack includes:

    • pandas for tabular data and cleaning
    • NumPy for numerical operations
    • matplotlib or seaborn for visual checks
    • scikit-learn for preprocessing, models, and metrics
    • Jupyter for exploration, followed by a Python script or pipeline for repeatable training

    Store the raw data separately from transformed data, keep a README, and record the dataset source, licence, target definition, and known limitations. Never commit API keys, personal data, or unredacted customer records to a public repository.

    3. Inspect the data before modelling

    Load the dataset and answer basic questions before selecting an algorithm:

    • How many rows and columns are present?
    • Which column is the target?
    • What are the data types?
    • Which values are missing or duplicated?
    • Are there impossible values, such as negative ages or dates in the future?
    • Is the target heavily imbalanced?

    Use df.info(), df.describe(include='all'), df.isna().sum(), and value counts as a starting point. Plot distributions and examine relationships between important features and the target. A model trained on a flawed label or a leaked feature can appear successful while failing in practice.

    Treat privacy as part of data preparation. For datasets involving Aadhaar-linked services, health records, financial information, student data, or call recordings, minimise collection, remove direct identifiers, restrict access, and document consent and retention requirements. A technically strong model is not ready for deployment if its data practices are unsafe.

    4. Split data without creating leakage

    Separate the data into training and test sets before fitting transformations or selecting features. A common first split is 80% for training and 20% for testing, with a fixed random_state so results can be reproduced.

    Use a validation set or cross-validation on the training data for model selection. Keep the test set untouched until the end. If the target classes are uneven, use a stratified split so both sets have a similar class distribution.

    Random splitting is not always correct. For time-based predictions, train on earlier records and test on later records. For users, patients, merchants, or devices with multiple rows, split by entity so information from the same entity does not appear in both training and testing. These choices often matter more than changing the algorithm.

    5. Build a preprocessing pipeline

    Machine learning models require numerical, consistently formatted inputs. Typical steps include:

    • Impute missing numerical values with a documented rule, such as the training-set median.
    • Encode categorical values with one-hot encoding.
    • Scale numerical features for algorithms such as logistic regression, K-nearest neighbours, and support vector machines.
    • Drop identifiers that have no predictive meaning or create leakage.

    Fit preprocessing only on training data. Scikit-learn’s Pipeline and ColumnTransformer make this safer by applying the same transformations during validation, testing, and inference. A pipeline also makes deployment easier because the model receives the same feature treatment it saw during training.

    6. Train a baseline before trying complex models

    For binary classification, start with a majority-class baseline and then try logistic regression or a decision tree. For regression, compare against predicting the training-set median or mean before using linear regression or a tree-based model.

    A simple classification workflow looks like this:

    from sklearn.compose import ColumnTransformer
    from sklearn.impute import SimpleImputer
    from sklearn.linear_model import LogisticRegression
    from sklearn.metrics import classification_report, confusion_matrix
    from sklearn.pipeline import Pipeline
    from sklearn.preprocessing import OneHotEncoder, StandardScaler
    from sklearn.model_selection import train_test_split
    
    X = df.drop(columns="target")
    y = df["target"]
    
    X_train, X_test, y_train, y_test = train_test_split(
        X, y, test_size=0.2, stratify=y, random_state=42
    )
    
    numeric_features = X.select_dtypes(include="number").columns
    categorical_features = X.select_dtypes(exclude="number").columns
    
    preprocess = ColumnTransformer([
        ("num", Pipeline([
            ("imputer", SimpleImputer(strategy="median")),
            ("scaler", StandardScaler())
        ]), numeric_features),
        ("cat", Pipeline([
            ("imputer", SimpleImputer(strategy="most_frequent")),
            ("onehot", OneHotEncoder(handle_unknown="ignore"))
        ]), categorical_features)
    ])
    
    model = Pipeline([
        ("preprocess", preprocess),
        ("classifier", LogisticRegression(max_iter=1000))
    ])
    
    model.fit(X_train, y_train)
    predictions = model.predict(X_test)
    print(classification_report(y_test, predictions))

    The code is only the beginning. Inspect the features, understand the error cases, and confirm that the target definition matches the product decision.

    7. Evaluate with the right metric

    Accuracy is useful only when mistakes have similar consequences and classes are reasonably balanced. Use metrics that match the decision:

    • Precision: Of the positive predictions, how many were correct?
    • Recall: Of the actual positives, how many did the model find?
    • F1 score: A combined measure of precision and recall.
    • ROC-AUC or PR-AUC: Useful for comparing ranking performance, with PR-AUC often more informative for rare positives.
    • MAE and RMSE: Common regression metrics; MAE is easier to interpret, while RMSE penalises large errors more heavily.

    Review the confusion matrix by segment: city, language, device, age group, or customer type. A strong overall score can hide poor performance for users in smaller Indian regions or for lower-resource languages. Also check calibration if the model outputs probabilities that will influence approvals, triage, or human review.

    8. Improve systematically

    Change one thing at a time and log the result. Useful improvements include better labels, additional representative data, feature cleanup, class weights, threshold tuning, and cross-validation. Compare every change with the baseline.

    Watch for overfitting: a large gap between training and validation performance usually means the model is memorising patterns that do not generalise. Regularisation, simpler features, more data, or a less flexible model can help. Hyperparameter search is valuable only after the data split and evaluation design are sound.

    When your project needs to move from a notebook to a service, learn the fundamentals of system design before adding infrastructure. Your model will need input validation, versioned artefacts, logging, monitoring, rollback, and a plan for retraining. For more advanced products involving autonomous workflows, distinguish ordinary prediction from distributed systems with AI agents; they have different reliability and evaluation requirements.

    9. Deploy a small, testable version

    Package the complete preprocessing-and-model pipeline, expose a narrow API, and validate inputs at the boundary. FastAPI is a practical option for a Python service, while batch predictions may be simpler and cheaper for many business workflows. Containerise the service only when it improves repeatability or deployment.

    Before launch, define latency, cost, uptime, and accuracy targets. Monitor input drift, missing fields, prediction distributions, feedback quality, and performance by important segments. Do not silently retrain on fresh production data; review labels, permissions, and data quality first.

    Common mistakes to avoid

    • Using the test set repeatedly: This turns the test set into another training signal.
    • Leaking future information: Features must be available at prediction time.
    • Optimising accuracy alone: The wrong metric can produce harmful decisions.
    • Ignoring a baseline: A complex model must beat a simple alternative.
    • Skipping documentation: Record data sources, assumptions, versions, and limitations.
    • Jumping to deep learning: Traditional models are faster to debug and often sufficient for tabular data.

    Once you have a dependable model, you can explore specialised directions such as building computer vision models on GitHub or compare your work against machine learning projects for computer science students. The next skill is not collecting more algorithms; it is learning to turn experiments into reliable software.

    FAQ

    Do I need advanced mathematics?
    No. Basic probability, statistics, vectors, and optimisation will help you reason about results, but you can begin with a clear workflow and deepen the theory as questions arise.

    Can I build a model on a laptop?
    Yes. Most first projects use datasets that run comfortably on a laptop. Cloud GPUs are unnecessary for many tabular problems.

    How much data do I need?
    There is no universal threshold. Data quality, label consistency, feature usefulness, and class balance matter more than a fixed row count. Start with a small, representative sample and measure learning curves.

    When should I use deep learning?
    Consider it when you have enough labelled data, a problem suited to images, audio, language, or complex sequences, and a clear reason simpler models are insufficient.

    Build beyond the tutorial

    A first model becomes valuable when it is reproducible, honestly evaluated, and connected to a real decision. If you are turning an ML prototype into an India-focused product, explore AI Grants India for potential funding and support for early-stage AI builders.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.