0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · integrating machine learning into web applications tutorial

Integrating Machine Learning Into Web Applications: A Practical Tutorial

  1. aigi

    A notebook prediction is not yet a product. To make machine learning useful inside a web application, you need a reliable path from user input to validated features, model inference, a clear response, and ongoing monitoring. This tutorial presents that path with Python, FastAPI, and a browser frontend, while also covering the decisions that matter when an Indian product moves from prototype to production.

    The examples suit applications such as sentiment analysis, document classification, fraud scoring, recommendations, and image recognition. If you are still building fundamentals, start with a small end-to-end project from this guide to machine learning portfolio projects for beginners in India, then apply the production patterns below.

    Choose the right inference architecture

    Use the simplest architecture that meets latency, privacy, and scaling requirements:

    • Synchronous server inference: The browser sends an HTTP request and receives a prediction immediately. This is appropriate for classification, tabular scoring, and short text analysis.
    • Asynchronous job inference: The API accepts a task and returns a job ID. A worker processes the request, while the frontend polls for status or receives a webhook. Use this for document extraction, video processing, image generation, and GPU-heavy workloads.
    • Client-side inference: Export the model to ONNX or TensorFlow.js and run it in the browser. This can improve privacy and reduce server cost, but increases bundle size and exposes more of the model.
    • Dedicated model serving: Place inference behind a specialised server when you need batching, GPU scheduling, multiple model versions, or independent autoscaling.

    For most first releases, keep the product API and inference service separate in code, even if they run in one container. This makes it easier to scale the model independently later. Teams planning larger systems should review this guide to scaling backend infrastructure for AI applications.

    1. Package training and preprocessing together

    The most damaging integration bug is training-serving skew: the production request is transformed differently from the data used during training. Do not rebuild tokenisation, scaling, encoding, or feature selection manually in the API. Save a single pipeline whenever possible.

    from sklearn.pipeline import Pipeline
    from sklearn.feature_extraction.text import TfidfVectorizer
    from sklearn.linear_model import LogisticRegression
    import joblib
    
    pipeline = Pipeline([
        ("vectorizer", TfidfVectorizer(max_features=20_000)),
        ("classifier", LogisticRegression(max_iter=1_000))
    ])
    
    pipeline.fit(train_texts, train_labels)
    joblib.dump(pipeline, "artifacts/sentiment_pipeline.joblib")

    Record the model version, Python version, dependency lockfile, training data snapshot, evaluation metrics, and expected input schema. Never load untrusted pickle or joblib files: deserialisation can execute code. Store artefacts in a controlled registry or object store and restrict who can publish them.

    For deep learning, use the framework’s stable export format or ONNX when portability matters. Test the exported model against the original model on a fixed validation set before deploying it.

    2. Build a typed inference API with FastAPI

    FastAPI provides request validation, OpenAPI documentation, and a straightforward deployment path. Load the model once at process startup, not once per request.

    from contextlib import asynccontextmanager
    from fastapi import FastAPI, HTTPException
    from pydantic import BaseModel, Field
    import joblib
    
    model = None
    
    @asynccontextmanager
    async def lifespan(app: FastAPI):
        global model
        model = joblib.load("artifacts/sentiment_pipeline.joblib")
        yield
        model = None
    
    app = FastAPI(title="Sentiment Inference API", lifespan=lifespan)
    
    class PredictionRequest(BaseModel):
        text: str = Field(min_length=1, max_length=5_000)
    
    class PredictionResponse(BaseModel):
        label: str
        confidence: float
        model_version: str
    
    @app.post("/v1/predict", response_model=PredictionResponse)
    def predict(request: PredictionRequest):
        if model is None:
            raise HTTPException(status_code=503, detail="Model is unavailable")
    
        probabilities = model.predict_proba([request.text])[0]
        index = probabilities.argmax()
        return PredictionResponse(
            label=str(model.classes_[index]),
            confidence=float(probabilities[index]),
            model_version="sentiment-2026-01"
        )

    Use a versioned route such as /v1/predict, explicit response schemas, bounded input sizes, and meaningful error codes. A confidence score is not automatically a calibrated probability; expose it only after validating calibration and explain what it means to users.

    For CPU-bound inference, async def does not make the model itself asynchronous. Run multiple worker processes, or move expensive work to a task queue. Avoid blocking database or model calls inside an event loop.

    3. Connect the frontend without creating a fragile UX

    The frontend should treat prediction as a network operation, not a local function call. Show loading, success, validation-error, timeout, and service-unavailable states. Disable duplicate submissions or attach an idempotency key where repeated requests could create cost or side effects.

    async function predict(text) {
      const response = await fetch("/api/ml/v1/predict", {
        method: "POST",
        headers: { "Content-Type": "application/json" },
        body: JSON.stringify({ text })
      });
    
      if (!response.ok) {
        throw new Error("Prediction service unavailable");
      }
      return response.json();
    }

    Keep secrets and provider credentials on the server. Configure CORS narrowly, use HTTPS, apply authentication and rate limits, and never trust a client-supplied user ID or model version. For long-running work, return 202 Accepted with a job ID and expose a status endpoint. WebSockets are useful for progress updates, but polling is often simpler and more reliable for an initial release.

    If your application is an education product, the same architecture can support adaptive recommendations; compare the design with personalized AI learning assistants for CBSE students and consider accessibility, multilingual input, and parental privacy from the start.

    4. Test the complete prediction path

    Unit tests for the model are not enough. Add tests for:

    • Schema validation, empty inputs, maximum lengths, malformed JSON, and unsupported file types.
    • Preprocessing parity between offline evaluation and the API.
    • Response shape, confidence ranges, error codes, and model version headers.
    • Authentication, rate limiting, CORS, and prompt or payload abuse where applicable.
    • Load behaviour at expected peak traffic and graceful failure when the model is unavailable.

    Maintain a small, versioned golden dataset. Every new model must produce acceptable predictions on this set before deployment. Track latency percentiles, not only averages: p95 and p99 reveal the experience of users who encounter slow requests.

    5. Deploy with repeatability and a rollback path

    Package the service in a container with a pinned Python version and lockfile. A typical production command is:

    uvicorn app:app --host 0.0.0.0 --port 8000 --workers 2

    Choose worker counts through load testing; more workers can exhaust memory when each process loads a large model. Add health endpoints that distinguish liveness from readiness: a live process may be running while the model artefact is unavailable. Store artefacts outside the image when appropriate, but pin the exact checksum and verify it at startup.

    For Indian users, select a region that meets your latency, availability, and data-residency requirements. Mumbai and Hyderabad regions are common options across major cloud providers, but measure performance from your actual user locations. Use a CDN for static assets, private networking for internal services, and autoscaling based on CPU, memory, queue depth, or GPU utilisation rather than request count alone.

    A model gateway or dedicated serving platform becomes worthwhile when you need dynamic batching, GPU sharing, canary releases, or several independently deployed models. For broader infrastructure decisions, see scalable machine learning infrastructure for developers.

    6. Monitor quality, cost, and privacy

    Log request IDs, model version, latency, status code, feature-shape metadata, and prediction outcomes. Avoid storing raw text, images, phone numbers, or other personal data unless there is a documented purpose, access control, retention period, and user notice. India’s Digital Personal Data Protection framework should be treated as an engineering requirement, not a final compliance check.

    Monitor:

    • Service health: error rate, timeouts, saturation, p50/p95/p99 latency, and queue depth.
    • Data quality: missing fields, unexpected categories, language mix, file sizes, and distribution changes.
    • Model quality: delayed labels, precision and recall by important cohorts, calibration, and drift.
    • Economics: cost per prediction, storage, GPU utilisation, and cache hit rate.

    Create alerts with owners and runbooks. A drift alert should lead to investigation, not automatic retraining. Retraining requires fresh labels, a reproducible dataset, bias checks, offline evaluation, staged rollout, and a tested rollback.

    A practical launch checklist

    Before exposing the feature to users, confirm that you have:

    • A versioned model pipeline and reproducible artefact.
    • Validated request and response schemas.
    • Authentication, rate limiting, HTTPS, and restricted CORS.
    • Unit, integration, load, and adversarial-input tests.
    • Health checks, structured logs, metrics, and alerts.
    • A canary or shadow deployment and a one-command rollback.
    • A documented data-retention and deletion policy.
    • A fallback experience when inference is slow or unavailable.

    The fastest route to a useful ML web application is not adding the largest model. It is making the full path dependable: consistent features, predictable APIs, understandable UX, measurable quality, and safe operations. Once that foundation works, you can upgrade models, add batching, or move to specialised serving without rewriting the product.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.