0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build full stack web apps with fastai

How to Build Full-Stack Web Apps with fastai

  1. aigi

    fastai makes model training unusually productive, but a notebook is not a product. To turn a trained Learner into a reliable web application, you need clear boundaries between inference, the API, the user interface, and operations. You also need to account for model versioning, input validation, privacy, latency, and cloud cost.

    This guide explains how to build full-stack web apps with fastai using a practical stack: fastai and PyTorch for inference, FastAPI for the backend, React or Next.js for the frontend, and Docker for repeatable deployment. The examples use image classification, but the same structure applies to tabular and text models. If your product is specifically computer-vision focused, pair this guide with how to build computer vision models on GitHub for dataset and collaboration practices.

    Choose the right application architecture

    A maintainable fastai application usually has four parts:

    • Model package: the exported learner, preprocessing code, labels, and metadata.
    • Inference service: a FastAPI application that validates requests and returns stable JSON.
    • Web client: a React or Next.js interface for uploads, progress states, errors, and results.
    • Operations layer: Docker, logs, metrics, authentication, rate limits, and deployment configuration.

    Keep the model package independent from HTTP concerns. Your prediction function should accept a validated Python object and return a predictable result; the FastAPI route should only handle transport, authentication, and error mapping. This separation makes local tests easier and lets you replace the frontend without retraining the model.

    For products that later add conversational workflows, keep inference services separate from agent orchestration. The design principles in how to build AI research assistant tools are useful when a simple prediction endpoint grows into a multi-step application.

    1. Export and verify the fastai learner

    Export the learner from the training environment:

    from pathlib import Path
    
    learn.export(Path("artifacts/model.pkl"))

    An exported learner includes the architecture, weights, vocabulary or labels, and much of the data-processing pipeline. It can still fail in production if the runtime does not contain custom transforms, classes, or compatible library versions. Record these details alongside the model:

    • Python, fastai, and PyTorch versions
    • Model name and training dataset version
    • Expected input size, channels, and file types
    • Label mapping and confidence interpretation
    • Evaluation metrics, including performance on Indian-language, regional, or low-quality inputs where relevant

    Load the artifact in a clean environment before integrating it with FastAPI. Test valid files, corrupt files, oversized uploads, and inputs outside the training distribution. A confidence score is not proof that a prediction is correct; expose an abstain or review state when the model is uncertain.

    2. Build a predictable FastAPI inference service

    Install only the dependencies needed at runtime. Keep training libraries and notebooks out of the production image where possible.

    pip install fastapi uvicorn[standard] python-multipart fastai pillow

    A minimal service can look like this:

    from contextlib import asynccontextmanager
    from io import BytesIO
    from pathlib import Path
    
    from fastapi import FastAPI, File, HTTPException, UploadFile
    from fastai.vision.all import PILImage, load_learner
    
    learner = None
    
    @asynccontextmanager
    async def lifespan(app: FastAPI):
        global learner
        learner = load_learner(Path("artifacts/model.pkl"), cpu=True)
        yield
        learner = None
    
    app = FastAPI(title="fastai inference API", lifespan=lifespan)
    
    @app.get("/health")
    def health():
        return {"status": "ok", "model_loaded": learner is not None}
    
    @app.post("/v1/predict")
    async def predict(file: UploadFile = File(...)):
        if file.content_type not in {"image/jpeg", "image/png", "image/webp"}:
            raise HTTPException(status_code=415, detail="Unsupported image type")
    
        content = await file.read()
        if len(content) > 5 * 1024 * 1024:
            raise HTTPException(status_code=413, detail="File is too large")
    
        try:
            image = PILImage.create(BytesIO(content))
            label, index, probabilities = learner.predict(image)
            return {
                "label": str(label),
                "confidence": float(probabilities[index]),
                "model_version": "2026-01"
            }
        except Exception as exc:
            raise HTTPException(status_code=400, detail="Could not process image") from exc

    The async route handles file I/O, but model inference itself is typically CPU-bound. Do not assume that adding async makes prediction parallel. For modest traffic, one process with a loaded model may be enough. For higher traffic, benchmark multiple workers carefully: each worker can load its own copy of the model and multiply RAM usage.

    Use versioned endpoints such as /v1/predict, structured errors, request IDs, and a health check that distinguishes a running process from a successfully loaded model. Never return raw exception messages to users.

    3. Design the React or Next.js frontend

    The frontend should handle the complete request lifecycle: file selection, client-side validation, upload progress, successful prediction, failure, and retry. A basic request looks like this:

    async function predict(file) {
      const formData = new FormData();
      formData.append("file", file);
    
      const response = await fetch(`${API_URL}/v1/predict`, {
        method: "POST",
        body: formData,
      });
    
      if (!response.ok) {
        const error = await response.json().catch(() => ({}));
        throw new Error(error.detail || "Prediction failed");
      }
    
      return response.json();
    }

    Validate file type and size in the browser for quick feedback, but repeat every check on the server. Compress large images before upload when appropriate, especially for users on mobile networks. Do not make the UI imply certainty from a probability value; show labels such as likely, needs review, or unsupported input based on a threshold chosen during evaluation.

    Keep the API URL in environment configuration rather than hard-coding it. For authenticated products, use short-lived tokens or secure cookies and avoid placing secrets in client-side JavaScript.

    4. Secure the application before deployment

    CORS is not an authentication mechanism. Configure it with an explicit allowlist instead of *:

    from fastapi.middleware.cors import CORSMiddleware
    
    app.add_middleware(
        CORSMiddleware,
        allow_origins=["https://app.example.in"],
        allow_credentials=True,
        allow_methods=["POST", "GET"],
        allow_headers=["Authorization", "Content-Type"],
    )

    Add authentication, per-user rate limits, upload limits, timeouts, and abuse monitoring. Strip or ignore untrusted filenames, scan uploads where the risk profile requires it, and do not store images by default. If the application handles health, legal, education, or workplace data, define retention and deletion rules before launch. For a privacy-sensitive legal workflow, the architecture in how to build a private AI chatbot for lawyers offers relevant principles around data control and deployment boundaries.

    5. Containerise and deploy efficiently

    A production Dockerfile should pin dependencies and avoid unnecessary build tools in the final image:

    FROM python:3.11-slim
    WORKDIR /app
    
    RUN apt-get update && apt-get install -y --no-install-recommends \
        libglib2.0-0 libgl1 && rm -rf /var/lib/apt/lists/*
    
    COPY requirements.txt .
    RUN pip install --no-cache-dir -r requirements.txt
    COPY app ./app
    COPY artifacts ./artifacts
    
    CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000"]

    Pin compatible versions, run the container as a non-root user, and add a .dockerignore so datasets and notebooks do not enter the image. Build separate CPU and GPU images only when benchmarks justify GPU infrastructure. A CPU service is often the better starting point for low or intermittent traffic.

    For Indian users, test Mumbai or Hyderabad regions where available and measure real latency rather than assuming proximity guarantees performance. Cloud Run, ECS, managed Kubernetes, and VM deployments can all work; choose based on traffic predictability, cold-start tolerance, and operational skill. Compress frontend assets, keep model downloads out of request paths, and monitor egress costs for image-heavy workloads.

    6. Test, observe, and improve the model

    Before launch, test at four levels:

    • Unit tests: preprocessing, label mapping, thresholds, and error handling.
    • API tests: valid uploads, invalid MIME types, oversized files, authentication, and timeouts.
    • UI tests: loading, retry, offline, and accessibility states.
    • Evaluation tests: accuracy, calibration, false positives, and performance across relevant Indian user segments and devices.

    Log request IDs, latency, model version, status code, and failure category. Avoid logging raw images or personally identifiable information. Track p50 and p95 latency, memory use, queue depth, error rates, and the percentage of low-confidence predictions. Store model artifacts in versioned object storage and make rollback a documented command, not an emergency improvisation.

    Common mistakes to avoid

    • Loading the learner inside every request instead of once at startup.
    • Running too many workers and exhausting RAM.
    • Treating async as a solution for CPU-bound inference.
    • Using permissive CORS in production.
    • Returning confidence without calibration or an uncertainty policy.
    • Requiring the frontend to understand model internals.
    • Shipping unpinned dependencies and an untested model artifact.
    • Ignoring regional latency, mobile bandwidth, and data-retention requirements.

    A fastai web app becomes production-ready when the prediction is only one part of a dependable system. Start with a narrow endpoint, a versioned artifact, explicit input contracts, and measured deployment targets. Then add authentication, queues, batch processing, or specialised hardware only when observed usage requires them.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.