fastai makes model training unusually productive, but a notebook is not a product. To turn a trained Learner into a reliable web application, you need clear boundaries between inference, the API, the user interface, and operations. You also need to account for model versioning, input validation, privacy, latency, and cloud cost.
This guide explains how to build full-stack web apps with fastai using a practical stack: fastai and PyTorch for inference, FastAPI for the backend, React or Next.js for the frontend, and Docker for repeatable deployment. The examples use image classification, but the same structure applies to tabular and text models. If your product is specifically computer-vision focused, pair this guide with how to build computer vision models on GitHub for dataset and collaboration practices.
Choose the right application architecture
A maintainable fastai application usually has four parts:
- Model package: the exported learner, preprocessing code, labels, and metadata.
- Inference service: a FastAPI application that validates requests and returns stable JSON.
- Web client: a React or Next.js interface for uploads, progress states, errors, and results.
- Operations layer: Docker, logs, metrics, authentication, rate limits, and deployment configuration.
Keep the model package independent from HTTP concerns. Your prediction function should accept a validated Python object and return a predictable result; the FastAPI route should only handle transport, authentication, and error mapping. This separation makes local tests easier and lets you replace the frontend without retraining the model.
For products that later add conversational workflows, keep inference services separate from agent orchestration. The design principles in how to build AI research assistant tools are useful when a simple prediction endpoint grows into a multi-step application.
1. Export and verify the fastai learner
Export the learner from the training environment:
from pathlib import Path
learn.export(Path("artifacts/model.pkl"))An exported learner includes the architecture, weights, vocabulary or labels, and much of the data-processing pipeline. It can still fail in production if the runtime does not contain custom transforms, classes, or compatible library versions. Record these details alongside the model:
- Python, fastai, and PyTorch versions
- Model name and training dataset version
- Expected input size, channels, and file types
- Label mapping and confidence interpretation
- Evaluation metrics, including performance on Indian-language, regional, or low-quality inputs where relevant
Load the artifact in a clean environment before integrating it with FastAPI. Test valid files, corrupt files, oversized uploads, and inputs outside the training distribution. A confidence score is not proof that a prediction is correct; expose an abstain or review state when the model is uncertain.
2. Build a predictable FastAPI inference service
Install only the dependencies needed at runtime. Keep training libraries and notebooks out of the production image where possible.
pip install fastapi uvicorn[standard] python-multipart fastai pillowA minimal service can look like this:
from contextlib import asynccontextmanager
from io import BytesIO
from pathlib import Path
from fastapi import FastAPI, File, HTTPException, UploadFile
from fastai.vision.all import PILImage, load_learner
learner = None
@asynccontextmanager
async def lifespan(app: FastAPI):
global learner
learner = load_learner(Path("artifacts/model.pkl"), cpu=True)
yield
learner = None
app = FastAPI(title="fastai inference API", lifespan=lifespan)
@app.get("/health")
def health():
return {"status": "ok", "model_loaded": learner is not None}
@app.post("/v1/predict")
async def predict(file: UploadFile = File(...)):
if file.content_type not in {"image/jpeg", "image/png", "image/webp"}:
raise HTTPException(status_code=415, detail="Unsupported image type")
content = await file.read()
if len(content) > 5 * 1024 * 1024:
raise HTTPException(status_code=413, detail="File is too large")
try:
image = PILImage.create(BytesIO(content))
label, index, probabilities = learner.predict(image)
return {
"label": str(label),
"confidence": float(probabilities[index]),
"model_version": "2026-01"
}
except Exception as exc:
raise HTTPException(status_code=400, detail="Could not process image") from excThe async route handles file I/O, but model inference itself is typically CPU-bound. Do not assume that adding async makes prediction parallel. For modest traffic, one process with a loaded model may be enough. For higher traffic, benchmark multiple workers carefully: each worker can load its own copy of the model and multiply RAM usage.
Use versioned endpoints such as /v1/predict, structured errors, request IDs, and a health check that distinguishes a running process from a successfully loaded model. Never return raw exception messages to users.
3. Design the React or Next.js frontend
The frontend should handle the complete request lifecycle: file selection, client-side validation, upload progress, successful prediction, failure, and retry. A basic request looks like this:
async function predict(file) {
const formData = new FormData();
formData.append("file", file);
const response = await fetch(`${API_URL}/v1/predict`, {
method: "POST",
body: formData,
});
if (!response.ok) {
const error = await response.json().catch(() => ({}));
throw new Error(error.detail || "Prediction failed");
}
return response.json();
}Validate file type and size in the browser for quick feedback, but repeat every check on the server. Compress large images before upload when appropriate, especially for users on mobile networks. Do not make the UI imply certainty from a probability value; show labels such as likely, needs review, or unsupported input based on a threshold chosen during evaluation.
Keep the API URL in environment configuration rather than hard-coding it. For authenticated products, use short-lived tokens or secure cookies and avoid placing secrets in client-side JavaScript.
4. Secure the application before deployment
CORS is not an authentication mechanism. Configure it with an explicit allowlist instead of *:
from fastapi.middleware.cors import CORSMiddleware
app.add_middleware(
CORSMiddleware,
allow_origins=["https://app.example.in"],
allow_credentials=True,
allow_methods=["POST", "GET"],
allow_headers=["Authorization", "Content-Type"],
)Add authentication, per-user rate limits, upload limits, timeouts, and abuse monitoring. Strip or ignore untrusted filenames, scan uploads where the risk profile requires it, and do not store images by default. If the application handles health, legal, education, or workplace data, define retention and deletion rules before launch. For a privacy-sensitive legal workflow, the architecture in how to build a private AI chatbot for lawyers offers relevant principles around data control and deployment boundaries.
5. Containerise and deploy efficiently
A production Dockerfile should pin dependencies and avoid unnecessary build tools in the final image:
FROM python:3.11-slim
WORKDIR /app
RUN apt-get update && apt-get install -y --no-install-recommends \
libglib2.0-0 libgl1 && rm -rf /var/lib/apt/lists/*
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY app ./app
COPY artifacts ./artifacts
CMD ["uvicorn", "app.main:app", "--host", "0.0.0.0", "--port", "8000"]Pin compatible versions, run the container as a non-root user, and add a .dockerignore so datasets and notebooks do not enter the image. Build separate CPU and GPU images only when benchmarks justify GPU infrastructure. A CPU service is often the better starting point for low or intermittent traffic.
For Indian users, test Mumbai or Hyderabad regions where available and measure real latency rather than assuming proximity guarantees performance. Cloud Run, ECS, managed Kubernetes, and VM deployments can all work; choose based on traffic predictability, cold-start tolerance, and operational skill. Compress frontend assets, keep model downloads out of request paths, and monitor egress costs for image-heavy workloads.
6. Test, observe, and improve the model
Before launch, test at four levels:
- Unit tests: preprocessing, label mapping, thresholds, and error handling.
- API tests: valid uploads, invalid MIME types, oversized files, authentication, and timeouts.
- UI tests: loading, retry, offline, and accessibility states.
- Evaluation tests: accuracy, calibration, false positives, and performance across relevant Indian user segments and devices.
Log request IDs, latency, model version, status code, and failure category. Avoid logging raw images or personally identifiable information. Track p50 and p95 latency, memory use, queue depth, error rates, and the percentage of low-confidence predictions. Store model artifacts in versioned object storage and make rollback a documented command, not an emergency improvisation.
Common mistakes to avoid
- Loading the learner inside every request instead of once at startup.
- Running too many workers and exhausting RAM.
- Treating
asyncas a solution for CPU-bound inference. - Using permissive CORS in production.
- Returning confidence without calibration or an uncertainty policy.
- Requiring the frontend to understand model internals.
- Shipping unpinned dependencies and an untested model artifact.
- Ignoring regional latency, mobile bandwidth, and data-retention requirements.
A fastai web app becomes production-ready when the prediction is only one part of a dependable system. Start with a narrow endpoint, a versioned artifact, explicit input contracts, and measured deployment targets. Then add authentication, queues, batch processing, or specialised hardware only when observed usage requires them.