Netlify is a strong fit for small, CPU-friendly machine learning inference services, not for training models or hosting GPU-heavy workloads. If your model can respond within a serverless function’s runtime, you can place a web interface and an inference API on the same platform, simplify operations, and scale from a prototype to a modest production workload.
This guide shows how to deploy machine learning models on Netlify in 2026 using ONNX Runtime and Node.js. It also explains when Netlify is the wrong choice, how to package model files safely, how to reduce cold-start latency, and how to build an API that is usable beyond a demo.
When Netlify is the right deployment choice
Netlify works well when your application has:
- A lightweight model, generally exported to ONNX or another compact JavaScript-compatible format.
- Short inference requests with predictable input and output sizes.
- CPU-only inference requirements.
- Traffic that is variable, seasonal, or too small to justify a continuously running server.
- A frontend that benefits from being deployed alongside the API.
Typical use cases include tabular classification, regression, recommendation scoring, image metadata classification, and compact computer-vision models. A crop-risk classifier, a local-language sentiment tool, or a document triage service for an Indian startup may fit this pattern.
Netlify is not suitable for model training, GPU inference, long-running batch jobs, streaming generation, or large language models that need substantial memory. For those workloads, use a dedicated inference provider and call it from a Netlify Function. The same separation is useful when building open-source AI agents in production: Netlify can host the application layer while model execution runs elsewhere.
Recommended architecture
A practical Netlify ML application has four layers:
1. Frontend: Collects input, performs safe client-side validation, and displays the result.
2. Netlify Function: Authenticates the request, validates the payload, invokes the model, and returns JSON.
3. Model runtime: ONNX Runtime for Node.js loads and executes the exported model.
4. Observability and controls: Logs latency and failures without exposing sensitive user data.
The model should be treated as a versioned application asset, not as an arbitrary file downloaded during every request. Load it from the function bundle where possible. A warm function container can reuse an initialized inference session, avoiding repeated model loading on subsequent requests.
Export and validate the model
Export the model in the training environment, then validate that its predictions match the original implementation. For a scikit-learn model, skl2onnx is a common option:
import skl2onnx
from skl2onnx.common.data_types import FloatTensorType
initial_type = [('float_input', FloatTensorType([None, 4]))]
onx = skl2onnx.convert_sklearn(
model,
initial_types=initial_type
)
with open("model.onnx", "wb") as file:
file.write(onx.SerializeToString())Do not assume that a successful export means the model is production-ready. Test the ONNX output against the original model using representative examples, boundary values, missing fields, and malformed input. Record the expected input shape, feature order, data type, preprocessing steps, and output names in a small model card or README.
If your project is still at the experimentation stage, compare deployment constraints before choosing the model. Guides to machine learning portfolio projects for beginners in India and machine learning projects for computer science students can help narrow a project to a model that is realistic to deploy.
Create the Netlify project
A minimal project can use this structure:
project/
├── netlify/
│ └── functions/
│ └── predict.js
├── models/
│ └── model.onnx
├── public/
│ └── index.html
├── package.json
└── netlify.tomlInstall the runtime:
npm install onnxruntime-nodeKeep development tools, training frameworks, notebooks, and unused assets out of the deployment package. Check the generated function bundle during CI so a new dependency does not silently push the deployment beyond Netlify’s limits.
Build a production-minded inference function
The following example loads the ONNX session once per warm container, validates the request, and returns controlled error messages:
const ort = require('onnxruntime-node');
const path = require('path');
let sessionPromise;
function getSession() {
if (!sessionPromise) {
const modelPath = path.join(__dirname, '../../models/model.onnx');
sessionPromise = ort.InferenceSession.create(modelPath);
}
return sessionPromise;
}
exports.handler = async (event) => {
if (event.httpMethod !== 'POST') {
return { statusCode: 405, body: JSON.stringify({ error: 'Method not allowed' }) };
}
try {
const body = JSON.parse(event.body || '{}');
const features = body.features;
if (!Array.isArray(features) || features.length !== 4) {
return { statusCode: 400, body: JSON.stringify({ error: 'Expected four numeric features' }) };
}
if (!features.every(Number.isFinite)) {
return { statusCode: 400, body: JSON.stringify({ error: 'Features must be finite numbers' }) };
}
const session = await getSession();
const input = new ort.Tensor('float32', Float32Array.from(features), [1, 4]);
const output = await session.run({ float_input: input });
return {
statusCode: 200,
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ prediction: Array.from(output.variable.data) })
};
} catch (error) {
console.error('Inference failed:', error.message);
return { statusCode: 500, body: JSON.stringify({ error: 'Inference failed' }) };
}
};Update the output key, input name, tensor shape, and preprocessing to match your exported model. Never return raw stack traces to users. If the endpoint receives personal, health, education, or financial data, minimise logs and define retention rules before launch.
Configure model bundling and routes
Use netlify.toml to identify the functions directory, include the model, and expose a readable API path:
[functions]
directory = "netlify/functions"
included_files = ["models/model.onnx"]
[[redirects]]
from = "/api/predict"
to = "/.netlify/functions/predict"
status = 200The exact path resolution can vary with your build setup, so test using a deployed preview rather than relying only on local execution. Confirm that the model is present in the generated function bundle and that the function can find it after deployment.
Reduce size, latency, and failure risk
Serverless ML performance depends as much on packaging as on model architecture. Apply these practices:
- Quantise where accuracy permits: INT8 weights can materially reduce model size and memory use.
- Remove unused graph components: Simplify or prune the model before export.
- Keep preprocessing consistent: A fast model with mismatched scaling produces unreliable predictions.
- Cache the session: Initialise
InferenceSessionoutside the handler’s request path. - Limit payloads: Reject oversized JSON and avoid sending raw images when a compact representation is enough.
- Use browser inference selectively: Small models may run in the browser through ONNX Runtime Web, reducing server cost and improving privacy.
- Measure cold and warm latency: Report both, because users experience both types of request.
For models intended for phones or low-connectivity environments, compare serverless inference with the techniques in AI model optimization for mobile devices. Browser or mobile execution may be a better fit than a remote function.
Test before production
Create automated tests for preprocessing, model outputs, invalid payloads, missing fields, oversized requests, and concurrent calls. Use a small golden dataset whose predictions are reviewed by the team. Then deploy a preview and test the real function bundle, not just the local development server.
Add an authentication layer or rate limit before exposing a costly endpoint publicly. Set CORS deliberately, validate content types, and never place API keys in frontend code. Track request count, error rate, p50 and p95 latency, cold-start frequency, and model version. A prediction endpoint without basic monitoring is difficult to debug when users are distributed across Indian networks and devices.
Common failure modes
The deployment exceeds the bundle limit: remove training libraries, quantise the model, or move inference to a dedicated service.
The function cannot find the model: verify included_files, the deployed path, and case-sensitive filenames.
Predictions are incorrect: check feature order, scaling, tensor names, shape, and ONNX-versus-training outputs.
Requests time out: reduce model complexity, shrink payloads, or switch to an asynchronous job architecture.
Cold starts are too slow: reduce dependencies and model size, cache the session, or move the model to a service designed for persistent workers.
Final decision checklist
Netlify is a sensible choice when your model is compact, CPU-compatible, quick to execute, and attached to a web product. It is a poor choice when the workload needs GPUs, persistent processes, large artefacts, or long-running inference. Make that decision using measured latency and bundle size—not assumptions.
For Indian builders, the platform is particularly useful for launching focused products without maintaining a full backend fleet. Start with a narrow model and a validated API contract, then move inference to specialised infrastructure when usage or model complexity demands it.