GitHub Pages can host an AI demo, but it cannot run a Python server, keep a GPU process alive, or protect private model files. The workable pattern is client-side inference: GitHub Pages serves HTML, JavaScript, and model assets, while the visitor’s browser performs prediction with WebAssembly, WebGL, or WebGPU.
That makes GitHub Pages a strong fit for public prototypes, portfolios, classroom projects, research demonstrations, and lightweight production tools. It is not a replacement for an API when you need confidential weights, centralised monitoring, predictable latency, or models too large for practical browser downloads.
Choose the right deployment architecture
Before converting a model, decide where inference should happen:
- Browser inference: Best for demos, privacy-sensitive inputs, offline-friendly tools, and zero-cost hosting. Use TensorFlow.js, ONNX Runtime Web, Transformers.js, or WebLLM.
- Remote inference API: Better for large models, proprietary weights, GPU workloads, billing controls, and consistent results across devices. GitHub Pages remains the frontend.
- Hybrid deployment: Run preprocessing or small models in the browser, then call an API only for expensive operations.
For example, a handwritten-digit classifier or a compact image model is a sensible browser workload. A multi-billion-parameter language model may technically run with WebGPU, but its download size, memory demand, and mobile performance can make the experience unusable. If your work involves heavier agents, compare this approach with how to deploy open-source AI agents in production.
Select a browser-compatible runtime
Your training framework does not determine your deployment format. Convert the model into a format supported by the browser runtime:
- TensorFlow.js: Useful for Keras and TensorFlow models. It supports browser backends including WebGL and WebGPU.
- ONNX Runtime Web: A practical choice for models exported from PyTorch, scikit-learn, or other frameworks. It can use Wasm, WebGL, or WebGPU.
- Transformers.js: Suitable for supported NLP, vision, audio, and multimodal models from the Hugging Face ecosystem.
- WebLLM: Designed for selected large language models using WebGPU, though model downloads can be substantial.
Check operator support, tokenizer compatibility, precision options, and browser requirements before committing to a format. A model that exports successfully may still fail in the browser because an operation is unsupported or because the input and output tensor shapes differ from your JavaScript code.
If your target is Indian-language AI, test the actual tokenizer and language coverage rather than assuming that a general model will perform well. The available open-source vision-language models for Indian languages can help you evaluate realistic alternatives.
Optimise the model before publishing it
The largest performance improvement usually comes before deployment. Measure both download size and inference time on representative Indian devices and networks.
Use these techniques where the runtime supports them:
- Quantisation: Convert float32 weights to float16 or int8. Validate accuracy after conversion, especially for OCR, medical, and low-light vision tasks.
- Pruning: Remove low-value weights when your architecture and runtime benefit from sparsity.
- Distillation: Train a smaller student model for browser use instead of shipping the full research model.
- Sharding: Split weights into manageable files so the browser can download them progressively.
- Lazy loading: Load the model only after the user selects the feature that needs it.
- Warm-up inference: Run one small prediction after loading to initialise the backend and reveal failures early.
For mobile-first applications, follow a structured AI model optimisation workflow for mobile devices. Set a practical target: a user on a mid-range Android phone should not wait minutes or lose the page because the browser runs out of memory.
Build a predictable project structure
A small static project can look like this:
ai-demo/
├── index.html
├── src/
│ ├── main.js
│ └── style.css
├── public/
│ └── models/
│ ├── model.onnx
│ └── tokenizer.json
├── package.json
└── vite.config.jsFor a plain JavaScript site, place files directly under the repository root. With Vite, set the repository path as the base URL when deploying to https://username.github.io/repository-name/:
// vite.config.js
import { defineConfig } from 'vite';
export default defineConfig({
base: '/repository-name/'
});Use URL-safe, case-sensitive paths and test the production build locally. A path that works on macOS can fail on GitHub Pages because the deployed environment treats Model.onnx and model.onnx as different files.
Implement inference with a loading state
A usable demo tells the visitor what is happening. Do not block the interface while weights download or inference runs. A minimal ONNX Runtime Web pattern is:
import * as ort from 'onnxruntime-web';
const status = document.querySelector('#status');
async function classify(inputData) {
status.textContent = 'Loading model…';
const session = await ort.InferenceSession.create(
`${import.meta.env.BASE_URL}models/model.onnx`,
{ executionProviders: ['wasm'] }
);
const input = new ort.Tensor('float32', inputData, [1, 3, 224, 224]);
const output = await session.run({ input });
status.textContent = 'Complete';
return output;
}In production, create the session once and reuse it rather than loading the model for every prediction. Add error handling for unsupported browsers, failed downloads, invalid input dimensions, and insufficient memory. If WebGPU is available, feature-detect it and provide Wasm as a fallback; never assume every phone or browser supports the fastest backend.
Deploy with GitHub Actions
A repeatable deployment should build the site in CI rather than relying on manually committed generated files.
1. Open Settings → Pages in the repository.
2. Select GitHub Actions as the deployment source.
3. Add a workflow that installs dependencies, runs the production build, and uploads the build directory.
4. Grant the workflow permission to publish Pages.
5. Test the final URL, including direct navigation to nested routes and model asset URLs.
Keep training data, API keys, local checkpoints, and experiment logs outside the published directory. GitHub Pages content is public, and browser-delivered weights can be downloaded by anyone. Git LFS does not make a model private, and it is not a dependable distribution strategy for large public browser assets. For bigger files, use a suitable model host or CDN and configure CORS correctly.
Handle browser and hosting constraints
GitHub Pages cannot configure arbitrary response headers. This matters for features requiring cross-origin isolation, including some multithreaded Wasm configurations. If your runtime needs SharedArrayBuffer, verify whether a service-worker workaround is acceptable and test it across browsers; otherwise select a single-threaded or less demanding backend.
Also account for:
- CORS: External model hosts must allow requests from your Pages origin.
- HTTPS: GitHub Pages provides HTTPS, which is required by several browser capabilities.
- Caching: Use versioned model filenames and cache headers where your asset host permits them.
- Consent and privacy: Explain that uploaded images, audio, or text are processed locally only if that is genuinely true.
- Accessibility: Provide keyboard controls, readable progress text, and a non-WebGPU fallback.
Test for real Indian users
A desktop developer machine is not a sufficient benchmark. Test on a mid-range Android phone, a lower-memory laptop, Chrome and Safari, and a slower mobile connection. Record first-load time, model download size, peak memory, inference latency, and failure rate.
Prefer compressed assets, resumable downloads where available, and a clear offline or retry experience. IndexedDB can cache model files after the first successful load, but storage quotas vary and users can clear them. A progress indicator should report download progress when the server exposes a usable Content-Length; otherwise show an honest indeterminate state.
When GitHub Pages is the wrong choice
Move inference to a backend when you need private weights, GPU-intensive generation, centralised rate limiting, audit logs, user accounts, or dependable performance on low-end devices. GitHub Pages is excellent for a transparent public demo, not for hiding intellectual property or operating a regulated inference service.
For a lightweight portfolio, however, it offers an effective path from trained model to shareable product: optimise the model, ship only browser-compatible assets, feature-detect acceleration, measure on real devices, and automate deployment. Developers building a portfolio can also explore how to contribute to AI GitHub repositories in India to turn a demo into a more credible open-source project.