0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · developing small language models for personal websites

Developing Small Language Models for Personal Websites

  1. aigi

    Small language models (SLMs) make it practical to add useful AI to a portfolio, blog, documentation site, or independent business website without sending every visitor query to a commercial API. The strongest use cases are narrow: answering questions about your work, finding relevant articles, explaining a project, drafting a contact response, or switching between English and an Indian language.

    The goal is not to reproduce a frontier model. It is to build a reliable, bounded assistant that knows your public content, responds quickly, protects private data, and remains affordable to operate.

    Start with a narrow website job

    Define the assistant before choosing a model. A personal website rarely needs open-ended reasoning. It needs a clear contract such as:

    • Answer questions about projects, services, publications, or experience.
    • Retrieve and summarise posts, documentation, or case studies.
    • Guide visitors to the right page or contact channel.
    • Translate or explain selected content in Hindi, Tamil, Bengali, Marathi, or another target language.
    • Draft—but never automatically send—an enquiry response.

    This focus reduces model size, prompt complexity, and evaluation cost. If your primary requirement is a personalised portfolio, first review patterns in building personalised portfolio websites using AI agents, then decide whether an SLM adds value beyond search and structured navigation.

    Choose the deployment boundary first

    There are three practical architectures.

    Browser inference

    A quantised model runs on the visitor’s device through WebGPU, WebLLM, or Transformers.js. This offers strong privacy and eliminates inference charges, but the initial download can be hundreds of megabytes. Some phones and older browsers will not support the required hardware.

    Use browser inference when your model is small, your audience accepts a download, and the assistant can work without confidential server-side data. Display download size, supported browsers, and a non-AI fallback clearly.

    Server-side inference

    A model runs on a VPS or dedicated machine behind your website API. Tools such as llama.cpp, Ollama, or compatible inference servers can support quantised models with modest RAM. This works across devices and makes it easier to update the model, but you must manage concurrency, abuse prevention, monitoring, and hosting costs.

    For a low-traffic personal site, a CPU-based 1B–3B model may be adequate. GPU hosting is useful only when latency, concurrent visitors, or longer context windows justify it. Do not select a large GPU instance before measuring actual demand.

    Hybrid routing

    Use retrieval and simple rules first, then invoke a local model only when needed. You can also route unsupported devices to server inference and return keyword search results when the model is unavailable. This is usually the best production design for a public website.

    Select a model for your language and task

    Model choice should reflect licence terms, language coverage, context length, hardware, and expected answer quality—not parameter count alone. Current compact families include Gemma, Phi, Qwen, Llama, and Mistral variants, but verify the exact 2026 release, licence, tokenizer, and Indian-language performance before deployment.

    A useful selection process is:

    1. Build a test set of 50–150 real visitor questions.
    2. Compare two or three compact models at the same quantisation level.
    3. Measure factuality, refusal behaviour, latency, memory use, and language quality.
    4. Choose the smallest model that passes your acceptance threshold.

    For multilingual websites, test code-switching and transliteration separately. A model that performs well in English may produce poor Hindi-English mixed responses. Work involving Indic language coverage can also benefit from the techniques described in this low-resource Indic natural language processing guide.

    Use retrieval before fine-tuning

    Most personal-site assistants do not need their facts baked into model weights. Build a retrieval-augmented generation (RAG) pipeline instead:

    • Extract clean text from Markdown, HTML, PDFs, and project repositories.
    • Remove navigation, cookie notices, duplicate footers, and outdated pages.
    • Split content into meaningful sections rather than arbitrary tiny fragments.
    • Store embeddings with the page title, URL, date, language, and access status.
    • Retrieve a small number of relevant passages for each question.
    • Instruct the model to answer only from supplied sources and cite the page URL.

    RAG lets you update a biography or project page without retraining. It also reduces hallucinations and makes answers auditable. For a small site, a local vector index or an embedded database is often sufficient; avoid adding a managed search service until traffic or content volume requires it.

    Fine-tune only for behaviour and style

    Fine-tuning is appropriate when you need a consistent tone, response format, classification behaviour, or language style that prompting cannot reliably produce. It is not the best way to memorise frequently changing facts.

    Use supervised examples that reflect real interactions: concise answers, source-aware refusals, multilingual replies, and clear escalation to a contact form. Parameter-efficient methods such as LoRA or QLoRA reduce compute and keep the adapter portable. Keep private drafts, unpublished work, client information, and personal identifiers out of training data unless you have a documented reason and proper consent.

    A strong system instruction should say what the assistant knows, what it must not claim, when it should say “I don’t know,” and how it should link to evidence. Treat prompt instructions as part of the product specification, not as a substitute for testing.

    Quantise and measure the real footprint

    Quantisation can reduce memory and improve inference speed, but quality loss varies by task and language. Common formats include GGUF for llama.cpp-based deployment and browser-compatible formats supported by WebLLM or Transformers.js. Benchmark the complete application, not only the model:

    • Peak RAM or VRAM during loading and generation.
    • First-token latency and tokens per second.
    • Download size and cache behaviour in browsers.
    • Performance on low-end Android devices and typical Indian broadband.
    • Accuracy after quantisation on your evaluation set.

    Keep the original model and conversion settings so you can reproduce a release. Pin versions and record the model licence in your repository or deployment notes.

    Build a safe website API

    A public AI endpoint will attract automated traffic. Add rate limits by IP or session, maximum input and output tokens, request timeouts, logging with redaction, and a daily usage budget. Never expose private retrieval documents through a client-side index. Separate public knowledge from administrator notes and unpublished content.

    Return citations or source links whenever the assistant makes factual claims. Provide a correction or contact route, and avoid presenting the model as a legal, medical, financial, or employment authority. If your website serves students or exam candidates, the design principles in this personalised AI mentor for competitive exam preparation in India are useful for thinking about boundaries and escalation.

    Design the fallback experience

    An AI feature should not block the core website. Always provide ordinary navigation, site search, a contact form, and a way to report an incorrect answer. Stream responses where possible, show a brief loading state, and preserve the visitor’s question if generation fails. For browser inference, explain why a model is downloading and let users cancel.

    Track anonymous operational metrics such as failure rate, response latency, unanswered-question categories, and citation clicks. Do not collect chat transcripts by default. If you retain them for improvement, state the purpose, retention period, access controls, and deletion process.

    A practical build sequence

    1. Define one high-value assistant task and its failure boundaries.
    2. Clean and version the public content corpus.
    3. Implement retrieval and citations without fine-tuning.
    4. Benchmark compact models on real questions and target languages.
    5. Quantise the best candidate and test on representative devices.
    6. Add rate limits, privacy controls, fallbacks, and monitoring.
    7. Fine-tune only if style or formatting remains inconsistent.
    8. Re-test after every content, model, or prompt change.

    For most personal websites, a small quantised model plus strong retrieval beats an oversized model trained on a messy corpus. Build the narrowest useful system, make its limits visible, and let the website’s content—not the model’s confidence—define what it can claim.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.