What you are building
A municipal FAQ model should do more than repeat training examples. It should identify a citizen’s intent, return an accurate answer, preserve essential conditions such as ward or document requirements, and clearly say when the official department must be contacted. For most teams, the best first version is a retrieval-augmented question-answering system, not a model that memorises an entire civic website.
Fine-tuning is useful when you need consistent intent classification, question rewriting, answer formatting, or multilingual behaviour. Before choosing an approach, review best practices for fine-tuning LLMs on custom data. A small, well-labelled dataset often delivers more value than an expensive training run on noisy FAQs.
Choose the right training objective
Indian municipal FAQs commonly contain questions about property tax, water connections, birth and death certificates, waste collection, building permissions, parking, trade licences, and grievance registration. These use cases map to different objectives:
- Retrieval or semantic search: Find the relevant official FAQ or service page.
- Extractive question answering: Select an answer span from a trusted passage.
- Intent classification: Route a request to property tax, sanitation, certificates, or another department.
- Generative supervised fine-tuning: Teach a chat model to produce a particular answer format.
- Translation or rewriting: Convert Hinglish or regional-language queries into a standard search form.
For a production civic assistant, start with retrieval plus citations. Fine-tune only where evaluation shows a repeatable gap. This reduces hallucination risk and makes updates easier when fees, deadlines, portals, or procedures change.
Build a reliable municipal FAQ dataset
Collect material from official municipal and state-government sources wherever possible. Record the source URL, department, city, publication or update date, language, and date collected. Do not treat social-media posts or unverified citizen answers as authoritative without checking them against an official source.
A useful JSONL record might look like this:
{"question":"How do I pay property tax online?","answer":"Use the municipality's official property-tax portal, enter the assessment number, verify the displayed property details, and complete payment. Keep the receipt.","source_url":"https://example.gov.in/property-tax","department":"Revenue","city":"Pune","language":"en","updated_at":"2026-01-15"}Add fields that help you audit and filter the data:
question_variants: spelling errors, Hinglish, abbreviations, and speech-like phrasing.answer_type: procedure, eligibility, fee, timeline, document list, or escalation.jurisdiction: city, zone, ward, or state.effective_fromandeffective_to: essential for changing rules.source_text: the passage supporting the answer.status: verified, needs review, or retired.
Remove duplicate FAQs, obsolete procedures, personal information, phone numbers copied from unofficial sources, and contradictory answers. Split long pages into small passages, but retain enough context to preserve exceptions and eligibility conditions. Keep a human-reviewed test set separate from training data; it should include realistic misspellings, code-switching, incomplete questions, and adversarial requests.
Prepare the Hugging Face environment
Use a current Python environment and pin package versions for reproducibility. A basic setup is:
python -m venv .venv
source .venv/bin/activate
pip install -U torch transformers datasets accelerate evaluate sentence-transformersOn Windows, activate the environment with .venv\\Scripts\\activate. A GPU is helpful for generative models, but a classifier, embedding model, or extractive QA model can often be trained on a modest cloud GPU or CPU for experimentation. Track the model licence, dataset licence, and source permissions before publishing to the Hugging Face Hub.
Load and split the dataset
from datasets import load_dataset
files = {"train": "data/train.jsonl", "validation": "data/validation.jsonl"}
dataset = load_dataset("json", data_files=files)
print(dataset)Use a city- and source-aware split. Randomly splitting near-identical FAQs can produce inflated scores because the same answer appears in both training and validation. Keep entire question families, wards, or source pages together where possible. If you support multiple languages, report results separately for English, Hindi, Hinglish, and each regional language represented in the test set.
Fine-tune an extractive QA model correctly
For extractive QA, the answer must be a span inside a context passage. A question-and-answer table alone is not enough; add a context field and character offsets for the answer. Then tokenize with overflow handling:
from transformers import AutoTokenizer
model_name = "deepset/roberta-base-squad2"
tokenizer = AutoTokenizer.from_pretrained(model_name)
def tokenize_batch(batch):
return tokenizer(
batch["question"],
batch["context"],
truncation="only_second",
max_length=384,
stride=96,
return_overflowing_tokens=True,
return_offsets_mapping=True,
padding="max_length",
)
tokenized = dataset.map(tokenize_batch, batched=True, remove_columns=dataset["train"].column_names)In a complete training pipeline, convert the answer’s character start and end positions to token start and end positions. Do not assign the tokenised answer IDs as start_positions or end_positions; that is a common implementation error and will not train an extractive QA model correctly.
Train with AutoModelForQuestionAnswering and Trainer, using a small learning rate, gradient accumulation when memory is limited, and early stopping based on validation performance. Preserve the preprocessing script and configuration alongside the model checkpoint.
Fine-tune a classifier or chat model only when needed
For intent routing, label each example with one canonical intent and include a fallback class. Evaluate confusing pairs such as “new water connection” versus “water bill correction”. For a generative model, format examples consistently and teach the model to:
- answer only from supplied municipal context;
- state the city or jurisdiction when relevant;
- list documents and steps without inventing requirements;
- provide the source and retrieval date;
- escalate uncertain or emergency cases.
Use parameter-efficient methods such as LoRA or QLoRA when adapting an open model. They lower memory requirements and make it easier to maintain separate adapters for different cities or departments. Do not fine-tune sensitive citizen records. Redact names, addresses, application numbers, Aadhaar details, phone numbers, and other personal data before training.
Evaluate civic usefulness, not just loss
Report exact match and token-level F1 for extractive QA, but also run task-specific checks:
- Grounding: Is every procedural claim supported by an official passage?
- Completeness: Are fees, documents, deadlines, and exceptions retained?
- Jurisdiction accuracy: Does the answer apply to the requested municipality?
- Language quality: Does it handle code-switching and regional-language phrasing?
- Abstention: Does it refuse to guess when the source is missing or outdated?
- Safety: Does it avoid exposing personal data or giving unsafe emergency guidance?
Create a review sheet with 100–300 realistic queries per major service area, then have municipal-domain reviewers score correctness and actionability. Track failure cases by source page and intent. A retrieval baseline should be your benchmark; fine-tuning is worthwhile only if it improves citizen outcomes without reducing factual reliability.
Deploy with updates and human escalation
Publish an internal model or adapter to the Hugging Face Hub only after checking licensing, redaction, and access controls. In production, pair the model with a vector index or document store, source citations, confidence thresholds, and logs that exclude sensitive content. Refresh the index when official procedures change rather than retraining the model for every update.
Expose the system through a small FastAPI service or an approved cloud endpoint. Add rate limits, monitoring, rollback support, and a feedback route to the relevant department. Voice interfaces can improve access for citizens who prefer phone or regional-language interaction; compare the design with guidance on voice agents for Indian businesses, while remembering that municipal answers require stronger source and escalation controls.
If your team needs a fast proof of concept, use a focused pilot—one city, three services, two languages, and a fixed evaluation set. Rapid AI prototyping services for startups can help structure that pilot, but ownership of source verification must remain with the project team.
Practical launch checklist
- Verify every answer against an official source.
- Separate training, validation, and untouched test questions.
- Add jurisdiction, language, date, and answer-type metadata.
- Benchmark retrieval before fine-tuning.
- Redact personal and application data.
- Measure abstention, citation accuracy, and human-rated usefulness.
- Version datasets, prompts, adapters, and evaluation reports.
- Provide a clear human escalation path.
The strongest municipal assistant is not the one with the largest model. It is the one that gives a resident the correct next step, shows where that information came from, and knows when not to answer.