What counts as an LLM hallucination?
An LLM hallucination is an answer that is false, unsupported by the available evidence, or presented with unjustified confidence. It may be a fabricated citation, an incorrect translation, an invented policy clause, or a plausible-looking answer to a question the model cannot answer.
For Indian deployments, the risk is amplified by multilingual inputs, code-switching, uneven digitisation of public records, regional names, and domain-specific terminology. A customer-support bot that confuses Hindi and Hinglish is inconvenient; a system that invents a medical instruction, loan condition, or government eligibility rule is a safety and compliance problem.
The practical goal is not to “eliminate” hallucinations. It is to make unsupported answers less likely, easier to detect, and safer when they occur.
Clarify what “classical foundation models” means
Classical foundation models are useful here as stable, general-purpose base models that are adapted with structured data, retrieval, classifiers, or smaller task-specific components rather than relying only on unconstrained free-form generation. This can include encoder models for classification and retrieval, sequence-to-sequence models for translation, and decoder models used with strict grounding and output controls.
The best architecture depends on the task:
- Use an encoder or reranker to classify intent, detect language, retrieve evidence, or flag risky content.
- Use a retrieval system to find authoritative passages before generation.
- Use a generative model only where synthesis or explanation is genuinely needed.
- Use rules, schemas, and deterministic code for calculations, eligibility checks, dates, and regulatory workflows.
For Indian-language applications, compare language coverage rather than assuming that a larger model is automatically better. The practical differences between Hindi, Marathi, Telugu, Sanskrit, and code-switched speech or text can be substantial. Guides to benchmarking NLP models for Telugu and Sanskrit and fine-tuning AI models for Marathi dialects can help shape a realistic evaluation plan.
Build a grounded answer pipeline
The most reliable pattern is retrieval-augmented generation with explicit evidence handling:
1. Classify the request. Detect language, intent, domain, sensitivity, and whether the user is asking for a fact, recommendation, calculation, or creative output.
2. Retrieve authoritative sources. Search approved documents, internal databases, legislation, product manuals, or verified FAQs. Apply metadata filters such as jurisdiction, department, language, and effective date.
3. Rerank the evidence. A cross-encoder or other classical relevance model can remove superficially similar but incorrect passages.
4. Generate only from retrieved evidence. Instruct the model to quote, cite, or explicitly state when the evidence is insufficient.
5. Verify the draft. Check claims against source passages, validate numbers and dates programmatically, and run a contradiction or entailment classifier where appropriate.
6. Choose an answer mode. Return a supported answer, ask a clarifying question, or abstain and route the case to a human.
Do not treat retrieval as proof of correctness. A system can retrieve the wrong document, use an outdated circular, or combine passages from incompatible jurisdictions. Store document version, source owner, publication date, and review status alongside every chunk.
Use constraints instead of hoping prompts will work
Prompting helps, but it should not be the primary control. Add constraints at the application layer:
- Require citations for factual claims and reject responses without valid source IDs.
- Restrict outputs to a JSON schema for structured workflows.
- Set maximum answer length and prohibit unsupported names, figures, or legal conclusions.
- Use temperature and sampling settings appropriate to the task; deterministic settings are preferable for extraction and classification.
- Configure an explicit abstention response, such as “I could not verify this from the approved sources.”
- Separate system instructions, retrieved content, and user text to reduce prompt-injection risk.
- Keep business rules outside the model. Let code calculate taxes, loan repayments, scores, dates, and eligibility.
If infrastructure or data residency requires local serving, review practical trade-offs in how to deploy large language models locally. A smaller model with a strong retrieval and verification layer often delivers more dependable results than a larger model with weak controls.
Improve the data before changing the model
Hallucination reduction starts with source quality. Create a document pipeline that:
- Removes duplicate, obsolete, and contradictory documents.
- Preserves tables, headings, footnotes, and page references during parsing.
- Separates authoritative text from commentary and user-generated material.
- Normalises Indian names, addresses, dates, units, and government department terminology.
- Adds language and script metadata, including Devanagari, Latin-script Hinglish, and transliterated regional languages.
- Builds question-and-evidence examples from real support tickets and operational queries.
For smaller Hindi deployments, compare open models and test them on your own prompts rather than selecting by parameter count alone; the guide to open-source small language models for Hindi is a useful starting point. Translation systems need separate checks for names, honorifics, legal terms, and code-switching; Sanskrit and other low-resource language projects may benefit from fine-tuning large language models for Sanskrit translation.
Evaluate hallucinations with a risk-based test set
A generic accuracy score hides the failures that matter. Build a test set covering:
- Questions answerable from one source.
- Questions requiring multiple documents.
- Unanswerable and deliberately misleading questions.
- Conflicting or outdated documents.
- Spelling variation, transliteration, mixed scripts, and code-switching.
- Tables, numbers, dates, names, and long-context inputs.
- Adversarial prompts that attempt to override retrieval instructions.
Track groundedness, citation validity, retrieval recall, answer completeness, abstention precision, contradiction rate, latency, and cost. Report results by language, domain, user type, and severity—not only as one aggregate score. Have domain reviewers assess high-impact samples, especially in healthcare, finance, education, employment, and public services.
Run regression tests whenever you change the model, prompt, embedding model, chunking strategy, or source corpus. Keep a human-review queue for low-confidence and high-risk cases, and log the retrieved passages, model version, prompt template, output, and final disposition for each production incident.
Design for Indian production conditions
Plan for intermittent connectivity, mobile-first interfaces, variable-quality scans, and users who switch languages mid-conversation. OCR errors can create hallucinations before the LLM is even called, so validate extracted text and retain the original document for review. Protect personal data through minimisation, access controls, encryption, retention limits, and redaction before indexing.
Give users a visible way to report an incorrect answer. Publish the source and date where possible, distinguish information from advice, and provide escalation to a trained operator. For sensitive workflows, require confirmation before an action is taken and never allow an unverified generation to update a citizen, patient, customer, or financial record automatically.
A practical implementation sequence
Start with one narrow, measurable workflow rather than a general chatbot:
- Define prohibited errors and acceptable abstention behaviour.
- Assemble a verified, versioned source set.
- Implement retrieval, reranking, citations, and schema validation.
- Add deterministic tools for calculations and database lookups.
- Create multilingual and adversarial evaluation sets.
- Pilot with human review and monitor incidents.
- Expand coverage only after quality holds across languages and user groups.
Classical foundation models are valuable because they make parts of the system easier to inspect and specialise. They are not a substitute for source governance, evaluation, or responsible product design. In 2026, the strongest Indian deployments will treat the LLM as one component in a verified pipeline—not as the final authority.
FAQs
Can classical foundation models completely prevent hallucinations?
No. They can reduce unsupported generation when paired with retrieval, classifiers, deterministic tools, and abstention. Every production system still needs monitoring and human escalation.
Is fine-tuning better than retrieval?
They solve different problems. Use retrieval for changing facts, policies, and proprietary documents. Use fine-tuning for stable behaviour, formatting, language adaptation, or specialised task performance. Many systems need both.
How should a team measure success?
Set thresholds by risk. Measure citation correctness, groundedness, abstention quality, contradiction rate, and severe-error frequency by language and domain. A lower answer rate can be a positive result if the system avoids unsafe guesses.
What should happen when evidence is missing?
The system should say that it cannot verify the answer, identify what information is missing, and offer a safe next step—such as a link to an official source or escalation to a human reviewer.