Machine translation usually fails on context for a simple reason: the system receives too little information, or it has no reliable way to use the information it receives. A sentence-level model may translate a word correctly in isolation but choose the wrong meaning when the surrounding paragraph identifies a product, person, organisation, or technical process.
For teams building multilingual products in India, these failures are costly. They can alter legal meaning, misstate medical instructions, break customer-support tone, or produce incorrect gender and honorific agreement in languages such as Hindi, Marathi, Bengali, Tamil, and Kannada. The practical fix is not to choose between NMT and LLMs. It is to build a translation pipeline that supplies context, controls terminology, tests difficult cases, and routes uncertain output for review.
Diagnose the context failure first
Before changing the model, classify the error. Save the source document, the model input, the output, and an approved reference translation so the team can reproduce the issue.
Common categories include:
- Lexical ambiguity: “bank”, “port”, “charge”, or “server” is translated according to the wrong sense.
- Pronoun and coreference errors: a model loses track of who “she”, “they”, or “this” refers to across sentences.
- Gender, number, and honorific errors: the target language requires information that is not explicit in the current sentence.
- Terminology drift: the same product, law, feature, or medical term receives multiple translations.
- Register changes: formal customer communication suddenly becomes casual, or a respectful Indian-language form is replaced with an inappropriate one.
- Document-structure errors: headings, tables, labels, and notes are translated without understanding their relationship to the main text.
Create a small error set with examples from your actual domain. A hundred difficult sentences from a banking, healthcare, education, or government workflow will be more useful than a generic benchmark that does not reflect your users.
Supply document-level context
The first engineering improvement is to translate with context rather than sending isolated strings to an API. Provide the current segment together with selected surrounding content, such as the previous and next sentences, section heading, document title, speaker identity, and relevant metadata.
A practical input structure is:
- Document: title, language, locale, and content type
- Section: heading and short summary
- Context: one to three preceding and following segments
- Current segment: the text to translate
- Constraints: audience, tone, terminology, and formatting rules
Do not blindly send an entire document on every request. Large context windows increase cost and can introduce irrelevant information. Use a retrieval or segmentation layer to select the context that changes meaning. For dialogue, include the speaker and conversation history. For product documentation, include the feature name and glossary entries. For forms, preserve field labels and allowed values.
If you are deploying at scale, treat context handling as part of your scalable machine learning infrastructure, not as a prompt-only feature. Log which context was supplied so failures can be audited and improved.
Control terminology with glossaries and translation memory
A glossary should be a machine-readable source of truth, not a document that translators consult only at the end. Each entry can include the source term, approved target term, part of speech, domain, forbidden alternatives, grammatical gender, and examples.
Use terminology controls for:
- Product and company names
- Government schemes and official titles
- Legal and medical terms
- UI labels and navigation elements
- Acronyms and abbreviations
- Place names and Indian personal names
Translation memory adds approved sentence- or phrase-level examples. Fuzzy matches can help the system reproduce established wording, but they must be reviewed when the surrounding context has changed. Hard constraints are appropriate for regulated terms and UI labels; softer preferences are safer for ordinary prose, where forcing a term can create an ungrammatical sentence.
Run automated checks after generation. Flag missing glossary terms, inconsistent translations, changed numbers, altered dates, broken placeholders, and suspicious variations in names. Keep a record of whether each correction came from a glossary, a human reviewer, or a model revision.
Use LLMs with structured instructions, not vague prompts
LLMs can resolve context well when the task is clearly bounded, but “translate accurately” is not a reliable specification. Ask for the target locale, audience, register, spelling convention, and terminology rules. State what must not be translated, including brand names, code, variables, HTML, and placeholders.
A robust workflow has two stages:
1. Contextual draft: translate the complete section or a contextually coherent batch.
2. Consistency review: compare the output against the glossary and source, then identify unresolved references, terminology drift, numbers, omissions, and tone changes.
Require the review stage to return structured findings rather than silently rewriting everything. This makes it easier to distinguish a genuine correction from an unnecessary stylistic change. For high-risk content, have a human approve the proposed edits.
For Indian-language work, domain adaptation may require more than prompting. Teams experimenting with Sanskrit or other low-resource languages can study approaches to fine-tuning large language models for Sanskrit translation, while keeping evaluation data separate from training data.
Fine-tune only after fixing the data and workflow
Fine-tuning is useful when the same contextual mistakes recur and you have high-quality parallel data. Start by collecting examples that contain the intended sense, antecedent, register, and approved terminology. Include hard negatives: similar sentences where the correct translation changes because the context changes.
Useful methods include:
- Domain fine-tuning on verified parallel documents
- Parameter-efficient methods such as LoRA for specialised adapters
- Synthetic examples for ambiguity, pronoun resolution, and terminology
- Retrieval of approved bilingual examples at inference time
- Separate adapters for domains such as legal, healthcare, finance, and education
Do not fine-tune to compensate for corrupted segmentation, inconsistent references, or an incomplete glossary. Those problems will remain—and may become harder to diagnose—inside the model.
Evaluate context, not just fluency
BLEU alone cannot tell you whether a translation preserved the intended meaning. Build a test suite that measures the errors users actually see:
- Meaning accuracy: correct sense for ambiguous words
- Coreference: consistent people, objects, gender, and number across sentences
- Terminology: adherence to approved terms and forbidden-term rates
- Consistency: repeated names, units, dates, and UI strings
- Adequacy: omissions, additions, and altered claims
- Fluency and register: natural target-language grammar and appropriate formality
- Latency and cost: performance under production context limits
Evaluate by language, domain, and content type. A model that performs well on Hindi news may still fail on Marathi legal notices or Tamil customer support. Add regression tests whenever a reviewer identifies a new failure.
Quality estimation models can prioritise segments for review, but they should not be treated as a replacement for reference-based testing or expert judgment. Human review remains essential for medical, legal, safety, financial, and public-service content.
Address Indian-language realities
Indian languages introduce context requirements that sentence-level systems often miss. Grammatical gender may depend on a person introduced earlier. Honorifics can depend on social relationship and institutional setting. Names and brands may need transliteration in one field and preservation in another. A word may be translated in a general sentence but retained in English when it refers to a product or software feature.
Define locale-specific rules for numerals, dates, currency, punctuation, transliteration, honorifics, and code-switching. Ask native-language reviewers to approve these rules with examples, not just labels such as “formal” or “neutral.” If your team is building an end-to-end multilingual product, test the translation component alongside speech, video, or UI rendering; guidance for real-time AI video translation apps illustrates why timing and formatting can matter as much as text accuracy.
A practical production checklist
Before release, confirm that your pipeline:
- Preserves document structure and metadata
- Supplies relevant context without exceeding practical limits
- Applies glossary and translation-memory rules
- Protects variables, markup, numbers, and names
- Runs automated consistency and terminology checks
- Measures context-specific errors by language and domain
- Escalates low-confidence or high-risk content to reviewers
- Stores inputs, outputs, corrections, and model versions for audits
- Tests every model or prompt change against a regression suite
The most reliable solution is layered: document-aware input, controlled terminology, domain data, targeted model adaptation, automated checks, and human review. That combination lets Indian builders reduce context errors without sacrificing the cost and speed that make machine translation valuable.