Indian-language summarization is not simply an English summarization model with a different tokenizer. Scripts, morphology, code-mixing, honorifics, transliteration, and uneven training data all affect quality. Quantization adds another engineering constraint: the model must remain useful after its weights and activations move from floating-point precision to INT8, INT4, or another lower-precision format.
This guide explains how to build a production-oriented quantized summarization model for Indic languages, with an emphasis on mobile, browser, edge, and cost-sensitive deployments in India.
1. Define the deployment target first
Start with the device and latency budget, not the quantization library. Record:
- Target languages and scripts, including whether Romanised input is supported.
- Maximum input length and expected summary length.
- Latency target on the actual CPU, GPU, or NPU.
- Memory and package-size limits.
- Whether inference is offline, on-premises, or served through an API.
- Privacy requirements for documents such as legal, healthcare, education, or government records.
For a low-cost Android deployment, INT8 encoder-decoder inference may be a practical first target. For server-side inference, weight-only INT4 can reduce memory while preserving flexibility. If the product serves several languages, benchmark a multilingual checkpoint against language-specific models before committing to one architecture. Teams building for India’s next wave of users should also account for intermittent connectivity and affordable devices, as discussed in this guide to building AI apps for the next billion users in India.
2. Build a trustworthy Indic summarization dataset
The dataset usually determines quality more than the final quantization method. Collect document-summary pairs from sources where redistribution and model training are permitted. Useful categories include news, public notices, court or policy documents, educational material, customer support records, and enterprise documents supplied with consent.
Create a data card covering:
- Language, script, region, domain, date, and source.
- Whether the text is native-script, transliterated, or code-mixed.
- Summary type: headline, extractive abstract, executive brief, or answer-style summary.
- Average document and summary length.
- Licensing, consent, and removal procedures.
Deduplicate at both document and near-duplicate levels. News syndication can otherwise leak the same article into training and evaluation. Split data by source and time where possible, rather than making random splits that inflate results. Preserve difficult examples—long sentences, named entities, numerals, tables converted to text, and mixed Hindi-English or Tamil-English input—in a dedicated challenge set.
For low-resource languages, combine supervised pairs with monolingual pretraining or continued pretraining. The low-resource Indic natural language processing guide provides a useful framework for deciding when additional language data is more valuable than a larger model.
3. Prepare text without destroying meaning
Indic preprocessing should be conservative. Normalise Unicode consistently, but do not remove punctuation or diacritics blindly. Retain sentence boundaries, paragraph structure, list markers, and numerals when they carry meaning. Handle zero-width characters, duplicated whitespace, OCR errors, and inconsistent danda punctuation explicitly.
Use a tokenizer that covers the target scripts efficiently. Measure unknown-token rates, average tokens per character, and sequence-length inflation for every language. A tokenizer that performs well for Devanagari may be inefficient for Malayalam or Bengali. If users type in Roman script, decide whether to transliterate before inference, train on Romanised examples, or support both paths with an input normalisation layer.
A practical preprocessing pipeline should include:
- Language and script identification.
- Unicode and whitespace normalisation.
- Sentence segmentation suited to the target language.
- PII filtering or redaction where required.
- Length filtering without discarding minority-language examples.
- Entity and number preservation checks.
4. Select and fine-tune the base model
Use an encoder-decoder model designed for generation, such as a multilingual or Indic-focused sequence-to-sequence checkpoint. BERT-style encoder-only models are useful for extractive ranking but are not a direct substitute for an abstractive summarizer.
Fine-tune in full or mixed precision before quantizing. Start with a stable baseline and log:
- Training and validation loss.
- ROUGE scores by language and domain.
- Faithfulness and entity-preservation errors.
- Output length and repetition.
- Latency and memory on the target hardware.
Use language-balanced sampling so high-resource languages do not dominate training. Curriculum training can begin with shorter documents and move to longer inputs. For limited compute, parameter-efficient fine-tuning—such as LoRA—can reduce training cost, but validate the merged checkpoint before conversion. Keep a full-precision reference model for comparison throughout the process.
5. Choose a quantization route
Quantization maps floating-point values to lower-precision representations. The best method depends on the runtime and hardware.
Dynamic post-training quantization
Dynamic quantization commonly quantizes weights while calculating activation ranges at runtime. It is simple and often a good baseline for CPU inference, especially for linear layers. It requires no calibration dataset, but speed gains vary by operator and hardware.
Static post-training quantization
Static quantization uses representative calibration data to estimate activation ranges before deployment. Build a calibration set that reflects real Indian-language traffic: all target scripts, typical lengths, code-mixed inputs, numerals, and domain terminology. Calibration data need not contain reference summaries, but it must resemble production input.
Quantization-aware training
QAT inserts simulated quantization during fine-tuning so the model learns to tolerate reduced precision. Use it when post-training quantization causes a material quality drop or when the target accelerator requires strict INT8 operators. QAT increases training complexity, so first identify which layers and languages are actually affected.
For generative models, do not assume every component should use the same precision. Embeddings, layer normalisation, attention projections, and the output head may have different sensitivity. Mixed-precision deployment often offers a better quality-speed trade-off than forcing the entire model to INT8 or INT4.
6. Evaluate quality, speed, and safety together
ROUGE-1, ROUGE-2, and ROUGE-L are useful for regression testing, but they should not be the sole release gate. Reference summaries can use different wording, particularly across languages. Add human review by native speakers and measure:
- Factual consistency and unsupported claims.
- Coverage of key points.
- Hallucinated names, dates, amounts, and locations.
- Grammaticality and naturalness.
- Script correctness and code-mixing behaviour.
- Bias or omission across regions, communities, and dialects.
Compare full precision, dynamic quantization, static quantization, and QAT on the same challenge set. Report quality by language rather than only a macro-average. Also record peak RAM, model size, cold-start time, tokens per second, energy use, and tail latency. A model that loses two points on ROUGE but halves memory may be acceptable for offline use; the same trade-off may be unacceptable for legal or medical summaries.
7. Deploy with reproducible, observable inference
Export to a runtime supported by the target device, such as ONNX Runtime, TensorFlow Lite, ExecuTorch, or another accelerator-specific stack. Confirm that all critical operators have quantized kernels. Silent fallback to floating-point execution can erase expected speed and memory gains.
Package the tokenizer and preprocessing rules with the model version. Test long inputs, empty or malformed text, mixed scripts, unsupported characters, and adversarial prompts. Add safeguards against copying sensitive source text into logs. For server deployments, monitor language distribution, latency, failure rates, output length, and quality samples after every model update.
If summaries feed a voice interface, test the handoff from text generation to speech and consider the latency requirements covered in this guide to real-time voice agents with fast barge-in. For privacy-sensitive organisations, an on-device or private deployment may be preferable; the design principles in building a private AI chatbot for lawyers are relevant to access controls, audit logs, and data retention.
8. A practical release checklist
Before shipping, verify that:
- The dataset licences and consent records are documented.
- Evaluation is split by language, script, domain, and input length.
- Quantized outputs are compared with a full-precision baseline.
- Named entities, dates, numbers, and negation are explicitly tested.
- The runtime uses the intended quantized kernels.
- Offline, low-memory, and poor-connectivity scenarios are tested.
- Users can report incorrect or harmful summaries.
- Model, tokenizer, calibration data, and preprocessing versions are pinned.
Quantization is an optimisation layer, not a replacement for language-specific data and evaluation. A smaller model can make Indic summarization affordable on Indian devices and private infrastructure, but only if the engineering process protects factuality and language coverage at every step.