Court-document summarisation is not a generic text-generation problem. Indian judgments can run to hundreds of pages, combine English with legal terminology and Indic names, cite statutes and precedents, and contain procedural details that cannot be casually omitted. A useful system must therefore be faithful, traceable, privacy-conscious, and affordable to run.
Quantization can help. By reducing model weights and activations from formats such as FP16 or FP32 to INT8, INT4, or related lower-precision formats, it can reduce memory use and inference cost. But quantization is an optimisation step—not a substitute for a well-designed legal NLP pipeline. This guide explains how to build one in 2026.
Define the task before choosing a model
Start with a narrow, testable output. “Summarise the judgment” is too vague for production. Choose one or more formats such as:
- Case brief: facts, issues, arguments, decision, and reasoning.
- Procedural summary: court, dates, applications, orders, and next steps.
- Issue-wise summary: each legal issue paired with the court’s finding.
- Extractive digest: important passages with page or paragraph references.
- Plain-language explanation: a reader-friendly version clearly labelled as non-legal advice.
Decide whether the system is extractive, abstractive, or hybrid. A hybrid design is often safer: retrieve relevant passages, generate a concise summary, and attach citations to the source paragraphs. For lawyer-facing workflows, this approach is generally more useful than a fluent paragraph with no evidence trail. Teams building a secure legal assistant can also review the architecture in How to Build a Private AI Chatbot for Lawyers.
Set operational requirements early: maximum document length, target latency, languages, deployment location, acceptable hallucination rate, and whether documents may leave an organisation’s network.
Build a legally usable dataset
Public judgments are not automatically suitable training data. Establish provenance, usage rights, redaction procedures, and a policy for personally identifiable information. For Indian workloads, expect variation in formatting, OCR quality, court abbreviations, citation styles, and language. If the target includes Hindi or other regional-language material, the recommendations in Low-Resource Indic Natural Language Processing: A Builder’s Guide are directly relevant.
Create high-quality examples rather than relying only on large volumes of weak summaries. Each record should ideally contain:
- The original document and a clean, versioned text representation.
- Court, date, jurisdiction, case type, and language metadata.
- Paragraph or page boundaries preserved from the source.
- A human-written reference summary with an agreed structure.
- Citation spans showing which passages support each claim.
- A split by case, not by paragraph, to prevent leakage between training and testing.
Keep a difficult holdout set containing long judgments, poor OCR, multiple parties, dissenting opinions, dense citations, and mixed-language text. This set should not influence model selection.
Prepare long judgments without destroying context
PDF extraction is a frequent source of silent errors. Test whether headings, paragraph numbers, footnotes, tables, annexures, and page order survive conversion. Run OCR only where needed, retain the original page image, and flag low-confidence text for review.
Do not simply truncate every document to the model’s context window. Use a hierarchical pipeline:
1. Segment the judgment by headings, paragraphs, or page ranges.
2. Remove repeated headers, footers, and irrelevant boilerplate while retaining legal citations.
3. Generate chunk-level notes or extractive selections.
4. Combine those intermediate representations into a structured case summary.
5. Run a final consistency and citation check against the source.
Chunk overlap should be modest and evaluated empirically. Excessive overlap increases cost and can duplicate facts. Store stable chunk IDs so every generated statement can be traced back to source text.
Select a base model and quantization route
For a first prototype, choose a model that supports the required language coverage, context length, licence, and inference stack. Encoder models can work well for ranking and extractive selection; decoder or encoder-decoder models are suited to generation. Do not assume that a popular general model understands Indian legal conventions without domain evaluation.
There are three practical quantization strategies:
- Post-training quantization (PTQ): fastest to try and appropriate when the full-precision model already performs well.
- Quantization-aware training (QAT): simulates lower precision during training and can preserve quality when PTQ causes a noticeable drop.
- Parameter-efficient fine-tuning plus quantization: methods such as LoRA or QLoRA reduce training memory and are useful for domain adaptation.
For GPU serving, INT8 or INT4 weight-only formats may provide a good memory-quality trade-off. For CPU or edge deployment, verify that the chosen runtime has genuinely optimised kernels; a smaller file does not automatically mean faster inference. Benchmark the complete pipeline, not just the model.
Fine-tune for faithful legal summaries
Use structured targets rather than unconstrained prose. A template might require: case background, legal issues, submissions, findings, final order, and unresolved points. Train the model to state when information is absent rather than infer it.
During fine-tuning:
- Normalise citation and date formats without changing meaning.
- Include hard negatives where similar parties or precedents could confuse the model.
- Balance short and very long judgments.
- Preserve section labels and paragraph references.
- Record model, tokenizer, data, and prompt versions for reproducibility.
A retrieval-augmented design can reduce the burden on the generator. Retrieve relevant passages for each section, then constrain generation to those passages. This is especially valuable for orders, statutory provisions, and numerical details.
Quantize, then test the right things
Create a baseline using the unquantized model. Compare it with INT8 and INT4 variants under the same prompts, decoding settings, context strategy, and hardware. Measure:
- Peak RAM or VRAM and model size.
- Tokens per second, end-to-end latency, and throughput.
- Cost per document.
- ROUGE or similar overlap metrics, used only as screening signals.
- Faithfulness: whether each material claim is supported by the judgment.
- Completeness: whether the decision, issues, and key reasoning are present.
- Citation accuracy: whether references point to the correct page or paragraph.
- Legal-risk errors: invented holdings, reversed outcomes, omitted conditions, and conflated parties.
Human review remains essential. Use at least two reviewers for a representative sample and adjudicate disagreements. Evaluate English and each supported Indic language separately. A model that achieves a strong average score while failing on Hindi names, dates, or dissenting opinions is not production-ready.
Deploy with privacy and review controls
For confidential filings or privileged material, prefer a private VPC or on-premises deployment, encryption in transit and at rest, strict access controls, audit logs, and configurable retention. Avoid sending documents to an external API by default. Redact unnecessary personal data before processing and document who can view source documents and generated outputs.
Expose the system through an API that returns the summary, confidence or review flags, source citations, model version, and processing timestamp. Build a human-in-the-loop interface where a lawyer can inspect evidence, edit the output, and report an error. The output should clearly state that it is an AI-generated summary and not legal advice.
For organisations serving users across India, latency, bandwidth, and device constraints matter. The broader deployment considerations in Building AI Apps for the Next Billion Users in India can help teams design for varied infrastructure rather than assuming a high-end workstation.
A practical build sequence
A sensible implementation plan is:
1. Build an extractive baseline with paragraph citations.
2. Add hierarchical chunking and structured summaries.
3. Fine-tune a permitted base model on reviewed examples.
4. Establish a full-precision quality and safety baseline.
5. Compare PTQ formats on a fixed evaluation set.
6. Try QAT or QLoRA only if the quality loss justifies extra complexity.
7. Pilot with legal reviewers before wider deployment.
8. Monitor drift as courts, document formats, and case types change.
Treat quantization as part of system engineering, not as the headline feature. The best court-summarisation system is the one that reduces reading time while making every important claim easy to verify.