GST invoice assistants are not ordinary chatbots. They must extract facts from inconsistent documents, distinguish taxable value from tax components, handle Indian numbering and date formats, and provide answers that finance teams can verify. Quantization can reduce latency and memory use, but it does not fix poor OCR, incomplete training data, or unsafe answers.
This guide explains how to build a quantized model for GST invoice queries as a production system: document ingestion, structured extraction, retrieval, answer generation, quantization, evaluation, and monitoring. The approach works for a small on-premise deployment as well as an API-backed application serving multiple businesses.
Define the task before choosing a model
Start by writing down the questions the system must answer. Typical examples include:
- What is the invoice number, date, or supplier GSTIN?
- What are the taxable value, CGST, SGST, IGST, cess, and grand total?
- Which HSN or SAC code appears on a line item?
- Is the place of supply consistent with the tax type?
- Which invoices are missing mandatory fields or contain arithmetic mismatches?
- Show all invoices from a supplier within a date or value range.
Separate document extraction from reasoning and search. A reliable architecture usually converts each invoice into a validated JSON record, stores the original image or PDF, and uses a language model only to interpret the user’s question and explain results. For multilingual or code-mixed queries, the principles in this guide to low-resource Indic natural language processing are directly relevant.
Do not allow the model to invent invoice values. Every numeric answer should cite the source page, table row, or extracted field and should be marked uncertain when OCR confidence is low.
Build a representative Indian invoice dataset
Collect documents from the environments where the product will operate: scanned PDFs, mobile photographs, thermal prints, digitally generated invoices, credit notes, debit notes, e-invoice JSON, and invoices with tables spanning multiple pages. Include suppliers from different states and sectors, along with English, Hindi, regional-language, and mixed-language text where applicable.
Create labels for:
- Header fields: invoice number, invoice date, supplier and buyer GSTINs, legal names, and addresses.
- Tax fields: taxable value, tax rates, CGST, SGST, IGST, cess, round-off, and total.
- Line items: description, HSN/SAC, quantity, unit, rate, discount, and line-level taxes.
- Relationships: place of supply, reverse-charge indicator, export or SEZ status, and purchase-order reference.
- Query answers: the correct response, supporting fields, and evidence location.
Split data by supplier and template, not randomly by page. Otherwise, nearly identical invoices may appear in both training and test sets, producing misleadingly high accuracy. Remove or mask personal information when it is not needed, restrict access to raw documents, and record consent and retention rules. For a broader product serving Indian users, also apply the deployment principles in building AI apps for the next billion users in India, especially around device constraints, language, and connectivity.
Use a pipeline instead of a single large model
A practical pipeline has five stages:
1. Ingestion: identify file type, rotate pages, detect duplicates, and virus-scan uploads.
2. OCR and layout analysis: extract text, tables, coordinates, and confidence scores.
3. Schema mapping: normalize dates, GSTINs, currency values, tax rates, units, and state codes.
4. Validation: recompute totals, check GSTIN format, compare tax type with place of supply, and flag contradictions.
5. Query answering: translate the question into filters or structured operations, retrieve evidence, and generate a concise response.
Store monetary values as decimal integers in paise where possible. Preserve both the original string and normalized value—for example, ₹1,25,000.00 and 12500000 paise—so users can audit transformations. Keep OCR confidence and page coordinates alongside every field.
For straightforward questions, use deterministic SQL or rules rather than a generative response. Reserve the language model for intent classification, schema selection, ambiguous phrasing, and evidence-based explanation. This reduces hallucination risk and makes costs easier to control.
Select a baseline and quantization strategy
Establish a full-precision baseline before optimization. Measure field extraction F1, exact match for invoice identifiers, numeric tolerance for amounts, query answer accuracy, evidence citation accuracy, latency, memory, and cost per 1,000 queries.
For a compact text model, consider a small instruction-tuned model capable of handling the languages and context length you need. For document understanding, evaluate a specialized OCR or layout model separately. Quantizing every component in the same way may damage table recognition while leaving the retrieval layer unaffected.
The main options are:
- Dynamic post-training quantization: easy to apply and useful for CPU inference, especially for linear layers.
- Static post-training quantization: calibrates activations on representative invoices and can improve runtime performance on supported hardware.
- Quantization-aware training: simulates lower precision during training and is preferable when post-training quantization causes a significant accuracy drop.
- Weight-only 4-bit or 8-bit quantization: reduces model memory, but actual speed gains depend on kernels, hardware, and batch size.
Use representative calibration data: clean invoices alone will not expose failures caused by faint scans, long addresses, decimal amounts, or dense tables. Compare FP32, FP16 or BF16, INT8, and lower-bit versions on the same held-out supplier templates. Treat quantization as an engineering trade-off, not an automatic quality improvement.
Train and evaluate for reliability
Fine-tune only after the pipeline and baseline are stable. Use supervised examples that include the user question, structured context, expected answer, and evidence references. Add difficult cases such as missing GSTINs, conflicting totals, duplicate invoice numbers, credit notes, and questions that cannot be answered from the supplied document.
Evaluate by slice rather than reporting one aggregate score:
- Invoice format and scan quality.
- Supplier, state, language, and script.
- Query type: lookup, aggregation, comparison, validation, or explanation.
- Numeric magnitude, decimal precision, and tax type.
- OCR confidence and number of pages.
- Full-precision versus quantized model.
For numeric answers, require exact or tolerance-based matching. For tax calculations, independently recompute the result in code. Test prompt injection inside uploaded documents: invoice text should never override system instructions or cause data from another tenant to be disclosed. Add a refusal or escalation path when evidence is missing.
Deploy with controls for finance workflows
Package the model in the runtime used by your target hardware, such as ONNX Runtime, TensorFlow Lite, or a compatible PyTorch serving stack. Benchmark cold starts, warm latency, concurrent requests, peak memory, CPU utilization, and batch behavior. A quantized model that is smaller but slower on your actual processor is not an optimization.
Use tenant isolation, encryption, role-based access, audit logs, rate limits, and configurable retention. Keep model outputs separate from accounting records until a user approves them. For workflows involving multiple services—OCR, retrieval, validation, and model serving—design explicit retries, idempotency, and failure handling; the patterns in building distributed systems with AI agents are useful even when the system is not agentic.
Monitor production performance using sampled, privacy-preserving logs:
- OCR confidence and extraction correction rates.
- Abstention and escalation rates.
- Quantized-versus-baseline disagreement.
- Latency, memory, and cost.
- New invoice templates and unseen languages.
- User corrections by field and query type.
Route low-confidence or high-value cases to human review. Retrain from verified corrections, but version datasets, prompts, schemas, quantization settings, and evaluation results together.
A practical launch checklist
Before going live, confirm that:
- The test set contains unseen suppliers and document templates.
- Every answer can link back to invoice evidence.
- Amounts are computed with decimal-safe logic.
- Quantized accuracy meets field-level and query-level thresholds.
- The system abstains when OCR or evidence is inadequate.
- Sensitive documents are isolated by tenant and retained only as required.
- Rollback to the baseline model is tested.
- Human reviewers can correct extracted fields without editing raw evidence.
A strong GST invoice assistant is therefore a document intelligence and validation system with a quantized model, not merely a smaller chatbot. Build the evidence and accounting logic first, then quantize the components that benefit from it. This sequence delivers faster inference without sacrificing the traceability Indian finance teams need.