Why GST evaluation needs more than one accuracy score
A Marathi GST assistant can produce fluent, confident Marathi and still give an incorrect tax rate, apply the wrong registration threshold, or miss a critical exception. For that reason, how to measure Marathi language model accuracy on GST queries is not a question of choosing BLEU or a single pass rate. It requires a task-specific benchmark that tests language understanding, tax reasoning, source use, and safe communication.
GST information also changes through notifications, circulars, advisories, and court decisions. A useful evaluation must therefore distinguish between a model’s ability to answer a stable concept—such as the difference between CGST and SGST—and its ability to retrieve the current rule for a date, state, product, or taxpayer category.
Define the evaluation scope first
Write an evaluation specification before collecting examples. Include:
- User profile: individual taxpayer, small retailer, accountant, e-commerce seller, or GST practitioner.
- Query types: registration, invoicing, returns, input tax credit, e-way bills, refunds, composition scheme, place of supply, and notices.
- Language forms: formal Marathi, colloquial Marathi, code-mixed Marathi-English, Devanagari spelling variants, numerals, abbreviations, and misspellings.
- Answer standard: direct answer, explanation, calculation, source citation, clarifying question, or refusal to guess.
- Time and jurisdiction: applicable financial year, state, transaction type, and whether the question concerns current law or a historical period.
This scope prevents a benchmark from becoming a collection of easy FAQ prompts. For low-resource language work, the guidance in Low-Resource Indic Natural Language Processing: A Builder’s Guide is useful when planning coverage, annotation, and quality control for Marathi data.
Build a representative Marathi GST test set
Start with real user intents, not translated English questions alone. Collect anonymised support tickets, practitioner-created examples, public FAQs, and synthetic variations reviewed by Marathi speakers with GST knowledge. Do not copy personal identifiers, GSTINs, invoices, phone numbers, or confidential business data into the dataset.
Create multiple variants for each intent:
- Native Marathi phrasing and Marathi-English code mixing.
- Short, incomplete questions such as “ITC कधी मिळेल?”
- Spelling and transliteration variations, including Roman Marathi.
- Questions containing amounts, dates, product descriptions, HSN or SAC codes, and state names.
- Adversarial prompts that omit a decisive fact or combine two tax issues.
- Follow-up turns where the user corrects or adds information.
Split data by scenario and source, not only by random rows. Near-duplicate questions in training and testing can make results look stronger than they are. Keep a private, periodically refreshed test set for release decisions, and stratify the results by intent, language style, complexity, and current-versus-historical rules.
Annotate what a correct answer must contain
A reference answer should not be a single “gold” paragraph. GST questions often have several acceptable Marathi explanations. Instead, annotate each item with structured fields:
- Intent and required entities.
- Facts the answer must state.
- Conditions, exclusions, and assumptions.
- Expected calculation or decision.
- Authoritative sources and relevant date.
- Whether the model should ask a clarifying question.
- Severity if the answer is wrong.
Use at least two independent annotators for difficult items and adjudicate disagreements with a GST expert. Record whether a response is fully correct, substantially correct, partially correct, misleading, or unsafe. A fluent but materially wrong answer should not receive a passing score.
Use metrics that match the task
1. Intent and entity accuracy
For routing or structured extraction, measure precision, recall, and F1 for intents and entities such as state, tax period, supply type, rate, amount, and return form. Report macro-F1 so rare but important categories are not hidden by common FAQ traffic. A confusion matrix can reveal whether the model routinely confuses registration with return filing or input tax credit with refund claims.
2. Answer correctness and completeness
For question answering, use a rubric rather than relying on string overlap. Score independently for:
- Correct conclusion.
- Correct computation and units.
- Coverage of necessary conditions.
- Marathi clarity and terminology.
- Appropriate uncertainty or clarification.
- Citation or traceability to an approved source.
Exact-match and token-level F1 remain useful for fields such as rates, dates, form numbers, and yes/no classifications. BLEU and ROUGE can describe wording overlap, but they are weak indicators of GST correctness because valid answers may use different Marathi phrasing.
3. Factuality and groundedness
Require the model to cite the relevant official source or retrieved document where the workflow supports retrieval. Evaluate whether the citation actually supports the claim, not merely whether a link is present. Track unsupported claims, invented notifications, stale rates, and citations to irrelevant sections. For high-risk deployments, use a claim-level audit: break an answer into atomic claims and mark each as supported, contradicted, unverifiable, or unnecessary.
4. Safety and clarification
Create cases where the correct behaviour is to ask for the financial year, state, supply type, or taxpayer status. Measure clarification precision: does the model ask only when a missing fact changes the answer? Also track refusal quality. The assistant should explain its limitation and direct the user to an authorised professional or official portal when the question requires case-specific advice.
Add human evaluation in Marathi
Automated metrics cannot reliably judge Marathi grammar, code mixing, politeness, or whether a tax explanation is understandable to a small-business owner. Use a blinded review panel containing native Marathi speakers and GST practitioners. Give reviewers the user query, model response, retrieved evidence, and rubric—but hide model identity.
Ask reviewers to rate correctness, completeness, clarity, source support, and actionability on a defined scale. Measure inter-rater agreement and investigate disagreements rather than averaging them away. Include task-based testing: can a participant identify the correct next action, calculate the amount, or recognise that more information is needed after reading the answer?
Evaluate retrieval and calculations separately
If the assistant uses retrieval-augmented generation, separate three failure points: retrieval recall, evidence quality, and answer generation. Test whether the system retrieves the right notification, handles Marathi queries against English source documents, and respects effective dates. Keep an approved, versioned document store with publication date, effective date, jurisdiction, and supersession status.
Run calculation tests independently with deterministic checks. For tax amounts, validate arithmetic, rounding, taxable value, rate application, and split between CGST and SGST or IGST. A language model should not be the only calculator in a production GST workflow.
Report results by risk, not only overall average
Publish a dashboard with overall scores and slices for each important segment. At minimum, report:
- Exact and tolerant accuracy for structured fields.
- Macro and weighted F1 by intent.
- Fully correct answer rate.
- Critical-error rate and unsupported-claim rate.
- Clarification accuracy.
- Citation support rate.
- Marathi language quality and user task success.
- Latency, cost, and failure rate in production conditions.
Define release gates in advance. For example, a model may require near-zero critical tax errors, a minimum grounded-answer rate, and no regression on Marathi code-mixed queries. A higher average score should not justify deployment if severe errors increase.
Monitor after launch and refresh the benchmark
GST updates, new user vocabulary, and seasonal filing cycles will change the error distribution. Sample anonymised conversations, route suspected failures to reviewers, and maintain a regression suite for every corrected issue. Re-run evaluations after model, prompt, retrieval, document, or translation changes.
As of 2026, a practical operating model is to version the benchmark, annotation guidelines, source corpus, and evaluation code together. Track drift by intent and language form, not just monthly satisfaction. When a new GST notification arrives, add date-sensitive tests immediately and mark older answers according to the rule that applied at that time.
For deployment on constrained devices or local infrastructure, latency and memory can affect answer quality through shorter context windows or smaller models. The principles in AI Model Optimization for Mobile Devices: 2026 Deployment Guide help teams evaluate those trade-offs without losing critical evidence.
A practical evaluation checklist
Before approving a Marathi GST model, confirm that you have:
- A realistic, contamination-resistant Marathi and code-mixed test set.
- Expert-reviewed labels, acceptable-answer criteria, and severity levels.
- Separate scores for intent, entities, calculations, factuality, language, and safety.
- Date-aware retrieval and citation checks.
- Human evaluation by Marathi speakers and GST specialists.
- Critical-error release gates and a documented rollback process.
- Continuous monitoring, drift analysis, and a regression suite.
A credible benchmark makes errors visible instead of hiding them behind fluent output. For teams building Marathi systems with limited data, combining structured evaluation with the methods discussed in Low-Resource Language Datasets for AI Training in India is a stronger path to dependable GST assistance than translating an English benchmark and reporting one aggregate score.