GST circulars are useful training material for AI systems that need to retrieve, classify, summarise, or extract information from India’s indirect-tax guidance. They are also difficult data: circulars contain legal qualifications, references to notifications and Acts, tables, dates, exceptions, and language that can change the meaning of a sentence.
Fine-tuning can improve a model’s performance on a defined task, but it does not automatically turn a general-purpose model into a reliable tax adviser. For most GST applications, the strongest architecture combines retrieval-augmented generation (RAG) over verified source documents with targeted fine-tuning for output format, classification, or domain language. Use the workflow below to build a system that is testable and easier to update.
Define the task before collecting data
“Fine-tuning on GST circulars” can mean several different things. Choose one primary objective before downloading documents:
- Document classification: identify the topic, tax period, issuing authority, or affected taxpayer category.
- Information extraction: capture circular number, date, section, commodity, rate, condition, or effective period.
- Question answering: answer questions with citations to the relevant circular and paragraph.
- Summarisation: produce a short, structured explanation for an accountant, finance team, or founder.
- Semantic search: retrieve the most relevant passages from a circular archive.
A model trained only on raw circular text learns domain vocabulary, but may not learn how you want it to respond. For answer generation or extraction, create labelled examples in the required input-output format. If you are still deciding between approaches, review these best practices for fine-tuning LLMs on custom data.
Build a trustworthy GST corpus
Start with primary sources wherever possible: official GST, CBIC, and government portals. Record the source URL, document title, circular number, issue date, language, download date, and whether the file is a correction, clarification, or replacement. Do not silently merge circulars with notifications, orders, press releases, or third-party commentary.
Create a document manifest such as:
{
"doc_id": "cbic_circular_214_2024",
"circular_number": "214/08/2024-GST",
"issued_on": "2024-09-26",
"source_url": "https://example.gov.in/document.pdf",
"status": "active_or_verify",
"language": "en"
}The status field should be verified by a tax professional or through your legal-content process. GST guidance can be superseded, interpreted by later documents, or affected by amendments. As of 2026, treat the model as an information tool, not as an authority on filing positions or litigation strategy.
Extract and clean PDFs without losing meaning
PDF extraction is often the hardest part. Tables, footnotes, headers, page numbers, scanned pages, and multi-column layouts can corrupt the training text. Use OCR for scanned documents and retain page-level metadata so every answer can be traced back to a source.
A basic extraction step with PyMuPDF looks like this:
import fitz
pages = []
with fitz.open("circular.pdf") as doc:
for page_number, page in enumerate(doc, start=1):
pages.append({
"page": page_number,
"text": page.get_text("text")
})Do not remove all punctuation, numbers, or stopwords. In GST material, dates, percentages, section references, notification numbers, and qualifiers such as “subject to” are essential. Clean repeated headers and broken whitespace, but preserve paragraph boundaries and table labels. Keep the original PDF alongside the processed text for auditability.
For multilingual use cases, decide whether to train separate language-specific models, use translated pairs, or rely on a multilingual base model. If your product serves Hindi or other Indian languages, compare the results against models discussed in this guide to fine-tuning Llama for Indian regional languages.
Create training examples, not just a document dump
For supervised fine-tuning, convert documents into examples that represent real user requests. A question-answer record should include the answer, supporting passage, document ID, and page or paragraph reference:
{
"messages": [
{"role": "user", "content": "What clarification does this circular provide on the stated issue?"},
{"role": "assistant", "content": "The circular clarifies ... [Source: circular 214/08/2024-GST, page 3]."}
],
"source_id": "cbic_circular_214_2024",
"page": 3
}Include difficult examples: ambiguous questions, outdated provisions, conflicting documents, missing information, and requests that require the model to say “I cannot verify this from the provided sources.” Avoid training the model to invent citations or give confident conclusions where the circular is silent.
Split data by document, not by random rows. If passages from the same circular appear in both training and test sets, your metrics will be misleading. A practical starting point is 80% training, 10% validation, and 10% test data, followed by a manually reviewed challenge set containing recent and legally complex circulars.
Select the model and fine-tuning method
Choose a model that supports your language, context length, licence, hardware budget, and deployment environment. For classification or extraction, an encoder model such as a multilingual BERT variant may be sufficient. For structured answers or summarisation, use a compact instruction-tuned language model and consider parameter-efficient fine-tuning.
LoRA and QLoRA reduce memory requirements by training adapter weights instead of updating every parameter. They are useful when working with limited GPU access or when maintaining separate adapters for different GST workflows. Before training, pin compatible versions of transformers, datasets, peft, accelerate, and your quantisation library.
A minimal dataset-loading pattern is:
from datasets import load_dataset
dataset = load_dataset("json", data_files={
"train": "train.jsonl",
"validation": "validation.jsonl",
"test": "test.jsonl"
})Use chat templates supplied by the tokenizer for conversational models. Set a maximum sequence length based on actual examples, and measure truncation. Quietly cutting off the paragraph containing an exception can create dangerous outputs.
Evaluate factuality, citations, and usefulness
Accuracy alone is inadequate for GST systems. Evaluate:
- Retrieval recall: does the system find the right circular and passage?
- Citation accuracy: does the cited page or paragraph support the claim?
- Extraction F1: are fields such as dates and rates captured correctly?
- Answer completeness: are conditions, exceptions, and effective dates included?
- Abstention quality: does the model refuse or request clarification when evidence is missing?
- Robustness: does performance hold across document formats, languages, and question styles?
Have tax professionals review a representative sample using a fixed rubric. Compare the fine-tuned model with a RAG baseline; fine-tuning may improve style while leaving factual retrieval unchanged. For production systems, store the retrieved passages and model version with every response.
Deploy with safeguards
A reliable Hugging Face deployment should separate the model from the source-of-truth knowledge base. Index cleaned circular passages in a search or vector database, apply metadata filters for date and document status, and require the generation step to cite retrieved evidence. Add a visible disclaimer that outputs need professional verification before filing, payment, or a dispute response.
Control access to taxpayer data, redact GSTINs and other personal or commercially sensitive information, and log prompts, retrieved documents, outputs, and user feedback. If inference costs matter, quantisation and smaller adapters can help; this guide to deploying large language models locally covers relevant deployment considerations.
A practical launch checklist
- Confirm that every source document is authentic, versioned, and traceable.
- Preserve page, paragraph, table, and effective-date metadata.
- Define an abstention policy before training.
- Split evaluation data by document and time period.
- Test outdated, conflicting, and incomplete guidance.
- Compare fine-tuning against a RAG-only baseline.
- Require citations for factual answers.
- Review outputs with qualified GST professionals.
- Monitor new circulars and retrain or re-index through a controlled process.
Fine-tuning can make a Hugging Face model more consistent with GST terminology and workflows, but reliable tax AI depends on disciplined data governance, retrieval, evaluation, and human review. Build the narrowest useful system first, prove it on a held-out set of Indian GST documents, and expand only when its evidence trail is dependable.