0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to optimize karpathy autoresearch for indexing gst council meeting minutes

How to Optimize Karpathy Autoresearch for GST Minutes

  1. aigi

    Why GST Council minutes need a better search layer

    GST Council meeting minutes are high-value public records for tax professionals, businesses, researchers, journalists, and government teams. Yet the information is often difficult to retrieve: documents may be PDFs with inconsistent layouts, scanned pages, tables, annexures, abbreviations, and references to earlier decisions. A search system that matches only exact keywords will miss important passages and may return a technically similar but legally irrelevant result.

    A practical implementation of Karpathy Autoresearch should therefore be treated as a document-research and evaluation workflow, not a button that automatically makes files searchable. The goal is to preserve the source record, extract usable text, attach India-specific metadata, and return cited answers that users can verify against the original minutes.

    For teams also extracting decisions and follow-ups, an AI bot for extracting meeting action items can complement the indexing pipeline. Keep action-item extraction separate from the authoritative archive so generated interpretations never overwrite the official record.

    1. Define the archive and its source-of-truth rules

    Start by listing the documents to be indexed and recording their provenance. For every file, capture:

    • Meeting number and meeting date
    • Publication date, if different
    • Official source URL and download timestamp
    • Document title and language
    • File hash and version number
    • Whether the file is an agenda, minutes, press release, circular, annexure, or corrigendum

    Do not combine minutes with press releases without labelling the distinction. A press release may summarise a decision, while minutes may contain the fuller discussion, conditions, dissenting views, or implementation context. Store the original PDF unchanged and generate a separate, versioned processing copy.

    Before ingestion, check the official GST Council or government publication source and define a correction policy. If a document changes, retain both versions and mark the later file as a replacement or corrigendum. This matters when a user needs to establish what was publicly available at a particular time.

    2. Build a document preparation pipeline

    The quality of retrieval is usually constrained by extraction quality. A sensible pipeline has four stages:

    1. File inspection: Detect whether the PDF contains selectable text, scanned images, tables, or embedded attachments.
    2. OCR: Run OCR on scanned pages, retaining page numbers and confidence scores. Use an Indian-language-capable model when documents include Hindi or other regional-language content.
    3. Layout recovery: Preserve headings, numbered agenda items, speaker labels, tables, footnotes, and annexure boundaries.
    4. Normalisation: Fix obvious OCR errors without silently changing the source. Keep both raw extracted text and a cleaned search representation.

    Never discard page coordinates. They allow the application to cite the exact page, display the source excerpt, and help a reviewer verify a result. OCR confidence should also be searchable or available as a warning signal; a low-confidence match deserves human review before it is used in a policy or compliance decision.

    Tables require special handling. Tax rates, exemption lists, dates, and implementation conditions can become misleading when columns collapse into plain text. Extract tables into structured records where possible, while retaining an image or PDF reference for visual verification.

    3. Create GST-aware metadata and terminology

    Generic metadata such as “tax” or “meeting” is not enough. Generate a controlled metadata schema that reflects how Indian GST research is actually conducted. Useful fields include:

    • GST Council meeting number and date
    • Agenda item number and subject
    • Decision, recommendation, clarification, or discussion status
    • Tax type: CGST, SGST, IGST, UTGST, or compensation cess
    • Relevant HSN or SAC codes
    • Rate, exemption, threshold, or place-of-supply reference
    • State or Union Territory mentioned
    • Effective date and implementation dependency
    • Related notification, circular, press release, or Finance Act reference
    • Language, page range, OCR confidence, and document version

    Maintain an alias dictionary. The same concept may appear as “GST Council,” “Council,” a meeting number, an abbreviated ministry name, or a product description that varies across documents. Add known spelling variants, punctuation variants, HSN/SAC formats, and common OCR substitutions. Do not treat the dictionary as a substitute for source review; use it to improve recall while preserving the original wording in results.

    4. Choose chunking and indexing for legal and policy context

    Chunking should follow the document’s meaning rather than a fixed character count alone. A useful hierarchy is:

    • Document
    • Agenda item
    • Subsection or paragraph group
    • Table or annexure
    • Individual decision or recommendation

    Include a small amount of surrounding context in each indexed chunk, but keep the agenda number and page reference attached. Overly large chunks reduce precision; overly small chunks lose qualifiers such as “subject to,” “with effect from,” or “as recommended by the committee.”

    Use hybrid retrieval: combine lexical search for exact terms, numbers, HSN/SAC codes, and dates with semantic search for paraphrased questions. Rerank the candidates using metadata, document authority, date, and agenda-item relevance. For example, a query about a rate change should prioritise a decision or recommendation containing the effective date over a passing mention in an introductory section.

    If the index will serve a public web application, review deployment constraints using guidance on optimizing system performance for web apps. Keep ingestion, embedding generation, and user search as separate jobs so a large reindex does not make the archive unavailable.

    5. Use Autoresearch as an experiment loop

    Treat each change to extraction, chunking, metadata, prompts, or ranking as an experiment. Keep a fixed evaluation set of representative questions, such as:

    • Which GST Council meeting discussed a particular commodity or service?
    • What rate or exemption was recommended, and from what date?
    • Which agenda item refers to a specified HSN or SAC code?
    • What conditions or implementation dependencies accompanied the decision?
    • Which document is the latest authoritative version?

    For every experiment, record the configuration, model version, corpus version, latency, cost, and sample outputs. This makes improvements reproducible and prevents a change that raises recall from quietly damaging citation accuracy. A privacy-sensitive deployment should also review how to optimize large language models for privacy apps, particularly when unpublished drafts or internal annotations are included in the workflow.

    6. Evaluate retrieval and answer quality separately

    Measure retrieval before judging generated answers. Recommended metrics include:

    • Recall@k: Whether the relevant page or agenda item appears in the top results
    • Precision@k: How many returned results are genuinely relevant
    • Mean reciprocal rank: How quickly the first correct result appears
    • Citation accuracy: Whether every claim points to the correct document and page
    • Effective-date accuracy: Whether the system preserves dates and conditions
    • OCR error rate: Especially for numbers, codes, and names
    • Latency and cost: Split by ingestion, retrieval, reranking, and generation

    Create a small gold set reviewed by a GST practitioner, tax researcher, or experienced policy analyst. Include difficult examples: scanned pages, tables, repeated agenda items, conflicting versions, and queries using informal business language. Test bilingual and code-mixed queries if your users need them. A “no reliable answer found” outcome is preferable to an unsupported conclusion.

    7. Add safeguards before public release

    Every generated response should display the document title, meeting date, page number, source link, and a short quoted passage. Label summaries as machine-generated and provide a direct route to the original PDF. Apply access controls if the corpus includes internal working documents, and log retrieval and answer events without storing unnecessary personal data.

    Set human-review triggers for low OCR confidence, conflicting versions, numerical claims, legal interpretation, and questions about current applicability. Meeting minutes record deliberation and decisions; they do not automatically replace notifications, circulars, or statutory text. The interface should say this clearly.

    Practical rollout plan

    A staged launch is safer than indexing the entire archive at once:

    • Pilot: Process a representative set of recent minutes and validate OCR, citations, and metadata.
    • Expand: Add historical documents, annexures, language variants, and correction handling.
    • Harden: Introduce hybrid retrieval, reranking, monitoring, and reviewer workflows.
    • Maintain: Re-run evaluations whenever the corpus, parser, embedding model, or prompt changes.

    The strongest system is not the one that produces the most fluent summary. It is the one that helps a user find the right GST Council passage quickly, shows exactly where it came from, preserves uncertainty, and makes correction straightforward.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.