Large language models can summarise filings, classify customer queries, extract entities from invoices, and support analysts. But a general-purpose model is not automatically a reliable financial system. Finance demands precise numbers, source traceability, strict access controls, and careful handling of time-sensitive information.
LLM fine-tuning on financial data works best when it is treated as an engineering and governance programme—not simply as uploading documents and running another training job. This guide covers the decisions that matter: when to fine-tune, what data to use, how to avoid leakage, how to evaluate results, and how to deploy safely in Indian financial environments.
Fine-tuning versus retrieval
Fine-tuning changes a model’s behaviour or capabilities by training it on examples. Retrieval-augmented generation (RAG) supplies relevant documents at query time without changing the model’s weights. The two approaches solve different problems:
- Use RAG for changing facts such as prices, RBI circulars, product terms, policies, and the latest company filings.
- Use fine-tuning for repeatable behaviour such as structured extraction, classification, tone, output formats, and domain-specific instruction following.
- Use both when a model must apply a consistent workflow to current, source-grounded information.
A strong baseline is often a capable base model plus RAG and deterministic tools for calculations. Fine-tuning should be justified by measurable gains in accuracy, latency, cost, or consistency. Teams new to the process can first review these best practices for fine-tuning LLMs on custom data.
What financial data should you prepare?
The right dataset depends on the task. Useful sources include:
- Regulatory and corporate documents: RBI, SEBI, IRDAI, PFRDA, MCA filings, annual reports, prospectuses, credit reports, and policy documents.
- Transaction and operations data: anonymised payment descriptions, ledger entries, invoices, reconciliation records, and support tickets.
- Market and economic data: prices, volumes, corporate actions, macroeconomic indicators, and analyst research, with timestamps and licensing records.
- Customer-service conversations: consented and redacted chats, call transcripts, complaint categories, and approved responses.
- Indian-language content: financial terms and customer requests in Hindi, Tamil, Marathi, Bengali, Telugu, and other relevant languages, including code-mixed text.
Do not treat all data as equally trustworthy. Preserve source, timestamp, jurisdiction, document version, and access classification for every record. For high-stakes applications, a data veracity infrastructure approach helps track provenance, contradictions, stale information, and human review.
Build a training-ready dataset
Dataset quality usually matters more than adding another billion tokens. Start with a clear task specification: input, expected output, acceptable alternatives, escalation conditions, and prohibited behaviour.
A practical pipeline should:
1. Inventory rights and consent. Confirm whether data may be used for training, and document contractual, regulatory, and customer-consent restrictions.
2. Remove sensitive information. Redact account numbers, PAN, Aadhaar, card details, credentials, phone numbers, addresses, and free-text identifiers. Replace values with consistent placeholders where relationships must be preserved.
3. Normalise documents. Repair OCR errors, preserve tables where possible, standardise dates and currencies, and distinguish rupees, lakhs, crores, and foreign currencies.
4. Create high-quality examples. Use expert-written instruction-response pairs, labelled classifications, extraction targets, or preference comparisons. Include difficult and ambiguous cases.
5. Record provenance. Store source identifiers, versions, timestamps, annotators, and review decisions alongside each example.
6. Split by time and entity. Keep related documents from the same issuer, customer, or event in one split. A random split can leak near-duplicates and inflate results.
For repeatable preprocessing, teams can use tested Python scripts for automating data preprocessing, with unit tests and sample-based review before processing the full corpus.
Choose the fine-tuning method
Full-parameter fine-tuning is expensive and can overwrite useful general capabilities. Parameter-efficient methods are usually more practical for Indian startups and regulated teams:
- LoRA and QLoRA: Train small adapter layers while keeping the base model largely frozen. They reduce memory use and make versioning easier.
- Supervised fine-tuning: Teach a model to produce the desired answer or structured output from curated examples.
- Preference optimisation: Improve choices between acceptable and unacceptable responses, especially for tone, escalation, and policy adherence.
- Continued pretraining: Adapt a model to large volumes of domain language, but use it carefully; it does not by itself teach reliable task behaviour.
Choose the smallest model that meets the quality, latency, and deployment requirements. Test multilingual and code-mixed performance separately rather than assuming English benchmarks transfer to Indian users.
Evaluate financial reliability
Accuracy alone is insufficient. Create a holdout set that the training team cannot inspect, and evaluate by task, product, language, customer segment, and risk level. Important measures include:
- Extraction: field-level precision, recall, and numeric accuracy.
- Classification: macro-F1, confusion matrices, and performance on rare but consequential classes.
- Generation: factuality, citation correctness, completeness, and human preference.
- Operations: latency, token cost, throughput, abstention rate, and escalation rate.
- Safety: privacy leakage, prompt injection resistance, unauthorised advice, discriminatory outputs, and behaviour under adversarial inputs.
Use temporal backtesting for market or risk applications. Never evaluate a model with information that would not have been available at the decision time. Compare against simple baselines, rules, and human reviewers. A model that sounds convincing but misreads a decimal, date, or unit should fail the gate.
Deployment controls for India
Financial AI should have clear boundaries. Keep calculations, eligibility rules, limits, and transaction execution in deterministic services wherever possible. The language model should call approved tools rather than invent figures. Require citations or source links for material claims, and make uncertainty visible.
Recommended controls include:
- role-based access and encryption for training and inference data;
- tenant isolation for banks, fintechs, lenders, and enterprise customers;
- immutable audit logs for prompts, retrieved sources, model versions, tool calls, and outputs;
- human approval for credit, investment, insurance, collections, or customer-impacting decisions;
- monitoring for drift, data leakage, hallucinations, and changes in refusal or escalation rates;
- documented retention, deletion, incident response, and vendor-risk processes.
Map controls to the organisation’s obligations under applicable Indian financial regulation, privacy requirements, outsourcing contracts, and internal model-risk policy. Do not assume that a hosted model provider offers the required data residency, retention, or isolation by default.
A practical pilot plan
Start with a narrow, measurable workflow such as classifying service requests, extracting fields from statements, or drafting analyst summaries with mandatory citations. Establish a baseline, create a reviewed dataset, fine-tune an adapter, and run a time-split evaluation. Then conduct a silent production trial where outputs are logged but not acted upon.
Move to limited release only when quality, safety, cost, and escalation targets are met. Give reviewers a fast way to correct outputs and feed approved corrections into a governed improvement cycle. Track every model and dataset version so that a result can be reproduced months later.
Common mistakes
- Fine-tuning to memorise live prices or regulations instead of using retrieval.
- Training on raw customer data without consent, redaction, or access controls.
- Randomly splitting documents and reporting inflated scores.
- Measuring only fluent answers instead of numerical and source accuracy.
- Allowing the model to execute financial actions without deterministic checks.
- Launching multilingual support without language-specific evaluation.
The goal is not a model that appears financially sophisticated. It is a system that is accurate within its defined scope, transparent about sources and uncertainty, economical to operate, and safe to override. For Indian builders, disciplined data governance and evaluation will create more durable advantage than fine-tuning alone.