Why legacy integration needs an engineering plan
Indian enterprises rarely have the option of replacing a core banking platform, factory ERP, government registry, or claims engine in one project. These systems contain decades of business rules, operational data, and integrations. They may run on mainframes, Oracle, older Java services, flat files, or tightly coupled vendor software—but they remain the system of record.
Integrating large language models with legacy systems is therefore an orchestration problem, not a prompt-writing exercise. An LLM can interpret language, summarise records, classify requests, and choose tools. It should not silently become the authority for balances, eligibility, pricing, compliance, or irreversible transactions.
The safest architecture keeps the legacy platform authoritative while placing the model behind controlled APIs, retrieval services, policy checks, and human approval.
Start with the right use case
Begin with workflows where the model improves access to existing capabilities without changing the underlying source of truth. Strong first candidates include:
- Read-only knowledge access: answer questions over SOPs, product manuals, circulars, and internal policies.
- Service-desk assistance: summarise tickets, suggest responses, and retrieve account or shipment status.
- Document intake: extract fields from invoices, applications, inspection reports, and claims for validation.
- Legacy-code understanding: explain COBOL, PL/SQL, or older Java modules and generate test cases.
- Operations copilots: translate natural-language requests into approved queries or workflow actions.
Avoid starting with autonomous write access to payroll, lending decisions, procurement, or production controls. A read-only pilot produces measurable value while exposing data quality, latency, and access-control issues early.
Teams building multilingual interfaces should also plan for India’s language diversity. A model may need to handle English, Hindi, and regional-language queries while returning structured identifiers exactly as stored. Work on low-resource Indic natural language processing and open-source small language models for Hindi can inform model selection, evaluation, and on-premise deployment.
Reference architecture
A production integration normally has six layers:
1. User and application layer: web, mobile, contact-centre, WhatsApp, or internal workstation.
2. LLM gateway: model routing, authentication, rate limits, prompt templates, logging, and provider controls.
3. Orchestration layer: decides whether to retrieve documents, call a system API, ask a clarification question, or escalate.
4. Data and tool layer: API facades, search indexes, vector stores, query services, and approved business functions.
5. Policy and validation layer: identity checks, authorisation, PII handling, schema validation, and transaction limits.
6. Legacy system of record: mainframe, ERP, CRM, database, or proprietary application.
This separation matters. It lets the organisation change models without rewriting core integrations, and it prevents an LLM from receiving unrestricted database credentials.
For complex workflows, treat the orchestrator as a controlled distributed system rather than a chain of ad hoc prompts. Explicit state, retries, timeouts, idempotency keys, and audit events are essential; the principles in building distributed systems with AI agents are directly relevant.
Expose legacy capabilities safely
If a system has no modern API, create a capability facade around specific operations. The facade can call stored procedures, message queues, terminal interfaces, or existing Java and COBOL services, but it should expose narrow, documented functions such as get_policy_status or check_inventory, not arbitrary SQL execution.
Every tool should define:
- accepted inputs and strict data types;
- the user roles allowed to call it;
- timeout and retry behaviour;
- whether the operation is read-only or mutating;
- expected error codes and fallback messages;
- an audit record containing actor, purpose, inputs, outputs, and approval status.
Use structured tool calls and validate them before execution. For write operations, require a preview, confirmation, and—where appropriate—human approval. Make actions idempotent so a retry cannot create duplicate payments, orders, or service requests.
Use RAG for documents, APIs for facts
Retrieval-Augmented Generation is useful for policy documents, manuals, circulars, and historical case notes. It is not a substitute for calling the authoritative system when the user asks for a current balance, stock count, claim status, or account attribute.
A practical retrieval pipeline includes document ingestion, OCR where necessary, layout-aware chunking, metadata tagging, hybrid keyword and vector search, reranking, and citation of source passages. Keep document versions and effective dates. In regulated environments, the answer should identify the source and indicate when no reliable evidence was found.
For live records, query the legacy system through an API or read replica and provide only the minimum fields required by the model. Change-data-capture pipelines can update search indexes, but they must account for deletes, corrections, schema changes, and delayed events. Never describe a stale vector index as real-time data.
Control hallucinations and unsafe output
Reliability comes from constraining the workflow, not from assuming a larger model will always behave correctly. Apply several controls together:
- require citations for knowledge answers;
- use JSON schemas or typed models for machine-readable output;
- reject missing, extra, or invalid fields;
- verify calculations and policy conditions in deterministic code;
- distinguish “not found” from “not authorised” and “system unavailable”;
- set confidence and escalation thresholds;
- log prompts, retrieved context, tool calls, and final responses with appropriate redaction.
Evaluate with a representative test set containing spelling mistakes, mixed languages, ambiguous identifiers, outdated policies, adversarial prompts, and partial system failures. Track grounded-answer rate, tool-call accuracy, refusal quality, latency, cost per task, and human correction rate—not just generic benchmark scores.
Security and compliance for Indian deployments
Map data flows before selecting a model provider. Classify Aadhaar-linked information, PAN details, financial records, health data, employee information, and proprietary code. Apply least-privilege access, encryption in transit and at rest, secrets management, tenant isolation, retention limits, and provider-level controls that prevent enterprise prompts from being used for training where required.
Under India’s DPDP framework and sector-specific rules, obligations depend on the data, purpose, organisation, and processing arrangement. Involve the security, legal, and compliance teams early. Tokenise or redact PII before inference where feasible, but preserve a secure mapping when the response must be attached to the correct record. Keep sensitive operations inside a controlled VPC, private cloud, or on-premise environment when risk and latency justify it.
Voice and contact-centre deployments need additional safeguards: consent, recording retention, speaker or account verification, and safe transfer to an agent. A voice interface should complement—not bypass—existing authentication controls; related design considerations appear in integrating a voice agent with Twilio telephony.
A phased delivery roadmap
Phase one: inventory and baseline. Map systems of record, data owners, interfaces, failure modes, and current handling time. Choose one workflow with a clear business metric.
Phase two: read-only pilot. Build retrieval or API access with no write permissions. Establish evaluation datasets, observability, red-team tests, and a human escalation path.
Phase three: governed actions. Add narrowly scoped tools, approval gates, transaction limits, and rollback procedures. Test outages, duplicate requests, prompt injection, and malformed outputs.
Phase four: scale deliberately. Add workloads only after measuring cost, latency, accuracy, and operational burden. Use model routing: smaller models for classification and extraction, stronger models for difficult reasoning, and deterministic services for calculations.
The business case should include reduced handling time, fewer manual lookups, faster onboarding, improved first-contact resolution, and avoided errors. Include integration maintenance, evaluation, security reviews, observability, and human review in the total cost—not only token fees.
Practical checklist
Before production, confirm that:
- the legacy system remains the authoritative source;
- every model action has a bounded tool and permission scope;
- sensitive data is classified, minimised, and logged safely;
- answers can cite evidence or clearly state uncertainty;
- write operations support approval, idempotency, and rollback;
- failures degrade to a human or deterministic workflow;
- evaluation covers Indian languages, real records, and adversarial cases;
- owners exist for the model, prompts, APIs, data, and incident response.
The strongest enterprise deployments do not hide the legacy estate. They wrap it with disciplined interfaces, preserve its controls, and use language models where interpretation is valuable. That approach delivers useful AI without betting critical operations on an unverified autonomous layer.