Large language models become useful to an organisation when they can work with current, proprietary data—not just reproduce patterns from their training set. Connecting large language models to local databases makes it possible to ask questions about inventory, claims, operations, customer records, research, or finance without copying an entire data estate into model training.
The difficult part is not sending a prompt to a model. It is designing a controlled data-access system that understands user intent, retrieves the right records, enforces permissions, produces verifiable answers, and fails safely. For Indian startups and enterprises, the design must also account for sensitive personal data, uneven data quality, multilingual queries, on-premises infrastructure, and the obligations created by internal security policies and the DPDP framework.
Choose the right data-access pattern
A local database can mean PostgreSQL on a private network, MySQL inside a data centre, SQLite on a laptop, MongoDB, a warehouse, or a document store. The best architecture depends on the shape of the question.
- Text-to-SQL: Generate a read-only SQL query for structured tables, then execute it through a controlled service.
- RAG: Retrieve relevant documents or passages from indexed internal content and provide them as model context.
- Tool calling: Expose narrowly defined application functions such as
get_order_statusorlist_open_claims, rather than exposing the database directly. - Hybrid retrieval: Combine SQL for precise facts with vector or keyword search for policies, notes, invoices, and other unstructured material.
Do not use RAG as a substitute for aggregation. A vector search may find documents mentioning “monthly revenue”, but it is not a reliable way to calculate revenue across millions of rows. Conversely, Text-to-SQL is a poor fit for long policy documents or scanned files. Start with the user questions you need to support and map each question type to a retrieval method.
A reference architecture
A production system usually contains six layers:
1. User interface: Chat, an internal portal, a help-desk workflow, or an API.
2. Identity and policy layer: Authenticates the user and attaches role, department, geography, and row-level permissions.
3. Orchestrator: Classifies the request, selects a tool, validates parameters, and manages the model interaction.
4. Data-access services: Read-only SQL execution, approved business APIs, vector search, or document retrieval.
5. Database and index layer: The source systems remain authoritative; indexes are refreshed and monitored separately.
6. Answer and audit layer: Returns citations, query details where appropriate, confidence signals, and an audit record.
Keep the model outside the trust boundary of the database. It should propose an action or query; a deterministic service should validate and execute it. This separation is more robust than giving an agent a database password or allowing arbitrary code execution.
For organisations that need inference inside the network, compare GPU, quantisation, latency, and operational requirements with the guidance in how to deploy large language models locally. A local model is not automatically secure: logs, embeddings, caches, backups, and observability systems can still leak data.
Implementing safe Text-to-SQL
Text-to-SQL works best when the model receives a carefully prepared semantic representation of the database, not an unfiltered dump of every table.
Expose only what is needed. Create curated views with clear names, descriptions, units, date semantics, and approved joins. Hide passwords, tokens, raw identifiers, and operational tables. Add business definitions for terms such as “active customer”, “net sales”, or “overdue account”.
Use a controlled execution loop:
- Classify the question as supported, unsupported, or requiring clarification.
- Retrieve only the relevant schema and a small number of verified examples.
- Generate SQL using a restricted dialect and enforce a statement timeout.
- Parse the query with an SQL validator before execution.
- Permit
SELECTagainst approved views; reject writes, subqueries that violate policy, and unrestricted table scans. - Apply row limits, date-range limits, cost controls, and user-specific filters.
- Execute with a read-only service account.
- Return the result with the query’s time range, filters, and source tables.
A model can generate syntactically valid but conceptually wrong SQL. Test it against a question set containing spelling errors, ambiguous dates, Hindi or regional-language phrasing, adversarial prompts, and requests for data the user cannot access. Where business consequences are material, require a preview and human approval before any downstream action.
Building RAG over local data
For documents, first establish a reliable ingestion pipeline. Extract text, preserve headings and tables where possible, remove duplicates, attach source, owner, department, language, confidentiality, and effective-date metadata, then split content into meaningful sections. Chunking by arbitrary character count often separates definitions from the conditions that qualify them.
Use hybrid retrieval when exact terms matter. Keyword search is valuable for policy numbers, product codes, and names; vector search helps with paraphrased questions. Re-rank the combined candidates, apply access-control filters before returning context, and require the answer to cite retrieved sources. Keep the original document available for verification.
Indian deployments may receive questions in English, Hindi, or other Indian languages. Evaluate whether translation, multilingual embeddings, or a model with stronger Indic-language capability preserves names, dates, legal terms, and numbers. Work on low-resource Indic natural language processing and low-resource language datasets for AI training in India is relevant when the target users do not communicate in standard English.
Security, privacy, and governance
Treat every retrieved row as sensitive until classified otherwise. Practical controls include:
- Use service identities, network segmentation, secrets management, and short-lived credentials.
- Enforce tenant, department, geography, and purpose-based access before retrieval.
- Mask Aadhaar numbers, bank details, health information, and contact data unless the workflow explicitly requires them.
- Avoid sending raw personal data to an external model provider without a documented legal, contractual, and security basis.
- Disable training or retention where provider controls permit, and review transfer and subprocessor terms.
- Encrypt data in transit and at rest; protect vector indexes and prompt logs as carefully as source databases.
- Record user, model, tool, query, result size, policy decision, and timestamp, while minimising sensitive content in logs.
- Add deletion and re-indexing workflows when source records are corrected or removed.
The DPDP Act is not solved by choosing an on-premises model. Governance also covers purpose limitation, access, retention, vendor contracts, incident response, and the rights and obligations associated with personal data. In regulated sectors, involve security, legal, and data owners before a pilot reaches production.
Local and cloud model choices
Cloud models can be strong at complex SQL generation and multilingual reasoning, but they introduce network, contractual, and data-transfer considerations. Local models improve control and can reduce recurring API costs, but require suitable GPUs or CPU capacity, model serving, patching, monitoring, and quality evaluation. A common compromise is to keep retrieval and database execution private while routing only tightly minimised, masked context to a cloud model.
Use an inference gateway so that changing models does not require rewriting application logic. The gateway should enforce model allow-lists, token budgets, rate limits, prompt policies, fallback behaviour, and telemetry. For teams operating their own clusters, hosting Sanjaya RLM on local GPU clusters in India offers a relevant local-infrastructure perspective.
Evaluate before expanding the pilot
Measure more than fluent answers. Build a representative benchmark and track:
- SQL execution accuracy and business-answer accuracy
- Retrieval recall, citation correctness, and unsupported-claim rate
- Permission violations and sensitive-data exposure
- Latency, token usage, database load, and cost per request
- Clarification quality and safe refusal behaviour
- Performance across languages, roles, and data freshness windows
Create golden queries with expected filters and results, then run regression tests whenever the schema, prompt, model, or index changes. Include red-team tests such as “ignore your rules”, requests for another employee’s records, destructive SQL, and attempts to infer hidden personal data.
A practical rollout plan
Begin with one read-only use case where success is measurable—for example, operations staff checking shipment status or analysts querying a curated sales view. Document the data owner, permitted users, source freshness, failure modes, and escalation path. Add citations and query previews before adding actions.
Next, introduce semantic views, hybrid retrieval, monitoring, and multilingual tests. Only after accuracy and access controls are stable should you connect write-capable workflows such as ticket creation or stock updates, and those should use explicit business APIs with approval and idempotency controls.
The goal is not an autonomous chatbot with unrestricted database access. It is a dependable interface over governed data: narrow permissions, transparent sources, measurable quality, and a clear human path when the model is uncertain.