A production GenAI application is not a prompt wrapped in an API. It is a software system that must remain useful when models change, documents are incomplete, traffic spikes, users enter adversarial inputs, and cloud costs rise. The engineering target is reliable behaviour under measurable constraints: quality, latency, availability, safety, privacy, and cost.
For Indian startups and enterprise teams, there is an additional constraint: build for uneven connectivity, multilingual users, strict data boundaries, and price-sensitive customers. The right approach is to start with a narrow workflow, define what “good” means, and then make every model call observable and replaceable.
1. Define the production contract before choosing a model
Write a short technical and product contract before selecting GPT, Claude, Gemini, an open-weight model, or a hosted Indian-language model. It should specify:
- Supported tasks: What the application does and, equally important, what it refuses to do.
- Quality targets: Accuracy, groundedness, citation quality, structured-output validity, and escalation rate.
- Performance targets: Time to first token, full-response latency, requests per second, and uptime.
- Cost limits: Maximum input and output tokens, cost per successful task, and monthly budget.
- Risk controls: Personally identifiable information, regulated advice, data retention, and human approval requirements.
Separate the model interface from business logic. A provider adapter should let you switch models, apply retries, record usage, and enforce timeouts without rewriting the application. This matters when an API changes pricing, suffers an outage, or performs poorly on a regional language.
If the product involves autonomous actions, treat it as a distributed system rather than a chatbot. The principles in building distributed systems with AI agents are especially relevant for queues, retries, idempotency, state, and failure recovery.
2. Build retrieval for evidence, not just similarity
Retrieval-Augmented Generation (RAG) is often the best first architecture for private or frequently changing information. A production pipeline should make the path from source document to final answer inspectable:
1. Ingest documents with source, owner, timestamp, language, and access metadata.
2. Parse PDFs, tables, scans, and web pages while preserving headings and page references.
3. Split content by semantic structure rather than an arbitrary character count.
4. Generate embeddings and store both vectors and searchable metadata.
5. Combine vector retrieval with keyword or lexical search for names, codes, policies, and exact phrases.
6. Rerank the candidate passages before sending context to the model.
7. Require citations or evidence spans where users need to verify an answer.
Apply document-level permissions at retrieval time. Filtering results after generation is too late: confidential text may already have entered the prompt or appeared in logs. For complex entities and relationships, a graph layer can complement vector search, but it should solve a demonstrated retrieval problem rather than add architectural ceremony.
Test ingestion as aggressively as generation. Broken OCR, stale indexes, duplicate chunks, and missing page references can create convincing but incorrect answers. Maintain a small golden set of questions tied to expected source passages, including Hindi, Tamil, Marathi, and mixed-script queries when those languages are in scope.
3. Treat evaluation as a release gate
LLM output is variable, so “it worked in testing” is not a quality strategy. Create an evaluation set from real user tasks, support tickets, domain examples, and known failure cases. Include both successful and adversarial inputs.
Use several evaluation layers:
- Deterministic checks: JSON schema validity, required fields, citation presence, prohibited content, and tool-call permissions.
- Reference-based checks: Exact or semantic comparison for classification, extraction, and known-answer tasks.
- RAG checks: Context relevance, answer faithfulness, citation correctness, and handling of insufficient evidence.
- Human review: Domain experts score ambiguous, high-risk, or multilingual cases.
- Operational metrics: Latency, token usage, retries, fallback frequency, and user feedback.
Frameworks such as Ragas, DeepEval, Promptfoo, and custom test harnesses can run in CI. Do not rely blindly on an LLM judge: calibrate it against human ratings, provide explicit rubrics, and track disagreement. A prompt, model, retriever, or chunking change should be compared against a fixed baseline before release.
4. Design latency, resilience, and cost together
Users experience the whole workflow, not just model generation. Stream safe responses, display progress for long jobs, and move indexing, analytics, bulk processing, and evaluation to background workers. Measure time to first token, time to useful evidence, and total completion time separately.
Use practical resilience patterns:
- Set connection, model, and total-request timeouts.
- Retry only transient failures, with exponential backoff and jitter.
- Make tool calls and queued jobs idempotent.
- Add circuit breakers and provider fallbacks where contracts permit.
- Degrade gracefully to search results, a smaller model, or human review.
- Apply per-user, tenant, and endpoint rate limits.
Control cost at the workflow level. Route classification and extraction to smaller models, cap context and output tokens, summarise long histories, cache stable results, and track cost per completed business outcome rather than cost per request alone. For infrastructure planning, pair these choices with guidance on scaling backend infrastructure for AI applications.
5. Add security and governance before launch
Prompt injection is an application-security problem, not merely a bad prompt. Assume that retrieved documents, web pages, uploaded files, and user messages may contain instructions designed to manipulate the system.
Use layered controls:
- Keep system instructions and untrusted content clearly separated.
- Allowlist tools, arguments, domains, and actions.
- Require confirmation for payments, messages, deletions, and irreversible changes.
- Validate model-generated structured data before execution.
- Redact or hash sensitive fields in logs.
- Encrypt data in transit and at rest, with tenant isolation.
- Define retention, deletion, access, and incident-response procedures.
- Record model version, prompt version, retrieved sources, tool calls, and policy decisions.
For legal, financial, health, or public-sector use cases, route uncertain answers to a human rather than presenting confidence as correctness. A private legal assistant, for example, needs access controls and source traceability in addition to a capable model; the architecture described in building a private AI chatbot for lawyers illustrates that distinction.
6. Make observability useful to engineers
Traditional uptime dashboards cannot explain why an answer was wrong. Use trace-based observability to connect a user request to retrieval, reranking, prompt construction, model generation, tool execution, and final response.
Track at least:
- Request and trace IDs across services.
- Model, provider, prompt, embedding, and index versions.
- Input and output token counts, cost, and cache hits.
- Retrieval scores, selected sources, and context length.
- Latency by stage, timeout, retry, and fallback reason.
- Safety intervention, refusal, escalation, and user-feedback rates.
Sample or redact payloads to protect privacy. A thumbs-up button is useful only when it captures the task, model version, evidence, and failure category needed for diagnosis. Feed reviewed failures back into the evaluation set; do not automatically fine-tune on unverified conversations.
7. Build for India’s languages and operating conditions
Benchmark multilingual performance with real user inputs, not translated English prompts alone. Test code-switching, transliteration, local names, dates, currency formats, noisy voice transcripts, and regional terminology. Indic scripts can have different tokenisation and retrieval behaviour, affecting both quality and cost.
For voice products, budget separately for speech recognition, reasoning, and text-to-speech latency. How to build a voice agent covers the broader architecture, while production teams should additionally test interruptions, weak connectivity, accents, consent, and transcript privacy.
Choose hosting based on data residency, support, latency, and total cost—not just benchmark scores. Keep a documented fallback plan for provider outages and model deprecations. Before launch, run load tests with realistic prompt lengths, concurrency, document sizes, and multilingual traffic.
8. A practical launch checklist
Before opening the system to real users, confirm that you can:
- Reproduce an answer from its trace and versioned inputs.
- Explain which evidence supported the answer.
- Block unauthorised data and tool actions.
- Detect quality regressions before deployment.
- Recover from provider, queue, database, and retrieval failures.
- Estimate and cap cost by tenant and workflow.
- Export, delete, and protect user data.
- Escalate uncertain or high-impact cases to humans.
Start with a narrow production slice, measure it for several weeks, and expand only when quality and operational targets hold. Production readiness is a continuing discipline: every incident should improve the tests, telemetry, safeguards, or architecture.