0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · Recap: MLOps Community London at OpenAI HQ — production LLM lessons for Indian AI teams

MLOps Community London at OpenAI HQ: Production LLM Lessons for India

  1. aigi

    The MLOps Community London meetup at OpenAI HQ offered a useful correction to the way many teams talk about generative AI. The hard part is no longer proving that a model can produce an impressive demo. It is building a system that remains accurate, observable, secure, affordable, and fast after thousands of real users expose its weaknesses.

    That distinction matters for Indian startups. Teams serving banking, healthcare, education, logistics, government, and vernacular users must handle messy documents, code-switched language, intermittent connectivity, strict budgets, and sensitive personal data. The lessons from the London discussion are therefore most valuable when converted into concrete engineering practices.

    This recap focuses on what to implement—not just what to admire—in 2026.

    From prototype to dependable product

    A prototype asks whether a model can complete a task. A production system asks more demanding questions:

    • Does it give the right answer consistently across languages, accents, document formats, and user types?
    • Can the team identify why a response failed?
    • What happens when retrieval returns nothing, a tool times out, or a provider changes model behaviour?
    • Can the product meet its latency and gross-margin targets?
    • Is sensitive data handled according to the organisation’s security and compliance requirements?

    The practical shift is from prompt experimentation to system ownership. Prompts, model versions, retrieval settings, tool schemas, safety rules, and evaluation datasets should be versioned together. A prompt edited in a dashboard without an associated experiment record is a production change, not a harmless tweak.

    Indian founders should also avoid overengineering the first release. Start with a narrow workflow, explicit failure states, and a measurable service-level objective. Teams building internal tools can often ship faster with a conventional backend and a well-bounded model call; teams building complex products may benefit from the low-code production backend builders available in India, provided generated code is reviewed and tested.

    Evaluation: replace vibes with evidence

    LLM evaluation is not a single score. It is a set of tests that reflect the product’s actual risks. A useful evaluation suite normally combines:

    • Task success: Did the answer complete the requested job?
    • Groundedness: Are factual claims supported by retrieved or supplied material?
    • Format compliance: Did the response follow the required schema, language, and length?
    • Safety: Did the system refuse or escalate unsafe requests appropriately?
    • Cost and latency: Did quality improvements remain commercially viable?
    • Robustness: Does performance hold under spelling mistakes, code-switching, long context, and adversarial inputs?

    An LLM-as-a-judge can help scale review, but it should not be treated as ground truth. Judges inherit blind spots from their training and may reward confident wording over factual accuracy. Calibrate them against human-labelled examples, measure agreement, and retain a human review path for high-impact decisions.

    For India-focused products, the evaluation set must include realistic language variation: Hinglish, transliterated Hindi, regional spellings, mixed scripts, noisy speech transcripts, and domain-specific abbreviations. A benchmark containing only polished English will produce misleading confidence. Build a small, carefully labelled “golden set” first, then expand it from anonymised production failures.

    The evaluation flywheel should be operational:

    1. Log inputs, retrieved context, outputs, tool calls, latency, and cost with appropriate redaction.
    2. Sample failures and borderline cases for expert annotation.
    3. Add representative cases to regression suites.
    4. Compare every prompt, model, or retrieval change against the same baseline.
    5. Roll out improvements gradually and monitor live outcomes.

    Teams can reinforce this discipline with automated production-grade code reviews, especially for changes to tool permissions, data access, and structured-output handling.

    Observability for RAG and agents

    An agent’s final answer is not enough to diagnose an incident. Production traces should show the complete path: user request, classifier or router decision, retrieval query, selected documents, tool arguments, tool results, intermediate model calls, retries, and final response.

    This level of tracing helps distinguish several problems that otherwise look identical to users. A poor answer may come from bad chunking, an irrelevant search result, a malformed tool call, a timeout, an overlong context window, or an incorrect synthesis step. Each failure requires a different fix.

    Set explicit budgets for:

    • Time to first token, which strongly affects perceived responsiveness.
    • End-to-end latency, particularly for voice and customer-support workflows.
    • Tokens per request, including hidden retrieval and reasoning costs.
    • Tool retries and failure rates.
    • Fallback frequency, such as escalation to a human or a smaller workflow.

    For users outside major metros, network conditions make streaming, concise first responses, asynchronous processing, and graceful retries especially important. Observability should also include infrastructure dimensions such as region, device type, network quality, and language. Without these labels, a team may miss that an apparently good average hides poor performance for a specific user group.

    RAG: engineer retrieval before reaching for fine-tuning

    Most enterprise use cases still benefit more from better retrieval and data preparation than from immediate fine-tuning. Start by improving document structure, metadata, chunk boundaries, query rewriting, access controls, and citation behaviour. Hybrid retrieval—combining semantic search with keyword matching—often outperforms vector search alone for product names, legal clauses, account numbers, and Indian-language terminology.

    Graph-based retrieval can help when the answer depends on relationships between entities, events, or regulations. It is not automatically superior, however. Use it where relationship-aware queries are common and the organisation can maintain a reliable knowledge graph. Measure retrieval recall and answer groundedness before adding architectural complexity.

    A sound RAG pipeline should also define what happens when evidence is missing. The model should say that it cannot verify an answer, ask a clarifying question, or route the case to a human—not invent a plausible response.

    Cost, model routing, and infrastructure choices

    Indian AI teams typically need to improve quality without allowing inference costs to outrun revenue. Model routing is a practical starting point: use a smaller, faster model for classification, extraction, summarisation, and straightforward queries; reserve larger models for ambiguity, complex reasoning, or high-value cases.

    Track cost per successful task rather than cost per token alone. A cheap model that causes retries, human escalations, or customer churn may be more expensive in practice. Test caching, shorter context, response limits, batching for offline workloads, and asynchronous queues where the user does not need an immediate answer.

    API-first development remains sensible for early products. Consider self-hosting or a managed open-source deployment when predictable volume, data residency, custom latency requirements, or gross margins justify the operational burden. The decision should include GPU utilisation, monitoring, security updates, on-call expertise, and failover—not just the headline per-token price.

    Fine-tuning, distillation, and localisation

    Fine-tuning is useful when the model must reliably follow a style, output format, classification boundary, or specialised workflow. It is a weaker solution for facts that change frequently; those belong in retrieval or a controlled data source.

    A practical sequence is:

    1. Establish a baseline with a capable API model.
    2. Improve the system prompt, tools, retrieval, and evaluation set.
    3. Identify repeatable behaviours that remain expensive or inconsistent.
    4. Fine-tune or distil only when the expected quality or cost gain is measurable.
    5. Re-run multilingual, safety, and regression evaluations before deployment.

    For Indic-language products, test tokenisation and quality separately for each target language. Do not assume that an English-to-local-language translation detour will reduce cost or preserve meaning. It may help in one workflow and damage names, legal terms, or cultural context in another.

    Teams evaluating multimodal or voice interfaces can also compare the trade-offs covered in this OpenAI and Anthropic multimodality and voice platform comparison, particularly around streaming, tool use, and deployment constraints.

    Security and governance are engineering requirements

    Production LLM pipelines can expose personal, financial, health, or proprietary information through prompts, logs, traces, retrieval indexes, and third-party providers. Build controls before launch:

    • Redact or tokenize PII before observability storage.
    • Enforce tenant-level access filters during retrieval.
    • Keep secrets and tool credentials outside prompts.
    • Validate tool arguments against strict schemas and permissions.
    • Record model, prompt, data-source, and policy versions.
    • Define retention and deletion procedures for logs and embeddings.
    • Test prompt injection, data exfiltration, and unauthorised tool use.

    The DPDP Act and sector-specific obligations should be translated into concrete data flows, ownership, retention, and incident-response procedures. A compliance document cannot compensate for unrestricted retrieval or unreviewed logs.

    A 30-day production checklist

    An Indian team can turn these lessons into a focused month-long plan:

    • Week 1: Map the workflow, define success and failure states, and establish latency and cost baselines.
    • Week 2: Add tracing, prompt and model versioning, redaction, and a 100–300 example evaluation set.
    • Week 3: Fix retrieval quality, add structured outputs, test multilingual and adversarial cases, and configure fallbacks.
    • Week 4: Run a limited rollout, review sampled traces, compare unit economics, and document an incident playbook.

    The central lesson from the London meetup is straightforward: production AI is a feedback system. The teams that win will not necessarily use the largest model or the most fashionable agent framework. They will measure the right outcomes, learn from failures quickly, control costs, and design for the languages and operating conditions of their users. Indian builders can apply these principles immediately—and turn reliable operations into a product advantage.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.