0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · Recap: MLOps Community London March 2026 — ops patterns for shipping LLM apps from India

MLOps Community London March 2026: Shipping LLM Apps from India

  1. aigi

    What the London meetup means for Indian builders

    The central lesson from the MLOps Community London March 2026 discussions was straightforward: shipping an LLM feature is easy; operating one reliably is the product. Production teams now need repeatable systems for quality, latency, cost, security, and change management—not another impressive prototype.

    That distinction matters in India. A startup may serve customers in multiple languages, operate with a lean platform team, depend on several model providers, and still need to meet enterprise procurement requirements in Europe or North America. The strongest architecture is therefore not necessarily the one with the largest model. It is the one that makes failures visible, limits spend, and lets engineers improve quality without breaking existing workflows.

    This recap turns the event’s operational themes into an India-ready implementation plan. It also complements the broader lessons in MLOps Community London at OpenAI HQ: Production LLM Lessons for India, where production architecture and deployment discipline are examined in more detail.

    1. Replace demos with evaluation-driven development

    The first operating pattern is to treat evaluations as a release gate. Manual spot-checking can help during discovery, but it cannot protect a customer-facing application after prompts, models, retrieval indexes, or tool calls change.

    Build an evaluation set before optimising the application. Include:

    • Golden examples: questions with reviewed reference answers or acceptable answer criteria.
    • Adversarial cases: prompt injection, ambiguous requests, unsupported claims, and malformed inputs.
    • India-specific coverage: Hinglish, regional names, Indian addresses, rupee values, date formats, legal terms, and code-switched conversations.
    • Business metrics: resolution rate, escalation rate, task completion, conversion, or time saved.
    • Operational metrics: latency, token use, error rate, retries, and provider failures.

    Use LLM judges carefully. They are useful for scale, but their scores should be calibrated against human reviewers and supplemented with deterministic checks. For example, a judge can assess helpfulness while code verifies JSON validity, citation presence, policy fields, or whether a response contains prohibited data.

    Run evaluations in CI whenever a prompt, model, retrieval configuration, or tool schema changes. Store the dataset version, model identifier, prompt version, judge configuration, and score distribution. A single average score hides regressions; slice results by language, customer segment, intent, and failure type.

    2. Design a model-routing layer before costs become a crisis

    A compound AI system uses several specialised components instead of forcing one expensive model to handle every request. This is particularly valuable for Indian startups competing on price while serving global customers.

    A practical routing flow looks like this:

    1. Classify the request. Detect intent, risk, language, required context, and expected output format.
    2. Apply policy first. Block or redirect unsafe, unsupported, or high-risk requests before invoking an expensive model.
    3. Choose the smallest suitable model. Use a hosted or self-managed small model for classification, extraction, rewriting, and simple FAQs.
    4. Escalate selectively. Send difficult reasoning, long-context synthesis, or high-value interactions to a stronger model.
    5. Validate and recover. Check the output schema, retry with a controlled policy, or route to a human when confidence is low.

    Do not judge routing only by per-token price. Track total cost per successful task, including retries, tool calls, embedding generation, storage, and human review. Keep a fallback provider or model, but test fallbacks in advance: a nominally compatible API may produce different formatting, safety behaviour, or tool-call semantics.

    Semantic caching can reduce repeated work, especially for support and documentation workloads. Cache only when the request, permissions, data freshness requirements, and response variability allow it. Never let a cache bypass tenant isolation or return a response generated for a different user.

    3. Move beyond basic RAG

    Retrieval-augmented generation remains useful even as context windows grow. A focused retrieval pipeline is usually cheaper and easier to audit than sending an entire knowledge base to a model.

    A robust RAG system should combine:

    • Hybrid retrieval: keyword search for product IDs, laws, names, and exact phrases alongside vector search for meaning.
    • Metadata filters: tenant, geography, document type, language, access level, and publication date.
    • Reranking: a second-stage model that improves the order of candidate passages.
    • Query transformation: spelling correction, acronym expansion, decomposition, or multilingual rewriting.
    • Citation checks: confirm that the answer is supported by the retrieved evidence.
    • Freshness controls: re-index changed documents and expose document timestamps to the application.

    Agentic or self-correcting retrieval can help with complex questions, but it adds latency and failure modes. Start with a measurable baseline: retrieval recall, answer groundedness, citation accuracy, latency, and cost. Add query rewriting or iterative retrieval only when evaluation data shows a clear gain.

    For teams building multilingual products, evaluate retrieval separately from generation. A fluent answer can still be wrong because the retriever missed an English document for a Hindi query or confused two similarly named Indian entities.

    4. Make observability useful to engineers

    Traditional infrastructure dashboards are not enough for LLM applications. Every request should be traceable across the user input, prompt template, retrieval calls, model calls, tool executions, validation steps, and final response.

    At minimum, capture:

    • prompt and application version, with secrets and personal data redacted;
    • model, provider, region, token counts, latency, and retry information;
    • retrieved document IDs and relevance scores;
    • tool-call arguments, outcomes, and timeouts;
    • safety events, refusals, fallbacks, and human escalations;
    • user feedback linked to the exact trace.

    Create dashboards that answer operational questions: Which customer segment is seeing regressions? Which model produces the most retries? Are Hindi queries slower? Did a new index reduce citation accuracy? Observability should lead to an action, not merely generate more logs.

    A small team can begin with structured traces and a weekly failure review. Classify incidents into retrieval, prompt, model, tool, data, infrastructure, and policy failures. This taxonomy becomes the roadmap for engineering work.

    5. Treat security and compliance as architecture

    Indian teams exporting AI products must design for privacy and governance early. The Digital Personal Data Protection framework, contractual data-residency requirements, and customer-specific security reviews can affect model choice and logging design.

    Use data minimisation by default. Redact or tokenise personal information before sending requests to external providers, define retention periods, encrypt stored traces, and separate customer tenants at the data and cache layers. Maintain an inventory of models, datasets, prompts, tools, and vendors so a security or procurement review does not become a last-minute scramble.

    Guardrails should operate at several points:

    • validate input and detect prompt injection before retrieval;
    • enforce document-level access controls;
    • constrain tool permissions and network access;
    • validate structured output after generation;
    • scan responses for sensitive data and unsafe actions;
    • require human approval for financial, legal, medical, or irreversible actions.

    Governance lessons from AI Salon Trustworthy AI Futures London: Governance Lessons for Indian Founders are relevant here: trust is not a policy document placed beside the product. It is a set of technical controls, ownership rules, and evidence collected during operation.

    6. A 30-day implementation plan

    Week 1: Establish the baseline. Instrument traces, define success metrics, collect 100–300 representative examples, and document current cost and latency.

    Week 2: Build the evaluation loop. Add deterministic checks, human review criteria, judge-model scoring, and CI evaluation for prompt and model changes.

    Week 3: Improve efficiency. Add hybrid retrieval, caching where safe, request classification, and routing between at least two model tiers. Measure cost per successful task rather than cost per call.

    Week 4: Harden production. Add redaction, access controls, rate limits, fallback behaviour, incident runbooks, and approval gates for high-risk actions. Review failures with product, engineering, and customer-facing teams.

    Founders can also compare these practices with the founder workflow lessons in Code with Claude Extended London 2026: Founder Workflows and the product-building guidance in How to Build Claude-Powered Products from India. The tools may differ, but the operating principle is the same: shorten the distance between an observed failure and a tested fix.

    The takeaway for India’s AI ecosystem

    The competitive advantage is shifting from access to a model toward disciplined execution around models. Indian teams can win with smaller systems, stronger retrieval, rigorous evaluations, efficient routing, and clear compliance evidence.

    The most credible 2026 LLM applications will not promise that models never fail. They will show how failures are detected, contained, corrected, and learned from. Build that operating layer early, and a lean Indian engineering team can ship globally without accepting uncontrolled cost or unreliable quality.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.