A production LLM application should not send every request to the same model. A short FAQ, a multilingual support reply, a 100-page document analysis task, and a tool-using coding agent have different latency, context, reasoning, and reliability requirements. Treating them identically increases cost and creates avoidable delays.
LLM query routing adds a decision layer between your application and model providers. That layer selects a model, endpoint, region, or execution strategy using signals such as request complexity, latency budget, context length, user tier, data sensitivity, and current provider health.
For Indian products, routing is especially useful when mobile connectivity varies, users expect support in multiple languages, and rupee-denominated unit economics must withstand high request volumes. The goal is not to choose the most powerful model every time. It is to meet a defined quality target with the lowest practical latency and cost.
Start with a routing policy, not a model list
Before comparing providers, define the service levels your application must meet. A useful policy has four parts:
- Quality requirement: What failure rate is acceptable for each task?
- Latency budget: Set separate targets for time to first token (TTFT) and total response time.
- Cost ceiling: Establish a maximum cost per request or per successful task.
- Data constraints: Decide whether prompts may leave India, use third-party APIs, or be processed by shared infrastructure.
For example, a voice support assistant might require TTFT below 500 milliseconds and tolerate a concise answer, while a compliance review can accept a longer response if citations and extraction accuracy are stronger. If your product depends on real-time interaction, study the architectural patterns in this low-latency conversational AI guide for Indian businesses.
Classify query complexity with observable signals
A router does not need to understand a prompt perfectly. It needs enough information to make a reliable routing decision. Combine deterministic rules with a small classifier rather than relying entirely on another expensive LLM call.
Useful complexity signals
- Intent: FAQ, translation, extraction, summarisation, coding, planning, or open-ended reasoning.
- Context size: Input tokens, attached files, retrieved passages, and conversation history.
- Constraint count: Required fields, formatting rules, citations, tone, or policy conditions.
- Tool requirements: Browsing, database queries, calculators, code execution, or external actions.
- Risk level: Medical, financial, legal, safety, or customer-impacting decisions need stronger verification.
- Output shape: Short text is generally easier than strict JSON, tables, or structured records with validation rules.
A practical first version can assign scores from zero to three for reasoning, context, tools, and risk. Sum the scores and map them to tiers:
- Tier A: Direct answers, classification, translation, and short extraction.
- Tier B: Summarisation, grounded question answering, moderate transformation, and routine coding.
- Tier C: Multi-step reasoning, long-context synthesis, complex code, planning, and high-risk review.
Do not infer complexity only from prompt length. A long retrieved document may support a simple lookup, while a short request can require difficult reasoning.
Build latency-aware model tiers
Each tier should have at least one primary model and one fallback. Record performance by task type rather than trusting a provider-wide average. A fast model for short English prompts may perform poorly on long Hindi documents or structured extraction.
A typical policy looks like this:
- Fast tier: Small or distilled model for low-risk, short requests.
- Balanced tier: General-purpose model for most production traffic.
- Reasoning tier: Stronger model for difficult, ambiguous, or high-impact requests.
- Private tier: Self-hosted or India-region deployment for sensitive workloads.
Latency has multiple components: queue time, network round trip, TTFT, token generation speed, tool execution, and retries. Track each separately. Streaming can improve perceived responsiveness, but it does not fix a slow first token or an overloaded provider. Teams building their own gateway can use this guide to building a low-latency LLM API for connection pooling, streaming, and concurrency patterns.
Use a staged routing architecture
A reliable router should make cheap decisions first and reserve expensive checks for uncertain cases.
Stage 1: Cache and deterministic rules
Check an exact or semantic cache before calling a model. Cache only responses that are stable, permission-safe, and appropriate for reuse. Include tenant, language, retrieval version, and policy version in the cache key. Never allow a cached response from one customer or user role to leak to another.
Rules can immediately route requests based on file type, language, maximum context, required tools, or compliance flags. They are fast, explainable, and easy to test.
Stage 2: Lightweight classification
A small classifier can predict intent, complexity, language, and risk. It may be a conventional text classifier, a compact open model, or a constrained LLM call. Keep its output structured and versioned. If the classifier is uncertain, route upward rather than pretending to know.
Stage 3: Model selection
Choose among eligible models using a score such as:
score = quality_weight * expected_quality
- latency_weight * predicted_latency
- cost_weight * estimated_cost
- risk_penalty * failure_riskThe weights should vary by product surface. A voice assistant should weight latency heavily; a back-office research workflow can weight quality and cost more strongly.
Stage 4: Validation and escalation
Validate output format, required fields, citations, safety conditions, and task-specific checks. If validation fails, retry with a corrected prompt or escalate to a stronger model. A generic confidence score is not enough: confidence should be calibrated against actual success on representative traffic.
Design fallbacks without creating retry storms
Provider failure and model failure are different. A timeout may justify another endpoint, while an invalid answer may require a stronger model or a revised prompt. Use bounded retries, exponential backoff with jitter, circuit breakers, and an overall deadline.
Pass a request budget through the entire chain. For example, if the user-facing deadline is two seconds, the router cannot spend 700 milliseconds classifying, wait 800 milliseconds for a provider, and then start an unrestricted fallback. Cancel work that can no longer meet the deadline and return a safe, useful degraded response where appropriate.
For high-priority workloads, maintain provider diversity, but normalise prompts and response schemas. Different APIs vary in tool-calling semantics, token accounting, safety behaviour, and streaming formats. An adapter layer prevents provider-specific details from spreading through your application.
Measure routing quality in production
A router is an optimisation system, so monitor both decisions and outcomes. At minimum, capture:
- TTFT, total latency, queue time, and timeout rate by model and route.
- Cost per request, successful task, user, and workflow.
- Escalation, fallback, retry, and cache-hit rates.
- Task accuracy, schema validity, groundedness, and human correction rate.
- Performance by language, geography, device, customer segment, and prompt size.
Create a labelled evaluation set from real Indian usage patterns, including English, Hindi, Hinglish, and other target languages if relevant. Replay the set against routing policies before changing thresholds. Use shadow traffic to evaluate a new policy without changing user-visible responses, then launch gradually with a rollback switch.
Common mistakes to avoid
- Using a large router model: Its cost and latency can erase the benefit of routing.
- Routing only by token count: Context length is a signal, not a complete complexity measure.
- Ignoring conversation state: Switching models may require carrying history, tool results, and system instructions safely.
- Optimising averages: Tail latency, especially p95 and p99, determines user experience during traffic spikes.
- Treating confidence as truth: Validate confidence scores against labelled outcomes and calibrate them.
- Hard-coding provider names: Route by capabilities and service levels so providers can change without a rewrite.
If the application runs on devices or near users, edge execution can reduce network delay and improve resilience. Compare that approach with patterns in this guide to low-latency AI agents on edge devices, particularly for offline or intermittently connected environments.
A practical rollout plan for Indian startups
Start with two or three models and three task tiers. Instrument every request before adding machine-learning-based routing. Establish a baseline for quality, p95 latency, and cost. Then introduce deterministic rules, followed by a small classifier and validation-based escalation.
Keep sensitive data handling explicit. Review provider retention policies, encryption, access controls, and data residency requirements before routing customer or regulated information to external APIs. For workloads with strict locality requirements, compare self-hosted inference with managed endpoints rather than assuming either option is automatically cheaper.
The strongest routing system is not the most elaborate one. It is a measurable control plane that makes predictable decisions, fails safely, and improves through real traffic. Revisit thresholds as models, provider prices, context windows, and user behaviour change through 2026.