0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · reducing llm token usage with intent layers

Reducing LLM Token Usage with Intent Layers

  1. aigi

    LLM costs rarely come from one oversized request. They accumulate through repeated system instructions, irrelevant retrieval results, long chat histories, duplicate questions, and workflows that send every task to the most capable model. An intent layer addresses this waste before it reaches the expensive inference step.

    For Indian startups, the pattern is especially useful when products serve multilingual users, operate on usage-sensitive budgets, or combine customer support, search, agents, and automation in one application. The goal is not to shorten every prompt blindly. It is to make each model call carry only the information and capability required for the user’s actual objective.

    What an intent layer does

    An intent layer is a small classification and routing service between the application interface and one or more downstream models. It converts an open-ended request into structured decisions such as:

    • Intent: refund status, order tracking, account recovery, document extraction, or general information.
    • Language: English, Hindi, Tamil, Bengali, Hinglish, or another supported language.
    • Complexity: deterministic, retrieval-based, tool-using, or reasoning-intensive.
    • Risk: routine, sensitive, financial, medical, employment-related, or requiring human review.
    • Context needs: no history, a short summary, selected records, or the full conversation.

    A useful output might look like this:

    {
      "intent": "order_status",
      "language": "hi",
      "route": "workflow_api",
      "confidence": 0.96,
      "needs_history": false,
      "risk": "low"
    }

    The application can then call an order-status API instead of an LLM, use a compact retrieval prompt, or escalate to a stronger model only when necessary. Teams designing the classifier should also study intent extraction in short text and test recognition across spelling variation, code-switching, and incomplete queries.

    Where the token savings come from

    1. Route deterministic requests away from the LLM

    Many requests do not need generative reasoning. “Track order 4812,” “download my invoice,” and “change my phone number” can be handled by authenticated APIs and fixed response templates. The intent layer identifies the operation, validates required fields, and invokes the correct workflow.

    This saves both input and output tokens while improving reliability. It also reduces the chance that a model invents a status that should come from a source-of-truth system.

    2. Replace one universal prompt with modular prompts

    A large system prompt often contains instructions for every department, policy, tool, and exception. Most of that text is irrelevant to an individual request. Create small prompt modules for specific routes instead:

    • Billing policy and invoice fields for finance queries.
    • Product documentation and diagnostic steps for technical support.
    • Eligibility rules and escalation procedures for claims.
    • A concise answer format for frequently asked questions.

    Keep shared instructions genuinely shared—such as safety, output schema, and citation requirements—and load route-specific guidance only when selected. This is more maintainable than continually expanding a “god prompt.”

    3. Retrieve less, but retrieve better

    Intent classification should influence retrieval. A refund-status request may require one transaction record, not ten pages of policy documentation. A technical issue may need a product manual section, while a general question needs only a short knowledge-base passage.

    Use metadata filters, structured databases, and small top-k values before adding more context. Measure answer quality after reducing retrieval volume; fewer tokens are valuable only if groundedness and task completion remain stable. For agent systems, a governance or knowledge layer can prevent low-confidence retrieval from being presented as fact; see building honestly calibrated knowledge layers for AI.

    4. Cache repeated work safely

    Use a layered cache rather than sending every request to a model:

    • Exact-match cache: useful for identical FAQs and fixed policy answers.
    • Normalized cache: handles differences in case, punctuation, and whitespace.
    • Semantic cache: matches meaning, but only within a defined intent and tenant.
    • Workflow cache: stores stable tool results for a short, explicit time-to-live.

    Do not return a cached answer merely because the embedding similarity is high. Include intent, product version, language, user permissions, and freshness requirements in the cache key. Avoid semantic caching for account-specific, financial, medical, or rapidly changing information unless the underlying data is validated again. A separate guide to reducing repetitive responses in LLM applications covers these failure modes in more detail.

    5. Compact conversation context without deleting facts

    Long histories are expensive and can reduce answer quality. Maintain three levels of memory:

    • Working context: the current request, relevant recent turns, and active tool results.
    • Structured memory: stable facts such as customer ID, preferences, or case number.
    • Conversation summary: decisions, unresolved issues, constraints, and commitments.

    Summaries should be generated against a fixed schema and refreshed when important facts change. Never summarise away permissions, quantities, dates, negations, or unresolved actions. For regulated or high-stakes workflows, retain the original transcript for audit while sending only the minimum necessary context to the model.

    A production architecture

    A practical request path is:

    1. Normalize: detect language, clean obvious noise, and preserve identifiers exactly.
    2. Classify: predict intent, entities, risk, and complexity using a local model, rules, embeddings, or a small hosted model.
    3. Validate: check confidence, required fields, authentication, and policy constraints.
    4. Route: choose an API, cache, retrieval template, small model, large model, or human queue.
    5. Assemble context: add only the records and instructions required for that route.
    6. Generate or execute: call the selected model or business workflow.
    7. Verify: validate JSON, citations, tool results, and policy-sensitive claims.
    8. Log: record token counts, route, latency, cache result, confidence, and outcome.

    For multi-intent messages, split only when the sub-requests are independent. Otherwise, preserve the original request and ask a targeted clarification question. Read how to improve intent recognition in conversational AI for approaches to ambiguity, fallback handling, and multi-label classification.

    Indian-language and multilingual considerations

    Token counts vary substantially by language, script, tokenizer, and model. Do not assume that translating every Indian-language query into English will always reduce cost or improve accuracy. Translation adds calls, can lose names and legal phrasing, and may create errors in Hinglish or regional spellings.

    Benchmark representative traffic in each supported language. Compare native processing with translation-assisted processing on:

    • Input and output tokens.
    • End-to-end latency.
    • Intent accuracy.
    • Entity preservation for names, addresses, and IDs.
    • User-rated answer quality.
    • Escalation and correction rates.

    Keep identifiers and quoted text untouched, and use language-specific examples in classifier evaluation. A fallback should preserve the user’s language even when the downstream reasoning path uses a different internal representation.

    Measuring savings without creating hidden costs

    Track cost per completed task, not only cost per request. At minimum, monitor:

    • Input, output, and cached tokens by route.
    • Percentage of requests resolved without a large model.
    • Classifier latency and confidence distribution.
    • Cache hit rate and invalidation errors.
    • Retrieval context size.
    • Fallback, correction, and human-escalation rates.
    • Task success and user satisfaction.

    Run an offline evaluation set covering common, ambiguous, adversarial, and multilingual requests. Then use shadow routing or a controlled rollout. Savings claims such as “50% cheaper” are meaningful only when quality, latency, translation, embedding, storage, and classifier costs are included. For SaaS teams, reducing AI token costs for SaaS startups provides a useful companion framework for unit economics.

    Common mistakes to avoid

    • Over-compressing user input: remove noise, not meaning or constraints.
    • Using confidence as truth: calibrate scores and test difficult examples.
    • Caching across tenants: isolate permissions, data, and policy versions.
    • Routing by keyword alone: combine rules with examples and structured entities.
    • Ignoring prompt-injection risk: treat retrieved text and user content as untrusted data.
    • Adding a classifier to every path: skip it when a deterministic application route is already obvious.
    • Optimising tokens before reliability: a cheap wrong answer costs more through retries, support tickets, and lost trust.

    Implementation checklist

    • Define 10–20 high-value intents before expanding the taxonomy.
    • Map each intent to the cheapest safe execution path.
    • Build confidence thresholds and an explicit fallback route.
    • Keep prompts modular and retrieval filters route-aware.
    • Add exact, normalized, and carefully scoped semantic caches.
    • Preserve structured facts and audit logs outside the prompt.
    • Evaluate native-language and translation-assisted routes separately.
    • Monitor cost per successful task alongside quality and latency.

    Intent layers work best as a control plane for model usage, not as a cosmetic prompt-shortening step. They help Indian builders make routing, context, safety, and cost decisions explicitly—so expensive models are reserved for the requests that genuinely need them.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.