0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to build a quantized model for multilingual sales assistants

How to Build a Quantized Model for Multilingual Sales Assistants

  1. aigi

    Start with the deployment target

    A quantized multilingual sales assistant should be designed around its operating environment—not treated as a full-precision model compressed at the end. Define the channels, languages, latency target, privacy requirements, and hardware before selecting a model.

    For an India-focused product, the language plan may include English, Hindi, Bengali, Marathi, Tamil, Telugu, Kannada, Malayalam, Gujarati, Punjabi, or Hinglish. Do not assume that one benchmark score represents all of them. Measure performance separately by language, script, code-switching pattern, customer segment, and sales workflow.

    Typical deployment choices include:

    • Cloud inference: easier to scale, but requires careful handling of customer and payment data.
    • Private VPC or on-premises inference: useful for regulated sectors and enterprise buyers.
    • Edge or device inference: reduces latency and connectivity dependence, but imposes stricter memory and compute limits.

    If the assistant will handle voice, quantization is only one part of the system. Speech recognition, language identification, retrieval, dialogue management, and text-to-speech each need latency and quality budgets. The architecture principles in this voice agent architecture and deployment guide are useful even when the final interface is chat.

    Choose the right model and quantization method

    Start with a multilingual instruction-tuned model whose licence permits commercial use and whose tokenizer performs reasonably on your target scripts. Smaller models are often preferable for sales workflows because the task can be constrained with retrieval, tools, and structured outputs. A large model that is rarely available within the latency budget is not a production advantage.

    The main quantization options are:

    • Weight-only quantization: converts model weights to formats such as INT8, INT4, GPTQ, AWQ, or bitsandbytes-supported representations. It is often the quickest route for decoder models.
    • Post-training dynamic quantization: quantizes selected operations at runtime. It is straightforward, but speedups depend on the hardware and operator support.
    • Post-training static quantization: uses calibration data to determine activation scales. It can deliver predictable inference performance when the runtime supports it well.
    • Quantization-aware training: exposes the model to simulated low-precision operations during fine-tuning. Use it when post-training quantization causes unacceptable quality loss.

    Do not select a format based only on model size. Compare end-to-end tokens per second, time to first token, peak RAM or VRAM, batch behaviour, power use, and cost per conversation. Validate the exact runtime—such as llama.cpp, TensorRT-LLM, ONNX Runtime, ExecuTorch, or a hardware vendor stack—because support for operators and multilingual tokenization varies.

    Build a representative multilingual dataset

    Sales data must reflect real customer intent, not merely translated English scripts. Assemble examples for product discovery, qualification, pricing, objections, comparisons, order status, returns, escalation, and handoff to a human agent. Include both typed and spoken-language transcripts if the assistant will support voice.

    For each example, preserve metadata such as language, script, region, channel, product category, and whether the customer code-switches. Include natural variations: spelling errors, Romanised Indic languages, mixed English and local-language phrases, abbreviations, and speech-recognition noise.

    A practical dataset should contain:

    • Supervised conversations with preferred responses and refusal or escalation labels.
    • Tool-use examples showing when to query inventory, pricing, CRM, or order systems.
    • Hard negatives where a plausible but incorrect product or policy response must be rejected.
    • Safety examples covering personal data, payment details, misleading claims, and unsupported discounts.
    • Evaluation-only conversations kept separate from training and prompt development.

    For low-resource Indic languages, data quality and coverage matter more than simply adding translated rows. Use native reviewers, and consult this low-resource Indic NLP builder’s guide when designing annotation and evaluation processes. Machine translation can accelerate bootstrapping, but it should not be the final authority for colloquial language, politeness, or commercial terminology.

    Fine-tune for the workflow, not generic conversation

    Begin with supervised fine-tuning or parameter-efficient fine-tuning such as LoRA or QLoRA. Teach the model the assistant’s response style, language behaviour, tool schemas, and escalation rules. Keep factual product information in a retrieval or business-system layer rather than embedding frequently changing catalogues into the model.

    A robust request path commonly looks like this:

    1. Detect the user’s language and script, while allowing the customer to correct it.
    2. Classify intent and identify entities such as product, location, budget, and order number.
    3. Retrieve current, permission-checked information from approved sources.
    4. Generate a concise answer in the customer’s chosen language.
    5. Validate structured fields, policy constraints, and tool results.
    6. Escalate when confidence is low or the request requires a human decision.

    For larger deployments, an agent or distributed orchestration layer may coordinate retrieval, CRM actions, and human handoff. Read more about building distributed systems with AI agents, but avoid adding agent complexity where a deterministic workflow is sufficient.

    Quantize, then measure quality loss

    Create a full-precision baseline before quantization. Quantize one configuration at a time and compare it with the baseline using the same prompts, retrieval documents, tool responses, decoding settings, and hardware.

    Evaluate at minimum:

    • Task success: correct intent, product recommendation, tool call, and escalation.
    • Factuality: answers agree with approved catalogue, policy, and inventory data.
    • Language quality: meaning preservation, grammar, script handling, and code-switching.
    • Safety: no invented discounts, fabricated availability, privacy violations, or unsafe persuasion.
    • Operational performance: latency, throughput, memory, error rate, and cost per session.

    BLEU alone is a weak measure for conversational sales quality. Combine rubric-based human review, targeted test suites, customer satisfaction, conversion, abandonment, and handoff rates. Report results by language; an aggregate score can hide severe degradation in a low-volume language.

    Use a calibration set that includes short and long conversations, common and rare languages, names, numbers, currency, dates, product codes, and noisy input. Quantization can disproportionately affect rare tokens and exact-value handling, so test prices, phone numbers, addresses, and policy thresholds explicitly.

    Deploy with guardrails and observability

    Before launch, add authentication, rate limits, prompt-injection defences, retrieval access controls, PII redaction, audit logs, and a clear human-handoff path. Never allow the model to invent stock, approve exceptions, or execute irreversible actions without server-side validation.

    Run a shadow or limited pilot first. Track quality by language, model version, quantization format, device, and channel. Log prompts and outputs only under an approved data policy, with sensitive fields masked. Establish rollback criteria—for example, a rise in incorrect tool calls, unsupported claims, or language-specific complaints.

    Quantization is not a one-time optimisation. Re-test after model updates, tokenizer changes, new products, retrieval changes, or runtime upgrades. Maintain a small regression suite that runs in CI and a larger evaluation set before each production release.

    A practical build sequence

    For most teams, the lowest-risk path is:

    • Define target languages, workflows, latency, privacy, and hardware.
    • Build a representative, consented dataset with native-language review.
    • Establish a full-precision baseline and deterministic evaluation harness.
    • Fine-tune a licence-compatible multilingual model using LoRA or QLoRA.
    • Add retrieval, tool validation, safety checks, and human escalation.
    • Benchmark FP16 or BF16, INT8, and INT4 variants on target hardware.
    • Choose the smallest configuration that meets language-specific quality thresholds.
    • Pilot with limited traffic, monitor outcomes, and iterate.

    For products serving India’s diverse users, the goal is not the lowest possible bit width. It is reliable multilingual assistance at a sustainable cost, with measurable quality for every language you promise to support. Builders designing for broader access can also review guidance on building AI apps for the next billion users in India.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.