0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · how to fine tune a model using non pii indian call center data on hugging face

How to Fine-Tune a Model on Non-PII Indian Call Data

  1. aigi

    What you will build

    Fine-tuning is useful when a general model understands language but does not reliably recognise your support intents, follow your escalation policy, or handle Indian-language and code-mixed conversations. This guide shows how to fine tune a model using non PII Indian call center data on Hugging Face for two common use cases:

    • Intent classification: route a call to billing, delivery, KYC, cancellation, technical support, or escalation.
    • Response generation: produce a draft reply grounded in approved support procedures.

    Choose classification when you need predictable labels and low latency. Use supervised fine-tuning for generation only when you have high-quality agent responses and a separate retrieval or policy layer to control factual answers. Teams building voice systems should also review BPO call automation with voice agents before moving a text model into live calls.

    Treat “non-PII” as a data-governance claim

    Removing names and phone numbers does not automatically make a transcript safe. Indian call-centre data can contain account numbers, addresses, email IDs, order references, health details, financial information, voice characteristics, or combinations of facts that identify a person. Under India’s Digital Personal Data Protection Act, 2023, obligations depend on the nature and processing of personal data; de-identification should therefore be documented, tested, and reviewed rather than assumed.

    Before uploading anything to Hugging Face:

    • Obtain a documented lawful basis, appropriate notices, and internal approval for using recordings or transcripts for model training.
    • Remove direct identifiers, account references, URLs, email addresses, phone numbers, exact addresses, government IDs, and free-text secrets.
    • Replace sensitive spans with typed placeholders such as <CUSTOMER_NAME>, <ORDER_ID>, and <CITY> rather than simply deleting context.
    • Check audio separately: anonymised text does not anonymise a speaker’s voice.
    • Keep the raw dataset in a restricted environment and publish only a synthetic, heavily redacted sample if you need a public repository.
    • Record provenance, transformation steps, retention limits, annotator access, and deletion procedures.

    Run automated detection followed by human review. Regex alone will miss indirect identifiers, Indian language variants, spelling errors, and code-mixed text. A privacy review should sample both training and validation splits to prevent leakage across near-duplicate calls.

    Design a useful dataset

    Start with a narrow task and a clear label policy. For intent classification, use one JSONL record per example:

    {"text":"Mera refund abhi tak nahi aaya","label":"refund_status","language":"hi-en","channel":"phone"}

    For supervised response tuning, use a conversation format that separates the customer, agent, and policy context:

    {"messages":[{"role":"user","content":"Mera connection kaam nahi kar raha"},{"role":"assistant","content":"Main basic checks batata hoon. Agar issue rahe, main technician visit raise karunga."}],"language":"hi-en","intent":"technical_support"}

    Create train, validation, and test sets by customer, case, and conversation, not by randomly splitting individual utterances. Otherwise the same caller, script, or issue can appear in every split and inflate results. Preserve representative variation across Hindi-English code switching, regional vocabulary, accents represented in transcripts, short calls, interruptions, ASR errors, and escalation cases.

    Remove low-value examples such as duplicated scripts, agent placeholders, empty transcripts, and conversations whose correct answer depends on unavailable backend data. Keep difficult examples: ambiguous intents, policy exceptions, abusive language, incomplete information, and requests that must be declined.

    For a deeper checklist on data quality, learning rates, and adapter methods, see best practices for fine-tuning LLMs on custom data.

    Set up Hugging Face securely

    Create a private dataset and model repository, enable organisation access controls, and use a token with the smallest required permissions. Do not commit tokens or raw files to Git. A practical environment is:

    python -m venv .venv
    source .venv/bin/activate
    pip install -U transformers datasets evaluate accelerate peft trl bitsandbytes scikit-learn
    huggingface-cli login

    Use an approved compute region and confirm where uploaded artefacts, logs, checkpoints, and experiment traces are stored. If your organisation cannot permit external processing, run the Hugging Face stack in an approved private environment rather than uploading restricted data to the public Hub.

    Load a private JSONL dataset as follows:

    from datasets import load_dataset
    
    data = load_dataset(
        "json",
        data_files={
            "train": "data/train.jsonl",
            "validation": "data/validation.jsonl",
            "test": "data/test.jsonl",
        },
    )
    print(data)

    Pick the smallest model that meets the task

    For intent classification, an encoder model such as multilingual BERT or XLM-R is often cheaper and easier to evaluate than a generative LLM. For response drafting, select an instruction-tuned multilingual model that supports your target languages and licence requirements. Check Hindi, Tamil, Telugu, Bengali, Marathi, and code-mixed performance rather than relying on an English benchmark.

    Parameter-efficient fine-tuning, especially LoRA or QLoRA, reduces memory use and makes rollback simpler. Keep the base model frozen, train adapters, and version the adapter separately from the base model. This is usually a better first experiment than full fine-tuning on a modest call-centre dataset.

    Example: intent classification with Trainer

    Tokenise consistently and retain the original text outside the model pipeline for audit purposes:

    from transformers import AutoTokenizer, AutoModelForSequenceClassification
    from transformers import TrainingArguments, Trainer
    
    base = "xlm-roberta-base"
    labels = sorted(set(data["train"]["label"]))
    label_to_id = {name: i for i, name in enumerate(labels)}
    
    def encode(batch):
        tokens = tokenizer(batch["text"], truncation=True, max_length=256)
        tokens["labels"] = [label_to_id[x] for x in batch["label"]]
        return tokens
    
    tokenizer = AutoTokenizer.from_pretrained(base)
    encoded = data.map(encode, batched=True)
    model = AutoModelForSequenceClassification.from_pretrained(
        base, num_labels=len(labels), id2label={i: x for i, x in enumerate(labels)},
        label2id=label_to_id,
    )
    
    args = TrainingArguments(
        output_dir="outputs/intent-model",
        eval_strategy="epoch",
        save_strategy="epoch",
        learning_rate=2e-5,
        per_device_train_batch_size=16,
        per_device_eval_batch_size=32,
        num_train_epochs=3,
        weight_decay=0.01,
        load_best_model_at_end=True,
        metric_for_best_model="f1",
        push_to_hub=False,
    )
    
    trainer = Trainer(model=model, args=args,
                      train_dataset=encoded["train"],
                      eval_dataset=encoded["validation"])
    trainer.train()

    For generative fine-tuning, use TRL with a chat template and LoRA, mask padding correctly, and cap sequence length after measuring the distribution of transcript lengths. Do not train the model to invent account status, refunds, or policy exceptions. Teach it to ask for missing information or call a tool instead.

    Evaluate for safety and business value

    Accuracy alone is inadequate. Report macro-F1, per-intent recall, confusion matrices, calibration, and abstention performance. For multilingual deployments, break results down by language, script, code-mixing, ASR quality, region, and call type. For generation, assess factuality, policy adherence, refusal quality, toxicity, unnecessary disclosure, and whether a human agent can use the answer without rewriting it.

    Use a locked test set and compare against a zero-shot or prompt-based baseline. Add adversarial tests for prompt injection, requests for another customer’s information, sensitive-data repetition, abusive calls, and unsupported policy claims. Sample errors with operations and compliance teams; a small number of high-risk false positives may matter more than a several-point average improvement.

    For post-deployment monitoring, track intent drift, abstention rate, escalation rate, latency, token cost, and complaints. AI call transcript analysis for sales teams offers a useful model for turning transcripts into operational signals, though support deployments still need stricter privacy controls.

    Publish and deploy responsibly

    Keep the repository private unless the dataset and model artefacts are genuinely shareable. The model card should state the base model, training data scope, languages, preprocessing, known limitations, evaluation slices, intended use, prohibited use, and contact for deletion or incident requests. Never publish raw transcripts, memorised examples, or logs containing customer text.

    For production, place the model behind authentication, rate limits, audit logging, encryption, and a policy or retrieval layer. Require human review for refunds, account changes, regulated advice, complaints, and vulnerable-customer cases. A voice agent should expose an immediate transfer path, disclose automation where required by your policy, and preserve only the minimum information needed for service.

    Practical launch checklist

    • Define one measurable task and an escalation policy.
    • Obtain approval and document lawful use, retention, and access.
    • De-identify text and audio, then validate the process with sampling.
    • Split by case or customer to prevent leakage.
    • Establish multilingual and code-mixed evaluation slices.
    • Start with LoRA or a small encoder model before increasing scale.
    • Compare against a baseline and test unsafe failure modes.
    • Keep Hub repositories private and publish a complete model card.
    • Monitor drift, privacy incidents, latency, cost, and human override rates.

    For teams selecting the surrounding product stack, compare these model capabilities with top-rated voice agent services for Indian businesses and define exactly where a fine-tuned model adds value over prompting, retrieval, or rules.

    Last updated 23 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.