Indian call centres rarely receive clean, single-language queries. A customer may switch between Hindi and English, use transliterated Hindi in Latin script, mention a product name incorrectly, or describe a billing issue without using the word “bill”. A useful intent classifier must learn these patterns rather than merely memorise English examples.
This guide explains how to use Hugging Face tools—often described loosely as the Hugging Face MCP workflow—to fine-tune a text-classification model for Indian call-centre intents. MCP is not a separate fine-tuning algorithm or a “Model Card Platform” that trains models. In practice, use Hugging Face Hub for model and dataset management, transformers for modelling, datasets for data pipelines, and an MCP-compatible assistant or tool server only if you want to automate those operations.
Define the intent taxonomy before training
Start with a stable set of intents that maps to an operational action. Avoid labels such as “other” becoming a dumping ground. For a telecom, fintech, ecommerce, or utility support desk, an initial taxonomy might include:
- Bill amount or bill dispute
- Failed payment or refund pending
- Account verification and KYC
- Plan, product, or service information
- Service outage or technical troubleshooting
- Cancellation, replacement, or upgrade
- Delivery status or address change
- Complaint escalation
- Fraud or unauthorised transaction
- Agent transfer or callback request
Document the difference between similar labels. “Payment failed” and “refund pending” should have separate definitions if they trigger different workflows. Include an escalation policy for safety-sensitive categories such as fraud, health-related services, or financial distress.
If the final system will power a voice bot, design the classifier alongside your BPO call automation workflow. Intent classification is only one component; speech recognition errors, interruptions, authentication, and human hand-off can determine whether the overall system succeeds.
Build representative Indian training data
Use real or realistically simulated utterances, with personally identifiable information removed. Each record should contain a customer utterance and one canonical label:
{"text":"Mera UPI payment deduct ho gaya, refund kab milega?","intent":"refund_pending"}
{"text":"My recharge failed but money was debited","intent":"payment_failed"}
{"text":"नेटवर्क नहीं आ रहा है, क्या समस्या है?","intent":"technical_outage"}Cover the variation your production system will see:
- English, Hindi, and relevant regional languages
- Code-mixed speech, such as Hinglish and Tanglish
- Romanised Indian-language text, including inconsistent spelling
- Short queries, incomplete sentences, and repeated words from transcripts
- Multiple accents and transcription artefacts from automatic speech recognition
- Different customer journeys, not just different phrasings of one example
Keep speakers, households, and support cases separated across training, validation, and test sets. Otherwise, near-duplicate calls can inflate accuracy. A useful starting point is several hundred labelled examples per intent, followed by active labelling of the categories where the model is uncertain or frequently wrong.
For recorded calls, establish consent, retention, access controls, and an anonymisation process. Replace phone numbers, account IDs, addresses, names, and payment details before data reaches a shared repository. Maintain a dataset card describing source, language coverage, labelling rules, known gaps, and permitted use.
Choose a multilingual base model
A multilingual encoder is usually a better starting point than bert-base-uncased for Indian call-centre data. Evaluate models such as multilingual BERT, XLM-R, or an Indian-language-focused encoder against your own validation set. The best choice depends on languages, latency, licence terms, sequence length, and available GPU capacity—not on benchmark reputation alone.
For practical selection, compare:
- Macro-F1 across all intents, not only overall accuracy
- Performance by language and script
- Confusion between operationally similar intents
- CPU latency and memory use at your expected traffic
- Licence and data-governance requirements
Review the best practices for fine-tuning LLMs on custom data, but remember that intent classification often needs a compact encoder rather than a large generative model. A smaller classifier can be cheaper, faster, and easier to audit.
Prepare the Hugging Face project
Install a current environment and pin versions for reproducibility:
python -m venv .venv
source .venv/bin/activate
pip install -U torch transformers datasets evaluate accelerate scikit-learnCreate a CSV with text and intent columns, then load and encode the labels:
from datasets import load_dataset, ClassLabel
raw = load_dataset("csv", data_files={
"train": "train.csv",
"validation": "validation.csv",
"test": "test.csv",
})
labels = sorted(set(raw["train"]["intent"]))
label_to_id = {name: i for i, name in enumerate(labels)}
id_to_label = {i: name for name, i in label_to_id.items()}
def add_label(row):
row["label"] = label_to_id[row["intent"]]
return row
raw = raw.map(add_label)Select a multilingual checkpoint and tokenize consistently:
from transformers import AutoTokenizer
checkpoint = "xlm-roberta-base"
tokenizer = AutoTokenizer.from_pretrained(checkpoint)
def tokenize(batch):
return tokenizer(batch["text"], truncation=True, max_length=128)
tokenized = raw.map(tokenize, batched=True)Do not blindly pad every example to a large fixed length. Dynamic padding generally reduces compute, while a measured max_length prevents unusually long transcripts from dominating inference.
Fine-tune and measure the right metrics
Load a sequence-classification head and train with class-aware evaluation:
import numpy as np
import evaluate
from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer
model = AutoModelForSequenceClassification.from_pretrained(
checkpoint,
num_labels=len(labels),
id2label=id_to_label,
label2id=label_to_id,
)
f1 = evaluate.load("f1")
def compute_metrics(eval_pred):
logits, y_true = eval_pred
y_pred = np.argmax(logits, axis=-1)
return {"macro_f1": f1.compute(
predictions=y_pred, references=y_true, average="macro"
)["f1"]}
args = TrainingArguments(
output_dir="./intent-model",
eval_strategy="epoch",
save_strategy="epoch",
learning_rate=2e-5,
per_device_train_batch_size=16,
per_device_eval_batch_size=32,
num_train_epochs=3,
load_best_model_at_end=True,
metric_for_best_model="macro_f1",
report_to="none",
)
trainer = Trainer(
model=model,
args=args,
train_dataset=tokenized["train"],
eval_dataset=tokenized["validation"],
processing_class=tokenizer,
compute_metrics=compute_metrics,
)
trainer.train()For imbalanced datasets, inspect per-class precision, recall, and F1. A high overall score can hide poor performance on fraud, cancellation, or regional-language categories. Add a confusion matrix and review false positives with operations staff. Set a confidence threshold: uncertain cases should trigger clarification or human transfer rather than an incorrect automated action.
Test production conditions, not just clean text
Create fixed test slices for language, script, code-mixing, noisy transcripts, location, channel, and intent. Track macro-F1, recall for high-risk intents, abstention rate, and hand-off rate. Test temporal drift after new plans, products, policies, or campaigns launch.
A useful release gate might require minimum recall for fraud and complaint escalation, no material regression in Hindi or regional-language slices, and an agreed latency ceiling. Keep the test set locked and version every label-policy change.
Call transcripts also support downstream quality work. Once consent and governance are in place, AI call transcript analysis for sales teams can help identify recurring objections and uncover new intents for labelling.
Publish and deploy safely
Push the model, tokenizer, label mapping, evaluation results, and dataset documentation to a private Hugging Face repository or your approved registry:
trainer.push_to_hub("indian-call-centre-intent-classifier")At inference time, return the predicted intent, confidence, model version, and an audit identifier—never sensitive transcript content by default. Add rate limits, access controls, encryption, monitoring, and a rollback path. Log representative errors with privacy safeguards, then retrain from reviewed examples rather than automatically feeding every production prediction into the dataset.
If the classifier feeds a voice agent, pair it with a clear transfer policy and human review. Explore top-rated voice agent services for Indian businesses only after confirming their language support, data residency, integration options, and escalation controls.
Practical checklist
- Define action-oriented intents and written labelling rules.
- Remove PII and document consent, provenance, and retention.
- Include code-mixed, transliterated, multilingual, and noisy ASR examples.
- Split data by speaker or case to prevent leakage.
- Compare multilingual encoders on your own slices.
- Report macro-F1, per-intent recall, confusion, latency, and abstention.
- Version the model, labels, dataset, and evaluation report.
- Launch with confidence thresholds, monitoring, and human hand-off.
Fine-tuning is the easy part. The durable advantage comes from a well-defined taxonomy, representative Indian data, disciplined evaluation, and an operational path for uncertain predictions.