Fine-tuning support-ticket data can improve routing, priority prediction, language detection, and suggested replies—but only when the dataset, objective, and privacy controls are designed carefully. This guide explains how to fine-tune a model using anonymized Indian support tickets on Hugging Face, with a classification workflow that can be adapted to multilingual customer-support operations.
The examples assume a ticket-classification task such as assigning categories including billing, delivery, account access, fraud, refund, or technical support. If your goal is response generation, use a smaller, carefully curated instruction dataset and apply stricter human review. For broader guidance on training custom models, see best practices for fine-tuning LLMs on custom data.
Define the task before choosing a model
Start with a narrow operational outcome. “Understand support tickets” is not a measurable training objective; “assign one of eight queue labels with at least 90% macro-F1” is.
Decide whether you need:
- Single-label classification: one primary queue or intent per ticket.
- Multi-label classification: several attributes, such as billing plus account access.
- Priority prediction: urgent, standard, or low priority, ideally with clear annotation rules.
- Language or script detection: useful for English, Hindi, Hinglish, and regional-language routing.
- Entity extraction: order IDs, product names, dates, or locations, with a separate privacy review.
Do not train a generative model when a classifier will solve the problem more cheaply and predictably. For Indian operations, preserve meaningful language variation—Hinglish, transliterated Hindi, Tamil-English, abbreviations, and spelling variation—but remove irrelevant personal information.
Prepare and anonymize the dataset
A ticket dataset should include the text, target label, and optional metadata that is genuinely available at prediction time. A simple JSONL record might look like this:
{"text":"Mera refund abhi tak nahi aaya", "label":"refund"}Anonymization is more than replacing email addresses. Build a repeatable pipeline that detects and removes or masks:
- Names, phone numbers, email addresses, postal addresses, and account numbers.
- Aadhaar, PAN, passport, bank, card, UPI, and transaction identifiers.
- Order IDs, policy numbers, vehicle registrations, and internal case references.
- Precise locations, dates of birth, health information, and free-text secrets.
- Agent notes or escalation comments that would not be present during inference.
Use consistent placeholders such as <PHONE> and <ORDER_ID> rather than deleting every entity. This lets the model learn useful ticket structure without memorising real values. Keep the raw dataset in a restricted system, separate from the training repository, and document who can access it, why it is processed, and how long it is retained. Obtain the required organisational approvals and apply India’s privacy and security requirements to the full data lifecycle; anonymization should not be treated as a substitute for governance.
Before uploading anything to the Hugging Face Hub, inspect samples manually and run automated scans for residual secrets. Never put production tickets, access tokens, or private dataset URLs in a public repository.
Split and label the data properly
Create train, validation, and test sets before training. A common starting point is 80/10/10, but the split must prevent leakage. Tickets from the same conversation, customer, order, or incident should remain in one partition. If near-duplicate tickets appear across splits, test performance will be misleading.
Review label quality before increasing model size. Write a short annotation guide with positive examples, difficult boundaries, escalation rules, and an “unknown” or “other” category. Measure agreement between annotators on a sample. If labels are heavily imbalanced, report per-class precision, recall, and F1 rather than accuracy alone. Maintain a test slice for Hindi, Hinglish, regional languages, transliteration, short messages, and code-mixed text.
Set up Hugging Face
Create an isolated environment and install current compatible packages:
python -m venv .venv
source .venv/bin/activate
pip install -U transformers datasets evaluate accelerate scikit-learnLoad a local CSV or JSONL file and split it explicitly:
from datasets import load_dataset
data = load_dataset("json", data_files={"train": "train.jsonl", "validation": "validation.jsonl", "test": "test.jsonl"})
print(data)Choose a multilingual encoder when tickets contain multiple Indian languages or scripts. A compact multilingual model is often a better first experiment than a large language model because classification is faster, cheaper, and easier to evaluate. Compare at least one multilingual checkpoint with a model suited to your dominant language mix, and record the model card, licence, tokenizer, and training-data constraints before using it commercially.
Tokenize and fine-tune a classifier
Map string labels to integer IDs and tokenize with a maximum length chosen from your ticket distribution. Avoid blindly padding every example to an unnecessarily large value.
from transformers import AutoTokenizer
model_name = "xlm-roberta-base"
tokenizer = AutoTokenizer.from_pretrained(model_name)
labels = sorted(set(data["train"]["label"]))
label2id = {label: i for i, label in enumerate(labels)}
id2label = {i: label for label, i in label2id.items()}
def encode(batch):
encoded = tokenizer(batch["text"], truncation=True, max_length=256)
encoded["labels"] = [label2id[x] for x in batch["label"]]
return encoded
tokenized = data.map(encode, batched=True, remove_columns=data["train"].column_names)Train with the modern eval_strategy argument and save the best checkpoint:
import evaluate
import numpy as np
from transformers import AutoModelForSequenceClassification, TrainingArguments, Trainer
metric = evaluate.load("f1")
model = AutoModelForSequenceClassification.from_pretrained(
model_name, num_labels=len(labels), label2id=label2id, id2label=id2label
)
def compute_metrics(pred):
predictions = np.argmax(pred.predictions, axis=-1)
return metric.compute(predictions=predictions, references=pred.label_ids, average="macro")
args = TrainingArguments(
output_dir="support-ticket-model",
eval_strategy="epoch",
save_strategy="epoch",
load_best_model_at_end=True,
metric_for_best_model="f1",
learning_rate=2e-5,
per_device_train_batch_size=16,
per_device_eval_batch_size=32,
num_train_epochs=3,
weight_decay=0.01,
report_to="none",
)
trainer = Trainer(
model=model,
args=args,
train_dataset=tokenized["train"],
eval_dataset=tokenized["validation"],
processing_class=tokenizer,
compute_metrics=compute_metrics,
)
trainer.train()If GPU memory is limited, reduce batch size and use gradient accumulation. For larger models, parameter-efficient methods such as LoRA can reduce compute and make iteration easier, but they still require the same privacy and evaluation discipline.
Evaluate beyond one score
Run the final test set only after decisions are complete. Report macro-F1, per-class recall, confusion matrices, calibration, and abstention behaviour. A model that confidently routes fraud tickets to a general queue is more dangerous than one that marks uncertain cases for review.
Evaluate separate slices for language, script, channel, ticket length, geography where appropriate, and new versus familiar issue types. Have support agents review false positives and false negatives. Check whether the model relies on shortcuts such as queue names, agent signatures, or leaked IDs. Test adversarial inputs containing mixed scripts, emojis, spelling variation, and deliberately incomplete descriptions.
For production, define a confidence threshold and a human fallback. Log model version, input schema, prediction, confidence, and reviewer outcome without retaining unnecessary personal data. Monitor drift as products, policies, and customer language change.
Publish and deploy safely
Save the model and tokenizer locally, then publish only an approved artefact:
trainer.save_model("support-ticket-model")
tokenizer.save_pretrained("support-ticket-model")A private Hugging Face repository can support controlled collaboration; a public repository should contain synthetic or fully approved data, a clear model card, licence information, intended use, limitations, evaluation slices, and known failure modes. Use access tokens through environment variables or secret managers, not notebooks.
A practical deployment pattern is to place the classifier behind an API that validates input length, strips unexpected fields, applies the same anonymization pipeline, and returns both label and confidence. Keep automated actions limited until the model demonstrates stable performance in a shadow or assisted-routing phase. If support is moving toward voice, compare the resulting workflow with voice agent versus IVR for customer support, especially where transcription errors affect downstream classification.
Common mistakes to avoid
- Training on raw tickets because anonymization was postponed.
- Randomly splitting messages from the same case across train and test sets.
- Using accuracy on an imbalanced dataset.
- Treating Hinglish and transliteration as noise rather than a test requirement.
- Selecting a model without checking its licence and language coverage.
- Uploading private data or secrets to the Hub.
- Deploying without confidence thresholds, audit logs, and human escalation.
Fine-tuning is one component of a dependable support system. Pair it with retrieval, workflow rules, monitoring, and agent feedback where those controls are more appropriate than additional model training. For multilingual claims operations, the same design principles apply to automated multilingual health insurance claims support.