First, clarify what “Hugging Face MCP” means
The phrase Hugging Face MCP is often used imprecisely. Hugging Face’s core fine-tuning stack is built around Transformers, Datasets, Accelerate, PEFT, TRL, and the Hub. MCP usually refers to the Model Context Protocol, a way for an AI application or coding agent to call tools and retrieve context—not to a standard Hugging Face “Model Customization Pipeline”.
That distinction matters. MCP can help an agent select datasets, inspect model cards, launch jobs, or evaluate checkpoints, while PEFT performs the parameter-efficient training. The guide below shows a reliable Python workflow and explains where an MCP server can sit around it. For broader dataset, hyperparameter, and evaluation guidance, see these best practices for fine-tuning LLMs on custom data.
What you need before training
A successful run begins with a clear task definition rather than a large model. Decide whether you are doing:
- Causal language modelling or instruction tuning for a generative model.
- Sequence classification for labels such as intent, sentiment, or risk.
- Token classification for named-entity recognition.
- Preference optimisation using a library such as TRL after supervised fine-tuning.
Prepare a training, validation, and test split. Remove duplicates, redact personal information, and ensure that examples in the test set cannot be reconstructed from training data. For Indian deployments, retain language, script, region, and domain metadata so you can measure performance separately for Hindi, Marathi, Tamil, Bengali, or transliterated text instead of reporting one blended score. Projects involving regional language adaptation can also review fine-tuning Llama for Indian regional languages.
For a first experiment, use a small instruct model that fits your available GPU. Quantisation plus LoRA can make a 7B-class model practical on a strong single-GPU workstation, but memory requirements still depend on sequence length, batch size, optimiser, and checkpointing.
Install the training stack
Create an isolated environment and install versions that work together:
python -m venv .venv
source .venv/bin/activate
pip install -U torch transformers datasets accelerate peft trl bitsandbytes evaluateRun accelerate config once, then authenticate only when you need private Hub assets or want to push a model:
huggingface-cli loginKeep secrets out of notebooks and MCP configuration files. Use environment variables or a managed secret store, and give automation the narrowest permissions required.
Load a model and dataset
For supervised instruction tuning, your dataset should contain a consistent prompt and response structure. A conversational dataset commonly uses a messages column:
from datasets import load_dataset
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "Qwen/Qwen2.5-3B-Instruct"
dataset = load_dataset("json", data_files={
"train": "data/train.jsonl",
"validation": "data/validation.jsonl",
})
tokenizer = AutoTokenizer.from_pretrained(model_id, use_fast=True)
if tokenizer.pad_token is None:
tokenizer.pad_token = tokenizer.eos_token
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype="auto",
device_map="auto",
)Use the model’s chat template rather than manually concatenating role labels. This prevents training and inference formats from drifting:
def format_example(example):
text = tokenizer.apply_chat_template(
example["messages"],
tokenize=False,
add_generation_prompt=False,
)
return {"text": text}
dataset = dataset.map(format_example)Inspect several rendered examples before training. Check for missing roles, unexpected HTML, excessively long records, and answers that contain instructions intended for the model rather than the user.
Add LoRA with PEFT
PEFT freezes the base model and trains a small adapter. LoRA is usually the best starting point because it reduces memory use, keeps experiments cheap, and produces a portable adapter that can be merged later.
from peft import LoraConfig
peft_config = LoraConfig(
r=16,
lora_alpha=32,
lora_dropout=0.05,
bias="none",
task_type="CAUSAL_LM",
target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
)Target module names differ across architectures. Inspect model.named_modules() or consult the model documentation before copying a configuration. If you use 4-bit QLoRA, load the model with a BitsAndBytesConfig, then call prepare_model_for_kbit_training(model) before attaching the adapter. Confirm that only adapter parameters are trainable:
from peft import get_peft_model
model = get_peft_model(model, peft_config)
model.print_trainable_parameters()If nearly every parameter is trainable, stop and fix the configuration before spending compute.
Train with TRL’s SFTTrainer
TRL simplifies supervised fine-tuning for causal language models. The exact argument names can change between releases, so pin versions for reproducibility and check the installed documentation.
from trl import SFTTrainer, SFTConfig
training_args = SFTConfig(
output_dir="outputs/adapter",
num_train_epochs=2,
per_device_train_batch_size=2,
gradient_accumulation_steps=8,
learning_rate=2e-4,
logging_steps=10,
eval_strategy="steps",
eval_steps=100,
save_steps=100,
save_total_limit=2,
gradient_checkpointing=True,
bf16=True,
max_seq_length=2048,
report_to="none",
)
trainer = SFTTrainer(
model=model,
args=training_args,
train_dataset=dataset["train"],
eval_dataset=dataset["validation"],
processing_class=tokenizer,
peft_config=peft_config,
dataset_text_field="text",
)
trainer.train()
trainer.save_model("outputs/adapter")Start with a short run and verify loss, sample generations, GPU memory, and checkpoint size. Training loss alone is not evidence of usefulness. Compare the tuned model with the original model on a held-out, task-specific test set.
Evaluate for quality, safety, and Indian deployment conditions
Measure task metrics such as exact match, F1, accuracy, or ROUGE where appropriate, but add human review for factuality, instruction following, refusal behaviour, and language quality. Test code-switching, spelling variation, transliteration, low-resource scripts, and domain terminology. For a Sanskrit or translation project, a specialised evaluation design is more informative than generic perplexity; see fine-tuning large language models for Sanskrit translation.
Create a small regression suite containing difficult and safety-sensitive prompts. Re-run it after every data or adapter change. Watch for:
- Memorisation of phone numbers, Aadhaar-like identifiers, or proprietary text.
- Hallucinated legal, medical, financial, or government guidance.
- Uneven performance across scripts, dialects, genders, and regions.
- Prompt injection through retrieved documents or tool responses.
Where MCP fits
An MCP server can expose controlled tools such as list_datasets, inspect_model_card, start_training_job, get_metrics, or publish_adapter. The MCP client—perhaps an internal developer assistant—can then coordinate the workflow using natural language. Keep training execution behind an allowlist and require approval for costly jobs, Hub publication, dataset deletion, or access to sensitive files.
Do not let an MCP tool accept arbitrary shell commands. Validate model IDs, dataset paths, GPU limits, output locations, and network access. Log tool calls, arguments, identity, duration, and resulting artefacts. MCP should orchestrate a tested training script, not replace reproducible configuration.
Save, publish, and deploy the adapter
Adapters are much smaller than full model checkpoints. Save the adapter, tokenizer, base-model revision, dataset version, PEFT configuration, training arguments, evaluation results, and software environment together. Push only after checking licensing and data governance:
model.push_to_hub("your-org/domain-assistant-lora")
tokenizer.push_to_hub("your-org/domain-assistant-lora")At inference time, load the base model and adapter, or merge them when your serving stack benefits from a single checkpoint. Keep the unmerged adapter when you need to serve several domain variants against one base model. Before choosing a hosting route, compare GPU, CPU, quantised, and private-cloud options in this guide to platforms for hosting custom fine-tuned models.
Common failures and practical fixes
- Out-of-memory errors: reduce sequence length, use gradient accumulation, enable checkpointing, quantise the base model, or choose a smaller model.
- Training loss falls but outputs worsen: inspect data quality, lower the learning rate, reduce epochs, and strengthen the validation set.
- Adapter has no visible effect: check target modules, trainable parameter counts, chat templates, and adapter loading paths.
- Poor regional-language results: balance scripts and dialects, retain native-script examples, and evaluate each language separately.
- Reproducibility problems: pin package versions, record the base-model commit, seed runs, version datasets, and store the complete configuration.
For teams constrained to Indian lab or edge hardware, the trade-offs covered in fine-tuning large language models on local hardware are especially relevant.
Final checklist
Before calling the project production-ready, confirm that you have a clean data card, a documented base model and licence, held-out evaluations, safety regression tests, cost and latency measurements, rollback capability, and a monitoring plan. PEFT makes iteration affordable, but it does not remove the need for disciplined data engineering and evaluation. Used with a tightly scoped MCP layer, it can give Indian builders a practical path from a reproducible experiment to a maintainable domain model.