Fine-tuning data can expose more than the model developer intended. Names, phone numbers, email addresses, Aadhaar-linked details, medical records, customer IDs, and private conversations may be memorised during training and reproduced later. Cleaning the dataset before it reaches a tokenizer or training job is therefore a privacy control—not just a formatting step.
This guide explains how to remove PII before fine-tuning a model on Hugging Face. It covers detection, redaction, pseudonymisation, validation, dataset governance, and the limits of automated tools. For the wider training workflow, pair this process with best practices for fine-tuning LLMs on custom data.
Define what counts as PII
Start with a project-specific data inventory. PII is broader than obvious identifiers and varies by context. A combination of harmless-looking fields can also identify someone.
Common examples include:
- Names, usernames, email addresses, phone numbers, and postal addresses
- Aadhaar numbers, PAN numbers, passport details, driving-licence numbers, and voter IDs
- Bank accounts, UPI handles, card numbers, and customer or policy identifiers
- Health records, prescriptions, diagnoses, voice recordings, and biometric information
- IP addresses, precise location, device IDs, and account metadata
- Free-text references to employers, schools, family members, or rare events
- Images, scanned documents, and screenshots containing readable personal details
For Indian datasets, define whether regional-language names, transliterated addresses, code-mixed text, and local identity-document formats need special handling. Record the decision in a data card before cleaning begins.
Choose removal, masking, or pseudonymisation
The safest default for general-purpose fine-tuning is remove the sensitive span entirely. This reduces the chance that the model learns an identifier and avoids creating a reversible mapping.
Use other treatments only when the information is needed for the task:
- Redaction: Replace content with labels such as
[EMAIL],[PHONE], or[PERSON]. - Generalisation: Convert exact ages to ranges or precise locations to district or state level.
- Pseudonymisation: Replace identities with stable synthetic tokens such as
CUSTOMER_001. Store the mapping separately, encrypt it, and restrict access. - Tokenisation or encryption: Appropriate for controlled operational systems, but not a substitute for removing secrets from a training corpus. A model can still memorise encrypted strings or learn relationships around them.
- Synthetic replacement: Generate realistic but non-identifying values, then test that they do not accidentally match real records.
Do not rely on hashing alone. Low-entropy values, predictable identifiers, and linkage with other datasets can make hashed data reversible or identifying.
Build a Hugging Face cleaning pipeline
Keep raw data outside the training repository and produce a versioned, cleaned derivative. Do not upload the raw dataset to a public Hugging Face Hub repository, issue tracker, notebook, or model card.
A simple datasets pipeline can apply a redaction function:
from datasets import load_dataset
import re
email_re = re.compile(r"\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b")
phone_re = re.compile(r"(?<!\d)(?:\+91[-\s]?)?[6-9]\d{9}(?!\d)")
def redact_pii(example):
text = example["text"]
text = email_re.sub("[EMAIL]", text)
text = phone_re.sub("[PHONE]", text)
return {"text": text}
raw = load_dataset("json", data_files="input.jsonl", split="train")
clean = raw.map(redact_pii, desc="Redacting obvious identifiers")
clean.save_to_disk("dataset_clean")Regular expressions are useful for structured values, but they will miss context-dependent PII, spelling variations, OCR errors, and Indian-language text. Add a named-entity recognition or commercial privacy detector for names, addresses, organisations, and free-text references. Treat detector output as candidates for review, not proof of anonymisation.
Run cleaning before tokenisation, and apply the same transformation to every relevant field: prompts, completions, chat messages, metadata, filenames, captions, and evaluation data. Inspect nested JSON structures rather than cleaning only a visible text column.
Handle multilingual and unstructured data
PII detection often fails on code-mixed and regional-language content. Names may appear in Devanagari, Tamil, Bengali, Kannada, or transliteration; phone numbers may contain spaces or local punctuation; addresses may be incomplete but still identifying.
Improve coverage by:
- Normalising Unicode while preserving a traceable original copy in restricted storage
- Testing detectors on English, Hindi, and the languages represented in the corpus
- Including local identifier formats and common transliteration variants
- Reviewing samples containing OCR text, voice transcripts, tables, and chat exports
- Running image redaction or OCR-specific checks for scanned documents
If your project involves regional-language fine-tuning, review the dataset design alongside resources such as fine-tuning Llama for Indian regional languages. Language coverage should include privacy coverage.
Validate that PII is actually gone
A successful cleaning job needs measurable checks. Create a held-out audit sample and run multiple passes after redaction:
- Scan for emails, phone numbers, identity-number patterns, URLs, account IDs, and secrets
- Run a second detector from a different method or vendor
- Search for known test records and synthetic canaries
- Sample redacted and untouched rows for human review
- Check columns, metadata, filenames, and document boundaries for leakage
- Compare entity counts before and after cleaning
- Confirm that placeholders do not preserve excessive context around a person
Use a small canary set containing deliberately inserted fake identifiers. The pipeline should remove every canary before the dataset is approved. Keep false-positive rates visible: over-redaction can reduce training quality, while under-redaction can create a privacy incident.
Also test the trained model. Prompt it with canary strings, partial identifiers, and completion requests designed to elicit memorised text. This does not prove privacy, but it can reveal failures before deployment. Model evaluation should be part of the same release gate as accuracy and safety testing.
Govern access and document decisions
Apply least-privilege access to raw data, intermediate files, logs, caches, and experiment trackers. Disable accidental logging of examples, encrypt storage, set retention periods, and remove raw files from temporary training environments after verification.
Publish a dataset card that records:
- Source, collection purpose, consent or legal basis, and geographic scope
- PII categories considered and detection methods used
- Redaction rules, known blind spots, and human-review procedure
- Dataset versions, hashes, access restrictions, and retention policy
- Contact and takedown process for privacy complaints
Do not claim that a dataset is “fully anonymised” solely because a recogniser found no entities. Anonymisation is an evidence-based risk assessment, and re-identification risk depends on the surrounding data and intended release.
Before launching fine-tuning
Use this release checklist:
- [ ] Raw data is isolated from the training and Hub repositories
- [ ] PII definitions cover Indian identifiers, multilingual text, images, and metadata
- [ ] Redaction runs before tokenisation and covers all fields
- [ ] A second detector and human sample review found no critical leakage
- [ ] Synthetic canaries are removed successfully
- [ ] Clean data, code, and reports are versioned
- [ ] Access, retention, and deletion controls are documented
- [ ] The model has been tested for memorisation before release
For privacy-sensitive workloads, consider private infrastructure and restricted model repositories rather than public sharing. If the final model must run on constrained devices, privacy testing should happen before compression and deployment; the AI model optimisation guide for mobile devices covers the engineering trade-offs that follow.
FAQ
Can I fine-tune directly on a private Hugging Face dataset?
Private access reduces exposure but does not remove the risk of memorisation, insider access, misconfiguration, or later publication. Clean the data first.
Is replacing names with `[NAME]` enough?
Not always. Addresses, dates, rare events, and linked fields can identify a person even after names are removed.
Should I use differential privacy?
Differential privacy can reduce memorisation risk, but it does not replace source-data minimisation or PII detection. It also affects utility and training configuration.
What if PII is required for the task?
Use the least specific representation possible, restrict access, pseudonymise consistently when necessary, and document why retention is justified.
Who should approve the dataset?
Assign a data owner and involve privacy, security, legal, or ethics reviewers according to the data’s sensitivity and the organisation’s obligations.