Indian factory safety information is difficult training data: it spans scanned forms, multilingual instructions, legal language, tables, incident narratives, and site-specific operating procedures. A useful model must do more than recognise keywords. It should identify hazards, retrieve the right clause, classify incidents consistently, and signal uncertainty when the source is incomplete.
This guide explains how to fine-tune a model using Indian factory safety documents on Hugging Face in a way that is practical for builders and defensible for safety teams. It focuses on classification and information extraction first, while showing when retrieval-augmented generation may be safer than fine-tuning.
Start with a clearly bounded task
Do not begin by uploading every document you can find. Define one measurable task and its users:
- Document classification: distinguish safety manuals, permits, inspection checklists, incident reports, standard operating procedures, and statutory notices.
- Hazard extraction: identify equipment, hazardous substances, unsafe actions, controls, severity, and responsible roles.
- Clause retrieval: return the relevant paragraph or checklist item for a question from a supervisor or auditor.
- Incident triage: classify reports by hazard type, location, severity, and whether immediate escalation is required.
- Structured extraction: convert inspection forms into fields such as date, machine, finding, corrective action, owner, and due date.
For regulatory answers that must remain traceable to current rules, consider retrieval before fine-tuning. Fine-tuning changes model behaviour; it does not reliably keep a model current when a regulation, state rule, or plant procedure changes. A hybrid system can retrieve approved source passages and use a tuned model to classify or structure them.
Assemble and govern the source data
Potential sources include public guidance, internal safety manuals, anonymised incident reports, maintenance procedures, permit-to-work records, and inspection checklists. Before training, create a source register containing:
- Document owner and permission to use it
- Factory, industry, state, and language context
- Publication or revision date
- Confidentiality and personally identifiable information status
- Applicable law, standard, or internal policy
- OCR quality and page-level provenance
Remove names, phone numbers, employee IDs, medical details, signatures, and commercially sensitive information unless there is a documented reason to retain them. Keep the original files in restricted storage and publish only a sanitised dataset. Do not treat publicly accessible documents as automatically reusable; check copyright, contractual restrictions, and organisational approvals.
Indian factories may use English, Hindi, Tamil, Marathi, Telugu, Bengali, or mixed-language instructions. Preserve the original language in a separate field and record whether text came from OCR, transcription, or manual entry. For multilingual deployments, assess an appropriate multilingual encoder rather than translating everything into English and losing operational terminology.
Design labels that safety experts can apply consistently
A small, reliable dataset is more valuable than a large set of ambiguous labels. Write a labelling guide with definitions, positive examples, negative examples, and escalation rules. For example, distinguish missing machine guarding, bypass of an interlock, and inadequate lockout/tagout rather than placing all three under “machine hazard”.
Use two trained reviewers for an initial sample. Measure agreement, discuss disagreements with a qualified safety professional, and revise the guide before scaling annotation. Retain a rationale or evidence span for extraction labels. This makes errors reviewable and supports audits.
Split data by site, document, and time, not by random pages. Pages from the same manual in both training and test sets create leakage and inflate scores. A better test set contains documents from plants, equipment types, or revision periods that were not seen during training.
Prepare a Hugging Face dataset
A simple tabular classification dataset might contain text, label, source_id, language, revision_date, and evidence_span. Keep metadata available for evaluation, but avoid passing sensitive or irrelevant fields to the model.
from datasets import load_dataset
from transformers import AutoTokenizer
raw = load_dataset("csv", data_files={
"train": "train.csv",
"validation": "validation.csv",
"test": "test.csv"
})
model_id = "distilbert-base-multilingual-cased"
tokenizer = AutoTokenizer.from_pretrained(model_id)
def tokenize(batch):
return tokenizer(
batch["text"],
truncation=True,
padding="max_length",
max_length=512
)
tokenized = raw.map(tokenize, batched=True)For long manuals, do not silently truncate the most important content. Chunk by headings or logical sections, retain page and section identifiers, and aggregate predictions at document level. For scanned material, inspect OCR output manually. Tables, warning symbols, units, and numbered steps often need specialised parsing.
If you need broader guidance on data quality, learning rates, and checkpoint selection, use these best practices for fine-tuning LLMs on custom data. Developers building a multilingual pipeline can also compare approaches in open-source vision-language models for Indian languages, especially when documents contain images, diagrams, or low-quality scans.
Choose the smallest model that meets the requirement
For document classification, start with a compact encoder such as multilingual DistilBERT or another suitable multilingual transformer. For extraction, consider token classification or a sequence-to-sequence model. If the task requires summarisation or question answering, first test retrieval with a general instruction model and grounded citations.
Fine-tuning a large language model is not automatically better. It raises GPU, evaluation, latency, and governance costs. Parameter-efficient methods such as LoRA or other PEFT techniques can reduce memory use and make experiments easier to reproduce. Pin model, dataset, and library versions; store the training configuration and random seeds.
A representative training setup is:
from transformers import TrainingArguments
args = TrainingArguments(
output_dir="./safety-classifier",
eval_strategy="epoch",
save_strategy="epoch",
learning_rate=2e-5,
per_device_train_batch_size=8,
num_train_epochs=3,
weight_decay=0.01,
load_best_model_at_end=True,
metric_for_best_model="f1_macro",
report_to="none"
)Treat these as starting points, not universal settings. Tune batch size, maximum length, learning rate, class weighting, and number of epochs against a held-out validation set. Rare but high-consequence labels may require oversampling or threshold tuning, but never hide poor recall behind an aggregate accuracy score.
Evaluate for safety, not just accuracy
Report macro-F1, per-class precision and recall, confusion matrices, and calibration. For hazard detection, false negatives may matter more than false positives. Set an escalation threshold with safety stakeholders and route uncertain cases to a human reviewer.
Evaluate across meaningful slices:
- Language and script
- OCR quality and document format
- Factory or industry type
- New versus old document revisions
- Common versus rare hazards
- Short forms versus long manuals
Review actual error examples. A model that confuses “recommended” with “mandatory”, drops a unit, or reverses a negation can create operational risk despite a good headline score. Test prompt injection and malicious text in uploaded documents if the system uses an LLM. The model should never invent a legal requirement or approve a hazardous action without an authoritative source and human sign-off.
Publish and deploy responsibly on Hugging Face
Use a private or organisation-controlled Hugging Face repository for proprietary material. The model card should state the intended use, excluded uses, training sources, languages, known limitations, evaluation splits, licences, and review requirements. Do not publish weights that enable reconstruction of confidential incident details.
In production, log the model version, source document revision, retrieved evidence, confidence, and human decision. Provide a feedback workflow for safety officers to correct labels and quarantine problematic documents. Re-train only after reviewing drift: regulatory updates, new machinery, changed terminology, and shifts in reporting behaviour can all reduce performance.
For teams building a wider open-source workflow, Indian open-source AI developer projects offers useful context on collaboration and model stewardship. If your safety product also analyses equipment photographs or warning signs, pair text models with the workflow described in how to build computer vision models on GitHub.
A practical launch checklist
Before a pilot, confirm that you have:
- A single defined task and approved label guide
- Permissioned, anonymised, versioned documents
- Site-separated train, validation, and test sets
- OCR and multilingual quality checks
- A baseline model and a retrieval-only comparison
- Per-class metrics, slice analysis, and human error review
- Escalation rules for low confidence and high-severity findings
- Model-card, access-control, logging, and rollback procedures
Fine-tuning can make factory safety systems more consistent, but it cannot replace statutory interpretation, competent safety professionals, or site-level controls. Use Hugging Face to build a traceable assistant that helps people find and structure evidence—not one that quietly turns uncertain predictions into safety decisions.