Semi-supervised learning is useful when a small, trusted dataset sits alongside a much larger collection of unlabeled examples. That pattern is common in Indian AI projects: a team may have carefully annotated medical images, crop images, speech clips, or classroom content, but far more raw data waiting to be used. The data loader is the component that turns those two sources into predictable training batches.
A good loader does more than concatenate datasets. It must preserve labels, expose unlabeled samples safely, apply the right transforms, control the labeled-to-unlabeled ratio, and make each batch reproducible. These details directly affect pseudo-label quality and model stability.
What a semi-supervised data loader should do
A practical loader should provide four capabilities:
- Separate supervision paths: labeled items return an input and ground-truth label; unlabeled items return one or more augmented views and an identifier.
- Controlled batch composition: every training step should receive enough labeled examples to calculate supervised loss, plus enough unlabeled examples to calculate a consistency or pseudo-label loss.
- Independent transforms: weak augmentation can generate a teacher prediction, while strong augmentation tests whether the student learns a stable representation.
- Traceability: sample IDs, source splits, and metadata should remain available for audits, error analysis, and data-quality checks.
This last point matters when the model is deployed in high-stakes settings. Teams working with sensitive domains should treat loader metadata as part of their data veracity infrastructure, not as an afterthought.
Recommended dataset structure
Keep labeled and unlabeled data in separate records rather than placing unlabeled samples in a fake class. A simple manifest might contain:
labeled/image_001.jpg,3
labeled/image_002.jpg,1
unlabeled/image_901.jpg,
unlabeled/image_902.jpg,For production work, use a structured format such as CSV, Parquet, or JSON Lines with fields for path, label, source, language, region, and sample_id. In India, fields such as state, script, dialect, device type, or collection setting can reveal distribution gaps that random shuffling would hide. For language projects, this is especially important when working with low-resource language datasets for AI training in India.
Create validation and test splits from labeled data before training. Do not use pseudo-labels to build the main test set. The test set should contain human-verified labels and remain untouched throughout training.
A robust PyTorch implementation
The following pattern returns two views for each unlabeled item and one transformed view for labeled data:
from PIL import Image
from torch.utils.data import Dataset, DataLoader
class LabeledDataset(Dataset):
def __init__(self, records, transform=None):
self.records = records
self.transform = transform
def __len__(self):
return len(self.records)
def __getitem__(self, index):
path, label, sample_id = self.records[index]
image = Image.open(path).convert("RGB")
if self.transform:
image = self.transform(image)
return {"image": image, "label": label, "id": sample_id}
class UnlabeledDataset(Dataset):
def __init__(self, records, weak_transform, strong_transform):
self.records = records
self.weak_transform = weak_transform
self.strong_transform = strong_transform
def __len__(self):
return len(self.records)
def __getitem__(self, index):
path, sample_id = self.records[index]
image = Image.open(path).convert("RGB")
return {
"weak": self.weak_transform(image),
"strong": self.strong_transform(image),
"id": sample_id,
}A standard DataLoader can then be created for each stream:
labeled_loader = DataLoader(
labeled_dataset, batch_size=16, shuffle=True,
num_workers=4, pin_memory=True, drop_last=True
)
unlabeled_loader = DataLoader(
unlabeled_dataset, batch_size=48, shuffle=True,
num_workers=4, pin_memory=True, drop_last=True
)The 1:3 ratio is only an example. Choose it based on label quality, class coverage, memory limits, and the confidence of pseudo-labels. A custom batch sampler is useful when both streams must be returned in one batch, but two iterators are often easier to debug and tune.
Training logic: supervised plus consistency loss
A common objective is:
total_loss = supervised_loss + lambda_u * unsupervised_lossFor labeled data, calculate cross-entropy or the task-specific supervised loss. For unlabeled data, send the weak view through a teacher or evaluation-mode copy of the model, then retain predictions only when confidence exceeds a threshold. Use those predictions as targets for the strong view.
with torch.no_grad():
teacher_logits = model(unlabeled_batch["weak"])
probabilities = teacher_logits.softmax(dim=-1)
confidence, pseudo_labels = probabilities.max(dim=-1)
mask = confidence.ge(0.90)
student_logits = model(unlabeled_batch["strong"])
unsup_loss = F.cross_entropy(
student_logits[mask], pseudo_labels[mask]
) if mask.any() else student_logits.sum() * 0.0Start with a small unsupervised-loss weight and increase it gradually. A fixed high value can let incorrect pseudo-labels overwhelm the verified labels. Track the proportion of accepted pseudo-labels, confidence distribution, and class distribution—not only total loss.
Sampling, imbalance, and reproducibility
Unlabeled data is not automatically representative. A loader can amplify collection bias if one source, geography, class, or device dominates the stream. Use weighted sampling, source-aware batches, or per-group quotas when necessary. For rare classes, oversample labeled examples carefully, but monitor whether repeated samples cause overfitting.
Use deterministic seeds during experiments and record:
- random seed and worker seed;
- dataset and manifest version;
- transform configuration;
- labeled-to-unlabeled batch ratio;
- confidence threshold and loss weight;
- number of accepted pseudo-labels per class.
In data-sensitive projects, also log rejected or low-confidence samples for review. This can create a targeted annotation queue instead of paying to label the entire corpus.
Common implementation mistakes
- Treating unlabeled data as a normal class: this teaches the model an artificial category rather than using the sample for consistency learning.
- Applying identical transforms to both views: without meaningful perturbation, the consistency objective provides limited additional signal.
- Allowing augmentation to change the label: aggressive crops or transformations can invalidate pseudo-labels, especially for documents and medical imagery.
- Evaluating on pseudo-labeled data: this produces optimistic metrics and hides confirmation bias.
- Ignoring empty masks: a batch may contain no predictions above the confidence threshold; return a zero-valued differentiable loss instead of crashing.
- Loading every image into memory: stream samples from disk or object storage, and tune
num_workers,pin_memory, and prefetching against actual hardware.
Before scaling, validate the pipeline on a small, fully labeled subset. Compare supervised-only training with semi-supervised training under the same test split. If performance does not improve, inspect pseudo-label precision before adding more unlabeled data.
Choosing the right starting point
For a first implementation, use a simple supervised baseline, a two-stream loader, weak/strong augmentation, confidence masking, and a fixed human-labeled validation set. More advanced methods—teacher-student momentum, distribution alignment, FixMatch-style thresholds, or contrastive objectives—should come after the data pipeline is measurable.
The project can also serve as a strong machine learning portfolio project for beginners in India if it includes ablations, data documentation, and error analysis rather than only a training script. For teams fine-tuning foundation models, the same discipline applies to fine-tuning LLMs on custom data: data quality, split integrity, and evaluation design matter more than loader complexity.
Practical checklist
Before training a serious model, confirm that:
- labeled and unlabeled manifests are versioned;
- validation and test labels are human-verified;
- each unlabeled item has a stable ID;
- weak and strong transforms are appropriate for the domain;
- batches contain enough labeled examples;
- pseudo-label acceptance is logged by class and source;
- class and geographic imbalance are measured;
- checkpoints include configuration and data versions;
- the final model is compared with a supervised-only baseline.
A semi supervised learning data loader is ultimately an experiment-control system. When it keeps data provenance visible, balances the two learning signals, and fails safely, unlabeled data becomes a measurable advantage rather than an uncontrolled source of noise.