Factory inspection is not a model-selection exercise alone. The camera, optics, lighting, labelling process, line speed, rejection mechanism, and quality workflow often determine performance more than the neural network. For Indian manufacturers, the right system must also handle dust, variable power and connectivity, mixed suppliers, changing SKUs, and limited defect examples.
The best AI models for factory inspection are therefore the ones that meet a defined inspection objective at an acceptable false-reject rate and cost. A production-ready system should make a reliable decision, explain or localise the defect, and trigger an action on the line.
Start with the inspection task
Define the output before comparing architectures:
- Classification: Is the complete product pass or fail?
- Object detection: Where are missing, misplaced, or damaged components?
- Instance segmentation: What is the pixel-level boundary of each defect or part?
- Anomaly detection: Does this product differ from a verified-good reference?
- Measurement: Is a dimension, gap, angle, or alignment within tolerance?
- Tracking and verification: Did the correct part follow the correct assembly sequence?
A narrow, well-lit inspection often needs a small detector or classifier. A variable surface or a novel defect may require anomaly detection. For teams building their first system, this guide to building computer vision models on GitHub is useful for structuring datasets, experiments, and reproducible deployment code.
Best model families by use case
YOLO and other real-time detectors
YOLO-family models remain strong choices for conveyor inspection, component presence checks, packaging errors, PPE detection, and localisation of visible defects. They provide a practical balance of speed, accuracy, and deployment support on industrial PCs and edge accelerators.
Use a real-time detector when the system must identify several objects in one frame and respond within a fixed cycle time. Select the smallest model that meets recall requirements; a larger model is not automatically better if it causes buffering or missed triggers. Evaluate the complete pipeline, including image capture and PLC communication, rather than reporting model latency alone.
CNN classifiers: ResNet, EfficientNet, and MobileNet
CNN classifiers are effective when the camera presents one normalised part at a time and the decision is simply pass or fail. ResNet is a dependable baseline, EfficientNet offers a useful accuracy-efficiency trade-off, and MobileNet is suitable for constrained edge hardware.
Classification is easy to deploy but can hide the reason for rejection. If operators need to see a scratch, missing screw, or contamination region, pair the classifier with a heatmap cautiously or choose detection or segmentation instead. Class activation maps are helpful for debugging, but they should not be treated as precise defect measurements.
PatchCore, PaDiM, and industrial anomaly detection
Anomaly detection is valuable when good products are abundant but defect samples are rare. PatchCore and PaDiM learn the distribution of normal visual features and flag regions that depart from it. They are useful for machined surfaces, castings, textiles, ceramics, and electronics where new defect types may appear after deployment.
These methods are not a substitute for controlled imaging. Reflections, focus changes, dust on the lens, and product-position shifts can look anomalous. Build a representative good dataset across shifts, machines, lots, and environmental conditions. Set thresholds using a validation set that includes hard-but-acceptable products, not only obvious failures.
Segmentation models: U-Net, Mask R-CNN, and modern mask heads
Choose segmentation when defect area, shape, or severity matters. U-Net is a strong option for pixel-level masks on a fixed region of interest. Mask R-CNN can separate individual parts and defects, while newer detector architectures with segmentation heads may provide better speed for production systems.
Segmentation supports weld-pore estimation, corrosion mapping, sealant coverage, paint defects, burr analysis, and dimensional measurement. It also requires more annotation effort. Before asking operators to draw detailed masks, confirm that a bounding box or line measurement cannot answer the quality question.
Vision Transformers and hybrid architectures
Vision Transformers, including hierarchical designs such as Swin, can capture broader context and are useful for high-resolution inspection or complex assembly relationships. In 2026, hybrid CNN-transformer models and transformer backbones are practical options, but they usually demand more data, memory, and tuning than compact CNNs.
Use them when a local patch is not enough—for example, when assessing a large panel, full assembly, or correlated defects across a component. Benchmark against a strong CNN baseline. A transformer that wins on a research dataset may lose in the factory because of latency, domain shift, or insufficient training diversity.
A practical selection matrix
| Factory requirement | Strong starting point | Why |
|---|---|---|
| Missing or misplaced components | YOLO-family detector | Fast localisation and simple integration |
| Pass/fail for a single aligned part | EfficientNet or ResNet | Efficient supervised classification |
| Rare or previously unseen surface defects | PatchCore or PaDiM | Trains primarily on good samples |
| Defect size and shape | U-Net or segmentation detector | Pixel-level localisation |
| Complex full-assembly context | Swin or hybrid transformer | Wider visual context |
| Tight edge-device limits | MobileNet or compact YOLO | Lower memory and latency |
| Dimensional tolerance | Vision model plus calibrated geometry | More reliable than pixels alone |
Treat this table as a starting hypothesis, not a procurement decision. If the project resembles infrastructure inspection, the lessons from AI-based railway track inspection software in India are relevant: combine visual detection with operating conditions, inspection frequency, and an escalation workflow.
Data and evaluation: the decisions that matter most
Split data by production lot, date, machine, and site, not by randomly distributing adjacent frames. Random frame splits can place near-identical images in both training and test sets and produce misleadingly high scores.
Track metrics that reflect factory economics:
- Recall for critical defects: How many failures are missed?
- Precision and false-reject rate: How many good products are unnecessarily stopped?
- Per-class performance: Rare defects should not disappear inside an overall average.
- Latency and throughput: Can the system keep pace with the fastest operating condition?
- Calibration: Does a confidence score of 0.9 mean roughly the same thing across shifts?
- Drift: Does performance change with new suppliers, tooling, lighting, or product variants?
For critical safety or compliance checks, use a human review path for uncertain predictions. Do not market or design around “100% accuracy”; instead, define an acceptance threshold, audit samples continuously, and quantify the cost of missed and false defects.
Deploying on an Indian factory floor
Control the image before scaling the model
Use fixed mounts, diffuse lighting, suitable lenses, polarising filters for reflective surfaces, and hardware triggers where possible. Enclose the inspection area if ambient light changes across the day. Better imaging frequently delivers a larger improvement than moving from one modern architecture to another.
Choose edge or cloud deliberately
Run inference at the edge when the line needs predictable latency, local operation, or protection for proprietary production images. Industrial PCs, NVIDIA Jetson devices, and accelerator-equipped gateways are common options. Cloud services remain useful for fleet monitoring, central retraining, dashboards, and cross-site analytics.
Quantisation, pruning, batching, and TensorRT-style optimisation can reduce cost, but validate the optimised model again: compression may affect small cracks or low-contrast defects disproportionately. Teams with an existing cloud stack can review how to deploy deep learning models on GKE, while edge-first teams should design offline buffering and safe recovery from network outages.
Integrate with operations
A model is production-ready only when it connects to the camera trigger, PLC, reject actuator, MES, and audit trail. Store the image, model version, prediction, operator action, and line context for disputed cases. Add a review queue so engineers can label new examples and retrain without rebuilding the entire system.
A phased implementation plan
1. Baseline the process: Record defect definitions, cycle time, acceptable variation, and current inspection cost.
2. Pilot one station: Choose a constrained use case with measurable defects and stable imaging.
3. Collect difficult negatives: Include glare, blur, contamination, cosmetic variation, and borderline-good parts.
4. Benchmark two or three families: Compare a compact detector, classifier, and anomaly method where applicable.
5. Run in shadow mode: Make predictions without controlling rejection until false positives and misses are understood.
6. Deploy with safeguards: Add confidence thresholds, manual review, rollback, and daily sampling.
7. Monitor drift: Retrain when tooling, supplier material, camera position, or product design changes.
For video-heavy lines, frame-level accuracy is insufficient; assess duplicate detections, missed events, and temporal stability. A useful reference is this comparison of vision models for video understanding, though factory deployment still requires controlled benchmarking on your own footage.
FAQ
How much data is needed? For a constrained classifier, a few hundred carefully selected examples per class may establish a baseline. Defect diversity matters more than raw image count. Anomaly methods can start with good samples, but those samples must cover normal variation.
Should we choose an open-source model or a managed API? An open-source model usually offers lower recurring latency and better control for on-premise production. A managed API can speed up prototyping, but check data residency, connectivity, predictable pricing, and whether the service supports your resolution and throughput.
Can one model inspect every product variant? Sometimes, but separate models or a first-stage SKU detector are often safer when geometry, materials, or acceptable appearance differ substantially.
What should a startup prove to a manufacturer? Demonstrate recall on critical defects, false rejects on good production, end-to-end cycle time, recovery behaviour, and the process for handling new defects—not only a benchmark score.
AI Grants India supports founders working on industrial AI, computer vision, and automation. Explore AI Grants India for funding and ecosystem opportunities as you move from a pilot to a production inspection system.