Start with the decision, not the algorithm
Learning how to train reinforcement learning models on historical data from the West Bengal manufacturing sector begins with defining the operational decision the model will support. Reinforcement learning (RL) is not simply a forecasting method. It learns a policy: what action to take in a particular state, subject to costs, constraints, and delayed outcomes.
Useful starting problems include:
- Production sequencing across machines, shifts, and product grades.
- Preventive-maintenance timing based on vibration, temperature, load, and failure history.
- Inventory replenishment when supplier lead times and demand vary.
- Energy scheduling for boilers, compressors, furnaces, and refrigeration systems.
- Workforce or job allocation under skill, safety, and shift constraints.
Choose one decision with a measurable outcome. A narrow pilot—such as reducing unplanned downtime on a production line—is easier to validate than a broad objective such as “optimise the factory”.
What historical data can and cannot do
Historical records can support offline reinforcement learning, where an agent learns from previously collected transitions rather than experimenting immediately on a live plant. A transition normally contains:
- State: machine condition, active order, inventory, shift, weather, energy tariff, and other information available at decision time.
- Action: the intervention taken, such as changing a schedule, setting a machine parameter, or ordering material.
- Reward: the resulting operational value, including output, quality, energy use, delay, maintenance cost, or safety penalties.
- Next state: the condition observed after the action.
- Termination flag: whether the episode ended because of a completed batch, breakdown, shift boundary, or other event.
The dataset must also record what was actually possible at each point. If a maintenance action was unavailable because a technician or spare part was missing, the model should not treat it as a valid historical alternative.
West Bengal’s industrial mix—jute and textiles, engineering, foundries, food processing, chemicals, and small and medium-sized manufacturing—creates substantial variation in equipment, data maturity, and operating practice. Do not pool records across plants without checking whether sensors, units, product definitions, and control policies are comparable.
Before modelling, establish data lineage and trust. The principles in data veracity infrastructure for high-stakes AI are especially relevant when maintenance, quality, or safety decisions depend on imperfect logs.
Build a decision-ready dataset
A practical preparation workflow is:
1. Create a canonical event timeline. Convert PLC events, MES records, ERP transactions, maintenance tickets, laboratory results, and operator logs to one timezone and one clock. Resolve duplicate events and clock drift.
2. Define the observation window. Specify how often the policy receives a state—every minute, batch, job, or shift. Avoid using information that became available only after the decision.
3. Join actions to outcomes. Link each scheduling or maintenance action to its downstream production, quality, cost, and downtime effects. Keep action timestamps and approval delays.
4. Handle missingness explicitly. Distinguish a zero measurement from an unrecorded measurement. Add data-quality indicators rather than silently imputing every gap.
5. Represent constraints. Encode machine capacity, minimum batch sizes, changeover rules, maintenance windows, labour availability, safety limits, and contractual delivery commitments.
6. Split by time. Train on earlier periods, validate on later periods, and reserve the most recent period for a final holdout. Random row splits leak future conditions into training.
For a first build, a tabular transition store in Parquet or a relational database is usually sufficient. Track dataset versions, feature definitions, source systems, and excluded records so results can be reproduced.
Select an offline RL approach carefully
Begin with a supervised baseline before RL. Compare against the current operator policy, a rules-based scheduler, linear optimisation, and a demand or failure predictor. If RL cannot beat a credible baseline in an offline test, it is not ready for deployment.
The algorithm should match the action space:
- Discrete actions: fitted Q-iteration or conservative Q-learning can suit choices such as selecting the next job or maintenance option.
- Continuous actions: offline actor-critic methods can handle set points, but require careful action-bound enforcement.
- Large or structured schedules: use a hierarchical policy, mixed-integer optimisation, or a learned ranking model rather than forcing the entire schedule into one monolithic agent.
- Limited historical coverage: prefer conservative methods that avoid assigning high value to actions rarely seen in the data.
A digital simulator or learned dynamics model is valuable, but it is not automatically accurate. Calibrate it against held-out episodes and report uncertainty. A model that performs well only inside a simulator trained on the same historical assumptions can create false confidence.
Design rewards that reflect factory economics
Reward design determines what the policy optimises. A useful reward may combine contribution margin, completed units, on-time delivery, scrap, energy, maintenance expenditure, and downtime. Use explicit penalties for safety violations, quality failures, missed delivery commitments, and infeasible schedules.
Keep the components visible rather than hiding everything in one unexplained score. For example:
reward = margin - energy_cost - scrap_cost - delay_penalty - downtime_cost
Review the scale and frequency of each term. A large but rare penalty may be ignored if frequent small rewards dominate learning. Reward the actual business outcome, not an easy proxy such as machine utilisation alone; maximising utilisation can increase wear, defects, or unsafe operating conditions.
Evaluate without pretending historical data shows counterfactuals
Offline evaluation cannot directly reveal what would have happened if the plant had taken a different action. Use several safeguards:
- Policy support checks: measure how often proposed actions resemble actions in the training data.
- Off-policy evaluation: apply importance sampling, doubly robust estimators, or model-based evaluation with confidence intervals.
- Temporal holdouts: test on later seasons, product mixes, supplier conditions, and maintenance regimes.
- Stress scenarios: simulate demand spikes, raw-material delays, sensor outages, equipment degradation, and power interruptions.
- Operational metrics: report throughput, on-time completion, scrap, energy per unit, downtime, maintenance cost, safety violations, and worst-case outcomes—not just cumulative reward.
- Human review: ask operators whether proposed actions are feasible, understandable, and consistent with known process behaviour.
Compare every result with the existing policy. A small average improvement may not justify deployment if it introduces unacceptable tail risk or unstable behaviour during unusual conditions.
Move from notebook to controlled pilot
Deploy in stages. First run the policy in shadow mode, where it observes live data and recommends actions without controlling equipment. Then use a human approval workflow for low-risk decisions such as job prioritisation or replenishment suggestions. Keep hard safety interlocks and deterministic constraint checks outside the model.
Log the state, recommendation, operator override, executed action, and resulting outcome. Monitor distribution shift, data freshness, action-support coverage, reward components, and override rates. Establish rollback criteria before the pilot starts. If the policy recommends actions outside its training distribution, default to the existing rule or escalate to an operator.
For teams building their first prototype, small reproducible projects such as those described in machine learning portfolio projects for beginners in India can help establish version control, evaluation, and deployment discipline before tackling plant-scale RL.
Governance for Indian manufacturing deployments
Manufacturing data may contain commercially sensitive production volumes, supplier information, worker identifiers, and proprietary process parameters. Apply role-based access, encryption, retention limits, and pseudonymisation. Separate worker performance monitoring from optimisation objectives unless there is a clear, lawful, and documented purpose.
Maintain a model card covering training periods, plants, excluded conditions, known failure modes, action limits, and human accountability. Document who can approve a recommendation and who can stop the system. For regulated or safety-critical use cases, involve process-safety, quality, IT security, and plant leadership from the beginning.
A practical 90-day pilot plan
- Weeks 1–2: select one decision, baseline current performance, and define safety and business constraints.
- Weeks 3–5: build the event timeline, audit data quality, and create an offline transition dataset.
- Weeks 6–8: train supervised and conservative offline-RL baselines; run temporal and scenario tests.
- Weeks 9–10: review recommendations with operators and validate feasibility in a simulator or replay environment.
- Weeks 11–12: launch shadow mode, monitor drift and overrides, and decide whether a human-approved pilot is justified.
The objective is not to place an autonomous agent on a factory floor quickly. It is to demonstrate a measurable improvement while preserving safety, traceability, and operator control.