Start with the investment problem, not the algorithm
Reinforcement learning (RL) can help test sequential portfolio decisions, but it is not a shortcut to reliable returns. For Odisha-linked steel and power companies, the useful question is not “which model predicts tomorrow’s price?” It is: how should a portfolio adjust exposure as commodity prices, electricity demand, operating costs, regulation, and market liquidity change?
This distinction matters. A model trained on clean historical prices can look impressive while ignoring brokerage, taxes, slippage, corporate actions, position limits, and the risk of learning from information that was unavailable at the time. Treat RL as a research and decision-support system, not an autonomous investment adviser. This article is educational, not financial advice.
Define the Odisha exposure carefully
“Odisha-based stocks” can mean a company headquartered in the state, operating major assets there, or gaining material revenue from Odisha’s steel, mining, logistics, and power ecosystem. Create a documented universe before collecting data. Depending on your research objective, it may include:
- Listed steel producers with Odisha plants or mineral linkages.
- Power generators, transmission businesses, and distribution-linked companies exposed to eastern India.
- Mining, pellets, ports, rail, and engineering firms whose earnings are tied to the regional industrial cycle.
- A broad market benchmark and sector benchmarks for comparison.
Avoid treating all companies as interchangeable. An integrated steel producer reacts differently to iron-ore prices, coking-coal costs, capacity utilisation, export demand, and domestic spreads than a regulated transmission company. Build company-level features and sector labels so the agent can learn these differences.
For a sound foundation, review the company’s exchange filings, annual reports, investor presentations, production disclosures, debt maturity profile, and environmental or regulatory updates. News sentiment can be useful, but it should supplement—rather than replace—auditable operating data.
Build a leakage-resistant data set
A practical data set should combine market, macroeconomic, operating, and event information. Useful inputs include:
- Adjusted OHLCV prices, returns, turnover, volatility, and corporate-action history.
- Nifty and sector-index returns, interest rates, inflation, INR movement, and market breadth.
- Domestic and international steel prices, iron ore, coking coal, aluminium, crude oil, and power-market indicators.
- Company revenue, EBITDA margin, debt, interest coverage, capacity, production, sales volume, and cash flow.
- Electricity demand, monsoon conditions, fuel availability, freight costs, and relevant policy announcements.
- Earnings dates, dividend events, exchange notices, credit-rating changes, and major project updates.
Use timestamps that reflect when information became public. If a quarterly result was released after market close, the agent must not use it for that day’s action. This is one of the most common sources of artificial performance.
Start with daily data. Intraday RL requires reliable tick data, a precise execution simulator, and much stronger controls for slippage and liquidity. For beginners building a research prototype, the machine learning portfolio projects for beginners in India offers a useful way to structure feature engineering, evaluation, and reproducible experiments.
Design the trading environment
Represent the environment as a portfolio simulator rather than a price chart. At each decision point, the agent observes a state, selects an action, receives a reward, and moves to the next time step.
State space
Include only information available at the decision time. A state might contain:
- Recent returns, rolling volatility, momentum, volume, and drawdown.
- Commodity spreads and changes in power or fuel indicators.
- Company fundamentals, valuation ratios, leverage, and earnings surprises.
- Current cash, holdings, portfolio weights, and remaining risk budget.
- Market regime variables such as trend, volatility, liquidity, and benchmark drawdown.
Normalise features using training-period statistics only. Refit transformations on a rolling basis if the strategy is intended to operate online.
Action space
For a small portfolio, actions can be target weights—for example, a vector assigning exposure to selected stocks, cash, and a benchmark. This is generally more realistic than “buy, sell, or hold” because it makes position sizing explicit. Add constraints such as:
- Maximum weight per company.
- Maximum combined steel or power exposure.
- Minimum cash allocation.
- Turnover and liquidity limits.
- No shorting or leverage unless the simulator supports the applicable Indian-market rules.
Reward function
Do not reward gross price gains alone. A more credible daily reward is risk-adjusted portfolio return after transaction costs, with penalties for excessive turnover, concentration, drawdown, and constraint violations. For example:
reward = net_return - cost_penalty - drawdown_penalty - concentration_penalty
Keep the reward interpretable. If penalties dominate, the agent may simply remain in cash; if they are too weak, it may overtrade. Test reward components separately and document every assumption.
Choose a simple baseline before deep RL
Begin with buy-and-hold, equal-weight, momentum, moving-average, and supervised return-prediction baselines. If an RL agent cannot beat a passive or rules-based strategy after costs and risk adjustment, adding a larger neural network is unlikely to solve the problem.
For discrete actions, tabular Q-learning can clarify the environment, though it scales poorly. DQN may handle larger discrete state representations, while policy-gradient or actor–critic methods suit continuous target weights. Libraries such as Stable-Baselines3 can accelerate experimentation, but the simulator and data pipeline deserve more scrutiny than the algorithm name.
Researchers who want to understand the implementation fundamentals can first work through best machine learning projects for computer science students, then move to an RL-specific portfolio environment. Keep experiments versioned, seed-controlled, and reproducible; open-source workflows can also benefit from the practices described in Indian open-source AI developer projects: 2026 guide.
Validate with walk-forward testing
Never rely on a single random train-test split for time-series trading. Use chronological evaluation:
1. Train on an initial historical window.
2. Validate on the next period without retraining on future information.
3. Roll the window forward and repeat.
4. Reserve a final untouched period for one-time evaluation.
Measure annualised return, volatility, Sharpe and Sortino ratios, maximum drawdown, Calmar ratio, turnover, hit rate, profit concentration, and performance by market regime. Compare against benchmarks and report results after brokerage, exchange charges, securities transaction tax, GST, stamp duty, slippage, and applicable taxes. Model liquidity conservatively, especially for smaller stocks.
Run stress tests for commodity-price shocks, abrupt policy changes, power shortages, market gaps, exchange outages, and prolonged sector drawdowns. A strategy that works only during one steel-price cycle is not robust.
Add governance and deployment controls
An RL system should produce an audit trail: input snapshot, model version, action, expected risk, executed trade, and reason for rejecting or modifying an action. Use paper trading before capital deployment, then introduce strict limits and human approval. Monitor feature drift, turnover, drawdown, data failures, and divergence between predicted and realised costs.
Retraining should be scheduled, not triggered impulsively by a losing week. Establish stop conditions—for example, a drawdown threshold, abnormal data quality, or repeated constraint violation. Keep the model as one input to an investment process, with fundamental review and independent risk oversight.
Common failure modes
- Look-ahead bias: using revised or late-published information too early.
- Survivorship bias: excluding delisted, merged, or poorly performing stocks.
- Overfitting: tuning rewards and hyperparameters against the test period.
- Unrealistic execution: assuming every trade fills at the closing price.
- Sparse data: treating sentiment or company disclosures as precise numerical signals.
- Regime dependence: mistaking a commodity boom for general skill.
- Unclear ownership: deploying a model without a responsible reviewer.
A practical build sequence
Start with a two- or three-asset paper portfolio and daily observations. Build a transparent simulator, add costs and constraints, establish passive baselines, and only then compare Q-learning with actor–critic methods. Publish a research log covering data dates, feature definitions, reward design, splits, failed experiments, and limitations. This makes the project easier to inspect and improves its value as a serious AI or quantitative-finance prototype.
RL can be useful for disciplined allocation research around Odisha’s industrial economy, but its credibility comes from realistic assumptions, leakage-resistant evaluation, and risk controls—not from model complexity.