Start with the right market question
A reinforcement learning (RL) system should not begin with “Can an agent beat the market?” Begin with a narrower, testable question: Can an agent improve risk-adjusted execution or portfolio decisions for a defined set of India-listed companies with Assam exposure? Assam’s petroleum and gas economy is important, but there is no single “Assam stock” universe. Companies may have assets, production, logistics, refineries, pipelines, or exploration interests connected to the state while being listed on the NSE or BSE.
Create a documented universe before modelling. Record each company’s ticker, exchange, corporate actions, liquidity, Assam connection, and inclusion date. Avoid implying that a company’s share price is driven only by local production: global crude prices, refining margins, rupee movements, interest rates, regulation, and broad Indian market sentiment can dominate returns.
This is a research and engineering project, not a promise of investment returns. Retail users should also review applicable SEBI rules, broker terms, taxes, and suitability requirements before automating any order flow.
Build an auditable data pipeline
Use adjusted daily OHLCV data from a licensed or reliable exchange-data provider. Preserve raw files and create versioned cleaned datasets so every experiment can be reproduced. Your pipeline should include:
- Prices and corporate actions: adjust for splits, bonuses, dividends, and symbol changes without leaking future information.
- Trading data: volume, turnover, bid–ask spread where available, delivery statistics, and market holidays.
- Energy variables: Brent or relevant crude benchmarks, natural-gas prices, refining margins, freight indicators, and inventory proxies.
- India-specific context: Nifty or sector benchmarks, USD/INR, RBI rates, inflation, policy announcements, and relevant government releases.
- Company fundamentals: results, debt, cash flow, production, reserves, and capex, using only information available on each historical date.
- Event and text signals: earnings announcements, regulatory changes, disruptions, and weather events, with publication timestamps retained.
Treat missingness as information to investigate, not a value to silently fill. Align all series to exchange trading sessions and use publication-time joins for fundamentals and news. A feature available after market close must not influence that same day’s action. These controls matter more than adding another neural-network layer.
For a beginner-friendly way to practise reproducible modelling, compare your pipeline with the workflow in machine learning portfolio projects for beginners in India.
Define the trading environment precisely
A useful environment mirrors the decisions an actual portfolio manager can make. At each step—normally one trading day—the agent receives an observation, chooses a portfolio action, pays realistic costs, and receives a next-period reward.
A practical observation can contain:
- Recent returns, volatility, volume, spreads, and momentum for each asset.
- Normalised crude, gas, currency, benchmark, and rates data.
- Portfolio weights, cash balance, turnover, drawdown, and current exposures.
- Position and liquidity limits, so the policy understands operational constraints.
Prefer continuous portfolio weights or bounded position changes over a simplistic buy/hold/sell action when managing several stocks. Enforce constraints directly: maximum weight per company, gross and net exposure, turnover limits, minimum cash, and no shorting unless permitted by the chosen account and product.
A basic net portfolio return is:
r_net,t = portfolio_return,t – brokerage_t – taxes_t – exchange_charges_t – slippage_t
A reward can combine return and risk, for example:
reward_t = r_net,t – λ × volatility_t – μ × drawdown_penalty_t
Do not hide risk inside an arbitrary reward. Report return, volatility, maximum drawdown, turnover, costs, hit rate, concentration, and downside measures separately. Reward shaping can make training easier, but the final evaluation must use economic metrics.
Choose algorithms for the problem, not the trend
Start with a transparent baseline before RL: buy-and-hold, equal weight, volatility targeting, moving-average rules, and supervised return forecasts. If an RL policy cannot beat a simple strategy after costs and risk adjustment, complexity is not justified.
For a discrete, small action space, DQN may be useful for experimentation. For bounded portfolio weights, PPO or another actor–critic approach is usually a more natural starting point. An offline or batch-constrained method may be safer when the only evidence is historical data, because the agent cannot genuinely explore markets during training. Libraries such as Gymnasium-compatible environment tooling, PyTorch, and Stable-Baselines3 can accelerate prototyping, but every library still needs financial validation.
Use multiple random seeds and keep the training configuration under version control. Track the data snapshot, feature definitions, reward formula, transaction-cost assumptions, and model checkpoint for every run.
Train without leaking the future
Financial time series are non-stationary and highly dependent across observations. Randomly shuffling rows into training and test sets creates leakage. Use chronological splits such as:
- Training window: earlier history for fitting the policy.
- Validation window: later history for reward and hyperparameter decisions.
- Test window: a sealed period used once for final reporting.
- Walk-forward evaluation: repeatedly retrain on the past and test on the next period.
Keep preprocessing fit within each training window. Test sensitivity to different market regimes, including oil-price shocks, weak liquidity, sharp rupee moves, and broad Indian sell-offs. Run ablations that remove news, fundamentals, or commodity inputs; this reveals whether the policy is learning a plausible signal or exploiting a data artefact.
A strong report includes confidence intervals or bootstrap ranges, not just the best backtest curve. Compare against the same universe, dates, rebalancing frequency, and cost assumptions. Include failed experiments and periods when the strategy underperformed.
Add risk controls before deployment
A model that produces attractive historical returns can still be unsuitable for live use. Add hard safeguards outside the policy:
- Reject orders that exceed position, turnover, price-band, or liquidity limits.
- Stop or reduce trading after abnormal data gaps, stale prices, broker errors, or excessive drawdown.
- Require human approval for new instruments, large orders, or policy changes.
- Log observations, actions, expected fills, actual fills, costs, and overrides.
- Separate research, paper trading, and production credentials.
- Monitor drift in spreads, volatility, feature distributions, and action frequency.
Begin with paper trading and shadow mode. Compare hypothetical fills with executable quotes, then use the smallest practical size. For an India-focused AI product, deployment discipline matters as much as model accuracy; the broader considerations are similar to those in building AI apps for the next billion users in India.
Suggested implementation stack
A lean research stack can use Python, pandas or Polars for data preparation, NumPy for numerical work, PyTorch for models, and Gymnasium-style interfaces for the environment. Store raw and processed data separately, use Parquet with schemas, and track experiments with Git plus an experiment tracker. Containerise training so dependencies are reproducible. A simple project structure might include data/, features/, env/, agents/, backtests/, reports/, and configs/.
Do not overbuild distributed infrastructure at the start. Scale only when repeated walk-forward experiments justify it. If you later run many agents or market simulations, principles from building distributed systems with AI agents can help with job isolation, observability, and fault handling.
A practical 2026 build plan
1. Define the investable universe and compliance boundaries.
2. Assemble timestamped, adjusted data and write data-quality tests.
3. Implement a deterministic backtesting environment with costs and constraints.
4. Establish non-RL baselines and publish their assumptions.
5. Train one simple PPO or offline-policy prototype with fixed seeds.
6. Run walk-forward tests, ablations, stress scenarios, and cost sensitivity checks.
7. Paper trade with monitoring, alerts, and manual approval.
8. Review performance and operational evidence before considering limited deployment.
The objective is not to manufacture a sophisticated equity curve. It is to build a reproducible decision system whose assumptions, limitations, and failure modes are visible. That standard is especially important when a geographically focused theme such as Assam’s petroleum and gas sector is being represented through nationally traded securities.
FAQ
Can I build this with free data?
You can prototype with openly available data, but licensing, survivorship bias, delayed prices, missing corporate actions, and inaccurate fundamentals can invalidate results. Use dependable, permitted data for serious research.
Should the agent trade only Assam-linked companies?
Not necessarily. Start with a clearly defined Assam-linked universe, then test whether adding an energy benchmark or broader Indian assets improves diversification without weakening the research question.
Is reinforcement learning better than supervised learning here?
Not automatically. RL is useful when actions, portfolio state, costs, and sequential trade-offs are central. For narrow prediction tasks, supervised models and transparent rules may be easier to validate.
How do I learn the required skills?
Build a small end-to-end portfolio project first, then study time-series validation, portfolio construction, market microstructure, and responsible deployment. Best AI frameworks for Indian student entrepreneurs offers a useful starting point for selecting practical tools.
Apply for AI Grants India
If you are building an India-focused AI system for financial research, risk analytics, or energy intelligence, explore support and funding opportunities through AI Grants India.