What machine learning can—and cannot—do
If you are asking how to optimize a stock portfolio using machine learning, start with the right objective. ML is useful for estimating returns, volatility, correlations, regime changes, and the probability that an investment signal will hold up out of sample. It is not a dependable crystal ball, and it does not remove market risk.
For Indian investors, the practical goal is usually to build a repeatable process for a universe such as Nifty 50, Nifty 500, sector indices, or liquid ETFs. The process should improve portfolio construction after accounting for brokerage, taxes, slippage, liquidity, concentration, and drawdowns. A model that predicts prices accurately but produces excessive turnover may be less useful than a simple low-cost strategy.
This is also a strong applied ML project. Beginners can first build a smaller experiment alongside machine learning portfolio projects for beginners in India, then add financial data and portfolio constraints once the evaluation pipeline is reliable.
Define the investment problem before choosing a model
Write down the decision your system must make. Common formulations include:
- Ranking: score stocks by expected risk-adjusted return and select the top group.
- Return forecasting: estimate next-week or next-month returns.
- Risk forecasting: estimate volatility, downside risk, or correlations.
- Allocation: convert forecasts into portfolio weights.
- Rebalancing: decide when expected benefits justify trading costs.
Avoid vague targets such as “predict the market.” A well-defined target might be the next 20-trading-day excess return relative to the stock’s sector or benchmark. Use a fixed investment horizon, define whether dividends are included, and specify when information becomes available.
Your data universe must also avoid survivorship bias. If you test only companies that are members of today’s index, you may exclude past failures and overstate performance. Preserve historical constituents where possible, include delisted securities in research datasets, and use adjusted prices consistently.
Build a clean Indian-market dataset
Useful inputs include daily OHLCV data, corporate actions, financial statements, benchmark returns, sector labels, interest rates, currency data, and broad macroeconomic indicators. News and alternative data can be added later; they introduce timestamp, licensing, language, and interpretation challenges.
At minimum, create features such as:
- Momentum over several lookback periods
- Volatility and downside deviation
- Trading volume and liquidity measures
- Moving-average distance and trend strength
- Value and quality ratios from reported financials
- Sector-relative returns
- Market beta and rolling correlations
Every feature needs an as-of timestamp. Do not use a quarterly result before its public release date, or a revised macroeconomic figure that was unavailable at the time. Align data by trading day, handle exchange holidays, and forward-fill only when that reflects how the information would actually have been available.
Do not randomly shuffle time-series observations. A random split can leak future market conditions into training. Use chronological train, validation, and test periods; better still, use walk-forward validation, where the model trains on an expanding or rolling window and is tested on the next period.
Choose models for robustness, not novelty
Begin with strong baselines: equal weighting, market-cap weighting, minimum variance, and a simple momentum or value rule. Then compare them with models such as regularised linear regression, random forests, gradient-boosted trees, or carefully constrained neural networks.
Tree models can capture nonlinear interactions among tabular features, while regularised linear models are easier to inspect and less prone to overfitting. Deep learning may be appropriate for large, high-quality datasets, but it is rarely the best first step for a small stock universe with noisy labels.
A useful architecture separates three tasks:
1. Signal model: estimates expected return or ranks securities.
2. Risk model: estimates volatility, covariance, liquidity, and drawdown exposure.
3. Portfolio optimiser: converts those estimates into weights subject to constraints.
The optimiser might maximise expected return minus a risk penalty, minimise predicted variance, or optimise a risk-adjusted objective such as the Sharpe ratio. Add constraints for maximum stock and sector weights, minimum liquidity, turnover, short-selling rules, and a cash allocation if required.
Backtest like a sceptical portfolio manager
A backtest should simulate the decisions and information available at each historical point. Include realistic brokerage, exchange charges, STT where applicable, stamp duty, GST, slippage, bid-ask spread, and market-impact assumptions. For Indian equities, costs vary by product, broker, turnover, and execution route, so document assumptions rather than copying a generic rate.
Measure more than cumulative return:
- Annualised return and volatility
- Sharpe and Sortino ratios
- Maximum drawdown and recovery time
- Win rate, turnover, and average holding period
- Worst month and worst rolling one-year period
- Exposure by sector, market capitalisation, and factor
- Performance after estimated costs and taxes
Use an untouched final test period. Do not repeatedly tune the strategy on it. If you try many feature sets, models, or rebalance frequencies, the best result may simply be the product of multiple testing. Record experiments, preserve failed runs, and prefer a modest edge that survives sensitivity checks.
A good next step is to turn the research into a reproducible pipeline. Developers working on larger datasets can study scalable machine learning infrastructure for developers and implementing scalable ML pipelines for predictive analytics.
Add risk controls before paper trading
Portfolio construction is where many attractive signals fail. Cap single-stock and sector exposure, limit turnover, impose liquidity filters, and use volatility targeting only if it improves behaviour out of sample. Monitor beta, factor concentration, correlation changes, and exposure to a single macro theme.
Paper trade the complete system before committing capital. Log the raw data, features, model version, orders, rejected orders, prices, costs, and portfolio state. Set alerts for stale data, missing prices, abnormal predictions, breached limits, and unexpected turnover. Retrain on a defined schedule rather than after every disappointing trade.
Do not treat stop-losses as a complete risk-management system. They can be affected by gaps and execution conditions. A diversified allocation, position sizing, liquidity discipline, and a maximum portfolio drawdown policy are more fundamental.
A practical Python project structure
A maintainable research repository can contain:
data/: ingestion metadata and local feature snapshotsfeatures/: point-in-time feature generationmodels/: training code and saved model versionsportfolio/: optimisation and constraint logicbacktest/: event-driven simulation and cost modelreports/: performance, risk, and exposure reportstests/: checks for leakage, dates, weights, and reproducibility
Use version control, configuration files, fixed random seeds where relevant, and unit tests for portfolio weights. Never place broker credentials in source code. If this is primarily a learning project, document the assumptions and results in a public repository; guidance on how to build a machine learning portfolio on GitHub can help structure that work.
Compliance and responsible use in India
This article is educational, not personalised investment advice. Automated or advisory activity may involve obligations under Indian securities regulations, exchange rules, broker agreements, data licences, and applicable tax requirements. Review the current requirements before offering signals to others, managing external money, or connecting an automated execution system to a broker.
The strongest ML portfolio system is not the one with the most complex algorithm. It is the one with clean point-in-time data, honest validation, realistic costs, explicit constraints, and a monitoring process that can identify when its assumptions no longer hold.