Actor-critic reinforcement learning can help a trading system choose among actions such as holding cash, entering a position, reducing exposure, or exiting. It does not predict the market with certainty. For Karnataka-linked biotech stocks, the difficult work is usually not selecting a neural-network architecture; it is defining the investable universe, preventing data leakage, modelling liquidity and costs, and enforcing risk limits.
This guide explains a research workflow for 2026. It is educational, not investment advice. Verify securities, filings, tax treatment, broker rules, and applicable SEBI requirements before deploying capital or automation.
Start with a precise trading problem
“Karnataka biotech stocks” is not a standard exchange classification. Build a transparent universe instead of assuming that every healthcare or pharmaceutical company belongs to the sector. Possible inclusion rules include:
- A listed company headquartered in Karnataka.
- Material revenue, facilities, partnerships, or research operations in the state.
- A defined biotechnology, diagnostics, life-sciences, or biopharma business segment.
- Minimum liquidity and listing-history thresholds.
Record the inclusion date and keep delisted, suspended, and acquired companies in historical datasets where appropriate. Otherwise, survivorship bias can make the strategy look better than it was. For practical research, start with daily data and a weekly decision frequency. Intraday systems require stronger data quality, execution modelling, and operational controls.
Define the action space before training. A simple first version might choose long, reduce, or hold, with position size determined separately by a risk layer. Allowing the policy to trade unlimited quantities often produces unrealistic turnover and unstable results.
Design the state and data pipeline
The state should contain only information available at the decision timestamp. Useful feature groups include:
- Adjusted prices, returns, volume, turnover, volatility, drawdown, and moving-average relationships.
- Market and sector benchmarks, interest-rate signals, currency movements, and broad risk sentiment.
- Earnings dates, corporate actions, promoter disclosures, fundraising, trial announcements, and material exchange filings.
- Liquidity measures such as traded value, bid-ask spread estimates, price impact, and days with no trades.
- Portfolio context: current holdings, cash, average entry price, unrealised risk, and recent turnover.
Use point-in-time fundamentals. Financial statements and announcements must enter the dataset only after they became public. Timestamp news and regulatory events carefully; a publication date is not always the same as the time a trader could act.
Clean splits, bonuses, dividends, symbol changes, missing observations, and stale prices. Normalise features using rolling or training-period statistics rather than full-sample statistics. Keep training, validation, and test periods chronological. Randomly shuffling financial observations can leak future regimes into the past.
If you are building a broader AI research stack, document datasets and experiments with the same discipline used when deploying deep learning models on GKE. Reproducible data versions matter more than an impressive model name.
Choose an actor-critic formulation
The actor represents the policy: it maps the current state to action probabilities or a continuous target allocation. The critic estimates expected future value and supplies a learning signal. For a discrete action space, advantage actor-critic is a clear baseline. For continuous allocations, algorithms such as PPO, SAC, or deterministic actor-critic variants may be tested, but complexity should follow a demonstrated need.
A practical architecture can use:
- A small multilayer perceptron for tabular features.
- A recurrent layer only when you can show that sequence memory adds value without creating leakage.
- Separate policy and value heads after shared feature layers.
- Action masking for unavailable actions, such as buying when cash or liquidity rules prohibit it.
Start with a supervised or rule-based benchmark. Compare the reinforcement-learning agent with buy-and-hold, equal-weight exposure, momentum, and a cash-aware sector benchmark. If the actor-critic model cannot beat simple baselines after costs and risk adjustment, adding layers is unlikely to solve the underlying problem.
Build a realistic reward and environment
A reward based only on next-day return encourages excessive turnover and ignores catastrophic losses. A more useful daily reward can be expressed as net portfolio return minus explicit penalties:
- Brokerage, exchange charges, taxes, slippage, spread, and market impact.
- Turnover and concentration penalties.
- Drawdown or volatility penalties.
- Penalties for breaching liquidity, leverage, or maximum-position constraints.
Use a portfolio simulator that processes orders, cash, holdings, corporate actions, and transaction delays. Do not let the agent trade at the same close used to calculate its features unless that execution assumption is genuinely available. Add conservative slippage, partial fills, price limits, and rejected orders where relevant.
For biotech, include event gaps in stress tests. Trial outcomes, approvals, safety announcements, litigation, and fundraising can move prices sharply. A model trained mainly on calm periods may fail precisely when risk is highest.
Train and evaluate without fooling yourself
Use walk-forward evaluation: train on an earlier window, validate on the next period, then roll the window forward. Keep a final untouched test period. Run multiple seeds and report the distribution of outcomes, not just the best run.
Track metrics that reflect deployable performance:
- Net cumulative return and annualised return.
- Maximum drawdown, downside deviation, and Sharpe or Sortino ratio.
- Turnover, average holding period, win rate, profit factor, and exposure.
- Slippage sensitivity and performance by liquidity bucket.
- Performance during market drawdowns, earnings seasons, and biotech-specific events.
Perform ablation tests: remove news, remove fundamentals, change costs, or restrict the universe. A strategy whose edge disappears under a modest cost increase is not ready for live use. Evaluate calibration too: if the actor assigns high confidence to actions that repeatedly fail, reduce position sizing or retrain rather than trusting the score.
Add a risk and execution layer
The model should propose actions; deterministic controls should decide whether those actions are permitted. Set maximum position and sector exposure, daily loss limits, turnover caps, liquidity thresholds, and a portfolio-level stop or review process. Include a kill switch for data outages, abnormal spreads, repeated order failures, model drift, or unexpected broker responses.
Paper trade before deployment. Compare intended orders with fills, latency, rejected orders, and actual costs. Begin with small, capped exposure and require human approval for unusual events. Store model versions, feature snapshots, decisions, orders, fills, and overrides so every trade is auditable.
For production infrastructure, a managed endpoint or serverless component may be useful, but trading systems also need durable state, monitoring, secret management, retries, and strict latency expectations. Review deployment patterns such as running ML models on AWS Lambda in India, while checking whether that architecture suits broker connectivity and order sequencing.
India-specific governance checklist
Before live deployment, clarify whether the activity falls under applicable algorithmic-trading, investment-advisory, research, brokerage, or portfolio-management obligations. Use authorised market infrastructure and broker APIs, protect API credentials, and never rely on scraped or unauthorised data for production decisions. Maintain a written model-risk policy covering validation, approval, retraining, incident response, and retirement.
Do not present historical performance as a promise. Biotech concentration, limited liquidity, corporate actions, and information asymmetry can make losses rapid and difficult to exit. A strong system is one that knows when not to trade.
A practical build sequence
1. Define the Karnataka-linked universe and point-in-time data rules.
2. Create a cost-aware simulator and deterministic risk constraints.
3. Establish simple, reproducible benchmarks.
4. Train a small actor-critic baseline with walk-forward validation.
5. Stress-test gaps, liquidity shocks, costs, and regime changes.
6. Paper trade, monitor drift, and audit every decision.
7. Scale only when live evidence supports the expected risk-adjusted edge.
Actor-critic models are best treated as experimental decision policies inside a controlled portfolio system—not as autonomous stock-picking machines. The quality of timestamps, assumptions, controls, and evaluation will determine usefulness far more than the choice between neighbouring reinforcement-learning algorithms.