What Karpathy Autoresearch can—and cannot—do
Karpathy Autoresearch is best treated as an experimental framework for iterating on research code and model configurations, not as an oracle that independently understands monetary policy. For RBI analysis, it can help you test repeatable pipelines: ingest official releases, construct features, run forecasts or classifications, compare experiments, and retain the code and metrics behind each result.
That distinction matters. The RBI’s decisions reflect inflation, growth, liquidity, exchange-rate conditions, financial stability, global events, and judgement. A model can identify historical relationships, but it cannot establish that a repo-rate change caused a particular outcome. Use its outputs to support research—not to replace policy documents, economic reasoning, or human review.
If you are new to AI-assisted economic research, first understand the data-cleaning principles used in an AI workflow for analysing bank statements in India. The same discipline—source validation, consistent schemas, and careful handling of dates—applies here.
Define a narrow research question
Avoid beginning with “predict the next RBI decision.” Start with a question that can be measured and audited, such as:
- How did CPI inflation behave in the six months following repo-rate changes?
- Which indicators most often appeared before a change in the policy stance?
- Can official policy statements be classified as hawkish, neutral, or dovish using a transparent rubric?
- How accurately can a model nowcast inflation using only information available at the time?
- Did transmission to lending rates differ across banks and loan categories?
Write down the target, forecast horizon, information cutoff, and success metric before running experiments. This prevents accidental use of data published after the forecast date—a form of look-ahead bias that can make an impressive model useless in practice.
Build a trustworthy RBI dataset
Prefer primary sources and preserve the original files or URLs. Useful inputs include:
- RBI monetary policy statements, resolutions, minutes, speeches, reports, and historical databases
- CPI, WPI, IIP, GDP and national accounts releases from the Ministry of Statistics and Programme Implementation
- Union Budget and fiscal data from the Ministry of Finance
- High-frequency indicators such as GST collections, fuel prices, credit growth, system liquidity, yields, and exchange rates
- Global variables including crude oil, US policy rates, and major-market volatility indices
Create a tidy table with one observation per date and explicit publication timestamps where available. Store the observation period separately from the release date. For example, an inflation number for January may have been unavailable when a January policy decision was made. Your training data must reflect that information set.
For text analysis, retain the full document, title, issuing body, publication date, document type, and page or paragraph references. Do not rely solely on scraped news summaries. News can be useful for context, but official wording should anchor any conclusion.
Researchers working with Indian public documents may also benefit from fine-tuning a model on Indian policy data, particularly when building a domain-specific classifier. Fine-tuning is optional; a carefully designed retrieval and scoring workflow is often easier to audit.
Structure the Autoresearch experiment
A practical project separates the repository into four layers:
1. Ingestion: Download or load source files, record provenance, and standardise dates, units, and frequency.
2. Feature construction: Create lags, rolling averages, surprises against forecasts, rate-change indicators, policy-stance labels, and event windows.
3. Evaluation: Run time-ordered validation, calculate metrics, and compare the experiment with a simple baseline.
4. Reporting: Save charts, configuration files, model outputs, and a plain-language interpretation.
Autoresearch-style iteration is most useful when each run changes one meaningful component—such as a lag window, text representation, or model family—and records the result automatically. Keep a README, data dictionary, experiment log, and environment lockfile. Pin package versions and use a fixed random seed where possible.
Begin with strong baselines: a previous-value forecast, a moving average, a policy-statement keyword score, or a logistic regression. A complex neural network should earn its place by outperforming these baselines out of sample, not by producing a more elaborate chart.
Analyse policy decisions and communication
For rate analysis, tag each policy meeting with the repo rate, standing deposit facility, marginal standing facility, cash reserve ratio, stance, vote details where available, and the relevant event date. Create windows around announcements—for example, trading-day and monthly windows—and avoid mixing announcement effects with unrelated events.
For document analysis, define a rubric before looking at results. Possible signals include references to persistent inflation, liquidity management, growth risks, transmission, food prices, external risks, and future action. Have a human reviewer label a sample, then measure agreement between reviewers before training or prompting a classifier.
Do not present correlation as transmission. A repo-rate increase may coincide with an oil shock, currency pressure, or fiscal news. Test alternative specifications, include controls where justified, and show confidence intervals. For market outcomes, examine abnormal changes relative to a benchmark rather than attributing every movement on policy day to the RBI.
A related workflow is predicting Nifty 50 trends with machine learning, but the same warning applies: market prediction is noisy, regime-dependent, and highly vulnerable to leakage.
Evaluate honestly
Use expanding-window or rolling-window validation rather than random train-test splits. Monetary-policy datasets are small, and observations are serially related. Report:
- Forecast error metrics such as MAE or RMSE for continuous outcomes
- Precision, recall, F1, and calibration for decision or stance classification
- Performance by regime, including high-inflation and crisis periods
- Results against naive and policy-relevant baselines
- Sensitivity to missing data, revised releases, and alternative lag structures
Backtest only with data that would genuinely have been available at each historical point. Treat revisions as a separate research choice: a real-time dataset answers a different question from a revised historical dataset.
Common failure modes and safeguards
Overfitting: A small number of RBI meetings can support many spurious patterns. Limit feature searches, use nested validation where feasible, and prefer simpler models.
Unverified scraping: RBI pages, PDFs, and archives can change. Cache source files, calculate checksums, and manually verify important observations.
False precision: A forecast interval is more informative than a single predicted number. Explain uncertainty and avoid investment or policy recommendations based on one model run.
Prompt-generated citations: If a research agent summarises documents, require document identifiers and quoted passages. An agent that cannot show its evidence should not support a factual claim. Guidance on building research agents for Indian policy transitions offers a useful parallel in renewable-energy policy report synthesis.
Security and governance: Do not upload confidential bank, customer, or proprietary market data to an external model without approval. Apply the same access controls and logging expected in any production AI system; real-time AI agent policy enforcement in India provides a relevant governance reference.
A practical output for 2026
A useful deliverable is not merely a forecast. Publish a compact research pack containing the question, data cutoff, source register, methodology, baseline, model comparison, uncertainty range, charts, failed experiments, and limitations. Add a “what would change my conclusion?” section covering new inflation data, revisions, global shocks, or a change in the RBI’s reaction function.
As of 2026, the strongest India-focused use of Autoresearch is reproducible evidence gathering and model evaluation. It can reduce repetitive experimentation and expose patterns across large policy archives, while analysts remain responsible for causal reasoning, context, and final claims. That combination is more credible than treating automation as a substitute for macroeconomic expertise.
FAQ
Is Karpathy Autoresearch a ready-made RBI forecasting tool?
No. It is a framework or workflow for automated experimentation. You must supply data, define the objective, write or adapt the research code, and validate every result.
Which model should I start with?
Start with transparent baselines and classical time-series or tabular models. Move to language models or neural networks only when they address a clear limitation and improve out-of-sample performance.
Can it predict the next RBI repo-rate decision?
It can estimate historical probabilities under stated assumptions, but decisions depend on information and judgement that are difficult to encode. Do not present a prediction as certainty.
Where should I publish the work?
Share a reproducible repository or notebook with source links, data definitions, cutoff dates, experiment logs, and limitations. Keep sensitive or licensed data out of public releases.
Apply for AI Grants India
If you are building an India-focused AI research or policy product, explore support and funding through AI Grants India.