0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · non-stationary time-series analysis

Non-Stationary Time-Series Analysis: A Practical Guide

  1. aigi

    Non-stationary time-series analysis deals with data whose statistical behaviour changes over time. The mean may drift, volatility may expand, seasonal patterns may shift, or relationships between variables may evolve. This is the default condition for many real-world systems: electricity demand, UPI transactions, traffic, rainfall, commodity prices, disease counts, and AI product usage rarely behave like fixed laboratory processes.

    Treating such data as stationary can produce spurious correlations, misleading confidence intervals, and forecasts that fail when conditions change. A sound workflow therefore separates structural change from noise, chooses transformations deliberately, validates forecasts chronologically, and monitors the model after deployment.

    What makes a time series non-stationary?

    A weakly stationary series has a stable expected value, variance, and dependence structure over time. A non-stationary series violates one or more of these assumptions.

    Common forms include:

    • Trend non-stationarity: a deterministic movement, such as steadily increasing demand.
    • Stochastic trend: a random walk or unit-root process in which shocks have persistent effects.
    • Seasonality: repeated patterns by hour, day, week, month, or festival period.
    • Changing variance: volatility that rises during market stress or falls after a policy intervention.
    • Structural breaks: abrupt changes caused by regulation, outages, pandemics, product launches, or pricing changes.
    • Time-varying relationships: predictors that become more or less useful as behaviour changes.

    These categories matter because they require different remedies. Differencing can remove a stochastic trend but may damage a series with a purely deterministic trend. Seasonal adjustment will not fix a one-time level shift, and a more complex model cannot compensate for poor data definitions.

    A practical diagnostic workflow

    Start with a time-ordered plot, not a statistical test. Mark known events such as a tariff revision, a new data pipeline, a lockdown, or a change in sampling frequency. Plot the raw series alongside rolling mean and rolling standard deviation. A rising rolling mean suggests drift; changing dispersion suggests heteroskedasticity; repeating peaks suggest seasonality.

    Then inspect:

    • ACF and PACF: slow ACF decay often signals non-stationarity; seasonal spikes indicate periodic dependence.
    • Grouped summaries: compare distributions by month, weekday, region, or operating regime.
    • Missingness and outliers: sensor failures and reporting delays can mimic structural change.
    • Change-point diagnostics: use them to locate possible breaks, then verify against domain events.

    Unit-root tests are useful but should not be treated as verdicts. The Augmented Dickey-Fuller test has a null hypothesis of a unit root, while KPSS starts from stationarity as the null. Using both provides a more balanced view. Results are sensitive to lag selection, deterministic trend terms, sample size, and breaks. In Indian economic or operational data, a policy change or festival effect can easily invalidate a simplistic test interpretation.

    Transformations that preserve the business question

    Choose a transformation based on the quantity you need to forecast and explain.

    Differencing

    First differences, \(y_t-y_{t-1}\), model changes rather than levels. Seasonal differences compare an observation with the equivalent prior season, such as this Monday with last Monday. Use the smallest order that removes persistent dependence; over-differencing creates unnecessary noise and weakens interpretability.

    Log and variance-stabilising transforms

    A log transform is useful when percentage changes are more meaningful than absolute changes, for example transaction volume or revenue. For zero or negative values, consider a carefully documented alternative such as Yeo-Johnson or a shifted scale. Always reverse transformations correctly when reporting forecasts; averaging on the transformed scale can bias level predictions.

    Detrending and decomposition

    A deterministic trend can be estimated and removed, while STL or related decomposition methods can separate trend, seasonality, and remainder. Decomposition is valuable for diagnostics, but components should be estimated inside each training window during evaluation to avoid leaking future information.

    Models for non-stationary series

    ARIMA represents non-stationarity through integration: the “I” term captures differencing before autoregressive and moving-average dynamics are modelled. Seasonal ARIMA extends this to recurring patterns. These models remain strong baselines when the series is reasonably regular and the forecast horizon is clear.

    For multiple related series, consider vector autoregression after establishing suitable stationarity, or vector error-correction models when variables are cointegrated. Cointegration means individual series may wander, but a combination of them remains stable over the long run. The Engle-Granger approach is practical for a small number of variables; Johansen methods support a richer multivariate setting.

    Modern forecasting systems can combine lag features, calendar variables, exogenous signals, and machine-learning models. Gradient-boosted trees are effective for tabular lag features; state-space models handle evolving level and trend; probabilistic models provide intervals rather than a single optimistic number. Deep sequence models are not automatically better: they need enough history, careful backtesting, and stable data pipelines. Teams building production systems should also review scaling backend infrastructure for AI applications before adding model complexity.

    Forecast evaluation without leakage

    Random train-test splits are inappropriate for time series. Use expanding-window or rolling-window validation:

    1. Train on the earliest period.
    2. Forecast the next horizon.
    3. Move the cutoff forward.
    4. Repeat across normal and stressed regimes.

    Compare against meaningful baselines, such as last value, seasonal naive, or drift. Report MAE or RMSE alongside business measures such as stockout rate, capacity shortfall, or interval coverage. Evaluate multiple horizons because a model that performs well tomorrow may fail next month.

    Do not difference, scale, impute, select features, or tune hyperparameters using the full dataset before splitting. Every preprocessing step must be fitted only on the training window. For operational dashboards, pair forecasts with uncertainty bands and clear freshness indicators; real-time data storytelling for non-technical users offers useful guidance on communicating changing signals responsibly.

    India-focused use cases

    In India, non-stationary analysis is particularly relevant to:

    • Power and renewables: demand varies with heat, monsoon conditions, local holidays, and distributed solar adoption.
    • Digital payments: transaction volumes shift with salary cycles, festivals, outages, and changing consumer behaviour.
    • Mobility and logistics: route demand changes with weather, events, fuel prices, and urban construction.
    • Agriculture and climate: rainfall, reservoir levels, crop prices, and heat stress show seasonality alongside long-term shifts.
    • Public health: reported cases are affected by testing capacity, reporting practices, interventions, and outbreaks.
    • Infrastructure: sensor streams from roads and bridges require drift detection as assets age; monitoring architectures can draw on real-time bridge health monitoring systems in India.

    The strongest projects combine statistical diagnostics with domain calendars, data-quality checks, and a documented response plan for regime changes.

    Common mistakes and a deployment checklist

    Avoid these failure modes:

    • Differencing every series without checking whether the trend is deterministic.
    • Declaring stationarity from one test with a convenient p-value.
    • Ignoring breaks, missing data, revisions, or changing measurement definitions.
    • Selecting a model on one historical period and assuming the regime will persist.
    • Reporting point forecasts without uncertainty or operational thresholds.
    • Monitoring accuracy but not drift in inputs, residuals, seasonality, and forecast coverage.

    Before deployment, document the target, forecast horizon, time zone, aggregation rule, transformation, validation design, baseline, and retraining trigger. Track residual autocorrelation and bias by region, product, or customer segment. If the series feeds a live AI product, connect monitoring to the broader real-time location intelligence platform approach where spatial context affects demand or risk.

    FAQ

    Is a trending series always non-stationary?
    Not necessarily. A deterministic trend can be modelled explicitly, while stochastic trends require different treatment. The distinction should be tested and validated rather than assumed.

    Should I always difference before using machine learning?
    No. Tree models can learn some level and trend effects from lag and calendar features, but transformations may still improve stability. Compare pipelines using chronological backtesting.

    How much data is enough?
    There is no universal threshold. You need enough observations to cover multiple seasonal cycles, relevant regimes, and the forecast horizon. Short, highly seasonal datasets require stronger domain assumptions.

    When should I use cointegration?
    Use it when several non-stationary variables have a plausible long-run equilibrium relationship. Validate that relationship economically or operationally; statistical cointegration alone does not establish causation.

    Last updated 24 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.