0tokens

Apply for AI Grants India

Financial support for innovators building the future of AI in India.

Apply now

Chat · causal discovery baselines

Causal Discovery Baselines: Methods, Benchmarks & Evaluation

  1. aigi

    Causal discovery baselines are the reference methods used to test whether a new algorithm can recover meaningful cause-and-effect structure from observational or interventional data. Choosing them well is not a checkbox exercise: the baseline set determines whether results are credible, comparable, and useful for real-world decision-making.

    A strong comparison typically combines constraint-based, score-based, functional-model, continuous-optimization, and modern neural methods. It also evaluates more than predictive fit. Researchers should report graph-recovery metrics, uncertainty, computational cost, sensitivity to assumptions, and performance under realistic violations such as latent confounding, cycles, measurement noise, and finite samples.

    What Are Causal Discovery Baselines?

    A causal discovery baseline is an established algorithm, model, or null strategy against which a proposed causal structure-learning method is compared. Given data, the method generally attempts to infer a directed acyclic graph (DAG), a partially directed graph, an equivalence class, or a more general causal graph.

    Baselines serve several purposes:

    • Scientific comparison: They show whether a new method improves on known approaches.
    • Ablation and diagnosis: They reveal whether gains come from a particular modeling choice, regularizer, or data source.
    • Reproducibility: Standard methods make results easier for other teams to verify.
    • Deployment assessment: They establish whether a complex method is preferable to a simpler, faster alternative.

    The right baseline depends on the estimand and data-generating assumptions. A DAG learner designed for causal sufficiency should not be judged only against methods that explicitly model latent variables. Likewise, a method intended for time series should include temporal baselines rather than relying exclusively on static i.i.d. algorithms.

    Core Causal Discovery Baseline Families

    Constraint-based methods

    Constraint-based algorithms infer conditional independence relationships and use them to remove or orient graph edges. The classic PC algorithm is the most widely recognized baseline in this family. It tests conditional independences, constructs a skeleton, and applies orientation rules to produce a CPDAG representing a Markov equivalence class.

    Common choices include:

    • PC: A standard baseline for causally sufficient, acyclic settings.
    • CPC: A conservative variant that avoids some potentially ambiguous orientations.
    • FCI: Extends constraint-based discovery to latent confounding and selection bias, returning a PAG rather than a DAG.
    • RFCI: A faster, often less complete alternative to FCI for high-dimensional settings.

    These methods are interpretable and theoretically grounded, but their performance can degrade when conditional independence tests are underpowered, variables are numerous, or the data are far from the assumed distribution. The choice of test matters: partial-correlation tests are suitable for linear-Gaussian relationships, while kernel, rank-based, discrete, or nonlinear tests may be more appropriate elsewhere.

    Score-based methods

    Score-based algorithms search over candidate graphs and optimize a statistical score. The Greedy Equivalence Search (GES) algorithm is a standard comparison point. It performs forward and backward phases over equivalence classes and is commonly paired with BIC or related penalized likelihood scores.

    Other useful baselines include:

    • Greedy DAG search: Directly optimizes a graph score through local edge operations.
    • Hill climbing: A practical heuristic that can be initialized in multiple ways.
    • MMHC: Combines constraint-based neighborhood selection with score-based orientation.
    • Bayesian structure learning: Samples or optimizes over graph structures using a posterior distribution.

    Score-based methods can perform strongly when the score matches the data-generating process. However, search complexity, local optima, equivalent graphs, and score misspecification should be reported. For a fair comparison, document the score, penalty, number of restarts, maximum indegree, and stopping criteria.

    Functional and additive-noise models

    Functional causal model baselines exploit asymmetries in how variables are generated. In an additive-noise model, for example, one may compare whether a relationship is better represented as:

    \[
    Y = f(X) + N_Y
    \]

    where the noise is independent of the cause. The reverse direction often fails to satisfy the same independence pattern.

    Representative approaches include additive-noise modeling, nonlinear causal additive models, and variants that use distributional asymmetry. These baselines are valuable when the proposed method relies on nonlinear mechanisms or non-Gaussian noise. Their assumptions must be stated clearly: identifiability may disappear under linear-Gaussian models, and flexible function classes can overfit in small samples.

    Continuous-optimization methods

    Continuous methods convert discrete DAG search into differentiable optimization. NOTEARS is the best-known baseline in this category. It represents the adjacency matrix as a continuous parameter and imposes an acyclicity constraint such as:

    \[
    h(W) = \operatorname{tr}(e^{W \circ W}) - d = 0,
    \]

    where \(W\) is the weighted adjacency matrix, \(\circ\) denotes the Hadamard product, and \(d\) is the number of variables.

    Frequently compared methods include:

    • NOTEARS-Linear: Linear structural equation models with differentiable acyclicity.
    • NOTEARS-MLP: Nonlinear neural structural equations.
    • DAG-GNN: A variational graph-learning approach with an acyclicity constraint.
    • GraN-DAG: Neural additive mechanisms for nonlinear DAG learning.
    • RL-based DAG learners: Use reinforcement learning or policy optimization to construct graphs.

    Continuous approaches are convenient for incorporating differentiable objectives, but they can be sensitive to scaling, regularization, optimization tolerances, and numerical violations of acyclicity. Always verify the final graph using an independent cycle check rather than assuming the optimizer reached a valid DAG.

    Time-series and dynamic baselines

    Static causal discovery baselines are inappropriate when temporal ordering is central. For time-series data, consider methods such as:

    • Granger causality: Tests whether past values of one series improve prediction of another.
    • VAR-based methods: Model multivariate lagged relationships.
    • PCMCI and PCMCI+: Constraint-based methods designed for time series and autocorrelated processes.
    • TiMINo: Uses temporal structural equation assumptions.
    • Dynamic Bayesian networks: Represent lagged dependencies explicitly.

    A temporal baseline should distinguish contemporaneous effects from lagged effects, account for autocorrelation, and use time-aware train/test splits. Randomly shuffling observations can leak future information and produce misleadingly strong results.

    How to Choose Causal Discovery Baselines

    Begin with the claims made by the proposed method. A compact but defensible baseline suite often contains:

    1. A classical constraint-based method, such as PC or FCI.
    2. A classical score-based method, such as GES.
    3. A functional-model method if nonlinear or non-Gaussian identifiability is relevant.
    4. A continuous-optimization method, such as NOTEARS.
    5. A domain-specific method for time series, text, images, genomics, or experiments.
    6. A simple null or heuristic baseline, such as an empty graph, prior graph, or correlation network where appropriate.

    Match each baseline to the same observed variables, preprocessing pipeline, train/test split, and available interventions. Do not silently give one method access to domain knowledge, latent-variable annotations, or tuned hyperparameters that are unavailable to the others.

    It is also useful to separate assumption-matched and assumption-challenging comparisons. Assumption-matched experiments answer whether a method works when its theory applies. Assumption-challenging experiments test robustness when relationships are nonlinear, noise is heteroscedastic, hidden confounders exist, or the graph is sparse only approximately.

    Benchmark Datasets and Synthetic Generators

    Synthetic DAG benchmarks

    Synthetic data are essential because the ground-truth graph is known. Common generators sample a DAG, assign edge weights, and generate observations using linear or nonlinear structural equations. Vary at least:

    • Number of nodes and edges
    • Expected degree and graph density
    • Sample size
    • Noise distribution and variance
    • Linear versus nonlinear mechanisms
    • Gaussian versus non-Gaussian noise
    • Presence of latent variables
    • Faithfulness violations
    • Cycles or feedback, if supported by the task

    Report the random seeds and generator configuration. A single synthetic setting can hide instability; use multiple graph draws and aggregate results with confidence intervals.

    Real-world benchmarks

    Real datasets provide ecological validity but often have incomplete or uncertain ground truth. Examples may come from biology, economics, climate science, healthcare, or industrial systems. In many cases, the reference graph is an expert consensus or a curated database rather than a fully verified causal DAG.

    For Indian AI and deep-tech teams, domain-specific data can be especially valuable: agricultural interventions, public-health programs, energy demand, mobility, and financial risk all introduce missingness, distribution shift, policy effects, and confounding that generic benchmarks may not capture. Clearly distinguish validated causal edges from associations or expert hypotheses.

    Metrics for Evaluating Causal Discovery

    No single metric captures causal graph quality. Report several complementary measures.

    Structural Hamming distance

    Structural Hamming distance (SHD) counts edge additions, deletions, and orientation errors needed to transform the estimated graph into the ground truth. It is intuitive, but implementations differ in how they treat reversals and partially directed edges. Define the convention explicitly.

    Structural intervention distance

    Structural intervention distance (SID) evaluates whether the estimated graph implies correct adjustment and intervention effects. It is often more causally meaningful than SHD because a graph can contain a minor structural error without changing key intervention queries—or can look close while producing incorrect adjustment sets.

    Precision, recall, and F1

    For skeleton recovery, report adjacency precision, recall, and F1. For directed edge recovery, compute orientation-aware metrics separately. If the method outputs a CPDAG or PAG, do not force every uncertain edge into an arbitrary direction; evaluate compelled and reversible orientations distinctly.

    Treatment-effect and intervention metrics

    When the application is decision-making, evaluate downstream causal quantities:

    • Average treatment effect error
    • Conditional treatment effect error
    • Interventional distribution distance
    • Policy value or regret
    • Calibration of uncertainty intervals

    These metrics connect structure learning to the business or scientific objective. A method that has slightly worse SHD but substantially better intervention estimates may be preferable in deployment.

    Efficiency and stability

    Include runtime, peak memory, scalability with node count, and failure rates. Measure stability across bootstrap samples or dataset perturbations. An edge appearing in 95% of resamples is more actionable than one appearing in 51%, even if both count equally in a single graph metric.

    Experimental Design and Reproducibility

    A credible causal discovery comparison should document:

    • Variable transformations, normalization, and missing-value treatment
    • Whether data were split by rows, time, environment, or intervention regime
    • Hyperparameter search space and validation procedure
    • Independence-test significance levels or score penalties
    • Neural architecture, optimizer, learning rate, and stopping rules
    • Hardware, software versions, and random seeds
    • How cycles, near-zero edges, and uncertain orientations are handled
    • The exact graph-to-metric conversion procedure

    Avoid selecting hyperparameters on the test graph. If ground truth is available only for evaluation, use validation simulations or a pre-registered configuration. For real data, use expert review or downstream validation rather than tuning toward an uncertain reference network.

    A useful reporting table includes each method's assumptions, graph type, runtime, SHD, SID, adjacency F1, orientation F1, and failure rate. Pair it with qualitative graphs and uncertainty summaries. Open-source code should include configuration files and scripts that reproduce every table and figure.

    Common Mistakes When Using Baselines

    Comparing incompatible graph outputs

    A DAG, CPDAG, PAG, and weighted adjacency matrix do not encode the same information. Convert outputs only through a documented, theoretically justified procedure.

    Treating correlation as causation

    Correlation networks can be useful descriptive baselines, but they are not causal discovery methods. Label them accordingly and never interpret their edges as intervention effects without additional assumptions.

    Ignoring hidden confounding

    If latent variables are present, PC or NOTEARS under causal sufficiency may return misleading directions. Include FCI, RFCI, or latent-variable-aware methods, and report whether the benchmark violates causal sufficiency.

    Using unrealistic data splits

    Random splits can leak temporal, geographical, or environmental information. Causal mechanisms may shift between training and deployment, so evaluate across environments whenever possible.

    Reporting only the best run

    Optimization-based methods can vary substantially across seeds. Report mean, standard deviation, confidence intervals, and the number of successful runs.

    A Practical Baseline Recipe

    For a new tabular causal discovery project, start with this reproducible sequence:

    1. Define the target graph type and causal estimand.
    2. Audit assumptions: acyclicity, causal sufficiency, faithfulness, noise model, and temporal ordering.
    3. Run PC and GES as classical references.
    4. Add FCI or RFCI if hidden confounding is plausible.
    5. Add NOTEARS or a comparable continuous method for optimization-based comparison.
    6. Add a nonlinear or domain-specific method if the data-generating process requires it.
    7. Evaluate SHD, SID, adjacency/orientation F1, runtime, and stability.
    8. Test multiple seeds, graph densities, sample sizes, and perturbations.
    9. Validate important intervention queries independently.
    10. Publish code, configurations, and exact preprocessing details.

    This recipe is a starting point, not a universal leaderboard. The strongest baseline is the one that reflects the scientific question, data-generating process, and deployment risk.

    Frequently Asked Questions

    Which causal discovery baseline should I use first?

    Use PC and GES as complementary classical references, then add NOTEARS for continuous optimization. Add FCI or RFCI when latent confounding is plausible and a temporal method for time-series data.

    Is NOTEARS better than PC or GES?

    Not universally. NOTEARS may be convenient for differentiable nonlinear models, while PC can be competitive when independence tests are appropriate and GES can perform well under a suitable score. Results depend on assumptions, sample size, scaling, and tuning.

    What is the best metric for causal discovery?

    Use multiple metrics. SHD measures graph edits, SID assesses intervention-related correctness, and precision/recall describe edge recovery. For applications, include treatment-effect or intervention-distribution error.

    Can causal discovery baselines be used on observational data?

    Yes, but observational data identify causal structure only under assumptions. State assumptions about confounding, faithfulness, functional form, interventions, and temporal order, and test robustness when those assumptions fail.

    How many baselines are enough for a research paper?

    There is no fixed number. A defensible set usually spans at least two major methodological families, includes an assumption-matched method, and contains a domain-specific or robustness baseline where relevant. Explain exclusions rather than optimizing for a large list.

    Apply for AI Grants India

    Building a causal AI product or research system in India? Apply for support through AI Grants India and connect your technical work with funding opportunities, mentorship, and ecosystem access.

    Last updated 24 September 2026

AIGI may be inaccurate. Replies seeded from the guide above.