Causal discovery is the task of inferring cause-and-effect structure from data, often represented as a directed acyclic graph (DAG), a partially directed graph, or an equivalence class of graphs. As AI systems increasingly claim to discover causal relationships, AI causal discovery benchmarks provide the discipline needed to distinguish genuine structural learning from correlation fitting, memorisation, or favourable experimental design.
A useful benchmark must test more than prediction accuracy. It should measure whether a method recovers the correct edges, respects statistical assumptions, handles hidden variables and cycles, scales to realistic datasets, and produces results that can be reproduced. This guide explains how to evaluate causal discovery systems and how researchers, enterprises, and Indian AI startups can build credible benchmark studies.
What Are AI Causal Discovery Benchmarks?
AI causal discovery benchmarks are standardised datasets, tasks, metrics, and protocols used to compare algorithms that infer causal structure. A benchmark may provide:
- Synthetic data generated from a known causal graph, where the ground truth is available exactly.
- Semi-synthetic data created by combining real covariates with a simulated or partially known mechanism.
- Real-world data from domains such as genomics, climate science, healthcare, economics, and industrial systems.
- Interventional data collected after deliberately changing one or more variables.
- Time-series data for discovering lagged, contemporaneous, or feedback relationships.
The central difficulty is that causal ground truth is rarely observable in ordinary observational datasets. As a result, strong evaluation usually combines synthetic experiments, carefully curated real data, and domain validation rather than relying on one leaderboard score.
Why Benchmarking Causal Discovery Is Difficult
Causal discovery is more assumption-sensitive than standard supervised learning. Two algorithms can produce different graphs because they make different assumptions about noise, sparsity, faithfulness, latent confounding, or the presence of cycles.
Important sources of difficulty include:
- Observational identifiability: Many causal graphs imply the same conditional independencies. Data may identify only a Markov equivalence class, not one unique DAG.
- Unmeasured confounding: Hidden common causes can make a direct causal edge appear where none exists.
- Finite samples: Conditional independence tests become unreliable as the number of variables grows.
- Distribution shift: A graph learned in one population may not transfer to another.
- Cycles and feedback: Many algorithms assume acyclicity, while biological, economic, and industrial systems often contain feedback loops.
- Mixed data types: Count, binary, categorical, continuous, missing, and censored variables require different modelling choices.
- Selection bias: Data collected only from a selected population can distort both dependencies and interventions.
A benchmark that ignores these issues may reward methods for exploiting artefacts rather than discovering robust causal structure.
Main Families of Causal Discovery Methods to Benchmark
Constraint-based methods
Constraint-based algorithms infer graph structure from conditional independence tests. The PC algorithm is a classic example for causally sufficient, acyclic settings. Variants such as FCI are designed to account for latent confounding and selection bias.
Benchmark considerations include:
- The conditional independence test used.
- Significance-level selection and multiple-testing behaviour.
- Sensitivity to sample size and graph density.
- Runtime as the number of variables increases.
- Whether the output is a DAG, CPDAG, or PAG.
Score-based methods
Score-based approaches search for a graph that optimises a criterion such as BIC, BDeu, or a differentiable acyclicity objective. Greedy equivalence search, hill climbing, and continuous optimisation methods fall into this broad category.
These methods should be tested for search stability, local optima, computational cost, and sensitivity to score misspecification. A high training score does not necessarily mean that the recovered graph is causally correct.
Functional and noise-based methods
Some algorithms exploit assumptions about the data-generating mechanism, such as additive noise, non-Gaussian noise, or asymmetric cause-effect relationships. Examples include LiNGAM-style models and additive-noise approaches.
Such methods may identify direction more strongly than purely conditional-independence-based approaches, but their assumptions must be tested explicitly. Benchmarks should include both matched and mismatched data-generating processes.
Neural and differentiable methods
Neural causal discovery systems often represent an adjacency matrix and impose an acyclicity constraint during optimisation. NOTEARS-inspired methods, graph neural networks, variational approaches, and reinforcement-learning systems are common examples.
Evaluation should report:
- Random seeds and the number of restarts.
- Optimisation failures and constraint violations.
- Regularisation settings.
- Hardware and wall-clock time.
- Performance as dimensionality and graph density increase.
- Robustness to scaling, noise, and missing values.
Temporal and dynamical methods
Time-series causal discovery methods use lagged observations, temporal ordering, or dynamical system assumptions. Common families include Granger-causality methods, VAR-based models, PCMCI-style methods, temporal constraint-based algorithms, and neural dynamical models.
A temporal benchmark must distinguish lagged causation from contemporaneous association and should specify sampling frequency, stationarity assumptions, intervention timing, and whether instantaneous effects are allowed.
Core Datasets for AI Causal Discovery Benchmarks
Synthetic DAG benchmarks
Synthetic benchmarks remain indispensable because the true graph is known. A robust suite should vary:
- Number of nodes, such as 10, 50, 100, and 1,000.
- Expected degree and graph sparsity.
- Erdős–Rényi, scale-free, and custom graph topologies.
- Linear and nonlinear structural equations.
- Gaussian, non-Gaussian, heteroscedastic, and heavy-tailed noise.
- Sample size and measurement noise.
- Latent variables, missingness, and interventions.
However, synthetic data can be too easy if the benchmark generator matches the assumptions of the algorithm. For example, evaluating a linear-Gaussian method only on linear-Gaussian data measures implementation quality more than general causal discovery ability.
Sachs protein-signalling data
The Sachs dataset is widely used in causal discovery because it contains protein-signalling measurements and known or partially validated biological relationships. It also includes intervention conditions, making it more informative than a purely observational dataset.
Researchers should avoid treating its published graph as unquestionable ground truth. Biological knowledge may be incomplete, measurements may be noisy, and the experimental design can impose selection effects.
DREAM challenges and gene regulatory data
DREAM challenge datasets and gene regulatory network benchmarks are useful for evaluating network reconstruction. They often provide simulated or experimentally informed regulatory relationships, allowing methods to be compared at different scales.
These datasets are especially valuable for testing sparse high-dimensional discovery, but metrics should account for incomplete ground truth. Missing edges in a reference network are not necessarily true negatives.
CauseMe and causal time-series suites
CauseMe-style resources provide simulated time-series tasks with known causal relationships and different dynamical regimes. They are useful for comparing temporal discovery methods under varying noise, nonlinearity, confounding, and sample-length conditions.
When using time-series benchmarks, report whether the model receives the correct temporal order and whether variables are sampled synchronously. These choices can materially affect results.
Real-world and domain-specific data
Real-world datasets can expose failure modes hidden by simulation. Relevant sources include:
- Electronic health records and longitudinal patient data.
- Climate and weather sensor networks.
- Manufacturing telemetry and predictive-maintenance logs.
- Energy demand and smart-grid measurements.
- Financial and macroeconomic time series.
- Agricultural, soil, and crop-monitoring data.
For Indian applications, potential domains include public-health surveillance, agriculture, air-quality monitoring, digital payments, logistics, manufacturing, and language or speech systems. Data governance is critical: de-identification, consent, access controls, and compliance with applicable Indian privacy requirements should be part of the benchmark design.
Metrics for Comparing Causal Discovery Systems
No single metric captures all aspects of graph quality. Use several complementary measures.
Structural Hamming Distance
Structural Hamming Distance, or SHD, counts graph-edit operations needed to transform the estimated graph into the reference graph. Depending on the implementation, an incorrect orientation may count as one or two errors.
SHD is intuitive but can penalise uncertain orientations when only a Markov equivalence class is identifiable. Always define the exact convention used.
Structural Intervention Distance
Structural Intervention Distance, or SID, evaluates whether the estimated graph would produce correct interventional distributions under graph-based adjustment. It is often more causally meaningful than edge-wise comparison, particularly when some graph differences do not change intervention predictions.
SID depends on the estimated and reference graph assumptions, so it should be reported alongside other measures rather than used in isolation.
Precision, recall, and F1 score
For edge recovery, report:
- Adjacency precision: the fraction of predicted adjacencies that are correct.
- Adjacency recall: the fraction of true adjacencies recovered.
- Orientation precision and recall: how accurately edge directions are identified.
- F1 score: the harmonic mean of precision and recall.
For sparse graphs, precision-recall curves are often more informative than accuracy because true non-edges dominate the comparison.
SHD, SID, and CPDAG-aware evaluation
If a method returns a CPDAG or PAG, compare it using an evaluation procedure that respects partially oriented edges. Penalising a method for refusing to orient an unidentifiable edge can create a misleading ranking.
A strong report should distinguish:
- Correct skeleton recovery.
- Correct compelled orientations.
- Correctly oriented but non-identifiable edges.
- False orientation commitments.
Interventional prediction quality
When intervention data is available, evaluate the predicted effect of interventions. Useful metrics include mean squared error, log loss, calibration error, and rank correlation for treatment effects or downstream outcomes.
This provides a practical test: does the discovered graph support better decisions under intervention, policy change, or process control?
Stability and reproducibility
Measure how often an edge appears across bootstrap samples, random seeds, and reasonable hyperparameter settings. A graph with slightly lower SHD but highly unstable edges may be less useful than a stable graph with modestly higher error.
Report confidence intervals, not only a single average. Include failed runs and computational-resource requirements.
A Reproducible Benchmark Protocol
A credible benchmark should follow a documented protocol.
1. Define the estimand. State whether the objective is graph recovery, causal effect estimation, intervention prediction, or scientific hypothesis generation.
2. Declare assumptions. Specify acyclicity, causal sufficiency, faithfulness, stationarity, linearity, and noise assumptions.
3. Separate data splits correctly. Avoid ordinary random splits when temporal or grouped leakage is possible.
4. Standardise preprocessing. Document scaling, imputation, feature selection, discretisation, and outlier treatment.
5. Tune fairly. Use validation data or pre-registered settings, and do not tune against the hidden test graph.
6. Use multiple seeds. Neural and stochastic methods should receive enough restarts to estimate variance.
7. Report compute. Include CPU/GPU type, memory, runtime, parallelism, and stopping criteria.
8. Evaluate failure modes. Test missing data, latent confounding, measurement error, distribution shift, and graph misspecification.
9. Release artefacts. Provide code, configuration files, environment specifications, processed-data hashes, and result logs.
Containerised environments and automated experiment tracking make it easier for other teams to reproduce results. For sensitive Indian datasets, release synthetic substitutes, feature schemas, evaluation code, and access procedures even when raw data cannot be published.
Common Benchmarking Mistakes
Treating correlation recovery as causal discovery
A model can reproduce a dependency network without identifying direction. Distinguish association, prediction, causal orientation, and intervention validity.
Reporting only one synthetic setting
A method may perform well under the exact assumptions used to generate its training data. Vary functional form, noise, confounding, sparsity, and sample size.
Ignoring the output graph type
DAGs, CPDAGs, PAGs, mixed graphs, and cyclic graphs require different comparison procedures. Converting every result into a fully directed DAG can introduce artificial errors.
Hiding computational failures
Out-of-memory errors, non-convergence, invalid graphs, and timeouts are part of practical performance. Report them transparently instead of excluding them silently.
Using incomplete real-world ground truth as absolute truth
A reference network is usually incomplete and uncertain. Use expert review, interventional evidence, confidence-weighted evaluation, and sensitivity analysis where appropriate.
How to Choose a Benchmark for Your AI Project
Select benchmarks according to the intended deployment setting:
- For research on graph recovery, use synthetic DAGs with controlled assumptions and CPDAG-aware metrics.
- For healthcare, combine longitudinal observational data with clinical knowledge and intervention or treatment-effect evaluation.
- For industrial AI, prioritise temporal data, regime changes, sensor faults, and actionable intervention tests.
- For agriculture and climate, evaluate spatial-temporal dependence, missing sensors, seasonal shifts, and policy or intervention scenarios.
- For early-stage startups, begin with a reproducible small benchmark suite before investing in a large proprietary dataset.
A practical minimum suite includes one synthetic linear benchmark, one nonlinear benchmark, one latent-confounding scenario, one temporal dataset, and one domain-relevant real-world dataset.
FAQ: AI Causal Discovery Benchmarks
What is the best benchmark for causal discovery?
There is no universal best benchmark. Use a combination of synthetic datasets with known graphs, intervention-aware datasets such as Sachs, temporal suites, and a domain-specific real-world dataset.
Which metrics should I report?
At minimum, report SHD, adjacency precision and recall, orientation performance, SID where applicable, runtime, failure rate, and variability across seeds or bootstrap samples.
Are synthetic datasets enough?
No. Synthetic data enables exact ground-truth evaluation, but real-world data reveals measurement error, missing variables, distribution shift, and operational constraints. Both are necessary.
Can large language models perform causal discovery?
Language models can help formulate hypotheses, select variables, interpret graphs, and integrate scientific literature. Their causal claims still require statistical testing, domain validation, and preferably interventional evidence.
How can an Indian AI startup build a credible benchmark?
Define the causal question, document assumptions, use privacy-safe data practices, compare against strong baselines, run multiple seeds, measure intervention usefulness, and publish reproducible evaluation artefacts.
Apply for AI Grants India
Building a causal AI system for healthcare, agriculture, climate, manufacturing, or another high-impact sector? Apply through AI Grants India to explore support and opportunities for Indian AI founders.