Causal discovery benchmarks are essential for testing whether an algorithm can recover cause-and-effect structure from observational or interventional data. They help researchers compare methods across graph sizes, noise models, data types, and levels of prior knowledge—but only when the benchmark is designed and interpreted carefully.
For Indian AI researchers, startups, and academic teams, benchmark selection also matters because compute budgets, access to real-world data, and regulatory constraints can shape what is feasible. This guide explains the main benchmark families, evaluation metrics, datasets, experimental protocols, and common failure modes.
What Are Causal Discovery Benchmarks?
A causal discovery benchmark is a standardized evaluation setup for algorithms that infer a causal graph from data. A typical benchmark specifies:
- A data-generating process or dataset
- The ground-truth causal graph, when available
- The discovery task, such as DAG learning or ancestral graph recovery
- Allowed inputs, including observational data, interventions, or background knowledge
- Evaluation metrics
- Compute and sample-size constraints
- Reporting requirements for reproducibility
The goal is not simply to identify the model with the highest score. A useful benchmark reveals when a method succeeds, how performance changes as conditions become harder, and whether the claimed causal structure is identifiable from the available data.
Causal discovery differs from prediction. A predictive model can achieve excellent accuracy while learning associations that do not represent causal relationships. Benchmarks therefore need to measure graph recovery, intervention quality, or causal effect estimation—not only downstream predictive performance.
Main Causal Discovery Benchmark Categories
Synthetic benchmarks
Synthetic data remains the most controlled option because the true graph and data-generating mechanism are known. Researchers can vary:
- Number of variables and edges
- Sparsity and degree distribution
- Linear versus nonlinear mechanisms
- Gaussian, non-Gaussian, count, binary, or mixed variables
- Homoscedastic versus heteroscedastic noise
- Hidden confounding
- Measurement error and missing values
- Observational versus interventional samples
Common synthetic graph families include Erdős–Rényi, scale-free, small-world, and application-inspired DAGs. Synthetic benchmarks are valuable for ablation studies, but they can overstate performance when their assumptions match an algorithm too closely.
Semi-synthetic benchmarks
Semi-synthetic benchmarks combine a real graph or real covariate distribution with simulated mechanisms. For example, a known biological network may be used as the graph while synthetic observations are generated under controlled interventions.
These benchmarks provide more realistic structure than purely random graphs while preserving access to ground truth. They are especially useful when real observational data lacks a fully verified causal graph.
Real-world benchmarks
Real-world benchmarks use datasets from biology, medicine, economics, climate science, manufacturing, or public policy. Their main advantage is ecological validity. Their main limitation is that the proposed ground truth is often incomplete, disputed, or based on expert consensus rather than exhaustive intervention experiments.
A real dataset should not be treated as definitive evidence merely because it comes from a respected domain. Researchers should document how the reference graph was created, which edges are uncertain, and whether the data-generating process changed over time.
Interventional benchmarks
Interventional benchmarks test whether a method can use data collected under controlled or semi-controlled interventions. These settings are important because interventions can resolve Markov-equivalent structures that observational data alone cannot distinguish.
A rigorous interventional benchmark should specify intervention targets, intervention strength, soft versus hard interventions, sample allocation, and whether the method receives intervention metadata.
Important Benchmark Datasets and Suites
No single dataset is sufficient for evaluating causal discovery. A strong study combines several benchmark families.
Sachs protein-signaling data
The Sachs dataset is a widely used example in causal discovery and biological network reconstruction. It contains flow-cytometry measurements with intervention-related information and is often used to compare recovered signaling relationships.
It is useful for testing small-scale biological discovery, but it should not be interpreted as a universally complete ground truth. Preprocessing choices, discretization, hidden variables, and edge-definition conventions can materially change results.
DREAM and gene-regulatory network challenges
DREAM challenges popularized systematic evaluation of gene-regulatory network inference. These datasets often combine simulated and biological settings, including perturbation information and network reconstruction tasks.
They are relevant for methods handling high-dimensional variables, sparse graphs, and intervention-like perturbations. Researchers should verify the exact challenge version and scoring protocol because datasets and evaluation rules differ across challenges.
CausalTime and temporal datasets
Time-series causal discovery requires benchmarks that preserve temporal ordering, lagged dependencies, autocorrelation, and possible regime changes. Suitable datasets may come from climate systems, industrial sensors, finance, or physiology.
Temporal benchmarks should distinguish contemporaneous edges from lagged edges. A method that discovers a statistically useful temporal dependency is not automatically identifying a valid structural causal relationship.
CauseMe and related time-series evaluation platforms
CauseMe-style platforms provide standardized time-series causal discovery tasks across synthetic and real-world systems. They can help compare algorithms under nonlinear dynamics, confounding, synchronization, and varying sample sizes.
When using these platforms, report the data-generation setting, allowed lags, preprocessing, stationarity assumptions, and whether hyperparameters were tuned on test instances.
Large-scale synthetic DAG suites
Modern benchmarking increasingly uses large synthetic suites with thousands of variables or many graph instances. These tests evaluate scalability, memory use, runtime, and degradation under high dimensionality.
Large graphs are valuable for engineering evaluation, but graph-recovery metrics can become misleading when the graph is extremely sparse. Always report both absolute edge counts and normalized metrics.
Causal Discovery Algorithms to Compare
A benchmark should include methods from different causal assumptions rather than comparing only variants of one model family.
Constraint-based methods
Algorithms such as PC and FCI use conditional independence tests. PC generally targets causal sufficiency under assumptions including acyclicity and faithfulness. FCI is designed to handle latent confounding and can return a partial ancestral graph.
Their performance depends strongly on sample size, conditioning-set complexity, variable type, and the quality of independence tests.
Score-based methods
GES and related approaches search over graph structures using a statistical score. Score-based methods can perform well when the score and data-generating assumptions are appropriate, but search complexity and local optima can become important in larger graphs.
Functional causal model methods
LiNGAM exploits non-Gaussianity, while additive-noise and post-nonlinear methods use asymmetries in the data-generating mechanism. These methods may orient edges that are ambiguous under purely observational Gaussian models, but only when their structural assumptions are credible.
Continuous optimization methods
NOTEARS and its successors formulate DAG learning as a continuous optimization problem with a differentiable acyclicity constraint. These methods are convenient for integration with neural networks and complex likelihoods, but regularization, thresholding, optimization stability, and scalability require careful testing.
Neural and deep causal discovery methods
Deep models can represent nonlinear relationships, sequences, images, and multimodal data. However, a neural architecture does not by itself guarantee causal identification. Benchmarks should separate representation quality from graph correctness and include strong non-neural baselines.
Hybrid and knowledge-informed methods
Hybrid algorithms combine statistical discovery with expert constraints, temporal ordering, known directions, or forbidden edges. These methods can be highly practical in healthcare, agriculture, finance, and industrial settings, but benchmarks must report the amount and reliability of prior knowledge supplied.
Core Evaluation Metrics
Structural Hamming Distance
Structural Hamming Distance, or SHD, counts graph-edit operations required to transform the estimated graph into the reference graph. Depending on the convention, a reversed edge may count as one or two operations.
Because conventions differ, define SHD explicitly and report the treatment of reversed and extra edges.
Structural Intervention Distance
Structural Intervention Distance evaluates differences in intervention consequences rather than only edge-level errors. It can better reflect whether a discovered graph supports correct decision-making, especially when some graph errors have little effect on downstream interventions.
Precision, recall, and F1
For directed edges:
- Precision measures the fraction of predicted edges that are correct.
- Recall measures the fraction of true edges that are recovered.
- F1 summarizes the two.
Report results separately for skeleton recovery and edge orientation. A method may identify adjacency accurately while orienting very few edges correctly.
Area under precision–recall curves
Many algorithms produce continuous edge scores or a regularization path. Precision–recall curves are often more informative than a single thresholded graph, especially for sparse networks. The threshold-selection rule must be fixed before evaluating the test set.
Markov equivalence-aware metrics
Observational data may identify a Markov equivalence class rather than one unique DAG. Penalizing a method for failing to choose an unidentifiable orientation is inappropriate. Use metrics designed for CPDAGs or partial ancestral graphs when the task only identifies equivalence classes.
Causal effect and intervention metrics
If the scientific goal is decision support, evaluate estimated intervention distributions, average treatment effects, policy value, or counterfactual accuracy. These metrics can reveal whether a graph that is imperfect structurally is still useful for the intended intervention task.
Runtime and resource use
Include wall-clock time, peak memory, hardware, parallelization, and failed runs. For Indian startups and research teams operating under constrained compute, efficiency can be as important as a small improvement in SHD.
Designing a Reliable Benchmark Protocol
Define the estimand and graph target
State whether the target is a DAG, CPDAG, PAG, dynamic graph, causal ordering, or intervention distribution. Do not compare outputs with incompatible semantics without a conversion procedure.
Match assumptions to data
Document assumptions about:
- Causal sufficiency
- Acyclicity
- Faithfulness
- No selection bias
- Stationarity
- Additive noise
- Variable measurement quality
- Intervention consistency and positivity
A benchmark should include both matched and violated-assumption settings. Reporting only favorable conditions produces inflated claims.
Use independent graph instances and seeds
One graph can produce a misleading result. Generate or select multiple graph instances, use multiple random seeds, and report mean, median, variance, and confidence intervals. Paired comparisons across the same instances are often more informative than comparing aggregate averages.
Separate tuning from testing
Hyperparameters should be selected using validation data, held-out graph instances, or a predeclared rule. Test graphs must remain untouched until final evaluation. This is particularly important for thresholding continuous adjacency scores.
Conduct scaling experiments
Vary one difficulty dimension at a time where possible:
- Sample size
- Number of variables
- Expected degree
- Noise level
- Nonlinearity
- Confounding strength
- Missingness
- Intervention coverage
Plot performance curves rather than reporting only one operating point.
Test robustness and distribution shift
Real deployments rarely match benchmark assumptions. Add perturbations such as measurement noise, missing-not-at-random data, changing mechanisms, outliers, temporal drift, and spurious proxy variables.
For Indian applications, useful stress tests may include multilingual or heterogeneous populations, district-level distribution shift, seasonal agricultural data, hospital-specific measurement practices, and changes caused by policy interventions.
Common Benchmarking Mistakes
- Treating an expert graph as complete ground truth
- Comparing a PAG directly with a DAG without defining equivalence
- Reporting only accuracy while hiding false discoveries
- Tuning hyperparameters on the test set
- Using a single synthetic graph or random seed
- Ignoring preprocessing and variable transformations
- Giving neural methods more compute or tuning effort than baselines
- Comparing runtime across different hardware without disclosure
- Evaluating only observational data when the claim concerns interventions
- Assuming correlation or predictive improvement proves causality
A credible paper or product evaluation should make the full benchmark configuration reproducible, including data versions, preprocessing code, software versions, seeds, hardware, and failed-run handling.
A Practical Benchmark Checklist
Before publishing causal discovery results, verify that you have:
- Defined the causal target and estimand
- Selected synthetic, semi-synthetic, and real-world tasks
- Included suitable baseline families
- Reported skeleton and orientation performance separately
- Used metrics compatible with partial identification
- Tuned parameters without test leakage
- Repeated experiments across graphs and seeds
- Reported uncertainty and computational cost
- Tested assumption violations and distribution shift
- Released code, configurations, and benchmark metadata
How to Choose a Benchmark for an AI Project
Start with the decision the system must support. If the objective is scientific network reconstruction, prioritize graph and uncertainty metrics. If the objective is treatment or policy selection, add intervention and counterfactual evaluation. If the product must operate online, include temporal drift, latency, and resource constraints.
For an early-stage Indian AI startup, a practical sequence is:
1. Validate the method on controlled synthetic graphs.
2. Test robustness using semi-synthetic domain structures.
3. Evaluate on one or more real datasets with carefully documented uncertainty.
4. Run a domain-expert review of discovered relationships.
5. Measure intervention or decision value before claiming business impact.
This staged approach prevents a benchmark score from being mistaken for deployment readiness.
FAQ: Causal Discovery Benchmarks
What is the best causal discovery benchmark?
There is no universal best benchmark. Use a portfolio combining synthetic graphs for controlled ground truth, semi-synthetic data for realism, and real or interventional datasets for domain validity.
Which metric is most important?
It depends on the task. SHD and precision–recall are useful for graph recovery, while intervention distance, causal-effect error, or policy value may be more relevant for decision-making systems.
Can causal discovery be evaluated on observational data alone?
Yes, but identifiability is limited. Observational data often supports only a Markov equivalence class, and hidden confounding can make the problem substantially harder. Claims should match what the data can identify.
Are synthetic benchmarks reliable?
They are reliable for testing behavior under known assumptions, not for proving real-world performance. Use varied mechanisms and include assumption-violation tests to reduce benchmark bias.
How should startups report causal discovery results?
Report the graph target, datasets, preprocessing, assumptions, metrics, baselines, tuning procedure, hardware, uncertainty, and failure cases. Also show whether the discovered relationships improve a real intervention or decision task.
Apply for AI Grants India
Building an AI system for causal inference, scientific discovery, healthcare, agriculture, or public-impact applications? Apply to AI Grants India for support and opportunities designed for Indian AI founders.