TMLR 2023 · Causal Learning

Learning Causal Discovery

Learn to discover causality inside a large complex system without human prior — outperforming human-designed, domain-agnostic methods on the MOS 6502 microprocessor, the NetSim fMRI dataset and the Dream3 gene dataset.

Xinyue Wang · Konrad Kording
Learning causal discovery in a data-driven way.
Learning causal discovery in a data-driven way. Rather than designing a discovery algorithm, we learn one from systems whose ground-truth causal structure is known.

Abstract

Causal discovery from time-varying data is important in neuroscience, medicine and machine learning. Techniques encompass randomized experiments, which are generally unbiased but expensive, and algorithms such as Granger causality, conditional-independence-based, structural-equation-based and score-based methods that are only accurate under strong assumptions made by human designers.

However, as demonstrated in other areas of machine learning, human expertise is often not entirely accurate and tends to be outperformed in domains with abundant data. In this study we examine whether we can enhance domain-specific causal discovery for time series using a data-driven approach.

Our findings indicate that this procedure significantly outperforms human-designed, domain-agnostic causal discovery methods — such as Mutual Information, VAR-LiNGAM and Granger Causality — on the MOS 6502 microprocessor, the NetSim fMRI dataset and the Dream3 gene dataset. We argue that, when feasible, the causality field should consider a supervised approach in which domain-specific procedures are learned from extensive datasets with known causal relationships, rather than being designed by human specialists.

Results

The learned procedure holds up across games and durations on the microprocessor benchmark, and against the classical baselines it was compared with.

Methods comparison.
Methods comparison. Across different games and durations.
Classical baselines.
Classical baselines. Domain-agnostic methods for reference.