Transformer Is Inherently a Causal Learner
We reveal that transformers trained autoregressively naturally encode causal structures — gradient attributions directly recover underlying causal graphs without any explicit causal objectives.


My recent work spans causal learning from time series at scale, causality-guided world modeling for reinforcement learning generalization, and agentic systems for end-to-end causal analysis. Before joining UCSD, I worked with Dr. Konrad Kording on meta-learning methods on domain-specific causal discovery for large complex systems (e.g., microprocessors). I am also interested in brain-computer interfaces and computational neuroscience, and previously worked on real-time neurofeedback systems advised by Dr. Gan Huang.
We reveal that transformers trained autoregressively naturally encode causal structures — gradient attributions directly recover underlying causal graphs without any explicit causal objectives.

An LLM-powered agent that runs the whole causal analysis loop — diagnosing the data, selecting and configuring the right method from 20+ options, checking its own results, and producing an inspectable report.

A reinforcement learning framework that generalizes to unseen environments by learning language-controlled causal components and recombining them — with identifiability guarantees.

Instead of designing a causal discovery algorithm, we learn one — from a microprocessor whose every causal edge can be established by intervention. It outperforms human-designed methods on silicon, simulated fMRI and gene networks.

A millisecond-level phase locked neural feedback system based on OpenBCI for real-time alpha wave regulation, integrating acquisition, phase estimation and stimulation on one chip.
