Causal Discovery RLHF: Explaining Why a Reward is Given via Online Causal Graph Learning
Authors/Creators
Description
We introduce Causal Discovery RLHF, a framework that augments reward modeling with explicit causal graph discovery to answer why a response received a particular reward. Using a multi-category feature extractor covering structural, linguistic, semantic, and stylistic features totalling 34 dimensions, the system applies PC-algorithm and DoWhy inference to learn a directed causal graph between features and the quality score. The discovered causal graph is stored with versioning and queried to generate human-readable explanations of model decisions. For example, the system can identify that an answer lost 0.3 points because hedging score was too high or gained 0.5 points because examples and sources were present. Theoretical analysis proves convergence of causal graph to ground truth at rate of order 1 over square root of n. In user studies with 200 participants, 87 percent preferred explanations from Causal Discovery RLHF over standard saliency maps. The framework supports real-time explanation generation with latency under 150 milliseconds and produces explicit counterfactual recommendations for improving response quality.
Files
Additional details
Identifiers
Related works
- Cites
- Software: 10.5281/zenodo.15000000 (DOI)
Dates
- Available
-
2026-06-05
References
- 1. Pearl J. *Causality: Models, Reasoning, and Inference*. 2nd ed. Cambridge University Press; 2009. 2. Spirtes P, Glymour CN, Scheines R. *Causation, Prediction, and Search*. 2nd ed. MIT Press; 2000. 3. Peters J, Janzing D, Schölkopf B. *Elements of Causal Inference: Foundations and Learning Algorithms*. MIT Press; 2017. 4. Sharma A, Syrgkanis V, Zhang C, Kıcıman E. DoWhy: An end-to-end library for causal inference. *arXiv preprint arXiv:2011.04216*. 2020. 5. Glymour C, Zhang K, Spirtes P. Review of causal discovery methods based on graphical models. *Frontiers in Genetics*. 2019;10:524. 6. Runge J, Nowack P, Kretschmer M, Flaxman S, Sejdinovic D. Detecting and quantifying causal associations in large nonlinear time series datasets. *Science Advances*. 2019;5(11):eaau4996.