CausalGDP:用于强化学习的因果引导扩散策略 / CausalGDP: Causality-Guided Diffusion Policies for Reinforcement Learning
1️⃣ 一句话总结
这篇论文提出了一种名为CausalGDP的新方法,它将因果推理融入基于扩散模型的强化学习中,通过识别并引导那些真正能带来高回报的关键动作,从而在复杂任务中取得了比现有方法更好的性能。
Reinforcement learning (RL) has achieved remarkable success in a wide range of sequential decision-making problems. Recent diffusion-based policies further improve RL by modeling complex, high-dimensional action distributions. However, existing diffusion policies primarily rely on statistical associations and fail to explicitly account for causal relationships among states, actions, and rewards, limiting their ability to identify which action components truly cause high returns. In this paper, we propose Causality-guided Diffusion Policy (CausalGDP), a unified framework that integrates causal reasoning into diffusion-based RL. CausalGDP first learns a base diffusion policy and an initial causal dynamical model from offline data, capturing causal dependencies among states, actions, and rewards. During real-time interaction, the causal information is continuously updated and incorporated as a guidance signal to steer the diffusion process toward actions that causally influence future states and rewards. By explicitly considering causality beyond association, CausalGDP focuses policy optimization on action components that genuinely drive performance improvements. Experimental results demonstrate that CausalGDP consistently achieves competitive or superior performance over state-of-the-art diffusion-based and offline RL methods, especially in complex, high-dimensional control tasks.
CausalGDP:用于强化学习的因果引导扩散策略 / CausalGDP: Causality-Guided Diffusion Policies for Reinforcement Learning
这篇论文提出了一种名为CausalGDP的新方法,它将因果推理融入基于扩散模型的强化学习中,通过识别并引导那些真正能带来高回报的关键动作,从而在复杂任务中取得了比现有方法更好的性能。
源自 arXiv: 2602.09207