Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines

TL;DR AI
2 min readKey summary
Researchers propose a bias-free action-dependent baseline that exploits policy structure to cut policy-gradient variance without extra MDP assumptions.
The method delivers both theoretical and experimental gains on standard RL, high-dimensional manipulation, and a synthetic 2000-dimensional control task.
It also extends to partially observed and multi-agent settings, where lower-variance gradients can improve training stability and speed.
Overall, the work strengthens policy-gradient methods for long-horizon problems and large action spaces.



