Switch language한국어
Back to the list

Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines

TL;DR AI

Key summary

2 min read
  1. Researchers propose a bias-free action-dependent baseline that exploits policy structure to cut policy-gradient variance without extra MDP assumptions.

  2. The method delivers both theoretical and experimental gains on standard RL, high-dimensional manipulation, and a synthetic 2000-dimensional control task.

  3. It also extends to partially observed and multi-agent settings, where lower-variance gradients can improve training stability and speed.

  4. Overall, the work strengthens policy-gradient methods for long-horizon problems and large action spaces.

Read the original