Switch language한국어
Back to the list

Rewarding Beliefs, Not Actions: Consistency-Guided Credit Assignment for Long-Horizon Agents

TL;DR AI

Key summary

2 min read
  1. Researchers introduced ReBel, a process-level reinforcement learning method for long-horizon LLM agents in partially observable tasks.

  2. ReBel models structured beliefs, adds belief-consistency supervision, and groups trajectories by similar belief states to make training signals denser and less noisy.

  3. On ALFWorld and WebShop, it reportedly beats the episode-level GRPO baseline by up to 20.4 percentage points and improves sample efficiency by 2.1x.

  4. The approach addresses delayed-reward credit assignment without requiring step-level labels, making it more practical for long-horizon agents.

Read the original