Policy and World Modeling Co-Training for Language Agents
TL;DR AI
2 min readKey summary
Researchers introduced PaW, a co-training framework that adds world-model supervision to policy training for language agents.
PaW reuses on-policy reinforcement-learning rollouts, avoiding extra simulators or inference overhead.
It combines action-entropy-based data selection, a noise-tolerant loss, and reward-adaptive loss balancing.
Across three agentic benchmarks, PaW outperformed strong RL baselines across models and algorithms.
