Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember

TL;DR AI
2 min readKey summary
Researchers introduced SESA, a self-evolving search agent framework where problem generation and skill memory co-develop.
In SESA, one agent creates questions while another solves them using retrieved procedural skills in a self-play loop.
Failures are turned into new skills and written back to memory, which then shapes future training data and policy learning.
Across seven open-domain and multi-hop QA benchmarks, SESA outperformed prior self-play and skill-augmented baselines.
The paper suggests external skill memory can improve not just inference, but also training dynamics and future task generation.
