RL orchestration: how a 7B model routes tasks across GPT-5, Claude, and Gemini

TL;DR AI
2 min readKey summary
Sakana AI introduced RL Conductor, a 7B model trained with reinforcement learning to break down tasks, route work to the right models, and manage information flow across multiple LLMs.
It reportedly outperforms frontier models and hand-built multi-agent systems on hard reasoning and coding benchmarks while making fewer API calls.
The system underpins Sakana’s Fugu orchestration product, pointing to a more flexible and cost-efficient alternative to rigid manual agent workflows.


