Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles
TL;DR AI
2 min readKey summary
Maestro is an RL-based orchestration framework that treats multimodal tasks as sequential decisions over expert models and skills.
Trained only with outcome-based feedback, its 4B policy achieved stronger average benchmark results than GPT-5 and Gemini-2.5-Pro.
The system also generalized to unseen experts and kept computational costs low, making routing more efficient than retraining.
The result suggests a small orchestrator can outperform much larger models by learning how to coordinate specialized tools and skills.
