Switch language한국어
Back to the list

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles

TL;DR AI

Key summary

2 min read
  1. Maestro is an RL-based orchestration framework that treats multimodal tasks as sequential decisions over expert models and skills.

  2. Trained only with outcome-based feedback, its 4B policy achieved stronger average benchmark results than GPT-5 and Gemini-2.5-Pro.

  3. The system also generalized to unseen experts and kept computational costs low, making routing more efficient than retraining.

  4. The result suggests a small orchestrator can outperform much larger models by learning how to coordinate specialized tools and skills.

Read the original