Alibaba's HDPO cuts AI agent tool overuse from 98% to 2%

TL;DR AI
2 min readKey summary
Alibaba researchers introduced Hierarchical Decoupled Policy Optimization (HDPO) to train AI agents with separate accuracy and efficiency objectives.
Their multimodal Metis model cut redundant tool invocations from 98% to 2%, showing far less tool overuse.
Metis also achieved state-of-the-art reasoning results on key benchmarks, suggesting the method can improve both performance and deployment efficiency.
The work targets a major AI-agent pain point: deciding when to rely on tools versus internal knowledge to reduce latency, cost, and failures.
