ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning
TL;DR AI
2 min readKey summary
Researchers introduced ParaVT, the first end-to-end RL-trained multi-agent system for parallel video tool calling.
ParaVT uses PARA-GRPO to address two training failures from pretrained tool priors: unstable tool-call formatting and weak incentives to actually use tools.
The framework is designed for long-video understanding, where parallel tool use can reduce sequential-call inefficiency and improve reasoning accuracy.
The work appears alongside LongVT-related research and is available via arXiv, GitHub, and Hugging Face.
