FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

TL;DR AI
2 min readKey summary
Researchers proposed FaithEyes, a multi-agent framework that checks whether tool-generated process images are actually useful for vision-language models.
The system feeds the model’s own judgment signal back into reasoning and down-weights rewards for unhelpful tool calls.
It is trained with a two-stage pipeline of supervised fine-tuning and reinforcement learning.
The approach aims to improve tool faithfulness, reliability, and interpretability without needing an external judge at inference time.
