Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning
TL;DR AI
2 min readKey summary
Researchers introduced MARS, a mono-anchored framework for multi-source visual reasoning.
It treats each modality as an independent source and uses mono-source rewards as dynamic anchors in advantage normalization.
This helps estimate information gain, reduce cross-modal interference, and decide when extra inputs like infrared or depth actually help.
The approach improved reinforcement learning with verifiable rewards on GRPO and DAPO benchmarks.
