A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation

TL;DR AI
2 min readKey summary
Researchers introduced VIABLE, a benchmark with 300,000+ judgment samples for three visually impaired assistance scenarios.
Tests of seven VLM judges found weak reliability in effectiveness, impartiality, and stability, plus strong self-preference and adversarial vulnerability.
The study also proposed VIA-Judge-Agent, an inference-time harness that improves judge performance and downstream assistance responses.
The findings raise concerns about using current vision-language model judges in sensitive accessibility settings for BLV users.
