DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models

TL;DR AI
2 min readKey summary
Researchers introduced DRBench, a 14,573-question benchmark over 2,943 images, to evaluate dense-scene reasoning in vision-language models.
They also proposed DRScaffold, a four-stage supervised fine-tuning method that teaches lightweight VLMs to reason with grounded evidence.
The approach delivers strong gains on dense-scene tasks while causing little to no drop on general benchmarks.
Notably, a smaller tuned model can outperform a much larger frozen model on the new benchmark.
