Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text
TL;DR AI
2 min readKey summary
Researchers introduced ProVisE, a benchmark-agnostic framework that evaluates image-generation models on spatial tasks using pixel-space answers.
They also released SpatialGen-Bench, a 470-sample benchmark covering 14 spatial subtasks.
Results show image-generation models can perform competitively when they respond visually, while text-based VLMs still lead on more compositional spatial reasoning.
The work offers a fairer way to measure spatial cognition and clarifies the strengths of visual vs. text-based systems.
