Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
TL;DR AI
2 min readKey summary
Researchers introduced Flat-Pack Bench, a furniture-assembly video benchmark for fine-grained spatio-temporal reasoning.
It tests temporal ordering, state localization, part mating, and tracking using multiple-choice questions and visual prompts.
Experiments showed state-of-the-art large vision-language models still perform poorly on these detailed video understanding tasks.
The benchmark highlights major gaps in LVLMs for real-world activities such as assembly and cooking.
