Switch language한국어
Back to the list

Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly

TL;DR AI

Key summary

2 min read
  1. Researchers introduced Flat-Pack Bench, a furniture-assembly video benchmark for fine-grained spatio-temporal reasoning.

  2. It tests temporal ordering, state localization, part mating, and tracking using multiple-choice questions and visual prompts.

  3. Experiments showed state-of-the-art large vision-language models still perform poorly on these detailed video understanding tasks.

  4. The benchmark highlights major gaps in LVLMs for real-world activities such as assembly and cooking.

Read the original