Switch language한국어
Back to the list

VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System

TL;DR AI

Key summary

2 min read
  1. VideoCoCo is a new text-to-video system that turns prompts into executable Blender code as an intermediate reasoning step.

  2. A simulator produces a deterministic draft video, and a generative editor refines it into a photorealistic final output.

  3. The team also released VideoCoCo-3K to train the editor on draft-instruction-target pairs and reported gains over a baseline on PhyGenBench and VBench-2.0.

  4. The approach improves temporal and physical consistency by making scene dynamics explicit, executable, and easier to inspect.

Read the original