Switch language한국어
Back to the list

Bernini: Latent Semantic Planning for Video Diffusion

TL;DR AI

Key summary

2 min read
  1. Researchers introduced Bernini, a unified video generation and editing framework.

  2. It uses an MLLM to plan high-level video semantics, then a diffusion renderer to synthesize frames.

  3. The two parts can be trained mostly separately, improving efficiency and modularity.

  4. A segment-aware 3D positional encoding helps the system handle video structure better.

  5. Bernini reports state-of-the-art results on several video generation and editing benchmarks.

Read the original