Bernini: Latent Semantic Planning for Video Diffusion
TL;DR AI
2 min readKey summary
Researchers introduced Bernini, a unified video generation and editing framework.
It uses an MLLM to plan high-level video semantics, then a diffusion renderer to synthesize frames.
The two parts can be trained mostly separately, improving efficiency and modularity.
A segment-aware 3D positional encoding helps the system handle video structure better.
Bernini reports state-of-the-art results on several video generation and editing benchmarks.
