FlowLong: Inference-time Long Video Generation via Manifold-constrained Tweedie Matching
TL;DR AI
2 min readKey summary
Researchers introduced FlowLong, an inference-only method for much longer video generation without extra training.
It stitches overlapping sliding-window predictions with Tweedie matching to preserve temporal consistency and visual stability.
A stochastic early-sampling phase followed by deterministic sampling helps reduce drift and improve quality over long sequences.
FlowLong is architecture-agnostic and can extend the native length of different video diffusion models, including audio-video and text-to-3DGS settings.
