One Model, Three Modalities: ByteDance Releases Lance for Image and Video Understanding, Generation, and Editing

TL;DR AI
2 min readKey summary
ByteDance researchers introduced Lance, a single multimodal model for image and video understanding, text generation, and visual generation/editing.
Lance uses shared multimodal context with separate understanding and generation paths, plus a new modality-aware positional encoding called MaPE.
The system aims to cover the full image-video pipeline, including text-to-image, text-to-video, image editing, and video editing.
Its broader scope could simplify multimodal AI stacks and improve consistency between understanding and generation tasks.
