Switch language한국어
Back to the list

Chimera: Designing and Chinchilla-Scaling Hybrid Visual Diffusion Transformers

TL;DR AI

Key summary

2 min read
  1. Researchers introduced Chimera, a hybrid diffusion backbone for text, image, and video tokens.

  2. It combines KDA, MLA, short convolutions, and sparse MoE layers to improve efficiency for long-context generation.

  3. They also propose HeteroP, a scaling recipe that tunes width and depth, and train an 11B model with 2B activated parameters.

  4. Experiments show better compute efficiency than a full-attention baseline and strong zero-shot video length extrapolation.

Read the original