Switch language한국어
Back to the list

Native Audio-Visual Alignment for Generation

TL;DR AI

Key summary

2 min read
  1. Researchers introduced NAVA, a native audio-visual alignment framework for joint audio-video generation.

  2. Its Align-then-Fuse MMDiT design first aligns audio and video, then applies context-conditioned denoising for better sync and control.

  3. Timbre-in-Context Conditioning helps match speech timbre to reference cues more accurately.

  4. On Verse-Bench and Seed-TTS, NAVA delivered stronger synchronization, competitive audio quality, and improved controllability with 6.3B parameters.

Read the original