Switch language한국어
Back to the list

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation

TL;DR AI

Key summary

2 min read
  1. Researchers introduced OmniVAE, a variational autoencoder for audio-video generation that learns shared semantics during training.

  2. It uses segment-level contrastive alignment and modality-specific distillation from pretrained teachers to capture both common events and modality details.

  3. The unified latent space improved text-to-audio-video generation quality and audio-video synchronization with little reconstruction loss.

Read the original