AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars

TL;DR AI
2 min readKey summary
AptAvatar is a 14B-parameter framework for generating long-form avatar videos from audio.
It combines endpoint-anchored distribution distillation with self-generated history replay to preserve identity and quality over long sequences.
With a two-step inference process and just 2 NFEs, it can produce 720p output and aims to make production-grade avatars much faster.
