Switch language한국어
Back to the list

ViDS: Video Diffusion Shader using 3D Face Tracking

TL;DR AI

Key summary

2 min read
  1. ViDS is a new portrait animation system that reconstructs a subject-specific 3D face model from a reference image and drives it with expression and pose from video.

  2. It feeds dense geometric cues into a diffusion model to produce more consistent, lifelike talking-head animations.

  3. The paper also adds autoregressive sampling to extend generation beyond the usual window while reducing clip-to-clip artifacts.

  4. Overall, the approach improves realistic portrait animation by using 3D face tracking for better motion control and stronger identity preservation.

Read the original