Less is More: Modality-Decoupling for General AIGC Audio-Video Detection

TL;DR AI
2 min readKey summary
Researchers proposed DAV-Det, a general audio-video AIGC detector that decouples modalities and analyzes audio and video separately.
Instead of depending on cross-modal consistency, the model learns independent visual and audio forensic cues, then fuses decisions at the end.
It uses multi-level visual representations and a dual-branch audio design to better detect forged or generated content.
The system ranked first in a benchmark challenge, highlighting the need for more robust detection beyond traditional deepfake methods.
