ThinkOmni: A Reasoning-Driven Omni-Modal LLM Framework for Audio Forgery Detection and Localization

TL;DR AI
2 min readKey summary
Researchers introduced ThinkOmni, a reasoning-driven omni-modal LLM for audio forgery detection and temporal localization.
It is trained with a 100K-sample forensic reasoning dataset and losses that align semantic, acoustic, and spectral cues.
The system is reported to improve both detection and segment-level localization of audio deepfakes.
ThinkOmni also generalizes well across datasets, aiming for more reliable forensic analysis across attack types.
