FGSVQA: Frequency-Guided Short-form Video Quality Assessment

TL;DR AI
2 min readKey summary
Researchers introduced FGSVQA, a short-form video quality assessment model for UGC content.
The end-to-end framework combines CLIP-based dense features with frequency-derived priors to build artifact- and structure-aware maps.
It then fuses artifact, structure, and original visual branches over time with a gating module for better quality prediction.
On short-form video datasets, FGSVQA achieved strong results, including SRCC 0.736 and PLCC 0.787, while remaining efficient at inference.
