A Hybrid Feature Fusion Approach for Cross-quality Deepfake Detection – F2CV

Abstract
The rapid advancement of the deep generative models has enabled the creation of highly realistic deepfake videos capable of manipulating facial content with high visual fidelity. Detecting such manipulations becomes more challenging when videos are stored or transmitted in compressed formats, as compression alters pixel-level artifacts and may obscure forensic traces. This paper proposes a compression-aware multimodal deepfake detection framework that jointly exploits spatial, frequency and motion information from compressed videos. The framework first extracts frames and localizes facial regions using the Multitask Cascaded Convolutional Network (MTCNN) detector. Spatial features are obtained from RGB frames using ConvNeXt, frequency-domain features are extracted from Discrete Cosine Transform (DCT) representations using a Swin Transformer and temporal motion patterns are captured from motion vectors using a 3D CNN combined with a Temporal Convolutional Network (TCN). In addition, codec-level artifacts such as motion vectors and DCT coefficients are utilized to capture compression-induced inconsistencies associated with manipulated content. The extracted features are integrated using an attention-based fusion module that adaptively weights modality-specific representations before classification. By jointly modeling visual and codecdomain information, the proposed framework achieves an accuracy (ACC) level of 99.63% and Area Under Curve (AUC) of 99.83% on the FaceForensics++ (FF++) dataset, improving robustness to compression artifacts and enhances deepfake detection performance across different compression levels.
Keywords: Attention-based Fusion, Compressed Video Forensics, DCT Features, Deepfake Detection, Motion Vectors (Mvs), RGB Features.

Author(s): Diya Garg*, Rupali Gill
Volume: 7 Issue: 3 Pages: 1462-1480
DOI: https://doi.org/10.47857/irjms.2026.v07i03.012628