Research Article

From Spatial CNNs to Multimodal Fusion: A Quantitative Survey of Cross-Dataset Generalization in Deepfake Detection

by  Neha Gupta, Rahul Kumar
journal cover
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
Volume 187 - Issue 122
Published: July 2026
Authors: Neha Gupta, Rahul Kumar
10.5120/ijca00d0437ece9c
PDF

Neha Gupta, Rahul Kumar . From Spatial CNNs to Multimodal Fusion: A Quantitative Survey of Cross-Dataset Generalization in Deepfake Detection. International Journal of Computer Applications. 187, 122 (July 2026), 55-62. DOI=10.5120/ijca00d0437ece9c

                        @article{ 10.5120/ijca00d0437ece9c,
                        author  = { Neha Gupta,Rahul Kumar },
                        title   = { From Spatial CNNs to Multimodal Fusion: A Quantitative Survey of Cross-Dataset Generalization in Deepfake Detection },
                        journal = { International Journal of Computer Applications },
                        year    = { 2026 },
                        volume  = { 187 },
                        number  = { 122 },
                        pages   = { 55-62 },
                        doi     = { 10.5120/ijca00d0437ece9c },
                        publisher = { Foundation of Computer Science (FCS), NY, USA }
                        }
                        %0 Journal Article
                        %D 2026
                        %A Neha Gupta
                        %A Rahul Kumar
                        %T From Spatial CNNs to Multimodal Fusion: A Quantitative Survey of Cross-Dataset Generalization in Deepfake Detection%T 
                        %J International Journal of Computer Applications
                        %V 187
                        %N 122
                        %P 55-62
                        %R 10.5120/ijca00d0437ece9c
                        %I Foundation of Computer Science (FCS), NY, USA
Abstract

With the advent of AI-generated synthetic media, including that produced by generative adversarial networks (GANs), variational autoencoders, and diffusion models, deepfake detection has become a formidable challenge in digital forensics and media integrity. The survey examines 30 representative papers published from 2020 to 2025, and includes every type of spatial detector: CNN-based, frequency-domain analysis, temporal and recurrent networks, vision transformers, contrastive and self-supervised learning and multimodal audio-visual fusion. This survey provides mathematical expressions of several important loss functions, comparison tables of performance across three benchmark datasets (FaceForensics++, Celeb-DF, DFDC), and cross-dataset generalization analysis. The analysis reveals that the gap between the transformer-based and spatial-frequency hybrid detectors is ~24% (relative to the CNN baseline) and highlights open challenges in adversarial robustness, diffusion-model deepfakes, and fairness.

References
  • Altuncu, E., Franqueira, V. N. L., & Li, S. (2024). Deepfake: Definitions, performance metrics and standards, datasets, and a meta-review. Frontiers in Big Data, 7, 1400024. https://doi.org/10.3389/fdata.2024.1400024
  • Amerini, I., Galteri, L., Caldelli, R., and Del Bimbo, A. 2019. Deepfake video detection through optical flow based CNN. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW). IEEE, 1205–1207. https://doi.org/10.1109/ICCVW.2019.00152
  • Aqiao, M., Tian, R., and Wang, Y. 2025. Towards Generalizable Deepfake Detection with Spatial-Frequency Collaborative Learning and Hierarchical Cross-Modal Fusion. arXiv preprint arXiv:2504.17223.
  • Chai, L., Bau, D., Lim, S. N., and Isola, P. 2020. What makes fake images detectable? Understanding properties that generalize. In Proceedings of the European Conference on Computer Vision (ECCV). Springer, Cham, 103–120.
  • Chen, L., Zhang, Y., Song, Y., Liu, L., and Wang, J. 2022. Self-supervised learning of adversarial example: Towards good generalizations for deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 18710–18719.
  • Coccomini, D. A., Messina, N., Gennaro, C., and Falchi, F. 2022. Combining EfficientNet and vision transformers for video deepfake detection. In Proceedings of the International Conference on Image Analysis and Processing (ICIAP). Springer, Cham, 219–229.
  • Fang, Z., Zhao, H., Wei, T., Zhou, W., Wan, M., Wang, Z., and Yu, N. 2025. UniForensics: Face forgery detection via general facial representation. IEEE Transactions on Dependable and Secure Computing (2025).
  • Haliassos, A., Vougioukas, K., Petridis, S., and Pantic, M. 2021. Lips don't lie: A generalisable and robust approach to face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 5039–5049.
  • Huang, B., Wang, Z., Yang, J., Ai, J., Zou, Q., Wang, Q., and Ye, D. 2023. Implicit identity driven deepfake face swapping detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4490–4499.
  • Kaur, A., Noori Hoshyar, A., Saikrishna, V., Firmin, S., and Xia, F. 2024. Deepfake video detection: Challenges and opportunities. Artificial Intelligence Review 57, 6 (2024), 159.
  • Khan, I., Khan, K., and Ahmad, A. 2025. A comprehensive survey of DeepFake generation and detection techniques in audio-visual media. ICCK Journal of Image Analysis and Processing 1, 2 (2025), 73–95.
  • Deepa, K. P., Lokesh, C. K., Umamaheswari, D., Ayshwarya, B., Yethish, P. V., and Suhaas, B. 2026. An enhanced deep learning framework for DeepFake detection using EfficientNet-B3: Comparative evaluation of deep and machine learning techniques. Discover Computing 29, 1 (2026), 18.
  • Li, Y., Chang, M. C., and Lyu, S. 2018. In ictu oculi: Exposing AI created fake videos by detecting eye blinking. In Proceedings of the 2018 IEEE International Workshop on Information Forensics and Security (WIFS). IEEE, 1–7.
  • Li, L. and Lyu, S. 2020. Face X-ray for more general face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 5001–5010.
  • Li, Z., Tang, W., Gao, S., Wang, Y., and Wang, S. 2026. Multiple-context and frequency aggregation network for deepfake detection. PLoS ONE 21, 1 (2026), e0337409.
  • Liu, H., Li, X., Zhou, W., Chen, Y., He, Y., Xue, H., and Yu, N. 2021. Spatial-phase shallow learning: Rethinking face forgery detection in frequency domain. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 772–781.
  • Masi, I., Killekar, A., Mascarenhas, R. M., Gurudatt, S. P., and AbdAlmageed, W. 2020. Two-branch recurrent network for isolating deepfakes in videos. In Proceedings of the European Conference on Computer Vision (ECCV). Springer, Cham, 667–684.
  • Nguyen, H. H., Yamagishi, J., and Echizen, I. 2021. Capsule-network-based method for detecting deepfake videos and images. arXiv preprint arXiv:1910.12467.
  • Ni, Y., Meng, D., Yu, C., Quan, C., Ren, D., and Zhao, Y. 2022. CORE: Consistent representation learning for face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 12–21.
  • Raza, A., Basit, A., Amin, A., Arfeen, Z. A., Masud, M. I., Fayyaz, U., and Jumani, T. A. 2026. A comprehensive review of deepfake detection techniques: From traditional machine learning to advanced deep learning architectures. AI 7, 2 (2026), 68.
  • Rössler, A., Cozzolino, D., Verdoliva, L., Riess, C., Thies, J., and Nießner, M. 2019. FaceForensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 1–11.
  • Shiohara, K. and Yamasaki, T. 2022. Detecting deepfakes with self-blended images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 18720–18729.
  • Sunil, R., Mer, P., Diwan, A., Mahadeva, R., and Sharma, A. 2025. Exploring autonomous methods for deepfake detection: A detailed survey on techniques and evaluation.
  • Tan, C., Zhao, Y., Wei, S., Gu, G., Liu, P., & Wei, Y. (2024, March). Frequency-aware deepfake detection: Improving generalizability through frequency space domain learning. In Proceedings of the AAAI Conference on Artificial Intelligence (Vol. 38, No. 5, pp. 5052-5060).
  • Tian, C., Hua, G., and Tang, H. 2023. Deepfake detection with masked attention. In Proceedings of the ACM International Conference on Multimedia (MM '23).
  • Wang, Z., Bao, J., Zhou, W., Wang, W., and Li, H. 2023. AltFreezing for more general video face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 4129–4138.
  • Wang, J., Wu, Z., Ouyang, W., Han, X., Chen, J., Jiang, Y. G., and Li, S. N. 2022. M2TR: Multi-modal multi-scale transformers for deepfake detection. In Proceedings of the 2022 International Conference on Multimedia Retrieval (ICMR '22). 615–623.
  • Wodajo, D., Atnafu, S., and Akhtar, Z. 2023. Deepfake video detection using generative ConvViT. arXiv preprint.
  • Xu, J., et al. 2023. Leveraging real-world data for deepfake detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (ICCVW).
  • Yang, Z., Liang, J., Xu, Y., Zhang, X. Y., and He, R. 2023. Masked relation learning for deepfake detection. IEEE Transactions on Information Forensics and Security 18 (2023), 1696–1708.
  • Yan, Z., et al. 2023. LSDA: Large scale deepfake detection algorithm. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
  • Zsai, C. C., Wu, T. H., and Lai, S. H. 2022. Multi-scale patch-based representation learning for image anomaly detection and segmentation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). 3992–4000.
  • Zhan, N., Nguyen, T., Bermak, A., and Khalil, I. 2025. CAMME: Adaptive deepfake image detection with multi-modal cross-attention. arXiv preprint arXiv:2505.18035.
  • Zheng, Y., Bao, J., Chen, D., Zeng, M., and Wen, F. 2021. Exploring temporal coherence for more general video face forgery detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 15044–15054.
Index Terms
Computer Science
Information Sciences
No index terms available.
Keywords

Deepfake Detection GAN Vision Transformer FaceForensics++ Frequency Domain Contrastive Learning Multimodal Fusion Digital Forensics

Powered by PhDFocusTM