Multimodal Audio Visual Fake Detection with Explainable Artificial Intelligence


Date Published : 2 August 2026

Contributors

D R Jiji Mol

SRM Arts and Science College, Tamilnadu, India
Author

Deepak Gupta

Maharaja Agrasen Institute of Technology, India.
Author

Keywords

Deepfake detection Multimodal approach Explainable AI Audio-Visual forensics

Proceeding

Track

Engineering, Sciences and Mathematics

License

Copyright (c) 2026 Sustainable Global Societies Initiative

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

Abstract

The development of deepfake generation methods has brought a number of challenges in the field of digital forensics, especially realism and multimodality of manipulated content. Most of the current detection methods are unimodal, which are not interpretable and not effective in real world situations.  To overcome these challenges, a multimodal and explainable deepfake detection framework AV-Fake detection is proposed. The AV-Fake Detection is an advanced deepfake detection framework developed by using multimodal approach and explainable AI to boost the accuracy and explainability of deepfake detection. The proposed framework mitigates the shortcomings of extent unimodal system and enhancing the reliability of forensic tasks, social media management and validation of digital content. This innovation architecture establishes a scalable framework for future intelligent multimedia forensic systems.

References

No References

Downloads

How to Cite

D R Jiji Mol, D. R. J. M., & Deepak Gupta, D. G. (2026). Multimodal Audio Visual Fake Detection with Explainable Artificial Intelligence. Sustainable Global Societies Initiative, 1(10). https://vectmag.com/sgsi/paper/view/1073