Multimodal Audio Visual Fake Detection with Explainable Artificial Intelligence
Contributors
D R Jiji Mol
Deepak Gupta
Keywords
Proceeding
Track
Engineering, Sciences and Mathematics
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
The development of deepfake generation methods has brought a number of challenges in the field of digital forensics, especially realism and multimodality of manipulated content. Most of the current detection methods are unimodal, which are not interpretable and not effective in real world situations. To overcome these challenges, a multimodal and explainable deepfake detection framework AV-Fake detection is proposed. The AV-Fake Detection is an advanced deepfake detection framework developed by using multimodal approach and explainable AI to boost the accuracy and explainability of deepfake detection. The proposed framework mitigates the shortcomings of extent unimodal system and enhancing the reliability of forensic tasks, social media management and validation of digital content. This innovation architecture establishes a scalable framework for future intelligent multimedia forensic systems.