A Comprehensive Review and Deep Learning Framework for Audio-Visual Zero-Shot Learning Leveraging Multiscale Temporal Dynamics and Semantic Feature Integration
Contributors
sowmya lakshmi bs
Sudhakar K
Keywords
Proceeding
Track
General Track
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
Audio-Visual Zero-Shot Learning (AV-ZSL) has emerged as a promising paradigm for recognizing unseen event categories without requiring labelled training samples for every class. Conventional supervised learning approaches depend heavily on large-scale annotated datasets and struggle to generalize to novel categories encountered in dynamic real-world environments. Recent advances in multimodal deep learning have demonstrated the effectiveness of integrating audio, visual, and semantic information to improve unseen-class recognition. However, existing approaches often focus on isolated aspects such as semantic alignment, temporal modeling, or cross-modal fusion, resulting in limited generalization capabilities. This paper presents a comprehensive review of recent developments in AV-ZSL and proposes a unified deep learning framework integrating multiscale temporal dynamics, semantic feature integration, and cross-modal attention mechanisms. The proposed framework combines hierarchical temporal representations from audio and visual streams with semantic embeddings derived from textual descriptions and class attributes. Through an extensive review of existing literature, key research gaps are identified and a generalized architecture is proposed to address challenges related to semantic inconsistency, modality imbalance, and domain shift. The framework provides a scalable foundation for intelligent surveillance, multimedia retrieval, assistive technologies, healthcare monitoring, and next-generation multimodal artificial intelligence systems.