A Unified Object-Driven Temporal Framework for Multi-Domain Video Summarization Using Deep Learning
Contributors
Dr. Rachit Adhvaryu
Shashi Kant Gupta
Prof. (Dr.) Shashi Kant Gupta
Keywords
Proceeding
Track
Engineering, Sciences and Mathematics
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
The explosion in multimedia repositories in healthcare, intelligent transportation systems, surveillance and sports analytics has created a growing demand for efficient video summarization techniques. Recent advances in deep learning have produced substantial improvements in the accuracy of summarization, yet most existing methods remain domain-specific, treating object detection, temporal event modeling and summary generation as independent tasks. These limitations impair their generalization ability across heterogeneous video domains, affecting semantic coherence and contextual understanding. In this paper, we propose a new unified hybrid deep learning framework for multi-domain video summarization that integrates object detection, temporal event modeling, and hybrid extractive-abstractive summarization into a single architecture. The proposed framework includes four sequential phases: video preprocessing, deep object detection, temporal event structuring and hybrid summary generation. The proposed methodology has been validated over benchmark datasets from medical, traffic and sports domains. The framework is evaluated using metrics for object detection, temporal localization, extractive summarization and abstractive summarization. The presented framework aims at a unified architecture for processing heterogeneous video datasets, which can enhance semantic understanding, temporal consistency and cross-domain generalization for intelligent multimedia applications.