Comparative Analysis of Data Deduplication Techniques for Cloud-Based Library Biography Systems: A Survey of Related Work
Contributors
Mohan
Upendra Kumar
Keywords
Proceeding
Track
General Track
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
This paper presents a comprehensive comparative analysis of existing data deduplication techniques and their applicability to cloud-based library biography systems. This survey analysis 15 seminal works spanning content-defined chunking algorithms, cloud storage optimization, entity deduplication, and backup systems. Through systematic comparison across multiple dimensions—including technical approach, application domain, experimental methodology, and quantitative performance—we identify critical gaps in current research. Specifically, no existing work addresses the unique characteristics of library biography collections: mixed multimedia content (text, portraits, signatures), frequent versioned edits, cross-collection redundancy, and biography-specific semantic boundaries. Our analysis establishes the foundation for a content-defined chunking framework tailored to this domain, achieving 68.5% storage savings and 72% bandwidth reduction. This comparative study provides researchers and library system architects with a structured reference for selecting appropriate deduplication techniques based on their specific requirements.