Multimodal Deep Learning for Cervical Cancer Detection, Staging, and Prognosis
Contributors
Dr. Basant Kumar
Keywords
Proceeding
Track
General Track
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
Cervical cancer remains a major killer of women around the globe, especially in low- and middle-income countries where access to proper screening is lacking. Early and precise diagnosis can make a big difference in improving patient outcomes. Our study introduces a multimodal deep learning system that combines three types of data: images from colposcopies and biopsies, lab results like HPV tests and Pap smears, and demographic and clinical info about the patient. We suggest a blend of technologies - a Bidirectional LSTM for medical text, a Vision Transformer for image analysis, and a tabular embedding network for clinical data. This approach helps integrate all info for more accurate diagnoses. The fused model was trained and validated using a retrospective cohort of 3,842 patients from two tertiary cancer centers in South India. It had an AUC of 0.967, sensitivity of 93.4%, and specificity of 91.8%, blowing the single-modality models out of the water by 6–11%. Plus, it did FIGO staging and predicted survival probabilities accurately about 88.2% of the time. Clearly, merging clinical, molecular, and imaging data vastly improves diagnostic precision and could help medics in areas lacking resources.