Multimodal Deep Learning for Cervical Cancer Detection, Staging, and Prognosis


Date Published : 2 August 2026

Contributors

Dr. Basant Kumar

Modern College of Business and Science, Muscat, Oman
Author

Keywords

Cervical cancer; multimodal deep learning; Vision Transformer; colposcopy; HPV; early detection; FIGO staging; fusion model.

Proceeding

Track

General Track

License

Copyright (c) 2026 Sustainable Global Societies Initiative

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

Abstract

Cervical cancer remains a major killer of women around the globe, especially in low- and middle-income countries where access to proper screening is lacking. Early and precise diagnosis can make a big difference in improving patient outcomes. Our study introduces a multimodal deep learning system that combines three types of data: images from colposcopies and biopsies, lab results like HPV tests and Pap smears, and demographic and clinical info about the patient. We suggest a blend of technologies - a Bidirectional LSTM for medical text, a Vision Transformer for image analysis, and a tabular embedding network for clinical data. This approach helps integrate all info for more accurate diagnoses. The fused model was trained and validated using a retrospective cohort of 3,842 patients from two tertiary cancer centers in South India. It had an AUC of 0.967, sensitivity of 93.4%, and specificity of 91.8%, blowing the single-modality models out of the water by 6–11%. Plus, it did FIGO staging and predicted survival probabilities accurately about 88.2% of the time. Clearly, merging clinical, molecular, and imaging data vastly improves diagnostic precision and could help medics in areas lacking resources.

References

No References

Downloads

How to Cite

Dr. Basant Kumar, D. B. K. (2026). Multimodal Deep Learning for Cervical Cancer Detection, Staging, and Prognosis. Sustainable Global Societies Initiative, 1(6). https://vectmag.com/sgsi/paper/view/725