An Uncertainty-Aware Hybrid Multimodal Deep Learning Framework for Skin Cancer Detection
Contributors
Dr. Sonia Arora
Prof. Dr. Shashi Kant Gupta
Keywords
Proceeding
Track
General Track
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
While automatic classification of skin lesions for computer aided assessment in dermatology has been proven to be successful, it is not sufficient to assume that it is good image classification performance, to achieve good clinical decision support. A hybrid multimodal deep learning approach that fuses dermoscopic appearance, clinical metadata and predictive uncertainty is proposed in this paper. The data set contains 10015 dermoscopic images of 7 diagnostic categories, age, gender and location of the lesion (HAM10000). Images are normalised, enhanced and metadata is cleaned, encoded and normalised. The feature extraction module used for the visual feature is based on efficientNetV2B0, the clinical information is expressed by a clinical encoder, and the attention/gated mechanism is used to integrate the two modalities. To estimate the predictive uncertainty, Monte Carlo Dropout is kept on for inference. The splitting of patients/lesions, controlled ablations, macro-averaged measures, calibration and uncertainty analysis are the foundations of the validation design. Two goals of the study are to see if clinical information helps to improve class-sensitive performance and to see if uncertainty estimates can detect potentially unreliable predictions.