A Comprehensive Survey of CNN and Transformer-Based Deep Learning Architectures for Brain Tumor MRI Analysis
Contributors
Dr PARA RAJESH
Keywords
Proceeding
Track
General Track
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
Brain tumor analysis from magnetic resonance imaging (MRI) is an important problem in medical image computing because accurate identification and delineation of tumor tissue support diagnosis, treatment planning, monitoring, and clinical decision-making. Conventional segmentation methods rely on handcrafted intensity, boundary, or region information and are sensitive to noise, intensity inhomogeneity, and anatomical variability [18],[20]. Deep convolutional neural networks (CNNs), particularly Fully Convolutional Networks, U-Net, U-Net++, and nnU-Net, have substantially advanced automated medical image segmentation through hierarchical feature learning and encoder-decoder representations [1]-[3]. However, convolutional operations predominantly model local neighborhoods, which can limit the representation of long-range contextual relationships in heterogeneous tumors [6],[7]. Transformer-based architectures, including Vision Transformer (ViT), Swin Transformer, Swin UNETR, and hybrid models such as CoAtNet, provide attention-based mechanisms for contextual representation learning [4]-[8]. This survey reviews the progression from conventional segmentation to CNN and transformer architectures for brain tumor MRI analysis [1],[2],[6],[7],[20]. It further examines MRI preprocessing, data augmentation, benchmark datasets, evaluation metrics, explainability, robustness, and cross-dataset generalization [12]-[18]. A taxonomy and comparative synthesis are presented to clarify the strengths and limitations of major architectural families. The survey identifies six persistent gaps: preprocessing sensitivity, fragmented multi-task analysis, insufficiently controlled comparisons, limited cross-dataset generalization, inadequate integration of interpretability, and emphasis on benchmark performance over clinical robustness [13]-[16]. Future directions include hybrid CNN-transformer models, self-supervised learning, efficient 3D attention, multimodal fusion, uncertainty-aware prediction, and clinically grounded explainability