Comparative Study of Fine-Tuned Vision Transformer and Convolutional Neural Network for Mammography Classification
Contributors
Dr.Leena Nesamani S
Dr.M.Malathi
Keywords
Proceeding
Track
General Track
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
Convolutional Neural Networks (CNNs) are efficient in mammography classification, but they inefficient in capturing global context data. CNNs have demonstrated excellent performance in mammography classification; nevertheless, they have limited capacity to collect global contextual information. A different deep learning technique for medical image processing has recently surfaced: Vision Transformers (ViTs). This research compares the ResNet50 model and the Swin Transformer models in the classifying mammogram images utilizing the CBIS-DDSM dataset and INbreast datasets. ImageNet pretrained weights were utilized in transfer learning to refine both models. The INbreast dataset was used for external validation in order to assess model generalization. A limited subset of target-domain samples was then used for lightweight domain adaption. According to experimental results, both models performed well in internal validation; however, due to domain shift between datasets, there was a noticeable decline in performance during external evaluation. After domain adaptation, ResNet50 showed greater resilience and significantly reduced the False Negative Rate, whereas Swin Transformer only slightly improved. The results emphasize the significance of lightweight domain adaption and cross-dataset evaluation in the development of clinically reliable AI systems for mammography-based computer-aided diagnosis and breast cancer screening.