A Hybrid Vision Transformer and ESA-Based Framework for Binary and Multi-Class Breast Cancer Histopathological Image Classification
Contributors
Amrutanshu Panigrahi
Subrata Chowdhury
Keywords
Proceeding
Track
General Track
License
Copyright (c) 2026 Sustainable Global Societies Initiative

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
Abstract
Breast cancer is a major cause of death among women throughout the world. Therefore, a correct and early diagnosis is essential for the effective treatment of the disease. Analysing histopathological images is essential for disease detection and classification of breast cancer. However, analysis of these images currently is a manual task that is lengthy, subjective, and reliant on expert opinion. The work presents a visionary transformer ViT and auto encoder network AEN based framework for breast histopathology image classification using deep learning. The Vision Transformer is applied as a deep feature extractor to capture global contextual information and long-range dependencies from histopathological images, while the Autoencoder Network is employed for feature selection to remove redundant and irrelevant features. Therefore, to perform classification tasks, optimized extracted features are used which may be binary or multi-class. The proposed framework improves classification efficiency, reduces computation complexity and the diagnostic accuracy by reducing the dimensionality of features and enhancing feature representation. Through experimental evaluation, the ViT and AEN integration demonstrates a well-robust and reliable performance predicting breast cancer histopathology image classification, making this a promising computer-aided diagnosis system in medical imaging applications.