An Explainable Vision Transformer Framework for Robust Disease Detection in Medical Imaging
Main Article Content
Abstract
Deep learning has achieved substantial progress in medical image analysis, but limited interpretability and sensitivity to domain shifts continue to hinder reliable clinical deployment. This paper proposes an explainable Vision Transformer framework for robust disease detection from medical images. The proposed model combines multi-scale visual tokenization with hierarchical self-attention to capture both localized pathological patterns and global anatomical context. A consistency-guided attention mechanism encourages the model to focus on clinically relevant regions across image perturbations, while an uncertainty-aware classification head estimates confidence for individual predictions. To improve generalization across imaging environments, the framework further incorporates feature-level domain regularization during training. Experiments across multiple medical imaging datasets demonstrate improved classification performance and robustness compared with conventional convolutional and Transformer-based baselines. Visualization and localization analyses further show that the learned attention patterns correspond more closely to disease-relevant regions. The framework provides an interpretable and generalizable approach for AI-assisted medical image analysis.
Article Details

This work is licensed under a Creative Commons Attribution 4.0 International License.
Mind forge Academia also operates under the Creative Commons Licence CC-BY 4.0. This allows for copy and redistribute the material in any medium or format for any purpose, even commercially. The premise is that you must provide appropriate citation information.
References
M. Xiao, Y. Li, X. Yan, M. Gao, and W. Wang, “Convolutional neural network classification of cancer cytopathology images: Taking breast cancer as an example,” pp. 145–149, 2024.
A. Dosovitskiy et al., “An image is worth 16×16 words: Transformers for image recognition at scale,” ICLR, 2021.
X. Yan, J. Du, X. Li, X. Wang, X. Sun, P. Li, and H. Zheng, “A Hierarchical Feature Fusion and Dynamic Collaboration Framework for Robust Small Target Detection,” IEEE Access, vol. 13, pp. 92953–92964, 2025, doi: 10.1109/ACCESS.2025.3570669.
R. R. Selvaraju et al., “Grad-CAM: Visual explanations from deep networks via gradient-based localization,” ICCV, pp. 618–626, 2017.
W. Wang, Y. Li, X. Yan, M. Xiao, and M. Gao, “Breast cancer image classification method based on deep transfer learning,” pp. 190–197, 2024.
X. Yan, J. Du, L. Wang, Y. Liang, J. Hu, and B. Wang, “The Synergistic Role of Deep Learning and Neural Architecture Search in Advancing Artificial Intelligence,” ICEDCS, pp. 452–456, 2024.
J. N. Acosta, G. J. Falcone, P. Rajpurkar, and E. J. Topol, “Multimodal biomedical AI,” Nature Medicine, vol. 28, pp. 1773–1784, 2022.
H. Zheng, Y. Ma, Y. Wang, G. Liu, Z. Qi, and X. Yan, “Structuring low-rank adaptation with semantic guidance for model fine-tuning,” ICECAI, pp. 731–735, 2025.
X. Yan, W. Wang, M. Xiao, Y. Li, and M. Gao, “Survival prediction across diverse cancer types using neural networks,” pp. 134–138, 2024.
M. J. Sheller et al., “Federated learning in medicine: Facilitating multi-institutional collaborations without sharing patient data,” Scientific Reports, vol. 10, Art. no. 12598, 2020.
Y. Li, W. Zhao, B. Dang, X. Yan, M. Gao, W. Wang, and M. Xiao, “Research on adverse drug reaction prediction model combining knowledge graph embedding and deep learning,” MLISE, pp. 322–329, 2024.
J. Wei, Y. Liu, X. Huang, X. Zhang, W. Liu, and X. Yan, “Self-Supervised Graph Neural Networks for Enhanced Feature Extraction in Heterogeneous Information Networks,” ICMLCA, pp. 272–276, 2024.
N. Rieke et al., “The future of digital health with federated learning,” npj Digital Medicine, vol. 3, Art. no. 119, 2020.
H. Zheng, L. Zhu, W. Cui, R. Pan, X. Yan, and Y. Xing, “Selective knowledge injection via adapter modules in large-scale language models,” ICAIDE, pp. 373–377, 2025.
Y. Li, X. Yan, M. Xiao, W. Wang, and F. Zhang, “Investigation of Creating Accessibility Linked Data Based on Publicly Available Accessibility Datasets,” pp. 77–81, 2024.
B. Shickel et al., “Deep EHR: A survey of recent advances in deep learning techniques for electronic health record analysis,” IEEE Journal of Biomedical and Health Informatics, vol. 22, no. 5, pp. 1589–1604, 2018.
X. Yan, Y. Jiang, W. Liu, D. Yi, and J. Wei, “Transforming Multidimensional Time Series into Interpretable Event Sequences for Advanced Data Mining,” ICHCI, pp. 126–130, 2024.
A. Rajkomar et al., “Scalable and accurate deep learning with electronic health records,” npj Digital Medicine, vol. 1, Art. no. 18, 2018.