|
International Journal of Computer Applications
Foundation of Computer Science (FCS), NY, USA
|
| Volume 187 - Issue 135 |
| Published: August 2026 |
| Authors: Tu Thanh Tri, Duong Thi Thuy Nga, Chau Phuong Toan, Dang Kim Lien |
10.5120/ijcac919f6ae2058
|
Tu Thanh Tri, Duong Thi Thuy Nga, Chau Phuong Toan, Dang Kim Lien . An Explainable Vision Transformer Framework for Skin Lesion Classification using Dual-Map Fusion Strategy. International Journal of Computer Applications. 187, 135 (August 2026), 35-42. DOI=10.5120/ijcac919f6ae2058
@article{ 10.5120/ijcac919f6ae2058,
author = { Tu Thanh Tri,Duong Thi Thuy Nga,Chau Phuong Toan,Dang Kim Lien },
title = { An Explainable Vision Transformer Framework for Skin Lesion Classification using Dual-Map Fusion Strategy },
journal = { International Journal of Computer Applications },
year = { 2026 },
volume = { 187 },
number = { 135 },
pages = { 35-42 },
doi = { 10.5120/ijcac919f6ae2058 },
publisher = { Foundation of Computer Science (FCS), NY, USA }
}
%0 Journal Article
%D 2026
%A Tu Thanh Tri
%A Duong Thi Thuy Nga
%A Chau Phuong Toan
%A Dang Kim Lien
%T An Explainable Vision Transformer Framework for Skin Lesion Classification using Dual-Map Fusion Strategy%T
%J International Journal of Computer Applications
%V 187
%N 135
%P 35-42
%R 10.5120/ijcac919f6ae2058
%I Foundation of Computer Science (FCS), NY, USA
Accurate and interpretable skin lesion classification is essential for early melanoma diagnosis and clinical decision support. This study proposes an explainable computer-aided diagnosis framework based on a Vision Transformer (ViT) for binary classification of benign and malignant skin lesions. To improve model transparency, Grad-CAM and transformer attention maps are integrated through a Dual-Map Fusion strategy, providing complementary local and global visual explanations of the model's predictions. The proposed framework was evaluated using a combined HAM10000 and ISIC dermoscopic image dataset. Experimental results demonstrated an overall classification accuracy of 92.55% and a malignant lesion recall of 94.68%, indicating reliable diagnostic performance with a reduced risk of missed malignant cases. In addition, the fused explanation maps provided more informative and interpretable visual evidence than Grad-CAM or attention maps alone, facilitating a better understanding of the model's decision-making process. The proposed framework combines the strong classification capability of Vision Transformer with enhanced explainability, providing an effective and transparent decision-support tool for automated skin lesion diagnosis. These findings demonstrate the potential of explainable transformer-based models for reliable clinical application in dermatological image analysis.