🔔 Scheduled Maintenance:The platform will undergo backend updates on July 1st, 2026, 2:00–4:00 AM UTC.
AJETAS
Back to Journal Table of Contents
Paper Code: AJETAS-2026-6622Open AccessResearch ArticlePeer Reviewed

Explainable Deep Learning for Gastrointestinal Endoscopy: A Comprehensive Review of XAI Methods for Computer-Aided Diagnosis

Vaibhav Santosh More;VISHAL DEEPAK KHAIRNAR
Received: May 15, 2026Accepted: June 20, 2026Published: July 2026
Published inAvrmitra Journal of Engineering, Technology and Applied Sciences
Volume & IssueVol. 4, Issue 2
DOI10.2694/ajetas.2026.8379

Abstract & Keywords

Gastrointestinal (GI) diseases — including polyps, colorectal neoplasia, ulcerative colitis, esophagitis, and Barrett's esophagus — represent a substantial global health burden, and endoscopy remains the clinical gold standard for their detection and characterization. Deep learning has driven rapid progress in endoscopy-based Computer-Aided Diagnosis (CAD), with convolutional neural networks (CNNs), Vision Transformers (ViTs), and hybrid CNN-Transformer architectures now approaching or exceeding expert-level classification accuracy on public benchmarks. However, the black-box nature of these models is a persistent barrier to clinical trust, regulatory approval, and safe deployment, motivating the integration of Explainable AI (XAI) methods that expose the visual evidence underlying a model's decision. This review consolidates and benchmarks the principal XAI techniques applied to endoscopy-based GI disease classification, spanning gradient-based class activation methods (Grad-CAM, Grad-CAM++, Score-CAM, Layer-CAM, Eigen-CAM), perturbation- and game-theoretic attribution methods (LIME, SHAP), gradient-attribution methods (Integrated Gradients, Occlusion Sensitivity), and architecture-native attention maps from Vision Transformers. We synthesize reported classification performance (accuracy, precision, recall, F1-score, AUC), computational cost (inference time, FLOPs, memory footprint), and explainability quality (faithfulness, localization accuracy against expert-annotated ground truth, robustness to perturbation, and human interpretability) across the major public endoscopy datasets — Kvasir, HyperKvasir, Kvasir-Capsule, CVC-ClinicDB, ETIS-Larib, the SUN colonoscopy dataset, and GastroVision. Findings indicate a consistent trade-off: gradient- and activation-based CAM variants (Grad-CAM++, Score-CAM, Layer-CAM) offer the best balance of localization fidelity and computational cost for real-time endoscopic use, while perturbation-based methods (LIME, SHAP) achieve competitive or superior fidelity in isolated benchmarks but incur substantially higher inference latency, limiting their use to offline audit and model validation rather than point-of-care deployment. We present comparative tables spanning model architectures, XAI methods, datasets, and computational cost, propose a standardized benchmark methodology, and outline open challenges — including the absence of pixel-level explainability ground truth for most disease classes, weak correlation between saliency-map plausibility and true model faithfulness, and the lack of clinician-in-the-loop validation studies — together with directions for future research toward clinically trustworthy, regulator-ready explainable CAD systems for gastroenterology.

Index Keywords
Index Terms — Explainable AIXAIGrad-CAMSHAPLIMEVision TransformerComputer-Aided DiagnosisGastrointestinal EndoscopyKvasirDeep LearningMedical Image Classification.

Author Affiliations

Vaibhav Santosh MoreCorresponding Author · Department of Computer Science & Engineering, IIT Bombay, India
VISHAL DEEPAK KHAIRNARResearch Partner · Department of Computer Science & Engineering, IIT Bombay, India

References Listing (3)

  1. [1]K. R. Rao and J. Doe, "High-resolution convolutional modeling in thermal imaging systems," IEEE Trans. Image Process., vol. 31, pp. 210–222, Jan. 2024.
  2. [2]W. Liu, R. Chen, and M. Rossi, "Deep residual neural networks for noise-reduction and super-resolution edge rebuilding," Pattern Recognition, vol. 142, pp. 110-125, Oct. 2025.
  3. [3]S. Jenkins, "Cooperative token-heuristics inside warehousing routing grids," Int. J. Rob. Res., vol. 18, no. 4, pp. 450–467, May 2025.