A systematic review of machine learning and deep learning approaches for gastrointestinal cancer diagnosis
Vijayalakshmi D., Bharanidharan Nagarajan
Globally, one of the prominent causes of cancer-related deaths is gastrointestinal cancer. It includes the tumour in the regions of the gastrointestinal tract, such as the esophagus, stomach, liver, pancreas, and colon. Improving patient outcomes requires an early and accurate diagnosis, but traditional diagnostic techniques are frequently laborious and subjective. Across various modalities, machine learning and deep learning techniques have become effective solutions for computerized diagnostics, categorization, and lesion segmentation. Through an emphasis on the larger category of gastrointestinal cancer rather than specific cancer types, this systematic survey offers a thorough review of machine learning and deep learning implementations in gastrointestinal cancer diagnosis. In this systematic review, 45 documents are selected for the qualitative synthesis while the inclusion criteria are majorly the usage of machine learning and deep learning models for diagnostic tasks such as classification, detection, or segmentation of gastrointestinal cancer. The included studies were systematically analyzed across multiple dimensions, including cancer subtype, imaging modality, model architecture, validation strategy, and reported performance metrics. In addition, publicly accessible datasets, important assessment metrics, and current challenges are highlighted in this article. It also describes emerging research trends that include multimodal data integration combining imaging with clinical or molecular data, the adoption of transformer-based architectures for improved contextual modeling, and increasing interest in federated learning frameworks to address data privacy and cross-institutional generalizability. Overall, transformer-based and hybrid deep learning architectures are emerging as the leading approaches for gastrointestinal cancer diagnosis, demonstrating enhanced contextual representation compared with conventional unimodal Convolutional Neural Network frameworks.