EyefusionAI: A Lightweight Multimodal and Secure Framework for Eye Disease Diagnostics
Loading...
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Nazarbayev University School of Engineering and Digital Sciences
Abstract
Eye diseases including Diabetic Retinopathy, Age-related Macular Degeneration, and Glaucoma are major contributors to preventable blindness worldwide, yet detection at early stages is hindered by the scarcity of clinical expertise and the complexity of interpreting retinal images. The current research introduces a deep learning framework for the concurrent segmentation and classification of eye diseases from both fundus images and optical coherence tomography angiography (OCTA) images. The framework is based on extending the Lightweight Bottleneck Narrowing with Attention U-Net (LWBNA-UNet) to develop a multi-task and multi-modal framework consisting of a shared encoder, dual decoders, and independent segmentation and classification heads for both modalities. The shared encoder facilitates knowledge transfer from the larger fundus image dataset (FIVES, containing 600 training images) to the smaller OCTA dataset (OCT500, containing 189 training images), and the bottleneck narrowing mechanism with squeeze and excitation attention helps to reduce false activations and increase segmentation accuracy from low-quality images. 3 different training methodologies have been utilized: two-phase training, joint training, and federated learning. Two-phase training and Joint training methods differ in the way how the total loss is computed: in the former method classification and segmentation layers are trained separately, while in the latter method we train both layers simultaneously. In the Federated learning approach, we have a completely different training methodology simulating the way the model would be training on 2 hospitals for 20 rounds using the Flower framework. As a result, we have achieved the dice scores of 0.8121 and 0.7996, 0.7663 and 0.6256, 0.7921 and 0.7595 during two-phase training, joint training, and Federated learning for OCTA and Fundus segmentations respectively. In terms of classification performance, we have achieved the Macro average F1 scores of 0.947 and 0.488, 0.137 and 0.247, 0.3333 and 0.1020 for OCTA Binary and Fundus 4-class classification tasks. Although all training approaches yielded comparable Dice scores, their classification heads did not learn properly reaching the maximum potential during two-phase training and completely collapsing in Joint training and Federated learning. The experiments demonstrate the efficacy of the two-phase training paradigm for multi-task retinal image analysis under small dataset conditions, providing a solid foundation for end-to-end fine-tuning and clinical evaluation.
Description
Citation
Zharas, G. (2026). EyefusionAI: A Lightweight Multimodal and Secure Framework for Eye Disease DiagnosticsEyefusionAI: A Lightweight Multimodal and Secure Framework for Eye Disease Diagnostics. Nazarbayev University School of Engineering and Digital Sciences
Collections
Endorsement
Review
Supplemented By
Referenced By
Creative Commons license
Except where otherwised noted, this item's license is described as Attribution-ShareAlike 3.0 United States
