Articles
| Open Access | Hybrid Deep Vision Framework for Diabetic Retinopathy Classification Using ResNet-ViT Feature Fusion
Abstract
Diabetic retinopathy (DR) is a major vision-threatening complication of diabetes that requires accurate and early diagnosis to prevent irreversible retinal damage. Conventional image-based diagnostic systems often struggle with complex retinal variations, including microvascular abnormalities, lesion diversity, and differences in disease severity. Recent advancements in deep learning have demonstrated significant potential for automated DR analysis; however, existing approaches frequently face limitations in feature representation, generalization capability, and the integration of local and global retinal characteristics. This research proposes a Hybrid Deep Vision Framework for Diabetic Retinopathy Classification Using ResNet-ViT Feature Fusion, combining convolutional feature extraction capabilities of Residual Networks (ResNet) with the global contextual modeling strength of Vision Transformers (ViT). The proposed framework utilizes complementary feature learning, where ResNet captures fine-grained retinal structures while ViT enhances long-range dependency modeling across image regions. The study is positioned within modern intelligent medical imaging systems that require scalable, accurate, and computationally efficient classification mechanisms. Existing research on hybrid neural networks, segmentation-assisted analysis, transformer-based classification, and optimized transfer learning demonstrates the importance of integrating multiple learning paradigms for improved DR recognition. The proposed framework addresses these gaps by introducing a feature fusion strategy capable of improving representation diversity and classification robustness. The findings indicate that hybrid deep vision architectures provide a promising direction for automated retinal screening systems, although challenges related to computational complexity, dataset variability, and clinical deployment remain important considerations.
Keywords
Diabetic Retinopathy, Deep Learning, ResNet, Vision Transformer
References
1. A. Rahman et al., “Diabetic retinopathy detection: A hybrid intelli-
2. gent approach,” Comput. Mater. Continua., vol. 80, no. 3, pp. 4561–
3. 4576, 2024.
4. D. Yi et al., “Compound scaling encoder-decoder (CoSED) network
5. for diabetic retinopathy related bio- marker detection,” IEEE J.
6. Biomed. Health. Inf., vol. 28, no. 4, pp. 1959–1970, 2023.
7. G. Sivapriya et al., “Automated diagnostic classification of diabetic
8. retinopathy with microvascular structure of fundus images using deep
9. learning method,” Biomed. Signal Process. Control., vol. 88,
10. p. 105616, 2024.
11. H. Xu et al., “A hybrid neural network ap- proach for classifying
12. diabetic retinopathy subtypes,” Front Med (Lausanne)., vol. 10,
13. p. 1293019, 2024.
14. I. K. Gupta et al., “Retinal fundus imaging based diabetic retinopathy
15. classification using transfer learning and fennec fox optimization,”
16. MethodsX., vol. 14, p. 103232, 2025.
17. Q. Zhang et al., “Recent advances and applications of optical
18. coherence tomography angiography in diabetic retinopathy,” Front
19. Endocrinol (Lausanne)., vol. 16, p. 1438739, 2025.
20. Y. Yang et al., “A novel transformer model with multiple instance
21. learning for diabetic retinopathy classification,” IEEE Access.,
22. Z. Ullah et al., “SSMD-UNet: Semi-supervised multi-task decoders
23. network for diabetic retinopathy segmentation,” Sci Rep., vol. 13, no.
24. K. S. Hebbar, "Evolving High-Volume Systems: Reactive Execution Models for Resilient Operations," Computer Fraud and Security, vol. 2024, no.04, pp. 49-58, Apr. 2024.
25. S. S. Sravanthi Valiveti, "Cloud Service Models and Execution Architectures: A Unified Survey of IaaS to Serverless Computing," 2025 5th International Conference on Emerging Research in Electronics, Computer Science and Technology (ICERECT), MANDYA, India, 2025, pp. 1-7, doi: 10.1109/ICERECT65215.2025.11375903.
Article Statistics
Downloads
Copyright License
Copyright (c) 2026 Farhan Alamsyah (Author)

This work is licensed under a Creative Commons Attribution 4.0 International License.