A MULTI-TASK VISION TRANSFORMER FOR JOINT PLANT DISEASE CLASSIFICATION AND LOCALIZATION
DOI:
https://doi.org/10.32955/neuaiit2026511328Özet
Plant diseases pose a serious threat to agricultural productivity and food security, particularly in regions with limited access to expert diagnosis. While deep learning–based methods have shown strong performance in plant disease recognition, most existing approaches focus solely on classification and provide limited spatial interpretability. In this paper, we present MTL-ViT, a multi-task Vision Transformer framework for joint plant disease classification and localization using leaf images. The proposed model employs a shared Vision Transformer backbone with task-specific heads for disease classification and bounding-box localization. Experiments are conducted on a curated subset of the PlantVillage dataset using an 80:20 train–validation split. The results show that the proposed framework achieves 88.89% classification accuracy while simultaneously localizing disease regions with a mean IoU of 0.99 under controlled imaging conditions. Ablation studies highlight the trade-off between predictive accuracy and spatial interpretability, demonstrating the practical value of multi-task learning for explainable plant disease analysis.
Keywords: Vision Transformer, Multi-Task Learning, Plant Disease Classification, Disease Localization, Agricultural Computer Vision.

