Dual-view Fusion for Fine-grained Image Recognition with Vision Transformer
Author:
Affiliation:

Clc Number:

TP391

Fund Project:

  • Article
  • |
  • Figures
  • |
  • Metrics
  • |
  • Reference
  • |
  • Related
  • |
  • Cited by
  • |
  • Materials
  • |
  • Comments
    Abstract:

    With the continuous advancement of computer vision technology, fine-grained image recognition plays a crucial role across various application domains. Unlike traditional coarse-grained image recognition, fine-grained image recognition aims to precisely distinguish subcategories with subtle visual differences within the same major category, making this task particularly challenging. In recent years, the vision Transformer has gained widespread adoption in image recognition due to its exceptional performance in modeling global contextual information. However, the vision Transformer exhibits certain limitations when applied to fine-grained image recognition, particularly in processing detailed features and mitigating background noise. To address these issues, this study proposes a dual-view recognition framework based on the vision Transformer. This framework effectively integrates global and local views to enhance recognition accuracy. In this framework, an attention-based fusion module is designed to filter redundant information and optimize the classification token embedding of global views by merging and filtering patch features through hierarchical attention weights within the encoder. In addition, an attention threshold-based key region localization module is introduced. This module dynamically selects and magnifies key patches in the global view using an adaptive threshold strategy, forming detailed local views for further analysis. Furthermore, an adaptive enhancement module for local region features is proposed to strengthen the focus on local details, thus enhancing the recognition capability of fine-grained features. To optimize the dual-view framework, a contrastive loss function based on dual-view similarity and an adaptive inference strategy based on dual-view confidence are proposed. These strategies aim to enhance the global and local feature discriminability of the vision Transformer model while efficiently saving computational resources and shortening inference time. Experimental results on the CUB-200-2011, Stanford Dogs, NABirds, and iNaturalist2017 public datasets demonstrate that the proposed method achieves significant improvements in recognition accuracy compared to the traditional vision Transformer model. These results validate the proposed method’s effectiveness and superiority in fine-grained image recognition tasks.

    Reference
    Related
    Cited by
Get Citation

唐昊,李泽超,蒋鑫,唐金辉.基于视觉Transformer的双视图融合细粒度图像识别.软件学报,2026,37(5):2286-2308

Copy
Share
Article Metrics
  • Abstract:
  • PDF:
  • HTML:
  • Cited by:
History
  • Received:October 15,2024
  • Revised:March 11,2025
  • Adopted:
  • Online: October 29,2025
  • Published: May 06,2026
You are the firstVisitors
Copyright: Institute of Software, Chinese Academy of Sciences Beijing ICP No. 05046678-4
Address:4# South Fourth Street, Zhong Guan Cun, Beijing 100190,Postal Code:100190
Phone:010-62562563 Fax:010-62562533 Email:jos@iscas.ac.cn
Technical Support:Beijing Qinyun Technology Development Co., Ltd.

Beijing Public Network Security No. 11040202500063