Dual Adaptive Redundancy Elimination for Training-free Video Question Answering
Author:
Affiliation:

Clc Number:

TP37

Fund Project:

  • Article
  • |
  • Figures
  • |
  • Metrics
  • |
  • Reference
  • |
  • Related
  • |
  • Cited by
  • |
  • Materials
  • |
  • Comments
    Abstract:

    In recent years, training-free video question answering (VQA) models have become a research hotspot for lightweight multimodal reasoning due to their plug-and-play nature. However, although high frame rate videos contain rich semantic information, their inherent redundancy leads to a balance problem between information density and computational efficiency in the temporal dimension, with traditional sampling strategies being susceptible to noise frame interference. Furthermore, in complex dynamic scenes, background clutter and local body parts, as non-target regions, introduce spatial feature bias, significantly affecting the reliability of answer generation. To address these two issues, this study proposes a dual adaptive redundancy elimination (DARE) framework, which aims to systematically improve the accuracy of video semantic understanding and answer quality in the training-free paradigm through a spatiotemporal redundancy collaborative optimization mechanism. First, a dual-relation temporal sampling method is proposed, based on text-visual alignment and inter-frame semantic consistency. This method selects key frame sequences through bidirectional interactive reasoning, while simultaneously eliminating redundant frames that conflict with the text context. Next, a dynamic spatial sampling method is introduced, which extracts the largest connected semantic region from candidate regions in the prompt-related heatmap, aiming to eliminate scattered non-target regions and enhance the compactness of spatial feature representations. Experiments are conducted on widely used datasets, including MSVD-QA, MSRVTT-QA, TGIF-QA, and ActivityNet-QA. The proposed method is evaluated in a zero-shot setting against 14 state-of-the-art models. The results show that the proposed approach achieves competitive performance with significantly fewer video feature sequences. Visual analysis confirms that the proposed method exhibits more accurate spatiotemporal localization abilities in challenging tasks, such as multi-person interactions and fine-grained action recognition in complex scenes. The proposed DARE-VQA framework achieves significant improvements in video question answering performance by collaboratively optimizing spatiotemporal redundancy. It can generate accurate and high-quality answers within the training-free paradigm, demonstrating its potential in multimodal video understanding.

    Reference
    Related
    Cited by
Get Citation

方承炀,朱畅,姜文晖,方玉明,鄢杰斌.面向免训练视频问答的双重自适应冗余消除.软件学报,2026,37(5):1950-1963

Copy
Share
Article Metrics
  • Abstract:
  • PDF:
  • HTML:
  • Cited by:
History
  • Received:May 26,2025
  • Revised:July 11,2025
  • Adopted:
  • Online: September 23,2025
  • Published: May 06,2026
You are the firstVisitors
Copyright: Institute of Software, Chinese Academy of Sciences Beijing ICP No. 05046678-4
Address:4# South Fourth Street, Zhong Guan Cun, Beijing 100190,Postal Code:100190
Phone:010-62562563 Fax:010-62562533 Email:jos@iscas.ac.cn
Technical Support:Beijing Qinyun Technology Development Co., Ltd.

Beijing Public Network Security No. 11040202500063