引用本文:吴 飞,刘亚楠,庄越挺.基于张量表示的直推式多模态视频语义概念检测.软件学报,2008,19(11):2853-2868
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 8602次   下载 12948 本文二维码信息
码上扫一扫!
分享到: 微信 更多
基于张量表示的直推式多模态视频语义概念检测
吴 飞1, 刘亚楠1, 庄越挺1
浙江大学 计算机科学与技术学院 数字媒体计算与设计实验室,浙江 杭州 310027
摘要:
提出了一种基于高阶张量表示的视频语义分析与理解框架.在此框架中,视频镜头首先被表示成由视频中所包含的文本、视觉和听觉等多模态数据构成的三阶张量;其次,基于此三阶张量表达及视频的时序关联共生特性设计了一种子空间嵌入降维方法,称为张量镜头;由于直推式学习从已知样本出发能对特定的未知样本进行学习和识别,最后在这个框架中提出了一种基于张量镜头的直推式支持张量机算法,它不仅保持了张量镜头所在的流形空间的本征结构,而且能够将训练集合外数据直接映射到流形子空间,同时充分利用未标记样本改善分类器的学习性能.实验结果表明,该方法能够有效地进行视频镜头的语义概念检测.
关键词:  多模态  张量镜头  时序关联共生  高阶SVD  降维  直推式支持张量机
DOI:
分类号:
基金项目:Supported by the National Natural Science Foundation of China under Grant Nos.60603096, 60533090 (国家自然科学基金); the National High-Tech Research and Development Plan of China under Grant No.2006AA010107 (国家高技术研究发展计划(863); the National Key Technology R&D Program of China under Grant No.2007BAH11B01 (国家科技支撑计划); the Program for Changjiang Scholars and Innovative Research Team in University of China under Grant Nos.IRT0652, PCSIRT (长江学者和创新团队发展计划)
Transductive Multi-Modality Video Semantic Concept Detection with Tensor Representation
WU Fei,LIU Ya-Nan,ZHUANG Yue-Ting
Abstract:
A higher-order tensor framework for video analysis and understanding is proposed in this paper. In this framework, image frame, audio and text are represented, which are the three modalities in video shots as data points by the 3rd-order tensor. Then a subspace embedding and dimension reduction method is proposed, which explicitly considers the manifold structure of the tensor space from temporal-sequenced associated co-occurring multimodal media data in video. It is called TensorShot approach. Transductive learning uses a large amount of unlabeled data together with the labeled data to build better classifiers. A transductive support tensor machines algorithm is proposed to train effective classifier. This algorithm preserves the intrinsic structure of the submanifold where tensorshots are sampled, and is also able to map out-of-sample data points directly. Moreover, the utilization of unlabeled data improves classification ability. Experimental results show that this method improves the performance of video semantic concept detection.
Key words:  multi-modality  TensorShot  temporal associated cooccurrence (TAC)  higher order SVD (HOSVD)  dimensionality reduction  transductive support tensor machine (TSTM)

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: