引用本文:赵海全,王续武,李金亮,李直旭,肖仰华.面向视频的细粒度多模态实体链接.软件学报,2024,35(3):1140-1153
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 2185次   下载 4916 本文二维码信息
码上扫一扫!
分享到: 微信 更多
面向视频的细粒度多模态实体链接
赵海全1,2, 王续武1,2, 李金亮3, 李直旭1,2, 肖仰华1,2
1.复旦大学 计算机科学技术学院, 上海 201203;2.上海市数据科学重点实验室(复旦大学), 上海 201203;3.苏州大学 计算机科学与技术学院, 江苏 苏州 215006
摘要:
随着互联网和大数据的飞速发展,数据规模越来越大,种类也越来越多.视频作为其中重要的一种信息方式,随着近期短视频的发展,占比越来越大.如何对这些大规模视频进行理解分析,成为学界关注的热点.实体链接作为一种背景知识补全方式,可以提供丰富的外部知识.视频上的实体链接可以有效地帮助理解视频内容,从而实现对视频内容的分类、检索、推荐等.但是现有的视频链接数据集和方法的粒度过粗,因此提出面向视频的细粒度实体链接,并立足于直播场景,构建了细粒度视频实体链接数据集.此外,依据细粒度视频链接任务的难点,提出利用大模型抽取视频中的实体及其属性,并利用对比学习得到视频和对应实体的更好表示.实验结果表明,该方法能够有效地处理视频上的细粒度实体链接任务.
关键词:  细粒度  视频实体链接  数据集  大语言模型  对比学习
DOI:10.13328/j.cnki.jos.007078
分类号:
基金项目:国家重点研发计划(2020AAA0109302);国家自然科学基金(62072323,62102095);上海市科技创新行动计划(22511105902,22511104700);上海市科技重大专项(2021SHZDZX0103);上海市科学技术委员会资助项目(22511105902)
Fine-grained Multimodal Entity Linking for Videos
ZHAO Hai-Quan1,2, WANG Xu-Wu1,2, LI Jin-Liang3, LI Zhi-Xu1,2, XIAO Yang-Hua1,2
1.School of Computer Science, Fudan University, Shanghai 201203, China;2.Shanghai Key Laboratory of Data Science, (Fudan University), Shanghai 201203, China;3.School of Computer Science and Technology, Soochow University, Suzhou 215006, China
Abstract:
With the rapid development of the Internet and big data, the scale and variety of data are increasing. Video, as an important form of information, is becoming increasingly prevalent, particularly with the recent growth of short videos. Understanding and analyzing large-scale videos has become a hot topic of research. Entity linking, as a way of enriching background knowledge, can provide a wealth of external information. Entity linking in videos can effectively assist in understanding the content of video, enabling classification, retrieval, and recommendation of video content. However, the granularity of existing video linking datasets and methods is too coarse. Therefore, this study proposes a video-based fine-grained entity linking approach, focusing on live streaming scenarios, and constructs a fine-grained video entity linking dataset. Additionally, based on the challenges of fine-grained video linking tasks, this study proposes the use of large models to extract entities and their attributes from videos, as well as utilizing contrastive learning to obtain better representations of videos and their corresponding entities. The results demonstrate that the proposed method can effectively handle fine-grained entity linking tasks in videos.
Key words:  fine-grained  video entity linking  dataset  large language model  contrastive learning

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: