引用本文:李芳芳,苏朴真,段俊文,张师超,毛星亮.多粒度信息关系增强的多标签文本分类.软件学报,2023,34(12):5686-5703
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 2101次   下载 4286 本文二维码信息
码上扫一扫!
分享到: 微信 更多
多粒度信息关系增强的多标签文本分类
李芳芳1, 苏朴真1, 段俊文1, 张师超1, 毛星亮2
1.中南大学 计算机学院, 湖南 长沙 410038;2.湖南工商大学 大数据与互联网创新研究院, 湖南 长沙 410205
摘要:
基于深度学习的多标签文本分类方法存在两个主要缺陷: 缺乏对文本信息多粒度的学习, 以及对标签间约束性关系的利用. 针对这些问题, 提出一种多粒度信息关系增强的多标签文本分类方法. 首先, 通过联合嵌入的方式将文本与标签嵌入到同一空间, 并利用BERT预训练模型获得文本和标签的隐向量特征表示. 然后, 构建3个多粒度信息关系增强模块: 文档级信息浅层标签注意力分类模块、词级信息深层标签注意力分类模块和标签约束性关系匹配辅助模块. 其中, 前两个模块针对共享特征表示进行多粒度学习: 文档级文本信息与标签信息浅层交互学习, 以及词级文本信息与标签信息深层交互学习. 辅助模块通过学习标签间关系来提升分类性能. 最后, 所提方法在3个代表性数据集上, 与当前主流的多标签文本分类算法进行了比较. 结果表明, 在主要指标Micro-F1、Macro-F1、nDCG@kP@k上均达到了最佳效果.
关键词:  注意力机制  多标签文本分类  标签关系  多粒度信息
DOI:10.13328/j.cnki.jos.006802
分类号:TP18
基金项目:国家自然科学基金(62172449, 61836016, 71790615, 62006251, 62172441); 湖南省自然科学基金(2021JJ30870, 2021JJ40783); 长沙市自然科学基金(kq2014134); 国防科技重点实验室基金(6142101190302)
Multi-label Text Classification with Enhancing Multi-granularity Information Relations
LI Fang-Fang1, SU Pu-Zhen1, DUAN Jun-Wen1, ZHANG Shi-Chao1, MAO Xing-Liang2
1.School of Computer Science and Engineering, Central South University, Changsha 410038, China;2.Institute of Big Data and Internet Innovation, Hunan University of Technology and Business, Changsha 410205, China
Abstract:
Multi-label text classification methods based on deep learning lack multi-granularity learning of text information and the utilization of constraint relations between labels. To solve these problems, this study proposes a multi-label text classification method with enhancing multi-granularity information relations. First, this method embeds text and labels in the same space by joint embedding and employs the BERT pre-trained model to obtain the implicit vector feature representation of text and labels. Then, three multi-granularity information relations enhancing modules including document-level information shallow label attention (DISLA) classification module, word-level information deep label attention (WIDLA) classification module, and label constraint relation matching auxiliary module are constructed. The first two modules carry out multi-granularity learning from shared feature representation: the shallow interactive learning between document-level text information and label information, and the deep interactive learning between word-level text information and label information. The auxiliary module improves the classification performance by learning the relation between labels. Finally, the comparison with current mainstream multi-label text classification algorithms on three representative datasets shows that the proposed method achieves the best performance on main indicators of Micro-F1, Macro-F1, nDCG@k, and P@k.
Key words:  attention mechanism  multi-label text classification  label relation  multi-granularity information