序列推荐启发式数据增强的修复和丰富
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金(62576083)


Repairing and Enriching Heuristic Data Augmentation for Sequential Recommendation
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    启发式数据增强通过扰动原始的用户行为序列并生成新的序列, 能够有效缓解序列推荐中的数据稀疏性. 相比于基于外部信息、利用复杂的增强规则或引入可学习生成模块的增强方法, 启发式方法通过基于人类先验知识、规则或直觉设计的简单操作来完成增强, 具备较高的效率及泛化性. 然而, 现有的启发式数据增强通常面临两个问题: 一是其产生的增强数据与原始数据相比存在语义偏差或重要用户行为丢失的问题, 导致模型学习到不准确甚至是错误的用户偏好; 二是现有方法专注于对稀疏的输入行为进行增强, 模型训练仍受限于标签行为的稀疏性和不匹配问题, 导致推荐性能提升受限. 为此, 提出一种用于修复和丰富序列推荐中启发式数据增强的框架(REDRec). 具体而言, 首先通过实证研究验证启发式推荐方法导致的序列偏好语义偏离, 并证实简单的詹森-香农散度过滤策略即可缓解这一问题. 在此基础上, 设计语义偏移程度感知的修复机制. REDRec以詹森-香农散度量化增强序列的语义偏离程度, 并为不同偏离程度的样本执行基于差异化贝塔分布的嵌入表示插值操作以实现语义修复. 为了丰富增强数据的标签, REDRec对修复后样本执行基于扰动的多轮推理, 通过自适应聚合这些推理结果生成新的软标签. 新的软标签将作为额外的数据添加到模型训练过程中. 在多个代表性骨干网络和真实世界数据集上的实验验证了REDRec的有效性和泛化性.

    Abstract:

    Heuristic data augmentation effectively mitigates data sparsity in sequential recommendation by perturbing the original user behavior sequences and generating new ones. Compared with augmentation methods that rely on external information, complex augmentation rules, or learnable generation modules, heuristic methods perform augmentation through simple operations based on prior human knowledge, rules, or intuition, offering high efficiency and generalizability. However, existing heuristic data augmentation methods typically face two challenges. First, the augmented data often exhibits semantic bias or loss of important user behaviors compared with the original data, leading models to learn inaccurate or even erroneous user preferences. Second, existing approaches focus on augmenting sparse input behaviors, yet the model training remains constrained by the sparsity and mismatch in labeled behaviors, limiting improvements in recommendation performance. To address this issue, this study proposes a framework for repairing and enriching heuristic data augmentation in sequential recommendation (REDRec). Specifically, this study first validates the semantic deviation in sequence preferences caused by heuristic recommendation methods through empirical studies and demonstrates that a simple Jensen-Shannon divergence filtering strategy can mitigate this issue. On this basis, a repair mechanism that is sensitive to the degree of semantic shift is designed. REDRec quantifies the semantic deviation of augmented sequences using Jensen-Shannon divergence and performs an embedding representation interpolation based on differentiated beta distributions for samples with varying levels of deviation to achieve semantic repair. To enrich the labels of augmented data, REDRec performs perturbation-based multi-round inference on repaired samples and generates new soft labels by adaptively aggregating these inference results. The new soft labels are added as additional data during model training. Experiments on multiple representative backbone networks and real-world datasets validate the effectiveness and generalization of REDRec.

    参考文献
    相似文献
    引证文献
引用本文

党翌洲,赵楚,姜琳颖,马连博,郭贵冰.序列推荐启发式数据增强的修复和丰富.软件学报,,():1-22

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-01-14
  • 最后修改日期:2026-02-20
  • 录用日期:
  • 在线发布日期: 2026-07-01
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号