引用本文:李玘芮,励雪巍,赵奇,李杰,李玺.姿态控制人物图像视频生成技术综述.软件学报,2026,37(5):1982-2005
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览 2127次   下载 1162 本文二维码信息
码上扫一扫!
分享到: 微信 更多
姿态控制人物图像视频生成技术综述
李玘芮1, 励雪巍2, 赵奇1, 李杰1, 李玺1
1.浙江大学 计算机科学与技术学院, 浙江 杭州 310013;2.上海电机学院 电子信息学院, 上海 201306
摘要:
生成技术的飞速发展揭示了相关技术在实际应用中的潜力, 姿态控制人物图像视频生成技术(pose-guided person image and video generation)的核心目标是将输入信息的人物转换为指定姿态, 同时保持人物外观的高度一致性. 该技术可以广泛应用于虚拟试穿与时尚行业、广告内容生成领域的视频生成与编辑以及多模态结合生成等多个应用场景, 推动用户体验和技术创新的进步. 尽管该技术已经取得了显著进展, 仍面临着多个挑战, 包括姿态迁移过程中外观信息的有效提取和重排、不可见信息的生成、一致性保持、模型的高效训练与使用等. 基于现有技术的挑战, 详细分析了当前主流的姿态控制生成方法应对挑战的策略, 并探讨了它们在实际应用中的可行性和局限性. 同时, 还讨论了姿态控制生成技术的常用生成模型以及不同的姿态信息表示方法. 此外, 整理讨论了该技术常用的数据集大小、特点等信息、各项测试基准, 并从虚拟试穿、视频生成与编辑、多模态结合生成等应用场景展开了讨论. 此外, 还揭示了目前方法仍遇到的个性化信息的保留、复杂场景的生成以及模型效率与实时性能等挑战, 并讨论姿态控制生成技术可能的未来发展趋势, 旨在为相关领域的研究人员提供系统的总结与参考, 以期推动该技术在各行业中的应用与创新.
关键词:  姿态控制生成  人物图像视频生成  生成对抗网络  扩散模型  可控生成
DOI:10.13328/j.cnki.jos.007539
分类号:
基金项目:
Review on Pose-guided Person Image and Video Generation Technologies
LI Qi-Rui1, LI Xue-Wei2, ZHAO Qi1, LI Jie1, LI Xi1
1.College of Computer Science and Technology, Zhejiang University, Hangzhou 310013, China;2.School of Electronic Information Engineering, Shanghai Dianji University, Shanghai 201306, China
Abstract:
The rapid development of generative technologies has revealed their potential for real-world applications. The core objective of pose-guided person image and video generation is to transform a person from inputs into a specified pose while maintaining a high level of appearance consistency. This technology can be widely applied in various fields such as virtual try-on and fashion, advertising video generation and editing, and multimodal content creation, driving advancements in user experience and technological innovation. However, despite significant progress, the technology still faces multiple challenges, including effective extraction and rearrangement of appearance information during pose transfer, generation of unseen information, consistency preservation, and efficient model training and deployment. Based on the existing challenges, this study provides a detailed analysis of the strategies employed by current mainstream pose-guided generation methods to address these issues, discussing their feasibility and limitations in practical applications. Moreover, it explores the commonly used generative models and pose representation methods in pose-guided generation. It also reviews the datasets, their sizes, characteristics, and evaluation benchmarks used in this field. Furthermore, this study discusses the applications of this technology in virtual try-on, video generation and editing, and multimodal content generation. It highlights the remaining challenges, such as the retention of personalized information, generation in complex scenes, and model efficiency and real-time performance. Finally, this study discusses potential future development trends of pose-guided generation technology, aiming to provide researchers with a systematic summary and reference to promote its application and innovation across industries.
Key words:  pose-guided generation  person image and video generation  generative adversarial network (GAN)  diffusion model  controllable generation

引用本文:
【打印本页】   【下载PDF全文】   查看/发表评论  【EndNote】   【RefMan】   【BibTex】
←前一篇|后一篇→ 过刊浏览    高级检索
本文已被:浏览次   下载  
分享到: 微信 更多
摘要:
关键词:  
DOI:
分类号:
基金项目:
Abstract:
Key words: