基于分块自回归与局部扩散的语音驱动动作生成
CSTR:
作者:
作者单位:

作者简介:

通讯作者:

中图分类号:

基金项目:

国家自然科学基金面上项目(62571369, 62572346)


Patch-level Autoregressive and Local Bidirectional Diffusion for Speech-driven Body Motion Generation
Author:
Affiliation:

Fund Project:

  • 摘要
  • |
  • 图/表
  • |
  • 访问统计
  • |
  • 参考文献
  • |
  • 相似文献
  • |
  • 引证文献
  • |
  • 资源附件
  • |
  • 文章评论
    摘要:

    生成与语音同步且自然连贯的身体动作是提升数字人交互表现力的重要基础.该任务要求模型既能建模长时序动作演化,又能恢复局部动作细节.针对纯自回归模型易出现长程误差累积、纯扩散模型在长序列生成中采样开销较大的问题,本文提出一种分块自回归与局部双向扩散相结合的混合生成框架.该方法先将动作潜在序列划分为固定长度分块,并由聚合编码器提取块级表示;再利用块级因果Transformer建模跨块长时依赖,输出分块摘要向量;最后通过结合历史上下文的局部双向扩散解码器在块内恢复动作细节.进一步地,本文引入基于自回归条件摘要的无分类器引导和基于温度控制的常微分方程采样机制,以增强语音条件约束并提高采样可控性.在BEAT2数据集上的实验结果表明,本文方法取得了4.925的FGD、7.932的BC和13.24的Diversity.相较于EMAGE,本文方法将FGD由5.512降至4.925,BC由7.724提升至7.932,表现出更优的动作质量与语音同步性.

    Abstract:

    Generating body motions that are synchronized with speech and temporally coherent is essential for enhancing the interactivity of digital humans. This task requires a model to capture both long-term motion evolution and fine-grained local motion details. To address the long-range error accumulation of purely autoregressive models and the high sampling cost of purely diffusion-based models for long-sequence generation, this paper proposes a hybrid generation framework that combines patch-level autoregression with local bidirectional diffusion. Specifically, the latent motion sequence is first divided into fixed-length patches, from which patch-level representations are extracted by an aggregation encoder. A patch-level causal Transformer is then employed to model long-range dependencies across patches and produce patch summary vectors. Finally, a local bidirectional diffusion decoder, conditioned on historical context, is used to recover motion details within each patch. Furthermore, classifier-free guidance based on autoregressive conditional summaries and a temperature-controlled ordinary differential equation (ODE) sampling mechanism are introduced to strengthen speech-conditioned generation and improve sampling controllability. Experimental results on the BEAT2 dataset show that the proposed method achieves an FGD of 4.925, a BC of 7.932, and a Diversity score of 13.24. Compared with EMAGE, the proposed method reduces FGD from 5.512 to 4.925 and raises BC from 7.724 to 7.932, demonstrating improved motion quality and speech synchronization.

    参考文献
    相似文献
    引证文献
引用本文

宋丹,赵山山,高志廷,王岚君,刘安安.基于分块自回归与局部扩散的语音驱动动作生成.软件学报,2027,38(5):

复制
相关视频

分享
文章指标
  • 点击次数:
  • 下载次数:
  • HTML阅读次数:
  • 引用次数:
历史
  • 收稿日期:2026-04-27
  • 最后修改日期:2026-06-18
  • 录用日期:
  • 在线发布日期: 2026-09-14
  • 出版日期:
文章二维码
您是第位访问者
版权所有:中国科学院软件研究所 京ICP备05046678号-3
地址:北京市海淀区中关村南四街4号,邮政编码:100190
电话:010-62562563 传真:010-62562533 Email:jos@iscas.ac.cn
技术支持:北京勤云科技发展有限公司

京公网安备 11040202500063号