Abstract:Generating body motions that are synchronized with speech and temporally coherent is essential for enhancing the interactivity of digital humans. This task requires a model to capture both long-term motion evolution and fine-grained local motion details. To address the long-range error accumulation of purely autoregressive models and the high sampling cost of purely diffusion-based models for long-sequence generation, this paper proposes a hybrid generation framework that combines patch-level autoregression with local bidirectional diffusion. Specifically, the latent motion sequence is first divided into fixed-length patches, from which patch-level representations are extracted by an aggregation encoder. A patch-level causal Transformer is then employed to model long-range dependencies across patches and produce patch summary vectors. Finally, a local bidirectional diffusion decoder, conditioned on historical context, is used to recover motion details within each patch. Furthermore, classifier-free guidance based on autoregressive conditional summaries and a temperature-controlled ordinary differential equation (ODE) sampling mechanism are introduced to strengthen speech-conditioned generation and improve sampling controllability. Experimental results on the BEAT2 dataset show that the proposed method achieves an FGD of 4.925, a BC of 7.932, and a Diversity score of 13.24. Compared with EMAGE, the proposed method reduces FGD from 5.512 to 4.925 and raises BC from 7.724 to 7.932, demonstrating improved motion quality and speech synchronization.