Abstract:Embodied intelligence requires agents to generate natural, coherent, and physically plausible motions from language instructions. Text-driven human motion generation has therefore become an important task for bridging language understanding and embodied behavior expression. Existing diffusion-based methods have achieved notable progress in motion naturalness and text-motion semantic consistency. However, they usually apply a unified training and inference strategy to samples with different complexity levels, limiting their ability to handle variations in sample complexity. To address this, this paper proposes CARM, a complexity-assessed routing method for human motion generation. CARM first measures the motion direction matching degree between pelvis features and text features, and the motion posture matching degree between limb features and text features. The two matching scores are then used to compute a sample complexity score and generate routing labels. During training, CARM assigns differentiated timestep sampling ranges according to the sample complexity score, enabling timestep sampling routing. During inference, the complexity score is used to select a shallow, intermediate, or deep output path. Samples with lower complexity scores can exit early, while samples whose complexity remains high at the current stage are forwarded to deeper paths for generation. In addition, auxiliary losses are introduced at the shallow and intermediate exits to improve the stability of multi-exit inference. Experiments on the HumanML3D and KIT-ML datasets show that the proposed method improves motion generation quality, demonstrating its effectiveness.