Abstract:Imitation learning for visuomotor policies has shown strong promise in robotic manipulation, yet it typically relies on a large amount of costly human demonstrations. While synthetic data generation in simulation can partially alleviate this dependence, it requires substantial engineering efforts and remains limited by the sim-to-real transfer gap. A lightweight, plug-and-play data-augmentation framework for robotics is still lacking. To address this gap, we propose a data augmentation framework for real-robot demonstrations, which takes 3D point clouds as its representation and rearranges scene elements via lightweight editing. Prior work has shown that motion-level augmentation (e.g., trajectory variation) is effective for improving spatial generalization; in this paper, we instead focus on skill-level augmentation to enhance stability and generalization during manipulation and physical interaction. Given only a single human demonstration per task, we introduce two mechanisms: Failure Recovery, which turns perturbed trajectories into corrective training signals, and Contact-Preserving Perturbations, which deform only the non-contact regions of a manipulated object while keeping the contact geometry intact. Extensive real-robot experiments across diverse manipulation tasks show that the proposed method significantly improves task success rates and generalizes better to out-of-distribution objects and perturbed executions, confirming the effectiveness of skill-level augmentation.