Andreas Geiger

Publications of Chuqiao Li

FrankenMotion: Part-level Human Motion Generation and Composition
C. Li, X. Xie, Y. Cao, A. Geiger and G. Pons-Moll
Conference on Computer Vision and Pattern Recognition (CVPR), 2026
Abstract: Human motion generation from text prompts has made remarkable progress in recent years. However, existing methods primarily rely on either sequence-level or action-level descriptions due to the absence of fine-grained, part-level motion annotations. This limits their controllability over individual body parts. In this work, we construct a high-quality motion dataset with atomic, temporally-aware part-level text annotations, leveraging the reasoning capabilities of large language models (LLMs). Unlike prior datasets that either provide synchronized part captions with fixed time segments or rely solely on global sequence labels, our dataset captures asynchronous and semantically distinct part movements at fine temporal resolution. Based on this dataset, we introduce a diffusion-based part-aware motion generation framework, namely FrankenMotion, where each body part is guided by its own temporally-structured textual prompt. This is, to our knowledge, the first work to provide atomic, temporally-aware part-level motion annotations and have a model that allows motion generation with both spatial (body part) and temporal (atomic action) control. Experiments demonstrate that FrankenMotion outperforms all previous baseline models adapted and retrained for our setting, and our model can compose motions unseen during training. Our code and dataset will be publicly available upon publication.
Latex Bibtex Citation:
@inproceedings{Li2026CVPR,
  author = {Chuqiao Li and Xianghui Xie and Yong Cao and Andreas Geiger and Gerard Pons-Moll},
  title = {FrankenMotion: Part-level Human Motion Generation and Composition},
  booktitle = {Conference on Computer Vision and Pattern Recognition (CVPR)},
  year = {2026}
}
Unimotion: Unifying 3D Human Motion Synthesis and Understanding (oral)
C. Li, J. Chibane, Y. He, N. Pearl, A. Geiger and G. Pons-Moll
International Conference on 3D Vision (3DV), 2025
Abstract: We introduce Unimotion, the first unified multi-task human motion model capable of both flexible motion control and frame-level motion understanding. While existing works control avatar motion with global text conditioning, or with fine-grained per frame scripts, none can do both at once. In addition, none of the existing works can output frame-level text paired with the generated poses. In contrast, Unimotion allows to control motion with global text, or local frame-level text, or both at once, providing more flexible control for users. Importantly, Unimotion is the first model which by design outputs local text paired with the generated poses, allowing users to know what motion happens and when, which is necessary for a wide range of applications. We show Unimotion opens up new applications: 1.) Hierarchical control, allowing users to specify motion at different levels of detail, 2.) Obtaining motion text descriptions for existing MoCap data or YouTube videos 3.) Allowing for editability, generating motion from text, and editing the motion via text edits. Moreover, Unimotion attains state-of-the-art results for the frame-level text-to-motion task on the established HumanML3D dataset.
Latex Bibtex Citation:
@inproceedings{Li2025THREEDV,
  author = {Chuqiao Li and Julian Chibane and Yannan He and Naama Pearl and Andreas Geiger and Gerard Pons-Moll},
  title = {Unimotion: Unifying 3D Human Motion Synthesis and Understanding},
  booktitle = {International Conference on 3D Vision (3DV)},
  year = {2025}
}


eXTReMe Tracker