Abstract
Text-to-motion generation and motion-to-text captioning are two fundamental tasks in human motion modeling, both grounded in the same underlying motion-text correspondence. Existing unified approaches mostly rely on autoregressive modeling, which imposes a fixed generation order and is therefore poorly suited to the bidirectional dependencies between language and motion, allowing early prediction errors to persist as fixed context and degrade both temporal coherence and cross-modal consistency. Masked discrete diffusion, which models sequences through iterative bidirectional prediction, offers a natural remedy. We therefore propose BiMoGen (Bidirectional Motion-text Generation), a unified masked discrete diffusion framework for bidirectional motion-text modeling. To stabilize training, we design Decoupled Uni- and Cross-Modal Training, in which masked pretraining first establishes cross-modal correspondence on paired motion-text sequences, after which supervised fine-tuning specializes the model for bidirectional generation. Masked diffusion nonetheless introduces its own source of error, as the model is trained on clean ground-truth context yet encounters self-generated and potentially erroneous context at inference, with errors committed under heavily masked states propagating through subsequent steps. We further introduce Generation-Aware Self-Correction that exposes the model to its own predictions during training and applies correction passes at early sampling steps to revise unreliably committed tokens. Extensive experiments on HumanML3D and KIT-ML demonstrate state-of-the-art performance on both tasks, validating the effectiveness of the proposed two-stage training and self-correction designs.
Method Overview
BiMoGen treats motion and text as discrete token sequences and performs iterative masked prediction over noisy states \(x_t\) for both directions. The same framework supports text-to-motion generation, motion-to-text captioning, and motion in-between generation.
Quantitative Results
The tables below summarize bidirectional motion-text generation performance on HumanML3D and KIT-ML, followed by training-strategy ablations and learning dynamics.
Benchmark tables
Qualitative Figures
Static visualizations highlight improvements from self-correction and show motion in-between behavior under different observed-frame conditions.
Text-to-Motion Videos
Paired videos compare generation without self-correction and with self-correction. Red terms in the captions indicate the motion concept emphasized by the prompt.
Motion-to-Text Caption
Each motion is paired with a generated text description. The videos below temporarily reuse the text-to-motion examples as placeholders and can be replaced with final motion-to-text samples.
Motion In-Between Videos
Results are grouped by observed-frame condition. Prefix 25% gives the first quarter of the motion and generates the remaining frames; Middle 50% gives the central half and generates both ends; Suffix 25% gives the final quarter and generates the preceding motion; Ends 25% gives short observed segments at both ends and generates the middle transition.
Given Prefix 25%
Samples 0-4
Given Middle 50%
Samples 0-4
Given Suffix 25%
Samples 0-4
Given Ends 25%
Samples 0-4
BibTeX
@inproceedings{weng2026bimogen,
title = {{BiMoGen}: Bidirectional Motion-Text Generation via Unified Masked Discrete Diffusion},
author = {Weng, Wanjiang and Wu, Yongliang and Tan, Xiaofeng and Zhu, Xingyu and Zhu, Wenbo and Wang, Hongsong},
booktitle = {Advances in Neural Information Processing Systems},
year = {2026}
}
Download BibTeX