arXiv 2026

Streaming Multi-Track Timeline Control
for 3D Human Motion Generation

Yangsong Zhang1 Anujith Muraleedharan1 Rikhat Akizhanov1 Gül Varol2 Fabio Pizzati1 Ivan Laptev1
1 Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) 2 LIGM, École des Ponts, IP Paris, Univ Gustave Eiffel, CNRS
Paper · Coming soon Code · Coming soon Data & Models · Coming soon
TimelineControl teaser showing streaming text control, spatial control, and continuous real-robot execution

Streaming multi-track motion generation with TimelineControl. (a) Sequential and overlapping body-part instructions. (b) Joint text, root-path, and heading control. (c) Humanoid robot execution. In (a) and (b), body-part colors correspond to instructions.

Abstract

Text-driven human motion generation has advanced substantially, yet most methods assume instructions are available before synthesis. Interactive applications require responding to new instructions while continuing ongoing actions, such as answering a phone while walking. Existing approaches address streaming generation or simultaneous composition without explicitly combining streaming instruction arrival with independently timed, overlapping actions. We introduce streaming multi-track timeline control and propose TimelineControl to incorporate new instructions alongside ongoing actions. Interval-aware conditioning preserves instruction timing, while causal part-structured representations and part-aware denoising coordinate concurrent actions across body regions. We also construct TimelineMotion, a dataset with overlapping instruction intervals and body-part annotations. Experiments on TimelineMotion and MTT demonstrate improved semantic alignment and temporal adherence over evaluated streaming baselines, including models retrained on the same data. Ablations and human evaluations validate our design, complemented by spatial conditioning and humanoid execution demonstrations.

Video

Method Overview

Architecture overview of TimelineControl

Overview of the TimelineControl architecture. We process part features and whole-body features separately (left). The part features undergo block-causal self-attention (yellow). Both use online multi-prompt conditioning (pink, OMPC) to allow for multi-track streaming inference. We add a part bias (orange) to enforce better mapping to body parts. Details (right) show examples of elements used in text

01

Block-causal attention

Motion features interact within the same generation block while future blocks remain masked, enabling causal streaming generation.

02

Part-aware guidance

A body-part relevance predictor identifies which of the six anatomical regions each instruction should guide.

03

Multi-track attention

Temporal prompt scores are combined with interval visibility masks, so each motion position attends only to applicable instructions.

Qualitative Comparisons

We compare TimelineControl with streaming and offline composition baselines on two multi-track motion benchmarks. Each video visualizes the instruction timeline together with the generated motion.

TimelineMotion

STMC [1] ActionPlan* [2] Ours

MTT

STMC [1] ActionPlan* [2] Ours

Applications

Online interaction

Interactive Demo

New text instructions can arrive during generation and be incorporated alongside compatible ongoing actions.

Long-horizon generation

Continuous Long Sequence

A causal sliding window supports continuous motion across many sequential and overlapping instructions.

Multimodal guidance

Text + Spatial Control

TimelineControl combines streamed language with root paths, waypoints, heading, body keyframes, end-effector targets, and foot contacts.

Physical deployment

Real-Robot Execution

Generated motions are executed on a G1 humanoid in real time through a motion-tracking controller.

BibTeX

If you find this work useful, please cite:

@article{zhang2026timelinecontrol,
  title     = {Streaming Multi-Track Timeline Control for 3D Human Motion Generation},
  author    = {Zhang, Yangsong and Muraleedharan, Anujith and Akizhanov, Rikhat and
               Varol, G{\"u}l and Pizzati, Fabio and Laptev, Ivan},
  journal   = {arXiv preprint},
  year      = {2026}
}

Acknowledgements

We thank helps from public code like ActionPlan, STMC, MotionStreamer, ARDY, KIMODO, SONIC, etc.