|
|
SUAVE: Unified Video-Action Models via Masked Diffusion
Rhythm Syed,
Jean Mercat,
Sedrick Keh,
Kushal Arora,
Paarth Shah,
Aykut Onol,
Mengchao Zhang,
Tony Dear
arXiv 2026
We present SUAVE, a Single vocabulary Unified Action-Video modEl in which a masked diffusion transformer generates video and actions conditioned on language, with all three modalities represented as discrete tokens in a shared sequence. Choosing which tokens to mask at inference turns the same network into a world model, a robot policy, or a video-action model. A single SUAVE model predicts long-horizon video and acts as a policy, competitive with dedicated world models and specialized action policies on static and dynamic manipulation tasks, while running closed-loop on a real robot at 2.5 actions per second on an RTX 5090. Pretraining on robot video and co-training on human video substantially improves policy performance and zero-shot robustness to distribution shift.
|
|
|
Interactive World Simulator for Robot Policy Training and Evaluation
Yixuan Wang,
Rhythm Syed,
Fangyu Wu,
Mengchao Zhang,
Aykut Onol,
Jose Barreiros,
Hooshang Nayyeri,
Tony Dear,
Huan Zhang,
Yunzhu Li
RSS 2026
Best Student Paper Award at the RSS 2026 Workshop on Robot World Models
project website /
arXiv /
code /
models & data
We introduce a world model that runs at 15 FPS for over 10 minutes on a single RTX 4090 GPU and predicts photorealistic and physically accurate future frames. Our model is interactive in real time and performs action-conditioned video prediction which enables high quality data generation for policy training and evaluation making it scalable, reproducible, and faithful.
|
Last updated October 2026. Template from here.
|
|