|
|
SUAVE: Unified Video-Action Models via Masked Diffusion
Rhythm Syed,
Jean Mercat,
Sedrick Keh,
Kushal Arora,
Paarth Shah,
Aykut Onol,
Mengchao Zhang,
Tony Dear
arXiv 2026
We present SUAVE, a Single vocabulary Unified Action-Video modEl in which a masked diffusion transformer generates video and actions conditioned on language, with all three modalities represented as discrete tokens in a shared sequence. Choosing which tokens to mask at inference turns the same network into a world model, a robot policy, or a video-action model. A single SUAVE model predicts long-horizon video and acts as a policy, competitive with dedicated world models and specialized action policies on static and dynamic manipulation tasks, while running closed-loop on a real robot at 2.5 actions per second on an RTX 5090. Pretraining on robot video and co-training on human video substantially improves policy performance and zero-shot robustness to distribution shift.
|
|
|
Interactive World Simulator for Robot Policy Training and Evaluation
Yixuan Wang,
Rhythm Syed,
Fangyu Wu,
Mengchao Zhang,
Aykut Onol,
Jose Barreiros,
Hooshang Nayyeri,
Tony Dear,
Huan Zhang,
Yunzhu Li
RSS 2026
Best Student Paper Award at the RSS 2026 Workshop on Robot World Models
project website /
arXiv /
code /
models & data
We introduce a world model that runs at 15 FPS for over 10 minutes on a single RTX 4090 GPU and predicts photorealistic and physically accurate future frames. Our model is interactive in real time and performs action-conditioned video prediction which enables high quality data generation for policy training and evaluation making it scalable, reproducible, and faithful.
|
|
|
Ontology Learning and Knowledge Graph Construction from Unstructured Text
Rhythm Syed,
Zach Welz,
Santiago Balestrini
GTRI Research Conference 2021
code
Knowledge graphs are valuable data models for natural language and AI systems, but they are time-intensive and expensive to build by hand. We propose a framework for automatically constructing and augmenting knowledge graphs from unstructured text corpora with minimal human supervision, and evaluate it on an existing ontology domain including an entity linking task for graph expansion.
|
|
|
Generalized Latency Performance Estimation for Once-For-All Neural Architecture Search
Rhythm Syed,
Arvind Akpuram Srinivasan
arXiv 2021
Once-For-All (OFA) neural architecture search relies on per-hardware latency lookup tables that are slow and manual to build. We instead train latency predictors for neural network architectures and introduce two generalization strategies: finetuning a predictor to new hardware, and GPU generalization that conditions on hardware parameters such as core count, RAM size, and memory bandwidth. Our predictors achieve over 50% lower RMSE than ProxylessNAS and match or exceed the NAS performance of the lookup table baseline.
|
Last updated October 2026. Template from here.
|
|