UniWAM is an 8B-parameter model that brings semantic understanding, visual prediction, and action generation into one MoT architecture. Its three experts exchange information through joint multimodal attention. During pretraining, physical-language supervision and complementary signals from robot, human egocentric, and VQA data build embodied knowledge. During post-training, future visual noise augmentation reduces reliance on precise future predictions, while encoded action history initializes action generation through flow matching.
- Unified world-action architecture: A Mixture-of-Transformers (MoT) connects a physical reasoner, a world generator, and an action predictor through joint multimodal attention.
- Physical-language grounding: Robot actions are represented in natural language, adapting the vision-language component to embodied tasks while retaining its language capabilities.
- Complementary multimodal supervision: The pretraining recipe combines robot demonstrations, human egocentric data, and visual question answering (VQA) data to train the appropriate experts.
- Efficient action generation: Future visual noise augmentation and history-conditioned flow matching support action generation with fewer denoising steps.
- Efficiency: Post-training on RoboTwin 2.0 takes 10 hours on 8 NVIDIA H100 GPUs.
models/,train/,utils/,bak/wan/: physical language model and training runtime.data/robotwin2/: RoboTwin loader and conversion utilities (code only).data/libero/,examples/libero_plus/,scripts/libero/: LIBERO training and LIBERO/LIBERO-plus evaluation code.configs/: RoboTwin physical language, IDM, history-flow, and future-noise examples.inference/robotwin/uniwam/: self-contained RoboTwin policy deployment.
Bridge, DROID, Fractal, and real-world inference are intentionally out of scope for this release.
Python 3.10 and a CUDA-capable PyTorch installation are recommended. Install PyTorch for your CUDA version first, then install the remaining dependencies:
conda create -n ola-sem python=3.10 -y
conda activate ola-sem
pip install torch==2.7.1 torchvision==0.22.1 --index-url https://download.pytorch.org/whl/cu128
pip install flash-attn --no-build-isolation
pip install -r requirements.txtStandard training uses the following pretrained assets:
| Pretrained asset | Link | Fine-tuning data or role |
|---|---|---|
UniWAM-base |
ModelScope | pretained checkpoint on mixed robot,human and VQA datasets. |
UniWAM-robotwin-clean |
ModelScope | post-trained checkpoint on the clean subset of RoboTwin 2.0. |
Wan-AI/Wan2.2-TI2V-5B |
Hugging Face | Video backbone, VAE. |
Qwen/Qwen3-VL-2B-Instruct |
Hugging Face | Vision-language backbone;. |
Keep the downloaded directory structure as follows so that it matches the default paths in the training configs:
pretrained_models/
├── UniWAM-base/
├── Qwen3-VL-2B-Instruct/
└── Wan2.2-TI2V-5B/
└── Wan2.2_VAE.pth
See the RoboTwin post-training guide for data preparation, training, and inference.
See the LIBERO guide for original LIBERO training and separate original LIBERO / LIBERO-plus evaluation.
@article{chen2026uniwam,
title = {UniWAM: Unified World-Action Model},
author={Chen, Jiayi and Song, Wenxuan and Wang, Jingbo and Zhou, Shuai and Gong, Xicheng and Fan, Zehua and Zhou, Ziyang and E, Junwu and Yan, Haodong and Li, Fuhao and Yu, Qize and Huang, Xu and Wang, Pengwei and Chen, Wen and Zhou, Shunbo and Li, Haoang},
journal={arXiv preprint arXiv:2610.02054},
year={2026}
}