Renping Zhou*, Zanlin Ni*, Zihao Fan, Guohao Fu, Zeyu Liu, Hao Shi, Jie Zhang, Chi Bene Chen, Yang Yue, Xueyang Fu, Gao Huang†
(* indicates equal contribution, † means Co-corresponding author)
This repository contains the official implementation of the paper "What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling".
conda create -n simple-wam python=3.10 -y
conda activate simple-wamCore model and data dependencies are pinned in requirements.txt; benchmark-specific environments are not included.
pip install torch==2.7.1+cu128 torchvision==0.22.1+cu128 \
--index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt
pip install -e .These trajectory datasets are needed for training, not for simulation evaluation with the released checkpoints.
Download the dataset from
yuanty/LIBERO-fastwam
into data/libero_mujoco3.3.2/, then extract each archive:
mkdir -p data/libero_mujoco3.3.2
# Place all four .tar.gz files in this directory first.
for archive in data/libero_mujoco3.3.2/*.tar.gz; do
tar -xzf "$archive" -C data/libero_mujoco3.3.2
doneThe existing data configs expect this layout:
data/libero_mujoco3.3.2/
libero_10_no_noops_lerobot/
libero_goal_no_noops_lerobot/
libero_object_no_noops_lerobot/
libero_spatial_no_noops_lerobot/
Download all robotwin2.0.tar.gz.part-* files from
yuanty/robotwin2.0-fastwam
into data/robotwin2.0/. Concatenate the parts in filename order and extract:
mkdir -p data/robotwin2.0
# Place every archive part in this directory first.
cat data/robotwin2.0/robotwin2.0.tar.gz.part-* | tar -xzf - -C data/robotwin2.0Expected layout:
data/robotwin2.0/
robotwin2.0/
data/
meta/
videos/
Download the released checkpoints and simulator assets from rpzhou/Simple-WAM:
huggingface-cli download rpzhou/Simple-WAM \
--repo-type model \
--include "simplewam_checkpoints/*" "benchmark_assets/*" \
--local-dir checkpoints
export PYTHONPATH="$(pwd)/src:${PYTHONPATH:-}"
export DIFFSYNTH_MODEL_BASE_PATH="$(pwd)/checkpoints"If the files are already downloaded, skip the download command. Expected layout:
checkpoints/
simplewam_checkpoints/
libero_joint_2cam224_1e-4.pt
libero_joint_2cam224_1e-4_fewshot10.pt
...
robotwin_joint_3cam_384_1e-4_fewshot10.pt
robotwin_joint_3cam_384_1e-4_taskgen.pt
robotwin_joint_3cam_384_1e-4_taskgen_video.pt
libero_dataset_stats.json
robotwin_dataset_stats.json
benchmark_assets/
libero_standard.tar.gz
libero_plus.tar.gz
robotwin_assets.tar.gz
Wan VAE, text encoder, and tokenizer files are downloaded automatically on first
use if missing. For offline evaluation, prepare those files in the same
checkpoints/ model-cache directory beforehand. The trained VideoDiT and
ActionDiT weights come from the released .pt file.
Install the simulator dependencies into the existing environment.
pip install -c requirements.txt \
mujoco==3.3.2 robosuite==1.4.0 bddl==1.0.1 gym==0.25.2 \
easydict==1.9 cloudpickle==2.1.0 future==0.18.2 \
opencv-python==4.6.0.66 matplotlib==3.5.3 \
h5py==3.16.0 scipy==1.13.1 PyOpenGL==3.1.10
tar -xzf checkpoints/benchmark_assets/libero_standard.tar.gz -C third_party/LIBERO
pip install -e third_party/LIBERO
export PYTHONPATH="$(pwd)/third_party/LIBERO:${PYTHONPATH:-}"
export LIBERO_CONFIG_PATH="$(pwd)/third_party/LIBERO/.libero"
export MUJOCO_GL=egl
export PYOPENGL_PLATFORM=eglThe archive restores the omitted third_party/LIBERO/libero/ directory, including
assets, BDDL files, and initial states. Its paths are resolved through the bundled
third_party/LIBERO/.libero/config.yaml.
Evaluate the full-data checkpoint on all four suites, with 50 trials per task:
config_name=libero_joint_2cam224_1e-4
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
python experiments/libero/run_libero_manager.py \
"task=${config_name}" \
"ckpt=checkpoints/simplewam_checkpoints/${config_name}.pt" \
EVALUATION.dataset_stats_path=checkpoints/simplewam_checkpoints/libero_dataset_stats.json \
'MULTIRUN.task_suite_names=[libero_spatial,libero_object,libero_goal,libero_10]' \
MULTIRUN.num_gpus=8 \
MULTIRUN.max_tasks_per_gpu=1 \
MULTIRUN.persistent_workers=true \
EVALUATION.num_trials=50 \
EVALUATION.num_inference_steps=10 \
EVALUATION.freeze_future_video_noise=true \
EVALUATION.torch_compile_action=true \
EVALUATION.offload_text_encoder=trueUse the same command with a different config_name to evaluate another released
checkpoint. For task generalization, replace MULTIRUN.task_suite_names with the
held-out suite:
| Checkpoint/config suffix | Evaluation suites |
|---|---|
_fewshot5, _fewshot10 |
All four suites |
_taskgen_spatial, _taskgen_video_spatial |
[libero_spatial] |
_taskgen_object, _taskgen_video_object |
[libero_object] |
_taskgen_goal, _taskgen_video_goal |
[libero_goal] |
_taskgen_long, _taskgen_video_long |
[libero_10] |
All suffixes above follow libero_joint_2cam224_1e-4. Results are saved under
evaluate_results/libero/<config_name>/.
Install the simulator dependencies listed in LIBERO above, then add the image-corruption dependencies and extract the LIBERO-plus archive:
conda install -c conda-forge imagemagick -y
pip install -c requirements.txt Wand==0.7.2 scikit-image==0.25.2
tar -xzf checkpoints/benchmark_assets/libero_plus.tar.gz -C third_party/LIBERO-plus
pip install -e third_party/LIBERO-plus
export PYTHONPATH="$(pwd)/third_party/LIBERO-plus:$(pwd)/src:${PYTHONPATH:-}"
export LIBERO_CONFIG_PATH="$(pwd)/third_party/LIBERO-plus/.libero"
export NUMBA_CACHE_DIR="$(pwd)/.numba_cache"
export MPLCONFIGDIR="$(pwd)/.matplotlib_cache"
export MUJOCO_GL=egl
export PYOPENGL_PLATFORM=egl
python -c "import libero.libero; print(libero.libero.__file__)"The printed path must be under third_party/LIBERO-plus. Both benchmarks use
the same Python package name, libero; installing LIBERO-plus replaces the
editable LIBERO installation. The explicit PYTHONPATH and LIBERO_CONFIG_PATH
above select the plus benchmark in this shell. To switch back to standard
LIBERO, rerun its editable install and environment exports from the previous
section.
Use the same released LIBERO checkpoint and normalization statistics. This command evaluates all four robustness suites with one rollout per task variant:
config_name=libero_joint_2cam224_1e-4
run_id=$(date +%Y%m%d_%H%M%S)
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
python experiments/libero/run_libero_manager.py \
"task=${config_name}" \
"ckpt=checkpoints/simplewam_checkpoints/${config_name}.pt" \
EVALUATION.dataset_stats_path=checkpoints/simplewam_checkpoints/libero_dataset_stats.json \
"EVALUATION.output_dir=./evaluate_results/libero_plus/${config_name}/${run_id}" \
'MULTIRUN.task_suite_names=[libero_spatial,libero_object,libero_goal,libero_10]' \
MULTIRUN.num_gpus=8 \
MULTIRUN.max_tasks_per_gpu=1 \
MULTIRUN.persistent_workers=true \
EVALUATION.num_trials=1 \
EVALUATION.num_inference_steps=10 \
EVALUATION.freeze_future_video_noise=true \
EVALUATION.torch_compile_action=true \
EVALUATION.offload_text_encoder=trueAfter evaluation, summarize Camera, Robot, Language, and the other components:
python experiments/libero/summarize_libero_plus_components.py \
--run-dir "evaluate_results/libero_plus/${config_name}/${run_id}"The summary CSV and JSON files are written into that run directory.
Our reference RoboTwin environment uses CUDA Toolkit 12.8.1 and PyTorch
2.7.1+cu128. The key simulator packages are SAPIEN 3.0.0b1, mplib
0.2.1, cuRobo 0.7.5, Warp 1.14.0, Open3D 0.19.0, and toppra
0.6.8. RoboTwin rendering also requires a working Vulkan setup and ffmpeg.
Machine and driver setups vary. Follow the official RoboTwin repository for installation instructions and environment-specific troubleshooting, while using the versions above as the reference configuration for reproducing our evaluation.
After configuring RoboTwin, extract the released assets and expose the Simple-WAM policy and checkpoints:
tar -xzf checkpoints/benchmark_assets/robotwin_assets.tar.gz -C third_party/RoboTwin
ln -sfn ../../../experiments/robotwin/simplewam_policy \
third_party/RoboTwin/policy/simplewam_policy
ln -sfn ../../checkpoints third_party/RoboTwin/checkpointsEvaluate the few-shot checkpoint on all tasks, in both clean and randomized phases:
config_name=robotwin_joint_3cam_384_1e-4_fewshot10
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
python experiments/robotwin/run_robotwin_manager.py \
"task=${config_name}" \
"ckpt=checkpoints/simplewam_checkpoints/${config_name}.pt" \
EVALUATION.dataset_stats_path=checkpoints/simplewam_checkpoints/robotwin_dataset_stats.json \
EVALUATION.eval_phase=both \
EVALUATION.eval_num_episodes=100 \
EVALUATION.num_inference_steps=10 \
EVALUATION.freeze_future_video_noise=true \
EVALUATION.torch_compile_action=true \
EVALUATION.offload_text_encoder=true \
MULTIRUN.num_gpus=8 \
MULTIRUN.max_tasks_per_gpu=1For the generalization checkpoints, change config_name and append the matching
override to evaluate only the held-out tasks:
config_name |
Additional override |
|---|---|
robotwin_joint_3cam_384_1e-4_taskgen |
'EVALUATION.task_names=${data.train.robotwin_drop_tasks}' |
robotwin_joint_3cam_384_1e-4_taskgen_video |
'EVALUATION.task_names=${data.train.robotwin_action_loss_drop_tasks}' |
Run the following commands from the repository root after installing Simple-WAM
and extracting the datasets. Use the full-data configs below, not a few-shot or
task-generalization config, so the caches cover all training subsets.
For RoboTwin, place the normalization statistics at
data/robotwin2.0/dataset_stats.json before computing video or metadata caches.
The preprocessing scripts download missing Wan weights and tokenizer files on first use. Set their local cache directory:
mkdir -p checkpoints
export DIFFSYNTH_MODEL_BASE_PATH="$(pwd)/checkpoints"Interpolate the pretrained Wan2.2 video backbone to the 1024-dimensional ActionDiT backbone, with alpha scaling enabled:
python scripts/preprocess_action_dit_backbone.py \
--model-config configs/model/simplewam.yaml \
--output checkpoints/ActionDiT_linear_interp_Wan22_alphascale_1024hdim.pt \
--device cuda \
--dtype bfloat16Encode the task instructions from each dataset's meta/tasks.jsonl:
# LIBERO
CUDA_VISIBLE_DEVICES=0 python scripts/precompute_text_embeds.py \
task=libero_joint_2cam224_1e-4 \
+overwrite=false
# RoboTwin
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
torchrun --standalone --nproc_per_node=8 scripts/precompute_text_embeds.py \
task=robotwin_joint_3cam_384_1e-4 \
+overwrite=falseOutputs are written to data/text_embeds_cache/libero/ and
data/text_embeds_cache/robotwin/, as configured by
data.train.text_embedding_cache_dir. Existing entries are skipped; use
+overwrite=true to regenerate them. The script loads the text encoder even
though training uses model.load_text_encoder=false.
Precomputed video latents matching the released configs are available from rpzhou/simplewam-precompute-cache. Download and install them from the repository root:
huggingface-cli download rpzhou/simplewam-precompute-cache \
--repo-type dataset \
--include "video_latents/libero.tar.gz" "video_latents/robotwin/*" \
--local-dir ./cache_downloads
mkdir -p ./data/video_latents/libero ./data/video_latents/robotwin
tar -xzf ./cache_downloads/video_latents/libero.tar.gz \
-C ./data/video_latents/libero --strip-components=1
set -o pipefail
cat ./cache_downloads/video_latents/robotwin/robotwin_part_{aa,ab,ac,ad,ae,af} | \
tar -xzf - -C ./data/video_latents/robotwin --strip-components=1Alternatively, compute the caches locally. The commands below use eight GPUs,
with one process per GPU. Adjust
CUDA_VISIBLE_DEVICES and --nproc_per_node together for your machine.
# LIBERO
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
torchrun --standalone --nproc_per_node=8 scripts/precompute_video_latents.py \
task=libero_joint_2cam224_1e-4 \
data.train.use_precomputed_metadata_cache=false \
data.train.metadata_cache_dir=null \
data.train.video_latent_cache_dir=./data/video_latents/libero \
data.train.video_latent_storage_format=npz_compressed \
overwrite_video_latents=false \
video_latent_save_dtype=fp16 \
video_latent_batch_size=8 \
video_latent_num_workers=8
# RoboTwin
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
torchrun --standalone --nproc_per_node=8 scripts/precompute_video_latents.py \
task=robotwin_joint_3cam_384_1e-4 \
data.train.robotwin_clean_only=false \
data.train.use_precomputed_metadata_cache=false \
data.train.metadata_cache_dir=null \
data.train.video_latent_cache_dir=./data/video_latents/robotwin \
data.train.video_latent_storage_format=npz_compressed \
overwrite_video_latents=false \
video_latent_save_dtype=fp16 \
video_latent_batch_size=8 \
video_latent_num_workers=8The output paths match the current training configs. ./data/video_latents/libero
and ./data/video_latents/robotwin are relative to the current working directory.
Run these commands from the repository root to store caches under its data/.
If you change the VAE, frame sampling, resolution, or camera layout, generate a
new cache and update data.train.video_latent_cache_dir accordingly.
Ensure data/robotwin2.0/dataset_stats.json is present first. This CPU-only step
caches normalized actions, proprioception, padding masks, and instructions; it
does not run the VAE or text encoder.
A compatible precomputed cache is also available from rpzhou/simplewam-precompute-cache:
huggingface-cli download rpzhou/simplewam-precompute-cache \
--repo-type dataset \
--include "metadata_cache/robotwin/*" \
--local-dir ./cache_downloads
mkdir -p ./data/metadata_cache/robotwin
cp -a ./cache_downloads/metadata_cache/robotwin/. \
./data/metadata_cache/robotwin/To recompute it locally, note that the metadata script expects a resolved configuration, not a task YAML with Hydra defaults. Export one without starting preprocessing:
mkdir -p ./data/metadata_cache
python scripts/precompute_video_latents.py \
task=robotwin_joint_3cam_384_1e-4 \
data.train.robotwin_clean_only=false \
data.train.use_precomputed_video_latents=false \
data.train.use_precomputed_metadata_cache=false \
data.train.metadata_cache_dir=null \
--cfg job --resolve > ./data/metadata_cache/robotwin_precompute.yaml
python scripts/precompute_metadata_cache.py \
--config ./data/metadata_cache/robotwin_precompute.yaml \
--output-cache-dir ./data/metadata_cache/robotwin \
--shard-size 10000 \
--batch-size 64 \
--num-workers 8The relative output path ./data/metadata_cache/robotwin matches
data.train.metadata_cache_dir in the RoboTwin training config. Run from the
repository root. Existing arrays are skipped unless
--overwrite is passed.
Regenerate the cache if normalization statistics, action processing, or sample
indexing changes.
Run from the repository root after preparing the datasets, ActionDiT
initialization, text cache, and video latents above. RoboTwin also uses the
metadata cache. Training uses DeepSpeed ZeRO-1 and BF16 through
scripts/train_zero1.sh; no simulator environment is needed for training.
Train on all four suites with eight GPUs:
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
bash scripts/train_zero1.sh 8 \
task=libero_joint_2cam224_1e-4 \
gradient_accumulation_steps=2 \
wandb.mode=offlineReplace task= with the corresponding config below; keep the other arguments
unchanged:
| Task config | Training setting |
|---|---|
libero_joint_2cam224_1e-4 |
Full data, all four suites |
libero_joint_2cam224_1e-4_fewshot5 |
5 demonstrations per task |
libero_joint_2cam224_1e-4_fewshot10 |
10 demonstrations per task |
libero_joint_2cam224_1e-4_taskgen_spatial |
Exclude the spatial suite |
libero_joint_2cam224_1e-4_taskgen_object |
Exclude the object suite |
libero_joint_2cam224_1e-4_taskgen_goal |
Exclude the goal suite |
libero_joint_2cam224_1e-4_taskgen_long |
Exclude the long-horizon suite (libero_10) |
libero_joint_2cam224_1e-4_taskgen_video_spatial |
Spatial suite contributes video supervision only |
libero_joint_2cam224_1e-4_taskgen_video_object |
Object suite contributes video supervision only |
libero_joint_2cam224_1e-4_taskgen_video_goal |
Goal suite contributes video supervision only |
libero_joint_2cam224_1e-4_taskgen_video_long |
Long-horizon suite contributes video supervision only |
LIBERO-plus evaluates the same trained LIBERO models; it does not require a separate training command.
Train the few-shot setting with eight GPUs:
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 \
bash scripts/train_zero1.sh 8 \
task=robotwin_joint_3cam_384_1e-4_fewshot10 \
gradient_accumulation_steps=16 \
wandb.mode=offlineReplace task= to select another training setting:
| Task config | Training setting |
|---|---|
robotwin_joint_3cam_384_1e-4 |
Full clean and randomized data |
robotwin_joint_3cam_384_1e-4_fewshot10 |
10 episodes per task sampled from the combined clean and randomized data |
robotwin_joint_3cam_384_1e-4_taskgen |
Exclude tasks listed in data.train.robotwin_drop_tasks |
robotwin_joint_3cam_384_1e-4_taskgen_video |
Tasks in data.train.robotwin_action_loss_drop_tasks contribute video supervision only |
The examples use global batch sizes of 128 (LIBERO) and 1024 (RoboTwin), giving a per-GPU microbatch of 8 in both commands:
per_gpu_batch_size = global_batch_size / (num_gpus * gradient_accumulation_steps)
When changing GPU count, update both CUDA_VISIBLE_DEVICES and the positional
argument to train_zero1.sh. Increase gradient_accumulation_steps to reduce
per-GPU memory while keeping global_batch_size fixed; the division above must
remain exact. Epoch counts, checkpoint intervals, and any step limits come from
the selected task config.
Each launch writes its resolved config.yaml, checkpoints/weights/step_XXXXXX.pt,
and checkpoints/state/ under runs/<task>/<timestamp>/. wandb.mode=offline keeps W&B
logging local; use wandb.enabled=false to disable it entirely.
We thank the authors and maintainers of the following open-source projects:
- FastWAM for the codebase on which Simple-WAM is built and the preprocessed datasets used in this repository.
- LIBERO for the benchmark and simulation environment.
- LIBERO-plus for the robustness benchmark.
- RoboTwin for the simulation environment, tasks, and assets.
@misc{zhou2026makesworldactionmodels,
title={What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling},
author={Renping Zhou and Zanlin Ni and Zihao Fan and Guohao Fu and Zeyu Liu and Hao Shi and Jie Zhang and Chi Bene Chen and Yang Yue and Xueyang Fu and Gao Huang},
year={2026},
eprint={2609.34981},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.34981},
}
