Minqi Wang2, Minghao Yang3, Dake Zhong3, Guorui Song3, Yu Zhang1, Xianming Liu1, Boyang Wang1†
1 XPeng Motors 2 The Chinese University of Hong Kong 3 Tsinghua University
- September 29, 2026: The VGGT-Diff paper is now available on arXiv. We also released the pose-free inference pipeline for generating camera-controlled videos directly from six RGB images.
- September 27, 2026: We released the VGGT-Diff paper, project page, and code.
- Release half-resolution and full-resolution VGGT-Diff checkpoints.
- Release the full-resolution checkpoint for continuous camera-trajectory generation.
- Release pose-free image-to-novel-view video inference from six RGB images.
VGGT-Diff is a geometry-routed multi-view diffusion model for sparse-view novel view synthesis. It combines visual evidence from VGGT-Omega with a pretrained video diffusion prior. Given six sparse-view images, it jointly synthesizes a camera-controlled novel-view sequence while preserving observed structure and completing unseen content. This release contains the main training and inference paths, six-view visual conditioning, per-frame camera-trajectory conditioning, and a ready-to-run garden example. Model checkpoints are not stored in this repository and will be released separately.
git clone https://github.com/chenkangjie1123/VGGT-Diff.git
cd VGGT-Diff
conda create -n vggtdiff python=3.10 -y
conda activate vggtdiff
pip install -e .
pip install git+https://github.com/facebookresearch/vggt-omega.git@399d4d62935deb71cedb1e1c35b7a90413a6bee4Request access to the VGGT-Omega checkpoint and download vggt_omega_1b_512.pt. Wan2.1 weights are downloaded automatically on first use, or can be supplied through --base-model-dir.
The default example includes six source views and an 80-frame camera trajectory:
python scripts/infer.py \
--checkpoint /path/to/vggtdiff.safetensors \
--omega-checkpoint /path/to/vggt_omega_1b_512.ptThe generated frames and video are written to outputs/garden/. Full-resolution inference is enabled with --height 480 --width 832; the lower default resolution is convenient for a first run.
The default inference path uses CPU offloading to minimize CUDA memory use. On a GPU with ample memory, set --vram-limit-gib 64 to keep more weights resident, or --vram-limit-gib -1 to disable offloading.
The pose-free entry point estimates the six source cameras with VGGT-Omega, then generates 80 novel-view frames with VGGT-Diff. It defaults to 480p (832 × 480), 12 fps, and a route through source views 0 → 1 → 2 → 3 → 4 → 5. Name the six input images in their physical walking order. The route interpolates camera centers and orientations; it does not check for collisions with scene geometry.
python scripts/infer_pose_free.py \
--checkpoint /path/to/full_resolution_vggtdiff.safetensors \
--omega-checkpoint /path/to/vggt_omega_1b_512.pt \
--source-dir /path/to/six_rgb_images \
--output outputs/pose_freeThe output contains prediction.mp4, camera_trajectory.mp4, a synchronized prediction_with_trajectory.mp4, all 80 PNG frames, and cameras.json with the estimated source cameras and the target path actually used. The camera view uses small antialiased frustums. No ground-truth poses or frames are read. You can reorder the default path with --route-order 0 2 1 3 4 5.
To provide your own 80 target cameras, pass --trajectory-json /path/to/targets.json. Give exactly one of target_w2c (OpenCV world-to-camera matrices in the VGGT-Omega-estimated world saved in cameras.json) or target_c2w_relative_to_source0 (camera-to-world matrices in the first source camera's coordinate frame). Both are arrays of 80 homogeneous 4 × 4 matrices. Optional target_intrinsics is an array of 80 pixel-space 3 × 3 matrices; if omitted, the median of the six estimated source intrinsics is used. The custom path should avoid scene collisions and keep a sensible distance from the observed views.
Create a directory with this layout:
my_scene/
├── source_views/
│ ├── 00.png
│ ├── 01.png
│ └── ...
└── trajectory.json
trajectory.json contains source_w2c, target_w2c, source_intrinsics, and target_intrinsics. Extrinsics use world-to-camera matrices in OpenCV convention; intrinsics are 3x3 pixel-space matrices at the source image resolution.
python scripts/infer.py \
--checkpoint /path/to/vggtdiff.safetensors \
--omega-checkpoint /path/to/vggt_omega_1b_512.pt \
--example /path/to/my_scene \
--output outputs/my_sceneTraining scenes follow the folder-form DL3DV convention:
dataset_root/
└── scene_id/
├── images_4/
└── transforms.json
First cache the frozen VGGT-Omega source features:
python scripts/prepare_omega_cache.py \
--dataset-root /path/to/dataset_root \
--cache-root /path/to/omega_cache \
--omega-checkpoint /path/to/vggt_omega_1b_512.ptBoth recipes freeze the Wan2.1 VAE and VGGT-Omega while training the DiT and conditioning modules. Their latent layouts differ, so their checkpoints are separate model families.
| Paper NVS model | Continuous-trajectory model | |
|---|---|---|
| Targets | Variable N, treated as a set of views |
80 ordered, contiguous frames |
| VAE | Every source and target is encoded independently | Sources are independent; targets use the causal temporal VAE |
| DiT slots | 6 + N |
6 + 21 for 80 targets |
| Plucker conditioning | One pose per target slot | Four ordered poses packed per compressed slot |
This is the independent-view protocol used for the quantitative results in the paper. Six sources and N targets produce 6 + N one-to-one image, camera, Plucker, and geometry slots without temporal VAE compression.
accelerate launch scripts/train.py \
--training-mode nvs \
--dataset-root /path/to/dataset_root \
--omega-cache /path/to/omega_cache \
--target-frames 1 2 4 \
--height 192 --width 336 \
--backbone-lr 1e-5 \
--condition-lr 1e-4 \
--output outputs/nvs_training_runThe paper progresses the half-resolution target curriculum through {1,2,4} -> {2,4,8} -> {4,8,12} -> {4,8,12,16}, then trains at 480x832 with {4,8} -> {4,8,12} -> {4,8,12,16} and a 5e-6 backbone learning rate. Resume the previous stage and change --target-frames, resolution, and learning rate accordingly.
An 80-frame target is padded to 81 frames and causally compressed into 21 target slots. Camera conditions bypass the VAE; the four ordered Plucker maps in each causal window are packed and projected by the temporal Plucker adapter.
accelerate launch scripts/train.py \
--training-mode trajectory \
--dataset-root /path/to/dataset_root \
--omega-cache /path/to/omega_cache \
--resume /path/to/vggtdiff_trajectory.safetensors \
--target-frames 80 \
--height 480 \
--width 832 \
--output outputs/trajectory_training_runThe paper's original trajectory adaptation uses 192x336; the command above selects the full-resolution continuation. Both recipes use source-anchor camera normalization and point-track residual consistency.
@article{chen2026vggtdiff,
title={VGGT-Diff: Visual Geometry Meets Diffusion for Sparse-View Novel View Synthesis},
author={Chen, Kangjie and Li, Xiangyu and Zhang, Dongbin and Zheng, Chaoda and Chen, Shijia and Deng, Jinhao and Lin, Hongbin and Choo, Sin Wai and Wang, Minqi and Yang, Minghao and Zhong, Dake and Song, Guorui and Zhang, Yu and Liu, Xianming and Wang, Boyang},
journal={arXiv preprint arXiv:2609.33253},
year={2026}
}This project builds on Wan2.1, VGGT-Omega, FrameCrafter, and DiffSynth-Studio. We thank their authors for releasing their work.
VGGT-Diff is released under the VGGT-Diff Research License for non-commercial research and educational use. Third-party components and dependencies remain subject to their respective licenses and terms; see NOTICE.