Bringing Any 3D World Alive as Diverse, Endlessly Looping 3D Cinemagraphs
You-Zhe Xie![]()
,
Ting-Wei Chou
,
Yu-Hsuan Li
,
Kaipeng Zhang
,
Zhixiang Wang
†,
Yu-Lun Liu
†
National Yang Ming Chiao Tung University
Alaya Lab
† Corresponding authors
OuroWorld turns a static 3D Gaussian Splatting (3DGS) scene from any source (a reconstruction, or a 3D world model such as HY-World 2.0, Marble or Lyra 2.0) into a 3D cinemagraph: a dynamic scene with vivid, diverse motion that loops endlessly and can be viewed from moving cameras. It is mask-free and captures general deformation, object motion and illumination change, not only fluid-like motion.
![]() |
![]() |
| HY-World 2.0 · volcano_lava | Marble · mountain_lake |
A vision-language model infers plausible dynamics for the scene and guides a video model to generate a looping reference video from one view. The video is lifted to 3D and completed into multi-view videos by a video inpainting model. These are imperfect supervision, so we fit an Inconsistency-Robust Periodic 4DGS: a Fourier-series deformation field P makes the scene loop by construction (t = T returns exactly to t = 0), while a Grounded Drift Field Δ, anchored at the reference view, absorbs the cross-view inconsistency and is discarded at render time.
static 3DGS ──► scene package ──► reference video ──► multi-view videos ──► periodic 4DGS ──► loop videos
(App. A) pivot + reference GPT-5.5 + Seedance VGGT-Ω + Trajectory- P + grounded Δ static / orbit
view 2.0 (§3.1) Crafter (§4.1) + SVCO (§4.2) (§4.3)
- Project Structure
- Installation
- Released Data
- Quick Start
- Full Pipeline on Your Own Scene
- Metrics
- Evaluation and the Paper's Tables
- Ablations
- Output Format
- Hardware and Run Times
- Tests
- License and Acknowledgements
- Citation
OuroWorld/
├── ouroworld/ # The library
│ ├── geometry/ # camera and SE(3) math (pure)
│ ├── io/ # artifact formats
│ ├── gaussians/ # canonical Gaussians, densification
│ ├── render/ # rasterizer wrapper and static renders
│ ├── fields/ # P, Δ and the cinemagraph model
│ ├── datasets/ # scene importers
│ ├── reference/ # pivot picking, reference candidates
│ ├── generation/ # reference video (prompt, video model) and multi-view videos
│ ├── training/ # supervision, losses, refinement, the optimisation loop
│ ├── rendering/ # loop videos
│ ├── evaluation/ # the metrics of §4.4 and the tables
│ ├── adapters/ # the only code that reaches into third_party/
│ ├── config/ # schema and loader
│ └── cli/ # entry points; the only code that reads the config
├── configs/ # default.yaml (paper setting), evaluation.yaml, scenes/, ablations/
├── scripts/ # download_weights.py, run_released.sh
├── third_party/ # rasterizer, TrajectoryCrafter, VGGT-Omega, VBench, SEA-RAFT
├── docs/ # scene_package.md
├── tests/ # unit / smoke
└── data/ # released data (see below)
The layering is enforced by lint-imports (pyproject.toml).
Linux, an NVIDIA GPU with 24 GB (multi-view generation needs ~35 GB of CPU RAM), CUDA 12.1 and gcc ≤ 12 for the rasterizer. One conda environment covers every step.
git clone https://github.com/youzhe0305/OuroWorld.git OuroWorld && cd OuroWorld
conda env create -f environment.yml
conda activate ouroworld
CC=gcc-11 CXX=g++-11 pip install --no-build-isolation ./third_party/diff-gaussian-rasterization
pip install --no-build-isolation -e ".[refinement,reference-video,dev]"Model weights (~55 GB) go to checkpoints/. VGGT-Omega is gated: request
access at https://huggingface.co/facebook/VGGT-Omega first, then log in
with huggingface-cli login or set HF_TOKEN.
python scripts/download_weights.py # or --only train / generation / evaluationcheckpoints/ |
Used by |
|---|---|
lcm-lora-sdv1-5/ |
4D optimisation (refinement; SD 1.5 is fetched by diffusers on first use) |
vggt_omega_1b_512.pt |
multi-view generation: depth lifting |
CogVideoX-Fun-V1.1-5b-InP/, TrajectoryCrafter/ |
multi-view generation: video inpainting |
blip2-opt-2.7b/ |
multi-view generation: the video model's caption |
vbench/, sea-raft-spring-M/, i3d_torchscript.pt |
evaluation |
API keys. Only the reference-video step calls paid services (GPT-5.5 for
the motion prompt, Seedance 2.0 for the video). Copy .env.example to .env
and fill in OPENAI_API_KEY (optionally OPENAI_BASE_URL for a compatible
endpoint) and ARK_API_KEY (BytePlus ModelArk). The released scenes need no
key.
The data is on Hugging Face. Download it into data/:
huggingface-cli download YouZhe/OuroWorld-dataset --repo-type dataset --local-dir dataThe 39 scenes of the paper (App. A), under data/<kind>/<dataset>/<scene>/:
| Dataset | Scenes | Source |
|---|---|---|
| HY_World_2.0 | 10 | 3D world model, public gallery |
| Marble | 10 | 3D world model, public gallery |
| Lyra_2.0 | 10 | 3D world model, generated with its official code |
| mip-nerf | 9 | Mip-NeRF 360, reconstructed with the official 3DGS code (30k iterations) |
data/ |
Files | Content |
|---|---|---|
3dgs/ |
metadata.yaml, pivot.json, reference_camera.json, reference.png (point_cloud.ply not included, see below) |
the clip planes, the clicked pivot and the reference view of each static 3DGS |
refvideo/ |
prompt_used.txt, raw.mp4 |
the GPT-5.5 motion prompt and the Seedance 2.0 video used in the paper |
render/ |
static.mp4, orbit.mp4 |
our results (Table 2, "Ours") |
naturalness/ |
minimax_h3_seed_0000.mp4 … 0019.mp4, prompt_used.txt, reference.png |
the MiniMax-H3 reference videos of KVD (§4.4), 20 seeds per scene, with the prompt and the conditioning image |
The input static 3DGS scenes (point_cloud.ply) are not included. The
scenes of HY-World 2.0, Marble and Lyra 2.0 come from 3D world models whose
licenses do not allow us to redistribute them, so we release the cameras,
pivots and reference renders of every scene but not the Gaussians themselves.
The scenes can be obtained from their original sources (the public galleries
of HY-World 2.0 and Marble, the official code of Lyra 2.0, and the Mip-NeRF 360
dataset trained with the official 3DGS code for 30k iterations) and imported
with the importers below. If you run into
any problem obtaining them, please
contact us.
Rerun a scene of the paper, from its static 3DGS to its loop videos. First put
the scene's static 3DGS at data/3dgs/<dataset>/<scene>/point_cloud.ply
(INRIA 3DGS layout, in the coordinates of the released cameras; see
Released Data), then:
bash scripts/run_released.sh Marble world # any <dataset> <scene> under data/3dgs/It imports the released scene (pivot and reference view included), reuses the
paper's motion prompt and reference video, so no API key is needed, then
generates the multi-view videos (~70 min), trains the periodic 4DGS (~1.5 h)
and renders output/Marble/world/render/iter_XXXXX/{static,orbit}.mp4.
The paper's videos are in data/render/Marble/world/.
You need two things: a 3DGS PLY (the INRIA layout of the original 3DGS code) and a camera JSON that sees the part of the scene you want to animate. The pivot is clicked on this camera's render, and the reference view, from which the reference video is generated, is chosen around that pivot.
{
"width": 1600,
"height": 900,
"K": [[800.0, 0.0, 800.0], [0.0, 800.0, 450.0], [0.0, 0.0, 1.0]],
"c2w": [[1.0, 0.0, 0.0, 0.0], [0.0, 1.0, 0.0, 0.0], [0.0, 0.0, 1.0, 0.0], [0.0, 0.0, 0.0, 1.0]]
}c2wis the 4×4 camera-to-world pose in the OpenCV convention (x right, y down, z forward), in the coordinates of the PLY.Kis in pixels, and the principal point must be the image centre (cx = width / 2,cy = height / 2).- The camera's up direction (−y) is taken as the world's up, around which the cameras orbit, so keep the camera level.
- This is the format of
data/3dgs/*/*/reference_camera.json, so a released camera is a working example.
Then run each step; each reads the previous step's output under
output/<dataset>/<scene>/:
ouroworld-import world --ply my_scene.ply --camera my_camera.json --out output/Custom/my_scene/scene
S="dataset=Custom scene=my_scene"
ouroworld-pivot $S # 1. click the pivot on the camera's render (a browser page)
ouroworld-reference $S # render the reference view
ouroworld-reference-video $S # 2. looping reference video (paid APIs, ~6 min)
ouroworld-multiview $S # 3. 20 more views (~70 min, resumable)
ouroworld-train $S # 4. periodic 4DGS (~1.5 h)
ouroworld-render $S # 5. static.mp4 and orbit.mp4 (20 s, 30 FPS)- Without a camera JSON,
ouroworld-import worldplaces the camera on a world axis:--forward=-z --up=+y --position=0,1.5,0 --fov 90 60 --size 1600 900(defaults-x,+z, the origin, 120°×90°, 1732×1000). --clip NEAR FARsets the clip planes (default 0.01 and 100). The rasterizer also culls everything closer than 0.2 units, so a world only a few units across should be scaled up first (docs/scene_package.md).- Every setting is in
configs/default.yaml(the paper setting) and can be overridden withkey=valueor another--configfile.
Importers for the paper's four sources, and more options
One example scene per source is configured in configs/scenes/
(ouroworld-pivot --config configs/scenes/Lyra_2.0/palace_rooftops.yaml, ...).
After importing from the original source, ouroworld-pivot --pixel X Y with
reference_pixel of the released pivot.json repeats the paper's click (the
pixel is on the source's camera, so it does not apply after import released).
# Mip-NeRF 360: the trained 3DGS and the COLMAP scene it was trained on
ouroworld-import mipnerf360 --ply <3dgs-run>/point_cloud/iteration_30000/point_cloud.ply \
--colmap <mipnerf360>/room --images images_2 --out output/MipNeRF360/room/scene
# HY-World 2.0: the gallery's scene.ply, seen from HY-World's own view0
ouroworld-import world --ply <hy-world>/mossy_forest/scene.ply --out output/HY_World_2.0/jungle/scene
# Marble: the downloaded splat file, from a camera on a world axis
ouroworld-import world --ply <marble>/Jurassic_Park.ply --forward=-z --up=+y --position=-1.5,0,-1 \
--out output/Marble/jurassic_park/scene
# Lyra 2.0: scene.ply + camera.json of a generated scene; the world is scaled out of the near cull
ouroworld-import lyra --source <lyra>/lyra2_3dgs_scenes/02 --out output/Lyra_2.0/palace_rooftops/sceneouroworld-pivot --pixel X Yskips the browser page;--host/--portset where it is served.ouroworld-referencealso renders candidate views toreference_candidates/candidates.png; pick one withreference.selected=candidate_XXXand run it again.- The reference-video step keeps
prompt.txtandraw.mp4and reuses them on a rerun, so nothing is paid twice; delete them to regenerate. Wan 2.2 is an open-source alternative:reference_video.backend=wan reference_video.frame_count=73. ouroworld-train --resume output/<dataset>/<scene>/modelcontinues an interrupted run.
The evaluation needs no ground truth (§4.4). Every method is rendered under a
static camera (static) and a moving one (orbit).
| Metric | Group | Description |
|---|---|---|
| Vividness Degree ↑ | Vividness | Mean of Motion, Illumination and Visual Variation (MV, IV, VV) between adjacent frames at 3 FPS. |
| KVD ↓ | Naturalness | Kernel Video Distance to the MiniMax-H3 videos (20 seeds per scene) generated from the same reference image and prompt. |
| MALF ↑ | Loop seam coherence | Motion-Aware Loop Fidelity: frames must match the frame one period later better than other frames, so a static scene does not score by looping trivially. |
| Seam SSIM ↑ | Loop seam coherence | SSIM between the last frame of one cycle and the first frame of the next. |
| Aesthetic Quality / Overall Consistency ↑ | Scene quality | VBench aesthetic score and text-video consistency with the motion prompt. |
| Subject / Background Consistency ↑ | Scene quality (Table 3) | VBench temporal consistency of the subject and the background. |
| VoL ↑ | Sharpness (Table 3) | Variance of the Laplacian. |
The metrics run on collected renders,
results/<group>/<dataset>/<scene>/{static,orbit}.mp4, for the scenes in
results/used_scene.json; the motion prompts are read from
results/prompts.json (Overall Consistency) and the MiniMax-H3 reference
videos of KVD from data/naturalness/<dataset>/<scene>/ (configs/evaluation.yaml).
ouroworld-evaluate vividness # Vividness Degree (MV, IV, VV)
ouroworld-evaluate naturalness # KVD against the MiniMax-H3 videos
ouroworld-evaluate loop-seam # MALF and Seam SSIM (static camera)
ouroworld-evaluate scene-quality # VBench aesthetic / overall / subject / background
ouroworld-evaluate sharpness # VoL
ouroworld-evaluate tables # output/evaluation/tables.md, table2.csv, table3.csvWithout --groups, each metric runs on every group that Tables 2 and 3 need
(the baselines, ours, and the ablations under ablation/); finished results
are kept, --overwrite recomputes them. On the paper's renders this
reproduces Tables 2 and 3 digit for digit.
Each ablation of Table 3 and §4.9 is a config file added to the run; give it its own output directories, e.g. on the scene of the Quick Start:
S="dataset=Marble scene=world"
ouroworld-train --config configs/ablations/wo_drift.yaml $S paths.run_dir=output/Marble/world/model_wo_drift
ouroworld-render --config configs/ablations/wo_drift.yaml $S paths.run_dir=output/Marble/world/model_wo_drift \
paths.render_dir=output/Marble/world/render_wo_drift| Config | Paper |
|---|---|
wo_drift.yaml |
w/o Drift (Table 3) |
wo_grounding.yaml |
w/o Grounding (Table 3) |
wo_svco.yaml |
w/o Scene-View Consistent Optimization (Table 3) |
period_2t.yaml |
period-2T basis (§4.9) |
output/<dataset>/<scene>/
├── scene/ gaussians.ply, scene.json (import camera, pivot, reference camera),
│ reference.png
├── reference_candidates/ candidates.png / .json, pivot.png
├── reference_video/ reference_video.json, frames/ (241, t = 0 .. 1), prompt.txt, raw.mp4
├── multiview/ cameras.json, observations.json, frames/, intermediate/
├── model/ checkpoints/iter_XXXXX/, metrics.jsonl, config.yaml
└── render/ iter_XXXXX/{static,orbit}.mp4, loop_report.json
observations.json gives every training frame a role: reference (the
reference video), t0 (another view at t = 0, a render of the input 3DGS) or
generated.
Measured on one RTX 4090 (24 GB), Lyra_2.0/palace_rooftops:
| Step | Time | Peak memory |
|---|---|---|
| Import, pivot, reference view | < 1 min | < 4 GB GPU |
| Reference video (GPT-5.5 + Seedance 2.0) | ~6 min | API |
| Multi-view videos | 68-74 min | 24 GB GPU, 35 GB CPU RAM |
| 4D optimisation (13k iterations) | ~1.4 h | 6.4 GB GPU |
| Loop videos | a few minutes | |
| Evaluation of one group (39 scenes, both cameras) | VBench ~8 min per dimension, Vividness ~20 min |
pytest tests/unit # CPU, seconds
pytest tests/smoke # GPU, tiny end-to-end runsOur code is released under the MIT License (LICENSE). Parts derived from
Free4D (S-Lab License 1.0) and 3D Gaussian Splatting (Inria / MPII), the
vendored code in third_party/ and the model weights keep their own
licenses, several of them non-commercial, so the full pipeline is for
non-commercial research use. See NOTICE and LICENSES/. The input 3DGS
scenes of the paper are not redistributed and remain under the licenses of
their original sources.
We build on 3D Gaussian Splatting, Free4D, TrajectoryCrafter, VGGT-Omega, VBench, SEA-RAFT and LCM-LoRA.
@article{xie2026ouroworld,
title = {{OuroWorld}: Bringing Any {3D} World Alive as Diverse, Endlessly Looping {3D} Cinemagraphs},
author = {Xie, You-Zhe and Chou, Ting-Wei and Li, Yu-Hsuan and Zhang, Kaipeng and Wang, Zhixiang and Liu, Yu-Lun},
journal = {arXiv preprint arXiv:2610.12461},
year = {2026}
}

