Stable Virtual Camera: Generative View Synthesis with Diffusion Models

[Jensen (Jinghao) Zhou](https://shallowtoil.github.io/)\*, [Hang Gao](https://hangg7.com/)\* [Vikram Voleti](https://voletiv.github.io/), [Aaryaman Vasishta](https://www.aaryaman.net/), [Chun-Han Yao](https://chhankyao.github.io/), [Mark Boss](https://markboss.me/) [Philip Torr](https://eng.ox.ac.uk/people/philip-torr/), [Christian Rupprecht](https://chrirupp.github.io/), [Varun Jampani](https://varunjampani.github.io/)
# Overview `Stable Virtual Camera (Seva)` is a 1.3B generalist diffusion model for Novel View Synthesis (NVS), generating 3D consistent novel views of a scene, given any number of input views and target cameras. # :tada: News - March 2025 - `Stable Virtual Camera` is out everywhere. # :wrench: Installation ```bash git clone --recursive https://github.com/Stability-AI/stable-virtual-camera cd stable-virtual-camera pip install -e . ``` Please note that you will need `python>=3.10` and `torch>=2.6.0`. Check [INSTALL.md](docs/INSTALL.md) for other dependencies if you want to use our demos or develop from this repo. For windows users, please use WSL as flash attention isn't supported on native Windows [yet](https://github.com/pytorch/pytorch/issues/108175). # :open_book: Usage We provide two demos for you to interative with `Stable Virtual Camera`. ### :rocket: Gradio demo This gradio demo is a GUI interface that requires no expertised knowledge, suitable for general users. Simply run ```bash python demo_gr.py ``` For a more detailed guide, follow [GR_USAGE.md](docs/GR_USAGE.md). ### :computer: CLI demo This cli demo allows you to pass in more options and control the model in a fine-grained way, suitable for power users and academic researchers. An examplar command line looks as simple as ```bash python demo.py --data_path [additional arguments] ``` For a more detailed guide, follow [CLI_USAGE.md](docs/CLI_USAGE.md). For users interested in benchmarking NVS models using command lines, check [`benchmark`](benchmark/) containing the details about scenes, splits, and input/target views we reported in the paper. # :books: Citing If you find this repository useful, please consider giving a star :star: and citation. ``` @article{zhou2025stable, title={Stable Virtual Camera: Generative View Synthesis with Diffusion Models}, author={Jensen (Jinghao) Zhou and Hang Gao and Vikram Voleti and Aaryaman Vasishta and Chun-Han Yao and Mark Boss and Philip Torr and Christian Rupprecht and Varun Jampani }, journal={arXiv preprint}, year={2025} } ```