Merge branch 'main' into patch-1
This commit is contained in:
@@ -15,24 +15,28 @@
|
||||
|
||||
# :wrench: Installation
|
||||
|
||||
## Windows Users: Please use WSL as flash attention isn't supported on native Windows yet: https://github.com/pytorch/pytorch/issues/108175
|
||||
|
||||
Clone the repository:
|
||||
|
||||
```bash
|
||||
git clone --recursive https://github.com/Stability-AI/stable-virtual-camera
|
||||
```
|
||||
|
||||
To setup the virtual environment and install all necessary model dependencies, simply run:
|
||||
|
||||
```bash
|
||||
cd stable-virtual-camera
|
||||
pip install -e .
|
||||
```
|
||||
|
||||
Please note that you will need `python>=3.10` and `torch>=2.6.0`.
|
||||
|
||||
Check [INSTALL.md](docs/INSTALL.md) for other dependencies if you want to use our demos or develop from this repo.
|
||||
For windows users, please use WSL as flash attention isn't supported on native Windows [yet](https://github.com/pytorch/pytorch/issues/108175).
|
||||
|
||||
# :open_book: Usage
|
||||
|
||||
You need to properly authenticate with Hugging Face to download our model weights. Once set up, our code will handle it automatically at your first run. You can authenticate by running
|
||||
|
||||
```bash
|
||||
# This will prompt you to enter your Hugging Face credentials.
|
||||
huggingface-cli login
|
||||
```
|
||||
|
||||
Once authenticated, go to our model card [here](https://huggingface.co/stabilityai/stable-virtual-camera) and enter your information for access.
|
||||
|
||||
We provide two demos for you to interact with `Stable Virtual Camera`.
|
||||
|
||||
### :rocket: Gradio demo
|
||||
|
||||
@@ -290,7 +290,7 @@ def main(
|
||||
VERSION_DICT["T"] = [int(t) for t in T.split(",")] if isinstance(T, str) else T
|
||||
|
||||
options = VERSION_DICT["options"]
|
||||
options["chunk_strategy"] = "interp"
|
||||
options["chunk_strategy"] = "nearest-gt"
|
||||
options["video_save_fps"] = 30.0
|
||||
options["beta_linear_start"] = 5e-6
|
||||
options["log_snr_shift"] = 2.4
|
||||
|
||||
+3
-4
@@ -53,7 +53,7 @@ We provide <a href="https://github.com/Stability-AI/stable-virtual-camera/releas
|
||||
└── scene_3
|
||||
```
|
||||
|
||||
You can specify which scene to run by passing in `--data_items scene_1 scene_2` to run, for example, `scene_1` and `scene_2`.
|
||||
You can specify which scene to run by passing in `--data_items scene_1,scene_2` to run, for example, `scene_1` and `scene_2`.
|
||||
|
||||
### Recommended Usage
|
||||
|
||||
@@ -80,14 +80,13 @@ python demo.py \
|
||||
- For the evaluation in semi-dense-view regime (i.e., DL3DV-140 and Tanks and Temples dataset) with `32` input views, we zero-shot extend `T` to fit all input and target views in one forward. Specifically, we set `--T 90` for the DL3DV-140 dataset and `--T 80` for the Tanks and Temples dataset.
|
||||
- For the evaluation on ViewCrafter split (including the RealEastate10K, CO3D, and Tanks and Temples dataset), we find zero-shot extending `T` to `25` to fit all input and target views in one forward is better. Also, the V split uses the original image resolutions: we therefore set `--T 25 --L_short 576`.
|
||||
|
||||
For example, you can run the following command on the example `dl3d140-165f5af8bfe32f70595a1c9393a6e442acf7af019998275144f605b89a306557` with 1 input view:
|
||||
For example, you can run the following command on the example `dl3d140-165f5af8bfe32f70595a1c9393a6e442acf7af019998275144f605b89a306557` with 3 input views:
|
||||
|
||||
```bash
|
||||
python demo.py \
|
||||
--data_path /path/to/assets_demo_cli/ \
|
||||
--data_items dl3d140-165f5af8bfe32f70595a1c9393a6e442acf7af019998275144f605b89a306557 \
|
||||
--num_inputs 1 \
|
||||
--chunk_strategy nearest-gt \
|
||||
--num_inputs 3 \
|
||||
--video_save_fps 10
|
||||
```
|
||||
|
||||
|
||||
Reference in New Issue
Block a user