diff --git a/__init__.py b/__init__.py
index e69de29..cc26a06 100644
--- a/__init__.py
+++ b/__init__.py
@@ -0,0 +1,2 @@
+# -*- coding: utf-8 -*-
+# Copyright (c) Alibaba, Inc. and its affiliates.
diff --git a/docs/zh_cn/scepter/utils/utils.md b/docs/zh_cn/scepter/utils/utils.md
index be8917a..3140b43 100644
--- a/docs/zh_cn/scepter/utils/utils.md
+++ b/docs/zh_cn/scepter/utils/utils.md
@@ -914,7 +914,7 @@ data = {
_model(data)
probe = _model.probe_data()
for key in probe:
- print(key, probe[key].to_log(prefix=f"xxx/dev_easytorch/{key}"))
+ print(key, probe[key].to_log(prefix=f"xxx/{key}"))
```
diff --git a/readme.md b/readme.md
index 7ed2242..a21623e 100644
--- a/readme.md
+++ b/readme.md
@@ -14,6 +14,7 @@
- [Installation](#-Installation)
- [Getting Started](#-getting-started)
- [SCEPTER Studio](#-scepter-studio)
+- [Gallery](#-gallery)
- [Features](#-features)
- [Learn More](#-learn-more)
- [License](#license)
@@ -43,6 +44,7 @@ Currently supported approaches (and counting):
3. Res-Tuning(TODO): [Res-Tuning: A Flexible and Efficient Tuning Paradigm via Unbinding Tuner from Backbone](https://arxiv.org/abs/2310.19859) [](https://arxiv.org/abs/2310.19859) [](https://res-tuning.github.io/)
## 🎉 News
+- [2024.02]: We release new SCEdit controllable image synthesis models for SD v2.1 and SD XL. Multiple strategies applied to accelerate inference time for SCEPTER Studio.
- [2024.01]: We release **SCEPTER Studio**, an integrated toolkit for data management, model training and inference based on [Gradio](https://www.gradio.app/).
- [2024.01]: [SCEdit](https://arxiv.org/abs/2312.11392) support controllable image synthesis for training and inference.
- [2023.12]: We propose [SCEdit](https://arxiv.org/abs/2312.11392), an efficient and controllable generation framework.
@@ -56,13 +58,17 @@ Currently supported approaches (and counting):
conda env create -f environment.yaml
conda activate scepter
```
+- We recommend installing the specific version of PyTorch and accelerate toolbox [xFormers](https://pypi.org/project/xformers/). You can install these recommended version by pip:
+
+```shell
+pip install -r requirements/recommended.txt
+```
- Install SCEPTER by the `pip` command:
```shell
pip install scepter
```
-- PS: We recommend installing PyTorch follwing [official documentation](https://pytorch.org/get-started/locally/)
## 🚀 Getting Started
@@ -159,7 +165,6 @@ python scepter/tools/run_inference.py --cfg scepter/methods/scedit/ctr/sd21_768_
python scepter/tools/run_inference.py --cfg scepter/methods/scedit/ctr/sd21_768_sce_ctr_pose.yaml --num_samples 1 --prompt 'super mario' --save_folder 'test_mario_pose' --image_size 768 --task control --image 'asset/images/pose_source.png' --control_mode source --pretrained_model ms://damo/scepter_scedit@controllable_model/SD2.1/pose_control/0_SwiftSCETuning/pytorch_model.bin # pose
```
-
## 🖥️ SCEPTER Studio
### Launch
@@ -167,24 +172,49 @@ python scepter/tools/run_inference.py --cfg scepter/methods/scedit/ctr/sd21_768_
To fully experience **SCEPTER Studio**, you can launch the following command line:
```shell
-pip install scepter
-python -m scepter.tools.webui
-```
-or run after clone repo code
-```shell
-git clone https://github.com/modelscope/scepter.git
PYTHONPATH=. python scepter/tools/webui.py --cfg scepter/methods/studio/scepter_ui.yaml
```
-The startup of **SCEPTER Studio** eliminates the need for manual downloading and organizing of models; it will automatically load the corresponding models and store them in a local directory.
-Depending on the network and hardware situation, the initial startup usually requires 15-60 minutes, primarily involving the download and processing of SDv1.5, SDv2.1, and SDXL models.
-Therefore, subsequent startups will become much faster (about one minute) as downloading is no longer required.
-
-
### Modelscope Studio
We deploy a work studio on Modelscope that includes only the inference tab, please refer to [ms_scepter_studio](https://www.modelscope.cn/studios/damo/scepter_studio/summary)
+## 🖼️ Gallery
+
+### Dragon Year Special: Dragon Tuner
+
+
+
+ | Gold Dragon Tuner |
+ Sloppy Dragon Tuner |
+ Red Dragon Tuner + Papercraft Mantra |
+ Azure Dragon Tuner + Pose Control |
+
+
+  |
+  |
+  |
+  |
+
+
+
+### Text Effect Image
+
+
+
+ | Conditional Image |
+ Midas Control "Race track, top view" |
+ Midas Control + Watercolor Mantra "white lilies" |
+ Midas Control + Dragon Tuner "Spring Festival, Chinese dragon" |
+
+
+  |
+  |
+  |
+  |
+
+
+
## ✨ Features
### Text-to-Image Generation
@@ -201,8 +231,8 @@ We deploy a work studio on Modelscope that includes only the inference tab, plea
| **Model** | **Canny** | **HED** | **Depth** | **Pose** | **Color** |
|:---------:|:---------:|:-------:|:---------:|:--------:|:---------:|
| SD 1.5 | ✅ | ✅ | ✅ | ✅ | ✅ |
-| SD 2.1 | 🪄 | ✅ | ✅ | 🪄 | 🪄 |
-| SD XL | ✅ | ✅ | ✅ | ✅ | ✅ |
+| SD 2.1 | 🪄 | 🪄 | 🪄 | 🪄 | 🪄 |
+| SD XL | 🪄 | 🪄 | 🪄 | 🪄 | 🪄 |
### Model URL
@@ -210,9 +240,9 @@ We deploy a work studio on Modelscope that includes only the inference tab, plea
- 🪄 denotes that the model has been published.
- More models will be released in the future.
-| Model | URL |
-|--------|-------------------------------------------------------------------------------------|
-| SCEdit | [ModelCard](https://modelscope.cn/models/damo/scepter_scedit/summary) |
+| Model | URL |
+|--------|------------------------------------------------------------------------------------------------------------------------------------------------|
+| SCEdit | [ModelScope](https://modelscope.cn/models/damo/scepter_scedit/summary) [HuggingFace](https://huggingface.co/scepter-studio/scepter_scedit) |
PS: Scripts running within the SCEPTER framework will automatically fetch and load models based on the required dependency files, eliminating the need for manual downloads.
diff --git a/requirements/framework.txt b/requirements/framework.txt
index 3e7a72e..3e12b56 100644
--- a/requirements/framework.txt
+++ b/requirements/framework.txt
@@ -7,5 +7,6 @@ opencv-python
opencv_transforms>=0.0.6
oss2>=2.15.0
pyyaml>=5.3.1
+scikit-image
+torchsde
transformers
-xformers>=0.0.21
diff --git a/requirements/recommended.txt b/requirements/recommended.txt
new file mode 100644
index 0000000..d25969c
--- /dev/null
+++ b/requirements/recommended.txt
@@ -0,0 +1,3 @@
+torch==2.0.1
+torchvision==0.15.2
+xformers==0.0.21
diff --git a/scepter/methods/examples/generation/stable_diffusion_1.5_512.yaml b/scepter/methods/examples/generation/stable_diffusion_1.5_512.yaml
index 7918211..207114d 100644
--- a/scepter/methods/examples/generation/stable_diffusion_1.5_512.yaml
+++ b/scepter/methods/examples/generation/stable_diffusion_1.5_512.yaml
@@ -117,7 +117,7 @@ SOLVER:
SAMPLE_STEPS: 50
SEED: 2023
GUIDE_SCALE: 7.5
- GUIDE_RESCALE:
+ GUIDE_RESCALE: 0.5
DISCRETIZATION: trailing
IMAGE_SIZE: [512, 512]
RUN_TRAIN_N: False
diff --git a/scepter/methods/examples/generation/stable_diffusion_1.5_512_lora.yaml b/scepter/methods/examples/generation/stable_diffusion_1.5_512_lora.yaml
index 4df6747..8e8af4d 100644
--- a/scepter/methods/examples/generation/stable_diffusion_1.5_512_lora.yaml
+++ b/scepter/methods/examples/generation/stable_diffusion_1.5_512_lora.yaml
@@ -125,7 +125,7 @@ SOLVER:
SAMPLE_STEPS: 50
SEED: 2023
GUIDE_SCALE: 7.5
- GUIDE_RESCALE:
+ GUIDE_RESCALE: 0.5
DISCRETIZATION: trailing
IMAGE_SIZE: [512, 512]
RUN_TRAIN_N: False
diff --git a/scepter/methods/examples/generation/stable_diffusion_2.1_512.yaml b/scepter/methods/examples/generation/stable_diffusion_2.1_512.yaml
new file mode 100644
index 0000000..f148649
--- /dev/null
+++ b/scepter/methods/examples/generation/stable_diffusion_2.1_512.yaml
@@ -0,0 +1,214 @@
+ENV:
+ BACKEND: nccl
+SOLVER:
+ NAME: LatentDiffusionSolver
+ RESUME_FROM:
+ LOAD_MODEL_ONLY: True
+ USE_FSDP: False
+ SHARDING_STRATEGY:
+ USE_AMP: True
+ DTYPE: float16
+ CHANNELS_LAST: True
+ MAX_STEPS: 2000
+ MAX_EPOCHS: -1
+ NUM_FOLDS: 1
+ ACCU_STEP: 1
+ EVAL_INTERVAL: 100
+ #
+ WORK_DIR: ./cache/save_data/sd21_512_full
+ LOG_FILE: std_log.txt
+ #
+ FILE_SYSTEM:
+ NAME: "ModelscopeFs"
+ TEMP_DIR: "./cache/data"
+ #
+ MODEL:
+ NAME: LatentDiffusion
+ PARAMETERIZATION: eps
+ TIMESTEPS: 1000
+ MIN_SNR_GAMMA:
+ ZERO_TERMINAL_SNR: False
+ PRETRAINED_MODEL: ms://AI-ModelScope/stable-diffusion-2-1-base@v2-1_512-ema-pruned.safetensors
+ IGNORE_KEYS: [ ]
+ SCALE_FACTOR: 0.18215
+ SIZE_FACTOR: 8
+ DEFAULT_N_PROMPT:
+ SCHEDULE_ARGS:
+ "NAME": "scaled_linear"
+ "BETA_MIN": 0.00085
+ "BETA_MAX": 0.012
+ USE_EMA: False
+ #
+ DIFFUSION_MODEL:
+ NAME: DiffusionUNet
+ IN_CHANNELS: 4
+ OUT_CHANNELS: 4
+ MODEL_CHANNELS: 320
+ NUM_HEADS_CHANNELS: 64
+ NUM_RES_BLOCKS: 2
+ ATTENTION_RESOLUTIONS: [ 4, 2, 1 ]
+ CHANNEL_MULT: [ 1, 2, 4, 4 ]
+ CONV_RESAMPLE: True
+ DIMS: 2
+ USE_CHECKPOINT: False
+ USE_SCALE_SHIFT_NORM: False
+ RESBLOCK_UPDOWN: False
+ USE_SPATIAL_TRANSFORMER: True
+ TRANSFORMER_DEPTH: 1
+ CONTEXT_DIM: 1024
+ DISABLE_MIDDLE_SELF_ATTN: False
+ USE_LINEAR_IN_TRANSFORMER: True
+ PRETRAINED_MODEL:
+ #
+ FIRST_STAGE_MODEL:
+ NAME: AutoencoderKL
+ EMBED_DIM: 4
+ PRETRAINED_MODEL:
+ IGNORE_KEYS: [ ]
+ BATCH_SIZE: 4
+ #
+ ENCODER:
+ NAME: Encoder
+ CH: 128
+ OUT_CH: 3
+ NUM_RES_BLOCKS: 2
+ IN_CHANNELS: 3
+ ATTN_RESOLUTIONS: [ ]
+ CH_MULT: [ 1, 2, 4, 4 ]
+ Z_CHANNELS: 4
+ DOUBLE_Z: True
+ DROPOUT: 0.0
+ RESAMP_WITH_CONV: True
+ #
+ DECODER:
+ NAME: Decoder
+ CH: 128
+ OUT_CH: 3
+ NUM_RES_BLOCKS: 2
+ IN_CHANNELS: 3
+ ATTN_RESOLUTIONS: [ ]
+ CH_MULT: [ 1, 2, 4, 4 ]
+ Z_CHANNELS: 4
+ DROPOUT: 0.0
+ RESAMP_WITH_CONV: True
+ GIVE_PRE_END: False
+ TANH_OUT: False
+ #
+ TOKENIZER:
+ NAME: OpenClipTokenizer
+ LENGTH: 77
+ #
+ COND_STAGE_MODEL:
+ NAME: FrozenOpenCLIPEmbedder
+ ARCH: ViT-H-14
+ PRETRAINED_MODEL:
+ LAYER: penultimate
+ #
+ LOSS:
+ NAME: ReconstructLoss
+ LOSS_TYPE: l2
+ #
+ SAMPLE_ARGS:
+ SAMPLER: ddim
+ SAMPLE_STEPS: 50
+ SEED: 2023
+ GUIDE_SCALE: 7.5
+ GUIDE_RESCALE:
+ DISCRETIZATION: trailing
+ IMAGE_SIZE: [512, 512]
+ RUN_TRAIN_N: False
+ #
+ OPTIMIZER:
+ NAME: AdamW
+ LEARNING_RATE: 0.0064
+ BETAS: [ 0.9, 0.999 ]
+ EPS: 1e-8
+ WEIGHT_DECAY: 1e-2
+ AMSGRAD: False
+ #
+ TRAIN_DATA:
+ NAME: ImageTextPairMSDataset
+ MODE: train
+ MS_DATASET_NAME: style_custom_dataset
+ MS_DATASET_NAMESPACE: damo
+ MS_DATASET_SUBNAME: 3D
+ PROMPT_PREFIX: ""
+ MS_DATASET_SPLIT: train_short
+ MS_REMAP_KEYS: { 'Image:FILE': 'Target:FILE' }
+ REPLACE_STYLE: False
+ PIN_MEMORY: True
+ BATCH_SIZE: 1
+ NUM_WORKERS: 4
+ SAMPLER:
+ NAME: LoopSampler
+ TRANSFORMS:
+ - NAME: LoadImageFromFile
+ RGB_ORDER: RGB
+ BACKEND: pillow
+ - NAME: Resize
+ SIZE: 512
+ INTERPOLATION: bilinear
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: pillow
+ - NAME: CenterCrop
+ SIZE: 512
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: pillow
+ - NAME: ImageToTensor
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: pillow
+ - NAME: Normalize
+ MEAN: [ 0.5, 0.5, 0.5 ]
+ STD: [ 0.5, 0.5, 0.5 ]
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'image' ]
+ BACKEND: torchvision
+ - NAME: Select
+ KEYS: [ 'image', 'prompt' ]
+ META_KEYS: [ 'data_key' ]
+ #
+ EVAL_DATA:
+ NAME: ImageTextPairMSDataset
+ MODE: eval
+ MS_DATASET_NAME: style_custom_dataset
+ MS_DATASET_NAMESPACE: damo
+ MS_DATASET_SUBNAME: 3D
+ PROMPT_PREFIX: ""
+ MS_REMAP_KEYS: { 'Image': 'Target:FILE' }
+ MS_DATASET_SPLIT: test_short
+ OUTPUT_SIZE: [512, 512]
+ REPLACE_STYLE: False
+ PIN_MEMORY: True
+ BATCH_SIZE: 4
+ NUM_WORKERS: 4
+ FILE_SYSTEM:
+ NAME: "ModelscopeFs"
+ TEMP_DIR: "./cache/data"
+ #
+ TRANSFORMS:
+ -
+ NAME: Select
+ KEYS: ['prompt']
+ META_KEYS: ['image_size']
+ #
+ TRAIN_HOOKS:
+ -
+ NAME: BackwardHook
+ PRIORITY: 0
+ -
+ NAME: LogHook
+ LOG_INTERVAL: 50
+ -
+ NAME: CheckpointHook
+ INTERVAL: 1000
+ -
+ NAME: ProbeDataHook
+ PROB_INTERVAL: 100
+ #
+ EVAL_HOOKS:
+ -
+ NAME: ProbeDataHook
+ PROB_INTERVAL: 100
diff --git a/scepter/methods/examples/generation/stable_diffusion_2.1_512_lora.yaml b/scepter/methods/examples/generation/stable_diffusion_2.1_512_lora.yaml
new file mode 100644
index 0000000..2ea41b4
--- /dev/null
+++ b/scepter/methods/examples/generation/stable_diffusion_2.1_512_lora.yaml
@@ -0,0 +1,223 @@
+ENV:
+ BACKEND: nccl
+SOLVER:
+ NAME: LatentDiffusionSolver
+ RESUME_FROM:
+ LOAD_MODEL_ONLY: True
+ USE_FSDP: False
+ SHARDING_STRATEGY:
+ USE_AMP: True
+ DTYPE: float16
+ CHANNELS_LAST: True
+ MAX_STEPS: 2000
+ MAX_EPOCHS: -1
+ NUM_FOLDS: 1
+ ACCU_STEP: 1
+ EVAL_INTERVAL: 100
+ #
+ WORK_DIR: ./cache/save_data/sd21_512_lora
+ LOG_FILE: std_log.txt
+ #
+ FILE_SYSTEM:
+ NAME: "ModelscopeFs"
+ TEMP_DIR: "./cache/data"
+ #
+ TUNER:
+ -
+ NAME: SwiftLoRA
+ R: 64
+ LORA_ALPHA: 64
+ LORA_DROPOUT: 0.0
+ BIAS: "none"
+ TARGET_MODULES: model.*(to_q|to_k|to_v|to_out.0|net.0.proj|net.2)$
+ #
+ MODEL:
+ NAME: LatentDiffusion
+ PARAMETERIZATION: eps
+ TIMESTEPS: 1000
+ MIN_SNR_GAMMA:
+ ZERO_TERMINAL_SNR: False
+ PRETRAINED_MODEL: ms://AI-ModelScope/stable-diffusion-2-1-base@v2-1_512-ema-pruned.safetensors
+ IGNORE_KEYS: [ ]
+ SCALE_FACTOR: 0.18215
+ SIZE_FACTOR: 8
+ DEFAULT_N_PROMPT:
+ SCHEDULE_ARGS:
+ "NAME": "scaled_linear"
+ "BETA_MIN": 0.00085
+ "BETA_MAX": 0.012
+ USE_EMA: False
+ #
+ DIFFUSION_MODEL:
+ NAME: DiffusionUNet
+ IN_CHANNELS: 4
+ OUT_CHANNELS: 4
+ MODEL_CHANNELS: 320
+ NUM_HEADS_CHANNELS: 64
+ NUM_RES_BLOCKS: 2
+ ATTENTION_RESOLUTIONS: [ 4, 2, 1 ]
+ CHANNEL_MULT: [ 1, 2, 4, 4 ]
+ CONV_RESAMPLE: True
+ DIMS: 2
+ USE_CHECKPOINT: False
+ USE_SCALE_SHIFT_NORM: False
+ RESBLOCK_UPDOWN: False
+ USE_SPATIAL_TRANSFORMER: True
+ TRANSFORMER_DEPTH: 1
+ CONTEXT_DIM: 1024
+ DISABLE_MIDDLE_SELF_ATTN: False
+ USE_LINEAR_IN_TRANSFORMER: True
+ PRETRAINED_MODEL:
+ #
+ FIRST_STAGE_MODEL:
+ NAME: AutoencoderKL
+ EMBED_DIM: 4
+ PRETRAINED_MODEL:
+ IGNORE_KEYS: [ ]
+ BATCH_SIZE: 4
+ #
+ ENCODER:
+ NAME: Encoder
+ CH: 128
+ OUT_CH: 3
+ NUM_RES_BLOCKS: 2
+ IN_CHANNELS: 3
+ ATTN_RESOLUTIONS: [ ]
+ CH_MULT: [ 1, 2, 4, 4 ]
+ Z_CHANNELS: 4
+ DOUBLE_Z: True
+ DROPOUT: 0.0
+ RESAMP_WITH_CONV: True
+ #
+ DECODER:
+ NAME: Decoder
+ CH: 128
+ OUT_CH: 3
+ NUM_RES_BLOCKS: 2
+ IN_CHANNELS: 3
+ ATTN_RESOLUTIONS: [ ]
+ CH_MULT: [ 1, 2, 4, 4 ]
+ Z_CHANNELS: 4
+ DROPOUT: 0.0
+ RESAMP_WITH_CONV: True
+ GIVE_PRE_END: False
+ TANH_OUT: False
+ #
+ TOKENIZER:
+ NAME: OpenClipTokenizer
+ LENGTH: 77
+ #
+ COND_STAGE_MODEL:
+ NAME: FrozenOpenCLIPEmbedder
+ ARCH: ViT-H-14
+ PRETRAINED_MODEL:
+ LAYER: penultimate
+ #
+ LOSS:
+ NAME: ReconstructLoss
+ LOSS_TYPE: l2
+ #
+ SAMPLE_ARGS:
+ SAMPLER: ddim
+ SAMPLE_STEPS: 50
+ SEED: 2023
+ GUIDE_SCALE: 7.5
+ GUIDE_RESCALE:
+ DISCRETIZATION: trailing
+ IMAGE_SIZE: [512, 512]
+ RUN_TRAIN_N: False
+ #
+ OPTIMIZER:
+ NAME: AdamW
+ LEARNING_RATE: 0.0064
+ BETAS: [ 0.9, 0.999 ]
+ EPS: 1e-8
+ WEIGHT_DECAY: 1e-2
+ AMSGRAD: False
+ #
+ TRAIN_DATA:
+ NAME: ImageTextPairMSDataset
+ MODE: train
+ MS_DATASET_NAME: style_custom_dataset
+ MS_DATASET_NAMESPACE: damo
+ MS_DATASET_SUBNAME: 3D
+ PROMPT_PREFIX: ""
+ MS_DATASET_SPLIT: train_short
+ MS_REMAP_KEYS: { 'Image:FILE': 'Target:FILE' }
+ REPLACE_STYLE: False
+ PIN_MEMORY: True
+ BATCH_SIZE: 1
+ NUM_WORKERS: 4
+ SAMPLER:
+ NAME: LoopSampler
+ TRANSFORMS:
+ - NAME: LoadImageFromFile
+ RGB_ORDER: RGB
+ BACKEND: pillow
+ - NAME: Resize
+ SIZE: 512
+ INTERPOLATION: bilinear
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: pillow
+ - NAME: CenterCrop
+ SIZE: 512
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: pillow
+ - NAME: ImageToTensor
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: pillow
+ - NAME: Normalize
+ MEAN: [ 0.5, 0.5, 0.5 ]
+ STD: [ 0.5, 0.5, 0.5 ]
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'image' ]
+ BACKEND: torchvision
+ - NAME: Select
+ KEYS: [ 'image', 'prompt' ]
+ META_KEYS: [ 'data_key' ]
+ #
+ EVAL_DATA:
+ NAME: ImageTextPairMSDataset
+ MODE: eval
+ MS_DATASET_NAME: style_custom_dataset
+ MS_DATASET_NAMESPACE: damo
+ MS_DATASET_SUBNAME: 3D
+ PROMPT_PREFIX: ""
+ MS_REMAP_KEYS: { 'Image': 'Target:FILE' }
+ MS_DATASET_SPLIT: test_short
+ OUTPUT_SIZE: [512, 512]
+ REPLACE_STYLE: False
+ PIN_MEMORY: True
+ BATCH_SIZE: 4
+ NUM_WORKERS: 4
+ FILE_SYSTEM:
+ NAME: "ModelscopeFs"
+ TEMP_DIR: "./cache/data"
+ #
+ TRANSFORMS:
+ -
+ NAME: Select
+ KEYS: ['prompt']
+ META_KEYS: ['image_size']
+ #
+ TRAIN_HOOKS:
+ -
+ NAME: BackwardHook
+ PRIORITY: 0
+ -
+ NAME: LogHook
+ LOG_INTERVAL: 50
+ -
+ NAME: CheckpointHook
+ INTERVAL: 1000
+ -
+ NAME: ProbeDataHook
+ PROB_INTERVAL: 100
+ #
+ EVAL_HOOKS:
+ -
+ NAME: ProbeDataHook
+ PROB_INTERVAL: 100
diff --git a/scepter/methods/examples/generation/stable_diffusion_2.1_768.yaml b/scepter/methods/examples/generation/stable_diffusion_2.1_768.yaml
index 0d9a1cf..08b3707 100644
--- a/scepter/methods/examples/generation/stable_diffusion_2.1_768.yaml
+++ b/scepter/methods/examples/generation/stable_diffusion_2.1_768.yaml
@@ -113,7 +113,7 @@ SOLVER:
SAMPLE_STEPS: 50
SEED: 2023
GUIDE_SCALE: 7.5
- GUIDE_RESCALE:
+ GUIDE_RESCALE: 0.5
DISCRETIZATION: trailing
IMAGE_SIZE: [768, 768]
RUN_TRAIN_N: False
diff --git a/scepter/methods/examples/generation/stable_diffusion_2.1_768_lora.yaml b/scepter/methods/examples/generation/stable_diffusion_2.1_768_lora.yaml
index 90becb2..e3434c7 100644
--- a/scepter/methods/examples/generation/stable_diffusion_2.1_768_lora.yaml
+++ b/scepter/methods/examples/generation/stable_diffusion_2.1_768_lora.yaml
@@ -122,7 +122,7 @@ SOLVER:
SAMPLE_STEPS: 50
SEED: 2023
GUIDE_SCALE: 7.5
- GUIDE_RESCALE:
+ GUIDE_RESCALE: 0.5
DISCRETIZATION: trailing
IMAGE_SIZE: [768, 768]
RUN_TRAIN_N: False
diff --git a/scepter/methods/scedit/ctr/sd21_768_sce_ctr_canny.yaml b/scepter/methods/scedit/ctr/sd21_768_sce_ctr_canny.yaml
index a40cb3e..af07087 100644
--- a/scepter/methods/scedit/ctr/sd21_768_sce_ctr_canny.yaml
+++ b/scepter/methods/scedit/ctr/sd21_768_sce_ctr_canny.yaml
@@ -31,7 +31,7 @@ SOLVER:
PARAMETERIZATION: v
TIMESTEPS: 1000
MIN_SNR_GAMMA:
- ZERO_TERMINAL_SNR: True
+ ZERO_TERMINAL_SNR: False
PRETRAINED_MODEL: ms://AI-ModelScope/stable-diffusion-2-1@v2-1_768-ema-pruned.safetensors
IGNORE_KEYS: [ ]
SCALE_FACTOR: 0.18215
diff --git a/scepter/methods/scedit/ctr/sd21_768_sce_ctr_pose.yaml b/scepter/methods/scedit/ctr/sd21_768_sce_ctr_pose.yaml
index 932e22f..bce9527 100644
--- a/scepter/methods/scedit/ctr/sd21_768_sce_ctr_pose.yaml
+++ b/scepter/methods/scedit/ctr/sd21_768_sce_ctr_pose.yaml
@@ -31,7 +31,7 @@ SOLVER:
PARAMETERIZATION: v
TIMESTEPS: 1000
MIN_SNR_GAMMA:
- ZERO_TERMINAL_SNR: True
+ ZERO_TERMINAL_SNR: False
PRETRAINED_MODEL: ms://AI-ModelScope/stable-diffusion-2-1@v2-1_768-ema-pruned.safetensors
IGNORE_KEYS: [ ]
SCALE_FACTOR: 0.18215
diff --git a/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_canny.yaml b/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_canny.yaml
new file mode 100644
index 0000000..d853348
--- /dev/null
+++ b/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_canny.yaml
@@ -0,0 +1,378 @@
+ENV:
+ BACKEND: nccl
+SOLVER:
+ NAME: LatentDiffusionSolver
+ RESUME_FROM:
+ LOAD_MODEL_ONLY: True
+ USE_FSDP: False
+ SHARDING_STRATEGY:
+ USE_AMP: True
+ DTYPE: float16
+ CHANNELS_LAST: True
+ MAX_STEPS: 200
+ MAX_EPOCHS: -1
+ NUM_FOLDS: 1
+ ACCU_STEP: 1
+ EVAL_INTERVAL: 100
+ #
+ WORK_DIR: ./cache/save_data/sdxl_1024_sce_ctr_canny
+ LOG_FILE: std_log.txt
+ #
+ FILE_SYSTEM:
+ NAME: "ModelscopeFs"
+ TEMP_DIR: "./cache/data"
+ #
+ FREEZE:
+ FREEZE_PART: [ "first_stage_model", "cond_stage_model", "model" ]
+ TRAIN_PART: [ "control_blocks" ]
+ #
+ MODEL:
+ NAME: LatentDiffusionXLSCEControl
+ PARAMETERIZATION: eps
+ TIMESTEPS: 1000
+ MIN_SNR_GAMMA:
+ ZERO_TERMINAL_SNR: False
+ PRETRAINED_MODEL: ms://AI-ModelScope/stable-diffusion-xl-base-1.0@sd_xl_base_1.0.safetensors
+ IGNORE_KEYS: [ ]
+ SCALE_FACTOR: 0.13025
+ SIZE_FACTOR: 8
+ DEFAULT_N_PROMPT:
+ SCHEDULE_ARGS:
+ "NAME": "scaled_linear"
+ "BETA_MIN": 0.00085
+ "BETA_MAX": 0.0120
+ USE_EMA: False
+ LOAD_REFINER: False
+ #
+ DIFFUSION_MODEL:
+ NAME: DiffusionUNetXL
+ PRETRAINED_MODEL:
+ IN_CHANNELS: 4
+ OUT_CHANNELS: 4
+ NUM_RES_BLOCKS: 2
+ MODEL_CHANNELS: 320
+ ATTENTION_RESOLUTIONS: [ 4, 2 ]
+ DROPOUT: 0
+ CHANNEL_MULT: [ 1, 2, 4 ]
+ CONV_RESAMPLE: True
+ DIMS: 2
+ NUM_CLASSES: sequential
+ USE_CHECKPOINT: False
+ NUM_HEADS: -1
+ NUM_HEADS_CHANNELS: 64
+ USE_SCALE_SHIFT_NORM: False
+ RESBLOCK_UPDOWN: False
+ USE_NEW_ATTENTION_ORDER: True
+ USE_SPATIAL_TRANSFORMER: True
+ TRANSFORMER_DEPTH: [ 1, 2, 10 ]
+ CONTEXT_DIM: 2048
+ DISABLE_MIDDLE_SELF_ATTN: False
+ USE_LINEAR_IN_TRANSFORMER: True
+ ADM_IN_CHANNELS: 2816
+ USE_SENTENCE_EMB: False
+ USE_WORD_MAPPING: False
+ #
+ FIRST_STAGE_MODEL:
+ NAME: AutoencoderKL
+ EMBED_DIM: 4
+ PRETRAINED_MODEL:
+ IGNORE_KEYS: []
+ BATCH_SIZE: 1
+ #
+ ENCODER:
+ NAME: Encoder
+ CH: 128
+ OUT_CH: 3
+ NUM_RES_BLOCKS: 2
+ IN_CHANNELS: 3
+ ATTN_RESOLUTIONS: [ ]
+ CH_MULT: [ 1, 2, 4, 4 ]
+ Z_CHANNELS: 4
+ DOUBLE_Z: True
+ DROPOUT: 0.0
+ RESAMP_WITH_CONV: True
+ #
+ DECODER:
+ NAME: Decoder
+ CH: 128
+ OUT_CH: 3
+ NUM_RES_BLOCKS: 2
+ IN_CHANNELS: 3
+ ATTN_RESOLUTIONS: [ ]
+ CH_MULT: [ 1, 2, 4, 4 ]
+ Z_CHANNELS: 4
+ DROPOUT: 0.0
+ RESAMP_WITH_CONV: True
+ GIVE_PRE_END: False
+ TANH_OUT: False
+ #
+ COND_STAGE_MODEL:
+ NAME: GeneralConditioner
+ PRETRAINED_MODEL:
+ EMBEDDERS:
+ -
+ NAME: FrozenCLIPEmbedder
+ PRETRAINED_MODEL: ms://AI-ModelScope/clip-vit-large-patch14
+ TOKENIZER_PATH: ms://AI-ModelScope/clip-vit-large-patch14
+ MAX_LENGTH: 77
+ FREEZE: True
+ LAYER: hidden
+ LAYER_IDX: 11
+ USE_FINAL_LAYER_NORM: False
+ IS_TRAINABLE: False
+ UCG_RATE: 0.0
+ INPUT_KEYS: [ "prompt" ]
+ LEGACY_UCG_VALUE:
+ -
+ NAME: FrozenOpenCLIPEmbedder2
+ ARCH: ViT-bigG-14
+ PRETRAINED_MODEL:
+ MAX_LENGTH: 77
+ FREEZE: True
+ ALWAYS_RETURN_POOLED: True
+ LEGACY: False
+ LAYER: penultimate
+ IS_TRAINABLE: False
+ UCG_RATE: 0.0
+ INPUT_KEYS: [ "prompt" ]
+ LEGACY_UCG_VALUE:
+ -
+ NAME: ConcatTimestepEmbedderND
+ OUT_DIM: 256
+ IS_TRAINABLE: False
+ UCG_RATE: 0.0
+ INPUT_KEYS: [ "original_size_as_tuple" ]
+ LEGACY_UCG_VALUE:
+ -
+ NAME: ConcatTimestepEmbedderND
+ OUT_DIM: 256
+ IS_TRAINABLE: False
+ UCG_RATE: 0.0
+ INPUT_KEYS: [ "crop_coords_top_left" ]
+ LEGACY_UCG_VALUE:
+ -
+ NAME: ConcatTimestepEmbedderND
+ OUT_DIM: 256
+ IS_TRAINABLE: False
+ UCG_RATE: 0.0
+ INPUT_KEYS: [ "target_size_as_tuple" ]
+ LEGACY_UCG_VALUE:
+ #
+ REFINER_MODEL:
+ NAME: DiffusionUNetXL
+ PRETRAINED_MODEL:
+ IN_CHANNELS: 4
+ OUT_CHANNELS: 4
+ NUM_RES_BLOCKS: 2
+ MODEL_CHANNELS: 384
+ ATTENTION_RESOLUTIONS: [ 4, 2 ]
+ DROPOUT: 0
+ CHANNEL_MULT: [ 1, 2, 4, 4 ]
+ CONV_RESAMPLE: True
+ DIMS: 2
+ NUM_CLASSES: sequential
+ USE_CHECKPOINT: False
+ NUM_HEADS: -1
+ NUM_HEADS_CHANNELS: 64
+ USE_SCALE_SHIFT_NORM: False
+ RESBLOCK_UPDOWN: False
+ USE_NEW_ATTENTION_ORDER: True
+ USE_SPATIAL_TRANSFORMER: True
+ TRANSFORMER_DEPTH: 4
+ CONTEXT_DIM: [ 1280, 1280, 1280, 1280 ]
+ DISABLE_MIDDLE_SELF_ATTN: False
+ USE_LINEAR_IN_TRANSFORMER: True
+ ADM_IN_CHANNELS: 2560
+ USE_SENTENCE_EMB: False
+ USE_WORD_MAPPING: False
+ REFINER_COND_MODEL:
+ NAME: GeneralConditioner
+ PRETRAINED_MODEL:
+ EMBEDDERS:
+ -
+ NAME: FrozenOpenCLIPEmbedder2
+ ARCH: ViT-bigG-14
+ PRETRAINED_MODEL:
+ MAX_LENGTH: 77
+ FREEZE: True
+ ALWAYS_RETURN_POOLED: True
+ LEGACY: False
+ LAYER: penultimate
+ IS_TRAINABLE: False
+ UCG_RATE: 0.0
+ INPUT_KEYS: [ "prompt" ]
+ LEGACY_UCG_VALUE:
+ -
+ NAME: ConcatTimestepEmbedderND
+ OUT_DIM: 256
+ IS_TRAINABLE: False
+ UCG_RATE: 0.0
+ INPUT_KEYS: [ "original_size_as_tuple" ]
+ LEGACY_UCG_VALUE:
+ -
+ NAME: ConcatTimestepEmbedderND
+ OUT_DIM: 256
+ IS_TRAINABLE: False
+ UCG_RATE: 0.0
+ INPUT_KEYS: [ "crop_coords_top_left" ]
+ LEGACY_UCG_VALUE:
+ -
+ NAME: ConcatTimestepEmbedderND
+ OUT_DIM: 256
+ IS_TRAINABLE: False
+ UCG_RATE: 0.0
+ INPUT_KEYS: [ "aesthetic_score" ]
+ LEGACY_UCG_VALUE:
+ #
+ LOSS:
+ NAME: ReconstructLoss
+ LOSS_TYPE: l2
+ #
+ CONTROL_MODEL:
+ NAME: CSCTuners
+ PRE_HINT_IN_CHANNELS: 3
+ PRE_HINT_OUT_CHANNELS: 320
+ DENSE_HINT_KERNAL: 3
+ PRE_HINT_DIM_RATIO: 2.0
+ SCALE: 1.0
+ SC_TUNER_CFG:
+ NAME: SCTuner
+ TUNER_NAME: SCEAdapter
+ DOWN_RATIO: 1.0
+ CONTROL_ANNO:
+ NAME: CannyAnnotator
+ #
+ SAMPLE_ARGS:
+ SAMPLER: ddim
+ SAMPLE_STEPS: 50
+ SEED: 2023
+ GUIDE_SCALE: 7.5
+ GUIDE_RESCALE: 0.5
+ DISCRETIZATION: trailing
+ IMAGE_SIZE: [1024, 1024]
+ RUN_TRAIN_N: False
+ #
+ OPTIMIZER:
+ NAME: AdamW
+ LEARNING_RATE: 0.064
+ BETAS: [ 0.9, 0.999 ]
+ EPS: 1e-8
+ WEIGHT_DECAY: 1e-2
+ AMSGRAD: False
+ #
+ TRAIN_DATA:
+ NAME: ImageTextPairMSDataset
+ MODE: train
+ MS_DATASET_NAME: style_custom_dataset
+ MS_DATASET_NAMESPACE: damo
+ MS_DATASET_SUBNAME: 3D
+ PROMPT_PREFIX: ""
+ MS_DATASET_SPLIT: train_short
+ MS_REMAP_KEYS: { 'Image:FILE': 'Target:FILE' }
+ REPLACE_STYLE: False
+ PIN_MEMORY: True
+ BATCH_SIZE: 1
+ NUM_WORKERS: 4
+ SAMPLER:
+ NAME: LoopSampler
+ TRANSFORMS:
+ - NAME: LoadImageFromFile
+ RGB_ORDER: RGB
+ BACKEND: pillow
+ - NAME: FlexibleResize
+ SIZE: 1024
+ INTERPOLATION: bilinear
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: pillow
+ - NAME: FlexibleCropXL
+ SIZE: 1024
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: pillow
+ - NAME: ToNumpy
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'image_preprocess' ]
+ - NAME: ImageToTensor
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: pillow
+ - NAME: Normalize
+ MEAN: [ 0.5, 0.5, 0.5 ]
+ STD: [ 0.5, 0.5, 0.5 ]
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: torchvision
+ - NAME: Rename
+ INPUT_KEY: [ 'img', 'image_preprocess', 'img_original_size_as_tuple', 'img_target_size_as_tuple', 'img_crop_coords_top_left' ]
+ OUTPUT_KEY: [ 'image', 'image_preprocess', 'original_size_as_tuple', 'target_size_as_tuple', 'crop_coords_top_left' ]
+ - NAME: Select
+ KEYS: [ 'image', 'prompt', 'image_preprocess', 'original_size_as_tuple', 'target_size_as_tuple', 'crop_coords_top_left' ]
+ META_KEYS: [ 'data_key' ]
+ #
+ EVAL_DATA:
+ NAME: ImageTextPairMSDataset
+ MODE: eval
+ MS_DATASET_NAME: style_custom_dataset
+ MS_DATASET_NAMESPACE: damo
+ MS_DATASET_SUBNAME: 3D
+ PROMPT_PREFIX: ""
+ MS_DATASET_SPLIT: train_short
+ MS_REMAP_KEYS: { 'Image:FILE': 'Target:FILE' }
+ REPLACE_STYLE: False
+ PIN_MEMORY: True
+ BATCH_SIZE: 10
+ NUM_WORKERS: 4
+ TRANSFORMS:
+ - NAME: LoadImageFromFile
+ RGB_ORDER: RGB
+ BACKEND: pillow
+ - NAME: Resize
+ SIZE: 1024
+ INTERPOLATION: bilinear
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: pillow
+ - NAME: CenterCrop
+ SIZE: 1024
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: pillow
+ - NAME: ToNumpy
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'image_preprocess' ]
+ - NAME: ImageToTensor
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: pillow
+ - NAME: Normalize
+ MEAN: [ 0.5, 0.5, 0.5 ]
+ STD: [ 0.5, 0.5, 0.5 ]
+ INPUT_KEY: [ 'img' ]
+ OUTPUT_KEY: [ 'img' ]
+ BACKEND: torchvision
+ - NAME: Rename
+ INPUT_KEY: [ 'img', 'image_preprocess' ]
+ OUTPUT_KEY: [ 'image', 'image_preprocess' ]
+ - NAME: Select
+ KEYS: [ 'image', 'prompt', 'image_preprocess' ]
+ META_KEYS: [ 'data_key' ]
+ #
+ TRAIN_HOOKS:
+ -
+ NAME: BackwardHook
+ PRIORITY: 0
+ -
+ NAME: LogHook
+ LOG_INTERVAL: 50
+ -
+ NAME: CheckpointHook
+ INTERVAL: 100
+ -
+ NAME: ProbeDataHook
+ PROB_INTERVAL: 100
+ #
+ EVAL_HOOKS:
+ -
+ NAME: ProbeDataHook
+ PROB_INTERVAL: 100
diff --git a/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_color.yaml b/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_color.yaml
index cdf7fa3..7a95e2d 100644
--- a/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_color.yaml
+++ b/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_color.yaml
@@ -231,8 +231,9 @@ SOLVER:
CONTROL_MODEL:
NAME: CSCTuners
PRE_HINT_IN_CHANNELS: 3
- PRE_HINT_OUT_CHANNELS: 256
+ PRE_HINT_OUT_CHANNELS: 320
DENSE_HINT_KERNAL: 3
+ PRE_HINT_DIM_RATIO: 2.0
SCALE: 1.0
SC_TUNER_CFG:
NAME: SCTuner
diff --git a/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_color_datatxt.yaml b/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_color_datatxt.yaml
index 4a15c00..40dddec 100644
--- a/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_color_datatxt.yaml
+++ b/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_color_datatxt.yaml
@@ -231,8 +231,9 @@ SOLVER:
CONTROL_MODEL:
NAME: CSCTuners
PRE_HINT_IN_CHANNELS: 3
- PRE_HINT_OUT_CHANNELS: 256
+ PRE_HINT_OUT_CHANNELS: 320
DENSE_HINT_KERNAL: 3
+ PRE_HINT_DIM_RATIO: 2.0
SCALE: 1.0
SC_TUNER_CFG:
NAME: SCTuner
diff --git a/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_depth.yaml b/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_depth.yaml
index e4b33ab..af556a6 100644
--- a/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_depth.yaml
+++ b/scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_depth.yaml
@@ -231,8 +231,9 @@ SOLVER:
CONTROL_MODEL:
NAME: CSCTuners
PRE_HINT_IN_CHANNELS: 3
- PRE_HINT_OUT_CHANNELS: 256
+ PRE_HINT_OUT_CHANNELS: 320
DENSE_HINT_KERNAL: 3
+ PRE_HINT_DIM_RATIO: 2.0
SCALE: 1.0
SC_TUNER_CFG:
NAME: SCTuner
diff --git a/scepter/methods/studio/extensions/controllers/official_controllers.yaml b/scepter/methods/studio/extensions/controllers/official_controllers.yaml
index 63a8c8d..836bfed 100644
--- a/scepter/methods/studio/extensions/controllers/official_controllers.yaml
+++ b/scepter/methods/studio/extensions/controllers/official_controllers.yaml
@@ -1,19 +1,63 @@
CONTROLLERS:
+ # SD2.1
- NAME: canny
NAME_ZH:
DESCRIPTION:
BASE_MODEL: SD2.1
TYPE: Canny
- MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD2.1/canny_control/0_SwiftSCETuning
+ MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD2.1/canny_control/
- NAME: openpose
NAME_ZH:
DESCRIPTION:
BASE_MODEL: SD2.1
TYPE: Openpose
- MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD2.1/pose_control/0_SwiftSCETuning
+ MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD2.1/pose_control/
- NAME: color
NAME_ZH:
DESCRIPTION:
BASE_MODEL: SD2.1
TYPE: Color
- MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD2.1/color_control/0_SwiftSCETuning
+ MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD2.1/color_control/
+ - NAME: hed
+ NAME_ZH:
+ DESCRIPTION:
+ BASE_MODEL: SD2.1
+ TYPE: Hed
+ MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD2.1/hed_control
+ - NAME: depth
+ NAME_ZH:
+ DESCRIPTION:
+ BASE_MODEL: SD2.1
+ TYPE: Midas
+ MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD2.1/depth_control
+ # SD_XL1.0
+ - NAME: canny
+ NAME_ZH:
+ DESCRIPTION:
+ BASE_MODEL: SD_XL1.0
+ TYPE: Canny
+ MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD_XL1.0/canny_control
+ - NAME: color
+ NAME_ZH:
+ DESCRIPTION:
+ BASE_MODEL: SD_XL1.0
+ TYPE: Color
+ MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD_XL1.0/color_control
+ - NAME: depth
+ NAME_ZH:
+ DESCRIPTION:
+ BASE_MODEL: SD_XL1.0
+ TYPE: Midas
+ MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD_XL1.0/depth_control
+ - NAME: hed
+ NAME_ZH:
+ DESCRIPTION:
+ BASE_MODEL: SD_XL1.0
+ TYPE: Hed
+ MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD_XL1.0/hed_control
+ - NAME: openpose
+ NAME_ZH:
+ DESCRIPTION:
+ BASE_MODEL: SD_XL1.0
+ TYPE: Openpose
+ MODEL_PATH: ms://damo/scepter_scedit@controllable_model/SD_XL1.0/pose_control
diff --git a/scepter/methods/studio/extensions/tuners/official_tuners.yaml b/scepter/methods/studio/extensions/tuners/official_tuners.yaml
index 99ee032..d195271 100644
--- a/scepter/methods/studio/extensions/tuners/official_tuners.yaml
+++ b/scepter/methods/studio/extensions/tuners/official_tuners.yaml
@@ -1,4 +1,76 @@
TUNERS:
+ - NAME: Azure-Dragon
+ NAME_ZH: 青龙
+ SOURCE: wanx
+ DESCRIPTION: None
+ BASE_MODEL: SD_XL1.0
+ MODEL_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/azure_dragon/
+ IMAGE_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/azure_dragon/xl_azure_dragon.png
+ TUNER_TYPE: SwiftSCE
+ PROMPT_EXAMPLE: Azure Dragon, 8K, high quality,Ultra High Detail.One of the Four Divine Creatures in Charge of Water.
+ - NAME: Gold-Dragon
+ NAME_ZH: 金龙
+ SOURCE: wanx
+ DESCRIPTION: None
+ BASE_MODEL: SD_XL1.0
+ MODEL_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/gold_dragon/
+ IMAGE_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/gold_dragon/xl_gold_dragon.png
+ TUNER_TYPE: SwiftSCE
+ PROMPT_EXAMPLE: Chinese Gold Dragon in the clouds. Translucent Texture. Zbrush. Fuzzy Art. Exquisite Craftsmanship. 3D. 8K. Ultra High Detail
+ - NAME: SpringFestival-Dragon
+ NAME_ZH: 春节龙
+ SOURCE: wanx
+ DESCRIPTION: None
+ BASE_MODEL: SD_XL1.0
+ MODEL_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/spring_festival_dragon/
+ IMAGE_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/spring_festival_dragon/xl_spring_festival_dragon.png
+ TUNER_TYPE: SwiftSCE
+ PROMPT_EXAMPLE: Chinese dragon. Spring Festival.Festive.Street.Lanterns.32K.High quality.expressive, dramatic, dreamlike and mysterious, Surrealism
+ - NAME: Red-Dragon
+ NAME_ZH: 红龙
+ SOURCE: wanx
+ DESCRIPTION: None
+ BASE_MODEL: SD_XL1.0
+ MODEL_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/red_dragon/
+ IMAGE_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/red_dragon/xl_red_dragon.png
+ TUNER_TYPE: SwiftSCE
+ PROMPT_EXAMPLE: Traditional Red Dragon of China. Low Water Level. Studio Ghibli Style. Mural Illustration. White Background. High Detail
+ - NAME: ChinesePunk-Dragon
+ NAME_ZH: 中国朋克龙
+ SOURCE: wanx
+ DESCRIPTION: None
+ BASE_MODEL: SD_XL1.0
+ MODEL_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/chinese_punk_dragon/
+ IMAGE_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/chinese_punk_dragon/xl_chinese_punk_dragon.png
+ TUNER_TYPE: SwiftSCE
+ PROMPT_EXAMPLE: uhd Image,Dragon,Chinese Dragon, Dunhuang Mural Style, Traditional Maritime Art Style
+ - NAME: Cute-Dragon
+ NAME_ZH: 喜庆龙
+ SOURCE: wanx
+ DESCRIPTION: None
+ BASE_MODEL: SD_XL1.0
+ MODEL_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/cute_dragon/
+ IMAGE_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/cute_dragon/xl_kawaii_dragon.png
+ TUNER_TYPE: SwiftSCE
+ PROMPT_EXAMPLE: China Kawaii Dragon. Contest Winner. Minimalist Illustration. White Background. Flat Style. Digital Painting Style. Red. 32k uhd. Fun Comics. Fuzzy Art. Bold. Comic-Inspired Characters
+ - NAME: Dragon-Baby
+ NAME_ZH: 龙宝宝
+ SOURCE: wanx
+ DESCRIPTION: None
+ BASE_MODEL: SD_XL1.0
+ MODEL_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/baby_dragon/
+ IMAGE_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/baby_dragon/xl_baby_dragon.png
+ TUNER_TYPE: SwiftSCE
+ PROMPT_EXAMPLE: Warm Colors, Soft,Chinese Dragon Baby, Felt Style,Dragon Baby, Best Quality, 3D Doll, Macaron Tones, Glittering Big Eyes, Winter,Dragon
+ - NAME: Sloppy-Dragon
+ NAME_ZH: 潦草龙
+ SOURCE: wanx
+ DESCRIPTION: None
+ BASE_MODEL: SD_XL1.0
+ MODEL_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/sloppy_dragon/
+ IMAGE_PATH: ms://damo/scepter_scedit@tuners_model/SD_XL1.0/sloppy_dragon/xl_sloppy_dragon.png
+ TUNER_TYPE: SwiftSCE
+ PROMPT_EXAMPLE: Messy Chinese Dragon,Cute, Wu Guanzhong, Rough
-
NAME: Caricature
NAME_ZH: 夸张漫画
diff --git a/scepter/methods/studio/home/home.yaml b/scepter/methods/studio/home/home.yaml
index e760d92..61ebabc 100644
--- a/scepter/methods/studio/home/home.yaml
+++ b/scepter/methods/studio/home/home.yaml
@@ -38,15 +38,15 @@ GUIDE_INFO:
width: 100%;
}
.video-wrapper {
- width: 75%;
+ width: 75%;
}
video {
- width: 100%;
- display: block;
+ width: 100%;
+ display: block;
}
.description {
- text-align: center;
- margin-top: 10px;
+ text-align: center;
+ margin-top: 10px;
font-size: 0.8em;
}
@@ -70,15 +70,15 @@ GUIDE_INFO:
width: 100%;
}
.video-wrapper {
- width: 75%;
+ width: 75%;
}
video {
- width: 100%;
- display: block;
+ width: 100%;
+ display: block;
}
.description {
- text-align: center;
- margin-top: 10px;
+ text-align: center;
+ margin-top: 10px;
font-size: 0.8em;
}
@@ -92,4 +92,4 @@ GUIDE_INFO:
Train & Inference Video
-