Merge pull request #24 from modelscope/v0.0.5_dev

update readme&OOM
This commit is contained in:
mcj
2024-04-22 20:24:57 +08:00
committed by GitHub
13 changed files with 43 additions and 34 deletions
+4 -6
View File
@@ -21,7 +21,7 @@
<a href='https://github.com/modelscope/swift'><img src='https://img.shields.io/badge/swift-SCEdit-blue'></a>
<br>
</p>
SCEdit is an efficient generative fine-tuning framework proposed by Alibaba TongYi Vision Intelligence Lab. This framework enhances the fine-tuning capabilities for text-to-image generation downstream tasks and enables quick adaptation to specific generative scenarios, **saving 30%-50% of training memory costs compared to LoRA**. Furthermore, it can be directly extended to controllable image generation tasks, **requiring only 7.9% of the parameters that ControlNet needs for conditional generation and saving 30% of memory usage**. It supports various conditional generation tasks including edge maps, depth maps, segmentation maps, poses, color maps, and image completion.
## Usage
@@ -48,7 +48,7 @@ python scepter/tools/run_train.py --cfg scepter/methods/scedit/ctr/sdxl_1024_sce
### Gradio
```shell
python -m scepter.tools.webui # Then click [Use Tuners] or [Use Controller]
python -m scepter.tools.webui # Then click [Use Tuners] or [Use Controller]
```
## Models
@@ -75,7 +75,7 @@ python -m scepter.tools.webui # Then click [Use Tuners] or [Use Controller]
| SD XL | 🪄 | 🪄 | 🪄 | 🪄 | 🪄 |
## Application Gallery
## Application Gallery
### Dragon Year Special: Dragon Tuner
@@ -120,8 +120,6 @@ python -m scepter.tools.webui # Then click [Use Tuners] or [Use Controller]
title = {SCEdit: Efficient and Controllable Image Diffusion Generation via Skip Connection Editing},
author = {Jiang, Zeyinzi and Mao, Chaojie and Pan, Yulin and Han, Zhen and Zhang, Jingfeng},
year = {2023},
journal = {arXiv preprint arXiv:2312.11392}
journal = {arXiv preprint arXiv:2312.11392}
}
```
+3 -3
View File
@@ -5,7 +5,7 @@ Zhen Han, Chaojie Mao, Zeyinzi Jiang, Yulin Pan, Jingfeng Zhang
Alibaba Group
[[paper](https://arxiv.org/abs/2404.12154)][[Model](https://modelscope.cn/models/iic/stylebooth/summary)] [[Dataset](https://modelscope.cn/models/iic/stylebooth/summary)]
[[paper](https://arxiv.org/abs/2404.12154)][[Model](https://modelscope.cn/models/iic/stylebooth/summary)] [[Dataset](https://modelscope.cn/models/iic/stylebooth/summary)]
## Abstract
@@ -45,7 +45,7 @@ Given an original image, image editing aims to generate an image that align with
## Run StyleBooth
- Code implementation: See model configuration and code based on [🪄SCEPTER](https://github.com/modelscope/scepter).
- Demo: Try [🖥️SCEPTER Studio](https://github.com/modelscope/scepter/tree/main?tab=readme-ov-file#%EF%B8%8F-scepter-studio).
- Demo: Try [🖥️SCEPTER Studio](https://github.com/modelscope/scepter/tree/main?tab=readme-ov-file#%EF%B8%8F-scepter-studio).
- Easy run:
Try the following example script to run StyleBooth modified from [tests/modules/test_diffusion_inference.py](https://github.com/modelscope/scepter/blob/main/tests/modules/test_diffusion_inference.py):
@@ -100,4 +100,4 @@ class DiffusionInferenceTest(unittest.TestCase):
if __name__ == '__main__':
unittest.main()
```
```
-1
View File
@@ -43,4 +43,3 @@ python scepter/tools/run_inference.py --cfg scepter/methods/scedit/t2i/sd15_512_
python scepter/tools/run_inference.py --cfg scepter/methods/scedit/ctr/sd21_768_sce_ctr_canny.yaml --num_samples 1 --prompt 'a single flower is shown in front of a tree' --save_folder 'test_flower_canny' --image_size 768 --task control --image 'asset/images/flower.jpg' --control_mode canny --pretrained_model ms://damo/scepter_scedit@controllable_model/SD2.1/canny_control/0_SwiftSCETuning/pytorch_model.bin # canny
python scepter/tools/run_inference.py --cfg scepter/methods/scedit/ctr/sd21_768_sce_ctr_pose.yaml --num_samples 1 --prompt 'super mario' --save_folder 'test_mario_pose' --image_size 768 --task control --image 'asset/images/pose_source.png' --control_mode source --pretrained_model ms://damo/scepter_scedit@controllable_model/SD2.1/pose_control/0_SwiftSCETuning/pytorch_model.bin # pose
```
+4 -4
View File
@@ -12,7 +12,7 @@
SCEPTER integrates popular community-driven implementations as well as proprietary methods by Tongyi Lab of Alibaba Group, offering a comprehensive toolkit for researchers and practitioners in the field of AIGC. This versatile library is designed to facilitate innovation and accelerate development in the rapidly evolving domain of generative models.
SCEPTER offers 3 core components:
- [Generative training and inference framework](#tutorials)
- [Generative training and inference framework](#tutorials)
- [Easy implementation of popular approaches](#currently-supported-approaches)
- [Interactive user interface: SCEPTER Studio](#launch)
@@ -105,8 +105,8 @@ pip install scepter
| Text-to-image generation | SD v1.5 | [![Hugging Face Repo](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Repo-blue)](https://huggingface.co/runwayml/stable-diffusion-v1-5) |
| Text-to-image generation | SD v2.1 | [![Hugging Face Repo](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Repo-blue)](https://huggingface.co/runwayml/stable-diffusion-v1-5) |
| Text-to-image generation | SD-XL | [![Hugging Face Repo](https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Repo-blue)](https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0) |
| Efficent Tuning | LoRA | [![Arxiv link](https://img.shields.io/static/v1?label=arXiv&message=LoRA&color=red&logo=arxiv)](https://arxiv.org/abs/2106.09685) |
| Efficent Tuning | Res-Tuning(NeurIPS23) | [![Arxiv link](https://img.shields.io/static/v1?label=arXiv&message=Res-Tuing&color=red&logo=arxiv)](https://arxiv.org/abs/2310.19859) [![Page link](https://img.shields.io/badge/Page-ResTuning-Gree)](https://res-tuning.github.io/) |
| Efficient Tuning | LoRA | [![Arxiv link](https://img.shields.io/static/v1?label=arXiv&message=LoRA&color=red&logo=arxiv)](https://arxiv.org/abs/2106.09685) |
| Efficient Tuning | Res-Tuning(NeurIPS23) | [![Arxiv link](https://img.shields.io/static/v1?label=arXiv&message=Res-Tuing&color=red&logo=arxiv)](https://arxiv.org/abs/2310.19859) [![Page link](https://img.shields.io/badge/Page-ResTuning-Gree)](https://res-tuning.github.io/) |
| Controllable image synthesis | [🌟SCEdit(CVPR24)](docs/en/tasks/scedit.md) | [![Arxiv link](https://img.shields.io/static/v1?label=arXiv&message=SCEdit&color=red&logo=arxiv)](https://arxiv.org/abs/2312.11392) [![Page link](https://img.shields.io/badge/Page-SCEdit-Gree)](https://scedit.github.io/) |
| Image editing | [🌟LAR-Gen](docs/en/tasks/largen.md) | [![Arxiv link](https://img.shields.io/static/v1?label=arXiv&message=LARGen&color=red&logo=arxiv)](https://arxiv.org/abs/2403.19534) [![Page link](https://img.shields.io/badge/Page-LARGen-Gree)](https://ali-vilab.github.io/largen-page/) |
| Image editing | [🌟StyleBooth](docs/en/tasks/stylebooth.md) | [![Arxiv link](https://img.shields.io/static/v1?label=arXiv&message=StyleBooth&color=red&logo=arxiv)](https://arxiv.org/abs/2404.12154) [![Page link](https://img.shields.io/badge/Page-StyleBooth-Gree)](https://ali-vilab.github.io/stylebooth-page/) |
@@ -175,4 +175,4 @@ This project is licensed under the [Apache License (Version 2.0)](https://github
## Acknowledgement
Thanks to [Stability-AI](https://github.com/Stability-AI), [SWIFT library](https://github.com/modelscope/swift/) and [Fooocus](https://github.com/lllyasviel/Fooocus) for their awesome work.
Thanks to [Stability-AI](https://github.com/Stability-AI), [SWIFT library](https://github.com/modelscope/swift/) and [Fooocus](https://github.com/lllyasviel/Fooocus) for their awesome work.
@@ -621,6 +621,7 @@ SOLVER:
PRIORITY: 200
SAVE_LAST: True
SAVE_NAME_PREFIX: 'step'
DISABLE_SNAPSHOT: True
#
EVAL_HOOKS:
-
@@ -332,6 +332,7 @@ SOLVER:
PRIORITY: 200
SAVE_LAST: True
SAVE_NAME_PREFIX: 'step'
DISABLE_SNAPSHOT: True
#
EVAL_HOOKS:
-
@@ -274,6 +274,7 @@ SOLVER:
PRIORITY: 200
SAVE_LAST: True
SAVE_NAME_PREFIX: 'step'
DISABLE_SNAPSHOT: True
#
EVAL_HOOKS:
-
@@ -5,10 +5,10 @@ import os.path
import random
from collections import OrderedDict
from PIL.Image import Image
import torch
import torch.nn.functional as F
from PIL.Image import Image
from scepter.modules.model.network.diffusion.diffusion import GaussianDiffusion
from scepter.modules.model.network.diffusion.schedules import noise_schedule
from scepter.modules.model.registry import (BACKBONES, EMBEDDERS, MODELS,
@@ -274,7 +274,7 @@ class DiffusionInference():
module = self.load(module)
self.loaded_model[name] = module
return module
elif module['device'] == 'cpu':
elif module['device'] == 'cpu' or module['device'] == "offline":
module = self.load(module)
return module
else:
@@ -276,7 +276,7 @@ class LargenInference():
module = self.load(module)
self.loaded_model[name] = module
return module
elif module['device'] == 'cpu':
elif module['device'] == 'cpu' or module['device'] == 'offline':
module = self.load(module)
return module
else:
@@ -498,7 +498,7 @@ class LargenInference():
self.dynamic_unload(self.first_stage_model,
'first_stage_model',
skip_loaded=True)
skip_loaded=False)
# cond stage
self.dynamic_load(self.cond_stage_model, 'cond_stage_model')
@@ -513,7 +513,7 @@ class LargenInference():
self.dynamic_unload(self.cond_stage_model,
'cond_stage_model',
skip_loaded=True)
skip_loaded=False)
# get noise
seed = kwargs.pop('seed', -1)
@@ -582,7 +582,7 @@ class LargenInference():
x_samples = self.decode_first_stage(latent).float()
self.dynamic_unload(self.first_stage_model,
'first_stage_model',
skip_loaded=True)
skip_loaded=False)
images = torch.clamp((x_samples + 1.0) / 2.0, min=0.0, max=1.0)
if base_image is not None:
stitch_images = []
@@ -5,11 +5,11 @@ import os.path
import random
from collections import OrderedDict
from PIL.Image import Image
import torch
import torch.nn.functional as F
import torchvision.transforms.functional as TF
from PIL.Image import Image
from scepter.modules.model.network.diffusion.diffusion import GaussianDiffusion
from scepter.modules.model.network.diffusion.schedules import noise_schedule
from scepter.modules.model.registry import (BACKBONES, EMBEDDERS, MODELS,
@@ -275,7 +275,7 @@ class StyleboothInference():
module = self.load(module)
self.loaded_model[name] = module
return module
elif module['device'] == 'cpu':
elif module['device'] == 'cpu' or module['device'] == "offline":
module = self.load(module)
return module
else:
+12 -4
View File
@@ -49,6 +49,12 @@ class CheckpointHook(Hook):
'',
'description':
'If save the best model, which order should be sorted, +/-!'
},
'DISABLE_SNAPSHOT': {
'value':
False,
'description':
'Skip to save snapshot checkpoint.'
}
}]
@@ -62,6 +68,7 @@ class CheckpointHook(Hook):
self.save_best_by = cfg.get('SAVE_BEST_BY', '')
self.push_to_hub = cfg.get('PUSH_TO_HUB', False)
self.hub_model_id = cfg.get('HUB_MODEL_ID', None)
self.disable_save_snapshot = cfg.get('DISABLE_SNAPSHOT', False)
self.last_ckpt = None
if self.save_best and not self.save_best_by:
warnings.warn(
@@ -110,10 +117,11 @@ class CheckpointHook(Hook):
solver.work_dir,
'checkpoints/{}-{}.pth'.format(self.save_name_prefix,
solver.total_iter + 1))
with FS.put_to(save_path) as local_path:
with open(local_path, 'wb') as f:
checkpoint = solver.save_checkpoint()
torch.save(checkpoint, f)
if not self.disable_save_snapshot:
with FS.put_to(save_path) as local_path:
with open(local_path, 'wb') as f:
checkpoint = solver.save_checkpoint()
torch.save(checkpoint, f)
from swift import SwiftModel
if isinstance(solver.model, SwiftModel):
@@ -8,8 +8,7 @@ from scepter.modules.utils.file_system import FS
def download_image(image):
if image is not None:
name = get_md5(image)
local_path = FS.get_from(image,
f'/tmp/gradio/scepter_examples/{name}')
local_path = FS.get_from(image, f'/tmp/gradio/scepter_examples/{name}')
return local_path
else:
return image
+6 -4
View File
@@ -161,10 +161,12 @@ class DiffusionInferenceTest(unittest.TestCase):
diff_infer = StyleboothInference(logger=self.logger)
diff_infer.init_from_cfg(cfg)
output = diff_infer({'prompt': 'Let this image be in the style of sai-lowpoly'},
style_edit_image=Image.open('asset/images/inpainting_text_ref/ex4_scene_im.jpg'),
style_guide_scale_text=7.5,
style_guide_scale_image=0.5)
output = diff_infer(
{'prompt': 'Let this image be in the style of sai-lowpoly'},
style_edit_image=Image.open(
'asset/images/inpainting_text_ref/ex4_scene_im.jpg'),
style_guide_scale_text=7.5,
style_guide_scale_image=0.5)
save_path = os.path.join(self.tmp_dir,
'stylebooth_test_lowpoly_cute_dog.png')
save_image(output['images'], save_path)