@@ -21,7 +21,7 @@
|
||||
<a href='https://github.com/modelscope/swift'><img src='https://img.shields.io/badge/swift-SCEdit-blue'></a>
|
||||
<br>
|
||||
</p>
|
||||
|
||||
|
||||
SCEdit is an efficient generative fine-tuning framework proposed by Alibaba TongYi Vision Intelligence Lab. This framework enhances the fine-tuning capabilities for text-to-image generation downstream tasks and enables quick adaptation to specific generative scenarios, **saving 30%-50% of training memory costs compared to LoRA**. Furthermore, it can be directly extended to controllable image generation tasks, **requiring only 7.9% of the parameters that ControlNet needs for conditional generation and saving 30% of memory usage**. It supports various conditional generation tasks including edge maps, depth maps, segmentation maps, poses, color maps, and image completion.
|
||||
|
||||
## Usage
|
||||
@@ -48,7 +48,7 @@ python scepter/tools/run_train.py --cfg scepter/methods/scedit/ctr/sdxl_1024_sce
|
||||
|
||||
### Gradio
|
||||
```shell
|
||||
python -m scepter.tools.webui # Then click [Use Tuners] or [Use Controller]
|
||||
python -m scepter.tools.webui # Then click [Use Tuners] or [Use Controller]
|
||||
```
|
||||
|
||||
## Models
|
||||
@@ -75,7 +75,7 @@ python -m scepter.tools.webui # Then click [Use Tuners] or [Use Controller]
|
||||
| SD XL | 🪄 | 🪄 | 🪄 | 🪄 | 🪄 |
|
||||
|
||||
|
||||
## Application Gallery
|
||||
## Application Gallery
|
||||
|
||||
### Dragon Year Special: Dragon Tuner
|
||||
|
||||
@@ -120,8 +120,6 @@ python -m scepter.tools.webui # Then click [Use Tuners] or [Use Controller]
|
||||
title = {SCEdit: Efficient and Controllable Image Diffusion Generation via Skip Connection Editing},
|
||||
author = {Jiang, Zeyinzi and Mao, Chaojie and Pan, Yulin and Han, Zhen and Zhang, Jingfeng},
|
||||
year = {2023},
|
||||
journal = {arXiv preprint arXiv:2312.11392}
|
||||
journal = {arXiv preprint arXiv:2312.11392}
|
||||
}
|
||||
```
|
||||
|
||||
|
||||
|
||||
@@ -5,7 +5,7 @@ Zhen Han, Chaojie Mao, Zeyinzi Jiang, Yulin Pan, Jingfeng Zhang
|
||||
|
||||
Alibaba Group
|
||||
|
||||
[[paper](https://arxiv.org/abs/2404.12154)][[Model](https://modelscope.cn/models/iic/stylebooth/summary)] [[Dataset](https://modelscope.cn/models/iic/stylebooth/summary)]
|
||||
[[paper](https://arxiv.org/abs/2404.12154)][[Model](https://modelscope.cn/models/iic/stylebooth/summary)] [[Dataset](https://modelscope.cn/models/iic/stylebooth/summary)]
|
||||
|
||||
## Abstract
|
||||
|
||||
@@ -45,7 +45,7 @@ Given an original image, image editing aims to generate an image that align with
|
||||
## Run StyleBooth
|
||||
- Code implementation: See model configuration and code based on [🪄SCEPTER](https://github.com/modelscope/scepter).
|
||||
|
||||
- Demo: Try [🖥️SCEPTER Studio](https://github.com/modelscope/scepter/tree/main?tab=readme-ov-file#%EF%B8%8F-scepter-studio).
|
||||
- Demo: Try [🖥️SCEPTER Studio](https://github.com/modelscope/scepter/tree/main?tab=readme-ov-file#%EF%B8%8F-scepter-studio).
|
||||
|
||||
- Easy run:
|
||||
Try the following example script to run StyleBooth modified from [tests/modules/test_diffusion_inference.py](https://github.com/modelscope/scepter/blob/main/tests/modules/test_diffusion_inference.py):
|
||||
@@ -100,4 +100,4 @@ class DiffusionInferenceTest(unittest.TestCase):
|
||||
|
||||
if __name__ == '__main__':
|
||||
unittest.main()
|
||||
```
|
||||
```
|
||||
|
||||
@@ -43,4 +43,3 @@ python scepter/tools/run_inference.py --cfg scepter/methods/scedit/t2i/sd15_512_
|
||||
python scepter/tools/run_inference.py --cfg scepter/methods/scedit/ctr/sd21_768_sce_ctr_canny.yaml --num_samples 1 --prompt 'a single flower is shown in front of a tree' --save_folder 'test_flower_canny' --image_size 768 --task control --image 'asset/images/flower.jpg' --control_mode canny --pretrained_model ms://damo/scepter_scedit@controllable_model/SD2.1/canny_control/0_SwiftSCETuning/pytorch_model.bin # canny
|
||||
python scepter/tools/run_inference.py --cfg scepter/methods/scedit/ctr/sd21_768_sce_ctr_pose.yaml --num_samples 1 --prompt 'super mario' --save_folder 'test_mario_pose' --image_size 768 --task control --image 'asset/images/pose_source.png' --control_mode source --pretrained_model ms://damo/scepter_scedit@controllable_model/SD2.1/pose_control/0_SwiftSCETuning/pytorch_model.bin # pose
|
||||
```
|
||||
|
||||
|
||||
@@ -12,7 +12,7 @@
|
||||
SCEPTER integrates popular community-driven implementations as well as proprietary methods by Tongyi Lab of Alibaba Group, offering a comprehensive toolkit for researchers and practitioners in the field of AIGC. This versatile library is designed to facilitate innovation and accelerate development in the rapidly evolving domain of generative models.
|
||||
|
||||
SCEPTER offers 3 core components:
|
||||
- [Generative training and inference framework](#tutorials)
|
||||
- [Generative training and inference framework](#tutorials)
|
||||
- [Easy implementation of popular approaches](#currently-supported-approaches)
|
||||
- [Interactive user interface: SCEPTER Studio](#launch)
|
||||
|
||||
@@ -105,8 +105,8 @@ pip install scepter
|
||||
| Text-to-image generation | SD v1.5 | [](https://huggingface.co/runwayml/stable-diffusion-v1-5) |
|
||||
| Text-to-image generation | SD v2.1 | [](https://huggingface.co/runwayml/stable-diffusion-v1-5) |
|
||||
| Text-to-image generation | SD-XL | [](https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0) |
|
||||
| Efficent Tuning | LoRA | [](https://arxiv.org/abs/2106.09685) |
|
||||
| Efficent Tuning | Res-Tuning(NeurIPS23) | [](https://arxiv.org/abs/2310.19859) [](https://res-tuning.github.io/) |
|
||||
| Efficient Tuning | LoRA | [](https://arxiv.org/abs/2106.09685) |
|
||||
| Efficient Tuning | Res-Tuning(NeurIPS23) | [](https://arxiv.org/abs/2310.19859) [](https://res-tuning.github.io/) |
|
||||
| Controllable image synthesis | [🌟SCEdit(CVPR24)](docs/en/tasks/scedit.md) | [](https://arxiv.org/abs/2312.11392) [](https://scedit.github.io/) |
|
||||
| Image editing | [🌟LAR-Gen](docs/en/tasks/largen.md) | [](https://arxiv.org/abs/2403.19534) [](https://ali-vilab.github.io/largen-page/) |
|
||||
| Image editing | [🌟StyleBooth](docs/en/tasks/stylebooth.md) | [](https://arxiv.org/abs/2404.12154) [](https://ali-vilab.github.io/stylebooth-page/) |
|
||||
@@ -175,4 +175,4 @@ This project is licensed under the [Apache License (Version 2.0)](https://github
|
||||
|
||||
|
||||
## Acknowledgement
|
||||
Thanks to [Stability-AI](https://github.com/Stability-AI), [SWIFT library](https://github.com/modelscope/swift/) and [Fooocus](https://github.com/lllyasviel/Fooocus) for their awesome work.
|
||||
Thanks to [Stability-AI](https://github.com/Stability-AI), [SWIFT library](https://github.com/modelscope/swift/) and [Fooocus](https://github.com/lllyasviel/Fooocus) for their awesome work.
|
||||
|
||||
@@ -621,6 +621,7 @@ SOLVER:
|
||||
PRIORITY: 200
|
||||
SAVE_LAST: True
|
||||
SAVE_NAME_PREFIX: 'step'
|
||||
DISABLE_SNAPSHOT: True
|
||||
#
|
||||
EVAL_HOOKS:
|
||||
-
|
||||
|
||||
@@ -332,6 +332,7 @@ SOLVER:
|
||||
PRIORITY: 200
|
||||
SAVE_LAST: True
|
||||
SAVE_NAME_PREFIX: 'step'
|
||||
DISABLE_SNAPSHOT: True
|
||||
#
|
||||
EVAL_HOOKS:
|
||||
-
|
||||
|
||||
@@ -274,6 +274,7 @@ SOLVER:
|
||||
PRIORITY: 200
|
||||
SAVE_LAST: True
|
||||
SAVE_NAME_PREFIX: 'step'
|
||||
DISABLE_SNAPSHOT: True
|
||||
#
|
||||
EVAL_HOOKS:
|
||||
-
|
||||
|
||||
@@ -5,10 +5,10 @@ import os.path
|
||||
import random
|
||||
from collections import OrderedDict
|
||||
|
||||
from PIL.Image import Image
|
||||
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
from PIL.Image import Image
|
||||
|
||||
from scepter.modules.model.network.diffusion.diffusion import GaussianDiffusion
|
||||
from scepter.modules.model.network.diffusion.schedules import noise_schedule
|
||||
from scepter.modules.model.registry import (BACKBONES, EMBEDDERS, MODELS,
|
||||
@@ -274,7 +274,7 @@ class DiffusionInference():
|
||||
module = self.load(module)
|
||||
self.loaded_model[name] = module
|
||||
return module
|
||||
elif module['device'] == 'cpu':
|
||||
elif module['device'] == 'cpu' or module['device'] == "offline":
|
||||
module = self.load(module)
|
||||
return module
|
||||
else:
|
||||
|
||||
@@ -276,7 +276,7 @@ class LargenInference():
|
||||
module = self.load(module)
|
||||
self.loaded_model[name] = module
|
||||
return module
|
||||
elif module['device'] == 'cpu':
|
||||
elif module['device'] == 'cpu' or module['device'] == 'offline':
|
||||
module = self.load(module)
|
||||
return module
|
||||
else:
|
||||
@@ -498,7 +498,7 @@ class LargenInference():
|
||||
|
||||
self.dynamic_unload(self.first_stage_model,
|
||||
'first_stage_model',
|
||||
skip_loaded=True)
|
||||
skip_loaded=False)
|
||||
|
||||
# cond stage
|
||||
self.dynamic_load(self.cond_stage_model, 'cond_stage_model')
|
||||
@@ -513,7 +513,7 @@ class LargenInference():
|
||||
|
||||
self.dynamic_unload(self.cond_stage_model,
|
||||
'cond_stage_model',
|
||||
skip_loaded=True)
|
||||
skip_loaded=False)
|
||||
|
||||
# get noise
|
||||
seed = kwargs.pop('seed', -1)
|
||||
@@ -582,7 +582,7 @@ class LargenInference():
|
||||
x_samples = self.decode_first_stage(latent).float()
|
||||
self.dynamic_unload(self.first_stage_model,
|
||||
'first_stage_model',
|
||||
skip_loaded=True)
|
||||
skip_loaded=False)
|
||||
images = torch.clamp((x_samples + 1.0) / 2.0, min=0.0, max=1.0)
|
||||
if base_image is not None:
|
||||
stitch_images = []
|
||||
|
||||
@@ -5,11 +5,11 @@ import os.path
|
||||
import random
|
||||
from collections import OrderedDict
|
||||
|
||||
from PIL.Image import Image
|
||||
|
||||
import torch
|
||||
import torch.nn.functional as F
|
||||
import torchvision.transforms.functional as TF
|
||||
from PIL.Image import Image
|
||||
|
||||
from scepter.modules.model.network.diffusion.diffusion import GaussianDiffusion
|
||||
from scepter.modules.model.network.diffusion.schedules import noise_schedule
|
||||
from scepter.modules.model.registry import (BACKBONES, EMBEDDERS, MODELS,
|
||||
@@ -275,7 +275,7 @@ class StyleboothInference():
|
||||
module = self.load(module)
|
||||
self.loaded_model[name] = module
|
||||
return module
|
||||
elif module['device'] == 'cpu':
|
||||
elif module['device'] == 'cpu' or module['device'] == "offline":
|
||||
module = self.load(module)
|
||||
return module
|
||||
else:
|
||||
|
||||
@@ -49,6 +49,12 @@ class CheckpointHook(Hook):
|
||||
'',
|
||||
'description':
|
||||
'If save the best model, which order should be sorted, +/-!'
|
||||
},
|
||||
'DISABLE_SNAPSHOT': {
|
||||
'value':
|
||||
False,
|
||||
'description':
|
||||
'Skip to save snapshot checkpoint.'
|
||||
}
|
||||
}]
|
||||
|
||||
@@ -62,6 +68,7 @@ class CheckpointHook(Hook):
|
||||
self.save_best_by = cfg.get('SAVE_BEST_BY', '')
|
||||
self.push_to_hub = cfg.get('PUSH_TO_HUB', False)
|
||||
self.hub_model_id = cfg.get('HUB_MODEL_ID', None)
|
||||
self.disable_save_snapshot = cfg.get('DISABLE_SNAPSHOT', False)
|
||||
self.last_ckpt = None
|
||||
if self.save_best and not self.save_best_by:
|
||||
warnings.warn(
|
||||
@@ -110,10 +117,11 @@ class CheckpointHook(Hook):
|
||||
solver.work_dir,
|
||||
'checkpoints/{}-{}.pth'.format(self.save_name_prefix,
|
||||
solver.total_iter + 1))
|
||||
with FS.put_to(save_path) as local_path:
|
||||
with open(local_path, 'wb') as f:
|
||||
checkpoint = solver.save_checkpoint()
|
||||
torch.save(checkpoint, f)
|
||||
if not self.disable_save_snapshot:
|
||||
with FS.put_to(save_path) as local_path:
|
||||
with open(local_path, 'wb') as f:
|
||||
checkpoint = solver.save_checkpoint()
|
||||
torch.save(checkpoint, f)
|
||||
|
||||
from swift import SwiftModel
|
||||
if isinstance(solver.model, SwiftModel):
|
||||
|
||||
@@ -8,8 +8,7 @@ from scepter.modules.utils.file_system import FS
|
||||
def download_image(image):
|
||||
if image is not None:
|
||||
name = get_md5(image)
|
||||
local_path = FS.get_from(image,
|
||||
f'/tmp/gradio/scepter_examples/{name}')
|
||||
local_path = FS.get_from(image, f'/tmp/gradio/scepter_examples/{name}')
|
||||
return local_path
|
||||
else:
|
||||
return image
|
||||
|
||||
@@ -161,10 +161,12 @@ class DiffusionInferenceTest(unittest.TestCase):
|
||||
diff_infer = StyleboothInference(logger=self.logger)
|
||||
diff_infer.init_from_cfg(cfg)
|
||||
|
||||
output = diff_infer({'prompt': 'Let this image be in the style of sai-lowpoly'},
|
||||
style_edit_image=Image.open('asset/images/inpainting_text_ref/ex4_scene_im.jpg'),
|
||||
style_guide_scale_text=7.5,
|
||||
style_guide_scale_image=0.5)
|
||||
output = diff_infer(
|
||||
{'prompt': 'Let this image be in the style of sai-lowpoly'},
|
||||
style_edit_image=Image.open(
|
||||
'asset/images/inpainting_text_ref/ex4_scene_im.jpg'),
|
||||
style_guide_scale_text=7.5,
|
||||
style_guide_scale_image=0.5)
|
||||
save_path = os.path.join(self.tmp_dir,
|
||||
'stylebooth_test_lowpoly_cute_dog.png')
|
||||
save_image(output['images'], save_path)
|
||||
|
||||
Reference in New Issue
Block a user