SCEdit: Efficient and Controllable Image Diffusion Generation via Skip Connection Editing

(CVPR 2024 Highlight)

Zeyinzi Jiang ยท Chaojie Mao ยท Yulin Pan ยท Zhen Han ยท Jingfeng Zhang
Alibaba Group
Paper PDF Project Page

SCEdit is an efficient generative fine-tuning framework proposed by Alibaba TongYi Vision Intelligence Lab. This framework enhances the fine-tuning capabilities for text-to-image generation downstream tasks and enables quick adaptation to specific generative scenarios, **saving 30%-50% of training memory costs compared to LoRA**. Furthermore, it can be directly extended to controllable image generation tasks, **requiring only 7.9% of the parameters that ControlNet needs for conditional generation and saving 30% of memory usage**. It supports various conditional generation tasks including edge maps, depth maps, segmentation maps, poses, color maps, and image completion. ## Usage ### Text-to-Image Generation ```shell # SD v1.5 python scepter/tools/run_train.py --cfg scepter/methods/scedit/t2i/sd15_512_sce_t2i.yaml # SD v2.1 python scepter/tools/run_train.py --cfg scepter/methods/scedit/t2i/sd21_768_sce_t2i.yaml # SD XL python scepter/tools/run_train.py --cfg scepter/methods/scedit/t2i/sdxl_1024_sce_t2i.yaml ``` ### Controllable Image Synthesis ```shell # SD v1.5 + hed python scepter/tools/run_train.py --cfg scepter/methods/scedit/ctr/sd15_512_sce_ctr_hed.yaml # SD v2.1 + canny python scepter/tools/run_train.py --cfg scepter/methods/scedit/ctr/sd21_768_sce_ctr_canny.yaml # SD XL + depth python scepter/tools/run_train.py --cfg scepter/methods/scedit/ctr/sdxl_1024_sce_ctr_depth.yaml ``` ### Gradio ```shell python -m scepter.tools.webui # Then click [Use Tuners] or [Use Controller] ``` ## Models ### Model URL | Model | URL | |--------|-------------------------------------------------------------------------------------------------------------------------------------------| | SCEdit | [ModelScope](https://modelscope.cn/models/iic/scepter_scedit/summary) [HuggingFace](https://huggingface.co/scepter-studio/scepter_scedit) | ### Text-to-Image Generation | **Model** | **SCEdit** | |:---------:|:----------:| | SD 1.5 | ๐Ÿช„ | | SD 2.1 | ๐Ÿช„ | | SD XL | ๐Ÿช„ | ### Controllable Image Synthesis | **Model** | **Canny** | **HED** | **Depth** | **Pose** | **Color** | |:---------:|:---------:|:-------:|:---------:|:--------:|:---------:| | SD 2.1 | ๐Ÿช„ | ๐Ÿช„ | ๐Ÿช„ | ๐Ÿช„ | ๐Ÿช„ | | SD XL | ๐Ÿช„ | ๐Ÿช„ | ๐Ÿช„ | ๐Ÿช„ | ๐Ÿช„ | ## Application Gallery ### Dragon Year Special: Dragon Tuner
Gold Dragon Tuner Sloppy Dragon Tuner Red Dragon Tuner
+ Papercraft Mantra
Azure Dragon Tuner
+ Pose Control
### Text Effect Image
Conditional Image Midas Control
"Race track, top view"
Midas Control
+ Watercolor Mantra
"white lilies"
Midas Control
+ Dragon Tuner
"Spring Festival, Chinese dragon"
## BibTeX ```bibtex @article{jiang2023scedit, title = {SCEdit: Efficient and Controllable Image Diffusion Generation via Skip Connection Editing}, author = {Jiang, Zeyinzi and Mao, Chaojie and Pan, Yulin and Han, Zhen and Zhang, Jingfeng}, year = {2023}, journal = {arXiv preprint arXiv:2312.11392} } ```