add workflow examples

This commit is contained in:
maochaojie
2025-03-22 09:47:22 +08:00
parent db78fe59f4
commit 21886da779
13 changed files with 6056 additions and 181 deletions
+1 -1
View File
@@ -1 +1 @@
.idea
.idea
+196 -178
View File
@@ -40,11 +40,9 @@
The original intention behind the design of ACE++ was to unify reference image generation, local editing,
and controllable generation into a single framework, and to enable one model to adapt to a wider range of tasks.
A more versatile model is often capable of handling more complex tasks. We have released three LoRA models for
specific vertical domains and a more versatile FFT model (the performance of the FFT model may decline compared
specific vertical domains and a more versatile FFT model (the performance of the FFT model declines compared
to the LoRA model across various tasks). Users can flexibly utilize these models and their
combinations for their own scenarios. Furthermore, many community members have found that using them
in conjunction with Redux modules significantly improves performance. We believe there are many more
use cases to explore.
combinations for their own scenarios.
## 📢 News
- [x] **[2025.01.06]** Release the code and models of ACE++.
@@ -52,8 +50,8 @@ use cases to explore.
- [x] **[2025.01.16]** Release the training code for lora.
- [x] **[2025.02.15]** Collection of workflows in Comfyui.
- [x] **[2025.02.15]** Release the config for fully fine-tuning.
- [x] **[2025.03.03]** Release a unified fft model for ACE++, support more image to image tasks.
- [x] **[2025.03.11]** Release the comfyui workflow for ACE++ FFT model.
- [x] **[2025.03.03]** Release the fft model for ACE++, support more image to image tasks.
- [x] **[2025.03.11]** Release some comfyui workflow examples for ACE++ model.
- We sincerely apologize
for the delayed responses and updates regarding ACE++ issues.
@@ -61,82 +59,118 @@ Further development of the ACE model through post-training on the FLUX model mus
We have identified several significant challenges in post-training on the FLUX foundation.
The primary issue is the high degree of heterogeneity between the training dataset and the FLUX model,
which results in highly unstable training. Moreover, FLUX-Dev is a distilled model, and the influence of its original negative prompts on its final performance is uncertain.
As a result, subsequent efforts will be focused on post-training the ACE model using the Wan series of foundational models.
- __I would like to emphasize that, due to the reasons mentioned earlier, the performance of the FFT model may decline compared
As a result, subsequent efforts will be focused on post-training the ACE model using the Wan series of foundational models. __Due to the reasons mentioned earlier, the performance of the FFT model may decline compared
to the LoRA model across various tasks. Therefore, we recommend continuing to use the LoRA model to achieve better results.
We provide the FFT model with the hope that it may facilitate academic exploration and research in this area.__
- We have been busy with other projects recently. Our new work in the video domain, [VACE](https://ali-vilab.github.io/VACE-Page/),
has also been released, and we welcome you to continue following our work.
## Models
### ACE++ Portrait LoRA
## 🔥The unified fft model for ACE++
Fully finetuning a composite model with ACE’s data to support various editing and reference generation tasks through an instructive approach.
Portrait-consistent generation to maintain the consistency of the portrait.
We found that there are conflicts between the repainting task and the editing task during the experimental process. This is because the edited image is concatenated with noise in the channel dimension, whereas the repainting task modifies the region using zero pixel values in the VAE's latent space. The editing task uses RGB pixel values in the modified region through the VAE's latent space, which is similar to the distribution of the non-modified part of the repainting task, making it a challenge for the model to distinguish between the two tasks.
To address this issue, we introduced 64 additional channels in the channel dimension to differentiate between these two tasks. In these channels, we place the latent representation of the pixel space from the edited image, while keeping other channels consistent with the repainting task. This approach significantly enhances the model's adaptability to different tasks.
One issue with this approach is that it changes the input channel number of the FLUX-Fill-Dev model from 384 to 448. The specific configuration can be referenced in the [configuration file](config/ace_plus_fft.yaml).
We used tools from [stella](https://gist.github.com/Stella2211/10f5bd870387ec1ddb9932235321068e) to convert the fft-fp16 model to fft-fp8.
But We have tested with some users and found that the FP8 version has precision loss, so we will not provide the FP8 version of the model for the time being.
### ComfyUI Workflow
Copy the workflow/ComfyUI-ACE_Plus folder into ComfyUI’s custom_nodes directory. Launch ComfyUI, and we have provided four example workflows in workflow_example_fft with the following explanations.
We provide a parameter to adjust the GPU memory usage. As shown in the figure below, max_seq_length controls the length of the token sequence during inference, thereby controlling the model's inference memory consumption.
The range of this value is from 1024 to 5120, and it correspondingly affects the clarity of the generated image. The smaller the value, the lower the image clarity.
<img src="./assets/comfyui/snapshot.jpg" width="800">
<table><tbody>
<tr>
<td>Workflow</td>
<td>Description</td>
<td>Setting</td>
<td>Tuning Method</td>
<td>Input</td>
<td>Output</td>
<td>Instruction</td>
<td>Models</td>
</tr>
<tr>
<td>ACE_Plus_FFT_workflow_no_preprocess.json</td>
<td>Use the preprocessed images, such as depth and contour, as input, or the super-resolution.</td>
<td>Task_type: no_preprocess (you don't need to install dependencies like scepter)</td>
<td>LoRA <br>+ ACE Data</td>
<td><img src="./assets/samples/portrait/human_1.jpg" width="200"></td>
<td><img src="./assets/samples/portrait/human_1_1.jpg" width="200"></td>
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Maintain the facial features. A girl is wearing a neat police uniform and sporting a badge. She is smiling with a friendly and confident demeanor. The background is blurred, featuring a cartoon logo."</td>
<td align="center" style="word-wrap:break-word;word-break:break-all;" width="200px";><a href="https://www.modelscope.cn/models/iic/ACE_Plus/"><img src="https://img.shields.io/badge/ModelScope-Model-blue" alt="ModelScope link"> </a> <a href="https://huggingface.co/ali-vilab/ACE_Plus/tree/main/portrait/"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model-yellow" alt="HuggingFace link"> </a> </td>
</tr>
</tbody>
</table>
Models' scepter_path:
- **ModelScope:** ms://iic/ACE_Plus@portrait/xxxx.safetensors
- **HuggingFace:** hf://ali-vilab/ACE_Plus@portrait/xxxx.safetensors
### ACE++ Subject LoRA
Subject-driven image generation task to maintain the consistency of a specific subject in different scenes.
<table><tbody>
<tr>
<td>Tuning Method</td>
<td>Input</td>
<td>Output</td>
<td>Instruction</td>
<td>Models</td>
</tr>
<tr>
<td>ACE_Plus_FFT_workflow_controlpreprocess.json</td>
<td>Controllable image-to-image translation capability. To preprocess depth and contour information from images,
we use externally-provided models that are typically downloaded from the ModelScope Hub. Because download success
can vary depending on the user's environment, we offer alternatives: users can either leverage existing community
nodes (depth extration node or contour extraction node) for this task (then choosing the 'no_preprocess' option),
or users can pre-download the required models
<a href="https://www.modelscope.cn/models/iic/scepter_annotator/file/view/master?fileName=annotator%252Fckpts%252Finformative_drawing_contour_style.pth&status=2">contour</a> and
<a href="https://www.modelscope.cn/models/iic/scepter_annotator/file/view/master?fileName=annotator%252Fckpts%252Fdpt_hybrid-midas-501f0c75.pt&status=2">depth</a>
and adjust
the configuration file 'workflow/ComfyUI-ACE_Plus/config/ace_plus_fft_processor.yaml' to
specify the models' local paths.</td>
<td>Task_type: contour_repainting/depth_repainting/recolorizing (you need to install dependencies like scepter)</td>
<td>LoRA <br>+ ACE Data</td>
<td><img src="./assets/samples/subject/subject_1.jpg" width="200"></td>
<td><img src="./assets/samples/subject/subject_1_1.jpg" width="200"></td>
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Display the logo in a minimalist style printed in white on a matte black ceramic coffee mug, alongside a steaming cup of coffee on a cozy cafe table."</td>
<td align="center" style="word-wrap:break-word;word-break:break-all;" width="200px";><a href="https://www.modelscope.cn/models/iic/ACE_Plus/"><img src="https://img.shields.io/badge/ModelScope-Model-blue" alt="ModelScope link"> </a> <a href="https://huggingface.co/ali-vilab/ACE_Plus/tree/main/subject/"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model-yellow" alt="HuggingFace link"> </a> </td>
</tr>
</tbody>
</table>
Models' scepter_path:
- **ModelScope:** ms://iic/ACE_Plus@subject/xxxx.safetensors
- **HuggingFace:** hf://ali-vilab/ACE_Plus@subject/xxxx.safetensors
### ACE++ LocalEditing LoRA
Redrawing the mask area of images while maintaining the original structural information of the edited area.
<table><tbody>
<tr>
<td>Tuning Method</td>
<td>Input</td>
<td>Output</td>
<td>Instruction</td>
<td>Models</td>
</tr>
<tr>
<td>ACE_Plus_FFT_workflow_reference_generation.json</td>
<td>Reference image generation capability for portrait or subject.</td>
<td>Task_type: repainting (you don't need to install dependencies like scepter)</td>
<td>LoRA <br>+ ACE Data</td>
<td><img src="./assets/samples/local/local_1.webp" width="200"><br><img src="./assets/samples/local/local_1_m.webp" width="200"></td>
<td><img src="./assets/samples/local/local_1_1.jpg" width="200"></td>
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"By referencing the mask, restore a partial image from the doodle {image} that aligns with the textual explanation: "1 white old owl"."</td>
<td align="center" style="word-wrap:break-word;word-break:break-all;" width="200px";><a href="https://www.modelscope.cn/models/iic/ACE_Plus/"><img src="https://img.shields.io/badge/ModelScope-Model-blue" alt="ModelScope link"> </a> <a href="https://huggingface.co/ali-vilab/ACE_Plus/tree/main/local_editing/"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model-yellow" alt="HuggingFace link"> </a> </td>
</tr>
<tr>
<td>ACE_Plus_FFT_workflow_referenceediting_generation.json</td>
<td>Reference image editing capability</td>
<td>Task_type: repainting (you don't need to install dependencies like scepter)</td>
</tr>
<tbody>
<table>
</tbody>
</table>
Models' scepter_path:
- **ModelScope:** ms://iic/ACE_Plus@local_editing/xxxx.safetensors
- **HuggingFace:** hf://ali-vilab/ACE_Plus@local_editing/xxxx.safetensors
### ACE++ FFT model
Fully finetuning a composite model with ACE’s data to support various editing and reference generation tasks through an instructive approach.
We introduced 64 additional channels in the channel dimension to differentiate between the repainting task and the editing task. In these channels, we place the latent representation of the pixel space from the edited image, while keeping other channels consistent with the repainting task. One issue with this approach is that it changes the input channel number of the FLUX-Fill-Dev model from 384 to 448. The specific configuration can be referenced in the [configuration file](config/ace_plus_fft.yaml).
### Examples
The ACE++ model supports a wide range of downstream tasks through simple adaptations. Here are some examples.
<table><tbody>
<tr>
<th align="center" colspan="1">ACE++ Model</th>
<td>Input Reference Image</td>
<td>Input Edit Image</td>
<td>Input Edit Mask</td>
@@ -145,6 +179,7 @@ The range of this value is from 1024 to 5120, and it correspondingly affects the
<td>Function</td>
</tr>
<tr>
<td>Portrait LoRA(recommended) / FFT model</td>
<td><img src="./assets/samples/portrait/human_1.jpg" width="200"></td>
<td></td>
<td></td>
@@ -153,14 +188,16 @@ The range of this value is from 1024 to 5120, and it correspondingly affects the
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Character ID Consistency Generation"</td>
</tr>
<tr>
<td>Subject LoRA(recommended) / FFT model</td>
<td><img src="./assets/samples/subject/subject_1.jpg" width="200"></td>
<td></td>
<td></td>
<td></td>
<td><img src="./assets/samples/subject/subject_1_fft.webp" width="200"></td>
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Display the logo in a minimalist style printed in white on a matte black ceramic coffee mug, alongside a steaming cup of coffee on a cozy cafe table."</td>
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Subject Consistency Generation"</td>
</tr>
<tr>
<td>Subject LoRA(recommended) / FFT model</td>
<td><img src="./assets/samples/application/photo_editing/1_ref.png" width="200"></td>
<td><img src="./assets/samples/application/photo_editing/1_2_edit.jpg" width="200"></td>
<td><img src="./assets/samples/application/photo_editing/1_2_m.webp" width="200"></td>
@@ -169,6 +206,7 @@ The range of this value is from 1024 to 5120, and it correspondingly affects the
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Subject Consistency Editing"</td>
</tr>
<tr>
<td>Subject LoRA(recommended) / FFT model</td>
<td><img src="./assets/samples/application/logo_paste/1_ref.png" width="200"></td>
<td><img src="./assets/samples/application/logo_paste/1_1_edit.png" width="200"></td>
<td><img src="./assets/samples/application/logo_paste/1_1_m.png" width="200"></td>
@@ -177,6 +215,7 @@ The range of this value is from 1024 to 5120, and it correspondingly affects the
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Subject Consistency Editing"</td>
</tr>
<tr>
<td>Subject LoRA(recommended) / FFT model</td>
<td><img src="./assets/samples/application/try_on/1_ref.png" width="200"></td>
<td><img src="./assets/samples/application/try_on/1_1_edit.png" width="200"></td>
<td><img src="./assets/samples/application/try_on/1_1_m.png" width="200"></td>
@@ -185,6 +224,7 @@ The range of this value is from 1024 to 5120, and it correspondingly affects the
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Try On"</td>
</tr>
<tr>
<td>Portrait LoRA(recommended) / FFT model</td>
<td><img src="./assets/samples/application/movie_poster/1_ref.png" width="200"></td>
<td><img src="./assets/samples/portrait/human_1.jpg" width="200"></td>
<td><img src="./assets/samples/application/movie_poster/1_2_m.webp" width="200"></td>
@@ -193,6 +233,7 @@ The range of this value is from 1024 to 5120, and it correspondingly affects the
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Face swap"</td>
</tr>
<tr>
<td>FFT model</td>
<td></td>
<td><img src="./assets/samples/application/sr/sr_tiger.png" width="200"></td>
<td><img src="./assets/samples/application/sr/sr_tiger_m.webp" width="200"></td>
@@ -201,6 +242,7 @@ The range of this value is from 1024 to 5120, and it correspondingly affects the
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Super-resolution"</td>
</tr>
<tr>
<td>FFT model</td>
<td></td>
<td><img src="./assets/samples/application/photo_editing/1_ref.png" width="200"></td>
<td><img src="./assets/samples/application/photo_editing/1_1_orm.webp" width="200"></td>
@@ -209,6 +251,7 @@ The range of this value is from 1024 to 5120, and it correspondingly affects the
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Regional Editing"</td>
</tr>
<tr>
<td>FFT model</td>
<td></td>
<td><img src="./assets/samples/application/photo_editing/1_ref.png" width="200"></td>
<td><img src="./assets/samples/application/photo_editing/1_1_rm.webp" width="200"></td>
@@ -217,6 +260,7 @@ The range of this value is from 1024 to 5120, and it correspondingly affects the
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Regional Editing"</td>
</tr>
<tr>
<td>Local Editing LoRA/FFT model</td>
<td></td>
<td><img src="./assets/samples/control/1_1_recolor.webp" width="200"></td>
<td><img src="./assets/samples/control/1_1_m.webp" width="200"></td>
@@ -224,7 +268,9 @@ The range of this value is from 1024 to 5120, and it correspondingly affects the
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"{image} Beautiful female portrait, Robot with smooth White transparent carbon shell, rococo detailing, Natural lighting, Highly detailed, Cinematic, 4K."</td>
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Recolorizing"</td>
</tr>
<tr>
<td>Local Editing LoRA/FFT model</td>
<td></td>
<td><img src="./assets/samples/control/1_1_depth.webp" width="200"></td>
<td><img src="./assets/samples/control/1_1_m.webp" width="200"></td>
@@ -233,6 +279,7 @@ The range of this value is from 1024 to 5120, and it correspondingly affects the
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Depth Guided Generation"</td>
</tr>
<tr>
<td>Local Editing LoRA/FFT model</td>
<td></td>
<td><img src="./assets/samples/control/1_1_contourc.webp" width="200"></td>
<td><img src="./assets/samples/control/1_1_m.webp" width="200"></td>
@@ -243,7 +290,6 @@ The range of this value is from 1024 to 5120, and it correspondingly affects the
</tbody>
</table>
## Comfyui Workflows in community
We are deeply grateful to the community developers for building many fascinating applications based on the ACE++ series of models.
During this process, we have received valuable feedback, particularly regarding artifacts in generated images and the stability of the results.
@@ -337,131 +383,103 @@ Additionally, many bloggers have published tutorials on how to use it, which are
</tbody>
</table>
## ComfyUI Workflow Examples
### ACE++ Portrait
Portrait-consistent generation to maintain the consistency of the portrait.
Copy the workflow/ComfyUI-ACE_Plus folder into ComfyUI’s custom_nodes directory. Launch ComfyUI, and we have provided some example workflows in workflow_example with the following explanations. It is recommended to use the LoRA model workflow, as it offers more stable results compared to the FFT model.
<table><tbody>
<tr>
<td>Tuning Method</td>
<td>Input</td>
<td>Output</td>
<td>Instruction</td>
<td>Models</td>
<td>Workflow</td>
<td>Description</td>
<td>Other dependency models</td>
<td>Setting</td>
</tr>
<tr>
<td>LoRA <br>+ ACE Data</td>
<td><img src="./assets/samples/portrait/human_1.jpg" width="200"></td>
<td><img src="./assets/samples/portrait/human_1_1.jpg" width="200"></td>
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Maintain the facial features. A girl is wearing a neat police uniform and sporting a badge. She is smiling with a friendly and confident demeanor. The background is blurred, featuring a cartoon logo."</td>
<td align="center" style="word-wrap:break-word;word-break:break-all;" width="200px";><a href="https://www.modelscope.cn/models/iic/ACE_Plus/"><img src="https://img.shields.io/badge/ModelScope-Model-blue" alt="ModelScope link"> </a> <a href="https://huggingface.co/ali-vilab/ACE_Plus/tree/main/portrait/"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model-yellow" alt="HuggingFace link"> </a> </td>
<td>ACE_Plus_LoRA_workflow_reference_generation.json</td>
<td>Reference image generation capability for portrait or subject.</td>
<td>Potrait or subject LoRA Model + FLUX.1-Fill-dev</td>
<td>Task_type: repainting (you don't need to install dependencies like scepter)</td>
</tr>
</tbody>
</table>
<tr>
<td>ACE_Plus_LoRA_workflow_redux_reference_generation.json</td>
<td>Reference image generation capability for portrait or subject used in conjunction with Redux.</td>
<td>Potrait or subject LoRA Model + FLUX.1-Fill-dev + FLUX.1-Redux</td>
<td>Task_type: repainting (you don't need to install dependencies like scepter)</td>
</tr>
<tr>
<td>ACE_Plus_LoRA_workflow_reference_editing.json</td>
<td>Reference image editing capability such as logo paste, face swap.</td>
<td>Potrait or subject LoRA Model + FLUX.1-Fill-dev</td>
<td>Task_type: repainting (you don't need to install dependencies like scepter)</td>
</tr>
<tr>
<td>ACE_Plus_LoRA_workflow_redux_reference_editing.json</td>
<td>Reference image editing capability such as logo paste, face swap used in conjunction with Redux.</td>
<td>Potrait or subject LoRA Model + FLUX.1-Fill-dev + FLUX.1-Redux</td>
<td>Task_type: repainting (you don't need to install dependencies like scepter)</td>
</tr>
<tr>
<td>ACE_Plus_LoRA_workflow_localcontrol_generation.json</td>
<td>Controllable image-to-image translation capability. To preprocess depth and contour information from images,
we use externally-provided models that are typically downloaded from the ModelScope Hub. Because download success
can vary depending on the user's environment, we offer alternatives: users can either leverage existing community
nodes (depth extration node or contour extraction node) for this task (then choosing the 'no_preprocess' option),
or users can pre-download the required models
<a href="https://www.modelscope.cn/models/iic/scepter_annotator/file/view/master?fileName=annotator%252Fckpts%252Finformative_drawing_contour_style.pth&status=2">contour</a> and
<a href="https://www.modelscope.cn/models/iic/scepter_annotator/file/view/master?fileName=annotator%252Fckpts%252Fdpt_hybrid-midas-501f0c75.pt&status=2">depth</a>
and adjust
the configuration file 'workflow/ComfyUI-ACE_Plus/config/ace_plus_fft_processor.yaml' to
specify the models' local paths.</td>
<td>Local editing LoRA Model + FLUX.1-Fill-dev + Preprocessing model (depth or contour) </td>
<td>Task_type: contour_repainting/depth_repainting/recolorizing (you need to install dependencies like scepter)</td>
</tr>
<tr>
<td>ACE_Plus_FFT_workflow_referenceediting_generation.json</td>
<td>Reference image editing capability</td>
<td>FFT model</td>
<td>Task_type: repainting (you don't need to install dependencies like scepter)</td>
</tr>
<tr>
<td>ACE_Plus_FFT_workflow_no_preprocess.json</td>
<td>Use the preprocessed images, such as depth and contour, as input, or the super-resolution.</td>
<td>FFT model</td>
<td>Task_type: no_preprocess (you don't need to install dependencies like scepter)</td>
</tr>
<tr>
<td>ACE_Plus_FFT_workflow_controlpreprocess.json</td>
<td>Controllable image-to-image translation capability. To preprocess depth and contour information from images,
we use externally-provided models that are typically downloaded from the ModelScope Hub. Because download success
can vary depending on the user's environment, we offer alternatives: users can either leverage existing community
nodes (depth extration node or contour extraction node) for this task (then choosing the 'no_preprocess' option),
or users can pre-download the required models
<a href="https://www.modelscope.cn/models/iic/scepter_annotator/file/view/master?fileName=annotator%252Fckpts%252Finformative_drawing_contour_style.pth&status=2">contour</a> and
<a href="https://www.modelscope.cn/models/iic/scepter_annotator/file/view/master?fileName=annotator%252Fckpts%252Fdpt_hybrid-midas-501f0c75.pt&status=2">depth</a>
and adjust
the configuration file 'workflow/ComfyUI-ACE_Plus/config/ace_plus_fft_processor.yaml' to
specify the models' local paths.</td>
<td>FFT model</td>
<td>Task_type: contour_repainting/depth_repainting/recolorizing (you need to install dependencies like scepter)</td>
</tr>
<tr>
<td>ACE_Plus_FFT_workflow_reference_generation.json</td>
<td>Reference image generation capability for portrait or subject.</td>
<td>FFT model</td>
<td>Task_type: repainting (you don't need to install dependencies like scepter)</td>
</tr>
<tr>
<td>ACE_Plus_FFT_workflow_referenceediting_generation.json</td>
<td>Reference image editing capability</td>
<td>FFT model</td>
<td>Task_type: repainting (you don't need to install dependencies like scepter)</td>
</tr>
<tbody>
<table>
Models' scepter_path:
- **ModelScope:** ms://iic/ACE_Plus@portrait/xxxx.safetensors
- **HuggingFace:** hf://ali-vilab/ACE_Plus@portrait/xxxx.safetensors
As shown in the figure below, max_seq_length controls the length of the token sequence during inference, thereby controlling the model's inference memory consumption.
The range of this value is from 1024 to 5120, and it correspondingly affects the clarity of the generated image. The smaller the value, the lower the image clarity.
<img src="./assets/comfyui/snapshot.jpg" width="200">
### ACE++ Subject
Subject-driven image generation task to maintain the consistency of a specific subject in different scenes.
<table><tbody>
<tr>
<td>Tuning Method</td>
<td>Input</td>
<td>Output</td>
<td>Instruction</td>
<td>Models</td>
</tr>
<tr>
<td>LoRA <br>+ ACE Data</td>
<td><img src="./assets/samples/subject/subject_1.jpg" width="200"></td>
<td><img src="./assets/samples/subject/subject_1_1.jpg" width="200"></td>
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"Display the logo in a minimalist style printed in white on a matte black ceramic coffee mug, alongside a steaming cup of coffee on a cozy cafe table."</td>
<td align="center" style="word-wrap:break-word;word-break:break-all;" width="200px";><a href="https://www.modelscope.cn/models/iic/ACE_Plus/"><img src="https://img.shields.io/badge/ModelScope-Model-blue" alt="ModelScope link"> </a> <a href="https://huggingface.co/ali-vilab/ACE_Plus/tree/main/subject/"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model-yellow" alt="HuggingFace link"> </a> </td>
</tr>
</tbody>
</table>
Models' scepter_path:
- **ModelScope:** ms://iic/ACE_Plus@subject/xxxx.safetensors
- **HuggingFace:** hf://ali-vilab/ACE_Plus@subject/xxxx.safetensors
### ACE++ LocalEditing
Redrawing the mask area of images while maintaining the original structural information of the edited area.
<table><tbody>
<tr>
<td>Tuning Method</td>
<td>Input</td>
<td>Output</td>
<td>Instruction</td>
<td>Models</td>
</tr>
<tr>
<td>LoRA <br>+ ACE Data</td>
<td><img src="./assets/samples/local/local_1.webp" width="200"><br><img src="./assets/samples/local/local_1_m.webp" width="200"></td>
<td><img src="./assets/samples/local/local_1_1.jpg" width="200"></td>
<td style="word-wrap:break-word;word-break:break-all;" width="250px";>"By referencing the mask, restore a partial image from the doodle {image} that aligns with the textual explanation: "1 white old owl"."</td>
<td align="center" style="word-wrap:break-word;word-break:break-all;" width="200px";><a href="https://www.modelscope.cn/models/iic/ACE_Plus/"><img src="https://img.shields.io/badge/ModelScope-Model-blue" alt="ModelScope link"> </a> <a href="https://huggingface.co/ali-vilab/ACE_Plus/tree/main/local_editing/"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-Model-yellow" alt="HuggingFace link"> </a> </td>
</tr>
</tbody>
</table>
Models' scepter_path:
- **ModelScope:** ms://iic/ACE_Plus@local_editing/xxxx.safetensors
- **HuggingFace:** hf://ali-vilab/ACE_Plus@local_editing/xxxx.safetensors
## 🔥 Applications
The ACE++ model supports a wide range of downstream tasks through simple adaptations. Here are some examples, and we look forward to seeing the community explore even more exciting applications utilizing the ACE++ model.
<table><tbody>
<tr>
<th align="center" colspan="1">Application</th>
<th align="center" colspan="1">ACE++ Model</th>
<th align="center" colspan="5">Examples</th>
</tr>
<tr>
<td>Try On</td>
<td>ACE++ Subject</td>
<td><img src="./assets/samples/application/try_on/1_ref.png" width="200"></td>
<td><img src="./assets/samples/application/try_on/1_1_edit.png" width="200"></td>
<td><img src="./assets/samples/application/try_on/1_1_m.png" width="200"></td>
<td><img src="./assets/samples/application/try_on/1_1_res.png" width="200"></td>
<td style="word-wrap:break-word;word-break:break-all;" width="100px";>"The woman dresses this skirt."</td>
</tr>
<tr>
<td>Logo Paste</td>
<td>ACE++ Subject</td>
<td><img src="./assets/samples/application/logo_paste/1_ref.png" width="200"></td>
<td><img src="./assets/samples/application/logo_paste/1_1_edit.png" width="200"></td>
<td><img src="./assets/samples/application/logo_paste/1_1_m.png" width="200"></td>
<td><img src="./assets/samples/application/logo_paste/1_1_res.webp" width="200"></td>
<td style="word-wrap:break-word;word-break:break-all;" width="100px";>"The logo is printed on the headphones."</td>
</tr>
<tr>
<td>Photo Editing</td>
<td>ACE++ Subject</td>
<td><img src="./assets/samples/application/photo_editing/1_ref.png" width="200"></td>
<td><img src="./assets/samples/application/photo_editing/1_1_edit.png" width="200"></td>
<td><img src="./assets/samples/application/photo_editing/1_1_m.png" width="200"></td>
<td><img src="./assets/samples/application/photo_editing/1_1_res.jpg" width="200"></td>
<td style="word-wrap:break-word;word-break:break-all;" width="100px";>"The item is put on the ground."</td>
</tr>
<tr>
<td>Movie Poster Editor</td>
<td>ACE++ Portrait</td>
<td><img src="./assets/samples/application/movie_poster/1_ref.png" width="200"></td>
<td><img src="./assets/samples/application/movie_poster/1_1_edit.png" width="200"></td>
<td><img src="./assets/samples/application/movie_poster/1_1_m.png" width="200"></td>
<td><img src="./assets/samples/application/movie_poster/1_1_res.webp" width="200"></td>
<td style="word-wrap:break-word;word-break:break-all;" width="100px";>"The man is facing the camera and is smiling."</td>
</tr>
</tbody>
</table>
## ⚙️️ Installation
Download the code using the following command:
@@ -537,7 +555,6 @@ python run_train.py --cfg train_config/ace_plus_fft.yaml
The models trained by ACE++ can be found in ./examples/exp_example/xxxx/checkpoints/xxxx/0_SwiftLoRA/comfyui_model.safetensors.
## 💻 Demo
We have built a GUI demo based on Gradio to help users better utilize the ACE++ model. Just execute the following command.
```bash
@@ -556,6 +573,7 @@ export ACE_PLUS_FFT_MODEL="ms://iic/ACE_Plus@ace_plus_fft.safetensors.safetensor
python demo_fft.py
```
## 📚 Limitations
* For certain tasks, such as deleting and adding objects, there are flaws in instruction following. For adding and replacing objects, we recommend trying the repainting method of the local editing model to achieve this.
* The generated results may contain artifacts, especially when it comes to the generation of hands, which still exhibit distortions.
+5 -2
View File
@@ -1,11 +1,14 @@
# -*- coding: utf-8 -*-
# Copyright (c) Alibaba, Inc. and its affiliates.
from .ace_plus_fft_node import ACEPlusFFTLoader, ACEPlusFFTConditioning, AcePlusFFTProcessor
from .ace_plus_fft_node import (ACEPlusFFTLoader, ACEPlusFFTConditioning, ACEPlusLoraConditioning,
AcePlusFFTProcessor, AcePlusLoraProcessor)
NODE_MAPPINGS = {
'ACEPlusLoader': ('ACEPlusFFTLoader~', ACEPlusFFTLoader),
'ACEPlusConditioning': ('ACEPlusFFTConditioning~', ACEPlusFFTConditioning),
'ACEPlusFFTProcessor': ('ACEPlusFFTProcessor~', AcePlusFFTProcessor)
'ACEPlusFFTProcessor': ('ACEPlusFFTProcessor~', AcePlusFFTProcessor),
'ACEPlusLoraProcessor': ('ACEPlusLoraProcessor~', AcePlusLoraProcessor),
'ACEPlusLoraConditioning': ('ACEPlusLoraConditioning~', ACEPlusLoraConditioning)
}
NODE_CLASS_MAPPINGS = {k: v[1] for k, v in NODE_MAPPINGS.items()}
@@ -123,6 +123,55 @@ class ACEPlusFFTConditioning:
return (out[0], out[1], out_latent)
class ACEPlusLoraConditioning:
@classmethod
def INPUT_TYPES(s):
return {"required": {"positive": ("CONDITIONING",),
"negative": ("CONDITIONING",),
"vae": ("VAE",),
"pixels": ("IMAGE",),
"mask": ("MASK",),
"noise_mask": ("BOOLEAN", {"default": True,
"tooltip": "Add a noise mask to the latent so sampling will only happen within the mask. Might improve results or completely break things depending on the model."}),
}}
RETURN_TYPES = ("CONDITIONING", "CONDITIONING", "LATENT")
RETURN_NAMES = ("positive", "negative", "latent")
FUNCTION = "encode"
CATEGORY = "ComfyUI-ACE_Plus"
def encode(self, positive, negative, pixels, vae, mask, noise_mask=True):
x = (pixels.shape[1] // 8) * 8
y = (pixels.shape[2] // 8) * 8
mask = torch.nn.functional.interpolate(mask.reshape((-1, 1, mask.shape[-2], mask.shape[-1])),
size=(pixels.shape[1], pixels.shape[2]), mode="bilinear")
orig_pixels = pixels
pixels = orig_pixels.clone()
if pixels.shape[1] != x or pixels.shape[2] != y:
x_offset = (pixels.shape[1] % 8) // 2
y_offset = (pixels.shape[2] % 8) // 2
pixels = pixels[:, x_offset:x + x_offset, y_offset:y + y_offset, :]
mask = mask[:, :, x_offset:x + x_offset, y_offset:y + y_offset]
concat_latent = vae.encode(pixels)
orig_latent = vae.encode(orig_pixels)
out_latent = {}
out_latent["samples"] = orig_latent
if noise_mask:
out_latent["noise_mask"] = mask
out = []
for conditioning in [positive, negative]:
c = node_helpers.conditioning_set_values(conditioning, {"concat_latent_image": concat_latent,
"concat_mask": mask})
out.append(c)
return (out[0], out[1], out_latent)
import torch
import math
@@ -345,3 +394,192 @@ class AcePlusFFTProcessor:
slice_w = slice_w if slice_w < 30 else slice_w + 30
return edit_image + 0.5, change_image + 0.5, edit_mask, out_h, out_w, slice_w
class AcePlusLoraProcessor:
def __init__(self,
max_aspect_ratio=4,
d=16,
max_seq_len=1024):
self.max_aspect_ratio = max_aspect_ratio
self.max_seq_len = max_seq_len
self.d = d
current_dir = os.path.dirname(os.path.abspath(__file__))
config_path = os.path.join(current_dir, 'config', 'ace_plus_fft_processor.yaml')
self.processor_cfg = self.load_yaml(config_path)
self.task_list = {}
for task in self.processor_cfg['PREPROCESSOR']:
self.task_list[task['TYPE']] = task
self.transforms = T.Compose([
T.ToTensor(),
T.Normalize(mean=[0, 0, 0], std=[1.0, 1.0, 1.0])
])
CATEGORY = 'ComfyUI-ACE_Plus'
@classmethod
def INPUT_TYPES(s):
return {
'required': {
'use_reference': ('BOOLEAN', {'default': True}),
'height': ('INT', {
'default': 1024,
'min': 256,
'max': 1436,
'step': 16
}),
'width': ('INT', {
'default': 1024,
'min': 256,
'max': 1436,
'step': 16
}),
'task_type': (list(s().task_list.keys()),),
'max_seq_length': ('INT', {
'default': 3072,
'min': 1024,
'max': 5120,
'step': 0.01
}),
},
'optional': {
'reference_image': ('IMAGE',),
'edit_image': ('IMAGE',),
'edit_mask': ('MASK',),
}
}
OUTPUT_NODE = True
RETURN_TYPES = ('IMAGE', 'MASK', 'INT', 'INT', 'INT')
RETURN_NAMES = ('IMAGE', 'MASK', 'OUT_H', 'OUT_W', 'SLICE_W')
FUNCTION = 'preprocess'
def load_yaml(self, cfg_file):
with open(cfg_file, 'r') as f:
cfg = yaml.load(f.read(), Loader=yaml.SafeLoader)
return cfg
def image_check(self, image):
if image is None:
return image
# preprocess
H, W = image.shape[1: 3]
image = image.permute(0, 3, 1, 2)
if H / W > self.max_aspect_ratio:
image[0] = T.CenterCrop([int(self.max_aspect_ratio * W), W])(image[0])
elif W / H > self.max_aspect_ratio:
image[0] = T.CenterCrop([H, int(self.max_aspect_ratio * H)])(image[0])
return image[0]
def trans_pil_tensor(self, pil_image):
transform = T.Compose([
T.ToTensor()
])
tensor_image = transform(pil_image)
return tensor_image
def edit_preprocess(self, processor, device, edit_image, edit_mask):
if edit_image is None or processor is None:
return edit_image
if not SCEPTER:
raise ImportError(f'Please install scepter to use edit processor {processor} by '
f'runing "pip install scepter" in the conda env')
processor = Config(cfg_dict=processor, load=False)
processor = ANNOTATORS.build(processor).to(device)
edit_image = Image.fromarray(np.array(edit_image[0] * 255).astype(np.uint8)).convert('RGB')
new_edit_image = processor(np.asarray(edit_image))
del processor
new_edit_image = Image.fromarray(new_edit_image)
if edit_mask is not None:
edit_mask = np.where(edit_mask > 0.5, 1, 0) * 255
edit_mask = Image.fromarray(np.array(edit_mask[0]).astype(np.uint8)).convert('L')
if new_edit_image.size != edit_image.size:
new_edit_image = T.Resize((edit_image.size[1], edit_image.size[0]),
interpolation=T.InterpolationMode.BILINEAR,
antialias=True)(new_edit_image)
image = Image.composite(new_edit_image, edit_image, edit_mask)
return self.trans_pil_tensor(image).unsqueeze(0).permute(0, 2, 3, 1)
def preprocess(self,
reference_image=None,
edit_image=None,
edit_mask=None,
use_reference=True,
task_type=None,
height=1024,
width=1024,
max_seq_length=4096):
self.max_seq_len = max_seq_length
if not use_reference and edit_image is not None:
reference_image = None
if edit_mask is not None and edit_image is not None:
iH, iW = edit_image.shape[1:3]
mH, mW = edit_mask.shape[1:3]
if iH != mH or iW != mW:
edit_mask = torch.ones(edit_image.shape[:3])
if task_type != 'repainting':
repainting_scale = 0.0
else:
repainting_scale = 1.0
if task_type in self.task_list:
edit_image = self.edit_preprocess(self.task_list[task_type]['ANNOTATOR'], 0,
edit_image, edit_mask)
if reference_image is not None:
reference_image = self.image_check(reference_image) - 0.5
if edit_image is not None:
edit_image = self.image_check(edit_image) - 0.5
# for reference generation
if edit_image is None:
edit_image = torch.zeros([3, height, width])
edit_mask = torch.ones([1, height, width])
else:
if edit_mask is None:
_, eH, eW = edit_image.shape
edit_mask = np.ones((eH, eW))
else:
edit_mask = np.asarray(edit_mask)[0]
edit_mask = np.where(edit_mask > 0.5, 1, 0)
edit_mask = edit_mask.astype(
np.float32) if np.any(edit_mask) else np.ones_like(edit_mask).astype(
np.float32)
edit_mask = torch.tensor(edit_mask).unsqueeze(0)
edit_image = edit_image * (1 - edit_mask * repainting_scale)
out_h, out_w = edit_image.shape[-2:]
assert edit_mask is not None
if reference_image is not None:
_, H, W = reference_image.shape
_, eH, eW = edit_image.shape
# align height with edit_image
scale = eH / H
tH, tW = eH, int(W * scale)
reference_image = T.Resize((tH, tW), interpolation=T.InterpolationMode.BILINEAR, antialias=True)(
reference_image)
edit_image = torch.cat([reference_image, edit_image], dim=-1)
edit_mask = torch.cat([torch.zeros([1, reference_image.shape[1], reference_image.shape[2]]), edit_mask],
dim=-1)
slice_w = reference_image.shape[-1]
else:
slice_w = 0
H, W = edit_image.shape[-2:]
scale = min(1.0, math.sqrt(self.max_seq_len * 2 / ((H / self.d) * (W / self.d))))
rH = int(H * scale) // self.d * self.d
rW = int(W * scale) // self.d * self.d
slice_w = int(slice_w * scale) // self.d * self.d
edit_image = T.Resize((rH, rW), interpolation=T.InterpolationMode.NEAREST_EXACT, antialias=True)(edit_image)
edit_mask = T.Resize((rH, rW), interpolation=T.InterpolationMode.NEAREST_EXACT, antialias=True)(edit_mask)
edit_image = edit_image.unsqueeze(0).permute(0, 2, 3, 1)
slice_w = slice_w if slice_w < 30 else slice_w + 30
return edit_image + 0.5, edit_mask, out_h, out_w, slice_w
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
@@ -0,0 +1,996 @@
{
"last_node_id": 381,
"last_link_id": 707,
"nodes": [
{
"id": 289,
"type": "DualCLIPLoader",
"pos": [
-3826.0087890625,
1734.926513671875
],
"size": [
315,
106
],
"flags": {},
"order": 0,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "CLIP",
"type": "CLIP",
"links": [
439,
440
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "DualCLIPLoader"
},
"widgets_values": [
"clip_l.safetensors",
"t5xxl_fp16.safetensors",
"flux"
]
},
{
"id": 287,
"type": "VAEDecode",
"pos": [
-2991.447509765625,
1874.02880859375
],
"size": [
210,
46
],
"flags": {},
"order": 15,
"mode": 0,
"inputs": [
{
"name": "samples",
"type": "LATENT",
"link": 442
},
{
"name": "vae",
"type": "VAE",
"link": 441
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
652
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAEDecode"
},
"widgets_values": []
},
{
"id": 373,
"type": "ImageCrop",
"pos": [
-3042.69140625,
2015.809326171875
],
"size": [
315,
130
],
"flags": {},
"order": 16,
"mode": 0,
"inputs": [
{
"name": "image",
"type": "IMAGE",
"link": 652
},
{
"name": "width",
"type": "INT",
"link": 706,
"widget": {
"name": "width"
}
},
{
"name": "height",
"type": "INT",
"link": 707,
"widget": {
"name": "height"
}
},
{
"name": "x",
"type": "INT",
"link": 680,
"widget": {
"name": "x"
}
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
658
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "ImageCrop"
},
"widgets_values": [
512,
512,
0,
0
]
},
{
"id": 374,
"type": "PreviewImage",
"pos": [
-2646.8134765625,
1807.708740234375
],
"size": [
297.7324523925781,
332.22821044921875
],
"flags": {},
"order": 17,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 658
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": [],
"color": "#232",
"bgcolor": "#353"
},
{
"id": 371,
"type": "MaskPreview+",
"pos": [
-4157.91162109375,
2039.6533203125
],
"size": [
210,
246
],
"flags": {},
"order": 11,
"mode": 0,
"inputs": [
{
"name": "mask",
"type": "MASK",
"link": 677
}
],
"outputs": [],
"properties": {
"Node name for S&R": "MaskPreview+"
},
"widgets_values": []
},
{
"id": 376,
"type": "LoadImage",
"pos": [
-4823.74267578125,
1608.7518310546875
],
"size": [
315,
314
],
"flags": {},
"order": 1,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
673
],
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": null
}
],
"properties": {
"Node name for S&R": "LoadImage"
},
"widgets_values": [
"human_1.jpg",
"image"
]
},
{
"id": 295,
"type": "DifferentialDiffusion",
"pos": [
-3524.858154296875,
2274.34521484375
],
"size": [
203.87991333007812,
28.94829559326172
],
"flags": {},
"order": 12,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 697
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
438
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "DifferentialDiffusion"
},
"widgets_values": []
},
{
"id": 380,
"type": "LoraLoaderModelOnly",
"pos": [
-3876.464599609375,
2339.6025390625
],
"size": [
315,
82
],
"flags": {},
"order": 8,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 696
}
],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
697
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "LoraLoaderModelOnly"
},
"widgets_values": [
"comfyui_face_lora64.safetensors",
1
]
},
{
"id": 367,
"type": "PreviewImage",
"pos": [
-4346.46435546875,
1542.8682861328125
],
"size": [
210,
246
],
"flags": {},
"order": 10,
"mode": 0,
"inputs": [
{
"name": "images",
"type": "IMAGE",
"link": 676
}
],
"outputs": [],
"properties": {
"Node name for S&R": "PreviewImage"
},
"widgets_values": []
},
{
"id": 290,
"type": "VAELoader",
"pos": [
-3826.618408203125,
1603.526611328125
],
"size": [
315,
58
],
"flags": {},
"order": 2,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "VAE",
"type": "VAE",
"links": [
441,
700
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "VAELoader"
},
"widgets_values": [
"ae.safetensors"
]
},
{
"id": 294,
"type": "FluxGuidance",
"pos": [
-3702.181396484375,
1909.62646484375
],
"size": [
210,
58
],
"flags": {},
"order": 9,
"mode": 0,
"inputs": [
{
"name": "conditioning",
"type": "CONDITIONING",
"link": 446
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
701
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "FluxGuidance"
},
"widgets_values": [
50
]
},
{
"id": 292,
"type": "CLIPTextEncode",
"pos": [
-4063.243896484375,
1959.1868896484375
],
"size": [
400,
200
],
"flags": {
"collapsed": true
},
"order": 6,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 440
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
702
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
""
],
"color": "#322",
"bgcolor": "#533"
},
{
"id": 381,
"type": "ACEPlusLoraConditioning",
"pos": [
-3861.20703125,
2016.699951171875
],
"size": [
315,
138
],
"flags": {},
"order": 13,
"mode": 0,
"inputs": [
{
"name": "positive",
"type": "CONDITIONING",
"link": 701
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 702
},
{
"name": "vae",
"type": "VAE",
"link": 700
},
{
"name": "pixels",
"type": "IMAGE",
"link": 698
},
{
"name": "mask",
"type": "MASK",
"link": 699
}
],
"outputs": [
{
"name": "positive",
"type": "CONDITIONING",
"links": [
703
],
"slot_index": 0
},
{
"name": "negative",
"type": "CONDITIONING",
"links": [
704
],
"slot_index": 1
},
{
"name": "latent",
"type": "LATENT",
"links": [
705
],
"slot_index": 2
}
],
"properties": {
"Node name for S&R": "ACEPlusLoraConditioning"
},
"widgets_values": [
false
]
},
{
"id": 323,
"type": "Note",
"pos": [
-4408.5859375,
2355.6123046875
],
"size": [
454.2545166015625,
174.908447265625
],
"flags": {},
"order": 3,
"mode": 0,
"inputs": [],
"outputs": [],
"properties": {},
"widgets_values": [
"ACE_Plus Model download: https://huggingface.co/ali-vilab/ACE_Plus/blob/main"
],
"color": "#332922",
"bgcolor": "#593930"
},
{
"id": 379,
"type": "UNETLoader",
"pos": [
-3880.09423828125,
2205.884033203125
],
"size": [
315,
82
],
"flags": {},
"order": 4,
"mode": 0,
"inputs": [],
"outputs": [
{
"name": "MODEL",
"type": "MODEL",
"links": [
696
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "UNETLoader"
},
"widgets_values": [
"flux1-fill-dev.safetensors",
"fp8_e4m3fn"
]
},
{
"id": 291,
"type": "CLIPTextEncode",
"pos": [
-3441.659423828125,
1648.312255859375
],
"size": [
389.0423278808594,
213.64186096191406
],
"flags": {},
"order": 5,
"mode": 0,
"inputs": [
{
"name": "clip",
"type": "CLIP",
"link": 439
}
],
"outputs": [
{
"name": "CONDITIONING",
"type": "CONDITIONING",
"links": [
446
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "CLIPTextEncode"
},
"widgets_values": [
"Maintain the facial features, A girl is wearing a neat police uniform and sporting a badge. She is smiling with a friendly and confident demeanor. The background is blurred, featuring a cartoon logo."
],
"color": "#232",
"bgcolor": "#353"
},
{
"id": 286,
"type": "KSampler",
"pos": [
-3473.733642578125,
2037.3380126953125
],
"size": [
317.2386474609375,
485.9841613769531
],
"flags": {},
"order": 14,
"mode": 0,
"inputs": [
{
"name": "model",
"type": "MODEL",
"link": 438
},
{
"name": "positive",
"type": "CONDITIONING",
"link": 703
},
{
"name": "negative",
"type": "CONDITIONING",
"link": 704
},
{
"name": "latent_image",
"type": "LATENT",
"link": 705
}
],
"outputs": [
{
"name": "LATENT",
"type": "LATENT",
"links": [
442
],
"slot_index": 0
}
],
"properties": {
"Node name for S&R": "KSampler"
},
"widgets_values": [
41446931693062,
"randomize",
28,
1,
"euler",
"normal",
1
]
},
{
"id": 377,
"type": "ACEPlusLoraProcessor",
"pos": [
-4383.654296875,
1751.32421875
],
"size": [
315,
234
],
"flags": {},
"order": 7,
"mode": 0,
"inputs": [
{
"name": "reference_image",
"type": "IMAGE",
"link": 673,
"shape": 7
},
{
"name": "edit_image",
"type": "IMAGE",
"link": null,
"shape": 7
},
{
"name": "edit_mask",
"type": "MASK",
"link": null,
"shape": 7
}
],
"outputs": [
{
"name": "IMAGE",
"type": "IMAGE",
"links": [
676,
698
],
"slot_index": 0
},
{
"name": "MASK",
"type": "MASK",
"links": [
677,
699
],
"slot_index": 1
},
{
"name": "OUT_H",
"type": "INT",
"links": [
707
],
"slot_index": 2
},
{
"name": "OUT_W",
"type": "INT",
"links": [
706
],
"slot_index": 3
},
{
"name": "SLICE_W",
"type": "INT",
"links": [
680
],
"slot_index": 4
}
],
"properties": {
"Node name for S&R": "ACEPlusLoraProcessor"
},
"widgets_values": [
true,
1024,
1024,
"repainting",
3072
]
}
],
"links": [
[
438,
295,
0,
286,
0,
"MODEL"
],
[
439,
289,
0,
291,
0,
"CLIP"
],
[
440,
289,
0,
292,
0,
"CLIP"
],
[
441,
290,
0,
287,
1,
"VAE"
],
[
442,
286,
0,
287,
0,
"LATENT"
],
[
446,
291,
0,
294,
0,
"CONDITIONING"
],
[
652,
287,
0,
373,
0,
"IMAGE"
],
[
658,
373,
0,
374,
0,
"IMAGE"
],
[
673,
376,
0,
377,
0,
"IMAGE"
],
[
676,
377,
0,
367,
0,
"IMAGE"
],
[
677,
377,
1,
371,
0,
"MASK"
],
[
680,
377,
4,
373,
3,
"INT"
],
[
696,
379,
0,
380,
0,
"MODEL"
],
[
697,
380,
0,
295,
0,
"MODEL"
],
[
698,
377,
0,
381,
3,
"IMAGE"
],
[
699,
377,
1,
381,
4,
"MASK"
],
[
700,
290,
0,
381,
2,
"VAE"
],
[
701,
294,
0,
381,
0,
"CONDITIONING"
],
[
702,
292,
0,
381,
1,
"CONDITIONING"
],
[
703,
381,
0,
286,
1,
"CONDITIONING"
],
[
704,
381,
1,
286,
2,
"CONDITIONING"
],
[
705,
381,
2,
286,
3,
"LATENT"
],
[
706,
377,
3,
373,
1,
"INT"
],
[
707,
377,
2,
373,
2,
"INT"
]
],
"groups": [
{
"id": 20,
"title": "ACE_Plus FFT Workflow",
"bounding": [
-4893.498046875,
1443.87890625,
2592.630126953125,
1121.891845703125
],
"color": "#3f789e",
"font_size": 40,
"flags": {}
}
],
"config": {},
"extra": {
"ds": {
"scale": 0.7627768444385477,
"offset": [
5138.725590463966,
-1489.7090030437157
]
},
"ue_links": [],
"groupNodes": {}
},
"version": 0.4
}