diff --git a/README.md b/README.md index 92f8948..d0b97cf 100644 --- a/README.md +++ b/README.md @@ -143,6 +143,76 @@ resulting latent as you normally would in your workflow. - **LATENT**: The latent image output by the Core ML model. This can be decoded using a VAE Decoder or used as input to the next node in your workflow. +#### Checkpoint Converter + +![CoreMLConverter](./assets/checkpoint_converter.png?raw=true) + +You can use this node to convert any **SD1.5** based checkpoint to a Core ML model. The converted model is stored in the +`models/unet` directory and can be used with the `Core ML UNet Loader`. The conversion parameters are encoded in +the node name, so if the model already exists, the node will not convert it again. + +- **Inputs**: + - **ckpt_name**: The name of the checkpoint to convert. This should be the name of the checkpoint file stored in the + `models/checkpoints` directory. + - **height**: The desired height of the image generated by the model. The default is 512. Must be a multiple of 8. + - **width**: The desired width of the image generated by the model. The default is 512. Must be a multiple of 8. + - **batch_size**: The batch size of generated images. If you're planning to generate batches of images, you can try + increasing this value to speed up the generation process. The default is 1. + - **attention_implementation**: The attention implementation used when converting the model. Choose SPLIT_EINSUM or + SPLIT_EINSUM_V2 for better ANE support. Choose ORIGINAL for better GPU support. + - **compute_unit**: The hardware on which the model should run. This is used only when loading the model and doesn't + affect the conversion process. + - **controlnet_support**: For the model to support ControlNet, it must be converted with this option set to True. + The + default is False. + - **lora_params** [optional]: Optional LoRA names and weights. If provided, the model will be converted with LoRA(s) + baked in. More on loading LoRAs below. +- **Outputs**: + - **coreml_model**: The converted Core ML model that can be used with Core ML Sampler. + +> [!NOTE] +> Some models use a custom config .yaml file. If you're using such a model, you'll need to place the config file in the +> `models/configs` directory. The config file should be named the same as the checkpoint file. For example, if the +> checkpoint file is named `juggernaut_aftermath.safetensors`, the config file should be +> named `juggernaut_aftermath.yaml`. +> The config file will be automatically loaded during conversion. + +> [!NOTE] +> For now, the converter relies heavilty on the model name to determine the conversion parameters. This means that if +> you change the model name, the node will convert the model again. Other than that, if you find the name too long or +> confusing, you can change it to anything you want. + +#### LoRA Loader + +![LoRALoader](./assets/lora_loader.png?raw=true) + +This node allows you to load LoRAs and bake them into a model. Since this is a workaround (as model weights can't be +modified +after conversion), there are a few caveats to keep in mind: + +- The LoRA weights and _strength_model_ parameter are baked into the model. This means that you can't change them + after conversion. This also means that you need to convert the model again if you want to change the LoRA weights. +- Loading LoRA affects CLIP, which is not a part of Core ML workflow, so you'll need to load CLIP separately, + either using `CLIPLoader` or `CheckpointLoaderSimple`. (See [example workflows](#example-workflows) for more details.) +- After conversion, if you want to load the model using `CoreMLUnetLoader`, you'll need to apply the same LoRAs to + CLIP manually. (See [example workflows](#example-workflows) for more details.) +- The LoRA names are encoded in the model name. This means that if you change the name of the LoRA file, + you'll need to change the model name as well, or the node will convert the model again. (Model strength is not + encoded, so if you want to change it, you'll need to delete the converted model manually) +- _strength_clip_ parameter only affects the CLIP model and is not baked into the converted model. This means that + you can change it after conversion. + +- **Inputs**: + - **lora_name**: The name of the LoRA to load. + - **strength_model**: The strength of the LoRA model. + - **strength_clip**: The strength of the LoRA CLIP. + - **lora_params** [optional]: Optional output from other LoRA Loaders. + - **clip**: The CLIP model to use with the LoRA. This can be either output of the + `CLIPLoader`/`CheckpointLoaderSimple` or other LoRA Loaders. +- **Outputs**: + - **lora_params**: The LoRA parameters that can be passed to the Core ML Converter or other LoRA Loaders. + - **CLIP**: The CLIP model with LoRA applied. + #### LCM Converter ![LCMConverter](./assets/lcm_converter.png?raw=true) @@ -225,6 +295,41 @@ The ControlNet model used in this workflow is available Once downloaded, place the model in the `models/controlnet` directory. ![coreml-unet+controlnet](./assets/unet+sampler+controlnet.png?raw=true) +#### Checkpoint conversion + +This workflow uses the Checkpoint Converter to convert the checkpoint file. See +[Checkpoint Converter](#checkpoint-converter) description for more details. + +![checkpoint-converter](./assets/basic_conversion.png?raw=true) + +#### Checkpoint conversion with LoRA + +This workflow uses the Checkpoint Converter to convert the checkpoint file with LoRA. See +[LoRA Loader](#lora-loader) description to read more about the caveats of using LoRA. + +![checkpoint-converter+lora](./assets/conversion+lora.png?raw=true) + +#### LCM LoRA conversion + +Please note that you can use multiple LoRAs with the same model. To do this, you'll need to use multiple LoRA Loaders. +> [!IMPORTANT] +> In this example, the model is passed through the adapter and `ModelSamplingDiscrete` nodes to a standard ComfyUI's +> KSampler (not Core ML Sampler). ModelSamplingDiscrete needs to be used to sample models with LCM LoRAs properly. + +![multiple-loras](./assets/conversion+lcm_lora.png?raw=true) + +#### Loader with LoRAs + +This workflow uses the Core ML UNet Loader to load a model with LoRAs. The CLIP must be loaded separately and passed +through the same LoRA nodes as during conversion. See [LoRA Loader](#lora-loader) description to read more about the +caveats of using LoRA. Since _lora_name_ and _strength_model_ are baked into the model, it is not necessary to pass +them as inputs to the loader. +> [!IMPORTANT] +> In this example, the model is passed through the adapter and `ModelSamplingDiscrete` nodes to a standard ComfyUI's +> KSampler (not Core ML Sampler). ModelSamplingDiscrete needs to be used to sample models with LCM LoRAs properly. + +![loader+lora](./assets/loader+lcm_lora.png?raw=true) + #### LCM conversion with ControlNet This workflow uses LCM converter to @@ -240,7 +345,6 @@ model to Core ML. The converted model can then be used with or without ControlNe However, you can convert the model to a different input size using tools available in the [apple/ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion) repository. - For now, only Stable Diffusion v1.5 is supported. -- LoRA is not supported yet. [^1]: Unless [EnumeratedShapes](https://apple.github.io/coremltools/docs-guides/source/flexible-inputs.html#select-from-predetermined-shapes) diff --git a/assets/checkpoint_converter.png b/assets/checkpoint_converter.png new file mode 100644 index 0000000..d7d52f1 Binary files /dev/null and b/assets/checkpoint_converter.png differ diff --git a/assets/lora_loader.png b/assets/lora_loader.png new file mode 100644 index 0000000..7c5ce60 Binary files /dev/null and b/assets/lora_loader.png differ