Merge pull request #2 from aszc-dev/dev
Make Core ML models incompatibile with default nodes
This commit is contained in:
@@ -1,2 +1,3 @@
|
||||
playground/
|
||||
experiments/
|
||||
__pycache__/
|
||||
|
||||
@@ -4,7 +4,7 @@
|
||||
|
||||
Welcome! I've developed a set of custom nodes for ComfyUI that allows you to use Core ML models in your ComfyUI
|
||||
workflows.
|
||||
These models are designed to leverage the Apple Neural Engine (ANE) on Apple Silicon (M1/M2) machines,
|
||||
These models are designed to leverage the Apple Neural Engine (ANE) on Apple Silicon (M1/M2) machines,
|
||||
thereby enhancing your workflows and improving performance.
|
||||
|
||||
If you're not sure how to obtain these models, you can download them
|
||||
@@ -24,11 +24,11 @@ To start using custom nodes in your ComfyUI, follow these simple steps:
|
||||
2. Install the dependencies: You'll need to use a package manager like pip to do this.
|
||||
|
||||
That's it! You're now ready to start enhancing your ComfyUI workflows with Core ML models.
|
||||
|
||||
- Check [Installation](#installation) for more details on installation.
|
||||
- Check [How to use](#how-to-use) for more details on how to use the custom nodes.
|
||||
- Check [Example Workflows](#example-workflows) for some example workflows.
|
||||
|
||||
|
||||
## Glossary
|
||||
|
||||
- **Core ML**: A machine learning framework developed by Apple. It's used to run machine learning models on Apple
|
||||
@@ -38,24 +38,27 @@ That's it! You're now ready to start enhancing your ComfyUI workflows with Core
|
||||
- **mlpackage**: A Core ML model packaged in a directory. This is the default format for Core ML models.
|
||||
- **ANE**: Apple Neural Engine. A hardware accelerator for machine learning tasks on Apple devices.
|
||||
- **Compute Unit**: A Core ML option that allows you to specify the hardware on which the model should run.
|
||||
- **CPU_AND_ANE**: A Core ML compute unit option that allows the model to run on both the CPU and ANE. This is the
|
||||
default option.
|
||||
- **CPU_AND_GPU**: A Core ML compute unit option that allows the model to run on both the CPU and GPU.
|
||||
- **CPU_ONLY**: A Core ML compute unit option that allows the model to run on the CPU only.
|
||||
- **ALL**: A Core ML compute unit option that allows the model to run on all available hardware.
|
||||
- **CLIP**: Contrastive Language-Image Pre-training. A model that learns visual concepts from natural language
|
||||
- **CPU_AND_ANE**: A Core ML compute unit option that allows the model to run on both the CPU and ANE. This is the
|
||||
default option.
|
||||
- **CPU_AND_GPU**: A Core ML compute unit option that allows the model to run on both the CPU and GPU.
|
||||
- **CPU_ONLY**: A Core ML compute unit option that allows the model to run on the CPU only.
|
||||
- **ALL**: A Core ML compute unit option that allows the model to run on all available hardware.
|
||||
- **CLIP**: Contrastive Language-Image Pre-training. A model that learns visual concepts from natural language
|
||||
supervision. It's used as a text encoder in Stable Diffusion.
|
||||
- **VAE**: Variational Autoencoder. A model that learns a latent representation of images. It's used as a prior in
|
||||
- **VAE**: Variational Autoencoder. A model that learns a latent representation of images. It's used as a prior in
|
||||
Stable Diffusion.
|
||||
- **Checkpoint**: A file that contains the weights of a model. It's used to load models in Stable Diffusion.
|
||||
|
||||
> [!NOTE]
|
||||
> Note on Compute Units:
|
||||
> For the model to run on the ANE, the model must be converted with the `--attention-implementation SPLIT_EINSUM` option.
|
||||
> For the model to run on the ANE, the model must be converted with the `--attention-implementation SPLIT_EINSUM`
|
||||
> option.
|
||||
> Models converted with `--attention-implementation ORIGINAL` will run on GPU instead of ANE.
|
||||
|
||||
## Features
|
||||
|
||||
These custom nodes come with a host of features, including:
|
||||
|
||||
- Loading Core ML Unet models
|
||||
- Support for ControlNet
|
||||
- Support for ANE (Apple Neural Engine)
|
||||
@@ -63,8 +66,8 @@ These custom nodes come with a host of features, including:
|
||||
- Support for `mlmodelc` and `mlpackage` files
|
||||
|
||||
> [!NOTE]
|
||||
> Please note that using Core ML models can take a bit longer to load initially.
|
||||
> For the best experience, I recommend using the compiled models
|
||||
> Please note that using Core ML models can take a bit longer to load initially.
|
||||
> For the best experience, I recommend using the compiled models
|
||||
> (.mlmodelc files) instead of the .mlpackage files.
|
||||
|
||||
> [!NOTE]
|
||||
@@ -90,73 +93,112 @@ The installation process is simple!
|
||||
## How to use
|
||||
|
||||
Once you've installed the custom nodes, you can start using them in your ComfyUI workflows.
|
||||
To do this, you need to add the nodes to your workflow. You can do this by double-clicking on the workflow canvas and
|
||||
selecting the nodes from the list of available nodes. You can also use the search bar to find the nodes.
|
||||
The list of available nodes is given below. You can also find the nodes in the `CoreML Suite` category.
|
||||
To do this, you need to add the nodes to your workflow. You can do this by right-clicking on the workflow canvas and
|
||||
selecting the nodes from the list of available nodes (the nodes are in the `Core ML Suite` category).
|
||||
You can also double-click the canvas and use the search bar to find the nodes. The list of available nodes is given
|
||||
below.
|
||||
|
||||
### Available Nodes
|
||||
|
||||
#### CoreML UNet Loader (`CoreMLUnetLoader`)
|
||||
#### Core ML UNet Loader (`CoreMLUnetLoader`)
|
||||
|
||||

|
||||

|
||||
|
||||
This node allows you to load a Core ML UNet model and use it in your ComfyUI workflow. Place the converted
|
||||
.mlpackage or .mlmodelc file in ComfyUI's `models/unet` directory and use the node to load the model. The output of the
|
||||
node is a `MODEL` object similar to standard ComfyUI models.
|
||||
node is a `coreml_model` object that can be used with the Core ML Sampler.
|
||||
|
||||
- **Inputs**:
|
||||
- **model_name**: The name of the model to load. This should be the name of the .mlpackage or .mlmodelc file.
|
||||
- **compute_unit**: The hardware on which the model should run. This can be one of the following:
|
||||
- `CPU_AND_ANE`: The model will run on both the CPU and ANE. This is the default option. It works best with
|
||||
models
|
||||
converted with `--attention-implementation SPLIT_EINSUM` or `--attention-implementation SPLIT_EINSUM_V2`.
|
||||
- `CPU_AND_GPU`: The model will run on both the CPU and GPU. It works best with models converted with
|
||||
`--attention-implementation ORIGINAL`.
|
||||
- `CPU_ONLY`: The model will run on the CPU only.
|
||||
- `ALL`: The model will run on all available hardware.
|
||||
- **Outputs**:
|
||||
- **coreml_model**: A Core ML model that can be used with the Core ML Sampler.
|
||||
|
||||
> [!NOTE]
|
||||
> Some models are designed to support ControlNet. If you're using such a model,
|
||||
> make sure to provide a ControlNet input; otherwise, the model will use random noise as ControlNet input.
|
||||
|
||||
#### Core ML Sampler (`CoreMLSampler`)
|
||||
|
||||

|
||||
|
||||
This node allows you to generate images using a Core ML model. The node takes a Core ML model as input and outputs a
|
||||
latent image similar to the latent image output by the KSampler. This means that you can use the
|
||||
resulting latent as you normally would in your workflow.
|
||||
|
||||
- **Inputs**:
|
||||
- **coreml_model**: The Core ML model to use for sampling. This should be the output of the Core ML UNet Loader.
|
||||
- **latent_image** [optional]: The latent image to use for sampling. If provided, should be of the same size as the
|
||||
input of the Core ML model. If not provided, the node will create a latent suitable for the Core ML model used.
|
||||
Useful in img2img workflows.
|
||||
- ... _(the rest of the inputs are the same as the KSampler)_
|
||||
- **Outputs**:
|
||||
- **LATENT**: The latent image output by the Core ML model. This can be decoded using a VAE Decoder or used as input
|
||||
to the next node in your workflow.
|
||||
|
||||
### Example Workflows
|
||||
|
||||
> [!NOTE]
|
||||
> The models used are just an example. Feel free to experiment with different models and see what works best for you.
|
||||
|
||||
|
||||
#### Basic txt2img with Core ML UNet loader
|
||||
This is a basic txt2img workflow that uses the Core ML UNet loader to load a Core ML UNet model. The CLIP and VAE models
|
||||
are loaded using the standard ComfyUI nodes. In the first example, the text encoder (CLIP) and VAE models are loaded
|
||||
separately. In the second example, the text encoder and VAE models are loaded from the checkpoint file. Note that you can use any CLIP or VAE model
|
||||
as long as it's compatible with Stable Diffusion v1.5.
|
||||
|
||||
This is a basic txt2img workflow that uses the Core ML UNet loader to load a model. The CLIP and VAE models
|
||||
are loaded using the standard ComfyUI nodes. In the first example, the text encoder (CLIP) and VAE models are loaded
|
||||
separately. In the second example, the text encoder and VAE models are loaded from the checkpoint file. Note that you
|
||||
can use any CLIP or VAE model as long as it's compatible with Stable Diffusion v1.5.
|
||||
|
||||
1. **Loading text encoder (CLIP) and VAE models separately**
|
||||
- This workflow uses CLIP and VAE models available
|
||||
[here](https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/main/text_encoder/model.safetensors) and
|
||||
[here](https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/main/vae/diffusion_pytorch_model.safetensors).
|
||||
Once downloaded, place the models in the`models/clip` and `models/vae` directories respectively.
|
||||
- The Core ML UNet model is available
|
||||
[here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip).
|
||||
Once downloaded, place the model in the `models/unet` directory.
|
||||

|
||||
|
||||
|
||||
- This workflow uses CLIP and VAE models available
|
||||
[here](https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/main/text_encoder/model.safetensors) and
|
||||
[here](https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/main/vae/diffusion_pytorch_model.safetensors).
|
||||
Once downloaded, place the models in the`models/clip` and `models/vae` directories respectively.
|
||||
- The Core ML UNet model is available
|
||||
[here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip).
|
||||
Once downloaded, place the model in the `models/unet` directory.
|
||||

|
||||
2. **Loading text encoder (CLIP) and VAE models from checkpoint file**
|
||||
- This workflow loads the CLIP and VAE models from the checkpoint file available
|
||||
[here](https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/main/v1-5-pruned-emaonly.safetensors).
|
||||
Once downloaded, place the model in the`models/checkpoints` directory.
|
||||
- The Core ML UNet model is available
|
||||
[here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip).
|
||||
Once downloaded, place the model in the `models/unet` directory.
|
||||

|
||||
- This workflow loads the CLIP and VAE models from the checkpoint file available
|
||||
[here](https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/main/v1-5-pruned-emaonly.safetensors).
|
||||
Once downloaded, place the model in the`models/checkpoints` directory.
|
||||
- The Core ML UNet model is available
|
||||
[here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip).
|
||||
Once downloaded, place the model in the `models/unet` directory.
|
||||

|
||||
|
||||
#### ControlNet with Core ML UNet loader
|
||||
|
||||
#### ControlNet with Core ML UNet loader
|
||||
|
||||
(Coming soon)
|
||||
This workflow uses the Core ML UNet loader to load a Core ML UNet model that supports ControlNet. The ControlNet is
|
||||
being loaded using the standard ComfyUI nodes. Please refer to
|
||||
the [basic txt2img workflow](#basic-txt2img-with-core-ml-unet-loader) for more details on how to load the CLIP and VAE
|
||||
models.
|
||||
The ControlNet model used in this workflow is available
|
||||
[here](https://huggingface.co/lllyasviel/ControlNet-v1-1/blob/main/control_v11p_sd15_lineart.pth).
|
||||
Once downloaded, place the model in the `models/controlnet` directory.
|
||||

|
||||
|
||||
## Limitations
|
||||
|
||||
- Core ML models are fixed in terms of their inputs and outputs.
|
||||
This means you'll need to use latent images of the same size as the input of the model (512x512 is the default for SD1.5).
|
||||
However, you can convert the model to a different input size using tools available
|
||||
in the [apple/ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion) repository.
|
||||
- Core ML models are fixed in terms of their inputs and outputs.
|
||||
This means you'll need to use latent images of the same size as the input of the model (512x512 is the default for
|
||||
SD1.5).
|
||||
However, you can convert the model to a different input size using tools available
|
||||
in the [apple/ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion) repository.
|
||||
- For now, only Stable Diffusion v1.5 is supported.
|
||||
- LoRA is not supported yet.
|
||||
|
||||
[^1]: Unless [EnumeratedShapes](https://apple.github.io/coremltools/docs-guides/source/flexible-inputs.html#select-from-predetermined-shapes)
|
||||
[^1]:
|
||||
Unless [EnumeratedShapes](https://apple.github.io/coremltools/docs-guides/source/flexible-inputs.html#select-from-predetermined-shapes)
|
||||
is used during conversion. Needs more testing.
|
||||
|
||||
## Support
|
||||
|
||||
I'm here to help! If you have any questions or suggestions, don't hesitate to open an issue and I'll do my best
|
||||
I'm here to help! If you have any questions or suggestions, don't hesitate to open an issue and I'll do my best
|
||||
to assist you.
|
||||
|
||||
+9
-2
@@ -1,8 +1,15 @@
|
||||
from .loaders import CoreMLLoaderUNet
|
||||
import os
|
||||
import sys
|
||||
|
||||
sys.path.append(os.path.dirname(__file__))
|
||||
|
||||
from coreml_suite import CoreMLLoaderUNet, CoreMLSampler
|
||||
|
||||
NODE_CLASS_MAPPINGS = {
|
||||
"CoreMLUNetLoader": CoreMLLoaderUNet,
|
||||
"CoreMLSampler": CoreMLSampler,
|
||||
}
|
||||
NODE_DISPLAY_NAME_MAPPINGS = {
|
||||
"CoreMLUNetLoader": "Load Core ML UNet"
|
||||
"CoreMLUNetLoader": "Load Core ML UNet",
|
||||
"CoreMLSampler": "Core ML Sampler",
|
||||
}
|
||||
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 83 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 44 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 449 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 446 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 469 KiB |
Binary file not shown.
|
After Width: | Height: | Size: 46 KiB |
@@ -0,0 +1,4 @@
|
||||
from coreml_suite.loaders import CoreMLLoaderUNet
|
||||
from coreml_suite.samplers import CoreMLSampler
|
||||
|
||||
__all__ = ["CoreMLLoaderUNet", "CoreMLSampler"]
|
||||
@@ -0,0 +1,84 @@
|
||||
import os.path
|
||||
|
||||
from coremltools import ComputeUnit
|
||||
from python_coreml_stable_diffusion.coreml_model import CoreMLModel
|
||||
|
||||
import folder_paths
|
||||
|
||||
from coreml_suite.logger import logger
|
||||
|
||||
|
||||
class CoreMLLoader:
|
||||
PACKAGE_DIRNAME = ""
|
||||
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
return {
|
||||
"required": {
|
||||
"coreml_name": (list(s.coreml_filenames().keys()),),
|
||||
"compute_unit": (
|
||||
[
|
||||
ComputeUnit.CPU_AND_NE.name,
|
||||
ComputeUnit.CPU_AND_GPU.name,
|
||||
ComputeUnit.ALL.name,
|
||||
ComputeUnit.CPU_ONLY.name,
|
||||
],
|
||||
),
|
||||
}
|
||||
}
|
||||
|
||||
FUNCTION = "load"
|
||||
CATEGORY = "Core ML Suite"
|
||||
|
||||
@classmethod
|
||||
def coreml_filenames(cls):
|
||||
extensions = (".mlmodelc", ".mlpackage")
|
||||
all_paths = folder_paths.get_filename_list_(cls.PACKAGE_DIRNAME)[1]
|
||||
coreml_paths = folder_paths.filter_files_extensions(all_paths, extensions)
|
||||
|
||||
return {os.path.split(p)[-1]: p for p in coreml_paths}
|
||||
|
||||
def load(self, coreml_name, compute_unit):
|
||||
logger.info(f"Loading {coreml_name} to {compute_unit}")
|
||||
|
||||
coreml_path = self.coreml_filenames()[coreml_name]
|
||||
|
||||
sources = "compiled" if coreml_name.endswith(".mlmodelc") else "packages"
|
||||
|
||||
return self._load(coreml_path, compute_unit, sources)
|
||||
|
||||
def _load(self, coreml_path, compute_unit, sources):
|
||||
return (CoreMLModel(coreml_path, compute_unit, sources),)
|
||||
|
||||
|
||||
class CoreMLLoaderCkpt(CoreMLLoader):
|
||||
PACKAGE_DIRNAME = "checkpoints"
|
||||
RETURN_TYPES = ("MODEL", "CLIP", "VAE")
|
||||
|
||||
def load(self, coreml_name, compute_unit):
|
||||
# TODO: Implement this
|
||||
pass
|
||||
|
||||
|
||||
class CoreMLLoaderTextEncoder(CoreMLLoader):
|
||||
PACKAGE_DIRNAME = "clip"
|
||||
RETURN_TYPES = ("CLIP",)
|
||||
|
||||
def load(self, coreml_name, compute_unit):
|
||||
# TODO: Implement this
|
||||
pass
|
||||
|
||||
|
||||
class CoreMLLoaderUNet(CoreMLLoader):
|
||||
PACKAGE_DIRNAME = "unet"
|
||||
RETURN_TYPES = ("COREML_UNET",)
|
||||
RETURN_NAMES = ("coreml_model",)
|
||||
|
||||
|
||||
class CoreMLLoaderVAE(CoreMLLoader):
|
||||
PACKAGE_DIRNAME = "vae"
|
||||
RETURN_TYPES = ("VAE",)
|
||||
|
||||
def load(self, coreml_name, compute_unit):
|
||||
# TODO: Implement this
|
||||
pass
|
||||
@@ -0,0 +1,63 @@
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
|
||||
from comfy import supported_models_base
|
||||
from comfy.latent_formats import SD15
|
||||
from comfy.model_base import BaseModel
|
||||
|
||||
from coreml_suite.utils import expand_inputs, extract_residual_kwargs
|
||||
|
||||
|
||||
def get_model_config():
|
||||
# TODO: This is a dummy model config, but it should be enough to
|
||||
# get the model to load - implement a proper model config
|
||||
model_config = supported_models_base.BASE({})
|
||||
model_config.latent_format = SD15()
|
||||
model_config.unet_config = {
|
||||
"disable_unet_model_creation": True,
|
||||
"num_res_blocks": 2,
|
||||
"attention_resolutions": [1, 2, 4],
|
||||
"channel_mult": [1, 2, 4, 4],
|
||||
"transformer_depth": [1, 1, 1, 0],
|
||||
}
|
||||
return model_config
|
||||
|
||||
|
||||
class CoreMLModelWrapper(BaseModel):
|
||||
def __init__(self, model_config, coreml_model):
|
||||
super().__init__(model_config)
|
||||
self.diffusion_model = coreml_model
|
||||
|
||||
def apply_model(
|
||||
self,
|
||||
x,
|
||||
t,
|
||||
c_concat=None,
|
||||
c_crossattn=None,
|
||||
c_adm=None,
|
||||
control=None,
|
||||
transformer_options={},
|
||||
):
|
||||
sample = x.cpu().numpy().astype(np.float16)
|
||||
|
||||
context = c_crossattn.cpu().numpy().astype(np.float16)
|
||||
context = context.transpose(0, 2, 1)[:, :, None, :]
|
||||
|
||||
t = t.cpu().numpy().astype(np.float16)
|
||||
|
||||
model_input_kwargs = {
|
||||
"sample": sample,
|
||||
"encoder_hidden_states": context,
|
||||
"timestep": t,
|
||||
}
|
||||
residual_kwargs = extract_residual_kwargs(self.diffusion_model, control)
|
||||
model_input_kwargs |= residual_kwargs
|
||||
model_input_kwargs = expand_inputs(model_input_kwargs)
|
||||
|
||||
np_out = self.diffusion_model(**model_input_kwargs)["noise_pred"]
|
||||
return torch.from_numpy(np_out).to(x.device)
|
||||
|
||||
def get_dtype(self):
|
||||
# Hardcoding torch-compatible dtype (used for memory allocation)
|
||||
return torch.float16
|
||||
@@ -0,0 +1,74 @@
|
||||
import torch
|
||||
from torchvision.transforms.functional import resize
|
||||
|
||||
from comfy.model_management import get_torch_device
|
||||
from comfy.model_patcher import ModelPatcher
|
||||
from coreml_suite.logger import logger
|
||||
from nodes import KSampler
|
||||
|
||||
from coreml_suite.models import CoreMLModelWrapper, get_model_config
|
||||
|
||||
|
||||
def reshape_latent_image(latent_image, target_shape):
|
||||
if latent_image is None:
|
||||
logger.warning("No latent image provided, using zeros.")
|
||||
return {"samples": torch.zeros(target_shape)}
|
||||
|
||||
if latent_image["samples"].shape == target_shape:
|
||||
return latent_image
|
||||
|
||||
logger.warning(
|
||||
"Latent image shape does not match model input shape,"
|
||||
" resizing to match models expected input shape."
|
||||
)
|
||||
resized = resize(latent_image["samples"], target_shape[-2:])
|
||||
return {"samples": resized}
|
||||
|
||||
|
||||
class CoreMLSampler(KSampler):
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
old_required = KSampler.INPUT_TYPES()["required"].copy()
|
||||
old_required.pop("model")
|
||||
old_required.pop("latent_image")
|
||||
new_required = {"coreml_model": ("COREML_UNET",)}
|
||||
return {
|
||||
"required": new_required | old_required,
|
||||
"optional": {"latent_image": ("LATENT",)},
|
||||
}
|
||||
|
||||
CATEGORY = "Core ML Suite"
|
||||
|
||||
def sample(
|
||||
self,
|
||||
coreml_model,
|
||||
seed,
|
||||
steps,
|
||||
cfg,
|
||||
sampler_name,
|
||||
scheduler,
|
||||
positive,
|
||||
negative,
|
||||
latent_image=None,
|
||||
denoise=1.0,
|
||||
):
|
||||
sample_shape = coreml_model.expected_inputs["sample"]["shape"]
|
||||
latent_image = reshape_latent_image(latent_image, sample_shape)
|
||||
latent_image["samples"] = latent_image["samples"][0:1]
|
||||
|
||||
model_config = get_model_config()
|
||||
wrapped_model = CoreMLModelWrapper(model_config, coreml_model)
|
||||
model = ModelPatcher(wrapped_model, get_torch_device(), None)
|
||||
|
||||
return super().sample(
|
||||
model,
|
||||
seed,
|
||||
steps,
|
||||
cfg,
|
||||
sampler_name,
|
||||
scheduler,
|
||||
positive,
|
||||
negative,
|
||||
latent_image,
|
||||
denoise,
|
||||
)
|
||||
@@ -3,7 +3,7 @@ from itertools import chain
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
from .logger import logger
|
||||
from coreml_suite.logger import logger
|
||||
|
||||
|
||||
def expand_inputs(inputs):
|
||||
@@ -21,16 +21,14 @@ def expand_inputs(inputs):
|
||||
|
||||
|
||||
def extract_residual_kwargs(model, control):
|
||||
if ("additional_residual_0" not in model.expected_inputs.keys()):
|
||||
if "additional_residual_0" not in model.expected_inputs.keys():
|
||||
return {}
|
||||
if control is None:
|
||||
return no_control(model)
|
||||
|
||||
residual_kwargs = {
|
||||
"additional_residual_{}".format(i): r.cpu().numpy().astype(
|
||||
np.float16)
|
||||
for i, r in
|
||||
enumerate(chain(control["output"], control["middle"]))
|
||||
"additional_residual_{}".format(i): r.cpu().numpy().astype(np.float16)
|
||||
for i, r in enumerate(chain(control["output"], control["middle"]))
|
||||
}
|
||||
return residual_kwargs
|
||||
|
||||
@@ -40,16 +38,25 @@ def no_control(model):
|
||||
# 0.18215 is the latent scale factor (IDK, it kinda works)
|
||||
# TODO: Find a better way to do this or tweak the values
|
||||
|
||||
logger.warning("No ControlNet input, despite the model supports it. "
|
||||
"Using random noise as ControlNet residuals. "
|
||||
"For better results, please use a ControlNet or a model "
|
||||
"that does not support ControlNet.")
|
||||
residuals_names = [name for name in model.expected_inputs.keys()
|
||||
if name.startswith("additional_residual")]
|
||||
logger.warning(
|
||||
"No ControlNet input, despite the model supports it. "
|
||||
"Using random noise as ControlNet residuals. "
|
||||
"For better results, please use a ControlNet or a model "
|
||||
"that does not support ControlNet."
|
||||
)
|
||||
residuals_names = [
|
||||
name
|
||||
for name in model.expected_inputs.keys()
|
||||
if name.startswith("additional_residual")
|
||||
]
|
||||
residual_kwargs = {
|
||||
"additional_residual_{}".format(i): 0.18215 * torch.randn(
|
||||
*model.expected_inputs["additional_residual_{}".format(i)][
|
||||
"shape"]).cpu().numpy().astype(dtype=np.float16)
|
||||
"additional_residual_{}".format(i): 0.18215
|
||||
* torch.randn(
|
||||
*model.expected_inputs["additional_residual_{}".format(i)]["shape"]
|
||||
)
|
||||
.cpu()
|
||||
.numpy()
|
||||
.astype(dtype=np.float16)
|
||||
for i in range(len(residuals_names))
|
||||
}
|
||||
return residual_kwargs
|
||||
-103
@@ -1,103 +0,0 @@
|
||||
from coremltools import ComputeUnit
|
||||
from python_coreml_stable_diffusion.coreml_model import CoreMLModel
|
||||
|
||||
import folder_paths
|
||||
from comfy import supported_models_base, model_management
|
||||
from comfy.latent_formats import SD15
|
||||
from comfy.model_patcher import ModelPatcher
|
||||
from .logger import logger
|
||||
from .model import CoreMLModelWrapper
|
||||
|
||||
|
||||
class CoreMLLoader:
|
||||
PACKAGE_DIRNAME = ""
|
||||
|
||||
@classmethod
|
||||
def INPUT_TYPES(s):
|
||||
return {
|
||||
"required": {
|
||||
"coreml_name": (list(s.coreml_filenames().keys()),),
|
||||
"compute_unit": ([
|
||||
ComputeUnit.CPU_AND_NE.name,
|
||||
ComputeUnit.CPU_AND_GPU.name,
|
||||
ComputeUnit.ALL.name,
|
||||
ComputeUnit.CPU_ONLY.name,
|
||||
],)
|
||||
}
|
||||
}
|
||||
|
||||
RETURN_TYPES = ("COREML_MODEL",)
|
||||
FUNCTION = "load"
|
||||
CATEGORY = "CoreML Suite"
|
||||
|
||||
@classmethod
|
||||
def coreml_filenames(cls):
|
||||
return {
|
||||
p.split('/')[-1]:
|
||||
p for p in
|
||||
folder_paths.get_filename_list_(cls.PACKAGE_DIRNAME)[1]
|
||||
if p.endswith((".mlpackage", ".mlmodelc"))
|
||||
}
|
||||
|
||||
def load(self, coreml_name, compute_unit):
|
||||
logger.info(f"Loading {coreml_name}")
|
||||
|
||||
coreml_path = self.coreml_filenames()[coreml_name]
|
||||
|
||||
sources = "compiled" if coreml_name.endswith(
|
||||
".mlmodelc") else "packages"
|
||||
|
||||
return self._load(coreml_path, compute_unit, sources)
|
||||
|
||||
def _load(self, coreml_path, compute_unit, sources):
|
||||
return (CoreMLModel(coreml_path, compute_unit, sources),)
|
||||
|
||||
|
||||
class CoreMLLoaderCkpt(CoreMLLoader):
|
||||
PACKAGE_DIRNAME = "checkpoints"
|
||||
RETURN_TYPES = ("MODEL", "CLIP", "VAE")
|
||||
|
||||
def load(self, coreml_name, compute_unit):
|
||||
# TODO: Implement this
|
||||
pass
|
||||
|
||||
|
||||
class CoreMLLoaderTextEncoder(CoreMLLoader):
|
||||
PACKAGE_DIRNAME = "clip"
|
||||
RETURN_TYPES = ("CLIP",)
|
||||
|
||||
def load(self, coreml_name, compute_unit):
|
||||
# TODO: Implement this
|
||||
pass
|
||||
|
||||
|
||||
class CoreMLLoaderUNet(CoreMLLoader):
|
||||
PACKAGE_DIRNAME = "unet"
|
||||
RETURN_TYPES = ("MODEL",)
|
||||
|
||||
def _load(self, coreml_path, compute_unit, sources):
|
||||
# TODO: This is a dummy model config, but it should be enough to
|
||||
# get the model to load - implement a proper model config
|
||||
model_config = supported_models_base.BASE({})
|
||||
model_config.latent_format = SD15()
|
||||
model_config.unet_config = {
|
||||
"disable_unet_model_creation": True,
|
||||
"num_res_blocks": 2,
|
||||
"attention_resolutions": [1, 2, 4],
|
||||
"channel_mult": [1, 2, 4, 4],
|
||||
"transformer_depth": [1, 1, 1, 0],
|
||||
}
|
||||
coreml_model = CoreMLModelWrapper(model_config, coreml_path,
|
||||
compute_unit, sources)
|
||||
|
||||
return (ModelPatcher(coreml_model, model_management.get_torch_device(),
|
||||
None),)
|
||||
|
||||
|
||||
class CoreMLLoaderVAE(CoreMLLoader):
|
||||
PACKAGE_DIRNAME = "vae"
|
||||
RETURN_TYPES = ("VAE",)
|
||||
|
||||
def load(self, coreml_name, compute_unit):
|
||||
# TODO: Implement this
|
||||
pass
|
||||
@@ -1,42 +0,0 @@
|
||||
import numpy as np
|
||||
import torch
|
||||
|
||||
from python_coreml_stable_diffusion.coreml_model import CoreMLModel
|
||||
|
||||
from comfy.model_base import BaseModel
|
||||
|
||||
from .utils import expand_inputs, extract_residual_kwargs
|
||||
|
||||
|
||||
class CoreMLModelWrapper(BaseModel):
|
||||
def __init__(self, model_config, mlpackage_path, compute_unit,
|
||||
sources="packages"):
|
||||
super().__init__(model_config)
|
||||
self.diffusion_model = CoreMLModel(mlpackage_path, compute_unit,
|
||||
sources)
|
||||
|
||||
def apply_model(self, x, t, c_concat=None, c_crossattn=None, c_adm=None,
|
||||
control=None, transformer_options={}):
|
||||
sample = x.cpu().numpy().astype(np.float16)
|
||||
|
||||
context = c_crossattn.cpu().numpy().astype(np.float16)
|
||||
context = context.transpose(0, 2, 1)[:, :, None, :]
|
||||
|
||||
t = t.cpu().numpy().astype(np.float16)
|
||||
|
||||
model_input_kwargs = {
|
||||
"sample": sample,
|
||||
"encoder_hidden_states": context,
|
||||
"timestep": t,
|
||||
}
|
||||
residual_kwargs = extract_residual_kwargs(self.diffusion_model,
|
||||
control)
|
||||
model_input_kwargs |= residual_kwargs
|
||||
model_input_kwargs = expand_inputs(model_input_kwargs)
|
||||
|
||||
np_out = self.diffusion_model(**model_input_kwargs)["noise_pred"]
|
||||
return torch.from_numpy(np_out).to(x.device)
|
||||
|
||||
def get_dtype(self):
|
||||
# Hardcoding torch-compatible dtype (used for memory allocation)
|
||||
return torch.float16
|
||||
@@ -0,0 +1,19 @@
|
||||
import pytest
|
||||
|
||||
import torch
|
||||
|
||||
from coreml_suite.samplers import reshape_latent_image
|
||||
|
||||
|
||||
def test_fix_latents_no_latent_image():
|
||||
reshaped = reshape_latent_image(None, (2, 4, 64, 64))
|
||||
assert reshaped["samples"].shape == (2, 4, 64, 64)
|
||||
|
||||
|
||||
@pytest.mark.parametrize(
|
||||
"latent_shape", [(2, 4, 64, 64), (2, 4, 128, 128), (2, 4, 32, 32), (2, 4, 128, 64)]
|
||||
)
|
||||
def test_reshape_latents(latent_shape):
|
||||
latent_image = {"samples": torch.zeros(latent_shape)}
|
||||
reshaped = reshape_latent_image(latent_image, (2, 4, 64, 64))
|
||||
assert reshaped["samples"].shape == (2, 4, 64, 64)
|
||||
Reference in New Issue
Block a user