Merge pull request #2 from aszc-dev/dev

Make Core ML models incompatibile with default nodes
This commit is contained in:
aszc
2023-10-30 00:51:12 +01:00
committed by GitHub
19 changed files with 367 additions and 211 deletions
+1
View File
@@ -1,2 +1,3 @@
playground/
experiments/
__pycache__/
+91 -49
View File
@@ -4,7 +4,7 @@
Welcome! I've developed a set of custom nodes for ComfyUI that allows you to use Core ML models in your ComfyUI
workflows.
These models are designed to leverage the Apple Neural Engine (ANE) on Apple Silicon (M1/M2) machines,
These models are designed to leverage the Apple Neural Engine (ANE) on Apple Silicon (M1/M2) machines,
thereby enhancing your workflows and improving performance.
If you're not sure how to obtain these models, you can download them
@@ -24,11 +24,11 @@ To start using custom nodes in your ComfyUI, follow these simple steps:
2. Install the dependencies: You'll need to use a package manager like pip to do this.
That's it! You're now ready to start enhancing your ComfyUI workflows with Core ML models.
- Check [Installation](#installation) for more details on installation.
- Check [How to use](#how-to-use) for more details on how to use the custom nodes.
- Check [Example Workflows](#example-workflows) for some example workflows.
## Glossary
- **Core ML**: A machine learning framework developed by Apple. It's used to run machine learning models on Apple
@@ -38,24 +38,27 @@ That's it! You're now ready to start enhancing your ComfyUI workflows with Core
- **mlpackage**: A Core ML model packaged in a directory. This is the default format for Core ML models.
- **ANE**: Apple Neural Engine. A hardware accelerator for machine learning tasks on Apple devices.
- **Compute Unit**: A Core ML option that allows you to specify the hardware on which the model should run.
- **CPU_AND_ANE**: A Core ML compute unit option that allows the model to run on both the CPU and ANE. This is the
default option.
- **CPU_AND_GPU**: A Core ML compute unit option that allows the model to run on both the CPU and GPU.
- **CPU_ONLY**: A Core ML compute unit option that allows the model to run on the CPU only.
- **ALL**: A Core ML compute unit option that allows the model to run on all available hardware.
- **CLIP**: Contrastive Language-Image Pre-training. A model that learns visual concepts from natural language
- **CPU_AND_ANE**: A Core ML compute unit option that allows the model to run on both the CPU and ANE. This is the
default option.
- **CPU_AND_GPU**: A Core ML compute unit option that allows the model to run on both the CPU and GPU.
- **CPU_ONLY**: A Core ML compute unit option that allows the model to run on the CPU only.
- **ALL**: A Core ML compute unit option that allows the model to run on all available hardware.
- **CLIP**: Contrastive Language-Image Pre-training. A model that learns visual concepts from natural language
supervision. It's used as a text encoder in Stable Diffusion.
- **VAE**: Variational Autoencoder. A model that learns a latent representation of images. It's used as a prior in
- **VAE**: Variational Autoencoder. A model that learns a latent representation of images. It's used as a prior in
Stable Diffusion.
- **Checkpoint**: A file that contains the weights of a model. It's used to load models in Stable Diffusion.
> [!NOTE]
> Note on Compute Units:
> For the model to run on the ANE, the model must be converted with the `--attention-implementation SPLIT_EINSUM` option.
> For the model to run on the ANE, the model must be converted with the `--attention-implementation SPLIT_EINSUM`
> option.
> Models converted with `--attention-implementation ORIGINAL` will run on GPU instead of ANE.
## Features
These custom nodes come with a host of features, including:
- Loading Core ML Unet models
- Support for ControlNet
- Support for ANE (Apple Neural Engine)
@@ -63,8 +66,8 @@ These custom nodes come with a host of features, including:
- Support for `mlmodelc` and `mlpackage` files
> [!NOTE]
> Please note that using Core ML models can take a bit longer to load initially.
> For the best experience, I recommend using the compiled models
> Please note that using Core ML models can take a bit longer to load initially.
> For the best experience, I recommend using the compiled models
> (.mlmodelc files) instead of the .mlpackage files.
> [!NOTE]
@@ -90,73 +93,112 @@ The installation process is simple!
## How to use
Once you've installed the custom nodes, you can start using them in your ComfyUI workflows.
To do this, you need to add the nodes to your workflow. You can do this by double-clicking on the workflow canvas and
selecting the nodes from the list of available nodes. You can also use the search bar to find the nodes.
The list of available nodes is given below. You can also find the nodes in the `CoreML Suite` category.
To do this, you need to add the nodes to your workflow. You can do this by right-clicking on the workflow canvas and
selecting the nodes from the list of available nodes (the nodes are in the `Core ML Suite` category).
You can also double-click the canvas and use the search bar to find the nodes. The list of available nodes is given
below.
### Available Nodes
#### CoreML UNet Loader (`CoreMLUnetLoader`)
#### Core ML UNet Loader (`CoreMLUnetLoader`)
![CoreMLUnetLoader](https://github.com/aszc-dev/ComfyUI-CoreMLSuite/assets/24932801/2bd10f73-4103-4860-894c-b6a6e56c6546)
![CoreMLUnetLoader](./assets/unet_loader.png?raw=true)
This node allows you to load a Core ML UNet model and use it in your ComfyUI workflow. Place the converted
.mlpackage or .mlmodelc file in ComfyUI's `models/unet` directory and use the node to load the model. The output of the
node is a `MODEL` object similar to standard ComfyUI models.
node is a `coreml_model` object that can be used with the Core ML Sampler.
- **Inputs**:
- **model_name**: The name of the model to load. This should be the name of the .mlpackage or .mlmodelc file.
- **compute_unit**: The hardware on which the model should run. This can be one of the following:
- `CPU_AND_ANE`: The model will run on both the CPU and ANE. This is the default option. It works best with
models
converted with `--attention-implementation SPLIT_EINSUM` or `--attention-implementation SPLIT_EINSUM_V2`.
- `CPU_AND_GPU`: The model will run on both the CPU and GPU. It works best with models converted with
`--attention-implementation ORIGINAL`.
- `CPU_ONLY`: The model will run on the CPU only.
- `ALL`: The model will run on all available hardware.
- **Outputs**:
- **coreml_model**: A Core ML model that can be used with the Core ML Sampler.
> [!NOTE]
> Some models are designed to support ControlNet. If you're using such a model,
> make sure to provide a ControlNet input; otherwise, the model will use random noise as ControlNet input.
#### Core ML Sampler (`CoreMLSampler`)
![CoreMLSampler](./assets/sampler.png?raw=true)
This node allows you to generate images using a Core ML model. The node takes a Core ML model as input and outputs a
latent image similar to the latent image output by the KSampler. This means that you can use the
resulting latent as you normally would in your workflow.
- **Inputs**:
- **coreml_model**: The Core ML model to use for sampling. This should be the output of the Core ML UNet Loader.
- **latent_image** [optional]: The latent image to use for sampling. If provided, should be of the same size as the
input of the Core ML model. If not provided, the node will create a latent suitable for the Core ML model used.
Useful in img2img workflows.
- ... _(the rest of the inputs are the same as the KSampler)_
- **Outputs**:
- **LATENT**: The latent image output by the Core ML model. This can be decoded using a VAE Decoder or used as input
to the next node in your workflow.
### Example Workflows
> [!NOTE]
> The models used are just an example. Feel free to experiment with different models and see what works best for you.
#### Basic txt2img with Core ML UNet loader
This is a basic txt2img workflow that uses the Core ML UNet loader to load a Core ML UNet model. The CLIP and VAE models
are loaded using the standard ComfyUI nodes. In the first example, the text encoder (CLIP) and VAE models are loaded
separately. In the second example, the text encoder and VAE models are loaded from the checkpoint file. Note that you can use any CLIP or VAE model
as long as it's compatible with Stable Diffusion v1.5.
This is a basic txt2img workflow that uses the Core ML UNet loader to load a model. The CLIP and VAE models
are loaded using the standard ComfyUI nodes. In the first example, the text encoder (CLIP) and VAE models are loaded
separately. In the second example, the text encoder and VAE models are loaded from the checkpoint file. Note that you
can use any CLIP or VAE model as long as it's compatible with Stable Diffusion v1.5.
1. **Loading text encoder (CLIP) and VAE models separately**
- This workflow uses CLIP and VAE models available
[here](https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/main/text_encoder/model.safetensors) and
[here](https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/main/vae/diffusion_pytorch_model.safetensors).
Once downloaded, place the models in the`models/clip` and `models/vae` directories respectively.
- The Core ML UNet model is available
[here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip).
Once downloaded, place the model in the `models/unet` directory.
![coreml-unet+clip+vae](https://github.com/aszc-dev/ComfyUI-CoreMLSuite/assets/24932801/ff7b8d75-37ea-4da9-a258-829edd6eb1b7)
- This workflow uses CLIP and VAE models available
[here](https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/main/text_encoder/model.safetensors) and
[here](https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/main/vae/diffusion_pytorch_model.safetensors).
Once downloaded, place the models in the`models/clip` and `models/vae` directories respectively.
- The Core ML UNet model is available
[here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip).
Once downloaded, place the model in the `models/unet` directory.
![coreml-unet+clip+vae](./assets/unet+sampler+clip+vae.png?raw=true)
2. **Loading text encoder (CLIP) and VAE models from checkpoint file**
- This workflow loads the CLIP and VAE models from the checkpoint file available
[here](https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/main/v1-5-pruned-emaonly.safetensors).
Once downloaded, place the model in the`models/checkpoints` directory.
- The Core ML UNet model is available
[here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip).
Once downloaded, place the model in the `models/unet` directory.
![coreml-unet+checkpoint](https://github.com/aszc-dev/ComfyUI-CoreMLSuite/assets/24932801/0c8b4f65-9bde-4b0d-936b-5bb27023d2ce)
- This workflow loads the CLIP and VAE models from the checkpoint file available
[here](https://huggingface.co/runwayml/stable-diffusion-v1-5/blob/main/v1-5-pruned-emaonly.safetensors).
Once downloaded, place the model in the`models/checkpoints` directory.
- The Core ML UNet model is available
[here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip).
Once downloaded, place the model in the `models/unet` directory.
![coreml-unet+checkpoint](./assets/unet+sampler+checkpoint.png?raw=true)
#### ControlNet with Core ML UNet loader
#### ControlNet with Core ML UNet loader
(Coming soon)
This workflow uses the Core ML UNet loader to load a Core ML UNet model that supports ControlNet. The ControlNet is
being loaded using the standard ComfyUI nodes. Please refer to
the [basic txt2img workflow](#basic-txt2img-with-core-ml-unet-loader) for more details on how to load the CLIP and VAE
models.
The ControlNet model used in this workflow is available
[here](https://huggingface.co/lllyasviel/ControlNet-v1-1/blob/main/control_v11p_sd15_lineart.pth).
Once downloaded, place the model in the `models/controlnet` directory.
![coreml-unet+controlnet](./assets/unet+sampler+controlnet.png?raw=true)
## Limitations
- Core ML models are fixed in terms of their inputs and outputs.
This means you'll need to use latent images of the same size as the input of the model (512x512 is the default for SD1.5).
However, you can convert the model to a different input size using tools available
in the [apple/ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion) repository.
- Core ML models are fixed in terms of their inputs and outputs.
This means you'll need to use latent images of the same size as the input of the model (512x512 is the default for
SD1.5).
However, you can convert the model to a different input size using tools available
in the [apple/ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion) repository.
- For now, only Stable Diffusion v1.5 is supported.
- LoRA is not supported yet.
[^1]: Unless [EnumeratedShapes](https://apple.github.io/coremltools/docs-guides/source/flexible-inputs.html#select-from-predetermined-shapes)
[^1]:
Unless [EnumeratedShapes](https://apple.github.io/coremltools/docs-guides/source/flexible-inputs.html#select-from-predetermined-shapes)
is used during conversion. Needs more testing.
## Support
I'm here to help! If you have any questions or suggestions, don't hesitate to open an issue and I'll do my best
I'm here to help! If you have any questions or suggestions, don't hesitate to open an issue and I'll do my best
to assist you.
+9 -2
View File
@@ -1,8 +1,15 @@
from .loaders import CoreMLLoaderUNet
import os
import sys
sys.path.append(os.path.dirname(__file__))
from coreml_suite import CoreMLLoaderUNet, CoreMLSampler
NODE_CLASS_MAPPINGS = {
"CoreMLUNetLoader": CoreMLLoaderUNet,
"CoreMLSampler": CoreMLSampler,
}
NODE_DISPLAY_NAME_MAPPINGS = {
"CoreMLUNetLoader": "Load Core ML UNet"
"CoreMLUNetLoader": "Load Core ML UNet",
"CoreMLSampler": "Core ML Sampler",
}
Binary file not shown.

After

Width:  |  Height:  |  Size: 83 KiB

BIN
View File
Binary file not shown.

After

Width:  |  Height:  |  Size: 44 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 449 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 446 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 469 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 46 KiB

+4
View File
@@ -0,0 +1,4 @@
from coreml_suite.loaders import CoreMLLoaderUNet
from coreml_suite.samplers import CoreMLSampler
__all__ = ["CoreMLLoaderUNet", "CoreMLSampler"]
+84
View File
@@ -0,0 +1,84 @@
import os.path
from coremltools import ComputeUnit
from python_coreml_stable_diffusion.coreml_model import CoreMLModel
import folder_paths
from coreml_suite.logger import logger
class CoreMLLoader:
PACKAGE_DIRNAME = ""
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"coreml_name": (list(s.coreml_filenames().keys()),),
"compute_unit": (
[
ComputeUnit.CPU_AND_NE.name,
ComputeUnit.CPU_AND_GPU.name,
ComputeUnit.ALL.name,
ComputeUnit.CPU_ONLY.name,
],
),
}
}
FUNCTION = "load"
CATEGORY = "Core ML Suite"
@classmethod
def coreml_filenames(cls):
extensions = (".mlmodelc", ".mlpackage")
all_paths = folder_paths.get_filename_list_(cls.PACKAGE_DIRNAME)[1]
coreml_paths = folder_paths.filter_files_extensions(all_paths, extensions)
return {os.path.split(p)[-1]: p for p in coreml_paths}
def load(self, coreml_name, compute_unit):
logger.info(f"Loading {coreml_name} to {compute_unit}")
coreml_path = self.coreml_filenames()[coreml_name]
sources = "compiled" if coreml_name.endswith(".mlmodelc") else "packages"
return self._load(coreml_path, compute_unit, sources)
def _load(self, coreml_path, compute_unit, sources):
return (CoreMLModel(coreml_path, compute_unit, sources),)
class CoreMLLoaderCkpt(CoreMLLoader):
PACKAGE_DIRNAME = "checkpoints"
RETURN_TYPES = ("MODEL", "CLIP", "VAE")
def load(self, coreml_name, compute_unit):
# TODO: Implement this
pass
class CoreMLLoaderTextEncoder(CoreMLLoader):
PACKAGE_DIRNAME = "clip"
RETURN_TYPES = ("CLIP",)
def load(self, coreml_name, compute_unit):
# TODO: Implement this
pass
class CoreMLLoaderUNet(CoreMLLoader):
PACKAGE_DIRNAME = "unet"
RETURN_TYPES = ("COREML_UNET",)
RETURN_NAMES = ("coreml_model",)
class CoreMLLoaderVAE(CoreMLLoader):
PACKAGE_DIRNAME = "vae"
RETURN_TYPES = ("VAE",)
def load(self, coreml_name, compute_unit):
# TODO: Implement this
pass
+63
View File
@@ -0,0 +1,63 @@
import numpy as np
import torch
from comfy import supported_models_base
from comfy.latent_formats import SD15
from comfy.model_base import BaseModel
from coreml_suite.utils import expand_inputs, extract_residual_kwargs
def get_model_config():
# TODO: This is a dummy model config, but it should be enough to
# get the model to load - implement a proper model config
model_config = supported_models_base.BASE({})
model_config.latent_format = SD15()
model_config.unet_config = {
"disable_unet_model_creation": True,
"num_res_blocks": 2,
"attention_resolutions": [1, 2, 4],
"channel_mult": [1, 2, 4, 4],
"transformer_depth": [1, 1, 1, 0],
}
return model_config
class CoreMLModelWrapper(BaseModel):
def __init__(self, model_config, coreml_model):
super().__init__(model_config)
self.diffusion_model = coreml_model
def apply_model(
self,
x,
t,
c_concat=None,
c_crossattn=None,
c_adm=None,
control=None,
transformer_options={},
):
sample = x.cpu().numpy().astype(np.float16)
context = c_crossattn.cpu().numpy().astype(np.float16)
context = context.transpose(0, 2, 1)[:, :, None, :]
t = t.cpu().numpy().astype(np.float16)
model_input_kwargs = {
"sample": sample,
"encoder_hidden_states": context,
"timestep": t,
}
residual_kwargs = extract_residual_kwargs(self.diffusion_model, control)
model_input_kwargs |= residual_kwargs
model_input_kwargs = expand_inputs(model_input_kwargs)
np_out = self.diffusion_model(**model_input_kwargs)["noise_pred"]
return torch.from_numpy(np_out).to(x.device)
def get_dtype(self):
# Hardcoding torch-compatible dtype (used for memory allocation)
return torch.float16
+74
View File
@@ -0,0 +1,74 @@
import torch
from torchvision.transforms.functional import resize
from comfy.model_management import get_torch_device
from comfy.model_patcher import ModelPatcher
from coreml_suite.logger import logger
from nodes import KSampler
from coreml_suite.models import CoreMLModelWrapper, get_model_config
def reshape_latent_image(latent_image, target_shape):
if latent_image is None:
logger.warning("No latent image provided, using zeros.")
return {"samples": torch.zeros(target_shape)}
if latent_image["samples"].shape == target_shape:
return latent_image
logger.warning(
"Latent image shape does not match model input shape,"
" resizing to match models expected input shape."
)
resized = resize(latent_image["samples"], target_shape[-2:])
return {"samples": resized}
class CoreMLSampler(KSampler):
@classmethod
def INPUT_TYPES(s):
old_required = KSampler.INPUT_TYPES()["required"].copy()
old_required.pop("model")
old_required.pop("latent_image")
new_required = {"coreml_model": ("COREML_UNET",)}
return {
"required": new_required | old_required,
"optional": {"latent_image": ("LATENT",)},
}
CATEGORY = "Core ML Suite"
def sample(
self,
coreml_model,
seed,
steps,
cfg,
sampler_name,
scheduler,
positive,
negative,
latent_image=None,
denoise=1.0,
):
sample_shape = coreml_model.expected_inputs["sample"]["shape"]
latent_image = reshape_latent_image(latent_image, sample_shape)
latent_image["samples"] = latent_image["samples"][0:1]
model_config = get_model_config()
wrapped_model = CoreMLModelWrapper(model_config, coreml_model)
model = ModelPatcher(wrapped_model, get_torch_device(), None)
return super().sample(
model,
seed,
steps,
cfg,
sampler_name,
scheduler,
positive,
negative,
latent_image,
denoise,
)
+22 -15
View File
@@ -3,7 +3,7 @@ from itertools import chain
import numpy as np
import torch
from .logger import logger
from coreml_suite.logger import logger
def expand_inputs(inputs):
@@ -21,16 +21,14 @@ def expand_inputs(inputs):
def extract_residual_kwargs(model, control):
if ("additional_residual_0" not in model.expected_inputs.keys()):
if "additional_residual_0" not in model.expected_inputs.keys():
return {}
if control is None:
return no_control(model)
residual_kwargs = {
"additional_residual_{}".format(i): r.cpu().numpy().astype(
np.float16)
for i, r in
enumerate(chain(control["output"], control["middle"]))
"additional_residual_{}".format(i): r.cpu().numpy().astype(np.float16)
for i, r in enumerate(chain(control["output"], control["middle"]))
}
return residual_kwargs
@@ -40,16 +38,25 @@ def no_control(model):
# 0.18215 is the latent scale factor (IDK, it kinda works)
# TODO: Find a better way to do this or tweak the values
logger.warning("No ControlNet input, despite the model supports it. "
"Using random noise as ControlNet residuals. "
"For better results, please use a ControlNet or a model "
"that does not support ControlNet.")
residuals_names = [name for name in model.expected_inputs.keys()
if name.startswith("additional_residual")]
logger.warning(
"No ControlNet input, despite the model supports it. "
"Using random noise as ControlNet residuals. "
"For better results, please use a ControlNet or a model "
"that does not support ControlNet."
)
residuals_names = [
name
for name in model.expected_inputs.keys()
if name.startswith("additional_residual")
]
residual_kwargs = {
"additional_residual_{}".format(i): 0.18215 * torch.randn(
*model.expected_inputs["additional_residual_{}".format(i)][
"shape"]).cpu().numpy().astype(dtype=np.float16)
"additional_residual_{}".format(i): 0.18215
* torch.randn(
*model.expected_inputs["additional_residual_{}".format(i)]["shape"]
)
.cpu()
.numpy()
.astype(dtype=np.float16)
for i in range(len(residuals_names))
}
return residual_kwargs
-103
View File
@@ -1,103 +0,0 @@
from coremltools import ComputeUnit
from python_coreml_stable_diffusion.coreml_model import CoreMLModel
import folder_paths
from comfy import supported_models_base, model_management
from comfy.latent_formats import SD15
from comfy.model_patcher import ModelPatcher
from .logger import logger
from .model import CoreMLModelWrapper
class CoreMLLoader:
PACKAGE_DIRNAME = ""
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"coreml_name": (list(s.coreml_filenames().keys()),),
"compute_unit": ([
ComputeUnit.CPU_AND_NE.name,
ComputeUnit.CPU_AND_GPU.name,
ComputeUnit.ALL.name,
ComputeUnit.CPU_ONLY.name,
],)
}
}
RETURN_TYPES = ("COREML_MODEL",)
FUNCTION = "load"
CATEGORY = "CoreML Suite"
@classmethod
def coreml_filenames(cls):
return {
p.split('/')[-1]:
p for p in
folder_paths.get_filename_list_(cls.PACKAGE_DIRNAME)[1]
if p.endswith((".mlpackage", ".mlmodelc"))
}
def load(self, coreml_name, compute_unit):
logger.info(f"Loading {coreml_name}")
coreml_path = self.coreml_filenames()[coreml_name]
sources = "compiled" if coreml_name.endswith(
".mlmodelc") else "packages"
return self._load(coreml_path, compute_unit, sources)
def _load(self, coreml_path, compute_unit, sources):
return (CoreMLModel(coreml_path, compute_unit, sources),)
class CoreMLLoaderCkpt(CoreMLLoader):
PACKAGE_DIRNAME = "checkpoints"
RETURN_TYPES = ("MODEL", "CLIP", "VAE")
def load(self, coreml_name, compute_unit):
# TODO: Implement this
pass
class CoreMLLoaderTextEncoder(CoreMLLoader):
PACKAGE_DIRNAME = "clip"
RETURN_TYPES = ("CLIP",)
def load(self, coreml_name, compute_unit):
# TODO: Implement this
pass
class CoreMLLoaderUNet(CoreMLLoader):
PACKAGE_DIRNAME = "unet"
RETURN_TYPES = ("MODEL",)
def _load(self, coreml_path, compute_unit, sources):
# TODO: This is a dummy model config, but it should be enough to
# get the model to load - implement a proper model config
model_config = supported_models_base.BASE({})
model_config.latent_format = SD15()
model_config.unet_config = {
"disable_unet_model_creation": True,
"num_res_blocks": 2,
"attention_resolutions": [1, 2, 4],
"channel_mult": [1, 2, 4, 4],
"transformer_depth": [1, 1, 1, 0],
}
coreml_model = CoreMLModelWrapper(model_config, coreml_path,
compute_unit, sources)
return (ModelPatcher(coreml_model, model_management.get_torch_device(),
None),)
class CoreMLLoaderVAE(CoreMLLoader):
PACKAGE_DIRNAME = "vae"
RETURN_TYPES = ("VAE",)
def load(self, coreml_name, compute_unit):
# TODO: Implement this
pass
-42
View File
@@ -1,42 +0,0 @@
import numpy as np
import torch
from python_coreml_stable_diffusion.coreml_model import CoreMLModel
from comfy.model_base import BaseModel
from .utils import expand_inputs, extract_residual_kwargs
class CoreMLModelWrapper(BaseModel):
def __init__(self, model_config, mlpackage_path, compute_unit,
sources="packages"):
super().__init__(model_config)
self.diffusion_model = CoreMLModel(mlpackage_path, compute_unit,
sources)
def apply_model(self, x, t, c_concat=None, c_crossattn=None, c_adm=None,
control=None, transformer_options={}):
sample = x.cpu().numpy().astype(np.float16)
context = c_crossattn.cpu().numpy().astype(np.float16)
context = context.transpose(0, 2, 1)[:, :, None, :]
t = t.cpu().numpy().astype(np.float16)
model_input_kwargs = {
"sample": sample,
"encoder_hidden_states": context,
"timestep": t,
}
residual_kwargs = extract_residual_kwargs(self.diffusion_model,
control)
model_input_kwargs |= residual_kwargs
model_input_kwargs = expand_inputs(model_input_kwargs)
np_out = self.diffusion_model(**model_input_kwargs)["noise_pred"]
return torch.from_numpy(np_out).to(x.device)
def get_dtype(self):
# Hardcoding torch-compatible dtype (used for memory allocation)
return torch.float16
View File
+19
View File
@@ -0,0 +1,19 @@
import pytest
import torch
from coreml_suite.samplers import reshape_latent_image
def test_fix_latents_no_latent_image():
reshaped = reshape_latent_image(None, (2, 4, 64, 64))
assert reshaped["samples"].shape == (2, 4, 64, 64)
@pytest.mark.parametrize(
"latent_shape", [(2, 4, 64, 64), (2, 4, 128, 128), (2, 4, 32, 32), (2, 4, 128, 64)]
)
def test_reshape_latents(latent_shape):
latent_image = {"samples": torch.zeros(latent_shape)}
reshaped = reshape_latent_image(latent_image, (2, 4, 64, 64))
assert reshaped["samples"].shape == (2, 4, 64, 64)