Compare commits
2
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
1008f144ad | ||
|
|
8f94f0eea5 |
@@ -3,4 +3,6 @@ __pycache__/
|
||||
models/
|
||||
.venv/
|
||||
test_results/
|
||||
*.log
|
||||
.DS_Store
|
||||
.claude/
|
||||
|
||||
@@ -1,618 +0,0 @@
|
||||
# ComfyUI-CoreMLSuite — Converter Extraction Spec for Claude Code
|
||||
|
||||
> **Companion to `MODERNIZATION_SPEC.md`.** That spec hardens the repo and (Phase 3)
|
||||
> splits the *inference* math from the framework. **This** spec splits the *conversion*
|
||||
> path (`safetensors → CoreML`) out into a standalone, `comfy`-free, pip-installable
|
||||
> package that CoreMLSuite then depends on — and that other projects (incl. on-device
|
||||
> iOS tooling) can reuse.
|
||||
>
|
||||
> **Same discipline as the modernization spec:** safety-net first, behavior-preserving
|
||||
> until told otherwise, one phase = one branch = one PR, `STOP — VALIDATE` gate between
|
||||
> every phase, golden-latent as the regression anchor. `[M2]` = needs macOS/Apple Silicon;
|
||||
> `[M2-ANE]` = needs the Neural Engine. Everything else must run on plain Linux/CI.
|
||||
|
||||
---
|
||||
|
||||
## 0. How to work (read first — non-negotiable)
|
||||
|
||||
1. **Behavior-preserving until Phase E6.** Phases E1–E5 must not change image output, node
|
||||
names, `INPUT_TYPES` field names, or `NODE_CLASS_MAPPINGS` keys. The node graph is the
|
||||
public contract; saved user-workflow JSON breaks if these change.
|
||||
2. **The conversion package produces an artifact and stops there.** Its job ends at a written
|
||||
`.mlpackage` / `.mlmodelc` on disk. It must NOT import `comfy`, `folder_paths`, or
|
||||
`comfy_extras`, and must NOT know ComfyUI's `models/unet` layout. Paths are *inputs*.
|
||||
3. **The runtime loader stays in the suite.** The loader is the **local** `coreml_suite.coreml_model.CoreMLModel`
|
||||
— a thin wrapper over `coremltools.models.MLModel` (NOT Apple's
|
||||
`python_coreml_stable_diffusion.coreml_model.CoreMLModel`, which is no longer used; see #58).
|
||||
It *runs* a compiled model in Python — a desktop/Python inference concern, not a conversion
|
||||
concern. It is NOT moved into the package. (On iOS the `.mlmodelc` is loaded natively; the
|
||||
package's output is the deliverable, not a Python runner.)
|
||||
4. **Decouple in-repo before splitting repos.** Phases E1–E4 create the package *inside this
|
||||
repo* and prove equivalence. The physical second-repo split is Phase E5, only after the
|
||||
golden latent is proven identical. Do not create a second repository before Gate E4 passes.
|
||||
5. **Reuse the existing regression anchor.** The golden latent / PSNR anchor from
|
||||
`MODERNIZATION_SPEC.md` Phase 2 is the cross-cutting proof for every gate here. If it is not
|
||||
yet captured, capture it first (it is a prerequisite for E2 onward).
|
||||
6. **No new runtime dependencies** without flagging in the gate report (name, why, license, size).
|
||||
7. **A failing gate means stop and report**, not work around into the next phase.
|
||||
8. **Tooling is `uv`, not bare `pip`/`venv`.** Every environment/install/lock step uses the
|
||||
project's `uv` toolchain: `uv venv`, `uv pip install`, `uv pip install -e .`, `uv lock`,
|
||||
`uv run pytest`, `uv export`/`uv pip freeze` for baselines. Where this spec says "fresh venv",
|
||||
read "`uv venv` + `uv pip install`". Reserve `uv pip` (not `pip`) inside that venv too.
|
||||
9. **The package is the single source of truth for *what is possible*; the node is a thin,
|
||||
discovery-driven frontend.** See the "Interface contract" pillar below — this is the
|
||||
maintainer's hard requirement and it overrides the earlier (now-rescinded) "freeze the
|
||||
dropdown list" instruction.
|
||||
|
||||
---
|
||||
|
||||
## Interface contract (the maintainer's hard requirement) — read before any phase
|
||||
|
||||
Two coupled guarantees must hold once the package is split out:
|
||||
|
||||
**(A) Updating the converter must NOT require updating CoreMLSuite.**
|
||||
This is satisfied by treating the package's public surface as a versioned contract:
|
||||
- `convert(...)` and `compile_model(...)` are **keyword-only with defaults** for everything
|
||||
past the genuinely-required positionals (`ckpt_path`, `model_version`, `out_path`). New
|
||||
capabilities are added as new keyword args with defaults, so an old Suite's call still
|
||||
validates against a newer package. **Never** reorder or rename existing parameters.
|
||||
- `compose_out_name` (the `.mlpackage` filename = the cache key) **moves into the package** and
|
||||
is versioned with it. The Suite must not carry its own copy; if the package changes the naming
|
||||
scheme that is a **major** bump (old cached artifacts stop resolving).
|
||||
|
||||
**(B) CoreMLSuite must be able to list *new* conversion types WITHOUT a Suite code change or
|
||||
version bump.** Today the node hardcodes its dropdowns:
|
||||
```python
|
||||
"model_version": ([ModelVersion.SD15.name, ModelVersion.SDXL.name],), # hand-typed, also INCOMPLETE (no LCM / SDXL_REFINER)
|
||||
"attention_implementation": (list(ATTENTION_IMPLEMENTATIONS),), # from coreml_suite.attention
|
||||
"quantize_nbits": (list(QUANT_NBITS_VALUES), {"default": "none"}), # from coreml_suite.core.naming
|
||||
```
|
||||
These are replaced by **runtime discovery calls into the package**, evaluated inside
|
||||
`INPUT_TYPES` (ComfyUI re-evaluates `INPUT_TYPES` on every plugin load):
|
||||
```python
|
||||
import coreml_diffusion
|
||||
"model_version": (coreml_diffusion.list_model_versions(),),
|
||||
"attention_implementation": (coreml_diffusion.list_attention_impls(),),
|
||||
"quantize_nbits": (coreml_diffusion.list_quant_modes(), {"default": "none"}),
|
||||
```
|
||||
Effect: `uv pip install -U coreml_diffusion` + ComfyUI restart surfaces any newly-added type in the old
|
||||
plugin's dropdown — **no Suite edit, no Suite version bump.** This is the requirement.
|
||||
|
||||
**The cost, stated honestly (accept this trade-off explicitly at Gate E0):**
|
||||
- The Suite becomes a "dumb" frontend; the package is the sole authority on what conversions
|
||||
exist. The Suite can no longer guarantee its saved workflows are valid against *arbitrary*
|
||||
future package versions.
|
||||
- Therefore the package's discovery identifiers (`ModelVersion` values, attn-impl strings, quant
|
||||
modes) are an **ADDITIVE-ONLY contract**: the package may *add* identifiers freely (minor bump,
|
||||
no Suite change); **removing or renaming an identifier is a breaking change requiring a MAJOR
|
||||
bump and a migration note**, because a saved workflow JSON references these strings verbatim.
|
||||
Without this rule, "no version bump" silently becomes "randomly broken workflows."
|
||||
- `INPUT_TYPES` must **fail soft** when the package is missing/old: wrap the discovery calls so a
|
||||
missing `coreml_diffusion` (or an old one lacking a `list_*` function) yields a sane fallback list and a
|
||||
logged warning, instead of the node failing to register and disappearing from the menu.
|
||||
|
||||
**Discovery API the package must expose (stable names):**
|
||||
```python
|
||||
coreml_diffusion.list_model_versions() -> list[str] # VERIFIED ones only, e.g. ["SD15","SDXL"] today (.name — see seam.md)
|
||||
coreml_diffusion.list_attention_impls() -> list[str] # ["SPLIT_EINSUM","SPLIT_EINSUM_V2","ORIGINAL"]
|
||||
coreml_diffusion.list_quant_modes() -> list[str] # ["none","8","6","4"]
|
||||
coreml_diffusion.CONTRACT_VERSION: str # bump rules above; Suite may log/compare it
|
||||
```
|
||||
These return the *display strings already used today*, so existing workflows keep validating.
|
||||
|
||||
**Verification status is a PACKAGE property, not a node hardcode (maintainer's intent).**
|
||||
The Suite wants to expose *every model the converter can verifiably convert*. Today `lcm` and
|
||||
`sdxl_refiner` are absent from the converter node not because the Suite chooses to hide them, but
|
||||
because they lack a full golden/PSNR verification. So the gating lives in the package as a status:
|
||||
```python
|
||||
from enum import Enum
|
||||
class Status(Enum):
|
||||
VERIFIED = "verified" # has a golden anchor + passing [M2-ANE] check
|
||||
EXPERIMENTAL = "experimental" # convertible but not yet anchored/verified
|
||||
|
||||
# internal registry, single source of truth.
|
||||
# KEY by ModelVersion enum MEMBER so list_* can emit .name. Keying by the lowercase
|
||||
# .value string returns ["sd15",...], which the node reverses via ModelVersion[...] -> KeyError.
|
||||
_MODEL_STATUS = {ModelVersion.SD15: Status.VERIFIED, ModelVersion.SDXL: Status.VERIFIED,
|
||||
ModelVersion.SDXL_REFINER: Status.EXPERIMENTAL, ModelVersion.LCM: Status.EXPERIMENTAL}
|
||||
|
||||
def list_model_versions(include_experimental: bool = False) -> list[str]:
|
||||
return [v.name for v, s in _MODEL_STATUS.items() # .name -> "SD15","SDXL"; node reverses with ModelVersion[...]
|
||||
if s is Status.VERIFIED or (include_experimental and s is Status.EXPERIMENTAL)]
|
||||
```
|
||||
Consequence: **promoting a model to VERIFIED in the package expands the Suite's dropdown with no
|
||||
Suite change and no Suite bump** — exactly the requirement. The act of verification (E-LCM
|
||||
produces an LCM golden anchor; same later for refiner) is what flips the status. The Suite's
|
||||
converter node calls `list_model_versions()` (verified-only); a power-user/CLI path may pass
|
||||
`include_experimental=True`. Promotion VERIFIED-from-EXPERIMENTAL is additive (minor bump);
|
||||
demotion or removal is breaking (major bump + note).
|
||||
|
||||
---
|
||||
|
||||
## Naming & layout (chosen — frozen at Gate E0)
|
||||
|
||||
**Distribution name (PyPI):** `coreml-diffusion`. **Import name (Python):** `coreml_diffusion`.
|
||||
(PyPI normalizes `-`/`_`; the distribution uses the hyphen, the importable module the underscore.)
|
||||
Availability checked: both `coreml-diffusion` and the near variants were free on PyPI at E0.
|
||||
|
||||
**Why this name (the positioning it encodes):** the project's niche is *diffusion models on Apple
|
||||
Neural Engine via CoreML, inside ComfyUI and on-device* — **not** Stable Diffusion specifically.
|
||||
`sd*` was rejected because it falsely narrows scope to SD; `coreml-diffusion` keeps `coreml` on the
|
||||
front for discoverability while `diffusion` honestly states the scope (SD/SDXL/LCM today, Flux and
|
||||
other diffusion architectures later) **without** promising arbitrary non-diffusion torch models,
|
||||
whose tracing/shape/sample-input pipeline differs. The name must not be re-narrowed to SD in
|
||||
future docs. ANE is the *differentiator* (documented in the README), but `coreml` was chosen over
|
||||
`ane` in the name for search discoverability per maintainer decision.
|
||||
|
||||
Target package layout (framework-free — zero `comfy` imports):
|
||||
|
||||
```
|
||||
coreml_diffusion/
|
||||
__init__.py # public API surface (see "Public API" below)
|
||||
model_version.py # ModelVersion enum — the SINGLE source of truth, no comfy
|
||||
attention.py # ATTENTION_IMPLEMENTATIONS tuple (from coreml_suite/attention.py) + apply_attention_implementation
|
||||
pipeline.py # get_pipeline (from_single_file), get_unet (cml UNet from ref unet)
|
||||
unet.py # UNet2DConditionModelLCM (moved from coreml_suite/lcm/unet.py)
|
||||
inputs.py # get_sample_input, lcm_inputs, sdxl_inputs,
|
||||
# get_encoder_hidden_states_shape, get_coreml_inputs, get_inputs_spec
|
||||
controlnet.py # add_cnet_support (conversion-side residual SHAPE calc only)
|
||||
convert.py # convert_unet, convert (orchestration), convert_to_coreml, load_coreml_model
|
||||
compile.py # compile_coreml_model
|
||||
quantize.py # (Phase E6 / MODERNIZATION Phase 6 lands here) palettization 4/6/8-bit
|
||||
cli.py # console entry point: `coreml-diffusion convert ...`
|
||||
pyproject.toml # standalone packaging (at E5)
|
||||
```
|
||||
|
||||
What stays in `coreml_suite/` (the ComfyUI side, thinned):
|
||||
- `nodes.py` — still owns **name-encoding** (`out_name` construction), path resolution via
|
||||
`folder_paths`, the node `INPUT_TYPES`/mappings, and wrapping the result in `CoreMLModel`.
|
||||
- `models.py`, `latents.py`, `controlnet.py` (inference parts), `lcm/utils.py`, `config.py`
|
||||
(inference config build) — untouched by this spec except the import-source of `ModelVersion`.
|
||||
|
||||
### Public API (the contract `coreml_diffusion` exposes)
|
||||
```python
|
||||
from coreml_diffusion import ModelVersion, convert, compile_model, compose_out_name
|
||||
from coreml_diffusion import list_model_versions, list_attention_impls, list_quant_modes, CONTRACT_VERSION
|
||||
|
||||
# Mirror the CURRENT converter.py signature, made keyword-only past the required positionals
|
||||
# and with paths/device injected (no folder_paths, no comfy.model_management):
|
||||
# convert(ckpt_path, model_version, out_path, *,
|
||||
# batch_size=1, sample_size=(64, 64), controlnet_support=False,
|
||||
# lora_weights=None, attn_impl=list_attention_impls()[0], config_path=None,
|
||||
# quantize_nbits="none", device=None) -> None # side effect: writes out_path
|
||||
# (current convert() returns None and writes via convert_unet → coreml_unet.save; keep that,
|
||||
# or change to `return out_path` as a deliberate, documented improvement — pick one at E0.)
|
||||
# compile_model(src_path, out_dir, final_name) -> str # returns compiled .mlmodelc path
|
||||
```
|
||||
Note: `convert` takes an **explicit `out_path`** — no `folder_paths`. `device` is injected
|
||||
(defaults to torch's default device). `compose_out_name` lives here (cache-key contract) and the
|
||||
node imports it from the package. The `list_*` discovery functions back the node's dropdowns.
|
||||
|
||||
---
|
||||
|
||||
## The import chains to cut (root cause inventory) — REVISED against current code
|
||||
|
||||
> **State note (verified):** the code moved on since the original draft. Several chains are
|
||||
> already cut. Re-verify each line by `grep` before acting; do not assume the original draft.
|
||||
|
||||
**Already done (verify, then skip):**
|
||||
- ✅ `converter.py` already imports `from coreml_suite.model_version import ModelVersion`, and
|
||||
`model_version.py` is **clean** (`from enum import Enum` only — zero comfy). The old
|
||||
"converter → config → comfy" chain is **already broken**. `config.py` still imports comfy, but
|
||||
it is **inference-side** (`get_model_config` via `supported_models_base`/`latent_formats`) —
|
||||
*not* on the conversion path. Do **not** treat `config.py` as a converter dependency.
|
||||
- ✅ `converter.py` now uses `diffusers.UNet2DConditionModel.from_single_file` and a local
|
||||
`CoreMLUNetWrapper` (in `coreml_suite/conversion/unet.py`) — it is **no longer** importing the
|
||||
Apple `python_coreml_stable_diffusion.unet.UNet2DConditionModel*` internals on the main path.
|
||||
A `coreml_suite/conversion/` subpackage already exists (`attention`, `shapes`, `trace`, `unet`).
|
||||
- ✅ Name-encoding already extracted to `coreml_suite/core/naming.py` (`compose_out_name`,
|
||||
`lora_names_from_params`, `ATTN_SUFFIX`, `QUANT_NBITS_VALUES`) **with characterization tests**
|
||||
(`tests/unit/test_characterization_out_name.py`). The pure-naming split is done.
|
||||
- ✅ Quantization is **already implemented** in `converter.py` (`quantize_nbits`, k-means
|
||||
`palettize_weights`) and surfaced as an optional node input. Phase E6 is therefore *move*, not
|
||||
*build* (see revised E6).
|
||||
|
||||
**Still to cut (the real remaining work):**
|
||||
1. `coreml_suite/converter.py::get_out_path` → `from folder_paths import get_folder_paths`.
|
||||
Main converter still reaches into ComfyUI's model dir. **Cut: `out_path` is an injected arg;
|
||||
`folder_paths` resolution moves up into the node** (the node already computes `out_name`).
|
||||
2. `coreml_suite/lcm/converter.py` → still has its **own** `from folder_paths import
|
||||
get_folder_paths` (`get_out_path`) and (per original draft) `comfy.model_management`. Verify
|
||||
the current LCM file and cut both: inject `out_path` and `device`.
|
||||
3. Global mutation of the attention impl: confirm where it now lives. Main path appears to route
|
||||
through `coreml_suite/conversion/attention.apply_attention_implementation` (cleaner than the
|
||||
old global), but `lcm/converter.py` may still set a module global at import. **Ensure the
|
||||
package sets attention per-call, never at import time.**
|
||||
4. **Duplication LCM vs main:** `lcm/converter.py` still carries its own copies of
|
||||
`convert_to_coreml`, `load_coreml_model`, `get_out_path`, `get_sample_input` (the LCM variant
|
||||
takes a `scheduler` arg), and hardcodes `SimianLuo/LCM_Dreamshaper_v7`. **Dedupe into the
|
||||
single `coreml_diffusion` implementation;** the HF-hardcode consolidation is the *behavior-changing*
|
||||
part → deferred to optional **E-LCM**, not E1–E5.
|
||||
5. **`compose_out_name` ownership:** currently in `coreml_suite/core/naming.py` and called by the
|
||||
node. Per the Interface-contract pillar it must **move into the package** (it is the cache-key
|
||||
contract) and the node must import it from `coreml_diffusion`, not keep a copy.
|
||||
|
||||
---
|
||||
|
||||
## Phase E0 — Seam decision & inventory (no code change)
|
||||
|
||||
**Objective:** lock the cut line, the interface contract, and naming so later phases don't drift.
|
||||
|
||||
### Tasks
|
||||
1. Produce `docs/extraction/seam.md`: a table of every symbol in `converter.py`,
|
||||
`lcm/converter.py`, `lcm/unet.py`, **plus the already-extracted `conversion/` subpackage
|
||||
(`attention`, `shapes`, `trace`, `unet`) and `core/naming.py`**, classified
|
||||
**CONVERSION → coreml_diffusion** vs **STAYS (comfy/node)**. Note which are already framework-free.
|
||||
2. ~~Confirm the current `python_coreml_stable_diffusion` footprint.~~ **DONE (seam.md §6):
|
||||
footprint is ZERO** — no runtime imports anywhere; only a docstring mention in
|
||||
`core/__init__.py:4`. Main path uses `diffusers` + local `CoreMLUNetWrapper`; the runtime
|
||||
`CoreMLModel` (STAYS in suite) is a local coremltools wrapper, not Apple's. No shape/attn helper
|
||||
comes from Apple (local `conversion/shapes.py`, `conversion/attention.py`).
|
||||
3. **Decide the interface contract concretely (the maintainer's hard requirement):**
|
||||
- Discovery functions `list_model_versions / list_attention_impls / list_quant_modes` live in
|
||||
the package and return today's display strings verbatim. Node `INPUT_TYPES` calls them.
|
||||
- `ModelVersion` values, attn-impl strings, quant modes are **ADDITIVE-ONLY** across package
|
||||
versions; removal/rename = MAJOR bump + migration note. Write this into the package's
|
||||
versioning policy doc now.
|
||||
- `compose_out_name` moves to the package; node imports it (no copy). Confirm the
|
||||
characterization tests in `test_characterization_out_name.py` will be re-pointed, not
|
||||
duplicated.
|
||||
- **Resolve the `model_version` dropdown question (maintainer decided):** the Suite exposes
|
||||
*every model the converter can verifiably convert*. `lcm` and `sdxl_refiner` are absent today
|
||||
only because they lack a golden/PSNR verification — **not** because the node hardcodes a
|
||||
short list. Encode this as a **status registry in the package** (`VERIFIED` vs
|
||||
`EXPERIMENTAL`); `list_model_versions()` returns VERIFIED-only by default. The converter node
|
||||
calls it plainly. Promoting LCM/refiner to VERIFIED (after E-LCM / a refiner anchor) expands
|
||||
the dropdown with **no Suite change**. Do NOT add permanent per-node filtering — the gate is
|
||||
verification status, owned by the package.
|
||||
4. ~~Confirm the `ml-stable-diffusion` git dep is pinned.~~ **N/A — already removed (#58).** Verified:
|
||||
zero `python_coreml_stable_diffusion` imports in the repo; `CoreMLModel` is now a local
|
||||
coremltools wrapper; the dep is absent from `pyproject.toml`/`requirements.txt`. No SHA to pin.
|
||||
|
||||
### STOP — VALIDATE (Gate E0)
|
||||
```
|
||||
## Gate E0 report
|
||||
- seam.md committed: <path>; symbol counts (move / stay / already-framework-free)
|
||||
- python_coreml_stable_diffusion usage (verified by grep): conversion=<list> runtime=<list>
|
||||
- Discovery API signatures frozen: list_model_versions (verified-only) / list_attention_impls / list_quant_modes
|
||||
- Status registry decided: sd15+sdxl=VERIFIED, lcm+sdxl_refiner=EXPERIMENTAL (gated, not hidden)
|
||||
- Additive-only contract policy doc written (incl. promotion=minor, demotion/removal=major): <path>
|
||||
- model_version dropdown: expose all (incl. LCM/REFINER) / filtered per node — DECISION: <...>
|
||||
- compose_out_name move-not-copy confirmed; tests re-point plan: <...>
|
||||
- LCM consolidation deferred to optional E-LCM: YES/NO
|
||||
- ml-stable-diffusion: N/A — already removed (#58), not a dependency (was: pin-or-BLOCKER)
|
||||
- Package name in-repo: coreml_diffusion (final PyPI name deferred to E5)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Phase E1 — Establish `coreml_diffusion` package + discovery API (mostly verification)
|
||||
|
||||
**Objective:** stand up the package namespace and the discovery surface. Much of the comfy-chain
|
||||
cut is **already done** — this phase mostly *verifies* that and adds the discovery functions.
|
||||
|
||||
### Tasks
|
||||
1. **Verify (don't redo):** `coreml_suite/model_version.py` is already clean (`Enum` only). Confirm
|
||||
`import coreml_suite.model_version` works with **no comfy** (`uv run python -c "..."` in a
|
||||
comfy-free `uv venv`). If true, E1's original "extract ModelVersion" task is already satisfied.
|
||||
2. Create the `coreml_diffusion/` package skeleton with `__init__.py` exporting the **discovery API**
|
||||
backed by the *existing* sources of truth for now (re-export `ModelVersion`, the
|
||||
`ATTENTION_IMPLEMENTATIONS` tuple, and `QUANT_NBITS_VALUES`) so values are byte-identical:
|
||||
```python
|
||||
def list_model_versions(): return [v.name for v in ModelVersion] # .name -> "SD15" (node reverses via ModelVersion[...]; .value KeyErrors)
|
||||
def list_attention_impls(): return list(ATTENTION_IMPLEMENTATIONS)
|
||||
def list_quant_modes(): return list(QUANT_NBITS_VALUES)
|
||||
CONTRACT_VERSION = "1.0"
|
||||
```
|
||||
(At this stage `coreml_diffusion` may live inside the repo and import from `coreml_suite.*`; the
|
||||
physical move of implementation happens in E2. The point of E1 is to freeze the *contract*.)
|
||||
3. **Decided (`.name`):** the node renders `ModelVersion.SD15.name` (`"SD15"`) and reverses the
|
||||
dropdown string via `ModelVersion[model_version]` (name lookup, `nodes.py:286`). Discovery API
|
||||
therefore returns `.name`; `.value` (`"sd15"`) would `KeyError`. Recorded in `seam.md` §5.
|
||||
|
||||
### Acceptance criteria
|
||||
- `uv run python -c "import coreml_diffusion; print(coreml_diffusion.list_model_versions(), coreml_diffusion.list_quant_modes())"`
|
||||
works in a **comfy-free** `uv venv` and prints today's exact strings.
|
||||
- Existing characterization tests pass unchanged.
|
||||
- No node behavior change yet (node still uses its current hardcoded lists in E1).
|
||||
|
||||
### STOP — VALIDATE (Gate E1)
|
||||
```
|
||||
## Gate E1 report
|
||||
- model_version.py confirmed comfy-free (uv, no comfy): PASS/FAIL
|
||||
- coreml_diffusion.list_* returns byte-identical strings to current dropdowns: YES/NO (show values)
|
||||
- .name vs .value decision for model_version discovery: <...>
|
||||
- CONTRACT_VERSION set; additive-only policy linked: <path>
|
||||
- Characterization tests unchanged & green (uv run pytest): YES/NO
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Phase E2 — Move conversion code into `coreml_diffusion` (in-repo, dedup, behavior-preserving)
|
||||
|
||||
**Objective:** physically relocate the conversion mechanics into the framework-free package,
|
||||
collapsing the two duplicate converters into one, with paths/device injected.
|
||||
|
||||
### Tasks
|
||||
1. Move into `coreml_diffusion/`: `pipeline.py` (`get_pipeline`, `get_unet`), `unet.py`
|
||||
(`UNet2DConditionModelLCM`), `inputs.py` (sample/lcm/sdxl input builders +
|
||||
`get_encoder_hidden_states_shape` + `get_coreml_inputs` + `get_inputs_spec`),
|
||||
`controlnet.py` (`add_cnet_support`), `convert.py` (`convert_unet`, `convert`,
|
||||
`convert_to_coreml`, `load_coreml_model`), `compile.py` (`compile_coreml_model`).
|
||||
2. **Dedupe LCM vs main** (the real remaining duplication): delete `lcm/converter.py`'s copies of
|
||||
`convert_to_coreml` / `load_coreml_model` / `get_out_path` / `get_sample_input` (LCM variant
|
||||
carries a `scheduler` arg — fold that into the shared `get_sample_input` as an optional param)
|
||||
in favor of the single `coreml_diffusion` implementation. The main path's helpers
|
||||
(`get_unet`/`get_encoder_hidden_states_shape`/`get_coreml_inputs`/`convert_unet`/`convert`) and
|
||||
the `conversion/` subpackage (`attention`, `shapes`, `trace`, `unet`) move as-is.
|
||||
3. **Inject paths**: replace `get_out_path`'s `folder_paths` reach-in with an injected `out_path`
|
||||
argument on `convert(...)`; `folder_paths` resolution moves up into the node (which already
|
||||
computes `out_name`). No `folder_paths` import anywhere in `coreml_diffusion`.
|
||||
4. **Inject device** where the LCM path used `comfy.model_management` (verify it still does):
|
||||
`convert(..., device=None)`, default to torch's default device.
|
||||
5. **Attention per-call, never at import:** main path already routes through
|
||||
`conversion/attention.apply_attention_implementation` — keep that. If `lcm/converter.py` still
|
||||
sets any module global at import, remove it; the package sets attention from the `attn_impl`
|
||||
arg inside `convert`.
|
||||
6. **Move `compose_out_name` into the package** (`coreml_diffusion/naming.py`); re-point
|
||||
`test_characterization_out_name.py` imports to `coreml_diffusion.naming` — assertions and values
|
||||
unchanged. The node will import it from the package in E3.
|
||||
7. Leave **thin shims** in `coreml_suite/converter.py` and `coreml_suite/lcm/converter.py` that
|
||||
re-export from `coreml_diffusion`, preserving the old call signatures the nodes use (nodes untouched
|
||||
this phase). Shims map comfy `folder_paths`/device into package args.
|
||||
|
||||
### Acceptance criteria
|
||||
- `uv run pytest -m unit` (Tier 0) imports `coreml_diffusion.*` with **no comfy / no MPS** and is green on Linux.
|
||||
- The dedup leaves exactly one implementation of each previously-duplicated function.
|
||||
- Characterization tests pass unchanged after the `compose_out_name` re-point.
|
||||
- `[M2]` A real SD1.5 conversion via the shim still produces a loadable model.
|
||||
- `[M2-ANE]` **Golden latent identical / within tolerance** to the MODERNIZATION Phase 2 anchor
|
||||
(same seed/prompt) — proves the move + dedup changed nothing.
|
||||
|
||||
### STOP — VALIDATE (Gate E2 — first regression gate)
|
||||
```
|
||||
## Gate E2 report
|
||||
- Tier 0 import of coreml_diffusion without comfy/MPS (uv run): PASS/FAIL
|
||||
- LCM/main duplicated funcs collapsed to one (list old→new): <map>
|
||||
- compose_out_name moved to package; char-tests re-pointed & green: YES/NO
|
||||
- Paths injected (no folder_paths in package): confirmed
|
||||
- Device injected (no comfy.model_management in package): confirmed
|
||||
- Attention set per-call, not at import (both main & lcm): confirmed
|
||||
- [M2-ANE] Golden latent vs Phase-2 anchor: identical / within tol <x> / DIVERGED (STOP)
|
||||
- Node INPUT_TYPES / mappings untouched: confirmed (diff)
|
||||
```
|
||||
**If the golden latent diverged at all, STOP and report — do not continue.**
|
||||
|
||||
---
|
||||
|
||||
## Phase E3 — Thin the nodes onto the package (behavior-preserving)
|
||||
|
||||
**Objective:** remove the shims; have the ComfyUI nodes call `coreml_diffusion` directly, keeping the
|
||||
node contract byte-identical.
|
||||
|
||||
### Tasks
|
||||
1. `CoreMLConverter.convert` (in `coreml_suite/nodes.py`): keep the `folder_paths`-based path
|
||||
resolution **in the node**; import `compose_out_name` from `coreml_diffusion` (not `coreml_suite.core`);
|
||||
call `coreml_diffusion.convert(...)` and `coreml_diffusion.compile_model(...)` directly; wrap the compiled path
|
||||
in `CoreMLModel`.
|
||||
2. **Wire the dropdowns to discovery (the maintainer's hard requirement).** Replace the hardcoded
|
||||
`INPUT_TYPES` lists with fail-soft discovery calls:
|
||||
```python
|
||||
def _discover(fn, fallback):
|
||||
try:
|
||||
import coreml_diffusion
|
||||
return getattr(coreml_diffusion, fn)()
|
||||
except Exception as e: # missing/old package, or import error
|
||||
logger.warning(f"coreml_diffusion.{fn} unavailable ({e}); using fallback {fallback}")
|
||||
return fallback
|
||||
...
|
||||
"model_version": (_discover("list_model_versions", ["SD15", "SDXL"]),),
|
||||
"attention_implementation": (_discover("list_attention_impls", ["SPLIT_EINSUM","SPLIT_EINSUM_V2","ORIGINAL"]),),
|
||||
"quantize_nbits": (_discover("list_quant_modes", ["none","8","6","4"]), {"default": "none"}),
|
||||
```
|
||||
This is what makes "update the package → new types appear in the old node, no Suite bump" true.
|
||||
3. `COREML_CONVERT_LCM` (in `coreml_suite/lcm/nodes.py`): route through `coreml_diffusion` for the shared
|
||||
mechanics. **Keep the existing LCM behavior/HF-hardcode for now** — consolidation is optional E-LCM.
|
||||
4. Delete the now-dead `coreml_suite/converter.py` / `coreml_suite/lcm/converter.py` shims (or
|
||||
reduce to a one-line re-export if anything external imports them — grep first).
|
||||
|
||||
### Acceptance criteria
|
||||
- `NODE_CLASS_MAPPINGS` / `NODE_DISPLAY_NAME_MAPPINGS` keys: **unchanged** (diff `__init__.py`).
|
||||
- Every `INPUT_TYPES` **field name** unchanged. Dropdown **values**: the discovery calls must
|
||||
return **a superset of** today's values, with every previously-present value still present and
|
||||
spelled identically (additive-only). *(This deliberately replaces the original spec's
|
||||
"values must be byte-identical/frozen" criterion — the maintainer requires the list be
|
||||
extensible at runtime. Frozen-field-names + additive-only-values is the new contract.)*
|
||||
- With `coreml_diffusion` **absent**, the node still registers and shows the fallback lists (fail-soft).
|
||||
- `[M2-ANE]` Golden latent still identical to the Phase-2 anchor.
|
||||
- `[M2-ANE]` The committed e2e workflow `tests/integration/...` still passes (PSNR > 25).
|
||||
|
||||
### STOP — VALIDATE (Gate E3)
|
||||
```
|
||||
## Gate E3 report
|
||||
- Node mappings diff: empty (confirmed)
|
||||
- INPUT_TYPES field-names diff: empty (confirmed)
|
||||
- Dropdown values: superset of prior, all prior values still present & identical: YES/NO (show)
|
||||
- Fail-soft with coreml_diffusion absent (node still registers): PASS/FAIL
|
||||
- compose_out_name now imported from coreml_diffusion (no node-side copy): confirmed
|
||||
- [M2-ANE] Golden latent vs anchor: identical / within tol / DIVERGED (STOP)
|
||||
- [M2-ANE] e2e workflow PSNR: <value> (> 25?)
|
||||
- Dead converter shims removed / reduced: <list>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Phase E4 — Standalone packaging & CLI (still in-repo)
|
||||
|
||||
**Objective:** make `coreml_diffusion` independently installable and usable without ComfyUI, with a CLI
|
||||
suitable for the planned article and for on-device/iOS conversion workflows.
|
||||
|
||||
### Tasks
|
||||
1. Add `coreml_diffusion/pyproject.toml`: name (working `coreml_diffusion`), `requires-python`, dependencies
|
||||
= `coremltools` (pinned to the MODERNIZATION-validated version), `diffusers`, `transformers`,
|
||||
`peft`, `omegaconf`, `numpy`, `torch`. **No `ml-stable-diffusion`** (already removed in #58, see
|
||||
§0.3) and **no comfy**. Suite pins `transformers>=4.44`/`peft>=0.13`/`omegaconf>=2.3` today;
|
||||
grep-confirm each is on the conversion path before listing it. A `[project.scripts]` entry:
|
||||
`coreml-diffusion = "coreml_diffusion.cli:main"`.
|
||||
2. `coreml_diffusion/cli.py`: `coreml-diffusion convert --ckpt PATH --model-version sd15 --out PATH
|
||||
[--height --width --batch-size --attn-impl --controlnet --lora NAME:STRENGTH ... --config PATH]`
|
||||
and `coreml-diffusion compile --src PATH --out-dir DIR --name NAME`. Mirrors `convert()`/`compile_model()`.
|
||||
3. Tier-0 Linux tests for the CLI **arg→call mapping** (mock the heavy `convert`); the real
|
||||
convert remains `[M2]`. Add a `[M2]` smoke test: convert a tiny synthetic UNet end-to-end.
|
||||
4. README for the package: install, CLI usage, "produce a `.mlpackage`/`.mlmodelc` for use in a
|
||||
Swift/iOS app", and the ANE positioning note (low-power, GPU-free, embeddable; SD1.5/SDXL on
|
||||
ANE, **not** a Flux-speed claim).
|
||||
|
||||
### Acceptance criteria
|
||||
- Fresh `python -m venv` + `uv pip install ./coreml-diffusion` (no ComfyUI present) imports and runs
|
||||
`coreml-diffusion --help` and the arg-mapping tests on Linux.
|
||||
- `[M2]` `coreml-diffusion convert` produces a model file identical (golden) to the node path.
|
||||
|
||||
### STOP — VALIDATE (Gate E4)
|
||||
```
|
||||
## Gate E4 report
|
||||
- uv pip install ./coreml-diffusion in comfy-free venv: PASS/FAIL (log)
|
||||
- CLI arg→call tests (Tier 0, Linux): green
|
||||
- [M2] CLI-produced model golden vs node-produced model: identical / DIVERGED
|
||||
- Package deps list (with pinned SHAs/versions + licenses):
|
||||
- New runtime deps vs suite before: <none / list>
|
||||
```
|
||||
**This is the gate that proves the package stands alone. Do not split repos before it passes.**
|
||||
|
||||
---
|
||||
|
||||
## Phase E5 — Physical split into a second repository
|
||||
|
||||
**Objective:** move `coreml_diffusion/` to its own repo; CoreMLSuite depends on it by pinned version.
|
||||
|
||||
### Tasks
|
||||
1. Create the new repo (maintainer action — agent prepares the tree, not the GitHub repo).
|
||||
Choose final distributable name; rename imports if changed (single sweep, recorded).
|
||||
2. CoreMLSuite `pyproject.toml` / `requirements.txt`: replace the conversion-only deps with a
|
||||
pinned dependency on the new package (`coreml_diffusion==<version>` from PyPI, or `git+...@<tag>`
|
||||
until first PyPI release). (There is no `git+...ml-stable-diffusion` line to remove — already
|
||||
gone since #58.)
|
||||
3. ~~Keep `python_coreml_stable_diffusion` for the loader.~~ **Void.** The loader is the local
|
||||
`coreml_suite/coreml_model.py` over `coremltools`; the suite keeps `coremltools` as a direct dep
|
||||
for it. No Apple lib involved.
|
||||
4. Set up the new repo's CI: Tier 0 on Linux (import + arg-mapping + input-shape math),
|
||||
`[M2]`/`[M2-ANE]` on a self-hosted/macOS-ARM runner reusing the golden-latent anchor.
|
||||
5. Versioning: SemVer; first release `0.1.0`. Document the compatibility matrix
|
||||
(coreml_diffusion ↔ coremltools version ↔ diffusers version). No ml-stable-diffusion axis.
|
||||
|
||||
### Acceptance criteria
|
||||
- CoreMLSuite installs in a fresh venv pulling the new package; e2e workflow still passes `[M2-ANE]`.
|
||||
- New repo CI green on Linux (Tier 0) and `[M2-ANE]` golden latent matches the anchor.
|
||||
- No conversion code remains in CoreMLSuite (grep: no `ct.convert`, no `from_single_file`,
|
||||
no `torch.jit.trace`).
|
||||
|
||||
### STOP — VALIDATE (Gate E5)
|
||||
```
|
||||
## Gate E5 report
|
||||
- New repo tree prepared: <path/branch>; final package name: <name>
|
||||
- Suite depends on package by pinned version: <spec>
|
||||
- Suite e2e [M2-ANE] PSNR after split: <value> (> 25?)
|
||||
- Conversion code fully absent from suite: confirmed (grep output)
|
||||
- Compatibility matrix documented: <link>
|
||||
- First release tag: 0.1.0
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Phase E6 — Quantization travels WITH the conversion code (already implemented → move)
|
||||
|
||||
**Objective:** quantization is **already implemented** (k-means `palettize_weights` in
|
||||
`converter.py`, `quantize_nbits` node input, `_q<bits>` filename suffix, README tradeoff table).
|
||||
There is nothing to *build*. It simply **moves with the conversion code in E2** as part of
|
||||
`convert_unet`. This phase is a checkpoint that it survived the extraction intact, plus exposing
|
||||
it through the CLI.
|
||||
|
||||
### Tasks
|
||||
1. Confirm the palettization block moved cleanly into `coreml_diffusion` (lives in `convert.py` or a
|
||||
`quantize.py` helper called from `convert_unet`). Default `"none"` stays byte-identical.
|
||||
2. Expose via CLI flag `--quantize {none,8,6,4}` (E4 already lists this) and via
|
||||
`list_quant_modes()` discovery (E1/E3).
|
||||
3. The existing README tradeoff table (SD1.5 1×512×512 SPLIT_EINSUM: none/8/6/4 → size/ms/PSNR)
|
||||
moves to the package README. Re-confirm one row `[M2-ANE]` so the article can cite a live number.
|
||||
|
||||
### Acceptance criteria
|
||||
- Default (`none`) output byte-identical to pre-extraction (covered by the E2/E3 golden latent).
|
||||
- `coreml-diffusion convert --quantize 4` produces a `_q4` artifact matching the node's `_q4` artifact `[M2]`.
|
||||
- `list_quant_modes()` drives the node dropdown (no hardcoded copy remains).
|
||||
|
||||
### STOP — VALIDATE (Gate E6)
|
||||
```
|
||||
## Gate E6 report
|
||||
- Palettization relocated into coreml_diffusion, called from convert_unet: confirmed
|
||||
- Default none output identical (golden): YES/NO
|
||||
- [M2] CLI --quantize {8,6,4} artifacts match node artifacts: YES/NO
|
||||
- Tradeoff table in package README with at least one re-confirmed [M2-ANE] row: <link>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Phase E-LCM — FIRST task after the split: clean up LCM + verify → promote (behavior-changing, gated)
|
||||
|
||||
> Promoted from "optional, someday" to **the first thing after E5**, per maintainer intent: the
|
||||
> Suite should expose every verifiably-convertible model, and LCM is the obvious first cleanup.
|
||||
|
||||
Two coupled goals:
|
||||
1. **Consolidate the LCM path.** Make the LCM node use the unified `from_single_file` path in
|
||||
`coreml_diffusion.convert(model_version=LCM, ...)` instead of the hardcoded `SimianLuo/LCM_Dreamshaper_v7`
|
||||
HF download; drop the duplicated LCM helpers (already deduped in E2). **Behavior change** ⇒
|
||||
capture an LCM golden anchor *before* the change, then prove within-tolerance after.
|
||||
2. **Verify → promote.** Once the LCM conversion has a passing `[M2-ANE]` golden anchor, flip
|
||||
`_MODEL_STATUS["lcm"] = Status.VERIFIED` **in the package** (minor bump). The Suite's dropdown
|
||||
gains `lcm` automatically — no Suite change, no Suite bump. This is the end-to-end proof that
|
||||
the discovery contract works as designed.
|
||||
|
||||
Repeat the same recipe for `sdxl_refiner` when it gets an anchor (separate small gate). Do NOT
|
||||
bundle E-LCM into E1–E5; it changes behavior and must stand on its own golden.
|
||||
|
||||
### STOP — VALIDATE (Gate E-LCM)
|
||||
```
|
||||
## Gate E-LCM report
|
||||
- LCM golden anchor captured BEFORE change: <path/hash>
|
||||
- LCM node now uses unified from_single_file path; HF hardcode removed: confirmed
|
||||
- [M2-ANE] LCM golden after change: identical / within tol <x> / DIVERGED (STOP)
|
||||
- Status flipped lcm→VERIFIED in package (minor bump <ver>): confirmed
|
||||
- Suite dropdown now lists lcm with NO Suite code change / NO Suite bump: confirmed (diff empty)
|
||||
- LCM node accepts a checkpoint arg now (documented breaking-ish UI note): <link>
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Article deliverable (after E4)
|
||||
|
||||
Once the CLI exists and stands alone, the "convert a Comfy/A1111 workflow into an on-device iOS
|
||||
app" write-up becomes a clean tutorial: `coreml-diffusion convert` → `.mlmodelc` → load in Swift/CoreML.
|
||||
Frame the niche honestly per the README note above (ANE feasibility & power, not raw Flux speed).
|
||||
|
||||
---
|
||||
|
||||
## Quick reference: extraction gate discipline
|
||||
|
||||
```
|
||||
E0 Seam decision, interface contract, discovery API frozen → Gate E0 (cut line + additive-only policy?)
|
||||
E1 Stand up coreml_diffusion + discovery API (mostly verify) → Gate E1 (list_* byte-identical, comfy-free?)
|
||||
E2 Move conversion code, dedup LCM/main, inject paths/device→ Gate E2 (golden identical? duplicates gone?) ← first regression gate
|
||||
E3 Thin nodes onto package + wire discovery dropdowns → Gate E3 (field-names frozen, values additive, fail-soft, golden identical?)
|
||||
E4 Standalone packaging + CLI (uv) → Gate E4 (uv pip install w/o comfy? CLI golden?) ← proves it stands alone
|
||||
E5 Physical second-repo split → Gate E5 (suite depends on pkg? conversion absent?)
|
||||
E-LCM FIRST post-split: clean up LCM, verify → promote → Gate E-LCM (LCM golden? dropdown gains lcm w/ no Suite bump?)
|
||||
E6 Quantization checkpoint (already built → moved in E2) → Gate E6 (default identical? CLI quant matches?)
|
||||
(refiner) same recipe as E-LCM when an anchor exists → own small gate (promote sdxl_refiner→VERIFIED)
|
||||
```
|
||||
|
||||
**Interface-contract invariants (the maintainer's hard requirement), restated:**
|
||||
- Package API is keyword-only-with-defaults past the required positionals → converter updates
|
||||
don't force Suite updates.
|
||||
- Node dropdowns are discovery-driven (`coreml_diffusion.list_*`) + fail-soft → new conversion types
|
||||
appear in the old plugin with `uv pip install -U coreml_diffusion`, **no Suite code change, no bump**.
|
||||
- Discovery identifiers are **additive-only**; removal/rename = MAJOR bump + migration note.
|
||||
- `compose_out_name` (cache key) lives in the package, single copy.
|
||||
|
||||
|
||||
**Golden rule (inherited): never cross a gate with a failing acceptance criterion.
|
||||
Stop, report, wait. The golden latent is the single source of truth that the extraction
|
||||
changed nothing.**
|
||||
@@ -1,452 +1,113 @@
|
||||
# Core ML Suite for ComfyUI
|
||||
|
||||
## Overview
|
||||
Custom nodes for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) that run
|
||||
Stable Diffusion UNets as [Core ML](https://developer.apple.com/documentation/coreml)
|
||||
models on Apple Silicon (M1/M2/M3). Core ML can use the Apple Neural Engine
|
||||
(ANE), which is unavailable to PyTorch — on an M2 Pro 32 GB, SD1.5 at 512×512
|
||||
generates roughly **1.5–2× faster** than the standard PyTorch/MPS path.
|
||||
|
||||
Welcome! In this repository you'll find a set of custom nodes for [ComfyUI](https://github.com/comfyanonymous/ComfyUI)
|
||||
that allows you to use Core ML models in your ComfyUI workflows.
|
||||
These models are designed to leverage the Apple Neural Engine (ANE) on Apple Silicon (M1/M2) machines,
|
||||
thereby enhancing your workflows and improving performance.
|
||||
You convert a Stable Diffusion checkpoint to a Core ML model with the nodes in
|
||||
this suite, then sample from it like any other ComfyUI workflow.
|
||||
|
||||
If you're not sure how to obtain these models, you can download them
|
||||
[here](https://huggingface.co/coreml-community) or convert your own checkpoints
|
||||
directly with the conversion nodes in this suite (see [How to use](#how-to-use)).
|
||||
> [!IMPORTANT]
|
||||
> **Convert your own checkpoints — that is the only supported path.** This
|
||||
> suite uses its own input dimensions, naming convention, and metadata
|
||||
> (produced by the [coreml-diffusion](https://github.com/aszc-dev/coreml-diffusion)
|
||||
> package). Pre-converted Core ML models from elsewhere (e.g. the
|
||||
> coreml-community Hugging Face org) are **not** supported. Conversion is cheap
|
||||
> and runs on your machine, so there is no need to download Core ML models.
|
||||
|
||||
In simple terms, think of Core ML models as a tool that can help your ComfyUI work faster and more efficiently.
|
||||
For instance, during my tests on an M2 Pro 32GB machine,
|
||||
the use of Core ML models sped up the generation of 512x512 images by a factor
|
||||
of approximately 1.5 to 2 times.
|
||||
## Installation
|
||||
|
||||
## Getting Started
|
||||
### ComfyUI-Manager (recommended)
|
||||
|
||||
To start using custom nodes in your ComfyUI, follow these simple steps:
|
||||
Open **Manager → Install Custom Nodes**, search for `Core ML`, click
|
||||
**Install**, and restart ComfyUI.
|
||||
|
||||
1. Clone or download this repository: You can do this directly into the custom_nodes directory of your ComfyUI.
|
||||
2. Install the dependencies: You'll need to use a package manager like pip to do this.
|
||||
### Manual
|
||||
|
||||
That's it! You're now ready to start enhancing your ComfyUI workflows with Core ML models.
|
||||
```bash
|
||||
cd /path/to/comfyui/custom_nodes
|
||||
git clone https://github.com/aszc-dev/ComfyUI-CoreMLSuite.git
|
||||
cd ComfyUI-CoreMLSuite
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
- Check [Installation](#installation) for more details on installation.
|
||||
- Check [How to use](#how-to-use) for more details on how to use the custom nodes.
|
||||
- Check [Example Workflows](#example-workflows) for some example workflows.
|
||||
Dependencies (`coreml-diffusion`, `coremltools`, `numpy`, `diffusers`) install
|
||||
from PyPI. PyTorch is intentionally **not** pinned — it is provided by your
|
||||
ComfyUI host, and a hard cap here would downgrade it and break ComfyUI.
|
||||
|
||||
## Quickstart
|
||||
|
||||
1. Put a SD1.5 checkpoint in `models/checkpoints`.
|
||||
2. Add the **Convert Checkpoint to Core ML** node, select the checkpoint, and
|
||||
queue once. It writes a `.mlpackage` to `models/unet` (cached by name — it
|
||||
won't reconvert next time).
|
||||
3. Sample with the **Core ML Sampler** node, decoding the latent with a normal
|
||||
VAE Decode. CLIP and VAE come from standard ComfyUI nodes.
|
||||
|
||||
See [docs/workflows.md](docs/workflows.md) for complete example graphs (txt2img,
|
||||
ControlNet, LoRA, LCM, SDXL).
|
||||
|
||||
## Which compute unit should I pick?
|
||||
|
||||
The **compute unit** selects the hardware Core ML runs on. Pair it with the
|
||||
attention implementation chosen at conversion time:
|
||||
|
||||
| Model | Convert with | Load with | Runs on |
|
||||
|---|---|---|---|
|
||||
| SD1.5 @ 512×512 | `SPLIT_EINSUM` | `CPU_AND_NE` | Neural Engine (fastest) |
|
||||
| SD1.5 @ larger sizes | `ORIGINAL` | `CPU_AND_GPU` | GPU |
|
||||
| SDXL | `ORIGINAL` | `CPU_AND_GPU` | GPU (ANE unsupported) |
|
||||
|
||||
`CPU_AND_NE` is usually the fastest option for SD1.5 — often faster than `ALL`.
|
||||
This suite uses Core ML compute units only; it never touches PyTorch MPS, so
|
||||
`PYTORCH_ENABLE_MPS_FALLBACK` is irrelevant to these nodes. Full reasoning and
|
||||
benchmarks: [docs/hardware.md](docs/hardware.md).
|
||||
|
||||
## Documentation
|
||||
|
||||
- [Hardware & compute units](docs/hardware.md) — ANE vs GPU vs MPS, attention
|
||||
implementations, which to choose.
|
||||
- [Nodes](docs/nodes.md) — full reference for every node.
|
||||
- [Conversion](docs/conversion.md) — how conversion works, caching,
|
||||
quantization.
|
||||
- [Example workflows](docs/workflows.md) — annotated example graphs.
|
||||
- [FAQ](docs/faq.md) — answers to common questions.
|
||||
- [Troubleshooting](docs/troubleshooting.md) — common errors and fixes.
|
||||
- [Limitations & support matrix](docs/limitations.md) — what is and isn't
|
||||
supported.
|
||||
|
||||
## Glossary
|
||||
|
||||
- **Core ML**: A machine learning framework developed by Apple. It's used to run machine learning models on Apple
|
||||
devices.
|
||||
- **Core ML Model**: A machine learning model that can be run on Apple devices using Core ML.
|
||||
- **mlmodelc**: A compiled Core ML model. This is the recommended format for Core ML models.
|
||||
- **mlpackage**: A Core ML model packaged in a directory. This is the default format for Core ML models.
|
||||
- **ANE**: Apple Neural Engine. A hardware accelerator for machine learning tasks on Apple devices.
|
||||
- **Compute Unit**: A Core ML option that allows you to specify the hardware on which the model should run.
|
||||
- **CPU_AND_ANE**: A Core ML compute unit option that allows the model to run on both the CPU and ANE. This is the
|
||||
default option.
|
||||
- **CPU_AND_GPU**: A Core ML compute unit option that allows the model to run on both the CPU and GPU.
|
||||
- **CPU_ONLY**: A Core ML compute unit option that allows the model to run on the CPU only.
|
||||
- **ALL**: A Core ML compute unit option that allows the model to run on all available hardware.
|
||||
- **CLIP**: Contrastive Language-Image Pre-training. A model that learns visual concepts from natural language
|
||||
supervision. It's used as a text encoder in Stable Diffusion.
|
||||
- **VAE**: Variational Autoencoder. A model that learns a latent representation of images. It's used as a prior in
|
||||
Stable Diffusion.
|
||||
- **Checkpoint**: A file that contains the weights of a model. It's used to load models in Stable Diffusion.
|
||||
- **LCM**: [Latent Consistency Model](https://latent-consistency-models.github.io/). A type of model designed to
|
||||
generate images with as few steps as possible.
|
||||
|
||||
> [!NOTE]
|
||||
> Note on Compute Units:
|
||||
> For the model to run on the ANE, the model must be converted with the `--attention-implementation SPLIT_EINSUM`
|
||||
> option.
|
||||
> Models converted with `--attention-implementation ORIGINAL` will run on GPU instead of ANE.
|
||||
|
||||
## Features
|
||||
|
||||
These custom nodes come with a host of features, including:
|
||||
|
||||
- Loading Core ML Unet models
|
||||
- Support for ControlNet
|
||||
- Support for ANE (Apple Neural Engine)
|
||||
- Support for CPU and GPU
|
||||
- Support for `mlmodelc` and `mlpackage` files
|
||||
- Support for SDXL models
|
||||
- Support for LCM models
|
||||
- Support for LoRAs
|
||||
- SD1.5 -> Core ML conversion
|
||||
- SDXL -> Core ML conversion
|
||||
- LCM -> Core ML conversion
|
||||
|
||||
> [!NOTE]
|
||||
> Please note that using Core ML models can take a bit longer to load initially.
|
||||
> For the best experience, I recommend using the compiled models
|
||||
> (.mlmodelc files) instead of the .mlpackage files.
|
||||
|
||||
> [!NOTE]
|
||||
> This repository will continue to be updated with more nodes and features over time.
|
||||
|
||||
## Conversion & Acknowledgements
|
||||
|
||||
The Core ML conversion pipeline in this repository began as an adaptation of
|
||||
Apple's [ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion),
|
||||
which pioneered running Stable Diffusion on the Apple Neural Engine. The
|
||||
implementation has since diverged and no longer depends on that package:
|
||||
|
||||
- UNet conversion runs natively on `diffusers`' `UNet2DConditionModel`.
|
||||
- The ANE-friendly attention path (`SPLIT_EINSUM`, `SPLIT_EINSUM_V2`) is
|
||||
reimplemented as standalone `diffusers` attention processors.
|
||||
- The toolchain tracks current ComfyUI (NumPy 2, Torch 2.7, coremltools 9,
|
||||
Python 3.12).
|
||||
|
||||
The goal is to keep iterating on these methods independently and to explore
|
||||
support beyond SD1.5.
|
||||
- **Core ML** — Apple's on-device machine-learning framework.
|
||||
- **`.mlpackage`** — the Core ML model format this suite produces and loads.
|
||||
- **ANE** — Apple Neural Engine, a hardware accelerator for ML.
|
||||
- **Compute unit** — which hardware Core ML uses (`CPU_AND_NE`, `CPU_AND_GPU`,
|
||||
`CPU_ONLY`, `ALL`).
|
||||
- **Attention implementation** — `SPLIT_EINSUM` / `SPLIT_EINSUM_V2` (ANE-friendly)
|
||||
or `ORIGINAL` (GPU-friendly), chosen at conversion.
|
||||
|
||||
> [!IMPORTANT]
|
||||
> **Breaking change in 2.0.0.** The converted Core ML UNet now takes
|
||||
> `encoder_hidden_states` in the native `diffusers` layout
|
||||
> `(batch, tokens, hidden)` instead of the previous
|
||||
> `(batch, hidden, 1, tokens)`. Core ML models converted with earlier versions
|
||||
> are not compatible with 2.0.0 and must be re-converted.
|
||||
|
||||
## Installation
|
||||
|
||||
### Using ComfyUI-Manager
|
||||
|
||||
The easiest way to install the custom nodes is to use the ComfyUI-Manager. You can find the installation instructions
|
||||
[here](https://github.com/ltdrdata/ComfyUI-Manager#installation). Once you've installed the ComfyUI-Manager, you can
|
||||
install the custom nodes by following these steps:
|
||||
|
||||
- Open the ComfyUI-Manager by clicking the `Manager` button in the ComfyUI toolbar.
|
||||
- Click the `Install Custom Nodes` button.
|
||||
- Search for `Core ML` and click the `Install` button.
|
||||
- Restart ComfyUI.
|
||||
|
||||
### Manual Installation
|
||||
|
||||
1. Clone this repository into the custom_nodes directory of your ComfyUI. If you're not sure how to do this, you can
|
||||
download the repository as a zip file and extract it into the same directory.
|
||||
```bash
|
||||
cd /path/to/comfyui/custom_nodes
|
||||
git clone https://github.com/aszc-dev/ComfyUI-CoreMLSuite.git
|
||||
```
|
||||
2. Next, install the required dependencies using pip or another package manager:
|
||||
|
||||
```bash
|
||||
cd /path/to/comfyui/custom_nodes/ComfyUI-CoreMLSuite
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
## How to use
|
||||
|
||||
Once you've installed the custom nodes, you can start using them in your ComfyUI workflows.
|
||||
To do this, you need to add the nodes to your workflow. You can do this by right-clicking on the workflow canvas and
|
||||
selecting the nodes from the list of available nodes (the nodes are in the `Core ML Suite` category).
|
||||
You can also double-click the canvas and use the search bar to find the nodes. The list of available nodes is given
|
||||
below.
|
||||
|
||||
### Available Nodes
|
||||
|
||||
#### Core ML UNet Loader (`CoreMLUnetLoader`)
|
||||
|
||||

|
||||
|
||||
This node allows you to load a Core ML UNet model and use it in your ComfyUI workflow. Place the converted
|
||||
.mlpackage or .mlmodelc file in ComfyUI's `models/unet` directory and use the node to load the model. The output of the
|
||||
node is a `coreml_model` object that can be used with the Core ML Sampler.
|
||||
|
||||
- **Inputs**:
|
||||
- **model_name**: The name of the model to load. This should be the name of the .mlpackage or .mlmodelc file.
|
||||
- **compute_unit**: The hardware on which the model should run. This can be one of the following:
|
||||
- `CPU_AND_ANE`: The model will run on both the CPU and ANE. This is the default option. It works best with
|
||||
models
|
||||
converted with `--attention-implementation SPLIT_EINSUM` or `--attention-implementation SPLIT_EINSUM_V2`.
|
||||
- `CPU_AND_GPU`: The model will run on both the CPU and GPU. It works best with models converted with
|
||||
`--attention-implementation ORIGINAL`.
|
||||
- `CPU_ONLY`: The model will run on the CPU only.
|
||||
- `ALL`: The model will run on all available hardware.
|
||||
- **Outputs**:
|
||||
- **coreml_model**: A Core ML model that can be used with the Core ML Sampler.
|
||||
|
||||
#### Core ML Sampler (`CoreMLSampler`)
|
||||
|
||||

|
||||
|
||||
This node allows you to generate images using a Core ML model. The node takes a Core ML model as input and outputs a
|
||||
latent image similar to the latent image output by the KSampler. This means that you can use the
|
||||
resulting latent as you normally would in your workflow.
|
||||
|
||||
- **Inputs**:
|
||||
- **coreml_model**: The Core ML model to use for sampling. This should be the output of the Core ML UNet Loader.
|
||||
- **latent_image** [optional]: The latent image to use for sampling. If provided, should be of the same size as the
|
||||
input of the Core ML model. If not provided, the node will create a latent suitable for the Core ML model used.
|
||||
Useful in img2img workflows.
|
||||
- ... _(the rest of the inputs are the same as the KSampler)_
|
||||
- **Outputs**:
|
||||
- **LATENT**: The latent image output by the Core ML model. This can be decoded using a VAE Decoder or used as input
|
||||
to the next node in your workflow.
|
||||
|
||||
#### Checkpoint Converter
|
||||
|
||||

|
||||
|
||||
You can use this node to convert any **SD1.5** based checkpoint to a Core ML model. The converted model is stored in the
|
||||
`models/unet` directory and can be used with the `Core ML UNet Loader`. The conversion parameters are encoded in
|
||||
the node name, so if the model already exists, the node will not convert it again.
|
||||
|
||||
- **Inputs**:
|
||||
- **ckpt_name**: The name of the checkpoint to convert. This should be the name of the checkpoint file stored in the
|
||||
`models/checkpoints` directory.
|
||||
- **model_version**: Whether the model is based on SD1.5 or SDXL.
|
||||
- **height**: The desired height of the image generated by the model. The default is 512. Any positive multiple of 8 is accepted.
|
||||
- **width**: The desired width of the image generated by the model. The default is 512. Any positive multiple of 8 is accepted.
|
||||
- **batch_size**: The batch size of generated images. If you're planning to generate batches of images, you can try
|
||||
increasing this value to speed up the generation process. The default is 1.
|
||||
- **attention_implementation**: The attention implementation used when converting the model. Choose SPLIT_EINSUM or
|
||||
SPLIT_EINSUM_V2 for better ANE support. Choose ORIGINAL for better GPU support.
|
||||
- **compute_unit**: The hardware on which the model should run. This is used only when loading the model and doesn't
|
||||
affect the conversion process.
|
||||
- **controlnet_support**: For the model to support ControlNet, it must be converted with this option set to True.
|
||||
The
|
||||
default is False.
|
||||
- **lora_params** [optional]: Optional LoRA names and weights. If provided, the model will be converted with LoRA(s)
|
||||
baked in. More on loading LoRAs below.
|
||||
- **Outputs**:
|
||||
- **coreml_model**: The converted Core ML model that can be used with Core ML Sampler.
|
||||
|
||||
> [!NOTE]
|
||||
> Some models use a custom config .yaml file. If you're using such a model, you'll need to place the config file in the
|
||||
> `models/configs` directory. The config file should be named the same as the checkpoint file. For example, if the
|
||||
> checkpoint file is named `juggernaut_aftermath.safetensors`, the config file should be
|
||||
> named `juggernaut_aftermath.yaml`.
|
||||
> The config file will be automatically loaded during conversion.
|
||||
|
||||
> [!NOTE]
|
||||
> For now, the converter relies heavilty on the model name to determine the conversion parameters. This means that if
|
||||
> you change the model name, the node will convert the model again. Other than that, if you find the name too long or
|
||||
> confusing, you can change it to anything you want.
|
||||
|
||||
#### LoRA Loader
|
||||
|
||||

|
||||
|
||||
This node allows you to load LoRAs and bake them into a model. Since this is a workaround (as model weights can't be
|
||||
modified
|
||||
after conversion), there are a few caveats to keep in mind:
|
||||
|
||||
- The LoRA weights and _strength_model_ parameter are baked into the model. This means that you can't change them
|
||||
after conversion. This also means that you need to convert the model again if you want to change the LoRA weights.
|
||||
- Loading LoRA affects CLIP, which is not a part of Core ML workflow, so you'll need to load CLIP separately,
|
||||
either using `CLIPLoader` or `CheckpointLoaderSimple`. (See [example workflows](#example-workflows) for more details.)
|
||||
- After conversion, if you want to load the model using `CoreMLUnetLoader`, you'll need to apply the same LoRAs to
|
||||
CLIP manually. (See [example workflows](#example-workflows) for more details.)
|
||||
- The LoRA names are encoded in the model name. This means that if you change the name of the LoRA file,
|
||||
you'll need to change the model name as well, or the node will convert the model again. (Model strength is not
|
||||
encoded, so if you want to change it, you'll need to delete the converted model manually)
|
||||
- _strength_clip_ parameter only affects the CLIP model and is not baked into the converted model. This means that
|
||||
you can change it after conversion.
|
||||
|
||||
- **Inputs**:
|
||||
- **lora_name**: The name of the LoRA to load.
|
||||
- **strength_model**: The strength of the LoRA model.
|
||||
- **strength_clip**: The strength of the LoRA CLIP.
|
||||
- **lora_params** [optional]: Optional output from other LoRA Loaders.
|
||||
- **clip**: The CLIP model to use with the LoRA. This can be either output of the
|
||||
`CLIPLoader`/`CheckpointLoaderSimple` or other LoRA Loaders.
|
||||
- **Outputs**:
|
||||
- **lora_params**: The LoRA parameters that can be passed to the Core ML Converter or other LoRA Loaders.
|
||||
- **CLIP**: The CLIP model with LoRA applied.
|
||||
|
||||
#### LCM Converter
|
||||
|
||||

|
||||
|
||||
This node converts [SimianLuo/LCM_Dreamshaper_v7](https://huggingface.co/SimianLuo/LCM_Dreamshaper_v7) model to Core
|
||||
ML. The converted model is stored in the `models/unet` directory and can be used with the Core ML UNet Loader. The
|
||||
conversion parameteres are encoded in the node name, so if the model already exists, the node will not convert it again.
|
||||
|
||||
- **Inputs**:
|
||||
- **height**: The desired height of the image generated by the model. The default is 512. Must be a multiple of 8.
|
||||
- **width**: The desired width of the image generated by the model. The default is 512. Must be a multiple of 8.
|
||||
- **batch_size**: The batch size of generated images. If you're planning to generate batches of images, you can try
|
||||
increasing this value to speed up the generation process. The default is 1.
|
||||
- **compute_unit**: The hardware on which the model should run. This is used only when loading the model and
|
||||
doesn't affect the conversion process.
|
||||
- **controlnet_support**: For the model to support ControlNet, it must be converted with this option set to True.
|
||||
The default is False.
|
||||
|
||||
> [!NOTE]
|
||||
> The conversion process can take a while, so please be patient.
|
||||
|
||||
> [!NOTE]
|
||||
> When using the LCM model with Core ML Sampler, please set _sampler_name_ to `lcm` and _scheduler_ to `sgm_uniform`.
|
||||
|
||||
#### Core ML Adapter (Experimental) (`CoreMLModelAdapter`)
|
||||
|
||||

|
||||
|
||||
This node allows you to use a Core ML as a standard ComfyUI model. This is an experimental node and may not work with
|
||||
all models and nodes. Please use with caution and pay attention to the expected inputs of the model.
|
||||
|
||||
- **Input**:
|
||||
- **coreml_model**: The Core ML model to use as a ComfyUI model.
|
||||
- **Output**:
|
||||
- **MODEL**: The Core ML model wrapped in a ComfyUI model.
|
||||
|
||||
> [!NOTE]
|
||||
> While this approach allows you to use Core ML models with many ComfyUI nodes (both standard and custom), the
|
||||
> expected inputs of the model will not be checked, which may cause errors. Please make sure to use a model compatible
|
||||
> with the expected parameters.
|
||||
|
||||
### Example Workflows
|
||||
|
||||
> [!NOTE]
|
||||
> The models used are just an example. Feel free to experiment with different models and see what works best for you.
|
||||
|
||||
#### Basic txt2img with Core ML UNet loader
|
||||
|
||||
This is a basic txt2img workflow that uses the Core ML UNet loader to load a model. The CLIP and VAE models
|
||||
are loaded using the standard ComfyUI nodes. In the first example, the text encoder (CLIP) and VAE models are loaded
|
||||
separately. In the second example, the text encoder and VAE models are loaded from the checkpoint file. Note that you
|
||||
can use any CLIP or VAE model as long as it's compatible with Stable Diffusion v1.5.
|
||||
|
||||
1. **Loading text encoder (CLIP) and VAE models separately**
|
||||
- This workflow uses CLIP and VAE models available
|
||||
[here](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/text_encoder/model.safetensors) and
|
||||
[here](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/vae/diffusion_pytorch_model.safetensors).
|
||||
Once downloaded, place the models in the`models/clip` and `models/vae` directories respectively.
|
||||
- The Core ML UNet model is available
|
||||
[here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip).
|
||||
Once downloaded, place the model in the `models/unet` directory.
|
||||

|
||||
2. **Loading text encoder (CLIP) and VAE models from checkpoint file**
|
||||
- This workflow loads the CLIP and VAE models from the checkpoint file available
|
||||
[here](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/v1-5-pruned-emaonly.safetensors).
|
||||
Once downloaded, place the model in the`models/checkpoints` directory.
|
||||
- The Core ML UNet model is available
|
||||
[here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip).
|
||||
Once downloaded, place the model in the `models/unet` directory.
|
||||

|
||||
|
||||
#### ControlNet with Core ML UNet loader
|
||||
|
||||
This workflow uses the Core ML UNet loader to load a Core ML UNet model that supports ControlNet. The ControlNet is
|
||||
being loaded using the standard ComfyUI nodes. Please refer to
|
||||
the [basic txt2img workflow](#basic-txt2img-with-core-ml-unet-loader) for more details on how to load the CLIP and VAE
|
||||
models.
|
||||
The ControlNet model used in this workflow is available
|
||||
[here](https://huggingface.co/lllyasviel/control_v11p_sd15_scribble/blob/main/diffusion_pytorch_model.fp16.safetensors).
|
||||
Once downloaded, place the model in the `models/controlnet` directory.
|
||||

|
||||
|
||||
#### Checkpoint conversion
|
||||
|
||||
This workflow uses the Checkpoint Converter to convert the checkpoint file. See
|
||||
[Checkpoint Converter](#checkpoint-converter) description for more details.
|
||||
|
||||

|
||||
|
||||
#### Checkpoint conversion with LoRA
|
||||
|
||||
This workflow uses the Checkpoint Converter to convert the checkpoint file with LoRA. See
|
||||
[LoRA Loader](#lora-loader) description to read more about the caveats of using LoRA.
|
||||
|
||||

|
||||
|
||||
#### LCM LoRA conversion
|
||||
|
||||
Please note that you can use multiple LoRAs with the same model. To do this, you'll need to use multiple LoRA Loaders.
|
||||
> [!IMPORTANT]
|
||||
> In this example, the model is passed through the adapter and `ModelSamplingDiscrete` nodes to a standard ComfyUI's
|
||||
> KSampler (not Core ML Sampler). ModelSamplingDiscrete needs to be used to sample models with LCM LoRAs properly.
|
||||
|
||||

|
||||
|
||||
#### Loader with LoRAs
|
||||
|
||||
This workflow uses the Core ML UNet Loader to load a model with LoRAs. The CLIP must be loaded separately and passed
|
||||
through the same LoRA nodes as during conversion. See [LoRA Loader](#lora-loader) description to read more about the
|
||||
caveats of using LoRA. Since _lora_name_ and _strength_model_ are baked into the model, it is not necessary to pass
|
||||
them as inputs to the loader.
|
||||
> [!IMPORTANT]
|
||||
> In this example, the model is passed through the adapter and `ModelSamplingDiscrete` nodes to a standard ComfyUI's
|
||||
> KSampler (not Core ML Sampler). ModelSamplingDiscrete needs to be used to sample models with LCM LoRAs properly.
|
||||
|
||||

|
||||
|
||||
#### LCM conversion with ControlNet
|
||||
|
||||
This workflow uses LCM converter to
|
||||
convert [SimianLuo/LCM_Dreamshaper_v7](https://huggingface.co/SimianLuo/LCM_Dreamshaper_v7)
|
||||
model to Core ML. The converted model can then be used with or without ControlNet to generate images.
|
||||

|
||||
|
||||
#### SDXL Base + Refiner conversion
|
||||
|
||||
This is a basic workflow for SDXL. You add LoRAs and ControlNets the same way as in the previous examples.
|
||||
You can also skip the refiner step.
|
||||
|
||||
The models used in this workflow are available at the following links:
|
||||
|
||||
- [Base model + text_encoder (clip) + text_encoder_2 (clip2)](https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0)
|
||||
- [Refiner model](https://huggingface.co/stabilityai/stable-diffusion-xl-refiner-1.0)
|
||||
- [VAE](https://huggingface.co/stabilityai/sdxl-vae)
|
||||
|
||||
> [!IMPORTANT]
|
||||
> SDXL on ANE is not supported. If loading of the model gets stuck, please try using CPU_AND_GPU or CPU_ONLY.
|
||||
> For best results, use ORIGINAL attention implementation.
|
||||
|
||||

|
||||
|
||||
## Quantization (opt-in)
|
||||
|
||||
The `Core ML Converter` and `Core ML LCM Converter` nodes accept an
|
||||
optional `quantize_nbits` dropdown that runs k-means weight palettization
|
||||
(`coremltools.optimize.coreml.palettize_weights`) on the UNet before save.
|
||||
|
||||
Values: `none` (default — no quantization, identical to unquantized
|
||||
behavior and filenames), `8`, `6`, `4`. The number is appended to the
|
||||
.mlpackage stem as `_q<bits>` so quantized and unquantized variants
|
||||
coexist on disk and in cache.
|
||||
|
||||
### SD1.5 1×512×512 SPLIT_EINSUM tradeoffs (M2 Pro, ANE)
|
||||
|
||||
Measured with 20 UNet forward passes at a fixed seed for the PSNR
|
||||
comparison:
|
||||
|
||||
| nbits | size (MB) | size vs none | fwd median (ms) | PSNR vs `none` (dB) |
|
||||
|---|---:|---:|---:|---:|
|
||||
| none | 1641 | 1.000 | 197.1 | — |
|
||||
| 8 | 822 | 0.501 | 186.6 | 53.5 |
|
||||
| 6 | 617 | 0.376 | 183.0 | 40.2 |
|
||||
| 4 | 412 | 0.251 | 179.8 | 27.5 |
|
||||
|
||||
PSNR here is computed on the raw `noise_pred` output of a single UNet
|
||||
forward at a fixed seed, not on the final decoded image — it isolates
|
||||
the quantization-induced drift from sampler / VAE noise. Final-image
|
||||
PSNR is comfortably higher (the sampler averages over 20 steps).
|
||||
|
||||
### Recommended settings per chip / RAM
|
||||
|
||||
- **8 GB RAM (M1 base, M2 base):** `nbits=4`. ~4× smaller model, still
|
||||
loads, PSNR 27 dB is visually identical at SD1.5 sizes.
|
||||
- **16 GB RAM (M1/M2/M3 Pro):** `nbits=6` is the sweet spot — ~2.7×
|
||||
smaller, PSNR 40 dB, no perceptible quality drop.
|
||||
- **32 GB+ RAM (Max / Ultra):** `nbits=8` if you want the safety
|
||||
margin, `none` if you want bit-identical output for golden testing.
|
||||
|
||||
The default stays `none` so existing workflows produce byte-for-byte
|
||||
identical output.
|
||||
|
||||
## Limitations
|
||||
|
||||
- Core ML models are fixed in terms of their inputs and outputs.
|
||||
This means you'll need to use latent images of the same size as the input of the model (512x512 is the default for
|
||||
SD1.5).
|
||||
However, you can re-convert the model to a different input size using the
|
||||
conversion nodes in this suite (set the desired width and height).
|
||||
- SD2.1 models are not supported.
|
||||
|
||||
[^1]:
|
||||
Unless [EnumeratedShapes](https://apple.github.io/coremltools/docs-guides/source/flexible-inputs.html#select-from-predetermined-shapes)
|
||||
is used during conversion. Needs more testing.
|
||||
> `(batch, tokens, hidden)` instead of the previous `(batch, hidden, 1, tokens)`.
|
||||
> Models converted with earlier versions are not compatible and must be
|
||||
> re-converted.
|
||||
|
||||
## Acknowledgements
|
||||
|
||||
The conversion pipeline began as an adaptation of Apple's
|
||||
[ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion), which
|
||||
pioneered running Stable Diffusion on the Neural Engine. It has since diverged
|
||||
and no longer depends on that package: UNet conversion runs natively on
|
||||
`diffusers`' `UNet2DConditionModel`, the ANE attention path (`SPLIT_EINSUM`,
|
||||
`SPLIT_EINSUM_V2`) is reimplemented as standalone `diffusers` attention
|
||||
processors, and the toolchain tracks current ComfyUI (NumPy 2, Torch 2.7+,
|
||||
coremltools 9, Python 3.12+). Conversion now lives in the separate
|
||||
[coreml-diffusion](https://github.com/aszc-dev/coreml-diffusion) package.
|
||||
|
||||
## Support
|
||||
|
||||
I'm here to help! If you have any questions or suggestions, don't hesitate to open an issue and I'll do my best
|
||||
to assist you.
|
||||
Questions or suggestions? Open an
|
||||
[issue](https://github.com/aszc-dev/ComfyUI-CoreMLSuite/issues).
|
||||
|
||||
@@ -0,0 +1,93 @@
|
||||
# Conversion
|
||||
|
||||
## Conversion is the only supported path
|
||||
|
||||
You always start from a Stable Diffusion checkpoint (`.safetensors` / `.ckpt`)
|
||||
and convert it with the **Convert Checkpoint to Core ML** node. The model
|
||||
version (SD1.5, SDXL, SDXL refiner, full-distill LCM) is auto-detected from the
|
||||
checkpoint. Pre-converted Core ML models from elsewhere are not supported,
|
||||
because:
|
||||
|
||||
- The suite uses its own input **dimensions**, **naming convention**, and
|
||||
**metadata**, all produced by the
|
||||
[coreml-diffusion](https://github.com/aszc-dev/coreml-diffusion) package.
|
||||
- Apple's `ml-stable-diffusion` (which most community Core ML models target) is
|
||||
effectively obsolete, and the layouts differ (see the 2.0.0
|
||||
`encoder_hidden_states` change in the README).
|
||||
- Conversion is cheap and one-time, so there is no value in maintaining
|
||||
backwards compatibility with foreign formats.
|
||||
|
||||
The output is always a **`.mlpackage`**. The suite no longer compiles to
|
||||
`.mlmodelc` (it didn't work with the inference backend), so there is **no Xcode
|
||||
or `coremlcompiler` dependency**.
|
||||
|
||||
## One-time conversion and name-based caching
|
||||
|
||||
Conversion runs **once**, not on every queue. The converter encodes all
|
||||
conversion parameters into the output filename (via `coreml_diffusion.compose_out_name`,
|
||||
called in `coreml_suite/nodes.py`):
|
||||
|
||||
- checkpoint name, `batch_size`, `width`, `height`
|
||||
- `controlnet_support`, `attention_implementation`
|
||||
- baked LoRA names, `quantize_nbits`
|
||||
|
||||
The result is written as `<encoded-name>_unet.mlpackage` in `models/unet`. If a
|
||||
file with that name already exists, it is reused and conversion is skipped. Change
|
||||
any parameter → new name → new conversion; keep them the same → the cached model
|
||||
is loaded instantly.
|
||||
|
||||
This is why the recommended workflow is **convert once, then load**: run the
|
||||
converter a single time, then in day-to-day use load the `.mlpackage` with the
|
||||
**Load Core ML UNet** node. (You can also leave the converter node in the graph;
|
||||
it short-circuits to the cached file.)
|
||||
|
||||
> [!NOTE]
|
||||
> The converter relies on the filename to decide whether to reconvert. If you
|
||||
> rename the `.mlpackage`, it will be converted again. You can otherwise rename
|
||||
> it freely if the auto-generated name is too long.
|
||||
|
||||
## Quantization
|
||||
|
||||
The converter node accepts an optional `quantize_nbits` dropdown that runs
|
||||
k-means weight palettization (`coremltools.optimize.coreml.palettize_weights`) on
|
||||
the UNet before saving.
|
||||
|
||||
Values: `none` (default — no quantization, identical output and filename to
|
||||
before), `8`, `6`, `4`. The number is appended to the `.mlpackage` stem as
|
||||
`_q<bits>`, so quantized and unquantized variants coexist on disk and in cache.
|
||||
|
||||
### SD1.5 1×512×512 SPLIT_EINSUM tradeoffs (M2 Pro, ANE)
|
||||
|
||||
Measured with 20 UNet forward passes at a fixed seed:
|
||||
|
||||
| nbits | size (MB) | size vs none | fwd median (ms) | PSNR vs `none` (dB) |
|
||||
|---|---:|---:|---:|---:|
|
||||
| none | 1641 | 1.000 | 197.1 | — |
|
||||
| 8 | 822 | 0.501 | 186.6 | 53.5 |
|
||||
| 6 | 617 | 0.376 | 183.0 | 40.2 |
|
||||
| 4 | 412 | 0.251 | 179.8 | 27.5 |
|
||||
|
||||
PSNR here is computed on the raw `noise_pred` output of a single UNet forward at a
|
||||
fixed seed, not on the final decoded image — it isolates quantization drift from
|
||||
sampler/VAE noise. Final-image PSNR is comfortably higher (the sampler averages
|
||||
over many steps).
|
||||
|
||||
### Recommended settings per chip / RAM
|
||||
|
||||
- **8 GB (M1/M2 base):** `nbits=4`. ~4× smaller, still loads, 27 dB is visually
|
||||
identical at SD1.5 sizes.
|
||||
- **16 GB (M1/M2/M3 Pro):** `nbits=6` — the sweet spot, ~2.7× smaller, 40 dB, no
|
||||
perceptible quality drop.
|
||||
- **32 GB+ (Max / Ultra):** `nbits=8` for a safety margin, or `none` for
|
||||
bit-identical output (golden testing).
|
||||
|
||||
The default stays `none`, so existing workflows produce byte-for-byte identical
|
||||
output.
|
||||
|
||||
## Where conversion lives
|
||||
|
||||
The conversion engine was extracted into the standalone
|
||||
[coreml-diffusion](https://github.com/aszc-dev/coreml-diffusion) PyPI package.
|
||||
The nodes in this suite resolve ComfyUI paths and call into it; node names,
|
||||
inputs, and outputs are unchanged, so the split has effectively no user-facing
|
||||
impact beyond `pip install` pulling one more dependency.
|
||||
+77
@@ -0,0 +1,77 @@
|
||||
# FAQ
|
||||
|
||||
## What's the difference between ANE, GPU, and MPS, and which do I pick?
|
||||
|
||||
ANE is the Neural Engine (Core ML only), GPU is the Metal GPU (Core ML or
|
||||
PyTorch), MPS is PyTorch's GPU backend. This suite uses **Core ML compute units
|
||||
only** and never touches MPS. Short answer: SD1.5 at 512×512 → convert
|
||||
`SPLIT_EINSUM`, load `CPU_AND_NE`; larger sizes or SDXL → convert `ORIGINAL`,
|
||||
load `CPU_AND_GPU`. Full reasoning: [hardware](hardware.md).
|
||||
|
||||
## Do I still need `PYTORCH_ENABLE_MPS_FALLBACK=1`?
|
||||
|
||||
Not for these nodes — Core ML inference doesn't use PyTorch MPS. It may still
|
||||
matter for other parts of your ComfyUI graph, but it has no effect on Core ML
|
||||
sampling.
|
||||
|
||||
## Why is my Core ML SDXL workflow no faster than the default nodes?
|
||||
|
||||
Because **SDXL can't run on the ANE** — the speedup comes from the Neural Engine,
|
||||
and SDXL falls back to the GPU, running at roughly MPS-equivalent speed. This is a
|
||||
known limitation, not a misconfiguration. The ANE benefit is real for SD1.5. See
|
||||
[limitations](limitations.md).
|
||||
|
||||
## Where do I get Core ML models?
|
||||
|
||||
You convert them yourself — that's the only supported path. See
|
||||
[conversion](conversion.md). Downloaded Core ML models (e.g. coreml-community) use
|
||||
different dimensions/metadata and are not supported.
|
||||
|
||||
## Is conversion run every time I queue, or once?
|
||||
|
||||
Once. Parameters are encoded in the output filename, so an already-converted model
|
||||
is reused and conversion is skipped. Convert once, then load the `.mlpackage`. See
|
||||
[conversion → caching](conversion.md#one-time-conversion-and-name-based-caching).
|
||||
|
||||
## Does a converted model produce the same output as the original?
|
||||
|
||||
With the default `quantize_nbits = none`, the converted UNet output matches the
|
||||
source within numerical rounding (the golden test in `tests/m2/test_golden_image.py`
|
||||
gates on PSNR ≥ 20 dB on the decoded image). Quantization (`8`/`6`/`4`) introduces
|
||||
measured, bounded drift — see the [PSNR table](conversion.md#quantization). For
|
||||
bit-identical output, keep `none`.
|
||||
|
||||
## Are `.mlpackage` models safe to use?
|
||||
|
||||
`.mlpackage` is a declarative Core ML model format — it carries weights and a
|
||||
compute graph, not arbitrary executable code or Python pickle, so its safety
|
||||
profile is comparable to `safetensors`. In practice this matters little here,
|
||||
since the only supported models are ones you convert locally from your own
|
||||
checkpoints.
|
||||
|
||||
## Are LoRAs reliable?
|
||||
|
||||
Partially. Some LoRAs convert cleanly; others produce poor or broken output —
|
||||
there's no firm rule, so test per-LoRA. LoRA weights and `strength_model` are
|
||||
baked in at conversion and can't be changed afterward; for some LCM-LoRA cases the
|
||||
[Core ML Adapter](nodes.md#core-ml-adapter-experimental-coremlmodeladapter) path
|
||||
is more reliable. Treat LoRA support as experimental. See
|
||||
[troubleshooting](troubleshooting.md).
|
||||
|
||||
## Does the experimental Adapter cost performance vs the Core ML Sampler?
|
||||
|
||||
Yes, a little. The Adapter wraps the model in a ComfyUI `ModelPatcher` so standard
|
||||
samplers work, which adds per-step interface overhead the native Core ML Sampler
|
||||
avoids. Use the native sampler unless you specifically need a `MODEL` (e.g.
|
||||
`ModelSamplingDiscrete` for LCM LoRAs).
|
||||
|
||||
## Which Python versions work?
|
||||
|
||||
Python 3.12 or newer (`requires-python >=3.12`). Older 3.12 install failures
|
||||
came from the now-removed `ml-stable-diffusion` build, not from this suite.
|
||||
|
||||
## Long prompts crash my workflow
|
||||
|
||||
Core ML has a hard **77-token** prompt limit and doesn't auto-chunk long prompts.
|
||||
Split the prompt across multiple CLIP Text Encode nodes and merge with Conditioning
|
||||
(Combine). See [troubleshooting](troubleshooting.md).
|
||||
@@ -0,0 +1,95 @@
|
||||
# Hardware & Compute Units
|
||||
|
||||
This page explains how the suite maps to Apple Silicon hardware, the difference
|
||||
between ANE, GPU, and MPS, and how to choose a compute unit and attention
|
||||
implementation.
|
||||
|
||||
## ANE vs GPU vs MPS
|
||||
|
||||
Three terms get conflated:
|
||||
|
||||
- **ANE (Apple Neural Engine)** — a dedicated ML accelerator on Apple Silicon.
|
||||
Only Core ML can target it; PyTorch cannot. This is the whole reason the suite
|
||||
exists.
|
||||
- **GPU** — the Metal GPU. Reachable both by Core ML (as a compute unit) and by
|
||||
PyTorch (via MPS).
|
||||
- **MPS (Metal Performance Shaders)** — PyTorch's GPU backend on macOS. This is
|
||||
the path standard ComfyUI nodes use.
|
||||
|
||||
**This suite uses Core ML compute units only — it never runs the UNet through
|
||||
PyTorch/MPS.** Consequently `PYTORCH_ENABLE_MPS_FALLBACK` has no effect on these
|
||||
nodes. It may still matter for the rest of your ComfyUI graph (CLIP, VAE,
|
||||
samplers on non-Core ML models), but not for Core ML inference itself.
|
||||
|
||||
Rough performance picture (SD1.5, maintainer- and user-reported):
|
||||
|
||||
- ANE is meaningfully faster than MPS — on the order of **50–100%** for SD1.5.
|
||||
- Core ML on the GPU is only marginally faster than PyTorch/MPS.
|
||||
|
||||
So the speedup comes from the Neural Engine, which means it depends on being able
|
||||
to actually run on the ANE (see [attention implementations](#attention-implementations)
|
||||
and the [SDXL caveat](#sdxl-and-the-ane)).
|
||||
|
||||
## Compute units
|
||||
|
||||
The **compute unit** is set on the loader/converter node and tells Core ML which
|
||||
hardware to use. It is applied when the model is loaded
|
||||
(`coreml_suite/coreml_model.py:22`), not during conversion.
|
||||
|
||||
| Value | Hardware | Best paired with |
|
||||
|---|---|---|
|
||||
| `CPU_AND_NE` (default) | CPU + Neural Engine | `SPLIT_EINSUM` / `SPLIT_EINSUM_V2` |
|
||||
| `CPU_AND_GPU` | CPU + Metal GPU | `ORIGINAL` |
|
||||
| `CPU_ONLY` | CPU only | fallback / debugging |
|
||||
| `ALL` | all available hardware | rarely optimal — see below |
|
||||
|
||||
Notes:
|
||||
|
||||
- Every option includes the CPU; there is no GPU-and-ANE-without-CPU combination.
|
||||
- `NE` in `CPU_AND_NE` is the Neural Engine (Apple's enum spells it `NE`, not
|
||||
`ANE`).
|
||||
- **`CPU_AND_NE` is often faster than `ALL`.** Letting Core ML use everything can
|
||||
be *slower* on non-Max chips, where memory bandwidth is the bottleneck. Try
|
||||
`CPU_AND_NE` first for SD1.5.
|
||||
|
||||
## Attention implementations
|
||||
|
||||
Chosen at conversion time on the **Convert Checkpoint to Core ML** node. It
|
||||
decides whether the model can run on the ANE:
|
||||
|
||||
- **`SPLIT_EINSUM`** — ANE-friendly attention. Use for the Neural Engine.
|
||||
- **`SPLIT_EINSUM_V2`** — a variant; in practice ≈ `SPLIT_EINSUM` for most users.
|
||||
- **`ORIGINAL`** — standard attention. Runs on the GPU, not the ANE.
|
||||
|
||||
The implementation and the compute unit must agree: a `SPLIT_EINSUM` model wants
|
||||
`CPU_AND_NE`; an `ORIGINAL` model wants `CPU_AND_GPU`.
|
||||
|
||||
## Which should I pick?
|
||||
|
||||
| Scenario | Attention | Compute unit |
|
||||
|---|---|---|
|
||||
| SD1.5 at 512×512 | `SPLIT_EINSUM` | `CPU_AND_NE` |
|
||||
| SD1.5 at larger sizes (e.g. 768) | `ORIGINAL` | `CPU_AND_GPU` |
|
||||
| SDXL / SDXL Turbo | `ORIGINAL` | `CPU_AND_GPU` |
|
||||
|
||||
### Resolution crossover
|
||||
|
||||
ANE shines at small latents; the GPU scales better as resolution grows. In user
|
||||
benchmarks:
|
||||
|
||||
- At **512×512**, ANE + `SPLIT_EINSUM` wins by roughly **10%** over the GPU path.
|
||||
- At **768×512**, GPU + `ORIGINAL` pulls ahead by roughly **10%**, and the larger
|
||||
image is about 2× slower overall.
|
||||
|
||||
If you mostly work at 512×512, convert with `SPLIT_EINSUM` and load on
|
||||
`CPU_AND_NE`. If you routinely go larger, an `ORIGINAL` + GPU model may be
|
||||
faster.
|
||||
|
||||
### SDXL and the ANE
|
||||
|
||||
SDXL (and SDXL Turbo) **cannot run on the ANE** — the dual-text-encoder UNet
|
||||
exceeds what the Neural Engine path supports. SDXL therefore runs at roughly
|
||||
MPS-equivalent speed with no ANE speedup. If a Core ML SDXL workflow feels no
|
||||
faster than the standard nodes, this is why. Convert SDXL with `ORIGINAL` and
|
||||
load with `CPU_AND_GPU` or `CPU_ONLY`. See
|
||||
[limitations](limitations.md) for the full picture.
|
||||
@@ -0,0 +1,53 @@
|
||||
# Limitations & Support Matrix
|
||||
|
||||
## Support matrix
|
||||
|
||||
| Feature | Status | Notes |
|
||||
|---|---|---|
|
||||
| SD1.5 | ✅ Full | ANE via `SPLIT_EINSUM`; the primary, fastest path |
|
||||
| SDXL / SDXL Turbo | ⚠️ Partial | GPU only (no ANE), no speedup; possible quality loss vs source. Don't run Turbo at 1024² |
|
||||
| SD2.1 | ❌ Unsupported | |
|
||||
| Inpainting checkpoints (9-channel) | ❌ Unsupported | |
|
||||
| ControlNet | ✅ Supported | Convert the checkpoint with `controlnet_support = True` |
|
||||
| LoRA | ⚠️ Experimental | Inconsistent per-LoRA; baked at conversion, immutable afterward |
|
||||
| LCM | ⚠️ Experimental | Full-distill LCM checkpoints auto-detected by the converter |
|
||||
| SVD | ❌ Not supported | |
|
||||
| AnimateDiff | ❌ Not supported | Motion modules need pre-conversion injection; not feasible today |
|
||||
| IPAdapter | ❌ Not supported | Needs a real `MODEL` the Core ML wrapper can't provide |
|
||||
| Core ML Adapter | ⚠️ Experimental | Works for many nodes; fails for merges/IPAdapter/etc. |
|
||||
|
||||
## Fixed input/output shapes
|
||||
|
||||
A Core ML model is converted for one specific resolution and batch size. To work
|
||||
at a different size, re-convert with the new width/height (conversion is cheap and
|
||||
cached by name). This is also why detailers and latent-upscale workflows that
|
||||
rescale mid-graph break — see [troubleshooting](troubleshooting.md).
|
||||
|
||||
There is experimental support for flexible shapes via
|
||||
[EnumeratedShapes](https://apple.github.io/coremltools/docs-guides/source/flexible-inputs.html#select-from-predetermined-shapes),
|
||||
but it is **much slower** — user benchmarks show roughly **5×** the per-iteration
|
||||
time on every run, not just the first. Fixed-shape models per resolution are the
|
||||
practical choice.
|
||||
|
||||
## SDXL on the Neural Engine
|
||||
|
||||
SDXL and SDXL Turbo cannot run on the ANE — the dual-text-encoder UNet exceeds the
|
||||
supported Neural Engine path. They run on the GPU at roughly MPS-equivalent speed,
|
||||
so Core ML offers no speed advantage for SDXL, and converted output may look
|
||||
degraded versus the safetensors original (an upstream conversion artifact). Use
|
||||
`ORIGINAL` + `CPU_AND_GPU`. See [hardware](hardware.md).
|
||||
|
||||
## Experimental Core ML Adapter
|
||||
|
||||
The Adapter wraps a Core ML model to look like a standard ComfyUI `MODEL`, which
|
||||
covers many standard and custom nodes. But it can't fully emulate a real model:
|
||||
operations that need genuine `MODEL` internals — model merges, IPAdapter, some
|
||||
LoRA flows, detailers — generally won't work, and the model's
|
||||
fixed input shapes aren't validated, so mismatches error at runtime. Prefer the
|
||||
native Core ML Sampler when you don't need the `MODEL` type.
|
||||
|
||||
## Prompt length
|
||||
|
||||
Core ML enforces a hard 77-token prompt limit with no auto-chunking. Split long
|
||||
prompts across multiple CLIP Text Encode nodes and merge with Conditioning
|
||||
(Combine).
|
||||
+157
@@ -0,0 +1,157 @@
|
||||
# Node Reference
|
||||
|
||||
All nodes live in the **Core ML Suite** category. Right-click the canvas →
|
||||
**Add Node → Core ML Suite**, or double-click and search.
|
||||
|
||||
| Display name | Class | Purpose |
|
||||
|---|---|---|
|
||||
| Load Core ML UNet | `CoreMLUNetLoader` | Load a converted `.mlpackage` |
|
||||
| Core ML Sampler | `CoreMLSampler` | Sample (KSampler-style) |
|
||||
| Core ML Sampler (Advanced) | `CoreMLSamplerAdvanced` | Sample (KSamplerAdvanced-style) |
|
||||
| Core ML Adapter (Experimental) | `CoreMLModelAdapter` | Wrap as a standard `MODEL` |
|
||||
| Load LoRA to use with Core ML | `Core ML LoRA Loader` | Bake LoRA(s) at conversion |
|
||||
| Convert Checkpoint to Core ML | `Core ML Converter` | Convert a checkpoint |
|
||||
|
||||
---
|
||||
|
||||
## Load Core ML UNet (`CoreMLUNetLoader`)
|
||||
|
||||

|
||||
|
||||
Loads a converted `.mlpackage` from `models/unet` and outputs a `coreml_model`
|
||||
for the samplers. Only `.mlpackage` files are listed — this suite no longer uses
|
||||
`.mlmodelc`.
|
||||
|
||||
- **Inputs**
|
||||
- `coreml_name` — the `.mlpackage` to load from `models/unet`.
|
||||
- `compute_unit` — hardware to run on: `CPU_AND_NE` (default), `CPU_AND_GPU`,
|
||||
`CPU_ONLY`, `ALL`. See [hardware](hardware.md).
|
||||
- **Output**
|
||||
- `coreml_model` — for the Core ML Sampler or Adapter.
|
||||
|
||||
---
|
||||
|
||||
## Core ML Sampler (`CoreMLSampler`)
|
||||
|
||||

|
||||
|
||||
Generates a latent from a Core ML model. Behaves like the standard KSampler and
|
||||
outputs a `LATENT` you can decode or feed downstream.
|
||||
|
||||
- **Inputs**
|
||||
- `coreml_model` — output of the loader or a converter.
|
||||
- `latent_image` *(optional)* — must match the model's input size. If omitted,
|
||||
a suitable empty latent is created. Provide one for img2img.
|
||||
- `negative` *(optional)* — required for normal models; optional for LCM.
|
||||
- Remaining inputs (`seed`, `steps`, `cfg`, `sampler_name`, `scheduler`,
|
||||
`positive`, `denoise`) match the KSampler.
|
||||
- **Output**
|
||||
- `LATENT` — decode with a VAE Decode, or use downstream.
|
||||
|
||||
---
|
||||
|
||||
## Core ML Sampler (Advanced) (`CoreMLSamplerAdvanced`)
|
||||
|
||||
The KSamplerAdvanced counterpart of the Core ML Sampler — same Core ML input,
|
||||
plus the advanced sampling controls. Use it for partial denoising, fixed noise,
|
||||
and multi-stage (e.g. SDXL base → refiner) workflows.
|
||||
|
||||
- **Inputs**
|
||||
- `coreml_model` — output of the loader or a converter.
|
||||
- `add_noise`, `noise_seed`, `start_at_step`, `end_at_step`,
|
||||
`return_with_leftover_noise` — as in KSamplerAdvanced.
|
||||
- `steps`, `cfg`, `sampler_name`, `scheduler`, `positive` — as usual.
|
||||
- `latent_image` *(optional)*, `negative` *(optional, required for non-LCM)*.
|
||||
- **Output**
|
||||
- `LATENT`.
|
||||
|
||||
---
|
||||
|
||||
## Core ML Adapter (Experimental) (`CoreMLModelAdapter`)
|
||||
|
||||

|
||||
|
||||
Wraps a Core ML model so it presents as a standard ComfyUI `MODEL`, letting you
|
||||
feed it to the normal KSampler and many other nodes (e.g. `ModelSamplingDiscrete`
|
||||
for LCM LoRAs).
|
||||
|
||||
- **Input**
|
||||
- `coreml_model`.
|
||||
- **Output**
|
||||
- `MODEL` — a Core ML model wrapped as a ComfyUI model.
|
||||
|
||||
> [!NOTE]
|
||||
> Experimental. The wrapper presents a `MODEL` interface but cannot fully
|
||||
> emulate one — model merges, IPAdapter, and similar advanced uses generally
|
||||
> won't work, and the model's fixed input shapes are not validated, so mismatched
|
||||
> inputs error at runtime. The native Core ML Sampler is faster when you don't
|
||||
> need the `MODEL` type. See the [FAQ](faq.md) and [limitations](limitations.md).
|
||||
|
||||
---
|
||||
|
||||
## Load LoRA to use with Core ML (`Core ML LoRA Loader`)
|
||||
|
||||

|
||||
|
||||
Collects LoRA name + `strength_model` to bake into the model at conversion, and
|
||||
applies the LoRA to CLIP (which is not part of the Core ML path). Chain multiple
|
||||
loaders for multiple LoRAs.
|
||||
|
||||
Because a converted model is immutable, the baked weights and `strength_model`
|
||||
**cannot** be changed afterward — changing them means re-converting. `strength_clip`
|
||||
only affects CLIP and can be changed freely. After conversion, when loading with
|
||||
`CoreMLUNetLoader`, apply the same LoRAs to CLIP manually (see
|
||||
[workflows](workflows.md)).
|
||||
|
||||
- **Inputs**
|
||||
- `lora_name`, `strength_model`, `strength_clip`.
|
||||
- `clip` — from `CLIPLoader` / `CheckpointLoaderSimple` or another LoRA loader.
|
||||
- `lora_params` *(optional)* — chain from another LoRA loader.
|
||||
- **Outputs**
|
||||
- `CLIP` — with the LoRA applied.
|
||||
- `lora_params` — pass to the converter or the next LoRA loader.
|
||||
|
||||
> [!NOTE]
|
||||
> LoRA support is experimental and inconsistent — some LoRAs convert cleanly,
|
||||
> others produce poor results. Test per-LoRA. See [troubleshooting](troubleshooting.md).
|
||||
|
||||
---
|
||||
|
||||
## Convert Checkpoint to Core ML (`Core ML Converter`)
|
||||
|
||||

|
||||
|
||||
Converts a checkpoint from `models/checkpoints` to a Core ML `.mlpackage` in
|
||||
`models/unet`. The model version (SD1.5, SDXL, SDXL refiner, or full-distill
|
||||
LCM) is auto-detected from the checkpoint's architecture — there is no version
|
||||
dropdown. The conversion parameters are encoded in the output name, so an
|
||||
already-converted model is reused instead of re-converted. See
|
||||
[conversion](conversion.md) for details.
|
||||
|
||||
- **Inputs**
|
||||
- `ckpt_name` — checkpoint in `models/checkpoints`.
|
||||
- `height`, `width` — target image size; any positive multiple of 8 (default
|
||||
512). The model's input size is fixed at these values.
|
||||
- `batch_size` — default 1; raise to convert a batch-capable model.
|
||||
- `attention_implementation` — `SPLIT_EINSUM` / `SPLIT_EINSUM_V2` (ANE) or
|
||||
`ORIGINAL` (GPU). See [hardware](hardware.md).
|
||||
- `compute_unit` — used only when loading the result; does not affect
|
||||
conversion.
|
||||
- `controlnet_support` — set `True` to make the model usable with ControlNet
|
||||
(default `False`).
|
||||
- `quantize_nbits` *(optional)* — `none` (default), `8`, `6`, `4`. See
|
||||
[conversion → quantization](conversion.md#quantization).
|
||||
- `lora_params` *(optional)* — from the LoRA loader, to bake LoRAs in.
|
||||
- **Output**
|
||||
- `coreml_model`.
|
||||
|
||||
> [!NOTE]
|
||||
> Some checkpoints need a custom config `.yaml`. Place it in `models/configs`
|
||||
> named like the checkpoint (e.g. `juggernaut.safetensors` →
|
||||
> `juggernaut.yaml`); it is loaded automatically during conversion.
|
||||
|
||||
> [!NOTE]
|
||||
> Full-distill LCM checkpoints (e.g.
|
||||
> [LCM_Dreamshaper_v7](https://huggingface.co/SimianLuo/LCM_Dreamshaper_v7)) are
|
||||
> detected and converted like any other checkpoint. When sampling an LCM model,
|
||||
> set `sampler_name` to `lcm` and `scheduler` to `sgm_uniform`.
|
||||
@@ -0,0 +1,86 @@
|
||||
# Troubleshooting
|
||||
|
||||
## `Expected shape … got …` / latent size mismatch
|
||||
|
||||
The most common error. A Core ML model has **fixed** input dimensions — a model
|
||||
converted for 512×512 expects a 64×64 latent and rejects any other size (batch
|
||||
size is handled and doesn't matter; only width/height are fixed).
|
||||
|
||||
**Fix:** set your Empty Latent (or upstream latent) to exactly the resolution the
|
||||
model was converted for, or re-convert at the size you want.
|
||||
|
||||
## Old `.mlmodelc` model, or `metadata.json` not found
|
||||
|
||||
This suite no longer produces or loads `.mlmodelc`; the loader lists `.mlpackage`
|
||||
only. Models from an older version (or downloaded community models) with a
|
||||
`.mlmodelc` structure won't load.
|
||||
|
||||
**Fix:** re-convert the checkpoint with **Convert Checkpoint to Core ML**. No
|
||||
Xcode or `coremlcompiler` is required — that dependency was removed.
|
||||
|
||||
## Prompt too long (`Expected size 154 but got 77`, or a crash)
|
||||
|
||||
Core ML enforces a hard **77-token** prompt limit and does not auto-chunk like
|
||||
A1111/ComfyUI.
|
||||
|
||||
**Fix:** split the prompt across multiple CLIP Text Encode nodes and merge them
|
||||
with **Conditioning (Combine)**.
|
||||
|
||||
## `cannot import name 'ModelSamplingDiscreteLCM'`
|
||||
|
||||
A ComfyUI refactor renamed this symbol.
|
||||
|
||||
**Fix:** update the suite (fixed in PR #29) and re-run
|
||||
`pip install -r requirements.txt`.
|
||||
|
||||
## LoRA loader `ImportError`
|
||||
|
||||
`peft` became a required dependency.
|
||||
|
||||
**Fix:** `pip install -r requirements.txt`. This recurs after ComfyUI-Manager
|
||||
updates if requirements aren't reinstalled.
|
||||
|
||||
## ControlNet has no effect
|
||||
|
||||
ControlNet support is baked at conversion. If the checkpoint was converted with
|
||||
`controlnet_support = False`, ControlNet does nothing.
|
||||
|
||||
**Fix:** re-convert with `controlnet_support = True`. The ControlNet model itself
|
||||
needs no conversion, and `.fp16.safetensors` vs `.safetensors` makes no
|
||||
difference.
|
||||
|
||||
## LoRAs produce garbage
|
||||
|
||||
LoRA support is inconsistent — some work, some don't, with no firm rule. Test
|
||||
per-LoRA. For some LCM-LoRA setups, routing through the
|
||||
[Core ML Adapter](nodes.md#core-ml-adapter-experimental-coremlmodeladapter) is
|
||||
more reliable than the basic loader path. Remember weights are baked at conversion
|
||||
and can't be changed afterward.
|
||||
|
||||
## FaceDetailer / detailers error on size
|
||||
|
||||
Detailers rescale latents internally (e.g. 512 → 1024), which breaks the model's
|
||||
fixed input shape. There is no workaround node — a Core ML model only accepts
|
||||
the resolution it was converted for.
|
||||
|
||||
**Fix:** convert a second model at the detailer's internal resolution and use it
|
||||
for the detailing pass, or run the detailer with a standard (non–Core ML) model.
|
||||
|
||||
## Inpainting checkpoint errors (`tensor size 9 vs 4`)
|
||||
|
||||
SD1.5 inpainting checkpoints use a 9-channel input and are **not supported**. This
|
||||
error is expected, not a bug.
|
||||
|
||||
## Errors mentioning `python_coreml_stable_diffusion` or `ml-stable-diffusion`
|
||||
|
||||
You're on a stale install. That dependency was removed; old install scripts tried
|
||||
`pip install git+…/ml-stable-diffusion.git`, which fails on modern Python.
|
||||
|
||||
**Fix:** reinstall the current suite (`pip install -r requirements.txt`, which
|
||||
pulls `coreml-diffusion` from PyPI).
|
||||
|
||||
## `all input tensors must be on the same device (mps:0 and cpu)` / ControlNet residual shape `(2,…) vs (1,…)`
|
||||
|
||||
Old bugs that have been fixed.
|
||||
|
||||
**Fix:** update to the latest version.
|
||||
@@ -0,0 +1,103 @@
|
||||
# Example Workflows
|
||||
|
||||
> [!NOTE]
|
||||
> The models referenced are examples — substitute your own. Every workflow
|
||||
> starts from a checkpoint you convert yourself (see [conversion](conversion.md));
|
||||
> there is no Core ML model to download.
|
||||
|
||||
## Basic txt2img
|
||||
|
||||
Convert a SD1.5 checkpoint, then sample from it. CLIP and VAE come from standard
|
||||
ComfyUI nodes — either loaded separately or pulled from the checkpoint.
|
||||
|
||||
1. Place a SD1.5 checkpoint in `models/checkpoints` (e.g.
|
||||
[v1-5-pruned-emaonly](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/v1-5-pruned-emaonly.safetensors)).
|
||||
2. **Convert Checkpoint to Core ML** → queue once → a `.mlpackage` lands in
|
||||
`models/unet`.
|
||||
3. **Load Core ML UNet** (or wire the converter output straight in) →
|
||||
**Core ML Sampler** → **VAE Decode**.
|
||||
|
||||
**CLIP and VAE from the checkpoint:**
|
||||
|
||||

|
||||
|
||||
**CLIP and VAE loaded separately** — use any SD1.5-compatible
|
||||
[CLIP](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/text_encoder/model.safetensors)
|
||||
and [VAE](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/vae/diffusion_pytorch_model.safetensors),
|
||||
placed in `models/clip` and `models/vae`:
|
||||
|
||||

|
||||
|
||||
## ControlNet
|
||||
|
||||
Convert the checkpoint with `controlnet_support = True`, then wire a standard
|
||||
ComfyUI ControlNet. The ControlNet model itself needs no conversion. Place it in
|
||||
`models/controlnet` (e.g.
|
||||
[control_v11p_sd15_scribble](https://huggingface.co/lllyasviel/control_v11p_sd15_scribble/blob/main/diffusion_pytorch_model.fp16.safetensors)).
|
||||
|
||||

|
||||
|
||||
## Checkpoint conversion
|
||||
|
||||
The minimal conversion graph. See
|
||||
[Convert Checkpoint to Core ML](nodes.md#convert-checkpoint-to-core-ml-core-ml-converter).
|
||||
|
||||

|
||||
|
||||
## Conversion with LoRA
|
||||
|
||||
Bake LoRA(s) into the model at conversion. Read the
|
||||
[LoRA caveats](nodes.md#load-lora-to-use-with-core-ml-core-ml-lora-loader) first
|
||||
— baked weights are immutable, and support is inconsistent per-LoRA.
|
||||
|
||||

|
||||
|
||||
## LCM LoRA conversion
|
||||
|
||||
Chain multiple LoRA loaders to use several LoRAs with one model.
|
||||
|
||||
> [!IMPORTANT]
|
||||
> Here the model goes through the **Core ML Adapter** and `ModelSamplingDiscrete`
|
||||
> into the standard ComfyUI KSampler (not the Core ML Sampler).
|
||||
> `ModelSamplingDiscrete` is required to sample LCM LoRAs correctly.
|
||||
|
||||

|
||||
|
||||
## Loading a model with baked LoRAs
|
||||
|
||||
Load a model that already has LoRAs baked in. CLIP must be loaded separately and
|
||||
passed through the same LoRA nodes used at conversion. Since `lora_name` and
|
||||
`strength_model` are baked in, they need not be passed to the loader.
|
||||
|
||||
> [!IMPORTANT]
|
||||
> As above, the model goes through the Core ML Adapter + `ModelSamplingDiscrete`
|
||||
> into the standard KSampler.
|
||||
|
||||

|
||||
|
||||
## LCM conversion with ControlNet
|
||||
|
||||
Convert a full-distill LCM checkpoint (e.g.
|
||||
[LCM_Dreamshaper_v7](https://huggingface.co/SimianLuo/LCM_Dreamshaper_v7)) with
|
||||
the standard **Convert Checkpoint to Core ML** node — the LCM architecture is
|
||||
auto-detected. Use it with or without ControlNet. When sampling, set
|
||||
`sampler_name` to `lcm` and `scheduler` to `sgm_uniform`.
|
||||
|
||||

|
||||
|
||||
## SDXL Base + Refiner
|
||||
|
||||
A basic SDXL graph. Add LoRAs and ControlNets as in the SD1.5 examples; the
|
||||
refiner step is optional.
|
||||
|
||||
Models:
|
||||
[base + text encoders](https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0),
|
||||
[refiner](https://huggingface.co/stabilityai/stable-diffusion-xl-refiner-1.0),
|
||||
[VAE](https://huggingface.co/stabilityai/sdxl-vae).
|
||||
|
||||
> [!IMPORTANT]
|
||||
> SDXL does not run on the ANE. Convert with `ORIGINAL` and load with
|
||||
> `CPU_AND_GPU` (or `CPU_ONLY`). If loading hangs on `CPU_AND_NE`, that is the
|
||||
> cause. See [limitations](limitations.md).
|
||||
|
||||

|
||||
@@ -1,208 +0,0 @@
|
||||
# Conversion Extraction — Seam Inventory (`docs/extraction/seam.md`)
|
||||
|
||||
> **Gate E0 deliverable.** Symbol-by-symbol cut line between the future `coreml_diffusion`
|
||||
> package (CONVERSION) and what stays in `coreml_suite` (the ComfyUI side).
|
||||
>
|
||||
> **Confidence legend:**
|
||||
> - ✅ **verified** — read directly from the current source in this repo.
|
||||
> - 🔍 **confirm** — inferred / partially seen; Claude Code must `grep`-verify before acting.
|
||||
>
|
||||
> **Cut rule:** a symbol goes to `coreml_diffusion` iff it participates in producing the `.mlpackage`
|
||||
> artifact AND can be made free of `comfy` / `folder_paths` / `comfy_extras`. The runtime
|
||||
> *loader* that **runs** a compiled model stays in the suite.
|
||||
|
||||
---
|
||||
|
||||
## 1. File-level map
|
||||
|
||||
| File | Side | Status | Note |
|
||||
|---|---|---|---|
|
||||
| `coreml_suite/model_version.py` | **coreml_diffusion** | ✅ | Already `Enum`-only, zero comfy. Becomes pkg source of truth. |
|
||||
| `coreml_suite/attention.py` | **coreml_diffusion** | ✅ | `ATTENTION_IMPLEMENTATIONS` tuple; pure constant. |
|
||||
| `coreml_suite/core/naming.py` | **coreml_diffusion** | ✅ | `compose_out_name` = cache-key contract. Move (not copy). |
|
||||
| `coreml_suite/converter.py` | **coreml_diffusion** (mostly) | ✅ | Main conversion. One symbol stays-adjacent: `get_out_path` (folder_paths) is replaced by injected `out_path`. |
|
||||
| `coreml_suite/conversion/attention.py` | **coreml_diffusion** | ✅ | `apply_attention_implementation`. Imports `logging`,`torch` only — no comfy. |
|
||||
| `coreml_suite/conversion/shapes.py` | **coreml_diffusion** | ✅ | `conv2d_output_shape`. Pure math, no imports. |
|
||||
| `coreml_suite/conversion/trace.py` | **coreml_diffusion** | ✅ | Imports `types.MethodType`, `diffusers...Transformer2DModel` only — torch/diffusers. |
|
||||
| `coreml_suite/conversion/unet.py` | **coreml_diffusion** | ✅ | `CoreMLUNetWrapper`. Imports `torch` only — no comfy. |
|
||||
| `coreml_suite/lcm/converter.py` | **coreml_diffusion** (after dedup) | ✅ | Dup helpers deleted; `MODEL_VERSION` HF-hardcode (L22) → E-LCM. `folder_paths` (L111) + `comfy.model_management` (L54) confirmed present → CUT. |
|
||||
| `coreml_suite/lcm/unet.py` | **coreml_diffusion** | ✅ | `UNet2DConditionModelLCM(UNet2DConditionModel)`. diffusers-only, no comfy. |
|
||||
| `coreml_suite/config.py` | **STAYS** | ✅ | Imports `comfy.supported_models_base`/`latent_formats`/`model_detection`. **Inference-side** (`get_model_config`), NOT conversion. |
|
||||
| `coreml_suite/coreml_model.py` | **STAYS** | ✅ | `CoreMLModel` = runtime loader (runs `.mlpackage`). Desktop/Python inference; not used on iOS. |
|
||||
| `coreml_suite/nodes.py` | **STAYS** | ✅ | Nodes; will call `coreml_diffusion` + own `folder_paths` path resolution + discovery dropdowns. |
|
||||
| `coreml_suite/lcm/nodes.py` | **STAYS** | ✅ | `COREML_CONVERT_LCM` node. |
|
||||
| `coreml_suite/models.py` | **STAYS** | ✅ | Inference: `add_sdxl_model_options`, `is_sdxl`, `get_model_patcher`, `get_latent_image`. |
|
||||
| `coreml_suite/latents.py` | **STAYS** | ✅ | Inference chunking (MODERNIZATION Phase 3 target, not this spec). |
|
||||
| `coreml_suite/controlnet.py` | **STAYS** | ✅ | Inference-side controlnet. Distinct from converter `add_cnet_support`. |
|
||||
| `coreml_suite/lcm/utils.py` | **STAYS** | ✅ | `add_lcm_model_options`, `lcm_patch`, `is_lcm`; imports `comfy_extras`. Inference. |
|
||||
| `coreml_suite/logger.py` | **both / copy** | ✅ | Trivial. Package gets its own logger; suite keeps its. |
|
||||
|
||||
---
|
||||
|
||||
## 2. Symbol-level: `coreml_suite/converter.py` (main conversion)
|
||||
|
||||
| Symbol | Side | Status | Cut action |
|
||||
|---|---|---|---|
|
||||
| `DEFAULT_TRACE_TIMESTEP`, `TEXT_TOKEN_SEQUENCE_LENGTH` | coreml_diffusion | ✅ | Move as-is (module constants). |
|
||||
| `get_unet(model_version, ref_unet, attention_implementation)` | coreml_diffusion | ✅ | Move. Uses `conversion.{trace,attention,unet}`. No comfy. |
|
||||
| `get_encoder_hidden_states_shape(ref_unet, batch_size)` | coreml_diffusion | ✅ | Move. Reads `ref_unet.config.cross_attention_dim`. Pure. |
|
||||
| `get_coreml_inputs(sample_inputs)` | coreml_diffusion | ✅ | Move. `ct.TensorType` build. |
|
||||
| `load_coreml_model(out_path)` | coreml_diffusion | ✅ | Move. `ct.models.MLModel(out_path)`. (Dedup target vs LCM copy.) |
|
||||
| `convert_to_coreml(submodule, ts_module, inputs, names, out_path)` | coreml_diffusion | ✅ | Move. `ct.convert(...)`. (Dedup target vs LCM copy.) |
|
||||
| `get_sample_input(batch, ehs_shape, sample_shape)` | coreml_diffusion | ✅ | Move. **Merge** with LCM variant (LCM passes extra `scheduler` → optional param). |
|
||||
| `lcm_inputs(sample_unet_inputs)` | coreml_diffusion | ✅ | Move. Adds `timestep_cond`. |
|
||||
| `sdxl_inputs(sample_unet_inputs, ref_unet, model_version)` | coreml_diffusion | ✅ | Move. `time_ids`/`text_embeds`/`add_embeds`. |
|
||||
| `add_cnet_support(sample_shape, ref_unet)` | coreml_diffusion | ✅ | Move. Builds `additional_residual_*` inputs from unet block channels. |
|
||||
| `convert_unet(ref_unet, model_version, unet_out_path, ...)` | coreml_diffusion | ✅ | Move. Orchestrates trace→convert→**quant (palettize)**→save. Quant travels here (E6). |
|
||||
| `convert(ckpt_path, model_version, unet_out_path, ...)` | coreml_diffusion | ✅ | Move. **Make kw-only past `ckpt_path,model_version,out_path`** (contract). Validates `attn_impl`. |
|
||||
| `load_unet(ckpt_path, config_path)` | coreml_diffusion | ✅ | Move. `UNet2DConditionModel.from_single_file`. |
|
||||
| `get_out_path(submodule_name, model_name)` | **STAYS (node)** | ✅ | Uses `folder_paths.get_folder_paths`. **Delete from converter; node resolves path and passes `out_path` in.** |
|
||||
|
||||
**Apple `python_coreml_stable_diffusion` footprint on this path:** ✅ **none.** Verified by grep:
|
||||
zero imports in `converter.py` / `conversion/*`. Main path uses `diffusers` +
|
||||
local `CoreMLUNetWrapper`. (And the runtime `CoreMLModel` is now a local coremltools wrapper too —
|
||||
see §6 stale-spec note.)
|
||||
|
||||
---
|
||||
|
||||
## 3. Symbol-level: `coreml_suite/lcm/converter.py` (LCM — dedup + defer)
|
||||
|
||||
| Symbol | Side | Status | Cut action |
|
||||
|---|---|---|---|
|
||||
| `load_coreml_model` (LCM copy) | DELETE | ✅ | Duplicate of main. Remove; use `coreml_diffusion.load_coreml_model`. |
|
||||
| `convert_to_coreml` (LCM copy) | DELETE | ✅ | Duplicate of main. Remove. |
|
||||
| `get_out_path` (LCM copy, folder_paths) | DELETE | ✅ | Duplicate + comfy. Remove; node injects `out_path`. |
|
||||
| `get_sample_input(..., scheduler)` (LCM copy) | MERGE → coreml_diffusion | ✅ | Fold `scheduler` into shared `get_sample_input` as optional param. |
|
||||
| `MODEL_NAME` (= LCM_Dreamshaper) | **E-LCM** | ✅ | HF hardcode. Removing it is the behavior change → E-LCM, not E2. |
|
||||
| `convert(out_path, sample_size, batch_size, controlnet_support)` (LCM, L190) | coreml_diffusion (via unified) | ✅ | Route through `coreml_diffusion.convert(model_version=LCM, ...)` in E-LCM. |
|
||||
| `from comfy.model_management import get_torch_device` (L54, in `get_scheduler`) | **CUT** | ✅ | Confirmed present. Inject `device`. |
|
||||
| module-global attention set at import | n/a | ✅ | **No module global.** Attention already per-call: `get_unets` (L36) calls `apply_attention_implementation(ref_unet, "SPLIT_EINSUM")`. No `ATTENTION_IMPLEMENTATION_IN_EFFECT` anywhere in repo. (Note: LCM hardcodes `"SPLIT_EINSUM"` — pass `attn_impl` through in dedup.) |
|
||||
|
||||
---
|
||||
|
||||
## 4. Symbol-level: `coreml_suite/core/naming.py` → `coreml_diffusion/naming.py`
|
||||
|
||||
| Symbol | Side | Status | Cut action |
|
||||
|---|---|---|---|
|
||||
| `compose_out_name(...)` | coreml_diffusion | ✅ | **Move** (cache-key contract). Node imports from pkg. |
|
||||
| `lora_names_from_params(...)` | coreml_diffusion | ✅ | Move. |
|
||||
| `ATTN_SUFFIX` dict | coreml_diffusion | ✅ | Move. |
|
||||
| `QUANT_NBITS_VALUES` | coreml_diffusion | ✅ | Move; backs `list_quant_modes()`. |
|
||||
| `tests/unit/test_characterization_out_name.py` | re-point | ✅ | Change import to `coreml_diffusion.naming`. Assertions/values **unchanged**. |
|
||||
|
||||
---
|
||||
|
||||
## 5. Discovery API + status registry (new in `coreml_diffusion/__init__.py`)
|
||||
|
||||
```python
|
||||
from enum import Enum
|
||||
|
||||
class Status(Enum):
|
||||
VERIFIED = "verified" # has a golden anchor + passing [M2-ANE] check
|
||||
EXPERIMENTAL = "experimental" # convertible, not yet anchored/verified
|
||||
|
||||
# Single source of truth. Suite gates on this, NOT on a hardcoded node list.
|
||||
# KEY by ModelVersion enum MEMBER (not a bare string) so list_model_versions can
|
||||
# emit .name — see the .name decision below. Keying by the lowercase .value string
|
||||
# (as an earlier draft of this block did) returns ["sd15",...], which the node then
|
||||
# reverses via ModelVersion[...] → KeyError. Do NOT key by .value.
|
||||
_MODEL_STATUS = {
|
||||
ModelVersion.SD15: Status.VERIFIED,
|
||||
ModelVersion.SDXL: Status.VERIFIED,
|
||||
ModelVersion.SDXL_REFINER: Status.EXPERIMENTAL, # → VERIFIED after a refiner golden anchor
|
||||
ModelVersion.LCM: Status.EXPERIMENTAL, # → VERIFIED after E-LCM golden anchor
|
||||
}
|
||||
|
||||
def list_model_versions(include_experimental: bool = False) -> list[str]:
|
||||
return [v.name for v, s in _MODEL_STATUS.items() # .name → "SD15","SDXL" (see decision)
|
||||
if s is Status.VERIFIED or (include_experimental and s is Status.EXPERIMENTAL)]
|
||||
|
||||
def list_attention_impls() -> list[str]: # from attention.ATTENTION_IMPLEMENTATIONS
|
||||
...
|
||||
def list_quant_modes() -> list[str]: # from naming.QUANT_NBITS_VALUES
|
||||
...
|
||||
|
||||
CONTRACT_VERSION = "1.0"
|
||||
# Additive-only: adding an id or promoting EXPERIMENTAL→VERIFIED = minor bump (Suite unaffected).
|
||||
# Removing/renaming an id, or demoting VERIFIED→EXPERIMENTAL = MAJOR bump + migration note.
|
||||
```
|
||||
|
||||
**Decision check (`.name` vs `.value`): RESOLVED → `.name`.** ✅
|
||||
Verified in current source:
|
||||
- Node renders `ModelVersion.SD15.name` / `ModelVersion.SDXL.name` → `"SD15"`, `"SDXL"`
|
||||
(`nodes.py:224-225`).
|
||||
- Node reverses the dropdown string with `model_version = ModelVersion[model_version]`
|
||||
(`nodes.py:286`) — i.e. **lookup by NAME**. Feeding it a `.value` (`"sd15"`) raises `KeyError`.
|
||||
- Enum values are lowercase (`model_version.py`: `SD15="sd15"`, `SDXL="sdxl"`,
|
||||
`SDXL_REFINER="sdxl_refiner"`, `LCM="lcm"`).
|
||||
- `compose_out_name` does NOT consume the model_version string (grep of `core/naming.py` empty) —
|
||||
no coupling there, so no constraint from that side.
|
||||
|
||||
**Decision:** `list_model_versions()` returns `.name` (uppercase). Saved workflows store `"SD15"`,
|
||||
node already validates them via `ModelVersion[...]`. The `_MODEL_STATUS` block above was corrected
|
||||
to key by enum member and emit `.name`. **The earlier `v.value` form was a latent bug.**
|
||||
|
||||
---
|
||||
|
||||
## 6. `python_coreml_stable_diffusion` split (Gate E0 line to fill by grep)
|
||||
|
||||
| Use | Side | Status |
|
||||
|---|---|---|
|
||||
| `coreml_model.CoreMLModel` (runs compiled model) | **STAYS** (suite runtime) | ✅ — **local class**, not Apple's |
|
||||
| `unet.UNet2DConditionModel*` internals | **gone** — `converter.py:319` uses `diffusers.UNet2DConditionModel.from_single_file` | ✅ |
|
||||
| `AttentionImplementations` enum | gone — local `apply_attention_implementation` + `attention.py` tuple | ✅ |
|
||||
| `calculate_conv2d_output_shape` | gone — replaced by `conversion/shapes.conv2d_output_shape` | ✅ |
|
||||
|
||||
> ### ⚠️ SPEC IS STALE: `ml-stable-diffusion` is already fully removed
|
||||
> Commit #58 ("replace apple/ml-stable-diffusion with native diffusers conversion") already did
|
||||
> the de-Apple work. Verified now:
|
||||
> - **Zero** `python_coreml_stable_diffusion` runtime imports anywhere in `coreml_suite` (only a
|
||||
> docstring mention at `core/__init__.py:4`).
|
||||
> - `coreml_suite/coreml_model.py:8` `CoreMLModel` is a **local** wrapper over
|
||||
> `coremltools.models.MLModel` (`coreml_model.py:22`) — it does **not** import Apple's class.
|
||||
> - `ml-stable-diffusion` / `python_coreml_stable_diffusion` appears in **neither** `pyproject.toml`
|
||||
> **nor** `requirements.txt`. It is not a dependency at all.
|
||||
>
|
||||
> **Consequences for the spec (correct these in CONVERTER_EXTRACTION_SPEC.md):**
|
||||
> - §0.3 premise ("runtime loader = `python_coreml_stable_diffusion.coreml_model.CoreMLModel`,
|
||||
> stays in suite") is **wrong**: the loader is already the local `coreml_model.CoreMLModel`. The
|
||||
> "stays in suite" conclusion still holds; the identity does not.
|
||||
> - **Gate E0 item "ml-stable-diffusion pinned SHA — BLOCKER if unpinned" is MOOT** — there is no
|
||||
> such dep to pin. Mark it N/A, not BLOCKER.
|
||||
> - **E4/E5 dependency lists must drop `git+...ml-stable-diffusion@<sha>`.** Package runtime deps
|
||||
> are: `coremltools`, `diffusers`, `peft` (LoRA), `omegaconf` (config), `numpy`, `torch`. Confirm
|
||||
> `peft`/`omegaconf` actually used before listing (grep at E4).
|
||||
> - The "keep `python_coreml_stable_diffusion` as a suite dep for the loader" instruction in E5 is
|
||||
> **void** — coremltools backs the loader.
|
||||
|
||||
---
|
||||
|
||||
## 7. Pre-flight checklist before E1 (run these greps)
|
||||
|
||||
```
|
||||
grep -rn "import comfy" coreml_suite/conversion coreml_suite/converter.py coreml_suite/lcm/converter.py coreml_suite/lcm/unet.py
|
||||
grep -rn "folder_paths" coreml_suite/converter.py coreml_suite/lcm/converter.py
|
||||
grep -rn "model_management" coreml_suite/lcm
|
||||
grep -rn "python_coreml_stable_diffusion" coreml_suite
|
||||
grep -rn "ATTENTION_IMPLEMENTATION_IN_EFFECT" coreml_suite
|
||||
grep -rn "SimianLuo\|LCM_Dreamshaper" coreml_suite/lcm
|
||||
```
|
||||
Every 🔍 above resolves to ✅ or a correction once these run. Do not start moving code (E2)
|
||||
with any 🔍 unresolved on the CONVERSION side.
|
||||
|
||||
**STATUS (run 2026-05-26): all 🔍 resolved.** Summary of what the greps found:
|
||||
- `conversion/*`, `lcm/unet.py`: comfy-free (torch/diffusers only). ✅
|
||||
- `converter.py`: only comfy reach-in is `folder_paths` in `get_out_path` (L91-94) → inject `out_path`.
|
||||
- `lcm/converter.py`: `folder_paths` (L111-114) + `comfy.model_management.get_torch_device` (L54)
|
||||
→ cut both. Dup helpers (`load_coreml_model`,`convert_to_coreml`,`get_out_path`,`get_sample_input`)
|
||||
confirmed → dedup E2. `MODEL_VERSION="SimianLuo/LCM_Dreamshaper_v7"` (L22) → E-LCM.
|
||||
- No attention module-global anywhere (`ATTENTION_IMPLEMENTATION_IN_EFFECT` absent); already per-call.
|
||||
LCM hardcodes `"SPLIT_EINSUM"` in `get_unets` — thread `attn_impl` through during dedup.
|
||||
- `.name` vs `.value`: **decided `.name`** (node reverses via `ModelVersion[...]`). §5 corrected.
|
||||
- `ml-stable-diffusion`: **already gone** (#58). §6 stale-spec note added — fix the spec's E0/E4/E5
|
||||
dep + pinning items.
|
||||
|
||||
Two grep blind-spots to note (the checklist above doesn't cover them, but cheap to add): the
|
||||
`folder_paths` grep only scans the two converter files — also grep `coreml_suite/lcm/utils.py`
|
||||
(it imports `comfy.model_management` at L3, but it's inference/STAYS, so fine) and confirm no other
|
||||
`conversion/` file grew a comfy import since.
|
||||
Reference in New Issue
Block a user