Author SHA1 Message Date
aszc-dev 1008f144ad docs: sync with converter consolidation
Docs described the pre-#67 suite: a removed standalone LCM converter
node, a removed model_version input, a nonexistent
CoreMLDetailerHookProvider node, and a deleted lcm/converter.py file.

- drop LCM converter node docs; LCM checkpoints are auto-detected by
  the consolidated CoreMLConverter
- remove model_version input and invented 512-768 resolution range
- replace phantom detailer-hook fix with real workarounds
- fix Python support claim (3.12+ per requires-python)
- update LCM support-matrix row, drop stale line reference

Refs #67, #68
2026-07-09 18:33:17 +02:00
aszc-dev 8f94f0eea5 docs: rewrite README and split into docs/ pages
Rewrite the README as a lean landing page and move depth into a docs/
folder. Correct the supported-model story and several stale facts, and
answer the recurring questions from issue #21.

- Convert-only is the supported path: suite-converted .mlpackage is the
  only supported input; drop coreml-community download guidance.
- Remove all .mlmodelc / Xcode references — compilation was dropped and
  the loader handles .mlpackage only.
- Fix compute-unit name (CPU_AND_NE, not CPU_AND_ANE) and the loader
  input name (coreml_name).
- Document CoreMLSamplerAdvanced (previously undocumented).
- Add docs/: hardware, nodes, conversion, workflows, faq,
  troubleshooting, limitations (with a support matrix).
- Note conversion now lives in the coreml-diffusion package.
- Remove dev scaffolding specs; ignore *.log, .DS_Store, .claude/.
2026-07-09 18:30:26 +02:00
11 changed files with 759 additions and 1258 deletions
+2
View File
@@ -3,4 +3,6 @@ __pycache__/
models/ models/
.venv/ .venv/
test_results/ test_results/
*.log
.DS_Store
.claude/ .claude/
-618
View File
@@ -1,618 +0,0 @@
# ComfyUI-CoreMLSuite — Converter Extraction Spec for Claude Code
> **Companion to `MODERNIZATION_SPEC.md`.** That spec hardens the repo and (Phase 3)
> splits the *inference* math from the framework. **This** spec splits the *conversion*
> path (`safetensors → CoreML`) out into a standalone, `comfy`-free, pip-installable
> package that CoreMLSuite then depends on — and that other projects (incl. on-device
> iOS tooling) can reuse.
>
> **Same discipline as the modernization spec:** safety-net first, behavior-preserving
> until told otherwise, one phase = one branch = one PR, `STOP — VALIDATE` gate between
> every phase, golden-latent as the regression anchor. `[M2]` = needs macOS/Apple Silicon;
> `[M2-ANE]` = needs the Neural Engine. Everything else must run on plain Linux/CI.
---
## 0. How to work (read first — non-negotiable)
1. **Behavior-preserving until Phase E6.** Phases E1–E5 must not change image output, node
names, `INPUT_TYPES` field names, or `NODE_CLASS_MAPPINGS` keys. The node graph is the
public contract; saved user-workflow JSON breaks if these change.
2. **The conversion package produces an artifact and stops there.** Its job ends at a written
`.mlpackage` / `.mlmodelc` on disk. It must NOT import `comfy`, `folder_paths`, or
`comfy_extras`, and must NOT know ComfyUI's `models/unet` layout. Paths are *inputs*.
3. **The runtime loader stays in the suite.** The loader is the **local** `coreml_suite.coreml_model.CoreMLModel`
— a thin wrapper over `coremltools.models.MLModel` (NOT Apple's
`python_coreml_stable_diffusion.coreml_model.CoreMLModel`, which is no longer used; see #58).
It *runs* a compiled model in Python — a desktop/Python inference concern, not a conversion
concern. It is NOT moved into the package. (On iOS the `.mlmodelc` is loaded natively; the
package's output is the deliverable, not a Python runner.)
4. **Decouple in-repo before splitting repos.** Phases E1–E4 create the package *inside this
repo* and prove equivalence. The physical second-repo split is Phase E5, only after the
golden latent is proven identical. Do not create a second repository before Gate E4 passes.
5. **Reuse the existing regression anchor.** The golden latent / PSNR anchor from
`MODERNIZATION_SPEC.md` Phase 2 is the cross-cutting proof for every gate here. If it is not
yet captured, capture it first (it is a prerequisite for E2 onward).
6. **No new runtime dependencies** without flagging in the gate report (name, why, license, size).
7. **A failing gate means stop and report**, not work around into the next phase.
8. **Tooling is `uv`, not bare `pip`/`venv`.** Every environment/install/lock step uses the
project's `uv` toolchain: `uv venv`, `uv pip install`, `uv pip install -e .`, `uv lock`,
`uv run pytest`, `uv export`/`uv pip freeze` for baselines. Where this spec says "fresh venv",
read "`uv venv` + `uv pip install`". Reserve `uv pip` (not `pip`) inside that venv too.
9. **The package is the single source of truth for *what is possible*; the node is a thin,
discovery-driven frontend.** See the "Interface contract" pillar below — this is the
maintainer's hard requirement and it overrides the earlier (now-rescinded) "freeze the
dropdown list" instruction.
---
## Interface contract (the maintainer's hard requirement) — read before any phase
Two coupled guarantees must hold once the package is split out:
**(A) Updating the converter must NOT require updating CoreMLSuite.**
This is satisfied by treating the package's public surface as a versioned contract:
- `convert(...)` and `compile_model(...)` are **keyword-only with defaults** for everything
past the genuinely-required positionals (`ckpt_path`, `model_version`, `out_path`). New
capabilities are added as new keyword args with defaults, so an old Suite's call still
validates against a newer package. **Never** reorder or rename existing parameters.
- `compose_out_name` (the `.mlpackage` filename = the cache key) **moves into the package** and
is versioned with it. The Suite must not carry its own copy; if the package changes the naming
scheme that is a **major** bump (old cached artifacts stop resolving).
**(B) CoreMLSuite must be able to list *new* conversion types WITHOUT a Suite code change or
version bump.** Today the node hardcodes its dropdowns:
```python
"model_version": ([ModelVersion.SD15.name, ModelVersion.SDXL.name],), # hand-typed, also INCOMPLETE (no LCM / SDXL_REFINER)
"attention_implementation": (list(ATTENTION_IMPLEMENTATIONS),), # from coreml_suite.attention
"quantize_nbits": (list(QUANT_NBITS_VALUES), {"default": "none"}), # from coreml_suite.core.naming
```
These are replaced by **runtime discovery calls into the package**, evaluated inside
`INPUT_TYPES` (ComfyUI re-evaluates `INPUT_TYPES` on every plugin load):
```python
import coreml_diffusion
"model_version": (coreml_diffusion.list_model_versions(),),
"attention_implementation": (coreml_diffusion.list_attention_impls(),),
"quantize_nbits": (coreml_diffusion.list_quant_modes(), {"default": "none"}),
```
Effect: `uv pip install -U coreml_diffusion` + ComfyUI restart surfaces any newly-added type in the old
plugin's dropdown — **no Suite edit, no Suite version bump.** This is the requirement.
**The cost, stated honestly (accept this trade-off explicitly at Gate E0):**
- The Suite becomes a "dumb" frontend; the package is the sole authority on what conversions
exist. The Suite can no longer guarantee its saved workflows are valid against *arbitrary*
future package versions.
- Therefore the package's discovery identifiers (`ModelVersion` values, attn-impl strings, quant
modes) are an **ADDITIVE-ONLY contract**: the package may *add* identifiers freely (minor bump,
no Suite change); **removing or renaming an identifier is a breaking change requiring a MAJOR
bump and a migration note**, because a saved workflow JSON references these strings verbatim.
Without this rule, "no version bump" silently becomes "randomly broken workflows."
- `INPUT_TYPES` must **fail soft** when the package is missing/old: wrap the discovery calls so a
missing `coreml_diffusion` (or an old one lacking a `list_*` function) yields a sane fallback list and a
logged warning, instead of the node failing to register and disappearing from the menu.
**Discovery API the package must expose (stable names):**
```python
coreml_diffusion.list_model_versions() -> list[str] # VERIFIED ones only, e.g. ["SD15","SDXL"] today (.name — see seam.md)
coreml_diffusion.list_attention_impls() -> list[str] # ["SPLIT_EINSUM","SPLIT_EINSUM_V2","ORIGINAL"]
coreml_diffusion.list_quant_modes() -> list[str] # ["none","8","6","4"]
coreml_diffusion.CONTRACT_VERSION: str # bump rules above; Suite may log/compare it
```
These return the *display strings already used today*, so existing workflows keep validating.
**Verification status is a PACKAGE property, not a node hardcode (maintainer's intent).**
The Suite wants to expose *every model the converter can verifiably convert*. Today `lcm` and
`sdxl_refiner` are absent from the converter node not because the Suite chooses to hide them, but
because they lack a full golden/PSNR verification. So the gating lives in the package as a status:
```python
from enum import Enum
class Status(Enum):
VERIFIED = "verified" # has a golden anchor + passing [M2-ANE] check
EXPERIMENTAL = "experimental" # convertible but not yet anchored/verified
# internal registry, single source of truth.
# KEY by ModelVersion enum MEMBER so list_* can emit .name. Keying by the lowercase
# .value string returns ["sd15",...], which the node reverses via ModelVersion[...] -> KeyError.
_MODEL_STATUS = {ModelVersion.SD15: Status.VERIFIED, ModelVersion.SDXL: Status.VERIFIED,
ModelVersion.SDXL_REFINER: Status.EXPERIMENTAL, ModelVersion.LCM: Status.EXPERIMENTAL}
def list_model_versions(include_experimental: bool = False) -> list[str]:
return [v.name for v, s in _MODEL_STATUS.items() # .name -> "SD15","SDXL"; node reverses with ModelVersion[...]
if s is Status.VERIFIED or (include_experimental and s is Status.EXPERIMENTAL)]
```
Consequence: **promoting a model to VERIFIED in the package expands the Suite's dropdown with no
Suite change and no Suite bump** — exactly the requirement. The act of verification (E-LCM
produces an LCM golden anchor; same later for refiner) is what flips the status. The Suite's
converter node calls `list_model_versions()` (verified-only); a power-user/CLI path may pass
`include_experimental=True`. Promotion VERIFIED-from-EXPERIMENTAL is additive (minor bump);
demotion or removal is breaking (major bump + note).
---
## Naming & layout (chosen — frozen at Gate E0)
**Distribution name (PyPI):** `coreml-diffusion`. **Import name (Python):** `coreml_diffusion`.
(PyPI normalizes `-`/`_`; the distribution uses the hyphen, the importable module the underscore.)
Availability checked: both `coreml-diffusion` and the near variants were free on PyPI at E0.
**Why this name (the positioning it encodes):** the project's niche is *diffusion models on Apple
Neural Engine via CoreML, inside ComfyUI and on-device* — **not** Stable Diffusion specifically.
`sd*` was rejected because it falsely narrows scope to SD; `coreml-diffusion` keeps `coreml` on the
front for discoverability while `diffusion` honestly states the scope (SD/SDXL/LCM today, Flux and
other diffusion architectures later) **without** promising arbitrary non-diffusion torch models,
whose tracing/shape/sample-input pipeline differs. The name must not be re-narrowed to SD in
future docs. ANE is the *differentiator* (documented in the README), but `coreml` was chosen over
`ane` in the name for search discoverability per maintainer decision.
Target package layout (framework-free — zero `comfy` imports):
```
coreml_diffusion/
__init__.py # public API surface (see "Public API" below)
model_version.py # ModelVersion enum — the SINGLE source of truth, no comfy
attention.py # ATTENTION_IMPLEMENTATIONS tuple (from coreml_suite/attention.py) + apply_attention_implementation
pipeline.py # get_pipeline (from_single_file), get_unet (cml UNet from ref unet)
unet.py # UNet2DConditionModelLCM (moved from coreml_suite/lcm/unet.py)
inputs.py # get_sample_input, lcm_inputs, sdxl_inputs,
# get_encoder_hidden_states_shape, get_coreml_inputs, get_inputs_spec
controlnet.py # add_cnet_support (conversion-side residual SHAPE calc only)
convert.py # convert_unet, convert (orchestration), convert_to_coreml, load_coreml_model
compile.py # compile_coreml_model
quantize.py # (Phase E6 / MODERNIZATION Phase 6 lands here) palettization 4/6/8-bit
cli.py # console entry point: `coreml-diffusion convert ...`
pyproject.toml # standalone packaging (at E5)
```
What stays in `coreml_suite/` (the ComfyUI side, thinned):
- `nodes.py` — still owns **name-encoding** (`out_name` construction), path resolution via
`folder_paths`, the node `INPUT_TYPES`/mappings, and wrapping the result in `CoreMLModel`.
- `models.py`, `latents.py`, `controlnet.py` (inference parts), `lcm/utils.py`, `config.py`
(inference config build) — untouched by this spec except the import-source of `ModelVersion`.
### Public API (the contract `coreml_diffusion` exposes)
```python
from coreml_diffusion import ModelVersion, convert, compile_model, compose_out_name
from coreml_diffusion import list_model_versions, list_attention_impls, list_quant_modes, CONTRACT_VERSION
# Mirror the CURRENT converter.py signature, made keyword-only past the required positionals
# and with paths/device injected (no folder_paths, no comfy.model_management):
# convert(ckpt_path, model_version, out_path, *,
# batch_size=1, sample_size=(64, 64), controlnet_support=False,
# lora_weights=None, attn_impl=list_attention_impls()[0], config_path=None,
# quantize_nbits="none", device=None) -> None # side effect: writes out_path
# (current convert() returns None and writes via convert_unet → coreml_unet.save; keep that,
# or change to `return out_path` as a deliberate, documented improvement — pick one at E0.)
# compile_model(src_path, out_dir, final_name) -> str # returns compiled .mlmodelc path
```
Note: `convert` takes an **explicit `out_path`** — no `folder_paths`. `device` is injected
(defaults to torch's default device). `compose_out_name` lives here (cache-key contract) and the
node imports it from the package. The `list_*` discovery functions back the node's dropdowns.
---
## The import chains to cut (root cause inventory) — REVISED against current code
> **State note (verified):** the code moved on since the original draft. Several chains are
> already cut. Re-verify each line by `grep` before acting; do not assume the original draft.
**Already done (verify, then skip):**
- ✅ `converter.py` already imports `from coreml_suite.model_version import ModelVersion`, and
`model_version.py` is **clean** (`from enum import Enum` only — zero comfy). The old
"converter → config → comfy" chain is **already broken**. `config.py` still imports comfy, but
it is **inference-side** (`get_model_config` via `supported_models_base`/`latent_formats`) —
*not* on the conversion path. Do **not** treat `config.py` as a converter dependency.
- ✅ `converter.py` now uses `diffusers.UNet2DConditionModel.from_single_file` and a local
`CoreMLUNetWrapper` (in `coreml_suite/conversion/unet.py`) — it is **no longer** importing the
Apple `python_coreml_stable_diffusion.unet.UNet2DConditionModel*` internals on the main path.
A `coreml_suite/conversion/` subpackage already exists (`attention`, `shapes`, `trace`, `unet`).
- ✅ Name-encoding already extracted to `coreml_suite/core/naming.py` (`compose_out_name`,
`lora_names_from_params`, `ATTN_SUFFIX`, `QUANT_NBITS_VALUES`) **with characterization tests**
(`tests/unit/test_characterization_out_name.py`). The pure-naming split is done.
- ✅ Quantization is **already implemented** in `converter.py` (`quantize_nbits`, k-means
`palettize_weights`) and surfaced as an optional node input. Phase E6 is therefore *move*, not
*build* (see revised E6).
**Still to cut (the real remaining work):**
1. `coreml_suite/converter.py::get_out_path` → `from folder_paths import get_folder_paths`.
Main converter still reaches into ComfyUI's model dir. **Cut: `out_path` is an injected arg;
`folder_paths` resolution moves up into the node** (the node already computes `out_name`).
2. `coreml_suite/lcm/converter.py` → still has its **own** `from folder_paths import
get_folder_paths` (`get_out_path`) and (per original draft) `comfy.model_management`. Verify
the current LCM file and cut both: inject `out_path` and `device`.
3. Global mutation of the attention impl: confirm where it now lives. Main path appears to route
through `coreml_suite/conversion/attention.apply_attention_implementation` (cleaner than the
old global), but `lcm/converter.py` may still set a module global at import. **Ensure the
package sets attention per-call, never at import time.**
4. **Duplication LCM vs main:** `lcm/converter.py` still carries its own copies of
`convert_to_coreml`, `load_coreml_model`, `get_out_path`, `get_sample_input` (the LCM variant
takes a `scheduler` arg), and hardcodes `SimianLuo/LCM_Dreamshaper_v7`. **Dedupe into the
single `coreml_diffusion` implementation;** the HF-hardcode consolidation is the *behavior-changing*
part → deferred to optional **E-LCM**, not E1–E5.
5. **`compose_out_name` ownership:** currently in `coreml_suite/core/naming.py` and called by the
node. Per the Interface-contract pillar it must **move into the package** (it is the cache-key
contract) and the node must import it from `coreml_diffusion`, not keep a copy.
---
## Phase E0 — Seam decision & inventory (no code change)
**Objective:** lock the cut line, the interface contract, and naming so later phases don't drift.
### Tasks
1. Produce `docs/extraction/seam.md`: a table of every symbol in `converter.py`,
`lcm/converter.py`, `lcm/unet.py`, **plus the already-extracted `conversion/` subpackage
(`attention`, `shapes`, `trace`, `unet`) and `core/naming.py`**, classified
**CONVERSION → coreml_diffusion** vs **STAYS (comfy/node)**. Note which are already framework-free.
2. ~~Confirm the current `python_coreml_stable_diffusion` footprint.~~ **DONE (seam.md §6):
footprint is ZERO** — no runtime imports anywhere; only a docstring mention in
`core/__init__.py:4`. Main path uses `diffusers` + local `CoreMLUNetWrapper`; the runtime
`CoreMLModel` (STAYS in suite) is a local coremltools wrapper, not Apple's. No shape/attn helper
comes from Apple (local `conversion/shapes.py`, `conversion/attention.py`).
3. **Decide the interface contract concretely (the maintainer's hard requirement):**
- Discovery functions `list_model_versions / list_attention_impls / list_quant_modes` live in
the package and return today's display strings verbatim. Node `INPUT_TYPES` calls them.
- `ModelVersion` values, attn-impl strings, quant modes are **ADDITIVE-ONLY** across package
versions; removal/rename = MAJOR bump + migration note. Write this into the package's
versioning policy doc now.
- `compose_out_name` moves to the package; node imports it (no copy). Confirm the
characterization tests in `test_characterization_out_name.py` will be re-pointed, not
duplicated.
- **Resolve the `model_version` dropdown question (maintainer decided):** the Suite exposes
*every model the converter can verifiably convert*. `lcm` and `sdxl_refiner` are absent today
only because they lack a golden/PSNR verification — **not** because the node hardcodes a
short list. Encode this as a **status registry in the package** (`VERIFIED` vs
`EXPERIMENTAL`); `list_model_versions()` returns VERIFIED-only by default. The converter node
calls it plainly. Promoting LCM/refiner to VERIFIED (after E-LCM / a refiner anchor) expands
the dropdown with **no Suite change**. Do NOT add permanent per-node filtering — the gate is
verification status, owned by the package.
4. ~~Confirm the `ml-stable-diffusion` git dep is pinned.~~ **N/A — already removed (#58).** Verified:
zero `python_coreml_stable_diffusion` imports in the repo; `CoreMLModel` is now a local
coremltools wrapper; the dep is absent from `pyproject.toml`/`requirements.txt`. No SHA to pin.
### STOP — VALIDATE (Gate E0)
```
## Gate E0 report
- seam.md committed: <path>; symbol counts (move / stay / already-framework-free)
- python_coreml_stable_diffusion usage (verified by grep): conversion=<list> runtime=<list>
- Discovery API signatures frozen: list_model_versions (verified-only) / list_attention_impls / list_quant_modes
- Status registry decided: sd15+sdxl=VERIFIED, lcm+sdxl_refiner=EXPERIMENTAL (gated, not hidden)
- Additive-only contract policy doc written (incl. promotion=minor, demotion/removal=major): <path>
- model_version dropdown: expose all (incl. LCM/REFINER) / filtered per node — DECISION: <...>
- compose_out_name move-not-copy confirmed; tests re-point plan: <...>
- LCM consolidation deferred to optional E-LCM: YES/NO
- ml-stable-diffusion: N/A — already removed (#58), not a dependency (was: pin-or-BLOCKER)
- Package name in-repo: coreml_diffusion (final PyPI name deferred to E5)
```
---
## Phase E1 — Establish `coreml_diffusion` package + discovery API (mostly verification)
**Objective:** stand up the package namespace and the discovery surface. Much of the comfy-chain
cut is **already done** — this phase mostly *verifies* that and adds the discovery functions.
### Tasks
1. **Verify (don't redo):** `coreml_suite/model_version.py` is already clean (`Enum` only). Confirm
`import coreml_suite.model_version` works with **no comfy** (`uv run python -c "..."` in a
comfy-free `uv venv`). If true, E1's original "extract ModelVersion" task is already satisfied.
2. Create the `coreml_diffusion/` package skeleton with `__init__.py` exporting the **discovery API**
backed by the *existing* sources of truth for now (re-export `ModelVersion`, the
`ATTENTION_IMPLEMENTATIONS` tuple, and `QUANT_NBITS_VALUES`) so values are byte-identical:
```python
def list_model_versions(): return [v.name for v in ModelVersion] # .name -> "SD15" (node reverses via ModelVersion[...]; .value KeyErrors)
def list_attention_impls(): return list(ATTENTION_IMPLEMENTATIONS)
def list_quant_modes(): return list(QUANT_NBITS_VALUES)
CONTRACT_VERSION = "1.0"
```
(At this stage `coreml_diffusion` may live inside the repo and import from `coreml_suite.*`; the
physical move of implementation happens in E2. The point of E1 is to freeze the *contract*.)
3. **Decided (`.name`):** the node renders `ModelVersion.SD15.name` (`"SD15"`) and reverses the
dropdown string via `ModelVersion[model_version]` (name lookup, `nodes.py:286`). Discovery API
therefore returns `.name`; `.value` (`"sd15"`) would `KeyError`. Recorded in `seam.md` §5.
### Acceptance criteria
- `uv run python -c "import coreml_diffusion; print(coreml_diffusion.list_model_versions(), coreml_diffusion.list_quant_modes())"`
works in a **comfy-free** `uv venv` and prints today's exact strings.
- Existing characterization tests pass unchanged.
- No node behavior change yet (node still uses its current hardcoded lists in E1).
### STOP — VALIDATE (Gate E1)
```
## Gate E1 report
- model_version.py confirmed comfy-free (uv, no comfy): PASS/FAIL
- coreml_diffusion.list_* returns byte-identical strings to current dropdowns: YES/NO (show values)
- .name vs .value decision for model_version discovery: <...>
- CONTRACT_VERSION set; additive-only policy linked: <path>
- Characterization tests unchanged & green (uv run pytest): YES/NO
```
---
## Phase E2 — Move conversion code into `coreml_diffusion` (in-repo, dedup, behavior-preserving)
**Objective:** physically relocate the conversion mechanics into the framework-free package,
collapsing the two duplicate converters into one, with paths/device injected.
### Tasks
1. Move into `coreml_diffusion/`: `pipeline.py` (`get_pipeline`, `get_unet`), `unet.py`
(`UNet2DConditionModelLCM`), `inputs.py` (sample/lcm/sdxl input builders +
`get_encoder_hidden_states_shape` + `get_coreml_inputs` + `get_inputs_spec`),
`controlnet.py` (`add_cnet_support`), `convert.py` (`convert_unet`, `convert`,
`convert_to_coreml`, `load_coreml_model`), `compile.py` (`compile_coreml_model`).
2. **Dedupe LCM vs main** (the real remaining duplication): delete `lcm/converter.py`'s copies of
`convert_to_coreml` / `load_coreml_model` / `get_out_path` / `get_sample_input` (LCM variant
carries a `scheduler` arg — fold that into the shared `get_sample_input` as an optional param)
in favor of the single `coreml_diffusion` implementation. The main path's helpers
(`get_unet`/`get_encoder_hidden_states_shape`/`get_coreml_inputs`/`convert_unet`/`convert`) and
the `conversion/` subpackage (`attention`, `shapes`, `trace`, `unet`) move as-is.
3. **Inject paths**: replace `get_out_path`'s `folder_paths` reach-in with an injected `out_path`
argument on `convert(...)`; `folder_paths` resolution moves up into the node (which already
computes `out_name`). No `folder_paths` import anywhere in `coreml_diffusion`.
4. **Inject device** where the LCM path used `comfy.model_management` (verify it still does):
`convert(..., device=None)`, default to torch's default device.
5. **Attention per-call, never at import:** main path already routes through
`conversion/attention.apply_attention_implementation` — keep that. If `lcm/converter.py` still
sets any module global at import, remove it; the package sets attention from the `attn_impl`
arg inside `convert`.
6. **Move `compose_out_name` into the package** (`coreml_diffusion/naming.py`); re-point
`test_characterization_out_name.py` imports to `coreml_diffusion.naming` — assertions and values
unchanged. The node will import it from the package in E3.
7. Leave **thin shims** in `coreml_suite/converter.py` and `coreml_suite/lcm/converter.py` that
re-export from `coreml_diffusion`, preserving the old call signatures the nodes use (nodes untouched
this phase). Shims map comfy `folder_paths`/device into package args.
### Acceptance criteria
- `uv run pytest -m unit` (Tier 0) imports `coreml_diffusion.*` with **no comfy / no MPS** and is green on Linux.
- The dedup leaves exactly one implementation of each previously-duplicated function.
- Characterization tests pass unchanged after the `compose_out_name` re-point.
- `[M2]` A real SD1.5 conversion via the shim still produces a loadable model.
- `[M2-ANE]` **Golden latent identical / within tolerance** to the MODERNIZATION Phase 2 anchor
(same seed/prompt) — proves the move + dedup changed nothing.
### STOP — VALIDATE (Gate E2 — first regression gate)
```
## Gate E2 report
- Tier 0 import of coreml_diffusion without comfy/MPS (uv run): PASS/FAIL
- LCM/main duplicated funcs collapsed to one (list old→new): <map>
- compose_out_name moved to package; char-tests re-pointed & green: YES/NO
- Paths injected (no folder_paths in package): confirmed
- Device injected (no comfy.model_management in package): confirmed
- Attention set per-call, not at import (both main & lcm): confirmed
- [M2-ANE] Golden latent vs Phase-2 anchor: identical / within tol <x> / DIVERGED (STOP)
- Node INPUT_TYPES / mappings untouched: confirmed (diff)
```
**If the golden latent diverged at all, STOP and report — do not continue.**
---
## Phase E3 — Thin the nodes onto the package (behavior-preserving)
**Objective:** remove the shims; have the ComfyUI nodes call `coreml_diffusion` directly, keeping the
node contract byte-identical.
### Tasks
1. `CoreMLConverter.convert` (in `coreml_suite/nodes.py`): keep the `folder_paths`-based path
resolution **in the node**; import `compose_out_name` from `coreml_diffusion` (not `coreml_suite.core`);
call `coreml_diffusion.convert(...)` and `coreml_diffusion.compile_model(...)` directly; wrap the compiled path
in `CoreMLModel`.
2. **Wire the dropdowns to discovery (the maintainer's hard requirement).** Replace the hardcoded
`INPUT_TYPES` lists with fail-soft discovery calls:
```python
def _discover(fn, fallback):
try:
import coreml_diffusion
return getattr(coreml_diffusion, fn)()
except Exception as e: # missing/old package, or import error
logger.warning(f"coreml_diffusion.{fn} unavailable ({e}); using fallback {fallback}")
return fallback
...
"model_version": (_discover("list_model_versions", ["SD15", "SDXL"]),),
"attention_implementation": (_discover("list_attention_impls", ["SPLIT_EINSUM","SPLIT_EINSUM_V2","ORIGINAL"]),),
"quantize_nbits": (_discover("list_quant_modes", ["none","8","6","4"]), {"default": "none"}),
```
This is what makes "update the package → new types appear in the old node, no Suite bump" true.
3. `COREML_CONVERT_LCM` (in `coreml_suite/lcm/nodes.py`): route through `coreml_diffusion` for the shared
mechanics. **Keep the existing LCM behavior/HF-hardcode for now** — consolidation is optional E-LCM.
4. Delete the now-dead `coreml_suite/converter.py` / `coreml_suite/lcm/converter.py` shims (or
reduce to a one-line re-export if anything external imports them — grep first).
### Acceptance criteria
- `NODE_CLASS_MAPPINGS` / `NODE_DISPLAY_NAME_MAPPINGS` keys: **unchanged** (diff `__init__.py`).
- Every `INPUT_TYPES` **field name** unchanged. Dropdown **values**: the discovery calls must
return **a superset of** today's values, with every previously-present value still present and
spelled identically (additive-only). *(This deliberately replaces the original spec's
"values must be byte-identical/frozen" criterion — the maintainer requires the list be
extensible at runtime. Frozen-field-names + additive-only-values is the new contract.)*
- With `coreml_diffusion` **absent**, the node still registers and shows the fallback lists (fail-soft).
- `[M2-ANE]` Golden latent still identical to the Phase-2 anchor.
- `[M2-ANE]` The committed e2e workflow `tests/integration/...` still passes (PSNR > 25).
### STOP — VALIDATE (Gate E3)
```
## Gate E3 report
- Node mappings diff: empty (confirmed)
- INPUT_TYPES field-names diff: empty (confirmed)
- Dropdown values: superset of prior, all prior values still present & identical: YES/NO (show)
- Fail-soft with coreml_diffusion absent (node still registers): PASS/FAIL
- compose_out_name now imported from coreml_diffusion (no node-side copy): confirmed
- [M2-ANE] Golden latent vs anchor: identical / within tol / DIVERGED (STOP)
- [M2-ANE] e2e workflow PSNR: <value> (> 25?)
- Dead converter shims removed / reduced: <list>
```
---
## Phase E4 — Standalone packaging & CLI (still in-repo)
**Objective:** make `coreml_diffusion` independently installable and usable without ComfyUI, with a CLI
suitable for the planned article and for on-device/iOS conversion workflows.
### Tasks
1. Add `coreml_diffusion/pyproject.toml`: name (working `coreml_diffusion`), `requires-python`, dependencies
= `coremltools` (pinned to the MODERNIZATION-validated version), `diffusers`, `transformers`,
`peft`, `omegaconf`, `numpy`, `torch`. **No `ml-stable-diffusion`** (already removed in #58, see
§0.3) and **no comfy**. Suite pins `transformers>=4.44`/`peft>=0.13`/`omegaconf>=2.3` today;
grep-confirm each is on the conversion path before listing it. A `[project.scripts]` entry:
`coreml-diffusion = "coreml_diffusion.cli:main"`.
2. `coreml_diffusion/cli.py`: `coreml-diffusion convert --ckpt PATH --model-version sd15 --out PATH
[--height --width --batch-size --attn-impl --controlnet --lora NAME:STRENGTH ... --config PATH]`
and `coreml-diffusion compile --src PATH --out-dir DIR --name NAME`. Mirrors `convert()`/`compile_model()`.
3. Tier-0 Linux tests for the CLI **arg→call mapping** (mock the heavy `convert`); the real
convert remains `[M2]`. Add a `[M2]` smoke test: convert a tiny synthetic UNet end-to-end.
4. README for the package: install, CLI usage, "produce a `.mlpackage`/`.mlmodelc` for use in a
Swift/iOS app", and the ANE positioning note (low-power, GPU-free, embeddable; SD1.5/SDXL on
ANE, **not** a Flux-speed claim).
### Acceptance criteria
- Fresh `python -m venv` + `uv pip install ./coreml-diffusion` (no ComfyUI present) imports and runs
`coreml-diffusion --help` and the arg-mapping tests on Linux.
- `[M2]` `coreml-diffusion convert` produces a model file identical (golden) to the node path.
### STOP — VALIDATE (Gate E4)
```
## Gate E4 report
- uv pip install ./coreml-diffusion in comfy-free venv: PASS/FAIL (log)
- CLI arg→call tests (Tier 0, Linux): green
- [M2] CLI-produced model golden vs node-produced model: identical / DIVERGED
- Package deps list (with pinned SHAs/versions + licenses):
- New runtime deps vs suite before: <none / list>
```
**This is the gate that proves the package stands alone. Do not split repos before it passes.**
---
## Phase E5 — Physical split into a second repository
**Objective:** move `coreml_diffusion/` to its own repo; CoreMLSuite depends on it by pinned version.
### Tasks
1. Create the new repo (maintainer action — agent prepares the tree, not the GitHub repo).
Choose final distributable name; rename imports if changed (single sweep, recorded).
2. CoreMLSuite `pyproject.toml` / `requirements.txt`: replace the conversion-only deps with a
pinned dependency on the new package (`coreml_diffusion==<version>` from PyPI, or `git+...@<tag>`
until first PyPI release). (There is no `git+...ml-stable-diffusion` line to remove — already
gone since #58.)
3. ~~Keep `python_coreml_stable_diffusion` for the loader.~~ **Void.** The loader is the local
`coreml_suite/coreml_model.py` over `coremltools`; the suite keeps `coremltools` as a direct dep
for it. No Apple lib involved.
4. Set up the new repo's CI: Tier 0 on Linux (import + arg-mapping + input-shape math),
`[M2]`/`[M2-ANE]` on a self-hosted/macOS-ARM runner reusing the golden-latent anchor.
5. Versioning: SemVer; first release `0.1.0`. Document the compatibility matrix
(coreml_diffusion ↔ coremltools version ↔ diffusers version). No ml-stable-diffusion axis.
### Acceptance criteria
- CoreMLSuite installs in a fresh venv pulling the new package; e2e workflow still passes `[M2-ANE]`.
- New repo CI green on Linux (Tier 0) and `[M2-ANE]` golden latent matches the anchor.
- No conversion code remains in CoreMLSuite (grep: no `ct.convert`, no `from_single_file`,
no `torch.jit.trace`).
### STOP — VALIDATE (Gate E5)
```
## Gate E5 report
- New repo tree prepared: <path/branch>; final package name: <name>
- Suite depends on package by pinned version: <spec>
- Suite e2e [M2-ANE] PSNR after split: <value> (> 25?)
- Conversion code fully absent from suite: confirmed (grep output)
- Compatibility matrix documented: <link>
- First release tag: 0.1.0
```
---
## Phase E6 — Quantization travels WITH the conversion code (already implemented → move)
**Objective:** quantization is **already implemented** (k-means `palettize_weights` in
`converter.py`, `quantize_nbits` node input, `_q<bits>` filename suffix, README tradeoff table).
There is nothing to *build*. It simply **moves with the conversion code in E2** as part of
`convert_unet`. This phase is a checkpoint that it survived the extraction intact, plus exposing
it through the CLI.
### Tasks
1. Confirm the palettization block moved cleanly into `coreml_diffusion` (lives in `convert.py` or a
`quantize.py` helper called from `convert_unet`). Default `"none"` stays byte-identical.
2. Expose via CLI flag `--quantize {none,8,6,4}` (E4 already lists this) and via
`list_quant_modes()` discovery (E1/E3).
3. The existing README tradeoff table (SD1.5 1×512×512 SPLIT_EINSUM: none/8/6/4 → size/ms/PSNR)
moves to the package README. Re-confirm one row `[M2-ANE]` so the article can cite a live number.
### Acceptance criteria
- Default (`none`) output byte-identical to pre-extraction (covered by the E2/E3 golden latent).
- `coreml-diffusion convert --quantize 4` produces a `_q4` artifact matching the node's `_q4` artifact `[M2]`.
- `list_quant_modes()` drives the node dropdown (no hardcoded copy remains).
### STOP — VALIDATE (Gate E6)
```
## Gate E6 report
- Palettization relocated into coreml_diffusion, called from convert_unet: confirmed
- Default none output identical (golden): YES/NO
- [M2] CLI --quantize {8,6,4} artifacts match node artifacts: YES/NO
- Tradeoff table in package README with at least one re-confirmed [M2-ANE] row: <link>
```
---
## Phase E-LCM — FIRST task after the split: clean up LCM + verify → promote (behavior-changing, gated)
> Promoted from "optional, someday" to **the first thing after E5**, per maintainer intent: the
> Suite should expose every verifiably-convertible model, and LCM is the obvious first cleanup.
Two coupled goals:
1. **Consolidate the LCM path.** Make the LCM node use the unified `from_single_file` path in
`coreml_diffusion.convert(model_version=LCM, ...)` instead of the hardcoded `SimianLuo/LCM_Dreamshaper_v7`
HF download; drop the duplicated LCM helpers (already deduped in E2). **Behavior change** ⇒
capture an LCM golden anchor *before* the change, then prove within-tolerance after.
2. **Verify → promote.** Once the LCM conversion has a passing `[M2-ANE]` golden anchor, flip
`_MODEL_STATUS["lcm"] = Status.VERIFIED` **in the package** (minor bump). The Suite's dropdown
gains `lcm` automatically — no Suite change, no Suite bump. This is the end-to-end proof that
the discovery contract works as designed.
Repeat the same recipe for `sdxl_refiner` when it gets an anchor (separate small gate). Do NOT
bundle E-LCM into E1–E5; it changes behavior and must stand on its own golden.
### STOP — VALIDATE (Gate E-LCM)
```
## Gate E-LCM report
- LCM golden anchor captured BEFORE change: <path/hash>
- LCM node now uses unified from_single_file path; HF hardcode removed: confirmed
- [M2-ANE] LCM golden after change: identical / within tol <x> / DIVERGED (STOP)
- Status flipped lcm→VERIFIED in package (minor bump <ver>): confirmed
- Suite dropdown now lists lcm with NO Suite code change / NO Suite bump: confirmed (diff empty)
- LCM node accepts a checkpoint arg now (documented breaking-ish UI note): <link>
```
---
## Article deliverable (after E4)
Once the CLI exists and stands alone, the "convert a Comfy/A1111 workflow into an on-device iOS
app" write-up becomes a clean tutorial: `coreml-diffusion convert` → `.mlmodelc` → load in Swift/CoreML.
Frame the niche honestly per the README note above (ANE feasibility & power, not raw Flux speed).
---
## Quick reference: extraction gate discipline
```
E0 Seam decision, interface contract, discovery API frozen → Gate E0 (cut line + additive-only policy?)
E1 Stand up coreml_diffusion + discovery API (mostly verify) → Gate E1 (list_* byte-identical, comfy-free?)
E2 Move conversion code, dedup LCM/main, inject paths/device→ Gate E2 (golden identical? duplicates gone?) ← first regression gate
E3 Thin nodes onto package + wire discovery dropdowns → Gate E3 (field-names frozen, values additive, fail-soft, golden identical?)
E4 Standalone packaging + CLI (uv) → Gate E4 (uv pip install w/o comfy? CLI golden?) ← proves it stands alone
E5 Physical second-repo split → Gate E5 (suite depends on pkg? conversion absent?)
E-LCM FIRST post-split: clean up LCM, verify → promote → Gate E-LCM (LCM golden? dropdown gains lcm w/ no Suite bump?)
E6 Quantization checkpoint (already built → moved in E2) → Gate E6 (default identical? CLI quant matches?)
(refiner) same recipe as E-LCM when an anchor exists → own small gate (promote sdxl_refiner→VERIFIED)
```
**Interface-contract invariants (the maintainer's hard requirement), restated:**
- Package API is keyword-only-with-defaults past the required positionals → converter updates
don't force Suite updates.
- Node dropdowns are discovery-driven (`coreml_diffusion.list_*`) + fail-soft → new conversion types
appear in the old plugin with `uv pip install -U coreml_diffusion`, **no Suite code change, no bump**.
- Discovery identifiers are **additive-only**; removal/rename = MAJOR bump + migration note.
- `compose_out_name` (cache key) lives in the package, single copy.
**Golden rule (inherited): never cross a gate with a failing acceptance criterion.
Stop, report, wait. The golden latent is the single source of truth that the extraction
changed nothing.**
+93 -432
View File
@@ -1,452 +1,113 @@
# Core ML Suite for ComfyUI # Core ML Suite for ComfyUI
## Overview Custom nodes for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) that run
Stable Diffusion UNets as [Core ML](https://developer.apple.com/documentation/coreml)
models on Apple Silicon (M1/M2/M3). Core ML can use the Apple Neural Engine
(ANE), which is unavailable to PyTorch — on an M2 Pro 32 GB, SD1.5 at 512×512
generates roughly **1.5–2× faster** than the standard PyTorch/MPS path.
Welcome! In this repository you'll find a set of custom nodes for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) You convert a Stable Diffusion checkpoint to a Core ML model with the nodes in
that allows you to use Core ML models in your ComfyUI workflows. this suite, then sample from it like any other ComfyUI workflow.
These models are designed to leverage the Apple Neural Engine (ANE) on Apple Silicon (M1/M2) machines,
thereby enhancing your workflows and improving performance.
If you're not sure how to obtain these models, you can download them > [!IMPORTANT]
[here](https://huggingface.co/coreml-community) or convert your own checkpoints > **Convert your own checkpoints — that is the only supported path.** This
directly with the conversion nodes in this suite (see [How to use](#how-to-use)). > suite uses its own input dimensions, naming convention, and metadata
> (produced by the [coreml-diffusion](https://github.com/aszc-dev/coreml-diffusion)
> package). Pre-converted Core ML models from elsewhere (e.g. the
> coreml-community Hugging Face org) are **not** supported. Conversion is cheap
> and runs on your machine, so there is no need to download Core ML models.
In simple terms, think of Core ML models as a tool that can help your ComfyUI work faster and more efficiently. ## Installation
For instance, during my tests on an M2 Pro 32GB machine,
the use of Core ML models sped up the generation of 512x512 images by a factor
of approximately 1.5 to 2 times.
## Getting Started ### ComfyUI-Manager (recommended)
To start using custom nodes in your ComfyUI, follow these simple steps: Open **Manager → Install Custom Nodes**, search for `Core ML`, click
**Install**, and restart ComfyUI.
1. Clone or download this repository: You can do this directly into the custom_nodes directory of your ComfyUI. ### Manual
2. Install the dependencies: You'll need to use a package manager like pip to do this.
That's it! You're now ready to start enhancing your ComfyUI workflows with Core ML models. ```bash
cd /path/to/comfyui/custom_nodes
git clone https://github.com/aszc-dev/ComfyUI-CoreMLSuite.git
cd ComfyUI-CoreMLSuite
pip install -r requirements.txt
```
- Check [Installation](#installation) for more details on installation. Dependencies (`coreml-diffusion`, `coremltools`, `numpy`, `diffusers`) install
- Check [How to use](#how-to-use) for more details on how to use the custom nodes. from PyPI. PyTorch is intentionally **not** pinned — it is provided by your
- Check [Example Workflows](#example-workflows) for some example workflows. ComfyUI host, and a hard cap here would downgrade it and break ComfyUI.
## Quickstart
1. Put a SD1.5 checkpoint in `models/checkpoints`.
2. Add the **Convert Checkpoint to Core ML** node, select the checkpoint, and
queue once. It writes a `.mlpackage` to `models/unet` (cached by name — it
won't reconvert next time).
3. Sample with the **Core ML Sampler** node, decoding the latent with a normal
VAE Decode. CLIP and VAE come from standard ComfyUI nodes.
See [docs/workflows.md](docs/workflows.md) for complete example graphs (txt2img,
ControlNet, LoRA, LCM, SDXL).
## Which compute unit should I pick?
The **compute unit** selects the hardware Core ML runs on. Pair it with the
attention implementation chosen at conversion time:
| Model | Convert with | Load with | Runs on |
|---|---|---|---|
| SD1.5 @ 512×512 | `SPLIT_EINSUM` | `CPU_AND_NE` | Neural Engine (fastest) |
| SD1.5 @ larger sizes | `ORIGINAL` | `CPU_AND_GPU` | GPU |
| SDXL | `ORIGINAL` | `CPU_AND_GPU` | GPU (ANE unsupported) |
`CPU_AND_NE` is usually the fastest option for SD1.5 — often faster than `ALL`.
This suite uses Core ML compute units only; it never touches PyTorch MPS, so
`PYTORCH_ENABLE_MPS_FALLBACK` is irrelevant to these nodes. Full reasoning and
benchmarks: [docs/hardware.md](docs/hardware.md).
## Documentation
- [Hardware & compute units](docs/hardware.md) — ANE vs GPU vs MPS, attention
implementations, which to choose.
- [Nodes](docs/nodes.md) — full reference for every node.
- [Conversion](docs/conversion.md) — how conversion works, caching,
quantization.
- [Example workflows](docs/workflows.md) — annotated example graphs.
- [FAQ](docs/faq.md) — answers to common questions.
- [Troubleshooting](docs/troubleshooting.md) — common errors and fixes.
- [Limitations & support matrix](docs/limitations.md) — what is and isn't
supported.
## Glossary ## Glossary
- **Core ML**: A machine learning framework developed by Apple. It's used to run machine learning models on Apple - **Core ML** — Apple's on-device machine-learning framework.
devices. - **`.mlpackage`** — the Core ML model format this suite produces and loads.
- **Core ML Model**: A machine learning model that can be run on Apple devices using Core ML. - **ANE** — Apple Neural Engine, a hardware accelerator for ML.
- **mlmodelc**: A compiled Core ML model. This is the recommended format for Core ML models. - **Compute unit** — which hardware Core ML uses (`CPU_AND_NE`, `CPU_AND_GPU`,
- **mlpackage**: A Core ML model packaged in a directory. This is the default format for Core ML models. `CPU_ONLY`, `ALL`).
- **ANE**: Apple Neural Engine. A hardware accelerator for machine learning tasks on Apple devices. - **Attention implementation** — `SPLIT_EINSUM` / `SPLIT_EINSUM_V2` (ANE-friendly)
- **Compute Unit**: A Core ML option that allows you to specify the hardware on which the model should run. or `ORIGINAL` (GPU-friendly), chosen at conversion.
- **CPU_AND_ANE**: A Core ML compute unit option that allows the model to run on both the CPU and ANE. This is the
default option.
- **CPU_AND_GPU**: A Core ML compute unit option that allows the model to run on both the CPU and GPU.
- **CPU_ONLY**: A Core ML compute unit option that allows the model to run on the CPU only.
- **ALL**: A Core ML compute unit option that allows the model to run on all available hardware.
- **CLIP**: Contrastive Language-Image Pre-training. A model that learns visual concepts from natural language
supervision. It's used as a text encoder in Stable Diffusion.
- **VAE**: Variational Autoencoder. A model that learns a latent representation of images. It's used as a prior in
Stable Diffusion.
- **Checkpoint**: A file that contains the weights of a model. It's used to load models in Stable Diffusion.
- **LCM**: [Latent Consistency Model](https://latent-consistency-models.github.io/). A type of model designed to
generate images with as few steps as possible.
> [!NOTE]
> Note on Compute Units:
> For the model to run on the ANE, the model must be converted with the `--attention-implementation SPLIT_EINSUM`
> option.
> Models converted with `--attention-implementation ORIGINAL` will run on GPU instead of ANE.
## Features
These custom nodes come with a host of features, including:
- Loading Core ML Unet models
- Support for ControlNet
- Support for ANE (Apple Neural Engine)
- Support for CPU and GPU
- Support for `mlmodelc` and `mlpackage` files
- Support for SDXL models
- Support for LCM models
- Support for LoRAs
- SD1.5 -> Core ML conversion
- SDXL -> Core ML conversion
- LCM -> Core ML conversion
> [!NOTE]
> Please note that using Core ML models can take a bit longer to load initially.
> For the best experience, I recommend using the compiled models
> (.mlmodelc files) instead of the .mlpackage files.
> [!NOTE]
> This repository will continue to be updated with more nodes and features over time.
## Conversion & Acknowledgements
The Core ML conversion pipeline in this repository began as an adaptation of
Apple's [ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion),
which pioneered running Stable Diffusion on the Apple Neural Engine. The
implementation has since diverged and no longer depends on that package:
- UNet conversion runs natively on `diffusers`' `UNet2DConditionModel`.
- The ANE-friendly attention path (`SPLIT_EINSUM`, `SPLIT_EINSUM_V2`) is
reimplemented as standalone `diffusers` attention processors.
- The toolchain tracks current ComfyUI (NumPy 2, Torch 2.7, coremltools 9,
Python 3.12).
The goal is to keep iterating on these methods independently and to explore
support beyond SD1.5.
> [!IMPORTANT] > [!IMPORTANT]
> **Breaking change in 2.0.0.** The converted Core ML UNet now takes > **Breaking change in 2.0.0.** The converted Core ML UNet now takes
> `encoder_hidden_states` in the native `diffusers` layout > `encoder_hidden_states` in the native `diffusers` layout
> `(batch, tokens, hidden)` instead of the previous > `(batch, tokens, hidden)` instead of the previous `(batch, hidden, 1, tokens)`.
> `(batch, hidden, 1, tokens)`. Core ML models converted with earlier versions > Models converted with earlier versions are not compatible and must be
> are not compatible with 2.0.0 and must be re-converted. > re-converted.
## Installation ## Acknowledgements
### Using ComfyUI-Manager The conversion pipeline began as an adaptation of Apple's
[ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion), which
The easiest way to install the custom nodes is to use the ComfyUI-Manager. You can find the installation instructions pioneered running Stable Diffusion on the Neural Engine. It has since diverged
[here](https://github.com/ltdrdata/ComfyUI-Manager#installation). Once you've installed the ComfyUI-Manager, you can and no longer depends on that package: UNet conversion runs natively on
install the custom nodes by following these steps: `diffusers`' `UNet2DConditionModel`, the ANE attention path (`SPLIT_EINSUM`,
`SPLIT_EINSUM_V2`) is reimplemented as standalone `diffusers` attention
- Open the ComfyUI-Manager by clicking the `Manager` button in the ComfyUI toolbar. processors, and the toolchain tracks current ComfyUI (NumPy 2, Torch 2.7+,
- Click the `Install Custom Nodes` button. coremltools 9, Python 3.12+). Conversion now lives in the separate
- Search for `Core ML` and click the `Install` button. [coreml-diffusion](https://github.com/aszc-dev/coreml-diffusion) package.
- Restart ComfyUI.
### Manual Installation
1. Clone this repository into the custom_nodes directory of your ComfyUI. If you're not sure how to do this, you can
download the repository as a zip file and extract it into the same directory.
```bash
cd /path/to/comfyui/custom_nodes
git clone https://github.com/aszc-dev/ComfyUI-CoreMLSuite.git
```
2. Next, install the required dependencies using pip or another package manager:
```bash
cd /path/to/comfyui/custom_nodes/ComfyUI-CoreMLSuite
pip install -r requirements.txt
```
## How to use
Once you've installed the custom nodes, you can start using them in your ComfyUI workflows.
To do this, you need to add the nodes to your workflow. You can do this by right-clicking on the workflow canvas and
selecting the nodes from the list of available nodes (the nodes are in the `Core ML Suite` category).
You can also double-click the canvas and use the search bar to find the nodes. The list of available nodes is given
below.
### Available Nodes
#### Core ML UNet Loader (`CoreMLUnetLoader`)
![CoreMLUnetLoader](./assets/unet_loader.png?raw=true)
This node allows you to load a Core ML UNet model and use it in your ComfyUI workflow. Place the converted
.mlpackage or .mlmodelc file in ComfyUI's `models/unet` directory and use the node to load the model. The output of the
node is a `coreml_model` object that can be used with the Core ML Sampler.
- **Inputs**:
- **model_name**: The name of the model to load. This should be the name of the .mlpackage or .mlmodelc file.
- **compute_unit**: The hardware on which the model should run. This can be one of the following:
- `CPU_AND_ANE`: The model will run on both the CPU and ANE. This is the default option. It works best with
models
converted with `--attention-implementation SPLIT_EINSUM` or `--attention-implementation SPLIT_EINSUM_V2`.
- `CPU_AND_GPU`: The model will run on both the CPU and GPU. It works best with models converted with
`--attention-implementation ORIGINAL`.
- `CPU_ONLY`: The model will run on the CPU only.
- `ALL`: The model will run on all available hardware.
- **Outputs**:
- **coreml_model**: A Core ML model that can be used with the Core ML Sampler.
#### Core ML Sampler (`CoreMLSampler`)
![CoreMLSampler](./assets/sampler.png?raw=true)
This node allows you to generate images using a Core ML model. The node takes a Core ML model as input and outputs a
latent image similar to the latent image output by the KSampler. This means that you can use the
resulting latent as you normally would in your workflow.
- **Inputs**:
- **coreml_model**: The Core ML model to use for sampling. This should be the output of the Core ML UNet Loader.
- **latent_image** [optional]: The latent image to use for sampling. If provided, should be of the same size as the
input of the Core ML model. If not provided, the node will create a latent suitable for the Core ML model used.
Useful in img2img workflows.
- ... _(the rest of the inputs are the same as the KSampler)_
- **Outputs**:
- **LATENT**: The latent image output by the Core ML model. This can be decoded using a VAE Decoder or used as input
to the next node in your workflow.
#### Checkpoint Converter
![CoreMLConverter](./assets/checkpoint_converter.png?raw=true)
You can use this node to convert any **SD1.5** based checkpoint to a Core ML model. The converted model is stored in the
`models/unet` directory and can be used with the `Core ML UNet Loader`. The conversion parameters are encoded in
the node name, so if the model already exists, the node will not convert it again.
- **Inputs**:
- **ckpt_name**: The name of the checkpoint to convert. This should be the name of the checkpoint file stored in the
`models/checkpoints` directory.
- **model_version**: Whether the model is based on SD1.5 or SDXL.
- **height**: The desired height of the image generated by the model. The default is 512. Any positive multiple of 8 is accepted.
- **width**: The desired width of the image generated by the model. The default is 512. Any positive multiple of 8 is accepted.
- **batch_size**: The batch size of generated images. If you're planning to generate batches of images, you can try
increasing this value to speed up the generation process. The default is 1.
- **attention_implementation**: The attention implementation used when converting the model. Choose SPLIT_EINSUM or
SPLIT_EINSUM_V2 for better ANE support. Choose ORIGINAL for better GPU support.
- **compute_unit**: The hardware on which the model should run. This is used only when loading the model and doesn't
affect the conversion process.
- **controlnet_support**: For the model to support ControlNet, it must be converted with this option set to True.
The
default is False.
- **lora_params** [optional]: Optional LoRA names and weights. If provided, the model will be converted with LoRA(s)
baked in. More on loading LoRAs below.
- **Outputs**:
- **coreml_model**: The converted Core ML model that can be used with Core ML Sampler.
> [!NOTE]
> Some models use a custom config .yaml file. If you're using such a model, you'll need to place the config file in the
> `models/configs` directory. The config file should be named the same as the checkpoint file. For example, if the
> checkpoint file is named `juggernaut_aftermath.safetensors`, the config file should be
> named `juggernaut_aftermath.yaml`.
> The config file will be automatically loaded during conversion.
> [!NOTE]
> For now, the converter relies heavilty on the model name to determine the conversion parameters. This means that if
> you change the model name, the node will convert the model again. Other than that, if you find the name too long or
> confusing, you can change it to anything you want.
#### LoRA Loader
![LoRALoader](./assets/lora_loader.png?raw=true)
This node allows you to load LoRAs and bake them into a model. Since this is a workaround (as model weights can't be
modified
after conversion), there are a few caveats to keep in mind:
- The LoRA weights and _strength_model_ parameter are baked into the model. This means that you can't change them
after conversion. This also means that you need to convert the model again if you want to change the LoRA weights.
- Loading LoRA affects CLIP, which is not a part of Core ML workflow, so you'll need to load CLIP separately,
either using `CLIPLoader` or `CheckpointLoaderSimple`. (See [example workflows](#example-workflows) for more details.)
- After conversion, if you want to load the model using `CoreMLUnetLoader`, you'll need to apply the same LoRAs to
CLIP manually. (See [example workflows](#example-workflows) for more details.)
- The LoRA names are encoded in the model name. This means that if you change the name of the LoRA file,
you'll need to change the model name as well, or the node will convert the model again. (Model strength is not
encoded, so if you want to change it, you'll need to delete the converted model manually)
- _strength_clip_ parameter only affects the CLIP model and is not baked into the converted model. This means that
you can change it after conversion.
- **Inputs**:
- **lora_name**: The name of the LoRA to load.
- **strength_model**: The strength of the LoRA model.
- **strength_clip**: The strength of the LoRA CLIP.
- **lora_params** [optional]: Optional output from other LoRA Loaders.
- **clip**: The CLIP model to use with the LoRA. This can be either output of the
`CLIPLoader`/`CheckpointLoaderSimple` or other LoRA Loaders.
- **Outputs**:
- **lora_params**: The LoRA parameters that can be passed to the Core ML Converter or other LoRA Loaders.
- **CLIP**: The CLIP model with LoRA applied.
#### LCM Converter
![LCMConverter](./assets/lcm_converter.png?raw=true)
This node converts [SimianLuo/LCM_Dreamshaper_v7](https://huggingface.co/SimianLuo/LCM_Dreamshaper_v7) model to Core
ML. The converted model is stored in the `models/unet` directory and can be used with the Core ML UNet Loader. The
conversion parameteres are encoded in the node name, so if the model already exists, the node will not convert it again.
- **Inputs**:
- **height**: The desired height of the image generated by the model. The default is 512. Must be a multiple of 8.
- **width**: The desired width of the image generated by the model. The default is 512. Must be a multiple of 8.
- **batch_size**: The batch size of generated images. If you're planning to generate batches of images, you can try
increasing this value to speed up the generation process. The default is 1.
- **compute_unit**: The hardware on which the model should run. This is used only when loading the model and
doesn't affect the conversion process.
- **controlnet_support**: For the model to support ControlNet, it must be converted with this option set to True.
The default is False.
> [!NOTE]
> The conversion process can take a while, so please be patient.
> [!NOTE]
> When using the LCM model with Core ML Sampler, please set _sampler_name_ to `lcm` and _scheduler_ to `sgm_uniform`.
#### Core ML Adapter (Experimental) (`CoreMLModelAdapter`)
![CoreMLModelAdapter](./assets/adapter.png?raw=true)
This node allows you to use a Core ML as a standard ComfyUI model. This is an experimental node and may not work with
all models and nodes. Please use with caution and pay attention to the expected inputs of the model.
- **Input**:
- **coreml_model**: The Core ML model to use as a ComfyUI model.
- **Output**:
- **MODEL**: The Core ML model wrapped in a ComfyUI model.
> [!NOTE]
> While this approach allows you to use Core ML models with many ComfyUI nodes (both standard and custom), the
> expected inputs of the model will not be checked, which may cause errors. Please make sure to use a model compatible
> with the expected parameters.
### Example Workflows
> [!NOTE]
> The models used are just an example. Feel free to experiment with different models and see what works best for you.
#### Basic txt2img with Core ML UNet loader
This is a basic txt2img workflow that uses the Core ML UNet loader to load a model. The CLIP and VAE models
are loaded using the standard ComfyUI nodes. In the first example, the text encoder (CLIP) and VAE models are loaded
separately. In the second example, the text encoder and VAE models are loaded from the checkpoint file. Note that you
can use any CLIP or VAE model as long as it's compatible with Stable Diffusion v1.5.
1. **Loading text encoder (CLIP) and VAE models separately**
- This workflow uses CLIP and VAE models available
[here](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/text_encoder/model.safetensors) and
[here](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/vae/diffusion_pytorch_model.safetensors).
Once downloaded, place the models in the`models/clip` and `models/vae` directories respectively.
- The Core ML UNet model is available
[here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip).
Once downloaded, place the model in the `models/unet` directory.
![coreml-unet+clip+vae](./assets/unet+sampler+clip+vae.png?raw=true)
2. **Loading text encoder (CLIP) and VAE models from checkpoint file**
- This workflow loads the CLIP and VAE models from the checkpoint file available
[here](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/v1-5-pruned-emaonly.safetensors).
Once downloaded, place the model in the`models/checkpoints` directory.
- The Core ML UNet model is available
[here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip).
Once downloaded, place the model in the `models/unet` directory.
![coreml-unet+checkpoint](./assets/unet+sampler+checkpoint.png?raw=true)
#### ControlNet with Core ML UNet loader
This workflow uses the Core ML UNet loader to load a Core ML UNet model that supports ControlNet. The ControlNet is
being loaded using the standard ComfyUI nodes. Please refer to
the [basic txt2img workflow](#basic-txt2img-with-core-ml-unet-loader) for more details on how to load the CLIP and VAE
models.
The ControlNet model used in this workflow is available
[here](https://huggingface.co/lllyasviel/control_v11p_sd15_scribble/blob/main/diffusion_pytorch_model.fp16.safetensors).
Once downloaded, place the model in the `models/controlnet` directory.
![coreml-unet+controlnet](./assets/unet+sampler+controlnet.png?raw=true)
#### Checkpoint conversion
This workflow uses the Checkpoint Converter to convert the checkpoint file. See
[Checkpoint Converter](#checkpoint-converter) description for more details.
![checkpoint-converter](./assets/basic_conversion.png?raw=true)
#### Checkpoint conversion with LoRA
This workflow uses the Checkpoint Converter to convert the checkpoint file with LoRA. See
[LoRA Loader](#lora-loader) description to read more about the caveats of using LoRA.
![checkpoint-converter+lora](./assets/conversion+lora.png?raw=true)
#### LCM LoRA conversion
Please note that you can use multiple LoRAs with the same model. To do this, you'll need to use multiple LoRA Loaders.
> [!IMPORTANT]
> In this example, the model is passed through the adapter and `ModelSamplingDiscrete` nodes to a standard ComfyUI's
> KSampler (not Core ML Sampler). ModelSamplingDiscrete needs to be used to sample models with LCM LoRAs properly.
![multiple-loras](./assets/conversion+lcm_lora.png?raw=true)
#### Loader with LoRAs
This workflow uses the Core ML UNet Loader to load a model with LoRAs. The CLIP must be loaded separately and passed
through the same LoRA nodes as during conversion. See [LoRA Loader](#lora-loader) description to read more about the
caveats of using LoRA. Since _lora_name_ and _strength_model_ are baked into the model, it is not necessary to pass
them as inputs to the loader.
> [!IMPORTANT]
> In this example, the model is passed through the adapter and `ModelSamplingDiscrete` nodes to a standard ComfyUI's
> KSampler (not Core ML Sampler). ModelSamplingDiscrete needs to be used to sample models with LCM LoRAs properly.
![loader+lora](./assets/loader+lcm_lora.png?raw=true)
#### LCM conversion with ControlNet
This workflow uses LCM converter to
convert [SimianLuo/LCM_Dreamshaper_v7](https://huggingface.co/SimianLuo/LCM_Dreamshaper_v7)
model to Core ML. The converted model can then be used with or without ControlNet to generate images.
![lcm+controlnet](./assets/lcm+controlnet.png?raw=true)
#### SDXL Base + Refiner conversion
This is a basic workflow for SDXL. You add LoRAs and ControlNets the same way as in the previous examples.
You can also skip the refiner step.
The models used in this workflow are available at the following links:
- [Base model + text_encoder (clip) + text_encoder_2 (clip2)](https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0)
- [Refiner model](https://huggingface.co/stabilityai/stable-diffusion-xl-refiner-1.0)
- [VAE](https://huggingface.co/stabilityai/sdxl-vae)
> [!IMPORTANT]
> SDXL on ANE is not supported. If loading of the model gets stuck, please try using CPU_AND_GPU or CPU_ONLY.
> For best results, use ORIGINAL attention implementation.
![sdxl](./assets/sdxl_conversion.png?raw=true)
## Quantization (opt-in)
The `Core ML Converter` and `Core ML LCM Converter` nodes accept an
optional `quantize_nbits` dropdown that runs k-means weight palettization
(`coremltools.optimize.coreml.palettize_weights`) on the UNet before save.
Values: `none` (default — no quantization, identical to unquantized
behavior and filenames), `8`, `6`, `4`. The number is appended to the
.mlpackage stem as `_q<bits>` so quantized and unquantized variants
coexist on disk and in cache.
### SD1.5 1×512×512 SPLIT_EINSUM tradeoffs (M2 Pro, ANE)
Measured with 20 UNet forward passes at a fixed seed for the PSNR
comparison:
| nbits | size (MB) | size vs none | fwd median (ms) | PSNR vs `none` (dB) |
|---|---:|---:|---:|---:|
| none | 1641 | 1.000 | 197.1 | — |
| 8 | 822 | 0.501 | 186.6 | 53.5 |
| 6 | 617 | 0.376 | 183.0 | 40.2 |
| 4 | 412 | 0.251 | 179.8 | 27.5 |
PSNR here is computed on the raw `noise_pred` output of a single UNet
forward at a fixed seed, not on the final decoded image — it isolates
the quantization-induced drift from sampler / VAE noise. Final-image
PSNR is comfortably higher (the sampler averages over 20 steps).
### Recommended settings per chip / RAM
- **8 GB RAM (M1 base, M2 base):** `nbits=4`. ~4× smaller model, still
loads, PSNR 27 dB is visually identical at SD1.5 sizes.
- **16 GB RAM (M1/M2/M3 Pro):** `nbits=6` is the sweet spot — ~2.7×
smaller, PSNR 40 dB, no perceptible quality drop.
- **32 GB+ RAM (Max / Ultra):** `nbits=8` if you want the safety
margin, `none` if you want bit-identical output for golden testing.
The default stays `none` so existing workflows produce byte-for-byte
identical output.
## Limitations
- Core ML models are fixed in terms of their inputs and outputs.
This means you'll need to use latent images of the same size as the input of the model (512x512 is the default for
SD1.5).
However, you can re-convert the model to a different input size using the
conversion nodes in this suite (set the desired width and height).
- SD2.1 models are not supported.
[^1]:
Unless [EnumeratedShapes](https://apple.github.io/coremltools/docs-guides/source/flexible-inputs.html#select-from-predetermined-shapes)
is used during conversion. Needs more testing.
## Support ## Support
I'm here to help! If you have any questions or suggestions, don't hesitate to open an issue and I'll do my best Questions or suggestions? Open an
to assist you. [issue](https://github.com/aszc-dev/ComfyUI-CoreMLSuite/issues).
+93
View File
@@ -0,0 +1,93 @@
# Conversion
## Conversion is the only supported path
You always start from a Stable Diffusion checkpoint (`.safetensors` / `.ckpt`)
and convert it with the **Convert Checkpoint to Core ML** node. The model
version (SD1.5, SDXL, SDXL refiner, full-distill LCM) is auto-detected from the
checkpoint. Pre-converted Core ML models from elsewhere are not supported,
because:
- The suite uses its own input **dimensions**, **naming convention**, and
**metadata**, all produced by the
[coreml-diffusion](https://github.com/aszc-dev/coreml-diffusion) package.
- Apple's `ml-stable-diffusion` (which most community Core ML models target) is
effectively obsolete, and the layouts differ (see the 2.0.0
`encoder_hidden_states` change in the README).
- Conversion is cheap and one-time, so there is no value in maintaining
backwards compatibility with foreign formats.
The output is always a **`.mlpackage`**. The suite no longer compiles to
`.mlmodelc` (it didn't work with the inference backend), so there is **no Xcode
or `coremlcompiler` dependency**.
## One-time conversion and name-based caching
Conversion runs **once**, not on every queue. The converter encodes all
conversion parameters into the output filename (via `coreml_diffusion.compose_out_name`,
called in `coreml_suite/nodes.py`):
- checkpoint name, `batch_size`, `width`, `height`
- `controlnet_support`, `attention_implementation`
- baked LoRA names, `quantize_nbits`
The result is written as `<encoded-name>_unet.mlpackage` in `models/unet`. If a
file with that name already exists, it is reused and conversion is skipped. Change
any parameter → new name → new conversion; keep them the same → the cached model
is loaded instantly.
This is why the recommended workflow is **convert once, then load**: run the
converter a single time, then in day-to-day use load the `.mlpackage` with the
**Load Core ML UNet** node. (You can also leave the converter node in the graph;
it short-circuits to the cached file.)
> [!NOTE]
> The converter relies on the filename to decide whether to reconvert. If you
> rename the `.mlpackage`, it will be converted again. You can otherwise rename
> it freely if the auto-generated name is too long.
## Quantization
The converter node accepts an optional `quantize_nbits` dropdown that runs
k-means weight palettization (`coremltools.optimize.coreml.palettize_weights`) on
the UNet before saving.
Values: `none` (default — no quantization, identical output and filename to
before), `8`, `6`, `4`. The number is appended to the `.mlpackage` stem as
`_q<bits>`, so quantized and unquantized variants coexist on disk and in cache.
### SD1.5 1×512×512 SPLIT_EINSUM tradeoffs (M2 Pro, ANE)
Measured with 20 UNet forward passes at a fixed seed:
| nbits | size (MB) | size vs none | fwd median (ms) | PSNR vs `none` (dB) |
|---|---:|---:|---:|---:|
| none | 1641 | 1.000 | 197.1 | — |
| 8 | 822 | 0.501 | 186.6 | 53.5 |
| 6 | 617 | 0.376 | 183.0 | 40.2 |
| 4 | 412 | 0.251 | 179.8 | 27.5 |
PSNR here is computed on the raw `noise_pred` output of a single UNet forward at a
fixed seed, not on the final decoded image — it isolates quantization drift from
sampler/VAE noise. Final-image PSNR is comfortably higher (the sampler averages
over many steps).
### Recommended settings per chip / RAM
- **8 GB (M1/M2 base):** `nbits=4`. ~4× smaller, still loads, 27 dB is visually
identical at SD1.5 sizes.
- **16 GB (M1/M2/M3 Pro):** `nbits=6` — the sweet spot, ~2.7× smaller, 40 dB, no
perceptible quality drop.
- **32 GB+ (Max / Ultra):** `nbits=8` for a safety margin, or `none` for
bit-identical output (golden testing).
The default stays `none`, so existing workflows produce byte-for-byte identical
output.
## Where conversion lives
The conversion engine was extracted into the standalone
[coreml-diffusion](https://github.com/aszc-dev/coreml-diffusion) PyPI package.
The nodes in this suite resolve ComfyUI paths and call into it; node names,
inputs, and outputs are unchanged, so the split has effectively no user-facing
impact beyond `pip install` pulling one more dependency.
+77
View File
@@ -0,0 +1,77 @@
# FAQ
## What's the difference between ANE, GPU, and MPS, and which do I pick?
ANE is the Neural Engine (Core ML only), GPU is the Metal GPU (Core ML or
PyTorch), MPS is PyTorch's GPU backend. This suite uses **Core ML compute units
only** and never touches MPS. Short answer: SD1.5 at 512×512 → convert
`SPLIT_EINSUM`, load `CPU_AND_NE`; larger sizes or SDXL → convert `ORIGINAL`,
load `CPU_AND_GPU`. Full reasoning: [hardware](hardware.md).
## Do I still need `PYTORCH_ENABLE_MPS_FALLBACK=1`?
Not for these nodes — Core ML inference doesn't use PyTorch MPS. It may still
matter for other parts of your ComfyUI graph, but it has no effect on Core ML
sampling.
## Why is my Core ML SDXL workflow no faster than the default nodes?
Because **SDXL can't run on the ANE** — the speedup comes from the Neural Engine,
and SDXL falls back to the GPU, running at roughly MPS-equivalent speed. This is a
known limitation, not a misconfiguration. The ANE benefit is real for SD1.5. See
[limitations](limitations.md).
## Where do I get Core ML models?
You convert them yourself — that's the only supported path. See
[conversion](conversion.md). Downloaded Core ML models (e.g. coreml-community) use
different dimensions/metadata and are not supported.
## Is conversion run every time I queue, or once?
Once. Parameters are encoded in the output filename, so an already-converted model
is reused and conversion is skipped. Convert once, then load the `.mlpackage`. See
[conversion → caching](conversion.md#one-time-conversion-and-name-based-caching).
## Does a converted model produce the same output as the original?
With the default `quantize_nbits = none`, the converted UNet output matches the
source within numerical rounding (the golden test in `tests/m2/test_golden_image.py`
gates on PSNR ≥ 20 dB on the decoded image). Quantization (`8`/`6`/`4`) introduces
measured, bounded drift — see the [PSNR table](conversion.md#quantization). For
bit-identical output, keep `none`.
## Are `.mlpackage` models safe to use?
`.mlpackage` is a declarative Core ML model format — it carries weights and a
compute graph, not arbitrary executable code or Python pickle, so its safety
profile is comparable to `safetensors`. In practice this matters little here,
since the only supported models are ones you convert locally from your own
checkpoints.
## Are LoRAs reliable?
Partially. Some LoRAs convert cleanly; others produce poor or broken output —
there's no firm rule, so test per-LoRA. LoRA weights and `strength_model` are
baked in at conversion and can't be changed afterward; for some LCM-LoRA cases the
[Core ML Adapter](nodes.md#core-ml-adapter-experimental-coremlmodeladapter) path
is more reliable. Treat LoRA support as experimental. See
[troubleshooting](troubleshooting.md).
## Does the experimental Adapter cost performance vs the Core ML Sampler?
Yes, a little. The Adapter wraps the model in a ComfyUI `ModelPatcher` so standard
samplers work, which adds per-step interface overhead the native Core ML Sampler
avoids. Use the native sampler unless you specifically need a `MODEL` (e.g.
`ModelSamplingDiscrete` for LCM LoRAs).
## Which Python versions work?
Python 3.12 or newer (`requires-python >=3.12`). Older 3.12 install failures
came from the now-removed `ml-stable-diffusion` build, not from this suite.
## Long prompts crash my workflow
Core ML has a hard **77-token** prompt limit and doesn't auto-chunk long prompts.
Split the prompt across multiple CLIP Text Encode nodes and merge with Conditioning
(Combine). See [troubleshooting](troubleshooting.md).
+95
View File
@@ -0,0 +1,95 @@
# Hardware & Compute Units
This page explains how the suite maps to Apple Silicon hardware, the difference
between ANE, GPU, and MPS, and how to choose a compute unit and attention
implementation.
## ANE vs GPU vs MPS
Three terms get conflated:
- **ANE (Apple Neural Engine)** — a dedicated ML accelerator on Apple Silicon.
Only Core ML can target it; PyTorch cannot. This is the whole reason the suite
exists.
- **GPU** — the Metal GPU. Reachable both by Core ML (as a compute unit) and by
PyTorch (via MPS).
- **MPS (Metal Performance Shaders)** — PyTorch's GPU backend on macOS. This is
the path standard ComfyUI nodes use.
**This suite uses Core ML compute units only — it never runs the UNet through
PyTorch/MPS.** Consequently `PYTORCH_ENABLE_MPS_FALLBACK` has no effect on these
nodes. It may still matter for the rest of your ComfyUI graph (CLIP, VAE,
samplers on non-Core ML models), but not for Core ML inference itself.
Rough performance picture (SD1.5, maintainer- and user-reported):
- ANE is meaningfully faster than MPS — on the order of **50–100%** for SD1.5.
- Core ML on the GPU is only marginally faster than PyTorch/MPS.
So the speedup comes from the Neural Engine, which means it depends on being able
to actually run on the ANE (see [attention implementations](#attention-implementations)
and the [SDXL caveat](#sdxl-and-the-ane)).
## Compute units
The **compute unit** is set on the loader/converter node and tells Core ML which
hardware to use. It is applied when the model is loaded
(`coreml_suite/coreml_model.py:22`), not during conversion.
| Value | Hardware | Best paired with |
|---|---|---|
| `CPU_AND_NE` (default) | CPU + Neural Engine | `SPLIT_EINSUM` / `SPLIT_EINSUM_V2` |
| `CPU_AND_GPU` | CPU + Metal GPU | `ORIGINAL` |
| `CPU_ONLY` | CPU only | fallback / debugging |
| `ALL` | all available hardware | rarely optimal — see below |
Notes:
- Every option includes the CPU; there is no GPU-and-ANE-without-CPU combination.
- `NE` in `CPU_AND_NE` is the Neural Engine (Apple's enum spells it `NE`, not
`ANE`).
- **`CPU_AND_NE` is often faster than `ALL`.** Letting Core ML use everything can
be *slower* on non-Max chips, where memory bandwidth is the bottleneck. Try
`CPU_AND_NE` first for SD1.5.
## Attention implementations
Chosen at conversion time on the **Convert Checkpoint to Core ML** node. It
decides whether the model can run on the ANE:
- **`SPLIT_EINSUM`** — ANE-friendly attention. Use for the Neural Engine.
- **`SPLIT_EINSUM_V2`** — a variant; in practice ≈ `SPLIT_EINSUM` for most users.
- **`ORIGINAL`** — standard attention. Runs on the GPU, not the ANE.
The implementation and the compute unit must agree: a `SPLIT_EINSUM` model wants
`CPU_AND_NE`; an `ORIGINAL` model wants `CPU_AND_GPU`.
## Which should I pick?
| Scenario | Attention | Compute unit |
|---|---|---|
| SD1.5 at 512×512 | `SPLIT_EINSUM` | `CPU_AND_NE` |
| SD1.5 at larger sizes (e.g. 768) | `ORIGINAL` | `CPU_AND_GPU` |
| SDXL / SDXL Turbo | `ORIGINAL` | `CPU_AND_GPU` |
### Resolution crossover
ANE shines at small latents; the GPU scales better as resolution grows. In user
benchmarks:
- At **512×512**, ANE + `SPLIT_EINSUM` wins by roughly **10%** over the GPU path.
- At **768×512**, GPU + `ORIGINAL` pulls ahead by roughly **10%**, and the larger
image is about 2× slower overall.
If you mostly work at 512×512, convert with `SPLIT_EINSUM` and load on
`CPU_AND_NE`. If you routinely go larger, an `ORIGINAL` + GPU model may be
faster.
### SDXL and the ANE
SDXL (and SDXL Turbo) **cannot run on the ANE** — the dual-text-encoder UNet
exceeds what the Neural Engine path supports. SDXL therefore runs at roughly
MPS-equivalent speed with no ANE speedup. If a Core ML SDXL workflow feels no
faster than the standard nodes, this is why. Convert SDXL with `ORIGINAL` and
load with `CPU_AND_GPU` or `CPU_ONLY`. See
[limitations](limitations.md) for the full picture.
+53
View File
@@ -0,0 +1,53 @@
# Limitations & Support Matrix
## Support matrix
| Feature | Status | Notes |
|---|---|---|
| SD1.5 | ✅ Full | ANE via `SPLIT_EINSUM`; the primary, fastest path |
| SDXL / SDXL Turbo | ⚠️ Partial | GPU only (no ANE), no speedup; possible quality loss vs source. Don't run Turbo at 1024² |
| SD2.1 | ❌ Unsupported | |
| Inpainting checkpoints (9-channel) | ❌ Unsupported | |
| ControlNet | ✅ Supported | Convert the checkpoint with `controlnet_support = True` |
| LoRA | ⚠️ Experimental | Inconsistent per-LoRA; baked at conversion, immutable afterward |
| LCM | ⚠️ Experimental | Full-distill LCM checkpoints auto-detected by the converter |
| SVD | ❌ Not supported | |
| AnimateDiff | ❌ Not supported | Motion modules need pre-conversion injection; not feasible today |
| IPAdapter | ❌ Not supported | Needs a real `MODEL` the Core ML wrapper can't provide |
| Core ML Adapter | ⚠️ Experimental | Works for many nodes; fails for merges/IPAdapter/etc. |
## Fixed input/output shapes
A Core ML model is converted for one specific resolution and batch size. To work
at a different size, re-convert with the new width/height (conversion is cheap and
cached by name). This is also why detailers and latent-upscale workflows that
rescale mid-graph break — see [troubleshooting](troubleshooting.md).
There is experimental support for flexible shapes via
[EnumeratedShapes](https://apple.github.io/coremltools/docs-guides/source/flexible-inputs.html#select-from-predetermined-shapes),
but it is **much slower** — user benchmarks show roughly **5×** the per-iteration
time on every run, not just the first. Fixed-shape models per resolution are the
practical choice.
## SDXL on the Neural Engine
SDXL and SDXL Turbo cannot run on the ANE — the dual-text-encoder UNet exceeds the
supported Neural Engine path. They run on the GPU at roughly MPS-equivalent speed,
so Core ML offers no speed advantage for SDXL, and converted output may look
degraded versus the safetensors original (an upstream conversion artifact). Use
`ORIGINAL` + `CPU_AND_GPU`. See [hardware](hardware.md).
## Experimental Core ML Adapter
The Adapter wraps a Core ML model to look like a standard ComfyUI `MODEL`, which
covers many standard and custom nodes. But it can't fully emulate a real model:
operations that need genuine `MODEL` internals — model merges, IPAdapter, some
LoRA flows, detailers — generally won't work, and the model's
fixed input shapes aren't validated, so mismatches error at runtime. Prefer the
native Core ML Sampler when you don't need the `MODEL` type.
## Prompt length
Core ML enforces a hard 77-token prompt limit with no auto-chunking. Split long
prompts across multiple CLIP Text Encode nodes and merge with Conditioning
(Combine).
+157
View File
@@ -0,0 +1,157 @@
# Node Reference
All nodes live in the **Core ML Suite** category. Right-click the canvas →
**Add Node → Core ML Suite**, or double-click and search.
| Display name | Class | Purpose |
|---|---|---|
| Load Core ML UNet | `CoreMLUNetLoader` | Load a converted `.mlpackage` |
| Core ML Sampler | `CoreMLSampler` | Sample (KSampler-style) |
| Core ML Sampler (Advanced) | `CoreMLSamplerAdvanced` | Sample (KSamplerAdvanced-style) |
| Core ML Adapter (Experimental) | `CoreMLModelAdapter` | Wrap as a standard `MODEL` |
| Load LoRA to use with Core ML | `Core ML LoRA Loader` | Bake LoRA(s) at conversion |
| Convert Checkpoint to Core ML | `Core ML Converter` | Convert a checkpoint |
---
## Load Core ML UNet (`CoreMLUNetLoader`)
![Load Core ML UNet](../assets/unet_loader.png?raw=true)
Loads a converted `.mlpackage` from `models/unet` and outputs a `coreml_model`
for the samplers. Only `.mlpackage` files are listed — this suite no longer uses
`.mlmodelc`.
- **Inputs**
- `coreml_name` — the `.mlpackage` to load from `models/unet`.
- `compute_unit` — hardware to run on: `CPU_AND_NE` (default), `CPU_AND_GPU`,
`CPU_ONLY`, `ALL`. See [hardware](hardware.md).
- **Output**
- `coreml_model` — for the Core ML Sampler or Adapter.
---
## Core ML Sampler (`CoreMLSampler`)
![Core ML Sampler](../assets/sampler.png?raw=true)
Generates a latent from a Core ML model. Behaves like the standard KSampler and
outputs a `LATENT` you can decode or feed downstream.
- **Inputs**
- `coreml_model` — output of the loader or a converter.
- `latent_image` *(optional)* — must match the model's input size. If omitted,
a suitable empty latent is created. Provide one for img2img.
- `negative` *(optional)* — required for normal models; optional for LCM.
- Remaining inputs (`seed`, `steps`, `cfg`, `sampler_name`, `scheduler`,
`positive`, `denoise`) match the KSampler.
- **Output**
- `LATENT` — decode with a VAE Decode, or use downstream.
---
## Core ML Sampler (Advanced) (`CoreMLSamplerAdvanced`)
The KSamplerAdvanced counterpart of the Core ML Sampler — same Core ML input,
plus the advanced sampling controls. Use it for partial denoising, fixed noise,
and multi-stage (e.g. SDXL base → refiner) workflows.
- **Inputs**
- `coreml_model` — output of the loader or a converter.
- `add_noise`, `noise_seed`, `start_at_step`, `end_at_step`,
`return_with_leftover_noise` — as in KSamplerAdvanced.
- `steps`, `cfg`, `sampler_name`, `scheduler`, `positive` — as usual.
- `latent_image` *(optional)*, `negative` *(optional, required for non-LCM)*.
- **Output**
- `LATENT`.
---
## Core ML Adapter (Experimental) (`CoreMLModelAdapter`)
![Core ML Adapter](../assets/adapter.png?raw=true)
Wraps a Core ML model so it presents as a standard ComfyUI `MODEL`, letting you
feed it to the normal KSampler and many other nodes (e.g. `ModelSamplingDiscrete`
for LCM LoRAs).
- **Input**
- `coreml_model`.
- **Output**
- `MODEL` — a Core ML model wrapped as a ComfyUI model.
> [!NOTE]
> Experimental. The wrapper presents a `MODEL` interface but cannot fully
> emulate one — model merges, IPAdapter, and similar advanced uses generally
> won't work, and the model's fixed input shapes are not validated, so mismatched
> inputs error at runtime. The native Core ML Sampler is faster when you don't
> need the `MODEL` type. See the [FAQ](faq.md) and [limitations](limitations.md).
---
## Load LoRA to use with Core ML (`Core ML LoRA Loader`)
![LoRA Loader](../assets/lora_loader.png?raw=true)
Collects LoRA name + `strength_model` to bake into the model at conversion, and
applies the LoRA to CLIP (which is not part of the Core ML path). Chain multiple
loaders for multiple LoRAs.
Because a converted model is immutable, the baked weights and `strength_model`
**cannot** be changed afterward — changing them means re-converting. `strength_clip`
only affects CLIP and can be changed freely. After conversion, when loading with
`CoreMLUNetLoader`, apply the same LoRAs to CLIP manually (see
[workflows](workflows.md)).
- **Inputs**
- `lora_name`, `strength_model`, `strength_clip`.
- `clip` — from `CLIPLoader` / `CheckpointLoaderSimple` or another LoRA loader.
- `lora_params` *(optional)* — chain from another LoRA loader.
- **Outputs**
- `CLIP` — with the LoRA applied.
- `lora_params` — pass to the converter or the next LoRA loader.
> [!NOTE]
> LoRA support is experimental and inconsistent — some LoRAs convert cleanly,
> others produce poor results. Test per-LoRA. See [troubleshooting](troubleshooting.md).
---
## Convert Checkpoint to Core ML (`Core ML Converter`)
![Checkpoint Converter](../assets/checkpoint_converter.png?raw=true)
Converts a checkpoint from `models/checkpoints` to a Core ML `.mlpackage` in
`models/unet`. The model version (SD1.5, SDXL, SDXL refiner, or full-distill
LCM) is auto-detected from the checkpoint's architecture — there is no version
dropdown. The conversion parameters are encoded in the output name, so an
already-converted model is reused instead of re-converted. See
[conversion](conversion.md) for details.
- **Inputs**
- `ckpt_name` — checkpoint in `models/checkpoints`.
- `height`, `width` — target image size; any positive multiple of 8 (default
512). The model's input size is fixed at these values.
- `batch_size` — default 1; raise to convert a batch-capable model.
- `attention_implementation` — `SPLIT_EINSUM` / `SPLIT_EINSUM_V2` (ANE) or
`ORIGINAL` (GPU). See [hardware](hardware.md).
- `compute_unit` — used only when loading the result; does not affect
conversion.
- `controlnet_support` — set `True` to make the model usable with ControlNet
(default `False`).
- `quantize_nbits` *(optional)* — `none` (default), `8`, `6`, `4`. See
[conversion → quantization](conversion.md#quantization).
- `lora_params` *(optional)* — from the LoRA loader, to bake LoRAs in.
- **Output**
- `coreml_model`.
> [!NOTE]
> Some checkpoints need a custom config `.yaml`. Place it in `models/configs`
> named like the checkpoint (e.g. `juggernaut.safetensors` →
> `juggernaut.yaml`); it is loaded automatically during conversion.
> [!NOTE]
> Full-distill LCM checkpoints (e.g.
> [LCM_Dreamshaper_v7](https://huggingface.co/SimianLuo/LCM_Dreamshaper_v7)) are
> detected and converted like any other checkpoint. When sampling an LCM model,
> set `sampler_name` to `lcm` and `scheduler` to `sgm_uniform`.
+86
View File
@@ -0,0 +1,86 @@
# Troubleshooting
## `Expected shape … got …` / latent size mismatch
The most common error. A Core ML model has **fixed** input dimensions — a model
converted for 512×512 expects a 64×64 latent and rejects any other size (batch
size is handled and doesn't matter; only width/height are fixed).
**Fix:** set your Empty Latent (or upstream latent) to exactly the resolution the
model was converted for, or re-convert at the size you want.
## Old `.mlmodelc` model, or `metadata.json` not found
This suite no longer produces or loads `.mlmodelc`; the loader lists `.mlpackage`
only. Models from an older version (or downloaded community models) with a
`.mlmodelc` structure won't load.
**Fix:** re-convert the checkpoint with **Convert Checkpoint to Core ML**. No
Xcode or `coremlcompiler` is required — that dependency was removed.
## Prompt too long (`Expected size 154 but got 77`, or a crash)
Core ML enforces a hard **77-token** prompt limit and does not auto-chunk like
A1111/ComfyUI.
**Fix:** split the prompt across multiple CLIP Text Encode nodes and merge them
with **Conditioning (Combine)**.
## `cannot import name 'ModelSamplingDiscreteLCM'`
A ComfyUI refactor renamed this symbol.
**Fix:** update the suite (fixed in PR #29) and re-run
`pip install -r requirements.txt`.
## LoRA loader `ImportError`
`peft` became a required dependency.
**Fix:** `pip install -r requirements.txt`. This recurs after ComfyUI-Manager
updates if requirements aren't reinstalled.
## ControlNet has no effect
ControlNet support is baked at conversion. If the checkpoint was converted with
`controlnet_support = False`, ControlNet does nothing.
**Fix:** re-convert with `controlnet_support = True`. The ControlNet model itself
needs no conversion, and `.fp16.safetensors` vs `.safetensors` makes no
difference.
## LoRAs produce garbage
LoRA support is inconsistent — some work, some don't, with no firm rule. Test
per-LoRA. For some LCM-LoRA setups, routing through the
[Core ML Adapter](nodes.md#core-ml-adapter-experimental-coremlmodeladapter) is
more reliable than the basic loader path. Remember weights are baked at conversion
and can't be changed afterward.
## FaceDetailer / detailers error on size
Detailers rescale latents internally (e.g. 512 → 1024), which breaks the model's
fixed input shape. There is no workaround node — a Core ML model only accepts
the resolution it was converted for.
**Fix:** convert a second model at the detailer's internal resolution and use it
for the detailing pass, or run the detailer with a standard (non–Core ML) model.
## Inpainting checkpoint errors (`tensor size 9 vs 4`)
SD1.5 inpainting checkpoints use a 9-channel input and are **not supported**. This
error is expected, not a bug.
## Errors mentioning `python_coreml_stable_diffusion` or `ml-stable-diffusion`
You're on a stale install. That dependency was removed; old install scripts tried
`pip install git+…/ml-stable-diffusion.git`, which fails on modern Python.
**Fix:** reinstall the current suite (`pip install -r requirements.txt`, which
pulls `coreml-diffusion` from PyPI).
## `all input tensors must be on the same device (mps:0 and cpu)` / ControlNet residual shape `(2,…) vs (1,…)`
Old bugs that have been fixed.
**Fix:** update to the latest version.
+103
View File
@@ -0,0 +1,103 @@
# Example Workflows
> [!NOTE]
> The models referenced are examples — substitute your own. Every workflow
> starts from a checkpoint you convert yourself (see [conversion](conversion.md));
> there is no Core ML model to download.
## Basic txt2img
Convert a SD1.5 checkpoint, then sample from it. CLIP and VAE come from standard
ComfyUI nodes — either loaded separately or pulled from the checkpoint.
1. Place a SD1.5 checkpoint in `models/checkpoints` (e.g.
[v1-5-pruned-emaonly](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/v1-5-pruned-emaonly.safetensors)).
2. **Convert Checkpoint to Core ML** → queue once → a `.mlpackage` lands in
`models/unet`.
3. **Load Core ML UNet** (or wire the converter output straight in) →
**Core ML Sampler** → **VAE Decode**.
**CLIP and VAE from the checkpoint:**
![Core ML UNet + checkpoint](../assets/unet+sampler+checkpoint.png?raw=true)
**CLIP and VAE loaded separately** — use any SD1.5-compatible
[CLIP](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/text_encoder/model.safetensors)
and [VAE](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/vae/diffusion_pytorch_model.safetensors),
placed in `models/clip` and `models/vae`:
![Core ML UNet + CLIP + VAE](../assets/unet+sampler+clip+vae.png?raw=true)
## ControlNet
Convert the checkpoint with `controlnet_support = True`, then wire a standard
ComfyUI ControlNet. The ControlNet model itself needs no conversion. Place it in
`models/controlnet` (e.g.
[control_v11p_sd15_scribble](https://huggingface.co/lllyasviel/control_v11p_sd15_scribble/blob/main/diffusion_pytorch_model.fp16.safetensors)).
![Core ML UNet + ControlNet](../assets/unet+sampler+controlnet.png?raw=true)
## Checkpoint conversion
The minimal conversion graph. See
[Convert Checkpoint to Core ML](nodes.md#convert-checkpoint-to-core-ml-core-ml-converter).
![Checkpoint converter](../assets/basic_conversion.png?raw=true)
## Conversion with LoRA
Bake LoRA(s) into the model at conversion. Read the
[LoRA caveats](nodes.md#load-lora-to-use-with-core-ml-core-ml-lora-loader) first
— baked weights are immutable, and support is inconsistent per-LoRA.
![Checkpoint converter + LoRA](../assets/conversion+lora.png?raw=true)
## LCM LoRA conversion
Chain multiple LoRA loaders to use several LoRAs with one model.
> [!IMPORTANT]
> Here the model goes through the **Core ML Adapter** and `ModelSamplingDiscrete`
> into the standard ComfyUI KSampler (not the Core ML Sampler).
> `ModelSamplingDiscrete` is required to sample LCM LoRAs correctly.
![Multiple LoRAs](../assets/conversion+lcm_lora.png?raw=true)
## Loading a model with baked LoRAs
Load a model that already has LoRAs baked in. CLIP must be loaded separately and
passed through the same LoRA nodes used at conversion. Since `lora_name` and
`strength_model` are baked in, they need not be passed to the loader.
> [!IMPORTANT]
> As above, the model goes through the Core ML Adapter + `ModelSamplingDiscrete`
> into the standard KSampler.
![Loader + LoRA](../assets/loader+lcm_lora.png?raw=true)
## LCM conversion with ControlNet
Convert a full-distill LCM checkpoint (e.g.
[LCM_Dreamshaper_v7](https://huggingface.co/SimianLuo/LCM_Dreamshaper_v7)) with
the standard **Convert Checkpoint to Core ML** node — the LCM architecture is
auto-detected. Use it with or without ControlNet. When sampling, set
`sampler_name` to `lcm` and `scheduler` to `sgm_uniform`.
![LCM + ControlNet](../assets/lcm+controlnet.png?raw=true)
## SDXL Base + Refiner
A basic SDXL graph. Add LoRAs and ControlNets as in the SD1.5 examples; the
refiner step is optional.
Models:
[base + text encoders](https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0),
[refiner](https://huggingface.co/stabilityai/stable-diffusion-xl-refiner-1.0),
[VAE](https://huggingface.co/stabilityai/sdxl-vae).
> [!IMPORTANT]
> SDXL does not run on the ANE. Convert with `ORIGINAL` and load with
> `CPU_AND_GPU` (or `CPU_ONLY`). If loading hangs on `CPU_AND_NE`, that is the
> cause. See [limitations](limitations.md).
![SDXL](../assets/sdxl_conversion.png?raw=true)
-208
View File
@@ -1,208 +0,0 @@
# Conversion Extraction — Seam Inventory (`docs/extraction/seam.md`)
> **Gate E0 deliverable.** Symbol-by-symbol cut line between the future `coreml_diffusion`
> package (CONVERSION) and what stays in `coreml_suite` (the ComfyUI side).
>
> **Confidence legend:**
> - ✅ **verified** — read directly from the current source in this repo.
> - 🔍 **confirm** — inferred / partially seen; Claude Code must `grep`-verify before acting.
>
> **Cut rule:** a symbol goes to `coreml_diffusion` iff it participates in producing the `.mlpackage`
> artifact AND can be made free of `comfy` / `folder_paths` / `comfy_extras`. The runtime
> *loader* that **runs** a compiled model stays in the suite.
---
## 1. File-level map
| File | Side | Status | Note |
|---|---|---|---|
| `coreml_suite/model_version.py` | **coreml_diffusion** | ✅ | Already `Enum`-only, zero comfy. Becomes pkg source of truth. |
| `coreml_suite/attention.py` | **coreml_diffusion** | ✅ | `ATTENTION_IMPLEMENTATIONS` tuple; pure constant. |
| `coreml_suite/core/naming.py` | **coreml_diffusion** | ✅ | `compose_out_name` = cache-key contract. Move (not copy). |
| `coreml_suite/converter.py` | **coreml_diffusion** (mostly) | ✅ | Main conversion. One symbol stays-adjacent: `get_out_path` (folder_paths) is replaced by injected `out_path`. |
| `coreml_suite/conversion/attention.py` | **coreml_diffusion** | ✅ | `apply_attention_implementation`. Imports `logging`,`torch` only — no comfy. |
| `coreml_suite/conversion/shapes.py` | **coreml_diffusion** | ✅ | `conv2d_output_shape`. Pure math, no imports. |
| `coreml_suite/conversion/trace.py` | **coreml_diffusion** | ✅ | Imports `types.MethodType`, `diffusers...Transformer2DModel` only — torch/diffusers. |
| `coreml_suite/conversion/unet.py` | **coreml_diffusion** | ✅ | `CoreMLUNetWrapper`. Imports `torch` only — no comfy. |
| `coreml_suite/lcm/converter.py` | **coreml_diffusion** (after dedup) | ✅ | Dup helpers deleted; `MODEL_VERSION` HF-hardcode (L22) → E-LCM. `folder_paths` (L111) + `comfy.model_management` (L54) confirmed present → CUT. |
| `coreml_suite/lcm/unet.py` | **coreml_diffusion** | ✅ | `UNet2DConditionModelLCM(UNet2DConditionModel)`. diffusers-only, no comfy. |
| `coreml_suite/config.py` | **STAYS** | ✅ | Imports `comfy.supported_models_base`/`latent_formats`/`model_detection`. **Inference-side** (`get_model_config`), NOT conversion. |
| `coreml_suite/coreml_model.py` | **STAYS** | ✅ | `CoreMLModel` = runtime loader (runs `.mlpackage`). Desktop/Python inference; not used on iOS. |
| `coreml_suite/nodes.py` | **STAYS** | ✅ | Nodes; will call `coreml_diffusion` + own `folder_paths` path resolution + discovery dropdowns. |
| `coreml_suite/lcm/nodes.py` | **STAYS** | ✅ | `COREML_CONVERT_LCM` node. |
| `coreml_suite/models.py` | **STAYS** | ✅ | Inference: `add_sdxl_model_options`, `is_sdxl`, `get_model_patcher`, `get_latent_image`. |
| `coreml_suite/latents.py` | **STAYS** | ✅ | Inference chunking (MODERNIZATION Phase 3 target, not this spec). |
| `coreml_suite/controlnet.py` | **STAYS** | ✅ | Inference-side controlnet. Distinct from converter `add_cnet_support`. |
| `coreml_suite/lcm/utils.py` | **STAYS** | ✅ | `add_lcm_model_options`, `lcm_patch`, `is_lcm`; imports `comfy_extras`. Inference. |
| `coreml_suite/logger.py` | **both / copy** | ✅ | Trivial. Package gets its own logger; suite keeps its. |
---
## 2. Symbol-level: `coreml_suite/converter.py` (main conversion)
| Symbol | Side | Status | Cut action |
|---|---|---|---|
| `DEFAULT_TRACE_TIMESTEP`, `TEXT_TOKEN_SEQUENCE_LENGTH` | coreml_diffusion | ✅ | Move as-is (module constants). |
| `get_unet(model_version, ref_unet, attention_implementation)` | coreml_diffusion | ✅ | Move. Uses `conversion.{trace,attention,unet}`. No comfy. |
| `get_encoder_hidden_states_shape(ref_unet, batch_size)` | coreml_diffusion | ✅ | Move. Reads `ref_unet.config.cross_attention_dim`. Pure. |
| `get_coreml_inputs(sample_inputs)` | coreml_diffusion | ✅ | Move. `ct.TensorType` build. |
| `load_coreml_model(out_path)` | coreml_diffusion | ✅ | Move. `ct.models.MLModel(out_path)`. (Dedup target vs LCM copy.) |
| `convert_to_coreml(submodule, ts_module, inputs, names, out_path)` | coreml_diffusion | ✅ | Move. `ct.convert(...)`. (Dedup target vs LCM copy.) |
| `get_sample_input(batch, ehs_shape, sample_shape)` | coreml_diffusion | ✅ | Move. **Merge** with LCM variant (LCM passes extra `scheduler` → optional param). |
| `lcm_inputs(sample_unet_inputs)` | coreml_diffusion | ✅ | Move. Adds `timestep_cond`. |
| `sdxl_inputs(sample_unet_inputs, ref_unet, model_version)` | coreml_diffusion | ✅ | Move. `time_ids`/`text_embeds`/`add_embeds`. |
| `add_cnet_support(sample_shape, ref_unet)` | coreml_diffusion | ✅ | Move. Builds `additional_residual_*` inputs from unet block channels. |
| `convert_unet(ref_unet, model_version, unet_out_path, ...)` | coreml_diffusion | ✅ | Move. Orchestrates trace→convert→**quant (palettize)**→save. Quant travels here (E6). |
| `convert(ckpt_path, model_version, unet_out_path, ...)` | coreml_diffusion | ✅ | Move. **Make kw-only past `ckpt_path,model_version,out_path`** (contract). Validates `attn_impl`. |
| `load_unet(ckpt_path, config_path)` | coreml_diffusion | ✅ | Move. `UNet2DConditionModel.from_single_file`. |
| `get_out_path(submodule_name, model_name)` | **STAYS (node)** | ✅ | Uses `folder_paths.get_folder_paths`. **Delete from converter; node resolves path and passes `out_path` in.** |
**Apple `python_coreml_stable_diffusion` footprint on this path:** ✅ **none.** Verified by grep:
zero imports in `converter.py` / `conversion/*`. Main path uses `diffusers` +
local `CoreMLUNetWrapper`. (And the runtime `CoreMLModel` is now a local coremltools wrapper too —
see §6 stale-spec note.)
---
## 3. Symbol-level: `coreml_suite/lcm/converter.py` (LCM — dedup + defer)
| Symbol | Side | Status | Cut action |
|---|---|---|---|
| `load_coreml_model` (LCM copy) | DELETE | ✅ | Duplicate of main. Remove; use `coreml_diffusion.load_coreml_model`. |
| `convert_to_coreml` (LCM copy) | DELETE | ✅ | Duplicate of main. Remove. |
| `get_out_path` (LCM copy, folder_paths) | DELETE | ✅ | Duplicate + comfy. Remove; node injects `out_path`. |
| `get_sample_input(..., scheduler)` (LCM copy) | MERGE → coreml_diffusion | ✅ | Fold `scheduler` into shared `get_sample_input` as optional param. |
| `MODEL_NAME` (= LCM_Dreamshaper) | **E-LCM** | ✅ | HF hardcode. Removing it is the behavior change → E-LCM, not E2. |
| `convert(out_path, sample_size, batch_size, controlnet_support)` (LCM, L190) | coreml_diffusion (via unified) | ✅ | Route through `coreml_diffusion.convert(model_version=LCM, ...)` in E-LCM. |
| `from comfy.model_management import get_torch_device` (L54, in `get_scheduler`) | **CUT** | ✅ | Confirmed present. Inject `device`. |
| module-global attention set at import | n/a | ✅ | **No module global.** Attention already per-call: `get_unets` (L36) calls `apply_attention_implementation(ref_unet, "SPLIT_EINSUM")`. No `ATTENTION_IMPLEMENTATION_IN_EFFECT` anywhere in repo. (Note: LCM hardcodes `"SPLIT_EINSUM"` — pass `attn_impl` through in dedup.) |
---
## 4. Symbol-level: `coreml_suite/core/naming.py` → `coreml_diffusion/naming.py`
| Symbol | Side | Status | Cut action |
|---|---|---|---|
| `compose_out_name(...)` | coreml_diffusion | ✅ | **Move** (cache-key contract). Node imports from pkg. |
| `lora_names_from_params(...)` | coreml_diffusion | ✅ | Move. |
| `ATTN_SUFFIX` dict | coreml_diffusion | ✅ | Move. |
| `QUANT_NBITS_VALUES` | coreml_diffusion | ✅ | Move; backs `list_quant_modes()`. |
| `tests/unit/test_characterization_out_name.py` | re-point | ✅ | Change import to `coreml_diffusion.naming`. Assertions/values **unchanged**. |
---
## 5. Discovery API + status registry (new in `coreml_diffusion/__init__.py`)
```python
from enum import Enum
class Status(Enum):
VERIFIED = "verified" # has a golden anchor + passing [M2-ANE] check
EXPERIMENTAL = "experimental" # convertible, not yet anchored/verified
# Single source of truth. Suite gates on this, NOT on a hardcoded node list.
# KEY by ModelVersion enum MEMBER (not a bare string) so list_model_versions can
# emit .name — see the .name decision below. Keying by the lowercase .value string
# (as an earlier draft of this block did) returns ["sd15",...], which the node then
# reverses via ModelVersion[...] → KeyError. Do NOT key by .value.
_MODEL_STATUS = {
ModelVersion.SD15: Status.VERIFIED,
ModelVersion.SDXL: Status.VERIFIED,
ModelVersion.SDXL_REFINER: Status.EXPERIMENTAL, # → VERIFIED after a refiner golden anchor
ModelVersion.LCM: Status.EXPERIMENTAL, # → VERIFIED after E-LCM golden anchor
}
def list_model_versions(include_experimental: bool = False) -> list[str]:
return [v.name for v, s in _MODEL_STATUS.items() # .name → "SD15","SDXL" (see decision)
if s is Status.VERIFIED or (include_experimental and s is Status.EXPERIMENTAL)]
def list_attention_impls() -> list[str]: # from attention.ATTENTION_IMPLEMENTATIONS
...
def list_quant_modes() -> list[str]: # from naming.QUANT_NBITS_VALUES
...
CONTRACT_VERSION = "1.0"
# Additive-only: adding an id or promoting EXPERIMENTAL→VERIFIED = minor bump (Suite unaffected).
# Removing/renaming an id, or demoting VERIFIED→EXPERIMENTAL = MAJOR bump + migration note.
```
**Decision check (`.name` vs `.value`): RESOLVED → `.name`.** ✅
Verified in current source:
- Node renders `ModelVersion.SD15.name` / `ModelVersion.SDXL.name` → `"SD15"`, `"SDXL"`
(`nodes.py:224-225`).
- Node reverses the dropdown string with `model_version = ModelVersion[model_version]`
(`nodes.py:286`) — i.e. **lookup by NAME**. Feeding it a `.value` (`"sd15"`) raises `KeyError`.
- Enum values are lowercase (`model_version.py`: `SD15="sd15"`, `SDXL="sdxl"`,
`SDXL_REFINER="sdxl_refiner"`, `LCM="lcm"`).
- `compose_out_name` does NOT consume the model_version string (grep of `core/naming.py` empty) —
no coupling there, so no constraint from that side.
**Decision:** `list_model_versions()` returns `.name` (uppercase). Saved workflows store `"SD15"`,
node already validates them via `ModelVersion[...]`. The `_MODEL_STATUS` block above was corrected
to key by enum member and emit `.name`. **The earlier `v.value` form was a latent bug.**
---
## 6. `python_coreml_stable_diffusion` split (Gate E0 line to fill by grep)
| Use | Side | Status |
|---|---|---|
| `coreml_model.CoreMLModel` (runs compiled model) | **STAYS** (suite runtime) | ✅ — **local class**, not Apple's |
| `unet.UNet2DConditionModel*` internals | **gone** — `converter.py:319` uses `diffusers.UNet2DConditionModel.from_single_file` | ✅ |
| `AttentionImplementations` enum | gone — local `apply_attention_implementation` + `attention.py` tuple | ✅ |
| `calculate_conv2d_output_shape` | gone — replaced by `conversion/shapes.conv2d_output_shape` | ✅ |
> ### ⚠️ SPEC IS STALE: `ml-stable-diffusion` is already fully removed
> Commit #58 ("replace apple/ml-stable-diffusion with native diffusers conversion") already did
> the de-Apple work. Verified now:
> - **Zero** `python_coreml_stable_diffusion` runtime imports anywhere in `coreml_suite` (only a
> docstring mention at `core/__init__.py:4`).
> - `coreml_suite/coreml_model.py:8` `CoreMLModel` is a **local** wrapper over
> `coremltools.models.MLModel` (`coreml_model.py:22`) — it does **not** import Apple's class.
> - `ml-stable-diffusion` / `python_coreml_stable_diffusion` appears in **neither** `pyproject.toml`
> **nor** `requirements.txt`. It is not a dependency at all.
>
> **Consequences for the spec (correct these in CONVERTER_EXTRACTION_SPEC.md):**
> - §0.3 premise ("runtime loader = `python_coreml_stable_diffusion.coreml_model.CoreMLModel`,
> stays in suite") is **wrong**: the loader is already the local `coreml_model.CoreMLModel`. The
> "stays in suite" conclusion still holds; the identity does not.
> - **Gate E0 item "ml-stable-diffusion pinned SHA — BLOCKER if unpinned" is MOOT** — there is no
> such dep to pin. Mark it N/A, not BLOCKER.
> - **E4/E5 dependency lists must drop `git+...ml-stable-diffusion@<sha>`.** Package runtime deps
> are: `coremltools`, `diffusers`, `peft` (LoRA), `omegaconf` (config), `numpy`, `torch`. Confirm
> `peft`/`omegaconf` actually used before listing (grep at E4).
> - The "keep `python_coreml_stable_diffusion` as a suite dep for the loader" instruction in E5 is
> **void** — coremltools backs the loader.
---
## 7. Pre-flight checklist before E1 (run these greps)
```
grep -rn "import comfy" coreml_suite/conversion coreml_suite/converter.py coreml_suite/lcm/converter.py coreml_suite/lcm/unet.py
grep -rn "folder_paths" coreml_suite/converter.py coreml_suite/lcm/converter.py
grep -rn "model_management" coreml_suite/lcm
grep -rn "python_coreml_stable_diffusion" coreml_suite
grep -rn "ATTENTION_IMPLEMENTATION_IN_EFFECT" coreml_suite
grep -rn "SimianLuo\|LCM_Dreamshaper" coreml_suite/lcm
```
Every 🔍 above resolves to ✅ or a correction once these run. Do not start moving code (E2)
with any 🔍 unresolved on the CONVERSION side.
**STATUS (run 2026-05-26): all 🔍 resolved.** Summary of what the greps found:
- `conversion/*`, `lcm/unet.py`: comfy-free (torch/diffusers only). ✅
- `converter.py`: only comfy reach-in is `folder_paths` in `get_out_path` (L91-94) → inject `out_path`.
- `lcm/converter.py`: `folder_paths` (L111-114) + `comfy.model_management.get_torch_device` (L54)
→ cut both. Dup helpers (`load_coreml_model`,`convert_to_coreml`,`get_out_path`,`get_sample_input`)
confirmed → dedup E2. `MODEL_VERSION="SimianLuo/LCM_Dreamshaper_v7"` (L22) → E-LCM.
- No attention module-global anywhere (`ATTENTION_IMPLEMENTATION_IN_EFFECT` absent); already per-call.
LCM hardcodes `"SPLIT_EINSUM"` in `get_unets` — thread `attn_impl` through during dedup.
- `.name` vs `.value`: **decided `.name`** (node reverses via `ModelVersion[...]`). §5 corrected.
- `ml-stable-diffusion`: **already gone** (#58). §6 stale-spec note added — fix the spec's E0/E4/E5
dep + pinning items.
Two grep blind-spots to note (the checklist above doesn't cover them, but cheap to add): the
`folder_paths` grep only scans the two converter files — also grep `coreml_suite/lcm/utils.py`
(it imports `comfy.model_management` at L3, but it's inference/STAYS, so fine) and confirm no other
`conversion/` file grew a comfy import since.