From 8f94f0eea5c680e0392165cd49c117d402d24c91 Mon Sep 17 00:00:00 2001 From: aszc-dev Date: Thu, 4 Jun 2026 17:29:41 +0200 Subject: [PATCH] docs: rewrite README and split into docs/ pages MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Rewrite the README as a lean landing page and move depth into a docs/ folder. Correct the supported-model story and several stale facts, and answer the recurring questions from issue #21. - Convert-only is the supported path: suite-converted .mlpackage is the only supported input; drop coreml-community download guidance. - Remove all .mlmodelc / Xcode references — compilation was dropped and the loader handles .mlpackage only. - Fix compute-unit name (CPU_AND_NE, not CPU_AND_ANE) and the loader input name (coreml_name). - Document CoreMLSamplerAdvanced (previously undocumented). - Add docs/: hardware, nodes, conversion, workflows, faq, troubleshooting, limitations (with a support matrix). - Note conversion now lives in the coreml-diffusion package. - Remove dev scaffolding specs; ignore *.log, .DS_Store, .claude/. --- .gitignore | 2 + CONVERTER_EXTRACTION_SPEC.md | 618 ----------------------------------- README.md | 525 ++++++----------------------- docs/conversion.md | 95 ++++++ docs/faq.md | 77 +++++ docs/hardware.md | 95 ++++++ docs/limitations.md | 53 +++ docs/nodes.md | 173 ++++++++++ docs/troubleshooting.md | 86 +++++ docs/workflows.md | 100 ++++++ seam.md | 208 ------------ 11 files changed, 774 insertions(+), 1258 deletions(-) delete mode 100644 CONVERTER_EXTRACTION_SPEC.md create mode 100644 docs/conversion.md create mode 100644 docs/faq.md create mode 100644 docs/hardware.md create mode 100644 docs/limitations.md create mode 100644 docs/nodes.md create mode 100644 docs/troubleshooting.md create mode 100644 docs/workflows.md delete mode 100644 seam.md diff --git a/.gitignore b/.gitignore index d60dca5..44a73e5 100644 --- a/.gitignore +++ b/.gitignore @@ -3,4 +3,6 @@ __pycache__/ models/ .venv/ test_results/ +*.log +.DS_Store .claude/ diff --git a/CONVERTER_EXTRACTION_SPEC.md b/CONVERTER_EXTRACTION_SPEC.md deleted file mode 100644 index 26bf3ff..0000000 --- a/CONVERTER_EXTRACTION_SPEC.md +++ /dev/null @@ -1,618 +0,0 @@ -# ComfyUI-CoreMLSuite — Converter Extraction Spec for Claude Code - -> **Companion to `MODERNIZATION_SPEC.md`.** That spec hardens the repo and (Phase 3) -> splits the *inference* math from the framework. **This** spec splits the *conversion* -> path (`safetensors → CoreML`) out into a standalone, `comfy`-free, pip-installable -> package that CoreMLSuite then depends on — and that other projects (incl. on-device -> iOS tooling) can reuse. -> -> **Same discipline as the modernization spec:** safety-net first, behavior-preserving -> until told otherwise, one phase = one branch = one PR, `STOP — VALIDATE` gate between -> every phase, golden-latent as the regression anchor. `[M2]` = needs macOS/Apple Silicon; -> `[M2-ANE]` = needs the Neural Engine. Everything else must run on plain Linux/CI. - ---- - -## 0. How to work (read first — non-negotiable) - -1. **Behavior-preserving until Phase E6.** Phases E1–E5 must not change image output, node - names, `INPUT_TYPES` field names, or `NODE_CLASS_MAPPINGS` keys. The node graph is the - public contract; saved user-workflow JSON breaks if these change. -2. **The conversion package produces an artifact and stops there.** Its job ends at a written - `.mlpackage` / `.mlmodelc` on disk. It must NOT import `comfy`, `folder_paths`, or - `comfy_extras`, and must NOT know ComfyUI's `models/unet` layout. Paths are *inputs*. -3. **The runtime loader stays in the suite.** The loader is the **local** `coreml_suite.coreml_model.CoreMLModel` - — a thin wrapper over `coremltools.models.MLModel` (NOT Apple's - `python_coreml_stable_diffusion.coreml_model.CoreMLModel`, which is no longer used; see #58). - It *runs* a compiled model in Python — a desktop/Python inference concern, not a conversion - concern. It is NOT moved into the package. (On iOS the `.mlmodelc` is loaded natively; the - package's output is the deliverable, not a Python runner.) -4. **Decouple in-repo before splitting repos.** Phases E1–E4 create the package *inside this - repo* and prove equivalence. The physical second-repo split is Phase E5, only after the - golden latent is proven identical. Do not create a second repository before Gate E4 passes. -5. **Reuse the existing regression anchor.** The golden latent / PSNR anchor from - `MODERNIZATION_SPEC.md` Phase 2 is the cross-cutting proof for every gate here. If it is not - yet captured, capture it first (it is a prerequisite for E2 onward). -6. **No new runtime dependencies** without flagging in the gate report (name, why, license, size). -7. **A failing gate means stop and report**, not work around into the next phase. -8. **Tooling is `uv`, not bare `pip`/`venv`.** Every environment/install/lock step uses the - project's `uv` toolchain: `uv venv`, `uv pip install`, `uv pip install -e .`, `uv lock`, - `uv run pytest`, `uv export`/`uv pip freeze` for baselines. Where this spec says "fresh venv", - read "`uv venv` + `uv pip install`". Reserve `uv pip` (not `pip`) inside that venv too. -9. **The package is the single source of truth for *what is possible*; the node is a thin, - discovery-driven frontend.** See the "Interface contract" pillar below — this is the - maintainer's hard requirement and it overrides the earlier (now-rescinded) "freeze the - dropdown list" instruction. - ---- - -## Interface contract (the maintainer's hard requirement) — read before any phase - -Two coupled guarantees must hold once the package is split out: - -**(A) Updating the converter must NOT require updating CoreMLSuite.** -This is satisfied by treating the package's public surface as a versioned contract: -- `convert(...)` and `compile_model(...)` are **keyword-only with defaults** for everything - past the genuinely-required positionals (`ckpt_path`, `model_version`, `out_path`). New - capabilities are added as new keyword args with defaults, so an old Suite's call still - validates against a newer package. **Never** reorder or rename existing parameters. -- `compose_out_name` (the `.mlpackage` filename = the cache key) **moves into the package** and - is versioned with it. The Suite must not carry its own copy; if the package changes the naming - scheme that is a **major** bump (old cached artifacts stop resolving). - -**(B) CoreMLSuite must be able to list *new* conversion types WITHOUT a Suite code change or -version bump.** Today the node hardcodes its dropdowns: -```python -"model_version": ([ModelVersion.SD15.name, ModelVersion.SDXL.name],), # hand-typed, also INCOMPLETE (no LCM / SDXL_REFINER) -"attention_implementation": (list(ATTENTION_IMPLEMENTATIONS),), # from coreml_suite.attention -"quantize_nbits": (list(QUANT_NBITS_VALUES), {"default": "none"}), # from coreml_suite.core.naming -``` -These are replaced by **runtime discovery calls into the package**, evaluated inside -`INPUT_TYPES` (ComfyUI re-evaluates `INPUT_TYPES` on every plugin load): -```python -import coreml_diffusion -"model_version": (coreml_diffusion.list_model_versions(),), -"attention_implementation": (coreml_diffusion.list_attention_impls(),), -"quantize_nbits": (coreml_diffusion.list_quant_modes(), {"default": "none"}), -``` -Effect: `uv pip install -U coreml_diffusion` + ComfyUI restart surfaces any newly-added type in the old -plugin's dropdown — **no Suite edit, no Suite version bump.** This is the requirement. - -**The cost, stated honestly (accept this trade-off explicitly at Gate E0):** -- The Suite becomes a "dumb" frontend; the package is the sole authority on what conversions - exist. The Suite can no longer guarantee its saved workflows are valid against *arbitrary* - future package versions. -- Therefore the package's discovery identifiers (`ModelVersion` values, attn-impl strings, quant - modes) are an **ADDITIVE-ONLY contract**: the package may *add* identifiers freely (minor bump, - no Suite change); **removing or renaming an identifier is a breaking change requiring a MAJOR - bump and a migration note**, because a saved workflow JSON references these strings verbatim. - Without this rule, "no version bump" silently becomes "randomly broken workflows." -- `INPUT_TYPES` must **fail soft** when the package is missing/old: wrap the discovery calls so a - missing `coreml_diffusion` (or an old one lacking a `list_*` function) yields a sane fallback list and a - logged warning, instead of the node failing to register and disappearing from the menu. - -**Discovery API the package must expose (stable names):** -```python -coreml_diffusion.list_model_versions() -> list[str] # VERIFIED ones only, e.g. ["SD15","SDXL"] today (.name — see seam.md) -coreml_diffusion.list_attention_impls() -> list[str] # ["SPLIT_EINSUM","SPLIT_EINSUM_V2","ORIGINAL"] -coreml_diffusion.list_quant_modes() -> list[str] # ["none","8","6","4"] -coreml_diffusion.CONTRACT_VERSION: str # bump rules above; Suite may log/compare it -``` -These return the *display strings already used today*, so existing workflows keep validating. - -**Verification status is a PACKAGE property, not a node hardcode (maintainer's intent).** -The Suite wants to expose *every model the converter can verifiably convert*. Today `lcm` and -`sdxl_refiner` are absent from the converter node not because the Suite chooses to hide them, but -because they lack a full golden/PSNR verification. So the gating lives in the package as a status: -```python -from enum import Enum -class Status(Enum): - VERIFIED = "verified" # has a golden anchor + passing [M2-ANE] check - EXPERIMENTAL = "experimental" # convertible but not yet anchored/verified - -# internal registry, single source of truth. -# KEY by ModelVersion enum MEMBER so list_* can emit .name. Keying by the lowercase -# .value string returns ["sd15",...], which the node reverses via ModelVersion[...] -> KeyError. -_MODEL_STATUS = {ModelVersion.SD15: Status.VERIFIED, ModelVersion.SDXL: Status.VERIFIED, - ModelVersion.SDXL_REFINER: Status.EXPERIMENTAL, ModelVersion.LCM: Status.EXPERIMENTAL} - -def list_model_versions(include_experimental: bool = False) -> list[str]: - return [v.name for v, s in _MODEL_STATUS.items() # .name -> "SD15","SDXL"; node reverses with ModelVersion[...] - if s is Status.VERIFIED or (include_experimental and s is Status.EXPERIMENTAL)] -``` -Consequence: **promoting a model to VERIFIED in the package expands the Suite's dropdown with no -Suite change and no Suite bump** — exactly the requirement. The act of verification (E-LCM -produces an LCM golden anchor; same later for refiner) is what flips the status. The Suite's -converter node calls `list_model_versions()` (verified-only); a power-user/CLI path may pass -`include_experimental=True`. Promotion VERIFIED-from-EXPERIMENTAL is additive (minor bump); -demotion or removal is breaking (major bump + note). - ---- - -## Naming & layout (chosen — frozen at Gate E0) - -**Distribution name (PyPI):** `coreml-diffusion`. **Import name (Python):** `coreml_diffusion`. -(PyPI normalizes `-`/`_`; the distribution uses the hyphen, the importable module the underscore.) -Availability checked: both `coreml-diffusion` and the near variants were free on PyPI at E0. - -**Why this name (the positioning it encodes):** the project's niche is *diffusion models on Apple -Neural Engine via CoreML, inside ComfyUI and on-device* — **not** Stable Diffusion specifically. -`sd*` was rejected because it falsely narrows scope to SD; `coreml-diffusion` keeps `coreml` on the -front for discoverability while `diffusion` honestly states the scope (SD/SDXL/LCM today, Flux and -other diffusion architectures later) **without** promising arbitrary non-diffusion torch models, -whose tracing/shape/sample-input pipeline differs. The name must not be re-narrowed to SD in -future docs. ANE is the *differentiator* (documented in the README), but `coreml` was chosen over -`ane` in the name for search discoverability per maintainer decision. - -Target package layout (framework-free — zero `comfy` imports): - -``` -coreml_diffusion/ - __init__.py # public API surface (see "Public API" below) - model_version.py # ModelVersion enum — the SINGLE source of truth, no comfy - attention.py # ATTENTION_IMPLEMENTATIONS tuple (from coreml_suite/attention.py) + apply_attention_implementation - pipeline.py # get_pipeline (from_single_file), get_unet (cml UNet from ref unet) - unet.py # UNet2DConditionModelLCM (moved from coreml_suite/lcm/unet.py) - inputs.py # get_sample_input, lcm_inputs, sdxl_inputs, - # get_encoder_hidden_states_shape, get_coreml_inputs, get_inputs_spec - controlnet.py # add_cnet_support (conversion-side residual SHAPE calc only) - convert.py # convert_unet, convert (orchestration), convert_to_coreml, load_coreml_model - compile.py # compile_coreml_model - quantize.py # (Phase E6 / MODERNIZATION Phase 6 lands here) palettization 4/6/8-bit - cli.py # console entry point: `coreml-diffusion convert ...` - pyproject.toml # standalone packaging (at E5) -``` - -What stays in `coreml_suite/` (the ComfyUI side, thinned): -- `nodes.py` — still owns **name-encoding** (`out_name` construction), path resolution via - `folder_paths`, the node `INPUT_TYPES`/mappings, and wrapping the result in `CoreMLModel`. -- `models.py`, `latents.py`, `controlnet.py` (inference parts), `lcm/utils.py`, `config.py` - (inference config build) — untouched by this spec except the import-source of `ModelVersion`. - -### Public API (the contract `coreml_diffusion` exposes) -```python -from coreml_diffusion import ModelVersion, convert, compile_model, compose_out_name -from coreml_diffusion import list_model_versions, list_attention_impls, list_quant_modes, CONTRACT_VERSION - -# Mirror the CURRENT converter.py signature, made keyword-only past the required positionals -# and with paths/device injected (no folder_paths, no comfy.model_management): -# convert(ckpt_path, model_version, out_path, *, -# batch_size=1, sample_size=(64, 64), controlnet_support=False, -# lora_weights=None, attn_impl=list_attention_impls()[0], config_path=None, -# quantize_nbits="none", device=None) -> None # side effect: writes out_path -# (current convert() returns None and writes via convert_unet → coreml_unet.save; keep that, -# or change to `return out_path` as a deliberate, documented improvement — pick one at E0.) -# compile_model(src_path, out_dir, final_name) -> str # returns compiled .mlmodelc path -``` -Note: `convert` takes an **explicit `out_path`** — no `folder_paths`. `device` is injected -(defaults to torch's default device). `compose_out_name` lives here (cache-key contract) and the -node imports it from the package. The `list_*` discovery functions back the node's dropdowns. - ---- - -## The import chains to cut (root cause inventory) — REVISED against current code - -> **State note (verified):** the code moved on since the original draft. Several chains are -> already cut. Re-verify each line by `grep` before acting; do not assume the original draft. - -**Already done (verify, then skip):** -- ✅ `converter.py` already imports `from coreml_suite.model_version import ModelVersion`, and - `model_version.py` is **clean** (`from enum import Enum` only — zero comfy). The old - "converter → config → comfy" chain is **already broken**. `config.py` still imports comfy, but - it is **inference-side** (`get_model_config` via `supported_models_base`/`latent_formats`) — - *not* on the conversion path. Do **not** treat `config.py` as a converter dependency. -- ✅ `converter.py` now uses `diffusers.UNet2DConditionModel.from_single_file` and a local - `CoreMLUNetWrapper` (in `coreml_suite/conversion/unet.py`) — it is **no longer** importing the - Apple `python_coreml_stable_diffusion.unet.UNet2DConditionModel*` internals on the main path. - A `coreml_suite/conversion/` subpackage already exists (`attention`, `shapes`, `trace`, `unet`). -- ✅ Name-encoding already extracted to `coreml_suite/core/naming.py` (`compose_out_name`, - `lora_names_from_params`, `ATTN_SUFFIX`, `QUANT_NBITS_VALUES`) **with characterization tests** - (`tests/unit/test_characterization_out_name.py`). The pure-naming split is done. -- ✅ Quantization is **already implemented** in `converter.py` (`quantize_nbits`, k-means - `palettize_weights`) and surfaced as an optional node input. Phase E6 is therefore *move*, not - *build* (see revised E6). - -**Still to cut (the real remaining work):** -1. `coreml_suite/converter.py::get_out_path` → `from folder_paths import get_folder_paths`. - Main converter still reaches into ComfyUI's model dir. **Cut: `out_path` is an injected arg; - `folder_paths` resolution moves up into the node** (the node already computes `out_name`). -2. `coreml_suite/lcm/converter.py` → still has its **own** `from folder_paths import - get_folder_paths` (`get_out_path`) and (per original draft) `comfy.model_management`. Verify - the current LCM file and cut both: inject `out_path` and `device`. -3. Global mutation of the attention impl: confirm where it now lives. Main path appears to route - through `coreml_suite/conversion/attention.apply_attention_implementation` (cleaner than the - old global), but `lcm/converter.py` may still set a module global at import. **Ensure the - package sets attention per-call, never at import time.** -4. **Duplication LCM vs main:** `lcm/converter.py` still carries its own copies of - `convert_to_coreml`, `load_coreml_model`, `get_out_path`, `get_sample_input` (the LCM variant - takes a `scheduler` arg), and hardcodes `SimianLuo/LCM_Dreamshaper_v7`. **Dedupe into the - single `coreml_diffusion` implementation;** the HF-hardcode consolidation is the *behavior-changing* - part → deferred to optional **E-LCM**, not E1–E5. -5. **`compose_out_name` ownership:** currently in `coreml_suite/core/naming.py` and called by the - node. Per the Interface-contract pillar it must **move into the package** (it is the cache-key - contract) and the node must import it from `coreml_diffusion`, not keep a copy. - ---- - -## Phase E0 — Seam decision & inventory (no code change) - -**Objective:** lock the cut line, the interface contract, and naming so later phases don't drift. - -### Tasks -1. Produce `docs/extraction/seam.md`: a table of every symbol in `converter.py`, - `lcm/converter.py`, `lcm/unet.py`, **plus the already-extracted `conversion/` subpackage - (`attention`, `shapes`, `trace`, `unet`) and `core/naming.py`**, classified - **CONVERSION → coreml_diffusion** vs **STAYS (comfy/node)**. Note which are already framework-free. -2. ~~Confirm the current `python_coreml_stable_diffusion` footprint.~~ **DONE (seam.md §6): - footprint is ZERO** — no runtime imports anywhere; only a docstring mention in - `core/__init__.py:4`. Main path uses `diffusers` + local `CoreMLUNetWrapper`; the runtime - `CoreMLModel` (STAYS in suite) is a local coremltools wrapper, not Apple's. No shape/attn helper - comes from Apple (local `conversion/shapes.py`, `conversion/attention.py`). -3. **Decide the interface contract concretely (the maintainer's hard requirement):** - - Discovery functions `list_model_versions / list_attention_impls / list_quant_modes` live in - the package and return today's display strings verbatim. Node `INPUT_TYPES` calls them. - - `ModelVersion` values, attn-impl strings, quant modes are **ADDITIVE-ONLY** across package - versions; removal/rename = MAJOR bump + migration note. Write this into the package's - versioning policy doc now. - - `compose_out_name` moves to the package; node imports it (no copy). Confirm the - characterization tests in `test_characterization_out_name.py` will be re-pointed, not - duplicated. - - **Resolve the `model_version` dropdown question (maintainer decided):** the Suite exposes - *every model the converter can verifiably convert*. `lcm` and `sdxl_refiner` are absent today - only because they lack a golden/PSNR verification — **not** because the node hardcodes a - short list. Encode this as a **status registry in the package** (`VERIFIED` vs - `EXPERIMENTAL`); `list_model_versions()` returns VERIFIED-only by default. The converter node - calls it plainly. Promoting LCM/refiner to VERIFIED (after E-LCM / a refiner anchor) expands - the dropdown with **no Suite change**. Do NOT add permanent per-node filtering — the gate is - verification status, owned by the package. -4. ~~Confirm the `ml-stable-diffusion` git dep is pinned.~~ **N/A — already removed (#58).** Verified: - zero `python_coreml_stable_diffusion` imports in the repo; `CoreMLModel` is now a local - coremltools wrapper; the dep is absent from `pyproject.toml`/`requirements.txt`. No SHA to pin. - -### STOP — VALIDATE (Gate E0) -``` -## Gate E0 report -- seam.md committed: ; symbol counts (move / stay / already-framework-free) -- python_coreml_stable_diffusion usage (verified by grep): conversion= runtime= -- Discovery API signatures frozen: list_model_versions (verified-only) / list_attention_impls / list_quant_modes -- Status registry decided: sd15+sdxl=VERIFIED, lcm+sdxl_refiner=EXPERIMENTAL (gated, not hidden) -- Additive-only contract policy doc written (incl. promotion=minor, demotion/removal=major): -- model_version dropdown: expose all (incl. LCM/REFINER) / filtered per node — DECISION: <...> -- compose_out_name move-not-copy confirmed; tests re-point plan: <...> -- LCM consolidation deferred to optional E-LCM: YES/NO -- ml-stable-diffusion: N/A — already removed (#58), not a dependency (was: pin-or-BLOCKER) -- Package name in-repo: coreml_diffusion (final PyPI name deferred to E5) -``` - ---- - -## Phase E1 — Establish `coreml_diffusion` package + discovery API (mostly verification) - -**Objective:** stand up the package namespace and the discovery surface. Much of the comfy-chain -cut is **already done** — this phase mostly *verifies* that and adds the discovery functions. - -### Tasks -1. **Verify (don't redo):** `coreml_suite/model_version.py` is already clean (`Enum` only). Confirm - `import coreml_suite.model_version` works with **no comfy** (`uv run python -c "..."` in a - comfy-free `uv venv`). If true, E1's original "extract ModelVersion" task is already satisfied. -2. Create the `coreml_diffusion/` package skeleton with `__init__.py` exporting the **discovery API** - backed by the *existing* sources of truth for now (re-export `ModelVersion`, the - `ATTENTION_IMPLEMENTATIONS` tuple, and `QUANT_NBITS_VALUES`) so values are byte-identical: - ```python - def list_model_versions(): return [v.name for v in ModelVersion] # .name -> "SD15" (node reverses via ModelVersion[...]; .value KeyErrors) - def list_attention_impls(): return list(ATTENTION_IMPLEMENTATIONS) - def list_quant_modes(): return list(QUANT_NBITS_VALUES) - CONTRACT_VERSION = "1.0" - ``` - (At this stage `coreml_diffusion` may live inside the repo and import from `coreml_suite.*`; the - physical move of implementation happens in E2. The point of E1 is to freeze the *contract*.) -3. **Decided (`.name`):** the node renders `ModelVersion.SD15.name` (`"SD15"`) and reverses the - dropdown string via `ModelVersion[model_version]` (name lookup, `nodes.py:286`). Discovery API - therefore returns `.name`; `.value` (`"sd15"`) would `KeyError`. Recorded in `seam.md` §5. - -### Acceptance criteria -- `uv run python -c "import coreml_diffusion; print(coreml_diffusion.list_model_versions(), coreml_diffusion.list_quant_modes())"` - works in a **comfy-free** `uv venv` and prints today's exact strings. -- Existing characterization tests pass unchanged. -- No node behavior change yet (node still uses its current hardcoded lists in E1). - -### STOP — VALIDATE (Gate E1) -``` -## Gate E1 report -- model_version.py confirmed comfy-free (uv, no comfy): PASS/FAIL -- coreml_diffusion.list_* returns byte-identical strings to current dropdowns: YES/NO (show values) -- .name vs .value decision for model_version discovery: <...> -- CONTRACT_VERSION set; additive-only policy linked: -- Characterization tests unchanged & green (uv run pytest): YES/NO -``` - ---- - -## Phase E2 — Move conversion code into `coreml_diffusion` (in-repo, dedup, behavior-preserving) - -**Objective:** physically relocate the conversion mechanics into the framework-free package, -collapsing the two duplicate converters into one, with paths/device injected. - -### Tasks -1. Move into `coreml_diffusion/`: `pipeline.py` (`get_pipeline`, `get_unet`), `unet.py` - (`UNet2DConditionModelLCM`), `inputs.py` (sample/lcm/sdxl input builders + - `get_encoder_hidden_states_shape` + `get_coreml_inputs` + `get_inputs_spec`), - `controlnet.py` (`add_cnet_support`), `convert.py` (`convert_unet`, `convert`, - `convert_to_coreml`, `load_coreml_model`), `compile.py` (`compile_coreml_model`). -2. **Dedupe LCM vs main** (the real remaining duplication): delete `lcm/converter.py`'s copies of - `convert_to_coreml` / `load_coreml_model` / `get_out_path` / `get_sample_input` (LCM variant - carries a `scheduler` arg — fold that into the shared `get_sample_input` as an optional param) - in favor of the single `coreml_diffusion` implementation. The main path's helpers - (`get_unet`/`get_encoder_hidden_states_shape`/`get_coreml_inputs`/`convert_unet`/`convert`) and - the `conversion/` subpackage (`attention`, `shapes`, `trace`, `unet`) move as-is. -3. **Inject paths**: replace `get_out_path`'s `folder_paths` reach-in with an injected `out_path` - argument on `convert(...)`; `folder_paths` resolution moves up into the node (which already - computes `out_name`). No `folder_paths` import anywhere in `coreml_diffusion`. -4. **Inject device** where the LCM path used `comfy.model_management` (verify it still does): - `convert(..., device=None)`, default to torch's default device. -5. **Attention per-call, never at import:** main path already routes through - `conversion/attention.apply_attention_implementation` — keep that. If `lcm/converter.py` still - sets any module global at import, remove it; the package sets attention from the `attn_impl` - arg inside `convert`. -6. **Move `compose_out_name` into the package** (`coreml_diffusion/naming.py`); re-point - `test_characterization_out_name.py` imports to `coreml_diffusion.naming` — assertions and values - unchanged. The node will import it from the package in E3. -7. Leave **thin shims** in `coreml_suite/converter.py` and `coreml_suite/lcm/converter.py` that - re-export from `coreml_diffusion`, preserving the old call signatures the nodes use (nodes untouched - this phase). Shims map comfy `folder_paths`/device into package args. - -### Acceptance criteria -- `uv run pytest -m unit` (Tier 0) imports `coreml_diffusion.*` with **no comfy / no MPS** and is green on Linux. -- The dedup leaves exactly one implementation of each previously-duplicated function. -- Characterization tests pass unchanged after the `compose_out_name` re-point. -- `[M2]` A real SD1.5 conversion via the shim still produces a loadable model. -- `[M2-ANE]` **Golden latent identical / within tolerance** to the MODERNIZATION Phase 2 anchor - (same seed/prompt) — proves the move + dedup changed nothing. - -### STOP — VALIDATE (Gate E2 — first regression gate) -``` -## Gate E2 report -- Tier 0 import of coreml_diffusion without comfy/MPS (uv run): PASS/FAIL -- LCM/main duplicated funcs collapsed to one (list old→new): -- compose_out_name moved to package; char-tests re-pointed & green: YES/NO -- Paths injected (no folder_paths in package): confirmed -- Device injected (no comfy.model_management in package): confirmed -- Attention set per-call, not at import (both main & lcm): confirmed -- [M2-ANE] Golden latent vs Phase-2 anchor: identical / within tol / DIVERGED (STOP) -- Node INPUT_TYPES / mappings untouched: confirmed (diff) -``` -**If the golden latent diverged at all, STOP and report — do not continue.** - ---- - -## Phase E3 — Thin the nodes onto the package (behavior-preserving) - -**Objective:** remove the shims; have the ComfyUI nodes call `coreml_diffusion` directly, keeping the -node contract byte-identical. - -### Tasks -1. `CoreMLConverter.convert` (in `coreml_suite/nodes.py`): keep the `folder_paths`-based path - resolution **in the node**; import `compose_out_name` from `coreml_diffusion` (not `coreml_suite.core`); - call `coreml_diffusion.convert(...)` and `coreml_diffusion.compile_model(...)` directly; wrap the compiled path - in `CoreMLModel`. -2. **Wire the dropdowns to discovery (the maintainer's hard requirement).** Replace the hardcoded - `INPUT_TYPES` lists with fail-soft discovery calls: - ```python - def _discover(fn, fallback): - try: - import coreml_diffusion - return getattr(coreml_diffusion, fn)() - except Exception as e: # missing/old package, or import error - logger.warning(f"coreml_diffusion.{fn} unavailable ({e}); using fallback {fallback}") - return fallback - ... - "model_version": (_discover("list_model_versions", ["SD15", "SDXL"]),), - "attention_implementation": (_discover("list_attention_impls", ["SPLIT_EINSUM","SPLIT_EINSUM_V2","ORIGINAL"]),), - "quantize_nbits": (_discover("list_quant_modes", ["none","8","6","4"]), {"default": "none"}), - ``` - This is what makes "update the package → new types appear in the old node, no Suite bump" true. -3. `COREML_CONVERT_LCM` (in `coreml_suite/lcm/nodes.py`): route through `coreml_diffusion` for the shared - mechanics. **Keep the existing LCM behavior/HF-hardcode for now** — consolidation is optional E-LCM. -4. Delete the now-dead `coreml_suite/converter.py` / `coreml_suite/lcm/converter.py` shims (or - reduce to a one-line re-export if anything external imports them — grep first). - -### Acceptance criteria -- `NODE_CLASS_MAPPINGS` / `NODE_DISPLAY_NAME_MAPPINGS` keys: **unchanged** (diff `__init__.py`). -- Every `INPUT_TYPES` **field name** unchanged. Dropdown **values**: the discovery calls must - return **a superset of** today's values, with every previously-present value still present and - spelled identically (additive-only). *(This deliberately replaces the original spec's - "values must be byte-identical/frozen" criterion — the maintainer requires the list be - extensible at runtime. Frozen-field-names + additive-only-values is the new contract.)* -- With `coreml_diffusion` **absent**, the node still registers and shows the fallback lists (fail-soft). -- `[M2-ANE]` Golden latent still identical to the Phase-2 anchor. -- `[M2-ANE]` The committed e2e workflow `tests/integration/...` still passes (PSNR > 25). - -### STOP — VALIDATE (Gate E3) -``` -## Gate E3 report -- Node mappings diff: empty (confirmed) -- INPUT_TYPES field-names diff: empty (confirmed) -- Dropdown values: superset of prior, all prior values still present & identical: YES/NO (show) -- Fail-soft with coreml_diffusion absent (node still registers): PASS/FAIL -- compose_out_name now imported from coreml_diffusion (no node-side copy): confirmed -- [M2-ANE] Golden latent vs anchor: identical / within tol / DIVERGED (STOP) -- [M2-ANE] e2e workflow PSNR: (> 25?) -- Dead converter shims removed / reduced: -``` - ---- - -## Phase E4 — Standalone packaging & CLI (still in-repo) - -**Objective:** make `coreml_diffusion` independently installable and usable without ComfyUI, with a CLI -suitable for the planned article and for on-device/iOS conversion workflows. - -### Tasks -1. Add `coreml_diffusion/pyproject.toml`: name (working `coreml_diffusion`), `requires-python`, dependencies - = `coremltools` (pinned to the MODERNIZATION-validated version), `diffusers`, `transformers`, - `peft`, `omegaconf`, `numpy`, `torch`. **No `ml-stable-diffusion`** (already removed in #58, see - §0.3) and **no comfy**. Suite pins `transformers>=4.44`/`peft>=0.13`/`omegaconf>=2.3` today; - grep-confirm each is on the conversion path before listing it. A `[project.scripts]` entry: - `coreml-diffusion = "coreml_diffusion.cli:main"`. -2. `coreml_diffusion/cli.py`: `coreml-diffusion convert --ckpt PATH --model-version sd15 --out PATH - [--height --width --batch-size --attn-impl --controlnet --lora NAME:STRENGTH ... --config PATH]` - and `coreml-diffusion compile --src PATH --out-dir DIR --name NAME`. Mirrors `convert()`/`compile_model()`. -3. Tier-0 Linux tests for the CLI **arg→call mapping** (mock the heavy `convert`); the real - convert remains `[M2]`. Add a `[M2]` smoke test: convert a tiny synthetic UNet end-to-end. -4. README for the package: install, CLI usage, "produce a `.mlpackage`/`.mlmodelc` for use in a - Swift/iOS app", and the ANE positioning note (low-power, GPU-free, embeddable; SD1.5/SDXL on - ANE, **not** a Flux-speed claim). - -### Acceptance criteria -- Fresh `python -m venv` + `uv pip install ./coreml-diffusion` (no ComfyUI present) imports and runs - `coreml-diffusion --help` and the arg-mapping tests on Linux. -- `[M2]` `coreml-diffusion convert` produces a model file identical (golden) to the node path. - -### STOP — VALIDATE (Gate E4) -``` -## Gate E4 report -- uv pip install ./coreml-diffusion in comfy-free venv: PASS/FAIL (log) -- CLI arg→call tests (Tier 0, Linux): green -- [M2] CLI-produced model golden vs node-produced model: identical / DIVERGED -- Package deps list (with pinned SHAs/versions + licenses): -- New runtime deps vs suite before: -``` -**This is the gate that proves the package stands alone. Do not split repos before it passes.** - ---- - -## Phase E5 — Physical split into a second repository - -**Objective:** move `coreml_diffusion/` to its own repo; CoreMLSuite depends on it by pinned version. - -### Tasks -1. Create the new repo (maintainer action — agent prepares the tree, not the GitHub repo). - Choose final distributable name; rename imports if changed (single sweep, recorded). -2. CoreMLSuite `pyproject.toml` / `requirements.txt`: replace the conversion-only deps with a - pinned dependency on the new package (`coreml_diffusion==` from PyPI, or `git+...@` - until first PyPI release). (There is no `git+...ml-stable-diffusion` line to remove — already - gone since #58.) -3. ~~Keep `python_coreml_stable_diffusion` for the loader.~~ **Void.** The loader is the local - `coreml_suite/coreml_model.py` over `coremltools`; the suite keeps `coremltools` as a direct dep - for it. No Apple lib involved. -4. Set up the new repo's CI: Tier 0 on Linux (import + arg-mapping + input-shape math), - `[M2]`/`[M2-ANE]` on a self-hosted/macOS-ARM runner reusing the golden-latent anchor. -5. Versioning: SemVer; first release `0.1.0`. Document the compatibility matrix - (coreml_diffusion ↔ coremltools version ↔ diffusers version). No ml-stable-diffusion axis. - -### Acceptance criteria -- CoreMLSuite installs in a fresh venv pulling the new package; e2e workflow still passes `[M2-ANE]`. -- New repo CI green on Linux (Tier 0) and `[M2-ANE]` golden latent matches the anchor. -- No conversion code remains in CoreMLSuite (grep: no `ct.convert`, no `from_single_file`, - no `torch.jit.trace`). - -### STOP — VALIDATE (Gate E5) -``` -## Gate E5 report -- New repo tree prepared: ; final package name: -- Suite depends on package by pinned version: -- Suite e2e [M2-ANE] PSNR after split: (> 25?) -- Conversion code fully absent from suite: confirmed (grep output) -- Compatibility matrix documented: -- First release tag: 0.1.0 -``` - ---- - -## Phase E6 — Quantization travels WITH the conversion code (already implemented → move) - -**Objective:** quantization is **already implemented** (k-means `palettize_weights` in -`converter.py`, `quantize_nbits` node input, `_q` filename suffix, README tradeoff table). -There is nothing to *build*. It simply **moves with the conversion code in E2** as part of -`convert_unet`. This phase is a checkpoint that it survived the extraction intact, plus exposing -it through the CLI. - -### Tasks -1. Confirm the palettization block moved cleanly into `coreml_diffusion` (lives in `convert.py` or a - `quantize.py` helper called from `convert_unet`). Default `"none"` stays byte-identical. -2. Expose via CLI flag `--quantize {none,8,6,4}` (E4 already lists this) and via - `list_quant_modes()` discovery (E1/E3). -3. The existing README tradeoff table (SD1.5 1×512×512 SPLIT_EINSUM: none/8/6/4 → size/ms/PSNR) - moves to the package README. Re-confirm one row `[M2-ANE]` so the article can cite a live number. - -### Acceptance criteria -- Default (`none`) output byte-identical to pre-extraction (covered by the E2/E3 golden latent). -- `coreml-diffusion convert --quantize 4` produces a `_q4` artifact matching the node's `_q4` artifact `[M2]`. -- `list_quant_modes()` drives the node dropdown (no hardcoded copy remains). - -### STOP — VALIDATE (Gate E6) -``` -## Gate E6 report -- Palettization relocated into coreml_diffusion, called from convert_unet: confirmed -- Default none output identical (golden): YES/NO -- [M2] CLI --quantize {8,6,4} artifacts match node artifacts: YES/NO -- Tradeoff table in package README with at least one re-confirmed [M2-ANE] row: -``` - ---- - -## Phase E-LCM — FIRST task after the split: clean up LCM + verify → promote (behavior-changing, gated) - -> Promoted from "optional, someday" to **the first thing after E5**, per maintainer intent: the -> Suite should expose every verifiably-convertible model, and LCM is the obvious first cleanup. - -Two coupled goals: -1. **Consolidate the LCM path.** Make the LCM node use the unified `from_single_file` path in - `coreml_diffusion.convert(model_version=LCM, ...)` instead of the hardcoded `SimianLuo/LCM_Dreamshaper_v7` - HF download; drop the duplicated LCM helpers (already deduped in E2). **Behavior change** ⇒ - capture an LCM golden anchor *before* the change, then prove within-tolerance after. -2. **Verify → promote.** Once the LCM conversion has a passing `[M2-ANE]` golden anchor, flip - `_MODEL_STATUS["lcm"] = Status.VERIFIED` **in the package** (minor bump). The Suite's dropdown - gains `lcm` automatically — no Suite change, no Suite bump. This is the end-to-end proof that - the discovery contract works as designed. - -Repeat the same recipe for `sdxl_refiner` when it gets an anchor (separate small gate). Do NOT -bundle E-LCM into E1–E5; it changes behavior and must stand on its own golden. - -### STOP — VALIDATE (Gate E-LCM) -``` -## Gate E-LCM report -- LCM golden anchor captured BEFORE change: -- LCM node now uses unified from_single_file path; HF hardcode removed: confirmed -- [M2-ANE] LCM golden after change: identical / within tol / DIVERGED (STOP) -- Status flipped lcm→VERIFIED in package (minor bump ): confirmed -- Suite dropdown now lists lcm with NO Suite code change / NO Suite bump: confirmed (diff empty) -- LCM node accepts a checkpoint arg now (documented breaking-ish UI note): -``` - ---- - -## Article deliverable (after E4) - -Once the CLI exists and stands alone, the "convert a Comfy/A1111 workflow into an on-device iOS -app" write-up becomes a clean tutorial: `coreml-diffusion convert` → `.mlmodelc` → load in Swift/CoreML. -Frame the niche honestly per the README note above (ANE feasibility & power, not raw Flux speed). - ---- - -## Quick reference: extraction gate discipline - -``` -E0 Seam decision, interface contract, discovery API frozen → Gate E0 (cut line + additive-only policy?) -E1 Stand up coreml_diffusion + discovery API (mostly verify) → Gate E1 (list_* byte-identical, comfy-free?) -E2 Move conversion code, dedup LCM/main, inject paths/device→ Gate E2 (golden identical? duplicates gone?) ← first regression gate -E3 Thin nodes onto package + wire discovery dropdowns → Gate E3 (field-names frozen, values additive, fail-soft, golden identical?) -E4 Standalone packaging + CLI (uv) → Gate E4 (uv pip install w/o comfy? CLI golden?) ← proves it stands alone -E5 Physical second-repo split → Gate E5 (suite depends on pkg? conversion absent?) -E-LCM FIRST post-split: clean up LCM, verify → promote → Gate E-LCM (LCM golden? dropdown gains lcm w/ no Suite bump?) -E6 Quantization checkpoint (already built → moved in E2) → Gate E6 (default identical? CLI quant matches?) -(refiner) same recipe as E-LCM when an anchor exists → own small gate (promote sdxl_refiner→VERIFIED) -``` - -**Interface-contract invariants (the maintainer's hard requirement), restated:** -- Package API is keyword-only-with-defaults past the required positionals → converter updates - don't force Suite updates. -- Node dropdowns are discovery-driven (`coreml_diffusion.list_*`) + fail-soft → new conversion types - appear in the old plugin with `uv pip install -U coreml_diffusion`, **no Suite code change, no bump**. -- Discovery identifiers are **additive-only**; removal/rename = MAJOR bump + migration note. -- `compose_out_name` (cache key) lives in the package, single copy. - - -**Golden rule (inherited): never cross a gate with a failing acceptance criterion. -Stop, report, wait. The golden latent is the single source of truth that the extraction -changed nothing.** diff --git a/README.md b/README.md index e97ca15..8290b9c 100644 --- a/README.md +++ b/README.md @@ -1,452 +1,113 @@ # Core ML Suite for ComfyUI -## Overview +Custom nodes for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) that run +Stable Diffusion UNets as [Core ML](https://developer.apple.com/documentation/coreml) +models on Apple Silicon (M1/M2/M3). Core ML can use the Apple Neural Engine +(ANE), which is unavailable to PyTorch — on an M2 Pro 32 GB, SD1.5 at 512×512 +generates roughly **1.5–2× faster** than the standard PyTorch/MPS path. -Welcome! In this repository you'll find a set of custom nodes for [ComfyUI](https://github.com/comfyanonymous/ComfyUI) -that allows you to use Core ML models in your ComfyUI workflows. -These models are designed to leverage the Apple Neural Engine (ANE) on Apple Silicon (M1/M2) machines, -thereby enhancing your workflows and improving performance. +You convert a Stable Diffusion checkpoint to a Core ML model with the nodes in +this suite, then sample from it like any other ComfyUI workflow. -If you're not sure how to obtain these models, you can download them -[here](https://huggingface.co/coreml-community) or convert your own checkpoints -directly with the conversion nodes in this suite (see [How to use](#how-to-use)). +> [!IMPORTANT] +> **Convert your own checkpoints — that is the only supported path.** This +> suite uses its own input dimensions, naming convention, and metadata +> (produced by the [coreml-diffusion](https://github.com/aszc-dev/coreml-diffusion) +> package). Pre-converted Core ML models from elsewhere (e.g. the +> coreml-community Hugging Face org) are **not** supported. Conversion is cheap +> and runs on your machine, so there is no need to download Core ML models. -In simple terms, think of Core ML models as a tool that can help your ComfyUI work faster and more efficiently. -For instance, during my tests on an M2 Pro 32GB machine, -the use of Core ML models sped up the generation of 512x512 images by a factor -of approximately 1.5 to 2 times. +## Installation -## Getting Started +### ComfyUI-Manager (recommended) -To start using custom nodes in your ComfyUI, follow these simple steps: +Open **Manager → Install Custom Nodes**, search for `Core ML`, click +**Install**, and restart ComfyUI. -1. Clone or download this repository: You can do this directly into the custom_nodes directory of your ComfyUI. -2. Install the dependencies: You'll need to use a package manager like pip to do this. +### Manual -That's it! You're now ready to start enhancing your ComfyUI workflows with Core ML models. +```bash +cd /path/to/comfyui/custom_nodes +git clone https://github.com/aszc-dev/ComfyUI-CoreMLSuite.git +cd ComfyUI-CoreMLSuite +pip install -r requirements.txt +``` -- Check [Installation](#installation) for more details on installation. -- Check [How to use](#how-to-use) for more details on how to use the custom nodes. -- Check [Example Workflows](#example-workflows) for some example workflows. +Dependencies (`coreml-diffusion`, `coremltools`, `numpy`, `diffusers`) install +from PyPI. PyTorch is intentionally **not** pinned — it is provided by your +ComfyUI host, and a hard cap here would downgrade it and break ComfyUI. + +## Quickstart + +1. Put a SD1.5 checkpoint in `models/checkpoints`. +2. Add the **Convert Checkpoint to Core ML** node, select the checkpoint, and + queue once. It writes a `.mlpackage` to `models/unet` (cached by name — it + won't reconvert next time). +3. Sample with the **Core ML Sampler** node, decoding the latent with a normal + VAE Decode. CLIP and VAE come from standard ComfyUI nodes. + +See [docs/workflows.md](docs/workflows.md) for complete example graphs (txt2img, +ControlNet, LoRA, LCM, SDXL). + +## Which compute unit should I pick? + +The **compute unit** selects the hardware Core ML runs on. Pair it with the +attention implementation chosen at conversion time: + +| Model | Convert with | Load with | Runs on | +|---|---|---|---| +| SD1.5 @ 512×512 | `SPLIT_EINSUM` | `CPU_AND_NE` | Neural Engine (fastest) | +| SD1.5 @ larger sizes | `ORIGINAL` | `CPU_AND_GPU` | GPU | +| SDXL | `ORIGINAL` | `CPU_AND_GPU` | GPU (ANE unsupported) | + +`CPU_AND_NE` is usually the fastest option for SD1.5 — often faster than `ALL`. +This suite uses Core ML compute units only; it never touches PyTorch MPS, so +`PYTORCH_ENABLE_MPS_FALLBACK` is irrelevant to these nodes. Full reasoning and +benchmarks: [docs/hardware.md](docs/hardware.md). + +## Documentation + +- [Hardware & compute units](docs/hardware.md) — ANE vs GPU vs MPS, attention + implementations, which to choose. +- [Nodes](docs/nodes.md) — full reference for every node. +- [Conversion](docs/conversion.md) — how conversion works, caching, + quantization. +- [Example workflows](docs/workflows.md) — annotated example graphs. +- [FAQ](docs/faq.md) — answers to common questions. +- [Troubleshooting](docs/troubleshooting.md) — common errors and fixes. +- [Limitations & support matrix](docs/limitations.md) — what is and isn't + supported. ## Glossary -- **Core ML**: A machine learning framework developed by Apple. It's used to run machine learning models on Apple - devices. -- **Core ML Model**: A machine learning model that can be run on Apple devices using Core ML. -- **mlmodelc**: A compiled Core ML model. This is the recommended format for Core ML models. -- **mlpackage**: A Core ML model packaged in a directory. This is the default format for Core ML models. -- **ANE**: Apple Neural Engine. A hardware accelerator for machine learning tasks on Apple devices. -- **Compute Unit**: A Core ML option that allows you to specify the hardware on which the model should run. - - **CPU_AND_ANE**: A Core ML compute unit option that allows the model to run on both the CPU and ANE. This is the - default option. - - **CPU_AND_GPU**: A Core ML compute unit option that allows the model to run on both the CPU and GPU. - - **CPU_ONLY**: A Core ML compute unit option that allows the model to run on the CPU only. - - **ALL**: A Core ML compute unit option that allows the model to run on all available hardware. -- **CLIP**: Contrastive Language-Image Pre-training. A model that learns visual concepts from natural language - supervision. It's used as a text encoder in Stable Diffusion. -- **VAE**: Variational Autoencoder. A model that learns a latent representation of images. It's used as a prior in - Stable Diffusion. -- **Checkpoint**: A file that contains the weights of a model. It's used to load models in Stable Diffusion. -- **LCM**: [Latent Consistency Model](https://latent-consistency-models.github.io/). A type of model designed to - generate images with as few steps as possible. - -> [!NOTE] -> Note on Compute Units: -> For the model to run on the ANE, the model must be converted with the `--attention-implementation SPLIT_EINSUM` -> option. -> Models converted with `--attention-implementation ORIGINAL` will run on GPU instead of ANE. - -## Features - -These custom nodes come with a host of features, including: - -- Loading Core ML Unet models -- Support for ControlNet -- Support for ANE (Apple Neural Engine) -- Support for CPU and GPU -- Support for `mlmodelc` and `mlpackage` files -- Support for SDXL models -- Support for LCM models -- Support for LoRAs -- SD1.5 -> Core ML conversion -- SDXL -> Core ML conversion -- LCM -> Core ML conversion - -> [!NOTE] -> Please note that using Core ML models can take a bit longer to load initially. -> For the best experience, I recommend using the compiled models -> (.mlmodelc files) instead of the .mlpackage files. - -> [!NOTE] -> This repository will continue to be updated with more nodes and features over time. - -## Conversion & Acknowledgements - -The Core ML conversion pipeline in this repository began as an adaptation of -Apple's [ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion), -which pioneered running Stable Diffusion on the Apple Neural Engine. The -implementation has since diverged and no longer depends on that package: - -- UNet conversion runs natively on `diffusers`' `UNet2DConditionModel`. -- The ANE-friendly attention path (`SPLIT_EINSUM`, `SPLIT_EINSUM_V2`) is - reimplemented as standalone `diffusers` attention processors. -- The toolchain tracks current ComfyUI (NumPy 2, Torch 2.7, coremltools 9, - Python 3.12). - -The goal is to keep iterating on these methods independently and to explore -support beyond SD1.5. +- **Core ML** — Apple's on-device machine-learning framework. +- **`.mlpackage`** — the Core ML model format this suite produces and loads. +- **ANE** — Apple Neural Engine, a hardware accelerator for ML. +- **Compute unit** — which hardware Core ML uses (`CPU_AND_NE`, `CPU_AND_GPU`, + `CPU_ONLY`, `ALL`). +- **Attention implementation** — `SPLIT_EINSUM` / `SPLIT_EINSUM_V2` (ANE-friendly) + or `ORIGINAL` (GPU-friendly), chosen at conversion. > [!IMPORTANT] > **Breaking change in 2.0.0.** The converted Core ML UNet now takes > `encoder_hidden_states` in the native `diffusers` layout -> `(batch, tokens, hidden)` instead of the previous -> `(batch, hidden, 1, tokens)`. Core ML models converted with earlier versions -> are not compatible with 2.0.0 and must be re-converted. - -## Installation - -### Using ComfyUI-Manager - -The easiest way to install the custom nodes is to use the ComfyUI-Manager. You can find the installation instructions -[here](https://github.com/ltdrdata/ComfyUI-Manager#installation). Once you've installed the ComfyUI-Manager, you can -install the custom nodes by following these steps: - -- Open the ComfyUI-Manager by clicking the `Manager` button in the ComfyUI toolbar. -- Click the `Install Custom Nodes` button. -- Search for `Core ML` and click the `Install` button. -- Restart ComfyUI. - -### Manual Installation - -1. Clone this repository into the custom_nodes directory of your ComfyUI. If you're not sure how to do this, you can - download the repository as a zip file and extract it into the same directory. - ```bash - cd /path/to/comfyui/custom_nodes - git clone https://github.com/aszc-dev/ComfyUI-CoreMLSuite.git - ``` -2. Next, install the required dependencies using pip or another package manager: - - ```bash - cd /path/to/comfyui/custom_nodes/ComfyUI-CoreMLSuite - pip install -r requirements.txt - ``` - -## How to use - -Once you've installed the custom nodes, you can start using them in your ComfyUI workflows. -To do this, you need to add the nodes to your workflow. You can do this by right-clicking on the workflow canvas and -selecting the nodes from the list of available nodes (the nodes are in the `Core ML Suite` category). -You can also double-click the canvas and use the search bar to find the nodes. The list of available nodes is given -below. - -### Available Nodes - -#### Core ML UNet Loader (`CoreMLUnetLoader`) - -![CoreMLUnetLoader](./assets/unet_loader.png?raw=true) - -This node allows you to load a Core ML UNet model and use it in your ComfyUI workflow. Place the converted -.mlpackage or .mlmodelc file in ComfyUI's `models/unet` directory and use the node to load the model. The output of the -node is a `coreml_model` object that can be used with the Core ML Sampler. - -- **Inputs**: - - **model_name**: The name of the model to load. This should be the name of the .mlpackage or .mlmodelc file. - - **compute_unit**: The hardware on which the model should run. This can be one of the following: - - `CPU_AND_ANE`: The model will run on both the CPU and ANE. This is the default option. It works best with - models - converted with `--attention-implementation SPLIT_EINSUM` or `--attention-implementation SPLIT_EINSUM_V2`. - - `CPU_AND_GPU`: The model will run on both the CPU and GPU. It works best with models converted with - `--attention-implementation ORIGINAL`. - - `CPU_ONLY`: The model will run on the CPU only. - - `ALL`: The model will run on all available hardware. -- **Outputs**: - - **coreml_model**: A Core ML model that can be used with the Core ML Sampler. - -#### Core ML Sampler (`CoreMLSampler`) - -![CoreMLSampler](./assets/sampler.png?raw=true) - -This node allows you to generate images using a Core ML model. The node takes a Core ML model as input and outputs a -latent image similar to the latent image output by the KSampler. This means that you can use the -resulting latent as you normally would in your workflow. - -- **Inputs**: - - **coreml_model**: The Core ML model to use for sampling. This should be the output of the Core ML UNet Loader. - - **latent_image** [optional]: The latent image to use for sampling. If provided, should be of the same size as the - input of the Core ML model. If not provided, the node will create a latent suitable for the Core ML model used. - Useful in img2img workflows. - - ... _(the rest of the inputs are the same as the KSampler)_ -- **Outputs**: - - **LATENT**: The latent image output by the Core ML model. This can be decoded using a VAE Decoder or used as input - to the next node in your workflow. - -#### Checkpoint Converter - -![CoreMLConverter](./assets/checkpoint_converter.png?raw=true) - -You can use this node to convert any **SD1.5** based checkpoint to a Core ML model. The converted model is stored in the -`models/unet` directory and can be used with the `Core ML UNet Loader`. The conversion parameters are encoded in -the node name, so if the model already exists, the node will not convert it again. - -- **Inputs**: - - **ckpt_name**: The name of the checkpoint to convert. This should be the name of the checkpoint file stored in the - `models/checkpoints` directory. - - **model_version**: Whether the model is based on SD1.5 or SDXL. - - **height**: The desired height of the image generated by the model. The default is 512. Any positive multiple of 8 is accepted. - - **width**: The desired width of the image generated by the model. The default is 512. Any positive multiple of 8 is accepted. - - **batch_size**: The batch size of generated images. If you're planning to generate batches of images, you can try - increasing this value to speed up the generation process. The default is 1. - - **attention_implementation**: The attention implementation used when converting the model. Choose SPLIT_EINSUM or - SPLIT_EINSUM_V2 for better ANE support. Choose ORIGINAL for better GPU support. - - **compute_unit**: The hardware on which the model should run. This is used only when loading the model and doesn't - affect the conversion process. - - **controlnet_support**: For the model to support ControlNet, it must be converted with this option set to True. - The - default is False. - - **lora_params** [optional]: Optional LoRA names and weights. If provided, the model will be converted with LoRA(s) - baked in. More on loading LoRAs below. -- **Outputs**: - - **coreml_model**: The converted Core ML model that can be used with Core ML Sampler. - -> [!NOTE] -> Some models use a custom config .yaml file. If you're using such a model, you'll need to place the config file in the -> `models/configs` directory. The config file should be named the same as the checkpoint file. For example, if the -> checkpoint file is named `juggernaut_aftermath.safetensors`, the config file should be -> named `juggernaut_aftermath.yaml`. -> The config file will be automatically loaded during conversion. - -> [!NOTE] -> For now, the converter relies heavilty on the model name to determine the conversion parameters. This means that if -> you change the model name, the node will convert the model again. Other than that, if you find the name too long or -> confusing, you can change it to anything you want. - -#### LoRA Loader - -![LoRALoader](./assets/lora_loader.png?raw=true) - -This node allows you to load LoRAs and bake them into a model. Since this is a workaround (as model weights can't be -modified -after conversion), there are a few caveats to keep in mind: - -- The LoRA weights and _strength_model_ parameter are baked into the model. This means that you can't change them - after conversion. This also means that you need to convert the model again if you want to change the LoRA weights. -- Loading LoRA affects CLIP, which is not a part of Core ML workflow, so you'll need to load CLIP separately, - either using `CLIPLoader` or `CheckpointLoaderSimple`. (See [example workflows](#example-workflows) for more details.) -- After conversion, if you want to load the model using `CoreMLUnetLoader`, you'll need to apply the same LoRAs to - CLIP manually. (See [example workflows](#example-workflows) for more details.) -- The LoRA names are encoded in the model name. This means that if you change the name of the LoRA file, - you'll need to change the model name as well, or the node will convert the model again. (Model strength is not - encoded, so if you want to change it, you'll need to delete the converted model manually) -- _strength_clip_ parameter only affects the CLIP model and is not baked into the converted model. This means that - you can change it after conversion. - -- **Inputs**: - - **lora_name**: The name of the LoRA to load. - - **strength_model**: The strength of the LoRA model. - - **strength_clip**: The strength of the LoRA CLIP. - - **lora_params** [optional]: Optional output from other LoRA Loaders. - - **clip**: The CLIP model to use with the LoRA. This can be either output of the - `CLIPLoader`/`CheckpointLoaderSimple` or other LoRA Loaders. -- **Outputs**: - - **lora_params**: The LoRA parameters that can be passed to the Core ML Converter or other LoRA Loaders. - - **CLIP**: The CLIP model with LoRA applied. - -#### LCM Converter - -![LCMConverter](./assets/lcm_converter.png?raw=true) - -This node converts [SimianLuo/LCM_Dreamshaper_v7](https://huggingface.co/SimianLuo/LCM_Dreamshaper_v7) model to Core -ML. The converted model is stored in the `models/unet` directory and can be used with the Core ML UNet Loader. The -conversion parameteres are encoded in the node name, so if the model already exists, the node will not convert it again. - -- **Inputs**: - - **height**: The desired height of the image generated by the model. The default is 512. Must be a multiple of 8. - - **width**: The desired width of the image generated by the model. The default is 512. Must be a multiple of 8. - - **batch_size**: The batch size of generated images. If you're planning to generate batches of images, you can try - increasing this value to speed up the generation process. The default is 1. - - **compute_unit**: The hardware on which the model should run. This is used only when loading the model and - doesn't affect the conversion process. - - **controlnet_support**: For the model to support ControlNet, it must be converted with this option set to True. - The default is False. - -> [!NOTE] -> The conversion process can take a while, so please be patient. - -> [!NOTE] -> When using the LCM model with Core ML Sampler, please set _sampler_name_ to `lcm` and _scheduler_ to `sgm_uniform`. - -#### Core ML Adapter (Experimental) (`CoreMLModelAdapter`) - -![CoreMLModelAdapter](./assets/adapter.png?raw=true) - -This node allows you to use a Core ML as a standard ComfyUI model. This is an experimental node and may not work with -all models and nodes. Please use with caution and pay attention to the expected inputs of the model. - -- **Input**: - - **coreml_model**: The Core ML model to use as a ComfyUI model. -- **Output**: - - **MODEL**: The Core ML model wrapped in a ComfyUI model. - -> [!NOTE] -> While this approach allows you to use Core ML models with many ComfyUI nodes (both standard and custom), the -> expected inputs of the model will not be checked, which may cause errors. Please make sure to use a model compatible -> with the expected parameters. - -### Example Workflows - -> [!NOTE] -> The models used are just an example. Feel free to experiment with different models and see what works best for you. - -#### Basic txt2img with Core ML UNet loader - -This is a basic txt2img workflow that uses the Core ML UNet loader to load a model. The CLIP and VAE models -are loaded using the standard ComfyUI nodes. In the first example, the text encoder (CLIP) and VAE models are loaded -separately. In the second example, the text encoder and VAE models are loaded from the checkpoint file. Note that you -can use any CLIP or VAE model as long as it's compatible with Stable Diffusion v1.5. - -1. **Loading text encoder (CLIP) and VAE models separately** - - This workflow uses CLIP and VAE models available - [here](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/text_encoder/model.safetensors) and - [here](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/vae/diffusion_pytorch_model.safetensors). - Once downloaded, place the models in the`models/clip` and `models/vae` directories respectively. - - The Core ML UNet model is available - [here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip). - Once downloaded, place the model in the `models/unet` directory. - ![coreml-unet+clip+vae](./assets/unet+sampler+clip+vae.png?raw=true) -2. **Loading text encoder (CLIP) and VAE models from checkpoint file** - - This workflow loads the CLIP and VAE models from the checkpoint file available - [here](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/v1-5-pruned-emaonly.safetensors). - Once downloaded, place the model in the`models/checkpoints` directory. - - The Core ML UNet model is available - [here](https://huggingface.co/coreml-community/coreml-stable-diffusion-v1-5_cn/blob/main/split_einsum/stable-diffusion-_v1-5_split-einsum_cn.zip). - Once downloaded, place the model in the `models/unet` directory. - ![coreml-unet+checkpoint](./assets/unet+sampler+checkpoint.png?raw=true) - -#### ControlNet with Core ML UNet loader - -This workflow uses the Core ML UNet loader to load a Core ML UNet model that supports ControlNet. The ControlNet is -being loaded using the standard ComfyUI nodes. Please refer to -the [basic txt2img workflow](#basic-txt2img-with-core-ml-unet-loader) for more details on how to load the CLIP and VAE -models. -The ControlNet model used in this workflow is available -[here](https://huggingface.co/lllyasviel/control_v11p_sd15_scribble/blob/main/diffusion_pytorch_model.fp16.safetensors). -Once downloaded, place the model in the `models/controlnet` directory. -![coreml-unet+controlnet](./assets/unet+sampler+controlnet.png?raw=true) - -#### Checkpoint conversion - -This workflow uses the Checkpoint Converter to convert the checkpoint file. See -[Checkpoint Converter](#checkpoint-converter) description for more details. - -![checkpoint-converter](./assets/basic_conversion.png?raw=true) - -#### Checkpoint conversion with LoRA - -This workflow uses the Checkpoint Converter to convert the checkpoint file with LoRA. See -[LoRA Loader](#lora-loader) description to read more about the caveats of using LoRA. - -![checkpoint-converter+lora](./assets/conversion+lora.png?raw=true) - -#### LCM LoRA conversion - -Please note that you can use multiple LoRAs with the same model. To do this, you'll need to use multiple LoRA Loaders. -> [!IMPORTANT] -> In this example, the model is passed through the adapter and `ModelSamplingDiscrete` nodes to a standard ComfyUI's -> KSampler (not Core ML Sampler). ModelSamplingDiscrete needs to be used to sample models with LCM LoRAs properly. - -![multiple-loras](./assets/conversion+lcm_lora.png?raw=true) - -#### Loader with LoRAs - -This workflow uses the Core ML UNet Loader to load a model with LoRAs. The CLIP must be loaded separately and passed -through the same LoRA nodes as during conversion. See [LoRA Loader](#lora-loader) description to read more about the -caveats of using LoRA. Since _lora_name_ and _strength_model_ are baked into the model, it is not necessary to pass -them as inputs to the loader. -> [!IMPORTANT] -> In this example, the model is passed through the adapter and `ModelSamplingDiscrete` nodes to a standard ComfyUI's -> KSampler (not Core ML Sampler). ModelSamplingDiscrete needs to be used to sample models with LCM LoRAs properly. - -![loader+lora](./assets/loader+lcm_lora.png?raw=true) - -#### LCM conversion with ControlNet - -This workflow uses LCM converter to -convert [SimianLuo/LCM_Dreamshaper_v7](https://huggingface.co/SimianLuo/LCM_Dreamshaper_v7) -model to Core ML. The converted model can then be used with or without ControlNet to generate images. -![lcm+controlnet](./assets/lcm+controlnet.png?raw=true) - -#### SDXL Base + Refiner conversion - -This is a basic workflow for SDXL. You add LoRAs and ControlNets the same way as in the previous examples. -You can also skip the refiner step. - -The models used in this workflow are available at the following links: - -- [Base model + text_encoder (clip) + text_encoder_2 (clip2)](https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0) -- [Refiner model](https://huggingface.co/stabilityai/stable-diffusion-xl-refiner-1.0) -- [VAE](https://huggingface.co/stabilityai/sdxl-vae) - -> [!IMPORTANT] -> SDXL on ANE is not supported. If loading of the model gets stuck, please try using CPU_AND_GPU or CPU_ONLY. -> For best results, use ORIGINAL attention implementation. - -![sdxl](./assets/sdxl_conversion.png?raw=true) - -## Quantization (opt-in) - -The `Core ML Converter` and `Core ML LCM Converter` nodes accept an -optional `quantize_nbits` dropdown that runs k-means weight palettization -(`coremltools.optimize.coreml.palettize_weights`) on the UNet before save. - -Values: `none` (default — no quantization, identical to unquantized -behavior and filenames), `8`, `6`, `4`. The number is appended to the -.mlpackage stem as `_q` so quantized and unquantized variants -coexist on disk and in cache. - -### SD1.5 1×512×512 SPLIT_EINSUM tradeoffs (M2 Pro, ANE) - -Measured with 20 UNet forward passes at a fixed seed for the PSNR -comparison: - -| nbits | size (MB) | size vs none | fwd median (ms) | PSNR vs `none` (dB) | -|---|---:|---:|---:|---:| -| none | 1641 | 1.000 | 197.1 | — | -| 8 | 822 | 0.501 | 186.6 | 53.5 | -| 6 | 617 | 0.376 | 183.0 | 40.2 | -| 4 | 412 | 0.251 | 179.8 | 27.5 | - -PSNR here is computed on the raw `noise_pred` output of a single UNet -forward at a fixed seed, not on the final decoded image — it isolates -the quantization-induced drift from sampler / VAE noise. Final-image -PSNR is comfortably higher (the sampler averages over 20 steps). - -### Recommended settings per chip / RAM - -- **8 GB RAM (M1 base, M2 base):** `nbits=4`. ~4× smaller model, still - loads, PSNR 27 dB is visually identical at SD1.5 sizes. -- **16 GB RAM (M1/M2/M3 Pro):** `nbits=6` is the sweet spot — ~2.7× - smaller, PSNR 40 dB, no perceptible quality drop. -- **32 GB+ RAM (Max / Ultra):** `nbits=8` if you want the safety - margin, `none` if you want bit-identical output for golden testing. - -The default stays `none` so existing workflows produce byte-for-byte -identical output. - -## Limitations - -- Core ML models are fixed in terms of their inputs and outputs. - This means you'll need to use latent images of the same size as the input of the model (512x512 is the default for - SD1.5). - However, you can re-convert the model to a different input size using the - conversion nodes in this suite (set the desired width and height). -- SD2.1 models are not supported. - -[^1]: -Unless [EnumeratedShapes](https://apple.github.io/coremltools/docs-guides/source/flexible-inputs.html#select-from-predetermined-shapes) -is used during conversion. Needs more testing. +> `(batch, tokens, hidden)` instead of the previous `(batch, hidden, 1, tokens)`. +> Models converted with earlier versions are not compatible and must be +> re-converted. + +## Acknowledgements + +The conversion pipeline began as an adaptation of Apple's +[ml-stable-diffusion](https://github.com/apple/ml-stable-diffusion), which +pioneered running Stable Diffusion on the Neural Engine. It has since diverged +and no longer depends on that package: UNet conversion runs natively on +`diffusers`' `UNet2DConditionModel`, the ANE attention path (`SPLIT_EINSUM`, +`SPLIT_EINSUM_V2`) is reimplemented as standalone `diffusers` attention +processors, and the toolchain tracks current ComfyUI (NumPy 2, Torch 2.7+, +coremltools 9, Python 3.11/3.12). Conversion now lives in the separate +[coreml-diffusion](https://github.com/aszc-dev/coreml-diffusion) package. ## Support -I'm here to help! If you have any questions or suggestions, don't hesitate to open an issue and I'll do my best -to assist you. +Questions or suggestions? Open an +[issue](https://github.com/aszc-dev/ComfyUI-CoreMLSuite/issues). diff --git a/docs/conversion.md b/docs/conversion.md new file mode 100644 index 0000000..d7ce600 --- /dev/null +++ b/docs/conversion.md @@ -0,0 +1,95 @@ +# Conversion + +## Conversion is the only supported path + +You always start from a Stable Diffusion checkpoint (`.safetensors` / `.ckpt`) +and convert it with the **Convert Checkpoint to Core ML** (or **Convert LCM**) +node. Pre-converted Core ML models from elsewhere are not supported, because: + +- The suite uses its own input **dimensions**, **naming convention**, and + **metadata**, all produced by the + [coreml-diffusion](https://github.com/aszc-dev/coreml-diffusion) package. +- Apple's `ml-stable-diffusion` (which most community Core ML models target) is + effectively obsolete, and the layouts differ (see the 2.0.0 + `encoder_hidden_states` change in the README). +- Conversion is cheap and one-time, so there is no value in maintaining + backwards compatibility with foreign formats. + +The output is always a **`.mlpackage`**. The suite no longer compiles to +`.mlmodelc` (it didn't work with the inference backend), so there is **no Xcode +or `coremlcompiler` dependency**. + +## One-time conversion and name-based caching + +Conversion runs **once**, not on every queue. The converter encodes all +conversion parameters into the output filename (via `coreml_diffusion.compose_out_name`, +called in `coreml_suite/nodes.py:313`): + +- checkpoint name, `batch_size`, `width`, `height` +- `controlnet_support`, `attention_implementation` +- baked LoRA names, `quantize_nbits` + +The result is written as `_unet.mlpackage` in `models/unet`. If a +file with that name already exists, it is reused and conversion is skipped. Change +any parameter → new name → new conversion; keep them the same → the cached model +is loaded instantly. + +This is why the recommended workflow is **convert once, then load**: run the +converter a single time, then in day-to-day use load the `.mlpackage` with the +**Load Core ML UNet** node. (You can also leave the converter node in the graph; +it short-circuits to the cached file.) + +> [!NOTE] +> The converter relies on the filename to decide whether to reconvert. If you +> rename the `.mlpackage`, it will be converted again. You can otherwise rename +> it freely if the auto-generated name is too long. + +## Quantization + +Both converter nodes accept an optional `quantize_nbits` dropdown that runs +k-means weight palettization (`coremltools.optimize.coreml.palettize_weights`) on +the UNet before saving. + +Values: `none` (default — no quantization, identical output and filename to +before), `8`, `6`, `4`. The number is appended to the `.mlpackage` stem as +`_q`, so quantized and unquantized variants coexist on disk and in cache. + +### SD1.5 1×512×512 SPLIT_EINSUM tradeoffs (M2 Pro, ANE) + +Measured with 20 UNet forward passes at a fixed seed: + +| nbits | size (MB) | size vs none | fwd median (ms) | PSNR vs `none` (dB) | +|---|---:|---:|---:|---:| +| none | 1641 | 1.000 | 197.1 | — | +| 8 | 822 | 0.501 | 186.6 | 53.5 | +| 6 | 617 | 0.376 | 183.0 | 40.2 | +| 4 | 412 | 0.251 | 179.8 | 27.5 | + +PSNR here is computed on the raw `noise_pred` output of a single UNet forward at a +fixed seed, not on the final decoded image — it isolates quantization drift from +sampler/VAE noise. Final-image PSNR is comfortably higher (the sampler averages +over many steps). + +### Recommended settings per chip / RAM + +- **8 GB (M1/M2 base):** `nbits=4`. ~4× smaller, still loads, 27 dB is visually + identical at SD1.5 sizes. +- **16 GB (M1/M2/M3 Pro):** `nbits=6` — the sweet spot, ~2.7× smaller, 40 dB, no + perceptible quality drop. +- **32 GB+ (Max / Ultra):** `nbits=8` for a safety margin, or `none` for + bit-identical output (golden testing). + +The default stays `none`, so existing workflows produce byte-for-byte identical +output. + +## Where conversion lives + +The conversion engine was extracted into the standalone +[coreml-diffusion](https://github.com/aszc-dev/coreml-diffusion) PyPI package. +The nodes in this suite resolve ComfyUI paths and call into it; node names, +inputs, and outputs are unchanged, so the split has effectively no user-facing +impact beyond `pip install` pulling one more dependency. + +One detail: the LCM converter still imports `diffusers` directly (in +`coreml_suite/lcm/converter.py`) to download the hardcoded LCM model from Hugging +Face. This is an internal note, not something you need to act on. diff --git a/docs/faq.md b/docs/faq.md new file mode 100644 index 0000000..74c1505 --- /dev/null +++ b/docs/faq.md @@ -0,0 +1,77 @@ +# FAQ + +## What's the difference between ANE, GPU, and MPS, and which do I pick? + +ANE is the Neural Engine (Core ML only), GPU is the Metal GPU (Core ML or +PyTorch), MPS is PyTorch's GPU backend. This suite uses **Core ML compute units +only** and never touches MPS. Short answer: SD1.5 at 512×512 → convert +`SPLIT_EINSUM`, load `CPU_AND_NE`; larger sizes or SDXL → convert `ORIGINAL`, +load `CPU_AND_GPU`. Full reasoning: [hardware](hardware.md). + +## Do I still need `PYTORCH_ENABLE_MPS_FALLBACK=1`? + +Not for these nodes — Core ML inference doesn't use PyTorch MPS. It may still +matter for other parts of your ComfyUI graph, but it has no effect on Core ML +sampling. + +## Why is my Core ML SDXL workflow no faster than the default nodes? + +Because **SDXL can't run on the ANE** — the speedup comes from the Neural Engine, +and SDXL falls back to the GPU, running at roughly MPS-equivalent speed. This is a +known limitation, not a misconfiguration. The ANE benefit is real for SD1.5. See +[limitations](limitations.md). + +## Where do I get Core ML models? + +You convert them yourself — that's the only supported path. See +[conversion](conversion.md). Downloaded Core ML models (e.g. coreml-community) use +different dimensions/metadata and are not supported. + +## Is conversion run every time I queue, or once? + +Once. Parameters are encoded in the output filename, so an already-converted model +is reused and conversion is skipped. Convert once, then load the `.mlpackage`. See +[conversion → caching](conversion.md#one-time-conversion-and-name-based-caching). + +## Does a converted model produce the same output as the original? + +With the default `quantize_nbits = none`, the converted UNet output matches the +source within numerical rounding (the golden test in `tests/m2/test_golden_image.py` +gates on PSNR ≥ 20 dB on the decoded image). Quantization (`8`/`6`/`4`) introduces +measured, bounded drift — see the [PSNR table](conversion.md#quantization). For +bit-identical output, keep `none`. + +## Are `.mlpackage` models safe to use? + +`.mlpackage` is a declarative Core ML model format — it carries weights and a +compute graph, not arbitrary executable code or Python pickle, so its safety +profile is comparable to `safetensors`. In practice this matters little here, +since the only supported models are ones you convert locally from your own +checkpoints. + +## Are LoRAs reliable? + +Partially. Some LoRAs convert cleanly; others produce poor or broken output — +there's no firm rule, so test per-LoRA. LoRA weights and `strength_model` are +baked in at conversion and can't be changed afterward; for some LCM-LoRA cases the +[Core ML Adapter](nodes.md#core-ml-adapter-experimental-coremlmodeladapter) path +is more reliable. Treat LoRA support as experimental. See +[troubleshooting](troubleshooting.md). + +## Does the experimental Adapter cost performance vs the Core ML Sampler? + +Yes, a little. The Adapter wraps the model in a ComfyUI `ModelPatcher` so standard +samplers work, which adds per-step interface overhead the native Core ML Sampler +avoids. Use the native sampler unless you specifically need a `MODEL` (e.g. +`ModelSamplingDiscrete` for LCM LoRAs). + +## Which Python versions work? + +Both 3.11 and 3.12. Older 3.12 install failures came from the now-removed +`ml-stable-diffusion` build, not from this suite. + +## Long prompts crash my workflow + +Core ML has a hard **77-token** prompt limit and doesn't auto-chunk long prompts. +Split the prompt across multiple CLIP Text Encode nodes and merge with Conditioning +(Combine). See [troubleshooting](troubleshooting.md). diff --git a/docs/hardware.md b/docs/hardware.md new file mode 100644 index 0000000..9ffb917 --- /dev/null +++ b/docs/hardware.md @@ -0,0 +1,95 @@ +# Hardware & Compute Units + +This page explains how the suite maps to Apple Silicon hardware, the difference +between ANE, GPU, and MPS, and how to choose a compute unit and attention +implementation. + +## ANE vs GPU vs MPS + +Three terms get conflated: + +- **ANE (Apple Neural Engine)** — a dedicated ML accelerator on Apple Silicon. + Only Core ML can target it; PyTorch cannot. This is the whole reason the suite + exists. +- **GPU** — the Metal GPU. Reachable both by Core ML (as a compute unit) and by + PyTorch (via MPS). +- **MPS (Metal Performance Shaders)** — PyTorch's GPU backend on macOS. This is + the path standard ComfyUI nodes use. + +**This suite uses Core ML compute units only — it never runs the UNet through +PyTorch/MPS.** Consequently `PYTORCH_ENABLE_MPS_FALLBACK` has no effect on these +nodes. It may still matter for the rest of your ComfyUI graph (CLIP, VAE, +samplers on non-Core ML models), but not for Core ML inference itself. + +Rough performance picture (SD1.5, maintainer- and user-reported): + +- ANE is meaningfully faster than MPS — on the order of **50–100%** for SD1.5. +- Core ML on the GPU is only marginally faster than PyTorch/MPS. + +So the speedup comes from the Neural Engine, which means it depends on being able +to actually run on the ANE (see [attention implementations](#attention-implementations) +and the [SDXL caveat](#sdxl-and-the-ane)). + +## Compute units + +The **compute unit** is set on the loader/converter node and tells Core ML which +hardware to use. It is applied when the model is loaded +(`coreml_suite/coreml_model.py:22`), not during conversion. + +| Value | Hardware | Best paired with | +|---|---|---| +| `CPU_AND_NE` (default) | CPU + Neural Engine | `SPLIT_EINSUM` / `SPLIT_EINSUM_V2` | +| `CPU_AND_GPU` | CPU + Metal GPU | `ORIGINAL` | +| `CPU_ONLY` | CPU only | fallback / debugging | +| `ALL` | all available hardware | rarely optimal — see below | + +Notes: + +- Every option includes the CPU; there is no GPU-and-ANE-without-CPU combination. +- `NE` in `CPU_AND_NE` is the Neural Engine (Apple's enum spells it `NE`, not + `ANE`). +- **`CPU_AND_NE` is often faster than `ALL`.** Letting Core ML use everything can + be *slower* on non-Max chips, where memory bandwidth is the bottleneck. Try + `CPU_AND_NE` first for SD1.5. + +## Attention implementations + +Chosen at conversion time on the **Convert Checkpoint to Core ML** node. It +decides whether the model can run on the ANE: + +- **`SPLIT_EINSUM`** — ANE-friendly attention. Use for the Neural Engine. +- **`SPLIT_EINSUM_V2`** — a variant; in practice ≈ `SPLIT_EINSUM` for most users. +- **`ORIGINAL`** — standard attention. Runs on the GPU, not the ANE. + +The implementation and the compute unit must agree: a `SPLIT_EINSUM` model wants +`CPU_AND_NE`; an `ORIGINAL` model wants `CPU_AND_GPU`. + +## Which should I pick? + +| Scenario | Attention | Compute unit | +|---|---|---| +| SD1.5 at 512×512 | `SPLIT_EINSUM` | `CPU_AND_NE` | +| SD1.5 at larger sizes (e.g. 768) | `ORIGINAL` | `CPU_AND_GPU` | +| SDXL / SDXL Turbo | `ORIGINAL` | `CPU_AND_GPU` | + +### Resolution crossover + +ANE shines at small latents; the GPU scales better as resolution grows. In user +benchmarks: + +- At **512×512**, ANE + `SPLIT_EINSUM` wins by roughly **10%** over the GPU path. +- At **768×512**, GPU + `ORIGINAL` pulls ahead by roughly **10%**, and the larger + image is about 2× slower overall. + +If you mostly work at 512×512, convert with `SPLIT_EINSUM` and load on +`CPU_AND_NE`. If you routinely go larger, an `ORIGINAL` + GPU model may be +faster. + +### SDXL and the ANE + +SDXL (and SDXL Turbo) **cannot run on the ANE** — the dual-text-encoder UNet +exceeds what the Neural Engine path supports. SDXL therefore runs at roughly +MPS-equivalent speed with no ANE speedup. If a Core ML SDXL workflow feels no +faster than the standard nodes, this is why. Convert SDXL with `ORIGINAL` and +load with `CPU_AND_GPU` or `CPU_ONLY`. See +[limitations](limitations.md) for the full picture. diff --git a/docs/limitations.md b/docs/limitations.md new file mode 100644 index 0000000..11fc80e --- /dev/null +++ b/docs/limitations.md @@ -0,0 +1,53 @@ +# Limitations & Support Matrix + +## Support matrix + +| Feature | Status | Notes | +|---|---|---| +| SD1.5 | ✅ Full | ANE via `SPLIT_EINSUM`; the primary, fastest path | +| SDXL / SDXL Turbo | ⚠️ Partial | GPU only (no ANE), no speedup; possible quality loss vs source. Don't run Turbo at 1024² | +| SD2.1 | ❌ Unsupported | | +| Inpainting checkpoints (9-channel) | ❌ Unsupported | | +| ControlNet | ✅ Supported | Convert the checkpoint with `controlnet_support = True` | +| LoRA | ⚠️ Experimental | Inconsistent per-LoRA; baked at conversion, immutable afterward | +| LCM | ⚠️ Experimental | Hardcoded to LCM Dreamshaper v7 | +| SVD | ❌ Not supported | | +| AnimateDiff | ❌ Not supported | Motion modules need pre-conversion injection; not feasible today | +| IPAdapter | ❌ Not supported | Needs a real `MODEL` the Core ML wrapper can't provide | +| Core ML Adapter | ⚠️ Experimental | Works for many nodes; fails for merges/IPAdapter/etc. | + +## Fixed input/output shapes + +A Core ML model is converted for one specific resolution and batch size. To work +at a different size, re-convert with the new width/height (conversion is cheap and +cached by name). This is also why detailers and latent-upscale workflows that +rescale mid-graph break — see [troubleshooting](troubleshooting.md). + +There is experimental support for flexible shapes via +[EnumeratedShapes](https://apple.github.io/coremltools/docs-guides/source/flexible-inputs.html#select-from-predetermined-shapes), +but it is **much slower** — user benchmarks show roughly **5×** the per-iteration +time on every run, not just the first. Fixed-shape models per resolution are the +practical choice. + +## SDXL on the Neural Engine + +SDXL and SDXL Turbo cannot run on the ANE — the dual-text-encoder UNet exceeds the +supported Neural Engine path. They run on the GPU at roughly MPS-equivalent speed, +so Core ML offers no speed advantage for SDXL, and converted output may look +degraded versus the safetensors original (an upstream conversion artifact). Use +`ORIGINAL` + `CPU_AND_GPU`. See [hardware](hardware.md). + +## Experimental Core ML Adapter + +The Adapter wraps a Core ML model to look like a standard ComfyUI `MODEL`, which +covers many standard and custom nodes. But it can't fully emulate a real model: +operations that need genuine `MODEL` internals — model merges, IPAdapter, some +LoRA flows, detailers without the size hook — generally won't work, and the model's +fixed input shapes aren't validated, so mismatches error at runtime. Prefer the +native Core ML Sampler when you don't need the `MODEL` type. + +## Prompt length + +Core ML enforces a hard 77-token prompt limit with no auto-chunking. Split long +prompts across multiple CLIP Text Encode nodes and merge with Conditioning +(Combine). diff --git a/docs/nodes.md b/docs/nodes.md new file mode 100644 index 0000000..f046cbf --- /dev/null +++ b/docs/nodes.md @@ -0,0 +1,173 @@ +# Node Reference + +All nodes live in the **Core ML Suite** category. Right-click the canvas → +**Add Node → Core ML Suite**, or double-click and search. + +| Display name | Class | Purpose | +|---|---|---| +| Load Core ML UNet | `CoreMLUNetLoader` | Load a converted `.mlpackage` | +| Core ML Sampler | `CoreMLSampler` | Sample (KSampler-style) | +| Core ML Sampler (Advanced) | `CoreMLSamplerAdvanced` | Sample (KSamplerAdvanced-style) | +| Core ML Adapter (Experimental) | `CoreMLModelAdapter` | Wrap as a standard `MODEL` | +| Load LoRA to use with Core ML | `Core ML LoRA Loader` | Bake LoRA(s) at conversion | +| Convert Checkpoint to Core ML | `Core ML Converter` | Convert a checkpoint | +| Convert LCM to Core ML | `Core ML LCM Converter` | Convert LCM Dreamshaper v7 | + +--- + +## Load Core ML UNet (`CoreMLUNetLoader`) + +![Load Core ML UNet](../assets/unet_loader.png?raw=true) + +Loads a converted `.mlpackage` from `models/unet` and outputs a `coreml_model` +for the samplers. Only `.mlpackage` files are listed — this suite no longer uses +`.mlmodelc`. + +- **Inputs** + - `coreml_name` — the `.mlpackage` to load from `models/unet`. + - `compute_unit` — hardware to run on: `CPU_AND_NE` (default), `CPU_AND_GPU`, + `CPU_ONLY`, `ALL`. See [hardware](hardware.md). +- **Output** + - `coreml_model` — for the Core ML Sampler or Adapter. + +--- + +## Core ML Sampler (`CoreMLSampler`) + +![Core ML Sampler](../assets/sampler.png?raw=true) + +Generates a latent from a Core ML model. Behaves like the standard KSampler and +outputs a `LATENT` you can decode or feed downstream. + +- **Inputs** + - `coreml_model` — output of the loader or a converter. + - `latent_image` *(optional)* — must match the model's input size. If omitted, + a suitable empty latent is created. Provide one for img2img. + - `negative` *(optional)* — required for normal models; optional for LCM. + - Remaining inputs (`seed`, `steps`, `cfg`, `sampler_name`, `scheduler`, + `positive`, `denoise`) match the KSampler. +- **Output** + - `LATENT` — decode with a VAE Decode, or use downstream. + +--- + +## Core ML Sampler (Advanced) (`CoreMLSamplerAdvanced`) + +The KSamplerAdvanced counterpart of the Core ML Sampler — same Core ML input, +plus the advanced sampling controls. Use it for partial denoising, fixed noise, +and multi-stage (e.g. SDXL base → refiner) workflows. + +- **Inputs** + - `coreml_model` — output of the loader or a converter. + - `add_noise`, `noise_seed`, `start_at_step`, `end_at_step`, + `return_with_leftover_noise` — as in KSamplerAdvanced. + - `steps`, `cfg`, `sampler_name`, `scheduler`, `positive` — as usual. + - `latent_image` *(optional)*, `negative` *(optional, required for non-LCM)*. +- **Output** + - `LATENT`. + +--- + +## Core ML Adapter (Experimental) (`CoreMLModelAdapter`) + +![Core ML Adapter](../assets/adapter.png?raw=true) + +Wraps a Core ML model so it presents as a standard ComfyUI `MODEL`, letting you +feed it to the normal KSampler and many other nodes (e.g. `ModelSamplingDiscrete` +for LCM LoRAs). + +- **Input** + - `coreml_model`. +- **Output** + - `MODEL` — a Core ML model wrapped as a ComfyUI model. + +> [!NOTE] +> Experimental. The wrapper presents a `MODEL` interface but cannot fully +> emulate one — model merges, IPAdapter, and similar advanced uses generally +> won't work, and the model's fixed input shapes are not validated, so mismatched +> inputs error at runtime. The native Core ML Sampler is faster when you don't +> need the `MODEL` type. See the [FAQ](faq.md) and [limitations](limitations.md). + +--- + +## Load LoRA to use with Core ML (`Core ML LoRA Loader`) + +![LoRA Loader](../assets/lora_loader.png?raw=true) + +Collects LoRA name + `strength_model` to bake into the model at conversion, and +applies the LoRA to CLIP (which is not part of the Core ML path). Chain multiple +loaders for multiple LoRAs. + +Because a converted model is immutable, the baked weights and `strength_model` +**cannot** be changed afterward — changing them means re-converting. `strength_clip` +only affects CLIP and can be changed freely. After conversion, when loading with +`CoreMLUNetLoader`, apply the same LoRAs to CLIP manually (see +[workflows](workflows.md)). + +- **Inputs** + - `lora_name`, `strength_model`, `strength_clip`. + - `clip` — from `CLIPLoader` / `CheckpointLoaderSimple` or another LoRA loader. + - `lora_params` *(optional)* — chain from another LoRA loader. +- **Outputs** + - `CLIP` — with the LoRA applied. + - `lora_params` — pass to the converter or the next LoRA loader. + +> [!NOTE] +> LoRA support is experimental and inconsistent — some LoRAs convert cleanly, +> others produce poor results. Test per-LoRA. See [troubleshooting](troubleshooting.md). + +--- + +## Convert Checkpoint to Core ML (`Core ML Converter`) + +![Checkpoint Converter](../assets/checkpoint_converter.png?raw=true) + +Converts a SD1.5- or SDXL-based checkpoint from `models/checkpoints` to a Core ML +`.mlpackage` in `models/unet`. The conversion parameters are encoded in the +output name, so an already-converted model is reused instead of re-converted. +See [conversion](conversion.md) for details. + +- **Inputs** + - `ckpt_name` — checkpoint in `models/checkpoints`. + - `model_version` — `SD15` or `SDXL` (list is discovered from `coreml-diffusion`). + - `height`, `width` — target image size; any positive multiple of 8 (default + 512). The model's input size is fixed at these values. + - `batch_size` — default 1; raise to convert a batch-capable model. + - `attention_implementation` — `SPLIT_EINSUM` / `SPLIT_EINSUM_V2` (ANE) or + `ORIGINAL` (GPU). See [hardware](hardware.md). + - `compute_unit` — used only when loading the result; does not affect + conversion. + - `controlnet_support` — set `True` to make the model usable with ControlNet + (default `False`). + - `quantize_nbits` *(optional)* — `none` (default), `8`, `6`, `4`. See + [conversion → quantization](conversion.md#quantization). + - `lora_params` *(optional)* — from the LoRA loader, to bake LoRAs in. +- **Output** + - `coreml_model`. + +> [!NOTE] +> Some checkpoints need a custom config `.yaml`. Place it in `models/configs` +> named like the checkpoint (e.g. `juggernaut.safetensors` → +> `juggernaut.yaml`); it is loaded automatically during conversion. + +--- + +## Convert LCM to Core ML (`Core ML LCM Converter`) + +![LCM Converter](../assets/lcm_converter.png?raw=true) + +Converts [SimianLuo/LCM_Dreamshaper_v7](https://huggingface.co/SimianLuo/LCM_Dreamshaper_v7) +to Core ML in `models/unet`. As with the checkpoint converter, the parameters are +encoded in the name and an existing model is reused. + +- **Inputs** + - `height`, `width` — 512–768, multiple of 8 (default 512). + - `batch_size` — default 1. + - `compute_unit` — used only when loading. + - `controlnet_support` — default `False`. +- **Output** + - `coreml_model`. + +> [!NOTE] +> When sampling an LCM model, set `sampler_name` to `lcm` and `scheduler` to +> `sgm_uniform`. Conversion can take a while. diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md new file mode 100644 index 0000000..bae081d --- /dev/null +++ b/docs/troubleshooting.md @@ -0,0 +1,86 @@ +# Troubleshooting + +## `Expected shape … got …` / latent size mismatch + +The most common error. A Core ML model has **fixed** input dimensions — a model +converted for 512×512 expects a 64×64 latent and rejects any other size (batch +size is handled and doesn't matter; only width/height are fixed). + +**Fix:** set your Empty Latent (or upstream latent) to exactly the resolution the +model was converted for, or re-convert at the size you want. + +## Old `.mlmodelc` model, or `metadata.json` not found + +This suite no longer produces or loads `.mlmodelc`; the loader lists `.mlpackage` +only. Models from an older version (or downloaded community models) with a +`.mlmodelc` structure won't load. + +**Fix:** re-convert the checkpoint with **Convert Checkpoint to Core ML**. No +Xcode or `coremlcompiler` is required — that dependency was removed. + +## Prompt too long (`Expected size 154 but got 77`, or a crash) + +Core ML enforces a hard **77-token** prompt limit and does not auto-chunk like +A1111/ComfyUI. + +**Fix:** split the prompt across multiple CLIP Text Encode nodes and merge them +with **Conditioning (Combine)**. + +## `cannot import name 'ModelSamplingDiscreteLCM'` + +A ComfyUI refactor renamed this symbol. + +**Fix:** update the suite (fixed in PR #29) and re-run +`pip install -r requirements.txt`. + +## LoRA loader `ImportError` + +`peft` became a required dependency. + +**Fix:** `pip install -r requirements.txt`. This recurs after ComfyUI-Manager +updates if requirements aren't reinstalled. + +## ControlNet has no effect + +ControlNet support is baked at conversion. If the checkpoint was converted with +`controlnet_support = False`, ControlNet does nothing. + +**Fix:** re-convert with `controlnet_support = True`. The ControlNet model itself +needs no conversion, and `.fp16.safetensors` vs `.safetensors` makes no +difference. + +## LoRAs produce garbage + +LoRA support is inconsistent — some work, some don't, with no firm rule. Test +per-LoRA. For some LCM-LoRA setups, routing through the +[Core ML Adapter](nodes.md#core-ml-adapter-experimental-coremlmodeladapter) is +more reliable than the basic loader path. Remember weights are baked at conversion +and can't be changed afterward. + +## FaceDetailer / detailers error on size + +Detailers rescale latents internally (e.g. 512 → 1024), which breaks the model's +fixed input shape. + +**Fix:** use the `CoreMLDetailerHookProvider` node to pin the detailer's internal +size to the model's converted resolution. Note it only offers preset sizes, so +non-standard resolutions may not be selectable. + +## Inpainting checkpoint errors (`tensor size 9 vs 4`) + +SD1.5 inpainting checkpoints use a 9-channel input and are **not supported**. This +error is expected, not a bug. + +## Errors mentioning `python_coreml_stable_diffusion` or `ml-stable-diffusion` + +You're on a stale install. That dependency was removed; old install scripts tried +`pip install git+…/ml-stable-diffusion.git`, which fails on modern Python. + +**Fix:** reinstall the current suite (`pip install -r requirements.txt`, which +pulls `coreml-diffusion` from PyPI). + +## `all input tensors must be on the same device (mps:0 and cpu)` / ControlNet residual shape `(2,…) vs (1,…)` + +Old bugs that have been fixed. + +**Fix:** update to the latest version. diff --git a/docs/workflows.md b/docs/workflows.md new file mode 100644 index 0000000..5c81032 --- /dev/null +++ b/docs/workflows.md @@ -0,0 +1,100 @@ +# Example Workflows + +> [!NOTE] +> The models referenced are examples — substitute your own. Every workflow +> starts from a checkpoint you convert yourself (see [conversion](conversion.md)); +> there is no Core ML model to download. + +## Basic txt2img + +Convert a SD1.5 checkpoint, then sample from it. CLIP and VAE come from standard +ComfyUI nodes — either loaded separately or pulled from the checkpoint. + +1. Place a SD1.5 checkpoint in `models/checkpoints` (e.g. + [v1-5-pruned-emaonly](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/v1-5-pruned-emaonly.safetensors)). +2. **Convert Checkpoint to Core ML** → queue once → a `.mlpackage` lands in + `models/unet`. +3. **Load Core ML UNet** (or wire the converter output straight in) → + **Core ML Sampler** → **VAE Decode**. + +**CLIP and VAE from the checkpoint:** + +![Core ML UNet + checkpoint](../assets/unet+sampler+checkpoint.png?raw=true) + +**CLIP and VAE loaded separately** — use any SD1.5-compatible +[CLIP](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/text_encoder/model.safetensors) +and [VAE](https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5/resolve/main/vae/diffusion_pytorch_model.safetensors), +placed in `models/clip` and `models/vae`: + +![Core ML UNet + CLIP + VAE](../assets/unet+sampler+clip+vae.png?raw=true) + +## ControlNet + +Convert the checkpoint with `controlnet_support = True`, then wire a standard +ComfyUI ControlNet. The ControlNet model itself needs no conversion. Place it in +`models/controlnet` (e.g. +[control_v11p_sd15_scribble](https://huggingface.co/lllyasviel/control_v11p_sd15_scribble/blob/main/diffusion_pytorch_model.fp16.safetensors)). + +![Core ML UNet + ControlNet](../assets/unet+sampler+controlnet.png?raw=true) + +## Checkpoint conversion + +The minimal conversion graph. See +[Convert Checkpoint to Core ML](nodes.md#convert-checkpoint-to-core-ml-core-ml-converter). + +![Checkpoint converter](../assets/basic_conversion.png?raw=true) + +## Conversion with LoRA + +Bake LoRA(s) into the model at conversion. Read the +[LoRA caveats](nodes.md#load-lora-to-use-with-core-ml-core-ml-lora-loader) first +— baked weights are immutable, and support is inconsistent per-LoRA. + +![Checkpoint converter + LoRA](../assets/conversion+lora.png?raw=true) + +## LCM LoRA conversion + +Chain multiple LoRA loaders to use several LoRAs with one model. + +> [!IMPORTANT] +> Here the model goes through the **Core ML Adapter** and `ModelSamplingDiscrete` +> into the standard ComfyUI KSampler (not the Core ML Sampler). +> `ModelSamplingDiscrete` is required to sample LCM LoRAs correctly. + +![Multiple LoRAs](../assets/conversion+lcm_lora.png?raw=true) + +## Loading a model with baked LoRAs + +Load a model that already has LoRAs baked in. CLIP must be loaded separately and +passed through the same LoRA nodes used at conversion. Since `lora_name` and +`strength_model` are baked in, they need not be passed to the loader. + +> [!IMPORTANT] +> As above, the model goes through the Core ML Adapter + `ModelSamplingDiscrete` +> into the standard KSampler. + +![Loader + LoRA](../assets/loader+lcm_lora.png?raw=true) + +## LCM conversion with ControlNet + +Convert [LCM_Dreamshaper_v7](https://huggingface.co/SimianLuo/LCM_Dreamshaper_v7) +with the LCM converter, then use it with or without ControlNet. + +![LCM + ControlNet](../assets/lcm+controlnet.png?raw=true) + +## SDXL Base + Refiner + +A basic SDXL graph. Add LoRAs and ControlNets as in the SD1.5 examples; the +refiner step is optional. + +Models: +[base + text encoders](https://huggingface.co/stabilityai/stable-diffusion-xl-base-1.0), +[refiner](https://huggingface.co/stabilityai/stable-diffusion-xl-refiner-1.0), +[VAE](https://huggingface.co/stabilityai/sdxl-vae). + +> [!IMPORTANT] +> SDXL does not run on the ANE. Convert with `ORIGINAL` and load with +> `CPU_AND_GPU` (or `CPU_ONLY`). If loading hangs on `CPU_AND_NE`, that is the +> cause. See [limitations](limitations.md). + +![SDXL](../assets/sdxl_conversion.png?raw=true) diff --git a/seam.md b/seam.md deleted file mode 100644 index 5d2d9e8..0000000 --- a/seam.md +++ /dev/null @@ -1,208 +0,0 @@ -# Conversion Extraction — Seam Inventory (`docs/extraction/seam.md`) - -> **Gate E0 deliverable.** Symbol-by-symbol cut line between the future `coreml_diffusion` -> package (CONVERSION) and what stays in `coreml_suite` (the ComfyUI side). -> -> **Confidence legend:** -> - ✅ **verified** — read directly from the current source in this repo. -> - 🔍 **confirm** — inferred / partially seen; Claude Code must `grep`-verify before acting. -> -> **Cut rule:** a symbol goes to `coreml_diffusion` iff it participates in producing the `.mlpackage` -> artifact AND can be made free of `comfy` / `folder_paths` / `comfy_extras`. The runtime -> *loader* that **runs** a compiled model stays in the suite. - ---- - -## 1. File-level map - -| File | Side | Status | Note | -|---|---|---|---| -| `coreml_suite/model_version.py` | **coreml_diffusion** | ✅ | Already `Enum`-only, zero comfy. Becomes pkg source of truth. | -| `coreml_suite/attention.py` | **coreml_diffusion** | ✅ | `ATTENTION_IMPLEMENTATIONS` tuple; pure constant. | -| `coreml_suite/core/naming.py` | **coreml_diffusion** | ✅ | `compose_out_name` = cache-key contract. Move (not copy). | -| `coreml_suite/converter.py` | **coreml_diffusion** (mostly) | ✅ | Main conversion. One symbol stays-adjacent: `get_out_path` (folder_paths) is replaced by injected `out_path`. | -| `coreml_suite/conversion/attention.py` | **coreml_diffusion** | ✅ | `apply_attention_implementation`. Imports `logging`,`torch` only — no comfy. | -| `coreml_suite/conversion/shapes.py` | **coreml_diffusion** | ✅ | `conv2d_output_shape`. Pure math, no imports. | -| `coreml_suite/conversion/trace.py` | **coreml_diffusion** | ✅ | Imports `types.MethodType`, `diffusers...Transformer2DModel` only — torch/diffusers. | -| `coreml_suite/conversion/unet.py` | **coreml_diffusion** | ✅ | `CoreMLUNetWrapper`. Imports `torch` only — no comfy. | -| `coreml_suite/lcm/converter.py` | **coreml_diffusion** (after dedup) | ✅ | Dup helpers deleted; `MODEL_VERSION` HF-hardcode (L22) → E-LCM. `folder_paths` (L111) + `comfy.model_management` (L54) confirmed present → CUT. | -| `coreml_suite/lcm/unet.py` | **coreml_diffusion** | ✅ | `UNet2DConditionModelLCM(UNet2DConditionModel)`. diffusers-only, no comfy. | -| `coreml_suite/config.py` | **STAYS** | ✅ | Imports `comfy.supported_models_base`/`latent_formats`/`model_detection`. **Inference-side** (`get_model_config`), NOT conversion. | -| `coreml_suite/coreml_model.py` | **STAYS** | ✅ | `CoreMLModel` = runtime loader (runs `.mlpackage`). Desktop/Python inference; not used on iOS. | -| `coreml_suite/nodes.py` | **STAYS** | ✅ | Nodes; will call `coreml_diffusion` + own `folder_paths` path resolution + discovery dropdowns. | -| `coreml_suite/lcm/nodes.py` | **STAYS** | ✅ | `COREML_CONVERT_LCM` node. | -| `coreml_suite/models.py` | **STAYS** | ✅ | Inference: `add_sdxl_model_options`, `is_sdxl`, `get_model_patcher`, `get_latent_image`. | -| `coreml_suite/latents.py` | **STAYS** | ✅ | Inference chunking (MODERNIZATION Phase 3 target, not this spec). | -| `coreml_suite/controlnet.py` | **STAYS** | ✅ | Inference-side controlnet. Distinct from converter `add_cnet_support`. | -| `coreml_suite/lcm/utils.py` | **STAYS** | ✅ | `add_lcm_model_options`, `lcm_patch`, `is_lcm`; imports `comfy_extras`. Inference. | -| `coreml_suite/logger.py` | **both / copy** | ✅ | Trivial. Package gets its own logger; suite keeps its. | - ---- - -## 2. Symbol-level: `coreml_suite/converter.py` (main conversion) - -| Symbol | Side | Status | Cut action | -|---|---|---|---| -| `DEFAULT_TRACE_TIMESTEP`, `TEXT_TOKEN_SEQUENCE_LENGTH` | coreml_diffusion | ✅ | Move as-is (module constants). | -| `get_unet(model_version, ref_unet, attention_implementation)` | coreml_diffusion | ✅ | Move. Uses `conversion.{trace,attention,unet}`. No comfy. | -| `get_encoder_hidden_states_shape(ref_unet, batch_size)` | coreml_diffusion | ✅ | Move. Reads `ref_unet.config.cross_attention_dim`. Pure. | -| `get_coreml_inputs(sample_inputs)` | coreml_diffusion | ✅ | Move. `ct.TensorType` build. | -| `load_coreml_model(out_path)` | coreml_diffusion | ✅ | Move. `ct.models.MLModel(out_path)`. (Dedup target vs LCM copy.) | -| `convert_to_coreml(submodule, ts_module, inputs, names, out_path)` | coreml_diffusion | ✅ | Move. `ct.convert(...)`. (Dedup target vs LCM copy.) | -| `get_sample_input(batch, ehs_shape, sample_shape)` | coreml_diffusion | ✅ | Move. **Merge** with LCM variant (LCM passes extra `scheduler` → optional param). | -| `lcm_inputs(sample_unet_inputs)` | coreml_diffusion | ✅ | Move. Adds `timestep_cond`. | -| `sdxl_inputs(sample_unet_inputs, ref_unet, model_version)` | coreml_diffusion | ✅ | Move. `time_ids`/`text_embeds`/`add_embeds`. | -| `add_cnet_support(sample_shape, ref_unet)` | coreml_diffusion | ✅ | Move. Builds `additional_residual_*` inputs from unet block channels. | -| `convert_unet(ref_unet, model_version, unet_out_path, ...)` | coreml_diffusion | ✅ | Move. Orchestrates trace→convert→**quant (palettize)**→save. Quant travels here (E6). | -| `convert(ckpt_path, model_version, unet_out_path, ...)` | coreml_diffusion | ✅ | Move. **Make kw-only past `ckpt_path,model_version,out_path`** (contract). Validates `attn_impl`. | -| `load_unet(ckpt_path, config_path)` | coreml_diffusion | ✅ | Move. `UNet2DConditionModel.from_single_file`. | -| `get_out_path(submodule_name, model_name)` | **STAYS (node)** | ✅ | Uses `folder_paths.get_folder_paths`. **Delete from converter; node resolves path and passes `out_path` in.** | - -**Apple `python_coreml_stable_diffusion` footprint on this path:** ✅ **none.** Verified by grep: -zero imports in `converter.py` / `conversion/*`. Main path uses `diffusers` + -local `CoreMLUNetWrapper`. (And the runtime `CoreMLModel` is now a local coremltools wrapper too — -see §6 stale-spec note.) - ---- - -## 3. Symbol-level: `coreml_suite/lcm/converter.py` (LCM — dedup + defer) - -| Symbol | Side | Status | Cut action | -|---|---|---|---| -| `load_coreml_model` (LCM copy) | DELETE | ✅ | Duplicate of main. Remove; use `coreml_diffusion.load_coreml_model`. | -| `convert_to_coreml` (LCM copy) | DELETE | ✅ | Duplicate of main. Remove. | -| `get_out_path` (LCM copy, folder_paths) | DELETE | ✅ | Duplicate + comfy. Remove; node injects `out_path`. | -| `get_sample_input(..., scheduler)` (LCM copy) | MERGE → coreml_diffusion | ✅ | Fold `scheduler` into shared `get_sample_input` as optional param. | -| `MODEL_NAME` (= LCM_Dreamshaper) | **E-LCM** | ✅ | HF hardcode. Removing it is the behavior change → E-LCM, not E2. | -| `convert(out_path, sample_size, batch_size, controlnet_support)` (LCM, L190) | coreml_diffusion (via unified) | ✅ | Route through `coreml_diffusion.convert(model_version=LCM, ...)` in E-LCM. | -| `from comfy.model_management import get_torch_device` (L54, in `get_scheduler`) | **CUT** | ✅ | Confirmed present. Inject `device`. | -| module-global attention set at import | n/a | ✅ | **No module global.** Attention already per-call: `get_unets` (L36) calls `apply_attention_implementation(ref_unet, "SPLIT_EINSUM")`. No `ATTENTION_IMPLEMENTATION_IN_EFFECT` anywhere in repo. (Note: LCM hardcodes `"SPLIT_EINSUM"` — pass `attn_impl` through in dedup.) | - ---- - -## 4. Symbol-level: `coreml_suite/core/naming.py` → `coreml_diffusion/naming.py` - -| Symbol | Side | Status | Cut action | -|---|---|---|---| -| `compose_out_name(...)` | coreml_diffusion | ✅ | **Move** (cache-key contract). Node imports from pkg. | -| `lora_names_from_params(...)` | coreml_diffusion | ✅ | Move. | -| `ATTN_SUFFIX` dict | coreml_diffusion | ✅ | Move. | -| `QUANT_NBITS_VALUES` | coreml_diffusion | ✅ | Move; backs `list_quant_modes()`. | -| `tests/unit/test_characterization_out_name.py` | re-point | ✅ | Change import to `coreml_diffusion.naming`. Assertions/values **unchanged**. | - ---- - -## 5. Discovery API + status registry (new in `coreml_diffusion/__init__.py`) - -```python -from enum import Enum - -class Status(Enum): - VERIFIED = "verified" # has a golden anchor + passing [M2-ANE] check - EXPERIMENTAL = "experimental" # convertible, not yet anchored/verified - -# Single source of truth. Suite gates on this, NOT on a hardcoded node list. -# KEY by ModelVersion enum MEMBER (not a bare string) so list_model_versions can -# emit .name — see the .name decision below. Keying by the lowercase .value string -# (as an earlier draft of this block did) returns ["sd15",...], which the node then -# reverses via ModelVersion[...] → KeyError. Do NOT key by .value. -_MODEL_STATUS = { - ModelVersion.SD15: Status.VERIFIED, - ModelVersion.SDXL: Status.VERIFIED, - ModelVersion.SDXL_REFINER: Status.EXPERIMENTAL, # → VERIFIED after a refiner golden anchor - ModelVersion.LCM: Status.EXPERIMENTAL, # → VERIFIED after E-LCM golden anchor -} - -def list_model_versions(include_experimental: bool = False) -> list[str]: - return [v.name for v, s in _MODEL_STATUS.items() # .name → "SD15","SDXL" (see decision) - if s is Status.VERIFIED or (include_experimental and s is Status.EXPERIMENTAL)] - -def list_attention_impls() -> list[str]: # from attention.ATTENTION_IMPLEMENTATIONS - ... -def list_quant_modes() -> list[str]: # from naming.QUANT_NBITS_VALUES - ... - -CONTRACT_VERSION = "1.0" -# Additive-only: adding an id or promoting EXPERIMENTAL→VERIFIED = minor bump (Suite unaffected). -# Removing/renaming an id, or demoting VERIFIED→EXPERIMENTAL = MAJOR bump + migration note. -``` - -**Decision check (`.name` vs `.value`): RESOLVED → `.name`.** ✅ -Verified in current source: -- Node renders `ModelVersion.SD15.name` / `ModelVersion.SDXL.name` → `"SD15"`, `"SDXL"` - (`nodes.py:224-225`). -- Node reverses the dropdown string with `model_version = ModelVersion[model_version]` - (`nodes.py:286`) — i.e. **lookup by NAME**. Feeding it a `.value` (`"sd15"`) raises `KeyError`. -- Enum values are lowercase (`model_version.py`: `SD15="sd15"`, `SDXL="sdxl"`, - `SDXL_REFINER="sdxl_refiner"`, `LCM="lcm"`). -- `compose_out_name` does NOT consume the model_version string (grep of `core/naming.py` empty) — - no coupling there, so no constraint from that side. - -**Decision:** `list_model_versions()` returns `.name` (uppercase). Saved workflows store `"SD15"`, -node already validates them via `ModelVersion[...]`. The `_MODEL_STATUS` block above was corrected -to key by enum member and emit `.name`. **The earlier `v.value` form was a latent bug.** - ---- - -## 6. `python_coreml_stable_diffusion` split (Gate E0 line to fill by grep) - -| Use | Side | Status | -|---|---|---| -| `coreml_model.CoreMLModel` (runs compiled model) | **STAYS** (suite runtime) | ✅ — **local class**, not Apple's | -| `unet.UNet2DConditionModel*` internals | **gone** — `converter.py:319` uses `diffusers.UNet2DConditionModel.from_single_file` | ✅ | -| `AttentionImplementations` enum | gone — local `apply_attention_implementation` + `attention.py` tuple | ✅ | -| `calculate_conv2d_output_shape` | gone — replaced by `conversion/shapes.conv2d_output_shape` | ✅ | - -> ### ⚠️ SPEC IS STALE: `ml-stable-diffusion` is already fully removed -> Commit #58 ("replace apple/ml-stable-diffusion with native diffusers conversion") already did -> the de-Apple work. Verified now: -> - **Zero** `python_coreml_stable_diffusion` runtime imports anywhere in `coreml_suite` (only a -> docstring mention at `core/__init__.py:4`). -> - `coreml_suite/coreml_model.py:8` `CoreMLModel` is a **local** wrapper over -> `coremltools.models.MLModel` (`coreml_model.py:22`) — it does **not** import Apple's class. -> - `ml-stable-diffusion` / `python_coreml_stable_diffusion` appears in **neither** `pyproject.toml` -> **nor** `requirements.txt`. It is not a dependency at all. -> -> **Consequences for the spec (correct these in CONVERTER_EXTRACTION_SPEC.md):** -> - §0.3 premise ("runtime loader = `python_coreml_stable_diffusion.coreml_model.CoreMLModel`, -> stays in suite") is **wrong**: the loader is already the local `coreml_model.CoreMLModel`. The -> "stays in suite" conclusion still holds; the identity does not. -> - **Gate E0 item "ml-stable-diffusion pinned SHA — BLOCKER if unpinned" is MOOT** — there is no -> such dep to pin. Mark it N/A, not BLOCKER. -> - **E4/E5 dependency lists must drop `git+...ml-stable-diffusion@`.** Package runtime deps -> are: `coremltools`, `diffusers`, `peft` (LoRA), `omegaconf` (config), `numpy`, `torch`. Confirm -> `peft`/`omegaconf` actually used before listing (grep at E4). -> - The "keep `python_coreml_stable_diffusion` as a suite dep for the loader" instruction in E5 is -> **void** — coremltools backs the loader. - ---- - -## 7. Pre-flight checklist before E1 (run these greps) - -``` -grep -rn "import comfy" coreml_suite/conversion coreml_suite/converter.py coreml_suite/lcm/converter.py coreml_suite/lcm/unet.py -grep -rn "folder_paths" coreml_suite/converter.py coreml_suite/lcm/converter.py -grep -rn "model_management" coreml_suite/lcm -grep -rn "python_coreml_stable_diffusion" coreml_suite -grep -rn "ATTENTION_IMPLEMENTATION_IN_EFFECT" coreml_suite -grep -rn "SimianLuo\|LCM_Dreamshaper" coreml_suite/lcm -``` -Every 🔍 above resolves to ✅ or a correction once these run. Do not start moving code (E2) -with any 🔍 unresolved on the CONVERSION side. - -**STATUS (run 2026-05-26): all 🔍 resolved.** Summary of what the greps found: -- `conversion/*`, `lcm/unet.py`: comfy-free (torch/diffusers only). ✅ -- `converter.py`: only comfy reach-in is `folder_paths` in `get_out_path` (L91-94) → inject `out_path`. -- `lcm/converter.py`: `folder_paths` (L111-114) + `comfy.model_management.get_torch_device` (L54) - → cut both. Dup helpers (`load_coreml_model`,`convert_to_coreml`,`get_out_path`,`get_sample_input`) - confirmed → dedup E2. `MODEL_VERSION="SimianLuo/LCM_Dreamshaper_v7"` (L22) → E-LCM. -- No attention module-global anywhere (`ATTENTION_IMPLEMENTATION_IN_EFFECT` absent); already per-call. - LCM hardcodes `"SPLIT_EINSUM"` in `get_unets` — thread `attn_impl` through during dedup. -- `.name` vs `.value`: **decided `.name`** (node reverses via `ModelVersion[...]`). §5 corrected. -- `ml-stable-diffusion`: **already gone** (#58). §6 stale-spec note added — fix the spec's E0/E4/E5 - dep + pinning items. - -Two grep blind-spots to note (the checklist above doesn't cover them, but cheap to add): the -`folder_paths` grep only scans the two converter files — also grep `coreml_suite/lcm/utils.py` -(it imports `comfy.model_management` at L3, but it's inference/STAYS, so fine) and confirm no other -`conversion/` file grew a comfy import since.