* docs(extraction): E0 seam inventory + correct stale spec assumptions Resolve all pre-flight greps for the converter-extraction seam: - conversion/* and lcm/unet.py confirmed comfy-free - converter.py: only folder_paths reach-in is get_out_path - lcm/converter.py: folder_paths + comfy.model_management to cut; dup helpers and SimianLuo HF hardcode confirmed - no attention module-global; already per-call Correct two stale assumptions verified against current source: - ml-stable-diffusion is already fully removed (#58); CoreMLModel is a local coremltools wrapper, not Apple's. Drop the dep-pinning blocker and the package/suite dep lines that assumed it. - model_version discovery must emit .name (node reverses via ModelVersion[...]); the .value form in the draft would KeyError on every saved workflow. * feat(extraction): E1 coreml_diffusion package + discovery API Stand up the framework-free coreml_diffusion namespace and freeze its versioned discovery contract. The package re-exports from already comfy-free coreml_suite sources (model_version, attention, core.naming); the conversion implementation moves in E2. - list_model_versions/list_attention_impls/list_quant_modes return today's exact dropdown strings, so wiring the node onto them (E3) changes no value and breaks no saved workflow. - Status/_MODEL_STATUS registry gates VERIFIED vs EXPERIMENTAL in the package, so promoting a model expands the node dropdown with no Suite change (additive-only contract; CONTRACT_VERSION=1.0). - Tier-0 test pins the contract and proves comfy/diffusers/coremltools are not pulled on import. Node untouched; zero behavior change. * refactor(extraction): E2 move conversion mechanics into coreml_diffusion Physically relocate the framework-free conversion code into the package and collapse the duplicated LCM/main helpers, behavior-preserving. - coreml_suite/conversion/ -> coreml_diffusion/conversion/ (attention, shapes, trace, unet) - coreml_suite/core/naming.py -> coreml_diffusion/naming.py (the cache-key contract now lives with the package; tests re-pointed) - coreml_suite/converter.py logic -> coreml_diffusion/convert.py, with convert() made keyword-only past (ckpt_path, model_version, out_path) per the interface contract; out_path is injected (no folder_paths) - dedup: load_coreml_model / convert_to_coreml / get_coreml_inputs / add_cnet_support / get_encoder_hidden_states_shape / inputs-spec now defined once in the package; get_sample_input gains an optional scheduler arg so the LCM path shares it (same keys/order/dtypes) - coreml_suite.{converter,lcm.converter} reduced to comfy-side shims: folder_paths path resolution and the LCM scheduler's comfy.model_management stay here; the package imports neither - __init__ keeps discovery + compose_out_name eager; convert is lazy via __getattr__ so 'import coreml_diffusion' stays Tier-0 pure Nodes untouched (E3 thins them onto the package). Tier-0 (109) and smoke (3, real coremltools conversion) green; [M2-ANE] golden pending a server. * refactor(extraction): E3 thin nodes onto coreml_diffusion + discovery dropdowns The CoreMLConverter node now calls coreml_diffusion directly instead of the coreml_suite.converter shim, and its dropdowns are populated at runtime from the package's discovery API. - INPUT_TYPES dropdowns (model_version / attention_implementation / quantize_nbits) now come from a fail-soft _discover() that calls coreml_diffusion.list_*; a missing/old package falls back to a literal list and logs a warning instead of de-registering the node. Installing a newer coreml_diffusion surfaces new conversion types with no Suite change. - folder_paths path resolution moved inline into the node; the package's convert() takes the output path as an injected positional. - compose_out_name / lora_names_from_params now imported from coreml_diffusion (lazily, inside convert) — no node-side copy. - deleted the dead coreml_suite/converter.py and coreml_suite/core/naming.py shims (no remaining importers). Field names, RETURN_TYPES/NAMES and NODE_*_MAPPINGS unchanged; dropdown values are a superset of the prior literals (additive-only). Tier-0 (109) and smoke (3) green; [M2-ANE] golden re-runs on push. * refactor(extraction): E5 depend on external coreml-diffusion package Conversion code now lives in the standalone coreml-diffusion repo. The Suite deletes its in-tree copy and depends on the package instead. - removed coreml_diffusion/ (whole package), coreml_suite/model_version.py and coreml_suite/attention.py (moved to the package as its source of truth), and the tests that moved with them (discovery, conversion_helpers, out_name; smoke synthetic_unet + split_einsum) - re-pointed ModelVersion imports (config.py, nodes.py, lcm/converter.py) to coreml_diffusion - pyproject: drop the coreml_diffusion package include and the conversion-only deps (peft/omegaconf/transformers, now transitive via coreml-diffusion); add coreml-diffusion as a dependency with a local path source until it is published (switch to git tag/PyPI once the repo exists, so CI can resolve it) Suite Tier-0 green (75); conversion code fully absent from the Suite. The comfy node still imports coreml_diffusion (installed package) for ModelVersion + the discovery dropdowns + convert. * build(extraction): pin coreml-diffusion to git tag v0.1.0 Switch the coreml-diffusion source from a local path to the published git tag so CI can resolve it. Suite Tier-0 green resolving from the tag. * ci(extraction): drop Suite smoke tier (moved to coreml-diffusion) The conversion smoke tests moved to the coreml-diffusion repo, which runs its own Tier 1. The Suite's smoke lane had no tests left (pytest exit 5). The Suite keeps Tier 0 (inference units) and the m2 golden e2e. * chore(release): v2.1.0; wire coreml-diffusion into requirements.txt Minor bump: the conversion path moved to the external coreml-diffusion package (node graph + artifact cache keys unchanged, golden-verified). requirements.txt (used by ComfyUI Manager) now installs coreml-diffusion from the v0.1.0 tag and drops the conversion-only deps now provided transitively.
38 KiB
ComfyUI-CoreMLSuite — Converter Extraction Spec for Claude Code
Companion to
MODERNIZATION_SPEC.md. That spec hardens the repo and (Phase 3) splits the inference math from the framework. This spec splits the conversion path (safetensors → CoreML) out into a standalone,comfy-free, pip-installable package that CoreMLSuite then depends on — and that other projects (incl. on-device iOS tooling) can reuse.Same discipline as the modernization spec: safety-net first, behavior-preserving until told otherwise, one phase = one branch = one PR,
STOP — VALIDATEgate between every phase, golden-latent as the regression anchor.[M2]= needs macOS/Apple Silicon;[M2-ANE]= needs the Neural Engine. Everything else must run on plain Linux/CI.
0. How to work (read first — non-negotiable)
- Behavior-preserving until Phase E6. Phases E1–E5 must not change image output, node
names,
INPUT_TYPESfield names, orNODE_CLASS_MAPPINGSkeys. The node graph is the public contract; saved user-workflow JSON breaks if these change. - The conversion package produces an artifact and stops there. Its job ends at a written
.mlpackage/.mlmodelcon disk. It must NOT importcomfy,folder_paths, orcomfy_extras, and must NOT know ComfyUI'smodels/unetlayout. Paths are inputs. - The runtime loader stays in the suite. The loader is the local
coreml_suite.coreml_model.CoreMLModel— a thin wrapper overcoremltools.models.MLModel(NOT Apple'spython_coreml_stable_diffusion.coreml_model.CoreMLModel, which is no longer used; see #58). It runs a compiled model in Python — a desktop/Python inference concern, not a conversion concern. It is NOT moved into the package. (On iOS the.mlmodelcis loaded natively; the package's output is the deliverable, not a Python runner.) - Decouple in-repo before splitting repos. Phases E1–E4 create the package inside this repo and prove equivalence. The physical second-repo split is Phase E5, only after the golden latent is proven identical. Do not create a second repository before Gate E4 passes.
- Reuse the existing regression anchor. The golden latent / PSNR anchor from
MODERNIZATION_SPEC.mdPhase 2 is the cross-cutting proof for every gate here. If it is not yet captured, capture it first (it is a prerequisite for E2 onward). - No new runtime dependencies without flagging in the gate report (name, why, license, size).
- A failing gate means stop and report, not work around into the next phase.
- Tooling is
uv, not barepip/venv. Every environment/install/lock step uses the project'suvtoolchain:uv venv,uv pip install,uv pip install -e .,uv lock,uv run pytest,uv export/uv pip freezefor baselines. Where this spec says "fresh venv", read "uv venv+uv pip install". Reserveuv pip(notpip) inside that venv too. - The package is the single source of truth for what is possible; the node is a thin, discovery-driven frontend. See the "Interface contract" pillar below — this is the maintainer's hard requirement and it overrides the earlier (now-rescinded) "freeze the dropdown list" instruction.
Interface contract (the maintainer's hard requirement) — read before any phase
Two coupled guarantees must hold once the package is split out:
(A) Updating the converter must NOT require updating CoreMLSuite. This is satisfied by treating the package's public surface as a versioned contract:
convert(...)andcompile_model(...)are keyword-only with defaults for everything past the genuinely-required positionals (ckpt_path,model_version,out_path). New capabilities are added as new keyword args with defaults, so an old Suite's call still validates against a newer package. Never reorder or rename existing parameters.compose_out_name(the.mlpackagefilename = the cache key) moves into the package and is versioned with it. The Suite must not carry its own copy; if the package changes the naming scheme that is a major bump (old cached artifacts stop resolving).
(B) CoreMLSuite must be able to list new conversion types WITHOUT a Suite code change or version bump. Today the node hardcodes its dropdowns:
"model_version": ([ModelVersion.SD15.name, ModelVersion.SDXL.name],), # hand-typed, also INCOMPLETE (no LCM / SDXL_REFINER)
"attention_implementation": (list(ATTENTION_IMPLEMENTATIONS),), # from coreml_suite.attention
"quantize_nbits": (list(QUANT_NBITS_VALUES), {"default": "none"}), # from coreml_suite.core.naming
These are replaced by runtime discovery calls into the package, evaluated inside
INPUT_TYPES (ComfyUI re-evaluates INPUT_TYPES on every plugin load):
import coreml_diffusion
"model_version": (coreml_diffusion.list_model_versions(),),
"attention_implementation": (coreml_diffusion.list_attention_impls(),),
"quantize_nbits": (coreml_diffusion.list_quant_modes(), {"default": "none"}),
Effect: uv pip install -U coreml_diffusion + ComfyUI restart surfaces any newly-added type in the old
plugin's dropdown — no Suite edit, no Suite version bump. This is the requirement.
The cost, stated honestly (accept this trade-off explicitly at Gate E0):
- The Suite becomes a "dumb" frontend; the package is the sole authority on what conversions exist. The Suite can no longer guarantee its saved workflows are valid against arbitrary future package versions.
- Therefore the package's discovery identifiers (
ModelVersionvalues, attn-impl strings, quant modes) are an ADDITIVE-ONLY contract: the package may add identifiers freely (minor bump, no Suite change); removing or renaming an identifier is a breaking change requiring a MAJOR bump and a migration note, because a saved workflow JSON references these strings verbatim. Without this rule, "no version bump" silently becomes "randomly broken workflows." INPUT_TYPESmust fail soft when the package is missing/old: wrap the discovery calls so a missingcoreml_diffusion(or an old one lacking alist_*function) yields a sane fallback list and a logged warning, instead of the node failing to register and disappearing from the menu.
Discovery API the package must expose (stable names):
coreml_diffusion.list_model_versions() -> list[str] # VERIFIED ones only, e.g. ["SD15","SDXL"] today (.name — see seam.md)
coreml_diffusion.list_attention_impls() -> list[str] # ["SPLIT_EINSUM","SPLIT_EINSUM_V2","ORIGINAL"]
coreml_diffusion.list_quant_modes() -> list[str] # ["none","8","6","4"]
coreml_diffusion.CONTRACT_VERSION: str # bump rules above; Suite may log/compare it
These return the display strings already used today, so existing workflows keep validating.
Verification status is a PACKAGE property, not a node hardcode (maintainer's intent).
The Suite wants to expose every model the converter can verifiably convert. Today lcm and
sdxl_refiner are absent from the converter node not because the Suite chooses to hide them, but
because they lack a full golden/PSNR verification. So the gating lives in the package as a status:
from enum import Enum
class Status(Enum):
VERIFIED = "verified" # has a golden anchor + passing [M2-ANE] check
EXPERIMENTAL = "experimental" # convertible but not yet anchored/verified
# internal registry, single source of truth.
# KEY by ModelVersion enum MEMBER so list_* can emit .name. Keying by the lowercase
# .value string returns ["sd15",...], which the node reverses via ModelVersion[...] -> KeyError.
_MODEL_STATUS = {ModelVersion.SD15: Status.VERIFIED, ModelVersion.SDXL: Status.VERIFIED,
ModelVersion.SDXL_REFINER: Status.EXPERIMENTAL, ModelVersion.LCM: Status.EXPERIMENTAL}
def list_model_versions(include_experimental: bool = False) -> list[str]:
return [v.name for v, s in _MODEL_STATUS.items() # .name -> "SD15","SDXL"; node reverses with ModelVersion[...]
if s is Status.VERIFIED or (include_experimental and s is Status.EXPERIMENTAL)]
Consequence: promoting a model to VERIFIED in the package expands the Suite's dropdown with no
Suite change and no Suite bump — exactly the requirement. The act of verification (E-LCM
produces an LCM golden anchor; same later for refiner) is what flips the status. The Suite's
converter node calls list_model_versions() (verified-only); a power-user/CLI path may pass
include_experimental=True. Promotion VERIFIED-from-EXPERIMENTAL is additive (minor bump);
demotion or removal is breaking (major bump + note).
Naming & layout (chosen — frozen at Gate E0)
Distribution name (PyPI): coreml-diffusion. Import name (Python): coreml_diffusion.
(PyPI normalizes -/_; the distribution uses the hyphen, the importable module the underscore.)
Availability checked: both coreml-diffusion and the near variants were free on PyPI at E0.
Why this name (the positioning it encodes): the project's niche is diffusion models on Apple
Neural Engine via CoreML, inside ComfyUI and on-device — not Stable Diffusion specifically.
sd* was rejected because it falsely narrows scope to SD; coreml-diffusion keeps coreml on the
front for discoverability while diffusion honestly states the scope (SD/SDXL/LCM today, Flux and
other diffusion architectures later) without promising arbitrary non-diffusion torch models,
whose tracing/shape/sample-input pipeline differs. The name must not be re-narrowed to SD in
future docs. ANE is the differentiator (documented in the README), but coreml was chosen over
ane in the name for search discoverability per maintainer decision.
Target package layout (framework-free — zero comfy imports):
coreml_diffusion/
__init__.py # public API surface (see "Public API" below)
model_version.py # ModelVersion enum — the SINGLE source of truth, no comfy
attention.py # ATTENTION_IMPLEMENTATIONS tuple (from coreml_suite/attention.py) + apply_attention_implementation
pipeline.py # get_pipeline (from_single_file), get_unet (cml UNet from ref unet)
unet.py # UNet2DConditionModelLCM (moved from coreml_suite/lcm/unet.py)
inputs.py # get_sample_input, lcm_inputs, sdxl_inputs,
# get_encoder_hidden_states_shape, get_coreml_inputs, get_inputs_spec
controlnet.py # add_cnet_support (conversion-side residual SHAPE calc only)
convert.py # convert_unet, convert (orchestration), convert_to_coreml, load_coreml_model
compile.py # compile_coreml_model
quantize.py # (Phase E6 / MODERNIZATION Phase 6 lands here) palettization 4/6/8-bit
cli.py # console entry point: `coreml-diffusion convert ...`
pyproject.toml # standalone packaging (at E5)
What stays in coreml_suite/ (the ComfyUI side, thinned):
nodes.py— still owns name-encoding (out_nameconstruction), path resolution viafolder_paths, the nodeINPUT_TYPES/mappings, and wrapping the result inCoreMLModel.models.py,latents.py,controlnet.py(inference parts),lcm/utils.py,config.py(inference config build) — untouched by this spec except the import-source ofModelVersion.
Public API (the contract coreml_diffusion exposes)
from coreml_diffusion import ModelVersion, convert, compile_model, compose_out_name
from coreml_diffusion import list_model_versions, list_attention_impls, list_quant_modes, CONTRACT_VERSION
# Mirror the CURRENT converter.py signature, made keyword-only past the required positionals
# and with paths/device injected (no folder_paths, no comfy.model_management):
# convert(ckpt_path, model_version, out_path, *,
# batch_size=1, sample_size=(64, 64), controlnet_support=False,
# lora_weights=None, attn_impl=list_attention_impls()[0], config_path=None,
# quantize_nbits="none", device=None) -> None # side effect: writes out_path
# (current convert() returns None and writes via convert_unet → coreml_unet.save; keep that,
# or change to `return out_path` as a deliberate, documented improvement — pick one at E0.)
# compile_model(src_path, out_dir, final_name) -> str # returns compiled .mlmodelc path
Note: convert takes an explicit out_path — no folder_paths. device is injected
(defaults to torch's default device). compose_out_name lives here (cache-key contract) and the
node imports it from the package. The list_* discovery functions back the node's dropdowns.
The import chains to cut (root cause inventory) — REVISED against current code
State note (verified): the code moved on since the original draft. Several chains are already cut. Re-verify each line by
grepbefore acting; do not assume the original draft.
Already done (verify, then skip):
- ✅
converter.pyalready importsfrom coreml_suite.model_version import ModelVersion, andmodel_version.pyis clean (from enum import Enumonly — zero comfy). The old "converter → config → comfy" chain is already broken.config.pystill imports comfy, but it is inference-side (get_model_configviasupported_models_base/latent_formats) — not on the conversion path. Do not treatconfig.pyas a converter dependency. - ✅
converter.pynow usesdiffusers.UNet2DConditionModel.from_single_fileand a localCoreMLUNetWrapper(incoreml_suite/conversion/unet.py) — it is no longer importing the Applepython_coreml_stable_diffusion.unet.UNet2DConditionModel*internals on the main path. Acoreml_suite/conversion/subpackage already exists (attention,shapes,trace,unet). - ✅ Name-encoding already extracted to
coreml_suite/core/naming.py(compose_out_name,lora_names_from_params,ATTN_SUFFIX,QUANT_NBITS_VALUES) with characterization tests (tests/unit/test_characterization_out_name.py). The pure-naming split is done. - ✅ Quantization is already implemented in
converter.py(quantize_nbits, k-meanspalettize_weights) and surfaced as an optional node input. Phase E6 is therefore move, not build (see revised E6).
Still to cut (the real remaining work):
coreml_suite/converter.py::get_out_path→from folder_paths import get_folder_paths. Main converter still reaches into ComfyUI's model dir. Cut:out_pathis an injected arg;folder_pathsresolution moves up into the node (the node already computesout_name).coreml_suite/lcm/converter.py→ still has its ownfrom folder_paths import get_folder_paths(get_out_path) and (per original draft)comfy.model_management. Verify the current LCM file and cut both: injectout_pathanddevice.- Global mutation of the attention impl: confirm where it now lives. Main path appears to route
through
coreml_suite/conversion/attention.apply_attention_implementation(cleaner than the old global), butlcm/converter.pymay still set a module global at import. Ensure the package sets attention per-call, never at import time. - Duplication LCM vs main:
lcm/converter.pystill carries its own copies ofconvert_to_coreml,load_coreml_model,get_out_path,get_sample_input(the LCM variant takes aschedulerarg), and hardcodesSimianLuo/LCM_Dreamshaper_v7. Dedupe into the singlecoreml_diffusionimplementation; the HF-hardcode consolidation is the behavior-changing part → deferred to optional E-LCM, not E1–E5. compose_out_nameownership: currently incoreml_suite/core/naming.pyand called by the node. Per the Interface-contract pillar it must move into the package (it is the cache-key contract) and the node must import it fromcoreml_diffusion, not keep a copy.
Phase E0 — Seam decision & inventory (no code change)
Objective: lock the cut line, the interface contract, and naming so later phases don't drift.
Tasks
- Produce
docs/extraction/seam.md: a table of every symbol inconverter.py,lcm/converter.py,lcm/unet.py, plus the already-extractedconversion/subpackage (attention,shapes,trace,unet) andcore/naming.py, classified CONVERSION → coreml_diffusion vs STAYS (comfy/node). Note which are already framework-free. Confirm the currentDONE (seam.md §6): footprint is ZERO — no runtime imports anywhere; only a docstring mention inpython_coreml_stable_diffusionfootprint.core/__init__.py:4. Main path usesdiffusers+ localCoreMLUNetWrapper; the runtimeCoreMLModel(STAYS in suite) is a local coremltools wrapper, not Apple's. No shape/attn helper comes from Apple (localconversion/shapes.py,conversion/attention.py).- Decide the interface contract concretely (the maintainer's hard requirement):
- Discovery functions
list_model_versions / list_attention_impls / list_quant_modeslive in the package and return today's display strings verbatim. NodeINPUT_TYPEScalls them. ModelVersionvalues, attn-impl strings, quant modes are ADDITIVE-ONLY across package versions; removal/rename = MAJOR bump + migration note. Write this into the package's versioning policy doc now.compose_out_namemoves to the package; node imports it (no copy). Confirm the characterization tests intest_characterization_out_name.pywill be re-pointed, not duplicated.- Resolve the
model_versiondropdown question (maintainer decided): the Suite exposes every model the converter can verifiably convert.lcmandsdxl_refinerare absent today only because they lack a golden/PSNR verification — not because the node hardcodes a short list. Encode this as a status registry in the package (VERIFIEDvsEXPERIMENTAL);list_model_versions()returns VERIFIED-only by default. The converter node calls it plainly. Promoting LCM/refiner to VERIFIED (after E-LCM / a refiner anchor) expands the dropdown with no Suite change. Do NOT add permanent per-node filtering — the gate is verification status, owned by the package.
- Discovery functions
Confirm theN/A — already removed (#58). Verified: zeroml-stable-diffusiongit dep is pinned.python_coreml_stable_diffusionimports in the repo;CoreMLModelis now a local coremltools wrapper; the dep is absent frompyproject.toml/requirements.txt. No SHA to pin.
STOP — VALIDATE (Gate E0)
## Gate E0 report
- seam.md committed: <path>; symbol counts (move / stay / already-framework-free)
- python_coreml_stable_diffusion usage (verified by grep): conversion=<list> runtime=<list>
- Discovery API signatures frozen: list_model_versions (verified-only) / list_attention_impls / list_quant_modes
- Status registry decided: sd15+sdxl=VERIFIED, lcm+sdxl_refiner=EXPERIMENTAL (gated, not hidden)
- Additive-only contract policy doc written (incl. promotion=minor, demotion/removal=major): <path>
- model_version dropdown: expose all (incl. LCM/REFINER) / filtered per node — DECISION: <...>
- compose_out_name move-not-copy confirmed; tests re-point plan: <...>
- LCM consolidation deferred to optional E-LCM: YES/NO
- ml-stable-diffusion: N/A — already removed (#58), not a dependency (was: pin-or-BLOCKER)
- Package name in-repo: coreml_diffusion (final PyPI name deferred to E5)
Phase E1 — Establish coreml_diffusion package + discovery API (mostly verification)
Objective: stand up the package namespace and the discovery surface. Much of the comfy-chain cut is already done — this phase mostly verifies that and adds the discovery functions.
Tasks
- Verify (don't redo):
coreml_suite/model_version.pyis already clean (Enumonly). Confirmimport coreml_suite.model_versionworks with no comfy (uv run python -c "..."in a comfy-freeuv venv). If true, E1's original "extract ModelVersion" task is already satisfied. - Create the
coreml_diffusion/package skeleton with__init__.pyexporting the discovery API backed by the existing sources of truth for now (re-exportModelVersion, theATTENTION_IMPLEMENTATIONStuple, andQUANT_NBITS_VALUES) so values are byte-identical:(At this stagedef list_model_versions(): return [v.name for v in ModelVersion] # .name -> "SD15" (node reverses via ModelVersion[...]; .value KeyErrors) def list_attention_impls(): return list(ATTENTION_IMPLEMENTATIONS) def list_quant_modes(): return list(QUANT_NBITS_VALUES) CONTRACT_VERSION = "1.0"coreml_diffusionmay live inside the repo and import fromcoreml_suite.*; the physical move of implementation happens in E2. The point of E1 is to freeze the contract.) - Decided (
.name): the node rendersModelVersion.SD15.name("SD15") and reverses the dropdown string viaModelVersion[model_version](name lookup,nodes.py:286). Discovery API therefore returns.name;.value("sd15") wouldKeyError. Recorded inseam.md§5.
Acceptance criteria
uv run python -c "import coreml_diffusion; print(coreml_diffusion.list_model_versions(), coreml_diffusion.list_quant_modes())"works in a comfy-freeuv venvand prints today's exact strings.- Existing characterization tests pass unchanged.
- No node behavior change yet (node still uses its current hardcoded lists in E1).
STOP — VALIDATE (Gate E1)
## Gate E1 report
- model_version.py confirmed comfy-free (uv, no comfy): PASS/FAIL
- coreml_diffusion.list_* returns byte-identical strings to current dropdowns: YES/NO (show values)
- .name vs .value decision for model_version discovery: <...>
- CONTRACT_VERSION set; additive-only policy linked: <path>
- Characterization tests unchanged & green (uv run pytest): YES/NO
Phase E2 — Move conversion code into coreml_diffusion (in-repo, dedup, behavior-preserving)
Objective: physically relocate the conversion mechanics into the framework-free package, collapsing the two duplicate converters into one, with paths/device injected.
Tasks
- Move into
coreml_diffusion/:pipeline.py(get_pipeline,get_unet),unet.py(UNet2DConditionModelLCM),inputs.py(sample/lcm/sdxl input builders +get_encoder_hidden_states_shape+get_coreml_inputs+get_inputs_spec),controlnet.py(add_cnet_support),convert.py(convert_unet,convert,convert_to_coreml,load_coreml_model),compile.py(compile_coreml_model). - Dedupe LCM vs main (the real remaining duplication): delete
lcm/converter.py's copies ofconvert_to_coreml/load_coreml_model/get_out_path/get_sample_input(LCM variant carries aschedulerarg — fold that into the sharedget_sample_inputas an optional param) in favor of the singlecoreml_diffusionimplementation. The main path's helpers (get_unet/get_encoder_hidden_states_shape/get_coreml_inputs/convert_unet/convert) and theconversion/subpackage (attention,shapes,trace,unet) move as-is. - Inject paths: replace
get_out_path'sfolder_pathsreach-in with an injectedout_pathargument onconvert(...);folder_pathsresolution moves up into the node (which already computesout_name). Nofolder_pathsimport anywhere incoreml_diffusion. - Inject device where the LCM path used
comfy.model_management(verify it still does):convert(..., device=None), default to torch's default device. - Attention per-call, never at import: main path already routes through
conversion/attention.apply_attention_implementation— keep that. Iflcm/converter.pystill sets any module global at import, remove it; the package sets attention from theattn_implarg insideconvert. - Move
compose_out_nameinto the package (coreml_diffusion/naming.py); re-pointtest_characterization_out_name.pyimports tocoreml_diffusion.naming— assertions and values unchanged. The node will import it from the package in E3. - Leave thin shims in
coreml_suite/converter.pyandcoreml_suite/lcm/converter.pythat re-export fromcoreml_diffusion, preserving the old call signatures the nodes use (nodes untouched this phase). Shims map comfyfolder_paths/device into package args.
Acceptance criteria
uv run pytest -m unit(Tier 0) importscoreml_diffusion.*with no comfy / no MPS and is green on Linux.- The dedup leaves exactly one implementation of each previously-duplicated function.
- Characterization tests pass unchanged after the
compose_out_namere-point. [M2]A real SD1.5 conversion via the shim still produces a loadable model.[M2-ANE]Golden latent identical / within tolerance to the MODERNIZATION Phase 2 anchor (same seed/prompt) — proves the move + dedup changed nothing.
STOP — VALIDATE (Gate E2 — first regression gate)
## Gate E2 report
- Tier 0 import of coreml_diffusion without comfy/MPS (uv run): PASS/FAIL
- LCM/main duplicated funcs collapsed to one (list old→new): <map>
- compose_out_name moved to package; char-tests re-pointed & green: YES/NO
- Paths injected (no folder_paths in package): confirmed
- Device injected (no comfy.model_management in package): confirmed
- Attention set per-call, not at import (both main & lcm): confirmed
- [M2-ANE] Golden latent vs Phase-2 anchor: identical / within tol <x> / DIVERGED (STOP)
- Node INPUT_TYPES / mappings untouched: confirmed (diff)
If the golden latent diverged at all, STOP and report — do not continue.
Phase E3 — Thin the nodes onto the package (behavior-preserving)
Objective: remove the shims; have the ComfyUI nodes call coreml_diffusion directly, keeping the
node contract byte-identical.
Tasks
CoreMLConverter.convert(incoreml_suite/nodes.py): keep thefolder_paths-based path resolution in the node; importcompose_out_namefromcoreml_diffusion(notcoreml_suite.core); callcoreml_diffusion.convert(...)andcoreml_diffusion.compile_model(...)directly; wrap the compiled path inCoreMLModel.- Wire the dropdowns to discovery (the maintainer's hard requirement). Replace the hardcoded
INPUT_TYPESlists with fail-soft discovery calls:This is what makes "update the package → new types appear in the old node, no Suite bump" true.def _discover(fn, fallback): try: import coreml_diffusion return getattr(coreml_diffusion, fn)() except Exception as e: # missing/old package, or import error logger.warning(f"coreml_diffusion.{fn} unavailable ({e}); using fallback {fallback}") return fallback ... "model_version": (_discover("list_model_versions", ["SD15", "SDXL"]),), "attention_implementation": (_discover("list_attention_impls", ["SPLIT_EINSUM","SPLIT_EINSUM_V2","ORIGINAL"]),), "quantize_nbits": (_discover("list_quant_modes", ["none","8","6","4"]), {"default": "none"}), COREML_CONVERT_LCM(incoreml_suite/lcm/nodes.py): route throughcoreml_diffusionfor the shared mechanics. Keep the existing LCM behavior/HF-hardcode for now — consolidation is optional E-LCM.- Delete the now-dead
coreml_suite/converter.py/coreml_suite/lcm/converter.pyshims (or reduce to a one-line re-export if anything external imports them — grep first).
Acceptance criteria
NODE_CLASS_MAPPINGS/NODE_DISPLAY_NAME_MAPPINGSkeys: unchanged (diff__init__.py).- Every
INPUT_TYPESfield name unchanged. Dropdown values: the discovery calls must return a superset of today's values, with every previously-present value still present and spelled identically (additive-only). (This deliberately replaces the original spec's "values must be byte-identical/frozen" criterion — the maintainer requires the list be extensible at runtime. Frozen-field-names + additive-only-values is the new contract.) - With
coreml_diffusionabsent, the node still registers and shows the fallback lists (fail-soft). [M2-ANE]Golden latent still identical to the Phase-2 anchor.[M2-ANE]The committed e2e workflowtests/integration/...still passes (PSNR > 25).
STOP — VALIDATE (Gate E3)
## Gate E3 report
- Node mappings diff: empty (confirmed)
- INPUT_TYPES field-names diff: empty (confirmed)
- Dropdown values: superset of prior, all prior values still present & identical: YES/NO (show)
- Fail-soft with coreml_diffusion absent (node still registers): PASS/FAIL
- compose_out_name now imported from coreml_diffusion (no node-side copy): confirmed
- [M2-ANE] Golden latent vs anchor: identical / within tol / DIVERGED (STOP)
- [M2-ANE] e2e workflow PSNR: <value> (> 25?)
- Dead converter shims removed / reduced: <list>
Phase E4 — Standalone packaging & CLI (still in-repo)
Objective: make coreml_diffusion independently installable and usable without ComfyUI, with a CLI
suitable for the planned article and for on-device/iOS conversion workflows.
Tasks
- Add
coreml_diffusion/pyproject.toml: name (workingcoreml_diffusion),requires-python, dependencies =coremltools(pinned to the MODERNIZATION-validated version),diffusers,transformers,peft,omegaconf,numpy,torch. Noml-stable-diffusion(already removed in #58, see §0.3) and no comfy. Suite pinstransformers>=4.44/peft>=0.13/omegaconf>=2.3today; grep-confirm each is on the conversion path before listing it. A[project.scripts]entry:coreml-diffusion = "coreml_diffusion.cli:main". coreml_diffusion/cli.py:coreml-diffusion convert --ckpt PATH --model-version sd15 --out PATH [--height --width --batch-size --attn-impl --controlnet --lora NAME:STRENGTH ... --config PATH]andcoreml-diffusion compile --src PATH --out-dir DIR --name NAME. Mirrorsconvert()/compile_model().- Tier-0 Linux tests for the CLI arg→call mapping (mock the heavy
convert); the real convert remains[M2]. Add a[M2]smoke test: convert a tiny synthetic UNet end-to-end. - README for the package: install, CLI usage, "produce a
.mlpackage/.mlmodelcfor use in a Swift/iOS app", and the ANE positioning note (low-power, GPU-free, embeddable; SD1.5/SDXL on ANE, not a Flux-speed claim).
Acceptance criteria
- Fresh
python -m venv+uv pip install ./coreml-diffusion(no ComfyUI present) imports and runscoreml-diffusion --helpand the arg-mapping tests on Linux. [M2]coreml-diffusion convertproduces a model file identical (golden) to the node path.
STOP — VALIDATE (Gate E4)
## Gate E4 report
- uv pip install ./coreml-diffusion in comfy-free venv: PASS/FAIL (log)
- CLI arg→call tests (Tier 0, Linux): green
- [M2] CLI-produced model golden vs node-produced model: identical / DIVERGED
- Package deps list (with pinned SHAs/versions + licenses):
- New runtime deps vs suite before: <none / list>
This is the gate that proves the package stands alone. Do not split repos before it passes.
Phase E5 — Physical split into a second repository
Objective: move coreml_diffusion/ to its own repo; CoreMLSuite depends on it by pinned version.
Tasks
- Create the new repo (maintainer action — agent prepares the tree, not the GitHub repo). Choose final distributable name; rename imports if changed (single sweep, recorded).
- CoreMLSuite
pyproject.toml/requirements.txt: replace the conversion-only deps with a pinned dependency on the new package (coreml_diffusion==<version>from PyPI, orgit+...@<tag>until first PyPI release). (There is nogit+...ml-stable-diffusionline to remove — already gone since #58.) KeepVoid. The loader is the localpython_coreml_stable_diffusionfor the loader.coreml_suite/coreml_model.pyovercoremltools; the suite keepscoremltoolsas a direct dep for it. No Apple lib involved.- Set up the new repo's CI: Tier 0 on Linux (import + arg-mapping + input-shape math),
[M2]/[M2-ANE]on a self-hosted/macOS-ARM runner reusing the golden-latent anchor. - Versioning: SemVer; first release
0.1.0. Document the compatibility matrix (coreml_diffusion ↔ coremltools version ↔ diffusers version). No ml-stable-diffusion axis.
Acceptance criteria
- CoreMLSuite installs in a fresh venv pulling the new package; e2e workflow still passes
[M2-ANE]. - New repo CI green on Linux (Tier 0) and
[M2-ANE]golden latent matches the anchor. - No conversion code remains in CoreMLSuite (grep: no
ct.convert, nofrom_single_file, notorch.jit.trace).
STOP — VALIDATE (Gate E5)
## Gate E5 report
- New repo tree prepared: <path/branch>; final package name: <name>
- Suite depends on package by pinned version: <spec>
- Suite e2e [M2-ANE] PSNR after split: <value> (> 25?)
- Conversion code fully absent from suite: confirmed (grep output)
- Compatibility matrix documented: <link>
- First release tag: 0.1.0
Phase E6 — Quantization travels WITH the conversion code (already implemented → move)
Objective: quantization is already implemented (k-means palettize_weights in
converter.py, quantize_nbits node input, _q<bits> filename suffix, README tradeoff table).
There is nothing to build. It simply moves with the conversion code in E2 as part of
convert_unet. This phase is a checkpoint that it survived the extraction intact, plus exposing
it through the CLI.
Tasks
- Confirm the palettization block moved cleanly into
coreml_diffusion(lives inconvert.pyor aquantize.pyhelper called fromconvert_unet). Default"none"stays byte-identical. - Expose via CLI flag
--quantize {none,8,6,4}(E4 already lists this) and vialist_quant_modes()discovery (E1/E3). - The existing README tradeoff table (SD1.5 1×512×512 SPLIT_EINSUM: none/8/6/4 → size/ms/PSNR)
moves to the package README. Re-confirm one row
[M2-ANE]so the article can cite a live number.
Acceptance criteria
- Default (
none) output byte-identical to pre-extraction (covered by the E2/E3 golden latent). coreml-diffusion convert --quantize 4produces a_q4artifact matching the node's_q4artifact[M2].list_quant_modes()drives the node dropdown (no hardcoded copy remains).
STOP — VALIDATE (Gate E6)
## Gate E6 report
- Palettization relocated into coreml_diffusion, called from convert_unet: confirmed
- Default none output identical (golden): YES/NO
- [M2] CLI --quantize {8,6,4} artifacts match node artifacts: YES/NO
- Tradeoff table in package README with at least one re-confirmed [M2-ANE] row: <link>
Phase E-LCM — FIRST task after the split: clean up LCM + verify → promote (behavior-changing, gated)
Promoted from "optional, someday" to the first thing after E5, per maintainer intent: the Suite should expose every verifiably-convertible model, and LCM is the obvious first cleanup.
Two coupled goals:
- Consolidate the LCM path. Make the LCM node use the unified
from_single_filepath incoreml_diffusion.convert(model_version=LCM, ...)instead of the hardcodedSimianLuo/LCM_Dreamshaper_v7HF download; drop the duplicated LCM helpers (already deduped in E2). Behavior change ⇒ capture an LCM golden anchor before the change, then prove within-tolerance after. - Verify → promote. Once the LCM conversion has a passing
[M2-ANE]golden anchor, flip_MODEL_STATUS["lcm"] = Status.VERIFIEDin the package (minor bump). The Suite's dropdown gainslcmautomatically — no Suite change, no Suite bump. This is the end-to-end proof that the discovery contract works as designed.
Repeat the same recipe for sdxl_refiner when it gets an anchor (separate small gate). Do NOT
bundle E-LCM into E1–E5; it changes behavior and must stand on its own golden.
STOP — VALIDATE (Gate E-LCM)
## Gate E-LCM report
- LCM golden anchor captured BEFORE change: <path/hash>
- LCM node now uses unified from_single_file path; HF hardcode removed: confirmed
- [M2-ANE] LCM golden after change: identical / within tol <x> / DIVERGED (STOP)
- Status flipped lcm→VERIFIED in package (minor bump <ver>): confirmed
- Suite dropdown now lists lcm with NO Suite code change / NO Suite bump: confirmed (diff empty)
- LCM node accepts a checkpoint arg now (documented breaking-ish UI note): <link>
Article deliverable (after E4)
Once the CLI exists and stands alone, the "convert a Comfy/A1111 workflow into an on-device iOS
app" write-up becomes a clean tutorial: coreml-diffusion convert → .mlmodelc → load in Swift/CoreML.
Frame the niche honestly per the README note above (ANE feasibility & power, not raw Flux speed).
Quick reference: extraction gate discipline
E0 Seam decision, interface contract, discovery API frozen → Gate E0 (cut line + additive-only policy?)
E1 Stand up coreml_diffusion + discovery API (mostly verify) → Gate E1 (list_* byte-identical, comfy-free?)
E2 Move conversion code, dedup LCM/main, inject paths/device→ Gate E2 (golden identical? duplicates gone?) ← first regression gate
E3 Thin nodes onto package + wire discovery dropdowns → Gate E3 (field-names frozen, values additive, fail-soft, golden identical?)
E4 Standalone packaging + CLI (uv) → Gate E4 (uv pip install w/o comfy? CLI golden?) ← proves it stands alone
E5 Physical second-repo split → Gate E5 (suite depends on pkg? conversion absent?)
E-LCM FIRST post-split: clean up LCM, verify → promote → Gate E-LCM (LCM golden? dropdown gains lcm w/ no Suite bump?)
E6 Quantization checkpoint (already built → moved in E2) → Gate E6 (default identical? CLI quant matches?)
(refiner) same recipe as E-LCM when an anchor exists → own small gate (promote sdxl_refiner→VERIFIED)
Interface-contract invariants (the maintainer's hard requirement), restated:
- Package API is keyword-only-with-defaults past the required positionals → converter updates don't force Suite updates.
- Node dropdowns are discovery-driven (
coreml_diffusion.list_*) + fail-soft → new conversion types appear in the old plugin withuv pip install -U coreml_diffusion, no Suite code change, no bump. - Discovery identifiers are additive-only; removal/rename = MAJOR bump + migration note.
compose_out_name(cache key) lives in the package, single copy.
Golden rule (inherited): never cross a gate with a failing acceptance criterion. Stop, report, wait. The golden latent is the single source of truth that the extraction changed nothing.