Commit Graph
23 Commits
Author SHA1 Message Date
aszc-dev 8569d5f796 fix: allow custom converter dimensions 2026-05-26 16:15:35 +02:00
aszc 65a2de2fab feat!: modernize toolchain and replace apple/ml-stable-diffusion with native diffusers conversion (#58)
* feat(deps): support ComfyUI's numpy 2 toolchain; make conversion optional

The runtime package now installs and runs under numpy 2 / coremltools 9 /
torch 2.7 — matching current ComfyUI — without apple/ml-stable-diffusion.

- Drop the heavy converter stack (ml-stable-diffusion, diffusers, peft,
  omegaconf, overrides, transformers) from runtime dependencies; require
  numpy>=2.
- Vendor the runtime pieces: a slim CoreMLModel wrapper around coremltools
  and the attention-implementation constants.
- Lazy-import the converters; the Convert nodes raise a clear error when the
  legacy conversion dependencies are absent. Loading and sampling existing
  Core ML models no longer needs them.
- CI: Tier 0 tracks the numpy 2 / torch 2.7 toolchain; drop the Tier 2
  golden-image lane (it converts at runtime, which now requires the legacy
  stack) and its fixtures.

* feat(conversion): replace apple/ml-stable-diffusion with native diffusers path

Reimplement Core ML UNet conversion on top of diffusers instead of the
apple/ml-stable-diffusion git dependency, so the full suite (including
conversion) installs through ComfyUI Manager without extras on the NumPy 2
toolchain.

- Add coreml_suite/conversion package: split-einsum attention processors,
  a conv2d output-shape helper, Transformer2D trace patches, and a UNet
  input-adapter wrapper preserving the historical Core ML I/O contract.
- Drop python_coreml_stable_diffusion and overrides; route SD15, SDXL,
  SDXL refiner, and LCM conversion through diffusers UNet2DConditionModel.
- Declare diffusers, peft, omegaconf, and transformers as runtime deps.
- Add characterization tests asserting split-einsum matches reference
  attention math; extend the synthetic-UNet smoke test for the wrapper.
- Bump to 1.1.0 and set requires-comfyui to a semver constraint (>=0.3.27)
  so the Comfy Registry publish succeeds.

* refactor(conversion)!: native diffusers context layout; drop legacy fallbacks

Address PR review feedback:

- Drop the legacy converter ImportError fallbacks and LEGACY_CONVERTER_MODULES
  guards in nodes.py and lcm/nodes.py. Conversion dependencies are mandatory in
  pyproject, so the indirection is dead code.
- Tier 0 CI resolves its toolchain from pyproject via uv (uv sync + uv run)
  instead of hand-pinned pip installs, removing duplicated version maintenance.
- Document the conversion lineage: credit apple/ml-stable-diffusion as the
  origin, note the implementation has diverged to a native diffusers path, and
  state the intent to iterate independently. Fix stale README links that pointed
  users to apple/ml-stable-diffusion for conversion.
- Drop the unused `sources` argument from CoreMLModel.

BREAKING CHANGE: the converted Core ML UNet now takes encoder_hidden_states in
the native diffusers layout (batch, tokens, hidden) instead of
(batch, hidden, 1, tokens). This removes the boundary transposes in
CoreMLUNetWrapper and CoreMLInputs. Core ML models converted with earlier
versions are incompatible and must be re-converted. Bump to 2.0.0.

* test: widen split-einsum allclose tolerance for cross-platform float drift

The split-einsum attention reorders float32 reductions relative to the
reference, so equality holds only up to rounding. The default allclose atol
(1e-8) is too tight on Linux x86 BLAS and failed Tier 0 CI; use atol=1e-6 to
match the existing chunked-path characterization test.

* ci: run macOS smoke tier on the self-hosted Apple Silicon runner

GitHub-hosted macOS carries a 10x minute multiplier and exhausts the included
Actions minutes too quickly. Move the Tier 1 smoke job onto the self-hosted
Apple Silicon runner ([self-hosted, macOS, ARM64, coreml]) so macOS coverage no
longer consumes hosted minutes. Tier 0 stays on hosted ubuntu (1x).

* ci: fix uv setup for both tiers

astral-sh/setup-uv@v3 was retagged and its old commit garbage-collected, so
codeload 404s when Actions resolves the stale SHA. Bump Tier 0 (ubuntu) to
setup-uv@v7, and drop the action entirely from Tier 1 since the self-hosted
runner already provides uv.

* test(ci): restore golden-image correctness gate on the self-hosted runner

The Tier 2 end-to-end correctness check (real SD1.5 -> Core ML -> image,
gated on SHA/PSNR vs a golden) was dropped during the modernization. With the
breaking 3D-context change, the synthetic smoke and shape/attention
characterization tests no longer cover real-model conversion correctness.

Restore tier2.yml (on the [self-hosted, macOS, ARM64, coreml] runner shared
with Tier 1), the golden-image test, and the e2e workflow. Adapt the pinned
ComfyUI resolution to the requires-comfyui semver tag (vX.Y.Z) instead of a
commit SHA, and re-register the m2 marker. The golden is intentionally not
committed: the first self-hosted run regenerates it under the new 3D contract
and fails for review, per the test's documented bootstrap.

* test(ci): add golden image for SD1.5 seed 42 under the 3D-context contract

Generated by the first Tier 2 self-hosted run after the native diffusers
conversion change. The decoded image is a coherent SD1.5 generation, confirming
the (batch, tokens, hidden) Core ML contract produces correct output
end-to-end. Subsequent runs gate on this golden (SHA-strict, PSNR fallback).

* test(ci): force fresh conversion in Tier 2; drop stale golden

The converter skips conversion when a same-named model already exists, keyed on
conversion parameters but not the conversion code/toolchain. The self-hosted
runner held a pre-existing v1-5 .mlmodelc (4D-context, old toolchain), so the
Tier 2 runs reused it (~8-30s) instead of converting — the gate validated a
stale model, not the new native diffusers path.

Purge the cached Core ML UNets before running so every Tier 2 run does a real
convert -> compile -> sample. Drop the golden generated from the stale cache;
the next run regenerates it from a genuine 3D-contract conversion and fails for
review.

* test(ci): add golden image from a genuine 3D-contract conversion

Regenerated by a Tier 2 run with the model cache purged, so the converter
actually ran (62s, not a cache hit). The fresh model exposes the new 3D
encoder_hidden_states input [1, 77, 768], and its decoded SD1.5 seed-42 image
is byte-identical to the prior baseline — confirming the native diffusers
conversion is behavior-preserving end to end.
2026-05-26 15:46:09 +02:00
aszc 02b6e8ece3 feat: modernize toolchain, refactor core, add tiered CI and opt-in quantization
Modernizes ComfyUI-CoreMLSuite onto Python 3.12 / torch 2.7 / coremltools 9
with a characterization-test safety net. The default conversion path is
unchanged; existing saved workflows produce identical output.

- Toolchain bump (Python 3.12, torch 2.7, coremltools 9, numpy <2) with the
  blocking upstream pins overridden.
- Framework-free logic moved into coreml_suite/core/ (no comfy/coremltools
  imports); old module paths re-export from there.
- Opt-in quantize_nbits dropdown (none|8|6|4) for k-means weight
  palettization; default none is byte-for-byte identical to before.
- Tiered CI: Tier 0 (Linux unit), Tier 1 (macOS-ARM smoke), Tier 2
  (self-hosted Apple Silicon golden-image check on the ANE).
2026-05-25 19:11:49 +02:00
aszc-dev aa60cda09b Add installation using ComfyUI-Manager instructions 2023-11-24 15:10:33 +01:00
aszc-dev 5f7fcd6df3 Add note on SD2.1 to readme 2023-11-24 12:42:33 +01:00
aszc-dev e89cff6d01 Update readme with SDXL info 2023-11-24 12:14:15 +01:00
aszc-dev 9f90083126 Update converter docs and workflows 2023-11-24 12:14:15 +01:00
aszc-dev bb44b4a35f Link to ComfyUI repo 2023-11-24 12:14:15 +01:00
aszc-dev 46d1124573 Update REAMDE.md (Conversion and LoRA) 2023-11-17 22:55:13 +01:00
aszc-dev f9f25fbeb7 Add LCM info to readme 2023-11-13 13:47:15 +01:00
aszc-dev d63df5b62f Remove the controlnet note in readme 2023-10-31 22:05:19 +01:00
aszc-dev 6319d2aedb Add model adapter for unstable compatibility 2023-10-30 11:49:07 +01:00
aszc-dev c043e1f9aa Update ControlNet workflow 2023-10-30 01:01:59 +01:00
aszc-dev 367c6beb19 Update README 2023-10-30 00:41:24 +01:00
aszc 1c85a0e397 Fix conversion options in readme 2023-10-26 12:44:54 +02:00
Adrian Szczepański 721cff74a7 Emphesize the usage of ANE in README 2023-10-26 02:37:57 +02:00
aszc 352809daa4 Add workflows to README 2023-10-26 02:29:20 +02:00
Adrian Szczepański 235ea70b3b Update README.md 2023-10-26 02:25:54 +02:00
aszc f66ea41148 Fix the footnote 2023-10-25 13:30:05 +02:00
Adrian Szczepański fdb75e2675 Update README.md with info on different input sizes 2023-10-24 23:26:30 +02:00
Adrian Szczepański 53ce35794f Update info on LoRA 2023-10-23 15:38:59 +02:00
aszc 5a89dffd2c Update README.md 2023-10-23 15:24:32 +02:00
Adrian Szczepański 556f47bfc5 Add README.md 2023-10-23 15:10:35 +02:00