Commit Graph
116 Commits
Author SHA1 Message Date
aszc-dev 0bbd8d8e0d feat(phase6): opt-in k-means weight palettization (quantize_nbits)
Phase 6 of the modernization plan: add weight palettization to the
Core ML converter as an opt-in knob, so the SD1.5 / SDXL UNet can
ship at 1/2, 1/2.7 or 1/4 of its current size with ANE-friendly
inference.

CoreMLConverter (and the LCM converter) gains a `quantize_nbits`
dropdown: `none` (default — identical to pre-Phase-6 behavior and
filenames, so existing cached .mlpackages still resolve) / `8` / `6` /
`4`. The value is encoded as `_q<bits>` after the attn suffix, so the
unquantized model and the three palettized variants coexist on disk
under distinct cache keys.

Implementation
- core/naming.compose_out_name: accepts `quantize_nbits`, validates
  against {none, 8, 6, 4}, appends `_q<bits>` (none = empty).
- converter.convert_unet: after ct.convert + before .save, runs
  coremltools.optimize.coreml.palettize_weights with
  OpPalettizerConfig(mode="kmeans", nbits=...) when the value is not
  "none". Adds a `Palettization took Xs` log line.
- converter.convert / nodes.CoreMLConverter.convert: pipe the new arg
  through; the ComfyUI node exposes it as a dropdown with default
  "none" so existing workflows are unchanged at load time.
- bench/scripts/convert_sd15.py: QUANT_NBITS env knob; uses the
  pure compose_out_name (replaces the inline string formatter).

Test infra
- tests/unit/test_characterization_out_name.py: 6 new tests pinning
  the `_q<bits>` suffix contract, the "none" passthrough (backward
  compat), the cn + lora + quant combination, and the invalid-value
  ValueError. Total Tier 0 now at 94.
- Makefile gains `bench-quant` (runs the matrix script) and
  `convert-quant` (converts q8, q6, q4 sequentially).
- bench/scripts/quant_matrix.py (new): loads each variant, runs
  REPEATS forward passes with a fixed seed, then computes the
  noise_pred PSNR of each quantized variant against the unquantized
  baseline. Writes bench/results/quant_matrix_<sha>.{json,md}.

README
- New "Quantization (Phase 6, opt-in)" section: tradeoff table
  measured on M2 Pro SD1.5 1x512x512 SPLIT_EINSUM (sizes 1641/822/
  617/412 MB; fwd 197/187/183/180 ms; PSNR 53.5 / 40.2 / 27.5 dB),
  plus per-chip/RAM recommendations.

Default-path safety
- "none" produces the same out_name as Phase 5 -> existing
  v1-5-pruned-emaonly_1x512x512_se_unet.mlmodelc is still picked up
  unchanged; the m2 golden image test continues to anchor.
2026-05-25 01:30:29 +02:00
aszc-dev 2b649e6606 chore(phase5,bench): record bumped-toolchain environment and results
Phase 5 reference numbers, captured against commit 1e5791d on macOS
26.1 with the bumped toolchain (torch 2.7.1, coremltools 9.0, numpy
1.26.4, Python 3.12.11).

- bench/env/baseline-1e5791d.txt: full uv-pip freeze + ComfyUI sha +
  macOS + resolved ml-stable-diffusion git metadata.
- bench/env/pytest-unit-1e5791d.txt: pytest -m unit 88/88 passing
  (Tier-0 purity gate confirms no comfy/coreml leak).
- bench/results/1e5791d.{json,md}: SD1.5 1x512x512 SPLIT_EINSUM UNet
  forward latency on the Apple Neural Engine and CPU+GPU. Held within
  noise of the Phase 1 baseline (ef2a18c.json) — bump is
  performance-neutral.
2026-05-24 16:37:26 +02:00
aszc-dev 1e5791d108 chore(phase5): bump Python 3.12 / torch 2.7 / coremltools 9
Phase 5 of the modernization plan: the intentional tooling upgrade
against the Phase 1 baseline. numpy 2 stays out of scope (decoupled —
see docs/deps.md).

Pyproject pins
- requires-python: ">=3.11,<3.12" -> ">=3.12,<3.13"
- torch: ==2.0.1 -> >=2.7,<2.8 (latest the coremltools 9 PyTorch
  frontend has been tested against)
- coremltools: ==8.2 -> >=9,<10
- numpy: <1.25 -> >=1.24,<2 (held below 2 — coremltools+numpy2 has
  known SD UNet trace bugs in `_cast` and `view`; none of our modules
  need numpy 2)
- ml-stable-diffusion SHA: unchanged at e5d960c4 (upstream main has
  the same restrictive pins; no working alternative)

uv overrides
- override-dependencies relaxes the four hard pins ml-stable-diffusion
  ships in setup.py: numpy<1.24, diffusers==0.30.2, transformers==4.44.2,
  huggingface-hub==0.24.6. The .unet / .coreml_model symbols we
  actually import (see docs/deps.md) are stable across the bumped
  versions.

[dependency-groups] comfy
- New group with ComfyUI's runtime deps (einops, torchvision, torchsde,
  comfyui-frontend-package, spandrel, ...). Replaces the Phase 1 / 4
  `uv pip install -r ComfyUI/requirements.txt` dance that floated torch
  to the latest version and broke the coremltools ceiling. `uv sync
  --group comfy` is the new contract; the Makefile already invokes the
  project venv directly.

Tier 2 golden re-anchored
- The toolchain bump is performance-neutral on SD1.5 (NE fwd median
  delta +0.2%, GPU +0.6% — within run-to-run noise) but bit-changes
  the Core ML UNet output (different MIL graph + kernel selection).
  The Phase 2 golden PNG hashes to a different SHA256 now and lands
  at ~29 dB PSNR against itself. Visually identical, just numerically
  different.
- tests/m2/goldens/sd15_seed42.{png,sha256} re-captured against the
  bumped toolchain.
- tests/m2/test_golden_image.py: GOLDEN_PSNR_MIN_DB lowered from 40
  to 25 (typical post-toolchain-bump tolerance). Header docstring
  updated to explain when to raise it back for refactor PRs.

docs/deps.md (new)
- ml-stable-diffusion compatibility decision (override vs vendor vs
  fork), why numpy 2 was punted, Tier 2 PSNR threshold reasoning,
  bench diff table, and explicit rollback instructions.

Local verification
- pytest -m unit  -> 88/88 passed in 1.87s
- pytest -m smoke -> 1/1 passed in 2.59s
- pytest -m m2    -> 1/1 passed (after re-anchor)
- bench/run.py    -> SD1.5 NE 197 ms / GPU 272 ms, perf-neutral vs
                     Phase 1 baseline (ef2a18c.json)
2026-05-24 16:36:37 +02:00
aszc-dev 8382b13598 ci(phase4): tiered test/CI infrastructure (Tier 0/1/2)
Phase 4 of the modernization plan: institutionalize the 3-tier strategy
so future changes are guarded automatically, and pin down the
self-hosted M2 path the maintainer's hardware needs.

Tier dispatch
- Makefile targets test-unit / test-smoke / test-m2 / bench (plus
  ci-tier0 / ci-tier1 wrappers that echo env first). check-macos-arm
  fails fast on non-Apple-Silicon hosts.

Tier 1 smoke
- tests/smoke/test_synthetic_unet.py: builds a TinyUNet (conv-in,
  time/text projections, conv-out), traces it, ct.convert to
  mlprogram + fp16 CPU_ONLY, loads back via CoreMLModel and asserts
  expected_inputs + named output. Runs in ~2s; auto-skips on
  non-Apple-Silicon. Catches coremltools / ml-stable-diffusion API
  drift without needing a real SD checkpoint or the ANE.

GitHub Actions
- .github/workflows/tier0.yml: ubuntu-latest on every push/PR, ~10
  min budget, minimal-deps install (torch==2.0.1, numpy<1.25, pytest)
  -> pytest -m unit.
- .github/workflows/tier1.yml: macos-14 (M1) on push/PR; opt-in via
  run-tier1 label on labeled PRs to spare external-doc PRs.
- .github/workflows/tier2.yml: self-hosted [macOS, ARM64, coreml] on
  PR label run-m2 / nightly cron / workflow_dispatch. Starts ComfyUI
  with --cpu-vae, runs pytest -m m2 + bench/run.py, uploads bench
  results.

Integration coverage moved
- Removed tests/integration/test_basic_conversion_1_5.py: it required
  an MPS reference image (broken on macOS 26 + torch 2.0.1, see
  Phase 1 Gate) and a checkpoint the maintainer doesn't have on disk
  (dreamshaper_8). The same coverage now lives in
  tests/m2/test_golden_image.py: deterministic numerical pass/fail
  (SHA256 + PSNR fallback) against a stored golden, Core ML pipeline
  only. No more human eyeballing.

Docs
- docs/ci-m2.md: one-time runner registration steps, COMFY_DIR
  persistence, baseline model pre-conversion, trigger semantics, what
  to do when the runner is offline, and the migration note from
  integration -> m2 golden.

Sanity check
- Temporarily set convert_to="BREAKAGE_CANARY_NOT_A_REAL_FORMAT" in
  the smoke test; Tier 1 surfaced
  NotImplementedError: Backend converter BREAKAGE_CANARY_NOT_A_REAL_FORMAT not implemented
  immediately. Reverted.

Local verification
- make test-unit -> 88/88 passed in 2.09s
- make test-smoke -> 1/1 passed in 1.99s
2026-05-23 23:11:35 +02:00
aszc-dev 5dafd261b7 refactor(phase3): split pure logic into coreml_suite.core
Phase 3 of the modernization plan: move the framework-free math out of
the comfy-coupled modules so Tier-0 tests can run on plain Linux without
ComfyUI, coremltools, or python_coreml_stable_diffusion.

New pure-core package (no comfy / coreml / mps imports):
- coreml_suite.core.latents: chunk_batch, merge_chunks
- coreml_suite.core.controlnet: expand_inputs, no_control,
  extract_residual_kwargs, chunk_control
- coreml_suite.core.inputs: CoreMLInputs (chunks + coreml_kwargs)
- coreml_suite.core.sdxl: is_sdxl / is_sdxl_base / is_sdxl_refiner,
  build_sdxl_time_ids (base len 6, refiner len 5), build_sdxl_text_embeds,
  sdxl_model_function_wrapper
- coreml_suite.core.naming: compose_out_name, lora_names_from_params

Thin adapters keep the public import paths:
- coreml_suite.latents / coreml_suite.controlnet: re-export from core
- coreml_suite.models: CoreMLModelWrapper, CoreMLModelWrapperLCM,
  add_sdxl_model_options (now uses the pure builders from core.sdxl),
  get_latent_image, get_model_patcher remain framework-coupled
- coreml_suite.nodes: CoreMLConverter.convert now delegates the out_name
  composition to core.naming.compose_out_name

Test infra:
- tests/unit/* re-pointed at coreml_suite.core.*
- test_chunks.py dropped `from comfy.model_management import ...` and
  the dead `model_config` fixture (Phase 1 left it broken; Phase 3
  removes it entirely)
- test_characterization_sdxl_options now targets the pure builders
  directly via inspect.getclosurevars on the wrapper closure
- test_characterization_out_name now calls compose_out_name without the
  heavy CoreMLConverter monkey-patching that Phase 2 needed
- tests/unit/test_tier0_purity.py: new gate that fails if comfy /
  coremltools / etc leak into sys.modules during a pure `-m unit` run
  (skipped in mixed runs where m2 / integration legitimately import them)
- tests/__init__.py + top-level conftest.py + pyproject addopts
  `--import-mode=importlib --confcutdir=tests` together stop pytest from
  importing the repo-root `__init__.py` (the ComfyUI custom-node entry
  pulls in comfy)
- tests/conftest.py adds tier-aware collect_ignore so `-m unit` skips
  tests/m2 + tests/integration at collection time

Verification:
- `pytest -m unit tests/` → 88 passed in ~2s; deterministic across runs
- Tier-0 purity gate confirms no comfy/coreml/etc in sys.modules
- m2 golden image (Phase 2 anchor) still hashes identical → refactor
  produced bit-for-bit unchanged output
- `git diff main -- __init__.py coreml_suite/nodes.py` shows zero churn
  to NODE_CLASS_MAPPINGS keys or INPUT_TYPES field names (public
  workflow contract intact)
2026-05-22 16:08:01 +02:00
aszc-dev 04911d0052 test(phase2): add characterization tests + M2 golden image anchor
Phase 2 of the modernization plan: lock the current behavior of the pure
math so the Phase 3 refactor cannot silently change it.

Unit characterization tests (Tier 0, 64 new):
- test_characterization_latents.py: chunk_batch / merge_chunks padding,
  truncation, and identity contracts.
- test_characterization_controlnet.py: expand_inputs / no_control /
  extract_residual_kwargs / chunk_control shape, dtype, and zero-fill
  behavior, including the [None]*target contract.
- test_characterization_inputs.py: CoreMLInputs.chunks / coreml_kwargs
  for SD1.5, SDXL base (time_ids len 6), SDXL refiner (time_ids len 5),
  and LCM (timestep_cond).
- test_characterization_sdxl_options.py: add_sdxl_model_options time_ids
  / text_embeds assembly via a SimpleNamespace fake ModelPatcher and
  inspect.getclosurevars on the returned model_function_wrapper.
- test_characterization_out_name.py: CoreMLConverter out_name encoding
  for attn_impl suffix, batch/size, ControlNet, LoRA (sorted), SDXL.

M2 [Tier 2] golden image anchor (1 new):
- test_golden_image.py: posts the SD1.5+CoreML workflow to a local
  ComfyUI server (auto-skips if unreachable), asserts SHA256 of the
  generated PNG against tests/m2/goldens/sd15_seed42.sha256; falls back
  to PSNR >= 40 dB if the hash drifts.

Test infra:
- pyproject.toml [tool.pytest.ini_options]: unit / m2 / smoke markers,
  testpaths=tests; rootdir is now this package (was ComfyUI's pytest.ini).
- tests/conftest.py: bootstraps sys.path for comfy imports, auto-marks
  tests by directory, and ignores the maintainer's WIP scaffolds
  (test_experiments / test_unet_conversion / standalone_test) so they
  don't break collection.

All 85 collected tests pass; two consecutive runs produced identical
results (run1: 3.55s, run2: 3.40s).
2026-05-22 15:38:09 +02:00
aszc-dev cf6d7c6855 chore(phase1,bench): record baseline environment and results
Phase 1 reference numbers, captured against commit ef2a18c on macOS 26.1
with the pinned toolchain (torch 2.0.1, coremltools 8.2, numpy 1.23.5,
python_coreml_stable_diffusion@e5d960c4).

- bench/env/baseline-ef2a18c.txt: full pip freeze + ComfyUI sha + macOS +
  resolved ml-stable-diffusion git metadata.
- bench/env/pytest-unit-ef2a18c.txt: pytest tests/unit (test_chunks +
  test_controlnet) 20/20 passing.
- bench/results/ef2a18c.{json,md}: SD1.5 1x512x512 SPLIT_EINSUM UNet
  forward latency on the Apple Neural Engine and CPU+GPU. NE median 197 ms
  / GPU median 270 ms; a second run reproduced both within 0.3% (noise).
- bench/results/smoke/ef2a18c/: end-to-end Core ML image (E2E-1.5-CoreML
  workflow, seed=42) saved by smoke_image.py. The MPS reference branch of
  the original workflow is omitted because torch 2.0.1's MPS backend on
  macOS 26.1 trips a BFloat16 conversion error in VAEDecode and an
  mps.add element-type mismatch in KSampler — both go away with newer
  torch and are tracked for the Phase 5 toolchain bump. The server was
  started with --cpu-vae to route the VAE through CPU; this is a runtime
  flag, not a pin change.
2026-05-22 15:08:31 +02:00
aszc-dev ef2a18cff3 chore(phase1): pin baseline toolchain and add bench harness scaffold
Phase 1 of the modernization plan: freeze the currently-working environment
so later refactors have a measured reference point.

- Pin python-coreml-stable-diffusion to commit e5d960c4 (the one already
  installed in the maintainer's apple_env), plus torch==2.0.1, coremltools==8.2
  and numpy<1.25 to match the only env that loads ComfyUI successfully
  (Comfy's checkpoint-safe-loading branch in utils.py is gated on torch>=2.4,
  so newer torch + numpy 1.23 breaks at import).
- Mirror the same pins in requirements.txt and commit uv.lock for
  reproducible installs.
- Add requires-comfyui pinning ComfyUI to ab541335 (the validated commit).
- Fix tests/unit/test_chunks.py fixture: get_model_config() now takes a
  ModelVersion argument; pass ModelVersion.SD15 (the previously-broken test
  was the only Phase 1 production-code change required).
- Add the Phase 1 baseline harness: bench/run.py (direct Core ML UNet
  latency, deterministic), bench/scripts/convert_sd15.py (one-command
  conversion bypassing the node graph), bench/scripts/smoke_image.py (POSTs
  the existing e2e workflow to a local ComfyUI server and saves the Core ML
  image), bench/env/capture.sh (env snapshot), bench/prompts.json (fixed
  prompt set).
- Ignore apple_env/, comfy_env/, and bench/scripts/*.log.

Tests: 20/20 unit pass (test_chunks + test_controlnet).
2026-05-22 15:06:24 +02:00
snomiao 7678a07ed5 chore(publish): update GitHub Actions workflow for node publishing
- Added permissions for issue writing
- Updated action version to v1 for publish-node-action
- Added condition to run job only for 'aszc-dev' repository owner
2025-04-01 23:45:31 +02:00
snomiao 43b77e8471 chore(licence-update): Update PyProject Toml - License 2024-08-15 20:37:19 +02:00
aszc-dev c96059ff0b Add basic conversion integration test 2024-07-04 08:44:37 +02:00
aszc-dev 3224d62342 Restructure tests directory 2024-07-04 08:44:37 +02:00
aszc-dev 2fb135df03 Fix set_timestamps for new LCMScheduler implementation 2024-07-04 08:44:37 +02:00
aszc-dev 66e83c2f2f Change syntax to support older Python versions 2024-07-04 08:44:37 +02:00
aszc fb7188e5a2 Update pyproject.toml to test registry workflow 2024-07-03 16:15:37 +02:00
haohaocreates 4096466f8c chore(publish): Add Github Action for Publishing to Comfy Registry 2024-07-03 16:13:36 +02:00
aszc b8c263b763 Update pyproject.toml 2024-07-03 16:08:02 +02:00
haohaocreates 56cff2bd91 chore(pyproject): Add pyproject.toml for Custom Node Registry 2024-07-03 16:08:02 +02:00
Chris Chance 7b3f8fc29e Update ModelSamplingDiscreteLCM to Distilled for latest comfyui 2023-12-01 01:09:14 +01:00
Chris Chance adaecd3f66 Lowered minimum CoreML Size to 256x256 2023-11-28 17:49:43 +01:00
aszc-dev aa60cda09b Add installation using ComfyUI-Manager instructions 2023-11-24 15:10:33 +01:00
aszc-dev 5f7fcd6df3 Add note on SD2.1 to readme 2023-11-24 12:42:33 +01:00
aszc-dev e89cff6d01 Update readme with SDXL info 2023-11-24 12:14:15 +01:00
aszc-dev 9f90083126 Update converter docs and workflows 2023-11-24 12:14:15 +01:00
aszc-dev 5c774ddc5e Remove LCM option from converter for now 2023-11-24 12:14:15 +01:00
aszc-dev b8197c21ef Converting refiner works 2023-11-24 12:14:15 +01:00
aszc-dev 0c78803b25 Base SDXL conversion works 2023-11-24 12:14:15 +01:00
aszc-dev 763ca3961b Handle SDXL config 2023-11-24 12:14:15 +01:00
aszc-dev ae9a9874c5 Add Advanced Sampler node 2023-11-24 12:14:15 +01:00
aszc-dev 67c902f761 Generating SDXL with Core ML Sampler works 2023-11-24 12:14:15 +01:00
aszc-dev bb44b4a35f Link to ComfyUI repo 2023-11-24 12:14:15 +01:00
aszc-dev 46d1124573 Update REAMDE.md (Conversion and LoRA) 2023-11-17 22:55:13 +01:00
aszc-dev ead01c08dd Remove lora.py 2023-11-17 22:55:13 +01:00
aszc-dev b10effc7c2 Add conversion/lora workflows 2023-11-17 22:55:13 +01:00
aszc-dev b1d2e82677 Add peft and omegaconf to requirements 2023-11-17 22:55:13 +01:00
aszc-dev 9f650acb79 Load .yaml config if present 2023-11-17 22:55:13 +01:00
aszc-dev c6d6917827 Setting LoRA model weights works 2023-11-17 22:55:13 +01:00
aszc-dev 63377ebd73 Store lora_params in dict 2023-11-17 22:55:13 +01:00
aszc-dev 42ff10cd43 Add node to load LoRAs 2023-11-17 22:55:13 +01:00
aszc-dev da3a8e13d3 Add logging during conversion 2023-11-17 22:55:13 +01:00
aszc-dev 8092a19173 Enable choosing attention implementation during conversion 2023-11-17 22:55:13 +01:00
aszc-dev 5477e3d71a Remove CLIP loader from nodes 2023-11-17 22:55:13 +01:00
aszc-dev a8d2d6ec46 Move lora related code around, remove clip stuff 2023-11-17 22:55:13 +01:00
aszc-dev 44cffbb8b8 Move load_lora to lora.py 2023-11-17 22:55:13 +01:00
aszc-dev 6907d4910f Remove ckpt loading when loading lora clip 2023-11-17 22:55:13 +01:00
aszc-dev 1930be5c98 Remove CLIP related code 2023-11-17 22:55:13 +01:00
aszc-dev 45be6761d1 Basic conversion + LoRA support works 2023-11-17 22:55:13 +01:00
aszc-dev fc1132a5d5 Fix category for all Core ML nodes 2023-11-17 22:55:13 +01:00
aszc-dev e440f725a4 Specify diffusers and coremltools versions in requirements.txt 2023-11-14 18:43:15 +01:00
aszc-dev f9f25fbeb7 Add LCM info to readme 2023-11-13 13:47:15 +01:00