Files
aszc-dev-ComfyUI-CoreMLSuite/Makefile
T
aszc-dev 0bbd8d8e0d feat(phase6): opt-in k-means weight palettization (quantize_nbits)
Phase 6 of the modernization plan: add weight palettization to the
Core ML converter as an opt-in knob, so the SD1.5 / SDXL UNet can
ship at 1/2, 1/2.7 or 1/4 of its current size with ANE-friendly
inference.

CoreMLConverter (and the LCM converter) gains a `quantize_nbits`
dropdown: `none` (default — identical to pre-Phase-6 behavior and
filenames, so existing cached .mlpackages still resolve) / `8` / `6` /
`4`. The value is encoded as `_q<bits>` after the attn suffix, so the
unquantized model and the three palettized variants coexist on disk
under distinct cache keys.

Implementation
- core/naming.compose_out_name: accepts `quantize_nbits`, validates
  against {none, 8, 6, 4}, appends `_q<bits>` (none = empty).
- converter.convert_unet: after ct.convert + before .save, runs
  coremltools.optimize.coreml.palettize_weights with
  OpPalettizerConfig(mode="kmeans", nbits=...) when the value is not
  "none". Adds a `Palettization took Xs` log line.
- converter.convert / nodes.CoreMLConverter.convert: pipe the new arg
  through; the ComfyUI node exposes it as a dropdown with default
  "none" so existing workflows are unchanged at load time.
- bench/scripts/convert_sd15.py: QUANT_NBITS env knob; uses the
  pure compose_out_name (replaces the inline string formatter).

Test infra
- tests/unit/test_characterization_out_name.py: 6 new tests pinning
  the `_q<bits>` suffix contract, the "none" passthrough (backward
  compat), the cn + lora + quant combination, and the invalid-value
  ValueError. Total Tier 0 now at 94.
- Makefile gains `bench-quant` (runs the matrix script) and
  `convert-quant` (converts q8, q6, q4 sequentially).
- bench/scripts/quant_matrix.py (new): loads each variant, runs
  REPEATS forward passes with a fixed seed, then computes the
  noise_pred PSNR of each quantized variant against the unquantized
  baseline. Writes bench/results/quant_matrix_<sha>.{json,md}.

README
- New "Quantization (Phase 6, opt-in)" section: tradeoff table
  measured on M2 Pro SD1.5 1x512x512 SPLIT_EINSUM (sizes 1641/822/
  617/412 MB; fwd 197/187/183/180 ms; PSNR 53.5 / 40.2 / 27.5 dB),
  plus per-chip/RAM recommendations.

Default-path safety
- "none" produces the same out_name as Phase 5 -> existing
  v1-5-pruned-emaonly_1x512x512_se_unet.mlmodelc is still picked up
  unchanged; the m2 golden image test continues to anchor.
2026-05-25 01:30:29 +02:00

89 lines
3.3 KiB
Makefile

# ComfyUI-CoreMLSuite — tiered test/bench dispatcher (Phase 4).
#
# Tiers (see MODERNIZATION_SPEC.md):
# Tier 0 (unit): framework-free pure-logic tests, run anywhere in seconds.
# Tier 1 (smoke): macOS-ARM, no ANE, no full model — converts a synthetic
# micro-UNet to catch coremltools / ml-stable-diffusion
# API breakage in minutes.
# Tier 2 (m2): real ANE on Apple Silicon; integration + bench.
#
# COMFY_DIR defaults to the canonical custom-node layout (two dirs up from here).
# PY defaults to the project's uv-managed venv interpreter.
COMFY_DIR ?= $(realpath $(CURDIR)/../..)
PY ?= $(CURDIR)/.venv/bin/python
PYTEST ?= $(PY) -m pytest
UNAME_S := $(shell uname -s)
UNAME_M := $(shell uname -m)
IS_MACOS_ARM := $(filter Darwin,$(UNAME_S))$(filter arm64,$(UNAME_M))
.PHONY: help test-unit test-smoke test-m2 bench bench-rerun ci-tier0 ci-tier1 clean check-macos-arm
help:
@echo "ComfyUI-CoreMLSuite — make targets"
@echo ""
@echo " test-unit Tier 0: pure-logic pytest, runs anywhere, seconds"
@echo " test-smoke Tier 1: synthetic micro-UNet ct.convert + load (macOS-ARM, minutes)"
@echo " test-m2 Tier 2: pytest -m m2 against a real Core ML UNet (Apple Silicon + ANE)"
@echo " bench Run bench/run.py against a converted .mlmodelc"
@echo ""
@echo "Vars: COMFY_DIR (default: $(COMFY_DIR)), PY (default: $(PY))"
check-macos-arm:
@if [ -z "$(IS_MACOS_ARM)" ]; then \
echo "this target requires macOS on Apple Silicon (got $(UNAME_S)/$(UNAME_M))"; \
exit 2; \
fi
# Tier 0 — Linux-safe pure logic. Should not import comfy/coremltools.
test-unit:
$(PYTEST) -m unit tests/
# Tier 1 — macOS-ARM smoke. Converts a synthetic UNet through coremltools to
# catch API breakage without needing a real SD checkpoint or the ANE.
test-smoke: check-macos-arm
$(PYTEST) -m smoke tests/
# Tier 2 — full Apple Silicon path: integration + m2 golden + bench.
# Requires a converted .mlmodelc (see bench/scripts/convert_sd15.py).
test-m2: check-macos-arm
$(PYTEST) -m m2 tests/
# Bench harness. Override MODEL=/path/to/.mlmodelc for an explicit model.
MODEL ?= $(COMFY_DIR)/models/unet/v1-5-pruned-emaonly_1x512x512_se_unet.mlmodelc
COMPUTE_UNITS ?= CPU_AND_NE CPU_AND_GPU
REPEATS ?= 30
ASSUMED_STEPS ?= 20
bench: check-macos-arm
$(PY) bench/run.py \
--model "$(MODEL)" \
--compute-units $(COMPUTE_UNITS) \
--repeats $(REPEATS) \
--assumed-steps $(ASSUMED_STEPS)
# Phase 6: quantization tradeoff matrix. Expects the {none, 8, 6, 4}
# variants to already exist on disk (run `make convert-quant` first or
# QUANT_NBITS=8 .venv/bin/python bench/scripts/convert_sd15.py).
bench-quant: check-macos-arm
$(PY) bench/scripts/quant_matrix.py
convert-quant: check-macos-arm
@for n in 8 6 4; do \
echo "=== converting nbits=$$n ==="; \
QUANT_NBITS=$$n $(PY) bench/scripts/convert_sd15.py || exit 1; \
done
# What CI actually invokes — same as test-unit but echoes the env capture
# alongside so failed runs land with diagnostics.
ci-tier0:
@echo "## env (Tier 0)" && $(PY) --version && uv pip freeze --python "$(PY)" 2>/dev/null | head -50 || true
$(MAKE) test-unit
ci-tier1: check-macos-arm
@echo "## env (Tier 1)" && $(PY) --version && uv pip freeze --python "$(PY)" 2>/dev/null | head -50 || true
$(MAKE) test-smoke
clean:
rm -rf .pytest_cache tests/m2/_latest_generated.png pytestdebug.log