Phase 6 of the modernization plan: add weight palettization to the
Core ML converter as an opt-in knob, so the SD1.5 / SDXL UNet can
ship at 1/2, 1/2.7 or 1/4 of its current size with ANE-friendly
inference.
CoreMLConverter (and the LCM converter) gains a `quantize_nbits`
dropdown: `none` (default — identical to pre-Phase-6 behavior and
filenames, so existing cached .mlpackages still resolve) / `8` / `6` /
`4`. The value is encoded as `_q<bits>` after the attn suffix, so the
unquantized model and the three palettized variants coexist on disk
under distinct cache keys.
Implementation
- core/naming.compose_out_name: accepts `quantize_nbits`, validates
against {none, 8, 6, 4}, appends `_q<bits>` (none = empty).
- converter.convert_unet: after ct.convert + before .save, runs
coremltools.optimize.coreml.palettize_weights with
OpPalettizerConfig(mode="kmeans", nbits=...) when the value is not
"none". Adds a `Palettization took Xs` log line.
- converter.convert / nodes.CoreMLConverter.convert: pipe the new arg
through; the ComfyUI node exposes it as a dropdown with default
"none" so existing workflows are unchanged at load time.
- bench/scripts/convert_sd15.py: QUANT_NBITS env knob; uses the
pure compose_out_name (replaces the inline string formatter).
Test infra
- tests/unit/test_characterization_out_name.py: 6 new tests pinning
the `_q<bits>` suffix contract, the "none" passthrough (backward
compat), the cn + lora + quant combination, and the invalid-value
ValueError. Total Tier 0 now at 94.
- Makefile gains `bench-quant` (runs the matrix script) and
`convert-quant` (converts q8, q6, q4 sequentially).
- bench/scripts/quant_matrix.py (new): loads each variant, runs
REPEATS forward passes with a fixed seed, then computes the
noise_pred PSNR of each quantized variant against the unquantized
baseline. Writes bench/results/quant_matrix_<sha>.{json,md}.
README
- New "Quantization (Phase 6, opt-in)" section: tradeoff table
measured on M2 Pro SD1.5 1x512x512 SPLIT_EINSUM (sizes 1641/822/
617/412 MB; fwd 197/187/183/180 ms; PSNR 53.5 / 40.2 / 27.5 dB),
plus per-chip/RAM recommendations.
Default-path safety
- "none" produces the same out_name as Phase 5 -> existing
v1-5-pruned-emaonly_1x512x512_se_unet.mlmodelc is still picked up
unchanged; the m2 golden image test continues to anchor.
89 lines
3.3 KiB
Makefile
89 lines
3.3 KiB
Makefile
# ComfyUI-CoreMLSuite — tiered test/bench dispatcher (Phase 4).
|
|
#
|
|
# Tiers (see MODERNIZATION_SPEC.md):
|
|
# Tier 0 (unit): framework-free pure-logic tests, run anywhere in seconds.
|
|
# Tier 1 (smoke): macOS-ARM, no ANE, no full model — converts a synthetic
|
|
# micro-UNet to catch coremltools / ml-stable-diffusion
|
|
# API breakage in minutes.
|
|
# Tier 2 (m2): real ANE on Apple Silicon; integration + bench.
|
|
#
|
|
# COMFY_DIR defaults to the canonical custom-node layout (two dirs up from here).
|
|
# PY defaults to the project's uv-managed venv interpreter.
|
|
|
|
COMFY_DIR ?= $(realpath $(CURDIR)/../..)
|
|
PY ?= $(CURDIR)/.venv/bin/python
|
|
PYTEST ?= $(PY) -m pytest
|
|
|
|
UNAME_S := $(shell uname -s)
|
|
UNAME_M := $(shell uname -m)
|
|
IS_MACOS_ARM := $(filter Darwin,$(UNAME_S))$(filter arm64,$(UNAME_M))
|
|
|
|
.PHONY: help test-unit test-smoke test-m2 bench bench-rerun ci-tier0 ci-tier1 clean check-macos-arm
|
|
|
|
help:
|
|
@echo "ComfyUI-CoreMLSuite — make targets"
|
|
@echo ""
|
|
@echo " test-unit Tier 0: pure-logic pytest, runs anywhere, seconds"
|
|
@echo " test-smoke Tier 1: synthetic micro-UNet ct.convert + load (macOS-ARM, minutes)"
|
|
@echo " test-m2 Tier 2: pytest -m m2 against a real Core ML UNet (Apple Silicon + ANE)"
|
|
@echo " bench Run bench/run.py against a converted .mlmodelc"
|
|
@echo ""
|
|
@echo "Vars: COMFY_DIR (default: $(COMFY_DIR)), PY (default: $(PY))"
|
|
|
|
check-macos-arm:
|
|
@if [ -z "$(IS_MACOS_ARM)" ]; then \
|
|
echo "this target requires macOS on Apple Silicon (got $(UNAME_S)/$(UNAME_M))"; \
|
|
exit 2; \
|
|
fi
|
|
|
|
# Tier 0 — Linux-safe pure logic. Should not import comfy/coremltools.
|
|
test-unit:
|
|
$(PYTEST) -m unit tests/
|
|
|
|
# Tier 1 — macOS-ARM smoke. Converts a synthetic UNet through coremltools to
|
|
# catch API breakage without needing a real SD checkpoint or the ANE.
|
|
test-smoke: check-macos-arm
|
|
$(PYTEST) -m smoke tests/
|
|
|
|
# Tier 2 — full Apple Silicon path: integration + m2 golden + bench.
|
|
# Requires a converted .mlmodelc (see bench/scripts/convert_sd15.py).
|
|
test-m2: check-macos-arm
|
|
$(PYTEST) -m m2 tests/
|
|
|
|
# Bench harness. Override MODEL=/path/to/.mlmodelc for an explicit model.
|
|
MODEL ?= $(COMFY_DIR)/models/unet/v1-5-pruned-emaonly_1x512x512_se_unet.mlmodelc
|
|
COMPUTE_UNITS ?= CPU_AND_NE CPU_AND_GPU
|
|
REPEATS ?= 30
|
|
ASSUMED_STEPS ?= 20
|
|
bench: check-macos-arm
|
|
$(PY) bench/run.py \
|
|
--model "$(MODEL)" \
|
|
--compute-units $(COMPUTE_UNITS) \
|
|
--repeats $(REPEATS) \
|
|
--assumed-steps $(ASSUMED_STEPS)
|
|
|
|
# Phase 6: quantization tradeoff matrix. Expects the {none, 8, 6, 4}
|
|
# variants to already exist on disk (run `make convert-quant` first or
|
|
# QUANT_NBITS=8 .venv/bin/python bench/scripts/convert_sd15.py).
|
|
bench-quant: check-macos-arm
|
|
$(PY) bench/scripts/quant_matrix.py
|
|
|
|
convert-quant: check-macos-arm
|
|
@for n in 8 6 4; do \
|
|
echo "=== converting nbits=$$n ==="; \
|
|
QUANT_NBITS=$$n $(PY) bench/scripts/convert_sd15.py || exit 1; \
|
|
done
|
|
|
|
# What CI actually invokes — same as test-unit but echoes the env capture
|
|
# alongside so failed runs land with diagnostics.
|
|
ci-tier0:
|
|
@echo "## env (Tier 0)" && $(PY) --version && uv pip freeze --python "$(PY)" 2>/dev/null | head -50 || true
|
|
$(MAKE) test-unit
|
|
|
|
ci-tier1: check-macos-arm
|
|
@echo "## env (Tier 1)" && $(PY) --version && uv pip freeze --python "$(PY)" 2>/dev/null | head -50 || true
|
|
$(MAKE) test-smoke
|
|
|
|
clean:
|
|
rm -rf .pytest_cache tests/m2/_latest_generated.png pytestdebug.log
|