* feat(deps): support ComfyUI's numpy 2 toolchain; make conversion optional The runtime package now installs and runs under numpy 2 / coremltools 9 / torch 2.7 — matching current ComfyUI — without apple/ml-stable-diffusion. - Drop the heavy converter stack (ml-stable-diffusion, diffusers, peft, omegaconf, overrides, transformers) from runtime dependencies; require numpy>=2. - Vendor the runtime pieces: a slim CoreMLModel wrapper around coremltools and the attention-implementation constants. - Lazy-import the converters; the Convert nodes raise a clear error when the legacy conversion dependencies are absent. Loading and sampling existing Core ML models no longer needs them. - CI: Tier 0 tracks the numpy 2 / torch 2.7 toolchain; drop the Tier 2 golden-image lane (it converts at runtime, which now requires the legacy stack) and its fixtures. * feat(conversion): replace apple/ml-stable-diffusion with native diffusers path Reimplement Core ML UNet conversion on top of diffusers instead of the apple/ml-stable-diffusion git dependency, so the full suite (including conversion) installs through ComfyUI Manager without extras on the NumPy 2 toolchain. - Add coreml_suite/conversion package: split-einsum attention processors, a conv2d output-shape helper, Transformer2D trace patches, and a UNet input-adapter wrapper preserving the historical Core ML I/O contract. - Drop python_coreml_stable_diffusion and overrides; route SD15, SDXL, SDXL refiner, and LCM conversion through diffusers UNet2DConditionModel. - Declare diffusers, peft, omegaconf, and transformers as runtime deps. - Add characterization tests asserting split-einsum matches reference attention math; extend the synthetic-UNet smoke test for the wrapper. - Bump to 1.1.0 and set requires-comfyui to a semver constraint (>=0.3.27) so the Comfy Registry publish succeeds. * refactor(conversion)!: native diffusers context layout; drop legacy fallbacks Address PR review feedback: - Drop the legacy converter ImportError fallbacks and LEGACY_CONVERTER_MODULES guards in nodes.py and lcm/nodes.py. Conversion dependencies are mandatory in pyproject, so the indirection is dead code. - Tier 0 CI resolves its toolchain from pyproject via uv (uv sync + uv run) instead of hand-pinned pip installs, removing duplicated version maintenance. - Document the conversion lineage: credit apple/ml-stable-diffusion as the origin, note the implementation has diverged to a native diffusers path, and state the intent to iterate independently. Fix stale README links that pointed users to apple/ml-stable-diffusion for conversion. - Drop the unused `sources` argument from CoreMLModel. BREAKING CHANGE: the converted Core ML UNet now takes encoder_hidden_states in the native diffusers layout (batch, tokens, hidden) instead of (batch, hidden, 1, tokens). This removes the boundary transposes in CoreMLUNetWrapper and CoreMLInputs. Core ML models converted with earlier versions are incompatible and must be re-converted. Bump to 2.0.0. * test: widen split-einsum allclose tolerance for cross-platform float drift The split-einsum attention reorders float32 reductions relative to the reference, so equality holds only up to rounding. The default allclose atol (1e-8) is too tight on Linux x86 BLAS and failed Tier 0 CI; use atol=1e-6 to match the existing chunked-path characterization test. * ci: run macOS smoke tier on the self-hosted Apple Silicon runner GitHub-hosted macOS carries a 10x minute multiplier and exhausts the included Actions minutes too quickly. Move the Tier 1 smoke job onto the self-hosted Apple Silicon runner ([self-hosted, macOS, ARM64, coreml]) so macOS coverage no longer consumes hosted minutes. Tier 0 stays on hosted ubuntu (1x). * ci: fix uv setup for both tiers astral-sh/setup-uv@v3 was retagged and its old commit garbage-collected, so codeload 404s when Actions resolves the stale SHA. Bump Tier 0 (ubuntu) to setup-uv@v7, and drop the action entirely from Tier 1 since the self-hosted runner already provides uv. * test(ci): restore golden-image correctness gate on the self-hosted runner The Tier 2 end-to-end correctness check (real SD1.5 -> Core ML -> image, gated on SHA/PSNR vs a golden) was dropped during the modernization. With the breaking 3D-context change, the synthetic smoke and shape/attention characterization tests no longer cover real-model conversion correctness. Restore tier2.yml (on the [self-hosted, macOS, ARM64, coreml] runner shared with Tier 1), the golden-image test, and the e2e workflow. Adapt the pinned ComfyUI resolution to the requires-comfyui semver tag (vX.Y.Z) instead of a commit SHA, and re-register the m2 marker. The golden is intentionally not committed: the first self-hosted run regenerates it under the new 3D contract and fails for review, per the test's documented bootstrap. * test(ci): add golden image for SD1.5 seed 42 under the 3D-context contract Generated by the first Tier 2 self-hosted run after the native diffusers conversion change. The decoded image is a coherent SD1.5 generation, confirming the (batch, tokens, hidden) Core ML contract produces correct output end-to-end. Subsequent runs gate on this golden (SHA-strict, PSNR fallback). * test(ci): force fresh conversion in Tier 2; drop stale golden The converter skips conversion when a same-named model already exists, keyed on conversion parameters but not the conversion code/toolchain. The self-hosted runner held a pre-existing v1-5 .mlmodelc (4D-context, old toolchain), so the Tier 2 runs reused it (~8-30s) instead of converting — the gate validated a stale model, not the new native diffusers path. Purge the cached Core ML UNets before running so every Tier 2 run does a real convert -> compile -> sample. Drop the golden generated from the stale cache; the next run regenerates it from a genuine 3D-contract conversion and fails for review. * test(ci): add golden image from a genuine 3D-contract conversion Regenerated by a Tier 2 run with the model cache purged, so the converter actually ran (62s, not a cache hit). The fresh model exposes the new 3D encoder_hidden_states input [1, 77, 768], and its decoded SD1.5 seed-42 image is byte-identical to the prior baseline — confirming the native diffusers conversion is behavior-preserving end to end.
187 lines
6.5 KiB
Python
187 lines
6.5 KiB
Python
"""Characterization tests for coreml_suite.controlnet.
|
|
|
|
Locks shapes + dtypes + zero-fill behavior of expand_inputs / no_control /
|
|
extract_residual_kwargs / chunk_control. These pure helpers feed the Core ML
|
|
UNet's additional_residual_N inputs; any drift here silently breaks
|
|
ControlNet-based workflows.
|
|
"""
|
|
import numpy as np
|
|
import pytest
|
|
import torch
|
|
|
|
from coreml_suite.core.controlnet import (
|
|
chunk_control,
|
|
expand_inputs,
|
|
extract_residual_kwargs,
|
|
no_control,
|
|
)
|
|
|
|
|
|
@pytest.fixture(autouse=True)
|
|
def _deterministic_seed():
|
|
torch.manual_seed(0)
|
|
np.random.seed(0)
|
|
|
|
|
|
SD15_RESIDUAL_SPEC = {
|
|
"additional_residual_0": {"shape": (2, 320, 64, 64)},
|
|
"additional_residual_1": {"shape": (2, 640, 32, 32)},
|
|
"additional_residual_2": {"shape": (2, 1280, 8, 8)},
|
|
}
|
|
NON_RESIDUAL_SPEC = {
|
|
"sample": {"shape": (2, 4, 64, 64)},
|
|
"encoder_hidden_states": {"shape": (2, 77, 768)},
|
|
}
|
|
|
|
|
|
# ---------- expand_inputs ----------------------------------------------------
|
|
|
|
|
|
def test_expand_inputs_doubles_singleton_numpy():
|
|
inputs = {"a": np.ones((1, 4), dtype=np.float32)}
|
|
out = expand_inputs(inputs)
|
|
assert out["a"].shape == (2, 4)
|
|
assert np.array_equal(out["a"], np.ones((2, 4)))
|
|
|
|
|
|
def test_expand_inputs_doubles_singleton_torch():
|
|
inputs = {"a": torch.ones(1, 4)}
|
|
out = expand_inputs(inputs)
|
|
assert out["a"].shape == (2, 4)
|
|
assert torch.equal(out["a"], torch.ones(2, 4))
|
|
|
|
|
|
def test_expand_inputs_doubles_singleton_list():
|
|
inputs = {"a": [42]}
|
|
out = expand_inputs(inputs)
|
|
assert out["a"] == [42, 42]
|
|
|
|
|
|
def test_expand_inputs_skips_already_batched():
|
|
"""batch > 1 inputs are returned unchanged (same object identity)."""
|
|
arr = np.ones((2, 4), dtype=np.float32)
|
|
tensor = torch.ones(3, 4)
|
|
lst = [1, 2]
|
|
out = expand_inputs({"a": arr, "b": tensor, "c": lst})
|
|
assert out["a"] is arr
|
|
assert out["b"] is tensor
|
|
assert out["c"] is lst
|
|
|
|
|
|
def test_expand_inputs_preserves_unknown_value_types():
|
|
# Strings/None pass through untouched — locks current permissive contract.
|
|
inputs = {"s": "hello", "none": None, "int": 7}
|
|
out = expand_inputs(inputs)
|
|
assert out == {"s": "hello", "none": None, "int": 7}
|
|
|
|
|
|
# ---------- no_control -------------------------------------------------------
|
|
|
|
|
|
def test_no_control_returns_zero_fp16_for_residuals():
|
|
out = no_control({**SD15_RESIDUAL_SPEC, **NON_RESIDUAL_SPEC})
|
|
# Only additional_residual_* keys are produced.
|
|
assert set(out.keys()) == set(SD15_RESIDUAL_SPEC.keys())
|
|
for key, spec in SD15_RESIDUAL_SPEC.items():
|
|
arr = out[key]
|
|
assert arr.shape == spec["shape"]
|
|
assert arr.dtype == np.float16
|
|
assert np.all(arr == 0)
|
|
|
|
|
|
def test_no_control_returns_empty_when_no_residuals():
|
|
out = no_control(NON_RESIDUAL_SPEC)
|
|
assert out == {}
|
|
|
|
|
|
# ---------- extract_residual_kwargs -----------------------------------------
|
|
|
|
|
|
def test_extract_residual_kwargs_empty_when_model_has_no_residual_inputs():
|
|
out = extract_residual_kwargs(NON_RESIDUAL_SPEC, control={"output": [], "middle": []})
|
|
assert out == {}
|
|
|
|
|
|
def test_extract_residual_kwargs_none_control_returns_no_control_shapes():
|
|
out = extract_residual_kwargs(SD15_RESIDUAL_SPEC, control=None)
|
|
assert set(out.keys()) == set(SD15_RESIDUAL_SPEC.keys())
|
|
for key, spec in SD15_RESIDUAL_SPEC.items():
|
|
assert out[key].shape == spec["shape"]
|
|
assert out[key].dtype == np.float16
|
|
assert np.all(out[key] == 0)
|
|
|
|
|
|
def test_extract_residual_kwargs_flattens_output_then_middle_and_casts_fp16():
|
|
"""output residuals come first (indexed 0..N-1), then middle residuals
|
|
(indexed N..M-1). Values come out of CPU as fp16 numpy arrays."""
|
|
control = {
|
|
"output": [torch.ones(2, 320, 64, 64) * 0.5, torch.ones(2, 640, 32, 32) * 2.0],
|
|
"middle": [torch.ones(2, 1280, 8, 8) * -1.0],
|
|
}
|
|
out = extract_residual_kwargs(SD15_RESIDUAL_SPEC, control)
|
|
assert set(out.keys()) == {"additional_residual_0", "additional_residual_1", "additional_residual_2"}
|
|
assert out["additional_residual_0"].shape == (2, 320, 64, 64)
|
|
assert out["additional_residual_1"].shape == (2, 640, 32, 32)
|
|
assert out["additional_residual_2"].shape == (2, 1280, 8, 8)
|
|
for arr in out.values():
|
|
assert arr.dtype == np.float16
|
|
# Locked order: index 0 == first output residual (0.5), index 2 == middle (-1.0).
|
|
assert np.allclose(out["additional_residual_0"], 0.5)
|
|
assert np.allclose(out["additional_residual_1"], 2.0)
|
|
assert np.allclose(out["additional_residual_2"], -1.0)
|
|
|
|
|
|
# ---------- chunk_control ----------------------------------------------------
|
|
|
|
|
|
def test_chunk_control_none_returns_list_of_nones_with_length_target():
|
|
"""`no_control` path: when there's no control, you get [None] * target_size
|
|
(NOT [None, None] regardless of target — this is the contract today)."""
|
|
assert chunk_control(None, 1) == [None]
|
|
assert chunk_control(None, 2) == [None, None]
|
|
assert chunk_control(None, 4) == [None, None, None, None]
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"batch,target,expected_chunks",
|
|
[(1, 2, 1), (2, 2, 1), (3, 2, 2), (4, 2, 2), (5, 3, 2), (9, 4, 3)],
|
|
)
|
|
def test_chunk_control_shapes_after_chunking(batch, target, expected_chunks):
|
|
cn = {
|
|
"output": [
|
|
torch.randn(batch, 320, 64, 64),
|
|
torch.randn(batch, 640, 32, 32),
|
|
],
|
|
"middle": [torch.randn(batch, 1280, 8, 8)],
|
|
}
|
|
chunks = chunk_control(cn, target)
|
|
assert len(chunks) == expected_chunks
|
|
for c in chunks:
|
|
assert c["output"][0].shape == (target, 320, 64, 64)
|
|
assert c["output"][1].shape == (target, 640, 32, 32)
|
|
assert c["middle"][0].shape == (target, 1280, 8, 8)
|
|
|
|
|
|
def test_chunk_control_preserves_keys_order():
|
|
"""Output dicts contain exactly {"output", "middle"} in that order."""
|
|
cn = {
|
|
"output": [torch.zeros(2, 4, 4, 4)],
|
|
"middle": [torch.zeros(2, 4, 4, 4)],
|
|
}
|
|
chunks = chunk_control(cn, 2)
|
|
assert list(chunks[0].keys()) == ["output", "middle"]
|
|
|
|
|
|
def test_chunk_control_zero_pads_remainder():
|
|
"""A batch=3, target=2 split puts the third row alongside a zero row."""
|
|
cn = {
|
|
"output": [torch.arange(3 * 4).reshape(3, 1, 2, 2).float()],
|
|
"middle": [torch.arange(3 * 4).reshape(3, 1, 2, 2).float()],
|
|
}
|
|
chunks = chunk_control(cn, 2)
|
|
assert len(chunks) == 2
|
|
last_out = chunks[-1]["output"][0]
|
|
# First row is the original third row; second row is padding zeros.
|
|
assert torch.equal(last_out[0], cn["output"][0][2])
|
|
assert torch.equal(last_out[1], torch.zeros(1, 2, 2))
|