Files
aszc-dev-ComfyUI-CoreMLSuite/tests/unit/test_characterization_inputs.py
aszc 65a2de2fab feat!: modernize toolchain and replace apple/ml-stable-diffusion with native diffusers conversion (#58)
* feat(deps): support ComfyUI's numpy 2 toolchain; make conversion optional

The runtime package now installs and runs under numpy 2 / coremltools 9 /
torch 2.7 — matching current ComfyUI — without apple/ml-stable-diffusion.

- Drop the heavy converter stack (ml-stable-diffusion, diffusers, peft,
  omegaconf, overrides, transformers) from runtime dependencies; require
  numpy>=2.
- Vendor the runtime pieces: a slim CoreMLModel wrapper around coremltools
  and the attention-implementation constants.
- Lazy-import the converters; the Convert nodes raise a clear error when the
  legacy conversion dependencies are absent. Loading and sampling existing
  Core ML models no longer needs them.
- CI: Tier 0 tracks the numpy 2 / torch 2.7 toolchain; drop the Tier 2
  golden-image lane (it converts at runtime, which now requires the legacy
  stack) and its fixtures.

* feat(conversion): replace apple/ml-stable-diffusion with native diffusers path

Reimplement Core ML UNet conversion on top of diffusers instead of the
apple/ml-stable-diffusion git dependency, so the full suite (including
conversion) installs through ComfyUI Manager without extras on the NumPy 2
toolchain.

- Add coreml_suite/conversion package: split-einsum attention processors,
  a conv2d output-shape helper, Transformer2D trace patches, and a UNet
  input-adapter wrapper preserving the historical Core ML I/O contract.
- Drop python_coreml_stable_diffusion and overrides; route SD15, SDXL,
  SDXL refiner, and LCM conversion through diffusers UNet2DConditionModel.
- Declare diffusers, peft, omegaconf, and transformers as runtime deps.
- Add characterization tests asserting split-einsum matches reference
  attention math; extend the synthetic-UNet smoke test for the wrapper.
- Bump to 1.1.0 and set requires-comfyui to a semver constraint (>=0.3.27)
  so the Comfy Registry publish succeeds.

* refactor(conversion)!: native diffusers context layout; drop legacy fallbacks

Address PR review feedback:

- Drop the legacy converter ImportError fallbacks and LEGACY_CONVERTER_MODULES
  guards in nodes.py and lcm/nodes.py. Conversion dependencies are mandatory in
  pyproject, so the indirection is dead code.
- Tier 0 CI resolves its toolchain from pyproject via uv (uv sync + uv run)
  instead of hand-pinned pip installs, removing duplicated version maintenance.
- Document the conversion lineage: credit apple/ml-stable-diffusion as the
  origin, note the implementation has diverged to a native diffusers path, and
  state the intent to iterate independently. Fix stale README links that pointed
  users to apple/ml-stable-diffusion for conversion.
- Drop the unused `sources` argument from CoreMLModel.

BREAKING CHANGE: the converted Core ML UNet now takes encoder_hidden_states in
the native diffusers layout (batch, tokens, hidden) instead of
(batch, hidden, 1, tokens). This removes the boundary transposes in
CoreMLUNetWrapper and CoreMLInputs. Core ML models converted with earlier
versions are incompatible and must be re-converted. Bump to 2.0.0.

* test: widen split-einsum allclose tolerance for cross-platform float drift

The split-einsum attention reorders float32 reductions relative to the
reference, so equality holds only up to rounding. The default allclose atol
(1e-8) is too tight on Linux x86 BLAS and failed Tier 0 CI; use atol=1e-6 to
match the existing chunked-path characterization test.

* ci: run macOS smoke tier on the self-hosted Apple Silicon runner

GitHub-hosted macOS carries a 10x minute multiplier and exhausts the included
Actions minutes too quickly. Move the Tier 1 smoke job onto the self-hosted
Apple Silicon runner ([self-hosted, macOS, ARM64, coreml]) so macOS coverage no
longer consumes hosted minutes. Tier 0 stays on hosted ubuntu (1x).

* ci: fix uv setup for both tiers

astral-sh/setup-uv@v3 was retagged and its old commit garbage-collected, so
codeload 404s when Actions resolves the stale SHA. Bump Tier 0 (ubuntu) to
setup-uv@v7, and drop the action entirely from Tier 1 since the self-hosted
runner already provides uv.

* test(ci): restore golden-image correctness gate on the self-hosted runner

The Tier 2 end-to-end correctness check (real SD1.5 -> Core ML -> image,
gated on SHA/PSNR vs a golden) was dropped during the modernization. With the
breaking 3D-context change, the synthetic smoke and shape/attention
characterization tests no longer cover real-model conversion correctness.

Restore tier2.yml (on the [self-hosted, macOS, ARM64, coreml] runner shared
with Tier 1), the golden-image test, and the e2e workflow. Adapt the pinned
ComfyUI resolution to the requires-comfyui semver tag (vX.Y.Z) instead of a
commit SHA, and re-register the m2 marker. The golden is intentionally not
committed: the first self-hosted run regenerates it under the new 3D contract
and fails for review, per the test's documented bootstrap.

* test(ci): add golden image for SD1.5 seed 42 under the 3D-context contract

Generated by the first Tier 2 self-hosted run after the native diffusers
conversion change. The decoded image is a coherent SD1.5 generation, confirming
the (batch, tokens, hidden) Core ML contract produces correct output
end-to-end. Subsequent runs gate on this golden (SHA-strict, PSNR fallback).

* test(ci): force fresh conversion in Tier 2; drop stale golden

The converter skips conversion when a same-named model already exists, keyed on
conversion parameters but not the conversion code/toolchain. The self-hosted
runner held a pre-existing v1-5 .mlmodelc (4D-context, old toolchain), so the
Tier 2 runs reused it (~8-30s) instead of converting — the gate validated a
stale model, not the new native diffusers path.

Purge the cached Core ML UNets before running so every Tier 2 run does a real
convert -> compile -> sample. Drop the golden generated from the stale cache;
the next run regenerates it from a genuine 3D-contract conversion and fails for
review.

* test(ci): add golden image from a genuine 3D-contract conversion

Regenerated by a Tier 2 run with the model cache purged, so the converter
actually ran (62s, not a cache hit). The fresh model exposes the new 3D
encoder_hidden_states input [1, 77, 768], and its decoded SD1.5 seed-42 image
is byte-identical to the prior baseline — confirming the native diffusers
conversion is behavior-preserving end to end.
2026-05-26 15:46:09 +02:00

229 lines
8.0 KiB
Python

"""Characterization tests for coreml_suite.models.CoreMLInputs.
Locks the shape transforms applied by chunks() and coreml_kwargs() for the
four model variants the suite supports: SD1.5, LCM (SD1.5 + timestep_cond),
SDXL base (time_ids len 6), and SDXL refiner (time_ids len 5).
These contracts feed the Core ML UNet at runtime; if a refactor silently
re-shapes them, generation breaks.
"""
import numpy as np
import pytest
import torch
from coreml_suite.core.inputs import CoreMLInputs
@pytest.fixture(autouse=True)
def _deterministic_seed():
torch.manual_seed(0)
np.random.seed(0)
# ---------- expected_inputs fixtures (mirror real model expectations) -------
SD15_EXPECTED = {
"sample": {"shape": (2, 4, 64, 64)},
"timestep": {"shape": (2,)},
"encoder_hidden_states": {"shape": (2, 77, 768)},
}
SD15_WITH_CN = {
**SD15_EXPECTED,
"additional_residual_0": {"shape": (2, 320, 64, 64)},
"additional_residual_1": {"shape": (2, 640, 32, 32)},
}
LCM_EXPECTED = {
**SD15_EXPECTED,
"timestep_cond": {"shape": (2, 256)},
}
SDXL_BASE_EXPECTED = {
"sample": {"shape": (2, 4, 128, 128)},
"timestep": {"shape": (2,)},
"encoder_hidden_states": {"shape": (2, 77, 2048)},
"time_ids": {"shape": (2, 6)},
"text_embeds": {"shape": (2, 1280)},
}
SDXL_REFINER_EXPECTED = {
"sample": {"shape": (2, 4, 128, 128)},
"timestep": {"shape": (2,)},
"encoder_hidden_states": {"shape": (2, 77, 1280)},
"time_ids": {"shape": (2, 5)},
"text_embeds": {"shape": (2, 1280)},
}
def _sd15_inputs(batch=1, with_control=False, with_ts_cond=False):
x = torch.randn(batch, 4, 64, 64)
t = torch.full((batch,), 999.0)
context = torch.randn(batch, 77, 768)
control = None
if with_control:
control = {
"output": [torch.randn(batch, 320, 64, 64), torch.randn(batch, 640, 32, 32)],
"middle": [],
}
kwargs = {}
if with_ts_cond:
kwargs["timestep_cond"] = torch.randn(batch, 256)
return CoreMLInputs(x, t, context, control, **kwargs)
def _sdxl_inputs(batch=1, refiner=False):
x = torch.randn(batch, 4, 128, 128)
t = torch.full((batch,), 999.0)
ctx_dim = 1280 if refiner else 2048
context = torch.randn(batch, 77, ctx_dim)
time_ids_dim = 5 if refiner else 6
time_ids = torch.randn(batch, time_ids_dim)
text_embeds = torch.randn(batch, 1280)
return CoreMLInputs(
x, t, context, control=None, time_ids=time_ids, text_embeds=text_embeds
)
# ---------- coreml_kwargs ---------------------------------------------------
def test_coreml_kwargs_sd15_shapes_and_fp16():
out = _sd15_inputs(batch=1).coreml_kwargs(SD15_EXPECTED)
assert set(out.keys()) == {"sample", "encoder_hidden_states", "timestep"}
assert out["sample"].shape == (1, 4, 64, 64)
assert out["sample"].dtype == np.float16
# encoder_hidden_states keeps Comfy's native (b, seq, dim) layout.
assert out["encoder_hidden_states"].shape == (1, 77, 768)
assert out["encoder_hidden_states"].dtype == np.float16
assert out["timestep"].shape == (1,)
assert out["timestep"].dtype == np.float16
def test_coreml_kwargs_sd15_with_controlnet_emits_residuals():
inputs = _sd15_inputs(batch=1, with_control=True)
out = inputs.coreml_kwargs(SD15_WITH_CN)
assert "additional_residual_0" in out
assert "additional_residual_1" in out
assert out["additional_residual_0"].shape == (1, 320, 64, 64)
assert out["additional_residual_1"].shape == (1, 640, 32, 32)
def test_coreml_kwargs_sd15_without_controlnet_zero_fills_residuals():
inputs = _sd15_inputs(batch=1, with_control=False)
out = inputs.coreml_kwargs(SD15_WITH_CN)
assert np.all(out["additional_residual_0"] == 0)
assert np.all(out["additional_residual_1"] == 0)
def test_coreml_kwargs_lcm_adds_timestep_cond():
inputs = _sd15_inputs(batch=1, with_ts_cond=True)
out = inputs.coreml_kwargs(LCM_EXPECTED)
assert "timestep_cond" in out
assert out["timestep_cond"].shape == (1, 256)
assert out["timestep_cond"].dtype == np.float16
def test_coreml_kwargs_lcm_skips_timestep_cond_when_not_provided():
"""timestep_cond is only forwarded when the input supplied one — even if
the model's expected_inputs lists it."""
inputs = _sd15_inputs(batch=1, with_ts_cond=False)
out = inputs.coreml_kwargs(LCM_EXPECTED)
assert "timestep_cond" not in out
def test_coreml_kwargs_sdxl_base_emits_time_ids_and_text_embeds():
out = _sdxl_inputs(batch=1, refiner=False).coreml_kwargs(SDXL_BASE_EXPECTED)
assert out["time_ids"].shape == (1, 6)
assert out["text_embeds"].shape == (1, 1280)
assert out["time_ids"].dtype == np.float16
assert out["text_embeds"].dtype == np.float16
def test_coreml_kwargs_sdxl_refiner_uses_len5_time_ids():
out = _sdxl_inputs(batch=1, refiner=True).coreml_kwargs(SDXL_REFINER_EXPECTED)
assert out["time_ids"].shape == (1, 5)
# ---------- chunks ----------------------------------------------------------
def test_chunks_sd15_pad_to_batch2_returns_one_chunk():
chunked = _sd15_inputs(batch=1).chunks(SD15_EXPECTED)
assert len(chunked) == 1
c = chunked[0]
assert c.x.shape == (2, 4, 64, 64)
assert c.t.shape == (2,)
# context shape: (b, seq, dim) padded along batch dim.
assert c.context.shape == (2, 77, 768)
assert c.control is None
assert c.ts_cond is None
assert c.time_ids is None
assert c.text_embeds is None
def test_chunks_sd15_with_controlnet_chunks_residuals_too():
chunked = _sd15_inputs(batch=1, with_control=True).chunks(SD15_EXPECTED)
assert len(chunked) == 1
cn = chunked[0].control
assert cn is not None
assert cn["output"][0].shape == (2, 320, 64, 64)
assert cn["output"][1].shape == (2, 640, 32, 32)
def test_chunks_lcm_carries_timestep_cond_per_chunk():
chunked = _sd15_inputs(batch=1, with_ts_cond=True).chunks(LCM_EXPECTED)
assert len(chunked) == 1
assert chunked[0].ts_cond is not None
assert chunked[0].ts_cond.shape == (2, 256)
def test_chunks_sdxl_base_propagates_time_ids_and_text_embeds():
chunked = _sdxl_inputs(batch=1, refiner=False).chunks(SDXL_BASE_EXPECTED)
assert len(chunked) == 1
c = chunked[0]
assert c.time_ids is not None and c.time_ids.shape == (2, 6)
assert c.text_embeds is not None and c.text_embeds.shape == (2, 1280)
def test_chunks_sdxl_refiner_uses_len5_time_ids():
chunked = _sdxl_inputs(batch=1, refiner=True).chunks(SDXL_REFINER_EXPECTED)
assert chunked[0].time_ids.shape == (2, 5)
def test_chunks_sdxl_synthesizes_zero_time_ids_when_caller_omits():
"""If the model expects time_ids but caller passed nothing, the suite
fabricates a zero-filled tensor. Lock that fallback."""
x = torch.randn(1, 4, 128, 128)
t = torch.full((1,), 999.0)
context = torch.randn(1, 77, 2048)
inputs = CoreMLInputs(x, t, context, control=None)
chunked = inputs.chunks(SDXL_BASE_EXPECTED)
assert chunked[0].time_ids.shape == (2, 6)
assert torch.equal(chunked[0].time_ids, torch.zeros(2, 6))
assert chunked[0].text_embeds.shape == (2, 1280)
assert torch.equal(chunked[0].text_embeds, torch.zeros(2, 1280))
def test_chunks_splits_batch_into_multiple_target2_chunks():
"""batch=5 with target_batch=2 -> 3 chunks (last padded)."""
chunked = _sd15_inputs(batch=5).chunks(SD15_EXPECTED)
assert len(chunked) == 3
for c in chunked:
assert c.x.shape == (2, 4, 64, 64)
assert c.context.shape == (2, 77, 768)
# Last chunk's second batch row is the zero-pad.
assert torch.equal(chunked[-1].x[1], torch.zeros(4, 64, 64))
def test_chunks_timestep_is_broadcast_from_first_value():
"""t is rebuilt from t[0] across all chunks: locks current behavior that
discards any per-row timestep variation."""
x = torch.randn(2, 4, 64, 64)
t = torch.tensor([42.0, 99.0]) # the second value will be lost
context = torch.randn(2, 77, 768)
inputs = CoreMLInputs(x, t, context, control=None)
chunked = inputs.chunks(SD15_EXPECTED)
assert chunked[0].t.shape == (2,)
assert torch.equal(chunked[0].t, torch.full((2,), 42.0))