* feat(deps): support ComfyUI's numpy 2 toolchain; make conversion optional The runtime package now installs and runs under numpy 2 / coremltools 9 / torch 2.7 — matching current ComfyUI — without apple/ml-stable-diffusion. - Drop the heavy converter stack (ml-stable-diffusion, diffusers, peft, omegaconf, overrides, transformers) from runtime dependencies; require numpy>=2. - Vendor the runtime pieces: a slim CoreMLModel wrapper around coremltools and the attention-implementation constants. - Lazy-import the converters; the Convert nodes raise a clear error when the legacy conversion dependencies are absent. Loading and sampling existing Core ML models no longer needs them. - CI: Tier 0 tracks the numpy 2 / torch 2.7 toolchain; drop the Tier 2 golden-image lane (it converts at runtime, which now requires the legacy stack) and its fixtures. * feat(conversion): replace apple/ml-stable-diffusion with native diffusers path Reimplement Core ML UNet conversion on top of diffusers instead of the apple/ml-stable-diffusion git dependency, so the full suite (including conversion) installs through ComfyUI Manager without extras on the NumPy 2 toolchain. - Add coreml_suite/conversion package: split-einsum attention processors, a conv2d output-shape helper, Transformer2D trace patches, and a UNet input-adapter wrapper preserving the historical Core ML I/O contract. - Drop python_coreml_stable_diffusion and overrides; route SD15, SDXL, SDXL refiner, and LCM conversion through diffusers UNet2DConditionModel. - Declare diffusers, peft, omegaconf, and transformers as runtime deps. - Add characterization tests asserting split-einsum matches reference attention math; extend the synthetic-UNet smoke test for the wrapper. - Bump to 1.1.0 and set requires-comfyui to a semver constraint (>=0.3.27) so the Comfy Registry publish succeeds. * refactor(conversion)!: native diffusers context layout; drop legacy fallbacks Address PR review feedback: - Drop the legacy converter ImportError fallbacks and LEGACY_CONVERTER_MODULES guards in nodes.py and lcm/nodes.py. Conversion dependencies are mandatory in pyproject, so the indirection is dead code. - Tier 0 CI resolves its toolchain from pyproject via uv (uv sync + uv run) instead of hand-pinned pip installs, removing duplicated version maintenance. - Document the conversion lineage: credit apple/ml-stable-diffusion as the origin, note the implementation has diverged to a native diffusers path, and state the intent to iterate independently. Fix stale README links that pointed users to apple/ml-stable-diffusion for conversion. - Drop the unused `sources` argument from CoreMLModel. BREAKING CHANGE: the converted Core ML UNet now takes encoder_hidden_states in the native diffusers layout (batch, tokens, hidden) instead of (batch, hidden, 1, tokens). This removes the boundary transposes in CoreMLUNetWrapper and CoreMLInputs. Core ML models converted with earlier versions are incompatible and must be re-converted. Bump to 2.0.0. * test: widen split-einsum allclose tolerance for cross-platform float drift The split-einsum attention reorders float32 reductions relative to the reference, so equality holds only up to rounding. The default allclose atol (1e-8) is too tight on Linux x86 BLAS and failed Tier 0 CI; use atol=1e-6 to match the existing chunked-path characterization test. * ci: run macOS smoke tier on the self-hosted Apple Silicon runner GitHub-hosted macOS carries a 10x minute multiplier and exhausts the included Actions minutes too quickly. Move the Tier 1 smoke job onto the self-hosted Apple Silicon runner ([self-hosted, macOS, ARM64, coreml]) so macOS coverage no longer consumes hosted minutes. Tier 0 stays on hosted ubuntu (1x). * ci: fix uv setup for both tiers astral-sh/setup-uv@v3 was retagged and its old commit garbage-collected, so codeload 404s when Actions resolves the stale SHA. Bump Tier 0 (ubuntu) to setup-uv@v7, and drop the action entirely from Tier 1 since the self-hosted runner already provides uv. * test(ci): restore golden-image correctness gate on the self-hosted runner The Tier 2 end-to-end correctness check (real SD1.5 -> Core ML -> image, gated on SHA/PSNR vs a golden) was dropped during the modernization. With the breaking 3D-context change, the synthetic smoke and shape/attention characterization tests no longer cover real-model conversion correctness. Restore tier2.yml (on the [self-hosted, macOS, ARM64, coreml] runner shared with Tier 1), the golden-image test, and the e2e workflow. Adapt the pinned ComfyUI resolution to the requires-comfyui semver tag (vX.Y.Z) instead of a commit SHA, and re-register the m2 marker. The golden is intentionally not committed: the first self-hosted run regenerates it under the new 3D contract and fails for review, per the test's documented bootstrap. * test(ci): add golden image for SD1.5 seed 42 under the 3D-context contract Generated by the first Tier 2 self-hosted run after the native diffusers conversion change. The decoded image is a coherent SD1.5 generation, confirming the (batch, tokens, hidden) Core ML contract produces correct output end-to-end. Subsequent runs gate on this golden (SHA-strict, PSNR fallback). * test(ci): force fresh conversion in Tier 2; drop stale golden The converter skips conversion when a same-named model already exists, keyed on conversion parameters but not the conversion code/toolchain. The self-hosted runner held a pre-existing v1-5 .mlmodelc (4D-context, old toolchain), so the Tier 2 runs reused it (~8-30s) instead of converting — the gate validated a stale model, not the new native diffusers path. Purge the cached Core ML UNets before running so every Tier 2 run does a real convert -> compile -> sample. Drop the golden generated from the stale cache; the next run regenerates it from a genuine 3D-contract conversion and fails for review. * test(ci): add golden image from a genuine 3D-contract conversion Regenerated by a Tier 2 run with the model cache purged, so the converter actually ran (62s, not a cache hit). The fresh model exposes the new 3D encoder_hidden_states input [1, 77, 768], and its decoded SD1.5 seed-42 image is byte-identical to the prior baseline — confirming the native diffusers conversion is behavior-preserving end to end.
62 lines
1.7 KiB
Python
62 lines
1.7 KiB
Python
from types import MethodType
|
|
|
|
from diffusers.models.transformers.transformer_2d import Transformer2DModel
|
|
|
|
|
|
def prepare_unet_for_coreml_trace(unet):
|
|
for module in unet.modules():
|
|
if isinstance(module, Transformer2DModel):
|
|
module._operate_on_continuous_inputs = MethodType(
|
|
_operate_on_continuous_inputs,
|
|
module,
|
|
)
|
|
module._get_output_for_continuous_inputs = MethodType(
|
|
_get_output_for_continuous_inputs,
|
|
module,
|
|
)
|
|
return unet
|
|
|
|
|
|
def _operate_on_continuous_inputs(self, hidden_states):
|
|
hidden_states = self.norm(hidden_states)
|
|
|
|
if not self.use_linear_projection:
|
|
hidden_states = self.proj_in(hidden_states)
|
|
inner_dim = self.inner_dim
|
|
hidden_states = hidden_states.flatten(2).transpose(1, 2)
|
|
else:
|
|
inner_dim = hidden_states.shape[1]
|
|
hidden_states = hidden_states.flatten(2).transpose(1, 2)
|
|
hidden_states = self.proj_in(hidden_states)
|
|
|
|
return hidden_states, inner_dim
|
|
|
|
|
|
def _get_output_for_continuous_inputs(
|
|
self,
|
|
hidden_states,
|
|
residual,
|
|
batch_size,
|
|
height,
|
|
width,
|
|
inner_dim,
|
|
):
|
|
if not self.use_linear_projection:
|
|
hidden_states = hidden_states.transpose(1, 2).reshape(
|
|
batch_size,
|
|
inner_dim,
|
|
height,
|
|
width,
|
|
)
|
|
hidden_states = self.proj_out(hidden_states)
|
|
else:
|
|
hidden_states = self.proj_out(hidden_states)
|
|
hidden_states = hidden_states.transpose(1, 2).reshape(
|
|
batch_size,
|
|
inner_dim,
|
|
height,
|
|
width,
|
|
)
|
|
|
|
return hidden_states + residual
|