* feat(deps): support ComfyUI's numpy 2 toolchain; make conversion optional The runtime package now installs and runs under numpy 2 / coremltools 9 / torch 2.7 — matching current ComfyUI — without apple/ml-stable-diffusion. - Drop the heavy converter stack (ml-stable-diffusion, diffusers, peft, omegaconf, overrides, transformers) from runtime dependencies; require numpy>=2. - Vendor the runtime pieces: a slim CoreMLModel wrapper around coremltools and the attention-implementation constants. - Lazy-import the converters; the Convert nodes raise a clear error when the legacy conversion dependencies are absent. Loading and sampling existing Core ML models no longer needs them. - CI: Tier 0 tracks the numpy 2 / torch 2.7 toolchain; drop the Tier 2 golden-image lane (it converts at runtime, which now requires the legacy stack) and its fixtures. * feat(conversion): replace apple/ml-stable-diffusion with native diffusers path Reimplement Core ML UNet conversion on top of diffusers instead of the apple/ml-stable-diffusion git dependency, so the full suite (including conversion) installs through ComfyUI Manager without extras on the NumPy 2 toolchain. - Add coreml_suite/conversion package: split-einsum attention processors, a conv2d output-shape helper, Transformer2D trace patches, and a UNet input-adapter wrapper preserving the historical Core ML I/O contract. - Drop python_coreml_stable_diffusion and overrides; route SD15, SDXL, SDXL refiner, and LCM conversion through diffusers UNet2DConditionModel. - Declare diffusers, peft, omegaconf, and transformers as runtime deps. - Add characterization tests asserting split-einsum matches reference attention math; extend the synthetic-UNet smoke test for the wrapper. - Bump to 1.1.0 and set requires-comfyui to a semver constraint (>=0.3.27) so the Comfy Registry publish succeeds. * refactor(conversion)!: native diffusers context layout; drop legacy fallbacks Address PR review feedback: - Drop the legacy converter ImportError fallbacks and LEGACY_CONVERTER_MODULES guards in nodes.py and lcm/nodes.py. Conversion dependencies are mandatory in pyproject, so the indirection is dead code. - Tier 0 CI resolves its toolchain from pyproject via uv (uv sync + uv run) instead of hand-pinned pip installs, removing duplicated version maintenance. - Document the conversion lineage: credit apple/ml-stable-diffusion as the origin, note the implementation has diverged to a native diffusers path, and state the intent to iterate independently. Fix stale README links that pointed users to apple/ml-stable-diffusion for conversion. - Drop the unused `sources` argument from CoreMLModel. BREAKING CHANGE: the converted Core ML UNet now takes encoder_hidden_states in the native diffusers layout (batch, tokens, hidden) instead of (batch, hidden, 1, tokens). This removes the boundary transposes in CoreMLUNetWrapper and CoreMLInputs. Core ML models converted with earlier versions are incompatible and must be re-converted. Bump to 2.0.0. * test: widen split-einsum allclose tolerance for cross-platform float drift The split-einsum attention reorders float32 reductions relative to the reference, so equality holds only up to rounding. The default allclose atol (1e-8) is too tight on Linux x86 BLAS and failed Tier 0 CI; use atol=1e-6 to match the existing chunked-path characterization test. * ci: run macOS smoke tier on the self-hosted Apple Silicon runner GitHub-hosted macOS carries a 10x minute multiplier and exhausts the included Actions minutes too quickly. Move the Tier 1 smoke job onto the self-hosted Apple Silicon runner ([self-hosted, macOS, ARM64, coreml]) so macOS coverage no longer consumes hosted minutes. Tier 0 stays on hosted ubuntu (1x). * ci: fix uv setup for both tiers astral-sh/setup-uv@v3 was retagged and its old commit garbage-collected, so codeload 404s when Actions resolves the stale SHA. Bump Tier 0 (ubuntu) to setup-uv@v7, and drop the action entirely from Tier 1 since the self-hosted runner already provides uv. * test(ci): restore golden-image correctness gate on the self-hosted runner The Tier 2 end-to-end correctness check (real SD1.5 -> Core ML -> image, gated on SHA/PSNR vs a golden) was dropped during the modernization. With the breaking 3D-context change, the synthetic smoke and shape/attention characterization tests no longer cover real-model conversion correctness. Restore tier2.yml (on the [self-hosted, macOS, ARM64, coreml] runner shared with Tier 1), the golden-image test, and the e2e workflow. Adapt the pinned ComfyUI resolution to the requires-comfyui semver tag (vX.Y.Z) instead of a commit SHA, and re-register the m2 marker. The golden is intentionally not committed: the first self-hosted run regenerates it under the new 3D contract and fails for review, per the test's documented bootstrap. * test(ci): add golden image for SD1.5 seed 42 under the 3D-context contract Generated by the first Tier 2 self-hosted run after the native diffusers conversion change. The decoded image is a coherent SD1.5 generation, confirming the (batch, tokens, hidden) Core ML contract produces correct output end-to-end. Subsequent runs gate on this golden (SHA-strict, PSNR fallback). * test(ci): force fresh conversion in Tier 2; drop stale golden The converter skips conversion when a same-named model already exists, keyed on conversion parameters but not the conversion code/toolchain. The self-hosted runner held a pre-existing v1-5 .mlmodelc (4D-context, old toolchain), so the Tier 2 runs reused it (~8-30s) instead of converting — the gate validated a stale model, not the new native diffusers path. Purge the cached Core ML UNets before running so every Tier 2 run does a real convert -> compile -> sample. Drop the golden generated from the stale cache; the next run regenerates it from a genuine 3D-contract conversion and fails for review. * test(ci): add golden image from a genuine 3D-contract conversion Regenerated by a Tier 2 run with the model cache purged, so the converter actually ran (62s, not a cache hit). The fresh model exposes the new 3D encoder_hidden_states input [1, 77, 768], and its decoded SD1.5 seed-42 image is byte-identical to the prior baseline — confirming the native diffusers conversion is behavior-preserving end to end.
135 lines
6.3 KiB
YAML
135 lines
6.3 KiB
YAML
name: Tier 2 — M2 / ANE (self-hosted)
|
|
|
|
on:
|
|
pull_request:
|
|
# `labeled` fires when run-m2 is first added; `synchronize`/`reopened`
|
|
# re-run on every subsequent push while the label is present, so the
|
|
# result tracks the PR head instead of going stale. The `if` below keeps
|
|
# the run gated on the run-m2 label for all pull_request events.
|
|
types: [labeled, synchronize, reopened]
|
|
schedule:
|
|
# Nightly at 04:00 UTC (~05/06 in PL). Keeps the M2 path honest
|
|
# without burning the runner on every PR.
|
|
- cron: "0 4 * * *"
|
|
workflow_dispatch:
|
|
|
|
jobs:
|
|
m2:
|
|
if: |
|
|
github.event_name == 'schedule' ||
|
|
github.event_name == 'workflow_dispatch' ||
|
|
(github.event_name == 'pull_request' &&
|
|
contains(github.event.pull_request.labels.*.name, 'run-m2'))
|
|
# Self-hosted Apple Silicon runner. Prerequisites: COMFY_DIR pointing at
|
|
# a runner-owned ComfyUI clone, plus a cached SD1.5 checkpoint.
|
|
runs-on: [self-hosted, macOS, ARM64, coreml]
|
|
timeout-minutes: 90
|
|
steps:
|
|
- uses: actions/checkout@v4
|
|
|
|
# Hybrid ComfyUI strategy:
|
|
# - schedule (nightly) -> latest origin/master + ComfyUI's own
|
|
# requirements.txt (constrained). Canary for upstream API breakage.
|
|
# - PR label / dispatch -> the requires-comfyui version tag + the frozen
|
|
# `comfy` uv group. Reproducible merge gate, immune to overnight drift.
|
|
- name: Resolve ComfyUI ref + mode
|
|
run: |
|
|
if [ "$GITHUB_EVENT_NAME" = "schedule" ]; then
|
|
echo "COMFY_MODE=latest" >> "$GITHUB_ENV"
|
|
echo "COMFY_REF=master" >> "$GITHUB_ENV"
|
|
else
|
|
# requires-comfyui is a semver constraint (e.g. ">=0.3.27"); pin the
|
|
# gate to the matching ComfyUI release tag (vX.Y.Z).
|
|
VERSION="$(sed -nE 's/^requires-comfyui *= *"[^0-9]*([0-9]+\.[0-9]+\.[0-9]+).*/\1/p' pyproject.toml)"
|
|
if [ -z "$VERSION" ]; then echo "could not parse requires-comfyui from pyproject.toml"; exit 1; fi
|
|
echo "COMFY_MODE=pinned" >> "$GITHUB_ENV"
|
|
echo "COMFY_REF=v$VERSION" >> "$GITHUB_ENV"
|
|
fi
|
|
|
|
- name: Set up ComfyUI checkout
|
|
# COMFY_DIR is exported by the self-hosted runner's .env and MUST be a
|
|
# runner-owned ComfyUI clone (never your dev checkout — this step does
|
|
# git reset --hard and rewrites custom_nodes). Cloned on first run.
|
|
run: |
|
|
set -euo pipefail
|
|
if [ -z "${COMFY_DIR:-}" ]; then echo "COMFY_DIR unset"; exit 1; fi
|
|
# Init-in-place rather than `git clone`: COMFY_DIR may already hold the
|
|
# cached checkpoint (models/checkpoints) or converted .mlmodelc, and
|
|
# `git clone` refuses a non-empty target. init + fetch + `checkout -f`
|
|
# populates the ComfyUI tree while leaving untracked files (the
|
|
# checkpoint, the cached models) untouched — so setup order is free.
|
|
if [ ! -d "$COMFY_DIR/.git" ]; then
|
|
echo "initialising ComfyUI repo in $COMFY_DIR"
|
|
mkdir -p "$COMFY_DIR"
|
|
git -C "$COMFY_DIR" init -q
|
|
fi
|
|
git -C "$COMFY_DIR" remote get-url origin >/dev/null 2>&1 \
|
|
|| git -C "$COMFY_DIR" remote add origin https://github.com/comfyanonymous/ComfyUI.git
|
|
git -C "$COMFY_DIR" fetch --quiet origin
|
|
if [ "$COMFY_MODE" = "latest" ]; then
|
|
git -C "$COMFY_DIR" checkout -f -B master origin/master
|
|
else
|
|
git -C "$COMFY_DIR" checkout -f "$COMFY_REF"
|
|
fi
|
|
COMFY_SHA="$(git -C "$COMFY_DIR" rev-parse HEAD)"
|
|
echo "COMFY_SHA=$COMFY_SHA" >> "$GITHUB_ENV"
|
|
echo "Tier 2 mode=$COMFY_MODE, ComfyUI \`$COMFY_SHA\`" >> "$GITHUB_STEP_SUMMARY"
|
|
|
|
# Point ComfyUI's custom-node loader at this checkout. Refresh the
|
|
# symlink only; refuse to clobber a real directory (guards against a
|
|
# COMFY_DIR that is accidentally a dev checkout).
|
|
NODE_LINK="$COMFY_DIR/custom_nodes/ComfyUI-CoreMLSuite"
|
|
if [ -e "$NODE_LINK" ] && [ ! -L "$NODE_LINK" ]; then
|
|
echo "ERROR: $NODE_LINK is a real directory, not a symlink."
|
|
echo "COMFY_DIR must be a runner-owned ComfyUI, not your dev checkout."
|
|
exit 1
|
|
fi
|
|
mkdir -p "$COMFY_DIR/custom_nodes"
|
|
ln -sfn "$GITHUB_WORKSPACE" "$NODE_LINK"
|
|
|
|
- name: Install dependencies
|
|
run: |
|
|
set -euo pipefail
|
|
if [ "$COMFY_MODE" = "latest" ]; then
|
|
# Node deps (our coremltools-9 toolchain), then ComfyUI's own
|
|
# requirements for the pulled SHA, capped by the toolchain ceiling.
|
|
uv sync
|
|
uv pip install -r "$COMFY_DIR/requirements.txt" \
|
|
-c constraints/comfy-ceiling.txt
|
|
else
|
|
# Pinned gate: the frozen group mirrors the known-good pinned SHA.
|
|
uv sync --group comfy
|
|
fi
|
|
|
|
- name: Start ComfyUI server (background)
|
|
run: |
|
|
cd "$COMFY_DIR"
|
|
nohup "$GITHUB_WORKSPACE/.venv/bin/python" main.py --port 8188 --cpu-vae > /tmp/comfyui-ci.log 2>&1 &
|
|
# Poll the HTTP endpoint for readiness — robust to startup-banner
|
|
# wording / colored-log changes in a floating-latest ComfyUI.
|
|
for _ in $(seq 1 90); do
|
|
if curl -sf -o /dev/null http://127.0.0.1:8188/system_stats; then
|
|
echo "comfy ready (ComfyUI ${COMFY_SHA:-unknown})"; exit 0
|
|
fi
|
|
sleep 2
|
|
done
|
|
echo "comfy failed to start"; tail -100 /tmp/comfyui-ci.log; exit 1
|
|
|
|
- name: Purge cached Core ML UNets (force fresh conversion)
|
|
# The converter skips when a model of the same name already exists. That
|
|
# cache key is conversion *parameters* only, not the conversion code or
|
|
# toolchain — so a stale model would let a conversion regression pass.
|
|
# Clear it so every Tier 2 run exercises the full convert -> compile ->
|
|
# sample path end to end.
|
|
run: |
|
|
rm -rf "$COMFY_DIR"/models/unet/*.mlpackage "$COMFY_DIR"/models/unet/*.mlmodelc || true
|
|
|
|
- name: Run Tier 2 (m2 marker)
|
|
# Drives the Core ML Converter node, which converts the UNet from the
|
|
# checkpoint on every run (cache purged above).
|
|
run: uv run --no-sync pytest -m m2 tests/ -v
|
|
|
|
- name: Stop ComfyUI server
|
|
if: always()
|
|
run: pkill -f "main.py.*8188" || true
|