Files
aszc-dev-ComfyUI-CoreMLSuite/docs/ci-m2.md
T
aszc-dev 8382b13598 ci(phase4): tiered test/CI infrastructure (Tier 0/1/2)
Phase 4 of the modernization plan: institutionalize the 3-tier strategy
so future changes are guarded automatically, and pin down the
self-hosted M2 path the maintainer's hardware needs.

Tier dispatch
- Makefile targets test-unit / test-smoke / test-m2 / bench (plus
  ci-tier0 / ci-tier1 wrappers that echo env first). check-macos-arm
  fails fast on non-Apple-Silicon hosts.

Tier 1 smoke
- tests/smoke/test_synthetic_unet.py: builds a TinyUNet (conv-in,
  time/text projections, conv-out), traces it, ct.convert to
  mlprogram + fp16 CPU_ONLY, loads back via CoreMLModel and asserts
  expected_inputs + named output. Runs in ~2s; auto-skips on
  non-Apple-Silicon. Catches coremltools / ml-stable-diffusion API
  drift without needing a real SD checkpoint or the ANE.

GitHub Actions
- .github/workflows/tier0.yml: ubuntu-latest on every push/PR, ~10
  min budget, minimal-deps install (torch==2.0.1, numpy<1.25, pytest)
  -> pytest -m unit.
- .github/workflows/tier1.yml: macos-14 (M1) on push/PR; opt-in via
  run-tier1 label on labeled PRs to spare external-doc PRs.
- .github/workflows/tier2.yml: self-hosted [macOS, ARM64, coreml] on
  PR label run-m2 / nightly cron / workflow_dispatch. Starts ComfyUI
  with --cpu-vae, runs pytest -m m2 + bench/run.py, uploads bench
  results.

Integration coverage moved
- Removed tests/integration/test_basic_conversion_1_5.py: it required
  an MPS reference image (broken on macOS 26 + torch 2.0.1, see
  Phase 1 Gate) and a checkpoint the maintainer doesn't have on disk
  (dreamshaper_8). The same coverage now lives in
  tests/m2/test_golden_image.py: deterministic numerical pass/fail
  (SHA256 + PSNR fallback) against a stored golden, Core ML pipeline
  only. No more human eyeballing.

Docs
- docs/ci-m2.md: one-time runner registration steps, COMFY_DIR
  persistence, baseline model pre-conversion, trigger semantics, what
  to do when the runner is offline, and the migration note from
  integration -> m2 golden.

Sanity check
- Temporarily set convert_to="BREAKAGE_CANARY_NOT_A_REAL_FORMAT" in
  the smoke test; Tier 1 surfaced
  NotImplementedError: Backend converter BREAKAGE_CANARY_NOT_A_REAL_FORMAT not implemented
  immediately. Reverted.

Local verification
- make test-unit -> 88/88 passed in 2.09s
- make test-smoke -> 1/1 passed in 1.99s
2026-05-23 23:11:35 +02:00

3.4 KiB

Self-hosted M2 runner — setup

Tier 2 (ANE + integration + bench) runs on a self-hosted GitHub Actions runner registered against the maintainer's M-series Mac. Hosted macOS runners on GitHub do not expose the Apple Neural Engine, so the ANE half of the matrix has to live on real hardware.

One-time runner setup

  1. Install dependencies on the Mac. Python 3.11.x (matching the requires-python pin), uv, git, plus the ComfyUI checkout at the path the workflow expects (default: $HOME/dev/ComfyUI). The workflow reads COMFY_DIR from the runner's env.

    brew install python@3.11 uv git
    
  2. Register the runner. From the repo Settings → Actions → Runners → New self-hosted runner, follow the macOS-ARM instructions. Add the labels exactly: self-hosted, macOS, ARM64, coreml (the workflow runs-on clause requires all four).

    mkdir ~/actions-runner && cd ~/actions-runner
    curl -O -L https://github.com/actions/runner/releases/download/v2.317.0/actions-runner-osx-arm64-2.317.0.tar.gz
    tar xzf actions-runner-osx-arm64-2.317.0.tar.gz
    ./config.sh --url https://github.com/<owner>/<repo> \
                --token <REGISTRATION_TOKEN> \
                --labels self-hosted,macOS,ARM64,coreml \
                --name "$(hostname)-m2"
    ./svc.sh install && ./svc.sh start   # run as a launchd service
    
  3. Persist COMFY_DIR for the runner. The workflow needs to know where the ComfyUI checkout lives. Add it to the runner's .env:

    echo 'COMFY_DIR=/Users/<you>/dev/ComfyUI' >> ~/actions-runner/.env
    
  4. Pre-convert the baseline SD1.5 model. The bench step expects $COMFY_DIR/models/unet/v1-5-pruned-emaonly_1x512x512_se_unet.mlmodelc. Run the conversion once manually:

    cd $GITHUB_WORKSPACE
    uv run python bench/scripts/convert_sd15.py
    

    Re-runs of the same combination are a no-op; the converter skips when the .mlmodelc already exists.

Triggers

The Tier 2 workflow (.github/workflows/tier2.yml) runs:

  • On PR label run-m2 — maintainers add the label to opt a PR into the ANE lane (the runner is not free; default off).
  • Nightly at 04:00 UTC via schedule:.
  • Manually via the workflow_dispatch button.

Artifacts

  • Bench JSON/MD are uploaded as bench-results.
  • M2 golden image regressions surface as test failures in tests/m2/test_golden_image.py; diff PNG is written next to the golden under tests/m2/_latest_generated.png (gitignored).

When the runner is down

If the maintainer's Mac is offline, the workflow queues until the runner comes back. Cancel a stuck run from the Actions UI; the gate is not blocking by default (Tier 0 + Tier 1 carry PR status). Tier 2 is "good to merge once it goes green," not "blocked until then."

Replacing the integration e2e

The legacy tests/integration/test_basic_conversion_1_5.py checked CoreML output against an MPS reference image at PSNR > 25 dB. That reference path is broken on macOS 26.x with torch 2.0.1 (see Phase 1 Gate report). Phase 4 moves the same coverage to tests/m2/test_golden_image.py, which:

  • runs the Core ML pipeline only (no MPS reference),
  • asserts SHA256 against tests/m2/goldens/sd15_seed42.sha256,
  • falls back to PSNR ≥ 40 dB if the hash drifts.

This removes the human-eyeball dependency: a regression is now a numerical fail, not a "looks different to me."