ci(phase4): tiered test/CI infrastructure (Tier 0/1/2)

Phase 4 of the modernization plan: institutionalize the 3-tier strategy
so future changes are guarded automatically, and pin down the
self-hosted M2 path the maintainer's hardware needs.

Tier dispatch
- Makefile targets test-unit / test-smoke / test-m2 / bench (plus
  ci-tier0 / ci-tier1 wrappers that echo env first). check-macos-arm
  fails fast on non-Apple-Silicon hosts.

Tier 1 smoke
- tests/smoke/test_synthetic_unet.py: builds a TinyUNet (conv-in,
  time/text projections, conv-out), traces it, ct.convert to
  mlprogram + fp16 CPU_ONLY, loads back via CoreMLModel and asserts
  expected_inputs + named output. Runs in ~2s; auto-skips on
  non-Apple-Silicon. Catches coremltools / ml-stable-diffusion API
  drift without needing a real SD checkpoint or the ANE.

GitHub Actions
- .github/workflows/tier0.yml: ubuntu-latest on every push/PR, ~10
  min budget, minimal-deps install (torch==2.0.1, numpy<1.25, pytest)
  -> pytest -m unit.
- .github/workflows/tier1.yml: macos-14 (M1) on push/PR; opt-in via
  run-tier1 label on labeled PRs to spare external-doc PRs.
- .github/workflows/tier2.yml: self-hosted [macOS, ARM64, coreml] on
  PR label run-m2 / nightly cron / workflow_dispatch. Starts ComfyUI
  with --cpu-vae, runs pytest -m m2 + bench/run.py, uploads bench
  results.

Integration coverage moved
- Removed tests/integration/test_basic_conversion_1_5.py: it required
  an MPS reference image (broken on macOS 26 + torch 2.0.1, see
  Phase 1 Gate) and a checkpoint the maintainer doesn't have on disk
  (dreamshaper_8). The same coverage now lives in
  tests/m2/test_golden_image.py: deterministic numerical pass/fail
  (SHA256 + PSNR fallback) against a stored golden, Core ML pipeline
  only. No more human eyeballing.

Docs
- docs/ci-m2.md: one-time runner registration steps, COMFY_DIR
  persistence, baseline model pre-conversion, trigger semantics, what
  to do when the runner is offline, and the migration note from
  integration -> m2 golden.

Sanity check
- Temporarily set convert_to="BREAKAGE_CANARY_NOT_A_REAL_FORMAT" in
  the smoke test; Tier 1 surfaced
  NotImplementedError: Backend converter BREAKAGE_CANARY_NOT_A_REAL_FORMAT not implemented
  immediately. Reverted.

Local verification
- make test-unit -> 88/88 passed in 2.09s
- make test-smoke -> 1/1 passed in 1.99s
This commit is contained in:
aszc-dev
2026-05-23 23:11:35 +02:00
parent 5dafd261b7
commit 8382b13598
8 changed files with 415 additions and 85 deletions
+33
View File
@@ -0,0 +1,33 @@
name: Tier 0 — Unit (Linux)
on:
push:
branches: [main]
pull_request:
# Minimal-deps run: Tier 0 must work without ComfyUI, coremltools, or
# python_coreml_stable_diffusion (Linux CI image won't have them). The
# in-tree purity gate (tests/unit/test_tier0_purity.py) double-checks
# that the suite hasn't started leaking framework imports.
jobs:
unit:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.11"
- name: Install Tier 0 deps
run: |
python -m pip install --upgrade pip
# Pins mirror pyproject (Phase 1 baseline). Tier 0 only needs
# torch + numpy + pytest; everything else is Mac-only.
python -m pip install \
"torch==2.0.1" "numpy<1.25" \
"pytest>=8" "pytest-xdist"
- name: Run Tier 0
run: pytest -m unit tests/ -v
+31
View File
@@ -0,0 +1,31 @@
name: Tier 1 — Smoke (macOS-ARM)
on:
push:
branches: [main]
pull_request:
# Gate behind the run-tier1 label too, so external PRs that touch
# only docs don't burn a minute of macOS-ARM time. Maintainers can
# always re-run via the run-tier1 label.
types: [opened, synchronize, reopened, labeled]
jobs:
smoke:
if: |
github.event_name == 'push' ||
github.event.action != 'labeled' ||
contains(github.event.pull_request.labels.*.name, 'run-tier1')
runs-on: macos-14 # M1, Apple Silicon hosted runner
timeout-minutes: 20
steps:
- uses: actions/checkout@v4
- uses: astral-sh/setup-uv@v3
with:
enable-cache: true
- name: uv sync
run: uv sync --no-install-project
- name: Run Tier 1 (synthetic micro-UNet smoke)
run: uv run pytest -m smoke tests/ -v
+63
View File
@@ -0,0 +1,63 @@
name: Tier 2 — M2 / ANE (self-hosted)
on:
pull_request:
types: [labeled]
schedule:
# Nightly at 04:00 UTC (~05/06 in PL). Keeps the M2 path honest
# without burning the runner on every PR.
- cron: "0 4 * * *"
workflow_dispatch:
jobs:
m2:
if: |
github.event_name == 'schedule' ||
github.event_name == 'workflow_dispatch' ||
(github.event_name == 'pull_request' &&
contains(github.event.pull_request.labels.*.name, 'run-m2'))
# Self-hosted Mac registered by the maintainer. See docs/ci-m2.md
# for runner setup and required model paths.
runs-on: [self-hosted, macOS, ARM64, coreml]
timeout-minutes: 60
steps:
- uses: actions/checkout@v4
- name: uv sync
run: uv sync
- name: Start ComfyUI server (background)
env:
COMFY_DIR: ${{ env.COMFY_DIR }}
run: |
cd "$COMFY_DIR"
nohup "$GITHUB_WORKSPACE/.venv/bin/python" main.py --port 8188 --cpu-vae > /tmp/comfyui-ci.log 2>&1 &
# Wait for readiness, fail fast if it never comes up.
for _ in $(seq 1 60); do
if grep -q "To see the GUI" /tmp/comfyui-ci.log 2>/dev/null; then
echo "comfy ready"; exit 0
fi
sleep 2
done
echo "comfy failed to start"; tail -100 /tmp/comfyui-ci.log; exit 1
- name: Run Tier 2 (m2 marker)
run: uv run pytest -m m2 tests/ -v
- name: Run bench harness
run: |
uv run python bench/run.py \
--model "$COMFY_DIR/models/unet/v1-5-pruned-emaonly_1x512x512_se_unet.mlmodelc" \
--compute-units CPU_AND_NE CPU_AND_GPU \
--repeats 30 --assumed-steps 20
- name: Upload bench results
uses: actions/upload-artifact@v4
if: always()
with:
name: bench-results
path: bench/results/*.json
- name: Stop ComfyUI server
if: always()
run: pkill -f "main.py.*8188" || true