Phase 1 reference numbers, captured against commit ef2a18c on macOS 26.1
with the pinned toolchain (torch 2.0.1, coremltools 8.2, numpy 1.23.5,
python_coreml_stable_diffusion@e5d960c4).
- bench/env/baseline-ef2a18c.txt: full pip freeze + ComfyUI sha + macOS +
resolved ml-stable-diffusion git metadata.
- bench/env/pytest-unit-ef2a18c.txt: pytest tests/unit (test_chunks +
test_controlnet) 20/20 passing.
- bench/results/ef2a18c.{json,md}: SD1.5 1x512x512 SPLIT_EINSUM UNet
forward latency on the Apple Neural Engine and CPU+GPU. NE median 197 ms
/ GPU median 270 ms; a second run reproduced both within 0.3% (noise).
- bench/results/smoke/ef2a18c/: end-to-end Core ML image (E2E-1.5-CoreML
workflow, seed=42) saved by smoke_image.py. The MPS reference branch of
the original workflow is omitted because torch 2.0.1's MPS backend on
macOS 26.1 trips a BFloat16 conversion error in VAEDecode and an
mps.add element-type mismatch in KSampler — both go away with newer
torch and are tracked for the Phase 5 toolchain bump. The server was
started with --cpu-vae to route the VAE through CPU; this is a runtime
flag, not a pin change.