Rewrite the README as a lean landing page and move depth into a docs/ folder. Correct the supported-model story and several stale facts, and answer the recurring questions from issue #21. - Convert-only is the supported path: suite-converted .mlpackage is the only supported input; drop coreml-community download guidance. - Remove all .mlmodelc / Xcode references — compilation was dropped and the loader handles .mlpackage only. - Fix compute-unit name (CPU_AND_NE, not CPU_AND_ANE) and the loader input name (coreml_name). - Document CoreMLSamplerAdvanced (previously undocumented). - Add docs/: hardware, nodes, conversion, workflows, faq, troubleshooting, limitations (with a support matrix). - Note conversion now lives in the coreml-diffusion package. - Remove dev scaffolding specs; ignore *.log, .DS_Store, .claude/.
3.9 KiB
Hardware & Compute Units
This page explains how the suite maps to Apple Silicon hardware, the difference between ANE, GPU, and MPS, and how to choose a compute unit and attention implementation.
ANE vs GPU vs MPS
Three terms get conflated:
- ANE (Apple Neural Engine) — a dedicated ML accelerator on Apple Silicon. Only Core ML can target it; PyTorch cannot. This is the whole reason the suite exists.
- GPU — the Metal GPU. Reachable both by Core ML (as a compute unit) and by PyTorch (via MPS).
- MPS (Metal Performance Shaders) — PyTorch's GPU backend on macOS. This is the path standard ComfyUI nodes use.
This suite uses Core ML compute units only — it never runs the UNet through
PyTorch/MPS. Consequently PYTORCH_ENABLE_MPS_FALLBACK has no effect on these
nodes. It may still matter for the rest of your ComfyUI graph (CLIP, VAE,
samplers on non-Core ML models), but not for Core ML inference itself.
Rough performance picture (SD1.5, maintainer- and user-reported):
- ANE is meaningfully faster than MPS — on the order of 50–100% for SD1.5.
- Core ML on the GPU is only marginally faster than PyTorch/MPS.
So the speedup comes from the Neural Engine, which means it depends on being able to actually run on the ANE (see attention implementations and the SDXL caveat).
Compute units
The compute unit is set on the loader/converter node and tells Core ML which
hardware to use. It is applied when the model is loaded
(coreml_suite/coreml_model.py:22), not during conversion.
| Value | Hardware | Best paired with |
|---|---|---|
CPU_AND_NE (default) |
CPU + Neural Engine | SPLIT_EINSUM / SPLIT_EINSUM_V2 |
CPU_AND_GPU |
CPU + Metal GPU | ORIGINAL |
CPU_ONLY |
CPU only | fallback / debugging |
ALL |
all available hardware | rarely optimal — see below |
Notes:
- Every option includes the CPU; there is no GPU-and-ANE-without-CPU combination.
NEinCPU_AND_NEis the Neural Engine (Apple's enum spells itNE, notANE).CPU_AND_NEis often faster thanALL. Letting Core ML use everything can be slower on non-Max chips, where memory bandwidth is the bottleneck. TryCPU_AND_NEfirst for SD1.5.
Attention implementations
Chosen at conversion time on the Convert Checkpoint to Core ML node. It decides whether the model can run on the ANE:
SPLIT_EINSUM— ANE-friendly attention. Use for the Neural Engine.SPLIT_EINSUM_V2— a variant; in practice ≈SPLIT_EINSUMfor most users.ORIGINAL— standard attention. Runs on the GPU, not the ANE.
The implementation and the compute unit must agree: a SPLIT_EINSUM model wants
CPU_AND_NE; an ORIGINAL model wants CPU_AND_GPU.
Which should I pick?
| Scenario | Attention | Compute unit |
|---|---|---|
| SD1.5 at 512×512 | SPLIT_EINSUM |
CPU_AND_NE |
| SD1.5 at larger sizes (e.g. 768) | ORIGINAL |
CPU_AND_GPU |
| SDXL / SDXL Turbo | ORIGINAL |
CPU_AND_GPU |
Resolution crossover
ANE shines at small latents; the GPU scales better as resolution grows. In user benchmarks:
- At 512×512, ANE +
SPLIT_EINSUMwins by roughly 10% over the GPU path. - At 768×512, GPU +
ORIGINALpulls ahead by roughly 10%, and the larger image is about 2× slower overall.
If you mostly work at 512×512, convert with SPLIT_EINSUM and load on
CPU_AND_NE. If you routinely go larger, an ORIGINAL + GPU model may be
faster.
SDXL and the ANE
SDXL (and SDXL Turbo) cannot run on the ANE — the dual-text-encoder UNet
exceeds what the Neural Engine path supports. SDXL therefore runs at roughly
MPS-equivalent speed with no ANE speedup. If a Core ML SDXL workflow feels no
faster than the standard nodes, this is why. Convert SDXL with ORIGINAL and
load with CPU_AND_GPU or CPU_ONLY. See
limitations for the full picture.