Commit Graph
21 Commits
Author SHA1 Message Date
aszc-dev 0bbd8d8e0d feat(phase6): opt-in k-means weight palettization (quantize_nbits)
Phase 6 of the modernization plan: add weight palettization to the
Core ML converter as an opt-in knob, so the SD1.5 / SDXL UNet can
ship at 1/2, 1/2.7 or 1/4 of its current size with ANE-friendly
inference.

CoreMLConverter (and the LCM converter) gains a `quantize_nbits`
dropdown: `none` (default — identical to pre-Phase-6 behavior and
filenames, so existing cached .mlpackages still resolve) / `8` / `6` /
`4`. The value is encoded as `_q<bits>` after the attn suffix, so the
unquantized model and the three palettized variants coexist on disk
under distinct cache keys.

Implementation
- core/naming.compose_out_name: accepts `quantize_nbits`, validates
  against {none, 8, 6, 4}, appends `_q<bits>` (none = empty).
- converter.convert_unet: after ct.convert + before .save, runs
  coremltools.optimize.coreml.palettize_weights with
  OpPalettizerConfig(mode="kmeans", nbits=...) when the value is not
  "none". Adds a `Palettization took Xs` log line.
- converter.convert / nodes.CoreMLConverter.convert: pipe the new arg
  through; the ComfyUI node exposes it as a dropdown with default
  "none" so existing workflows are unchanged at load time.
- bench/scripts/convert_sd15.py: QUANT_NBITS env knob; uses the
  pure compose_out_name (replaces the inline string formatter).

Test infra
- tests/unit/test_characterization_out_name.py: 6 new tests pinning
  the `_q<bits>` suffix contract, the "none" passthrough (backward
  compat), the cn + lora + quant combination, and the invalid-value
  ValueError. Total Tier 0 now at 94.
- Makefile gains `bench-quant` (runs the matrix script) and
  `convert-quant` (converts q8, q6, q4 sequentially).
- bench/scripts/quant_matrix.py (new): loads each variant, runs
  REPEATS forward passes with a fixed seed, then computes the
  noise_pred PSNR of each quantized variant against the unquantized
  baseline. Writes bench/results/quant_matrix_<sha>.{json,md}.

README
- New "Quantization (Phase 6, opt-in)" section: tradeoff table
  measured on M2 Pro SD1.5 1x512x512 SPLIT_EINSUM (sizes 1641/822/
  617/412 MB; fwd 197/187/183/180 ms; PSNR 53.5 / 40.2 / 27.5 dB),
  plus per-chip/RAM recommendations.

Default-path safety
- "none" produces the same out_name as Phase 5 -> existing
  v1-5-pruned-emaonly_1x512x512_se_unet.mlmodelc is still picked up
  unchanged; the m2 golden image test continues to anchor.
2026-05-25 01:30:29 +02:00
aszc-dev aa60cda09b Add installation using ComfyUI-Manager instructions 2023-11-24 15:10:33 +01:00
aszc-dev 5f7fcd6df3 Add note on SD2.1 to readme 2023-11-24 12:42:33 +01:00
aszc-dev e89cff6d01 Update readme with SDXL info 2023-11-24 12:14:15 +01:00
aszc-dev 9f90083126 Update converter docs and workflows 2023-11-24 12:14:15 +01:00
aszc-dev bb44b4a35f Link to ComfyUI repo 2023-11-24 12:14:15 +01:00
aszc-dev 46d1124573 Update REAMDE.md (Conversion and LoRA) 2023-11-17 22:55:13 +01:00
aszc-dev f9f25fbeb7 Add LCM info to readme 2023-11-13 13:47:15 +01:00
aszc-dev d63df5b62f Remove the controlnet note in readme 2023-10-31 22:05:19 +01:00
aszc-dev 6319d2aedb Add model adapter for unstable compatibility 2023-10-30 11:49:07 +01:00
aszc-dev c043e1f9aa Update ControlNet workflow 2023-10-30 01:01:59 +01:00
aszc-dev 367c6beb19 Update README 2023-10-30 00:41:24 +01:00
aszc 1c85a0e397 Fix conversion options in readme 2023-10-26 12:44:54 +02:00
Adrian Szczepański 721cff74a7 Emphesize the usage of ANE in README 2023-10-26 02:37:57 +02:00
aszc 352809daa4 Add workflows to README 2023-10-26 02:29:20 +02:00
Adrian Szczepański 235ea70b3b Update README.md 2023-10-26 02:25:54 +02:00
aszc f66ea41148 Fix the footnote 2023-10-25 13:30:05 +02:00
Adrian Szczepański fdb75e2675 Update README.md with info on different input sizes 2023-10-24 23:26:30 +02:00
Adrian Szczepański 53ce35794f Update info on LoRA 2023-10-23 15:38:59 +02:00
aszc 5a89dffd2c Update README.md 2023-10-23 15:24:32 +02:00
Adrian Szczepański 556f47bfc5 Add README.md 2023-10-23 15:10:35 +02:00