User report (Z-Image, FLUX VAE): blurred over-vibrant output
and pixel-mode crash (kernel > padded input).
CFG: an empty-string CLIPTextEncode negative is a real
encoding, not an uncond — scale amplified a meaningless
(cond-uncond) gap on a guidance-free model. Force
cond_scale=1.0 (ComfyUI's cfg1 skip) when the negative
carries no tokens.
Layout: ComfyUI's VAE boundary is channels-last
(decode -> [B,H,W,3], encode expects it and movedims
internally). Adapters converted to channels-first, so
encode moved the WIDTH axis into channels; pixel_up's
interpolate resized W and C axes instead. Adapters now
pass channels-last through; interpolate and sharpen
convert around their channels-first kernels.