S5 of the 2026-09-08 layout plan: tests/test_structure.py grows the full
guard set — nodes/ exclusively node classes, src/ schema-free, category
WMNodes/image pinned per module, no legacy *_node.py in src/, git-ls-files
grep guard for old category strings, and the ComfyUI loader regression
(replicates load_custom_node spec_from_file_location mechanics, executes the
real entry, asserts comfy_entrypoint → 8-class get_node_list + on_load
installs the Qwen2D patch).
S1 of the 2026-09-08 layout plan: DyPE/SEGA/SPA/HAP node classes now live
in nodes/{dype,sega,spa,hap}.py with engine imports via dual-form try/except
(loader-relative ..src / flat src). Entry __init__.py keeps extension
registration only. nodes/__init__.py re-exports the full 8-class surface
(transitional src/*_node re-exports until S2/S3 land). All patch nodes
category -> WMNodes/image. New tests/test_structure.py guards the layout.
User report: torch.cat size mismatch in krea2 model forward —
'Expected size 1 but got size 16'. Root cause: v2.12.1's
model-space noising called process_latent_in on the 4D core
tensor; Wan21 mean/std stats are [1,C,1,1,1], so 4D-vs-5D
broadcast SILENTLY makes [B,C,C,H,W] garbage (T=16=channels) —
exactly the plan risk-register trap; my S3 bridge covered the
model-call adapter but not the cascade noising conversions.
Fix: the node wraps process_in/out NDIM-TRANSPARENTLY for
latent_dimensions=3 — unsqueeze to 5D, convert in true model
space, squeeze back. The cascade sigma-mix stays 4D with
correctly-normalized values (first wrapper version returned
5D and re-triggered the broadcast in the mix — caught by the
strengthened crown test).
Mock process_latent_in now uses Wan21-faithful [1,C,1,1,1]
stats (the old affine mock masked this bug class). 171 hiflow
tests; suite 1323 passed / 4 skipped; ruff F,E9 green; v2.14.1.
S6 of plan 2026-09-07: README method table + node-family line +
HiFlow section list Krea2/Qwen-Image honestly (Qwen-Image was
gate-blocked since v2.10.0); tip notes Wan21 single-frame support
and multi-frame rejection; changelog v2.14.0; pyproject bumped.
Full suite 1323 passed / 4 skipped; ruff F,E9 green.
S4 of plan 2026-09-07: _make_vae_adapters speaks the latent_dim=3
Qwen-VAE boundary — decode unsqueezes 4D latents to 5D for the raw
vae.decode and slices the channels-last [B,T,H,W,3] image to its
first frame; encode feeds the channels-last 4D image (the VAE
unsqueezes internally per not_video, sd.py:1342-1346) and slices
the 5D latent to 4D for the core. _sharpen gains a defensive 5D
frame-slice backstop. Fake 3D-VAE mirrors the REAL sd.py shapes;
4 new tests; 51 node tests green.
S3 of plan 2026-09-07: _make_predict_x0 gains latent_dimensions;
with 3 (Wan21) it unsqueezes 4D->5D [B,C,1,H,W] before
process_latent_in (the Wan21 mean/std stats are [1,C,1,1,1] views —
5D broadcast only) and the sampling_function call, squeezes the x0
back to 4D for the core (PixelRush predict_eps was_4d pattern).
2 new tests (5D pin + inert-at-2D pin); 47 node tests green.
S1 of plan 2026-09-07 (Krea2/Qwen-Image 3D-latent support):
- hiflow_cascade asserts 4D [B,C,H,W] input; 5D (Wan21-style or
multi-frame) raises a clear error pointing at the node-layer bridge
- stage seed built from the current chain latent's shape, not the
input's (construction-local; batch/channels pinned by test)
- 3 new tests; 118 hiflow core tests green
Fixes NotImplementedError: compute_index_ranges_weights not implemented for 'Half'
when running FreeScale on FP16/BF16 models.
Same pattern as the PixelRush float16 fix from v2.8.2.
User: absolute pixel target unintuitive and vague. New param is
relative to the input latent: 2 = double each side, 1 = unchanged,
0.5 = half. Semantics:
- scale applies PER SIDE (old absolute form over-upscaled the short
side of non-square inputs — both sides doubled until the smaller
hit the pixel target)
- up >1 keeps paper 2x-stage quantization: 1.2-2.0 -> one 2x stage,
4 -> two; final may land above exact scale
- scale <1 (0.25-1): single guided refinement stage at the smaller
size; reference trajectory bicubic-resized down, walk unchanged
- range [0.25, 8] validated at node + core
_stage_latent_sizes rewritten (nearest-snap helper split out);
cascade + node params renamed; workflow JSON updated; 158 hiflow
tests (new: downscale stage, per-side non-square, range, scale-2
single stage); suite 1310 passed / 4 skipped; ruff F,E9 green;
v2.13.0.
Z-Image round 3: img2img drastically changed the image at any usable
denoise; only denoise=0.05 looked right. Root cause: the sigma mix
ran in VAE space. ComfyUI noises in MODEL space (samplers.py:1223
process_latent_in on content, :993 sigma mix) — mixing in VAE space
and converting the mixture scales noise by the latent format's
scale_factor (Flux/Z-Image 0.3611 -> 2.77x under-noising) plus shift
offsets. Model sees an input far noisier than the sigma claims, so
it 'corrects' aggressively. Same G8-class bug as PixelRush v2.9.0.
Both noise-injection points fixed: base img2img start and guided
stage init now take process_latent_in/out pairs, mix in model space,
convert back. Conversions injected from the node's inner_model;
None falls back to the plain VAE-space mix (identity formats).
Sampler/scheduler were already faithful (rectified-flow Euler,
model-table simple spacing) — recorded in plan addendum 2.
148 hiflow tests; suite 1300 passed / 4 skipped; ruff F,E9 green;
v2.12.1.
Z-Image report: connecting a real sampler latent did nothing. Root
cause: full flow schedule starts at sigma 1, so the noised start
sigma*eps + (1-sigma)*content zeroes the content weight entirely.
Fix follows comfy KSampler set_steps convention: denoise<1 computes
new_steps=int(steps/denoise) on a denser schedule and keeps the last
(steps+1) sigmas — entry lands below 1 and the content survives with
weight (1-sigma_start). denoise_sigmas densifies by interpolating the
sigma grid in flow-time (node passes one concrete schedule, not the
model table). Empty latents always run the full schedule (zeros carry
no content); node warns on the content+denoise=1 foot-gun.
16 new tests (145 hiflow total); suite 1297 passed / 4 skipped;
ruff F,E9 gate green; v2.12.0.
Z-Image report: burned blurred output identical in both upsampling
modes; tau=0.95 better than tau=0.05 (inverted). Root cause: base
stage sampled EmptySD3LatentImage zeros verbatim as noise — whole
reference trajectory off-distribution. Code-vs-theory audit of
flux_pipeline_hiflow.py found 4 more divergences from plan 2026-09-03:
- D1 base start noised: sigma[0]*eps + (1-sigma[0])*latent (sigma[0]=1
-> pure noise, matches reference randn start); noise_seed input
- D2 init anchors on previous chain's FINAL image, always pixel
round-tripped (decode->bicubic->sharpen->encode) regardless of
upsampling mode; per-step refs stay time-matched
- D3 v_ref = (x - ref_x0)/sigma from the walk's own state (reference
model_output_ref), not a separately-integrated chain
- D4 trajectories store RAW pre-correction x0 (original_pred_x0) —
guidance no longer compounds across cascade stages
- D6 alpha/beta linear-in-index (n-i)/n like the code, not paper
sigma/sigma_entry (over-locks low freqs late on shifted schedules)
129 hiflow tests rewritten/pinned; full suite 1281 passed / 4
skipped; ruff F,E9 gate green; v2.11.0.
User report (Z-Image, FLUX VAE): blurred over-vibrant output
and pixel-mode crash (kernel > padded input).
CFG: an empty-string CLIPTextEncode negative is a real
encoding, not an uncond — scale amplified a meaningless
(cond-uncond) gap on a guidance-free model. Force
cond_scale=1.0 (ComfyUI's cfg1 skip) when the negative
carries no tokens.
Layout: ComfyUI's VAE boundary is channels-last
(decode -> [B,H,W,3], encode expects it and movedims
internally). Adapters converted to channels-first, so
encode moved the WIDTH axis into channels; pixel_up's
interpolate resized W and C axes instead. Adapters now
pass channels-last through; interpolate and sharpen
convert around their channels-first kernels.
Step 7 of plan 2026-09-03. Gate runs before any model call; 5D
(video) latents are rejected with a PixelRush pointer. Pixel mode
sharpens via freescale's Gaussian blur (zero-pad aware).
Step 6 of plan 2026-09-03. x0 comes from sampling_function
(denoised output with CFG/areas/hooks — apply_model applies
calculate_denoised), not a hand-rolled cond runner; conditioning
is prepared once per latent shape. Non-flow and 3D-latent models
are rejected with a pointer to PixelRush.
Step 5 of plan 2026-09-03. Stage sizes double per stage until the
pixel target (16px-multiple snap for odd bases); steps_per_stage is
an upper bound — the stage only walks schedule sigmas below tau.
Step 4 of plan 2026-09-03. Init seeds from the time-matched
reference x0 (theory form; the authors' code seeds from the final
image encode). v_ref integrates the reference chain's own state,
not the high-res state. Init noise takes an explicit generator so
seeds are reproducible regardless of RNG-stream leftovers.
Step 3 of plan 2026-09-03. Entries park on CPU (a 30-step 4K
trajectory is ~1.5 GB fp32); time matching uses nearest-sigma
within tolerance, robust across differently-spaced schedules.
Step 2 of plan 2026-09-03. Entry clamps to the largest schedule
sigma <= tau so the stage never starts above tau (the reference's
[-n:] slice can); tail keeps the schedule's own spacing.
Dead leftover from the Step 8 test scaffolding (TestEmptyConditioningCFG)
tripped the CI correctness lint gate (ruff --select F,E9) on both
Python matrix entries. Verified locally: ruff clean + the CI pytest
marker expression green (1132 passed).
Succinct "Which method when?" table in the Nodes section: models,
mechanism, takes-your-image, output character for all six methods,
plus a two-family framing (model patches for native generation vs
cascades for refinement) and a quick picker tip.
Also adds ImageCompare to the workflow-test core-node allowlist (a real
ComfyUI core node in comfy_extras/nodes_image_compare.py that the new
DyPE-SDXL example workflow uses).
User report: "structure similar to the original raw image, but
completely noisy - soft non-uniform patches all over", unchanged since
pre-2.9 and across VAEs. Root cause: the corrected doc reference order
slerp(eps_pred, eps_random, 0.95) makes the injected eps 99.6% pure
random at real scales (per-pixel noise std 1.17 vs signal 1.0; mock
correlation with clean signal: 0.65 - structure visible through heavy
Gaussian-feathered noise). This is the exact caveat the doc itself
flagged for verification against the authors implementation. The
ablation only makes sense with lambda as prediction weight.
Fix: lambda weights the REFINER PREDICTION -
slerp(eps_random, eps_refined, lambda) (noise std 0.07, correlation
0.98); additive legacy mode flipped likewise to
eps_refined + (1-lambda)*eps_random. New TestLambdaConvention guards
(injected noise < 0.2 vs signal; end-to-end correlation >= 0.95).
Formula pins flipped (lambda=1 -> pure prediction). HF bounds
recalibrated from measurements (both modes ~0.583 -> >= 0.5).
Full repo suite: 1148 passed.
Document the refiner_model input (paper SDXL + SDXL-Turbo setup),
noise_injection modes, sigma=24 default and the gaussian_kernel_size
removal (migration note for old workflows), plus the 2.9.0 changelog:
corrected SLERP, adapter space fix (SDXL 7.7x under-noising root
cause), generic DDIM, analytic Gaussian mask, empty-negative CFG fix.
Version bumped 2.8.3 -> 2.9.0. Full repo suite green (1146 passed).
Step 11 of plan 2026-09-02: measured the corrected algorithm on the
SDXL-realistic mock (VAE std 7.7 = unit model space, unit-std structured
eps, space-converting adapters). Results: refinement HF ratio 0.748
(slerp) / 0.955 (additive); inversion rel norm 0.87 (= sigma_K, analytically
consistent); dominance 0.85; seam 0.83. The historical "noise dominance"
(6.3x) reproduces only with mismatched magnitudes - confirming the
2026-08-12 bug was the magnitude/space mismatch. slerp HF guard tightened
from provisional >=0.5 to >=0.6 (measured 0.748). Results recorded in
the plan.
Add an execute()-signature-matches-schema test (regex over multi-line
Input() calls) so removed inputs (gaussian_kernel_size) and new ones
(refiner_model, noise_injection) can never drift from the execute
parameters; assert sigma max=128 covers the paper default 24; forbid
gaussian_kernel_size in the node source entirely.
Per pixelrush-correct.txt the refiner should be a distinct model from
the base generator (paper: SDXL base + SDXL-Turbo ADD-distilled
refiner). New optional node input refiner_model: when provided (and
distinct from the base), execute builds refiner_eps from it via its own
model_sampling, while the schedule functions (alpha_bar_at, sigma_at,
forward/reverse steps) stay bound to the base model. When absent, the
base model is reused for both - the documented intentional choice.
Acceptance test: two recording adapters verify base at t=0 and
refiner at t=K per patch.
run_cond('negative') returned zeros for an empty negative list, so CFG
degenerated to eps = cfg_scale * eps_cond (7x amplification at the
default 7.0). Now an empty negative returns the conditional eps
unchanged (CFG undefined without an unconditional branch), and an
empty positive raises ValueError. Tests build a real _make_predict_eps
against new conftest stubs (comfy.model_management/samplers/
sampler_helpers/utils) and pin both paths exactly. Full repo suite
green (1139 passed).
Root fix for the SDXL 'compressed look': in VAE-space mode the
forward/reverse adapters previously applied noise_scaling directly to
VAE-space x with model-space eps, so for SDXL (scale_factor 0.13025)
the model saw noise 7.7x too small for the claimed timestep - the
input SNR never matched sigma and the refiner's prediction washed out.
Now forward_step converts x via process_latent_in, applies noise_scaling
in model space, and converts back via process_latent_out (reverse_step
mirrors it). predict_eps keeps returning model-space eps - principled
under the corrected theory since the slerp mixes it with std-1 random
noise. For pure-scaling formats the composition equals running the whole
algorithm in model space (exactness pinned by
process_latent_in(forward(x,e,s)) == s*x + s*e... i.e. noise at full
model-space scale).
The operate_in_vae_space flag is removed everywhere: the pipeline is
always VAE-space at the interfaces (ComfyUI LATENT convention); VAE
adapters no longer touch process_latent_out/in. Dominance guard gains
a lower bound (no-op detection) plus a refinement-changes-latent test.
Default injection is now the paper's slerp(eps_refined, eps_random,
noise_lambda) per pixelrush-correct.txt. The 2026-08-13 additive mode
(eps_pred + lambda*eps_rand) remains available via
noise_injection='additive' (node combo input) for workflows tuned
against it. Tests pin both formulas exactly via seeded randn streams
(lambda=0 -> pure eps, lambda=1 -> pure random, additive formula),
the slerp default, and invalid-mode rejection. The legacy-calibrated
HF tests are pinned to additive; a new slerp-mode HF guard holds the
provisional >=0.5 bound pending Step 11 calibration.