Document the refiner_model input (paper SDXL + SDXL-Turbo setup),
noise_injection modes, sigma=24 default and the gaussian_kernel_size
removal (migration note for old workflows), plus the 2.9.0 changelog:
corrected SLERP, adapter space fix (SDXL 7.7x under-noising root
cause), generic DDIM, analytic Gaussian mask, empty-negative CFG fix.
Version bumped 2.8.3 -> 2.9.0. Full repo suite green (1146 passed).
Step 11 of plan 2026-09-02: measured the corrected algorithm on the
SDXL-realistic mock (VAE std 7.7 = unit model space, unit-std structured
eps, space-converting adapters). Results: refinement HF ratio 0.748
(slerp) / 0.955 (additive); inversion rel norm 0.87 (= sigma_K, analytically
consistent); dominance 0.85; seam 0.83. The historical "noise dominance"
(6.3x) reproduces only with mismatched magnitudes - confirming the
2026-08-12 bug was the magnitude/space mismatch. slerp HF guard tightened
from provisional >=0.5 to >=0.6 (measured 0.748). Results recorded in
the plan.
Add an execute()-signature-matches-schema test (regex over multi-line
Input() calls) so removed inputs (gaussian_kernel_size) and new ones
(refiner_model, noise_injection) can never drift from the execute
parameters; assert sigma max=128 covers the paper default 24; forbid
gaussian_kernel_size in the node source entirely.
Per pixelrush-correct.txt the refiner should be a distinct model from
the base generator (paper: SDXL base + SDXL-Turbo ADD-distilled
refiner). New optional node input refiner_model: when provided (and
distinct from the base), execute builds refiner_eps from it via its own
model_sampling, while the schedule functions (alpha_bar_at, sigma_at,
forward/reverse steps) stay bound to the base model. When absent, the
base model is reused for both - the documented intentional choice.
Acceptance test: two recording adapters verify base at t=0 and
refiner at t=K per patch.
run_cond('negative') returned zeros for an empty negative list, so CFG
degenerated to eps = cfg_scale * eps_cond (7x amplification at the
default 7.0). Now an empty negative returns the conditional eps
unchanged (CFG undefined without an unconditional branch), and an
empty positive raises ValueError. Tests build a real _make_predict_eps
against new conftest stubs (comfy.model_management/samplers/
sampler_helpers/utils) and pin both paths exactly. Full repo suite
green (1139 passed).
Root fix for the SDXL 'compressed look': in VAE-space mode the
forward/reverse adapters previously applied noise_scaling directly to
VAE-space x with model-space eps, so for SDXL (scale_factor 0.13025)
the model saw noise 7.7x too small for the claimed timestep - the
input SNR never matched sigma and the refiner's prediction washed out.
Now forward_step converts x via process_latent_in, applies noise_scaling
in model space, and converts back via process_latent_out (reverse_step
mirrors it). predict_eps keeps returning model-space eps - principled
under the corrected theory since the slerp mixes it with std-1 random
noise. For pure-scaling formats the composition equals running the whole
algorithm in model space (exactness pinned by
process_latent_in(forward(x,e,s)) == s*x + s*e... i.e. noise at full
model-space scale).
The operate_in_vae_space flag is removed everywhere: the pipeline is
always VAE-space at the interfaces (ComfyUI LATENT convention); VAE
adapters no longer touch process_latent_out/in. Dominance guard gains
a lower bound (no-op detection) plus a refinement-changes-latent test.
Default injection is now the paper's slerp(eps_refined, eps_random,
noise_lambda) per pixelrush-correct.txt. The 2026-08-13 additive mode
(eps_pred + lambda*eps_rand) remains available via
noise_injection='additive' (node combo input) for workflows tuned
against it. Tests pin both formulas exactly via seeded randn streams
(lambda=0 -> pure eps, lambda=1 -> pure random, additive formula),
the slerp default, and invalid-mode rejection. The legacy-calibrated
HF tests are pinned to additive; a new slerp-mode HF guard holds the
provisional >=0.5 bound pending Step 11 calibration.
Per pixelrush-correct.txt the core must accept distinct epsilon
adapters: inversion_eps (base generator, drives the 0->K partial DDIM
inversion) and refiner_eps (distilled one-step refiner, drives the K->0
prediction). The node currently passes the same base-model predict_eps
for both - an intentional choice the corrected theory allows; a
separate refiner_model input is added in a later step. Tests pin the
call pattern (inversion at t=0, refiner at t=K) and that a distinct
refiner changes the output.
Per pixelrush-correct.txt: the mask is exp(-(xx^2+yy^2)/(2 sigma^2))
centered on the patch and peak-normalized to 1 - an explicit encoding of
the paper's Gaussian-filtered patch mask. Replaces the blurred-all-ones
approximation (sigma=8, kernel=41); gaussian_kernel_size is removed from
config, node schema and execute. Node sigma range widened to 1..128 so
the corrected default 24 is reachable from the UI.
alpha_k was only computed when sigma_at was None, so sigma_at +
DDIM fallback (either adapter missing) raised NameError. Compute it
whenever any fallback branch can execute; when both adapters are
provided alpha_bar_at is never called (pinned by test).
Add predict_x0_from_epsilon + ddim_deterministic_step (eta=0, source ->
dest schedule values, same epsilon) per pixelrush-correct.txt. The old
0->K / K->0 functions become thin wrappers over the generic form.
Tests pin the formula, wrapper equivalence, and eta=0 path-independence.
Per pixelrush-correct.txt: the 2026-08-12 unit-vector x linear-magnitude
form carried magnitudes separately; the paper convention is standard
vector slerp sin((1-t)w)/sin(w)*a + sin(tw)/sin(w)*b on raw flattened
vectors, with a lerp fallback when sin_omega < 1e-4. The conflicting
grep-tests (which forbade the corrected form) are rewritten to require it.
G1: wrapper now mirrors the real ComfyUI optimized_attention signature (mask/attn_precision/skip_reshape/skip_output_reshape in slots 5-8) and forwards orig() with the correct positional order - the pre-fix wrapper fed skip_reshape into the mask slot on the unmasked Anima/cosmos path (AttributeError: bool.ndim).
G2: pytest mock _sdpa flipped to the real ComfyUI signature (it matched the wrapper's inverted convention, hiding G1); conformance tripwire added (tests/test_orig_call_convention.py).
G3: HAP decline-guards - non-square attention (cross-attn) and plan/model head-count mismatch decline to plain attention with one-time logs instead of crashing or computing wrong math.
G4: _hrdit_carry_state() re-applies HRDiT patcher attrs after ModelPatcher.clone() in both apply functions (+ _hrdit_state_ref indirection), so SPA->HAP / HAP->SPA chaining is order-independent.
SPA non-square guard: averaged passes apply spatial RoPE to both q and k (only valid for square self-attention); Anima cross-attention (q_len != k_len) crashed einsum - now declined to plain attention.
README: v2.7.1 changelog + HAP model-specificity/v1-limitation notes.
The DyPE extrapolation methods (ntk/yarn/vision_yarn/pi) never
affected SPA output: SPA always applies the model's native
no-extrapolation RoPE (ntk_factor=1.0) on the bundled coords
(HRDiT "nor" RoPE). The knob was inherited UI plumbing from the
DyPE base-class constructor chain and only invited misleading
A/B testing.
- drop the method combo from the SPA node schema + execute()
- keep method/yarn_alt_scaling params in apply_spa_to_model for
the DyPEBasePosEmbed constructor chain (documented as no-ops)
- add a schema guard test asserting the input stays removed
- update README parameter table + changelog
HRDiT-style bundle+slide positional alignment for high-res DiTs
(Krea-2, Z-Image, FLUX, Anima, Qwen, Nunchaku). Fixes the
ripple/mosaic output and ~10x inference time at bundle_size>2.
- bundle_size is the paper's tokens-per-bundle N (0=auto, 1=off,
2..8 explicit); legacy group_num-scale values (>=32) migrate to
auto with a one-time warning
- trained-extent gate: identity at <=1024px (no big-patch mosaic)
- spa_steps (default 3) gates SPA to leading denoise steps;
~1.3-1.8x overhead instead of 10x
- delta-RoPE cache: inv(base)@variant composed once per grid
- averages attention OUTPUTS across 2s-1 slide variants (not
rotation matrices), killing the period-s ripple