G1: wrapper now mirrors the real ComfyUI optimized_attention signature (mask/attn_precision/skip_reshape/skip_output_reshape in slots 5-8) and forwards orig() with the correct positional order - the pre-fix wrapper fed skip_reshape into the mask slot on the unmasked Anima/cosmos path (AttributeError: bool.ndim).
G2: pytest mock _sdpa flipped to the real ComfyUI signature (it matched the wrapper's inverted convention, hiding G1); conformance tripwire added (tests/test_orig_call_convention.py).
G3: HAP decline-guards - non-square attention (cross-attn) and plan/model head-count mismatch decline to plain attention with one-time logs instead of crashing or computing wrong math.
G4: _hrdit_carry_state() re-applies HRDiT patcher attrs after ModelPatcher.clone() in both apply functions (+ _hrdit_state_ref indirection), so SPA->HAP / HAP->SPA chaining is order-independent.
SPA non-square guard: averaged passes apply spatial RoPE to both q and k (only valid for square self-attention); Anima cross-attention (q_len != k_len) crashed einsum - now declined to plain attention.
README: v2.7.1 changelog + HAP model-specificity/v1-limitation notes.
The DyPE extrapolation methods (ntk/yarn/vision_yarn/pi) never
affected SPA output: SPA always applies the model's native
no-extrapolation RoPE (ntk_factor=1.0) on the bundled coords
(HRDiT "nor" RoPE). The knob was inherited UI plumbing from the
DyPE base-class constructor chain and only invited misleading
A/B testing.
- drop the method combo from the SPA node schema + execute()
- keep method/yarn_alt_scaling params in apply_spa_to_model for
the DyPEBasePosEmbed constructor chain (documented as no-ops)
- add a schema guard test asserting the input stays removed
- update README parameter table + changelog
HRDiT-style bundle+slide positional alignment for high-res DiTs
(Krea-2, Z-Image, FLUX, Anima, Qwen, Nunchaku). Fixes the
ripple/mosaic output and ~10x inference time at bundle_size>2.
- bundle_size is the paper's tokens-per-bundle N (0=auto, 1=off,
2..8 explicit); legacy group_num-scale values (>=32) migrate to
auto with a one-time warning
- trained-extent gate: identity at <=1024px (no big-patch mosaic)
- spa_steps (default 3) gates SPA to leading denoise steps;
~1.3-1.8x overhead instead of 10x
- delta-RoPE cache: inv(base)@variant composed once per grid
- averages attention OUTPUTS across 2s-1 slide variants (not
rotation matrices), killing the period-s ripple