Succinct "Which method when?" table in the Nodes section: models, mechanism, takes-your-image, output character for all six methods, plus a two-family framing (model patches for native generation vs cascades for refinement) and a quick picker tip. Also adds ImageCompare to the workflow-test core-node allowlist (a real ComfyUI core node in comfy_extras/nodes_image_compare.py that the new DyPE-SDXL example workflow uses).
23 KiB
ComfyUI-DyPE
ComfyUI custom node pack for ultra-high-resolution generation (4K and beyond) with Diffusion Transformers — FLUX, Qwen Image, Z-Image, Anima/Cosmos, Krea-2.
▷ About
Training-free methods that push pre-trained DiT models far beyond their native resolution — no retraining, no workflow changes. Patch the model once after your loader and generate at 2K, 4K and above.
❖ Highlights
- Multi-Architecture — FLUX, Nunchaku, Qwen Image, Krea-2, Z-Image, Anima/Cosmos
- High-Resolution Generation — 4096×4096 and beyond
- Single-Node Integration — place after your model loader, done
- Full Compatibility — works with existing workflows, samplers and optimization nodes
- Zero Overhead — adjustments happen on-the-fly with negligible performance impact
▓ Nodes
| Node | What it does |
|---|---|
| ❖ DyPE | Dynamic Position Extrapolation — the core high-res method. |
| ❖ SEGA | Content-aware spectral sharpening as an alternative to DyPE. |
| ❖ SPA (HRDiT) | Fixes spatial disorder (repeated/collapsed structures) at high res. |
| ❖ HAP (HRDiT) | Sparse-attention acceleration — the speed half of HRDiT. |
| ❖ PixelRush | Cascade patch refinement of an existing base image. |
| ❖ FreeScale | Tuning-free self-cascade upscaling. |
Which method when?
Two families: model patches alter how your own KSampler run attends (no image input) — best for native high-res generation; cascades consume an existing latent/image and refine it.
| Method | Models | Mechanism | Takes your image | Output character |
|---|---|---|---|---|
| DyPE | FLUX, Nunchaku, Qwen/Krea-2, Z-Image, Anima | Dynamic position-encoding extrapolation | ✗ | Native high-res generation |
| SEGA | FLUX, Nunchaku, Qwen/Krea-2, Z-Image, Anima | Spectral-energy RoPE sharpening | ✗ | Native high-res generation |
| SPA | FLUX, Qwen/Krea-2, Z-Image, Anima | Position-bundle attention alignment | ✗ | Native high-res generation |
| HAP | FLUX, Qwen/Krea-2, Z-Image, Anima | Calibrated sparse attention (speed) | ✗ | Native high-res generation |
| PixelRush | Any (SDXL, SD1.5, FLUX, Qwen, …) | Patch-wise low-denoise img2img cascade | ✓ | Faithful upscale + refinement |
| FreeScale | FLUX-family DiTs | Scale-fused attention + self-cascade | ✓ | Regenerative hi-res, mostly new content |
Tip
Quick picker: starting from noise → DyPE (or SEGA), add SPA if you see repeated/collapsed structures, add HAP for speed. Starting from an existing image → PixelRush to keep it faithful, FreeScale to re-imagine it at high res (lower its
noise_timestepfor more fidelity).
❖ DyPE
Dynamic Position Extrapolation (paper, code). Adjusts positional encodings at each denoising step to match the current stage of generation — low-frequency structure early, fine detail later. Training-free, no additional sampling cost.
Usage: Load model → add DyPE for FLUX (under model_patches/unet) → connect MODEL → set width/height to match your latent → connect to KSampler.
Inputs & Parameters
Model Configuration
model_typeauto— auto-detects the architecture. Recommended.flux— Standard Flux.nunchaku— Quantized Flux.qwen— Qwen Image (also used for Krea-2).zimage— Z-Image (Lumina 2).anima— Anima/Cosmos.
base_resolution— native training resolution of the model.- Flux / Z-Image:
1024 - Qwen / Krea-2:
1328 - Anima/Cosmos:
1920(auto-detected)
- Flux / Z-Image:
Method Selection (method)
vision_yarn— decouples structure from texture; best aspect-ratio robustness. Recommended default.yarn— standard YaRN; good general performance.ntk— very stable, but softer at high resolutions.pi— Position Interpolation; preserves local structure well.base— no interpolation.
Scaling Options
yarn_alt_scaling(only affectsyarn): Anisotropic scales H/W independently (may stretch); Isotropic (default) is stable. Ignored byvision_yarn.
Dynamic Control
enable_dype— full dynamic algorithm (on), or schedule shift only (off).dype_scale— magnitude of the modulation (default2.0).dype_exponent— strength over time:2.0for 4K+,1.0for ~2K–3K,0.5just above native.
Advanced Noise Scheduling
base_shift/max_shift— noise-schedule shift control (max_shiftdefault1.15).
Tip
Z-Image: isotropic scaling is enforced automatically. Prefer
vision_yarnorntk. Anima/Cosmos: prefervision_yarn; other methods may produce speckle noise above 2K.
❖ SEGA
Spectral-Energy Guided Attention (code). Content-aware RoPE sharpening derived from the latent's frequency spectrum. Use as an alternative to DyPE on FLUX/Qwen.
Usage: Add the SEGA node after your model loader → set width/height to match your latent → tune mscale_alpha and spread_min/spread_max.
Inputs & Parameters
| Parameter | Default | Description |
|---|---|---|
method |
sega | sega = NTK + spectral mscale, ntk = NTK only |
mscale_alpha |
0.15 | Spectral redistribution amplitude |
mscale_beta |
1.5 | tanh sharpness |
mscale_min |
1.0 | Floor for per-frequency mscale |
spread_min |
0.0 | Min spectral spread (early steps) |
spread_max |
1.0 | Max spectral spread (late steps) |
spread_alpha |
1.5 | Spread schedule non-linearity |
base_mscale_formula |
power_res | power_res or log_res |
base_mscale_coefficient |
0.08 | κ (paper default) |
Note
SEGA builds on NTK. If NTK doesn't work for your model (e.g. Anima), use DyPE
vision_yarninstead.
❖ SPA (HRDiT)
Spatial Position Alignment, from the HRDiT paper (arXiv 2608.07003). A static, training-free patch that fixes high-resolution spatial disorder — repeated structures and positional collisions when pushing past native resolution. Resolution-aware (automatic no-op ≤ 1024px) with bounded overhead at 2K/4K. Mechanism: bundles token positions into groups of N, slides the bundle boundary per axis (2s − 1 variants), and averages the attention outputs across variants — never the RoPE matrices themselves.
Usage: Add the SPA (HRDiT) node after your model loader → set width/height → leave model_type: auto → connect to KSampler. Recommended bundle_size: 3 at 2K, 5 at 4K (0 = auto).
Inputs & Parameters
| Parameter | Default | Description |
|---|---|---|
model_type |
auto | Same detection as DyPE. Reads theta & axes_dim from the model. |
enable_spa |
True | Disable to pass the model through unchanged. |
bundle_size |
0 (auto) | Tokens per bundle (paper's N). 0 = auto, 1 = off, 2..8 explicit. Auto no-op inside the model's trained extent (≤ 1024px). |
spa_steps |
3 | SPA runs only on the first 3 denoising steps; later steps run at baseline speed. 0 = all steps. |
spa_start_sigma |
1.0 | Optional sigma-threshold gate (combined AND with spa_steps). |
spa_layer_filter |
"" | Restrict SPA to a subset of layers, e.g. "0-18,38-57". Empty = every layer. |
proportional_attention |
False | HRDiT proportional attention scaling for long sequences. No-op at/below 1024px. |
Performance: ~zero overhead at ≤ 1024px; roughly 1.3–1.8× total inference time at 2K/4K with defaults.
Model support: FLUX, Qwen/Krea-2, Z-Image, Anima/Cosmos. Nunchaku not supported (logs a warning, returns the model unchanged).
Warning
SPA and DyPE/SEGA are mutually exclusive — apply only one.
- SPA — fix spatial disorder with small, bounded overhead.
- DyPE/SEGA — full dynamic extrapolation far beyond native resolution.
❖ HAP (HRDiT)
Head-Adaptive attention Pruning, from the same HRDiT paper — the speed half complementing SPA (the quality half). Each attention head only sees the keys it actually needs, via a pre-calibrated scope plan executed through block-sparse attention. Composable with SPA in any order.
A ready-to-use FLUX scope plan ships at configs/scope_plan_flux.json.
Usage: Add the HAP (HRDiT) node after your model loader → point scope_plan_path at a plan JSON → connect to KSampler (optionally through an SPA node first).
Inputs & Parameters
| Parameter | Default | Description |
|---|---|---|
scope_plan_path |
configs/scope_plan_flux.json |
Path to the scope-plan JSON. Relative paths resolve against the repo root. Also accepts a linked scope_plan input. |
model_type |
auto | Architecture detection. Nunchaku unsupported. |
anchor_stride |
0 | Every Nth image key block stays globally visible. 0 = off. |
text_len |
512 | Leading text tokens always kept visible. |
enable_hap |
True | Disable to pass the model through unchanged. |
proportional_attention |
False | See SPA. Either node may enable it. |
Backends: fast path needs CUDA + PyTorch ≥ 2.5; otherwise falls back automatically to a correct dense-mask backend.
Calibration
Scope plans are model-specific. Calibrate a custom plan with the HAP Calibrate (HRDiT) node in-graph, or via the calibration/calibrate_hap.py CLI:
# Self-contained dry run (no GPU needed):
python calibration/calibrate_hap.py --dry_run --out tmp/scope_plan_toy.json
# Real-model calibration:
python calibration/calibrate_hap.py --model_path /path/to/flux.safetensors \
--model_type flux --width 4096 --height 4096 --num_prompts 30 \
--out configs/scope_plan_flux_4k.json
Calibrate once per model, then reuse the plan across resolutions and prompts.
From the paper (FLUX, budget 0.1): ~2.9× faster attention at 2K, ~5.5× at 4K.
❖ PixelRush
Cascade-based refinement node. Generates at native resolution first, then progressively adds detail through coarse-to-fine cascade refinements — producing crisp 4K output without regenerating the whole image from noise. Works with any ComfyUI model (SDXL, SD1.5, FLUX, Qwen, …).
Usage: Generate a base latent at native resolution → connect model, vae, positive, negative and the base latent_image → set num_cascade_stages (1 = 2× upscale, 2 = 4×, 3 = 8×) → decode the output latent.
Inputs & Parameters
| Parameter | Description |
|---|---|
num_cascade_stages |
Number of cascade stages — each doubles the resolution. |
refiner_model |
Optional separate refiner model (paper setup: SDXL base + SDXL-Turbo). When not connected, the base model refines too. |
noise_lambda |
Noise injection coefficient — the weight of the model's prediction (paper default 0.95 = 95% prediction + 5% random noise). |
noise_injection |
slerp (paper default) or additive (legacy pre-2.9 behavior, kept for workflows tuned against it). |
overlap |
Overlap between adjacent patches (blends seams). |
gaussian_sigma |
Analytic Gaussian feather sigma (paper default 24; rule of thumb: σ ≈ patch_size / 5). |
patch_h / patch_w |
Latent patch size (~native spatial size keeps VRAM flat). |
Note
PixelRush calls the diffusion model directly (not through ComfyUI's sampler), performing its own CFG and prediction-type handling for EPS, flow, V-prediction and X0 models.
Important
2.9 migration notes: the noise injection now uses the paper's SLERP with λ weighting the model's prediction (set
noise_injectiontoadditivefor the legacy formula);gaussian_sigmadefault moved 8 → 24 and its range extends to 128; thegaussian_kernel_sizeinput was removed (the mask is now the paper's analytic Gaussian — old workflows simply ignore the stale value).
❖ FreeScale
Tuning-free higher-resolution generation via scale-fused attention and self-cascade upscaling (paper, code). Supports FLUX-family DiTs (auto-detected); base-resolution inputs pass through untouched.
Inputs & Parameters
| Input | Default | Notes |
|---|---|---|
width / height |
2048 | Target resolution (snapped to multiples of 16). |
steps |
20 | Sampler steps per cascade stage. |
cfg |
1.0 | Classifier-free guidance scale. |
cascade_stages |
1 | Number of self-cascade stages (each doubles resolution). |
▓ Node Reference
All nodes registered by this pack (V3 schema ids):
| Node id | Display name | Purpose |
|---|---|---|
DyPE_FLUX |
DyPE | Dynamic Position Extrapolation for ultra-high-res generation. |
SEGA |
SEGA | Spectral-Energy Guided Attention (content-aware sharpening). |
SPA |
SPA (HRDiT) | Spatial Position Alignment — fixes spatial disorder. |
HAP |
HAP (HRDiT) | Head-Adaptive attention Pruning — the speed half. |
HAPCalibrate |
HAP Calibrate (HRDiT) | In-graph scope-plan calibration for HAP. |
PixelRushNode |
PixelRush | Cascade refinement for existing latents. |
FreeScaleNode |
FreeScale | Tuning-free scale-fusion + self-cascade upscaling. |
▓ Getting Started
Via ComfyUI Manager: Search ComfyUI-DyPE → Install.
Manual install:
cd ComfyUI/custom_nodes/
git clone https://github.com/wildminder/ComfyUI-DyPE.git
Restart ComfyUI. No further dependency installation is required.
▓ Tips & Best Practices
Important
Limitations at Extreme Resolutions (4K): you are pushing a model trained on ~1 megapixel toward 16 megapixels — minor artifacts can still appear even with these methods.
Tip
Speckle noise at 4K+: increase
dype_exponent(e.g.3.0–4.0) or apply smoothing / detailer LoRAs.
Tip
Experiment: there is no single magic setting — try different methods and adjust
dype_exponentfor the best sharpness/artifact balance.
▓ Changelog
v2.9.1 — 2026-09-02
- Fixed the PixelRush noise-injection λ convention (user-reported "structure visible but completely noisy, soft blurred patches"). The injection now uses
slerp(eps_random, eps_refined, λ)— λ weights the model's prediction (0.95 = 95% prediction + 5% noise). The previous order (slerp(eps_pred, eps_random, λ)) made λ=0.95 mean 99.6% pure random noise: at real scales per-pixel noise std ≈ 1.17 vs signal ≈ 1.0, which rendered through the Gaussian feather as the reported soft-patch noise. Theadditivelegacy mode uses the same convention (eps_refined + (1−λ)·eps_random). This was exactly the argument-order caveatpixelrush-correct.txtflagged for verification against the authors' implementation.
v2.9.0 — 2026-09-02
- PixelRush realigned with the corrected theory (plan 2026-09-02): standard raw-vector SLERP (with collinear lerp fallback) for the noise injection — the paper's
slerp(eps_pred, eps_random, λ)is now the default, with the 2026-08-13 additive injection kept as an opt-in (noise_injection). - Fixed the VAE/model space mixing in the forward/reverse steps: adapters now convert via
process_latent_in/out, so the model sees noise at the scale its timestep claims. For SDXL the previous code under-noised 7.7× — the root cause behind the "compressed look" that the additive hack had papered over. - Generic DDIM transitions (
ddim_deterministic_stepbetween arbitrary timesteps,predict_x0_from_epsilon); analytic Gaussian feather mask (σ default 24,gaussian_kernel_sizeinput removed). - Optional
refiner_modelinput — use a separate distilled refiner (e.g. SDXL-Turbo) as in the paper; the base model drives the partial inversion. - Bug fixes: empty-negative conditioning no longer amplifies eps by
cfg_scale(CFG is skipped);alpha_kNameError with partially-provided adapters; empty positive now raises a clear error.
v2.8.3 — 2026-08-31
- Qwen2D VAE support disabled by default. User reports showed that with the Qwen2D VAE interception installed, loading certain non-Qwen2D (video-style) VAE checkpoints crashed with a size-mismatch error whose traceback passed through this pack's delegation frame — breaking workflows that never used the Qwen2D VAE. The patch now installs only when the environment variable
DYPE_ENABLE_QWEN2D_VAE=1is set. If you relied on the Qwen2D VAE (Anzhc/Qwen2D-VAE checkpoint with FreeScale/PixelRush on Krea-2/Qwen/Anima), set that variable in your ComfyUI environment to restore the previous behavior.
v2.8.2 — 2026-08-31
- Fixed graph-build and execution crashes when resolution inputs are
None(validate_inputs now passes through uninitialized state; execute falls back to 1024) - Fixed PixelRush crash on float16: antialiased bicubic upsample casts to float32 and restores the original dtype
v2.8.1 — 2026-08-25
- Fixed valid resolutions being rejected at graph build
- Validation errors are now reported once, for the right input
v2.8.0 — 2026-08-16
- New HAP Calibrate node: calibrate HAP directly in-graph
- HAP accepts calibrated plans either by file or by direct connection
- CLI calibration tooling completed
v2.7.1 — 2026-08-16
- Fixed crashes on Anima/Cosmos models
- Safer automatic fallbacks instead of hard errors
- SPA and HAP nodes now work in any order
v2.7.0 — 2026-08-15
- New HAP node: sparse-attention acceleration (up to ~5× faster attention at 4K)
- One-click scope-plan calibration pipeline (in-graph + CLI)
- New optional attention scaling and per-layer filtering controls
- SPA and HAP can be composed together
v2.6.1 — 2026-08-15
- Reworked SPA bundle-size control to match the paper
- Much faster SPA runs (up to ~10× less overhead at strong settings)
- Automatic no-op at/below native resolution
v2.6.0 — 2026-08-15
- New SPA node (HRDiT)
PixelRush update
- Fixed "totally noisy" output on SDXL models
v2.5.0
- New SEGA node
- Video-model latent support
v2.4.0
- Anima/Cosmos support
- Krea-2 support
- Stability fixes and new example workflows
v2.3.0
- Z-Image quality improvements
v2.2.0
- Experimental Z-Image support
v2.1.0
- Qwen Image and Nunchaku support
- Modular codebase refactor for easier future model support
v2.0.0
- New
vision_yarnmethod for better aspect-ratio handling - Sharper results with fewer artifacts
- New start-sigma control
v1.0.0
- Initial release: core DyPE for FLUX with
yarnandntkmethods
▓ Acknowledgments
- Noam Issachar, Guy Yariv and co-authors — DyPE (paper)
- The SEGA authors — SEGA
- The HRDiT team — HRDiT (code) — basis for SPA & HAP
- The PixelRush authors — PixelRush
- Yanhong Zeng et al. — FreeScale (paper)
- The ComfyUI team — for the platform
══════════════════════════════════