ComfyUI-TIDE
ComfyUI-TIDE is a ComfyUI custom node implementation of inference-time mechanisms from TIDE: Text-Informed Dynamic Extrapolation with Step-Aware Temperature Control for Diffusion Transformers.
The primary implementation targets Flux-style DiT attention in ComfyUI. The repository also includes a native MiniMax H3 packed audio-video path, a WAN 2.1/2.2 path for ComfyUI WAN video DiTs, and an experimental SDXL/UNet adaptation that applies the usable attention-temperature part of the method to SDXL-style SpatialTransformer attention.
The nodes patch a cloned ComfyUI MODEL object. They do not add extra sampling steps, replace the sampler, replace the scheduler, or fork ComfyUI core.
Credits and attribution
The algorithmic method implemented here is based on the TIDE paper:
TIDE: Text-Informed Dynamic Extrapolation with Step-Aware Temperature Control for Diffusion Transformers
Yihua Liu, Fanjiang Ye, Bowen Lin, Rongyu Fang, Chengming Zhang
arXiv:2603.08928, 2026
https://arxiv.org/abs/2603.08928
Credit for the TIDE method, including Text Anchoring, Dynamic Temperature Control, and the paper's analysis of attention dilution in high-resolution Diffusion Transformer generation, belongs to the paper authors.
This repository is an independent ComfyUI custom-node implementation. It is not the official TIDE repository, not affiliated with the paper authors, and should not be cited as the original method. If this node is useful in your work, cite the TIDE paper.
@misc{liu2026tide,
title = {TIDE: Text-Informed Dynamic Extrapolation with Step-Aware Temperature Control for Diffusion Transformers},
author = {Yihua Liu and Fanjiang Ye and Bowen Lin and Rongyu Fang and Chengming Zhang},
year = {2026},
eprint = {2603.08928},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2603.08928}
}
Scope
This repository implements practical ComfyUI attention-side mechanisms inspired by TIDE.
For Flux-style joint text/image attention, it implements:
- text-token additive bias for Text Anchoring;
- step-aware RoPE temperature scaling for Dynamic Temperature Control;
- a lightweight ComfyUI model wrapper to pass the current denoising timestep into the attention patch;
- a small PyTorch SDPA fallback used only when an additive TIDE attention mask is active.
For WAN 2.1/2.2-style video DiT attention, it implements:
- step-aware RoPE temperature scaling for WAN self-attention;
- lazy wrapping of ComfyUI WAN
rope_encode/forward_origpaths through a cloned-model diffusion wrapper; - chaining through ComfyUI
WrappersMP.DIFFUSION_MODELwithout modifying ComfyUI core.
For native MiniMax H3 packed audio-video attention, it implements:
- Text Anchoring for genuine text and textual-label keys in H3's joint packed self-attention;
- live-shape derivation for T2VA, first/last-frame conditioning, and Ref2VA;
- opt-in experimental Dynamic Temperature Control for H3's temporal, height, and width RoPE axes;
- model-local diffusion wrappers, attention overrides, and block replacements that compose with Spectrum MiniMax H3.
For SDXL-style UNet attention, it implements:
- step-aware attention-temperature scaling through ComfyUI's
optimized_attention_override; - optional application to SDXL cross-attention, self-attention, or both;
- model-local patching that chains with a pre-existing attention override when present.
It does not implement the full official Diffusers pipeline, benchmark harness, datasets, metric evaluation scripts, Qwen pipeline, or complete YaRN/DyPE/NTK positional interpolation stack.
Supported model paths
Flux / Flux.2-style DiT path
The main target path is FLUX-family DiT models in ComfyUI, including FLUX.2-style paths if they use the same practical structure as ComfyUI's Flux implementation:
- joint text/image attention;
- text tokens before image tokens;
attn1_patchsupport;extra_options["img_slice"]available at the attention patch site;- RoPE matrix passed as
pe.
Other DiT models may require model-specific patch paths. They should not be assumed to work unless their ComfyUI implementation exposes the same attention-patch contract.
WAN 2.1 / 2.2 video DiT path
The WAN node targets ComfyUI WAN-family implementations based on comfy.ldm.wan.model.WanModel and close subclasses, including WAN 2.1 and WAN 2.2 paths that expose:
rope_encode;forward_orig;rope_embedder.axes_dim;- ComfyUI diffusion-model wrapper execution.
The WAN path applies Dynamic Temperature Control to the WAN self-attention RoPE matrices. It does not apply TIDE Text Anchoring because ComfyUI WAN uses separate self-attention over video tokens and cross-attention over text/context tokens. In that architecture, text keys and image/video keys are not competing inside one joint softmax, and adding the same positive bias to every text key in pure cross-attention would be cancelled by softmax shift invariance.
Native MiniMax H3 path
The MiniMax H3 node targets the current native comfy.ldm.minimax.model.MiniMaxH3Model contract:
- a packed joint sequence ordered as
[presentation | conditions/references | target audio | target video]; - exactly one contiguous target-audio span followed by one target-video span at the packed tail;
minimax_payload["text_token_tags"], where tag1is genuine text or a textual label and tag0is a Qwen vision-block position;- three-axis RoPE backed by the live model's
rope.inv_freq; - transformer blocks exposed through
patches_replace["dit"][("double_block", index)].
Supported conditioning paths are plain T2VA, first-frame I2V, first-and-last-frame video conditioning, and Ref2VA image/video/audio references. Target dimensions and target row counts are derived from the live native video/audio latent pair. Users do not enter a duplicate target width, height, or duration.
Only tag-1 keys in the presentation prefix receive Text Anchoring. Qwen vision positions, keyframe rows, reference rows, target audio rows, and target video rows are excluded. If tags are absent in plain T2VA, the native presentation prefix can safely fall back to all text. If tags are absent or invalid with keyframes or references, anchoring is skipped for that call.
The TIDE paper validates its mechanisms on high-resolution text-to-image DiTs. It does not validate this MiniMax H3 adaptation. This implementation independently applies the paper's attention-side ideas to H3's native packed audio-video architecture.
SDXL / UNet path
The SDXL node targets ComfyUI's UNet SpatialTransformer attention path:
- SDXL-style cross-attention and self-attention;
optimized_attention_overridesupport;transformer_options["activations_shape"]present at the attention site.
The SDXL path is intentionally separate from the Flux path.
Important: SDXL support is not a full paper-faithful TIDE implementation. TIDE Text Anchoring is defined for joint text/image attention, where text keys and image keys compete inside one softmax. In SDXL UNet cross-attention, the keys/values are text-only. Adding the same positive bias to every text key would be cancelled by softmax shift invariance and would not change the output. SDXL self-attention has image tokens but no text keys. Therefore the SDXL node implements the usable part: step-aware attention-temperature control.
Nodes
TIDE High-Resolution Extrapolation
Use this node for FLUX-family DiT models.
Implemented mechanisms:
- Text Anchoring;
- Dynamic Temperature Control;
- optional PyTorch SDPA fallback when the additive text-anchor mask is active.
TIDE WAN High-Resolution Extrapolation
Use this node for WAN 2.1 / WAN 2.2-style video DiT models in ComfyUI.
Implemented mechanism:
- Dynamic Temperature Control on WAN self-attention RoPE.
Not implemented for WAN:
- Text Anchoring, because WAN text conditioning is separate cross-attention rather than Flux-style joint text/image attention.
TIDE MiniMax H3 Extrapolation
Use this node for ComfyUI's native MiniMax H3 model.
Implemented mechanisms:
- packed-sequence Text Anchoring on tag-
1genuine-text keys; - optional temporal, spatial, or all-axis H3 RoPE Dynamic Temperature Control;
- mask-free SageAttention LSE Text Anchoring with automatic PyTorch SDPA fallback;
- existing attention-override and H3 block-replacement chaining;
- Spectrum MiniMax H3 actual-step composition.
Dynamic Temperature Control is disabled by default for H3. Spatial and all-axis modes are experimental because native H3 uses area-normalized spatial position coordinates.
TIDE SDXL High-Resolution Extrapolation
Use this node for SDXL-style UNet models.
Implemented mechanism:
- step-aware attention-temperature scaling.
Not implemented for SDXL:
- Text Anchoring;
- Flux-style RoPE temperature scaling;
- MM-DiT text/image token balancing.
What is implemented
1. Text Anchoring for Flux-style joint attention
TIDE identifies text-token influence decay as a core failure mode at high resolution: image token count grows with resolution while text token count stays fixed. Text Anchoring counteracts this by adding a positive bias to attention logits whose keys are text tokens.
This node computes the default adaptive bias as:
beta = log((target_width * target_height) / (base_width * base_height))
With the default base_width=1024 and base_height=1024, this is equivalent to:
log(width / 1024) + log(height / 1024)
The final applied value is:
applied_beta = text_anchor_strength * beta
By default, the node applies no Text Anchoring at native-or-smaller token counts unless apply_to_native_or_smaller=True.
2. Dynamic Temperature Control for Flux-style RoPE attention
TIDE uses a step-aware temperature curve so attention sharpening is stronger in the early/global part of denoising and relaxes toward the late/detail part of denoising. For Flux-style models, this repository applies that idea as a RoPE temperature multiplier, matching the reference implementation strategy rather than inserting a new attention kernel for every backend.
Default curve:
tau(t, f) = tau_max - (tau_max - tau_min) * t ** alpha(f)
alpha(f) = alpha_low + (alpha_high - alpha_low) * f
Defaults:
| Parameter | Default |
|---|---|
tau_max |
1.0 |
alpha_low |
0.6 |
alpha_high |
0.2 |
temperature_strength |
1.0 |
frequency_mode |
official_raw |
frequency_mode=official_raw uses raw RoPE frequencies, matching the released implementation behavior this port was written against. paper_normalized is exposed for comparison because the paper notation describes a normalized frequency variable.
3. Text Anchoring and experimental Dynamic Temperature Control for MiniMax H3
H3 packs text presentation, visual/audio conditions, references, target audio, and target video into one joint self-attention sequence. Its Text Anchoring bias uses actual competing-key budgets instead of the Flux pixel-area shortcut.
Definitions:
fixed_competing_rows = packed_rows - true_text_rows - target_video_rows - target_audio_rows
video_scale = (current_width * current_height * current_duration)
/ (base_width * base_height * base_duration)
audio_scale = current_duration / base_duration
base_video_rows = target_video_rows / video_scale
base_audio_rows = target_audio_rows / audio_scale
current_competing_rows = fixed_competing_rows + target_video_rows + target_audio_rows
base_competing_rows = fixed_competing_rows + base_video_rows + base_audio_rows
raw_beta = log(current_competing_rows / base_competing_rows)
applied_beta = text_anchor_strength * raw_beta
References and visual conditions remain in fixed_competing_rows on both sides of the comparison. They are competing packed keys, and they are never reclassified as target video or text.
When the optional sageattention package is available on CUDA, H3 Text Anchoring uses two unmasked SageAttention calls. The first is the normal full packed attention. The second reuses the same queries with only the small selected-text key/value set. Their returned log-sum-exp values recover the text probability mass, allowing the implementation to apply the algebraically exact exp(beta) reweighting without an additive mask. Numerical results remain subject to SageAttention's normal quantization. The second call is proportional to the number of genuine text keys rather than the full packed key count.
If that exact SageAttention LSE path is unavailable or an existing attention mask is already present, TIDE uses the compact additive bias [1, 1, 1, packed_sequence_length] with PyTorch SDPA. No dense query-by-key bias is allocated. Set force_pytorch_attention_with_mask=True only to force this fallback for diagnosis or compatibility testing.
H3 Dynamic Temperature Control uses the checkpoint's live rope.inv_freq and builds one scale vector in temporal, height, width order. The denoising phase is the normalized video sigma timestep / 1000, clamped to [0, 1]. The scaled RoPE tensor is created without mutating the native source and is reused by every H3 block in that model call.
4. Dynamic Temperature Control for WAN 2.1/2.2 RoPE attention
The WAN node applies the same step-aware RoPE temperature multiplier used by the Flux path, but at ComfyUI WAN's rope_encode / forward_orig boundary. WAN uses a three-axis RoPE layout (time, height, width), discovered from rope_embedder.axes_dim at runtime. Spatial scaling uses the node's width, height, base_width, and base_height; the temporal axis is left unscaled.
Because WAN does not expose a Flux-style joint text/image attention softmax, the WAN node sets text_anchor_strength=0.0 internally and only applies Dynamic Temperature Control.
5. Dynamic attention temperature for SDXL/UNet attention
The SDXL node applies the attention-temperature part of the method by scaling the attention query tensor before ComfyUI's optimized attention function:
attention_logits = (Q * inv_tau) K^T / sqrt(d)
This is equivalent to applying the temperature factor to the attention logits:
attention_logits = Q K^T / (tau * sqrt(d))
The SDXL curve uses a single exponent:
tau(t) = tau_max - (tau_max - tau_min) * t ** alpha
The minimum temperature is derived from the YaRN-style extrapolation scale:
scale = sqrt((width * height) / (base_width * base_height))
sqrt(1 / tau_min) = 0.1 * log(scale) + 1
The node blends from no-op to full temperature control with:
applied_inv_tau = 1 + (inv_tau - 1) * temperature_strength
SDXL defaults:
| Parameter | Default |
|---|---|
base_width / base_height |
1024 / 1024 |
temperature_strength |
1.0 |
alpha |
0.6 |
tau_max |
1.0 |
apply_to |
both |
Implementation notes
Flux path
- The node clones and patches the incoming ComfyUI
MODEL. - TIDE state is stored on the cloned model through ComfyUI model options.
TIDEModelWrapperinjects current timestep metadata intotransformer_options["tide"].TIDEAttentionPatchapplies Text Anchoring and Dynamic Temperature Control through ComfyUI'sattn1_patchhook.TIDEAttentionOverrideforces a compact PyTorch SDPA path only when an additive TIDE mask is active. This avoids attention backends that reject additive masks or try to materialize a dense high-resolution mask.- No global monkey-patching is used.
- No sampler or scheduler rewrite is performed.
WAN path
- The WAN node clones and patches the incoming ComfyUI
MODEL. - It installs a ComfyUI
WrappersMP.DIFFUSION_MODELwrapper on the cloned model. - The wrapper injects current timestep metadata into
transformer_options["tide"]. - The wrapper lazily wraps the live WAN model's
rope_encodeandforward_origmethods. rope_encodeoutput or externally suppliedfreqsare multiplied by the TIDE per-frequency temperature scale once per forward path.- The implementation preserves ComfyUI's native WAN block loop, sampler, scheduler, attention backend, and dynamic-VRAM lifetime.
MiniMax H3 path
- The node validates and clones the incoming native H3
MODEL. - A keyed
WrappersMP.DIFFUSION_MODELwrapper derives the live packed layout, target geometry, duration, tag-1text selection, Text Anchoring beta, and current video sigma. - Per-call tags, topology, backend state, fallback masks, and scaled RoPE tensors live in that call's
transformer_optionsruntime contract. - The H3-specific attention override activates only for attention calls whose query and key lengths match the packed H3 sequence. Active anchoring prefers the exact unmasked SageAttention LSE decomposition; token-refiner and unrelated attention calls pass through unchanged.
- Optional H3 block replacements change only
rope_freqs, call any prior replacement, and preserve the native{"img": tensor}contract. - Zero Text Anchoring and zero Dynamic Temperature Control return a cloned model without installing output-changing behavior or changing the attention backend.
SDXL path
-
The SDXL node clones and patches the incoming ComfyUI
MODEL. -
It installs an
optimized_attention_overrideon the cloned model only. -
The override detects SDXL/UNet
SpatialTransformerattention by checking fortransformer_options["activations_shape"]. -
It avoids Flux-style paths by skipping attention calls that expose Flux-specific
block_typeorimg_slicemetadata. -
It classifies attention as:
selfwhen query-token count equals key-token count;crosswhen query-token count differs from key-token count.
-
It applies query scaling only to the selected attention kind:
cross,self, orboth. -
If another attention override already exists, the SDXL node delegates to it after applying its query scaling.
Paper / reference-code / local-module mapping
| Paper component | Reference implementation behavior | This repository |
|---|---|---|
| Text-token influence decay | Additive text-token attention mask | tide_core.math.adaptive_text_bias, tide_core.patches.TIDEAttentionPatch |
Text Anchoring, beta = log(lambda) |
log(width / 1024) + log(height / 1024) for FLUX-sized text prefix |
Adaptive beta from node width, height, base_width, base_height; text-token count inferred from ComfyUI img_slice |
| YaRN temperature baseline | get_mscale, default temperature from extrapolation scale |
tide_core.math.get_mscale, get_default_temperature |
| Dynamic Temperature Control | dyheating() / temperature-aware RoPE scaling |
tide_core.math.rope_temperature_scale, applied to ComfyUI pe |
| Denoising-step-aware behavior | Update position embedding state from current timestep | TIDEModelWrapper injects normalized timestep into transformer_options |
| FLUX attention integration | Modified Diffusers FLUX transformer/processor | ComfyUI attn1_patch plus optional optimized_attention_override |
| WAN attention integration | Not part of the paper's main Flux/MM-DiT implementation | tide_core.wan, ComfyUI diffusion-model wrapper, RoPE scaling for WAN self-attention |
| MiniMax H3 packed attention integration | Not evaluated by the TIDE paper | tide_core.minimax_h3, tag-aware packed Text Anchoring and opt-in live-frequency H3 DTC |
| SDXL attention integration | Not part of the paper's main Flux/MM-DiT implementation | nodes_sdxl.py, ComfyUI optimized_attention_override, query scaling for UNet attention |
| Logarithmic FLUX scheduler shift | Scheduler/pipeline-level change | Not implemented by this node |
| DyPE / NTK-by-parts / YaRN positional interpolation | Custom positional interpolation stack | Not fully implemented; this node implements the TIDE attention-side mechanisms only |
Installation
Clone this repository into ComfyUI's custom node directory:
cd ComfyUI/custom_nodes
git clone <this-repo-url> ComfyUI-TIDE
Restart ComfyUI.
No extra runtime dependency is required beyond the PyTorch/ComfyUI environment. requirements.txt lists torch for standalone tests.
Usage
Flux / Flux.2 usage
- Load a FLUX-family model as usual.
- Add TIDE High-Resolution Extrapolation after the model loader.
- Connect the patched
modeloutput to your sampler. - Set
widthandheightto the final generated image dimensions used by your latent node. - Keep
base_width=1024andbase_height=1024for FLUX-family models unless you know the model's native training target differs.
Recommended starting values:
| Setting | Value |
|---|---|
width / height |
final image dimensions |
base_width / base_height |
1024 / 1024 |
text_anchor_strength |
1.0 |
temperature_strength |
1.0 |
alpha_low |
0.6 |
alpha_high |
0.2 |
tau_max |
1.0 |
frequency_mode |
official_raw |
force_pytorch_attention_with_mask |
True |
Ablation settings:
| Test | Settings |
|---|---|
| Text Anchoring only | text_anchor_strength=1.0, temperature_strength=0.0 |
| Dynamic Temperature only | text_anchor_strength=0.0, temperature_strength=1.0 |
| Disabled | text_anchor_strength=0.0, temperature_strength=0.0 |
WAN 2.1 / 2.2 usage
- Load a WAN 2.1 or WAN 2.2 model as usual.
- Add TIDE WAN High-Resolution Extrapolation after the model loader.
- Connect the patched
modeloutput to your sampler. - Set
widthandheightto the final video frame dimensions in pixels. - Set
base_widthandbase_heightto the resolution you want to treat as the model's native/reference resolution for this workflow. The node defaults to640x640because ComfyUI's WAN 2.2 text-to-video blueprint currently uses that size, but WAN checkpoints and workflows vary.
Recommended starting values:
| Setting | Value |
|---|---|
width / height |
final video frame dimensions |
base_width / base_height |
workflow/model reference |
temperature_strength |
1.0 |
alpha_low |
0.6 |
alpha_high |
0.2 |
tau_max |
1.0 |
frequency_mode |
official_raw |
WAN ablation settings:
| Test | Settings |
|---|---|
| Dynamic Temperature | temperature_strength=1.0 |
| Disabled | temperature_strength=0.0 |
| Force native-size patch | apply_to_native_or_smaller=True |
MiniMax H3 usage
Recommended node order with Spectrum:
Load Diffusion Model
-> TIDE MiniMax H3 Extrapolation
-> Spectrum Apply MiniMax H3
-> guider / scheduler
The reverse patch order remains supported because both integrations use ComfyUI's wrapper and block-replacement chains. On Spectrum actual steps, H3 executes with TIDE and Spectrum captures the resulting final hidden state. Forecast steps execute no H3 blocks; they inherit TIDE's effect indirectly through the fitted actual-step feature trajectory.
Recommended starting values:
| Setting | Value |
|---|---|
base_width / base_height |
1344 / 768 |
base_length |
124 frames |
text_anchor_strength |
1.0 |
temperature_strength |
0.0 |
temperature_axes |
temporal_only |
alpha_low / alpha_high |
0.6 / 0.2 |
tau_max |
1.0 |
frequency_mode |
official_raw |
force_pytorch_attention_with_mask |
False |
base_length is a 24 fps frame-count reference. Set it to 362 if you want duration extrapolation to begin only beyond the currently described trained range. Spatial and all-axis DTC modes require visual testing and remain experimental.
H3 ablation settings:
| Test | Settings |
|---|---|
| Text Anchoring only | text_anchor_strength=1.0, temperature_strength=0.0 |
| Temporal DTC only | text_anchor_strength=0.0, temperature_strength=1.0, temperature_axes=temporal_only |
| Experimental spatial DTC | text_anchor_strength=0.0, temperature_strength=1.0, temperature_axes=spatial_only |
| Numerical no-op | text_anchor_strength=0.0, temperature_strength=0.0 |
SDXL usage
- Load an SDXL checkpoint as usual.
- Add TIDE SDXL High-Resolution Extrapolation after the model loader.
- Connect the patched
modeloutput to your sampler. - Set
widthandheightto the final generated image dimensions used by your latent node. - Keep
base_width=1024andbase_height=1024for SDXL unless you intentionally want a different native-resolution reference. - Start with
apply_to=both. If results are unstable, testcrossandselfseparately.
Recommended starting values:
| Setting | Value |
|---|---|
width / height |
final image dimensions |
base_width / base_height |
1024 / 1024 |
temperature_strength |
1.0 |
alpha |
0.6 |
tau_max |
1.0 |
apply_to |
both |
SDXL ablation settings:
| Test | Settings |
|---|---|
| Cross-attention only | apply_to=cross, temperature_strength=1.0 |
| Self-attention only | apply_to=self, temperature_strength=1.0 |
| Both | apply_to=both, temperature_strength=1.0 |
| Disabled | temperature_strength=0.0 |
Node inputs
TIDE High-Resolution Extrapolation
Required
| Input | Description |
|---|---|
model |
ComfyUI MODEL object to patch. |
width, height |
Final target generation dimensions in pixels. Must match the latent/image size used by the workflow. |
text_anchor_strength |
Multiplier on adaptive beta. 1.0 follows the paper/reference behavior. 0.0 disables Text Anchoring. |
temperature_strength |
Multiplier on Dynamic Temperature Control. 1.0 follows the reference curve. 0.0 disables DTC. |
Optional
| Input | Default | Description |
|---|---|---|
base_width, base_height |
1024, 1024 |
Native/training resolution used for adaptive scaling. |
alpha_low, alpha_high |
0.6, 0.2 |
DTC exponents for low/high RoPE frequency behavior. |
tau_max |
1.0 |
Maximum temperature reached near the end of denoising. |
frequency_mode |
official_raw |
official_raw or paper_normalized. |
apply_to_double_blocks |
True |
Apply patch to FLUX double-stream blocks. |
apply_to_single_blocks |
True |
Apply patch to FLUX single-stream blocks. |
apply_to_native_or_smaller |
False |
Allow patching even when target token count is not above base token count. |
force_pytorch_attention_with_mask |
True |
Use internal PyTorch SDPA only when TIDE's additive mask is active. |
preserve_existing_wrapper |
True |
Delegate to an existing ComfyUI model wrapper after injecting TIDE metadata. |
debug |
False |
Log skipped DTC shape mismatches and exceptions. |
TIDE WAN High-Resolution Extrapolation
| Input | Default | Description |
|---|---|---|
model |
required | ComfyUI MODEL object to patch. |
width, height |
1280, 720 |
Final target video frame dimensions in pixels. Must match the latent/video size used by the workflow. |
temperature_strength |
1.0 |
Strength of WAN RoPE Dynamic Temperature Control. 0.0 disables the WAN patch. |
base_width, base_height |
640, 640 |
Reference/native resolution used for adaptive scaling. Adjust for the checkpoint/workflow. |
alpha_low, alpha_high |
0.6, 0.2 |
DTC exponents for low/high RoPE frequency behavior. |
tau_max |
1.0 |
Maximum temperature reached near the end of denoising. |
frequency_mode |
official_raw |
official_raw or paper_normalized. |
apply_to_native_or_smaller |
False |
Allow patching even when target token count is not above base token count. |
preserve_existing_wrapper |
True |
Delegate to an existing ComfyUI model wrapper after injecting TIDE metadata. |
debug |
False |
Log skipped WAN wrapping/shape cases. |
TIDE MiniMax H3 Extrapolation
| Input | Default | Description |
|---|---|---|
model |
required | ComfyUI MODEL containing the exact native MiniMaxH3Model. |
text_anchor_strength |
1.0 |
Multiplier on packed competing-key beta. 0.0 disables H3 Text Anchoring. |
temperature_strength |
0.0 |
Strength of experimental H3 RoPE DTC. Disabled by default. |
base_width, base_height |
1344, 768 |
Native/reference target canvas used for packed-budget and axis scaling. |
base_length |
124 |
Reference frame count at 24 fps. Use 362 to gate duration extrapolation beyond the described trained range. |
temperature_axes |
temporal_only |
temporal_only, spatial_only, or all. Spatial modes are experimental. |
alpha_low, alpha_high |
0.6, 0.2 |
DTC exponents for low/high live H3 RoPE frequencies. |
tau_max |
1.0 |
Maximum temperature near the end of denoising. |
frequency_mode |
official_raw |
Raw live inverse frequencies or the paper-normalized comparison mode. |
apply_to_native_or_smaller |
False |
Preserve signed Text Anchoring beta and permit DTC evaluation at native/smaller scales. |
force_pytorch_attention_with_mask |
False |
Force compact masked PyTorch SDPA instead of the exact unmasked SageAttention LSE path. |
debug |
False |
Emit a bounded topology summary; no per-block logging. |
TIDE SDXL High-Resolution Extrapolation
| Input | Default | Description |
|---|---|---|
model |
required | ComfyUI MODEL object to patch. |
width, height |
1536, 1536 |
Final target generation dimensions in pixels. Must match the latent/image size used by the workflow. |
temperature_strength |
1.0 |
Strength of SDXL attention-temperature scaling. 0.0 disables the patch. |
base_width, base_height |
1024, 1024 |
Native/training resolution used for adaptive scaling. |
alpha |
0.6 |
Single exponent for the SDXL step-aware temperature curve. |
tau_max |
1.0 |
Maximum temperature reached near the end of denoising. |
apply_to |
both |
cross, self, or both. Controls which SDXL attention calls are patched. |
Repository structure
ComfyUI-TIDE/
├── __init__.py
├── nodes.py
├── nodes_sdxl.py
├── requirements.txt
├── README.md
├── examples/
│ └── README.md
├── tide_core/
│ ├── __init__.py
│ ├── config.py
│ ├── math.py
│ ├── minimax_h3.py
│ ├── patches.py
│ └── wan.py
└── tests/
├── test_attention_patch.py
├── test_math.py
├── test_minimax_h3.py
└── test_wan.py
Tests
Run the standalone tests from the repository root:
python -m pip install pytest torch
python -m pytest -q
python -m compileall .
git diff --check
The tests cover:
- adaptive Text Anchoring beta computation;
- YaRN/default temperature formula;
- RoPE temperature scale shape and timestep progression;
- additive attention-mask creation;
- masked SDPA override behavior;
- native H3 model-contract validation and zero-strength identity;
- H3 tag-
1selection, safe fallback behavior, and packed-tail validation; - H3 competing-key Text Anchoring math, exact Sage LSE recombination, and compact fallback-mask broadcasting;
- live-frequency H3 DTC axis selection, block chaining, and one-scale-per-call caching;
- clone/call state isolation and Spectrum-style final-block composition;
- WAN RoPE temperature scaling helper behavior.
The tests do not validate visual quality or live ComfyUI execution.
For the SDXL module, a basic syntax check can be run with:
python -m py_compile nodes_sdxl.py
MiniMax H3 runtime validation
Automated tests do not establish visual or audio quality. For a live H3 checkpoint, compare matching seeds, samplers, schedules, dimensions, lengths, prompts, and conditioning across:
- native baseline;
- H3 TIDE only;
- Spectrum only;
- H3 TIDE followed by Spectrum.
Run these workflows:
- T2VA without references.
- First-frame I2V.
- First-and-last-frame video conditioning.
- Ref2VA with an image reference.
- Ref2VA with a video reference.
- Ref2VA with audio or reference-audio content.
- TIDE H3 followed by Spectrum H3.
- Zero-strength identity.
- Text Anchoring only.
- Temporal DTC only.
- Spatial DTC only as an experimental check.
- Text Anchoring plus Spectrum.
Check prompt adherence, identity preservation, motion stability, fast-action deviations, audio quality, dialogue/sound synchronization, AV synchronization, VRAM, per-step latency, total generation time, and attention-backend fallback behavior. Do not infer a quality improvement from successful execution alone.
Paper vs implementation differences
MiniMax H3 support is a packed-architecture adaptation
The TIDE paper does not evaluate MiniMax H3, audio-video joint denoising, reference conditioning, or H3's area-normalized spatial coordinates. The H3 path retains the paper's additive key-logit anchoring and step-aware frequency curve while deriving dilution from H3's full packed competing-key budget. H3 Dynamic Temperature Control remains opt-in, and spatial modes remain experimental.
WAN support is an adaptation, not full Flux/MM-DiT TIDE
The full Text Anchoring mechanism is defined for MM-DiT joint attention where text keys and image keys compete inside one softmax. ComfyUI WAN uses self-attention for video/image tokens and separate cross-attention for text/context tokens. Therefore this repository applies the TIDE Dynamic Temperature Control mechanism to WAN self-attention RoPE, but does not apply Text Anchoring to WAN cross-attention.
SDXL support is an adaptation, not full TIDE
The full TIDE method is designed around DiT/MM-DiT attention where text tokens and image tokens are present in the same attention sequence. That makes Text Anchoring meaningful because a positive bias on text-key logits changes the balance between text keys and image keys.
SDXL uses a UNet architecture with separate attention patterns. In cross-attention, the keys and values are text-only, so adding the same bias to all text logits would not change the softmax result. In self-attention, the keys are image tokens and there are no text keys to anchor. For that reason, the SDXL node only applies Dynamic Temperature Control as attention-logit sharpening.
Scheduler time shifting
The paper appendix describes a logarithmic FLUX time-shift schedule for high resolutions. This node receives an already-built ComfyUI sampler schedule and does not silently rewrite it. Use a scheduler setup that does not over-shift high-resolution FLUX timesteps.
Positional interpolation
The paper evaluates TIDE in combination with positional extrapolation/interpolation methods such as YaRN/DyPE-style handling. This repository does not port the full positional interpolation stack. It applies the TIDE attention-side mechanisms through ComfyUI's patch system.
Frequency variable
The paper notation describes a normalized frequency variable for alpha(f). The reference implementation behavior used by this port applies the curve using raw RoPE frequencies. This mismatch is exposed as frequency_mode so the default can follow the reference behavior while still allowing controlled comparison.
Text-token count
The paper writes the text-token length abstractly as L_T. Some FLUX implementations use a fixed text-token prefix length. This repository infers the text prefix from ComfyUI's img_slice rather than hard-coding a token count.
General DiT support
The method is architecture-relevant to DiTs, but this implementation is tied to ComfyUI's Flux-style patch interface. Non-Flux DiTs need compatible attention hooks and token-layout metadata.
Assumptions
Flux path
- The model uses Flux-style joint attention with text tokens before image tokens.
extra_options["img_slice"]identifies the split between text and image tokens.- The ComfyUI attention patch receives a RoPE matrix through
pe. - The node
widthandheightmatch the actual generated dimensions. - The timestep passed through the model wrapper is normalized or sigma-like in
[0, 1]; values outside the interval are clamped. - FLUX-family image token granularity is 16 pixels per transformer token.
WAN path
- The model uses a ComfyUI WAN implementation with
rope_encodeandforward_orig. - The node
widthandheightmatch the actual generated video frame dimensions. base_widthandbase_heightare chosen as the intended native/reference dimensions for the specific WAN checkpoint/workflow.- WAN RoPE axes are ordered as
(time, height, width), matching ComfyUI'sWanModel.rope_encode. - The temporal RoPE axis is not scaled by this node.
- The timestep passed through the model wrapper is normalized or sigma-like in
[0, 1]; values outside the interval are clamped.
MiniMax H3 path
- The model is ComfyUI's native
comfy.ldm.minimax.model.MiniMaxH3Modelwith patch size(1, 2, 2)and the current publicPackedLayouthelper. - Native video latents use a 16-pixel spatial VAE factor and native audio latents use a 40 Hz time axis.
- The H3 presentation tag contract remains
1for genuine text/labels and0for Qwen vision positions. - Target audio and video remain the final two contiguous packed segments.
- The three H3 RoPE axes remain ordered temporal, height, width.
- The model receives video sigma as
timestep / 1000; the DTC curve uses that normalized sigma rather than H3's internal1 - sigmavalue.
SDXL path
- The model uses ComfyUI's UNet
SpatialTransformerattention path. optimized_attention_overrideis honored by the active ComfyUI attention backend.transformer_options["activations_shape"]is present for the attention calls that should be patched.- The node
widthandheightmatch the actual generated dimensions. - SDXL's native-resolution reference is treated as
1024x1024unless overridden. - Sigma metadata is available through
transformer_options; if not, the SDXL patch falls back to the noisy/start side of the curve.
Limitations
- Not an official TIDE release.
- Not a full reproduction of the paper's complete experimental pipeline.
- Does not include paper benchmark scripts, datasets, generated result images, or metric evaluation.
- Does not modify the sampler's high-resolution time-shift schedule.
- Does not fully implement NTK-by-parts, YaRN positional interpolation, or DyPE positional interpolation.
- Flux support is tied to ComfyUI's Flux-style attention patch contract.
- WAN support applies Dynamic Temperature Control only; WAN Text Anchoring is intentionally not implemented.
- WAN visual quality needs live workflow testing across WAN 2.1/2.2 variants, frame counts, resolutions, samplers, and attention backends.
- MiniMax H3 quality, motion, audio, and AV synchronization require live checkpoint testing across T2VA, keyframe, and Ref2VA paths.
- H3 spatial and all-axis DTC are experimental and disabled by default.
- Active H3 Text Anchoring prefers the optional unmasked SageAttention LSE path. Environments without a compatible
sageattentionpackage fall back to PyTorch SDPA and can remain substantially slower. - Native ComfyUI H3 layout, token-tag, block-replacement, or RoPE contract changes can require a compatibility update and are rejected when detected.
- SDXL support is an experimental attention-temperature adaptation, not full Text Anchoring.
- SDXL visual quality needs live workflow testing across checkpoints, resolutions, samplers, and attention backends.
- Very large resolutions still require sufficient VRAM for the selected model, sampler, attention path, latent size, and VAE path.
License
This repository is released under the MIT License; see LICENSE.
ComfyUI is GPL-3.0 licensed. This custom node is distributed as a separate plugin, but it imports and runs inside ComfyUI. Review license compatibility before redistributing this node as part of a larger bundled package.