docs: correct the ceilings for decorrelation, which inverts the ordering
The tooltips quoted ceilings measured before decorrelate_channels existed, so
they were understating the usable range by roughly three times for the default
generator. Re-derived on real H3 at the current basis of 64:
generator without with
domain_warp ~0.25 ~0.75 (degrades at 1.0)
temporal_coherent ~0.35 ~0.75
tensor_field ~0.5 ~0.5 unchanged, skipped
curl_noise ~0.2 ~0.2 unchanged, skipped
This reverses the earlier finding that tensor_field tolerated the most strength.
Rank was the only thing limiting domain_warp; once that is fixed it overtakes the
generators decorrelation skips, which are limited by their spatial character
instead and gain nothing.
The sweep ran with zero conditioning, so content comes from the model's prior
rather than a prompt. Coherence is still unambiguous, and the control in the same
batch reproduced the green quilt seen earlier under real prompts, so the
comparison holds.
Two handoff items close with this; a third opens. Rank still tops out near 60% of
channels because the basis draws are not independent of each other either, so
mixing genuinely orthogonal fields would close the remaining gap.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 5
parent
c938137645
commit
f16390cff7
+34
-10
@@ -41,8 +41,27 @@ step count. See item 2. Ceilings measured on H3 (608x352, 8 steps,
|
||||
| shape masks | ~0.2 | 0.6 (mask drawn into the picture) |
|
||||
| `use_temporal_coherence` | ~0.2 | 0.5 (swamps the frame) |
|
||||
|
||||
With `decorrelate_channels` on, `domain_warp` is clean at 0.5 on H3 and 0.25 on
|
||||
SD 1.5 where both previously failed.
|
||||
Re-derived on H3 with `decorrelate_channels` on at the current basis of 64, which
|
||||
**inverts the ordering**:
|
||||
|
||||
| Generator | Ceiling without | With |
|
||||
|---|---|---|
|
||||
| `domain_warp` | ~0.25 | **~0.75** (degrades at 1.0) |
|
||||
| `temporal_coherent` | ~0.35 | **~0.75** |
|
||||
| `tensor_field` | ~0.5 | ~0.5 — unchanged, decorrelation skips it |
|
||||
| `curl_noise` | ~0.2 | ~0.2 — unchanged, skipped |
|
||||
|
||||
The two generators decorrelation fixes now beat the two it skips, reversing the
|
||||
earlier finding that `tensor_field` was the most tolerant: rank was what limited
|
||||
`domain_warp`, and nothing else was. The generators that already span their
|
||||
channels are limited by their spatial character instead, which decorrelation
|
||||
does not touch.
|
||||
|
||||
That sweep ran with zero conditioning (no text encoder resident), so the content
|
||||
comes from the model's prior rather than a prompt. Coherence is still
|
||||
unambiguous, and the control in the same batch — `domain_warp` at 0.75 with
|
||||
decorrelation off — reproduced the same green quilt seen earlier under real
|
||||
prompts, so the comparison holds.
|
||||
|
||||
**2. The audio stream moves even when nothing touches it.** Shader noise reaches
|
||||
only the spatial stream by default and the audio keeps its Gaussian noise
|
||||
@@ -144,12 +163,17 @@ other caller, including legacy. Doing it properly changes legacy output, which
|
||||
`tests/golden_cases.py` pins deliberately, so it needs the same gating
|
||||
discussion as item 1.
|
||||
|
||||
**Tune `DECORRELATION_BASIS`.** It is 8, chosen to bound cost at LTXV's 128
|
||||
channels. Nobody has tested whether 4 is as good or 16 better. Cost is roughly
|
||||
linear: 2-12x the noise-generation time, small against sampling but not free.
|
||||
~~Tune `DECORRELATION_BASIS`.~~ Done (`7661d7c`): it shipped at 8 on a guess and
|
||||
is now 64. On H3 at strength 0.75, basis 8 is still mostly destroyed, 16 is
|
||||
coherent and 32 is clean. The cost the old value guarded against did not exist —
|
||||
worst measured case is about a second per draw, one draw per stage boundary, on
|
||||
runs of thirty to fifty seconds.
|
||||
|
||||
**Re-derive the ceilings table with it on.** Every number in finding (1) was
|
||||
measured with it off, and the knobs interact.
|
||||
~~Re-derive the ceilings table with it on.~~ Done — see finding (1).
|
||||
|
||||
**Rank still tops out around 60% of channels**, because the basis draws are not
|
||||
fully independent of each other either. Mixing genuinely orthogonal fields rather
|
||||
than random combinations of correlated ones would close the rest of the gap.
|
||||
|
||||
---
|
||||
|
||||
@@ -223,9 +247,9 @@ anyone claims a direction. Nobody has *listened* to the output.
|
||||
correct shapes and finite unit-variance noise, and the pipeline accepts it, but
|
||||
no Stable Audio or ACE-Step checkpoint has been run through it.
|
||||
|
||||
**TripoSplat's second stream is camera parameters**, not audio. Enabling this
|
||||
there paints the camera. A per-stream opt-in would be safer than the current
|
||||
all-or-nothing.
|
||||
~~TripoSplat's second stream is camera parameters.~~ Handled (`c938137`):
|
||||
streams carrying fewer than 64 cells per batch item are skipped, which catches
|
||||
the `[B, 1, 5]` camera while leaving H3's 414-cell audio painted.
|
||||
|
||||
---
|
||||
|
||||
|
||||
@@ -21,13 +21,13 @@ class DirectShaderNoiseKSampler(ShaderNoiseKSampler):
|
||||
"denoise": ("FLOAT", {"default": 1.0, "min": 0.0, "max": 1.0, "step": 0.01, "tooltip": "Denoising strength. Lower values preserve more of the original image"}),
|
||||
"sequential_stages": ("INT", {"default": 1, "min": 0, "max": 10, "step": 1, "tooltip": "Number of sequential shader stages to apply before injection stages"}),
|
||||
"injection_stages": ("INT", {"default": 0, "min": 0, "max": 10, "step": 1, "tooltip": "Number of injection shader stages to apply after sequential stages"}),
|
||||
"shader_strength": ("FLOAT", {"default": 0.3, "min": 0.0, "max": 1.0, "step": 0.01, "tooltip": "How much shader noise replaces the base noise. 0.0 disables it. Raising it walks further from the seed, but past a point the shader's own structure survives denoising and shows up in the image. Video models reach that point early: on MiniMax H3, domain_warp stays photoreal to about 0.25 and is gone by 0.75. Start low and climb. curl_noise, shape masks and temporal coherence all need roughly half the value you would use otherwise."}),
|
||||
"shader_strength": ("FLOAT", {"default": 0.3, "min": 0.0, "max": 1.0, "step": 0.01, "tooltip": "How much shader noise replaces the base noise. 0.0 disables it. Raising it walks further from the seed, but past a point the shader's own structure survives denoising and shows up in the image. Video models reach that point early: on MiniMax H3, domain_warp stays photoreal to about 0.25 and is gone by 0.75. Turning on decorrelate_channels roughly triples that, to about 0.75. Start low and climb. curl_noise, shape masks and temporal coherence all need roughly half the value you would use otherwise."}),
|
||||
"blend_mode": (["normal", "add", "multiply", "screen", "overlay", "soft_light", "hard_light", "difference"], {"default": "multiply", "tooltip": "How shader noise is mixed into the base noise. Gentlest first: difference and soft_light tolerate the highest strength, then multiply and normal, then overlay, screen and hard_light; add is the most aggressive and needs the lowest strength. All of them keep mean 0 / std 1, so the sampler still gets the distribution it expects."}),
|
||||
"noise_transform": (["none", "reverse", "inverse", "absolute", "square", "sqrt", "log", "sin", "cos"], {"default": "none", "tooltip": "Apply mathematical transformations to the noise for creative effects"}),
|
||||
"use_temporal_coherence": ("BOOLEAN", {"default": False, "tooltip": "Hold one seed across every video frame so the shader pattern evolves only through time, instead of redrawing per frame. Ties frames together, but because the pattern no longer varies between them it reinforces rather than averages out: halve your shader_strength when you turn this on. On MiniMax H3 at 0.5 it swamps the picture, while 0.2 is clean. No effect on single images."}),
|
||||
|
||||
# New direct shader parameters
|
||||
"shader_type": (["domain_warp", "tensor_field", "curl_noise", "temporal_coherent"], {"default": "domain_warp", "tooltip": "Which noise pattern to walk with. domain_warp: flowing, intricate distortions, the most even-handed default. tensor_field: structured and directional, the most tolerant of high strength. curl_noise: smooth fluid motion, but the most aggressive -- use roughly half the strength. temporal_coherent: 4D simplex with time as a real axis, built for smooth animation and the best suited to video. The live preview only draws the first three; picking temporal_coherent leaves the preview on its last pattern, which does not affect sampling."}),
|
||||
"shader_type": (["domain_warp", "tensor_field", "curl_noise", "temporal_coherent"], {"default": "domain_warp", "tooltip": "Which noise pattern to walk with. domain_warp: flowing, intricate distortions, the most even-handed default. tensor_field: structured and directional. curl_noise: smooth fluid motion, but the most aggressive -- use roughly half the strength. temporal_coherent: 4D simplex with time as a real axis, built for smooth animation and the best suited to video. Which tolerates the most strength depends on decorrelate_channels: without it tensor_field leads, because the others return noise spanning only one or two channels. With it, domain_warp and temporal_coherent stay clean to about 0.75 on MiniMax H3 while tensor_field and curl_noise are unchanged -- they already span their channels, so decorrelation skips them. The live preview only draws the first three; picking temporal_coherent leaves the preview on its last pattern, which does not affect sampling."}),
|
||||
"shape_type": (["none", "radial", "linear", "spiral", "checkerboard", "spots", "hexgrid", "stripes", "gradient", "vignette", "cross", "stars", "triangles", "concentric", "rays", "zigzag"], {"default": "none", "tooltip": "Mask the shader noise into a shape before it reaches the sampler (not post-processing). A mask concentrates the noise into hard geometry, so it survives denoising far more readily than plain shader noise -- on MiniMax H3 at strength 0.6 the mask itself is drawn into the picture. Keep strength at or below about 0.2 when a shape is active."}),
|
||||
"color_scheme": (["none", "blue_red", "viridis", "plasma", "inferno", "magma", "turbo", "jet", "rainbow", "cool", "hot", "parula", "hsv", "autumn", "winter", "spring", "summer", "copper", "pink", "bone", "ocean", "terrain", "neon", "fire"], {"default": "none", "tooltip": "Choose a color palette to apply to the shader noise visualization [not post processing - is applied to the shader noise pattern before rendering]"}),
|
||||
"noise_scale": ("FLOAT", {"default": 1.0, "min": 0.1, "max": 10.0, "step": 0.001, "tooltip": "Adjust the scale of the shader noise pattern - lower values create larger, zoomed-in features; higher values create smaller, zoomed-out features [small value shifts can lead to larger variations]"}),
|
||||
@@ -47,7 +47,7 @@ class DirectShaderNoiseKSampler(ShaderNoiseKSampler):
|
||||
"fast_high_channel_noise": ("BOOLEAN", {"default": False, "tooltip": "Use a faster, simplified noise generation method for models with many channels (>16), like LTXV"}),
|
||||
"stage_progression": (["uniform", "coarse_to_fine", "fine_to_coarse"], {"default": "uniform", "tooltip": "Vary the shader across the run instead of drawing the same one at every stage. The trajectory is not uniform -- early steps settle composition, late steps settle detail -- but every stage has always used the same zoom. coarse_to_fine starts zoomed in on large features with fewer octaves and ends zoomed out on small ones with more, so the noise matches what each part of the run is deciding; fine_to_coarse reverses it. The adjustment spans 0.5x to 2x your noise_scale and plus or minus one octave, centred on your widget values, so uniform is unchanged. Needs more than one stage to do anything. Standard sampling only."}),
|
||||
"shade_non_spatial": ("BOOLEAN", {"default": False, "tooltip": "Also paint the streams that have no picture in them. Off, the shader touches only the spatial stream and everything else keeps the Gaussian noise ComfyUI gave it -- on MiniMax H3 and LTXAV that means the audio is left alone, and sequence latents (Stable Audio, ACE-Step 1.5, MiniMax Music 3, Hunyuan3D, TripoSplat) are refused outright. On, an audio stream is painted across stereo x time, and a sequence latent is painted as a single row. Video and audio are denoised together on H3, so this reaches the picture too. Unexplored and easy to overdo: audio has no busy scene to hide structure in, so start near 0.05. Streams too small to be content are skipped, so TripoSplat's camera parameters are left alone. Standard sampling only."}),
|
||||
"decorrelate_channels": ("BOOLEAN", {"default": False, "tooltip": "Give every latent channel its own shader draw instead of copies of one. The generators build extra channels as pointwise functions of the first one or two, so domain_warp returns noise spanning a single channel at SD's four and about two at any larger count, and temporal_coherent returns literally identical channels. Samplers expect independent noise, and that collapse is the main reason the shader's own pattern surfaces so readily: on SD 1.5 it moves the usable ceiling from below 0.25 to around 0.5. Costs a few extra noise renders. Generators that already span their channels, such as tensor_field, are detected and left untouched. Off by default so existing workflows reproduce; standard sampling only."}),
|
||||
"decorrelate_channels": ("BOOLEAN", {"default": False, "tooltip": "Give every latent channel its own shader draw instead of copies of one. The generators build extra channels as pointwise functions of the first one or two, so domain_warp returns noise spanning a single channel at SD's four and about two at any larger count, and temporal_coherent returns literally identical channels. Samplers expect independent noise, and that collapse is the main reason the shader's own pattern surfaces so readily: on SD 1.5 it moves the usable ceiling from below 0.25 to around 0.5, and on MiniMax H3 from about 0.25 to about 0.75. Costs a few extra noise renders. Generators that already span their channels, such as tensor_field, are detected and left untouched. Off by default so existing workflows reproduce; standard sampling only."}),
|
||||
"normalize_strength": ("BOOLEAN", {"default": False, "tooltip": "Make shader_strength mean the same thing in every blend mode. Untouched, the modes differ by up to twenty-three times at the same setting: at 0.5 normal hands the sampler 0.71 of the shader and difference only 0.03. With this on, strength is read on multiply's scale, so the default mode is unchanged and the others are rescaled to match -- soft_light needs about 1.6x its old number, add and hard_light about half. difference cannot reach the top of the scale at all and saturates. Off by default so existing workflows reproduce; standard sampling only."}),
|
||||
},
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user