H3's latent is a NestedTensor of a video stream [B,24,T,H,W] and an audio
stream [B,32,2,T], and the model -- not its latent format -- carries audio
scaled onto the video sigma schedule (audio_scale = shift / audio_shift = 4.0).
_split_noise inverted a segment boundary through latent_format.process_in,
which for MiniMaxH3AV is an identity, so the audio residual handed to the next
segment was wrong by that factor of 4. It now inverts through the model's own
process_latent_in / process_latent_out, which is what CFGGuider.inner_sample
actually applies. Every other model is unaffected: BaseModel.process_latent_in
just calls the format.
Measured on real H3 weights, two stages at shader_strength 0, where a segmented
run must reproduce an uninterrupted one:
video max error audio max error
before 9.3e-01 (stream max 4.80) 1.3e+00 (stream max 1.35)
after 4.8e-07 2.4e-07
The audio stream was almost entirely wrong, and because H3 denoises both
streams in one packed sequence the error reached the video through the DiT's
joint attention -- so this degraded picture as well as sound. Verified
bit-identical output on SD 1.5 and Wan 2.1, confirming it is a no-op elsewhere.
Also:
- shader noise at a boundary reads its shape from the noise it is about to
paint, rather than a shape captured before the run started
- latents with no spatial grid ([B,C,L]: Stable Audio, ACE-Step 1.5, MiniMax
Music 3, Hunyuan3D, TripoSplat) raise UnsupportedLatentError naming the
shape, before sampling starts, instead of a bare ValueError from inside noise
generation. At shader_strength 0 they sample through as a plain KSampler.
- delete core/model_compat.py and its three stale tables. Nothing called it;
its tables stopped at LTXV, its model_type == "FLOW" branch was unreachable
(str(ModelType.FLOW).upper() is "MODELTYPE.FLOW"), and its 5-D layout guess
defaulted to [B,F,C,H,W], which ComfyUI never produces. The legacy mode's own
detector is untouched, so pre-2.0 workflows still reproduce their seeds.
Noise generation is now exercised at every channel count ComfyUI ships -- 3, 4,
8, 12, 16, 24, 32, 48, 64, 128 and 256 -- for both image and video latents.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
96 lines
1.8 KiB
Python
96 lines
1.8 KiB
Python
"""
|
|
Constants and magic numbers for shader noise generation.
|
|
|
|
This module centralizes all constants used across the codebase to
|
|
ensure consistency and make values easy to tune.
|
|
"""
|
|
|
|
# Hash constant used in simplex noise and pseudo-random functions
|
|
HASH_CONSTANT = 43758.5453
|
|
|
|
# Default number of channels for latent space
|
|
DEFAULT_CHANNELS = 4
|
|
|
|
# Maximum octaves for FBM noise to prevent DoS
|
|
MAX_OCTAVES = 20
|
|
|
|
# Minimum and maximum scale values
|
|
MIN_SCALE = 0.001
|
|
MAX_SCALE = 100.0
|
|
|
|
# Default parameter values
|
|
DEFAULT_SCALE = 1.0
|
|
DEFAULT_OCTAVES = 3.0
|
|
DEFAULT_WARP_STRENGTH = 0.5
|
|
DEFAULT_PHASE_SHIFT = 0.5
|
|
DEFAULT_SHAPE_STRENGTH = 1.0
|
|
DEFAULT_COLOR_INTENSITY = 0.8
|
|
DEFAULT_SHADER_STRENGTH = 0.3
|
|
DEFAULT_TIME = 0.0
|
|
|
|
# Supported blend modes for combining shader noise with base noise
|
|
SUPPORTED_BLEND_MODES = [
|
|
"normal",
|
|
"add",
|
|
"multiply",
|
|
"screen",
|
|
"overlay",
|
|
"soft_light",
|
|
"hard_light",
|
|
"difference",
|
|
]
|
|
|
|
# Supported noise transforms
|
|
SUPPORTED_TRANSFORMS = [
|
|
"none",
|
|
"reverse",
|
|
"inverse",
|
|
"absolute",
|
|
"square",
|
|
"sqrt",
|
|
"log",
|
|
"sin",
|
|
"cos",
|
|
]
|
|
|
|
# Supported shader types
|
|
SUPPORTED_SHADER_TYPES = [
|
|
"domain_warp",
|
|
"tensor_field",
|
|
"curl_noise",
|
|
"temporal_coherent",
|
|
"fbm_noise",
|
|
"perlin",
|
|
"waves",
|
|
"gaussian",
|
|
"heterogeneous_fbm",
|
|
"interference",
|
|
"spectral",
|
|
"projection_3d",
|
|
]
|
|
|
|
# Supported stage distributions for multi-stage sampling
|
|
SUPPORTED_DISTRIBUTIONS = [
|
|
"uniform",
|
|
"linear_decrease",
|
|
"linear_increase",
|
|
"gaussian",
|
|
"first_stronger",
|
|
"last_stronger",
|
|
]
|
|
|
|
# High channel threshold for fast mode
|
|
HIGH_CHANNEL_THRESHOLD = 16
|
|
|
|
# Visualization types
|
|
VISUALIZATION_TYPES = {
|
|
0: "arrows",
|
|
1: "lines",
|
|
2: "dots",
|
|
3: "ellipses",
|
|
4: "streamlines",
|
|
}
|
|
|
|
# Default visualization type
|
|
DEFAULT_VISUALIZATION_TYPE = 3 # ellipses
|