Files
AEmotionStudio-ComfyUI-Shad…/pyproject.toml
T
Æmotion StudioandClaude Opus 5 363ac1e657 fix: sample MiniMax H3's paired audio+video latent correctly
H3's latent is a NestedTensor of a video stream [B,24,T,H,W] and an audio
stream [B,32,2,T], and the model -- not its latent format -- carries audio
scaled onto the video sigma schedule (audio_scale = shift / audio_shift = 4.0).

_split_noise inverted a segment boundary through latent_format.process_in,
which for MiniMaxH3AV is an identity, so the audio residual handed to the next
segment was wrong by that factor of 4. It now inverts through the model's own
process_latent_in / process_latent_out, which is what CFGGuider.inner_sample
actually applies. Every other model is unaffected: BaseModel.process_latent_in
just calls the format.

Measured on real H3 weights, two stages at shader_strength 0, where a segmented
run must reproduce an uninterrupted one:

                video max error            audio max error
    before      9.3e-01 (stream max 4.80)  1.3e+00 (stream max 1.35)
    after       4.8e-07                    2.4e-07

The audio stream was almost entirely wrong, and because H3 denoises both
streams in one packed sequence the error reached the video through the DiT's
joint attention -- so this degraded picture as well as sound. Verified
bit-identical output on SD 1.5 and Wan 2.1, confirming it is a no-op elsewhere.

Also:
- shader noise at a boundary reads its shape from the noise it is about to
  paint, rather than a shape captured before the run started
- latents with no spatial grid ([B,C,L]: Stable Audio, ACE-Step 1.5, MiniMax
  Music 3, Hunyuan3D, TripoSplat) raise UnsupportedLatentError naming the
  shape, before sampling starts, instead of a bare ValueError from inside noise
  generation. At shader_strength 0 they sample through as a plain KSampler.
- delete core/model_compat.py and its three stale tables. Nothing called it;
  its tables stopped at LTXV, its model_type == "FLOW" branch was unreachable
  (str(ModelType.FLOW).upper() is "MODELTYPE.FLOW"), and its 5-D layout guess
  defaulted to [B,F,C,H,W], which ComfyUI never produces. The legacy mode's own
  detector is untouched, so pre-2.0 workflows still reproduce their seeds.

Noise generation is now exercised at every channel count ComfyUI ships -- 3, 4,
8, 12, 16, 24, 32, 48, 64, 128 and 256 -- for both image and video latents.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-12 01:32:10 -07:00

15 lines
637 B
TOML

[project]
name = "comfyui-shadernoiseksampler"
description = "Transform AI image generation from random exploration into deliberate artistic navigation. This advanced KSampler replacement blends traditional noise with shader noise. Navigate latent space with intention using adjustable noise parameters, shape masks, and colors transformations."
version = "2.1.0"
license = {file = "LICENSE"}
[project.urls]
Repository = "https://github.com/AEmotionStudio/ComfyUI-ShaderNoiseKSampler"
# Used by Comfy Registry https://comfyregistry.org
[tool.comfy]
PublisherId = "aemotionstudio"
DisplayName = "ComfyUI-ShaderNoiseKSampler"
Icon = ""