Files
AEmotionStudio-ComfyUI-Shad…/tests
Æmotion StudioandClaude Opus 5 2c7047a0e3 perf: draw the channel axis in one call, not one per channel
At MiniMax H3's default latent the draw performed 37 frames x 24 channels =
888 separate renders of 4032 pixels each, and cProfile put the time in per-op
dispatch rather than arithmetic -- a 4032-element torch.floor was costing
0.3ms. Generators now offer fill_channels a batched draw and it takes it
whenever more than one extra channel is wanted.

At 1344x768/124 frames: domain_warp 5.31s -> 1.97s, temporal_coherent
6.52s -> 1.86s, curl_noise -> 2.11s.

curl_noise and temporal_coherent are byte-identical; their fixtures did not
move. domain_warp is not, and the reason is worth recording. It turns its
coordinates by an angle drawn from the seed, so unlike the others the
coordinates genuinely differ per channel and the batched draw has to
materialise them. Every downstream elementwise op then runs over N times as
many elements, and where the per-slice count does not suit the vector width
the tail is handled differently. Five fixtures moved by 1.2e-07 to 2.4e-07 --
float32 rounding, with effective channel rank identical to three decimals --
and the previous behaviour is tagged pre-batched-noise. Exact at 22x38, 48x84
and 64x64; not at 22x39, 23x38, 37x53. tests/test_simplex.py pins both sides
of that boundary so it cannot quietly get worse.

jump and stamp are unaffected by construction: fill_channels takes the
batched path only for more than one channel, and both travel modes build from
one-channel draws. 32 draws across four generators and four latent shapes
were checked byte-identical against the tag.

Also folded in, because the batching needed it:

- One simplex primitive instead of seven. The 2D hash was duplicated across
  three generators and the 3D across four, with real differences hidden
  between them -- three of the 3D copies sum a single corner rather than
  four, and one of those computed three more corner hashes and threw them
  away. shaders/simplex.py names them for what they do and keeps the
  differences rather than unifying them. temporal_coherent's four-corner
  version stays with that generator: it picks corners by the real simplex
  ordering and reads gradients from a table, so it is a different function
  and not a parameterisation of the others. I had written it as one, guessed
  wrong, and the equivalence check caught it.
- The accumulators in get_velocity_field, fbm_noise and temporal_spectral_noise
  no longer preallocate a fixed [B,H,W,C] and write into it in place, so the
  same code serves a shared coordinate grid and a per-channel one. The adds
  happen in the same order, so the values are unchanged.
- tensor_field draws its shape mask once instead of once per channel -- 127
  identical masks at LTXV's 128 channels -- and stops cloning the coordinate
  grid per channel. Byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 18:54:02 -07:00
..