The last generator still drew one channel at a time, and it drew them
differently from the others: each channel takes one of four visualisations,
its own scale, warp and time, and its own random perturbation of the
coordinate grid. The expensive part -- five simplex evaluations inside
compute_tensor_properties -- does not depend on the visualisation, so it now
runs once for the whole channel axis and each channel takes the cheap
visualisation its index asks for. Computing all four and selecting is less
work than grouping the channels and making four calls.
At 1344x768/124 frames, measured against the pre-batched-noise tag in one
sitting: 5.37s -> 2.08s. That completes the set -- domain_warp 6.76 -> 1.90,
curl_noise 6.54 -> 1.94, temporal_coherent 5.76 -> 1.66 -- with effective
channel rank unchanged in all four.
Two things this cost, both recorded in tests:
- Its fixtures moved. Exact at vector-aligned shapes, up to 1.3e-05 at
others. The amplification is larger than domain_warp's for a reason worth
knowing: tensor_field perturbs the coordinates per channel, and an ulp of
coordinate can push a point across a simplex cell boundary and flip a
gradient, so a 1e-07 input difference is not a 1e-07 output difference.
It was 5.6e-04 until the per-channel `time` was held in float64 and cast at
the point of use, which is where the scalar path's Python float was rounded.
Rounding it any earlier moves the coordinate.
- Channel 0 must stay on the scalar path. Batching it along with the rest
broke the identity between channel 0 of a wide draw and a one-channel draw
at 22x38, by 6e-07 -- small, but that identity is what every travel-mode
basis rests on, and jump and stamp are pinned byte-identical to
pre-collapse-fix through it. The other three generators get this for free by
handing channel 0 to fill_channels as `base`; tensor_field does not use
fill_channels, so it needed saying explicitly.
The new test that found the second one now runs all four generators across
nine latent shapes, ragged and aligned, so the invariant is checked where it
actually breaks rather than only where it holds. 32 jump/drift draws remain
byte-identical to the tag.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
None of this changes what a seed produces. All eleven golden fixtures are
byte-identical and the suite passes untouched; that was the acceptance
criterion for every change here.
`travel_mode: jump` -- and the `jump` and `stamp` presets -- rebuilds every
channel from one-channel draws and never reads the wide draw, but the wide
draw was rendered first anyway and dropped on the next line. The guard that
discards it turns only on values known before rendering, so it moves up into
generate(). Measured on domain_warp at MiniMax H3's latent: 5.99s to 0.22s
at 1344x768/124 frames, 1.44s to 0.05s at 608x352/56 frames. `drift` is
unaffected -- it needs the wide draw to measure its rank, so it still pays.
The draw now runs under torch.inference_mode(), worth about a tenth of it,
and clones on the way out. That clone is not optional: an inference tensor
raises "Inference tensors cannot be saved for backward" the moment it
reaches a grad-recording region, and this noise is handed to whatever the
workflow does next.
domain_warp, curl_noise and temporal_coherent stopped reseeding the global
RNG. All three are coordinate hashes of their seed argument, so the calls
changed nothing -- verified by neutering torch.manual_seed and re-drawing --
and domain_warp made one per channel render, 888 times per draw at H3's
default latent, each reseeding every CUDA device as well as the CPU.
tensor_field keeps its reseed: it draws torch.randn_like inside its channel
loop and that value reaches its output.
Also here, because the batching work needs it:
- shaders/simplex.py holds one copy of the 2D primitive, which was duplicated
across three generators, and takes a tensor of seeds as well as an int so a
whole channel axis can be drawn in one call. domain_warp delegates to it,
bit-identical and at performance parity. The variants are kept rather than
unified -- domain_warp rotates its coordinates by the seed and the others do
not -- because collapsing them would silently change three generators.
- tests/test_simplex.py pins how far batching stays exact. Without the
rotation it is exact everywhere. With it, the coordinates differ per slice
and have to be materialised, every downstream op then runs over N times as
many elements, and where the per-slice count does not suit the vector width
the tail differs by 3e-08 to 1.2e-07. So batching domain_warp will move the
fixtures, and batching the other three will not.
- verification/benchmark_draw.py, because none of the above was reproducible:
the only per-draw timing in the repo predated the per-channel fill and
understated the cost roughly tenfold. It refuses to let you raise the thread
count quietly -- 8 threads on this 4-core box is 2.5x slower than 4.
- SNK_TEST_CPU=1 pins the suite to the CPU, so it runs while a ComfyUI server
on the same box is holding the GPU.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>