None of this changes what a seed produces. All eleven golden fixtures are
byte-identical and the suite passes untouched; that was the acceptance
criterion for every change here.
`travel_mode: jump` -- and the `jump` and `stamp` presets -- rebuilds every
channel from one-channel draws and never reads the wide draw, but the wide
draw was rendered first anyway and dropped on the next line. The guard that
discards it turns only on values known before rendering, so it moves up into
generate(). Measured on domain_warp at MiniMax H3's latent: 5.99s to 0.22s
at 1344x768/124 frames, 1.44s to 0.05s at 608x352/56 frames. `drift` is
unaffected -- it needs the wide draw to measure its rank, so it still pays.
The draw now runs under torch.inference_mode(), worth about a tenth of it,
and clones on the way out. That clone is not optional: an inference tensor
raises "Inference tensors cannot be saved for backward" the moment it
reaches a grad-recording region, and this noise is handed to whatever the
workflow does next.
domain_warp, curl_noise and temporal_coherent stopped reseeding the global
RNG. All three are coordinate hashes of their seed argument, so the calls
changed nothing -- verified by neutering torch.manual_seed and re-drawing --
and domain_warp made one per channel render, 888 times per draw at H3's
default latent, each reseeding every CUDA device as well as the CPU.
tensor_field keeps its reseed: it draws torch.randn_like inside its channel
loop and that value reaches its output.
Also here, because the batching work needs it:
- shaders/simplex.py holds one copy of the 2D primitive, which was duplicated
across three generators, and takes a tensor of seeds as well as an int so a
whole channel axis can be drawn in one call. domain_warp delegates to it,
bit-identical and at performance parity. The variants are kept rather than
unified -- domain_warp rotates its coordinates by the seed and the others do
not -- because collapsing them would silently change three generators.
- tests/test_simplex.py pins how far batching stays exact. Without the
rotation it is exact everywhere. With it, the coordinates differ per slice
and have to be materialised, every downstream op then runs over N times as
many elements, and where the per-slice count does not suit the vector width
the tail differs by 3e-08 to 1.2e-07. So batching domain_warp will move the
fixtures, and batching the other three will not.
- verification/benchmark_draw.py, because none of the above was reproducible:
the only per-draw timing in the repo predated the per-channel fill and
understated the cost roughly tenfold. It refuses to let you raise the thread
count quietly -- 8 threads on this 4-core box is 2.5x slower than 4.
- SNK_TEST_CPU=1 pins the suite to the CPU, so it runs while a ComfyUI server
on the same box is holding the GPU.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>