Commit Graph
3 Commits
Author SHA1 Message Date
Æmotion StudioandClaude Fable 5.1 f666d4e908 feat: gaussian shader type, white noise as the control
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-22 01:23:31 -07:00
Æmotion StudioandClaude Opus 5 298e8bdc72 perf: batch tensor_field's channel axis too
The last generator still drew one channel at a time, and it drew them
differently from the others: each channel takes one of four visualisations,
its own scale, warp and time, and its own random perturbation of the
coordinate grid. The expensive part -- five simplex evaluations inside
compute_tensor_properties -- does not depend on the visualisation, so it now
runs once for the whole channel axis and each channel takes the cheap
visualisation its index asks for. Computing all four and selecting is less
work than grouping the channels and making four calls.

At 1344x768/124 frames, measured against the pre-batched-noise tag in one
sitting: 5.37s -> 2.08s. That completes the set -- domain_warp 6.76 -> 1.90,
curl_noise 6.54 -> 1.94, temporal_coherent 5.76 -> 1.66 -- with effective
channel rank unchanged in all four.

Two things this cost, both recorded in tests:

- Its fixtures moved. Exact at vector-aligned shapes, up to 1.3e-05 at
  others. The amplification is larger than domain_warp's for a reason worth
  knowing: tensor_field perturbs the coordinates per channel, and an ulp of
  coordinate can push a point across a simplex cell boundary and flip a
  gradient, so a 1e-07 input difference is not a 1e-07 output difference.
  It was 5.6e-04 until the per-channel `time` was held in float64 and cast at
  the point of use, which is where the scalar path's Python float was rounded.
  Rounding it any earlier moves the coordinate.
- Channel 0 must stay on the scalar path. Batching it along with the rest
  broke the identity between channel 0 of a wide draw and a one-channel draw
  at 22x38, by 6e-07 -- small, but that identity is what every travel-mode
  basis rests on, and jump and stamp are pinned byte-identical to
  pre-collapse-fix through it. The other three generators get this for free by
  handing channel 0 to fill_channels as `base`; tensor_field does not use
  fill_channels, so it needed saying explicitly.

The new test that found the second one now runs all four generators across
nine latent shapes, ragged and aligned, so the invariant is checked where it
actually breaks rather than only where it holds. 32 jump/drift draws remain
byte-identical to the tag.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 19:06:33 -07:00
Æmotion StudioandClaude Opus 5 ee7c7bc31d perf: stop rendering the draw a collapse throws away
None of this changes what a seed produces. All eleven golden fixtures are
byte-identical and the suite passes untouched; that was the acceptance
criterion for every change here.

`travel_mode: jump` -- and the `jump` and `stamp` presets -- rebuilds every
channel from one-channel draws and never reads the wide draw, but the wide
draw was rendered first anyway and dropped on the next line. The guard that
discards it turns only on values known before rendering, so it moves up into
generate(). Measured on domain_warp at MiniMax H3's latent: 5.99s to 0.22s
at 1344x768/124 frames, 1.44s to 0.05s at 608x352/56 frames. `drift` is
unaffected -- it needs the wide draw to measure its rank, so it still pays.

The draw now runs under torch.inference_mode(), worth about a tenth of it,
and clones on the way out. That clone is not optional: an inference tensor
raises "Inference tensors cannot be saved for backward" the moment it
reaches a grad-recording region, and this noise is handed to whatever the
workflow does next.

domain_warp, curl_noise and temporal_coherent stopped reseeding the global
RNG. All three are coordinate hashes of their seed argument, so the calls
changed nothing -- verified by neutering torch.manual_seed and re-drawing --
and domain_warp made one per channel render, 888 times per draw at H3's
default latent, each reseeding every CUDA device as well as the CPU.
tensor_field keeps its reseed: it draws torch.randn_like inside its channel
loop and that value reaches its output.

Also here, because the batching work needs it:

- shaders/simplex.py holds one copy of the 2D primitive, which was duplicated
  across three generators, and takes a tensor of seeds as well as an int so a
  whole channel axis can be drawn in one call. domain_warp delegates to it,
  bit-identical and at performance parity. The variants are kept rather than
  unified -- domain_warp rotates its coordinates by the seed and the others do
  not -- because collapsing them would silently change three generators.
- tests/test_simplex.py pins how far batching stays exact. Without the
  rotation it is exact everywhere. With it, the coordinates differ per slice
  and have to be materialised, every downstream op then runs over N times as
  many elements, and where the per-slice count does not suit the vector width
  the tail differs by 3e-08 to 1.2e-07. So batching domain_warp will move the
  fixtures, and batching the other three will not.
- verification/benchmark_draw.py, because none of the above was reproducible:
  the only per-draw timing in the repo predated the per-channel fill and
  understated the cost roughly tenfold. It refuses to let you raise the thread
  count quietly -- 8 threads on this 4-core box is 2.5x slower than 4.
- SNK_TEST_CPU=1 pins the suite to the CPU, so it runs while a ComfyUI server
  on the same box is holding the GPU.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 18:08:46 -07:00