Compare commits

..
Author SHA1 Message Date
Will LinandClaude Fable 5 0a861031c7 [docs] H3 parallel VAE: document the compiled-decoder cross-process determinism caveat
With enable_torch_compile_vae (#1734, opt-in) inductor autotunes kernels per
process, so chunk decodes on other ranks differ from the serial rank's decode
the way two serial processes differ (GB200 @124f: max 63/255 on <0.5% of
pixels, mean ~1e-2/255; audio and chunk 0 bit-identical). Eager decoder (the
default) stays bitwise-equal to serial decode_to_pixels - measured, both
strategies, x3, 124f+345f.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 18:17:28 +00:00
Will LinandClaude Fable 5 f08c5ee8af [perf] H3 parallel VAE: overlap first decodes with the meta rendezvous; assembly on a side stream
Two schedule fixes sized from the first GB200 tray measurements (job 2659,
serial 7.7s/21.2s at 124f/345f):

1. Every rank now decodes its round-0 chunk BEFORE the metadata broadcast.
   Non-leader ranks previously blocked on the broadcast until the leader
   finished chunk 0, serializing a full extra chunk-decode into round 0
   (visible as 1.9x instead of ~2.6x at 7 chunks / 4 ranks). Same reorder
   on the encode path.

2. The leader's per-chunk joining work (blend, denormalize, clamp, output
   copies) moves to a dedicated CUDA side stream. It depends only on
   already-gathered segments, but on the main stream it delayed the
   leader's next-round decode and therefore every rank's next collective
   (~0.1s/chunk on the critical path). Gathered storage is pinned to the
   assembly stream via record_stream; the driver drains the stream in a
   finally so an exception cannot leave an in-flight DMA into the output
   buffer. Stream placement does not change op order or values, so the
   bitwise-parity contract is untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 17:59:18 +00:00
Will LinandClaude Fable 5 755f4a7967 [docs] schema parity inventory: classify vae_parallel_* (+ the stack's unclassified VSA_tile_size)
vae_parallel_decode / vae_parallel_encode / vae_parallel_decode_strategy are
model-specific optimization knobs (compatibility_only, like VSA_sparsity).
VSA_tile_size came in with the merged tile-64 route without an inventory
entry and failed test_fastvideo_args_fields_are_classified on the whole
stack; classify it the same way. The remaining pipeline_config inventory
gaps (image_encoder_precisions, ...) predate this branch and are left for
the owning PRs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 17:53:41 +00:00
Will LinandClaude Fable 5 d543a67b10 [perf] MiniMax-H3 VAE: SP-rank-parallel chunk decode + reference-clip encode (opt-in)
Under SP>1 the H3 video VAE decoded all temporal chunks serially on the
output rank while the other ranks idled (#1703's gate), and every rank
encoded the full reference video redundantly. Chunk decodes and clip
encodes have no cross-chunk data dependency - only the joining (overlap
blend, trim, denormalize, moment concat) is sequential - so both are
round-robined across the sequence-parallel ranks:

- fastvideo/models/vaes/minimax_h3_parallel.py: decode_to_pixels_parallel
  gathers each round's decoded segments (body+halo tail, one contiguous
  slice per chunk) to the SP group's first rank via NCCL gather (or
  all_gather, FASTVIDEO_VAE_PARALLEL_DECODE_STRATEGY), which replays the
  serial blend/trim/denormalize/copy semantics with the same VAE methods -
  bitwise-equal to serial decode_to_pixels by construction. Placeholder
  rounds keep collective participation uniform; the leader decodes chunk 0
  first and broadcasts dtype/shape metadata so placeholders never guess the
  autocast dtype. encode_pixels_parallel all-gathers per-clip moments
  (latent-sized) so every rank keeps the identical full posterior,
  preserving the all-ranks-hold-latents contract.
- decoding stage: output gate moves from world rank 0 to the SP group's
  first rank (identical in the single-group e2e case; correct for the
  trainer validation callback, which consumes each group leader's batch);
  with vae_parallel_decode every rank enters the decode body so no
  rank-dependent branch guards the collectives.
- latent preparation: opt-in clip-parallel reference encode on the same
  seam (vae_parallel_encode).
- knobs: FastVideoArgs.vae_parallel_decode/encode (+ --vae-parallel-decode,
  --vae-parallel-encode, FASTVIDEO_VAE_PARALLEL_DECODE/ENCODE env
  parse-once adapters), default OFF.
- _copy_chunk_pixels factored out of _decode_to_pixels so serial and
  parallel share one output-copy path (behavior unchanged).

Tests: threaded fake-group CPU suite drives the real SPMD functions
end-to-end (world sizes 2-5, both strategies, pad/blend/trim geometries,
token_drop=0, batched slicing, placeholder rounds) bit-exact vs the serial
APIs; GPU regression (torchrun world>1 gated) asserts bitwise parity under
fp16 autocast with real NCCL plus repeat-determinism x3.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 17:38:38 +00:00
Will LinandClaude Fable 5 741aa8d289 [bugfix] logger: info_once crashed on the patched process-aware info (duplicate stacklevel)
_print_info_once passes stacklevel=2 into logger.info, and init_logger's
patched _info passed its own stacklevel=2 positionally into logger.log on
top of the caller's kwarg -> TypeError on every info_once call. Honor an
explicit stacklevel instead of passing the keyword twice.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 17:38:18 +00:00
Will LinandClaude Fable 5 ad9cd63122 [bugfix] H3 conditioner: no_grad instead of inference_mode - FSDP2-sharded encode crashed (#1732 default-path gap)
Carried blocker fix discovered while benchmarking this stack (GB200 job 2594,
all four legs): with text_encoder_cpu_offload=True - the FastVideoArgs DEFAULT
and the shipped H3 example configuration - TextEncoderLoader FSDP2-fully_shards
the H3 conditioner, and #1732's @torch.inference_mode() on
MiniMaxH3Qwen3VLConditioner.encode_ids then kills the very first encode:

  File torch/distributed/fsdp/_fully_shard/_fsdp_param_group.py, in
  wait_for_unshard: with torch.autograd._unsafe_preserve_version_counter(t):
  RuntimeError: Inference tensors do not track version counter.

FSDP2's lazy unshard runs inside the inference-mode region, so its all-gather
tensors are inference tensors, and the version-counter preservation hook
cannot read t._version. The PR's own benchmarks ran with offload disabled,
which is exactly the unexercised-default-path gap called out as finding 2 of
the pr1732 review (there for FP8; the same gap bites plain bf16 via
inference_mode).

Fix: @torch.no_grad() instead. It frees the same activation memory (the
-262 MiB claim comes from the early-return slim contract, not from
inference_mode's bookkeeping), is fully FSDP2-compatible, and as a bonus
retires review finding 7: prompt_embeds are ordinary tensors again, so any
future on-the-fly-conditioning training can backprop through them without a
clone at the stage boundary.

Verified: 1x GPU FastH3 leg boots and generates after this change (leg reruns
on GB200); the 1732 unit suites (truncation + checkpoint-fp8) still pass.
Note: fastvideo/tests/stages/test_text_encoding.py has 7 pre-existing failures
that reproduce byte-identically on plain origin/main 56d4a6074 (unrelated
generic-stage tests; not introduced by the stack, verified by A/B).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 10:49:47 +00:00
Will LinandClaude Fable 5 68e6ffca9e [bugfix] H3 VAE: clone the reduce-overhead stitched canvas at the tile-driver returns (#1734 review F1)
Carried blocker fix for PR #1734 on this integration stack, per pr1734.md
finding F1 (reproduced on GB200/torch 2.12): with
@torch.compile(mode="reduce-overhead") on _stitch_tiles, the stitched canvas
is a CUDA-graph static buffer, and the collect-then-torch.cat consumers -
_decode (via _decode_chunks), _encode, _encode_pixels, encode_keyframe - hold
each chunk/clip result across the next _stitch_tiles replay, which overwrites
the pooled storage. First tiled decode() with >=2 temporal chunks (any real
video) and tiled encode() of >17 frames raised:

  RuntimeError: Error: accessing tensor output of CUDAGraphs that has been
  overwritten by a subsequent run. ... line ..., in _stitch_tiles:
  return torch.cat(result_rows, dim=-2)

Fix: .clone() the stitch output at the two eager call sites (_encode_clip
tiled return, _decode_clip tiled return) so no cudagraph-owned storage
escapes the tile driver; the clone is read before the next replay, so it is
race-free by construction. The streaming _decode_to_pixels path was already
safe (full copy-out per chunk before the next decode) and stays correct at
one extra D2D copy per chunk (~tens of microseconds vs the decode compute).
This follows the option (a) recommendation in the review; reduce-overhead is
retained on _stitch_tiles/_project_decoder_tile.

Tests:
- test_decode_clip_emits_tiled_stage_ranges updated: the tile driver now
  returns a caller-owned copy (is-not + equal), pinning the ownership
  contract at the mock level.
- New CUDA regression gate (the review's must-add test):
  test_tiled_decode_and_encode_survive_cudagraph_buffer_reuse_on_cuda - real
  tiled decode() (2 temporal chunks, 2x2 spatial tiles) and encode()/
  encode_pixels() (2 clips) with _stitch_tiles unmocked, asserting bitwise
  repeat-consistency plus value parity against the fully eager tile helpers
  (via _torchdynamo_orig_callable). Red on the unfixed merge (cudagraph
  overwrite RuntimeError on the first decode), green with this fix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 10:33:06 +00:00
Will Lin 64cdcf6be4 Merge PR #1731 head (9713ea127) on top of the stack (the target of this branch)
[feat] FastH3 few-step preview: VSA-H3 64-token tile option on the native
Triton block-sparse path, run-level --VSA-tile-size plumbing, opt-in sm_100a
CUDA forward route (FASTVIDEO_VSA_SM100A=1), basic_fasth3.py example with the
corrected 5-point-grid = 4-forward schedule (num_inference_steps=5), and the
FastVideo-Minimax-FastH3-Preview-v0.1 release name.

Clean textual merge; the predicted 1734<->1731 conflict in
minimax_h3_denoising.py resolved automatically and was verified by hand:
1731's vsa_tile_size plumbing sits in the metadata-builder preamble
(L149-152, L177) while 1734's edits are the import line, the stage docstring,
and the profiler_region+nvtx_range context line (L157) - the merged loop keeps
both, plus main's cudagraph_mark_step_begin contract (L184).
video_sparse_attn_h3.py and the metadata tests are 1731-only files on this
stack (no sibling edits).
2026-08-21 10:27:49 +00:00
Will Lin aa0d98a6b8 Merge PR #1734 head (dca423fd3) into integration/h3-perf-stack
[perf] MiniMax-H3 VAE decode optimization: _compile_conditions for the video +
audio VAE decoders (makes enable_torch_compile_vae effective for H3),
torch.compile on _stitch_tiles/_project_decoder_tile, VAE attention through the
FastVideo selector (TORCH_SDPA/FLASH_ATTN), opt-in FASTVIDEO_NVTX_PROFILE
ranges, per-worker attention-backend receipt.

Clean textual merge; semantic overlaps verified by hand:
- minimax_h3_conditioning.py: 1732's rewritten _encode_fl2va/_encode_ref2va
  kept their names, 1734's nvtx_range wrap composes over them (checked).
- minimax_h3.py: 1735 fusion routing + 1734 per-block nvtx_range coexist
  (fusions at attention/mlp/modulate seams, nvtx at the block loop).
- component_loader.py: 1732 quant plumbing at load_model (~L369) vs 1734
  backend receipt (~L1103) - disjoint.
- envs.py: FASTVIDEO_NVTX_PROFILE (1734) + FASTVIDEO_MINIMAX_H3_FUSIONS (1735)
  are distinct additions.
- minimax_h3_video.py: 1734 was authored on a base that already contains the
  merged #1703 (incl. pr1703-fixes content), so the 1734-vs-1703 conflict the
  reviews predicted was pre-resolved by the author's rebase.

KNOWN CARRIED DEFECT at this point in the stack: review finding F1 of
pr1734.md - reduce-overhead CUDA-graph outputs of _stitch_tiles escape into
collect-then-cat consumers (_decode/_encode/_encode_pixels), crashing tiled
eager decode/encode on CUDA. Fixed in a follow-up commit on this branch.
2026-08-21 10:26:23 +00:00
Will Lin 93b03bc14d Merge PR #1735 head (cbab605ef) into integration/h3-perf-stack
[perf] Opt-in MiniMax-H3 Sol-Engine Triton fusions (FASTVIDEO_MINIMAX_H3_FUSIONS):
fused rmsnorm+modulate, residual+gate+rmsnorm+modulate, per-head qknorm+partial
RoPE, and packed SwiGLU. Default-off; fused path is disclosed NON-PARITY
numerics (same-seed decoded SSIM 0.7403 vs eager per the PR body).

Head cbab605ef already carries the review's F1 blocker fix upstream
(ac98869aa: int64 row offsets in the fused qknorm+RoPE kernel - verified
present at qknorm_rope.py:43, tl.program_id(0).to(tl.int64)) plus the
engagement-test/logging hardening (cbab605ef), so no fix needs to be carried
by this stack for #1735.

Clean merge, no conflicts. Note: #1732 does not touch minimax_h3.py or envs.py
(its true merge-base is 0462e1b0e; earlier apparent overlap was #1712/#1290
noise from diffing against the wrong base), so 1735's DiT + envs.py edits had
no sibling edits to reconcile.
2026-08-21 10:25:09 +00:00
Will Lin 160f0c9ccf Merge PR #1732 head (ac56806af) into integration/h3-perf-stack
[perf] Optimize MiniMax-H3 text encoder memory: slim single-tensor forward
contract for the Qwen3-VL conditioner (early return at the layer-50 tap,
-262 MiB peak on top of merged #1711), TextEncoder base relaxed to
Generic[TextEncoderOutputT], and opt-in checkpoint-serialized block-FP8
(new text_encoder_quantization.py + minimax_h3_checkpoint_fp8.py).

Clean merge, no conflicts (PR base c4ad4227c == main's state for all touched
files; the tests/local_tests/minimax_h3/README.md hunk applied cleanly on top
of the #1703 dedupe).

Review status (pr1732.md): the two blockers are sm12x/FP8-only - the cutlass
m%4 crash is on the sm12x route (GB200/sm100 uses trtllm, verified clean at
m=559), and the FP8+cpu-offload gap only matters with FP8 enabled. This stack
keeps text-encoder FP8 OFF for all benches, so neither blocker is reachable.
2026-08-21 10:22:26 +00:00
Will Lin e5d1110a0f Merge PR #1703 head (pr1703-fixes @ 942f7db3d) into integration/h3-perf-stack
#1703 (H3 VAE peak-memory streaming) was already squash-merged into main as
e0a3db565 INCLUDING the pr1703-fixes review commits (aadb23f40 per-plane async
pinned copies, 942f7db3d legacy-decode oracle + slicing + pinned-buffer tests) -
the fix-branch tip is byte-identical to main for minimax_h3_video.py, the H3
stages, and the streaming tests. Content no-op recording the reviewed head.

Resolution: the auto-merge textually duplicated the 'Video VAE memory benchmark'
section in tests/local_tests/minimax_h3/README.md (main's squash placed the same
block at a slightly different anchor); deduplicated to a single copy - final
README is byte-identical to origin/main's.
2026-08-21 10:22:00 +00:00
Will Lin 622217ff2a Merge PR #1362 head (pr1362-fixes @ aa95a4c18) into integration/h3-perf-stack
#1362 (on-device uint8 post-decode) was already squash-merged into main as
fca45bc8e INCLUDING the pr1362-fixes review commits (15a164a05 size-regression
fix, f56f56704 comment accuracy, aa95a4c18 CPU regression tests) - the fix-branch
tip is byte-identical to main for video_generator.py and its tests. This merge is
therefore a content no-op recording the reviewed head in the stack history.
No conflicts.
2026-08-21 10:20:46 +00:00
Will LinandClaude Fable 5 cbab605eff [misc] MiniMax-H3 fusions: exact eager fallback, enable/inert logging, engagement test
- _can_run_minimax_h3_fusion now also requires Triton availability, so an
  enabled fusion on a CUDA build without a working Triton falls back to
  eager instead of hitting the strict wrappers' RuntimeError mid-forward.
- One-time logger.info of the resolved fusion set at model init (and a
  warning when the set is requested without Triton), plus a
  prepare_for_compile hook warning that torch.compile capture makes the
  fusions inert inside compiled block forwards.
- Positive-engagement routing test: counts 1/1/2/1 fused-kernel calls in
  one CUDA inference block forward and asserts a grad-enabled forward
  leaves the counters unchanged, so a guard regression to always-eager
  can no longer pass silently.
- Document the index-bounds contract ([0, table_rows)) on the modulation
  wrappers; a device-side check would synchronize, and in-model callers
  are safe by construction.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 10:06:23 +00:00
Will LinandClaude Fable 5 ac98869aa1 [bugfix] MiniMax-H3 fusions: int64 row offsets in the fused qknorm+RoPE kernel
_qknorm_partial_rope_kernel left tl.program_id(0) in int32, so
row * head_dim wrapped once the flattened input reached 2**31 elements
and the kernel read/wrote out of bounds (CUDA illegal memory access).
The PR's other two kernels (modulation.py, swiglu.py) already cast
tl.program_id(0).to(tl.int64); this one now matches, and seq_index /
table_offset inherit int64 from row.

Confirmed on GB200: (1, 8_500_000, 2, 128) bf16 (2.176e9 elements)
crashed before the cast and matches eager after it; the just-under-2**31
control shape matched all along. For H3 (56 heads x 128 head_dim) the
boundary is batch*seq >= 299_593 tokens per rank, reachable at SP=1.
Adds a GPU regression test at the over-2**31 shape that compares the
head and tail rows against eager.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 10:06:12 +00:00
William LinandClaude Fable 5 56d4a6074f [bugfix] fastvideo-kernel: fix Triton block-sparse backward logit scaling (bf16 K pre-scaling) (#1730)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 04:50:06 -05:00
Will LinandClaude Fable 5 9713ea1275 [misc] use the full release name FastVideo-Minimax-FastH3-Preview-v0.1
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 09:27:38 +00:00
Will LinandClaude Fable 5 37aa382cce [bugfix] example: the distilled FastH3 grid is 4 forwards = a 5-point sigma grid
MiniMaxH3Scheduler.set_timesteps(N) builds an N-point sigma grid ending at
0 and runs N-1 transformer forwards (the base model's '50 steps' preset is
49 forwards). The student was distilled on a 4-FORWARD grid
(t = 1000/750/500/250 -> 0 on the shift-12 schedule), so --steps 4 was
silently running a 3-forward, off-distribution grid. Default is now 5
grid points = the distilled 4-forward grid, with the convention documented
on the flag.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 09:27:38 +00:00
Will LinandClaude Fable 5 3f00983287 [feat] attention: opt-in sm_100a CUDA forward route for VSA-H3 tile-64
FASTVIDEO_VSA_SM100A=1 (default off) sends no-grad tile-64 forwards through
fastvideo_kernel.block_sparse_attn_sm100a (the Blackwell block-sparse
forward merged in #1719, which handles per-q-tile NON-uniform q2k_num rows
and zero-count rows). Route preconditions: module importable,
is_supported() (sm_100 device, bf16, head_dim 128, even tile count), no
grad tracking; any failure with the env set logs one warning and falls
back to the Triton-64 kernels, which also keep the entire grad path
unchanged.

The bool block map is compacted with the same map_to_index the Triton bool
entry uses, so the sm_100a kernel sees H3's true NON-uniform per-row counts
(prefix query tiles dense, video tiles prefix+top-k).

The FastH3 example surfaces the route as --vsa-kernel {triton,sm100a}
(default triton), which sets the env before pipeline boot so spawned GPU
workers inherit it; documented in the basic README.

CPU route-selection tests: default-off, engage-on-env, grad fallback,
warn-once fallback for missing module/unsupported geometry, no-grad-context
detection.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 09:27:38 +00:00
Will LinandClaude Fable 5 2dc57f4070 [misc] rebrand the few-step preview to FastH3 (model string, example name, docs)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 09:27:38 +00:00
Will LinandClaude Fable 5 9df19be719 [feat] example: few-step FastVideo-Minimax-H3-Preview inference (4-step DMD2 student)
basic_fast_minimax_h3.py runs FastVideo/FastVideo-Minimax-H3-Preview-v0.1,
the data-free-DMD2 distillation of MiniMax-H3: 4 denoising steps on the
release sampler's shift-12 schedule (vs the base model's 50), synchronized
video+audio in one pipeline call, guidance_scale 1.0.

The script always requests the VSA-H3 attention backend through the typed
boot-time route (pipeline.experimental -> FastVideoArgs.attention_backend):
the student checkpoint carries trained to_gate_compress gates, which only
exist under that backend. --vsa-sparsity defaults to 0.0 (every tile
selected — exactly dense attention); --vsa-tile-size defaults to 64, the
geometry the student was trained with, and is forwarded even at sparsity 0
because the gate-compress branch pools per tile. The HF repo is private
while the MiniMax H3 Community License review completes; --model-path
accepts a local snapshot meanwhile (noted in the script and README).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 09:27:38 +00:00
Will LinandClaude Fable 5 089eea3970 [feat] inference: expose the VSA-H3 tile size on the run-level route
FastVideoArgs gains VSA_tile_size (default 256, CLI --VSA-tile-size). It
rides the same boot-time route as run-level sparsity
(pipeline.experimental -> FastVideoArgs) and the H3 denoising stage
forwards it to MiniMaxH3VSAMetadataBuilder.build(tile_size=...), which
validates the value against VSA_H3_TILE_SHAPES. 256 keeps today's
behavior everywhere; 64 selects the native 64-token Triton block-sparse
path (FASTVIDEO_VSA_CUTEDSL does not apply there). Only the H3 stage
consumes it; Wan/LTX-2 VSA paths are untouched.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 09:27:38 +00:00
Will LinandClaude Fable 5 dd8447ecc5 [feat] VSA-H3: 64-token tile option on the native Triton block-sparse path
The tile size becomes selectable at metadata build time:
MiniMaxH3VSAMetadataBuilder.build(tile_size=...) accepts 256 (default,
unchanged (4,8,8) tiles and VSA-256 CuTe/Triton routing) or 64. At 64 the
tiles are (4,4,4) and the block map is already at the Triton kernels'
native granularity, so forward and backward run
fastvideo_kernel.block_sparse_attn directly (BHSD, transposed around the
call like the 256 wrapper's Triton branch) with no 256->64 mask expansion;
FASTVIDEO_VSA_CUTEDSL does not apply at 64. Tile geometry, pooled scoring,
prefix chunking, the probe, and the gate_compress views all follow the
configured tile element count.

Also adds _validate_h3_tile_geometry: a synchronous, lru-cache-scoped
bounds check on every built geometry (per-tile sizes in (0, tile_elems],
sizes sum to the packed length, untile index injective into non-pad
slots). Malformed geometry now raises at build time with the numbers in
hand instead of surfacing as an unattributable async device fault at some
later kernel or collective.

CPU tests: hand-computed (4,4,4) oracle on a grid ragged in all three
dims, the production packed shape (768x1344, 124 frames) under both tile
sizes, sparsity-0 SDPA equivalence at tile 64, the guard's 64-bound, and
builder rejection of unknown tile sizes. GPU parity of the 64 route was
validated out of band on GB200: sparsity-0 forward+input-grad parity vs
dense SDPA 3.1e-3 rel-L2 at the production packed shape (matching the 256
route to the third digit); sparsity-0.9 out/dq bitwise same-seed
deterministic, dk/dv ~3e-5 reduction-order drift (same profile as the
existing 256 Triton route).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 09:27:38 +00:00
Davids048 dca423fd31 Enable H3 VAE attention backend selection. 2026-08-21 08:33:09 +00:00
Davids048 1b43af8e8e Custom compile H3 vae decoding. 2026-08-21 08:33:09 +00:00
Davids048 628591b620 Add some logging. 2026-08-21 08:33:09 +00:00
Davids048 0980ca563f Add nvtx profiling support. 2026-08-21 08:33:09 +00:00
H1yori233 b158388733 [perf] Add opt-in MiniMax-H3 Sol-Engine fusions 2026-08-21 00:20:38 -07:00
Shao Duan c4ad4227c0 [misc] MiniMax-H3: move the AdaLN converter into scripts/checkpoint_conversion (#1712) 2026-08-21 02:15:00 -05:00
KyleNeverGivesUpandSolitaryThinker a63ccce73d [docs]: add a maintained inference cookbook (#1290)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-08-21 01:53:44 -05:00
H1yori233andWill Lin ac56806aff [perf] Optimize MiniMax-H3 text encoder
Co-authored-by: Will Lin <160547796+KyleNeverGivesUp@users.noreply.github.com>
2026-08-20 21:42:53 -07:00
KyleNeverGivesUp 0462e1b0e7 [perf]: MiniMax H3 - build the Qwen3-VL encoder only as far as it is read (-13.7 GB) (#1711) 2026-08-20 23:19:25 -05:00
lpc0220 907f2100ec [kernel] sm_100a CUDA block-sparse VSA forward (Blackwell), 64- and 128-token blocks (#1719) 2026-08-20 23:16:25 -05:00
Kaiqin Kong e0a3db5651 [perf] Reduce MiniMax-H3 VAE peak memory (#1703) 2026-08-20 23:13:23 -05:00
Raghav K fca45bc8e1 [perf] Quantize frames to uint8 on-device before the post-decode D->H copy (#1362) 2026-08-20 21:55:55 -05:00
Will LinandClaude Fable 5 aa95a4c18e [misc]: add CPU regression tests for the on-device uint8 frame path
Two tests certifying #1362's frame semantics without a GPU:

- frames_match_legacy_cpu_loop: the quantize-then-grid path reproduces
  the legacy make_grid -> *255 -> uint8 per-frame loop bit-exactly for
  in-range fp32 pixels (batch>1 nrow=6 grid layout, odd frame count,
  uint8 HWC contract). Verified to pass against main's legacy loop too,
  so it pins both sides of the equivalence.
- frames_clamp_out_of_range_pixels: out-of-[0,1] VAE output saturates
  at 0/255 instead of wrapping mod 256; fails on the pre-#1362 loop.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:39:09 +00:00
Will LinandClaude Fable 5 942f7db3db [test] H3 VAE streaming: legacy-decode oracle, batch slicing, pinned-buffer coverage
Pin the chunk-iterator refactor to the pre-streaming _decode implementation
bit-for-bit across the seam and pad-trim geometries (one padded chunk, a pad
hitting the intra-clip tail, two blended chunks, three chunks plus trim), and
cover the use_slicing batch paths for encode_pixels/decode_to_pixels and the
CUDA pinned-buffer async copy path (skipped without a GPU).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:38:44 +00:00
Will LinandClaude Fable 5 aadb23f409 [perf] H3 VAE streamed decode: direct per-plane copies, async with pinned buffers
The temporal slice of the CPU output buffer is strided across channels, so
each finalized-chunk copy_ staged through a pageable CPU temporary plus a
CPU-side scatter, which is where the streamed path's decode-time regression
came from and why the pinned buffer bought nothing. Copy per (batch, channel)
plane instead - contiguous on both sides, memcpy-eligible - and make the
copies non_blocking when the destination is pinned, draining the stream once
in decode_to_pixels before the buffer can be read or released.

Also: raise on a non-positive decode plan before allocating the output
buffer, annotate _decode_chunks as an Iterator, and document the
encode_pixels CPU dtype/range contract.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:38:35 +00:00
Will LinandClaude Fable 5 f56f567042 [misc]: correct the post-decode rationale comments
The pre-#1362 `samples.copy_(output_batch.output)` never passed
`non_blocking=True` (git log -S confirms), so drop the deferred
non-blocking-transfer claim and describe the measured costs instead:
a full fp32 D->H copy plus a single-threaded per-frame CPU loop. Also
drop the reintroduced hardcoded "~50 MB" size estimate (same class of
comment Copilot flagged and commit 0399713e7 removed elsewhere) and
scope the "typical flow" claim to the CLI, since the SamplingParam
API default is return_frames=True.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:34:43 +00:00
Will LinandClaude Fable 5 15a164a052 [bugfix]: derive result size from the decoded output when the samples mirror is skipped
PR #1362 commit 3 leaves `samples` as an empty placeholder when
`return_frames=False`, but main picked up #1595's refiner size
reporting in the meantime, and `_resolve_output_size(samples, ...)`
silently fell back to the requested geometry in exactly the common
save flow the PR optimizes. Read the geometry from
`output_batch.output` instead (shape-only access, no D->H copy),
gated on `needs_frame_output` so metadata-only and audio-only calls
still never inspect the (possibly dropped) worker output.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 02:34:16 +00:00
Will Lin aaaa7a14a3 Merge remote-tracking branch 'origin/main' into pr1362-fixes 2026-08-21 02:21:44 +00:00
Aryan Kumar 86d639c848 [docs] Announce FastMetal-QAD (#1721) 2026-08-19 13:45:49 -07:00
00338aa9ca [perf] Add FA4 CuTe backward support for VSA-256 (#1639)
Co-authored-by: Hyunsung Lee <hyunsungl@sizigistudios.com>
Co-authored-by: alexzms <3036648523@qq.com>
2026-08-19 11:46:49 -07:00
H1yori233 74b409d7cf [test]: add H3 VAE parity and memory benchmark 2026-08-19 02:45:23 -07:00
Aryan KumarandAryan Kumar 8537dcd6de [feat]: Apple Silicon MLX runtime — INT8 Wan2.1 and Wan2.2 inference (#1638)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-18 15:08:00 -07:00
H1yori233 528cef02c4 optimize VAE memory 2026-08-12 02:48:43 -07:00
William Lin 8208536cd1 [bugfix] profiler region system: record + export actually work, usability roll-up (#1691) 2026-08-09 20:35:53 -07:00
Kai 0653f8f3af [new-model] Add V2A: native MMAudio inference pipeline (#1622) 2026-08-09 17:29:27 -07:00
William Lin e0d702decb [feat] VSA for MiniMax H3: packed mixed-modality sparse attention (#1695) 2026-08-09 13:10:51 -07:00
Shao Duan 541ef014ee [perf] MiniMax-H3: rank-reduced AdaLN pruned model option (-39% params, -23 GiB VRAM) (#1699) 2026-08-09 12:31:57 -07:00
William Lin ffc1a7a58b [refactor] H3 pipeline cleanup: shared helpers, dead machinery, loop-invariant hoists (#1698) 2026-08-09 04:51:31 -07:00
William Lin 9028953625 [misc] yapf pass under CI's interpreter (3.12) + pin hook language_version (#1702) 2026-08-08 22:24:17 -07:00
KyleNeverGivesUpandClaude Opus 5 6eb95693a1 [misc]: re-run yapf on main so pre-commit passes again (#1700)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:17:52 -07:00
Junda Su c3567eb468 [feat] add Minimax H3 sft pipeline (#1688) 2026-08-07 16:12:04 -07:00
William Lin 15568f27db [perf]: H3 torch.compile + CUDA graphs (1.2-1.3x) with denoising step marking (#1689) 2026-08-06 16:23:33 -07:00
Kaiqin Kong 126a75ad63 [misc] Support partial Hugging Face model downloads (#1684) 2026-08-06 16:09:46 -07:00
Raghav KandSolitaryThinker a2bfc7cdb2 [docs] DGX Spark (GB10) performance & tuning guide + reproduction examples (#1631)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-08-06 14:36:38 -07:00
William Lin b963a24612 [bugfix]: hard-fail when ATTN_QAT_INFER is selected but the kernel is unusable (#1690) 2026-08-06 13:12:27 -07:00
William Lin fb7be2fe2c [ci]: add golden-gate lane — single-layer bitwise DiT fingerprints for all SSIM-covered families (#1682) 2026-08-05 12:47:03 -07:00
William Lin ab00392664 [bugfix]: add GB200 to the inline SSIM device tables #1676 missed (#1681) 2026-08-05 10:33:46 -07:00
Shao Duan 9f1e7c19d2 [bugfix] Wan I2V: CLIP image conditioning silently dropped when passed as a tensor during training (#1673) 2026-08-05 01:35:09 -07:00
Kaiqin Kong e8b0e4c61e [feat] Add MiniMax H3 (#1674) 2026-08-04 13:54:39 -07:00
Haochen Jiang 9145ffdc46 [bugfix]: keep _resolved_attention_backend out of the positional config signature (#1678) 2026-08-03 17:09:15 -07:00
William Lin c3d07c870b [bugfix]: give GB200 its own SSIM reference folder instead of B200's (#1676) 2026-08-03 15:14:25 -07:00
William Lin e8812bef0b [docs]: batched docs cleanup (landing page, links, requirements, nav) (#1644) 2026-08-02 18:03:26 -07:00
William Lin b9be2449dc [refactor]: delete the dead global attention-backend override (#1672) 2026-08-02 18:00:42 -07:00
Adhvay IyerandSolitaryThinker bc7a804618 [bugfix]: harden FastVideo Studio UI reliability and accessibility (#1659)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-08-02 15:26:18 -07:00
William Lin 7b094c945b [refactor]: resolve attention backend once per component at load time (#1657) 2026-08-02 15:07:42 -07:00
William Lin eeb3e8a597 [bugfix]: left-align Gemma connector tokens per batch row (#1664) 2026-08-02 15:01:34 -07:00
Suhaan Khurana 05406c5d1b [misc]: consolidate dataset download scripts under examples/datasets/ (#1667) 2026-08-02 14:21:24 -07:00
KyleNeverGivesUpandClaude Opus 5 99d04a7f98 [bugfix] keep loader-populated text encoder configs in the validation pipeline (#1669)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 14:10:16 -07:00
William Lin 1b2b2a0161 [bugfix]: copy text encoder outputs out of CUDAGraph static buffers (#1650) 2026-07-27 21:52:53 -07:00
William Lin 98d65835b5 [misc]: refresh stale sm_120-only validation notes in QAT recipes (#1655) 2026-07-27 19:35:29 -07:00
William Lin 422585d08f [misc]: add LTX-2.3 fine-tuning example recipes (#1651) 2026-07-27 18:42:44 -07:00
ryanM154 e59a1ce16a [bugfix] Report actual package version in fastvideo --version (#1652) 2026-07-27 16:32:01 -07:00
William Lin d71acc0eb5 [bugfix]: fix stale imports in LTX-2.3 gradio local demo (#1640) 2026-07-27 16:31:07 -07:00
William Lin af2934dd6b [feat]: FA4-FP4 ATTN_QAT_INFER on sm_100/sm_103 + NVFP4 weight purge (#1647) 2026-07-27 15:46:52 -07:00
William Lin 1801512818 [docs]: cover all registered models in the support matrix (#1641) 2026-07-27 11:13:53 -07:00
William Lin 5ae05b032e [misc]: add LTX-2 fine-tuning example recipes (#1645) 2026-07-27 11:11:49 -07:00
Mac Lee 7a592ff09a [ci]: cache FastVideo kernel builds in Modal (#1562) 2026-07-26 04:21:09 -07:00
Lev NovitskiyandClaude Sonnet 5 8b23984c79 [feat] Add Kandinsky5 QAD training pipeline: data preprocessing, QAT finetune, QAT-aware DMD distillation (#1601)
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 02:59:21 -07:00
William Lin bf18371afe [feat] Enable LTX-2 NVFP4 linear and attention QAT fine-tuning (#1626) 2026-07-26 02:19:34 -07:00
William Lin 69349dd2aa [docs]: unify community links on the README Slack invite (#1643) 2026-07-25 17:26:58 -07:00
Yogya MehrotraandClaude Sonnet 5 8d89f30d3f [bugfix] Skip CUDA-only fastvideo-kernel/flashinfer-python deps on non-Linux (#1574)
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-23 14:23:45 -07:00
William Lin 10546353da [ci]: extend Full Suite training lane timeouts (#1616) 2026-07-23 13:00:00 -07:00
pkisfaludi-nvandClaude Opus 4.8 9fb74b9732 Make LTX-2 RMSNorm out-of-place so torch_tensorrt + Ulysses SP compiles (#1623)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 04:32:04 -07:00
Adriel FungandSolitaryThinker 521dee0e82 [perf]: enable per-block torch.compile for LTX2 with persistent cache (#1602)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-20 15:59:13 -07:00
William Lin 229419208e [ci]: opt in to fork-PR head checkout after actions/checkout guard change (#1625) 2026-07-20 13:12:22 -07:00
William Lin 65f3b946b9 [feat] Add LTX-2 and LTX-2.3 fine-tuning to the modular trainer (#1624) 2026-07-20 11:58:39 -07:00
Zhang Peiyuan 191fcbf46c [feat] add qat docs (#1621) 2026-07-19 18:37:47 -07:00
Junda Su 755a4e4470 [new-model] Add LingBot-Video Dense and MoE/refiner T2V inference (#1595) 2026-07-18 20:18:16 -07:00
9709b7513b [feat] Port NVFP4 QAT/QAD to modular train framework (#1619)
Co-authored-by: Peiyuan Zhang <email>

Co-authored-by: Peiyuan <a>
2026-07-18 15:01:09 -07:00
William Lin 32cd603515 Revert docs trusted-branch-only workflow (#1618) 2026-07-17 00:22:56 -07:00
Junda Su d4bdd3621a [new-model] Port LingBot-World-v2 (#1579) 2026-07-16 19:51:53 -07:00
William Lin e2f8322842 [ci]: skip unused Buildkite submodule checkout (#1614) 2026-07-16 18:37:57 -07:00
6966f9e0bc [fix] Z-Image (#1236) draft port: rebase + strict-load contract + bf16 encoder parity + PORT_STATUS (#1339)
Co-authored-by: Mrinaal Dogra <mdogra@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-16 18:18:25 -07:00
Mac Lee 743f4ed5f9 [ci]: harden Modal repository checkout (#1590) 2026-07-16 18:05:50 -07:00
RaghavandClaude Opus 4.7 3c3da4d057 [perf]: skip the fp32 samples D->H copy when return_frames=False
The previous commit rebuilt the post-decode frames path to read
`output_batch.output` directly via the GPU `vid_u8` cast, so `samples`
is now consumed in exactly one place — the result dict's `samples`
field, gated on `batch.return_frames`. When the caller doesn't ask
for `samples`, the pinned ~50 MB fp32 alloc and its D->H copy (and
the latent fallthrough `.cpu()`) are dead weight; the typical
generate flow (save_video=True, return_frames=False) hits this on
every call.

Extend `skip_pixel_prealloc` to include `not return_frames` so the
pinned buffer is allocated only when needed, and short-circuit the
copy/`.cpu()` on the same condition. No effect when
`return_frames=True` or for latent callers that read `samples`;
correctness is unchanged (SSIM-gated, same as the parent commit).

Removes the residual ~2 s the PR text already calls out for the
"only saving to disk" case.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-07-16 03:02:49 -07:00
Raghav 0399713e7b [perf]: clarify samples/output_batch.output equivalence; drop hardcoded transfer size
Address review comments on #1362:
- Document that `samples` is just the pinned-CPU mirror of
  `output_batch.output` with no intervening preprocessing, so sourcing
  from `output_batch.output` is the same data (Copilot).
- Replace the hardcoded ~0.3-1.4 GB estimate with a description of the
  scaling relationship; it varies with resolution/frames/batch/dtype
  and would rot (Copilot).
- Note the SSIM-gated (not bit-exact) equivalence inline.

No behavior change.
2026-07-16 02:59:59 -07:00
Raghav 0af2e9e8ef [perf]: quantize frames to uint8 on-device before the D->H copy
PostDecodeFrameProcessStage was charged ~6s/run (~25% of e2e on Cosmos
2.5, scaling with frames/resolution). Profiling (nsys + microbench)
showed the stage's own compute is only ~0.5-1.9s; the rest is the
non-blocking pinned-CPU samples.copy_(output) D->H of the full fp32
video (~0.3-1.4 GB) completing lazily and blocking the first
postprocess op, plus a single-threaded per-frame CPU *255/cast loop.

Cast to uint8 on the source device (typically CUDA) before the copy:
the transfer becomes 4x smaller (fp32 -> uint8) and the elementwise
work runs on the GPU. Microbench: 2.056s -> 0.034s (T=29), 2.304s ->
0.132s (T=125), 17-60x on the measurable cost.

clamp_(0, 255) additionally fixes a latent overflow: VAE output
slightly outside [0, 1] previously wrapped mod 256 in the unclamped
(x * 255).to(uint8) cast.

Output is not bit-identical to the old CPU cast (float->uint8 differs
<=1 LSB between CPU and GPU on boundary pixels); gate via SSIM rather
than exact equality. Scoped to the pixel-video path only; latent,
audio-only, and return-samples paths are unchanged.
2026-07-16 02:59:59 -07:00
Mac Lee 1c04ace573 [ci]: pin VSA training regression to H100 (#1591) 2026-07-15 22:04:57 -07:00
Mac LeeandSatyam Srivastava 6cbff73687 [bugfix]: skip unused output materialization (#1567)
Co-authored-by: Satyam Srivastava <srivastavasatyam53@gmail.com>
2026-07-16 04:03:07 +00:00
William Lin dec8b10939 [docs]: document automatic Docker image builds (#1608) 2026-07-15 18:52:21 -07:00
William Lin 133a5278af [ci] Run docs only for trusted PR branches (#1610) 2026-07-15 18:51:59 -07:00
William Lin da856274cc [feat]: FastVideo Studio — SvelteKit → Next.js port + review fixes (#1612) 2026-07-15 18:51:33 -07:00
Mac LeeandSolitaryThinker a253856147 [ci] Add exact identity performance statuses (#1560)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-15 14:59:09 -07:00
Satyam Srivastavaandgemini-code-assist[bot] 6e25d94ebc [ci]: enable scheduled perf runs to update rolling baseline (#1599)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-07-14 16:55:12 -07:00
Mac Lee cae8fa18dc [bugfix]: propagate Qwen2.5-VL visual dtype (#1580) 2026-07-13 18:37:11 -07:00
William Lin 821e5a0832 [bugfix]: fix FlashAttention resolver tests after tuple return (#1597) 2026-07-13 16:29:03 -07:00
William Lin c1abc42782 [bugfix]: allow unrestricted head sizes in SDPA (#1596) 2026-07-13 16:04:51 -07:00
William Lin ef15ea2391 [bugfix]: keep LTX2 rms_norm outputs bf16 under torch 2.12 autocast (#1587) 2026-07-13 16:04:21 -07:00
1166 changed files with 107262 additions and 31812 deletions
@@ -1,77 +0,0 @@
# MagiHuman — Codebase Map Entry
**Family:** daVinci-MagiHuman (joint audio-visual generative model)
**Reference:** [GAIR-NLP/daVinci-MagiHuman](https://github.com/GAIR-NLP/daVinci-MagiHuman)
**Architecture:** 15B-param single-stream DiT, 40 layers, hidden=5120,
head_dim=128, GQA num_query_groups=8. Joint AV denoising in a unified token
sequence; **no cross-attention**.
## Variant Matrix
| Variant | T2V | TI2V | DiT | Steps | CFG | Resolution |
|---|---|---|---|---|---|---|
| `base` | yes | yes | base | 32 | 2 | 480x256 |
| `distill` | yes | yes | distill (DMD-2) | 8 | 1 (no CFG) | 480x256 |
| `sr_540p` | yes | yes | base + sr_540p | 32 + 5 | 2 + cfg-trick | 896x512 |
| `sr_1080p` | yes | yes | base + sr_1080p | 32 + 5 | 2 + cfg-trick | 1920x1056 |
SR-1080p uses block-sparse video→video local-window attention on 32 of 40
SR DiT layers (`frame_receptive_field=11`), implemented as a 3-block SDPA
accumulator that mirrors upstream `flex_flash_attn_func`.
## File Locations
| Role | Path |
|---|---|
| Pipeline class | `fastvideo/pipelines/basic/magi_human/magi_human_pipeline.py` |
| Pipeline package AGENTS.md | `fastvideo/pipelines/basic/magi_human/AGENTS.md` |
| Stages | `fastvideo/pipelines/basic/magi_human/stages/*.py` |
| DiT | `fastvideo/models/dits/magi_human.py` |
| Text encoder (T5-Gemma) | `fastvideo/models/encoders/t5gemma.py` |
| Audio VAE wrapper | `fastvideo/models/vaes/sa_audio.py` (shared with `stable_audio` pipeline) |
| Conversion script | `scripts/checkpoint_conversion/convert_magi_human_to_diffusers.py` |
| Examples | `examples/inference/basic/basic_magi_human*.py` (8 files, one per variant × mode) |
| SSIM regression | `fastvideo/tests/ssim/test_magi_human_similarity.py` |
| Local parity battery | `tests/local_tests/magi_human/` (14 tests, GPU-gated) |
| Port journal | `fastvideo/pipelines/basic/magi_human/JOURNAL.md` |
## Canonical HF Repo
[FastVideo/MagiHuman-Diffusers](https://huggingface.co/FastVideo/MagiHuman-Diffusers)
— umbrella repo with sibling subfolders per variant.
```python
from fastvideo import VideoGenerator
gen = VideoGenerator.from_pretrained("FastVideo/MagiHuman-Diffusers/base")
gen.generate_video(prompt="...", output_path="out.mp4", save_video=True)
```
Four shared upstream components are lazy-loaded by `MagiHumanPipeline.load_modules`:
| Component | Upstream repo |
|---|---|
| Wan 2.2 VAE | `Wan-AI/Wan2.2-TI2V-5B` |
| T5-Gemma 9B UL2 | `google/t5gemma-9b-9b-ul2` |
| Stable Audio VAE | `stabilityai/stable-audio-open-1.0` |
| MagiHuman DiT weights | `GAIR/daVinci-MagiHuman` (gated) |
## Parity Invariants
Three load-bearing invariants. See `fastvideo/pipelines/basic/magi_human/AGENTS.md`
for the full discussion.
1. **Channel-major video token packing** (`stages/latent_preparation.py`)
2. **DiT dtype boundary**: residual stream stays fp32 across blocks
3. **Conversion `_FP32_KEEP_SUFFIXES`** (`scripts/checkpoint_conversion/convert_magi_human_to_diffusers.py`)
## Lessons
- `.agents/lessons/2026-05-07_silent-channel-major-packing-bugs.md`
- `.agents/lessons/2026-05-07_dit-dtype-boundary-with-flash-attn.md`
- `.agents/lessons/2026-05-07_conversion-cast-bf16-suffix-allowlist.md`
## Provenance
Decomposed from PR [#1280](https://github.com/hao-ai-lab/FastVideo/pull/1280)
(`will/magi` @ `4e1603634d27c8e1b5c4cc5d9387f046547f5c49`). See the package
AGENTS.md for the full PR-stack table.
@@ -10,9 +10,7 @@ from pathlib import Path
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Clone a reference repo for FastVideo parity tests."
)
parser = argparse.ArgumentParser(description="Clone a reference repo for FastVideo parity tests.")
parser.add_argument("repo_url", help="Official reference repository URL")
parser.add_argument("target_dir", help="Directory to clone into")
parser.add_argument("--branch", help="Branch or tag to clone")
@@ -62,9 +60,7 @@ def gitignore_entry_for(target: Path) -> str:
try:
relative = resolved.relative_to(root)
except ValueError as exc:
raise ValueError(
"--update-gitignore requires target_dir to be under the current directory"
) from exc
raise ValueError("--update-gitignore requires target_dir to be under the current directory") from exc
text = relative.as_posix().rstrip("/")
return "/" + text + "/"
@@ -8,14 +8,12 @@ import os
import sys
from pathlib import Path
HF_TOKEN_ENV_KEYS = ("HF_TOKEN", "HUGGINGFACE_HUB_TOKEN", "HF_API_KEY")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Download a HF model snapshot or selected files into a local directory."
)
description="Download a HF model snapshot or selected files into a local directory.")
parser.add_argument("repo_id", help="HF repo id, for example Org/Model")
parser.add_argument("local_dir", help="Destination directory")
parser.add_argument("--repo-type", default="model", help="HF repo type (default: model)")
@@ -10,7 +10,6 @@ import sys
from pathlib import Path
from typing import Any
HF_TOKEN_ENV_KEYS = ("HF_TOKEN", "HUGGINGFACE_HUB_TOKEN", "HF_API_KEY")
RAW_WEIGHT_SUFFIXES = (".safetensors", ".pt", ".pth", ".ckpt", ".bin")
KNOWN_COMPONENTS = {
@@ -34,8 +33,7 @@ KNOWN_COMPONENTS = {
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Classify a HF repo or local directory as Diffusers, raw, custom, or unknown."
)
description="Classify a HF repo or local directory as Diffusers, raw, custom, or unknown.")
parser.add_argument("source", help="HF repo id or local weights directory")
parser.add_argument("--repo-type", default="model", help="HF repo type (default: model)")
parser.add_argument("--revision", help="HF revision to inspect")
@@ -94,14 +92,12 @@ def load_remote_files(
) -> list[str]:
from huggingface_hub import list_repo_files
return sorted(
list_repo_files(
repo_id,
repo_type=repo_type,
revision=revision,
token=token,
)
)
return sorted(list_repo_files(
repo_id,
repo_type=repo_type,
revision=revision,
token=token,
))
def load_remote_model_index(
@@ -215,24 +211,24 @@ def build_result(args: argparse.Namespace) -> dict[str, Any]:
"components_seen": components,
"file_count": len(files),
"file_scan_truncated": truncated,
"files_sample": files[: args.sample_limit],
"files_sample": files[:args.sample_limit],
}
def print_human(result: dict[str, Any]) -> None:
for key in (
"source",
"source_kind",
"repo_type",
"revision",
"token_env",
"source_layout",
"needs_conversion",
"model_index_class",
"model_index_diffusers_version",
"model_index_error",
"file_count",
"file_scan_truncated",
"source",
"source_kind",
"repo_type",
"revision",
"token_env",
"source_layout",
"needs_conversion",
"model_index_class",
"model_index_diffusers_version",
"model_index_error",
"file_count",
"file_scan_truncated",
):
value = result.get(key)
if value is not None:
@@ -18,7 +18,6 @@ import pytest
import torch
from torch.testing import assert_close
os.environ.setdefault("MASTER_ADDR", "localhost")
os.environ.setdefault("MASTER_PORT", "29519")
os.environ.setdefault("DISABLE_SP", "1")
@@ -35,15 +34,10 @@ FASTVIDEO_CONFIG_CLASS = "<FastVideoConfig>" # TODO.
FASTVIDEO_MODEL_MODULE = "fastvideo.models.<bucket>.<module>" # TODO.
FASTVIDEO_MODEL_CLASS = "<FastVideoModel>" # TODO.
OFFICIAL_REF_DIR = Path(
os.getenv("<FAMILY_UPPER>_OFFICIAL_REF_DIR", REPO_ROOT / "<ReferenceDir>")
)
LOCAL_WEIGHTS_DIR = Path(
os.getenv("<FAMILY_UPPER>_LOCAL_WEIGHTS_DIR", REPO_ROOT / "official_weights" / FAMILY)
)
CONVERTED_WEIGHTS_DIR = Path(
os.getenv("<FAMILY_UPPER>_CONVERTED_WEIGHTS_DIR", REPO_ROOT / "converted_weights" / FAMILY)
)
OFFICIAL_REF_DIR = Path(os.getenv("<FAMILY_UPPER>_OFFICIAL_REF_DIR", REPO_ROOT / "<ReferenceDir>"))
LOCAL_WEIGHTS_DIR = Path(os.getenv("<FAMILY_UPPER>_LOCAL_WEIGHTS_DIR", REPO_ROOT / "official_weights" / FAMILY))
CONVERTED_WEIGHTS_DIR = Path(os.getenv("<FAMILY_UPPER>_CONVERTED_WEIGHTS_DIR",
REPO_ROOT / "converted_weights" / FAMILY))
def _resolve_hf_token() -> str | None:
@@ -99,18 +93,14 @@ def _load_official_model(device: torch.device, dtype: torch.dtype) -> torch.nn.M
model = OfficialClass() # TODO: pass official config kwargs.
state_dict = {} # TODO: load official state dict from LOCAL_WEIGHTS_DIR.
missing, unexpected = model.load_state_dict(state_dict, strict=True)
assert not missing and not unexpected, (
f"official load mismatch missing={missing[:5]} unexpected={unexpected[:5]}"
)
assert not missing and not unexpected, (f"official load mismatch missing={missing[:5]} unexpected={unexpected[:5]}")
return model.to(device=device, dtype=dtype).eval()
def _load_fastvideo_model(device: torch.device, dtype: torch.dtype) -> torch.nn.Module:
"""Load the FastVideo component with the same tensor content."""
if not CONVERTED_WEIGHTS_DIR.exists() and not LOCAL_WEIGHTS_DIR.exists():
pytest.skip(
f"No FastVideo loadable weights: {CONVERTED_WEIGHTS_DIR} or {LOCAL_WEIGHTS_DIR}"
)
pytest.skip(f"No FastVideo loadable weights: {CONVERTED_WEIGHTS_DIR} or {LOCAL_WEIGHTS_DIR}")
# TODO: replace with the bucket-specific FastVideo config/class/loader.
# DiT examples:
@@ -127,8 +117,7 @@ def _load_fastvideo_model(device: torch.device, dtype: torch.dtype) -> torch.nn.
state_dict = {} # TODO: load converted or directly mapped state dict.
missing, unexpected = model.load_state_dict(state_dict, strict=True)
assert not missing and not unexpected, (
f"FastVideo load mismatch missing={missing[:5]} unexpected={unexpected[:5]}"
)
f"FastVideo load mismatch missing={missing[:5]} unexpected={unexpected[:5]}")
return model.to(device=device, dtype=dtype).eval()
@@ -187,11 +176,9 @@ def test_component_parity():
assert official_out.shape == fastvideo_out.shape
diff = (official_out - fastvideo_out).abs()
print(
f"official abs_mean={official_out.abs().mean().item():.6f} "
f"fastvideo abs_mean={fastvideo_out.abs().mean().item():.6f} "
f"diff_max={diff.max().item():.6f} diff_mean={diff.mean().item():.6f}"
)
print(f"official abs_mean={official_out.abs().mean().item():.6f} "
f"fastvideo abs_mean={fastvideo_out.abs().mean().item():.6f} "
f"diff_max={diff.max().item():.6f} diff_mean={diff.mean().item():.6f}")
# TODO: pick tolerance by scope:
# - single block / same kernel: 1e-4
@@ -27,7 +27,6 @@ try:
except ImportError: # pragma: no cover - optional local conversion dependency
snapshot_download = None
# TODO: fill with authoritative component prefixes for monolithic checkpoints.
# Example: {"model.model.": "transformer", "pretransform.model.": "vae"}
COMPONENT_PREFIXES: dict[str, str] = {}
@@ -47,10 +46,7 @@ SKIP_PATTERNS: tuple[str, ...] = ()
def _hf_token() -> str | None:
return (
os.environ.get("HF_TOKEN") or os.environ.get("HUGGINGFACE_HUB_TOKEN")
or os.environ.get("HF_API_KEY")
)
return (os.environ.get("HF_TOKEN") or os.environ.get("HUGGINGFACE_HUB_TOKEN") or os.environ.get("HF_API_KEY"))
def resolve_src(src: str, revision: str | None) -> Path:
@@ -95,11 +91,10 @@ def apply_mapping(key: str) -> str | None:
return key
def split_monolithic(
state: dict[str, torch.Tensor],
) -> dict[str, OrderedDict[str, torch.Tensor]]:
def split_monolithic(state: dict[str, torch.Tensor], ) -> dict[str, OrderedDict[str, torch.Tensor]]:
components: dict[str, OrderedDict[str, torch.Tensor]] = {
name: OrderedDict() for name in set(COMPONENT_PREFIXES.values())
name: OrderedDict()
for name in set(COMPONENT_PREFIXES.values())
}
intentionally_skipped: list[str] = []
unowned: list[str] = []
@@ -117,10 +112,8 @@ def split_monolithic(
unowned.append(key)
if unowned:
sample = ", ".join(unowned[:10])
raise ValueError(
f"Unowned monolithic keys: {len(unowned)}. "
f"Add COMPONENT_PREFIXES or SKIP_PATTERNS entries. Sample: {sample}"
)
raise ValueError(f"Unowned monolithic keys: {len(unowned)}. "
f"Add COMPONENT_PREFIXES or SKIP_PATTERNS entries. Sample: {sample}")
if intentionally_skipped:
print(f"Intentionally skipped {len(intentionally_skipped)} keys")
return {name: weights for name, weights in components.items() if weights}
@@ -143,8 +136,12 @@ def build_component_configs(_src_dir: Path) -> dict[str, dict[str, Any]]:
# TODO: emit config content accepted by FastVideo loaders. Most components use
# config.json; schedulers use scheduler_config.json.
return {
"transformer": {"_class_name": "<FastVideoTransformerClass>"},
"vae": {"_class_name": "<FastVideoVAEClass>"},
"transformer": {
"_class_name": "<FastVideoTransformerClass>"
},
"vae": {
"_class_name": "<FastVideoVAEClass>"
},
}
@@ -177,19 +174,13 @@ def build_model_index(
}
if revision:
index["_fastvideo_converted_revision"] = revision
return {
key: value
for key, value in index.items()
if key.startswith("_") or key in available_components
}
return {key: value for key, value in index.items() if key.startswith("_") or key in available_components}
def validate_component_configs(configs: dict[str, dict[str, Any]]) -> None:
# TODO: instantiate each FastVideo config and call update_model_arch(...) or
# update_model_config(...) with this JSON so unknown emitted keys fail here.
placeholder_configs = [
name for name, config in configs.items() if "<" in json.dumps(config)
]
placeholder_configs = [name for name, config in configs.items() if "<" in json.dumps(config)]
if placeholder_configs:
raise ValueError(f"Replace config placeholders for: {placeholder_configs}")
@@ -201,9 +192,7 @@ def verify_conversion(
del dst_dir, components
# TODO: load each emitted stateful component through its production loader and
# assert strict load, or document exact allowed missing/unexpected keys.
raise NotImplementedError(
"Implement production config validation and strict-load checks"
)
raise NotImplementedError("Implement production config validation and strict-load checks")
def write_component(
@@ -216,9 +205,7 @@ def write_component(
if component_dir.exists() and any(component_dir.iterdir()):
shutil.rmtree(component_dir)
component_dir.mkdir(parents=True, exist_ok=True)
save_file(
dict(state), str(component_dir / "diffusion_pytorch_model.safetensors")
)
save_file(dict(state), str(component_dir / "diffusion_pytorch_model.safetensors"))
if config is not None:
config_path = component_dir / config_filename(name)
with config_path.open("w", encoding="utf-8") as f:
@@ -261,9 +248,7 @@ def convert(
if layout in {"monolithic", "raw_official"}:
# TODO: replace model.safetensors with the official monolithic file name.
components = split_monolithic(
load_checkpoint(default_monolithic_checkpoint(src_path))
)
components = split_monolithic(load_checkpoint(default_monolithic_checkpoint(src_path)))
elif layout in {"separate_components", "mixed"}:
if not src_path.is_dir():
raise ValueError(f"{layout} layout requires a source directory: {src_path}")
@@ -271,9 +256,7 @@ def convert(
else:
raise ValueError(f"Unsupported template layout: {layout}")
copied = (
copy_passthrough(src_path, dst_dir) if src_path.is_dir() else []
)
copied = (copy_passthrough(src_path, dst_dir) if src_path.is_dir() else [])
configs = build_component_configs(src_path if src_path.is_dir() else src_path.parent)
validate_component_configs(configs)
for name, state in components.items():
@@ -289,9 +272,7 @@ def convert(
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--src", required=True, help="HF repo id, local dir, or checkpoint path"
)
parser.add_argument("--src", required=True, help="HF repo id, local dir, or checkpoint path")
parser.add_argument("--revision", help="HF branch, tag, or commit for repo sources")
parser.add_argument(
"--dst",
@@ -24,8 +24,8 @@ from typing import Any
import torch
FAMILY: str = "<family>" # e.g. "magi_human", "ltx2", "wan"
COMPONENT: str = "<component>" # e.g. "dit", "vae", "encoder"
FAMILY: str = "<family>" # e.g. "magi_human", "ltx2", "wan"
COMPONENT: str = "<component>" # e.g. "dit", "vae", "encoder"
DRILL_LAYER_ENV: str = "<FAMILY>_DEBUG_DRILL_LAYER"
HYPOTHESIS_ENV: str = "<FAMILY>_DEBUG_PATCH_<HYPOTHESIS>"
REL_THRESHOLD: float = 0.005 # 0.5% abs_mean drift flags a block as divergent
@@ -94,6 +94,7 @@ def _attach_block_hooks(
handles: list[Any] = []
def _hook(name: str):
def fn(_module, _inputs, outputs):
t = outputs[0] if isinstance(outputs, tuple) else outputs
if not torch.is_tensor(t):
@@ -101,6 +102,7 @@ def _attach_block_hooks(
log.append({"side": label, **_stat(name, t)})
if tensors is not None:
tensors[name] = t.detach().float().cpu()
return fn
def _pre_hook(name: str):
@@ -114,6 +116,7 @@ def _attach_block_hooks(
log.append({"side": label, **_stat(key, t)})
if tensors is not None:
tensors[key] = t.detach().float().cpu()
return fn
# TODO: adapt attribute paths to your model. Remove adapter block if absent.
@@ -131,43 +134,21 @@ def _attach_block_hooks(
# magi-human uses: attention, mlp.pre_norm, mlp.up_gate_proj,
# mlp.down_proj (pre+post), mlp, attn_post_norm, mlp_post_norm.
if hasattr(layer, "attention"):
handles.append(
layer.attention.register_forward_hook(_hook(f"{tag}.attention"))
)
handles.append(layer.attention.register_forward_hook(_hook(f"{tag}.attention")))
if hasattr(layer, "mlp"):
mlp = layer.mlp
if hasattr(mlp, "pre_norm"):
handles.append(
mlp.pre_norm.register_forward_hook(_hook(f"{tag}.mlp.pre_norm"))
)
handles.append(mlp.pre_norm.register_forward_hook(_hook(f"{tag}.mlp.pre_norm")))
if hasattr(mlp, "up_gate_proj"):
handles.append(
mlp.up_gate_proj.register_forward_hook(
_hook(f"{tag}.mlp.up_gate_proj")
)
)
handles.append(mlp.up_gate_proj.register_forward_hook(_hook(f"{tag}.mlp.up_gate_proj")))
if hasattr(mlp, "down_proj"):
handles.append(
mlp.down_proj.register_forward_pre_hook(
_pre_hook(f"{tag}.mlp.down_proj")
)
)
handles.append(
mlp.down_proj.register_forward_hook(_hook(f"{tag}.mlp.down_proj"))
)
handles.append(mlp.down_proj.register_forward_pre_hook(_pre_hook(f"{tag}.mlp.down_proj")))
handles.append(mlp.down_proj.register_forward_hook(_hook(f"{tag}.mlp.down_proj")))
handles.append(mlp.register_forward_hook(_hook(f"{tag}.mlp")))
if hasattr(layer, "attn_post_norm"):
handles.append(
layer.attn_post_norm.register_forward_hook(
_hook(f"{tag}.attn_post_norm")
)
)
handles.append(layer.attn_post_norm.register_forward_hook(_hook(f"{tag}.attn_post_norm")))
if hasattr(layer, "mlp_post_norm"):
handles.append(
layer.mlp_post_norm.register_forward_hook(
_hook(f"{tag}.mlp_post_norm")
)
)
handles.append(layer.mlp_post_norm.register_forward_hook(_hook(f"{tag}.mlp_post_norm")))
return handles
@@ -193,11 +174,9 @@ def _write_log(entries: list[dict], path: Path) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
with open(path, "w") as f:
for e in entries:
f.write(
f"{e['name']} {e['shape']} "
f"{e['abs_mean']:.8f} {e['sum']:.4f} "
f"{e['min']:.6f} {e['max']:.6f}\n"
)
f.write(f"{e['name']} {e['shape']} "
f"{e['abs_mean']:.8f} {e['sum']:.4f} "
f"{e['min']:.6f} {e['max']:.6f}\n")
def _sort_key(name: str, drill_layer: int) -> tuple:
@@ -205,9 +184,14 @@ def _sort_key(name: str, drill_layer: int) -> tuple:
return (0, "")
if name.startswith(f"L{drill_layer:02d}."):
sub_order = {
"attention": 0, "attn_post_norm": 1, "mlp.pre_norm": 2,
"mlp.up_gate_proj": 3, "mlp.down_proj<in>": 4,
"mlp.down_proj": 5, "mlp": 6, "mlp_post_norm": 7,
"attention": 0,
"attn_post_norm": 1,
"mlp.pre_norm": 2,
"mlp.up_gate_proj": 3,
"mlp.down_proj<in>": 4,
"mlp.down_proj": 5,
"mlp": 6,
"mlp_post_norm": 7,
}.get(name.split(".", 1)[1], 9)
return (1, f"block[{drill_layer:02d}]", sub_order)
if name.startswith("block["):
@@ -216,10 +200,8 @@ def _sort_key(name: str, drill_layer: int) -> tuple:
def _print_table(by_name: dict[str, dict], drill_layer: int) -> int | None:
hdr = (
f"{'name':<18} {'up_shape':<22} {'up_absmean':>12} {'fv_absmean':>12} "
f"{'absmean_diff':>14} {'rel%':>8} {'up_sum':>14} {'fv_sum':>14} {'sum_diff':>12}"
)
hdr = (f"{'name':<18} {'up_shape':<22} {'up_absmean':>12} {'fv_absmean':>12} "
f"{'absmean_diff':>14} {'rel%':>8} {'up_sum':>14} {'fv_sum':>14} {'sum_diff':>12}")
print(f"\n{hdr}\n{'-' * len(hdr)}")
first_div: int | None = None
for name in sorted(by_name.keys(), key=lambda n: _sort_key(n, drill_layer)):
@@ -235,11 +217,9 @@ def _print_table(by_name: dict[str, dict], drill_layer: int) -> int | None:
flag = " <<< DIVERGE"
if first_div is None:
first_div = int(name[len("block["):-1])
print(
f"{name:<18} {str(up['shape']):<22} {up['abs_mean']:>12.6f} "
f"{fv['abs_mean']:>12.6f} {am_diff:>14.6f} {am_rel * 100:>7.3f}% "
f"{up['sum']:>14.4f} {fv['sum']:>14.4f} {sum_diff:>12.4f}{flag}"
)
print(f"{name:<18} {str(up['shape']):<22} {up['abs_mean']:>12.6f} "
f"{fv['abs_mean']:>12.6f} {am_diff:>14.6f} {am_rel * 100:>7.3f}% "
f"{up['sum']:>14.4f} {fv['sum']:>14.4f} {sum_diff:>12.4f}{flag}")
return first_div
@@ -255,10 +235,8 @@ def _print_elementwise(up_t: dict[str, torch.Tensor], fv_t: dict[str, torch.Tens
continue
diff = (a - b).abs()
rel = (diff.mean().item() / max(a.abs().mean().item(), 1e-9)) * 100
print(
f"{name:<30} {str(tuple(a.shape)):<22} "
f"{diff.max().item():>12.6f} {diff.mean().item():>12.6f} {rel:>9.4f}%"
)
print(f"{name:<30} {str(tuple(a.shape)):<22} "
f"{diff.max().item():>12.6f} {diff.mean().item():>12.6f} {rel:>9.4f}%")
def main() -> None:
@@ -43,12 +43,10 @@ def _add_official_to_path() -> Path:
def _log_tensor_stats(label: str, tensor: torch.Tensor) -> None:
value = tensor.detach().float()
print(
f"[{_MODEL_FAMILY} PIPELINE] {label}: shape={tuple(tensor.shape)} "
f"dtype={tensor.dtype} device={tensor.device} "
f"min={value.min().item():.6f} max={value.max().item():.6f} "
f"mean={value.mean().item():.6f} std={value.std().item():.6f}"
)
print(f"[{_MODEL_FAMILY} PIPELINE] {label}: shape={tuple(tensor.shape)} "
f"dtype={tensor.dtype} device={tensor.device} "
f"min={value.min().item():.6f} max={value.max().item():.6f} "
f"mean={value.mean().item():.6f} std={value.std().item():.6f}")
def _extract_tensor(output: Any, key: str) -> torch.Tensor:
@@ -73,10 +71,8 @@ def _run_official_pipeline(
device: torch.device,
) -> Any:
del official_path, params, device
pytest.skip(
"TODO: import the official pipeline/factory, load official weights, "
"run with params, and return the comparison target."
)
pytest.skip("TODO: import the official pipeline/factory, load official weights, "
"run with params, and return the comparison target.")
def _run_fastvideo_pipeline(model_path: Path, params: dict[str, Any]) -> Any:
@@ -146,8 +142,6 @@ def test_todo_model_family_pipeline_official_parity() -> None:
assert official_tensor.shape == fastvideo_tensor.shape
diff = (official_tensor - fastvideo_tensor).abs()
print(
f"diff max={diff.max().item():.6f} "
f"mean={diff.mean().item():.6f} median={diff.median().item():.6f}"
)
print(f"diff max={diff.max().item():.6f} "
f"mean={diff.mean().item():.6f} median={diff.median().item():.6f}")
assert_close(fastvideo_tensor, official_tensor, atol=1e-2, rtol=1e-2)
@@ -1,20 +1,23 @@
---
name: reseed-performance-baseline
description: Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, or environment-caused benchmark shift using one or more reviewed normalized performance JSONs. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline in FastVideo/performance-tracking must be advanced from a consistent batch of reviewed source results. The workflow backs up existing history under /tmp, validates all source JSONs for the same (model_id, gpu_type), rejects internally inconsistent source batches, uploads one success=true reseed record per accepted source JSON, and offers to clean local temp state after a successful upload.
description: Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, environment-caused benchmark shift, or reviewed v2 calibration using one or more reviewed normalized performance JSONs. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline in FastVideo/performance-tracking must be advanced, or when a new v2 exact comparable identity needs its first approved baseline. The workflow backs up existing history under /tmp, validates all source JSONs for the same legacy (model_id, gpu_type) target or the same v2 exact identity, rejects internally inconsistent source batches, uploads one success=true baseline record per accepted source JSON, and offers to clean local temp state after a successful upload.
---
# Re-seed Performance Baseline
## Purpose
Replace or advance the rolling performance baseline for a single
`(model_id, gpu_type)` pair in the HF dataset
`FastVideo/performance-tracking`.
Replace or advance the rolling performance baseline in the HF dataset
`FastVideo/performance-tracking`. Legacy targets are scoped by
`(model_id, gpu_type)`. V2 targets are scoped by exact comparable identity:
`workload_id`, `variant_id`, `benchmark_version`, `hardware_profile_id`,
`software_profile_id`, and `recipe_fingerprint`.
Performance comparison uses the median of up to the last 5 successful records
for the same model and GPU. Failed records are useful audit history, but they
do not move the future baseline because `compare_baseline.py` loads records
with `successful_only=True`.
Performance comparison uses the median of up to the last 5 successful,
baseline-eligible records for the same target. Failed or calibration-only
records are useful audit history, but they do not move the future baseline
because `compare_baseline.py` loads records with `successful_only=True` and
`baseline_eligible_only=True`.
This skill now reseeds from a reviewed batch of one or more source performance
JSONs. It uploads one new `success=true` record per accepted source JSON; it
@@ -22,11 +25,13 @@ does not blindly replicate one measurement into 3 or 5 records. The effective
reseed size is therefore dynamic and equals the number of provided, validated,
internally consistent source JSONs.
If the operator provides fewer than 3 records, call out that the last-5 rolling
median may not move immediately. If the operator provides 3 consistent shifted
records, the rolling median usually moves immediately. If the operator provides
5 consistent shifted records, the last-5 window is effectively reset to the new
runtime profile.
For baseline shifts with existing history, if the operator provides fewer than
3 records, call out that the last-5 rolling median may not move immediately. If
the operator provides 3 consistent shifted records, the rolling median usually
moves immediately. If the operator provides 5 consistent shifted records, the
last-5 window is effectively reset to the new runtime profile. For the first
approved v2 baseline of a new exact identity, one reviewed calibration seed is
enough for the next comparable run to leave `CALIBRATION_NEEDED`.
These records are intentional operator-approved baseline resets, not ordinary
independent main-branch persistence. Mark them clearly with provenance fields
@@ -66,8 +71,8 @@ approval, then upload reviewed accepted baseline records.
| Parameter | Required | Description |
|-----------|----------|-------------|
| `model_id` | Yes | Benchmark id, e.g. `wan-t2v-1.3b-2gpu`. This maps to the HF subdirectory after `sanitize(model_id)`. |
| `gpu_type` | Yes | Exact GPU device string from the performance record, e.g. the L40S device name emitted by CI. Baselines are GPU-specific. |
| `model_id` | Legacy required; v2 inferred | Benchmark id, e.g. `wan-t2v-1.3b-2gpu`. This maps to the HF subdirectory after `sanitize(model_id)`. For v2 records, use the `model_id` from each source artifact only as the upload directory; comparison is by exact identity. |
| `gpu_type` | Legacy required; v2 inferred | Exact GPU device string from the performance record, e.g. the L40S device name emitted by CI. V2 hardware matching uses `hardware_profile_id`; preserve `gpu_type` as display metadata. |
| `source_results` | Yes | One or more local paths or Buildkite artifact URLs for accepted shifted performance JSONs. Prefer normalized `normalized_perf_*.json` artifacts emitted by `compare_baseline.py`. Accept `source_result` as an alias only for a single JSON. |
| `max_intra_batch_regression` | No | Maximum allowed regression of any source JSON against the source batch median. Default: `0.05` (5%). |
| `intent_rationale` | Yes | One-line explanation for why the baseline shift is legitimate. This is written into provenance and should be reused in the PR. |
@@ -78,10 +83,14 @@ Hardcoded defaults:
supported by the code, but use the default unless the user explicitly asks).
- Local sync root: `/tmp/perf-tracking` (`PERFORMANCE_TRACKING_ROOT` override
is supported).
- Prepared-record staging root: `/tmp/performance_reseed_prepared`
(`PERFORMANCE_RESEED_STAGING_ROOT` override is supported). Keep it separate
and non-nested from the sync root.
- Backup root: `/tmp/performance_reseed_backup`.
- Download scratch root for source artifact URLs: `/tmp/performance_reseed_source`.
- Baseline window: last 5 `success=true` records for the same
`(model_id, gpu_type)`.
- Baseline window: last 5 `success=true`, `baseline_eligible=true` records
for the same legacy `(model_id, gpu_type)` target or the same v2 exact
comparable identity.
- Reseed count: dynamic. Upload exactly one accepted seed record per validated
source JSON.
@@ -115,12 +124,24 @@ with open(source_result, encoding="utf-8") as f:
record = json.load(f)
```
Stop if any normalized record's `model_id` or `gpu_type` does not match the
requested `model_id` and `gpu_type`.
Classify the source batch before continuing:
- **Legacy source records** have no v2 exact identity fields. Stop if any
normalized record's `model_id` or `gpu_type` does not match the requested
`model_id` and `gpu_type`.
- **V2 source records** have exact identity fields. Stop unless every source
record has all six comparable identity fields and they are identical across
the batch: `workload_id`, `variant_id`, `benchmark_version`,
`hardware_profile_id`, `software_profile_id`, and `recipe_fingerprint`.
Do not fall back to legacy `(model_id, gpu_type)` matching for v2 records.
The source records may have `success: false` when they came from failed
rolling baseline comparisons. That is expected; only the reviewed reseed
records become new `success: true` baseline records after explicit approval.
For a first v2 baseline seed, the source records must instead be successful
scheduled-main full-suite `CALIBRATION_NEEDED` normalized artifacts. Reject PR,
local, direct-run, non-main-branch, or non-full-suite calibration artifacts as
seed sources.
Sort validated source records by their original `timestamp` ascending before
preparing the seed records. If a source timestamp is missing or unparsable,
@@ -194,7 +215,7 @@ export HF_REPO_ID="${HF_REPO_ID:-FastVideo/performance-tracking}"
python -c 'from fastvideo.performance.hf_store import sync_from_hf; import os; sync_from_hf(os.environ["PERFORMANCE_TRACKING_ROOT"], strict=True)'
```
Then back up only the sanitized model directory under `/tmp`:
For legacy records, back up the sanitized model directory under `/tmp`:
```bash
SHORT_COMMIT=$(git rev-parse --short=12 HEAD)
@@ -209,6 +230,16 @@ mkdir -p "$BACKUP_DIR"
cp -R "${PERFORMANCE_TRACKING_ROOT}/${MODEL_SAFE}" "$BACKUP_DIR/" 2>/dev/null || true
```
For v2 records, back up the full local tracking root after sync. Exact identity
lookup scans across model directories, so a benchmark rename may have relevant
history outside the current source artifact's `model_id` directory:
```bash
BACKUP_DIR="/tmp/performance_reseed_backup/${TIMESTAMP}_${SHORT_COMMIT}_v2_exact_identity"
mkdir -p "$BACKUP_DIR"
cp -R "${PERFORMANCE_TRACKING_ROOT}" "$BACKUP_DIR/tracking-root"
```
Write provenance next to the backup:
```bash
@@ -231,7 +262,9 @@ first baseline seed. Continue, but report that baseline history was empty.
### 3. Compute old baseline and candidate shift
Load the last 5 successful records for the target:
Load the last 5 successful baseline records for the target.
For legacy targets:
```python
from fastvideo.performance.hf_store import load_records_for_model
@@ -242,6 +275,28 @@ records = load_records_for_model(
"<gpu_type>",
last_n=5,
successful_only=True,
baseline_eligible_only=True,
)
```
For v2 exact-identity targets:
```python
from fastvideo.performance.hf_store import load_records_for_identity
records = load_records_for_identity(
"/tmp/perf-tracking",
{
"workload_id": "<workload_id>",
"variant_id": "<variant_id>",
"benchmark_version": "<benchmark_version>",
"hardware_profile_id": "<hardware_profile_id>",
"software_profile_id": "<software_profile_id>",
"recipe_fingerprint": "<recipe_fingerprint>",
},
last_n=5,
successful_only=True,
baseline_eligible_only=True,
)
```
@@ -257,7 +312,8 @@ medians after appending the proposed seed records, and source batch spread for:
Also print how many successful old records exist. Make clear:
- 1 seed record usually does not move a last-5 median by itself.
- 1 seed record usually does not move an existing last-5 median by itself, but
it is enough to establish the first v2 baseline for a new exact identity.
- 3 consistent seed records usually move the last-5 median immediately.
- 5 consistent seed records effectively reset the last-5 window.
- The records are intentional approved baseline resets and must be labeled
@@ -267,10 +323,10 @@ Also print how many successful old records exist. Make clear:
Require an explicit confirmation phrase before preparing the upload:
> About to RE-SEED performance baseline for `<model_id>` on `<gpu_type>`.
> About to RE-SEED performance baseline for `<target description>`.
> This will upload `<N>` new `success=true` records to
> `FastVideo/performance-tracking/<sanitize(model_id)>/`, one per accepted
> source JSON.
> `FastVideo/performance-tracking/<sanitize(model_id)>/` or the source
> artifact's v2 model directory, one per accepted source JSON.
>
> Reason: `<intent_rationale>`
> Source results: `<source_results>`
@@ -288,25 +344,63 @@ Do not continue unless the user types exactly `confirm performance reseed`.
### 5. Create the accepted seed records
Create one seed record from each normalized source result. Do not copy the
Create one seed record from each normalized source result.
For first v2 baseline seeds, use the scoped utility. It validates exact
identity, requires successful scheduled-main full-suite `CALIBRATION_NEEDED`
source artifacts, preserves the normalized v2 identity and metadata fields,
and writes seed records with `success=true`, `baseline_eligible=true`, and
`comparison_status=PASS`:
```bash
python fastvideo/tests/performance/seed_baseline.py \
--source-result <normalized_perf_1.json> \
--source-result <normalized_perf_2.json> \
--intent-rationale "<intent_rationale>" \
--max-intra-batch-regression 0.05 \
--tracking-root "${PERFORMANCE_TRACKING_ROOT}" \
--staging-root "${PERFORMANCE_RESEED_STAGING_ROOT:-/tmp/performance_reseed_prepared}"
```
The utility is prepare-only and intentionally has no upload option. Upload the
scoped records only after the separate confirmation in step 6.
The utility validates against an isolated fresh HF snapshot and leaves
`PERFORMANCE_TRACKING_ROOT` untouched; that argument only proves the staging
root is separate from the operator's tracking mirror. Before writing, it stops
if the exact identity already has a successful baseline-eligible record or if
the workload/variant/version already trusts another recipe. It atomically
reserves the exact identity and writes a digest-protected upload manifest bound
to the current HF endpoint, repository id, and repository type. Keep the
prepared records, manifest, source files, and reservation unchanged until the
operation is uploaded or explicitly cleaned up.
If the prepared seed records look correct, upload only those scoped records in
step 7. Do not rerun the utility with a different source list after approval.
For legacy reseeds or accepted v2 baseline shifts from regression artifacts,
create one seed record from each normalized source result. Do not copy the
source JSON wholesale.
Infer the baseline field allowlist from all existing HF records for the target
`(model_id, gpu_type)` after syncing, including both `success=true` and
`success=false` records. Use the union of non-provenance keys present in those
target records, preserving only fields that also exist in the normalized
source record or are explicitly set by the reseed workflow. Always include
`model_id`, `timestamp`, and `success` because the upload path and baseline
loader depend on them. Always set `timestamp` to a fresh reseed timestamp and
`success` to `true`. Do not include unrelated source-only fields that are
absent from existing HF records.
after syncing, including both `success=true` and `success=false` records. For
legacy targets the target is `(model_id, gpu_type)`. For v2 baseline-shift
reseeds the target is the exact comparable identity. Use the union of
non-provenance keys present in those target records, preserving only fields
that also exist in the normalized source record or are explicitly set by the
reseed workflow. Always include `model_id`, `timestamp`, `success`,
`baseline_eligible`, and `comparison_status` because the upload path and
baseline loader depend on them. Always set `timestamp` to a fresh reseed
timestamp, `success` to `true`, `baseline_eligible` to `true`, and
`comparison_status` to `PASS`. Do not include unrelated source-only fields
that are absent from existing HF records.
Exclude existing provenance or operator metadata from the inferred baseline
field allowlist. At minimum, exclude keys prefixed with `baseline_reseed` and
any fields known to be local-only audit metadata.
If there are no previous HF records for the target model/GPU, fall back to this
default baseline field list:
If there are no previous HF records for the target, fall back to this default
baseline field list:
- `model_id`
- `timestamp`
@@ -319,6 +413,22 @@ default baseline field list:
- `dit_time_s`
- `vae_decode_time_s`
- `success`
- `baseline_eligible`
- `comparison_status`
For v2 baseline-shift reseeds with no previous HF records for the exact
identity, also preserve:
- `workload_id`
- `variant_id`
- `benchmark_version`
- `recipe_fingerprint`
- `hardware_profile_id`
- `software_profile_id`
- `recipe`
- `hardware_profile`
- `software_profile`
- `software_comparison_profile`
Do not upload extra fields from the source artifact.
@@ -334,6 +444,22 @@ Optional provenance fields are allowed and useful:
- `baseline_reseed_operator`
- `baseline_reseed_max_intra_batch_regression`
The v2 calibration seed utility writes analogous first-seed provenance:
- `baseline_seed: true`
- `baseline_seed_reason`
- `baseline_seed_source_result`
- `baseline_seed_source_status`
- `baseline_seed_source_timestamp`
- `baseline_seed_source_success`
- `baseline_seed_source_run_source`
- `baseline_seed_source_branch`
- `baseline_seed_source_test_scope`
- `baseline_seed_source_pr_number`
- `baseline_seed_batch_size`
- `baseline_seed_batch_index`
- `baseline_seed_operator`
Use a fresh reseed timestamp for each seed record, not the original source
result timestamp. This is required because
`load_records_for_model(..., last_n=5)` keeps the last records after loading
@@ -356,7 +482,8 @@ Prefer uploading new accepted seed records so failed history remains visible.
Print:
- Backup directory path under `/tmp`.
- Prepared local record paths under `PERFORMANCE_TRACKING_ROOT`.
- Prepared local record paths under `PERFORMANCE_RESEED_STAGING_ROOT`.
- Prepared upload-manifest path under the identity reservation.
- HF paths that will receive the new records.
- Old rolling medians.
- Source batch medians, source batch spread, reseed count, and candidate
@@ -368,22 +495,36 @@ prepared records plus backup on disk.
### 7. Upload only the scoped records
Use the shared storage helper so the path and repo type match CI:
For a first v2 calibration seed, use the manifest uploader after the user
replies exactly `upload`:
```python
from fastvideo.performance.hf_store import upload_record
upload_record("<local_record_path>", record, strict=True)
```bash
python -c 'from fastvideo.tests.performance.seed_baseline import upload_prepared_seed_manifest; print(upload_prepared_seed_manifest("<prepared_manifest>"))'
```
Run it once per prepared record. Each upload goes to:
The uploader verifies the source and prepared-record digests, pins and scans
the current HF revision, rechecks exact-identity and recipe-cohort conflicts,
and writes the entire batch in one commit whose `parent_commit` must still be
current. A concurrent Hub update makes the commit fail. Do not retry
automatically: preserve staging, refresh/review remote state, and request a new
explicit `upload` after the conflict is understood. Each record goes to:
```text
FastVideo/performance-tracking/<sanitize(model_id)>/<record_filename>.json
```
Never bulk upload the whole tracking root. Never modify another model's
directory in the same operation.
Never call `upload_record()` once per first-seed record: that can partially
land the batch and has no compare-and-swap guard.
For a legacy reseed or an accepted v2 baseline shift, the first-seed manifest
validator does not apply because an eligible baseline already exists. Upload
only the individually reviewed records prepared in step 5 with the shared
`upload_record(local_path, record, strict=True)` helper. Stop on the first
failure and report exactly which records reached HF; do not silently rerun or
replicate the remainder.
Never bulk upload the tracking or staging root, and never modify another
model's directory in the same operation.
### 8. Report outcome and offer cleanup
@@ -405,9 +546,14 @@ distinguish an accepted baseline shift from a hidden regression.
After the upload is verified, ask whether the user wants to clear temporary
local state. Explain what each directory is for:
- `PERFORMANCE_TRACKING_ROOT`, usually `/tmp/perf-tracking`: local synced
mirror of `FastVideo/performance-tracking` plus the prepared local seed
records used for scoped upload.
- `PERFORMANCE_TRACKING_ROOT`, usually `/tmp/perf-tracking`: read-only local
synced mirror used for operator review and reporting. First-v2 preparation
independently proves remote state from a fresh temporary HF snapshot.
- `PERFORMANCE_RESEED_STAGING_ROOT`, usually
`/tmp/performance_reseed_prepared`: prepared local seed records used for the
scoped upload, plus the identity reservation and digest manifest. Keeping
this separate prevents aborted preparations from appearing in later
baseline reads.
- `/tmp/performance_reseed_backup/<...>`: local backup of the target model's
pre-reseed HF history plus `PROVENANCE.txt`, kept so a bad reseed can be
audited or corrected.
@@ -417,14 +563,18 @@ local state. Explain what each directory is for:
Ask:
> Reseed succeeded. Do you want me to delete the local temp tracking mirror,
> source downloads, and reseed backup under `/tmp`? These files are local
> safety/audit artifacts only; HF already has the uploaded records.
> this reseed's prepared staging records, source downloads, and reseed backup
> under `/tmp`? These files are local safety/audit artifacts only; HF already
> has the uploaded records.
>
> Reply `cleanup reseed temp` to delete them, anything else to keep them.
Do not delete anything unless the user replies exactly
`cleanup reseed temp`. If cleanup is requested, remove only the specific
directories created for this reseed. Never remove unrelated `/tmp` contents.
directories and prepared record paths created for this reseed. Do not remove
the shared staging root when it contains other records. Remove this operation's
identity reservation only with its prepared records and manifest, and never
remove unrelated `/tmp` contents.
## Failure modes and handling
@@ -436,19 +586,34 @@ directories created for this reseed. Never remove unrelated `/tmp` contents.
against the source batch median by more than `max_intra_batch_regression`.
Ask for cleaner sources or a reviewed explanation before continuing.
- **Too few source records to move the median.** Continue only after making
clear that one or two records may not immediately move the last-5 median.
clear that one or two records may not immediately move an existing last-5
median. This warning does not block a first v2 calibration seed for an exact
identity with no eligible baseline yet.
- **The source results are noisy or suspicious.** Stop. Reseeding amplifies
those measurements into the baseline, so they must be reviewed first.
- **HF sync fails.** Stop for destructive reseeds. A stale or empty sync can
make the old baseline look missing.
- **The exact v2 identity already has an eligible baseline.** Stop. The
`CALIBRATION_NEEDED` artifact is stale; use the reviewed baseline-shift path
instead of the first-seed utility.
- **The workload/variant/version trusts another recipe.** Stop. The source is
stale relative to the current recipe cohort and must not bypass
`RECIPE_MISMATCH` by creating a second trusted recipe.
- **The staging root already has a prepared seed for the exact identity.**
Stop and reuse, upload, or explicitly clean that preparation. Do not prepare
another copy of the same measurement.
- **The conditional Hub commit loses its parent race.** Stop without retrying.
Keep the preparation, refresh and review the new remote state, then request
a new explicit `upload` only if the seed is still valid.
- **Candidate still violates fixed thresholds.** Report that this skill only
handles the rolling HF baseline; update benchmark JSON thresholds in code
review if maintainers accept the new absolute limit.
- **The user aborts at either confirmation.** Leave the backup and prepared
records on disk. Nothing should be uploaded.
- **The user declines cleanup.** Keep `/tmp/perf-tracking`, the source
download directory if any, and `/tmp/performance_reseed_backup/<...>` in
place for audit/debugging.
- **The user declines cleanup.** Keep `/tmp/perf-tracking`, the prepared seed
records under `/tmp/performance_reseed_prepared`, the source download
directory if any, and `/tmp/performance_reseed_backup/<...>` in place for
audit/debugging.
- **A bad seed was uploaded.** Use the backup and HF history to identify the
uploaded file, then remove or supersede it with an explicitly reviewed
corrective record. Do not silently rewrite unrelated history.
@@ -459,8 +624,9 @@ directories created for this reseed. Never remove unrelated `/tmp` contents.
intentional baseline replacement.
- `fastvideo/tests/performance/compare_baseline.py` — normalization, rolling
median comparison, and persistence rules.
- `fastvideo/performance/hf_store.py` — HF sync, record loading,
`sanitize()`, and `upload_record()`.
- `fastvideo/performance/hf_store.py` — HF sync and record loading helpers.
- `fastvideo/tests/performance/seed_baseline.py` — first-seed preparation,
staging reservation, manifest validation, and conditional batch upload.
- `fastvideo/tests/performance/test_inference_performance.py` — source result
JSON schema.
- `.buildkite/performance-benchmarks/tests/*.json` — fixed absolute benchmark
@@ -473,3 +639,4 @@ directories created for this reseed. Never remove unrelated `/tmp` contents.
| 2026-05-03 | Initial version. Sister workflow to `reseed-ssim-references`, scoped to one performance `(model_id, gpu_type)` baseline seed with backup, confirmation, provenance, and `success=true` upload. |
| 2026-05-03 | Previous policy: replicate one approved shifted source result into 3 success records by default, or 5 only when explicitly requested. Add provenance marker for replicated-source reseeds. Superseded by the 2026-05-08 dynamic multi-source policy. |
| 2026-05-08 | Replace fixed 3/5 replication with dynamic multi-source reseeding: upload one seed record per reviewed source JSON, validate intra-batch consistency, move backup/source scratch under `/tmp`, and ask whether to clean temp state after successful upload. |
| 2026-07-13 | Keep first-v2-seed preparation outside the canonical mirror, reserve staging identities atomically, reject stale or replayed calibration seeds, and upload reviewed manifests with a single parent-guarded Hub commit. |
@@ -3,7 +3,7 @@
"config_schema_version": 2,
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp2",
"benchmark_version": 2,
"benchmark_version": 3,
"description": "Wan2.1 T2V 1.3B inference performance",
"model": {
"model_path": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
+15 -2
View File
@@ -1,6 +1,8 @@
env:
IMAGE_VERSION: "py3.12-latest"
BUILDKITE_CLEAN_CHECKOUT: true
# Buildkite only launches Modal; remote jobs initialize their own submodules.
BUILDKITE_GIT_SUBMODULES: false
notify:
- github_commit_status:
@@ -66,6 +68,17 @@ steps:
limit: 2
agents:
queue: "default"
- label: ":vertical_traffic_light: Golden-Gate Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "golden_gate"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Unit Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "unit_test"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
@@ -402,7 +415,7 @@ steps:
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
command: "timeout 25m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Training Tests"
env:
- TEST_TYPE=training
@@ -413,7 +426,7 @@ steps:
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
command: "timeout 25m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Distillation DMD Tests"
env:
- TEST_TYPE=distillation_dmd
+5 -1
View File
@@ -76,7 +76,7 @@ EFFECTIVE_PR=${BUILDKITE_PULL_REQUEST:-false}
if [ "$EFFECTIVE_PR" = "false" ] && [ -n "${PR_NUMBER:-}" ]; then
EFFECTIVE_PR=$PR_NUMBER
fi
MODAL_ENV="BUILDKITE_REPO=$BUILDKITE_REPO BUILDKITE_COMMIT=$BUILDKITE_COMMIT BUILDKITE_PULL_REQUEST=$EFFECTIVE_PR BUILDKITE_BRANCH=${BUILDKITE_BRANCH:-} TEST_SCOPE=${TEST_SCOPE:-} BUILDKITE_BUILD_URL=${BUILDKITE_BUILD_URL:-} BUILDKITE_BUILD_ID=${BUILDKITE_BUILD_ID:-} BUILDKITE_JOB_ID=${BUILDKITE_JOB_ID:-} IMAGE_VERSION=$IMAGE_VERSION"
MODAL_ENV="BUILDKITE_REPO=$BUILDKITE_REPO BUILDKITE_COMMIT=$BUILDKITE_COMMIT BUILDKITE_PULL_REQUEST=$EFFECTIVE_PR BUILDKITE_BRANCH=${BUILDKITE_BRANCH:-} BUILDKITE_SOURCE=${BUILDKITE_SOURCE:-} TEST_SCOPE=${TEST_SCOPE:-} BUILDKITE_BUILD_URL=${BUILDKITE_BUILD_URL:-} BUILDKITE_BUILD_ID=${BUILDKITE_BUILD_ID:-} BUILDKITE_JOB_ID=${BUILDKITE_JOB_ID:-} IMAGE_VERSION=$IMAGE_VERSION"
POST_RUN_HOOK=""
@@ -187,6 +187,10 @@ case "$TEST_TYPE" in
log "Running transformer tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_transformer_tests"
;;
"golden_gate")
log "Running golden-gate tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_golden_gate_tests"
;;
"ssim")
log "Running SSIM tests..."
SSIM_BOOTSTRAP_ARGS=$(ssim_bootstrap_args)
+149
View File
@@ -0,0 +1,149 @@
name: macOS MLX Smoke
on:
pull_request:
branches: [main]
paths:
- ".github/workflows/ci-macos-mlx.yml"
- "fastvideo/mlx_runtime/**"
- "fastvideo/tests/mlx/**"
- "fastvideo/tests/platforms/test_mps_vsa_error.py"
- "fastvideo/platforms/mps.py"
- "fastvideo/platforms/__init__.py"
- "fastvideo/__init__.py"
- "examples/inference/basic/mlx_*.py"
- "fastvideo/benchmarks/mlx_*.py"
- "pyproject.toml"
workflow_dispatch:
permissions:
contents: read
concurrency:
group: macos-mlx-${{ github.ref }}
cancel-in-progress: true
jobs:
mlx-smoke:
if: github.event_name == 'workflow_dispatch' || github.event.pull_request.draft != true
runs-on: macos-15
timeout-minutes: 25
env:
FASTVIDEO_ATTENTION_BACKEND: TORCH_SDPA
TOKENIZERS_PARALLELISM: "false"
MASTER_ADDR: localhost
MASTER_PORT: "29513"
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
- uses: astral-sh/setup-uv@v3
- name: Install lightweight MLX smoke dependencies
run: |
uv pip install --system \
--index-url https://download.pytorch.org/whl/cpu \
torch==2.11.0 torchvision torchaudio
uv pip install --system \
pytest numpy scipy pillow imageio einops cloudpickle filelock \
PyYAML diffusers huggingface_hub remote-pdb safetensors loguru mlx \
"ftfy>=6.3.1" "opencv-python>=4.10.0.84" psutil "transformers>=5.0.0"
- name: Show Apple runtime
run: |
python - <<'PY'
import platform
import mlx.core as mx
import torch
print("machine:", platform.machine())
print("processor:", platform.processor())
print("mlx default device:", mx.default_device())
memory_size = mx.metal.device_info().get("memory_size") if mx.metal.is_available() else "metal unavailable"
print("mlx memory_size:", memory_size)
print("torch:", torch.__version__)
print("torch mps available:", torch.backends.mps.is_available())
PY
- name: Run MLX smoke tests
run: |
python -m pytest \
fastvideo/tests/mlx/test_dmd_sampling.py \
fastvideo/tests/mlx/test_memory_limits.py \
fastvideo/tests/mlx/test_quant_capability.py \
fastvideo/tests/mlx/test_mlx_dit_parity.py \
fastvideo/tests/mlx/test_mlx_compile_parity.py \
fastvideo/tests/mlx/test_mlx_checkpoint.py \
fastvideo/tests/mlx/test_mlx_fastwan_benchmark.py \
fastvideo/tests/mlx/test_taehv_decode.py \
fastvideo/tests/mlx/test_frame_upsample.py \
fastvideo/tests/mlx/test_mlx_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_refine.py \
fastvideo/tests/mlx/test_mlx_prompt_to_video_decode.py \
fastvideo/tests/mlx/test_mlx_wan22_prompt_cache_fingerprint.py \
fastvideo/tests/mlx/test_wan22_sample.py \
fastvideo/tests/mlx/test_windowed_attention.py \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_download_unavailable_has_specific_error \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
fastvideo/tests/platforms/test_mps_vsa_error.py \
-q
# Same tests on MLX's CPU backend. Hosted macOS runners are scarce and
# slower to schedule; this Linux job gives fast PR signal on the identical
# graph (the parity tests were designed to be backend-agnostic), while the
# macOS job above stays the source of truth for Metal behavior.
mlx-smoke-linux-cpu:
if: github.event_name == 'workflow_dispatch' || github.event.pull_request.draft != true
runs-on: ubuntu-latest
timeout-minutes: 20
env:
FASTVIDEO_ATTENTION_BACKEND: TORCH_SDPA
TOKENIZERS_PARALLELISM: "false"
MASTER_ADDR: localhost
MASTER_PORT: "29513"
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
cache: pip
- uses: astral-sh/setup-uv@v3
- name: Install lightweight MLX smoke dependencies (CPU backend)
run: |
uv pip install --system \
--index-url https://download.pytorch.org/whl/cpu \
torch==2.11.0 torchvision torchaudio
uv pip install --system \
pytest numpy scipy pillow imageio einops cloudpickle filelock \
PyYAML diffusers huggingface_hub remote-pdb safetensors loguru "mlx[cpu]" \
"ftfy>=6.3.1" "opencv-python>=4.10.0.84" psutil "transformers>=5.0.0"
- name: Run MLX smoke tests (CPU backend)
run: |
python -m pytest \
fastvideo/tests/mlx/test_dmd_sampling.py \
fastvideo/tests/mlx/test_memory_limits.py \
fastvideo/tests/mlx/test_quant_capability.py \
fastvideo/tests/mlx/test_mlx_dit_parity.py \
fastvideo/tests/mlx/test_mlx_compile_parity.py \
fastvideo/tests/mlx/test_mlx_checkpoint.py \
fastvideo/tests/mlx/test_mlx_fastwan_benchmark.py \
fastvideo/tests/mlx/test_taehv_decode.py \
fastvideo/tests/mlx/test_frame_upsample.py \
fastvideo/tests/mlx/test_mlx_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_refine.py \
fastvideo/tests/mlx/test_mlx_prompt_to_video_decode.py \
fastvideo/tests/mlx/test_mlx_wan22_prompt_cache_fingerprint.py \
fastvideo/tests/mlx/test_wan22_sample.py \
fastvideo/tests/mlx/test_windowed_attention.py \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_download_unavailable_has_specific_error \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
fastvideo/tests/platforms/test_mps_vsa_error.py \
-q
+17 -4
View File
@@ -27,14 +27,25 @@ jobs:
ref: ${{ inputs.ref || '' }}
# For PR events, lint the PR head — but keep the hook definitions from
# the base branch so an untrusted PR cannot alter what gets executed.
- name: Save trusted hook config
# The gate scripts are saved too: the self-test step below executes them,
# so it must run the base-branch copies, not the PR head's.
- name: Save trusted hook config and gate scripts
if: github.event_name == 'pull_request_target'
run: cp .pre-commit-config.yaml "$RUNNER_TEMP/trusted-pre-commit-config.yaml"
- uses: actions/checkout@v4
run: |
cp .pre-commit-config.yaml "$RUNNER_TEMP/trusted-pre-commit-config.yaml"
cp -a .github/scripts "$RUNNER_TEMP/trusted-scripts"
echo "GATE_SCRIPTS_DIR=$RUNNER_TEMP/trusted-scripts" >> "$GITHUB_ENV"
# allow-unsafe-pr-checkout acknowledges checkout's pull_request_target
# guard: the head is data for the trusted hooks to lint; nothing from it
# is executed (config and gate scripts are pinned to the base branch
# above) and credentials are not persisted. SHA-pinned to v4.4.0 because
# actionlint's action schema does not know the new input yet.
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
if: github.event_name == 'pull_request_target'
with:
ref: ${{ github.event.pull_request.head.sha }}
persist-credentials: false
allow-unsafe-pr-checkout: true
- name: Restore trusted hook config
if: github.event_name == 'pull_request_target'
run: cp "$RUNNER_TEMP/trusted-pre-commit-config.yaml" .pre-commit-config.yaml
@@ -48,5 +59,7 @@ jobs:
with:
extra_args: --all-files --hook-stage manual
# After pre-commit so a self-test failure cannot mask lint failures.
# GATE_SCRIPTS_DIR points at the base-branch copy on fork PRs (set above);
# push / workflow_call runs use the checked-out tree directly.
- name: Full-suite gate self-test
run: bash .github/scripts/test_gate_full_suite.sh
run: bash "${GATE_SCRIPTS_DIR:-.github/scripts}/test_gate_full_suite.sh"
+2 -2
View File
@@ -129,7 +129,7 @@ jobs:
set -euo pipefail
TEST_NAME=$(echo "$COMMENT" | grep -oP '(?<=/test\s)\S+' | head -1 || true)
VALID="encoder vae transformer kernel unit dreamverse ssim training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval full fastcheck pre-commit"
VALID="encoder vae transformer kernel unit dreamverse ssim golden-gate training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval full fastcheck pre-commit"
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
echo "Unknown test: '$TEST_NAME'. Valid: $VALID"
exit 1
@@ -138,7 +138,7 @@ jobs:
declare -A MAP=(
[encoder]=encoder [vae]=vae [transformer]=transformer
[kernel]=kernel_tests [unit]=unit_test [dreamverse]=dreamverse_app
[ssim]=ssim [training]=training
[ssim]=ssim [golden-gate]=golden_gate [training]=training
[lora-inference]=inference_lora [lora-training]=training_lora
[lora-extraction]=lora_extraction
[distillation]=distillation_dmd [self-forcing]=self_forcing
+14 -7
View File
@@ -13,16 +13,23 @@ on:
required: false
default: false
type: boolean
# Auto-rebuild the CUDA images when their Dockerfile changes on main. The CUDA
# matrix is the only lane that builds from docker/Dockerfile, so a path-scoped
# push trigger is a sufficient change detector on its own -- no separate
# detect-changes/paths-filter job is needed now that there is a single
# in-scope Dockerfile. Dreamverse (apps/dreamverse/docker/Dockerfile) and the
# rocm Dockerfile stay manual-dispatch only.
# Auto-rebuild the CUDA images when a repository-controlled image input
# changes on main. This includes the trusted SM89 kernel artifact's source,
# metadata/key helper, ABI dependency metadata, and build orchestration.
# Dreamverse (apps/dreamverse/docker/Dockerfile) and the ROCm Dockerfile stay
# manual-dispatch only.
push:
branches: [main]
paths:
- '.dockerignore'
- '.github/workflows/_template-build-image.yml'
- '.github/workflows/infra-build-image.yml'
- '.gitmodules'
- 'docker/Dockerfile'
- 'docker/uv-excludes'
- 'fastvideo-kernel/**'
- 'fastvideo/tests/modal/kernel_build_cache.py'
- 'pyproject.toml'
permissions:
@@ -50,7 +57,7 @@ jobs:
# 2.8.3 comes from the architecture-specific prebuilt releases.
build-cuda-images:
# Runs on a manual dispatch when build_cuda_matrix is set, or automatically
# on a push that changed docker/Dockerfile (inputs are null on push). The
# on an in-scope main push (inputs are null on push). The
# repository guard keeps fork syncs from auto-building; manual dispatch
# still works in forks.
if: ${{ (github.event_name == 'push' && github.repository == 'hao-ai-lab/FastVideo') || github.event.inputs.build_cuda_matrix == 'true' }}
+2
View File
@@ -6,6 +6,7 @@ on:
paths:
- 'docs/**'
- 'examples/**'
- 'scripts/inference/**'
- 'mkdocs.yml'
- 'requirements-mkdocs.in'
- 'requirements-mkdocs.txt'
@@ -16,6 +17,7 @@ on:
paths:
- 'docs/**'
- 'examples/**'
- 'scripts/inference/**'
- 'mkdocs.yml'
- 'requirements-mkdocs.in'
- 'requirements-mkdocs.txt'
+4
View File
@@ -6,6 +6,7 @@ results/
wandb/
*.ipynb
*.jpg
!examples/datasets/lingbotworld2/image.jpg
*.safetensors
*.mp4
*.png
@@ -22,6 +23,7 @@ Miniconda3-latest-Linux-x86_64.sh
*validation/
data/
outputs/
outputs_audio/
outputs_video
checkpoints/
sbatch.sh
@@ -34,6 +36,7 @@ env
*.log
weights/
logs/
/Z-Image/
official_weights/
converted_weights/
@@ -73,6 +76,7 @@ docs/distillation/examples/
*.pkl
# Reference videos (negations must come after the catch-all on line below)
!fastvideo/tests/nightly/reference_video_*.mp4
# Static images
!docs/assets/images/**/*.png
+1
View File
@@ -22,6 +22,7 @@ repos:
hooks:
- id: yapf
args: [--in-place, --verbose]
language_version: python3.12
additional_dependencies: [toml] # TODO: Remove when yapf is upgraded
- repo: https://github.com/astral-sh/ruff-pre-commit
rev: v0.11.12
+8 -2
View File
@@ -9,7 +9,8 @@
**FastVideo is a unified post-training and real-time inference framework for accelerated video generation.**
## NEWS
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), check out the [Blog](https://haoailab.com/blogs/fastwan-qad/).
- `2026/08/19`: FastVideo now supports MLX on Apple Silicon with [FastMetal-QAD](https://huggingface.co/collections/FastVideo/fastmetal), a family of 1.3B, 5B, and 14B models optimized for Mac—follow the [Apple Silicon guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/) and read the [Blog](https://haoailab.com/blogs/fastmetal/).
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. See the [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), [Attn-QAT training guide](https://haoailab.com/FastVideo/training/attn_qat/), and [blog](https://haoailab.com/blogs/fastwan-qad/).
- `2026/03/17`: Release demo: Into the Dreamverse: Vibe Directing in FastVideo, check out the [Blog](https://haoailab.com/blogs/dreamverse/).
- `2026/03/13`: Release demo: Create a 5s 1080p Video in 4.5s with FastVideo on a Single GPU, check out the [Blog](https://haoailab.com/blogs/fastvideo_realtime_1080p/).
- `2025/11/19`: Release [CausalWan2.2 I2V A14B Preview](https://huggingface.co/FastVideo/CausalWan2.2-I2V-A14B-Preview-Diffusers) models, [Blog](https://hao-ai-lab.github.io/blogs/fastvideo_causalwan_preview/) and [Inference Code!](https://github.com/hao-ai-lab/FastVideo/blob/main/examples/inference/basic/basic_self_forcing_causal_wan2_2_i2v.py).
@@ -33,7 +34,7 @@ FastVideo has the following features:
- [Sparse distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/) to achieve >50x denoising speedup
- Scalable training with FSDP2, sequence parallelism, and selective activation checkpointing.
- Causal distillation through Self-Forcing
- See this [page](https://hao-ai-lab.github.io/FastVideo/training/overview/) for full list of supported models and recipes.
- See this [page](https://hao-ai-lab.github.io/FastVideo/training/overview/) for the supported training workflows, and the [support matrix](https://hao-ai-lab.github.io/FastVideo/inference/support_matrix/) for supported models.
- State-of-the-art performance optimizations for inference
- Sequence Parallelism for distributed inference
- Multiple state-of-the-art attention backends
@@ -62,6 +63,11 @@ UV_TORCH_BACKEND=cu126 uv pip install fastvideo
Use `UV_TORCH_BACKEND=cu130` on CUDA 13. Apple silicon users should follow the
[MPS installation guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
> **On an Apple Silicon Mac?** FastVideo runs FastWan text-to-video natively
> through an MLX runtime — a 5-second 480p clip generated locally, no cloud,
> no discrete GPU. Install with `uv pip install -e '.[mlx]'` and follow the
> [Apple Silicon guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
Please see our [docs](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/) for more detailed installation instructions.
> **On an NVIDIA DGX Spark (GB10 / ARM64 + CUDA 13)?** There's no prebuilt ARM wheel for the FastVideo CUDA kernel, so it's an editable from-source install (`UV_TORCH_BACKEND=cu130 uv pip install -e .`, which compiles that kernel for you) rather than `UV_TORCH_BACKEND=cu130 uv pip install fastvideo`. A compatible prebuilt ARM64 FlashAttention wheel is available separately. Follow the [DGX Spark install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/spark/).
@@ -3,7 +3,6 @@ from __future__ import annotations
import sys
from pathlib import Path
TESTS_DIR = Path(__file__).resolve().parent
DREAMVERSE_PACKAGE_DIR = TESTS_DIR.parent
DREAMVERSE_APP_DIR = DREAMVERSE_PACKAGE_DIR.parent
@@ -5,7 +5,6 @@ from pathlib import Path
import pytest
SERVER_DIR = Path(__file__).resolve().parents[1]
@@ -53,9 +52,7 @@ def test_config_defaults_to_cerebras_with_parallel_groq_fallback_stage(monkeypat
module = _load_config_module()
assert module.PROMPT_PROVIDER == "cerebras"
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (
("cerebras", "groq"),
)
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (("cerebras", "groq"), )
assert module.PROMPT_PROVIDER_PRIORITY == (
"cerebras",
"groq",
@@ -86,9 +83,7 @@ def test_config_ignores_legacy_groq_primary_override(monkeypatch):
module = _load_config_module()
assert module.PROMPT_PROVIDER == "cerebras"
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (
("cerebras", "groq"),
)
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (("cerebras", "groq"), )
assert module.PROMPT_PROVIDER_PRIORITY == (
"cerebras",
"groq",
@@ -106,24 +101,17 @@ def test_config_uses_local_overlay_paths_when_devtools_enabled(monkeypatch, tmp_
assert module.DEVTOOLS_ENABLED is True
assert module.FRONTEND_ROOT.as_posix().endswith("apps/dreamverse/web")
assert module.PROMPT_ENHANCE_SYSTEM_PROMPT_PATH.endswith(
"dreamverse/prompts.local/next_segment_system_prompt.md"
)
assert module.PROMPT_ENHANCE_SYSTEM_PROMPT_PATH.endswith("dreamverse/prompts.local/next_segment_system_prompt.md")
assert module.PROMPT_ENHANCE_SYSTEM_PROMPT_FALLBACK_PATH.endswith(
"dreamverse/prompts/next_segment_system_prompt.md"
)
"dreamverse/prompts/next_segment_system_prompt.md")
assert module.PROMPT_REWRITE_USER_SYSTEM_PROMPT_PATH.endswith(
"dreamverse/prompts.local/rewrite_user_system_prompt.md"
)
"dreamverse/prompts.local/rewrite_user_system_prompt.md")
assert module.PROMPT_REWRITE_USER_SYSTEM_PROMPT_FALLBACK_PATH.endswith(
"dreamverse/prompts/rewrite_user_system_prompt.md"
)
"dreamverse/prompts/rewrite_user_system_prompt.md")
assert module.CURATED_PRESETS_FILE_PATH.endswith(
"apps/dreamverse/web/prompts.local/selected_ltx2_continuation_story_presets.json"
)
"apps/dreamverse/web/prompts.local/selected_ltx2_continuation_story_presets.json")
assert module.CURATED_PRESETS_FALLBACK_FILE_PATH.endswith(
"apps/dreamverse/web/prompts/selected_ltx2_continuation_story_presets.json"
)
"apps/dreamverse/web/prompts/selected_ltx2_continuation_story_presets.json")
assert module.FRONTEND_STATIC_DIR_CANDIDATES[:2] == (
str(module.FRONTEND_ROOT / "out"),
str(module.FRONTEND_ROOT / "dist"),
@@ -9,6 +9,7 @@ from fastapi.testclient import TestClient
import fastvideo.entrypoints.streaming as streaming_entrypoints
import pytest
def _install_stack03_import_stubs(monkeypatch):
"""Keep entrypoint tests focused while later-stack runtime modules are absent."""
if not hasattr(streaming_entrypoints, "build_health_router"):
@@ -17,6 +18,7 @@ def _install_stack03_import_stubs(monkeypatch):
gpu_pool_stub = types.ModuleType("dreamverse.gpu_pool")
class GPUPool:
def __init__(self, _gpu_ids):
pass
@@ -49,6 +51,7 @@ def _install_stack03_import_stubs(monkeypatch):
controller_stub = types.ModuleType("dreamverse.session.controller")
class SessionController:
def __init__(self, **_kwargs):
pass
@@ -76,13 +79,11 @@ def _run_cli(module, monkeypatch, argv: list[str]) -> list[dict[str, object]]:
uvicorn_stub = types.ModuleType("uvicorn")
def run(app, host: str, port: int) -> None:
calls.append(
{
"app": app,
"host": host,
"port": port,
}
)
calls.append({
"app": app,
"host": host,
"port": port,
})
uvicorn_stub.run = run
monkeypatch.setitem(sys.modules, "uvicorn", uvicorn_stub)
@@ -99,13 +100,11 @@ def test_server_cli_defaults_to_local_web_port(monkeypatch):
server_main = _import_server_main(monkeypatch)
calls = _run_cli(server_main, monkeypatch, ["dreamverse-server"])
assert calls == [
{
"app": server_main.app,
"host": "0.0.0.0",
"port": 8009,
}
]
assert calls == [{
"app": server_main.app,
"host": "0.0.0.0",
"port": 8009,
}]
def test_server_cli_allows_explicit_host_and_port(monkeypatch):
@@ -116,13 +115,11 @@ def test_server_cli_allows_explicit_host_and_port(monkeypatch):
["dreamverse-server", "--host", "127.0.0.1", "--port", "8123"],
)
assert calls == [
{
"app": server_main.app,
"host": "127.0.0.1",
"port": 8123,
}
]
assert calls == [{
"app": server_main.app,
"host": "127.0.0.1",
"port": 8123,
}]
def test_server_does_not_expose_backend_source_as_static_assets(monkeypatch):
@@ -142,13 +139,11 @@ def test_mock_server_cli_defaults_to_local_web_port(monkeypatch):
["dreamverse-mock-server"],
)
assert calls == [
{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8009,
}
]
assert calls == [{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8009,
}]
def test_mock_server_cli_updates_latency(monkeypatch):
@@ -161,13 +156,11 @@ def test_mock_server_cli_updates_latency(monkeypatch):
["dreamverse-mock-server", "--latency", "321", "--port", "8111"],
)
assert calls == [
{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8111,
}
]
assert calls == [{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8111,
}]
assert mock_server.LATENCY_MS == 321
finally:
mock_server.LATENCY_MS = old_latency_ms
@@ -7,7 +7,6 @@ from types import SimpleNamespace
import pytest
import dreamverse.gpu_pool as gpu_pool
@@ -85,9 +84,7 @@ def test_send_command_raises_on_worker_death():
cmd_q = ctx.Queue()
resp_q = ctx.Queue()
proc = ctx.Process(
target=_child_consume_and_exit, args=(cmd_q, resp_q)
)
proc = ctx.Process(target=_child_consume_and_exit, args=(cmd_q, resp_q))
proc.start()
# Wait for the spawn child to fully boot. Allow generous time —
@@ -7,7 +7,7 @@ ALLOWED_PREFIXES = (
"fastvideo.entrypoints.video_generator",
"fastvideo.configs",
)
ALLOWED_EXACT = ("fastvideo",)
ALLOWED_EXACT = ("fastvideo", )
FORBIDDEN_PREFIXES = (
"fastvideo.pipelines",
"fastvideo.models",
@@ -38,19 +38,13 @@ def test_dreamverse_server_imports_only_public_fastvideo_surfaces() -> None:
except SyntaxError as task_exc:
raise AssertionError(f"Failed to parse {path}") from task_exc
for node in ast.walk(tree):
names = (
[a.name for a in node.names] if isinstance(node, ast.Import)
else [node.module] if isinstance(node, ast.ImportFrom) and node.module
else []
)
names = ([a.name for a in node.names] if isinstance(node, ast.Import) else
[node.module] if isinstance(node, ast.ImportFrom) and node.module else [])
for name in names:
if not name:
continue
rel_path = str(path.relative_to(root))
if (
name.startswith(FORBIDDEN_PREFIXES)
and (rel_path, name) not in ALLOWED_INTERNAL_IMPORTS
):
if (name.startswith(FORBIDDEN_PREFIXES) and (rel_path, name) not in ALLOWED_INTERNAL_IMPORTS):
bad.append((str(path.relative_to(root)), getattr(node, "lineno", 0), name))
assert bad == [], f"Forbidden internal imports: {bad}"
@@ -6,7 +6,6 @@ import os
from fastapi import WebSocketDisconnect
os.environ.setdefault("CEREBRAS_API_KEY", "dummy")
os.environ.setdefault("GROQ_API_KEY", "dummy")
@@ -14,6 +13,7 @@ import dreamverse.mock_server as mock_server
class _FakeWebSocket:
def __init__(self, messages: list[tuple[float, dict[str, object]]]):
self._messages = messages
self._index = 0
@@ -49,34 +49,34 @@ def test_mock_server_matches_current_single5s_protocol():
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 1
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "simple_prompt_1",
"curated_prompts": ["selected prompt"],
"single_clip_mode": True,
"enhancement_enabled": False,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.01,
{
"type": "simple_generate",
"preset_id": "simple_custom_prompt",
"prompt_id": "simple_custom_prompt",
"prompt": "custom prompt",
"enhancement_enabled": True,
"initial_image": None,
},
),
(0.20, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "simple_prompt_1",
"curated_prompts": ["selected prompt"],
"single_clip_mode": True,
"enhancement_enabled": False,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.01,
{
"type": "simple_generate",
"preset_id": "simple_custom_prompt",
"prompt_id": "simple_custom_prompt",
"prompt": "custom prompt",
"enhancement_enabled": True,
"initial_image": None,
},
),
(0.20, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
@@ -92,24 +92,14 @@ def test_mock_server_matches_current_single5s_protocol():
assert message_types.count("ltx2_stream_complete") == 2
assert "prompt_sources_blocked" not in message_types
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
gpu_assigned_event = next(
payload for payload in ws.sent_json if payload["type"] == "gpu_assigned"
)
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
gpu_assigned_event = next(payload for payload in ws.sent_json if payload["type"] == "gpu_assigned")
assert gpu_assigned_event["session_timeout"] == mock_server.SESSION_TIMEOUT_SECONDS
assert [payload["segment_idx"] for payload in segment_start_events] == [1, 1]
assert segment_start_events[0]["prompt"] == "selected prompt"
assert segment_start_events[1]["prompt"] == "custom prompt"
step_complete_events = [
payload
for payload in ws.sent_json
if payload["type"] == "step_complete"
]
step_complete_events = [payload for payload in ws.sent_json if payload["type"] == "step_complete"]
assert len(step_complete_events) == 2
assert step_complete_events[0]["latency_ms"] == {
"total": 121.0,
@@ -134,29 +124,29 @@ def test_mock_server_regular_cap_waits_for_rewrite_rollout():
mock_server.LATENCY_MS = 1
mock_server.GENERATION_SEGMENT_CAP = 1
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "start a new rollout",
},
),
(0.20, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "start a new rollout",
},
),
(0.20, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
@@ -166,11 +156,7 @@ def test_mock_server_regular_cap_waits_for_rewrite_rollout():
assert "generation_cap_reached" not in message_types
assert "prompt_sources_blocked" not in message_types
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
assert [payload["segment_idx"] for payload in segment_start_events] == [1, 1]
assert segment_start_events[0]["prompt"] == "segment one"
assert segment_start_events[1]["prompt"] == "segment one [start a new rollout]"
@@ -187,54 +173,40 @@ def test_mock_server_rewrite_during_active_segment_restarts_from_first_rewritten
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 100
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one", "segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "restart from rewrite",
},
),
(0.40, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one", "segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "restart from rewrite",
},
),
(0.40, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
assert [payload["prompt"] for payload in segment_start_events[:2]] == [
"segment one",
"segment one [restart from rewrite]",
]
assert all(
payload["prompt"] != "segment two"
for payload in segment_start_events[1:]
)
reset_events = [
payload
for payload in ws.sent_json
if payload.get("type") == "seed_prompts_reset_applied"
]
assert any(
payload.get("reason") == "rewrite_during_generation"
for payload in reset_events
)
assert all(payload["prompt"] != "segment two" for payload in segment_start_events[1:])
reset_events = [payload for payload in ws.sent_json if payload.get("type") == "seed_prompts_reset_applied"]
assert any(payload.get("reason") == "rewrite_during_generation" for payload in reset_events)
finally:
mock_server.MOCK_SEGMENT_BYTES = old_segment_bytes
mock_server.LATENCY_MS = old_latency_ms
@@ -247,24 +219,24 @@ def test_mock_server_supports_initial_custom_rollout_prompt():
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 1
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "custom_editable",
"preset_label": "Custom rollout",
"curated_prompts": [],
"initial_rollout_prompt": "A moonbase corridor thriller with flooding",
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.20, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "custom_editable",
"preset_label": "Custom rollout",
"curated_prompts": [],
"initial_rollout_prompt": "A moonbase corridor thriller with flooding",
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.20, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
@@ -276,15 +248,9 @@ def test_mock_server_supports_initial_custom_rollout_prompt():
assert "ltx2_stream_start" in message_types
assert "prompt_sources_blocked" not in message_types
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
assert segment_start_events
assert segment_start_events[0]["prompt"] == (
"A moonbase corridor thriller with flooding [segment 1]"
)
assert segment_start_events[0]["prompt"] == ("A moonbase corridor thriller with flooding [segment 1]")
finally:
mock_server.MOCK_SEGMENT_BYTES = old_segment_bytes
mock_server.LATENCY_MS = old_latency_ms
@@ -297,35 +263,37 @@ def test_mock_server_can_start_new_project_without_reconnecting():
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 40
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.02, {"type": "end_project_keep_session"}),
(
0.20,
{
"type": "project_init_v1",
"preset_id": "test_preset_2",
"preset_label": "Test Preset 2",
"curated_prompts": ["segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.40, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.02, {
"type": "end_project_keep_session"
}),
(
0.20,
{
"type": "project_init_v1",
"preset_id": "test_preset_2",
"preset_label": "Test Preset 2",
"curated_prompts": ["segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.40, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
@@ -336,16 +304,11 @@ def test_mock_server_can_start_new_project_without_reconnecting():
project_idle_index = message_types.index("project_idle")
stream_start_indexes = [
index for index, message_type in enumerate(message_types)
if message_type == "ltx2_stream_start"
index for index, message_type in enumerate(message_types) if message_type == "ltx2_stream_start"
]
assert stream_start_indexes[0] < project_idle_index < stream_start_indexes[1]
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
assert [payload["prompt"] for payload in segment_start_events[:2]] == [
"segment one",
"segment two",
@@ -6,7 +6,6 @@ import os
import re
import time
os.environ.setdefault("CEREBRAS_API_KEY", "dummy")
os.environ.setdefault("GROQ_API_KEY", "dummy")
@@ -22,6 +21,7 @@ from dreamverse.prompt_enhancer import (
class _FakeResponse:
def __init__(self, payload: dict):
self._payload = payload
@@ -30,6 +30,7 @@ class _FakeResponse:
class _FakeSyncCompletions:
def __init__(self, payload: dict):
self._payload = payload
@@ -38,6 +39,7 @@ class _FakeSyncCompletions:
class _FakeSyncClient:
def __init__(self, payload: dict):
self.chat = type(
"_FakeChat",
@@ -47,6 +49,7 @@ class _FakeSyncClient:
class _DelayedSyncCompletions:
def __init__(self, payload: dict, delay_s: float = 0.0, exc: Exception | None = None):
self._payload = payload
self._delay_s = delay_s
@@ -61,29 +64,26 @@ class _DelayedSyncCompletions:
class _DelayedSyncClient:
def __init__(self, payload: dict, delay_s: float = 0.0, exc: Exception | None = None):
self.chat = type(
"_FakeChat",
(),
{
"completions": _DelayedSyncCompletions(
payload,
delay_s=delay_s,
exc=exc,
)
},
{"completions": _DelayedSyncCompletions(
payload,
delay_s=delay_s,
exc=exc,
)},
)()
def _chat_payload_with_content(content: str) -> dict:
return {
"choices": [
{
"message": {
"content": content,
}
"choices": [{
"message": {
"content": content,
}
]
}]
}
@@ -172,6 +172,7 @@ def _build_staged_enhancer(
class _FakeOpenAIClient:
def __init__(self, **kwargs):
self.kwargs = kwargs
self.chat = type(
@@ -182,6 +183,7 @@ class _FakeOpenAIClient:
class _FakeCerebrasClient:
def __init__(self, **kwargs):
self.kwargs = kwargs
self.chat = type(
@@ -192,16 +194,12 @@ class _FakeCerebrasClient:
def test_parse_json_response_accepts_fenced_json_with_prose():
parsed = _parse_json_response(
"Here is the rewrite:\n```json\n{\"segment_prompts\":[\"A\",\"B\"]}\n```\nThanks."
)
parsed = _parse_json_response("Here is the rewrite:\n```json\n{\"segment_prompts\":[\"A\",\"B\"]}\n```\nThanks.")
assert parsed == {"segment_prompts": ["A", "B"]}
def test_parse_json_response_extracts_first_embedded_object():
parsed = _parse_json_response(
"Model output:\n{\"segment_prompts\":[\"A\",\"B\"]}\n(complete)"
)
parsed = _parse_json_response("Model output:\n{\"segment_prompts\":[\"A\",\"B\"]}\n(complete)")
assert parsed == {"segment_prompts": ["A", "B"]}
@@ -268,16 +266,12 @@ def test_build_client_supports_groq_provider(monkeypatch):
def test_rewrite_prompt_sequence_accepts_segment_prompts_output():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.rollout_id == "preset_a"
@@ -286,15 +280,12 @@ def test_rewrite_prompt_sequence_accepts_segment_prompts_output():
def test_rewrite_prompt_sequence_accepts_legacy_rewritten_prompts_output():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"rewritten_prompts":["A","B"]}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"rewritten_prompts":["A","B"]}'))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.rollout_id == "current_rollout"
@@ -303,19 +294,14 @@ def test_rewrite_prompt_sequence_accepts_legacy_rewritten_prompts_output():
def test_rewrite_prompt_sequence_accepts_segment_dicts_without_top_level_rollout_metadata():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"segments":[{"prompt":"A"},{"text":"B"}]}'
)
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"segments":[{"prompt":"A"},{"text":"B"}]}'))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
preset_id="preset_a",
preset_label="Preset A",
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.rollout_id == "preset_a"
@@ -329,14 +315,12 @@ def test_rewrite_prompt_sequence_accepts_numbered_prose_output():
"The user is asking for a cinematic rewrite.\n\n"
"1. A dog bounds across the moon's dusty surface, kicking up silver regolith as it chases a rabbit beneath the black sky.\n"
"2. The rabbit darts around a crater rim while the dog lunges after it, Earth glowing blue in the distance.\n"
)
)
))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.rollout_id == "current_rollout"
@@ -355,12 +339,10 @@ def test_enhance_prompt_prefers_cerebras_before_groq_fallback():
groq_delay_s=0.01,
)
result = asyncio.run(
enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
)
)
result = asyncio.run(enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
))
assert result.fallback_used is False
assert result.error is None
@@ -381,12 +363,10 @@ def test_enhance_prompt_uses_groq_when_cerebras_fails():
groq_delay_s=0.01,
)
result = asyncio.run(
enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
)
)
result = asyncio.run(enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
))
assert result.fallback_used is False
assert result.error is None
@@ -408,13 +388,11 @@ def test_enhance_prompt_can_use_groq_when_cerebras_times_out():
enhancer.http_timeout_ms = 50
enhancer.default_timeout_ms = 50
result = asyncio.run(
enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
timeout_ms=50,
)
)
result = asyncio.run(enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
timeout_ms=50,
))
assert result.fallback_used is False
assert result.error is None
@@ -434,12 +412,10 @@ def test_enhance_prompt_can_use_cerebras_when_it_returns_first():
groq_delay_s=0.08,
)
result = asyncio.run(
enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
)
)
result = asyncio.run(enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
))
assert result.fallback_used is False
assert result.error is None
@@ -453,15 +429,12 @@ def test_enhance_prompt_can_use_cerebras_when_it_returns_first():
def test_rewrite_prompt_sequence_keeps_raw_output_on_parse_error():
enhancer = _build_test_enhancer(
_chat_payload_with_content("I cannot comply with JSON right now.")
)
enhancer = _build_test_enhancer(_chat_payload_with_content("I cannot comply with JSON right now."))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is True
assert "No JSON object found in assistant response." in (result.error or "")
assert result.raw_response_text == "I cannot comply with JSON right now."
@@ -473,9 +446,7 @@ def test_rewrite_prompt_sequence_keeps_raw_output_on_parse_error():
def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}'
)
)
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}'))
captured = {
"body": None,
"timeout_seconds": None,
@@ -486,8 +457,7 @@ def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
captured["timeout_seconds"] = timeout_seconds
return (
_chat_payload_with_content(
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}'
),
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}'),
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}',
)
@@ -502,8 +472,7 @@ def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
rewrite_model="gpt-test",
rewrite_temperature=0.2,
timeout_ms=800,
)
)
))
assert result.fallback_used is False
assert captured["body"]["messages"][0] == {
@@ -512,12 +481,12 @@ def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
}
assert captured["body"]["messages"][1]["role"] == "user"
assert prompt_enhancer_module.json.loads(captured["body"]["messages"][1]["content"]) == {
"mode": "edit_existing_rollout",
"request": (
"Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical."
),
"user_instruction": "make it cinematic",
"mode":
"edit_existing_rollout",
"request": ("Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical."),
"user_instruction":
"make it cinematic",
"current_rollout": {
"id": "preset_a",
"label": "Preset A",
@@ -528,11 +497,8 @@ def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
def test_rewrite_prompt_sequence_supports_new_rollout_mode():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}'
)
)
_chat_payload_with_content('{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}'))
captured = {
"body": None,
}
@@ -541,10 +507,8 @@ def test_rewrite_prompt_sequence_supports_new_rollout_mode():
del timeout_seconds
captured["body"] = body
return (
_chat_payload_with_content(
'{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}'
),
_chat_payload_with_content('{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}'),
'{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}',
)
@@ -560,30 +524,29 @@ def test_rewrite_prompt_sequence_supports_new_rollout_mode():
rewrite_model="gpt-test",
rewrite_temperature=0.2,
timeout_ms=800,
)
)
))
assert result.fallback_used is False
assert result.prompts == ["A", "B", "C", "D", "E", "F"]
assert prompt_enhancer_module.json.loads(captured["body"]["messages"][1]["content"]) == {
"mode": "new_rollout",
"request": (
"Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical."
),
"user_instruction": "A moonbase corridor thriller with flooding and red alarms",
"desired_segment_count": 6,
"rollout_id_hint": "custom_editable",
"rollout_label_hint": "Custom rollout",
"mode":
"new_rollout",
"request": ("Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical."),
"user_instruction":
"A moonbase corridor thriller with flooding and red alarms",
"desired_segment_count":
6,
"rollout_id_hint":
"custom_editable",
"rollout_label_hint":
"Custom rollout",
}
def test_rewrite_prompt_sequence_uses_session_override_system_prompt():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.rewrite_all_system_prompt = "shared system prompt"
captured = {
"body": None,
@@ -593,9 +556,7 @@ def test_rewrite_prompt_sequence_uses_session_override_system_prompt():
del timeout_seconds
captured["body"] = body
return (
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
),
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'),
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}',
)
@@ -609,8 +570,7 @@ def test_rewrite_prompt_sequence_uses_session_override_system_prompt():
rewrite_instruction="make it cinematic",
rewrite_model="gpt-test",
system_prompt_override="session specific system prompt",
)
)
))
assert result.fallback_used is False
assert captured["body"]["messages"][0] == {
@@ -621,10 +581,7 @@ def test_rewrite_prompt_sequence_uses_session_override_system_prompt():
def test_resolve_rewrite_new_rollout_system_prompt_uses_dedicated_prompt():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.rewrite_all_system_prompt = "shared rewrite system prompt"
enhancer.rewrite_user_system_prompt = "new rollout rewrite system prompt"
@@ -635,24 +592,17 @@ def test_resolve_rewrite_new_rollout_system_prompt_uses_dedicated_prompt():
def test_resolve_rewrite_new_rollout_system_prompt_prefers_override():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.rewrite_all_system_prompt = "shared rewrite system prompt"
enhancer.rewrite_user_system_prompt = "new rollout rewrite system prompt"
resolved = enhancer.resolve_rewrite_new_rollout_system_prompt(
"session specific system prompt"
)
resolved = enhancer.resolve_rewrite_new_rollout_system_prompt("session specific system prompt")
assert resolved == "session specific system prompt"
def test_generate_auto_prompt_uses_selected_model():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"next_prompt":"Auto next"}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"next_prompt":"Auto next"}'))
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test", "gpt-alt"]
enhancer.rewrite_default_model = "gpt-test"
@@ -680,8 +630,7 @@ def test_generate_auto_prompt_uses_selected_model():
next_segment_idx=2,
model="gpt-alt",
timeout_ms=800,
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.prompt == "Auto next"
@@ -690,9 +639,7 @@ def test_generate_auto_prompt_uses_selected_model():
def test_enhance_prompt_uses_selected_model():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"next_prompt":"Enhanced next"}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"next_prompt":"Enhanced next"}'))
enhancer.enhance_system_prompt = "enhance system prompt"
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test", "gpt-alt"]
@@ -722,8 +669,7 @@ def test_enhance_prompt_uses_selected_model():
next_segment_idx=2,
model="gpt-alt",
timeout_ms=800,
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.prompt == "Enhanced next"
@@ -732,9 +678,7 @@ def test_enhance_prompt_uses_selected_model():
def test_enhance_prompt_single_clip_uses_auto_extension_prompt_and_prompt_field():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"prompt":"Extended single clip"}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"prompt":"Extended single clip"}'))
enhancer.enhance_system_prompt = "enhance system prompt"
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test", "gpt-alt"]
@@ -764,14 +708,12 @@ def test_enhance_prompt_single_clip_uses_auto_extension_prompt_and_prompt_field(
enhancer._request_content = _fake_request_content # type: ignore[attr-defined]
result = asyncio.run(
enhancer.enhance_prompt(
"short 5s idea",
mode="single_clip",
model="gpt-alt",
timeout_ms=800,
)
)
result = asyncio.run(enhancer.enhance_prompt(
"short 5s idea",
mode="single_clip",
model="gpt-alt",
timeout_ms=800,
))
assert result.fallback_used is False
assert result.error is None
assert result.prompt == "Extended single clip"
@@ -784,17 +726,15 @@ def test_enhance_prompt_single_clip_uses_auto_extension_prompt_and_prompt_field(
"single 5-second LTX-2.3 video clip. Respond with "
'valid JSON only as {"prompt": "..."}.' # noqa: E501
),
"user_prompt": "short 5s idea",
"user_prompt":
"short 5s idea",
}
def test_enhance_prompt_single_clip_rejects_plain_text_response():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
"Medium shot of a woman by a rainy cafe window as she lifts her "
"phone, exhales softly, and the camera makes a slow push in."
)
)
_chat_payload_with_content("Medium shot of a woman by a rainy cafe window as she lifts her "
"phone, exhales softly, and the camera makes a slow push in."))
enhancer.auto_system_prompt = "auto system prompt"
result = asyncio.run(
@@ -803,17 +743,14 @@ def test_enhance_prompt_single_clip_rejects_plain_text_response():
mode="single_clip",
model="gpt-test",
timeout_ms=800,
)
)
))
assert result.fallback_used is True
assert "No JSON object found in assistant response." in result.error
assert result.prompt == ""
def test_enhance_prompt_single_clip_rejects_segment_prompts_json():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"segment_prompts":["A","B"]}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"segment_prompts":["A","B"]}'))
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test"]
enhancer.rewrite_default_model = "gpt-test"
@@ -824,17 +761,14 @@ def test_enhance_prompt_single_clip_rejects_segment_prompts_json():
mode="single_clip",
model="gpt-test",
timeout_ms=800,
)
)
))
assert result.fallback_used is True
assert result.prompt == ""
assert "Missing prompt string." in (result.error or "")
def test_enhance_prompt_requires_json_and_does_not_fallback_to_raw_text():
enhancer = _build_test_enhancer(
_chat_payload_with_content("A cinematic continuation with slow dolly movement.")
)
enhancer = _build_test_enhancer(_chat_payload_with_content("A cinematic continuation with slow dolly movement."))
enhancer.enhance_system_prompt = "enhance system prompt"
enhancer.rewrite_model_options = ["gpt-test"]
enhancer.rewrite_default_model = "gpt-test"
@@ -846,17 +780,14 @@ def test_enhance_prompt_requires_json_and_does_not_fallback_to_raw_text():
next_segment_idx=2,
model="gpt-test",
timeout_ms=800,
)
)
))
assert result.fallback_used is True
assert result.prompt == ""
assert "No JSON object found in assistant response." in (result.error or "")
def test_generate_auto_prompt_requires_json_and_does_not_fallback_to_raw_text():
enhancer = _build_test_enhancer(
_chat_payload_with_content("A calm, grounded continuation with subtle motion.")
)
enhancer = _build_test_enhancer(_chat_payload_with_content("A calm, grounded continuation with subtle motion."))
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test"]
enhancer.rewrite_default_model = "gpt-test"
@@ -867,34 +798,30 @@ def test_generate_auto_prompt_requires_json_and_does_not_fallback_to_raw_text():
next_segment_idx=2,
model="gpt-test",
timeout_ms=800,
)
)
))
assert result.fallback_used is True
assert result.prompt == ""
assert "No JSON object found in assistant response." in (result.error or "")
def test_rewrite_prompt_sequence_includes_raw_json_when_content_empty():
enhancer = _build_test_enhancer(
{
"choices": [
{
"finish_reason": "length",
"message": {
"content": [],
"refusal": None,
},
}
],
"usage": {"completion_tokens": 0},
}
)
enhancer = _build_test_enhancer({
"choices": [{
"finish_reason": "length",
"message": {
"content": [],
"refusal": None,
},
}],
"usage": {
"completion_tokens": 0
},
})
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is True
assert "No rewrite segment prompts found in assistant response." in (result.error or "")
assert isinstance(result.raw_response_text, str)
@@ -903,10 +830,7 @@ def test_rewrite_prompt_sequence_includes_raw_json_when_content_empty():
def test_get_rewrite_model_config_returns_fixed_defaults():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.rewrite_default_model = "gpt-oss-120b"
enhancer.rewrite_model_options = ["gpt-oss-120b"]
@@ -918,10 +842,7 @@ def test_get_rewrite_model_config_returns_fixed_defaults():
def test_get_prompt_config_includes_auto_extension_prompt():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.enhance_system_prompt_path = "/tmp/next.md"
enhancer.auto_system_prompt_path = "/tmp/auto.md"
enhancer.rewrite_all_system_prompt_path = "/tmp/rewrite.md"
@@ -948,19 +869,14 @@ def test_get_prompt_config_includes_auto_extension_prompt():
def test_get_prompt_config_reports_loaded_fallback_prompt_path(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
rewrite_fallback_path = tmp_path / "rewrite_window_system_prompt.md"
rewrite_fallback_path.write_text("rewrite prompt\n", encoding="utf-8")
next_path = tmp_path / "next.md"
next_path.write_text("next prompt\n", encoding="utf-8")
auto_path = tmp_path / "auto.md"
auto_path.write_text("auto prompt\n", encoding="utf-8")
enhancer.rewrite_all_system_prompt_path = str(
tmp_path / "prompts.local" / "rewrite_window_system_prompt.md"
)
enhancer.rewrite_all_system_prompt_path = str(tmp_path / "prompts.local" / "rewrite_window_system_prompt.md")
enhancer.rewrite_all_system_prompt_fallback_path = str(rewrite_fallback_path)
enhancer.enhance_system_prompt_path = str(next_path)
enhancer.auto_system_prompt_path = str(auto_path)
@@ -973,14 +889,9 @@ def test_get_prompt_config_reports_loaded_fallback_prompt_path(tmp_path):
assert config["rewrite_window_system_prompt_path"] == str(rewrite_fallback_path)
def test_reload_system_prompts_falls_back_to_rewrite_window_when_user_prompt_empty(
tmp_path,
):
def test_reload_system_prompts_falls_back_to_rewrite_window_when_user_prompt_empty(tmp_path, ):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite_window_system_prompt.md"
@@ -1007,10 +918,7 @@ def test_reload_system_prompts_falls_back_to_rewrite_window_when_user_prompt_emp
def test_save_prompt_config_updates_auto_extension_prompt(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite.md"
@@ -1021,9 +929,7 @@ def test_save_prompt_config_updates_auto_extension_prompt(tmp_path):
enhancer.auto_system_prompt_path = str(auto_path)
enhancer.rewrite_all_system_prompt_path = str(rewrite_path)
config = enhancer.save_prompt_config(
auto_extension_system_prompt="auto updated",
)
config = enhancer.save_prompt_config(auto_extension_system_prompt="auto updated", )
assert auto_path.read_text(encoding="utf-8").strip() == "auto updated"
assert config["auto_extension_system_prompt"] == "auto updated"
@@ -1031,10 +937,7 @@ def test_save_prompt_config_updates_auto_extension_prompt(tmp_path):
def test_save_prompt_config_updates_rewrite_user_prompt(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite.md"
@@ -1052,9 +955,7 @@ def test_save_prompt_config_updates_rewrite_user_prompt(tmp_path):
enhancer.rewrite_all_system_prompt_fallback_path = None
enhancer.rewrite_user_system_prompt_fallback_path = None
config = enhancer.save_prompt_config(
rewrite_user_system_prompt="rewrite user updated",
)
config = enhancer.save_prompt_config(rewrite_user_system_prompt="rewrite user updated", )
assert rewrite_user_path.read_text(encoding="utf-8").strip() == "rewrite user updated"
assert config["rewrite_user_system_prompt"] == "rewrite user updated"
@@ -1062,10 +963,7 @@ def test_save_prompt_config_updates_rewrite_user_prompt(tmp_path):
def test_save_prompt_config_updates_rewrite_model(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite.md"
@@ -1081,9 +979,7 @@ def test_save_prompt_config_updates_rewrite_model(tmp_path):
enhancer.rewrite_default_model = "gpt-test"
enhancer.rewrite_model_options = ["gpt-test", "gpt-alt"]
config = enhancer.save_prompt_config(
rewrite_model="gpt-alt",
)
config = enhancer.save_prompt_config(rewrite_model="gpt-alt", )
assert enhancer.rewrite_default_model == "gpt-alt"
assert config["rewrite_model"] == "gpt-alt"
@@ -1092,10 +988,7 @@ def test_save_prompt_config_updates_rewrite_model(tmp_path):
def test_save_prompt_config_updates_rewrite_temperature(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite.md"
@@ -1109,9 +1002,7 @@ def test_save_prompt_config_updates_rewrite_temperature(tmp_path):
enhancer.auto_system_prompt_fallback_path = None
enhancer.rewrite_all_system_prompt_fallback_path = None
config = enhancer.save_prompt_config(
rewrite_temperature=1.3,
)
config = enhancer.save_prompt_config(rewrite_temperature=1.3, )
assert enhancer.rewrite_default_temperature == 1.3
assert config["rewrite_temperature"] == 1.3
@@ -1119,10 +1010,7 @@ def test_save_prompt_config_updates_rewrite_temperature(tmp_path):
def test_save_prompt_config_creates_versioned_backup_for_existing_prompt(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite_window_system_prompt.md"
@@ -1136,13 +1024,9 @@ def test_save_prompt_config_creates_versioned_backup_for_existing_prompt(tmp_pat
enhancer.auto_system_prompt_fallback_path = None
enhancer.rewrite_all_system_prompt_fallback_path = None
enhancer.save_prompt_config(
rewrite_window_system_prompt="rewrite updated",
)
enhancer.save_prompt_config(rewrite_window_system_prompt="rewrite updated", )
backup_paths = sorted(
tmp_path.glob("rewrite_window_system_prompt.*.bak.md")
)
backup_paths = sorted(tmp_path.glob("rewrite_window_system_prompt.*.bak.md"))
assert rewrite_path.read_text(encoding="utf-8").strip() == "rewrite updated"
assert len(backup_paths) == 1
@@ -27,13 +27,8 @@ try:
except ModuleNotFoundError:
websockets = None # type: ignore[assignment]
DEFAULT_PRESET_FILE = (
Path(__file__).resolve().parents[2]
/ "web"
/ "prompts"
/ "selected_ltx2_continuation_story_presets.json"
)
DEFAULT_PRESET_FILE = (Path(__file__).resolve().parents[2] / "web" / "prompts" /
"selected_ltx2_continuation_story_presets.json")
def utc_now_iso() -> str:
@@ -65,10 +60,7 @@ def safe_percentile(values: list[float], percentile: float) -> float | None:
if lower == upper:
return sorted_values[lower]
fraction = rank - lower
return (
sorted_values[lower]
+ (sorted_values[upper] - sorted_values[lower]) * fraction
)
return (sorted_values[lower] + (sorted_values[upper] - sorted_values[lower]) * fraction)
def summarize_series(values: list[float]) -> dict[str, float | int | None]:
@@ -145,24 +137,16 @@ def load_curated_prompts(
selected_id = str(selected.get("id", "")).strip() or "unknown_preset"
raw_prompts = selected.get("segment_prompts", [])
if not isinstance(raw_prompts, list):
raise ValueError(
f"Preset {selected_id} has invalid segment_prompts (must be list)."
)
raise ValueError(f"Preset {selected_id} has invalid segment_prompts (must be list).")
prompts = [
str(prompt).strip()
for prompt in raw_prompts
if isinstance(prompt, str) and str(prompt).strip()
]
prompts = [str(prompt).strip() for prompt in raw_prompts if isinstance(prompt, str) and str(prompt).strip()]
if not prompts:
raise ValueError(f"Preset {selected_id} has no non-empty prompts.")
limited = prompts[:curated_limit]
if not limited:
raise ValueError(
f"curated_limit={curated_limit} produced no prompts for preset "
f"{selected_id}."
)
raise ValueError(f"curated_limit={curated_limit} produced no prompts for preset "
f"{selected_id}.")
return selected_id, limited, len(prompts)
@@ -224,11 +208,11 @@ async def run_single_session(
try:
async with websockets.connect(
url,
max_size=None,
ping_interval=None,
open_timeout=connect_timeout_s,
close_timeout=2.0,
url,
max_size=None,
ping_interval=None,
open_timeout=connect_timeout_s,
close_timeout=2.0,
) as ws:
connect_finish_monotonic = time.monotonic()
session_data["connect_finish_ts_utc"] = utc_now_iso()
@@ -249,9 +233,7 @@ async def run_single_session(
timeout_remaining = session_timeout_s - elapsed_s
if timeout_remaining <= 0:
session_data["status"] = "timeout"
session_data["error"] = (
f"Session timed out after {session_timeout_s:.1f}s."
)
session_data["error"] = (f"Session timed out after {session_timeout_s:.1f}s.")
break
recv_start_epoch = time.time()
@@ -265,9 +247,7 @@ async def run_single_session(
)
except asyncio.TimeoutError:
session_data["status"] = "timeout"
session_data["error"] = (
"Timed out waiting for websocket message."
)
session_data["error"] = ("Timed out waiting for websocket message.")
break
except Exception as exc:
session_data["status"] = "failed"
@@ -288,20 +268,16 @@ async def run_single_session(
chunk_gap_ms: float | None = None
if last_chunk_finish_monotonic is not None:
chunk_gap_ms = (
recv_finish_monotonic - last_chunk_finish_monotonic
) * 1000.0
chunk_gap_ms = (recv_finish_monotonic - last_chunk_finish_monotonic) * 1000.0
session_data["chunks"].append(
{
"segment_idx": current_segment_idx,
"chunk_idx": session_data["total_chunks"],
"size_bytes": len(message),
"chunk_start_ts_utc": recv_start_iso,
"chunk_finish_ts_utc": recv_finish_iso,
"chunk_gap_ms": chunk_gap_ms,
}
)
session_data["chunks"].append({
"segment_idx": current_segment_idx,
"chunk_idx": session_data["total_chunks"],
"size_bytes": len(message),
"chunk_start_ts_utc": recv_start_iso,
"chunk_finish_ts_utc": recv_finish_iso,
"chunk_gap_ms": chunk_gap_ms,
})
last_chunk_finish_monotonic = recv_finish_monotonic
last_chunk_finish_epoch = recv_finish_epoch
session_data["last_chunk_finish_ts_utc"] = recv_finish_iso
@@ -321,9 +297,7 @@ async def run_single_session(
if msg_type == "gpu_assigned":
session_data["gpu_assigned_ts_utc"] = recv_finish_iso
if connect_finish_monotonic is not None:
session_data["queue_wait_ms"] = (
recv_finish_monotonic - connect_finish_monotonic
) * 1000.0
session_data["queue_wait_ms"] = (recv_finish_monotonic - connect_finish_monotonic) * 1000.0
elif msg_type == "ltx2_stream_start":
if initial_total_segments is None:
parsed_total = parse_int(data.get("total_segments"))
@@ -338,20 +312,13 @@ async def run_single_session(
session_data["media_segments_completed"] += 1
if first_media_segment_complete_epoch is None:
first_media_segment_complete_epoch = recv_finish_epoch
session_data[
"first_media_segment_complete_ts_utc"
] = recv_finish_iso
session_data["first_media_segment_complete_ts_utc"] = recv_finish_iso
elif msg_type == "ltx2_segment_complete":
session_data["segments_completed"] += 1
seg_idx = parse_int(data.get("segment_idx"))
if (
initial_total_segments is not None
and seg_idx is not None
and seg_idx >= initial_total_segments
):
session_data[
"target_segment_complete_ts_utc"
] = recv_finish_iso
if (initial_total_segments is not None and seg_idx is not None
and seg_idx >= initial_total_segments):
session_data["target_segment_complete_ts_utc"] = recv_finish_iso
await asyncio.sleep(post_complete_wait_s)
session_data["leave_sent_ts_utc"] = utc_now_iso()
try:
@@ -362,15 +329,11 @@ async def run_single_session(
break
elif msg_type == "session_timeout":
session_data["status"] = "timeout"
session_data["error"] = str(
data.get("message") or "Backend session timeout"
)
session_data["error"] = str(data.get("message") or "Backend session timeout")
break
elif msg_type == "error":
session_data["status"] = "failed"
session_data["error"] = str(
data.get("message") or "Backend error message"
)
session_data["error"] = str(data.get("message") or "Backend error message")
break
if session_data["status"] == "failed" and session_data["error"] is None:
@@ -379,29 +342,18 @@ async def run_single_session(
session_data["status"] = "failed"
session_data["error"] = f"WebSocket connect/run failed: {exc}"
if (
first_chunk_finish_epoch is not None
and last_chunk_finish_epoch is not None
and session_data["total_chunk_bytes"] > 0
):
if (first_chunk_finish_epoch is not None and last_chunk_finish_epoch is not None
and session_data["total_chunk_bytes"] > 0):
duration_s = last_chunk_finish_epoch - first_chunk_finish_epoch
if duration_s > 0:
session_data["session_goodput_mbps"] = (
session_data["total_chunk_bytes"] * 8.0 / duration_s / 1_000_000.0
)
session_data["session_goodput_mbps"] = (session_data["total_chunk_bytes"] * 8.0 / duration_s / 1_000_000.0)
if (
first_chunk_finish_epoch is not None
and first_media_segment_complete_epoch is not None
):
session_data["first_chunk_before_first_media_complete"] = (
first_chunk_finish_epoch < first_media_segment_complete_epoch
)
if (first_chunk_finish_epoch is not None and first_media_segment_complete_epoch is not None):
session_data["first_chunk_before_first_media_complete"] = (first_chunk_finish_epoch
< first_media_segment_complete_epoch)
session_data["close_ts_utc"] = utc_now_iso()
session_data["duration_ms"] = (
time.monotonic() - session_start_monotonic
) * 1000.0
session_data["duration_ms"] = (time.monotonic() - session_start_monotonic) * 1000.0
return session_data
@@ -412,14 +364,11 @@ async def run_worker_sessions(
config: dict[str, Any],
) -> list[dict[str, Any]]:
tasks = [
asyncio.create_task(
run_single_session(
worker_id=worker_id,
worker_session_idx=idx,
config=config,
)
)
for idx in range(session_count)
asyncio.create_task(run_single_session(
worker_id=worker_id,
worker_session_idx=idx,
config=config,
)) for idx in range(session_count)
]
if not tasks:
return []
@@ -437,29 +386,23 @@ def worker_entry(
try:
ready_queue.put({"worker_id": worker_id, "status": "ready"})
start_event.wait()
sessions = asyncio.run(
run_worker_sessions(
worker_id=worker_id,
session_count=session_count,
config=config,
)
)
result_queue.put(
{
"worker_id": worker_id,
"status": "ok",
"sessions": sessions,
}
)
sessions = asyncio.run(run_worker_sessions(
worker_id=worker_id,
session_count=session_count,
config=config,
))
result_queue.put({
"worker_id": worker_id,
"status": "ok",
"sessions": sessions,
})
except Exception as exc:
result_queue.put(
{
"worker_id": worker_id,
"status": "error",
"error": str(exc),
"traceback": traceback.format_exc(),
}
)
result_queue.put({
"worker_id": worker_id,
"status": "error",
"error": str(exc),
"traceback": traceback.format_exc(),
})
def build_summary(
@@ -517,33 +460,22 @@ def build_summary(
if len(all_chunk_finish_epochs) >= 2 and total_chunk_bytes > 0:
duration_s = max(all_chunk_finish_epochs) - min(all_chunk_finish_epochs)
if duration_s > 0:
global_goodput_mbps = (
total_chunk_bytes * 8.0 / duration_s / 1_000_000.0
)
global_goodput_mbps = (total_chunk_bytes * 8.0 / duration_s / 1_000_000.0)
bucket_throughputs_mbps = [
(bytes_count * 8.0) / 1_000_000.0
for _, bytes_count in sorted(bucket_bytes.items())
]
bucket_throughputs_mbps = [(bytes_count * 8.0) / 1_000_000.0 for _, bytes_count in sorted(bucket_bytes.items())]
bucket_stats = summarize_series(bucket_throughputs_mbps)
chunk_gap_threshold_breaches = [
value for value in chunk_gaps if value >= chunk_gap_threshold_ms
]
chunk_gap_threshold_breaches = [value for value in chunk_gaps if value >= chunk_gap_threshold_ms]
non_success = len(sessions) - status_counts.get("success", 0)
fail_reasons: list[str] = []
if non_success > 0:
fail_reasons.append(
f"{non_success} session(s) did not complete successfully."
)
fail_reasons.append(f"{non_success} session(s) did not complete successfully.")
if not chunk_gaps:
fail_reasons.append("No chunk gap data collected.")
if chunk_gap_threshold_breaches:
fail_reasons.append(
f"{len(chunk_gap_threshold_breaches)} chunk gap(s) were >= "
f"{chunk_gap_threshold_ms:.0f}ms."
)
fail_reasons.append(f"{len(chunk_gap_threshold_breaches)} chunk gap(s) were >= "
f"{chunk_gap_threshold_ms:.0f}ms.")
passed = len(fail_reasons) == 0
progressive_ratio = None
@@ -554,20 +486,18 @@ def build_summary(
"passed": passed,
"fail_reasons": fail_reasons,
"sessions": {
"total": len(sessions),
"success": status_counts.get("success", 0),
"failed": status_counts.get("failed", 0),
"timeout": status_counts.get("timeout", 0),
"protocol_error": status_counts.get("protocol_error", 0),
"other": (
len(sessions)
- (
status_counts.get("success", 0)
+ status_counts.get("failed", 0)
+ status_counts.get("timeout", 0)
+ status_counts.get("protocol_error", 0)
)
),
"total":
len(sessions),
"success":
status_counts.get("success", 0),
"failed":
status_counts.get("failed", 0),
"timeout":
status_counts.get("timeout", 0),
"protocol_error":
status_counts.get("protocol_error", 0),
"other": (len(sessions) - (status_counts.get("success", 0) + status_counts.get("failed", 0) +
status_counts.get("timeout", 0) + status_counts.get("protocol_error", 0))),
},
"chunk_gap_ms": {
**chunk_gap_stats,
@@ -606,51 +536,39 @@ def print_summary(
bucket_bw = bandwidth["bucketed_1s"]
print("=== LTX2 Realtime Stress Test Summary ===")
print(
"Run: "
f"url={run_info['url']} clients={run_info['clients']} "
f"processes={run_info['processes']} "
f"preset={run_info['preset_id']} "
f"curated_limit={run_info['curated_limit']}"
)
print(
"Sessions: "
f"total={sessions['total']} success={sessions['success']} "
f"failed={sessions['failed']} timeout={sessions['timeout']} "
f"protocol_error={sessions['protocol_error']}"
)
print(
"Chunk gap ms: "
f"min={format_num(chunk_gap['min'])} "
f"p50={format_num(chunk_gap['p50'])} "
f"p95={format_num(chunk_gap['p95'])} "
f"p99={format_num(chunk_gap['p99'])} "
f"max={format_num(chunk_gap['max'])} "
f"threshold={format_num(chunk_gap['threshold_ms'])} "
f"breaches={chunk_gap['breach_count']}"
)
print(
"Queue wait ms: "
f"min={format_num(queue_wait['min'])} "
f"p50={format_num(queue_wait['p50'])} "
f"p95={format_num(queue_wait['p95'])} "
f"max={format_num(queue_wait['max'])}"
)
print("Run: "
f"url={run_info['url']} clients={run_info['clients']} "
f"processes={run_info['processes']} "
f"preset={run_info['preset_id']} "
f"curated_limit={run_info['curated_limit']}")
print("Sessions: "
f"total={sessions['total']} success={sessions['success']} "
f"failed={sessions['failed']} timeout={sessions['timeout']} "
f"protocol_error={sessions['protocol_error']}")
print("Chunk gap ms: "
f"min={format_num(chunk_gap['min'])} "
f"p50={format_num(chunk_gap['p50'])} "
f"p95={format_num(chunk_gap['p95'])} "
f"p99={format_num(chunk_gap['p99'])} "
f"max={format_num(chunk_gap['max'])} "
f"threshold={format_num(chunk_gap['threshold_ms'])} "
f"breaches={chunk_gap['breach_count']}")
print("Queue wait ms: "
f"min={format_num(queue_wait['min'])} "
f"p50={format_num(queue_wait['p50'])} "
f"p95={format_num(queue_wait['p95'])} "
f"max={format_num(queue_wait['max'])}")
ratio = progressive["ratio"]
ratio_text = "n/a" if ratio is None else f"{ratio * 100:.2f}%"
print(
"Progressive streaming: "
f"{progressive['success_sessions']}/"
f"{progressive['eligible_sessions']} ({ratio_text})"
)
print(
"Bandwidth Mbps: "
f"per_session_avg={format_num(per_session_bw['avg'])} "
f"per_session_p95={format_num(per_session_bw['p95'])} "
f"global={format_num(bandwidth['global_goodput_mbps'])} "
f"bucket_avg={format_num(bucket_bw['avg_mbps'])} "
f"bucket_peak={format_num(bucket_bw['peak_mbps'])}"
)
print("Progressive streaming: "
f"{progressive['success_sessions']}/"
f"{progressive['eligible_sessions']} ({ratio_text})")
print("Bandwidth Mbps: "
f"per_session_avg={format_num(per_session_bw['avg'])} "
f"per_session_p95={format_num(per_session_bw['p95'])} "
f"global={format_num(bandwidth['global_goodput_mbps'])} "
f"bucket_avg={format_num(bucket_bw['avg_mbps'])} "
f"bucket_peak={format_num(bucket_bw['peak_mbps'])}")
print(f"VERDICT: {'PASS' if summary['passed'] else 'FAIL'}")
if summary["fail_reasons"]:
print("Fail reasons:")
@@ -670,10 +588,8 @@ def distribute_sessions(total_clients: int, process_count: int) -> list[int]:
def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
if websockets is None:
raise RuntimeError(
"Missing dependency: websockets. Install it before running this "
"stress test."
)
raise RuntimeError("Missing dependency: websockets. Install it before running this "
"stress test.")
preset_file = Path(args.preset_file).expanduser().resolve()
selected_preset_id, curated_prompts, total_prompt_count = load_curated_prompts(
@@ -735,13 +651,8 @@ def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
start_event.set()
result_deadline = (
time.monotonic()
+ args.connect_timeout_s
+ args.session_timeout_s
+ args.post_complete_wait_s
+ 180.0
)
result_deadline = (time.monotonic() + args.connect_timeout_s + args.session_timeout_s +
args.post_complete_wait_s + 180.0)
worker_results: list[dict[str, Any]] = []
while len(worker_results) < len(processes):
timeout_s = max(0.1, result_deadline - time.monotonic())
@@ -765,24 +676,20 @@ def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
if result.get("status") == "ok":
sessions.extend(result.get("sessions", []))
else:
worker_errors.append(
{
"worker_id": result.get("worker_id"),
"error": result.get("error"),
"traceback": result.get("traceback"),
}
)
worker_errors.append({
"worker_id": result.get("worker_id"),
"error": result.get("error"),
"traceback": result.get("traceback"),
})
received_workers = {result.get("worker_id") for result in worker_results}
expected_workers = set(range(len(processes)))
missing_workers = sorted(expected_workers - received_workers)
for worker_id in missing_workers:
worker_errors.append(
{
"worker_id": worker_id,
"error": "No worker result received.",
}
)
worker_errors.append({
"worker_id": worker_id,
"error": "No worker result received.",
})
run_end_epoch = time.time()
run_end_iso = iso_from_epoch(run_end_epoch)
@@ -795,9 +702,8 @@ def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
if worker_errors:
summary["passed"] = False
summary["fail_reasons"] = list(summary["fail_reasons"]) + [
f"{len(worker_errors)} worker error(s) occurred."
]
summary["fail_reasons"] = list(
summary["fail_reasons"]) + [f"{len(worker_errors)} worker error(s) occurred."]
output_payload = {
"run_info": {
@@ -833,9 +739,7 @@ def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Multiprocess realtime stress test for LTX2 streaming.",
)
parser = argparse.ArgumentParser(description="Multiprocess realtime stress test for LTX2 streaming.", )
parser.add_argument(
"-u",
"--url",
@@ -47,13 +47,11 @@ def test_persist_session_init_image_returns_none_when_missing_data():
def test_persist_session_init_image_rejects_unsupported_mime():
with pytest.raises(ValueError, match="PNG, JPEG, or WebP"):
persist_session_init_image(
{
"name": "frame.gif",
"mime_type": "image/gif",
"data_url": "data:image/gif;base64,R0lGODlhAQABAAAAACw=",
}
)
persist_session_init_image({
"name": "frame.gif",
"mime_type": "image/gif",
"data_url": "data:image/gif;base64,R0lGODlhAQABAAAAACw=",
})
def test_persist_session_init_image_rejects_large_payload(monkeypatch):
@@ -66,10 +64,8 @@ def test_persist_session_init_image_rejects_large_payload(monkeypatch):
monkeypatch.setattr(base64, "b64decode", fake_b64decode)
with pytest.raises(ValueError, match="15 MB or smaller"):
persist_session_init_image(
{
"name": "frame.png",
"mime_type": "image/png",
"data_url": data_url,
}
)
persist_session_init_image({
"name": "frame.png",
"mime_type": "image/png",
"data_url": data_url,
})
File diff suppressed because it is too large Load Diff
@@ -39,6 +39,17 @@ export FASTVIDEO_GENERATION_SEGMENT_CAP="${FASTVIDEO_GENERATION_SEGMENT_CAP:-6}"
export FASTVIDEO_PROMPT_AUTO_SLEEP_MS="${FASTVIDEO_PROMPT_AUTO_SLEEP_MS:-120}"
export FASTVIDEO_PROMPT_AUTO_TIMEOUT_MS="${FASTVIDEO_PROMPT_AUTO_TIMEOUT_MS:-1800}"
if [[ "${ENABLE_TORCH_COMPILE}" == "1" ]]; then
# Persist Inductor, AOTAutograd, and Triton artifacts across launches.
export DREAMVERSE_TORCH_COMPILE_CACHE_ROOT="${DREAMVERSE_TORCH_COMPILE_CACHE_ROOT:-${HOME}/.cache/dreamverse/torch_compile}"
export TORCHINDUCTOR_CACHE_DIR="${TORCHINDUCTOR_CACHE_DIR:-${DREAMVERSE_TORCH_COMPILE_CACHE_ROOT}/inductor}"
export TRITON_CACHE_DIR="${TRITON_CACHE_DIR:-${DREAMVERSE_TORCH_COMPILE_CACHE_ROOT}/triton}"
export TORCHINDUCTOR_FX_GRAPH_CACHE="${TORCHINDUCTOR_FX_GRAPH_CACHE:-1}"
export TORCHINDUCTOR_AUTOGRAD_CACHE="${TORCHINDUCTOR_AUTOGRAD_CACHE:-1}"
mkdir -p "${TORCHINDUCTOR_CACHE_DIR}" "${TRITON_CACHE_DIR}"
echo "[launch-demo] torch.compile cache: ${DREAMVERSE_TORCH_COMPILE_CACHE_ROOT}"
fi
cd "${DREAMVERSE_ROOT}"
if ! command -v dreamverse-server >/dev/null 2>&1; then
+9 -15
View File
@@ -8,12 +8,10 @@ import modal
IMAGE = os.environ.get("DREAMVERSE_IMAGE")
if not IMAGE:
raise RuntimeError(
"DREAMVERSE_IMAGE is required. Set it to a published SHA-specific Dreamverse image, "
"for example a dreamverse-backend-cuda13.0.0-sha-* tag or a "
"dreamverse-ui-cuda13.0.0-sha-* tag if serving the static UI. "
"CUDA 12 / cu126 images use the corresponding cuda12.6.3 tag."
)
raise RuntimeError("DREAMVERSE_IMAGE is required. Set it to a published SHA-specific Dreamverse image, "
"for example a dreamverse-backend-cuda13.0.0-sha-* tag or a "
"dreamverse-ui-cuda13.0.0-sha-* tag if serving the static UI. "
"CUDA 12 / cu126 images use the corresponding cuda12.6.3 tag.")
# ``@modal.web_server`` invokes ``serve()`` directly and bypasses the image
# ENTRYPOINT (``docker/docker_entrypoint.sh``). That entrypoint normally
@@ -65,14 +63,10 @@ def serve():
# ``or ""`` collapses ``None`` (unset) into an empty string, ``.strip()``
# collapses whitespace-only values (e.g. ``" "``) — both should be
# treated as missing.
missing = [
k for k in _REQUIRED_SECRET_KEYS
if not (os.environ.get(k) or "").strip()
]
missing = [k for k in _REQUIRED_SECRET_KEYS if not (os.environ.get(k) or "").strip()]
if missing:
raise RuntimeError(
"dreamverse-api-keys secret is missing required entries: "
f"{', '.join(missing)}. Add them with `modal secret create "
"dreamverse-api-keys ... --force` and redeploy "
"(see apps/dreamverse/scripts/modal/README.md).")
raise RuntimeError("dreamverse-api-keys secret is missing required entries: "
f"{', '.join(missing)}. Add them with `modal secret create "
"dreamverse-api-keys ... --force` and redeploy "
"(see apps/dreamverse/scripts/modal/README.md).")
subprocess.Popen(["dreamverse-server", "--host", "0.0.0.0", "--port", "8009"])
+1 -1
View File
@@ -1 +1 @@
PUBLIC_API_BASE_URL=http://localhost:8189/api
NEXT_PUBLIC_API_BASE_URL=http://localhost:8189/api
+6 -1
View File
@@ -13,6 +13,12 @@
# testing
/coverage
# playwright
/test-results/
/playwright-report/
/blob-report/
/.playwright/
# next.js
/.next/
/out/
@@ -39,4 +45,3 @@ yarn-error.log*
# typescript
*.tsbuildinfo
next-env.d.ts
.svelte-kit
-4
View File
@@ -1,4 +0,0 @@
{
"useTabs": true,
"tabWidth": 4
}
+49 -5
View File
@@ -11,19 +11,63 @@ The UI currently supports:
- Datasets
- Gallery View
## Project Structure
```
apps/fastvideo_studio/
├── server.py / job_runner.py / database.py # FastAPI backend + job lifecycle
├── mock_server.py # In-memory API mock for e2e tests
├── models/ # Pydantic request models (shared with the mock)
├── training_config.py # Studio workloads → fastvideo/train YAML configs
├── tests/ # Backend unit tests (pytest)
├── e2e/ # Playwright specs (run against the mock)
└── src/
├── app/ # Next.js App Router pages (thin routes)
├── components/
│ ├── shell/ # App chrome: header, sidebars, layout
│ ├── jobs/ # Job queue, cards, create-job modal, log sidebar
│ ├── datasets/ # Dataset cards, upload, captions
│ └── ui/ # shadcn-style primitives (shared theme)
├── stores/ # Framework-agnostic state + React bridge (hooks/)
├── lib/ # API client, types, option persistence
└── test/ # Vitest setup + factories
```
The visual theme (slate light/dark palettes, IBM Plex type, `#356cff` accent)
is shared with `apps/dreamverse`; the toggle in the header persists the choice
per browser.
## Testing
```bash
npm run typecheck # tsc
npm test # vitest unit tests
npm run e2e # Playwright against the in-memory mock backend
python -m pytest tests/ # backend unit tests (from apps/fastvideo_studio)
```
## Quick Start
The easiest way to run the application is:
For local development, install dependencies and start the Next.js dev server:
```bash
cd apps/fastvideo_studio
npm install
npm run dev
```
You can then access the app at [http://localhost:3000](http://localhost:3000).
Start the Python API server (default port 8189) in a separate terminal — see below.
For a production build of the web app together with the API server:
```bash
cd apps/fastvideo_studio
npm install
npm run build
npm run start
npm run start:all
```
You can then access the app at [http://localhost:3000](http://localhost:3000).
### Running API and Web Separately
The UI is composed of two separate components:
@@ -31,7 +75,7 @@ The UI is composed of two separate components:
- API Server
- Web Server
To run each component separately, you can use the commands `npm run start:web` and `npm run start:api`.
To run each component separately, you can use the commands `npm run start:web` (after `npm run build`) and `npm run start:api`.
You can also run the API server with this command from the `apps/` directory:
+4 -28
View File
@@ -251,11 +251,10 @@ class Database:
sp_size, negative_prompt,
data_path, max_train_steps, train_batch_size, learning_rate,
num_latent_t, validation_dataset_file, lora_rank,
ltx2_first_frame_conditioning_p,
dmd_use_vsa, dmd_vsa_sparsity, dmd_denoising_steps,
min_timestep_ratio, max_timestep_ratio, real_score_guidance_scale,
real_score_guidance_scale,
generator_update_interval, real_score_model_path, fake_score_model_path
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
job["id"],
@@ -297,12 +296,9 @@ class Database:
job.get("num_latent_t", 20),
job.get("validation_dataset_file", ""),
job.get("lora_rank", 32),
job.get("ltx2_first_frame_conditioning_p"),
1 if job.get("dmd_use_vsa") else 0,
job.get("dmd_vsa_sparsity", 0.8),
job.get("dmd_denoising_steps", "1000,757,522"),
job.get("min_timestep_ratio", 0.02),
job.get("max_timestep_ratio", 0.98),
job.get("real_score_guidance_scale", 3.5),
job.get("generator_update_interval", 5),
job.get("real_score_model_path", ""),
@@ -401,24 +397,6 @@ class Database:
for file_name, caption in captions.items():
self.upsert_dataset_caption(dataset_id, file_name, caption)
def update_dataset(self, dataset_id: str, updates: dict[str, Any]) -> None:
"""Update dataset fields. Only provided keys are updated."""
if not updates:
return
allowed: set[str] = set()
cols = []
vals = []
for k, v in updates.items():
if k in allowed:
cols.append(f"{k} = ?")
vals.append(v)
if not cols:
return
vals.append(dataset_id)
sql = f"UPDATE datasets SET {', '.join(cols)} WHERE id = ?"
self._execute(sql, tuple(vals))
self._commit()
def delete_dataset(self, dataset_id: str) -> bool:
"""Delete a dataset. Returns True if a row was deleted."""
cur = self._execute("DELETE FROM datasets WHERE id = ?", (dataset_id, ))
@@ -443,7 +421,7 @@ class Database:
cur = self._execute("SELECT * FROM settings WHERE id = 1")
row = cur.fetchone()
if not row:
return _default_settings_dict()
return default_settings_dict()
t2v = ((row["default_model_id_t2v"] or row["default_model_id"] or "") if "default_model_id_t2v" in row else
(row["default_model_id"] or ""))
i2v = ((row["default_model_id_i2v"] or "") if "default_model_id_i2v" in row else "")
@@ -557,8 +535,6 @@ def _row_to_job(row: sqlite3.Row) -> dict[str, Any]:
}
float_defaults = {
"dmd_vsa_sparsity": 0.8,
"min_timestep_ratio": 0.02,
"max_timestep_ratio": 0.98,
"real_score_guidance_scale": 3.5,
}
result = {
@@ -625,7 +601,7 @@ def _row_to_job(row: sqlite3.Row) -> dict[str, Any]:
return result
def _default_settings_dict() -> dict[str, Any]:
def default_settings_dict() -> dict[str, Any]:
"""Return default settings as API-style dict."""
return {
"defaultModelId": DEFAULT_SETTINGS["default_model_id"],
@@ -0,0 +1,42 @@
import { expect, test } from '@playwright/test';
import { skipWithoutMock } from './helpers';
/**
* Create-job flow: open the Create Job modal on /inference, fill the prompt
* (the model auto-selects once the mock's /api/models loads), submit, and
* confirm the new job lands in the queue.
*/
test.describe('create inference job', () => {
skipWithoutMock();
test('creates a T2V job and shows it in the queue', async ({ page }) => {
await page.goto('/inference');
// The trigger opens a real menu on click, so this path works for touch,
// mouse, and keyboard users.
await page.getByRole('button', { name: /create job/i }).click();
const t2vItem = page.getByRole('menuitem', { name: /T2V/i });
await expect(t2vItem).toBeVisible();
await t2vItem.click();
const dialog = page.getByRole('dialog');
await expect(dialog).toBeVisible();
// Wait for the mock's model catalogue to populate the dropdown (more than
// just the disabled placeholder), then pick one explicitly — the app's
// auto-selection is racy.
const modelSelect = dialog.getByLabel('Model', { exact: true });
await expect(modelSelect.locator('option')).not.toHaveCount(1);
await modelSelect.selectOption({ index: 1 });
const prompt = `e2e raccoon in sunflowers ${Date.now()}`;
await dialog.getByLabel('Prompt', { exact: true }).fill(prompt);
await dialog.getByRole('button', { name: 'Create Job' }).click();
// Modal closes and the queue refreshes with the newly created job.
await expect(dialog).toBeHidden();
await expect(page.getByText(prompt)).toBeVisible();
});
});
@@ -0,0 +1,37 @@
import { expect, test } from '@playwright/test';
import { API_BASE, skipWithoutMock } from './helpers';
/**
* Datasets page: the seeded datasets render as cards, and the Create Dataset
* modal opens from the header action.
*/
test.describe('datasets', () => {
skipWithoutMock();
test('lists the seeded datasets', async ({ page, request }) => {
// Read the seeded names from the mock so the assertion tracks the fixture.
const res = await request.get(`${API_BASE}/datasets`);
const datasets = (await res.json()) as Array<{ name: string }>;
expect(datasets.length).toBeGreaterThan(0);
await page.goto('/datasets');
for (const ds of datasets) {
await expect(page.getByText(ds.name, { exact: true })).toBeVisible();
}
});
test('opens the Create Dataset modal', async ({ page }) => {
await page.goto('/datasets');
await page.getByRole('button', { name: 'Add Dataset' }).click();
const dialog = page.getByRole('dialog');
await expect(dialog).toBeVisible();
await expect(
dialog.getByRole('heading', { name: /Add Dataset/i }),
).toBeVisible();
await expect(dialog.getByLabel('Name', { exact: true })).toBeVisible();
});
});
+44
View File
@@ -0,0 +1,44 @@
import { expect, test } from '@playwright/test';
import { API_BASE, skipWithoutMock } from './helpers';
/**
* Gallery page: the seeded completed inference job surfaces as a media tile
* with playback controls or an explicit media-error fallback.
*/
test.describe('gallery', () => {
skipWithoutMock();
test('shows a media tile for the seeded completed job', async ({
page,
request,
}) => {
const res = await request.get(`${API_BASE}/jobs?job_type=inference`);
const jobs = (await res.json()) as Array<{
status: string;
output_path: string | null;
prompt: string;
}>;
const completed = jobs.find(
(j) => j.status === 'completed' && j.output_path,
);
expect(completed, 'mock should seed a completed inference job').toBeTruthy();
await page.goto('/gallery');
await expect(
page.getByRole('heading', { level: 1, name: 'Gallery' }),
).toBeVisible();
const tile = page.locator('article').filter({ hasText: completed!.prompt });
await expect(tile).toBeVisible();
await expect(
tile.locator('video').or(tile.getByText('Preview unavailable')),
).toBeVisible();
const video = tile.locator('video');
if (await video.isVisible()) {
await expect(video).toHaveAttribute('controls', '');
}
});
});
+30
View File
@@ -0,0 +1,30 @@
import { expect, test } from '@playwright/test';
import { skipWithoutMock } from './helpers';
/**
* GPU status page: the mock's fake GPUs render as cards with meters, and the
* page is reachable from the primary sidebar.
*/
test.describe('gpus', () => {
skipWithoutMock();
test('lists the mock GPUs with utilization meters', async ({ page }) => {
await page.goto('/gpus');
await expect(
page.getByRole('heading', { level: 1, name: 'GPUs' }),
).toBeVisible();
await expect(page.getByText('GPU 0')).toBeVisible();
await expect(page.getByText('GPU 1')).toBeVisible();
expect(await page.getByRole('meter').count()).toBe(4);
});
test('is reachable from the sidebar', async ({ page }) => {
await page.goto('/inference');
await page.getByRole('link', { name: 'GPUs' }).click();
await expect(page).toHaveURL(/\/gpus$/);
await expect(page.getByText('GPU 0')).toBeVisible();
});
});
+44
View File
@@ -0,0 +1,44 @@
import { test, type APIRequestContext } from '@playwright/test';
/**
* Port and base URL of the mock API — the single source of truth shared by
* playwright.config.ts (which boots the mock and points `npm run dev` at it
* via NEXT_PUBLIC_API_BASE_URL) and the specs (which probe/read mock state
* directly through the Playwright `request` fixture).
*
* The default port is deliberately NOT the real server's default (8189): the
* e2e specs create and (if autostart is on) launch jobs, so running them
* against a real backend would mutate the developer's database and fire real
* GPU runs. API_BASE is derived solely from MOCK_API_PORT (it does not inherit
* NEXT_PUBLIC_API_BASE_URL) so the specs always talk to the mock.
*/
export const MOCK_API_PORT = process.env.MOCK_API_PORT || '8190';
export const API_BASE = `http://127.0.0.1:${MOCK_API_PORT}/api`;
export const MOCK_SKIP_MESSAGE =
`Mock backend not reachable at ${API_BASE}/__mock__ ` +
`(start fastvideo_studio.mock_server on port ${MOCK_API_PORT}).`;
/**
* True only when the mock server answers on the configured port. Probes the
* mock-only /api/__mock__ sentinel (which the real server does not serve) so
* the suite refuses to run its mutating specs against a real backend.
*/
export async function mockIsUp(request: APIRequestContext): Promise<boolean> {
try {
const res = await request.get(`${API_BASE}/__mock__`);
if (!res.ok()) return false;
const body = await res.json();
return body?.mock === true;
} catch {
return false;
}
}
/** Self-skip every test in the enclosing describe when the mock is down. */
export function skipWithoutMock(): void {
test.beforeEach(async ({ request }) => {
test.skip(!(await mockIsUp(request)), MOCK_SKIP_MESSAGE);
});
}
+115
View File
@@ -0,0 +1,115 @@
import { expect, test } from '@playwright/test';
import { skipWithoutMock } from './helpers';
/**
* App-shell smoke: the Next.js frontend hydrates, renders the FastVideo logo
* and the primary-sidebar navigation, and routing between the top-level
* sections works. Each spec self-skips when the mock backend isn't reachable.
*/
test.describe('app shell', () => {
skipWithoutMock();
test('loads with the logo and primary-sidebar nav', async ({ page }) => {
await page.goto('/');
// The root route redirects to /inference.
await expect(page).toHaveURL(/\/inference$/);
await expect(page.getByRole('img', { name: /fastvideo/i })).toBeVisible();
for (const label of ['Inference', 'Datasets', 'Gallery', 'Settings']) {
await expect(page.getByRole('link', { name: label })).toBeVisible();
}
});
test('navigates between the primary sections', async ({ page }) => {
await page.goto('/inference');
await expect(
page.getByRole('heading', { level: 1, name: 'Jobs' }),
).toBeVisible();
const sections: Array<{ link: string; url: RegExp; title: string }> = [
{ link: 'Datasets', url: /\/datasets$/, title: 'Datasets' },
{ link: 'Gallery', url: /\/gallery$/, title: 'Gallery' },
{ link: 'Settings', url: /\/settings$/, title: 'Settings' },
{ link: 'Inference', url: /\/inference$/, title: 'Jobs' },
];
for (const section of sections) {
await page.getByRole('link', { name: section.link }).click();
await expect(page).toHaveURL(section.url);
// The header <h1> is the only level-1 heading and reflects the route.
await expect(
page.getByRole('heading', { level: 1, name: section.title }),
).toBeVisible();
await expect(page.getByRole('main')).toHaveCount(1);
}
});
test('keeps navigation and content usable at responsive breakpoints', async ({
page,
}) => {
for (const width of [320, 375, 414, 768]) {
await page.setViewportSize({ width, height: 800 });
await page.goto('/inference');
const main = page.getByRole('main');
await expect(main).toBeVisible();
await expect(
page.getByRole('button', { name: /Create Job/i }),
).toBeVisible();
const initialBox = await main.boundingBox();
expect(initialBox?.x).toBe(width < 768 ? 0 : 220);
expect(initialBox?.width).toBe(width < 768 ? width : width - 220);
const navigation = page.getByRole('navigation', {
name: 'Primary navigation',
});
if (width < 768) {
await expect(
page.getByRole('button', { name: 'Open navigation' }),
).toBeVisible();
await page.getByRole('button', { name: 'Open navigation' }).click();
}
await expect(navigation).toBeVisible();
await navigation.getByRole('link', { name: 'Datasets' }).click();
await expect(page).toHaveURL(/\/datasets$/);
expect(
await page.evaluate(
() => document.documentElement.scrollWidth <= window.innerWidth,
),
).toBe(true);
}
});
test('uses full-width detail drawers on mobile', async ({ page }) => {
await page.setViewportSize({ width: 320, height: 800 });
await page.goto('/inference');
await page
.locator('article button[aria-pressed="false"]')
.first()
.click();
const jobDrawer = page.getByRole('dialog', { name: 'Job details' });
await expect(jobDrawer).toBeVisible();
expect(await jobDrawer.boundingBox()).toMatchObject({ x: 0, width: 320 });
await jobDrawer.getByRole('button', { name: 'Close' }).click();
await page.goto('/datasets');
await page
.locator('article button[aria-pressed="false"]')
.first()
.click();
const datasetDrawer = page.getByRole('dialog', {
name: /dataset details$/,
});
await expect(datasetDrawer).toBeVisible();
expect(await datasetDrawer.boundingBox()).toMatchObject({
x: 0,
width: 320,
});
});
});
-2
View File
@@ -1,2 +0,0 @@
// Minimal ESLint config for SvelteKit (add typescript-eslint / eslint-plugin-svelte as needed)
export default [];
+74
View File
@@ -0,0 +1,74 @@
# SPDX-License-Identifier: Apache-2.0
"""GPU telemetry for the studio status page, via NVML (nvidia-ml-py)."""
from __future__ import annotations
import contextlib
import logging
from typing import Any
logger = logging.getLogger("fastvideo.studio.gpu")
_nvml_initialized = False
def _ensure_nvml() -> Any:
"""Import and initialize NVML once; raises on machines without it."""
global _nvml_initialized # noqa: PLW0603
import pynvml
if not _nvml_initialized:
pynvml.nvmlInit()
_nvml_initialized = True
return pynvml
def _device_snapshot(pynvml: Any, index: int) -> dict[str, Any]:
handle = pynvml.nvmlDeviceGetHandleByIndex(index)
name = pynvml.nvmlDeviceGetName(handle)
if isinstance(name, bytes):
name = name.decode()
mem = pynvml.nvmlDeviceGetMemoryInfo(handle)
util = pynvml.nvmlDeviceGetUtilizationRates(handle)
# Optional sensors: not every GPU/driver exposes them.
temperature: int | None = None
power_watts: float | None = None
power_limit_watts: float | None = None
with contextlib.suppress(pynvml.NVMLError):
temperature = int(pynvml.nvmlDeviceGetTemperature(handle, pynvml.NVML_TEMPERATURE_GPU))
with contextlib.suppress(pynvml.NVMLError):
power_watts = pynvml.nvmlDeviceGetPowerUsage(handle) / 1000.0
power_limit_watts = (pynvml.nvmlDeviceGetEnforcedPowerLimit(handle) / 1000.0)
return {
"index": index,
"name": name,
"utilization": int(util.gpu),
"memory_used_mib": int(mem.used / (1024 * 1024)),
"memory_total_mib": int(mem.total / (1024 * 1024)),
"temperature_c": temperature,
"power_watts": power_watts,
"power_limit_watts": power_limit_watts,
}
def get_gpu_snapshot() -> dict[str, Any]:
"""Return {"available", "gpus", "error"} — never raises.
``available: False`` covers both "no NVIDIA driver/library on this host"
and transient NVML failures; the frontend shows ``error`` as-is.
"""
try:
pynvml = _ensure_nvml()
count = pynvml.nvmlDeviceGetCount()
gpus = [_device_snapshot(pynvml, i) for i in range(count)]
return {"available": True, "gpus": gpus, "error": None}
except ImportError:
return {
"available": False,
"gpus": [],
"error": "nvidia-ml-py is not installed on the API server host.",
}
except Exception as exc: # NVMLError, driver issues, …
logger.warning("GPU snapshot failed: %s", exc)
return {"available": False, "gpus": [], "error": str(exc)}
+33 -58
View File
@@ -23,12 +23,13 @@ import time
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
import yaml
from fastvideo.utils import get_mp_context
from fastvideo_studio.database import Database
from fastvideo_studio.training_config import (
build_training_args,
build_training_config,
get_training_env,
get_training_module_info,
)
logger = logging.getLogger("fastvideo.studio.job_runner")
@@ -160,13 +161,10 @@ class Job:
num_latent_t: int = 20
validation_dataset_file: str = ""
lora_rank: int = 32
ltx2_first_frame_conditioning_p: float | None = None
# DMD options
dmd_use_vsa: bool = False
dmd_vsa_sparsity: float = 0.8
dmd_denoising_steps: str = "1000,757,522"
min_timestep_ratio: float = 0.02
max_timestep_ratio: float = 0.98
real_score_guidance_scale: float = 3.5
generator_update_interval: int = 5
real_score_model_path: str = ""
@@ -221,12 +219,9 @@ class Job:
"num_width": self.width,
"validation_dataset_file": self.validation_dataset_file,
"lora_rank": self.lora_rank,
"ltx2_first_frame_conditioning_p": self.ltx2_first_frame_conditioning_p,
"dmd_use_vsa": self.dmd_use_vsa,
"dmd_vsa_sparsity": self.dmd_vsa_sparsity,
"dmd_denoising_steps": self.dmd_denoising_steps,
"min_timestep_ratio": self.min_timestep_ratio,
"max_timestep_ratio": self.max_timestep_ratio,
"real_score_guidance_scale": self.real_score_guidance_scale,
"generator_update_interval": self.generator_update_interval,
"real_score_model_path": self.real_score_model_path or "",
@@ -330,12 +325,9 @@ class JobRunner:
num_latent_t=row.get("num_latent_t", 20),
validation_dataset_file=row.get("validation_dataset_file", "") or "",
lora_rank=row.get("lora_rank", 32),
ltx2_first_frame_conditioning_p=row.get("ltx2_first_frame_conditioning_p"),
dmd_use_vsa=row.get("dmd_use_vsa", False),
dmd_vsa_sparsity=float(row.get("dmd_vsa_sparsity", 0.8)),
dmd_denoising_steps=row.get("dmd_denoising_steps", "1000,757,522") or "1000,757,522",
min_timestep_ratio=float(row.get("min_timestep_ratio", 0.02)),
max_timestep_ratio=float(row.get("max_timestep_ratio", 0.98)),
real_score_guidance_scale=float(row.get("real_score_guidance_scale", 3.5)),
generator_update_interval=int(row.get("generator_update_interval", 5)),
real_score_model_path=row.get("real_score_model_path", "") or "",
@@ -412,12 +404,9 @@ class JobRunner:
num_latent_t: int = 20,
validation_dataset_file: str = "",
lora_rank: int = 32,
ltx2_first_frame_conditioning_p: float | None = None,
dmd_use_vsa: bool = False,
dmd_vsa_sparsity: float = 0.8,
dmd_denoising_steps: str = "1000,757,522",
min_timestep_ratio: float = 0.02,
max_timestep_ratio: float = 0.98,
real_score_guidance_scale: float = 3.5,
generator_update_interval: int = 5,
real_score_model_path: str = "",
@@ -457,12 +446,9 @@ class JobRunner:
num_latent_t=num_latent_t,
validation_dataset_file=validation_dataset_file or "",
lora_rank=lora_rank,
ltx2_first_frame_conditioning_p=ltx2_first_frame_conditioning_p,
dmd_use_vsa=dmd_use_vsa,
dmd_vsa_sparsity=dmd_vsa_sparsity,
dmd_denoising_steps=dmd_denoising_steps,
min_timestep_ratio=min_timestep_ratio,
max_timestep_ratio=max_timestep_ratio,
real_score_guidance_scale=real_score_guidance_scale,
generator_update_interval=generator_update_interval,
real_score_model_path=real_score_model_path or "",
@@ -733,46 +719,23 @@ class JobRunner:
self._save_job(job)
return
module_info = get_training_module_info(job.workload_type, job.model_id)
if not module_info:
env = os.environ.copy()
env.update(get_training_env())
try:
# Job.to_dict() carries every key the config builder reads
# (extra keys are ignored by its .get() lookups).
train_config = build_training_config(job.to_dict(), job_output_dir)
except ValueError as exc:
job.status = JobStatus.FAILED
job.error = f"Unknown workload type: {job.workload_type}"
job.error = str(exc)
job.finished_at = time.time()
self._save_job(job)
return
module_path, _pipeline_workload, use_vsa, _is_lora = module_info
dmd_use_vsa = (job.workload_type.startswith("dmd_") and getattr(job, "dmd_use_vsa", False))
env = os.environ.copy()
env.update(get_training_env(use_vsa or dmd_use_vsa))
job_dict = {
"model_id": job.model_id,
"data_path": job.data_path,
"workload_type": job.workload_type,
"num_gpus": job.num_gpus,
"max_train_steps": job.max_train_steps,
"train_batch_size": job.train_batch_size,
"learning_rate": job.learning_rate,
"num_latent_t": job.num_latent_t,
"num_height": job.height,
"num_width": job.width,
"num_frames": job.num_frames,
"validation_dataset_file": job.validation_dataset_file,
"lora_rank": job.lora_rank,
"ltx2_first_frame_conditioning_p": job.ltx2_first_frame_conditioning_p,
}
if job.workload_type.startswith("dmd_") or job.workload_type.startswith("self_forcing_"):
job_dict["dmd_use_vsa"] = getattr(job, "dmd_use_vsa", False)
job_dict["dmd_vsa_sparsity"] = getattr(job, "dmd_vsa_sparsity", 0.8)
job_dict["dmd_denoising_steps"] = getattr(job, "dmd_denoising_steps", "1000,757,522")
job_dict["min_timestep_ratio"] = getattr(job, "min_timestep_ratio", 0.02)
job_dict["max_timestep_ratio"] = getattr(job, "max_timestep_ratio", 0.98)
job_dict["real_score_guidance_scale"] = getattr(job, "real_score_guidance_scale", 3.5)
job_dict["generator_update_interval"] = getattr(job, "generator_update_interval", 5)
job_dict["real_score_model_path"] = (getattr(job, "real_score_model_path", "") or job.model_id)
job_dict["fake_score_model_path"] = (getattr(job, "fake_score_model_path", "") or job.model_id)
train_args = build_training_args(job_dict, job_output_dir)
config_path = os.path.join(job_output_dir, "train_config.yaml")
with open(config_path, "w", encoding="utf-8") as f:
yaml.safe_dump(train_config, f, sort_keys=False)
repo_root = Path(__file__).resolve().parent.parent
torchrun_cmd = [
@@ -783,9 +746,12 @@ class JobRunner:
str(job.num_gpus),
"--nnodes",
"1",
str(repo_root / module_path),
] + train_args
buf.write(f"Starting training: {' '.join(torchrun_cmd[:12])}...")
"-m",
"fastvideo.train.entrypoint.train",
"--config",
config_path,
]
buf.write(f"Starting training: {' '.join(torchrun_cmd)}")
buf.phase = "starting"
try:
@@ -825,12 +791,21 @@ class JobRunner:
job._process.wait()
exit_code = job._process.returncode or 0
if exit_code == 0:
if job._stop_event.is_set():
# Terminated by stop_job() without the reader loop observing the
# flag (e.g. the process died between log lines).
job.status = JobStatus.STOPPED
buf.phase = "stopped"
elif exit_code == 0:
job.status = JobStatus.COMPLETED
buf.progress = 100.0
buf.phase = "done"
# Training outputs checkpoints, not video
ckpt_dirs = sorted(Path(job_output_dir).glob("checkpoint-*"))
# Training outputs checkpoints, not video. Sort by step number,
# not lexically ("checkpoint-1000" < "checkpoint-500" as strings).
ckpt_dirs = sorted(
Path(job_output_dir).glob("checkpoint-*"),
key=lambda p: int(m.group(1)) if (m := re.fullmatch(r"checkpoint-(\d+)", p.name)) else -1,
)
if ckpt_dirs:
job.output_path = str(ckpt_dirs[-1])
else:
+660
View File
@@ -0,0 +1,660 @@
# SPDX-License-Identifier: Apache-2.0
"""
In-memory mock of the FastVideo Studio API (``server.py``).
This mock implements the same ``/api`` routes and response shapes as the real
FastAPI server but keeps everything in memory and never touches the FastVideo
library, a GPU, or a database. It is meant to back the Playwright e2e suite so
the Next.js frontend can be exercised end-to-end without a real backend.
Job lifecycle is simulated by recording a start timestamp and *computing* the
status on read: a started job reports ``running`` for a few seconds and then
flips to ``completed`` with an ``output_path``, so polling the job list / logs
shows progression. Generated media is a tiny 1s ``testsrc`` MP4 built lazily
with ffmpeg and cached.
Usage (from the ``apps/`` directory)::
PYTHONPATH=.. python -m fastvideo_studio.mock_server --port 8190
"""
from __future__ import annotations
import contextlib
import os
import random
import shutil
import subprocess
import tempfile
import threading
import time
import uuid
from typing import Annotated, Any
from fastapi import FastAPI, File, HTTPException, UploadFile
from fastapi.middleware.cors import CORSMiddleware
from fastapi.responses import FileResponse, PlainTextResponse
from fastvideo_studio.database import default_settings_dict
from fastvideo_studio.models import (CreateDatasetRequest, CreateJobRequest, SettingsUpdate, UpdateCaptionRequest,
model_label)
# --- Config -----------------------------------------------------------------
# How long a started job stays "running" before it flips to "completed".
COMPLETE_AFTER_SECONDS = 3.0
FFMPEG_BIN = shutil.which(os.getenv("FASTVIDEO_FFMPEG_BIN", "ffmpeg"))
# A small catalogue of fake models keyed by workload type. Mirrors the real
# server's {id, label} shape (label derived from the HF-style path).
_MODELS_BY_WORKLOAD: dict[str, list[str]] = {
"t2v": [
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
"FastVideo/FastHunyuan-diffusers",
],
"i2v": [
"Wan-AI/Wan2.1-I2V-14B-480P-Diffusers",
],
"t2i": [
"black-forest-labs/FLUX.1-schnell",
],
}
def _models_for(workload_type: str | None) -> list[dict[str, str]]:
if workload_type:
paths = _MODELS_BY_WORKLOAD.get(workload_type, [])
else:
seen: dict[str, None] = {}
for paths_for_workload in _MODELS_BY_WORKLOAD.values():
for path in paths_for_workload:
seen.setdefault(path, None)
paths = list(seen)
return [{"id": path, "label": model_label(path)} for path in paths]
# The real settings catalogue lives in database.py; the mock only pre-fills
# the default model ids so the Create Job modal auto-selects one.
_DEFAULT_SETTINGS: dict[str, Any] = {
**default_settings_dict(),
"defaultModelId": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
"defaultModelIdT2v": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
"defaultModelIdI2v": "Wan-AI/Wan2.1-I2V-14B-480P-Diffusers",
"defaultModelIdT2i": "black-forest-labs/FLUX.1-schnell",
}
# --- In-memory state --------------------------------------------------------
_state_lock = threading.Lock()
_settings: dict[str, Any] = dict(_DEFAULT_SETTINGS)
_jobs: dict[str, dict[str, Any]] = {}
_datasets: dict[str, dict[str, Any]] = {}
# dataset_id -> {"file_names": [...], "captions": {file_name: caption}}
_dataset_files: dict[str, dict[str, Any]] = {}
# Lazily-built, cached media clips (shared by every job/dataset media response),
# keyed by extension: "mp4" (1s testsrc video) or "png" (single testsrc frame).
_mock_media_cache: dict[str, str] = {}
_mock_media_lock = threading.Lock()
def _build_mock_media(kind: str) -> str:
"""Build (once) a tiny testsrc clip of the given ``kind`` with ffmpeg; cache the path."""
with _mock_media_lock:
cached = _mock_media_cache.get(kind)
if cached and os.path.isfile(cached):
return cached
if not FFMPEG_BIN:
raise HTTPException(status_code=500, detail="ffmpeg is required to build mock media")
fd, path = tempfile.mkstemp(prefix="fvstudio_mock_", suffix=f".{kind}")
os.close(fd)
if kind == "png":
command = [
FFMPEG_BIN,
"-hide_banner",
"-loglevel",
"error",
"-y",
"-f",
"lavfi",
"-i",
"testsrc=size=320x240:rate=1",
"-frames:v",
"1",
"-f",
"image2",
path,
]
else:
command = [
FFMPEG_BIN,
"-hide_banner",
"-loglevel",
"error",
"-y",
"-f",
"lavfi",
"-i",
"testsrc=size=320x240:rate=24",
"-t",
"1",
"-c:v",
"libx264",
"-preset",
"ultrafast",
"-pix_fmt",
"yuv420p",
"-movflags",
"+faststart",
"-f",
"mp4",
path,
]
try:
subprocess.run(command, check=True, capture_output=True)
except (subprocess.CalledProcessError, OSError) as exc:
with contextlib.suppress(OSError):
os.remove(path)
raise HTTPException(status_code=500, detail=f"ffmpeg failed to build mock {kind}: {exc}") from exc
if not os.path.isfile(path) or os.path.getsize(path) == 0:
raise HTTPException(status_code=500, detail=f"ffmpeg produced no mock {kind} bytes")
_mock_media_cache[kind] = path
return path
# --- Job helpers ------------------------------------------------------------
def _new_job_dict(req: CreateJobRequest) -> dict[str, Any]:
"""Build a job dict (mirrors job_runner.Job.to_dict()) in the pending state."""
job_id = str(uuid.uuid4())
return {
"id": job_id,
"model_id": req.model_id,
"prompt": req.prompt,
"workload_type": req.workload_type or "t2v",
"job_type": req.job_type or "inference",
"image_path": req.image_path or "",
"status": "pending",
"created_at": time.time(),
"started_at": None,
"finished_at": None,
"error": None,
"output_path": None,
"log_file_path": None,
"num_inference_steps": req.num_inference_steps,
"num_frames": req.num_frames,
"height": req.height,
"width": req.width,
"guidance_scale": req.guidance_scale,
"guidance_rescale": req.guidance_rescale,
"fps": req.fps,
"seed": req.seed,
"negative_prompt": req.negative_prompt or "",
"num_gpus": req.num_gpus,
"data_path": req.data_path or "",
"progress": 0.0,
"progress_msg": "",
"phase": "pending",
}
def _advance_job(job: dict[str, Any]) -> None:
"""Flip a running job to completed once enough wall-clock time has passed.
Status is *computed on read* from the recorded start timestamp, so polling
the job list / logs naturally shows pending -> running -> completed.
"""
if job["status"] != "running" or not job.get("started_at"):
return
elapsed = time.time() - job["started_at"]
if elapsed >= COMPLETE_AFTER_SECONDS:
job["status"] = "completed"
job["finished_at"] = job["started_at"] + COMPLETE_AFTER_SECONDS
job["progress"] = 100.0
job["progress_msg"] = "50/50 steps"
job["phase"] = "done"
ext = "png" if job.get("workload_type") == "t2i" else "mp4"
job["output_path"] = f"/mock/outputs/{job['id']}/output.{ext}"
def _public_job(job: dict[str, Any]) -> dict[str, Any]:
"""Advance + return a copy safe to serialize."""
_advance_job(job)
return dict(job)
_LOG_TAIL = [
"Loading model...",
"Model loaded.",
"Starting generation...",
"Denoising step 10/50",
"Denoising step 20/50",
"Denoising step 30/50",
"Denoising step 40/50",
"Denoising step 50/50",
"Generation complete. Saving output...",
"Saved output file.",
"Job completed successfully.",
]
def _log_sequence(job: dict[str, Any]) -> list[str]:
return [
f"Job {job['id']} started",
f"Model: {job['model_id']}",
f"Prompt: {job['prompt']}",
*_LOG_TAIL,
]
def _compute_logs(job: dict[str, Any]) -> dict[str, Any]:
"""Return JobLogs-shaped data, growing the visible lines as time passes."""
seq = _log_sequence(job)
status = job["status"]
if status == "pending":
return {"lines": [], "progress": 0.0, "progress_msg": "", "phase": "pending"}
if status == "completed":
return {"lines": seq, "progress": 100.0, "progress_msg": "50/50 steps", "phase": "done"}
if status in ("failed", "stopped"):
# Keep the total line count monotonic across running -> terminal (the
# `after` cursor relies on it): show every non-final line plus a terminal
# notice, which is always >= any prefix a running job revealed.
lines = seq[:-1] + [f"Job {status}."]
return {"lines": lines, "progress": job.get("progress", 0.0), "progress_msg": "", "phase": status}
# running: reveal a prefix proportional to elapsed time, but never the final
# "completed" line — that appears only once the job is actually completed.
elapsed = time.time() - (job.get("started_at") or time.time())
frac = max(0.0, min(elapsed / COMPLETE_AFTER_SECONDS, 0.99))
reveal = min(max(4, round(frac * len(seq))), len(seq) - 1)
lines = seq[:reveal]
if frac < 0.2:
phase = "loading model"
elif frac < 0.85:
phase = "denoising"
else:
phase = "saving"
return {
"lines": lines,
"progress": round(frac * 100.0, 1),
"progress_msg": f"{int(frac * 50)}/50 steps",
"phase": phase,
}
# --- Dataset helpers --------------------------------------------------------
def _dataset_stats(dataset_id: str) -> tuple[int, int]:
files = _dataset_files.get(dataset_id, {}).get("file_names", [])
count = len(files)
# Fake but stable per-file size so the UI shows a non-zero footprint.
return count, count * 1_048_576
def _public_dataset(dataset: dict[str, Any]) -> dict[str, Any]:
count, size = _dataset_stats(dataset["id"])
return {**dataset, "file_count": count, "size_bytes": size}
def _seed() -> None:
"""Seed a couple of datasets and one completed inference job."""
now = time.time()
seeds = [
("Sunset Clips", ["sunset_01.mp4", "sunset_02.mp4"], {
"sunset_01.mp4": "A sunset over the ocean",
"sunset_02.mp4": "A sunset over the mountains"
}),
("City Timelapse", ["city_01.mp4"], {
"city_01.mp4": "A busy city intersection at night"
}),
]
for idx, (name, file_names, captions) in enumerate(seeds):
dataset_id = str(uuid.uuid4())
_datasets[dataset_id] = {"id": dataset_id, "name": name, "created_at": now - 600 + idx}
_dataset_files[dataset_id] = {"file_names": list(file_names), "captions": dict(captions)}
job_id = str(uuid.uuid4())
_jobs[job_id] = {
"id": job_id,
"model_id": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
"prompt": "A curious raccoon peers through a field of yellow sunflowers",
"workload_type": "t2v",
"job_type": "inference",
"image_path": "",
"status": "completed",
"created_at": now - 300,
"started_at": now - 280,
"finished_at": now - 250,
"error": None,
"output_path": f"/mock/outputs/{job_id}/output.mp4",
"log_file_path": f"/mock/logs/{job_id}.log",
"num_inference_steps": 50,
"num_frames": 81,
"height": 480,
"width": 832,
"guidance_scale": 5.0,
"guidance_rescale": 0.0,
"fps": 24,
"seed": 1024,
"negative_prompt": "",
"num_gpus": 1,
"data_path": "",
"progress": 100.0,
"progress_msg": "50/50 steps",
"phase": "done",
}
# --- App --------------------------------------------------------------------
app = FastAPI(title="FastVideo Studio Mock API", version="0.1.0")
app.add_middleware(
CORSMiddleware,
allow_origins=["*"],
allow_methods=["*"],
allow_headers=["*"],
)
_seed()
@app.get("/api/__mock__")
def mock_sentinel() -> dict[str, bool]:
"""Mock-only marker so the e2e suite can prove it isn't hitting a real
backend before running its mutating specs (the real server has no such
route)."""
return {"mock": True}
# --- Settings ---------------------------------------------------------------
@app.get("/api/settings")
def get_settings() -> dict[str, Any]:
with _state_lock:
return dict(_settings)
@app.put("/api/settings")
def update_settings(settings: SettingsUpdate) -> dict[str, Any]:
updates = settings.model_dump(exclude_unset=True)
with _state_lock:
_settings.update({k: v for k, v in updates.items() if v is not None})
return dict(_settings)
# --- Models -----------------------------------------------------------------
@app.get("/api/models")
def list_models(workload_type: str | None = None) -> list[dict[str, Any]]:
return _models_for(workload_type)
# --- GPUs -------------------------------------------------------------------
@app.get("/api/gpus")
def list_gpus() -> dict[str, Any]:
"""Two fake GPUs with slight per-request jitter so the page looks live."""
gpus = []
for index, (base_util, used_mib) in enumerate([(62, 61_440), (7, 4_096)]):
gpus.append({
"index": index,
"name": "NVIDIA Mock GPU 80GB",
"utilization": max(0, min(100, base_util + random.randint(-5, 5))),
"memory_used_mib": used_mib + random.randint(-256, 256),
"memory_total_mib": 81_920,
"temperature_c": 55 + random.randint(-3, 3),
"power_watts": 310.0 + random.randint(-20, 20),
"power_limit_watts": 700.0,
})
return {"available": True, "gpus": gpus, "error": None}
# --- Uploads ----------------------------------------------------------------
@app.post("/api/upload-image")
async def upload_image(file: Annotated[UploadFile, File()]) -> dict[str, str]:
name = file.filename or "image.png"
return {"path": f"/mock/uploads/{uuid.uuid4().hex}_{os.path.basename(name)}"}
@app.post("/api/upload-raw-dataset")
async def upload_raw_dataset(files: Annotated[list[UploadFile], File()]) -> dict[str, Any]:
video_exts = {".mp4", ".webm", ".avi", ".mov", ".mkv"}
file_names: list[str] = []
for uf in files:
name = os.path.basename(uf.filename or f"{uuid.uuid4().hex}.mp4")
if os.path.splitext(name)[1].lower() in video_exts:
file_names.append(name)
if not file_names:
raise HTTPException(status_code=400, detail="No video files found.")
upload_id = uuid.uuid4().hex
path = f"/mock/uploads/{upload_id}"
return {"path": path, "upload_id": upload_id, "file_names": file_names}
# --- Jobs -------------------------------------------------------------------
@app.get("/api/jobs")
def list_jobs(job_type: str | None = None) -> list[dict[str, Any]]:
# _public_job mutates the shared job dict via _advance_job, so it must run
# under the lock (matching every other job route) to avoid racing start/stop.
with _state_lock:
jobs = list(_jobs.values())
if job_type:
jobs = [j for j in jobs if j.get("job_type") == job_type]
jobs.sort(key=lambda j: j["created_at"], reverse=True)
return [_public_job(j) for j in jobs]
@app.get("/api/jobs/{job_id}")
def get_job(job_id: str) -> dict[str, Any]:
with _state_lock:
job = _jobs.get(job_id)
if job is None:
raise HTTPException(status_code=404, detail="Job not found")
return _public_job(job)
@app.post("/api/jobs", status_code=201)
def create_job(req: CreateJobRequest) -> dict[str, Any]:
job = _new_job_dict(req)
with _state_lock:
_jobs[job["id"]] = job
if _settings.get("autoStartJob"):
_start(job)
return _public_job(job)
def _start(job: dict[str, Any]) -> None:
job["status"] = "running"
job["started_at"] = time.time()
job["finished_at"] = None
job["error"] = None
job["output_path"] = None
# The real server assigns the log path once the job starts running; mirror
# that so the Job Details "Download Log" button (gated on log_file_path) works.
job["log_file_path"] = f"/mock/logs/{job['id']}.log"
job["progress"] = 0.0
job["progress_msg"] = ""
job["phase"] = "starting"
@app.post("/api/jobs/{job_id}/start")
def start_job(job_id: str) -> dict[str, Any]:
with _state_lock:
job = _jobs.get(job_id)
if job is None:
raise HTTPException(status_code=404, detail=f"Job {job_id} not found")
_advance_job(job)
if job["status"] == "running":
raise HTTPException(status_code=409, detail="Job is already running")
if job["status"] == "completed":
raise HTTPException(status_code=409, detail="Job already completed. Delete and re-create to run again.")
_start(job)
return _public_job(job)
@app.post("/api/jobs/{job_id}/stop")
def stop_job(job_id: str) -> dict[str, Any]:
with _state_lock:
job = _jobs.get(job_id)
if job is None:
raise HTTPException(status_code=404, detail=f"Job {job_id} not found")
_advance_job(job)
if job["status"] != "running":
raise HTTPException(status_code=409, detail=f"Job is not running (status={job['status']})")
job["status"] = "stopped"
job["finished_at"] = time.time()
job["phase"] = "stopped"
return dict(job)
@app.delete("/api/jobs/{job_id}")
def delete_job(job_id: str) -> dict[str, str]:
with _state_lock:
if _jobs.pop(job_id, None) is None:
raise HTTPException(status_code=404, detail="Job not found")
return {"detail": f"Job {job_id} deleted"}
@app.get("/api/jobs/{job_id}/logs")
def get_job_logs(job_id: str, after: int = 0) -> dict[str, Any]:
with _state_lock:
job = _jobs.get(job_id)
if job is None:
raise HTTPException(status_code=404, detail=f"Job {job_id} not found")
_advance_job(job)
result = _compute_logs(job)
all_lines = result["lines"]
return {
"lines": all_lines[after:],
"total": len(all_lines),
"progress": result["progress"],
"progress_msg": result["progress_msg"],
"phase": result["phase"],
}
@app.get("/api/jobs/{job_id}/video")
def get_video(job_id: str) -> FileResponse:
with _state_lock:
job = _jobs.get(job_id)
if job is None:
raise HTTPException(status_code=404, detail="Job not found")
_advance_job(job)
output_path = job.get("output_path")
if job["status"] != "completed" or not output_path:
raise HTTPException(status_code=404, detail="No output available for this job")
# Mirror the real server: image workloads (t2i) output a .png served as an
# image; everything else is a video.
if output_path.endswith(".png"):
return FileResponse(_build_mock_media("png"), media_type="image/png", filename=f"job_{job_id}.png")
return FileResponse(_build_mock_media("mp4"), media_type="video/mp4", filename=f"job_{job_id}.mp4")
@app.get("/api/jobs/{job_id}/download_log")
def download_log(job_id: str) -> PlainTextResponse:
with _state_lock:
job = _jobs.get(job_id)
if job is None:
raise HTTPException(status_code=404, detail="Job not found")
_advance_job(job)
lines = _compute_logs(job)["lines"]
return PlainTextResponse("\n".join(lines) + "\n", media_type="text/plain")
# --- Datasets ---------------------------------------------------------------
@app.get("/api/datasets")
def list_datasets() -> list[dict[str, Any]]:
with _state_lock:
datasets = sorted(_datasets.values(), key=lambda d: d["created_at"], reverse=True)
return [_public_dataset(d) for d in datasets]
@app.get("/api/datasets/{dataset_id}")
def get_dataset(dataset_id: str) -> dict[str, Any]:
with _state_lock:
dataset = _datasets.get(dataset_id)
if dataset is None:
raise HTTPException(status_code=404, detail="Dataset not found")
return _public_dataset(dataset)
@app.post("/api/datasets", status_code=201)
def create_dataset(req: CreateDatasetRequest) -> dict[str, Any]:
if not req.upload_path:
raise HTTPException(status_code=400, detail="upload_path is required. Upload media files first.")
if not req.file_names:
raise HTTPException(status_code=400, detail="No media files found.")
dataset_id = str(uuid.uuid4())
dataset = {"id": dataset_id, "name": req.name, "created_at": time.time()}
captions = {fn: (req.captions.get(fn, "") if req.captions else "") for fn in req.file_names}
with _state_lock:
_datasets[dataset_id] = dataset
_dataset_files[dataset_id] = {"file_names": list(req.file_names), "captions": captions}
return _public_dataset(dataset)
@app.get("/api/datasets/{dataset_id}/files")
def get_dataset_files(dataset_id: str) -> dict[str, Any]:
with _state_lock:
if dataset_id not in _datasets:
raise HTTPException(status_code=404, detail="Dataset not found")
files = _dataset_files.get(dataset_id, {"file_names": [], "captions": {}})
return {"file_names": list(files["file_names"]), "captions": dict(files["captions"])}
@app.put("/api/datasets/{dataset_id}/captions")
def update_dataset_caption(dataset_id: str, req: UpdateCaptionRequest) -> dict[str, str]:
with _state_lock:
if dataset_id not in _datasets:
raise HTTPException(status_code=404, detail="Dataset not found")
files = _dataset_files.setdefault(dataset_id, {"file_names": [], "captions": {}})
files["captions"][req.file_name] = req.caption
return {"detail": "Caption updated"}
@app.get("/api/datasets/{dataset_id}/media/{file_name:path}")
def serve_dataset_media(dataset_id: str, file_name: str) -> FileResponse:
with _state_lock:
if dataset_id not in _datasets:
raise HTTPException(status_code=404, detail="Dataset not found")
if file_name not in _dataset_files.get(dataset_id, {}).get("file_names", []):
raise HTTPException(status_code=404, detail="File not found")
return FileResponse(_build_mock_media("mp4"), media_type="video/mp4")
@app.delete("/api/datasets/{dataset_id}")
def delete_dataset(dataset_id: str) -> dict[str, str]:
with _state_lock:
if _datasets.pop(dataset_id, None) is None:
raise HTTPException(status_code=404, detail="Dataset not found")
_dataset_files.pop(dataset_id, None)
return {"detail": f"Dataset {dataset_id} deleted"}
def main() -> None:
import argparse
import uvicorn
parser = argparse.ArgumentParser(description="FastVideo Studio mock API server")
parser.add_argument("--host", default="127.0.0.1", help="Bind address (default: 127.0.0.1)")
# Default off the real server's 8189 so the mock never shadows a real API.
parser.add_argument("--port", type=int, default=8190, help="Port number (default: 8190)")
args = parser.parse_args()
uvicorn.run(app, host=args.host, port=args.port, log_level="info")
if __name__ == "__main__":
main()
+9 -1
View File
@@ -1,14 +1,22 @@
# SPDX-License-Identifier: Apache-2.0
"""Pydantic request/response models for the API."""
"""Pydantic request/response models for the API, plus tiny shared helpers
usable by both the real server and the dependency-light mock server."""
from fastvideo_studio.models.create_job_request import CreateJobRequest
from fastvideo_studio.models.settings_update import SettingsUpdate
from fastvideo_studio.models.create_dataset_request import CreateDatasetRequest
from fastvideo_studio.models.update_caption_request import UpdateCaptionRequest
def model_label(model_path: str) -> str:
"""Derive a readable label from an HF-style model path."""
return model_path.split("/")[-1].replace("-", " ").replace("_", " ")
__all__ = [
"CreateJobRequest",
"SettingsUpdate",
"CreateDatasetRequest",
"UpdateCaptionRequest",
"model_label",
]
@@ -17,7 +17,6 @@ class CreateJobRequest(BaseModel):
num_latent_t: int = 20
validation_dataset_file: str = ""
lora_rank: int = 32
ltx2_first_frame_conditioning_p: float | None = None
negative_prompt: str = ""
num_inference_steps: int = 50
num_frames: int = 81
@@ -41,8 +40,6 @@ class CreateJobRequest(BaseModel):
dmd_use_vsa: bool = False
dmd_vsa_sparsity: float = 0.8
dmd_denoising_steps: str = "1000,757,522"
min_timestep_ratio: float = 0.02
max_timestep_ratio: float = 0.98
real_score_guidance_scale: float = 3.5
generator_update_interval: int = 5
real_score_model_path: str = ""
+13
View File
@@ -0,0 +1,13 @@
import path from 'node:path';
import { fileURLToPath } from 'node:url';
import type { NextConfig } from 'next';
const configDir = path.dirname(fileURLToPath(import.meta.url));
const nextConfig: NextConfig = {
// Point tracing at the monorepo root (apps/fastvideo_studio -> apps -> repo
// root) so Next stops warning about multiple lockfiles in the workspace.
outputFileTracingRoot: path.join(configDir, '..', '..'),
};
export default nextConfig;
+5299 -1038
View File
File diff suppressed because it is too large Load Diff
+44 -22
View File
@@ -4,33 +4,55 @@
"private": true,
"type": "module",
"scripts": {
"dev": "vite dev",
"build": "vite build",
"preview": "vite preview",
"start": "concurrently --kill-others-on-fail \"npm:start:api\" \"npm:start:web\"",
"dev": "next dev --port 3000",
"build": "next build",
"start": "next start --port 3000",
"typecheck": "tsc --noEmit",
"test": "vitest run",
"test:watch": "vitest",
"test:coverage": "vitest run --coverage",
"e2e": "playwright test",
"start:api": "cd .. && python -m fastvideo_studio.server",
"start:web": "node build",
"lint": "eslint ."
"start:web": "next start --port 3000",
"start:all": "concurrently --kill-others-on-fail \"npm:start:api\" \"npm:start:web\""
},
"dependencies": {
"@sveltejs/kit": "2.57.1",
"svelte": "5.55.7"
"@radix-ui/react-dialog": "^1.1.0",
"@radix-ui/react-dropdown-menu": "^2.1.24",
"@radix-ui/react-label": "^2.1.8",
"@radix-ui/react-scroll-area": "^1.2.10",
"@radix-ui/react-select": "^2.2.6",
"@radix-ui/react-separator": "^1.1.8",
"@radix-ui/react-slider": "^1.2.0",
"@radix-ui/react-slot": "^1.2.4",
"@radix-ui/react-switch": "^1.1.0",
"@radix-ui/react-tabs": "^1.1.0",
"class-variance-authority": "^0.7.1",
"clsx": "^2.1.1",
"lucide-react": "^0.577.0",
"next": "15.5.18",
"react": "^19.1.0",
"react-dom": "^19.1.0",
"sonner": "^2.0.7",
"tailwind-merge": "^3.5.0"
},
"devDependencies": {
"@sveltejs/adapter-node": "^5.5.4",
"@sveltejs/vite-plugin-svelte": "^5.0.0",
"@types/node": "^20",
"@playwright/test": "^1.59.1",
"@tailwindcss/postcss": "^4.2.1",
"@testing-library/dom": "^10.4.1",
"@testing-library/jest-dom": "^6.9.1",
"@testing-library/react": "^16.3.0",
"@testing-library/user-event": "^14.6.1",
"@types/node": "^22",
"@types/react": "^19.1.0",
"@types/react-dom": "^19.1.0",
"@vitejs/plugin-react": "^4.5.2",
"@vitest/coverage-v8": "^3.2.4",
"concurrently": "^9",
"eslint": "^9",
"typescript": "^5",
"vite": "6.4.2"
},
"overrides": {
"cookie": "0.7.0",
"devalue": "5.8.1",
"flatted": "3.4.2",
"minimatch": "3.1.4",
"picomatch": "4.0.4",
"postcss": "8.5.10"
"jsdom": "^26.1.0",
"postcss": "8.5.10",
"tailwindcss": "^4.2.1",
"typescript": "^5.8.3",
"vitest": "^3.2.4"
}
}
@@ -0,0 +1,54 @@
import { defineConfig, devices } from '@playwright/test';
import { API_BASE, MOCK_API_PORT } from './e2e/helpers';
/**
* Playwright config for FastVideo Studio end-to-end tests.
*
* The Next.js frontend runs on port 3000 and talks to the in-memory mock API
* (fastvideo_studio.mock_server) on port 8189. The webServer block boots both
* the mock backend and `npm run dev` (pointed at the mock via
* NEXT_PUBLIC_API_BASE_URL) if nothing is already listening, so the suite works
* both locally and in CI. Set PLAYWRIGHT_SKIP_WEBSERVER to reuse externally
* managed servers.
*/
export default defineConfig({
testDir: './e2e',
fullyParallel: false,
workers: 1,
forbidOnly: !!process.env.CI,
retries: process.env.CI ? 2 : 0,
reporter: process.env.CI ? 'github' : 'list',
timeout: 120_000,
expect: { timeout: 30_000 },
use: {
baseURL: process.env.PLAYWRIGHT_BASE_URL ?? 'http://127.0.0.1:3000',
headless: true,
viewport: { width: 1280, height: 720 },
screenshot: 'only-on-failure',
trace: 'retain-on-failure',
},
projects: [{ name: 'chromium', use: { ...devices['Desktop Chrome'] } }],
webServer: process.env.PLAYWRIGHT_SKIP_WEBSERVER
? undefined
: [
{
command: `PYTHONPATH=.. python -m fastvideo_studio.mock_server --port ${MOCK_API_PORT}`,
// Probe the mock-only sentinel, not /api/models (which a real server
// also serves), so a real backend on this port is never mistaken for
// the mock. Always boot a fresh mock in CI.
url: `http://127.0.0.1:${MOCK_API_PORT}/api/__mock__`,
reuseExistingServer: !process.env.CI,
timeout: 120_000,
},
{
command: 'npm run dev',
url: 'http://127.0.0.1:3000',
reuseExistingServer: true,
timeout: 120_000,
env: {
NEXT_PUBLIC_API_BASE_URL: API_BASE,
},
},
],
});
+5
View File
@@ -0,0 +1,5 @@
export default {
plugins: {
'@tailwindcss/postcss': {},
},
};
-142
View File
@@ -1,142 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
"""
Runs preprocessing subprocess for datasets.
"""
from __future__ import annotations
import logging
import os
import subprocess
import sys
import threading
from pathlib import Path
from collections.abc import Callable
logger = logging.getLogger("fastvideo.studio.preprocess_runner")
def build_preprocess_args(
dataset_id: str,
raw_path: str,
output_dir: str,
workload_type: str,
model_path: str,
dataset_type: str = "merged",
num_gpus: int = 1,
) -> list[str]:
"""Build CLI args for v1_preprocessing_new."""
return [
"--model-path",
model_path,
"--mode",
"preprocess",
"--workload-type",
workload_type,
"--preprocess.video_loader_type",
"torchvision",
"--preprocess.dataset_type",
dataset_type,
"--preprocess.dataset_path",
raw_path,
"--preprocess.dataset_output_dir",
output_dir,
"--preprocess.preprocess_video_batch_size",
"2",
"--preprocess.dataloader_num_workers",
"0",
"--preprocess.max_height",
"480",
"--preprocess.max_width",
"832",
"--preprocess.num_frames",
"77",
"--preprocess.train_fps",
"16",
"--preprocess.samples_per_file",
"8",
"--preprocess.flush_frequency",
"8",
"--preprocess.video_length_tolerance_range",
"5",
]
def run_preprocess(
dataset_id: str,
raw_path: str,
output_dir: str,
workload_type: str,
model_path: str,
dataset_type: str,
num_gpus: int,
log_file_path: str,
on_status_change: Callable[[str, str | None], None],
stop_event: threading.Event,
) -> None:
"""Run preprocessing subprocess. Calls on_status_change(status, error)."""
repo_root = Path(__file__).resolve().parent.parent
preprocess_module = "fastvideo.pipelines.preprocess.v1_preprocessing_new"
args = build_preprocess_args(
dataset_id=dataset_id,
raw_path=raw_path,
output_dir=output_dir,
workload_type=workload_type,
model_path=model_path,
dataset_type=dataset_type,
num_gpus=num_gpus,
)
cmd = [
sys.executable,
"-m",
"torch.distributed.run",
"--nproc_per_node",
str(num_gpus),
"--nnodes",
"1",
"-m",
preprocess_module,
] + args
env = os.environ.copy()
env["TOKENIZERS_PARALLELISM"] = "false"
try:
on_status_change("preprocessing", None)
with open(log_file_path, "w", encoding="utf-8") as log_file:
proc = subprocess.Popen(
cmd,
cwd=str(repo_root),
env=env,
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
text=True,
bufsize=1,
)
assert proc.stdout is not None
for line in iter(proc.stdout.readline, ""):
if stop_event.is_set():
proc.terminate()
try:
proc.wait(timeout=30)
except subprocess.TimeoutExpired:
proc.kill()
on_status_change("stopped", "Preprocessing stopped by user")
return
line = line.rstrip()
if line:
log_file.write(line + "\n")
log_file.flush()
proc.wait()
exit_code = proc.returncode or 0
if exit_code == 0:
on_status_change("ready", None)
else:
on_status_change("failed", f"Preprocessing exited with code {exit_code}")
except Exception as exc:
on_status_change("failed", f"{type(exc).__name__}: {exc}")
logger.exception("Preprocessing failed for dataset %s", dataset_id)
+54 -20
View File
@@ -32,8 +32,10 @@ from fastapi.responses import FileResponse
from fastvideo.registry import (get_registered_model_paths, get_registered_models_with_workloads)
from fastvideo_studio.database import Database, _get_db_path
from fastvideo_studio.gpu import get_gpu_snapshot
from fastvideo_studio.job_runner import JobRunner, JobStatus
from fastvideo_studio.models import (CreateDatasetRequest, CreateJobRequest, SettingsUpdate, UpdateCaptionRequest)
from fastvideo_studio.models import (CreateDatasetRequest, CreateJobRequest, SettingsUpdate, UpdateCaptionRequest,
model_label)
logging.basicConfig(
level=logging.INFO,
@@ -43,15 +45,9 @@ logger = logging.getLogger("fastvideo.studio.api")
DEFAULT_OUTPUT_DIR = os.path.join(os.path.dirname(__file__), "..", "outputs", "ui_jobs")
def _get_model_label(model_path: str) -> str:
"""Derive a readable label from a HF model path."""
return model_path.split("/")[-1].replace("-", " ").replace("_", " ")
_available_models: list[dict[str, str]] = [{
"id": path,
"label": _get_model_label(path)
"label": model_label(path)
} for path in get_registered_model_paths()]
job_runner: JobRunner
@@ -100,6 +96,12 @@ def update_settings(settings: SettingsUpdate) -> dict[str, Any]:
return database.get_settings()
@app.get("/api/gpus")
def list_gpus() -> dict[str, Any]:
"""Return an NVML snapshot of every GPU on the API server host."""
return get_gpu_snapshot()
@app.get("/api/models")
def list_models(workload_type: str | None = None) -> list[dict[str, Any]]:
"""Return the catalogue of available video-generation models.
@@ -148,7 +150,7 @@ async def upload_image(file: Annotated[UploadFile, File()], ) -> dict[str, str]:
return {"path": os.path.abspath(dest_path)}
ALLOWED_VIDEO_EXTENSIONS = {".mp4", ".webm", ".avi", ".mov"}
ALLOWED_VIDEO_EXTENSIONS = {".mp4", ".webm", ".avi", ".mov", ".mkv"}
def _filter_video_files(files: list[UploadFile]) -> list[UploadFile]:
@@ -156,6 +158,23 @@ def _filter_video_files(files: list[UploadFile]) -> list[UploadFile]:
return [f for f in files if Path(f.filename or "").suffix.lower() in ALLOWED_VIDEO_EXTENSIONS]
def _path_is_within(child: str, parent: str) -> bool:
"""True if ``child`` resolves to a location inside ``parent``."""
try:
parent_real = os.path.realpath(parent)
return os.path.commonpath([os.path.realpath(child), parent_real]) == parent_real
except (ValueError, OSError):
return False
def _staging_base_path() -> str:
"""Root under which raw uploads are staged (settings override or default)."""
settings = database.get_settings() if database is not None else {}
raw_path = (settings.get("datasetUploadPath") or settings.get("dataset_upload_path") or "")
base_path = (raw_path.strip() if raw_path and isinstance(raw_path, str) else "")
return datasets_upload_dir if not base_path else os.path.abspath(base_path)
@app.post("/api/upload-raw-dataset")
async def upload_raw_dataset(files: Annotated[list[UploadFile], File()], ) -> dict[str, Any]:
"""
@@ -167,10 +186,7 @@ async def upload_raw_dataset(files: Annotated[list[UploadFile], File()], ) -> di
status_code=503,
detail="Database not initialized",
)
settings = database.get_settings()
raw_path = (settings.get("datasetUploadPath") or settings.get("dataset_upload_path") or "")
base_path = (raw_path.strip() if raw_path and isinstance(raw_path, str) else "")
base_path = (datasets_upload_dir if not base_path else os.path.abspath(base_path))
base_path = _staging_base_path()
if not base_path:
raise HTTPException(
status_code=503,
@@ -250,6 +266,19 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
f"Valid options: {sorted(valid_ids)}"),
)
# Training jobs reference a dataset by id; resolve it to the on-disk media
# directory the trainer reads (the UI has no free-text path field). Falls
# through unchanged for inference and for anything already a real path.
def _resolve_dataset_path(value: str) -> str:
media_dir = _dataset_media_dir(value) if value else ""
return media_dir if media_dir and os.path.isdir(media_dir) else value
data_path = req.data_path or ""
validation_dataset_file = req.validation_dataset_file or ""
if job_type != "inference":
data_path = _resolve_dataset_path(data_path)
validation_dataset_file = _resolve_dataset_path(validation_dataset_file)
job = job_runner.create_job(
job_id=str(uuid.uuid4()),
model_id=req.model_id,
@@ -257,14 +286,13 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
workload_type=req.workload_type or "t2v",
job_type=job_type,
image_path=req.image_path or "",
data_path=req.data_path or "",
data_path=data_path,
max_train_steps=req.max_train_steps,
train_batch_size=req.train_batch_size,
learning_rate=req.learning_rate,
num_latent_t=req.num_latent_t,
validation_dataset_file=req.validation_dataset_file or "",
validation_dataset_file=validation_dataset_file,
lora_rank=req.lora_rank,
ltx2_first_frame_conditioning_p=req.ltx2_first_frame_conditioning_p,
negative_prompt=req.negative_prompt,
num_inference_steps=req.num_inference_steps,
num_frames=req.num_frames,
@@ -287,8 +315,6 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
dmd_use_vsa=req.dmd_use_vsa,
dmd_vsa_sparsity=req.dmd_vsa_sparsity,
dmd_denoising_steps=req.dmd_denoising_steps or "1000,757,522",
min_timestep_ratio=req.min_timestep_ratio,
max_timestep_ratio=req.max_timestep_ratio,
real_score_guidance_scale=req.real_score_guidance_scale,
generator_update_interval=req.generator_update_interval,
real_score_model_path=req.real_score_model_path or "",
@@ -418,6 +444,13 @@ def create_dataset(req: CreateDatasetRequest) -> dict[str, Any]:
status_code=400,
detail="No media files found. Ensure at least one image or video.",
)
# upload_path must be a staging dir produced by /api/upload-raw-dataset,
# not an arbitrary client path — the code below copies then rmtrees it.
if not _path_is_within(req.upload_path, _staging_base_path()):
raise HTTPException(
status_code=400,
detail="upload_path must be a staged upload directory.",
)
dataset_id = str(uuid.uuid4())
created_at = time.time()
dest_dir = os.path.join(datasets_upload_dir, dataset_id)
@@ -491,8 +524,9 @@ def serve_dataset_media(dataset_id: str, file_name: str) -> FileResponse:
ds = database.get_dataset(dataset_id)
if ds is None:
raise HTTPException(status_code=404, detail="Dataset not found")
media_path = os.path.join(_dataset_media_dir(dataset_id), file_name)
if not os.path.isfile(media_path):
media_dir = _dataset_media_dir(dataset_id)
media_path = os.path.realpath(os.path.join(media_dir, file_name))
if not (_path_is_within(media_path, media_dir) and os.path.isfile(media_path)):
raise HTTPException(status_code=404, detail="File not found")
import mimetypes
mime, _ = mimetypes.guess_type(media_path)
-66
View File
@@ -1,66 +0,0 @@
/* CSS Variables */
:root {
--header-height: 74px;
--bg: #0f1117;
--surface: #1a1d27;
--border: #2a2d3a;
--text: #e4e4e7;
--text-dim: #9ca3af;
--accent: #356cff;
--accent-h: #818cf8;
--green: #22c55e;
--red: #ef4444;
--yellow: #eab308;
--blue: #3b82f6;
--radius: 8px;
--shadow: 0 2px 8px rgba(0, 0, 0, 0.35);
}
/* Reset */
*,
*::before,
*::after {
box-sizing: border-box;
margin: 0;
padding: 0;
}
/* Base body styles */
body {
font-family:
"Inter",
system-ui,
-apple-system,
sans-serif;
background: var(--bg);
color: var(--text);
line-height: 1.6;
min-height: 100vh;
display: flex;
flex-direction: column;
}
/* Base form input styles (used globally) */
select,
input,
textarea {
background: var(--bg);
color: var(--text);
border: 1px solid var(--border);
border-radius: var(--radius);
padding: 0.55rem 0.75rem;
font-size: 0.9rem;
font-family: inherit;
outline: none;
transition: border-color 0.15s;
}
select:focus,
input:focus,
textarea:focus {
border-color: var(--accent);
}
textarea {
resize: vertical;
}
-13
View File
@@ -1,13 +0,0 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<link rel="icon" href="%sveltekit.assets%/fastvideo.ico" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<title>FastVideo</title>
%sveltekit.head%
</head>
<body class="antialiased" data-sveltekit-preload-data="hover">
<div style="display: contents">%sveltekit.body%</div>
</body>
</html>
@@ -0,0 +1,64 @@
import { act, fireEvent, render, screen } from '@testing-library/react';
import { describe, expect, it, vi } from 'vitest';
import { HeaderActionsProvider } from '@/components/shell/HeaderActionsContext';
import { getDatasets, type Dataset } from '@/lib/api';
import DatasetsPage from './page';
vi.mock('@/lib/api', () => ({
getDatasets: vi.fn(),
}));
vi.mock('@/components/datasets/AddDatasetButton', () => ({
default: () => null,
}));
vi.mock('@/components/datasets/CreateDatasetModal', () => ({
default: () => null,
}));
vi.mock('@/components/datasets/DatasetCard', () => ({
default: ({ dataset }: { dataset: Dataset }) => <div>{dataset.name}</div>,
}));
function renderPage() {
return render(
<HeaderActionsProvider>
<DatasetsPage />
</HeaderActionsProvider>,
);
}
describe('DatasetsPage', () => {
it('shows loading content before the initial request settles', async () => {
let resolveDatasets: (datasets: Dataset[]) => void = () => {};
vi.mocked(getDatasets).mockReturnValue(
new Promise<Dataset[]>((resolve) => {
resolveDatasets = resolve;
}),
);
renderPage();
expect(screen.getByLabelText('Loading datasets')).toBeInTheDocument();
expect(screen.queryByText('No datasets yet.')).not.toBeInTheDocument();
act(() => resolveDatasets([]));
expect(await screen.findByText('No datasets yet.')).toBeInTheDocument();
});
it('shows API failures separately from an empty list and retries', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
vi.mocked(getDatasets).mockRejectedValueOnce(new Error('network down'));
renderPage();
expect(
await screen.findByText(/Could not load datasets from the Studio API/),
).toBeInTheDocument();
expect(screen.queryByText('No datasets yet.')).not.toBeInTheDocument();
vi.mocked(getDatasets).mockResolvedValueOnce([]);
fireEvent.click(screen.getByRole('button', { name: 'Try Again' }));
expect(await screen.findByText('No datasets yet.')).toBeInTheDocument();
});
});
@@ -0,0 +1,136 @@
'use client';
import * as React from 'react';
import { AlertTriangle } from 'lucide-react';
import AddDatasetButton from '@/components/datasets/AddDatasetButton';
import CreateDatasetModal from '@/components/datasets/CreateDatasetModal';
import DatasetCard from '@/components/datasets/DatasetCard';
import { HeaderActions } from '@/components/shell/HeaderActionsContext';
import { Card } from '@/components/ui/card';
import { Button } from '@/components/ui/button';
import { useStore } from '@/hooks/useStore';
import { getDatasets } from '@/lib/api';
import type { Dataset } from '@/lib/api';
import {
setActiveDataset,
setActiveDatasetId,
} from '@/stores/activeDataset';
import {
createDatasetModalStore,
setCreateDatasetModalOpen,
} from '@/stores/createDatasetModalOpen';
export default function DatasetsPage() {
const [datasets, setDatasets] = React.useState<Dataset[]>([]);
const [isInitialLoading, setIsInitialLoading] = React.useState(true);
const [error, setError] = React.useState<string | null>(null);
const { open } = useStore(createDatasetModalStore);
const fetchSequence = React.useRef(0);
const fetchDatasets = React.useCallback(async () => {
const sequence = ++fetchSequence.current;
try {
const next = await getDatasets();
if (sequence === fetchSequence.current) {
setDatasets(next);
setError(null);
}
} catch (err) {
console.error('Failed to fetch datasets:', err);
if (sequence === fetchSequence.current) {
setError(
'Could not load datasets from the Studio API. Check the server and try again.',
);
}
} finally {
if (sequence === fetchSequence.current) setIsInitialLoading(false);
}
}, []);
React.useEffect(() => {
fetchDatasets();
}, [fetchDatasets]);
function handleSelectDataset(ds: Dataset) {
setActiveDataset(ds);
setActiveDatasetId(ds.id);
}
return (
<>
<HeaderActions>
<AddDatasetButton />
</HeaderActions>
<div className="mx-auto flex w-full max-w-[850px] flex-col gap-6 px-4 pb-12">
<Card className="p-6">
<div aria-busy={isInitialLoading}>
{isInitialLoading ? (
<div
aria-label="Loading datasets"
className="flex flex-col gap-3 py-2"
>
{[0, 1, 2].map((item) => (
<div
key={item}
className="h-24 animate-pulse rounded-lg border border-border bg-muted/50"
/>
))}
</div>
) : error && datasets.length === 0 ? (
<div
role="alert"
className="flex flex-col items-center gap-3 py-8 text-center"
>
<AlertTriangle
className="size-6 text-destructive"
aria-hidden
/>
<p className="max-w-md text-sm text-muted-foreground">
{error}
</p>
<Button type="button" variant="outline" onClick={fetchDatasets}>
Try Again
</Button>
</div>
) : (
<>
{error && (
<p
role="status"
className="mb-3 rounded-lg border border-amber-500/50 bg-amber-500/10 px-3 py-2 text-sm text-foreground"
>
Dataset updates are temporarily unavailable. Showing the
most recent results.
</p>
)}
{datasets.length === 0 ? (
<p className="py-8 text-center text-muted-foreground">
No datasets yet.
</p>
) : (
datasets.map((ds) => (
<DatasetCard
key={ds.id}
dataset={ds}
onUpdated={fetchDatasets}
onSelect={() => handleSelectDataset(ds)}
/>
))
)}
</>
)}
</div>
</Card>
</div>
<CreateDatasetModal
isOpen={open}
onClose={() => setCreateDatasetModalOpen(false)}
onSuccess={() => {
fetchDatasets();
setCreateDatasetModalOpen(false);
}}
/>
</>
);
}
@@ -0,0 +1,16 @@
'use client';
import CreateJobButton from '@/components/jobs/CreateJobButton';
import { HeaderActions } from '@/components/shell/HeaderActionsContext';
import JobQueue from '@/components/jobs/JobQueue';
export default function DistillationPage() {
return (
<>
<HeaderActions>
<CreateJobButton jobType="distillation" />
</HeaderActions>
<JobQueue jobType="distillation" />
</>
);
}
@@ -0,0 +1,21 @@
'use client';
import CreateJobButton from '@/components/jobs/CreateJobButton';
import { HeaderActions } from '@/components/shell/HeaderActionsContext';
import JobQueue from '@/components/jobs/JobQueue';
import type { JobType } from '@/lib/types';
// 'lora' is a backend job_type (LoRA finetunes) that isn't part of the
// JobType union; the finetuning queue lists both alongside full finetunes.
const FINETUNING_LIST = ['finetuning', 'lora'] as JobType[];
export default function FinetuningPage() {
return (
<>
<HeaderActions>
<CreateJobButton jobType="finetuning" />
</HeaderActions>
<JobQueue jobType="finetuning" jobTypesForList={FINETUNING_LIST} />
</>
);
}
@@ -0,0 +1,76 @@
import { fireEvent, render, screen } from '@testing-library/react';
import { describe, expect, it, vi } from 'vitest';
import GalleryPage from './page';
import { HeaderActionsProvider } from '@/components/shell/HeaderActionsContext';
import { getJobsList } from '@/lib/api';
import type { Job } from '@/lib/types';
import { makeJob as makeBaseJob } from '@/test/factories';
vi.mock('@/lib/api', () => ({
getJobsList: vi.fn(),
getJobVideoUrl: (id: string) => `http://test.local/api/jobs/${id}/video`,
}));
const makeJob = (overrides: Partial<Job> = {}): Job =>
makeBaseJob({
model_id: 'wan',
prompt: 'a cat surfing a wave',
status: 'completed',
created_at: 1,
finished_at: 2,
output_path: '/out/clip.mp4',
...overrides,
});
function renderGallery() {
return render(
<HeaderActionsProvider>
<GalleryPage />
</HeaderActionsProvider>,
);
}
describe('GalleryPage', () => {
it('renders a grid item for a completed inference job', async () => {
vi.mocked(getJobsList).mockResolvedValue([
makeJob({ prompt: 'a cat surfing a wave' }),
]);
renderGallery();
expect(await screen.findByText('a cat surfing a wave')).toBeInTheDocument();
expect(getJobsList).toHaveBeenCalledWith('inference');
});
it('provides video controls and a visible fallback when media fails', async () => {
vi.mocked(getJobsList).mockResolvedValue([makeJob()]);
renderGallery();
const video = await screen.findByLabelText(
'Generated video: a cat surfing a wave',
);
expect(video).toHaveAttribute('controls');
fireEvent.error(video);
expect(screen.getByText('Preview unavailable')).toBeInTheDocument();
expect(
screen.getByText('The generated file could not be loaded.'),
).toBeInTheDocument();
});
it('shows the empty state when no completed videos exist', async () => {
vi.mocked(getJobsList).mockResolvedValue([
makeJob({ status: 'running', output_path: null }),
]);
renderGallery();
expect(
await screen.findByText('No completed videos yet'),
).toBeInTheDocument();
expect(
screen.queryByText('a cat surfing a wave'),
).not.toBeInTheDocument();
});
});
@@ -0,0 +1,161 @@
'use client';
import { AlertTriangle, ImageOff, Loader2 } from 'lucide-react';
import { useEffect, useState } from 'react';
import { Button } from '@/components/ui/button';
import { Card } from '@/components/ui/card';
import { getJobVideoUrl, getJobsList } from '@/lib/api';
import type { Job } from '@/lib/types';
function isImage(job: Job): boolean {
return job.output_path?.toLowerCase().endsWith('.png') ?? false;
}
function GalleryMedia({ job }: { job: Job }) {
const [failed, setFailed] = useState(false);
if (failed) {
return (
<div
role="status"
className="flex h-full flex-col items-center justify-center gap-2 px-4 text-center text-muted-foreground"
>
<ImageOff className="size-7" aria-hidden />
<span className="text-sm font-medium">Preview unavailable</span>
<span className="text-xs">
The generated file could not be loaded.
</span>
</div>
);
}
if (isImage(job)) {
return (
// eslint-disable-next-line @next/next/no-img-element
<img
src={getJobVideoUrl(job.id)}
alt={job.prompt}
className="block h-full w-full object-contain"
loading="lazy"
onError={() => setFailed(true)}
/>
);
}
return (
<video
src={getJobVideoUrl(job.id)}
aria-label={
job.prompt ? `Generated video: ${job.prompt}` : 'Generated video'
}
className="block h-full w-full object-contain"
controls
muted
loop
playsInline
preload="metadata"
onError={() => setFailed(true)}
/>
);
}
export default function GalleryPage() {
const [jobs, setJobs] = useState<Job[]>([]);
const [isLoading, setIsLoading] = useState(true);
const [error, setError] = useState<string | null>(null);
const [reloadKey, setReloadKey] = useState(0);
useEffect(() => {
let cancelled = false;
async function load() {
try {
const list = await getJobsList('inference');
if (cancelled) return;
const sorted = [...list].sort(
(a, b) =>
(b.finished_at ?? b.created_at ?? 0) -
(a.finished_at ?? a.created_at ?? 0),
);
setJobs(sorted);
} catch (e) {
if (!cancelled) {
setError(e instanceof Error ? e.message : 'Failed to load jobs');
}
} finally {
if (!cancelled) setIsLoading(false);
}
}
void load();
return () => {
cancelled = true;
};
}, [reloadKey]);
function retry() {
setError(null);
setIsLoading(true);
setReloadKey((k) => k + 1);
}
const galleryJobs = jobs.filter(
(j) =>
j.status === 'completed' &&
j.output_path &&
(j.job_type === 'inference' || !j.job_type),
);
return (
<div className="mx-auto w-full max-w-[1200px] px-4 pb-12 pt-4">
<Card className="p-6">
<h2 className="mb-1 text-2xl font-semibold text-foreground">Gallery</h2>
<p className="mb-6 text-sm text-muted-foreground">
Generated videos from completed inference jobs. Captions show the
prompt used for each generation.
</p>
{isLoading ? (
<div className="flex items-center gap-3 p-8 text-muted-foreground">
<Loader2 className="h-6 w-6 animate-spin text-primary" />
<span>Loading gallery…</span>
</div>
) : error ? (
<div
role="alert"
className="flex flex-col items-center gap-3 py-8 text-center"
>
<AlertTriangle className="size-6 text-destructive" aria-hidden />
<p className="max-w-md text-sm text-muted-foreground">{error}</p>
<Button type="button" variant="outline" onClick={retry}>
Try Again
</Button>
</div>
) : galleryJobs.length === 0 ? (
<p className="py-8 text-center text-muted-foreground">
No completed videos yet
</p>
) : (
<div className="grid grid-cols-[repeat(auto-fill,minmax(280px,1fr))] gap-5">
{galleryJobs.map((job) => (
<article
key={job.id}
className="flex flex-col overflow-hidden rounded-lg border border-border bg-background"
>
<div className="relative aspect-video overflow-hidden bg-muted">
<GalleryMedia job={job} />
</div>
<p
className="line-clamp-3 border-t border-border px-4 py-3 text-sm text-muted-foreground"
title={job.prompt}
>
{job.prompt || '—'}
</p>
</article>
))}
</div>
)}
</Card>
</div>
);
}
+238
View File
@@ -0,0 +1,238 @@
@import "tailwindcss";
@custom-variant dark (&:is(.dark *));
/*
* FastVideo Studio theme — shared with the Dreamverse design system:
* slate light/dark palettes, #356cff accent blue, IBM Plex type. The studio
* defaults to dark (see the theme init script in layout.tsx); the toggle in
* the header persists the choice to localStorage.
*/
:root {
color-scheme: light;
--header-height: 74px;
--accent-blue: #356cff;
--background: #f5f4f4;
--foreground: #0f172a;
--card: #ffffff;
--card-foreground: #0f172a;
--popover: #ffffff;
--popover-foreground: #0f172a;
--primary: #0f172a;
--primary-foreground: #f8fafc;
--secondary: #f1f5f9;
--secondary-foreground: #1e293b;
--muted: #f1f5f9;
--muted-foreground: #64748b;
--accent: #e2e8f0;
--accent-foreground: #0f172a;
--destructive: #ef4444;
--destructive-foreground: #ffffff;
--border: #e2e8f0;
--input: #cbd5e1;
--ring: #1d4ed8;
--radius: 0.5rem;
}
.dark {
color-scheme: dark;
--accent-blue: #356cff;
--background: #0f172a;
--foreground: #e2e8f0;
--card: #0f172a;
--card-foreground: #f1f5f9;
--popover: #0f172a;
--popover-foreground: #f1f5f9;
--primary: #f1f5f9;
--primary-foreground: #0f172a;
--secondary: #1e293b;
--secondary-foreground: #f1f5f9;
--muted: #1e293b;
--muted-foreground: #94a3b8;
--accent: #1e293b;
--accent-foreground: #f1f5f9;
--destructive: #991b1b;
--destructive-foreground: #fecaca;
--border: #334155;
--input: #334155;
--ring: #7dd3fc;
}
@theme inline {
--font-sans: var(--font-plex-sans), ui-sans-serif, system-ui, sans-serif;
--font-mono: var(--font-plex-mono), ui-monospace, monospace;
--color-accent-blue: var(--accent-blue);
--color-background: var(--background);
--color-foreground: var(--foreground);
--color-card: var(--card);
--color-card-foreground: var(--card-foreground);
--color-popover: var(--popover);
--color-popover-foreground: var(--popover-foreground);
--color-primary: var(--primary);
--color-primary-foreground: var(--primary-foreground);
--color-secondary: var(--secondary);
--color-secondary-foreground: var(--secondary-foreground);
--color-muted: var(--muted);
--color-muted-foreground: var(--muted-foreground);
--color-accent: var(--accent);
--color-accent-foreground: var(--accent-foreground);
--color-destructive: var(--destructive);
--color-destructive-foreground: var(--destructive-foreground);
--color-border: var(--border);
--color-input: var(--input);
--color-ring: var(--ring);
--radius-sm: calc(var(--radius) - 4px);
--radius-md: calc(var(--radius) - 2px);
--radius-lg: var(--radius);
--radius-xl: calc(var(--radius) + 4px);
}
/* ——— Resets ——— */
html,
body {
min-height: 100%;
overflow-x: clip;
}
html {
background: var(--background);
}
body {
margin: 0;
line-height: 1.6;
color: var(--foreground);
font-family: var(--font-plex-sans), ui-sans-serif, system-ui, sans-serif;
background: var(--background);
}
/* Dreamverse's dark-mode backdrop: subtle radial glows over a deep fade. */
body::after {
content: "";
position: fixed;
inset: 0;
z-index: -1;
opacity: 0;
pointer-events: none;
transition: opacity 300ms ease;
background:
radial-gradient(circle at top, rgba(56, 189, 248, 0.14), transparent 36%),
radial-gradient(circle at right top, rgba(129, 140, 248, 0.12), transparent 28%),
linear-gradient(180deg, #020617 0%, #000000 100%);
background-attachment: fixed;
}
.dark body::after {
opacity: 1;
}
a {
color: inherit;
}
:where(
a,
button,
input,
textarea,
select,
summary,
[role="button"],
[role="menuitem"],
[role="slider"],
[tabindex]
):focus-visible {
outline: 3px solid var(--ring) !important;
outline-offset: 2px !important;
}
summary {
list-style: none;
}
summary::-webkit-details-marker {
display: none;
}
::selection {
background: rgba(56, 189, 248, 0.25);
}
.dark ::selection {
background: rgba(56, 189, 248, 0.35);
color: #ffffff;
}
/* Smooth the light/dark switch (class applied briefly by the theme toggle). */
html.theme-transition,
html.theme-transition *,
html.theme-transition *::before,
html.theme-transition *::after {
transition:
background-color 300ms ease,
color 300ms ease,
border-color 300ms ease,
box-shadow 300ms ease,
fill 300ms ease,
stroke 300ms ease !important;
}
/* ——— Base layer ——— */
@layer base {
button,
input,
textarea,
select {
font: inherit;
}
img,
svg,
video,
canvas {
display: block;
max-width: 100%;
}
* {
@apply border-border;
}
body {
@apply bg-background text-foreground;
}
}
@@ -0,0 +1,55 @@
import { readFileSync } from 'node:fs';
import { join } from 'node:path';
import { describe, expect, it } from 'vitest';
const css = readFileSync(join(process.cwd(), 'src/app/globals.css'), 'utf8');
function token(block: string, name: string): string {
const match = block.match(new RegExp(`--${name}:\\s*(#[0-9a-fA-F]{6})`));
if (!match) throw new Error(`Missing --${name} token`);
return match[1];
}
function luminance(hex: string): number {
const channels = hex
.slice(1)
.match(/.{2}/g)!
.map((channel) => parseInt(channel, 16) / 255)
.map((channel) =>
channel <= 0.04045
? channel / 12.92
: ((channel + 0.055) / 1.055) ** 2.4,
);
return (
0.2126 * channels[0] + 0.7152 * channels[1] + 0.0722 * channels[2]
);
}
function contrast(first: string, second: string): number {
const firstLuminance = luminance(first);
const secondLuminance = luminance(second);
return (
(Math.max(firstLuminance, secondLuminance) + 0.05) /
(Math.min(firstLuminance, secondLuminance) + 0.05)
);
}
describe('global focus styles', () => {
it('keeps focus tokens above 3:1 against both page themes', () => {
const light = css.match(/:root\s*{([\s\S]*?)\n}/)?.[1] ?? '';
const dark = css.match(/\.dark\s*{([\s\S]*?)\n}/)?.[1] ?? '';
expect(contrast(token(light, 'ring'), token(light, 'background'))).toBeGreaterThanOrEqual(
3,
);
expect(contrast(token(dark, 'ring'), token(dark, 'background'))).toBeGreaterThanOrEqual(
3,
);
});
it('applies a non-animated three-pixel outline to focus-visible controls', () => {
expect(css).toContain('):focus-visible {');
expect(css).toContain('outline: 3px solid var(--ring) !important;');
expect(css).toContain('outline-offset: 2px !important;');
});
});
@@ -0,0 +1,11 @@
'use client';
import GpuGrid from '@/components/system/GpuGrid';
export default function GpusPage() {
return (
<div className="mx-auto flex w-full max-w-[1100px] flex-col gap-6 px-4 pb-12 pt-6">
<GpuGrid />
</div>
);
}
@@ -0,0 +1,16 @@
'use client';
import CreateJobButton from '@/components/jobs/CreateJobButton';
import { HeaderActions } from '@/components/shell/HeaderActionsContext';
import JobQueue from '@/components/jobs/JobQueue';
export default function InferencePage() {
return (
<>
<HeaderActions>
<CreateJobButton jobType="inference" />
</HeaderActions>
<JobQueue jobType="inference" />
</>
);
}
+52
View File
@@ -0,0 +1,52 @@
import type { Metadata } from 'next';
import { IBM_Plex_Mono, IBM_Plex_Sans } from 'next/font/google';
import { AppShell } from '@/components/shell/AppShell';
import './globals.css';
const plexSans = IBM_Plex_Sans({
subsets: ['latin'],
weight: ['400', '500', '600', '700'],
variable: '--font-plex-sans',
display: 'swap',
});
const plexMono = IBM_Plex_Mono({
subsets: ['latin'],
weight: ['400', '500', '600'],
variable: '--font-plex-mono',
display: 'swap',
});
export const metadata: Metadata = {
title: 'FastVideo Studio',
icons: { icon: '/fastvideo.ico' },
};
export default function RootLayout({
children,
}: {
children: React.ReactNode;
}) {
return (
<html
lang="en"
suppressHydrationWarning
className={`${plexSans.variable} ${plexMono.variable}`}
>
<head>
{/* The studio defaults to dark; apply the stored choice before paint. */}
<script
id="theme-init-script"
suppressHydrationWarning
dangerouslySetInnerHTML={{
__html: `(function(){var dark=true;try{dark=localStorage.getItem('theme')!=='light'}catch(e){}if(dark)document.documentElement.classList.add('dark')})()`,
}}
/>
</head>
<body className="antialiased">
<AppShell>{children}</AppShell>
</body>
</html>
);
}
@@ -0,0 +1,5 @@
import { redirect } from 'next/navigation';
export default function Page() {
redirect('/finetuning');
}
+5
View File
@@ -0,0 +1,5 @@
import { redirect } from 'next/navigation';
export default function Page() {
redirect('/inference');
}
@@ -0,0 +1,74 @@
import { fireEvent, render, screen } from '@testing-library/react';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import { HeaderActionsProvider } from '@/components/shell/HeaderActionsContext';
import { resetToDefaults, updateOption } from '@/stores/defaultOptions';
import SettingsPage from './page';
vi.mock('@/lib/api', () => ({
getModels: vi.fn().mockResolvedValue([]),
getSettings: vi.fn().mockResolvedValue({}),
updateSettings: vi.fn().mockResolvedValue({}),
}));
// Keep the real store (so `useStore` resolves) but spy on the mutators.
vi.mock('@/stores/defaultOptions', async (importOriginal) => {
const actual =
await importOriginal<typeof import('@/stores/defaultOptions')>();
return {
...actual,
updateOption: vi.fn(),
resetToDefaults: vi.fn(),
};
});
function renderPage() {
return render(
<HeaderActionsProvider>
<SettingsPage />
</HeaderActionsProvider>,
);
}
describe('Settings page', () => {
beforeEach(() => {
vi.clearAllMocks();
});
it('calls updateOption when a toggle changes', () => {
renderPage();
fireEvent.click(
screen.getByRole('switch', { name: 'Auto Start Job on Create' }),
);
expect(updateOption).toHaveBeenCalledWith('autoStartJob', true);
});
it('calls updateOption when a slider changes', () => {
renderPage();
// First slider in DOM order is "Frames".
const slider = screen.getAllByRole('slider')[0];
fireEvent.keyDown(slider, { key: 'ArrowRight' });
expect(updateOption).toHaveBeenCalledWith('numFrames', expect.any(Number));
});
it('gives every slider an accessible name', () => {
renderPage();
const sliders = screen.getAllByRole('slider');
expect(sliders).toHaveLength(11);
for (const slider of sliders) {
expect(slider).toHaveAccessibleName();
}
expect(screen.getByRole('slider', { name: 'Frames' })).toBeInTheDocument();
expect(
screen.getByRole('slider', { name: 'Guidance Scale' }),
).toBeInTheDocument();
});
it('calls resetToDefaults when Reset to Defaults is clicked', () => {
renderPage();
fireEvent.click(screen.getByRole('button', { name: 'Reset to Defaults' }));
expect(resetToDefaults).toHaveBeenCalledTimes(1);
});
});
@@ -0,0 +1,358 @@
'use client';
import * as React from 'react';
import {
FieldRow,
NumberRow,
SliderRow,
ToggleRow,
} from '@/components/form-rows';
import { Button } from '@/components/ui/button';
import { Card, CardContent } from '@/components/ui/card';
import { Input } from '@/components/ui/input';
import { Label } from '@/components/ui/label';
import { NativeSelect } from '@/components/ui/native-select';
import { Separator } from '@/components/ui/separator';
import { Switch } from '@/components/ui/switch';
import { useStore } from '@/hooks/useStore';
import { getModels, type Model } from '@/lib/api';
import {
defaultOptionsStore,
resetToDefaults,
updateOption,
} from '@/stores/defaultOptions';
function TextSettingRow({
id,
label,
value,
placeholder,
onCommit,
}: {
id: string;
label: string;
value: string;
placeholder?: string;
onCommit: (v: string) => void;
}) {
// Buffer keystrokes locally and persist on blur/Enter so each character
// doesn't fire a settings PUT (or a localStorage write).
const [draft, setDraft] = React.useState<string | null>(null);
return (
<FieldRow htmlFor={id} label={label}>
<Input
id={id}
type="text"
className="font-mono text-sm"
value={draft ?? value}
onChange={(e) => setDraft(e.target.value)}
onBlur={() => {
if (draft !== null && draft !== value) onCommit(draft);
setDraft(null);
}}
onKeyDown={(e) => {
if (e.key === 'Enter') e.currentTarget.blur();
}}
placeholder={placeholder}
/>
</FieldRow>
);
}
function SelectRow({
id,
label,
value,
onChange,
children,
}: {
id: string;
label: string;
value: string;
onChange: (v: string) => void;
children: React.ReactNode;
}) {
return (
<FieldRow htmlFor={id} label={label}>
<NativeSelect
id={id}
value={value}
onChange={(e) => onChange(e.target.value)}
>
{children}
</NativeSelect>
</FieldRow>
);
}
export default function SettingsPage() {
const { options } = useStore(defaultOptionsStore);
const [models, setModels] = React.useState<
Record<'t2v' | 'i2v' | 't2i', Model[]>
>({ t2v: [], i2v: [], t2i: [] });
React.useEffect(() => {
Promise.all([getModels('t2v'), getModels('i2v'), getModels('t2i')])
.then(([t2v, i2v, t2i]) => setModels({ t2v, i2v, t2i }))
.catch((e) => console.error('Failed to load models:', e));
}, []);
return (
<div className="mx-auto flex w-full max-w-[850px] flex-col gap-6 px-4 pb-12 pt-6">
<Card>
<CardContent className="space-y-4 p-6">
<h2 className="text-lg font-semibold">Behavior</h2>
<div className="flex items-center justify-between gap-4">
<Label
htmlFor="settings-auto-start-job"
className="pl-0.5 text-xs font-normal tracking-wide text-muted-foreground"
>
Auto Start Job on Create
</Label>
<Switch
id="settings-auto-start-job"
checked={options.autoStartJob}
onCheckedChange={(v) => updateOption('autoStartJob', v)}
/>
</div>
<Separator />
<h2 className="text-lg font-semibold">Paths</h2>
<div className="space-y-4">
<TextSettingRow
id="settings-api-server-base-url"
label="API Server Base URL"
value={options.apiServerBaseUrl ?? ''}
onCommit={(v) => updateOption('apiServerBaseUrl', v)}
placeholder="http://localhost:8189/api"
/>
<TextSettingRow
id="settings-dataset-upload-path"
label="Dataset Upload Path"
value={options.datasetUploadPath ?? ''}
onCommit={(v) => updateOption('datasetUploadPath', v)}
placeholder="outputs/ui_data/uploads/datasets"
/>
</div>
<Separator />
<div className="flex items-center justify-between">
<h2 className="text-lg font-semibold">Default Options</h2>
<Button
type="button"
variant="outline"
size="sm"
onClick={resetToDefaults}
>
Reset to Defaults
</Button>
</div>
<p className="text-sm text-muted-foreground">
These values are used as defaults when creating new jobs.
</p>
<div className="grid gap-x-3 gap-y-2 [grid-template-columns:repeat(auto-fill,minmax(160px,1fr))]">
<SelectRow
id="settings-default-model-t2v"
label="Default Model (T2V)"
value={options.defaultModelIdT2v}
onChange={(v) => updateOption('defaultModelIdT2v', v)}
>
<option value="">None (select when creating job)</option>
{models.t2v.map((model) => (
<option key={model.id} value={model.id}>
{model.label} ({model.id})
</option>
))}
</SelectRow>
<SelectRow
id="settings-default-model-i2v"
label="Default Model (I2V)"
value={options.defaultModelIdI2v}
onChange={(v) => updateOption('defaultModelIdI2v', v)}
>
<option value="">None</option>
{models.i2v.map((model) => (
<option key={model.id} value={model.id}>
{model.label} ({model.id})
</option>
))}
</SelectRow>
<SelectRow
id="settings-default-model-t2i"
label="Default Model (T2I)"
value={options.defaultModelIdT2i}
onChange={(v) => updateOption('defaultModelIdT2i', v)}
>
<option value="">None</option>
{models.t2i.map((model) => (
<option key={model.id} value={model.id}>
{model.label} ({model.id})
</option>
))}
</SelectRow>
<SliderRow
id="settings-num-frames"
label="Frames"
min={1}
max={500}
step={1}
value={options.numFrames}
onChange={(v) => updateOption('numFrames', v)}
/>
<SliderRow
id="settings-height"
label="Height"
min={64}
max={1080}
step={16}
value={options.height}
onChange={(v) => updateOption('height', v)}
/>
<SliderRow
id="settings-width"
label="Width"
min={64}
max={1920}
step={16}
value={options.width}
onChange={(v) => updateOption('width', v)}
/>
<SliderRow
id="settings-num-steps"
label="Inference Steps"
min={1}
max={200}
step={1}
value={options.numInferenceSteps}
onChange={(v) => updateOption('numInferenceSteps', v)}
/>
<SliderRow
id="settings-vsa-sparsity"
label="VSA Sparsity"
title="VSA sparsity (0–1)"
min={0}
max={1}
step={0.05}
value={options.vsaSparsity}
onChange={(v) => updateOption('vsaSparsity', v)}
format={(v) => v.toFixed(2)}
/>
<SliderRow
id="settings-guidance"
label="Guidance Scale"
min={0}
max={20}
step={0.1}
value={options.guidanceScale}
onChange={(v) => updateOption('guidanceScale', v)}
format={(v) => v.toFixed(1)}
/>
<SliderRow
id="settings-guidance-rescale"
label="Guidance Rescale"
title="0 = disabled"
min={0}
max={1}
step={0.05}
value={options.guidanceRescale ?? 0}
onChange={(v) => updateOption('guidanceRescale', v)}
format={(v) => v.toFixed(2)}
/>
<SliderRow
id="settings-tp-size"
label="TP Size"
title="-1 = auto"
min={-1}
max={8}
step={1}
value={options.tpSize}
onChange={(v) => updateOption('tpSize', v)}
format={(v) => (v === -1 ? 'Auto' : String(v))}
/>
<SliderRow
id="settings-sp-size"
label="SP Size"
title="-1 = auto"
min={-1}
max={8}
step={1}
value={options.spSize}
onChange={(v) => updateOption('spSize', v)}
format={(v) => (v === -1 ? 'Auto' : String(v))}
/>
<SliderRow
id="settings-fps"
label="FPS"
min={1}
max={60}
step={1}
value={options.fps ?? 24}
onChange={(v) => updateOption('fps', v)}
/>
<ToggleRow
id="settings-dit-cpu-offload"
label="DiT CPU Offload"
checked={options.ditCpuOffload}
onChange={(v) => updateOption('ditCpuOffload', v)}
/>
<ToggleRow
id="settings-text-encoder-cpu-offload"
label="Text Encoder CPU Offload"
checked={options.textEncoderCpuOffload}
onChange={(v) => updateOption('textEncoderCpuOffload', v)}
/>
<ToggleRow
id="settings-use-fsdp-inference"
label="Use FSDP Inference"
checked={options.useFsdpInference}
onChange={(v) => updateOption('useFsdpInference', v)}
/>
<ToggleRow
id="settings-vae-cpu-offload"
label="VAE CPU Offload"
checked={options.vaeCpuOffload}
onChange={(v) => updateOption('vaeCpuOffload', v)}
/>
<ToggleRow
id="settings-image-encoder-cpu-offload"
label="Image Encoder CPU Offload"
checked={options.imageEncoderCpuOffload}
onChange={(v) => updateOption('imageEncoderCpuOffload', v)}
/>
<ToggleRow
id="settings-enable-torch-compile"
label="Torch Compile"
checked={options.enableTorchCompile}
onChange={(v) => updateOption('enableTorchCompile', v)}
/>
<SliderRow
id="settings-num-gpus"
label="GPUs"
min={1}
max={8}
step={1}
value={options.numGpus}
onChange={(v) => updateOption('numGpus', v)}
/>
<NumberRow
id="settings-seed"
label="Seed"
min={0}
value={options.seed}
onChange={(v) => updateOption('seed', v)}
/>
</div>
</CardContent>
</Card>
</div>
);
}
@@ -0,0 +1,12 @@
'use client';
import { Button } from '@/components/ui/button';
import { setCreateDatasetModalOpen } from '@/stores/createDatasetModalOpen';
export default function AddDatasetButton() {
return (
<Button type="button" onClick={() => setCreateDatasetModalOpen(true)}>
Add Dataset
</Button>
);
}
@@ -0,0 +1,43 @@
import { describe, expect, it, vi } from 'vitest';
import { fireEvent, render, screen } from '@testing-library/react';
import CreateDatasetModal from '@/components/datasets/CreateDatasetModal';
import { createDataset } from '@/lib/api';
vi.mock('@/lib/api', () => ({
createDataset: vi.fn(),
uploadRawDataset: vi.fn(),
}));
const mockedCreateDataset = vi.mocked(createDataset);
describe('CreateDatasetModal', () => {
it('renders nothing when closed', () => {
render(
<CreateDatasetModal isOpen={false} onClose={() => {}} onSuccess={() => {}} />,
);
expect(screen.queryByText('Add Dataset — Raw')).not.toBeInTheDocument();
});
it('renders the form and the default JSON caption upload when open', () => {
render(
<CreateDatasetModal isOpen onClose={() => {}} onSuccess={() => {}} />,
);
expect(screen.getByText('Add Dataset — Raw')).toBeInTheDocument();
expect(screen.getByText('Upload video files')).toBeInTheDocument();
expect(screen.getByText('Upload videos2caption.json')).toBeInTheDocument();
expect(
screen.getByRole('button', { name: 'Create Dataset' }),
).toBeInTheDocument();
});
it('does not create a dataset when the name is empty', () => {
const onSuccess = vi.fn();
render(
<CreateDatasetModal isOpen onClose={() => {}} onSuccess={onSuccess} />,
);
fireEvent.click(screen.getByRole('button', { name: 'Create Dataset' }));
expect(mockedCreateDataset).not.toHaveBeenCalled();
expect(onSuccess).not.toHaveBeenCalled();
});
});
@@ -0,0 +1,465 @@
'use client';
import * as React from 'react';
import UploadZone from '@/components/datasets/UploadZone';
import { Button } from '@/components/ui/button';
import {
Dialog,
DialogContent,
DialogHeader,
DialogTitle,
} from '@/components/ui/dialog';
import { Input } from '@/components/ui/input';
import { Label } from '@/components/ui/label';
import { Tabs, TabsContent, TabsList, TabsTrigger } from '@/components/ui/tabs';
import { createDataset, uploadRawDataset } from '@/lib/api';
import {
parseCaptionCsv,
parseVideos2Caption,
parseVideosCaptionsTxt,
} from '@/lib/captionParsing';
const ALLOWED_VIDEO_EXT = '.mp4,.webm,.avi,.mov,.mkv';
type CaptionFormat = 'json' | 'txt' | 'csv';
export interface CreateDatasetModalProps {
isOpen: boolean;
onClose: () => void;
onSuccess: () => void;
}
export default function CreateDatasetModal({
isOpen,
onClose,
onSuccess,
}: CreateDatasetModalProps) {
const [name, setName] = React.useState('');
const [isSubmitting, setIsSubmitting] = React.useState(false);
const [rawPath, setRawPath] = React.useState('');
const [fileNames, setFileNames] = React.useState<string[]>([]);
const [isUploading, setIsUploading] = React.useState(false);
const [validationError, setValidationError] = React.useState<string | null>(
null,
);
const [captionFormat, setCaptionFormat] = React.useState<CaptionFormat>('json');
const [captionMap, setCaptionMap] = React.useState<Record<
string,
string
> | null>(null);
const [captionFileName, setCaptionFileName] = React.useState<string | null>(
null,
);
const [videosTxtLines, setVideosTxtLines] = React.useState<string[] | null>(
null,
);
const [videosTxtFileName, setVideosTxtFileName] = React.useState<
string | null
>(null);
const [captionsTxtLines, setCaptionsTxtLines] = React.useState<
string[] | null
>(null);
const [captionsTxtFileName, setCaptionsTxtFileName] = React.useState<
string | null
>(null);
// Bumped on every new media selection and on reset/clear; an in-flight
// upload whose generation no longer matches is discarded, so a superseded or
// abandoned upload can't repopulate/overwrite state.
const uploadGeneration = React.useRef(0);
const txtCaptionMap = React.useMemo(() => {
if (!captionsTxtLines || fileNames.length === 0) return null;
const { captions, error } = parseVideosCaptionsTxt(
videosTxtLines ?? null,
captionsTxtLines,
fileNames,
);
if (error) return null;
return Object.keys(captions).length > 0 ? captions : null;
}, [captionsTxtLines, fileNames, videosTxtLines]);
const effectiveCaptionMap =
captionFormat === 'txt' ? txtCaptionMap : captionMap;
// Keep the validation message in sync with the TXT inputs.
React.useEffect(() => {
if (captionFormat === 'txt' && captionsTxtLines && fileNames.length > 0) {
const hasVideosTxt =
!!videosTxtLines &&
videosTxtLines.length > 0 &&
videosTxtLines.some((s) => s.trim());
if (hasVideosTxt && videosTxtLines) {
const { error } = parseVideosCaptionsTxt(
videosTxtLines,
captionsTxtLines,
fileNames,
);
setValidationError(error);
} else {
setValidationError(null);
}
}
}, [captionFormat, captionsTxtLines, fileNames, videosTxtLines]);
function resetState() {
uploadGeneration.current += 1;
setName('');
setRawPath('');
setFileNames([]);
setValidationError(null);
setCaptionFormat('json');
setCaptionMap(null);
setCaptionFileName(null);
setVideosTxtLines(null);
setVideosTxtFileName(null);
setCaptionsTxtLines(null);
setCaptionsTxtFileName(null);
}
function handleClose() {
if (isSubmitting) return;
resetState();
onClose();
}
async function handleMediaChange(files: File[]) {
const gen = (uploadGeneration.current += 1);
setValidationError(null);
if (files.length === 0) {
setRawPath('');
setFileNames([]);
return;
}
setIsUploading(true);
try {
const res = await uploadRawDataset(files);
if (gen !== uploadGeneration.current) return; // superseded or abandoned
setRawPath(res.path);
setFileNames(res.file_names);
if (res.file_names.length === 0) {
setValidationError(`No video files found. Allowed: ${ALLOWED_VIDEO_EXT}`);
}
} catch (err) {
if (gen !== uploadGeneration.current) return;
setRawPath('');
setFileNames([]);
setValidationError(err instanceof Error ? err.message : 'Upload failed');
} finally {
if (gen === uploadGeneration.current) setIsUploading(false);
}
}
async function handleCaptionJsonChange(files: File[]) {
setValidationError(null);
setCaptionMap(null);
setCaptionFileName(null);
if (files.length === 0) return;
const file = files[0];
try {
const text = await file.text();
const { captions, error } = parseVideos2Caption(text, fileNames);
if (error) {
setValidationError(error);
return;
}
setCaptionMap(captions);
setCaptionFileName(file.name);
} catch {
setValidationError('Could not read the file.');
}
}
async function handleCaptionCsvChange(files: File[]) {
setValidationError(null);
setCaptionMap(null);
setCaptionFileName(null);
if (files.length === 0) return;
const file = files[0];
try {
const text = await file.text();
const { captions, error } = parseCaptionCsv(text, fileNames);
if (error) {
setValidationError(error);
return;
}
setCaptionMap(captions);
setCaptionFileName(file.name);
} catch {
setValidationError('Could not read the file.');
}
}
async function handleVideosTxtChange(files: File[]) {
setValidationError(null);
setVideosTxtLines(null);
setVideosTxtFileName(null);
if (files.length === 0) return;
try {
const text = await files[0].text();
setVideosTxtLines(text.split(/\r?\n/).map((s) => s.trim()));
setVideosTxtFileName(files[0].name);
} catch {
setValidationError('Could not read videos.txt.');
}
}
async function handleCaptionsTxtChange(files: File[]) {
setValidationError(null);
setCaptionsTxtLines(null);
setCaptionsTxtFileName(null);
if (files.length === 0) return;
try {
const text = await files[0].text();
setCaptionsTxtLines(text.split(/\r?\n/).map((s) => s.trim()));
setCaptionsTxtFileName(files[0].name);
} catch {
setValidationError('Could not read captions.txt.');
}
}
function handleCaptionFormatChange(format: CaptionFormat) {
setCaptionFormat(format);
setValidationError(null);
setCaptionMap(null);
setCaptionFileName(null);
setVideosTxtLines(null);
setVideosTxtFileName(null);
setCaptionsTxtLines(null);
setCaptionsTxtFileName(null);
}
async function handleSubmit(e: React.FormEvent) {
e.preventDefault();
setValidationError(null);
if (!name.trim()) return;
if (!rawPath || fileNames.length === 0) {
setValidationError('No data was found. Upload at least one video.');
return;
}
if (captionFormat === 'json' || captionFormat === 'csv') {
if (captionFileName && !captionMap) {
setValidationError(
'Caption file has errors. Fix or remove it before creating the dataset.',
);
return;
}
} else if (captionFormat === 'txt') {
if (videosTxtFileName || captionsTxtFileName) {
if (!captionsTxtFileName) {
setValidationError('Upload captions.txt to use TXT captions.');
return;
}
// Re-validate against current state rather than the stale render-time
// `validationError`, so a videos.txt mismatch both blocks submission
// and keeps its message visible.
const hasVideosTxt =
!!videosTxtLines &&
videosTxtLines.length > 0 &&
videosTxtLines.some((s) => s.trim());
if (hasVideosTxt && videosTxtLines) {
const { error } = parseVideosCaptionsTxt(
videosTxtLines,
captionsTxtLines ?? [],
fileNames,
);
if (error) {
setValidationError(error);
return;
}
}
}
}
const finalCaptionMap =
effectiveCaptionMap && Object.keys(effectiveCaptionMap).length > 0
? effectiveCaptionMap
: null;
if (finalCaptionMap) {
const missing = fileNames.filter((fn) => !(fn in finalCaptionMap));
if (missing.length > 0) {
const list =
missing.length <= 5
? missing.join(', ')
: `${missing.slice(0, 5).join(', ')} and ${missing.length - 5} more`;
const ok = window.confirm(
`The caption file does not include captions for ${missing.length} video(s): ${list}. They will get empty captions. Continue?`,
);
if (!ok) return;
}
}
setIsSubmitting(true);
try {
await createDataset({
name: name.trim(),
upload_path: rawPath,
file_names: fileNames,
...(finalCaptionMap ? { captions: finalCaptionMap } : {}),
});
onSuccess();
resetState();
onClose();
} catch (err) {
setValidationError(
err instanceof Error ? err.message : 'Failed to create dataset',
);
} finally {
setIsSubmitting(false);
}
}
return (
<Dialog
open={isOpen}
onOpenChange={(open) => {
if (!open) handleClose();
}}
>
<DialogContent
aria-describedby={undefined}
className="max-h-[90vh] w-[90vw] max-w-[850px] overflow-y-auto"
>
<DialogHeader>
<DialogTitle>Add Dataset — Raw</DialogTitle>
</DialogHeader>
<form onSubmit={handleSubmit} autoComplete="off">
<div className="mb-3.5 flex flex-col gap-1.5">
<Label htmlFor="add-dataset-name">Name</Label>
<Input
id="add-dataset-name"
type="text"
value={name}
onChange={(e) => setName(e.target.value)}
placeholder="My dataset"
required
disabled={isSubmitting}
/>
</div>
<div className="mb-3.5 flex flex-col gap-1.5">
<Label>Videos</Label>
<UploadZone
label="Upload video files"
hint="Select files or a folder (.mp4, .webm, .avi, .mov, .mkv)"
accept={ALLOWED_VIDEO_EXT}
multiple
directory
allowBothFileAndDirectory
value={rawPath}
fileName={
fileNames.length > 0 ? `${fileNames.length} file(s)` : undefined
}
onFiles={handleMediaChange}
onClear={() => {
uploadGeneration.current += 1;
setRawPath('');
setFileNames([]);
setCaptionMap(null);
setCaptionFileName(null);
setValidationError(null);
}}
disabled={isSubmitting}
uploading={isUploading}
/>
</div>
<div className="mb-3.5 flex flex-col gap-1.5">
<Tabs
value={captionFormat}
onValueChange={(value) =>
handleCaptionFormatChange(value as CaptionFormat)
}
>
<div className="mb-1.5 flex flex-wrap items-center gap-4">
<Label>Captions (optional)</Label>
<TabsList>
<TabsTrigger value="json" disabled={isSubmitting}>
JSON
</TabsTrigger>
<TabsTrigger value="txt" disabled={isSubmitting}>
TXT
</TabsTrigger>
<TabsTrigger value="csv" disabled={isSubmitting}>
CSV
</TabsTrigger>
</TabsList>
</div>
<TabsContent value="json">
<UploadZone
label="Upload videos2caption.json"
hint="Array of { path, cap } or object mapping file names to captions"
accept=".json,application/json"
value={captionFileName ? '1' : ''}
fileName={captionFileName ?? undefined}
onFiles={handleCaptionJsonChange}
onClear={() => {
setCaptionMap(null);
setCaptionFileName(null);
setValidationError(null);
}}
disabled={isSubmitting}
/>
</TabsContent>
<TabsContent value="txt">
<div className="flex flex-wrap gap-[15px] [&>*]:min-w-[200px] [&>*]:flex-1">
<UploadZone
label="Upload videos.txt (optional)"
hint="One video path per line, or leave empty to match captions to videos in alphabetical order"
accept=".txt,text/plain"
value={videosTxtFileName ? '1' : ''}
fileName={videosTxtFileName ?? undefined}
onFiles={handleVideosTxtChange}
onClear={() => {
setVideosTxtLines(null);
setVideosTxtFileName(null);
setValidationError(null);
}}
disabled={isSubmitting}
/>
<UploadZone
label="Upload captions.txt"
hint="One caption per line (same order as videos.txt or alphabetical)"
accept=".txt,text/plain"
value={captionsTxtFileName ? '1' : ''}
fileName={captionsTxtFileName ?? undefined}
onFiles={handleCaptionsTxtChange}
onClear={() => {
setCaptionsTxtLines(null);
setCaptionsTxtFileName(null);
setValidationError(null);
}}
disabled={isSubmitting}
/>
</div>
</TabsContent>
<TabsContent value="csv">
<UploadZone
label="Upload captions CSV"
hint="Header: video_name, caption"
accept=".csv,text/csv"
value={captionFileName ? '1' : ''}
fileName={captionFileName ?? undefined}
onFiles={handleCaptionCsvChange}
onClear={() => {
setCaptionMap(null);
setCaptionFileName(null);
setValidationError(null);
}}
disabled={isSubmitting}
/>
</TabsContent>
</Tabs>
</div>
{validationError && (
<p className="mb-2 text-sm text-destructive">{validationError}</p>
)}
<Button type="submit" disabled={isSubmitting}>
{isSubmitting ? 'Creating…' : 'Create Dataset'}
</Button>
</form>
</DialogContent>
</Dialog>
);
}
@@ -0,0 +1,119 @@
import { beforeEach, describe, expect, it, vi } from 'vitest';
import { fireEvent, render, screen, waitFor } from '@testing-library/react';
import DatasetCard from '@/components/datasets/DatasetCard';
import { deleteDataset } from '@/lib/api';
import type { Dataset } from '@/lib/api';
import { setActiveDatasetId } from '@/stores/activeDataset';
vi.mock('@/lib/api', () => ({
deleteDataset: vi.fn(),
}));
const mockedDeleteDataset = vi.mocked(deleteDataset);
const dataset: Dataset = {
id: 'ds-1',
name: 'My Dataset',
created_at: 0,
file_count: 3,
size_bytes: 2048,
};
beforeEach(() => {
setActiveDatasetId(null);
});
describe('DatasetCard', () => {
it('keeps selection and delete buttons as semantic siblings', () => {
render(<DatasetCard dataset={dataset} onUpdated={() => {}} />);
const selectButton = screen.getByRole('button', { pressed: false });
const deleteButton = screen.getByRole('button', { name: 'Delete' });
expect(selectButton).toHaveTextContent('My Dataset');
expect(selectButton).not.toContainElement(deleteButton);
});
it('renders the name, file count and human-readable size', () => {
render(<DatasetCard dataset={dataset} onUpdated={() => {}} />);
expect(screen.getByText('My Dataset')).toBeInTheDocument();
expect(screen.getByText('3 files · 2.0 KB')).toBeInTheDocument();
});
it('uses the singular "file" label and byte units for a small dataset', () => {
render(
<DatasetCard
dataset={{ ...dataset, file_count: 1, size_bytes: 512 }}
onUpdated={() => {}}
/>,
);
expect(screen.getByText('1 file · 512 B')).toBeInTheDocument();
});
it('calls onSelect when the card body is clicked', () => {
const onSelect = vi.fn();
render(
<DatasetCard dataset={dataset} onUpdated={() => {}} onSelect={onSelect} />,
);
fireEvent.click(screen.getByText('My Dataset'));
expect(onSelect).toHaveBeenCalledTimes(1);
});
it('deletes after confirmation and notifies the parent without selecting', async () => {
const onUpdated = vi.fn();
const onSelect = vi.fn();
const confirmSpy = vi.spyOn(window, 'confirm').mockReturnValue(true);
mockedDeleteDataset.mockResolvedValue(undefined);
render(
<DatasetCard
dataset={dataset}
onUpdated={onUpdated}
onSelect={onSelect}
/>,
);
fireEvent.click(screen.getByRole('button', { name: 'Delete' }));
expect(confirmSpy).toHaveBeenCalledWith('Delete dataset "My Dataset"?');
expect(mockedDeleteDataset).toHaveBeenCalledWith('ds-1');
await waitFor(() => expect(onUpdated).toHaveBeenCalledTimes(1));
expect(onSelect).not.toHaveBeenCalled();
});
it('keeps the selection and delete actions separate', () => {
const onSelect = vi.fn();
render(
<DatasetCard dataset={dataset} onUpdated={() => {}} onSelect={onSelect} />,
);
// Activating the Delete button must not bubble into a card selection.
fireEvent.keyDown(screen.getByRole('button', { name: 'Delete' }), {
key: 'Enter',
});
expect(onSelect).not.toHaveBeenCalled();
// Activating the dedicated selection button selects the dataset.
fireEvent.click(
screen.getByRole('button', {
name: /My Dataset.*3 files.*2.0 KB/,
}),
);
expect(onSelect).toHaveBeenCalledTimes(1);
});
it('does not delete when the user cancels the confirm dialog', () => {
vi.spyOn(window, 'confirm').mockReturnValue(false);
render(<DatasetCard dataset={dataset} onUpdated={() => {}} />);
fireEvent.click(screen.getByRole('button', { name: 'Delete' }));
expect(mockedDeleteDataset).not.toHaveBeenCalled();
});
it('applies selected styling when it is the active dataset', () => {
setActiveDatasetId('ds-1');
const { container } = render(
<DatasetCard dataset={dataset} onUpdated={() => {}} />,
);
expect(container.firstChild).toHaveClass('border-accent-blue', 'bg-accent-blue/5');
});
});
@@ -0,0 +1,87 @@
'use client';
import * as React from 'react';
import { Button } from '@/components/ui/button';
import { deleteDataset } from '@/lib/api';
import type { Dataset } from '@/lib/api';
import { cn } from '@/lib/utils';
import { activeDatasetStore, setActiveDatasetId } from '@/stores/activeDataset';
import { useStore } from '@/hooks/useStore';
function formatSize(sizeBytes: number): string {
if (sizeBytes < 1024) return `${sizeBytes} B`;
if (sizeBytes < 1024 * 1024) return `${(sizeBytes / 1024).toFixed(1)} KB`;
if (sizeBytes < 1024 * 1024 * 1024) {
return `${(sizeBytes / (1024 * 1024)).toFixed(1)} MB`;
}
return `${(sizeBytes / (1024 * 1024 * 1024)).toFixed(1)} GB`;
}
export default function DatasetCard({
dataset,
onUpdated,
onSelect = () => {},
}: {
dataset: Dataset;
onUpdated: () => void;
onSelect?: () => void;
}) {
const { activeDatasetId } = useStore(activeDatasetStore);
const [isLoading, setIsLoading] = React.useState(false);
const isSelected = activeDatasetId === dataset.id;
const fileCount = dataset.file_count ?? 0;
const sizeLabel = formatSize(dataset.size_bytes ?? 0);
async function handleDelete(e: React.MouseEvent) {
e.preventDefault();
e.stopPropagation();
if (isLoading) return;
if (!window.confirm(`Delete dataset "${dataset.name}"?`)) return;
setIsLoading(true);
try {
await deleteDataset(dataset.id);
if (activeDatasetStore.get().activeDatasetId === dataset.id) {
setActiveDatasetId(null);
}
onUpdated();
} catch (err) {
window.alert(
err instanceof Error ? err.message : 'Failed to delete dataset',
);
} finally {
setIsLoading(false);
}
}
return (
<article
className={cn(
'mb-3 flex items-start gap-3 rounded-lg border border-border bg-background px-[1.15rem] py-4',
isSelected && 'border-accent-blue bg-accent-blue/5',
)}
>
<button
type="button"
aria-pressed={isSelected}
onClick={onSelect}
className="flex min-w-0 flex-1 cursor-pointer flex-col gap-[0.6rem] rounded-md text-left"
>
<span className="text-[0.95rem] font-semibold">{dataset.name}</span>
<span className="text-sm text-muted-foreground">
{fileCount} {fileCount === 1 ? 'file' : 'files'} · {sizeLabel}
</span>
</button>
<Button
type="button"
variant="destructive"
size="sm"
onClick={handleDelete}
disabled={isLoading}
>
Delete
</Button>
</article>
);
}
@@ -0,0 +1,223 @@
import { describe, it, expect, vi, beforeEach } from 'vitest';
import { render, screen, fireEvent, act } from '@testing-library/react';
import DatasetSidebar from '@/components/datasets/DatasetSidebar';
import * as api from '@/lib/api';
import type { Dataset } from '@/lib/api';
vi.mock('@/lib/api');
vi.mock('sonner', () => ({
toast: { error: vi.fn() },
}));
const mockedApi = vi.mocked(api);
const dataset: Dataset = {
id: 'ds-1',
name: 'My Dataset',
created_at: 0,
};
beforeEach(() => {
mockedApi.getDatasetFiles.mockResolvedValue({
file_names: ['a.mp4', 'b.mp4'],
captions: { 'a.mp4': 'cap a', 'b.mp4': '' },
});
mockedApi.getDatasetMediaUrl.mockImplementation(
(id, fileName) => `http://test/${id}/${fileName}`,
);
mockedApi.updateDatasetCaption.mockResolvedValue(undefined);
});
describe('DatasetSidebar', () => {
it('fills the mobile viewport without reserving main-content width', async () => {
const onWidthChange = vi.fn();
render(
<DatasetSidebar
dataset={dataset}
isMobile
onClose={() => {}}
onWidthChange={onWidthChange}
/>,
);
const drawer = screen.getByRole('dialog', {
name: 'My Dataset dataset details',
});
expect(drawer).toHaveStyle({ width: '100%', maxWidth: 'none' });
expect(drawer).toHaveAttribute('aria-modal', 'true');
expect(drawer).toHaveFocus();
expect(onWidthChange).toHaveBeenCalledWith(0);
});
it('lists dataset files after loading', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
// The first file's caption is rendered once loading resolves.
expect(await screen.findByDisplayValue('cap a')).toBeInTheDocument();
expect(mockedApi.getDatasetFiles).toHaveBeenCalledWith('ds-1');
// One caption editor per returned file.
const captionFields = screen.getAllByPlaceholderText('Caption');
expect(captionFields).toHaveLength(2);
// Media URLs are requested per visible file.
expect(mockedApi.getDatasetMediaUrl).toHaveBeenCalledWith('ds-1', 'a.mp4');
expect(mockedApi.getDatasetMediaUrl).toHaveBeenCalledWith('ds-1', 'b.mp4');
});
it('shows a fallback when a dataset preview cannot load', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
const preview = await screen.findByLabelText('Preview of a.mp4');
fireEvent.error(preview);
expect(screen.getByText('Preview unavailable')).toBeInTheDocument();
expect(screen.queryByLabelText('Preview of a.mp4')).not.toBeInTheDocument();
});
it('debounces caption save by 500ms', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
const textarea = await screen.findByDisplayValue('cap a');
vi.useFakeTimers();
try {
fireEvent.change(textarea, { target: { value: 'updated caption' } });
// Nothing saved immediately or just before the debounce window closes.
expect(mockedApi.updateDatasetCaption).not.toHaveBeenCalled();
act(() => {
vi.advanceTimersByTime(499);
});
expect(mockedApi.updateDatasetCaption).not.toHaveBeenCalled();
// Saved exactly once after the full 500ms.
act(() => {
vi.advanceTimersByTime(1);
});
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledTimes(1);
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledWith(
'ds-1',
'a.mp4',
'updated caption',
);
} finally {
vi.useRealTimers();
}
});
it('coalesces rapid edits into a single debounced save', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
const textarea = await screen.findByDisplayValue('cap a');
vi.useFakeTimers();
try {
fireEvent.change(textarea, { target: { value: 'one' } });
act(() => {
vi.advanceTimersByTime(300);
});
fireEvent.change(textarea, { target: { value: 'two' } });
act(() => {
vi.advanceTimersByTime(500);
});
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledTimes(1);
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledWith(
'ds-1',
'a.mp4',
'two',
);
} finally {
vi.useRealTimers();
}
});
it('shows a failed save and lets the user retry it', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
mockedApi.updateDatasetCaption
.mockRejectedValueOnce(new Error('network down'))
.mockResolvedValueOnce(undefined);
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
const textarea = await screen.findByDisplayValue('cap a');
vi.useFakeTimers();
try {
fireEvent.change(textarea, { target: { value: 'needs retry' } });
await act(async () => {
await vi.advanceTimersByTimeAsync(500);
});
expect(screen.getByText(/Not saved/)).toBeInTheDocument();
fireEvent.click(screen.getByRole('button', { name: 'Retry' }));
await act(async () => {
await Promise.resolve();
});
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledTimes(2);
expect(mockedApi.updateDatasetCaption).toHaveBeenLastCalledWith(
'ds-1',
'a.mp4',
'needs retry',
);
expect(screen.getByText('Saved')).toBeInTheDocument();
} finally {
vi.useRealTimers();
}
});
it('debounces per file: editing another caption does not cancel a pending save', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
await screen.findByDisplayValue('cap a');
const [textareaA, textareaB] = screen.getAllByPlaceholderText('Caption');
vi.useFakeTimers();
try {
fireEvent.change(textareaA, { target: { value: 'new cap a' } });
act(() => {
vi.advanceTimersByTime(300);
});
// Editing b.mp4 inside a.mp4's debounce window must not drop a's save.
fireEvent.change(textareaB, { target: { value: 'new cap b' } });
act(() => {
vi.advanceTimersByTime(500);
});
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledTimes(2);
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledWith(
'ds-1',
'a.mp4',
'new cap a',
);
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledWith(
'ds-1',
'b.mp4',
'new cap b',
);
} finally {
vi.useRealTimers();
}
});
it('flushes a pending save on unmount instead of dropping it', async () => {
const { unmount } = render(
<DatasetSidebar dataset={dataset} onClose={() => {}} />,
);
const textarea = await screen.findByDisplayValue('cap a');
vi.useFakeTimers();
try {
fireEvent.change(textarea, { target: { value: 'edited just before close' } });
unmount();
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledTimes(1);
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledWith(
'ds-1',
'a.mp4',
'edited just before close',
);
} finally {
vi.useRealTimers();
}
});
});
@@ -0,0 +1,347 @@
'use client';
import * as React from 'react';
import { ImageOff, X } from 'lucide-react';
import { toast } from 'sonner';
import DownloadCaptions from '@/components/datasets/DownloadCaptions';
import { Textarea } from '@/components/ui/textarea';
import { useResizable } from '@/hooks/useResizable';
import {
getDatasetFiles,
getDatasetMediaUrl,
updateDatasetCaption,
type Dataset,
} from '@/lib/api';
import { cn } from '@/lib/utils';
import { useDrawerFocus } from '@/hooks/useDrawerFocus';
const SIDEBAR_MIN_WIDTH = 320;
const SIDEBAR_MAX_WIDTH = 900;
const INITIAL_PAGE_SIZE = 24;
const PAGE_SIZE = 24;
const SCROLL_THRESHOLD = 200;
type CaptionSaveState = 'idle' | 'saving' | 'saved' | 'error';
// Memoized so a caption keystroke re-renders only the edited card, not every
// visible <video> in the grid (visibleCount grows unbounded with scrolling).
const DatasetFileCard = React.memo(function DatasetFileCard({
fileName,
mediaUrl,
caption,
thumbLoaded,
saveState,
onCaptionChange,
onCaptionRetry,
onThumbLoaded,
}: {
fileName: string;
mediaUrl: string;
caption: string;
thumbLoaded: boolean;
saveState: CaptionSaveState;
onCaptionChange: (fileName: string, value: string) => void;
onCaptionRetry: (fileName: string, value: string) => void;
onThumbLoaded: (fileName: string) => void;
}) {
const [mediaFailed, setMediaFailed] = React.useState(false);
React.useEffect(() => {
setMediaFailed(false);
}, [mediaUrl]);
return (
<div className="relative flex flex-col overflow-hidden rounded-lg border border-border bg-background">
{!thumbLoaded && !mediaFailed && (
<div className="pointer-events-none absolute inset-0 flex items-center justify-center bg-background/70">
<div className="h-6 w-6 animate-spin rounded-full border-2 border-muted-foreground/40 border-t-accent-blue" />
</div>
)}
{mediaFailed ? (
<div
role="status"
className="flex aspect-video w-full flex-col items-center justify-center gap-1 bg-muted px-2 text-center text-muted-foreground"
>
<ImageOff className="size-5" aria-hidden />
<span className="text-xs">Preview unavailable</span>
</div>
) : (
// eslint-disable-next-line jsx-a11y/media-has-caption
<video
src={mediaUrl}
aria-label={`Preview of ${fileName}`}
className="aspect-video w-full bg-border object-cover"
muted
autoPlay
loop
playsInline
onLoadedData={() => onThumbLoaded(fileName)}
onError={() => {
setMediaFailed(true);
onThumbLoaded(fileName);
}}
/>
)}
<Textarea
aria-label={`Caption for ${fileName}`}
value={caption}
onChange={(e) => onCaptionChange(fileName, e.target.value)}
placeholder="Caption"
rows={2}
className="min-h-[2.5rem] resize-y rounded-none border-0 bg-transparent p-1.5 text-xs shadow-none focus-visible:border-transparent focus-visible:ring-0"
/>
<div
aria-live="polite"
className="flex min-h-6 items-center px-1.5 pb-1 text-[0.7rem] text-muted-foreground"
>
{saveState === 'saving' && <span>Saving…</span>}
{saveState === 'saved' && <span>Saved</span>}
{saveState === 'error' && (
<span role="alert" className="text-destructive">
Not saved.{' '}
<button
type="button"
onClick={() => onCaptionRetry(fileName, caption)}
className="inline-flex min-h-11 items-center font-medium underline underline-offset-2"
>
Retry
</button>
</span>
)}
</div>
</div>
);
});
export default function DatasetSidebar({
dataset,
isMobile = false,
onClose,
onWidthChange,
}: {
dataset: Dataset;
isMobile?: boolean;
onClose: () => void;
onWidthChange?: (w: number) => void;
}) {
const drawerRef = useDrawerFocus<HTMLElement>(isMobile);
const [width, setWidth] = React.useState(400);
const [isDragging, setIsDragging] = React.useState(false);
const [fileNames, setFileNames] = React.useState<string[]>([]);
const [captions, setCaptions] = React.useState<Record<string, string>>({});
const [visibleCount, setVisibleCount] = React.useState(INITIAL_PAGE_SIZE);
const [isLoading, setIsLoading] = React.useState(true);
const [thumbLoaded, setThumbLoaded] = React.useState<
Record<string, boolean>
>({});
const [captionSaveStates, setCaptionSaveStates] = React.useState<
Record<string, CaptionSaveState>
>({});
// Pending debounced caption saves, keyed per file so editing one caption
// can't cancel another file's pending save.
const pendingSaves = React.useRef(
new Map<string, { timer: ReturnType<typeof setTimeout>; save: () => void }>(),
);
const captionVersions = React.useRef(new Map<string, number>());
const scrollRef = React.useRef<HTMLDivElement>(null);
React.useEffect(() => {
onWidthChange?.(isMobile ? 0 : width);
}, [isMobile, width, onWidthChange]);
React.useEffect(() => {
let cancelled = false;
setIsLoading(true);
getDatasetFiles(dataset.id)
.then((data) => {
if (cancelled) return;
setFileNames(data.file_names);
setCaptions(data.captions);
setVisibleCount(INITIAL_PAGE_SIZE);
setThumbLoaded({});
setCaptionSaveStates({});
captionVersions.current.clear();
})
.catch((err) => console.error('Failed to load dataset files:', err))
.finally(() => {
if (!cancelled) setIsLoading(false);
});
return () => {
cancelled = true;
};
}, [dataset.id]);
const flushPendingSaves = React.useCallback(() => {
for (const { timer, save } of pendingSaves.current.values()) {
clearTimeout(timer);
save();
}
pendingSaves.current.clear();
}, []);
// Flush (not drop) pending saves when the dataset changes or on unmount, so
// an edit made within the debounce window of closing isn't lost.
React.useEffect(() => flushPendingSaves, [dataset.id, flushPendingSaves]);
const { onMouseDown } = useResizable({
edge: 'right',
minWidth: SIDEBAR_MIN_WIDTH,
maxWidth: SIDEBAR_MAX_WIDTH,
getWidth: () => width,
onWidth: setWidth,
onDragChange: setIsDragging,
});
const datasetId = dataset.id;
const persistCaption = React.useCallback(
(fileName: string, value: string, version: number) => {
setCaptionSaveStates((prev) => ({ ...prev, [fileName]: 'saving' }));
void updateDatasetCaption(datasetId, fileName, value)
.then(() => {
if (captionVersions.current.get(fileName) !== version) return;
setCaptionSaveStates((prev) => ({ ...prev, [fileName]: 'saved' }));
})
.catch((error) => {
if (captionVersions.current.get(fileName) !== version) return;
console.error('Failed to save caption:', error);
setCaptionSaveStates((prev) => ({ ...prev, [fileName]: 'error' }));
toast.error('Caption was not saved', {
description: `${fileName}: check the Studio API, then retry.`,
});
});
},
[datasetId],
);
const handleCaptionChange = React.useCallback(
(fileName: string, value: string) => {
setCaptions((prev) => ({ ...prev, [fileName]: value }));
setCaptionSaveStates((prev) => ({ ...prev, [fileName]: 'idle' }));
const pending = pendingSaves.current.get(fileName);
if (pending) clearTimeout(pending.timer);
const version = (captionVersions.current.get(fileName) ?? 0) + 1;
captionVersions.current.set(fileName, version);
const save = () => persistCaption(fileName, value, version);
const timer = setTimeout(() => {
pendingSaves.current.delete(fileName);
save();
}, 500);
pendingSaves.current.set(fileName, { timer, save });
},
[persistCaption],
);
const handleCaptionRetry = React.useCallback(
(fileName: string, value: string) => {
const version = (captionVersions.current.get(fileName) ?? 0) + 1;
captionVersions.current.set(fileName, version);
persistCaption(fileName, value, version);
},
[persistCaption],
);
function handleScroll() {
const el = scrollRef.current;
if (!el || isLoading || visibleCount >= fileNames.length) return;
const { scrollTop, scrollHeight, clientHeight } = el;
const distanceFromBottom = scrollHeight - (scrollTop + clientHeight);
if (distanceFromBottom < SCROLL_THRESHOLD) {
setVisibleCount((c) => Math.min(c + PAGE_SIZE, fileNames.length));
}
}
// When the visible grid doesn't overflow (a wide/tall sidebar can fit the
// first page without a scrollbar), no scroll event ever fires — so top up
// visibleCount until it overflows or every file is shown. Re-runs on resize
// (width) and after each page grows, otherwise files 25..N are unreachable.
React.useEffect(() => {
if (isLoading || visibleCount >= fileNames.length) return;
const el = scrollRef.current;
if (el && el.scrollHeight <= el.clientHeight) {
setVisibleCount((c) => Math.min(c + PAGE_SIZE, fileNames.length));
}
}, [visibleCount, fileNames.length, isLoading, width]);
const markThumbLoaded = React.useCallback((fileName: string) => {
setThumbLoaded((prev) =>
prev[fileName] ? prev : { ...prev, [fileName]: true },
);
}, []);
const visibleFiles = fileNames.slice(0, visibleCount);
return (
<aside
ref={drawerRef}
tabIndex={-1}
role="dialog"
aria-label={`${dataset.name} dataset details`}
aria-modal={isMobile || undefined}
className="fixed bottom-0 right-0 top-[var(--header-height)] z-50 flex max-h-[calc(100dvh-var(--header-height))] min-w-0 shrink-0 flex-col border-l border-border bg-card md:min-w-[320px]"
style={{
width: isMobile ? '100%' : width,
maxWidth: isMobile ? 'none' : SIDEBAR_MAX_WIDTH,
}}
>
<div className="flex shrink-0 items-center justify-between border-b border-border px-5 py-4">
<h2 className="m-0 min-w-0 truncate text-base font-semibold text-foreground">
{dataset.name}
</h2>
<div className="flex shrink-0 items-center gap-2">
<DownloadCaptions fileNames={fileNames} captions={captions} />
<button
type="button"
onClick={onClose}
title="Close"
aria-label="Close"
className="flex size-11 items-center justify-center rounded-lg text-muted-foreground transition-colors hover:bg-accent hover:text-foreground"
>
<X className="h-[18px] w-[18px]" />
</button>
</div>
</div>
<div className="flex min-h-0 flex-1 flex-col overflow-hidden">
<div
ref={scrollRef}
onScroll={handleScroll}
className="flex-1 overflow-y-auto p-4"
>
{isLoading ? (
<p className="p-8 text-center text-muted-foreground">Loading…</p>
) : fileNames.length === 0 ? (
<p className="p-8 text-center text-muted-foreground">
No media files
</p>
) : (
<div className="grid gap-4 [grid-template-columns:repeat(auto-fill,minmax(140px,1fr))]">
{visibleFiles.map((fileName) => (
<DatasetFileCard
key={fileName}
fileName={fileName}
mediaUrl={getDatasetMediaUrl(dataset.id, fileName)}
caption={captions[fileName] ?? ''}
thumbLoaded={!!thumbLoaded[fileName]}
saveState={captionSaveStates[fileName] ?? 'idle'}
onCaptionChange={handleCaptionChange}
onCaptionRetry={handleCaptionRetry}
onThumbLoaded={markThumbLoaded}
/>
))}
</div>
)}
</div>
</div>
{!isMobile && <div
role="presentation"
onMouseDown={onMouseDown}
className={cn(
'absolute bottom-0 left-0 top-0 z-[1] w-1.5 cursor-col-resize hover:bg-accent-blue/25',
isDragging && 'bg-accent-blue/25',
)}
/>}
</aside>
);
}
@@ -0,0 +1,113 @@
'use client';
import { ChevronDown } from 'lucide-react';
import { Button } from '@/components/ui/button';
import { downloadBlob } from '@/lib/utils';
const MENU_ITEM =
'block min-h-11 w-full cursor-pointer px-4 py-2 text-left text-sm font-medium text-foreground transition-colors hover:bg-muted disabled:cursor-not-allowed disabled:opacity-50';
export default function DownloadCaptions({
fileNames,
captions,
}: {
fileNames: string[];
captions: Record<string, string>;
}) {
const disabled = fileNames.length === 0;
// Sorted lazily in the click handlers: this component re-renders with the
// sidebar on every caption keystroke, and the sort is only needed on click.
const sortedFileNames = () => [...fileNames].sort();
function handleDownloadJson() {
const data = sortedFileNames().map((path) => ({
path,
cap: captions[path] ?? '',
}));
const blob = new Blob([JSON.stringify(data, null, 2)], {
type: 'application/json',
});
downloadBlob(blob, 'videos2caption.json');
}
function handleDownloadTxt() {
const sortedNames = sortedFileNames();
const videosContent = sortedNames.join('\n');
const promptContent = sortedNames
.map((fn) => captions[fn] ?? '')
.join('\n');
downloadBlob(
new Blob([videosContent], { type: 'text/plain' }),
'videos.txt',
);
setTimeout(() => {
downloadBlob(
new Blob([promptContent], { type: 'text/plain' }),
'captions.txt',
);
}, 100);
}
function handleDownloadCsv() {
const escape = (s: string) =>
s.includes('"') || s.includes(',') || s.includes('\n')
? `"${s.replace(/"/g, '""')}"`
: s;
const rows = sortedFileNames().map(
(fn) => `${escape(fn)},${escape(captions[fn] ?? '')}`,
);
const csv = ['video_name,caption', ...rows].join('\n');
downloadBlob(new Blob([csv], { type: 'text/csv' }), 'captions.csv');
}
return (
<div className="group relative inline-block">
<Button
type="button"
variant="outline"
size="sm"
disabled={disabled}
aria-haspopup="menu"
className="gap-1.5"
>
Download Captions
<ChevronDown className="h-3.5 w-3.5 opacity-85" />
</Button>
{!disabled && (
<div
role="menu"
className="invisible absolute right-0 top-full z-[200] min-w-full -translate-y-1 pt-1 opacity-0 transition-all group-focus-within:visible group-focus-within:translate-y-0 group-focus-within:opacity-100 group-hover:visible group-hover:translate-y-0 group-hover:opacity-100"
>
<div className="overflow-hidden rounded-lg border border-border bg-popover py-1 shadow-lg">
<button
type="button"
role="menuitem"
className={MENU_ITEM}
onClick={handleDownloadJson}
>
JSON
</button>
<button
type="button"
role="menuitem"
className={MENU_ITEM}
onClick={handleDownloadTxt}
>
TXT
</button>
<button
type="button"
role="menuitem"
className={MENU_ITEM}
onClick={handleDownloadCsv}
>
CSV
</button>
</div>
</div>
)}
</div>
);
}
@@ -0,0 +1,205 @@
'use client';
import * as React from 'react';
import { cn } from '@/lib/utils';
export interface UploadZoneProps {
label: string;
hint?: string;
accept?: string;
multiple?: boolean;
directory?: boolean;
allowBothFileAndDirectory?: boolean;
value?: string;
fileName?: string;
onFiles?: (files: File[]) => void;
onClear?: () => void;
disabled?: boolean;
uploading?: boolean;
}
const linkButtonClass =
'cursor-pointer bg-transparent p-0 text-accent-blue hover:underline disabled:cursor-not-allowed disabled:no-underline disabled:opacity-50';
export default function UploadZone({
label,
hint,
accept,
multiple = false,
directory = false,
allowBothFileAndDirectory = false,
value = '',
fileName,
onFiles,
onClear,
disabled = false,
uploading = false,
}: UploadZoneProps) {
const fileInputRef = React.useRef<HTMLInputElement>(null);
const directoryInputRef = React.useRef<HTMLInputElement>(null);
const useBoth = directory && allowBothFileAndDirectory;
const clickable = !useBoth;
const hasContent = !!(value || fileName);
// `webkitdirectory` is a DOM property without a typed JSX prop, so set it
// imperatively. The primary input only selects folders when it is the sole
// input; the dedicated directory input always does.
React.useEffect(() => {
if (fileInputRef.current) {
fileInputRef.current.webkitdirectory = directory && !allowBothFileAndDirectory;
}
if (directoryInputRef.current) {
directoryInputRef.current.webkitdirectory = true;
}
}, [directory, allowBothFileAndDirectory]);
function handleChange(e: React.ChangeEvent<HTMLInputElement>) {
const files = e.target.files;
if (files && files.length > 0) {
onFiles?.(Array.from(files));
}
e.target.value = '';
}
function handleClick() {
if (!disabled) {
fileInputRef.current?.click();
}
}
function handleKeyActivate(e: React.KeyboardEvent) {
if ((e.target as HTMLElement).closest('button')) return;
if (e.key === 'Enter' || e.key === ' ') {
e.preventDefault();
handleClick();
}
}
function handleDrop(e: React.DragEvent) {
if (disabled) return;
e.preventDefault();
const files = e.dataTransfer.files;
if (files && files.length > 0) {
onFiles?.(Array.from(files));
}
}
function handleDragOver(e: React.DragEvent) {
if (disabled) return;
e.preventDefault();
}
function clearInputs() {
if (fileInputRef.current) fileInputRef.current.value = '';
if (directoryInputRef.current) directoryInputRef.current.value = '';
}
return (
<div
className={cn(
'flex min-h-[150px] grow flex-col items-center justify-center rounded-lg border-2 border-dashed border-border bg-muted/40 px-5 py-6 text-center transition-colors hover:border-accent-blue hover:bg-muted/60',
clickable ? 'cursor-pointer' : 'cursor-default',
hasContent && 'border-solid border-accent-blue',
)}
onClick={clickable ? handleClick : undefined}
onKeyDown={clickable ? handleKeyActivate : undefined}
onDragOver={handleDragOver}
onDrop={handleDrop}
role={clickable ? 'button' : undefined}
tabIndex={clickable ? 0 : undefined}
>
<input
ref={fileInputRef}
type="file"
className="hidden"
accept={accept}
multiple={multiple}
onChange={handleChange}
disabled={disabled}
/>
{useBoth && (
<input
ref={directoryInputRef}
type="file"
className="hidden"
multiple
onChange={handleChange}
disabled={disabled}
/>
)}
<div className="mb-2 text-sm text-muted-foreground">{label}</div>
{!hasContent && (
<span className="mt-1.5 text-xs text-muted-foreground">
{uploading ? (
'Uploading…'
) : useBoth ? (
<>
<span
role="button"
tabIndex={0}
className="cursor-pointer text-accent-blue hover:underline"
onClick={(e) => {
e.stopPropagation();
handleClick();
}}
onKeyDown={(e) => {
if (e.key === 'Enter' || e.key === ' ') {
e.preventDefault();
handleClick();
}
}}
>
Select files
</span>
{' · '}
<button
type="button"
className={linkButtonClass}
onClick={(e) => {
e.preventDefault();
e.stopPropagation();
if (!disabled) directoryInputRef.current?.click();
}}
disabled={disabled}
>
Select folder
</button>
</>
) : directory ? (
'Click or drop folder'
) : (
'Click or drop file(s)'
)}
</span>
)}
{fileName && (
<div className="mt-2 text-sm text-foreground">
{fileName}
{onClear && (
<>
{' · '}
<button
type="button"
className={linkButtonClass}
onClick={(e) => {
e.stopPropagation();
onClear();
clearInputs();
}}
disabled={disabled || uploading}
>
Clear
</button>
</>
)}
</div>
)}
{hint && <div className="mt-1.5 text-xs text-muted-foreground">{hint}</div>}
</div>
);
}
@@ -0,0 +1,174 @@
'use client';
import * as React from 'react';
import { Input } from '@/components/ui/input';
import { Label } from '@/components/ui/label';
import { Slider } from '@/components/ui/slider';
import { Switch } from '@/components/ui/switch';
import { cn } from '@/lib/utils';
/**
* Labeled form-field rows shared by the Settings page and the Create Job
* modal. They mirror the Svelte `Toggle`/`Slider` UX on top of the shadcn
* `Switch`/`Slider`/`Input` primitives.
*/
export function FieldRow({
htmlFor,
label,
title,
className,
children,
}: {
htmlFor: string;
label: string;
title?: string;
className?: string;
children: React.ReactNode;
}) {
return (
<div className={cn('flex flex-col gap-1.5', className)}>
<Label
htmlFor={htmlFor}
title={title}
className="pl-0.5 text-xs font-normal tracking-wide text-muted-foreground"
>
{label}
</Label>
{children}
</div>
);
}
export function SliderRow({
id,
label,
title,
min,
max,
step,
value,
onChange,
disabled,
format = (v) => String(v),
}: {
id: string;
label: string;
title?: string;
min: number;
max: number;
step: number;
value: number;
/** Called once per gesture (pointer release / key press), not per drag tick. */
onChange: (v: number) => void;
disabled?: boolean;
format?: (v: number) => string;
}) {
// Track the in-progress drag locally so `onChange` only fires on commit;
// any external `value` change (commit landing, reset button) takes over.
const [dragValue, setDragValue] = React.useState<number | null>(null);
React.useEffect(() => {
setDragValue(null);
}, [value]);
const shown = dragValue ?? value;
return (
<FieldRow htmlFor={id} label={label} title={title}>
<div className="flex items-center gap-2">
<Slider
id={id}
min={min}
max={max}
step={step}
value={[shown]}
onValueChange={(v) => setDragValue(v[0])}
onValueCommit={(v) => onChange(v[0])}
disabled={disabled}
aria-label={label}
className="min-w-0 flex-1"
/>
<span
aria-hidden="true"
className="min-w-10 shrink-0 text-right text-sm tabular-nums text-muted-foreground"
>
{format(shown)}
</span>
</div>
</FieldRow>
);
}
export function ToggleRow({
id,
label,
title,
checked,
onChange,
disabled,
}: {
id: string;
label: string;
title?: string;
checked: boolean;
onChange: (v: boolean) => void;
disabled?: boolean;
}) {
return (
<FieldRow htmlFor={id} label={label} title={title}>
<Switch
id={id}
checked={checked}
onCheckedChange={onChange}
disabled={disabled}
/>
</FieldRow>
);
}
export function NumberRow({
id,
label,
title,
min,
max,
step,
value,
onChange,
disabled,
}: {
id: string;
label: string;
title?: string;
min?: number;
max?: number;
step?: number | string;
value: number;
onChange: (v: number) => void;
disabled?: boolean;
}) {
// Buffer the raw text so the field can be emptied while retyping; only
// valid numbers are committed, and blur restores the last committed value.
const [draft, setDraft] = React.useState<string | null>(null);
React.useEffect(() => {
setDraft(null);
}, [value]);
return (
<FieldRow htmlFor={id} label={label} title={title}>
<Input
id={id}
type="number"
min={min}
max={max}
step={step}
value={draft ?? value}
onChange={(e) => {
setDraft(e.target.value);
const v = e.target.valueAsNumber;
if (!Number.isNaN(v)) onChange(v);
}}
onBlur={() => setDraft(null)}
disabled={disabled}
/>
</FieldRow>
);
}
@@ -0,0 +1,53 @@
import { render, screen } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { describe, expect, it, vi } from 'vitest';
import CreateJobButton from './CreateJobButton';
vi.mock('./CreateJobModal', () => ({
default: ({
isOpen,
workloadType,
}: {
isOpen: boolean;
workloadType: string;
}) =>
isOpen ? (
<div role="dialog" data-workload-type={workloadType}>
Create job form
</div>
) : null,
}));
describe('CreateJobButton', () => {
it('opens the workload menu on click and selects an item', async () => {
const user = userEvent.setup();
render(<CreateJobButton jobType="inference" />);
await user.click(screen.getByRole('button', { name: 'Create Job' }));
await user.click(screen.getByRole('menuitem', { name: /I2V/i }));
expect(screen.getByRole('dialog')).toHaveAttribute(
'data-workload-type',
'i2v',
);
});
it('opens and operates the workload menu from the keyboard', async () => {
const user = userEvent.setup();
render(<CreateJobButton jobType="inference" />);
const trigger = screen.getByRole('button', { name: 'Create Job' });
trigger.focus();
await user.keyboard('{Enter}');
const firstItem = await screen.findByRole('menuitem', { name: /T2V/i });
expect(firstItem).toHaveFocus();
await user.keyboard('{Enter}');
expect(screen.getByRole('dialog')).toHaveAttribute(
'data-workload-type',
't2v',
);
});
});
@@ -0,0 +1,75 @@
'use client';
import * as React from 'react';
import { ChevronDown } from 'lucide-react';
import * as DropdownMenu from '@radix-ui/react-dropdown-menu';
import CreateJobModal from '@/components/jobs/CreateJobModal';
import { Button } from '@/components/ui/button';
import { WORKLOAD_OPTIONS } from '@/lib/jobConfig';
import type { JobType } from '@/lib/types';
import { triggerRefresh } from '@/stores/jobsRefresh';
interface CreateJobButtonProps {
jobType: JobType;
}
export default function CreateJobButton({ jobType }: CreateJobButtonProps) {
const options = WORKLOAD_OPTIONS[jobType] ?? [];
const [modalOpen, setModalOpen] = React.useState(false);
const [workloadType, setWorkloadType] = React.useState(
options[0]?.type ?? 't2v',
);
function openModal(type: string) {
setWorkloadType(type);
setModalOpen(true);
}
function handleSuccess() {
triggerRefresh();
setModalOpen(false);
}
return (
<>
<DropdownMenu.Root>
<DropdownMenu.Trigger asChild>
<Button type="button" className="gap-1.5">
Create Job
<ChevronDown className="size-3.5 opacity-85" aria-hidden />
</Button>
</DropdownMenu.Trigger>
<DropdownMenu.Portal>
<DropdownMenu.Content
align="end"
sideOffset={4}
collisionPadding={8}
className="z-[200] min-w-48 overflow-hidden rounded-lg border border-border bg-popover py-1 text-popover-foreground shadow-lg"
>
{options.map((opt) => (
<DropdownMenu.Item
key={opt.type}
onSelect={() => openModal(opt.type)}
className="flex min-h-11 cursor-pointer select-none flex-col justify-center px-4 py-2 text-left text-sm font-medium outline-none data-[highlighted]:bg-secondary"
>
{opt.label}
<span className="mt-0.5 block text-xs font-normal text-muted-foreground">
{opt.desc}
</span>
</DropdownMenu.Item>
))}
</DropdownMenu.Content>
</DropdownMenu.Portal>
</DropdownMenu.Root>
<CreateJobModal
isOpen={modalOpen}
onClose={() => setModalOpen(false)}
onSuccess={handleSuccess}
jobType={jobType}
workloadType={workloadType}
/>
</>
);
}
@@ -0,0 +1,209 @@
import * as React from 'react';
import { render, screen, waitFor, within } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import CreateJobModal from './CreateJobModal';
import { createJob, getDatasets, getModels, uploadImage } from '@/lib/api';
import { defaultOptionsStore } from '@/stores/defaultOptions';
import { DEFAULT_OPTIONS } from '@/lib/defaultOptions';
vi.mock('@/lib/api', () => ({
createJob: vi.fn(),
getModels: vi.fn(),
getDatasets: vi.fn(),
uploadImage: vi.fn(),
getSettings: vi.fn(),
updateSettings: vi.fn(),
}));
const MODELS = [
{ id: 'wan/t2v-1.3b', label: 'Wan T2V', type: 't2v' },
{ id: 'wan/t2v-14b', label: 'Wan T2V Large', type: 't2v' },
];
beforeEach(() => {
// Reset the shared options store to a known baseline for test isolation.
defaultOptionsStore.set({ options: DEFAULT_OPTIONS });
vi.mocked(getModels).mockResolvedValue(MODELS);
vi.mocked(getDatasets).mockResolvedValue([]);
vi.mocked(uploadImage).mockResolvedValue({ path: '/uploads/x.png' });
vi.mocked(createJob).mockResolvedValue({ id: 'job-1' } as never);
});
function renderModal(
overrides: Partial<React.ComponentProps<typeof CreateJobModal>> = {},
) {
const onClose = vi.fn();
const onSuccess = vi.fn();
render(
<CreateJobModal
isOpen
onClose={onClose}
onSuccess={onSuccess}
jobType="inference"
workloadType="t2v"
{...overrides}
/>,
);
return { onClose, onSuccess };
}
describe('CreateJobModal', () => {
it('shows a model loading error instead of an empty model list', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
vi.mocked(getModels).mockRejectedValueOnce(new Error('network down'));
renderModal();
expect(
await screen.findByText(/Models could not be loaded/),
).toBeInTheDocument();
expect(screen.getByLabelText('Model')).toHaveAttribute(
'aria-invalid',
'true',
);
});
it('keeps the form open and reports job creation failures', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
vi.mocked(createJob).mockRejectedValueOnce(new Error('API rejected job'));
const user = userEvent.setup();
const { onClose, onSuccess } = renderModal();
await screen.findByRole('option', { name: 'Wan T2V (wan/t2v-1.3b)' });
await user.type(screen.getByLabelText('Prompt'), 'a careful test prompt');
await user.click(screen.getByRole('button', { name: 'Create Job' }));
expect(
await screen.findByText(/API rejected job.*then try again/),
).toBeInTheDocument();
expect(onSuccess).not.toHaveBeenCalled();
expect(onClose).not.toHaveBeenCalled();
});
it('reports image upload failures next to the file input', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
vi.mocked(uploadImage).mockRejectedValueOnce(new Error('Upload failed'));
const user = userEvent.setup();
renderModal({ workloadType: 'i2v' });
await screen.findByRole('option', { name: 'Wan T2V (wan/t2v-1.3b)' });
const input = screen.getByLabelText('Image');
await user.upload(
input,
new File(['image'], 'input.png', { type: 'image/png' }),
);
expect(
await screen.findByText(/Upload failed.*Choose the image again/),
).toBeInTheDocument();
expect(input).toHaveAttribute('aria-invalid', 'true');
});
it('renders the form fields for an inference job', async () => {
renderModal();
expect(
await screen.findByText('New Inference Job (T2V)'),
).toBeInTheDocument();
expect(screen.getByLabelText('Model')).toBeInTheDocument();
expect(screen.getByLabelText('Prompt')).toBeInTheDocument();
expect(screen.getByLabelText('Negative Prompt')).toBeInTheDocument();
expect(
screen.getByRole('button', { name: 'Create Job' }),
).toBeInTheDocument();
// The model dropdown is populated once getModels resolves.
expect(
await screen.findByRole('option', {
name: 'Wan T2V (wan/t2v-1.3b)',
}),
).toBeInTheDocument();
});
it('seeds fields from the options store and submits an inference payload', async () => {
// Non-default store values prove the open-time seeding effect ran (the
// useState defaults are 50 / 480).
defaultOptionsStore.set({
options: { ...DEFAULT_OPTIONS, numInferenceSteps: 25, height: 720 },
});
const user = userEvent.setup();
const { onClose, onSuccess } = renderModal();
// Wait for models to load so the default model is selected.
await screen.findByRole('option', { name: 'Wan T2V (wan/t2v-1.3b)' });
await user.type(
screen.getByLabelText('Prompt'),
'a raccoon in sunflowers',
);
await user.click(screen.getByRole('button', { name: 'Create Job' }));
await waitFor(() => expect(createJob).toHaveBeenCalledTimes(1));
const payload = vi.mocked(createJob).mock.calls[0][0];
expect(payload).toMatchObject({
model_id: 'wan/t2v-1.3b',
prompt: 'a raccoon in sunflowers',
workload_type: 't2v',
job_type: 'inference',
num_inference_steps: 25,
height: 720,
num_frames: 81,
width: 832,
guidance_scale: 5,
seed: 1024,
});
await waitFor(() => expect(onSuccess).toHaveBeenCalledTimes(1));
expect(onClose).toHaveBeenCalledTimes(1);
});
it('submits a dmd_t2v distillation payload including the DMD fields', async () => {
vi.mocked(getDatasets).mockResolvedValue([
{ id: 'ds1', name: 'My Dataset', created_at: 0 },
]);
const user = userEvent.setup();
const { onSuccess } = renderModal({
jobType: 'distillation',
workloadType: 'dmd_t2v',
});
// Models + datasets load asynchronously on open. The model/dataset option
// labels also appear in the Real/Fake Score Model and Validation Dataset
// selects, so scope each wait to the relevant select.
await within(screen.getByLabelText('Model')).findByRole('option', {
name: 'Wan T2V (wan/t2v-1.3b)',
});
const datasetSelect = screen.getByLabelText('Dataset *');
await within(datasetSelect).findByRole('option', { name: 'My Dataset' });
await user.type(screen.getByLabelText('Description'), 'distill run');
await user.selectOptions(datasetSelect, 'ds1');
await user.click(screen.getByRole('button', { name: 'Create Job' }));
await waitFor(() => expect(createJob).toHaveBeenCalledTimes(1));
const payload = vi.mocked(createJob).mock.calls[0][0];
expect(payload).toMatchObject({
workload_type: 'dmd_t2v',
job_type: 'distillation',
// The dataset id is sent; the backend resolves it to the on-disk dir.
data_path: 'ds1',
lora_rank: 32,
// DMD-specific fields added to CreateJobRequest for this modal.
dmd_use_vsa: false,
dmd_vsa_sparsity: 0.8,
dmd_denoising_steps: '1000,757,522',
real_score_guidance_scale: 3.5,
generator_update_interval: 5,
real_score_model_path: 'wan/t2v-1.3b',
fake_score_model_path: 'wan/t2v-1.3b',
});
// Inference-only keys must be absent for a training job.
expect(payload).not.toHaveProperty('num_inference_steps');
await waitFor(() => expect(onSuccess).toHaveBeenCalledTimes(1));
});
});
@@ -0,0 +1,931 @@
'use client';
import * as React from 'react';
import {
FieldRow,
NumberRow,
SliderRow,
ToggleRow,
} from '@/components/form-rows';
import {
Dialog,
DialogContent,
DialogHeader,
DialogTitle,
} from '@/components/ui/dialog';
import { Button } from '@/components/ui/button';
import { Input } from '@/components/ui/input';
import { NativeSelect } from '@/components/ui/native-select';
import { Textarea } from '@/components/ui/textarea';
import { useStore } from '@/hooks/useStore';
import { defaultOptionsStore } from '@/stores/defaultOptions';
import {
createJob,
getDatasets,
getModels,
uploadImage,
type CreateJobRequest,
type Model,
} from '@/lib/api';
import { getDefaultModelForWorkload } from '@/lib/defaultOptions';
import { WORKLOAD_OPTIONS } from '@/lib/jobConfig';
import type { JobType } from '@/lib/types';
export interface CreateJobModalProps {
isOpen: boolean;
onClose: () => void;
onSuccess: () => void;
jobType: JobType;
workloadType: string;
}
export default function CreateJobModal({
isOpen,
onClose,
onSuccess,
jobType,
workloadType,
}: CreateJobModalProps) {
const { options } = useStore(defaultOptionsStore);
const isInference = jobType === 'inference';
// Training jobs always pick from the t2v model catalogue.
const inferenceWorkload = isInference ? workloadType : 't2v';
const [models, setModels] = React.useState<Model[]>([]);
const [modelId, setModelId] = React.useState('');
const [prompt, setPrompt] = React.useState('');
const [imagePath, setImagePath] = React.useState('');
const [imageFileName, setImageFileName] = React.useState('');
const [isUploadingImage, setIsUploadingImage] = React.useState(false);
const [negativePrompt, setNegativePrompt] = React.useState('');
const [numInferenceSteps, setNumInferenceSteps] = React.useState(50);
const [numFrames, setNumFrames] = React.useState(81);
const [height, setHeight] = React.useState(480);
const [width, setWidth] = React.useState(832);
const [guidanceScale, setGuidanceScale] = React.useState(5);
const [guidanceRescale, setGuidanceRescale] = React.useState(0);
const [fps, setFps] = React.useState(24);
const [seed, setSeed] = React.useState(1024);
const [numGpus, setNumGpus] = React.useState(1);
const [ditCpuOffload, setDitCpuOffload] = React.useState(false);
const [textEncoderCpuOffload, setTextEncoderCpuOffload] =
React.useState(false);
const [vaeCpuOffload, setVaeCpuOffload] = React.useState(false);
const [imageEncoderCpuOffload, setImageEncoderCpuOffload] =
React.useState(false);
const [useFsdpInference, setUseFsdpInference] = React.useState(false);
const [enableTorchCompile, setEnableTorchCompile] = React.useState(false);
const [vsaSparsity, setVsaSparsity] = React.useState(0);
const [tpSize, setTpSize] = React.useState(-1);
const [spSize, setSpSize] = React.useState(-1);
const [selectedDatasetId, setSelectedDatasetId] = React.useState('');
const [readyDatasets, setReadyDatasets] = React.useState<
Awaited<ReturnType<typeof getDatasets>>
>([]);
const [maxTrainSteps, setMaxTrainSteps] = React.useState(1000);
const [trainBatchSize, setTrainBatchSize] = React.useState(1);
const [learningRate, setLearningRate] = React.useState(5e-5);
const [numLatentT, setNumLatentT] = React.useState(20);
const [selectedValidationDatasetId, setSelectedValidationDatasetId] =
React.useState('');
const [loraRank, setLoraRank] = React.useState(32);
const [dmdUseVsa, setDmdUseVsa] = React.useState(false);
const [dmdVsaSparsity, setDmdVsaSparsity] = React.useState(0.8);
const [dmdDenoisingSteps, setDmdDenoisingSteps] =
React.useState('1000,757,522');
const [realScoreGuidanceScale, setRealScoreGuidanceScale] =
React.useState(3.5);
const [generatorUpdateInterval, setGeneratorUpdateInterval] =
React.useState(5);
const [realScoreModelPath, setRealScoreModelPath] = React.useState('');
const [fakeScoreModelPath, setFakeScoreModelPath] = React.useState('');
const [isSubmitting, setIsSubmitting] = React.useState(false);
const [isLoadingModels, setIsLoadingModels] = React.useState(false);
const [isLoadingDatasets, setIsLoadingDatasets] = React.useState(false);
const [modelLoadError, setModelLoadError] = React.useState<string | null>(
null,
);
const [datasetLoadError, setDatasetLoadError] = React.useState<string | null>(
null,
);
const [imageUploadError, setImageUploadError] = React.useState<string | null>(
null,
);
const [submitError, setSubmitError] = React.useState<string | null>(null);
const imageInputRef = React.useRef<HTMLInputElement>(null);
// Seed field values from the persisted default options each time the modal
// OPENS. A naive port of the Svelte `$effect` would re-seed on every
// `$defaultOptions` change; we deliberately seed only on the open transition
// so a late `initDefaultOptions()` settings refresh can't clobber the user's
// in-progress edits or desync the (already validated) model selection.
const justOpenedRef = React.useRef(false);
React.useEffect(() => {
const justOpened = isOpen && !justOpenedRef.current;
justOpenedRef.current = isOpen;
if (!justOpened) return;
const opts = options;
setNumInferenceSteps(opts.numInferenceSteps);
setNumFrames(workloadType === 't2i' ? 1 : opts.numFrames);
setHeight(opts.height);
setWidth(opts.width);
setGuidanceScale(opts.guidanceScale);
setGuidanceRescale(opts.guidanceRescale);
setFps(opts.fps);
setSeed(opts.seed);
setNumGpus(opts.numGpus);
setDitCpuOffload(opts.ditCpuOffload);
setTextEncoderCpuOffload(opts.textEncoderCpuOffload);
setVaeCpuOffload(opts.vaeCpuOffload);
setImageEncoderCpuOffload(opts.imageEncoderCpuOffload);
setUseFsdpInference(opts.useFsdpInference);
setEnableTorchCompile(opts.enableTorchCompile);
setVsaSparsity(opts.vsaSparsity);
setTpSize(opts.tpSize);
setSpSize(opts.spSize);
setModelId(
getDefaultModelForWorkload(
opts,
inferenceWorkload as 't2v' | 'i2v' | 't2i',
),
);
setImagePath('');
setImageFileName('');
setSelectedDatasetId('');
setSelectedValidationDatasetId('');
setModelLoadError(null);
setDatasetLoadError(null);
setImageUploadError(null);
setSubmitError(null);
if (workloadType === 'dmd_t2v') {
setDmdUseVsa(false);
setDmdVsaSparsity(0.8);
setDmdDenoisingSteps('1000,757,522');
setRealScoreGuidanceScale(3.5);
setGeneratorUpdateInterval(5);
setRealScoreModelPath('');
setFakeScoreModelPath('');
}
}, [isOpen, workloadType, inferenceWorkload, options]);
// Load the models available for this workload.
React.useEffect(() => {
if (!isOpen) return;
// Ignore a superseded response so a slow fetch for a previous workload
// can't overwrite the current workload's model list/selection.
let stale = false;
setIsLoadingModels(true);
setModelLoadError(null);
getModels(inferenceWorkload)
.then((list) => {
if (stale) return;
setModels(list);
const ids = list.map((m) => m.id);
const opts = defaultOptionsStore.get().options;
const defaultId = getDefaultModelForWorkload(
opts,
inferenceWorkload as 't2v' | 'i2v' | 't2i',
);
const chosen = ids.includes(defaultId) ? defaultId : (list[0]?.id ?? '');
setModelId(chosen);
if (workloadType === 'dmd_t2v') {
setRealScoreModelPath(chosen);
setFakeScoreModelPath(chosen);
}
})
.catch((e) => {
if (stale) return;
console.error('Failed to load models:', e);
setModels([]);
setModelId('');
setModelLoadError(
'Models could not be loaded. Check the Studio API and reopen this form to try again.',
);
})
.finally(() => {
if (!stale) setIsLoadingModels(false);
});
return () => {
stale = true;
};
}, [isOpen, inferenceWorkload, workloadType]);
// Training jobs need a dataset; load the ready datasets when relevant.
React.useEffect(() => {
if (isOpen && !isInference) {
setIsLoadingDatasets(true);
setDatasetLoadError(null);
getDatasets()
.then(setReadyDatasets)
.catch((error) => {
console.error('Failed to load datasets:', error);
setReadyDatasets([]);
setDatasetLoadError(
'Datasets could not be loaded. Check the Studio API and reopen this form to try again.',
);
})
.finally(() => setIsLoadingDatasets(false));
} else {
setReadyDatasets([]);
setIsLoadingDatasets(false);
setDatasetLoadError(null);
}
}, [isOpen, isInference]);
async function handleImageChange(e: React.ChangeEvent<HTMLInputElement>) {
const file = e.target.files?.[0];
if (!file) {
setImagePath('');
setImageFileName('');
setImageUploadError(null);
return;
}
setIsUploadingImage(true);
setImageFileName(file.name);
setImageUploadError(null);
try {
const { path } = await uploadImage(file);
setImagePath(path);
} catch (error) {
console.error('Failed to upload image:', error);
setImagePath('');
setImageFileName('');
setImageUploadError(
error instanceof Error
? `${error.message}. Choose the image again to retry.`
: 'The image could not be uploaded. Choose it again to retry.',
);
} finally {
setIsUploadingImage(false);
}
}
function clearImage() {
setImagePath('');
setImageFileName('');
setImageUploadError(null);
if (imageInputRef.current) imageInputRef.current.value = '';
}
async function handleSubmit(e: React.FormEvent<HTMLFormElement>) {
e.preventDefault();
if (isInference && workloadType === 'i2v' && !imagePath) return;
// Send the dataset id; the backend resolves it to the on-disk media dir.
const effectiveDataPath = selectedDatasetId ?? '';
if (!isInference && !selectedDatasetId) return;
// `lora_t2v` jobs are persisted with a dedicated backend job_type that the
// front-end JobType enum does not model; cast to keep payload parity.
const effectiveJobType = (
workloadType === 'lora_t2v' ? 'lora' : jobType
) as JobType;
setIsSubmitting(true);
setSubmitError(null);
try {
const payload: CreateJobRequest = {
model_id: modelId,
prompt,
workload_type: workloadType,
job_type: effectiveJobType,
...(isInference
? {
...(workloadType === 'i2v' && imagePath
? { image_path: imagePath }
: {}),
negative_prompt: negativePrompt,
num_inference_steps: numInferenceSteps,
num_frames: numFrames,
height,
width,
guidance_scale: guidanceScale,
guidance_rescale: guidanceRescale,
fps,
seed,
num_gpus: numGpus,
dit_cpu_offload: ditCpuOffload,
text_encoder_cpu_offload: textEncoderCpuOffload,
vae_cpu_offload: vaeCpuOffload,
image_encoder_cpu_offload: imageEncoderCpuOffload,
use_fsdp_inference: useFsdpInference,
enable_torch_compile: enableTorchCompile,
vsa_sparsity: vsaSparsity,
tp_size: tpSize,
sp_size: spSize,
}
: {
data_path: effectiveDataPath.trim(),
max_train_steps: maxTrainSteps,
train_batch_size: trainBatchSize,
learning_rate: learningRate,
num_latent_t: numLatentT,
validation_dataset_file: selectedValidationDatasetId || undefined,
lora_rank: loraRank,
...(workloadType === 'dmd_t2v'
? {
dmd_use_vsa: dmdUseVsa,
dmd_vsa_sparsity: dmdVsaSparsity,
dmd_denoising_steps: dmdDenoisingSteps,
real_score_guidance_scale: realScoreGuidanceScale,
generator_update_interval: generatorUpdateInterval,
real_score_model_path: realScoreModelPath || modelId,
fake_score_model_path: fakeScoreModelPath || modelId,
}
: {}),
}),
};
await createJob(payload);
onSuccess();
onClose();
} catch (err) {
console.error('Failed to create job:', err);
setSubmitError(
err instanceof Error
? `${err.message}. Check the form and Studio API, then try again.`
: 'The job could not be created. Check the form and Studio API, then try again.',
);
} finally {
setIsSubmitting(false);
}
}
function handleClose() {
if (isSubmitting) return;
onClose();
}
const workloadLabel =
WORKLOAD_OPTIONS[jobType]?.find((o) => o.type === workloadType)?.label ?? '';
const title = `New ${jobType.charAt(0).toUpperCase() + jobType.slice(1)} Job${
workloadLabel ? ` (${workloadLabel})` : ''
}`;
return (
<Dialog
open={isOpen}
onOpenChange={(open) => {
if (!open) handleClose();
}}
>
<DialogContent
className="max-h-[90vh] w-[90vw] max-w-[850px] overflow-y-auto"
onEscapeKeyDown={(e) => {
if (isSubmitting) e.preventDefault();
}}
onInteractOutside={(e) => {
if (isSubmitting) e.preventDefault();
}}
>
<DialogHeader>
<DialogTitle>{title}</DialogTitle>
</DialogHeader>
<form
onSubmit={handleSubmit}
autoComplete="off"
className="flex flex-col gap-3.5"
>
<FieldRow htmlFor="modal-modelId" label="Model">
<NativeSelect
id="modal-modelId"
value={modelId}
onChange={(e) => setModelId(e.target.value)}
required
aria-describedby={
modelLoadError ? 'modal-model-error' : undefined
}
aria-invalid={modelLoadError ? true : undefined}
disabled={isSubmitting || isLoadingModels || !!modelLoadError}
>
<option value="" disabled>
{isLoadingModels
? 'Loading models…'
: models.length === 0
? 'No models available for this workload'
: 'Select a model…'}
</option>
{models.map((model) => (
<option key={model.id} value={model.id}>
{model.label} ({model.id})
</option>
))}
</NativeSelect>
{modelLoadError && (
<p
id="modal-model-error"
role="alert"
className="text-sm text-destructive"
>
{modelLoadError}
</p>
)}
</FieldRow>
{isInference && workloadType === 'i2v' && (
<FieldRow htmlFor="modal-image" label="Image">
<Input
ref={imageInputRef}
id="modal-image"
type="file"
accept=".png,.jpg,.jpeg,.webp,.bmp"
onChange={handleImageChange}
disabled={isSubmitting || isUploadingImage}
aria-describedby={
imageUploadError ? 'modal-image-error' : undefined
}
aria-invalid={imageUploadError ? true : undefined}
required
className="h-auto py-2 file:mr-3 file:cursor-pointer file:rounded-md file:border-0 file:bg-secondary file:px-2 file:py-1 file:text-sm file:text-secondary-foreground"
/>
{imageFileName && (
<span className="mt-0.5 text-xs text-muted-foreground">
{isUploadingImage ? 'Uploading…' : imageFileName} ·{' '}
<button
type="button"
onClick={clearImage}
disabled={isSubmitting || isUploadingImage}
className="text-accent-blue underline-offset-2 hover:underline disabled:cursor-not-allowed disabled:opacity-50"
>
Clear
</button>
</span>
)}
{imageUploadError && (
<p
id="modal-image-error"
role="alert"
className="text-sm text-destructive"
>
{imageUploadError}
</p>
)}
</FieldRow>
)}
<FieldRow
htmlFor="modal-prompt"
label={isInference ? 'Prompt' : 'Description'}
>
<Textarea
id="modal-prompt"
value={prompt}
onChange={(e) => setPrompt(e.target.value)}
rows={isInference ? 3 : 2}
placeholder={
isInference
? 'A curious raccoon peers through a vibrant field of yellow sunflowers…'
: 'Brief description of this training job…'
}
required
disabled={isSubmitting}
/>
</FieldRow>
{isInference && (
<FieldRow htmlFor="modal-negative-prompt" label="Negative Prompt">
<Textarea
id="modal-negative-prompt"
value={negativePrompt}
onChange={(e) => setNegativePrompt(e.target.value)}
rows={2}
placeholder="Optional: things to avoid in the output…"
disabled={isSubmitting}
/>
</FieldRow>
)}
{!isInference && (
<>
<div className="flex gap-4">
<FieldRow
htmlFor="modal-dataset"
label="Dataset *"
className="min-w-0 flex-1"
>
<NativeSelect
id="modal-dataset"
value={selectedDatasetId}
onChange={(e) => setSelectedDatasetId(e.target.value)}
aria-describedby={
datasetLoadError ? 'modal-dataset-error' : undefined
}
aria-invalid={datasetLoadError ? true : undefined}
disabled={
isSubmitting || isLoadingDatasets || !!datasetLoadError
}
>
<option value="" disabled>
{isLoadingDatasets
? 'Loading datasets…'
: datasetLoadError
? 'Datasets unavailable'
: readyDatasets.length === 0
? 'No datasets (add in Datasets tab)'
: 'Select a dataset…'}
</option>
{readyDatasets.map((d) => (
<option key={d.id} value={d.id}>
{d.name}
</option>
))}
</NativeSelect>
{datasetLoadError && (
<p
id="modal-dataset-error"
role="alert"
className="text-sm text-destructive"
>
{datasetLoadError}
</p>
)}
</FieldRow>
<FieldRow
htmlFor="modal-validation-dataset"
label="Validation Dataset (optional)"
className="min-w-0 flex-1"
>
<NativeSelect
id="modal-validation-dataset"
value={selectedValidationDatasetId}
onChange={(e) =>
setSelectedValidationDatasetId(e.target.value)
}
disabled={
isSubmitting || isLoadingDatasets || !!datasetLoadError
}
>
<option value="">None</option>
{readyDatasets.map((d) => (
<option key={d.id} value={d.id}>
{d.name}
</option>
))}
</NativeSelect>
</FieldRow>
</div>
<details>
<summary className="mb-2 cursor-pointer select-none text-sm font-medium text-accent-blue">
Options
</summary>
<div className="grid grid-cols-[repeat(auto-fill,minmax(160px,1fr))] gap-x-3 gap-y-2">
<SliderRow
id="modal-max-train-steps"
label="Max Train Steps"
min={100}
max={50000}
step={100}
value={maxTrainSteps}
onChange={setMaxTrainSteps}
disabled={isSubmitting}
/>
<SliderRow
id="modal-train-batch-size"
label="Train Batch Size"
min={1}
max={8}
step={1}
value={trainBatchSize}
onChange={setTrainBatchSize}
disabled={isSubmitting}
/>
<NumberRow
id="modal-learning-rate"
label="Learning Rate"
step="1e-6"
min={1e-6}
max={1}
value={learningRate}
onChange={setLearningRate}
disabled={isSubmitting}
/>
<SliderRow
id="modal-num-latent-t"
label="Num Latent T"
min={8}
max={40}
step={1}
value={numLatentT}
onChange={setNumLatentT}
disabled={isSubmitting}
/>
{workloadType === 'lora_t2v' && (
<SliderRow
id="modal-lora-rank"
label="LoRA Rank"
min={8}
max={128}
step={8}
value={loraRank}
onChange={setLoraRank}
disabled={isSubmitting}
/>
)}
{workloadType === 'dmd_t2v' && (
<>
<ToggleRow
id="modal-dmd-use-vsa"
label="Video Sparse Attention (VSA)"
title="Use Video Sparse Attention for DMD"
checked={dmdUseVsa}
onChange={setDmdUseVsa}
disabled={isSubmitting}
/>
{dmdUseVsa && (
<SliderRow
id="modal-dmd-vsa-sparsity"
label="VSA Sparsity"
title="VSA sparsity (0–1)"
min={0}
max={1}
step={0.05}
value={dmdVsaSparsity}
onChange={setDmdVsaSparsity}
disabled={isSubmitting}
format={(v) => v.toFixed(2)}
/>
)}
<FieldRow
htmlFor="modal-dmd-denoising-steps"
label="DMD Denoising Steps"
title="Comma-separated denoising steps, e.g. 1000,757,522"
>
<Input
id="modal-dmd-denoising-steps"
type="text"
value={dmdDenoisingSteps}
onChange={(e) => setDmdDenoisingSteps(e.target.value)}
placeholder="1000,757,522"
disabled={isSubmitting}
/>
</FieldRow>
<SliderRow
id="modal-real-score-guidance-scale"
label="Real Score Guidance Scale"
min={1}
max={10}
step={0.1}
value={realScoreGuidanceScale}
onChange={setRealScoreGuidanceScale}
disabled={isSubmitting}
format={(v) => v.toFixed(1)}
/>
<SliderRow
id="modal-generator-update-interval"
label="Generator Update Interval"
min={1}
max={20}
step={1}
value={generatorUpdateInterval}
onChange={setGeneratorUpdateInterval}
disabled={isSubmitting}
/>
{(
[
{
id: 'modal-real-score-model',
label: 'Real Score Model',
value: realScoreModelPath,
onChange: setRealScoreModelPath,
},
{
id: 'modal-fake-score-model',
label: 'Fake Score Model',
value: fakeScoreModelPath,
onChange: setFakeScoreModelPath,
},
] as const
).map((select) => (
<FieldRow
key={select.id}
htmlFor={select.id}
label={select.label}
>
<NativeSelect
id={select.id}
value={select.value}
onChange={(e) => select.onChange(e.target.value)}
disabled={isSubmitting || isLoadingModels}
>
<option value="">Same as main model</option>
{models.map((model) => (
<option key={model.id} value={model.id}>
{model.label} ({model.id})
</option>
))}
</NativeSelect>
</FieldRow>
))}
</>
)}
</div>
</details>
</>
)}
{isInference && (
<details>
<summary className="mb-2 cursor-pointer select-none text-sm font-medium text-accent-blue">
Options
</summary>
<div className="grid grid-cols-[repeat(auto-fill,minmax(160px,1fr))] gap-x-3 gap-y-2">
{workloadType !== 't2i' && (
<SliderRow
id="modal-num-frames"
label="Frames"
min={1}
max={500}
step={1}
value={numFrames}
onChange={setNumFrames}
disabled={isSubmitting}
/>
)}
<SliderRow
id="modal-height"
label="Height"
min={64}
max={1080}
step={16}
value={height}
onChange={setHeight}
disabled={isSubmitting}
/>
<SliderRow
id="modal-width"
label="Width"
min={64}
max={1920}
step={16}
value={width}
onChange={setWidth}
disabled={isSubmitting}
/>
<SliderRow
id="modal-num-steps"
label="Inference Steps"
min={1}
max={200}
step={1}
value={numInferenceSteps}
onChange={setNumInferenceSteps}
disabled={isSubmitting}
/>
<SliderRow
id="modal-vsa-sparsity"
label="VSA Sparsity"
title="VSA sparsity (0–1)"
min={0}
max={1}
step={0.05}
value={vsaSparsity}
onChange={setVsaSparsity}
disabled={isSubmitting}
format={(v) => v.toFixed(2)}
/>
<SliderRow
id="modal-guidance"
label="Guidance Scale"
min={0}
max={20}
step={0.1}
value={guidanceScale}
onChange={setGuidanceScale}
disabled={isSubmitting}
format={(v) => v.toFixed(1)}
/>
<SliderRow
id="modal-guidance-rescale"
label="Guidance Rescale"
title="0 = disabled"
min={0}
max={1}
step={0.05}
value={guidanceRescale}
onChange={setGuidanceRescale}
disabled={isSubmitting}
format={(v) => v.toFixed(2)}
/>
<SliderRow
id="modal-tp-size"
label="TP Size"
title="-1 = auto"
min={-1}
max={8}
step={1}
value={tpSize}
onChange={setTpSize}
disabled={isSubmitting}
format={(v) => (v === -1 ? 'Auto' : String(v))}
/>
<SliderRow
id="modal-sp-size"
label="SP Size"
title="-1 = auto"
min={-1}
max={8}
step={1}
value={spSize}
onChange={setSpSize}
disabled={isSubmitting}
format={(v) => (v === -1 ? 'Auto' : String(v))}
/>
{workloadType !== 't2i' && (
<SliderRow
id="modal-fps"
label="FPS"
min={1}
max={60}
step={1}
value={fps}
onChange={setFps}
disabled={isSubmitting}
/>
)}
<ToggleRow
id="modal-dit-cpu-offload"
label="DiT CPU Offload"
checked={ditCpuOffload}
onChange={setDitCpuOffload}
disabled={isSubmitting}
/>
<ToggleRow
id="modal-text-encoder-cpu-offload"
label="Text Encoder CPU Offload"
checked={textEncoderCpuOffload}
onChange={setTextEncoderCpuOffload}
disabled={isSubmitting}
/>
<ToggleRow
id="modal-use-fsdp-inference"
label="Use FSDP Inference"
checked={useFsdpInference}
onChange={setUseFsdpInference}
disabled={isSubmitting}
/>
<ToggleRow
id="modal-vae-cpu-offload"
label="VAE CPU Offload"
checked={vaeCpuOffload}
onChange={setVaeCpuOffload}
disabled={isSubmitting}
/>
<ToggleRow
id="modal-image-encoder-cpu-offload"
label="Image Encoder CPU Offload"
checked={imageEncoderCpuOffload}
onChange={setImageEncoderCpuOffload}
disabled={isSubmitting}
/>
<ToggleRow
id="modal-enable-torch-compile"
label="Torch Compile"
checked={enableTorchCompile}
onChange={setEnableTorchCompile}
disabled={isSubmitting}
/>
<SliderRow
id="modal-num-gpus"
label="GPUs"
min={1}
max={8}
step={1}
value={numGpus}
onChange={setNumGpus}
disabled={isSubmitting}
/>
<NumberRow
id="modal-seed"
label="Seed"
min={0}
value={seed}
onChange={setSeed}
disabled={isSubmitting}
/>
</div>
</details>
)}
<div className="flex flex-col items-start gap-2">
{submitError && (
<p role="alert" className="text-sm text-destructive">
{submitError}
</p>
)}
<Button
type="submit"
disabled={
isSubmitting ||
isUploadingImage ||
!!modelLoadError ||
!!datasetLoadError
}
>
{isSubmitting ? 'Creating…' : 'Create Job'}
</Button>
</div>
</form>
</DialogContent>
</Dialog>
);
}
@@ -0,0 +1,135 @@
import { render, screen, waitFor } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import JobCard from '@/components/jobs/JobCard';
import {
deleteJob,
downloadJobVideo,
startJob,
stopJob,
} from '@/lib/api';
import type { Job } from '@/lib/types';
import { activeJobStore, setActiveJobId } from '@/stores/activeJob';
import { makeJob as makeBaseJob } from '@/test/factories';
vi.mock('@/lib/api', () => ({
startJob: vi.fn(),
stopJob: vi.fn(),
deleteJob: vi.fn(),
downloadJobVideo: vi.fn(),
}));
vi.mock('@/lib/utils', async (importOriginal) => {
const actual = await importOriginal<typeof import('@/lib/utils')>();
return {
...actual,
downloadBlob: vi.fn(),
};
});
const makeJob = (overrides: Partial<Job> = {}): Job =>
makeBaseJob({
model_id: 'Wan2.1-T2V',
prompt: 'a cat surfing a wave',
created_at: 1_700_000_000,
num_inference_steps: 50,
num_frames: 81,
height: 480,
width: 832,
guidance_scale: 5,
seed: 42,
num_gpus: 1,
...overrides,
});
beforeEach(() => {
setActiveJobId(null);
vi.mocked(startJob).mockResolvedValue({} as Job);
vi.mocked(stopJob).mockResolvedValue({} as Job);
vi.mocked(deleteJob).mockResolvedValue(undefined);
vi.mocked(downloadJobVideo).mockResolvedValue(new Blob());
vi.spyOn(window, 'confirm').mockReturnValue(true);
vi.spyOn(window, 'alert').mockImplementation(() => {});
});
describe('JobCard', () => {
it('keeps selection and job action buttons as semantic siblings', () => {
render(<JobCard job={makeJob()} />);
const selectButton = screen.getByRole('button', { pressed: false });
const deleteButton = screen.getByRole('button', { name: 'Delete' });
expect(selectButton).toHaveTextContent('Wan2.1-T2V');
expect(selectButton).not.toContainElement(deleteButton);
});
it('renders the model, prompt, status and inference meta', () => {
render(<JobCard job={makeJob()} />);
expect(screen.getByText('Wan2.1-T2V')).toBeInTheDocument();
expect(screen.getByText('a cat surfing a wave')).toBeInTheDocument();
expect(screen.getByText('pending')).toBeInTheDocument();
expect(screen.getByText('81 frames')).toBeInTheDocument();
expect(screen.getByText('480×832')).toBeInTheDocument();
});
it('shows the workload type (not frames) for non-inference jobs', () => {
render(
<JobCard
job={makeJob({ job_type: 'finetuning', workload_type: 'lora_t2v' })}
/>,
);
expect(screen.getByText('lora t2v')).toBeInTheDocument();
expect(screen.queryByText('81 frames')).not.toBeInTheDocument();
});
it('starts a pending job and notifies the parent', async () => {
const onJobUpdated = vi.fn();
render(<JobCard job={makeJob({ status: 'pending' })} onJobUpdated={onJobUpdated} />);
await userEvent.click(screen.getByRole('button', { name: 'Start' }));
await waitFor(() => expect(startJob).toHaveBeenCalledWith('job-1'));
expect(onJobUpdated).toHaveBeenCalled();
});
it('stops a running job', async () => {
render(<JobCard job={makeJob({ status: 'running', started_at: Date.now() })} />);
await userEvent.click(screen.getByRole('button', { name: 'Stop' }));
await waitFor(() => expect(stopJob).toHaveBeenCalledWith('job-1'));
});
it('restarts a failed job via startJob', async () => {
render(<JobCard job={makeJob({ status: 'failed' })} />);
await userEvent.click(screen.getByRole('button', { name: 'Restart' }));
await waitFor(() => expect(startJob).toHaveBeenCalledWith('job-1'));
});
it('deletes when confirmed and skips when cancelled', async () => {
const { rerender } = render(<JobCard job={makeJob()} />);
await userEvent.click(screen.getByRole('button', { name: 'Delete' }));
await waitFor(() => expect(deleteJob).toHaveBeenCalledWith('job-1'));
vi.mocked(deleteJob).mockClear();
vi.mocked(window.confirm).mockReturnValue(false);
rerender(<JobCard job={makeJob()} />);
await userEvent.click(screen.getByRole('button', { name: 'Delete' }));
expect(deleteJob).not.toHaveBeenCalled();
});
it('downloads the video for a completed inference job', async () => {
render(
<JobCard
job={makeJob({ status: 'completed', output_path: '/out/video.mp4' })}
/>,
);
await userEvent.click(
screen.getByRole('button', { name: 'Download Video' }),
);
await waitFor(() => expect(downloadJobVideo).toHaveBeenCalledWith('job-1'));
});
it('selects the job when the card body is clicked', async () => {
render(<JobCard job={makeJob()} />);
await userEvent.click(screen.getByText('Wan2.1-T2V'));
expect(activeJobStore.get().activeJobId).toBe('job-1');
});
});
@@ -0,0 +1,240 @@
'use client';
import * as React from 'react';
import { Timer } from 'lucide-react';
import { Badge, type BadgeProps } from '@/components/ui/badge';
import { Button } from '@/components/ui/button';
import { useStore } from '@/hooks/useStore';
import {
deleteJob,
downloadJobVideo,
startJob,
stopJob,
} from '@/lib/api';
import type { Job } from '@/lib/types';
import { cn, downloadBlob } from '@/lib/utils';
import { activeJobStore, setActiveJobId } from '@/stores/activeJob';
interface JobCardProps {
job: Job;
onJobUpdated?: () => void;
}
function formatDuration(seconds: number): string {
const roundedSeconds = Math.round(seconds);
if (roundedSeconds < 60) return `${roundedSeconds}s`;
if (roundedSeconds < 3600) {
const mins = Math.floor(roundedSeconds / 60);
const secs = roundedSeconds % 60;
return secs > 0 ? `${mins}m ${secs}s` : `${mins}m`;
}
const hours = Math.floor(roundedSeconds / 3600);
const mins = Math.floor((roundedSeconds % 3600) / 60);
return mins > 0 ? `${hours}h ${mins}m` : `${hours}h`;
}
function computeElapsed(job: Job, currentTime: number): string | null {
if (!job.started_at) return null;
const endTime =
job.status === 'running' ? currentTime : (job.finished_at ?? 0);
if (!endTime && job.status !== 'running') return null;
const startedAtMs =
job.started_at < 1e12 ? job.started_at * 1000 : job.started_at;
const endTimeMs = endTime < 1e12 ? endTime * 1000 : endTime;
const elapsedSeconds = (endTimeMs - startedAtMs) / 1000;
if (elapsedSeconds <= 0) return null;
return formatDuration(elapsedSeconds);
}
const BADGE_VARIANTS: Record<string, BadgeProps['variant']> = {
pending: 'secondary',
running: 'warning',
completed: 'success',
ready: 'success',
failed: 'destructive',
stopped: 'secondary',
preprocessing: 'default',
};
export default function JobCard({ job, onJobUpdated }: JobCardProps) {
const { activeJobId } = useStore(activeJobStore);
const isSelected = activeJobId === job.id;
const [isLoading, setIsLoading] = React.useState(false);
const [currentTime, setCurrentTime] = React.useState(() => Date.now());
const elapsedTime = computeElapsed(job, currentTime);
React.useEffect(() => {
if (job.status !== 'running' || !job.started_at) return;
const interval = setInterval(() => setCurrentTime(Date.now()), 1000);
return () => clearInterval(interval);
}, [job.status, job.started_at]);
async function handleStart(e: React.MouseEvent) {
e.preventDefault();
e.stopPropagation();
if (isLoading || job.status === 'running' || job.status === 'completed')
return;
setIsLoading(true);
try {
await startJob(job.id);
onJobUpdated?.();
} catch (err) {
alert(err instanceof Error ? err.message : 'Failed to start job');
} finally {
setIsLoading(false);
}
}
async function handleStop(e: React.MouseEvent) {
e.preventDefault();
e.stopPropagation();
if (isLoading || job.status !== 'running') return;
setIsLoading(true);
try {
await stopJob(job.id);
onJobUpdated?.();
} catch (err) {
alert(err instanceof Error ? err.message : 'Failed to stop job');
} finally {
setIsLoading(false);
}
}
async function handleDelete(e: React.MouseEvent) {
e.preventDefault();
e.stopPropagation();
if (isLoading) return;
if (!confirm('Delete this job?')) return;
setIsLoading(true);
try {
await deleteJob(job.id);
onJobUpdated?.();
} catch (err) {
alert(err instanceof Error ? err.message : 'Failed to delete job');
} finally {
setIsLoading(false);
}
}
function handleSelectJob() {
setActiveJobId(isSelected ? null : job.id);
}
async function handleDownloadVideo(e: React.MouseEvent) {
e.preventDefault();
e.stopPropagation();
if (isLoading || !job.output_path) return;
setIsLoading(true);
try {
const blob = await downloadJobVideo(job.id);
const ext = job.output_path.endsWith('.png') ? 'png' : 'mp4';
downloadBlob(blob, `job_${job.id}.${ext}`);
} catch (err) {
alert(err instanceof Error ? err.message : 'Failed to download video');
} finally {
setIsLoading(false);
}
}
return (
<article
className={cn(
'mb-3 flex cursor-pointer flex-col gap-2.5 rounded-lg border bg-background p-4 transition-colors last:mb-0',
isSelected
? 'border-accent-blue bg-accent-blue/5'
: 'border-border hover:border-muted-foreground/40',
)}
>
<button
type="button"
aria-pressed={isSelected}
onClick={handleSelectJob}
className="flex w-full flex-col gap-2.5 rounded-md text-left"
>
<span className="flex flex-wrap items-center justify-between gap-2">
<span className="text-[0.95rem] font-semibold text-foreground">
{job.model_id}
</span>
<Badge variant={BADGE_VARIANTS[job.status] ?? 'secondary'}>
{job.status}
</Badge>
</span>
<span className="max-w-full overflow-hidden text-ellipsis whitespace-nowrap text-sm text-muted-foreground">
{job.prompt}
</span>
<span className="flex flex-wrap items-center gap-4 text-xs text-muted-foreground">
{job.job_type === 'inference' ? (
<>
<span>{job.num_frames} frames</span>
<span>
{job.height}×{job.width}
</span>
</>
) : (
<span>{job.workload_type?.replace(/_/g, ' ') ?? job.job_type}</span>
)}
{elapsedTime && (
<span className="inline-flex items-center gap-1">
<Timer className="size-3.5" aria-hidden />
{elapsedTime}
</span>
)}
</span>
</button>
<div className="flex flex-wrap items-center gap-1.5">
{job.status === 'running' ? (
<Button
size="sm"
onClick={handleStop}
disabled={isLoading}
className="border-transparent bg-amber-500 text-black shadow-md hover:bg-amber-400"
>
Stop
</Button>
) : job.status === 'failed' ? (
<Button
size="sm"
onClick={handleStart}
disabled={isLoading}
className="border-transparent bg-emerald-600 text-white shadow-md hover:bg-emerald-500"
>
Restart
</Button>
) : job.status === 'pending' || job.status === 'stopped' ? (
<Button
size="sm"
onClick={handleStart}
disabled={isLoading}
className="border-transparent bg-emerald-600 text-white shadow-md hover:bg-emerald-500"
>
Start
</Button>
) : null}
{job.status === 'completed' &&
job.output_path &&
job.job_type === 'inference' && (
<Button
size="sm"
variant="outline"
onClick={handleDownloadVideo}
disabled={isLoading}
title="Download video"
>
Download Video
</Button>
)}
<Button
size="sm"
variant="destructive"
onClick={handleDelete}
disabled={isLoading}
>
Delete
</Button>
</div>
</article>
);
}
@@ -0,0 +1,136 @@
import { act, render, screen } from '@testing-library/react';
import { describe, expect, it, vi } from 'vitest';
import JobDetailsSidebar from './JobDetailsSidebar';
import { getJobLogs } from '@/lib/api';
import type { Job } from '@/lib/types';
import { makeJob as makeBaseJob } from '@/test/factories';
vi.mock('@/lib/api', () => ({
getJobLogs: vi.fn(),
downloadJobLog: vi.fn(),
}));
const makeJob = (overrides: Partial<Job> = {}): Job =>
makeBaseJob({
status: 'running',
log_file_path: '/logs/job-1.log',
...overrides,
});
describe('JobDetailsSidebar', () => {
it('fills the mobile viewport without reserving main-content width', async () => {
vi.mocked(getJobLogs).mockResolvedValue({
lines: [],
total: 0,
progress: 0,
progress_msg: '',
phase: '',
});
const onWidthChange = vi.fn();
render(
<JobDetailsSidebar
job={makeJob({ status: 'completed' })}
isMobile
onClose={vi.fn()}
onWidthChange={onWidthChange}
/>,
);
const drawer = screen.getByRole('dialog', { name: 'Job details' });
expect(drawer).toHaveStyle({ width: '100%', maxWidth: 'none' });
expect(drawer).toHaveAttribute('aria-modal', 'true');
expect(drawer).toHaveFocus();
expect(onWidthChange).toHaveBeenCalledWith(0);
});
it('renders log lines streamed from the job log poll', async () => {
vi.mocked(getJobLogs).mockResolvedValue({
lines: ['boot sequence started', 'loading model weights'],
total: 2,
progress: 0,
progress_msg: '',
phase: '',
});
render(
<JobDetailsSidebar job={makeJob({ status: 'running' })} onClose={vi.fn()} />,
);
expect(await screen.findByText(/boot sequence started/)).toBeInTheDocument();
expect(screen.getByText(/loading model weights/)).toBeInTheDocument();
expect(getJobLogs).toHaveBeenCalledWith('job-1', 0);
});
it('keeps polling while the job is running', async () => {
vi.useFakeTimers();
try {
vi.mocked(getJobLogs).mockResolvedValue({
lines: [],
total: 0,
progress: 0,
progress_msg: '',
phase: '',
});
render(
<JobDetailsSidebar
job={makeJob({ status: 'running' })}
onClose={vi.fn()}
/>,
);
// Flush the immediate poll fired on mount.
await act(async () => {
await vi.advanceTimersByTimeAsync(0);
});
const initialCalls = vi.mocked(getJobLogs).mock.calls.length;
// Two 2s interval ticks should fire while the job is running.
await act(async () => {
await vi.advanceTimersByTimeAsync(4000);
});
expect(vi.mocked(getJobLogs).mock.calls.length).toBeGreaterThan(
initialCalls,
);
} finally {
vi.useRealTimers();
}
});
it('stops polling once the job is completed', async () => {
vi.useFakeTimers();
try {
vi.mocked(getJobLogs).mockResolvedValue({
lines: ['final line'],
total: 1,
progress: 0,
progress_msg: '',
phase: '',
});
render(
<JobDetailsSidebar
job={makeJob({ status: 'completed' })}
onClose={vi.fn()}
/>,
);
// The component fetches once on mount even for terminal jobs.
await act(async () => {
await vi.advanceTimersByTimeAsync(0);
});
expect(getJobLogs).toHaveBeenCalledTimes(1);
// No interval is registered, so advancing the clock must not re-poll.
await act(async () => {
await vi.advanceTimersByTimeAsync(10000);
});
expect(getJobLogs).toHaveBeenCalledTimes(1);
} finally {
vi.useRealTimers();
}
});
});
@@ -0,0 +1,220 @@
'use client';
import * as React from 'react';
import { X } from 'lucide-react';
import { Button } from '@/components/ui/button';
import { useDrawerFocus } from '@/hooks/useDrawerFocus';
import { useResizable } from '@/hooks/useResizable';
import { downloadJobLog, getJobLogs } from '@/lib/api';
import type { Job } from '@/lib/types';
import { cn, downloadBlob } from '@/lib/utils';
const SIDEBAR_MIN_WIDTH = 280;
const SIDEBAR_MAX_WIDTH = 750;
const POLL_INTERVAL_MS = 2000;
export default function JobDetailsSidebar({
job,
isMobile = false,
onClose,
onWidthChange,
}: {
job: Job;
isMobile?: boolean;
onClose: () => void;
onWidthChange?: (w: number) => void;
}) {
const drawerRef = useDrawerFocus<HTMLElement>(isMobile);
const [width, setWidth] = React.useState(360);
const [isDragging, setIsDragging] = React.useState(false);
const [isLoading, setIsLoading] = React.useState(false);
const [logs, setLogs] = React.useState('');
// Race-guard ref (mirrors the Svelte original): the cursor + accumulated
// text live here so in-flight polls don't read stale React state. Do NOT
// swap these for effect deps — they must update synchronously, outside
// React's render cycle. The log is one growing string (amortized O(new)
// appends), not an array re-copied and re-joined on every poll tick.
const stateRef = React.useRef<{ text: string; logAfter: number }>({
text: '',
logAfter: 0,
});
const previousJobId = React.useRef<string | null>(null);
const previousStatus = React.useRef<string | null>(null);
const consoleRef = React.useRef<HTMLPreElement | null>(null);
const { onMouseDown } = useResizable({
// Right-docked panel with the drag handle on its left edge: dragging the
// handle left must grow the panel, which is `edge: 'right'` (matches the
// Svelte original). `edge: 'left'` would invert the drag.
edge: 'right',
minWidth: SIDEBAR_MIN_WIDTH,
maxWidth: SIDEBAR_MAX_WIDTH,
getWidth: () => width,
onWidth: setWidth,
onDragChange: setIsDragging,
});
React.useEffect(() => {
onWidthChange?.(isMobile ? 0 : width);
}, [isMobile, width, onWidthChange]);
// Auto-scroll the console to the bottom whenever new lines land. Runs after
// commit so scrollHeight reflects the freshly-rendered output.
React.useEffect(() => {
const el = consoleRef.current;
if (el) el.scrollTop = el.scrollHeight;
}, [logs]);
// Reset on job switch or restart. Declared before the polling effect so the
// cursor is cleared before the next poll reads it.
React.useEffect(() => {
const wasTerminal =
previousStatus.current === 'failed' ||
previousStatus.current === 'stopped' ||
previousStatus.current === 'completed';
const isRestarting =
previousJobId.current === job.id &&
wasTerminal &&
(job.status === 'pending' || job.status === 'running');
if (previousJobId.current !== job.id || isRestarting) {
stateRef.current.text = '';
stateRef.current.logAfter = 0;
setLogs('');
}
previousJobId.current = job.id;
previousStatus.current = job.status;
}, [job.id, job.status]);
React.useEffect(() => {
const shouldPoll = job.status === 'running' || job.status === 'pending';
let pollInterval: ReturnType<typeof setInterval> | null = null;
let mounted = true;
// Effect-local lock so this job's first fetch is never blocked by a
// previous job's in-flight poll (stale writes are dropped via `mounted`).
let locked = false;
async function pollLogs() {
if (!mounted || locked) return;
locked = true;
try {
const logData = await getJobLogs(job.id, stateRef.current.logAfter);
if (mounted && logData.lines.length > 0) {
const chunk = logData.lines.join('\n');
stateRef.current.text = stateRef.current.text
? `${stateRef.current.text}\n${chunk}`
: chunk;
stateRef.current.logAfter = logData.total;
setLogs(stateRef.current.text);
}
} catch (e) {
console.error('Failed to fetch logs:', e);
} finally {
locked = false;
}
}
pollLogs();
if (shouldPoll) pollInterval = setInterval(pollLogs, POLL_INTERVAL_MS);
return () => {
mounted = false;
if (pollInterval) clearInterval(pollInterval);
};
}, [job.id, job.status]);
async function handleDownloadLog() {
if (isLoading) return;
setIsLoading(true);
try {
const blob = await downloadJobLog(job.id);
downloadBlob(blob, `job_${job.id}.log`);
} catch (err) {
console.error('Failed to download log:', err);
alert(err instanceof Error ? err.message : 'Failed to download log');
} finally {
setIsLoading(false);
}
}
return (
<aside
ref={drawerRef}
tabIndex={-1}
role="dialog"
aria-label="Job details"
aria-modal={isMobile || undefined}
className="fixed bottom-0 right-0 top-[var(--header-height)] z-50 flex max-h-[calc(100dvh-var(--header-height))] min-w-0 shrink-0 flex-col border-l border-border bg-card md:min-w-[280px]"
style={{
width: isMobile ? '100%' : width,
maxWidth: isMobile ? 'none' : SIDEBAR_MAX_WIDTH,
}}
>
<div className="flex items-center justify-between border-b border-border px-5 py-4">
<h2 className="m-0 text-base font-semibold text-foreground">
Job Details
</h2>
<div className="flex items-center gap-2">
<Button
type="button"
variant="outline"
size="sm"
onClick={handleDownloadLog}
disabled={isLoading || !job.log_file_path}
title="Download log file"
>
Download Log
</Button>
<Button
type="button"
variant="ghost"
size="icon-sm"
onClick={onClose}
title="Close"
aria-label="Close"
>
<X className="h-[18px] w-[18px]" />
</Button>
</div>
</div>
<div className="flex min-h-0 flex-1 flex-col px-5 py-4">
<div className="mb-2 flex items-center justify-between">
<span className="text-xs font-semibold uppercase tracking-wider text-muted-foreground">
Console Output
</span>
{job.status === 'running' && (
<span className="text-[0.7rem] font-medium text-emerald-600 dark:text-emerald-400">
● Live
</span>
)}
</div>
<pre
ref={consoleRef}
className="m-0 min-h-0 flex-1 overflow-auto whitespace-pre-wrap break-words rounded-lg border border-border bg-background p-3 font-mono text-xs leading-normal text-foreground"
>
{logs === '' ? (
<span className="italic text-muted-foreground">
{job.status === 'running'
? 'Waiting for logs...'
: 'No logs available'}
</span>
) : (
logs
)}
</pre>
</div>
{!isMobile && <div
role="presentation"
onMouseDown={onMouseDown}
className={cn(
'absolute bottom-0 left-0 top-0 z-[1] w-1.5 cursor-col-resize hover:bg-accent-blue/25',
isDragging && 'bg-accent-blue/25',
)}
/>}
</aside>
);
}
@@ -0,0 +1,122 @@
import { act, fireEvent, render, screen, waitFor } from '@testing-library/react';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import JobQueue from '@/components/jobs/JobQueue';
import { getJobsList } from '@/lib/api';
import type { Job, JobType } from '@/lib/types';
import { setActiveJobId } from '@/stores/activeJob';
import { triggerRefresh } from '@/stores/jobsRefresh';
import { makeJob as makeBaseJob } from '@/test/factories';
vi.mock('@/lib/api', () => ({
getJobsList: vi.fn(),
startJob: vi.fn(),
stopJob: vi.fn(),
deleteJob: vi.fn(),
downloadJobVideo: vi.fn(),
}));
const makeJob = (overrides: Partial<Job> = {}): Job =>
makeBaseJob({
model_id: 'Wan2.1-T2V',
status: 'completed',
created_at: 1_700_000_000,
num_inference_steps: 50,
num_frames: 81,
height: 480,
width: 832,
guidance_scale: 5,
seed: 42,
num_gpus: 1,
...overrides,
});
beforeEach(() => {
setActiveJobId(null);
vi.mocked(getJobsList).mockResolvedValue([]);
});
describe('JobQueue', () => {
it('shows a loading placeholder before the initial request settles', async () => {
let resolveJobs: (jobs: Job[]) => void = () => {};
vi.mocked(getJobsList).mockReturnValue(
new Promise<Job[]>((resolve) => {
resolveJobs = resolve;
}),
);
render(<JobQueue jobType="inference" />);
expect(screen.getByLabelText('Loading jobs')).toBeInTheDocument();
expect(
screen.queryByText('No inference jobs yet. Create one above.'),
).not.toBeInTheDocument();
act(() => resolveJobs([]));
expect(
await screen.findByText('No inference jobs yet. Create one above.'),
).toBeInTheDocument();
});
it('shows request failures separately from an empty queue and retries', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
vi.mocked(getJobsList).mockRejectedValueOnce(new Error('network down'));
render(<JobQueue jobType="inference" />);
expect(
await screen.findByText(/Could not load jobs from the Studio API/),
).toBeInTheDocument();
expect(
screen.queryByText('No inference jobs yet. Create one above.'),
).not.toBeInTheDocument();
vi.mocked(getJobsList).mockResolvedValueOnce([]);
fireEvent.click(screen.getByRole('button', { name: 'Try Again' }));
expect(
await screen.findByText('No inference jobs yet. Create one above.'),
).toBeInTheDocument();
});
it('shows an empty placeholder and fetches for the single job type', async () => {
render(<JobQueue jobType="inference" />);
expect(
await screen.findByText('No inference jobs yet. Create one above.'),
).toBeInTheDocument();
expect(getJobsList).toHaveBeenCalledWith('inference');
});
it('renders a JobCard for each fetched job', async () => {
vi.mocked(getJobsList).mockResolvedValue([
makeJob({ id: 'a', model_id: 'Model-A' }),
]);
render(<JobQueue jobType="distillation" />);
expect(await screen.findByText('Model-A')).toBeInTheDocument();
});
it('merges all job types when jobTypesForList is provided', async () => {
vi.mocked(getJobsList).mockImplementation((t?: JobType) =>
Promise.resolve(
t === ('lora' as JobType)
? [makeJob({ id: 'l', model_id: 'Lora-Model', created_at: 2 })]
: [makeJob({ id: 'f', model_id: 'Full-Model', created_at: 1 })],
),
);
render(
<JobQueue
jobType="finetuning"
jobTypesForList={['finetuning', 'lora'] as JobType[]}
/>,
);
expect(await screen.findByText('Lora-Model')).toBeInTheDocument();
expect(screen.getByText('Full-Model')).toBeInTheDocument();
expect(getJobsList).toHaveBeenCalledWith('finetuning');
expect(getJobsList).toHaveBeenCalledWith('lora');
});
it('refetches when a refresh is triggered', async () => {
render(<JobQueue jobType="inference" />);
await waitFor(() => expect(getJobsList).toHaveBeenCalledTimes(1));
act(() => triggerRefresh());
await waitFor(() => expect(getJobsList).toHaveBeenCalledTimes(2));
});
});
@@ -0,0 +1,195 @@
'use client';
import * as React from 'react';
import { AlertTriangle } from 'lucide-react';
import JobCard from '@/components/jobs/JobCard';
import { Button } from '@/components/ui/button';
import { useStore } from '@/hooks/useStore';
import { getJobsList } from '@/lib/api';
import type { Job, JobType } from '@/lib/types';
import {
activeJobStore,
setActiveJob,
setActiveJobId,
} from '@/stores/activeJob';
import { jobsRefreshStore } from '@/stores/jobsRefresh';
interface JobQueueProps {
jobType: JobType;
jobTypesForList?: JobType[];
}
// Jobs are flat objects of primitives, so a shallow compare detects "nothing
// changed" across poll responses (which are referentially fresh every fetch).
function jobsShallowEqual(a: Job | null, b: Job | null): boolean {
if (a === b) return true;
if (!a || !b) return false;
const aKeys = Object.keys(a) as (keyof Job)[];
return (
aKeys.length === Object.keys(b).length &&
aKeys.every((k) => a[k] === b[k])
);
}
export default function JobQueue({ jobType, jobTypesForList }: JobQueueProps) {
const [jobs, setJobs] = React.useState<Job[]>([]);
const [isInitialLoading, setIsInitialLoading] = React.useState(true);
const [error, setError] = React.useState<string | null>(null);
const { nonce } = useStore(jobsRefreshStore);
const { activeJobId } = useStore(activeJobStore);
// Stable primitive key so the memoized list below keeps its identity when an
// inline array prop (e.g. ['finetuning', 'lora']) gets a fresh reference each
// render, which would otherwise re-run the fetch effects needlessly.
const typesKey = (jobTypesForList ?? [jobType]).join(',');
const typesToFetch = React.useMemo<JobType[]>(
() => jobTypesForList ?? [jobType],
// eslint-disable-next-line react-hooks/exhaustive-deps
[typesKey],
);
// Guard the poll against slow responses: `inFlight` lets the interval skip
// a tick instead of stacking requests, and the sequence counter drops
// out-of-order responses so a delayed older payload can't overwrite a newer
// job list (direct refetches always run and supersede in-flight polls).
const fetchSeq = React.useRef(0);
const inFlight = React.useRef(false);
const fetchJobs = React.useCallback(async () => {
const seq = ++fetchSeq.current;
inFlight.current = true;
try {
let next: Job[];
if (typesToFetch.length === 1) {
next = await getJobsList(typesToFetch[0]);
} else {
const results = await Promise.all(
typesToFetch.map((t) => getJobsList(t)),
);
next = results
.flat()
.sort(
(a, b) =>
new Date(b.created_at ?? 0).getTime() -
new Date(a.created_at ?? 0).getTime(),
);
}
if (seq === fetchSeq.current) {
setJobs(next);
setError(null);
}
} catch (e) {
console.error('Failed to fetch jobs:', e);
if (seq === fetchSeq.current) {
setError(
'Could not load jobs from the Studio API. Check the server and try again.',
);
}
} finally {
if (seq === fetchSeq.current) {
inFlight.current = false;
setIsInitialLoading(false);
}
}
}, [typesKey]);
// Keep the latest fetchJobs reachable from the polling interval without
// tearing it down/recreating it on every fetchJobs identity change.
const fetchJobsRef = React.useRef(fetchJobs);
React.useEffect(() => {
fetchJobsRef.current = fetchJobs;
}, [fetchJobs]);
// Initial fetch + refetch whenever an external refresh is triggered.
React.useEffect(() => {
fetchJobs();
}, [fetchJobs, nonce]);
// Poll every second while any job is running/pending; stop otherwise.
const hasActive = jobs.some(
(j) => j.status === 'running' || j.status === 'pending',
);
React.useEffect(() => {
if (!hasActive) return;
const interval = setInterval(() => {
if (!inFlight.current) fetchJobsRef.current();
}, 1000);
return () => clearInterval(interval);
}, [hasActive]);
// Keep the active job in sync with the selected id, and void the selection
// if its job disappeared (e.g. deleted). Skip the store write when nothing
// changed — each poll returns fresh objects, and an unconditional write
// would re-render every subscriber (shell, sidebars, cards) once a second.
React.useEffect(() => {
if (!activeJobId) {
if (activeJobStore.get().activeJob) setActiveJob(null);
return;
}
const activeJob = jobs.find((j) => j.id === activeJobId) ?? null;
if (!jobsShallowEqual(activeJob, activeJobStore.get().activeJob)) {
setActiveJob(activeJob);
}
if (!activeJob) setActiveJobId(null);
}, [activeJobId, jobs]);
const multiType = typesToFetch.length > 1;
return (
<div className="mx-auto flex w-full max-w-[850px] flex-col gap-6 px-4 pb-12">
<section className="p-6">
<div aria-busy={isInitialLoading}>
{isInitialLoading ? (
<div
aria-label="Loading jobs"
className="flex flex-col gap-3 py-2"
>
{[0, 1, 2].map((item) => (
<div
key={item}
className="h-32 animate-pulse rounded-lg border border-border bg-muted/50"
/>
))}
</div>
) : error && jobs.length === 0 ? (
<div
role="alert"
className="flex flex-col items-center gap-3 py-8 text-center"
>
<AlertTriangle
className="size-6 text-destructive"
aria-hidden
/>
<p className="max-w-md text-sm text-muted-foreground">{error}</p>
<Button type="button" variant="outline" onClick={fetchJobs}>
Try Again
</Button>
</div>
) : (
<>
{error && (
<p
role="status"
className="mb-3 rounded-lg border border-amber-500/50 bg-amber-500/10 px-3 py-2 text-sm text-foreground"
>
Job updates are temporarily unavailable. Showing the most
recent results.
</p>
)}
{jobs.length === 0 ? (
<p className="py-8 text-center text-muted-foreground">
No {multiType ? 'jobs' : `${jobType} jobs`} yet. Create one above.
</p>
) : (
jobs.map((job) => (
<JobCard key={job.id} job={job} onJobUpdated={fetchJobs} />
))
)}
</>
)}
</div>
</section>
</div>
);
}
@@ -0,0 +1,125 @@
'use client';
import * as React from 'react';
import { usePathname } from 'next/navigation';
import DatasetSidebar from '@/components/datasets/DatasetSidebar';
import Header from '@/components/shell/Header';
import { HeaderActionsProvider } from '@/components/shell/HeaderActionsContext';
import PrimarySidebar from '@/components/shell/PrimarySidebar';
import JobDetailsSidebar from '@/components/jobs/JobDetailsSidebar';
import { Toaster } from '@/components/ui/sonner';
import { useMediaQuery } from '@/hooks/useMediaQuery';
import { useStore } from '@/hooks/useStore';
import {
activeDatasetStore,
setActiveDatasetId,
} from '@/stores/activeDataset';
import { activeJobStore, setActiveJobId } from '@/stores/activeJob';
import { initDefaultOptions } from '@/stores/defaultOptions';
const JOB_ROUTES = ['/inference', '/finetuning', '/distillation'];
export function AppShell({ children }: { children: React.ReactNode }) {
const pathname = usePathname();
const { activeJob } = useStore(activeJobStore);
const { activeDataset } = useStore(activeDatasetStore);
const isMobile = useMediaQuery('(max-width: 767px)');
const [primaryWidth, setPrimaryWidth] = React.useState(220);
const [secondaryWidth, setSecondaryWidth] = React.useState(0);
const [primaryOpen, setPrimaryOpen] = React.useState(false);
const jobSidebarOpen = JOB_ROUTES.includes(pathname) && activeJob != null;
const datasetSidebarOpen =
pathname === '/datasets' && activeDataset != null;
const secondaryOpen = jobSidebarOpen || datasetSidebarOpen;
// Mobile detail drawers claim aria-modal, so everything behind them must
// actually be inert — the platform enforces what the ARIA claims.
const drawerModal = isMobile && secondaryOpen;
React.useEffect(() => {
initDefaultOptions();
}, []);
React.useEffect(() => {
setPrimaryOpen(false);
}, [pathname]);
React.useEffect(() => {
function handleKeyDown(e: KeyboardEvent) {
if (e.key !== 'Escape' || document.querySelector('[data-modal]')) return;
if (primaryOpen) {
setPrimaryOpen(false);
return;
}
if (activeJobStore.get().activeJob) setActiveJobId(null);
if (activeDatasetStore.get().activeDataset) setActiveDatasetId(null);
}
document.addEventListener('keydown', handleKeyDown);
return () => document.removeEventListener('keydown', handleKeyDown);
}, [primaryOpen]);
return (
<HeaderActionsProvider>
<div
style={{ display: 'contents' }}
inert={drawerModal ? true : undefined}
>
<Header
navigationOpen={primaryOpen}
onNavigationToggle={() => setPrimaryOpen((open) => !open)}
/>
</div>
<div
className="flex overflow-hidden"
style={{
marginTop: 'var(--header-height)',
height: 'calc(100dvh - var(--header-height))',
}}
>
<PrimarySidebar
isMobile={isMobile}
mobileOpen={primaryOpen}
onMobileClose={() => setPrimaryOpen(false)}
onWidthChange={setPrimaryWidth}
/>
{primaryOpen && (
<button
type="button"
aria-label="Close navigation"
onClick={() => setPrimaryOpen(false)}
className="fixed inset-x-0 bottom-0 top-[var(--header-height)] z-40 bg-black/55 md:hidden"
/>
)}
<main
className="flex min-w-0 flex-1 flex-col overflow-auto"
inert={drawerModal ? true : undefined}
style={{
marginLeft: isMobile ? 0 : primaryWidth,
marginRight: isMobile || !secondaryOpen ? 0 : secondaryWidth,
}}
>
{children}
</main>
{jobSidebarOpen && activeJob && (
<JobDetailsSidebar
job={activeJob}
isMobile={isMobile}
onClose={() => setActiveJobId(null)}
onWidthChange={setSecondaryWidth}
/>
)}
{datasetSidebarOpen && activeDataset && (
<DatasetSidebar
dataset={activeDataset}
isMobile={isMobile}
onClose={() => setActiveDatasetId(null)}
onWidthChange={setSecondaryWidth}
/>
)}
</div>
<Toaster />
</HeaderActionsProvider>
);
}
@@ -0,0 +1,66 @@
'use client';
import { Menu, X } from 'lucide-react';
import { usePathname } from 'next/navigation';
import { useHeaderActions } from '@/components/shell/HeaderActionsContext';
import { Button } from '@/components/ui/button';
import { ThemeToggle } from '@/components/ui/theme-toggle';
const TAB_TITLES: Record<string, string> = {
'/inference': 'Jobs',
'/finetuning': 'Jobs',
'/distillation': 'Jobs',
'/datasets': 'Datasets',
'/gallery': 'Gallery',
'/gpus': 'GPUs',
'/settings': 'Settings',
};
export default function Header({
navigationOpen,
onNavigationToggle,
}: {
navigationOpen: boolean;
onNavigationToggle: () => void;
}) {
const pathname = usePathname();
const { actions } = useHeaderActions();
const title = TAB_TITLES[pathname] ?? 'FastVideo';
return (
<header className="fixed inset-x-0 top-0 z-[100] flex h-[var(--header-height)] items-center gap-2 border-b border-border bg-background/80 px-2 backdrop-blur sm:px-4 md:gap-6 md:px-6">
<Button
type="button"
variant="outline"
size="icon"
aria-label={navigationOpen ? 'Close navigation' : 'Open navigation'}
aria-controls="primary-navigation"
aria-expanded={navigationOpen}
onClick={onNavigationToggle}
className="shrink-0 md:hidden"
>
{navigationOpen ? (
<X className="size-5" aria-hidden />
) : (
<Menu className="size-5" aria-hidden />
)}
</Button>
{/* eslint-disable-next-line @next/next/no-img-element */}
<img
src="/logo.svg"
alt="FastVideo Logo"
width={100}
height={42}
className="hidden h-[42px] w-[78px] shrink-0 object-contain min-[361px]:block md:w-[100px]"
/>
<h1 className="sr-only m-0 flex-1 text-xl font-semibold tracking-tight md:not-sr-only">
{title}
</h1>
<div className="ml-auto flex min-w-0 items-center gap-2 md:gap-3">
{actions}
<ThemeToggle />
</div>
</header>
);
}
@@ -0,0 +1,52 @@
'use client';
import * as React from 'react';
interface HeaderActionsContextValue {
actions: React.ReactNode;
setActions: (node: React.ReactNode) => void;
}
const HeaderActionsContext =
React.createContext<HeaderActionsContextValue | null>(null);
export function HeaderActionsProvider({
children,
}: {
children: React.ReactNode;
}) {
const [actions, setActions] = React.useState<React.ReactNode>(null);
const value = React.useMemo(() => ({ actions, setActions }), [actions]);
return (
<HeaderActionsContext.Provider value={value}>
{children}
</HeaderActionsContext.Provider>
);
}
export function useHeaderActions(): HeaderActionsContextValue {
const ctx = React.useContext(HeaderActionsContext);
if (!ctx) {
throw new Error(
'useHeaderActions must be used within a HeaderActionsProvider',
);
}
return ctx;
}
/**
* Declaratively publish the current page's header actions: renders nothing,
* registers `children` in the header on mount and clears them on unmount.
* Pages without actions simply don't render it.
*/
export function HeaderActions({ children }: { children: React.ReactNode }) {
const { setActions } = useHeaderActions();
// Mount-only on purpose: pages pass inline JSX, which is referentially new
// every render and would re-register per render if it were a dependency.
React.useEffect(() => {
setActions(children);
return () => setActions(null);
// eslint-disable-next-line react-hooks/exhaustive-deps
}, [setActions]);
return null;
}
@@ -0,0 +1,218 @@
'use client';
import * as React from 'react';
import { X } from 'lucide-react';
import Link from 'next/link';
import { usePathname } from 'next/navigation';
import { useResizable } from '@/hooks/useResizable';
import { cn } from '@/lib/utils';
const SIDEBAR_MIN_WIDTH = 100;
const SIDEBAR_MAX_WIDTH = 300;
const SIDEBAR_COLLAPSED_WIDTH = 0;
const SIDEBAR_COLLAPSED_VISIBLE_WIDTH = 60;
const JOB_ROUTES = [
{ href: '/inference', label: 'Inference' },
{ href: '/finetuning', label: 'Finetuning' },
{ href: '/distillation', label: 'Distillation' },
] as const;
const TAB_BASE =
'block min-h-11 px-5 py-[0.65rem] text-left text-sm text-muted-foreground transition-colors hover:bg-accent/60 hover:text-foreground';
const TAB_ACTIVE = 'bg-accent-blue/10 font-medium text-accent-blue';
export default function PrimarySidebar({
isMobile,
mobileOpen,
onMobileClose,
onWidthChange,
}: {
isMobile: boolean;
mobileOpen: boolean;
onMobileClose: () => void;
onWidthChange?: (w: number) => void;
}) {
const pathname = usePathname();
const [width, setWidth] = React.useState(220);
const [isCollapsed, setIsCollapsed] = React.useState(false);
const [isDragging, setIsDragging] = React.useState(false);
const [jobsOpen, setJobsOpen] = React.useState(true);
const effectiveWidth = isCollapsed ? SIDEBAR_COLLAPSED_WIDTH : width;
const layoutWidth = isCollapsed ? SIDEBAR_COLLAPSED_VISIBLE_WIDTH : width;
const isJobsActive = JOB_ROUTES.some((r) => pathname === r.href);
React.useEffect(() => {
onWidthChange?.(isMobile ? 0 : layoutWidth);
}, [isMobile, layoutWidth, onWidthChange]);
React.useEffect(() => {
if (JOB_ROUTES.some((r) => pathname === r.href)) {
setJobsOpen(true);
}
}, [pathname]);
const { onMouseDown } = useResizable({
edge: 'left',
minWidth: SIDEBAR_MIN_WIDTH,
maxWidth: SIDEBAR_MAX_WIDTH,
getWidth: () => width,
onWidth: setWidth,
onDragChange: setIsDragging,
});
return (
<aside
id="primary-navigation"
aria-hidden={isMobile && !mobileOpen}
inert={isMobile && !mobileOpen ? true : undefined}
className={cn(
'fixed bottom-0 left-0 top-[var(--header-height)] z-50 flex max-h-[calc(100dvh-var(--header-height))] shrink-0 flex-col border-r border-border bg-card transition-transform duration-200 md:translate-x-0',
mobileOpen ? 'translate-x-0' : '-translate-x-full',
)}
style={{
width: isMobile
? 'min(18rem, calc(100vw - 3rem))'
: effectiveWidth,
}}
>
{isMobile && (
<div className="flex h-14 items-center justify-between border-b border-border px-4">
<span className="text-sm font-semibold">Navigation</span>
<button
type="button"
onClick={onMobileClose}
aria-label="Close navigation"
className="flex size-11 items-center justify-center rounded-lg text-muted-foreground hover:bg-accent hover:text-foreground"
>
<X className="size-5" aria-hidden />
</button>
</div>
)}
{!isCollapsed && (
<nav
aria-label="Primary navigation"
className="flex flex-col overflow-y-auto py-2"
>
<div className="flex flex-col">
<button
type="button"
onClick={() => setJobsOpen((v) => !v)}
aria-expanded={jobsOpen}
aria-haspopup="true"
className={cn(
TAB_BASE,
'flex w-full cursor-pointer items-center justify-between',
isJobsActive && TAB_ACTIVE,
)}
>
<span>Jobs</span>
<svg
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
strokeWidth={2}
className={cn(
'h-4 w-4 shrink-0 opacity-85 transition-transform',
jobsOpen && 'rotate-180',
)}
>
<path d="M6 9l6 6 6-6" />
</svg>
</button>
{jobsOpen && (
<div className="mb-1 ml-4 flex flex-col border-l-2 border-border pl-2">
{JOB_ROUTES.map((route) => (
<Link
key={route.href}
href={route.href}
aria-current={pathname === route.href ? 'page' : undefined}
onClick={onMobileClose}
className={cn(
TAB_BASE,
'px-4 py-2 text-[0.85rem]',
pathname === route.href && TAB_ACTIVE,
)}
>
{route.label}
</Link>
))}
</div>
)}
</div>
<Link
href="/datasets"
aria-current={pathname === '/datasets' ? 'page' : undefined}
onClick={onMobileClose}
className={cn(TAB_BASE, pathname === '/datasets' && TAB_ACTIVE)}
>
Datasets
</Link>
<Link
href="/gallery"
aria-current={pathname === '/gallery' ? 'page' : undefined}
onClick={onMobileClose}
className={cn(TAB_BASE, pathname === '/gallery' && TAB_ACTIVE)}
>
Gallery
</Link>
<Link
href="/gpus"
aria-current={pathname === '/gpus' ? 'page' : undefined}
onClick={onMobileClose}
className={cn(TAB_BASE, pathname === '/gpus' && TAB_ACTIVE)}
>
GPUs
</Link>
<Link
href="/settings"
aria-current={pathname === '/settings' ? 'page' : undefined}
onClick={onMobileClose}
className={cn(TAB_BASE, pathname === '/settings' && TAB_ACTIVE)}
>
Settings
</Link>
</nav>
)}
{!isMobile && <div
className={cn(
'absolute bottom-0 p-2',
isCollapsed ? '-right-[60px] top-0' : 'right-0',
)}
>
<button
type="button"
onClick={() => setIsCollapsed((v) => !v)}
title={isCollapsed ? 'Expand sidebar' : 'Collapse sidebar'}
className={cn(
'flex size-11 items-center justify-center rounded-lg text-muted-foreground transition-colors hover:bg-accent hover:text-foreground',
)}
>
<svg
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
strokeWidth={2}
className="h-[18px] w-[18px]"
>
<path d={isCollapsed ? 'M9 18l6-6-6-6' : 'M15 18l-6-6 6-6'} />
</svg>
</button>
</div>}
{!isMobile && !isCollapsed && (
<div
role="presentation"
onMouseDown={onMouseDown}
className={cn(
'absolute bottom-0 right-0 top-0 z-[1] w-1.5 cursor-col-resize hover:bg-accent-blue/25',
isDragging && 'bg-accent-blue/25',
)}
/>
)}
</aside>
);
}
@@ -0,0 +1,102 @@
import { act, render, screen } from '@testing-library/react';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import GpuGrid from './GpuGrid';
import { getGpus } from '@/lib/api';
import type { GpuSnapshot } from '@/lib/api';
vi.mock('@/lib/api', () => ({
getGpus: vi.fn(),
}));
const SNAPSHOT: GpuSnapshot = {
available: true,
error: null,
gpus: [
{
index: 0,
name: 'NVIDIA B200',
utilization: 62,
memory_used_mib: 61_440,
memory_total_mib: 183_359,
temperature_c: 41,
power_watts: 312.4,
power_limit_watts: 1000,
},
{
index: 1,
name: 'NVIDIA B200',
utilization: 0,
memory_used_mib: 1_024,
memory_total_mib: 183_359,
temperature_c: null,
power_watts: null,
power_limit_watts: null,
},
],
};
beforeEach(() => {
vi.mocked(getGpus).mockResolvedValue(SNAPSHOT);
});
describe('GpuGrid', () => {
it('renders a card per GPU with utilization and memory', async () => {
render(<GpuGrid />);
expect(await screen.findAllByText('NVIDIA B200')).toHaveLength(2);
expect(screen.getByText('GPU 0')).toBeInTheDocument();
expect(screen.getByText('GPU 1')).toBeInTheDocument();
expect(screen.getByText('62%')).toBeInTheDocument();
expect(screen.getByText('60.0 GiB / 179.1 GiB')).toBeInTheDocument();
// Optional sensors render only when present.
expect(screen.getByText('41°C')).toBeInTheDocument();
expect(screen.getByText('312 W / 1000 W')).toBeInTheDocument();
});
it('shows the backend-reported reason when telemetry is unavailable', async () => {
vi.mocked(getGpus).mockResolvedValue({
available: false,
gpus: [],
error: 'NVML Shared Library Not Found',
});
render(<GpuGrid />);
expect(
await screen.findByText(/GPU telemetry unavailable: NVML/),
).toBeInTheDocument();
});
it('explains when the API server is unreachable', async () => {
vi.mocked(getGpus).mockRejectedValue(new Error('network down'));
render(<GpuGrid />);
expect(
await screen.findByText(/Could not reach the API server/),
).toBeInTheDocument();
});
it('keeps the last snapshot visible and warns when a refresh fails', async () => {
vi.useFakeTimers();
try {
vi.mocked(getGpus)
.mockResolvedValueOnce(SNAPSHOT)
.mockRejectedValueOnce(new Error('network down'));
render(<GpuGrid />);
await act(async () => {
await vi.advanceTimersByTimeAsync(0);
});
expect(screen.getAllByText('NVIDIA B200')).toHaveLength(2);
await act(async () => {
await vi.advanceTimersByTimeAsync(3000);
});
expect(
screen.getByText(/values below may be stale/),
).toBeInTheDocument();
expect(screen.getAllByText('NVIDIA B200')).toHaveLength(2);
} finally {
vi.useRealTimers();
}
});
});
@@ -0,0 +1,197 @@
'use client';
import * as React from 'react';
import { AlertTriangle } from 'lucide-react';
import { Button } from '@/components/ui/button';
import { Card, CardContent } from '@/components/ui/card';
import { getGpus, type GpuInfo, type GpuSnapshot } from '@/lib/api';
import { cn } from '@/lib/utils';
const POLL_INTERVAL_MS = 3000;
function formatGib(mib: number): string {
return `${(mib / 1024).toFixed(1)} GiB`;
}
function Meter({
label,
percent,
detail,
warnAt,
}: {
label: string;
percent: number;
detail: string;
/** Turn the fill rose at this percentage (e.g. VRAM pressure). */
warnAt?: number;
}) {
const clamped = Math.max(0, Math.min(100, percent));
return (
<div className="flex flex-col gap-1">
<div className="flex items-baseline justify-between gap-2 text-xs">
<span className="text-muted-foreground">{label}</span>
<span className="font-medium tabular-nums text-foreground">
{detail}
</span>
</div>
<div
role="meter"
aria-label={label}
aria-valuenow={Math.round(clamped)}
aria-valuemin={0}
aria-valuemax={100}
className="h-1.5 overflow-hidden rounded-full bg-muted"
>
<div
className={cn(
'h-full rounded-full bg-accent-blue transition-[width] duration-500',
warnAt !== undefined && clamped >= warnAt && 'bg-rose-500',
)}
style={{ width: `${clamped}%` }}
/>
</div>
</div>
);
}
function GpuCard({ gpu }: { gpu: GpuInfo }) {
const memPercent =
gpu.memory_total_mib > 0
? (gpu.memory_used_mib / gpu.memory_total_mib) * 100
: 0;
return (
<Card>
<CardContent className="flex flex-col gap-4 p-5">
<div className="flex items-baseline justify-between gap-2">
<span className="min-w-0 truncate text-sm font-semibold">
{gpu.name}
</span>
<span className="shrink-0 text-xs font-medium uppercase tracking-wider text-muted-foreground">
GPU {gpu.index}
</span>
</div>
<Meter
label="Utilization"
percent={gpu.utilization}
detail={`${gpu.utilization}%`}
/>
<Meter
label="Memory"
percent={memPercent}
warnAt={90}
detail={`${formatGib(gpu.memory_used_mib)} / ${formatGib(gpu.memory_total_mib)}`}
/>
<div className="flex flex-wrap gap-x-5 gap-y-1 text-xs tabular-nums text-muted-foreground">
{gpu.temperature_c != null && <span>{gpu.temperature_c}°C</span>}
{gpu.power_watts != null && (
<span>
{Math.round(gpu.power_watts)} W
{gpu.power_limit_watts != null &&
` / ${Math.round(gpu.power_limit_watts)} W`}
</span>
)}
</div>
</CardContent>
</Card>
);
}
export default function GpuGrid() {
const [snapshot, setSnapshot] = React.useState<GpuSnapshot | null>(null);
const [fetchError, setFetchError] = React.useState<string | null>(null);
const [retryToken, setRetryToken] = React.useState(0);
React.useEffect(() => {
let mounted = true;
let inFlight = false;
async function poll() {
if (inFlight) return;
inFlight = true;
try {
const next = await getGpus();
if (mounted) {
setSnapshot(next);
setFetchError(null);
}
} catch {
if (mounted) {
setFetchError(
'GPU status could not be refreshed. The values below may be stale.',
);
}
} finally {
inFlight = false;
}
}
poll();
const interval = setInterval(poll, POLL_INTERVAL_MS);
return () => {
mounted = false;
clearInterval(interval);
};
}, [retryToken]);
if (fetchError && !snapshot) {
return (
<div
role="alert"
className="flex flex-col items-center gap-3 py-8 text-center"
>
<AlertTriangle className="size-6 text-destructive" aria-hidden />
<p className="text-muted-foreground">
Could not reach the API server. GPU status needs the Studio API server
running.
</p>
<Button
type="button"
variant="outline"
onClick={() => setRetryToken((token) => token + 1)}
>
Try Again
</Button>
</div>
);
}
if (!snapshot) {
return <p className="py-8 text-center text-muted-foreground">Loading…</p>;
}
if (!snapshot.available) {
return (
<p className="py-8 text-center text-muted-foreground">
GPU telemetry unavailable
{snapshot.error ? `: ${snapshot.error}` : '.'}
</p>
);
}
return (
<div className="flex flex-col gap-4">
{fetchError && (
<div
role="status"
aria-live="polite"
className="flex flex-wrap items-center gap-3 rounded-lg border border-amber-500/50 bg-amber-500/10 px-3 py-2 text-sm"
>
<AlertTriangle className="size-4 text-amber-600" aria-hidden />
<span className="min-w-0 flex-1">{fetchError}</span>
<Button
type="button"
variant="outline"
size="sm"
onClick={() => setRetryToken((token) => token + 1)}
>
Refresh Now
</Button>
</div>
)}
<div className="grid gap-4 [grid-template-columns:repeat(auto-fill,minmax(280px,1fr))]">
{snapshot.gpus.map((gpu) => (
<GpuCard key={gpu.index} gpu={gpu} />
))}
</div>
</div>
);
}
@@ -0,0 +1,48 @@
import { render, screen } from '@testing-library/react';
import { describe, expect, it } from 'vitest';
import { Button } from './button';
import { Input } from './input';
import { NativeSelect } from './native-select';
import { Slider } from './slider';
import { Switch } from './switch';
describe('shared control accessibility', () => {
it('keeps button, input, and select targets at least 44px tall', () => {
render(
<>
<Button size="sm">Small action</Button>
<Input aria-label="Text value" />
<NativeSelect aria-label="Choice" defaultValue="one">
<option value="one">One</option>
</NativeSelect>
</>,
);
expect(screen.getByRole('button', { name: 'Small action' })).toHaveClass(
'h-11',
);
expect(screen.getByRole('textbox', { name: 'Text value' })).toHaveClass(
'h-11',
);
expect(screen.getByRole('combobox', { name: 'Choice' })).toHaveClass(
'h-11',
);
});
it('uses 44px switch and slider interaction surfaces', () => {
render(
<>
<Switch aria-label="Enabled" />
<Slider aria-label="Amount" defaultValue={[50]} />
</>,
);
expect(screen.getByRole('switch', { name: 'Enabled' })).toHaveClass(
'size-11',
);
expect(screen.getByRole('slider', { name: 'Amount' })).toHaveClass(
'size-11',
);
});
});

Some files were not shown because too many files have changed in this diff Show More