Compare commits

..
Author SHA1 Message Date
Peiyuan Zhang 67f5b53595 [fix]: keep v2 examples runnable 2026-07-04 20:22:18 +00:00
Peiyuan Zhang aa6bf5079b [docs]: scope v2 docs to inference runtime 2026-07-04 19:33:05 +00:00
Peiyuan Zhang 4ca56d8951 [fix]: route v2 serving tasks by capabilities 2026-07-04 19:32:46 +00:00
Peiyuan Zhang 1e19650eb3 [misc]: simplify v2 program execution model 2026-07-04 19:32:30 +00:00
Peiyuan Zhang de354f804c [feat]: add FastWan QAD FP8 card 2026-07-04 19:32:18 +00:00
Peiyuan Zhang 6feee4e4fa [feat]: add FP8 vendor support for Wan weights 2026-07-04 19:32:05 +00:00
Peiyuan Zhang 82e00615e5 [feat]: wire v2 VideoGenerator to real torch backend 2026-07-04 19:31:55 +00:00
Peiyuan Zhang 30cfbfc4f6 [refactor]: remove training semantics from v2 inference contracts 2026-07-04 19:31:35 +00:00
Peiyuan Zhang a4d8978c37 [refactor]: remove v2 training package 2026-07-04 19:31:12 +00:00
Will Lin 803a6c99ae [refactor] v2: drop unwired ConditioningInjector policy
ConditioningInjector (ABC) + PassthroughConditioning were defined and exported
but never instantiated or called — zero call sites anywhere. Unlike the policies
the §5 thesis actually names (CFG / flow-shift / precision / expert-routing),
which every recipe wires (cfg=ClassicCFG(), expert=NoRouting(), ...), conditioning
is injected INLINE by the loops: `st.cond["prompt_embeds"] = ctx.slots.get(...)`
(wan21/loop.py, ltx2/loop.py); the qwen_omni cascade conditions via loop wiring.
PassthroughConditioning even described a dataflow (state.scratch["cond"]) that is
not how conditioning actually flows (loops read ctx.slots). So it was a designed-
but-bypassed seam, not forward-design — and conditioning isn't in the README §5
policy list. Removed the ABC + impl + exports; kept the live `cond` field (the
loops fill it directly) and fixed its comment.

Tests: 143 passed, 2 skipped.
2026-07-04 17:15:35 +00:00
Will Lin 8a637c3215 [refactor] v2: remove unused DataRef provenance spec
DataRef (dataset_id/revision/description — "what a recipe trained on") was only
the type of RecipeSpec.data_contract, which no recipe authored and no code read
(0/0). It is absent from the README RecipeSpec contract (§2.1: method, parents,
assumes_loop, assumes_precision, consistency_required) and not among the §18
"wire the inert metadata" roadmap items — i.e. unwired governance metadata, not
deliberate forward-design. RecipeSpec now matches §2.1 exactly.

Kept CheckpointManifest: unlike DataRef it IS wired (ModelCard.checkpoint), the
declarative "explicit components + key maps, no name-detector guessing" load
contract — declared intent, not dead.

Tests: 143 passed, 2 skipped.
2026-07-04 17:15:35 +00:00
Will Lin 51bae33c7b [refactor] v2: drop unused typed-schema slots from Component/LoopSpec
state_schema / step_schema / result_schema (LoopSpec) and config_schema /
io_schema (ComponentSpec), plus valid_parallel_plans and parallel_constraints,
were authored by no recipe and read by no executor (audited: 0 reads / 0 sets
across native + tests, no dynamic dataclasses.fields/asdict/__dict__ access, no
consumer in scripts/_vendor/examples). They duplicated mechanisms that already
exist: the concrete LoopState/WorkPlan/StepResult classes the driver uses
directly, and the card-level ParallelismContract. Removing them slims the core
spec surface with zero behavior change.

Thesis untouched: cards still own components/loops/recipe/parity; kept the
parity, precision/placement, behavior-capture (behavior_schema), wired
extension_schema, and roadmap required_for/optional_for/resident_for fields.
Also drop the stale README step_cost_model mention (that field went with the
cost mechanism in 352c1b28).

Tests: 143 passed, 2 skipped (toy backend).
2026-07-04 17:15:34 +00:00
Will Lin 54b7fe3a55 [refactor] v2: drop dead toy components + unused Karras schedule
ToyLoRA / ToyControlNet / ToyTargetModel / ToyDraftModel / ToyRewardModel (+ the
_spec_target_next helper) in the toy backend, and build_karras_sigmas in the
sampler, were defined but referenced nowhere — no recipe, card, loop, test, or
example used them. They were toy stand-ins for capabilities (adapter plane,
speculative decode, served reward) and an EDM/Karras noise schedule that were
written ahead of being wired. Remove them (-134 LOC); a toy can be re-added when
the capability is actually wired. No thesis impact, no behavior change.

Suite green: 143 passed, 2 skipped (toy backend); import v2 stays torch-free.
2026-07-04 17:15:34 +00:00
Will Lin 1fe50c0092 [refactor] v2: group flat top-level into planes; isolate vendored under _vendor/
The v2 top level had grown to ~27 dirs + 13 loose files — half of them
single-concept abstraction shells (memory/ 86 LOC, transport/ 222, parity/ 178,
extend/ 226) and half vendored fastvideo code sitting as peers to the actual v2
design. Regroup to mirror README section 3 "Planes & dependency order":

  core/     enums+types, card, loop, program, parity, request, parallel
            (the model-native contracts; no kernels)
  runtime/  + folded-in substrate: cache, memory, transport, extend
            (the import graph shows only runtime consumes them; compile/cudagraph
            already lived here)
  serving/  + deploy/ (products / fleet)
  _vendor/  all copied fastvideo: models, layers, attention, configs, distributed,
            platforms, api, hooks, logging_utils, third_party + fastvideo_args/
            utils/logger/envs/forward_context — internal layout unchanged, still
            mirrors upstream for diffing

28 dirs + 13 root files -> 8 dirs + 6 root files. Pure mechanical move: 947
absolute-import paths rewritten (v2.X -> v2.{core,runtime,serving,_vendor}.X),
boundary-anchored so platform/ (native dispatch) and platforms/ (vendored CUDA
detect) no longer collide and neither does hooks/ vs extend/. Deleted the empty
v2/loader/. README section 3 + 16 updated; stale test count corrected
(34 files/216 tests -> 22 files/143 tests).

Validated: `import v2` stays torch-free; `v2/run_tests.py` and `pytest v2/tests/`
-> 143 passed, 2 skipped on the numpy toy backend. (On a GPU box force the toy
backend with CUDA_VISIBLE_DEVICES="" or detect() picks cuda.) No external importer
changed — examples use the re-exported `from v2 import VideoGenerator`.
2026-07-04 17:15:34 +00:00
SolitaryThinker 290795daf8 [refactor] v2: remove interleave; pooled run-to-completion serving (P2)
Second step of the runtime simplification (after cost removal). Drop the
coordinated step-interleave scheduler + the interleave-parity gate; serving is
now pooled run-to-completion.

- Engine: remove run_interleaved + the WorkUnit/BatchScheduler imports; run /
  run_serial drive each request to completion (tick/run_to_completion kept as the
  per-request stepper).
- scheduler.py: remove BatchScheduler + WorkUnit + the batches metric; the
  AdmissionController is now a pure refundable memory/OOM guard.
- AsyncEngine: bound concurrency with a serving pool (asyncio.Semaphore,
  max_concurrent) — each request waits for a slot, then runs to completion.
- parity: remove assert_interleave_parity (the run_serial==run_interleaved gate);
  rename interleave_gate.py -> compare.py (compare_outputs stays — bit-parity
  between execution paths, e.g. disaggregated==inline).
- card specs: drop ParitySpec.interleave_required + LoopSpec.allows_interleaving;
  stripped interleave_required from all cards.
- tests: delete the interleave-gate/parity tests; refocus the ones that exercised
  real behavior (residual-skip, compare_outputs symmetric-empty).
- README: removed the design-doc references at the top; simplified the thesis /
  scheduler (§6) / parity (§9) / package-layout / comparison sections to pooled
  run-to-completion (no cost model, no interleave gate).

CPU mini: 143 passed / 2 skipped. Native omni port (P3) still to come. (pyproject
kernel hack excluded.)
2026-07-04 17:15:34 +00:00
SolitaryThinker 0047d54a1b [refactor] v2: remove the cost mechanism (P1 of runtime simplification)
First step toward lean pooled run-to-completion serving: rip out the GPU-time
cost/budget machinery entirely (it priced nothing useful for the target design).

Removed: CostModel + LoopSpec.step_cost_model; ResourceRequest.compute_seconds;
StepResult.actual_seconds; AdmissionController's compute budget + SchedulerMetrics
.gpu_seconds (the memory/OOM reservation guard stays); the Profiler observer
(cost calibration); per-step timing in RuntimeLoopContext; cost-based fleet/Dynamo
routing (now a coarse step-count load proxy); DeploymentCard.cost_model. Stripped
cost from all 9 recipe cards + their loops.

Tests: dropped the 3 cost-specific tests (cost routing, cost_model aliasing,
compute-budget gate); refocused 2 (loop cache validation, NaNWatch-clean).

CPU mini: 151 passed / 2 skipped. Interleave removal + pooled serving (P2) and the
native omni port (P3) follow. (pyproject kernel hack excluded as always.)
2026-07-04 17:15:34 +00:00
SolitaryThinker b2d55a7ba0 [refactor] v2: full vendor cutover — copy fastvideo modeling + layer code into v2 (zero fastvideo imports)
Replace the re-export stubs with real vendored copies of the fastvideo modeling
+ layer + supporting infra, for the kept diffusion models (wan21, wan_causal,
ltx2, flux2, matrixgame2). v2 now imports ZERO `fastvideo.*` — it is
self-contained. (bagel/qwen_omni load from vllm_omni, an external pkg, not
fastvideo; cosmos3's load_id was already dangling — both out of scope here.)

Vendored (cp + `sed fastvideo. -> v2.`):
- models/  the 5 models' nn.Module dits/vaes/encoders/audio/upsamplers + the
           component loader/ + the lazy class registry (other families' rows are
           dormant/lazy — only the 5 resolve).
- layers/ attention/ platforms/ distributed/ configs/ logging_utils/ hooks/
  third_party/pynvml + top-level forward_context/fastvideo_args/envs/logger/
  utils/version — copied verbatim (layers et al. 'as is').
- api/  slimmed to schema + results (the VideoGenerator's config dataclasses);
  the fastvideo parser/presets/overrides (which pull the pipeline runtime) are
  intentionally NOT vendored — v2 has its own runtime/loop.

Decoupling surgery (cut the loader's coupling to the fastvideo runtime):
- configs/pipeline_registry.py (vendored from fastvideo/registry.py, renamed to
  avoid colliding with v2/registry.py): dropped the _register_presets() auto-call
  and matrixgame3 (removed model); config-class resolution preserved.
- configs/pipelines/__init__.py: dropped the registry back-edge (fixes an import
  cycle) — base.py imports the registry lazily where used.
- torch_backend: load_component now uses v2.models.loader.

The only remaining external 'fastvideo*' refs are `fastvideo_kernel` (the
separate optional CUDA-kernel pkg for sparse/MoBA attention) — guarded; the dense
TORCH_SDPA path v2 uses never imports it.

Vendored subtrees added to the pre-commit exclude (faithful copies, mirroring the
existing fastvideo/models exclusion — not re-linted, to stay re-syncable).

Verified: grep finds zero fastvideo-package imports in v2/; Wan2.1 T2V on H100 is
BIT-IDENTICAL to the fastvideo-backed path (same .npy SHA256, byte-for-byte); CPU
mini 156 (154 passed + 2 env-skipped on x86, torch present). Backup: branch v2_backup.
2026-07-04 17:15:34 +00:00
SolitaryThinker 24cfe281b3 [refactor] v2: prune recipes to 8 models (+ omni shared infra)
Keep: bagel, cosmos3, flux2, matrixgame2, ltx2, qwen_omni, wan21, wan_causal
(plus the shared omni/ package that bagel/cosmos3/qwen_omni depend on). Remove
the other 26 recipe packages.

- Delete 26 recipe dirs (adapters, adaptive, cosmos2, cosmos25, fastwan, gen3c,
  hunyuangamecraft, hunyuan_video(15), hyworld, image_video, kandinsky5,
  lingbotworld, longcat, lucy_edit, matrixgame3, multi_expert, reward, sd35,
  sfwan22, speculative, stable_audio, tiled, turbowan, unified, wan_fun_control).
- recipes/__init__.py: keep-closure builders + build_default_engine /
  build_omni_engine (dropped the workflow/tiled/unified/image_video engine helpers).
- registry.py: _BUCKET_C pruned to flux2 + matrixgame2; removed the cosmos2
  ModelEntry + the CosmosTransformer3DModel arch branch + the TurboWan-14B entry.
- Delete 11 tests for removed recipes/features; patch test_bucket_c_ports to drop
  the cosmos2 reference (it now auto-derives from the pruned _BUCKET_C).
- README: correct the recipes/ roster to the kept families.

Backup of the full pre-prune tree is on branch v2_backup. CPU mini 156 passed / 0
failed (the prior 5 torch-absent bucket_c failures are gone with sd35/stable_audio);
no dangling references to any removed recipe; all kept model ids still resolve.
2026-07-04 17:15:34 +00:00
SolitaryThinker dc086b207f [perf] v2: on-device denoise loop — kill the per-step numpy<->torch round-trip (Wan2.1)
The torch adapter boundary marshalled the latent host<->device on EVERY denoise
step: _t uploaded the latent (and re-uploaded the text embeds) and _n downloaded
the velocity with a forced CUDA sync — 2*N PCIe copies + N syncs per generation,
buying nothing, since the latent could stay resident on the GPU the whole loop.

Root cause was a numpy loop surface. But the loop MATH is already array-agnostic
(CFG combine + flow-match Euler are pure arithmetic; the solver kernel already
passes torch through). So introduce a per-platform array namespace (v2/platform/
array_ns): numpy on CPU (torch-free — the parity mini is unchanged), torch-on-
device on cuda. The latent is seeded with numpy and uploaded ONCE; it then stays
resident through forward -> CFG combine -> solver -> next step; a single host
marshal happens at the request/output boundary (engine._to_artifact).

Opt-in per recipe via ModelCard.device_io (set on the Wan cards). When set on a
GPU box, build_component flips the components' TorchComponent.device_io so _out
keeps tensors on-device (in fp32, matching the old _to_numpy cast so the combine
dtype is unchanged). Un-migrated families and the CPU toy keep numpy in/out.
Also: PrecisionPolicy.cast is array-preserving; _t accepts resident tensors.

Verified BIT-IDENTICAL on real Wan2.1-1.3B / H100: the on-device latent equals
the pre-change numpy-path latent exactly (max_abs_diff 0.0, np.array_equal True).
CPU mini holds 237 passed / 5 pre-existing; pre-commit clean.

Other WanDenoiseLoop families can flip device_io next (per-family GPU re-verify);
non-Wan loops migrate to the xp namespace later.
2026-07-04 17:15:34 +00:00
SolitaryThinker b9db151658 [refactor] v2: co-locate per-model torch adapters into their recipe packages
platform/backends/ had become a flat dump of 15 per-model torch_<model>.py
adapters next to the genuinely-shared infra. Each adapter is referenced from
its card by a plain 'module:Class' string loaded via importlib, so there was
no real coupling forcing it into platform/ — the Cosmos/Flux/etc adapter
belongs WITH its recipe (card/loop/program).

Move each torch_<model>.py -> v2/recipes/<model>/adapter.py and flip the card
strings to v2.recipes.<model>.adapter:<Class>. backends/ now holds only the
shared substrate (torch_backend base, torch_cuda registration, torch_kernels,
toy/cpu/accel). Each recipe is now a self-contained package.

Cross-refs updated: gen3c/adapter imports CosmosT5Encoder from cosmos2/adapter;
sd35/program + stable_audio/card import from their own package. Recipes still
import torch-free (adapters pulled only via the string on a GPU box) — CPU mini
holds 237 passed / 5 pre-existing (bucket_c torch-absent). pre-commit clean.
2026-07-04 17:15:34 +00:00
SolitaryThinker 9124963238 [misc] v2: full pre-commit clean (ruff UP038/SIM/UP031 + mypy annotations)
Sweep all of v2/ through pre-commit (was previously only run on changed
files). Fixes surfaced across untouched modules:

- ruff: isinstance-tuple -> X | Y (UP038), try/except/pass ->
  contextlib.suppress (SIM105), negated-return (SIM103), %-format ->
  f-string (UP031).
- mypy: add annotations for no-untyped-call + var-annotated across recipes,
  training methods, torch/toy backends, and serving.
- yapf reflow of the SF-Wan KV-cache call sites (semantics unchanged).

yapf/ruff/codespell/mypy all pass; v2 tests 237 passed / 5 pre-existing
(bucket_c torch-absent on CPU venv).
2026-07-04 17:15:34 +00:00
SolitaryThinker f90e8f3e76 [bugfix] v2: SF-Wan cross-chunk KV cache — condition each chunk on prior clean chunks
The causal adapter never passed a kv_cache, so CausalWanTransformer3DModel.forward routed to
_forward_train (no cross-chunk KV) on every chunk instead of _forward_inference (the CausVid
Alg-2 KV-cache path). Each chunk denoised blind to the previous ones; the loop's cross-chunk
"context" was a toy mean(prior_latents) the adapter ignored. Result: hard discontinuities at
every chunk boundary (frame-to-frame absdiff spikes 43-55 every ~12 frames).

Fix (cuda path only; toy/CPU path and the 237-test suite untouched):
- WanDiT.alloc_causal_caches(): allocate the persistent per-block KV + cross-attn caches sized
  from the model config (mirrors CausalDenoisingStage._initialize_kv_cache).
- WanDiT.__call__: thread kv_cache/crossattn_cache/current_start/cache_start/start_frame/
  frame_seqlen so the model runs _forward_inference.
- wan_causal/loop.py: own the caches in LoopState (per-request -> interleave-safe); pass
  current_start = chunk_idx*chunk_size*frame_seqlen per chunk; do the clean-KV write
  (timestep ~0) after each chunk so the next attends to it.

Verified on H100: frame-to-frame absdiff mean 12.8->4.4, max 55.4->8.3; chunk-boundary spikes
eliminated; coherent across all 7 chunks. CPU causal toy tests unchanged (24 passed).
2026-07-04 17:15:34 +00:00
SolitaryThinker 4ff33a28b7 [misc] v2: simplify docstrings/comments + drop deleted-design-doc citations
Sweep all 296 v2 modules: simplify verbose docstrings/comments and remove 376 dangling
"(design_vN §X)" citations to the now-deleted design docs (v2/README.md is the source of
truth). Comment/docstring-only — AST-verified code-identical; the CPU suite holds at 237
passed / 5 pre-existing. Also applies yapf + ruff --fix auto-fixes (import ordering,
forward-ref annotation de-quoting under `from __future__ import annotations`; behavior-
neutral, suite-confirmed) and adds the legitimate domain terms mot/clen/te to the codespell
ignore-list. Remaining ruff (24) + mypy (68 no-untyped-call) findings are pre-existing v2
debt, untouched here.
2026-07-04 17:15:34 +00:00
SolitaryThinker 940f94f435 [docs] v2: make v2/README.md the design source of truth + one-page philosophy + M* roadmap
Unify the four design docs (design.md, designv2.md, design_v3.md, designv4.md) into a single
authoritative v2/README.md: the (recipe, runtime) thesis, driven loops, planes, one-WorkUnit
scheduler, the parity ladder + interleave gate, training-on-shared-loops, the weight-sharing
topology catalog, the current GPU status (20+ models + the BAGEL/Qwen-Omni/Cosmos3 trio
verified), and a prioritized roadmap. Recast design_summary.md as a one-page design philosophy
pointing to it. Add .agents/exploration/mstar-v2-roadmap.md (the adversarially-verified M*
Walk-Graph gap analysis driving the roadmap). Delete the four superseded design docs.
2026-07-04 17:15:34 +00:00
SolitaryThinker ac9dcb63ea [bugfix] v2: MatrixGame2/3 causal-loop progress counter + MG3 patch alignment
MatrixGame2 (causal DMD loop): bump st.step_idx on every executed work unit (each
DMD step and each clean-context pass). The loop drives its own control flow off
block_idx/dmd_idx/phase, but the runtime's no-progress watchdog keys on
st.step_idx, so a multi-block causal rollout was seen as stalled. Mirrors what
every other recipe loop does.

MatrixGame3 (5B WanModel): patch-align the latent H/W (patch_size (1,2,2)) before
denoise. The model folds (H/2, W/2) tokens, so an odd latent dim made the
unpatchified velocity come back one row/col short of the noise latent. Crop to
(latent // patch) * patch, faithful to MatrixGame3DenoisingStage.

v2 mini: 240 passed.
2026-07-04 17:15:34 +00:00
SolitaryThinker 861f87e843 [bugfix] v2: correct SF-Wan + LTX2 2-stage SR sampling defaults (GPU frame-verified)
Two distilled few-step video models rendered incorrectly on GPU; root-caused via
dense frame sampling (contact sheets) and fixed in the recipe cards.

SF-Wan2.1 (self-forcing causal, wan_causal card): was oversaturated/overcooked.
The distilled student is CFG-FREE (guidance 1.0, single forward/step) and denoises
with the 4-step DMD schedule [1000,750,500,250] (warped by FlowShiftPolicy(5.0)),
at a native causal block of 3 latent frames. Defaults were ClassicCFG@6.0 + 2 steps
+ block 2 -> overcooked AND under-denoised. Fixed: num_chunks=7, chunk_size=3,
steps_per_chunk=4; SamplingDefaults num_steps=4, guidance_scale=1.0 (7x3=21 latent
-> 81 frames). Renders a clean raccoon-in-sunflowers across all 81 frames.

LTX2-Distilled 2-stage SR (ltx2 card + LTX2VAE): was temporally blocky. Root cause
was an OOM-forced 57-frame reduction (only 8 latent temporal frames); the model is
designed for 121 (16 latent frames). Enable VAE tiling in LTX2VAE so the 121-frame
full-res decode fits the 80 GiB GPU; keep base cfg_scale=3.0 (drives brightness;
cfg=1 washed out) with stg_scale=0.0 (v2's drop-text perturbation is not real
skip-layer STG). Renders the on-prompt backyard shot, bright + temporally coherent.

torch_backend.py: enable LTX2VAE tiling; clear pre-existing mypy no-untyped-call /
yapf debt on the file (surfaced once per-file linting bypassed the duplicate-module
flakiness) by annotating the helper/constructor/maker signatures.

Tests: update the 3 affected CPU defaults/chunk-count tests. v2 mini: 240 passed.
2026-07-04 17:15:34 +00:00
SolitaryThinker 7ef2d9083e [docs] v2: GPU bring-up results — 20 models verified on H100
V2_PORTING_STATUS.md now records the GPU bring-up outcome: 20 models generate
real video/audio on H100 (the 7 prior + 13 newly-ported), with the remaining
split into fastvideo/env-blocked (SLA/VSA kernels, transformers incompat, fastvideo
registry/flash_attn gaps) and HF-access-blocked (gated cosmos2/flux2/sd35) — none
a v2 recipe bug.
2026-07-04 17:15:34 +00:00
SolitaryThinker 0ac5367b54 [feat] v2 GPU bring-up: 5 more models verified on H100 (huge MoE/world + DMD)
Second GPU pass (distilled + huge dense, 2-wide). 5 verified end-to-end with real
weights; CPU toy path kept green (240 passed, 2 skipped).

VERIFIED:
  * matrixgame3   — mp4 (9,256,256,3), zero fixes. 6.47B, standard Wan attn (NOT
                    sparse-attn-blocked, like its mg2 sibling); degenerate single-clip.
  * fastwan       — FastWan2.2-TI2V-5B-FullAttn DMD 3-step, mp4 (17,256,256,3), zero
                    fixes. The FULL-ATTENTION variant has no VSA params -> the generic
                    Wan loader maps it cleanly (reuses WanDiT via load_id, no adapter).
  * longcat       — LongCat-Video-T2V 13.58B, mp4 (17,256,256,3), zero fixes, CPU
                    offload (~40GB peak).
  * sfwan22       — Self-Forcing Wan2.2-A14B causal+MoE (2x14B), mp4 (29,288,288,3),
                    CPU expert offload (~80GB peak).
  * lingbotworld  — Wan2.2-class 2x14B camera world model, mp4 (9,256,256,3), offload
                    (~98GB peak transient).

BLOCKED: turbowan-i2v-a14b — SLA sparse-attn params (attn1.attn_impl.proj_l on all
40 layers of both experts) cannot load into the dense Wan build + needs the
fastvideo-kernel SLA Triton kernels (no nvcc here). Confirmed via the safetensors
header (no 60GB download). Same SLA family as turbowan-1.3b.

Fixes (own-port only): sfwan22/loop.py, lingbotworld/{card,program}.py +
torch_lingbotworld.py. pre-commit clean per-package.
2026-07-04 17:15:34 +00:00
SolitaryThinker 56f84018df [bugfix] v2 VideoGenerator: modality-aware result path (audio/image, not only video)
VideoGenerator._result hardcoded out.artifacts['video'].frames, so an audio-only
(Stable Audio) or image-only (SD3.5 / FLUX.2 T2I) generation crashed with
KeyError 'video' even though the engine had correctly produced the AudioArtifact /
image TensorArtifact (surfaced during stable_audio GPU bring-up). _result now
guards on the artifact present: video -> mp4 (unchanged), else image -> png
([C,H,W]/[B,..] normalized to [H,W,C]), plus the existing audio -> sibling .wav;
image_path recorded in result.extra. The video path is byte-for-byte unchanged.

v2 mini 240 passed, 2 skipped; pre-commit clean.
2026-07-04 17:15:34 +00:00
SolitaryThinker d04127fe6f [feat] v2 GPU bring-up: 8 ported models verified end-to-end on H100 (+CPU-safe fixes)
Ran each dense public port through the real VideoGenerator on H100 (2-wide across
both GPUs). 8 produce real finite output end-to-end; per-model adapter/loop fixes
landed in each port's OWN files (no shared/fastvideo edits). CPU toy path kept
green (240 passed, 2 skipped) — GPU-only conditioning gated to the cuda backend.

VERIFIED (real GPU output):
  * stable_audio    — stereo audio (2, 441000) @44.1kHz. Fixes: dedicated
                      'conditioner' component kind (SA owns its T5, empty
                      text_encoder_configs) + ConditionerLoader from conditioner/
                      + VDenoiser c_noise = atan(sigma)/(pi/2).
  * matrixgame2     — mp4 (9,256,256,3). Loads CLEAN (not sparse-attn-blocked).
                      Fixes: 20ch cond_concat (4ch mask + 16ch img), mandatory i2v
                      (synth blank first frame), pre-sized kv_cache/crossattn_cache
                      (SDPA inference path, avoids flex_attention compile), bf16
                      autocast, per-request reset_caches.
  * gen3c           — video (9,256,256,3), zero fixes (worked first try).
  * wan_fun_control — video (9,256,256,3).
  * lucy_edit       — video (17,256,256,3).
  * hunyuangamecraft— video (9,256,256,3).
  * hunyuan_video   — video (3,9,256,256), dual LLaMA+CLIP + Hunyuan VAE.
  * hunyuan_video15 — video (9,256,256,3), dual Qwen+ByT5 (gated to cuda; CPU passes
                      single embed).

BLOCKED (not v2 recipe bugs — load/run reached, then a fastvideo/env wall):
  * cosmos25 — DiT + VAE ran finite on GPU; the Qwen2.5-VL Reason1 encoder hits a
               transformers 5.12.1 incompat in fastvideo shared code
               (Qwen2_5_VLConfig.pad_token_id).
  * kandinsky5 — fastvideo registry.py registers it with a bare PipelineConfig (no
                 Kandinsky5 config) -> load fails fastvideo-side. (latent z=16 +
                 visual_cond adapter corrected; toy decoupled to its own channels.)
  * hyworld — fastvideo's hyworld DiT hardcodes flash_attn (no SDPA fallback);
              flash_attn kernel not built here.

CPU-safety fixes (mine): kandinsky5 toy ToyDiT/ToyVAE use the toy LATENT_CHANNELS
(not the real z=16); hunyuan_video15 dual-encoder packing gated to cuda.
Sparse-attn distilled (turbowan/SLA, fastwan/VSA) + gated (cosmos2/flux2/sd35) +
huge (>80GB) handled separately. pre-commit clean per-package.
2026-07-04 17:15:34 +00:00
SolitaryThinker ae0cf3e2d1 [docs] v2: porting status — ALL fastvideo models ported (63/64 by-id; VSA env-blocked)
Rewrites V2_PORTING_STATUS.md to reflect completion: the scope is now ALL
fastvideo models (not Wan+LTX-2 only). Documents the self-contained recipe-package
porting mechanism (ComponentSpec.adapter), the 15 net-new architectures + 5
Wan-family variants newly ported (CPU-verified end-to-end; GPU=BRINGUP), the 7
GPU-verified models, and the single env-blocked id (VSA-14B, needs nvcc).
2026-07-04 17:15:34 +00:00
SolitaryThinker f4af3cf886 [feat] v2 registry: LTX-2/2.3 repo aliases -> by-id resolution 63/64
Adds explicit ModelEntry aliases for the LTX-2 (FastVideo/LTX2-Diffusers,
LTX2-base, Lightricks/LTX-2 -> single-stage base) and LTX-2.3 (LTX2.3-Diffusers,
LTX2.3-Distilled-Diffusers, LTX2.3-base, Lightricks/LTX-2.3, lightricks/ltx-2.3
-> the distilled joint-A/V card) naming variants of already-ported LTX
checkpoints (the arch fallback also resolves LTX2Transformer3DModel from a root).

v2 now resolves 63/64 fastvideo registry ids by exact id; the only remaining id,
FastVideo/Wan2.1-VSA-T2V-14B-720P-Diffusers, is ENV-BLOCKED (VSA Sparse-Linear
Attention kernels require nvcc, not built in this bring-up; it arch-resolves to
the base Wan card but needs the VSA kernel build to run faithfully).

v2 mini 240 passed, 2 skipped; pre-commit clean.
2026-07-04 17:15:34 +00:00
SolitaryThinker 39580ad73a [feat] v2: port the residual Wan-family variants (rCM/DMD/v2v/control/causal-MoE)
Closes the bucket-B sampler/conditioning gap — each reuses the Wan/Causal ARCH
(no new torch adapter) with a new in-package loop/sampler/conditioning, declared
in _BUCKET_C as explicit-HF-id-only (transformer_cls="" so the generic Wan/Causal
arch fallback is NOT hijacked — only the exact id distinguishes the capability
variant from a base Wan of the same class).

  * turbowan      — TurboWan rCM (Reparameterized Consistency Model) few-step: a
                    faithful in-package RCMScheduler port (TrigFlow->RectifiedFlow
                    schedule + stochastic consistency SDE step), 1.3B/14B T2V +
                    TurboWan2.2-I2V-A14B (MoE i2v, boundary 0.9 in raw-sigma space)
  * lucy_edit     — Lucy-Edit v2v editor: a video_vae_encode node (the input video
                    -> 48ch cond latent) threaded via the shared i2v_cond hook ->
                    96ch Lucy DiT input (faithful to denoising.py is_lucy_edit)
  * wan_fun_control — Wan2.1-Fun-Control: control-video conditioning (reuses the
                    i2v [mask|cond] concat pattern)
  * sfwan22       — Self-Forcing Wan2.2-A14B: causal chunk_rollout + Wan2.2 MoE
                    boundary routing, i2v (boundary 0.9) + t2v (boundary 0.875)
  * fastwan       — FastWan DMD 3-step: TI2V-5B-FullAttn loadable; the VSA-trained
                    variants + non-strict to_gate_compress load are BRINGUP

All 12 residual ids resolve+build; base Wan/Causal resolution unchanged (arch
fallback not hijacked). The _BUCKET_C regression test auto-extended -> v2 mini
240 passed, 2 skipped; pre-commit clean (per-package + registry).
2026-07-04 17:15:34 +00:00
SolitaryThinker 521c2845e6 [test] v2: end-to-end CPU regression guard for the bucket-C ports
Data-driven from registry._BUCKET_C (+ cosmos2): each net-new ported arch
resolves through the registry (exact id + arch fallback) AND runs end-to-end on
the CPU toy backend via the public Engine path (resolve -> build card+program ->
load_card -> Engine.run), emitting exactly one modality-correct artifact
(video / image / audio) + latents. Auto-covers future _BUCKET_C rows.

21 tests pass; full v2 mini 232 passed, 2 skipped.
2026-07-04 17:15:34 +00:00
SolitaryThinker 0467edbd07 [feat] v2: port the 14 remaining bucket-C archs as self-contained recipe packages
Completes the bucket-C porting backlog. Each arch is a self-contained recipe
package (card-declared torch adapter via ComponentSpec.adapter + a new/forked
loop + program, NO edit to the shared torch_backend dispatch), following the
cosmos2 reference pattern. One _BUCKET_C table in v2/registry.py drives both the
explicit HF-id registry (PRIMARY) and the select_by_architecture fallback.

Ported (CPU-verified: import + card/program build + registry resolve + denoise
loop runs end-to-end on the CPU toy backend; GPU load/run is BRINGUP):
  * cosmos25       — Cosmos-Predict2.5 (flow-match, per-frame plain-sigma timestep;
                     reuse FLOW_MATCH_STEP; Reason1/Qwen2.5-VL encoder adapter)
  * hunyuan_video  — HunyuanVideo (reuses WanDenoiseLoop; dual LLaMA+CLIP encoders;
                     Hunyuan VAE scaling_factor) + FastHunyuan variant
  * hunyuan_video15 — HunyuanVideo 1.5 (480p/720p cards)
  * longcat        — LongCat-Video T2V/I2V/VC
  * sd35           — SD3.5 MMDiT (flow-match, image; triple-encoder joint embed +
                     pooled_projections)
  * gen3c          — GEN3C (EDM; 82ch pose-buffer DiT; camera/depth -> BRINGUP)
  * kandinsky5     — Kandinsky 5.0 T2V Lite
  * flux2          — FLUX.2 dev/klein (MMDiT, image; gated weights -> BRINGUP)
  * stable_audio   — Stable Audio Open (audio modality)
  * hunyuangamecraft, hyworld, lingbotworld, matrixgame2, matrixgame3 — interactive
                     world models; t2v/degenerate path CPU-verified, action/camera/
                     memory conditioning is BRINGUP (needs request-API extension)

Adapters declared via ComponentSpec.adapter (the ac29750b enabler) so each port
adds only NEW files (recipe package + per-arch torch_<arch>.py + optional facade
stub) — zero shared-file edits. Registry resolves all 31 bucket-C HF ids by exact
id + 14 architecture fallbacks; no regression (cosmos2/wan/ltx2 unchanged).
v2 mini green (211 passed, 2 skipped); pre-commit clean (per-file/registry).
2026-07-04 17:15:34 +00:00
SolitaryThinker bfbd90ea3a [feat] v2: Cosmos-Predict2-2B-Video2World port (EDM-Karras) — reference bucket-C recipe
First net-new architecture ported via the self-contained recipe-package pattern
(card-declared adapter + new loop, no shared-dispatch edit):

* CosmosDenoiseLoop (v2/recipes/cosmos2/loop.py): EDM preconditioning folded into
  a flow-match Euler integrator. Faithful port of CosmosDenoisingStage — Karras
  sigma schedule (rho=7, sigma_max=80 -> sigma_min=0.002, terminal clamp), latent
  init randn*sigma_max, per-step c_in/c_skip/c_out (sigma_data=1) -> x0, CFG in x0
  space, x0 -> velocity (x-x0)/sigma, FLOW_MATCH_STEP. video2world frame-replace
  conditioning threaded but inert for the t2v preset.
* build_karras_sigmas helper added to v2/loop/sampler.py.
* CosmosDiT + CosmosT5Encoder adapters (v2/platform/backends/torch_cosmos.py),
  declared on the card via ComponentSpec.adapter (the ac29750b enabler) — DiT
  returns raw EDM output + builds the mandatory zero condition/padding masks + fps;
  T5 uses the raw last_hidden_state (no Wan zero-pad). Reuses the WanVAE adapter.
* card/program/registry (HF id nvidia/Cosmos-Predict2-2B-Video2World + arch
  fallback on CosmosTransformer3DModel) + COSMOS_NEG prompt + SamplingDefaults
  (35 steps, gs 7, 704x1280, 93f, 16fps).

CPU-verified: Karras schedule, card/program build, registry resolve (id + arch),
EDM loop runs end-to-end on the CPU toy backend. GPU load/run is BRINGUP.
v2 mini green (211 passed, 2 skipped); pre-commit clean.
2026-07-04 17:15:34 +00:00
SolitaryThinker 5fd6e23e30 [feat] v2 torch backend: ComponentSpec.adapter — card-declared per-arch TorchComponent
A new architecture can declare its own torch adapter on the card
(ComponentSpec.adapter="module:Class") instead of editing the shared _make_dit/
_make_vae/_make_text_encoder dispatch. _explicit_adapter() constructs it as
cls(module, *extra, device=, dtype=) and short-circuits the built-in Wan/LTX2
class-name dispatch when set. This makes each bucket-C port a self-contained
recipe package (card + adapter module + loop + program) with no shared-file edit
-> conflict-free parallel porting. Unset -> unchanged built-in dispatch.

CPU mini green (211 passed, 2 skipped).
2026-07-04 17:15:34 +00:00
Will Lin fc3332550c [feat] v2: Wan2.2-I2V-A14B (MoE i2v) — combine boundary-routed experts + i2v conditioning
Reuses everything: 2 WanTransformer3DModel experts + BoundaryTimestepRouting (from the A14B MoE pattern),
the CLIP image encoder + first-frame [mask|cond] conditioning + the i2v program (from the Fun-InP i2v
port), and the shared WanDenoiseLoop (i2v hooks + the boundary expert). No new adapter. CPU-verified: the
toy MoE i2v runs end-to-end (2 experts + boundary + conditioning -> finite video); resolves with i2v caps.
Structural (GPU-pending: 2x14B, like the A14B T2V). Wan family now largely covered (T2V 1.3B/14B/TI2V-5B/
A14B, causal SF, i2v 1.3B/14B/A14B). CPU mini 211/2.
2026-07-04 17:15:34 +00:00
Will Lin 7464ef8308 [feat] v2: register Wan2.1-I2V-14B 480P/720P (reuse the GPU-verified i2v card)
The 14B i2v variants reuse the Wan2.1 i2v card/path proven on Fun-1.3B-InP — just per-variant params
(480P flow_shift 3.0 / 480x832, 720P flow_shift 5.0 / 720x1280). Registry resolution + caps verified;
specific 14B weights GPU-pending (same generic Wan i2v loader path that Fun-InP validated). i2v cluster
now supported; roadmap updated (11 models ported).
2026-07-04 17:15:34 +00:00
Will Lin 4a56274d10 [feat] v2: Wan2.1 i2v port (Wan2.1-Fun-1.3B-InP) — CLIP encoder + first-frame conditioning, GPU-verified
Real image-to-video, unlocking the i2v cluster. v2/recipes/wan21/i2v.py: CLIP image-encode -> the DiT's
encoder_hidden_states_image; first-frame VAE conditioning + a 4-channel mask -> the 20ch [mask|cond] that
the Wan adapter concatenates with the 16ch noise -> the 36ch i2v DiT input (mirrors fastvideo's
ImageEncodingStage + ImageVAEEncodingStage; v2's WanVAE.encode already applies the matching (z-mean)/std).
Reuses the shared WanDenoiseLoop (its None-default i2v hooks) + the Wan torch adapter unchanged. Adds
ToyImageEncoder + the image_encoder checkpoint subfolder stamp. Registered Wan2.1-Fun-1.3B-InP.

GPU-verified: loads via the generic Wan loader (1.56B, no param-mapping issue), runs the full i2v
conditioning, produces real video (3,9,256,384, std 0.44, finite, motion 0.041). CPU mini 211/2 (T2V
unregressed). BRINGUP: visual confirmation that the output follows the conditioning image is
human-in-the-loop.
2026-07-04 17:15:34 +00:00
Will Lin 9ede9af123 [feat] v2 Wan loop+adapter: optional i2v conditioning hooks (cond concat + CLIP context); T2V unchanged
Threads i2v conditioning through the SHARED WanDenoiseLoop with zero T2V risk: init() reads optional
slots i2v_cond (the [mask|cond] latent) + i2v_img_embeds (CLIP) into scratch; _velocity passes them to
the dit (context=, cond=); WanDiT concats cond (16->36ch) and uses the embeds as
encoder_hidden_states_image; capture is disabled only when i2v conditioning is present. For T2V both are
None -> the dit call, CFG, and cudagraph capture are byte-identical (CPU mini 211 pass, 2 skip — no
regression). ToyDiT accepts+ignores cond (image-conditioning is a GPU-path concern). Completes the i2v
backend seam; the program's mask+cond construction + the Wan i2v card/registry + GPU verify follow.
2026-07-04 17:15:34 +00:00
Will Lin 3d405f6dfa [feat] v2 torch backend: CLIP image-encoder adapter + image_encoder component kind (i2v groundwork)
Adds CLIPImageEncoder (encode_image -> the DiT's encoder_hidden_states_image) + the generic builder's
image_encoder maker (ImageEncoderLoader + ImageProcessorLoader), registered as the cuda 'image_encoder'
kind. Mirrors fastvideo's ImageEncodingStage. CPU-verified (component-kinds + lazy invariant, 211/2);
GPU path marked BRINGUP/written-not-run (processor subfolder + dtype to confirm on a real i2v checkpoint),
matching how the rest of the torch backend was originally landed. Reusable by the Wan i2v cluster + many
bucket-C models (Hunyuan/Cosmos i2v). Next i2v increments: the mask+cond latent construction + the
concat-into-DiT-input loop, then the card + registry + GPU-verify with Wan2.1-Fun-1.3B-InP.
2026-07-04 17:15:34 +00:00
Will Lin d891771ba3 [revert] v2: drop FastWan/VSA registry entries — generic Wan loader can't map their gated-attn params
GPU verification (Wan2.1-Fun... no: FastWan2.1-T2V-1.3B) failed at load: 'Parameter blocks.0.to_gate_compress.bias
not found in custom model state dict' — the FastVideo/* DMD-distilled checkpoints carry gated-attention
params (to_gate_compress) that the generic WanTransformer3DModel loader can't map. v2/registry.py's
select_by_architecture ALREADY rejects WanDMDPipeline for exactly this reason; my explicit ModelEntry
wrongly bypassed it. Reverted FastWan (1.3B + 14B-480P) and the unverified VSA-14B alias (same FastVideo/*
risk). Kept the official Wan2.1-T2V-14B (standard weights, same loader path as the GPU-verified 1.3B).

Lesson recorded in V2_PORTING_STATUS.md: FastWan/Turbo/VSA need a param-mapping fix (like LTX-2.3 did),
not just a schedule — bucket-C-effort. GPU-verify every port before claiming support. CPU mini 211/2.
2026-07-04 17:15:34 +00:00
Will Lin 67dea39052 [feat] v2: alias FastVideo/Wan2.1-VSA-T2V-14B-720P to the Wan-14B card (bucket B)
Same WanTransformer3DModel arch (VSA is an attention-backend choice, not a weight/arch difference); v2
runs dense TORCH_SDPA, so it resolves to build_wan_t2v_14b_card. Registry resolution verified; the GPU
forward path is the 1.3B-proven Wan adapter (specific 14B/VSA weights not separately GPU-run).
2026-07-04 17:15:34 +00:00
Will Lin 5f1d2ef7d2 [feat] v2: port Wan2.1-T2V-14B (bucket B) + document the all-models backlog
First bucket-B port toward 'support every fastvideo model': Wan2.1-T2V-14B reuses the Wan recipe +
torch adapter unchanged (same WanTransformer3DModel/AutoencoderKLWan/UMT5) — only a registry entry +
build_wan_t2v_14b_card (720p, flow_shift 5.0) + SamplingDefaults differ. Without the entry the arch
fallback would give it the 1.3B 480p defaults; the explicit entry gives 50 steps / 720x1280.

Also recorded the full backlog in V2_PORTING_STATUS.md: 63 fastvideo models = 8 ported / 21 bucket-B
(reuse Wan/Causal/LTX2 arch — registry+recipe+defaults, no new adapter) / 34 bucket-C (13 new
architectures needing a TorchComponent adapter). Updated the stale 'how to add a model' steps to the
post-redesign structure (v2/recipes/, torch_backend.py, SamplingDefaults). CPU mini 211 pass, 2 skip.
2026-07-04 17:15:34 +00:00
Will Lin 440b99523e [refactor] v2 torch backend: TorchComponent base + one generic builder + v2.* facade (Phase 1b/1c)
Addresses the adapter-setup pains: collapses the torch_cuda(trampolines)/torch_adapters/torch_ltx2 split
+ 11 near-identical adapter classes + 6 build_torch_* builders into:
- v2/platform/backends/torch_backend.py: a TorchComponent base centralizing .to/.eval, the numpy<->torch
  marshalling (ONE place), the set_forward_context wrap, and the weight surface; thin per-model subclasses
  (WanDiT/LTX2DiT/WanVAE/LTX2VAE/T5Encoder/Gemma/LTX2Upsampler/LTX2AudioVAE/LTX2Vocoder) carrying only
  forward semantics; and ONE build_component(spec) dispatching by spec.kind via _MAKERS.
- torch_cuda.py: registers that single generic builder for all 6 cuda kinds (no per-kind trampolines).
- v2 owns its namespace via re-export STUBS (facade, marked '# STUB'): v2/forward_context, v2/fastvideo_args,
  v2/distributed, v2/loader (the load_component seam), v2/api, v2/models/{dits,audio,upsamplers}/*. All v2
  code imports v2.*; 'from fastvideo' now lives ONLY in those 8 stub files -> a future per-module vendored
  cutover swaps a stub body, no caller changes. No divergence (stubs run fastvideo's live code).
- Deleted torch_adapters.py + torch_ltx2.py.

Verified: CPU mini 210 pass/2 skip; lazy invariant (platform load imports no torch); GPU bit-parity LTX-2.3
T2VS (audio std 0.04304, identical to pre-redesign) + Wan2.1 (video std 31.94, motion 5.997).
2026-07-04 17:15:34 +00:00
Will Lin 1cb2b4e84c [feat] v2: per-model sampling defaults on ModelCard (Phase 1a)
v2 had no per-model defaults — generate_video hardcoded 30 steps/25 frames/480x832/cfg5/16fps for
every model, badly wrong for e.g. LTX-2 distilled (wants 8 steps @1024x1536) or Wan2.2-TI2V (704x1280@24fps).

- New SamplingDefaults dataclass + ModelCard.sampling_defaults (v2/card/specs.py), exported from v2.card.
- Populated all 7 supported cards from fastvideo's InferencePreset defaults (steps/guidance/HxW/frames/fps
  + per-modality guidance for LTX-2.3 A/V). Negative prompts copied verbatim into v2/recipes/_prompts.py
  (recipe DATA, not model code -> v2-owned, no fastvideo import).
- VideoGenerator stores the resolved card; generate_video applies card defaults with precedence
  kwargs > SamplingParam > card > generic fallback (pure _resolve_default helper, unit-tested).
- test_sampling_defaults.py: per-card values + precedence (incl. empty-neg-prompt edge). CPU mini 210 pass, 2 skip.
2026-07-04 17:15:34 +00:00
Will Lin 51898f48e9 [refactor] v2: rename models/ (recipe layer) -> recipes/; move toy backend -> platform/backends/toy.py
Frees v2/models/ to become the vendored-architecture namespace that mirrors fastvideo/models
(part of making v2 self-contained / able to replace fastvideo). The v2 recipe layer (per-family
card.py/loop.py/program.py + common.py + the build_*_engine re-exports) is the recipe, not the
architectures, so it moves to v2/recipes/. The pure-numpy toy/parity implementations (ToyDiT etc.)
move from v2/models/backend.py to v2/platform/backends/toy.py (alongside cpu.py/accel.py/torch_*).

Mechanical: all imports are absolute, so v2.models.<x> -> v2.recipes.<x> and v2.models.backend ->
v2.platform.backends.toy across v2/ + examples/ (89 files, 178 refs). No behavior change.
CPU mini green (202 passed, 2 skipped); all v2 files compile.
2026-07-04 17:15:34 +00:00
SolitaryThinker 4e331eb7ff [refactor] v2: use absolute imports (v2.*) everywhere instead of relative
Mechanical conversion of every relative import under v2/ to an absolute v2.* path
(from .x / ..x / ...x -> from v2.<pkg>.x) so imports are unambiguous, grep-able, and
stable when code is copied/moved between entrypoints (VideoGenerator, CLI, server).

Surgical prefix-only rewrite: only the 'from <dots><module>' prefix changed — import
names, parentheses, multi-line formatting, comments, and ordering are byte-for-byte
preserved (no collapsing, no reorder, no unrelated reformatting).

- 471 imports across 125 files; v2/tests/ was already absolute (untouched).
- Validated: all 125 files compile, every 'from v2.* import' target resolves to a real
  module/package, zero relative imports remain (full sweep), CPU mini suite green
  (202 passed, 2 skipped).
2026-07-04 17:15:34 +00:00
SolitaryThinker c6d2976fc2 [feat] v2 VideoGenerator: A/V convenience path (generate_video -> T2VS -> mp4 + 24kHz wav)
Makes the 'Full A/V' LTX-2.3 deliverable reachable from the user-facing entrypoint, not just the
engine. A model advertising TEXT_TO_VIDEO_SOUND (LTX-2.3) now auto-issues a T2VS request, so generate()
/ generate_video() return BOTH modalities in one joint pass:
- VideoGenerator stores the resident instance + a supports_av flag (from card.capabilities); generate()
  gains want_audio (None=auto-by-capability, True/False to force) and routes T2V vs T2VS+{video,audio}.
- _result saves the stereo waveform as a sibling .wav at the vocoder's REAL rate (24000) — read off the
  built audio_vae adapter (TorchLTX2AudioVAE.sample_rate = Vocoder.output_sample_rate), since the
  AudioArtifact default rate is a placeholder. Populates GenerationResult.audio/.audio_sample_rate and
  extra['audio_path']. scipy IEEE-float WAV; [channels,samples] auto-transposed.
- Rewrote v2_basic_ltx2_3_distilled.py: registry routes to build_ltx2_3_card (its own joint T2VS A/V
  card, not the LTX-2 base/2-stage card); the example prints both the mp4 and the wav.
- GPU-verified via the convenience API: ev.mp4 + ev.wav (24000 Hz, stereo 61920x2, nonzero, std 0.043).
  CPU mini green (202 passed, 2 skipped); engine/program/toy paths untouched (test_ltx2_av pins 44100).
2026-07-04 17:15:34 +00:00
SolitaryThinker 6096b00aeb [feat] v2 LTX-2.3 T2VS GPU-verified: audio VAE/vocoder wiring + dual-connector audio fix
The full joint text->video+audio LTX-2.3 path now generates on the real 18.99B model:
- GPU audio components: build_torch_audio_vae (AudioDecoderLoader -> LTX2AudioDecoder, chains the
  vocoder) + build_torch_vocoder (VocoderLoader -> LTX2Vocoder); registered the 'audio_vae'/'vocoder'
  cuda component kinds; stamped their checkpoint subfolders (_WAN21_SUBFOLDERS).
- Fix: TorchGemma.encode_av must pass output_hidden_states=True — the 2.3 connector's SEPARATE audio
  projection lives in hidden_states[0] only then (gemma.py:703); without it the audio text fell back to
  the video embedding (4096 vs 2048 -> audio cross-attn shape mismatch).
- GPU-verified T2VS: video (3,33,256,384, std 0.68) + audio (stereo 2x61920 @24kHz, nonzero, std 0.059).
- test_torch_backend: cuda component kinds now include audio_vae + vocoder. CPU suite green (202+2).
2026-07-04 17:15:34 +00:00
SolitaryThinker 1d23399d81 [feat] v2 LTX-2.3 T2VS: single-stage joint audio+video card/loop/program (CPU-verified)
Makes LTX-2.3 a first-class, faithful card (was wrongly merged into the single-stage base):
- LTX23DenoiseLoop (loop.py): single-pass joint A/V denoise — one DiT forward per step cross-attends
  video<->audio via the adapter's (v_vel,a_vel) return; full-res video latent + a [8,T,16] audio latent;
  distilled few-step schedule (BASE_SIGMAS). Video-only when no audio requested.
- build_ltx2_3_card (model_id 'ltx2.3-distilled' — the name now correctly names the REAL 2.3): 5
  components incl. audio_vae (AudioDecoder) + vocoder (required_for t2vs, optional_for t2v); caps
  T2V + T2VS. build_ltx2_3_program: dual-connector text-encode -> joint denoise -> video + audio decode.
- registry.py routes FastVideo/LTX-2.3-Distilled-Diffusers -> this card (split from the base entry).
- Toy support: ToyTextEncoder.encode_av (separate video/audio text), ToyDiT joint A/V (audio now a
  keyword-only arg so positional  callers like the talker are unaffected), channel-agnostic
  ToyAudioVAE (np.resize identity for the existing 2-stage T2VS).
- CPU-verified: toy T2VS -> video+audio (8 steps), T2V -> video-only; CPU suite green (202+2).
GPU audio-VAE/vocoder loaders + the real-T2VS GPU verify are the next step.
2026-07-04 17:15:34 +00:00
SolitaryThinker fefcd415ff [feat] v2 LTX-2 adapters: A/V foundation (joint DiT forward + dual text connector + audio decode)
Foundation for the LTX-2.3 T2VS port (card/loop/program wiring + GPU verify to follow):
- TorchLTX2DiT.__call__ gains an optional joint audio path: pass audio_latent[8,T,16] + audio_text and it
  feeds audio_hidden_states/audio_encoder_hidden_states/audio_timestep/audio_sigma in ONE forward
  (LTX-2.3 cross-attends video<->audio) and returns (video_velocity, audio_velocity). Video-only call is
  byte-for-byte unchanged (audio_latent=None).
- TorchGemma.encode_av returns the SEPARATE (video_text, audio_text) projections from the 2.3 connector
  (video=last_hidden_state, audio=hidden_states[0]); 2.0 returns them equal.
- TorchLTX2AudioVAE (AudioDecoder -> Vocoder -> waveform@24kHz) + TorchLTX2Vocoder wrapper.
Additive + backward-compatible; CPU suite unaffected (adapters are GPU-lazy).
2026-07-04 17:15:34 +00:00
SolitaryThinker 7ac2ff0d1c [refactor] v2: shared model registry (HF-id primary + arch fallback) for all entrypoints
Per review: dispatch should be a directly-mapped HF-string -> card registry (like fastvideo), shared
by every entrypoint (VideoGenerator + a future CLI / server), not buried in the generator.

- New v2/registry.py mirrors fastvideo's fastvideo/registry.py hybrid resolution: (1) exact HF repo id
  in an explicit ModelEntry registry [PRIMARY — correct per-model card/capabilities, and the only way to
  split same-architecture capability variants like Wan2.1 T2V vs the i2v 'InP' 1.3B], (2) short repo-name
  match, (3) architecture inference [FALLBACK — local paths / unregistered repos]. resolve(model_path[,
  root]) is the single shared entry point.
- video_generator.py: moved _read_arch_signature/_select_builders into the registry; from_config now
  calls resolve() — registered ids resolve with no config read, else arch inference on a cheap *.json
  snapshot. Reconciles the earlier 'no brittle table' refactor with the 'map hf string -> card' ask: one
  clean registry + a fallback, not three coupled structures.
- Verified: registry resolves exact-id / short-name / arch-fallback / unregistered correctly; CPU suite
  green (202+2); wan21 GPU smoke generates via the new resolve path.
2026-07-04 17:15:34 +00:00
SolitaryThinker fa9c58b419 [fix] v2: name LTX-2 cards by architecture (2-stage vs single-stage) + Wan2.1 is T2V-only
Addresses the 'how is ltx2 separate from ltx2.3' confusion + a wrong capability:

- LTX-2 cards renamed by ARCHITECTURE (the version labels did not map to it): build_ltx2_card model_id
  'ltx2.3-distilled' -> 'ltx2-2stage-distilled' (two-stage base->upsample->refine; serves the
  upsampler-having FastVideo/LTX2-Distilled-Diffusers); build_ltx2_base_card 'ltx2.base' ->
  'ltx2-single-stage' (one loop; serves Davids048 base + the single-stage FastVideo/LTX-2.3-Distilled,
  which has NO spatial_upsampler). Dispatch already splits on has_spatial_upsampler. Updated the
  model-id refs in the mini's tests/examples.
- Wan2.1 base is T2V-only: dropped the wrong Capability.IMAGE_TO_VIDEO + narrowed components'
  required_for to {t2v} (i2v is the separate InP variant; v2 has no i2v path yet). build_wan21_card is
  shared by wan21 + wan2.2-ti2v; the A14B card was already T2V-only.
- CPU suite green (202 passed, 2 skipped).
2026-07-04 17:15:34 +00:00
SolitaryThinker 0662b42510 [docs] v2: LTX-2 base/2.3 GPU-verified on rebuilt x86 stack + remaining-port mechanisms
- LTX-2 base (Davids048) and LTX-2.3-Distilled both generate real video (inter-frame motion 4.5 / 6.6)
  via the single-stage base card — moved to Working (7 models now verified).
- Environment: the aarch64 venv was rebuilt for x86 (torch 2.11.0+cu128) + re-validated (CPU suite green,
  wan21 + LTX-2 base/2.3 generate).
- Remaining Wan+LTX-2 ports documented with concrete mechanisms: Wan2.2-i2v (SigLIP image_encoder +
  VAE-encode first-frame + concat-mask -> larger-in_channels i2v DiT), TurboWan (RCMScheduler consistency
  loop), Lucy-Edit (Wan v2v via VideoVAEEncodingStage), FastWan (VSA, env-blocked: no nvcc).
2026-07-04 17:15:34 +00:00
SolitaryThinker ae6d8085de [docs] v2: LTX-2.3 example + roadmap (A14B offload working, base/2.3 ported, env status)
- v2_basic_ltx2_3_distilled.py: LTX-2.3-Distilled routes to the single-stage base card (no
  spatial_upsampler) via the arch dispatch; pass few steps for the distilled schedule.
- V2_PORTING_STATUS.md: A14B moved to Working (CPU expert offload, 60GB peak); LTX-2 base + 2.3 added
  (code-complete, GPU re-verify pending); Environment-status note on the mid-session aarch64->x86 host
  reschedule that blocks GPU re-verify.
2026-07-04 17:15:34 +00:00
SolitaryThinker d0648ba8d9 [feat] v2: Wan2.2-A14B MoE CPU offload (fits 1 GPU) + LTX-2 base/2.3 single-stage port
Within the bounded Wan+LTX-2 scope:

- Wan2.2-A14B MoE now GENERATES on a single 80GB GPU via CPU offload (TorchWanDiT offload_group):
  the two 14B experts live on CPU and only the active one is swapped onto the GPU at the boundary-
  timestep transition (a single swap, not per-step). GPU-verified: 60GB peak (vs 79GB OOM), produced
  wan22_a14b_lion.mp4 (17x480x832, std 54.7, motion 8.36). Single-expert Wan stays resident.

- LTX-2 base (single-stage) port: build_ltx2_base_card + build_ltx2_base_program reuse the LTX-2
  adapters at FULL latent res with a request-driven many-step flow-match (LTX2DenoiseLoop full_res/
  request_steps/base_flow_sigmas; distilled base/refine path preserved via False defaults). The SAME
  single-stage card serves LTX-2.3-Distilled (also single-stage: no spatial_upsampler) — dispatched by
  the new has_spatial_upsampler discriminator in _select_builders. v2_basic_ltx2.py added; VideoGenerator
  gains shutdown() for API parity.

- Fixes: from_config 'os' scoping (shadowed module import); LTX-2 upsampler per_channel_statistics
  source (the AE's .decoder, not the top-level module).

Verification status: A14B offload, the upsampler, and the arch-dispatch refactor were GPU-verified
earlier this session. LTX-2 base/2.3 are CPU-verified (cards/programs build, dispatch routes, schedule
correct); their GPU smoke tests were pending when the box was rescheduled aarch64->x86 mid-session,
which broke the aarch64 venv (numpy/torch unrunnable on x86) — GPU re-verify blocked on the env.
2026-07-04 17:15:34 +00:00
SolitaryThinker e10828346f [feat] v2: architecture-driven dispatch + real LTX-2 upsampler + Wan2.2-A14B MoE card
Two reviewer asks + the next Wan port, all within the bounded Wan+LTX-2 scope:

1. Architecture-driven dispatch (replaces the HF-id table + substring fallback): from_config reads the
   checkpoint's pipeline/transformer/VAE class names (+ z_dim, transformer_2) and picks the v2 card via
   _select_builders — mirroring fastvideo's get_pipeline_config_cls_from_name. Resolves local paths /
   renamed repos / new distilled variants of a known arch with no table edits, and cleanly REJECTS
   FastWan (detected by WanDMDPipeline) with a precise message instead of a confusing load crash.

2. Real LTX-2 spatial upsampler (was a nearest-neighbor np.repeat stand-in): new 'upsampler' component
   kind -> TorchLTX2Upsampler wraps the real LTX2LatentUpsampler and applies the repo's upsample_video
   (un_normalize via the VAE decoder's per_channel_statistics -> learned 2x upsample -> normalize). CPU
   keeps ToyUpsampler (np.repeat) via the factory terminal, so the program calls
   component('spatial_upsampler').upsample(...) on both backends with no device branch. GPU-verified:
   9x512x768, std 70.5, motion 6.51.

3. Wan2.2-T2V-A14B MoE card (build_wan22_a14b_card): two WanTransformer3DModel experts +
   BoundaryTimestepRouting @0.875. GPU-verified that both experts denoise; OOMs in VAE decode on one
   80GB GPU (~70GB resident) — upstream offloads the DiT for MoE; documented as offload-blocked.

CPU suite green (202 passed, 2 skipped). FastWan root-caused (non-strict load of VSA gate_compress +
VSA not built); roadmap (V2_PORTING_STATUS.md) updated with the bounded scope + per-model status.
2026-07-04 17:15:34 +00:00
SolitaryThinker 655f362cf4 [feat] v2 port: Wan2.2-TI2V-5B (T2V) — 4th GPU-verified model
- Wan2.2-TI2V-5B reuses the Wan adapters (WanTransformer3DModel / AutoencoderKLWan / UMT5); deltas
  are the higher-compression VAE geometry (z_dim=48, 16x spatial, 4x temporal) and 480p flow-shift 5.0.
  The DiT forward accepts a scalar timestep (1D path), so no per-frame expand_timesteps for pure t2v.
- WanDenoiseLoop / build_wan21_card gain optional geometry params (latent_channels/spatial_ratio/
  temporal_ratio) defaulting to Wan2.1 (16/8/4) -> wan21 path unchanged; build_wan22_ti2v_card sets
  48/16/4. Registered as family 'wan2.2-ti2v' in VideoGenerator; v2_basic_wan2_2_ti2v.py added (T2V).
- Verified: 25x448x768 mp4, std 62.3, inter-frame motion 4.89 (coherent). CPU suite green (202+2skip).
- Corrected the now-disproven FastWan='wan21 reuse' mapping (its DMD checkpoint to_gate_compress param
  mapping differs); roadmap updated (TI2V-5B working; A14B MoE + I2V remain).
2026-07-04 17:15:34 +00:00
SolitaryThinker 3541e81d66 [feat] v2 VideoGenerator: convenience API (from_pretrained/generate_video) + porting roadmap
- VideoGenerator gains the convenience surface most basic examples use: from_pretrained(model,
  num_gpus/*_cpu_offload/...) and generate_video(prompt, sampling_param=, **kwargs) -> result, on top
  of the typed from_config/generate. Accepts SamplingParam.
- v2 examples matching the upstream convenience-API examples for the verified models: v2_basic.py
  (Wan2.1) and v2_basic_self_forcing_causal.py (SF-causal).
- V2_PORTING_STATUS.md: honest per-family roadmap. Working: wan21, wan_causal, ltx2-distilled. Each
  further model needs per-model work (FastWan: WanDMD to_gate_compress param mapping; TurboWan: RCM
  consistency sampler; Wan2.2: MoE card; LTX2 base/i2v; new families: new cards/adapters; gated Flux2 /
  local GEN3C / audio StableAudio / interactive MatrixGame blocked in this env).
2026-07-04 17:15:34 +00:00
SolitaryThinker f51497ee6d [feat] v2 VideoGenerator: typed fastvideo.api entrypoint over the v2 engine
Mirrors fastvideo.entrypoints.VideoGenerator (from_config(GeneratorConfig) -> generate(
GenerationRequest) -> GenerationResult.video_path), reusing the OFFICIAL fastvideo.api config classes
so a basic_dmd_new_api.py-style script differs only by importing VideoGenerator from v2.
- model_path -> v2 card registry (Wan2.1 / FastWan -> wan21; SFWan -> wan_causal; LTX2 -> ltx2);
  snapshot_download + stamp_wan21_checkpoints + Engine(cuda) + program; generate maps SamplingConfig
  -> DiffusionParams -> make_request -> eng.run, saves the [C,T,H,W] decode as an mp4.
- Lazy v2.__getattr__ keeps 'import v2' torch-free (verified) so the CPU mini stays green (202+2).
- examples/inference/basic/v2_basic_new_api.py runs all three GPU models through this API.
- Single-GPU, resident, TORCH_SDPA (EngineConfig offload/num_gpus>1/VSA accepted for parity, not applied).
verified: wan21 from_config->generate->mp4 (frames (5,256,256,3) uint8, video_path written).
2026-07-04 17:15:34 +00:00
SolitaryThinker d8af2e60d2 [feat] v2 ltx2: two-stage distilled GPU bring-up (LTX2Transformer3DModel 18.88B)
Official FastVideo/LTX2-Distilled-Diffusers. New torch_ltx2.py adapters (build_torch_* dispatch on
class):
- TorchLTX2DiT: patchify-internal; per-token timestep ones(B,tok,1)*sigma (sigma direct) + per-sample
  video_sigma; DiT predicts x0 so the adapter returns velocity=(x_t-x0)/sigma for the v2 flow-match step.
- TorchLTX2VAE: CausalVideoAutoencoder decode (internal per-channel un_normalize).
- TorchGemma: LTX2GemmaTextEncoderModel (Gemma + feature-extractor + connectors) -> last_hidden_state.
- ltx2 loop: real 128-ch latent geometry on cuda (32x spatial / 8x temporal; half-res base, 2x upsample).
e2e two-stage (8+3 steps) -> coherent, high-quality video (surfers at sunset), (3,9,512,768). NOTE:
still uses the v2 program's np.repeat upsampler between stages (the refine regenerates from noise so
output is faithful-quality); real LTX2LatentUpsampler swap-in is a follow-up.
2026-07-04 17:15:34 +00:00
SolitaryThinker f79919ba8a [feat] v2 wan_causal: causal DiT (CausalWanTransformer3DModel) GPU bring-up
Official SF checkpoint wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers (reuses TorchWanVAE + TorchT5Encoder).
- TorchWanDiT detects the causal transformer: ignores the chunk_rollout loop's latent `context` (the
  real model conditions across chunks via an internal kv_cache, not a forward arg) -> dispatches to
  full-attention _forward_train, and passes a per-latent-frame timestep [B, num_frames] (the causal
  block asserts a per-frame temb), uniform per chunk.
- chunk_rollout real geometry on cuda (16ch; chunk_size latent frames; 8x spatial).
e2e: chunk_rollout over the SF student -> coherent video (cat in a garden), (3,21,480,832). Fidelity
gap (artifacts): the v2 loop's per-chunk few-step sampling != the official kv-cache streaming + SF
schedule (a follow-up).
2026-07-04 17:15:34 +00:00
SolitaryThinker 1594f8e6be [fix] v2 tests: make no-torch-import guards GPU-aware (skipif torch installed)
The cuda-availability probe imports torch by design; the no-torch-import invariant is only
verifiable when torch is absent. skipif torch installed -> green on GPU box (202 passed, 2 skipped),
still enforced in torchless CI.
2026-07-04 17:15:34 +00:00
SolitaryThinker 64cadaa0bf [feat] v2 wan21: backend-aware latent geometry + checkpoint stamping
- latent_shape(req, model): real Wan geometry (16ch; (T-1)//4+1, H/8, W/8) on the cuda backend;
  toy stand-in stays on accel/cpu.
- stamp_wan21_checkpoints(card, model_root): map a root (local dir or HF id) onto the 3 components'
  ComponentSpec.checkpoint; build_wan21_card(checkpoint_root=...) optional.
2026-07-04 17:15:34 +00:00
SolitaryThinker 750fc1245b [feat] v2 cuda backend: real fastvideo construction + risk A-E fixes (Wan2.1 verified on H100)
Take the written-not-run torch adapters to runs-and-generates on 1x H100 (aarch64):
- A: FastVideoArgs.from_kwargs(model_path=root) builds the real pipeline_config; single-GPU dist
  init; load each component from its subfolder; tokenizer from the sibling <root>/tokenizer.
- B/C: DiT forward wrapped in set_forward_context(attn_metadata=None) (SDPA dense path);
  timestep=sigma*1000 + bare-velocity output confirmed.
- D: VAE decode denormalizes z*std+mean; removed the double-mean (it re-added
  shift_factor==latents_mean) that washed out the video.
- E: UMT5 from config; text embeds zero-padded to text_len (Wan t5_postprocess_text) - the fix
  that took output from a dark blur to a coherent prompt-matching scene.
- Components run at native precision (DiT bf16, VAE/text fp32); checkpoint check before dist init.
2026-07-04 17:15:34 +00:00
SolitaryThinker 7725998b0c [docs] add v2/HANDOFF.md for GPU-side bring-up of the torch backend
Orientation + process doc for an agent on a GPU branch: the 6 commits already
landed, the files to touch, the gating tasks (Risk A FastVideoArgs/checkpoint),
the verification bar (CPU suite stays 204; GPU generation matches a reference via
the SSIM harness), commit/push rules (no Claude co-author; don't rewrite history;
wandb token referenced not embedded), and the gotchas. Points to
GPU_BRINGUP.md for the detailed checklist + risk table.
2026-07-04 17:15:34 +00:00
SolitaryThinker 4b61dedc43 [fix] correct GPU adapters against real fastvideo API (cross-check findings)
Adversarial cross-check of the written-not-run torch adapters against the real
fastvideo source confirmed the interface contracts (DiT returns bare velocity
tensor; timestep=sigma*1000; encode().mode() + bare decode; .last_hidden_state;
no fused solver kernel) but caught a wrong construction layer. Fixed in code:

- Construction: WanTransformer3DModel / AutoencoderKLWan have NO from_pretrained.
  Replace it with the real FastVideo loaders (TransformerLoader / VAELoader /
  TextEncoderLoader + TokenizerLoader, each load(model_path, fastvideo_args)).
  The loader resolves the class from the checkpoint config — so UMT5-vs-T5 is
  chosen correctly instead of hardcoded (was BLOCKER #1/#4/#5).
- Text encoder: wrap the forward in set_forward_context(...) — the (U)MT5
  attention reads global state via get_forward_context(); a bare call mis-encodes
  (was BLOCKER #2). Drop the wrong padding="max_length".
- VAE: apply latent normalization the DiT expects — (z-mean)*inv_std on encode,
  inverse on decode, with latents_std stored as its reciprocal; shift_factor
  before decode (was BLOCKER #3). Skipping it yields washed-out video, not error.

The remaining unknowns are genuinely box-dependent (FastVideoArgs fields,
shift_factor placement, exact tokenizer kwargs, FSDP) — GPU_BRINGUP.md reconciled
to mark what's now fixed-in-code vs what still needs the box. 204 CPU tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 1e58d90d02 [feat] real torch/CUDA backend (written-not-run) behind the cuda cells
Implement the GPU backend the substrate was built for: Platform.detect() ->
cuda resolves real torch adapters + torch solver ops instead of the numpy
rungs, with the existing loops/policies/scheduler/training unchanged.

WRITTEN-NOT-RUN: this environment has no GPU/torch, so the torch code is
grounded in the verbatim real fastvideo APIs (DiT forward signature confirmed
from source) but cannot be executed/verified here. It is gated available=False
(CPU mini stays green; importing the backends never imports torch), with every
on-box confirm point marked `# BRINGUP` and an ordered checklist in
platform/backends/GPU_BRINGUP.md.

- torch_adapters.py: TorchWanDiT / TorchWanVAE / TorchT5Encoder wrap the real
  module named by each card's load_id and bridge it to the mini's duck-typed
  surface (numpy<->torch at the boundary; loop math stays numpy fp32). DiT
  weight-surface (copy_from/blend_from/clone) for serving sync; mse_grad_step
  raises (GPU training is a separate workstream).
- torch_kernels.py: flow_match_step / flow_sde_step as plain torch elementwise.
  Grounded conclusion from the kernel audit: fastvideo-kernel ships NO fused
  solver kernel (only attention/norm/quant primitives), so the cuda solver is
  torch, registered at arch generic with an honest source string.
- torch_cuda.py: rewritten as lazy trampolines (torch imported only inside
  builder/kernel bodies). Adds the missing vae + text_encoder cuda components
  (they'd otherwise silently fall back to the toy) and corrects the dishonest
  "fastvideo-kernel:flow_*" labels.
- ComponentSpec.checkpoint: the weights source for the torch adapter (risk A;
  the one field the cards didn't carry). Empty on the CPU toys.
- 7 CPU-verifiable wiring tests: honest registration/sources, torch-free import,
  cuda-resolves-real-cells-not-toy, build-fails-loudly-without-torch.

204 tests pass (CPU). The torch path needs a GPU box to verify (GPU_BRINGUP.md).
2026-07-04 17:15:34 +00:00
SolitaryThinker 47f7a04e09 [feat] static-buffer capture form for the cudagraph step body (Path A)
Close the loudest deferred gap from the cudagraph audit: the capturable step
now binds its I/O to address-stable static buffers (modeling real CUDA static
I/O buffers), instead of allocating fresh arrays per call.

- StaticWorkspace: address-stable buffers allocated once per capture key; bind()
  copies the current step's inputs in place via np.copyto, which RAISES on a
  shape/dtype mismatch — turning the weak peak_activation_bytes proxy into a real
  key-soundness backstop (a step whose shape doesn't fit can't replay an
  incompatible graph). Output written into a static buffer too.
- WorkPlan.graph_fn / graph_inputs: a capturable step exposes its deterministic
  op-structure as graph_fn(model, workspace) reading EVERY per-step input (latent,
  sigmas, conditioning, scale) from the workspace — never from closure over
  per-step data — plus the dict of current values. Loops without both stay on the
  eager path. wan21 factors a shared _velocity() so graph_fn and the eager run
  stay bit-identical.
- Capturer dispatch captures/replays via graph_fn against the keyed workspace;
  the workspace is shared per key on the instance. Correct under the engine's
  synchronous step execution (bind+graph_fn atomic per dispatch, output returned
  as a copy) — proven by the batch-of-N interleave gate running two same-key
  requests through the shared workspace bit-identically. A concurrent/multi-stream
  executor would need a per-stream pool (documented).
- 2 new tests (no-static-form eager-break, static-buffer shape-mismatch raises);
  workspace collision test split into bytes-proxy vs shape-backstop.

197 tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 9b8838834c [feat] piecewise CUDA-graph capture/replay at the step boundary (Path A)
Wire the capture/replay lifecycle into the driven-loop step boundary — the
other half of Path A (hand-fused kernels behind the registry + piecewise
cudagraphs, no compiler). Models and tests the correctness-critical control
logic; replay re-runs the current step thunk (CPU models the lifecycle, not the
GPU speedup).

- GraphCapturer on the instance (cross-request cache), wired into
  RuntimeLoopContext.execute and gated by LoopSpec.graph_capture ==
  "breakable_cudagraph". Capture key = (device, arch, loop, shape_sig,
  resident-weight-versions, graph_key). Eager-break for non-capturable (SDE) and
  interceptor-overridden steps. Never executes a stored thunk, so interleaved
  requests can't smear state.
- WorkPlan.capturable / graph_key. wan21 sets capturable=not sde and folds
  compute dtype (shape_sig.dtype) + CFG branch set + expert + scheduler-precision
  into the key — closing a cross-precision key-collision corruption path a
  non-fp32 build would otherwise hit on a real GPU (audit finding).
- Version-in-key auto-invalidation + real eviction: set_weights_version evicts
  the synced component's graphs (duck-typed, so card/ imports no runtime),
  preventing a GPU graph leak across FlowGRPO's per-iteration syncs.
- register_kernel gains the workspace_bytes capture-safety contract (declared in
  the matrix; cuda cells declare real scratch, numpy reference is 0).
- 11 tests: capture-once/replay-many, eager-break (SDE + override), capture ≡
  pure-eager bit-identical, recapture on shape + weight-version change with
  eviction, eager-loop gating, accel-backend capture, and capturer unit tests
  (key discrimination, eager-break, workspace-collision safety net, eviction).

Honestly deferred (bite a real GPU, not the CPU tests): the static-buffer
refactor of the step body, admission budgeting of capture cost (GRAPH_CAPTURE),
and per-card opt-in beyond wan21 — all documented in cudagraph.py + README.

195 tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 634f0828a2 [feat] route all diffusion loops + RL recompute through the kernel table
Finish the kernel seam across the board so the platform's KernelTable is the
universal solver-dispatch path, not just wan21.

- Loops: ltx2, wan_causal, adapters, adaptive now resolve the flow-match solver
  via model.platform.kernels.get(FLOW_MATCH_STEP) instead of importing the numpy
  sampler directly (wan21 already did). On CPU this is bit-identical (the cpu
  kernel IS the old function); a GPU/accel backend now overrides every loop.
- RL: the FlowGRPO log-prob recompute in unified_rl / joint_multi_rl /
  workflow_rl dispatches FLOW_SDE_STEP through the platform, pinned to the SAME
  kernel the rollout used (C2 kernel-pinning — otherwise the PPO ratio biases on
  a real GPU where rollout and recompute kernels could differ).
- accel backend: add a kind-generic AccelComponent wrapper and override the vae
  component too (text_encoder left unregistered to keep the device->cpu fallback
  demonstrated), closing the "only dit is overridden" gap.
- tests: vae override assertion; a second-loop (wan-causal chunk rollout) parity
  oracle proving accel == cpu bit-identical beyond wan21. README scope updated.

184 tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 2f044c02dd [feat] multi-backend dispatch substrate (device/arch/kernel registries)
Add v2/platform/: the (recipe, runtime) backend membrane that lets CPU, GPU,
and other devices coexist behind one dispatch substrate.

- Two tuple-keyed registries: COMPONENTS(kind, device, variant) for
  weight-bearing components and KERNELS(op, device, arch, variant) for
  stateless primitives, each with an availability predicate and an enumerable
  manifest (declared-but-unavailable cells listed without importing torch).
- Platform: detected (device, arch) owning the device + arch fallback chains
  and a per-platform cached KernelTable; detect() is honest (CPU/numpy unless
  torch+CUDA are actually present). Arch fallback is monotonic — only degrades
  to older/portable archs, never a newer binary-incompatible one.
- Seams wired: ModelInstance.component() -> platform.build_component(spec,self)
  with spec.factory as the numpy terminal rung (existing cards untouched); the
  wan21 denoise thunk dispatches solver ops through model.platform.kernels.
- Three backends: cpu (numpy terminal + parity oracle), accel (pure-python
  stand-in proving cross-device resolution, arch fallback, and the oracle),
  torch_cuda (declared-but-unavailable; no faked GPU).
- test_platform.py: 16 tests — detection, terminal rung, arch-fallback walk +
  monotonicity, device precedence, variant fallback, component override +
  per-kind device fallback, the parity oracle (accel == cpu, bit_identical via
  the C1 ladder), and matrix enumeration without importing torch.

Scope is honestly bounded in the README/docstrings: only wan21-denoise routes
through the kernel table and only the dit kind is overridden today (each a
one-line adoption); the torch/CUDA path and a cudagraph workspace-safety
contract are declared/deferred, not implemented. 183 tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 431f4daddb [feat] Adapter plane, non-linear workflows, RL→distill flywheel
Three more capabilities on distinct untested surfaces (167 tests pass, no new runtime primitive).

A — Adapter plane (§9.19): one base + swappable LoRA/ControlNet adapters, selected per request
(DiffusionParams.adapters); AdapterDenoiseLoop applies each active adapter's velocity delta. Per-request
selection changes output, multi-LoRA composes, ControlNet conditions on a control image, mixed-adapter
requests interleave without smearing, hot-swap changes generation, cache key partitions by adapter stack
(the adapter_versions field, previously declared-only). ToyLoRA/ToyControlNet; models/adapters/. 6 tests.

B — Non-linear workflows (§9.17): ParallelWorkflow (fan-out: one input → N models → merged) and
BestOfNWorkflow (generate N → score with the served reward card → return best; inference-time scaling).
The shapes a linear chain can't express. 4 tests.

D — RL→distill flywheel (§9.18): run_flywheel RL-improves the base (NFT), then distills FROM the RL'd model
(DMD2 teacher = RL'd policy) into a faster card, recording the base→rl→distilled provenance chain in
RecipeSpec.parents. The distilled student is measurably closer to the RL'd teacher than the base; the
distilled card serves few-step. training/flywheel.py. 4 tests.

designv4 §9.17–§9.19 + layout/closing/counts (167 tests, 29 files).
2026-07-04 17:15:34 +00:00
SolitaryThinker 6220d02746 [feat] #8b speculative (draft-verify) decoding — exact + lower-latency AR
The last audit stress test. A cheap draft model proposes K tokens, the target verifies
them in one batched step, and SpeculativeARLoop accepts the matching prefix + one target
correction — a variable accepted-length per round (a ragged AR loop the model owns).

- Exactness: the emitted sequence equals the target's OWN greedy decode for any draft
  quality (every accepted token is one the target would produce; the correction is the
  target's token) — the speedup is free.
- Speedup scales with accept rate: draft-agree 0.3→1x, 0.7→3x, 1.0→4x=K tokens/round
  (fewer verify_rounds, the expensive model's latency steps, for the same output).
- Two components (draft + target) co-scheduled on one resident instance; each round an
  AR_TOKEN WorkUnit.

models/speculative/ (loop+card+program); backend ToyTargetModel/ToyDraftModel with a
shared length-dependent target formula (no degenerate fixed point). 5 tests; full suite
153 passed. designv4 §9.16 + counts (153 tests, 26 files).
2026-07-04 17:15:34 +00:00
SolitaryThinker 9489c6c1dd [feat] Five more stress tests: LTX-2 A/V, weight-sync, served reward, cache-dit, nested workflows
The remaining design_v3 probes (all except 8b speculative decoding). All fit with no new runtime
primitive (148 tests pass).

#6 LTX-2 joint audio+video (§9.11) — LTX-2 declared an audio_vae required_for t2vs but never used
   it; now a single 2-stage denoise carries a synchronized audio latent (conditioned on video),
   applies per-modality CFG (guidance_per_modality), and decodes via video VAE + audio VAE → video +
   audio. Gated on requesting audio, so the T2V path is byte-identical (existing tests untouched).
   ToyAudioVAE; build_ltx2_av_program. 5 tests.

#4 Live weight-sync under in-flight serving (§9.14) — WeightSyncController makes the freeze → drain →
   transfer → bump version + invalidate → resume lifecycle explicit. Tests: a mid-flight swap corrupts
   (the hazard); draining first leaves the in-flight request bit-identical to baseline while a
   post-sync request reflects new weights; transformer-only sync, so the frozen text-encoder cache
   survives. The RL flywheel's hardest correctness. 3 tests.

#5 Reward-model-as-a-served-card (§9.15) — a reward model is a card (scorer + a score loop emitting
   REWARD_BATCH units); ServedRewardScorer drop-in-replaces the numpy scorer so any RL method becomes
   RLHF/RLAIF with no method change. ToyRewardModel; models/reward/. 4 tests.

#7 Content-adaptive control flow (§9.12) — CacheDiTDenoiseLoop (isolated WanDenoiseLoop subclass)
   reuses the cached velocity when predictions barely change (cache-dit skip) and early-exits on
   convergence — variable step count; interleave parity holds across ragged loops. models/adaptive/. 4 tests.

#8a Nested workflows (§9.13) — a workflow stage can invoke another workflow (engine.run routes ids);
   requires/validate recurse; cycles caught at registration + a run-time guard (engine._wf_running).
   build_t2i_i2v_extend_workflow. 5 tests.

designv4 §9.11–§9.15 + falsifier/layout/closing updates (148 tests, 25 files). Also removed a
pre-existing unused import in ltx2/loop.py.
2026-07-04 17:15:34 +00:00
SolitaryThinker 69c9871154 [examples] Add v2_examples/{training,omni,workflows}/ — runnable examples
Three more example folders alongside inference/, all CPU/numpy, self-contained
(sys.path bootstrap), public API only, every script verified to run green.

training/ (7) — one per method, all via the uniform method.train_step seam:
  01 finetune · 02 dmd2 distillation · 03 diffusion_nft (likelihood-free RL,
  samples from the old policy, feature-cache reuse) · 04 self-forcing (causal
  chunk_rollout) · 05 joint LM+generator RL (UniRL; joint + prompt-only) ·
  06 N-way joint RL (per_expert vs shared credit) · 07 end-to-end workflow RL
  (T2I+I2V from one final-video reward).

omni/ (4) — 01 Cosmos3 (reason→joint denoise, shared MoT) · 02 BAGEL
  (text→image, shared MoT; scheduler prices both WorkUnit kinds) · 03 Qwen-Omni
  (thinker→talker→vocoder, three separate experts, text+audio) · 04 interleave
  parity across AR + diffusion loop types.

workflows/ (2) — 01 cross-model T2I→I2V workflow (image provably conditions the
  video) · 02 workflow as a first-class servable (requires/validate, address by
  id, register_workflows catalog, WorkflowRegistry).

Each folder has a README indexing its scripts.
2026-07-04 17:15:34 +00:00
SolitaryThinker 18dd295e8d [examples] Add v2_examples/inference/ — runnable Wan2.1 inference examples
Five self-contained, runnable scripts (CPU/numpy) for the Wan2.1-1.3B card on the
v2 runtime, each bootstrapping the repo onto sys.path so they run from anywhere:

- 01_basic_t2v.py                  minimal path: build engine → T2V request → run → video
- 02_params_and_reproducibility.py DiffusionParams knobs + seeded bit-identical reproducibility
- 03_streaming.py                  per-denoise-step preview chunks (OutputSpec stream)
- 04_concurrent_interleaved.py     step-interleaved batching + interleave parity gate + cache reuse
- 05_async_serving.py              AsyncEngine: concurrent generate, event stream, step-boundary cancel

+ README.md indexing them. All five run green; use only the public API.
2026-07-04 17:15:34 +00:00
SolitaryThinker 4c333e0509 [docs] Update v2/README to designv4 + current scope (127 tests)
- v2/README.md: point to designv4.md as the unified design (design_v3 as
  north star); refresh the scope table (joint/N-way/workflow RL, Qwen-Omni
  cascade, cross-model Workflow, tiled VAE co-scheduling, WorldModelSession),
  package layout (program/Workflow, runtime/session, the 7 methods, new model
  dirs), the demonstrated-stress-tests list (§9.3–§9.10), and counts (49→127,
  20 files). Sessions moved out of "out of scope"; WebRTC wire stays out.
- designv4.md: drop two intermediate absolute suite totals (milestone "91/97
  passed") in favor of "zero regressions" so the only absolute count is the
  current 127 (intro + layout) — no stale numbers.
2026-07-04 17:15:34 +00:00
SolitaryThinker 32a7a6b87b [feat] Three stress tests: interactive sessions, workflow RL, heterogeneous co-scheduling
Targets the three design_v3 claims that were most load-bearing AND least
exercised (sessions/realtime, training-plane boundary, the WorkUnit-generality
falsifier). All fit with no new runtime primitive (127 tests pass).

1. Interactive world-model session (runtime/session.py, §9.8) — the Session
   plane had ZERO coverage. WorldModelSession drives the causal chunk_rollout
   loop as a long-lived session: persistent cross-request world state on the
   Session.kv_handle, frame streaming, transactional step-boundary cancellation
   (a cancelled act leaves the world resumable), no cross-session smearing.
   Added only a continuation seam to the chunk loop (init seeds context from a
   world_context slot; default empty = unchanged one-shot path). 5 tests.

2. End-to-end RL over a cross-model workflow (training/methods/workflow_rl.py,
   §9.9) — trains BOTH flux-t2i and wan-i2v from ONE final-video reward. Rolls
   out the whole workflow with SDE capture in both instances; the same final
   advantage drives FlowGRPO PPO on each stage's transformer; two WeightSyncPlans
   on two instances. The earlier model (T2I) is trained by a reward on the final
   video — end-to-end credit across a model boundary — proven causal by a control
   (constant reward => zero advantage => nothing moves). 4 tests.

3. Heterogeneous WorkUnit co-scheduling (models/tiled/, §9.10) — the §17
   falsifier. VAETileLoop makes VAE decode a loop of VAE_TILE units; tiling is
   exact (== one-shot, C0), and VAE_TILE + DIFFUSION_STEP pipelines interleave
   bit-identically and co-run in one batch. Validates the mechanism; the
   economic half (does it pay) stays a port-time measurement. 4 tests.

designv4 §9.8–§9.10 + falsifier/layout/closing updates (127 tests, 20 files).
2026-07-04 17:15:34 +00:00
SolitaryThinker 58223c0c41 [feat] Register cross-model workflows as first-class named servables
Answers "what's the right way to name/register custom pipelines like T2I→I2V":
treat a Workflow like a card — a stable namespaced id in the same servable
namespace, declared dependencies, and a two-level registry. No new concepts;
mirrors how cards are registered (and vllm-omni's pipeline_registry).

- program/workflow.py: Workflow gains `requires` (the cards it composes, derived
  from stages) and `validate(engine)` (fail-fast if a required card is absent,
  P7). New WorkflowRegistry: declarative workflow_id -> builder catalog for
  out-of-tree/ad hoc use.
- runtime/engine.py: `_workflows` registry + register_workflow (validates deps,
  rejects id collision with a model_id) + serves(); engine.run routes a request
  whose model_id is a workflow to workflow.run — addressable exactly like a model.
  Single-model hot path untouched.
- runtime/async_engine.py + serving/server.py: serves() and /models include
  workflows (discoverable as servables).
- models/__init__.py: declarative `_WORKFLOWS` catalog (the cross-model analog of
  _BUILDERS) + register_workflows() helper; build_image_video_engine now registers
  the workflow too. Adding a custom pipeline = one catalog line.
- Naming convention: dotted/namespaced workflow_id (`image_video.t2i_i2v`),
  distinct from kebab model ids, collision-checked. Renamed from `t2i_then_i2v`.
- tests (+5, 12 total in the file): addressable by id, requires/validate, id
  collision, registry catalog, register_workflows helper. Full suite 114 passed.
- designv4 §9.6: the naming & registration convention documented.
2026-07-04 17:15:34 +00:00
SolitaryThinker b254d1affe [feat] Cross-model T2I→I2V workflow + N-way joint RL over arbitrary experts
Two more pipelines stress-testing the design, plus a BAGEL-placement note in
designv4. Both fit with no new runtime primitive (109 tests pass).

Pipeline 1 — cross-model T2I→I2V (program/workflow.py, models/image_video/):
- Realizes ProgramKind.WORKFLOW as a thin multi-instance orchestrator ABOVE the
  engine (the hot path stays single-instance). A Program composes one model's
  loops; a Workflow chains full engine.run calls across distinct cards, threading
  artifacts. (LTX-2 already covers same-card multi-stage; cross-model — FLUX→Wan
  — is the new capability the single-instance runner can't express.)
- flux-t2i (text→image) and wan-i2v (text+image→video) cards; the I2V program
  folds the conditioning image into text_embeds so WanDenoiseLoop is unchanged.
- Each model keeps its own interleave-parity guarantee (crossing instances is a
  Workflow boundary, not a loop step). 7 tests incl. video-depends-on-image.

Pipeline 2 — N-way joint RL (training/methods/joint_multi_rl.py, models/multi_expert/):
- JointMultiExpertRL generalizes UnifiedRLMethod (N=2) to N refiner LMs + a
  generator: one reward → one group advantage → N token-PG updates + 1 FlowGRPO
  PPO update, N+1 independent WeightSyncPlans. Proves the substrate was already
  N-ready (card holds N components/loops; per-component weight-sync; dict grad
  targets) — only the method body looped over two; now it loops over a list.
- credit="per_expert" learns all N cleanly; credit="shared" (faithful to UniRL)
  works but is noisier — the honest multi-agent credit-assignment result, a
  reward-shaping choice, not a substrate limit. 6 tests (N=1,3,4; prompt-only).

Fix — flow_sde_ml_velocity (loop/sampler.py): the toy FlowGRPO generator update
targeted the velocity the model already produced (a no-op once guidance_scale=1
was set for the C2 identity; the unified generator moved only on ~1e-7 noise).
The correct PG surrogate targets the max-likelihood velocity of the realized
sample — nonzero at ratio==1. Both UniRL and N-way generators now learn for real;
the C2 ratio==1 identity still holds (measured before the update).

BAGEL: MoT/shared-weight (one transformer on both loops), same row as Cosmos3;
real BAGEL's co-resident experts are expressible via the expert-routing policy
(partial sharing) — captured in designv4 §2.3.
2026-07-04 17:15:34 +00:00
SolitaryThinker c7e0a8e894 [feat] Add Qwen-Omni thinker→talker→vocoder model (3 experts, 3 loops)
Ports vllm-omni's canonical qwen2_5_omni omni-speech cascade as a v2 card:
a third weight-sharing topology — three disjoint experts (thinker, talker,
vocoder) on three loop types (ar_decode → ar_decode → audio_decode) in one
request, with chained cross-stage conditioning and streaming codec→waveform.
vllm-omni runs these as three opaque request-scheduled stages; v2 makes every
thinker token, talker token, and vocoder chunk a runtime-visible WorkUnit.

- models/backend.py: ToyTalker (a genuinely distinct AR expert, weight-salted)
  + ToyVocoder (streaming code2wav: codec tokens → waveform chunks).
- models/omni/vocoder_loop.py: VocoderLoop filling the pre-declared
  LoopKind.AUDIO_DECODE / WorkUnitKind.AUDIO_CHUNK slot.
- models/omni/ar_loop.py: ARDecodeLoop gains a configurable prompt_slot so two
  chained AR loops don't collide on the prefill slot (thinker vs talker).
- models/qwen_omni/: card (3 experts/3 loops) + program (tokenize → thinker →
  emit_text → thinker→talker full-payload hand-off → talker → talker→vocoder →
  vocoder → emit_audio). Cross-stage hand-offs are explicit Program nodes, the
  model-native form of vllm-omni's custom_process_input_func.
- _enums.py: Capability.TEXT_TO_SPEECH.
- tests/test_thinker_talker.py: 6 tests incl. three-loop interleave parity,
  cascade conditioning, AUDIO_CHUNK streaming. Full suite 97 passed.

designv4.md: §2.3 topology table extended to four topologies; new §9.5 on the
cascade; reference-synthesis + package layout updated.
2026-07-04 17:15:34 +00:00
SolitaryThinker b3ddf6014d [feat] UniRL/PromptRL joint LM+generator RL stress test + designv4
Stress-tests the v2 Card/Loop/Program design with a UniRL/PromptRL-style
joint RL recipe: a prompt-refiner LM expert and a flow generator expert,
two separate experts driven by two loop types in one request, both updated
simultaneously from a single RL reward.

- loop/sampler.py: flow_sde_step_with_logprob — FlowGRPO SDE rollout sampler
  (per-step Gaussian log-prob), distinct from the deterministic ODE serve step.
- request/params.py: gated sde_rollout/sde_noise_scale on DiffusionParams so
  the serve path stays byte-identical (default ODE).
- models/wan21/loop.py: gated SDE-rollout capture in WanDenoiseLoop.advance.
- models/backend.py: ToyPromptRefiner — a real REINFORCE categorical policy
  (the Qwen role), separate weights from the generator.
- models/unified/: the unified card+program — two disjoint experts (llm +
  transformer) on ar_decode + diffusion_denoise; the topological opposite of
  the Cosmos3 MoT card, same vocabulary.
- training/methods/unified_rl.py: joint GRPO — one reward -> group advantage
  -> LM token policy gradient + DiT FlowGRPO PPO; two LRs; prompt-only/joint
  flag; reuses the shared diffusion loop for rollout.
- training/weight_sync.py: WeightSyncPlan gains a component scope so the two
  experts version + cache-invalidate independently (LM sync never flushes the
  frozen text-encoder feature cache).
- tests/test_unified_rl.py: 9 tests incl. likelihood-based C2 identity, the
  two-loop interleave parity gate, joint vs prompt-only. Full suite 91 passed.

designv4.md: unified design doc reflecting v2 as built+tested, with the joint
RL stress test as the validating case study (the design held — new card +
new method, no new runtime primitive).
2026-07-04 17:15:34 +00:00
SolitaryThinker a9e5f6ee7a [refactor] rename package mini_fastvideo → v2
Directory rename (git mv, history preserved) plus rewrite of all references: absolute imports in
tests, the zero-dep runner, docstrings, comments, and the README. No behavior change.

Run: python3 -m pytest v2/tests/ -q ; python3 v2/run_tests.py ; python3 -m v2.examples
2026-07-04 17:15:34 +00:00
SolitaryThinker 7467076d72 [fix] mini-fastvideo serving: address adversarial-review findings (capacity/credit leaks, robustness)
Review confirmed the core bets (concurrent disaggregation is bit-identical, design conformance holds,
Dynamo genuinely optional, engine stays step-scheduled). Fixes for the untested failure paths:

- HIGH: pool capacity (RolePool.in_flight) no longer leaks when a disaggregated request is cancelled
  or errors mid-occupancy — DisaggregatedRunner.close() releases the occupied pool and AsyncEngine._run
  calls it in a finally (a cancel on a capacity-1 denoiser no longer bricks the pool).
- credit flow-control: cross-pool transfer wraps acquire/release in try/finally (no credit leak on a
  failing transfer); slot is re-homed only on a successful fetch.
- AsyncEngine: duplicate in-flight request_id is rejected (was a deadlock); submit() cancels the driver
  task when the consumer abandons the stream (client disconnect → no orphaned compute); bounded
  per-request history (no unbounded _events/_states/_results/_runners growth).
- cancellation is common-path on the OFFLINE path too (cancel check at the top of every runner.tick()).
- HTTP server: read timeout (slowloris guard → 408), body-size cap (→ 413), invalid Content-Length
  (→ 400), explicit StreamReader limit, and aclose() of the SSE generator on client disconnect.
- build_deployment_card no longer aliases one mutable CostModel across replica cards (dataclasses.replace),
  so online calibration of one worker's cost doesn't mutate another's.
- video-job tasks tracked (not fire-and-forget); server.close() cancels/drains them; jobs dict bounded;
  fleet affinity map bounded.

6 regression tests added for these paths. 82 tests pass (pytest + zero-dep runner).
2026-07-04 17:15:33 +00:00
SolitaryThinker 01dc0c3377 [feat] mini-fastvideo serving + fleet (our own version, Dynamo-optional)
Builds the full serving layer the design files specify, instead of deferring it to Dynamo:

- transport/ (§7.3): pluggable Connectors (in-proc zero-copy / SHM-fake copy) with chunk_ready
  readiness (vllm-omni) AND credit-based flow control (sglang-omni Relay); KVConnector protocol shape;
  TransferManifest.
- runtime/ (§6, §13; plan M3/M4): AsyncEngine — request queue, lifecycle state machine
  (waiting→running→completed/cancelled/failed), live AsyncIterator[OmniEvent] streaming, common-path
  cancellation, step-level concurrency. RolePool + DisaggregatedRunner (encoder→denoiser→decoder,
  capacity-aware dispatch, cross-pool transfers via connectors); disaggregated output is bit-identical
  to inline. No-progress detection (no busy-spin).
- deploy/ (§14, §6.3.5-6): DeploymentCard; OUR OWN LocalFleet (discovery, health/drain, least-loaded
  / cost-model / sticky-affinity routing) so we never rely on Dynamo; DynamoWorkerAdapter +
  FakeDynamoRuntime export the SAME card + cost model so Dynamo CAN front us — one object, two consumers.
- serving/ (§6.3.5, §12): framework-free stdlib-asyncio OpenAI server (our own version of the
  vllm-omni pattern): /v1/chat/completions (SSE), /v1/images/generations, /v1/videos (async job+poll)
  + /v1/videos/sync, /v1/models, /health, /metrics. A thin shim over the STEP-scheduled engine — the
  runtime-visible loop scheduler vllm-omni's request-scheduled opaque DIFFUSION stage lacks.

15 serving tests (real-socket HTTP+SSE via stdlib asyncio, disagg==inline, fleet routing, Dynamo
contract, cancellation); 76 tests pass total (pytest + zero-dep runner). ~7900 LOC.
2026-07-04 17:15:33 +00:00
SolitaryThinker f1dc587c74 [feat] mini-fastvideo phase 2: omni/MoT — Cosmos3 + canonical vllm-omni (BAGEL/lance)
One resident MoT instance runs BOTH an ar_decode loop and a diffusion_denoise loop on shared weights
(the §16 claim no DAG-of-engines can express), with both loops runtime-visible: the scheduler prices
ar_token AND diffusion_step WorkUnits — unlike vllm-omni's opaque DIFFUSION stage the scheduler never
sees inside.

- ARDecodeLoop: token decode until EOS/max_tokens, paged text-KV — the omni AR pathway (loop/§5).
- ToyMoTDiT: one module exposing an und pathway (ar_forward) AND a gen pathway (denoise __call__);
  ToyTokenizer. Binding both loops to one instance = shared weights, no duplication.
- models/cosmos3/: tokenize → reason(ar_decode) → pack(tokens→conditioning) → diffusion_denoise →
  vae_decode; sound_vae declared optional_for non-t2vs (the lazy-component P8 fix, not an env-var hack).
- models/bagel/: the canonical vllm-omni model — generate_text(ar_decode) → generate_image(diffusion),
  text+image outputs, both loops step-scheduled.
- The diffusion loop is WanDenoiseLoop reused (one loop definition bound to the MoT module).
- build_omni_engine() + an omni worked example; 7 omni tests (shared-instance, both-kinds-scheduled,
  interleave parity across loop types, lazy sound_vae). 61 tests pass (pytest + zero-dep runner).
2026-07-04 17:15:33 +00:00
SolitaryThinker 098bcf014a [fix] mini-fastvideo: address adversarial-review findings (admission liveness + §7.1 cache key)
- Admission fails fast with AdmissionInfeasible on infeasible/deadlocked reservations instead of a
  10M-iteration busy-spin: no-progress detection in run_to_completion/run_interleaved via a real
  progress token, plus feasibility pre-checks (need > pool capacity).
- Compute budget is now a refundable concurrency gate (release() refunds spent), not a
  never-refunded lifetime cap that silently deadlocks.
- §7.1: the text-encoder feature CacheKey carries adapter_versions + precision (no stale serve across
  te-LoRA stacks); per-component weight versions mean a transformer-only RL weight sync no longer
  flushes the frozen text-encoder cache (component-scoped invalidation, not wholesale).
- Interleave gate flags symmetric-empty output instead of passing it vacuously.
- skipped_steps counted only when the override is actually consumed; BatchScheduler wired for round
  batch-accounting (metric renamed stepped_units); ResidualCache.get cleanup; stream chunks carry a
  latent preview payload; dead progress-vars removed.
- 5 regression tests added for the previously-untested paths. 54 tests pass (pytest + zero-dep runner).
2026-07-04 17:15:33 +00:00
SolitaryThinker 270fae959d [feat] mini-fastvideo: model-native runtime per design_v3 (Wan2.1/LTX2 + 4 training methods)
A scoped, CPU-testable realization of design_v3.md — the architecture where the atomic unit
is a typed (recipe, runtime) ModelCard, the model owns loop semantics while the runtime owns
loop lifecycle, one resident instance runs many loops, and training records behavior on the
same loops it serves.

Implements:
- card/ loop/ runtime/ cache/ memory/ parallel/ parity/ extend/ program/ request/ training/
  spanning design_v3 §4-§13: ModelCard + validate(); driven loops (init/next/advance/finalize);
  step-interleaving Engine with reservation-before-admission + per-class caches keyed by CacheKey;
  the C0-C4 consistency ladder + the non-negotiable batch-of-N interleave parity gate.
- Inference: Wan2.1-1.3B (T2V), LTX2.3 (two-stage distilled, shared transformer), Wan-causal
  (chunk rollout + slab-KV streaming).
- Training (Wan2.1-1.3B), each driving the SAME loops the engine serves: finetune (flow-match),
  DMD2 (teacher/critic distribution matching), DiffusionNFT (likelihood-free C2, samples from the
  decay-blended old policy, group-relative advantages, shared-prompt cache reuse), self-forcing
  (causal chunk loop). The engine never imports training (the §10 dependency rule, grep-verified).

numpy-only core (no torch/GPU here); heavy Wan/LTX forwards are deterministic toy stand-ins with
lazy torch-adapter seams (ComponentSpec.load_id/factory) for a GPU box. 49 tests pass via pytest
and a zero-dependency runner; interleave parity verified (and a buggy module-global interceptor
provably breaks it). Omni-ready spine (ar_decode/chunk_step loop kinds, multi-loop instances,
LoopState.extension) for the phase-2 Cosmos3 + vllm-omni omni ports.

Run: python3 -m pytest mini_fastvideo/tests/ -q ; python3 -m mini_fastvideo.examples
2026-07-04 17:15:33 +00:00
SolitaryThinker 4a14c1afa3 update 2026-07-04 17:15:33 +00:00
SolitaryThinker f1c19050c3 design 2026-07-04 17:15:33 +00:00
1114 changed files with 80527 additions and 88611 deletions
+207
View File
@@ -0,0 +1,207 @@
# v2 ← M\*: Architecture Gap-Analysis & Improvement Roadmap
**Status:** exploration, flagged for review. **Date:** 2026-06-19.
**Source paper:** *M\*: A Modular, Extensible, Serving System for Multimodal Models* (arXiv 2606.12688,
Stanford/UW/CMU; Jha, Sagan, Kamahori, …, Kasikci, S. Wang). It is a universal serving runtime for composite
multimodal models built on the **Walk Graph** abstraction (a model is a dataflow graph `G`; a request is a
*Walk* — a labeled subgraph — and the runtime executes walks). It beats vLLM-Omni (~20% lower T2I latency on
**BAGEL**, up to 2.64× on I2I), SGLang-Omni (2.7× TTS throughput on **Qwen3-Omni**), and native V-JEPA2
rollout (12.5×). It explicitly names **FastVideo's own** sparse/sliding-tile attention, xDiT/PipeFusion/USP,
Inferix, and FlashDrive as techniques integratable into the graph runtime.
**Method:** a 28-agent workflow — 6 parallel v2-subsystem maps → 10 M\*-dimension analyses, each
*adversarially verified against the actual v2 code* → synthesis + a completeness critic. The critic's
corrections and three P0 claims were then **spot-verified by hand** (file:line below). This doc folds those
corrections in; it is the corrected, authoritative synthesis.
---
## 1. Executive summary
v2 already implements the **harder half** of M\*'s thesis and in several axes **exceeds** it:
- v2's `Program` *is* M\*'s graph `G` (typed `ComponentNode`/`ModelLoopNode` + edges).
- v2's `shared_weight_components` *is* M\*'s cross-Walk node sharing — BAGEL/Cosmos3/LTX2 each bind two
`ModelLoopNode`s to **one resident transformer** (`instance.component()` returns the same live object). This
is the exact MoT serving property the omni cards in this repo already express.
- v2 adds three things M\* (serving-only) has **no equivalent for**: a required+validated per-loop **cost
model**, a non-negotiable **interleave bit-parity gate**, and an **integrated training plane** (RL→distill
flywheel driving the *same* serving Loop).
- The `extend/` plugin seam (interceptors/observers/registry with capability negotiation) is precisely the
hook M\*'s "extensible / integrate FastVideo-STA, xDiT, Inferix, FlashDrive" call-out asks for — **v2
already has the seam M\* only gestures at.**
What v2 lacks is M\*'s **declarative authoring layer above the substrate**, and — the key insight — *much of
that substrate is already authored but inert*: v2 has declared the metadata for "minimum components per
request" (`required_for`/`optional_for` on every omni card) and "branch as a cache axis" (`guidance_sig`,
`CacheKey`) but **never wired it to an executor**. The substrate is ~80% built and switched off.
**Highest-leverage cluster:** three small, parity-safe wires that turn on inert substrate and unblock the
BAGEL/Qwen-Omni/Cosmos3 latency wins M\* measured **on the exact models this repo already runs** — plus one
P1 that aligns v2 with the paper's headline "extensible" claim using a seam v2 already has.
### Verified P0 correctness findings (spot-checked by hand)
1. **Runner divergence (real bug).** `v2/runtime/engine.py:88` → `nodes = self.program.nodes`;
`v2/runtime/disaggregated.py:96` → `nodes = self.program.active_nodes(self.request)`. The inline and
disaggregated runners execute *different node sets*. ✅ confirmed.
2. **EOS is faked.** `v2/recipes/omni/ar_loop.py` docstring says "done on EOS/max_tokens"; `next()` (`:46-48`)
checks **only** `max_tokens`. M\*'s marquee `DynamicLoop` use case (EOS) is unimplemented in the loop that
serves the Qwen-Omni Thinker/Talker and Cosmos3 reasoner. ✅ confirmed.
3. **`required_for`/`optional_for` have zero runtime consumers** (grep outside `specs.py`/recipes/tests is
empty). The min-components metadata is declared on every card and never read. ✅ confirmed.
---
## 2. Dimension table (corrected)
| # | Dimension | v2 status | Gap | Priority | Effort | Payoff | Action |
|---|---|---|---|---|---|---|---|
| 1 | Min-components per request (`required_for` + `when_task`) | substrate built, **inert** | real, cheap | **P0** | S | Consume `required_for` in `active_nodes`; unify `engine.py:88` onto `active_nodes`; deliver via registry/card builder so all ~40 cards inherit it |
| 2 | Real EOS + declarative `DynamicLoop` | early-exit emergent; **EOS faked** | real | **P0** | S | `ARDecodeLoop` honors `eos_id` + `req.sampling.stop`; add `LoopSpec.dynamic_stop` + `register_loop_stop`. **Training-enabling** (world-model rollout horizon) |
| 3 | CFG/branch as label over one paged KV pool | absent (`PagedKVCache` is a counter) | real | **P1** | L | `(namespace,label)` paged store w/ one budget; reuse `guidance_sig` for hash (NOT `partition_field`); by-ref via existing `InProcKVConnector`. AR path only (diffusion has no KV) |
| 4 | `extend/` plugin seam → integrate FastVideo-STA / Inferix | **seam exists, unused for attn** | real (paper headline) | **P1** | M | Expose FastVideo sparse/sliding-tile attention + Inferix block-diffusion as `Interceptor`/`EngineKind` plugins — the paper's named integration targets, on this repo's own code |
| 5 | `ParitySpec.output_determinism` (C3 distributional) | C3 rung defined, **0 users** | real, dormant | **P1** | S | Add field; `compare_outputs` consults it. **Training-enabling** (SDE/FlowGRPO stochastic rollouts) |
| 6 | Registry-driven delivery of #1 | present, not leveraged | integration | **P1** | S | Express `when_task`/min-components through `WorkflowRegistry`/card builders, not 3 bespoke recipe patches |
| 7 | Serving conductor + pluggable data plane | conductor exists (`serving/http.py`); **single-process transport** | real | **P2** | L | v2 already has the step-scheduled worker surface; gap is ZeroMQ/Mooncake + direct worker→worker tensor routing (today `InProcKVConnector` only) |
| 8 | Fleet/Dynamo placement + replicas | **live** (`deploy/fleet.py`,`dynamo.py`) | partial | **P2** | M | Fleet-level placement/affinity/replica is real & ≥M\*; missing piece is only the intra-engine `(node,Walk)→rank` map decoupled from model code |
| 9 | Per-node TP / SP + cross-rank transport | axis vocab **exists** (`sp` incl.); not wired to runtime | partial | **P2** | XL | Wire declarative degrees into runtime; Wan/LTX are **SP-native** (TP is a no-op there); populate `parallel_plan_hash` on the serving cache path |
| 10 | Named Walks + per-model state machine | `Program`=G, sharing real; no Walk/SM | real | **P2** | M | Defer until a *re-entrant* phase graph (Thinker↔Talker, rollout) needs it; #1 captures the min-components win without it |
| 11 | Declarative `Parallel/Sequential/Loop` IR | imperative loop classes | real (authoring) | **P2** | M | Thin Section IR lowering to flat `Program`; scope to one AR recipe |
| 12 | Streaming `ChunkPolicy` + `StreamBuffer` | causal-chunk emit **already ships** (`wan_causal`); `EdgeKind.STREAM` inert | real | **P2** | L | Declarative `ChunkPolicy` vocab over the existing chunk mechanism; needs concurrent producer/consumer runner (= pipelined scheduling). Inferix integration point |
| 13 | Speculative deferred-termination; loop-spanning CUDA graphs; N+1 prefetch; attn double-buffer | absent / per-step capture (14 cards) | real | **P3** | L | Gate behind a real GPU executor; unobservable on CPU-toy CI; loop-span needs an `allows_interleaving=False` carve-out |
| — | Cost model + interleave/consistency parity | **exceeds M\*** | none | **guard** | — | Do not regress; keep `step_cost_model` mandatory + `bit_identical` default |
| — | Integrated training plane (flywheel, weight-sync) | **exceeds M\*** | none | **guard** | — | Protect train==serve loop identity with a toy fixture |
---
## 3. P0/P1 deep-dives (sequenced)
```
PR-1 (P0) min-components ──┐
PR-2 (P0) real EOS ─┼─► prereqs for honest "DynamicLoop" + min-component claims; both training-enabling
PR-3 (P1) output_determinism (independent)
PR-5 (P1) extend/ plugin: FastVideo-STA / Inferix as Interceptors (independent; highest paper-alignment)
PR-4 (P1) CFG-as-label paged pool ──► depends on PR-2 (AR loop is the only KV consumer)
```
PR-1, PR-2, PR-3, PR-5 are mutually independent; PR-4 depends on PR-2.
### PR-1 (P0) — Turn on the inert min-components substrate + fix runner divergence
- **Change.** Extend `Program.active_nodes(request)` (`v2/program/specs.py`) to also drop any node whose bound
`ComponentSpec.required_for` (`v2/card/specs.py:144`) excludes `request.task` (and isn't in `optional_for`).
**Fix the bug:** change `v2/runtime/engine.py:88` to `nodes = self.program.active_nodes(self.request)` so the
inline `ProgramRunner` matches `DisaggregatedRunner` (`disaggregated.py:96`). Deliver the `when_task` gating
through the **registry/card builder** (`recipes/__init__.py`, `program/workflow.py:WorkflowRegistry`) so all
~40 cards inherit it uniformly — not three bespoke `program.py` patches.
- **Why (this repo's models).** BAGEL T2I currently steps the AR-text loop and Cosmos3 t2v materializes the
reasoner even though the cards declare `transformer required_for={'reason','t2i'}`, `vae required_for={'t2i'}`.
On the GPU backend that is wasted resident-weight load + wasted steps on every single-modality request —
exactly M\*'s "execute the MINIMUM components per request," delivered by consuming existing metadata.
- **Risk/invariant.** Validate in `ModelCard.validate()` that every active node's `reads` are produced by an
active node for each declared `TaskType` (avoid dropping a producer). Pure node-id filtering ⇒ serial and
interleaved still walk the same filtered list ⇒ §9.3 interleave bit-parity holds by construction. CPU-toy clean.
### PR-2 (P0) — Real EOS + declarative `dynamic_stop` *(also training-enabling)*
- **Change.** In `v2/recipes/omni/ar_loop.py`, `advance()` reads the emitted token; if it equals the model
`eos_id` (toy backend exposes `EOS=0`) or matches `req.sampling.stop` (`params.py:21`, currently dead),
register termination; `next()` returns `Done()` on stop OR `max_tokens`. Add `StopRegistry` to `LoopState` +
`register_loop_stop(name)` to the `LoopContext` protocol (`contracts.py:204`) and to
`DisaggregatedRunner`'s `RuntimeLoopContext`. Add `LoopSpec.dynamic_stop: bool=False`, opt the AR cards in.
- **Why.** The docstring-vs-code lie sits in the loop serving Qwen-Omni Thinker/Talker and the Cosmos3 reasoner;
M\*'s second named `DynamicLoop` use case (world-model **rollout horizon**) is exactly what `self_forcing` RL
needs — so this is both a serving-credibility fix and a training enabler (raise its payoff accordingly).
- **Risk/invariant.** `dynamic_stop=False` is byte-identical back-compat. Must pass **all three** parity gates:
serial==interleaved AND disaggregated==inline. **Not** in this PR: speculative deferred-termination (unobservable
on CPU-toy, fights the interleave invariant — P3, gated on GPU executor).
### PR-3 (P1) — `ParitySpec.output_determinism` (close the dormant C3 hole) *(training-enabling)*
- **Change.** Add `output_determinism: str = "bit_identical"` to `ParitySpec` (`card/specs.py:88`); make
`compare_outputs` (`parity/interleave_gate.py:54`) consult it (`bit_identical` → today's exact check;
`distributional` → a moment/tolerance check — land a simple moment match first; a real KS test is new code).
- **Why.** `ConsistencyLevel.C3` is defined and used by zero recipes; an SDE/FlowGRPO stochastic rollout cannot
honestly declare its parity contract and would falsely fail the bit-identical gate. Additive; default unchanged.
### PR-5 (P1) — Expose FastVideo's own attention + Inferix as `extend/` plugins *(highest paper-alignment)*
- **Change.** Use the existing `extend/{interceptors,observers,registry}.py` seam (capability-negotiated, with
per-(request,branch) `plugin_state` that already passes the interleave gate) to register FastVideo's
sparse/sliding-tile attention and Inferix-style block-diffusion as `Interceptor`s / an `EngineKind` plugin.
- **Why.** M\*'s title is "Modular, **Extensible**" and it explicitly lists FastVideo-STA, xDiT/PipeFusion/USP,
Inferix, FlashDrive as integratable. v2 already has the seam M\* only describes — this is where v2 most
directly answers the paper, using this repo's own attention code. Low risk (the seam + capability negotiation
already exist and are tested).
### PR-4 (P1) — CFG/branch as a LABEL over one paged KV pool
- **Change.** Rewrite `PagedKVCache` (`cache/classes.py:155-172`) from a block *counter* into a real
`(namespace,label)->[block-handle]` store with **one shared `total_blocks` budget** (M\*'s single-pool
property). Reuse the existing-but-unpopulated `CacheKey.guidance_sig` (`keys.py:53`) for the hash. Thread the
label through `ar_loop.py` (alloc/append/get per `(request_id, branch)`; prefill once per shared-prefix label;
combine via `CFGPolicy.combine`). Wire `ResourceRequest.cache_blocks` (`contracts.py:64`, zero consumers) into
admission per (class,label).
- **Why.** The dossier-identified driver of M\*'s BAGEL win (3 CFG contexts as 3 labels over ONE pool vs dense
per-context). Targets AR_DECODE (BAGEL `generate_text`, omni Thinker); **correctly excludes diffusion**
(Wan/LTX are bidirectional, no KV — their CFG stays dense-but-batched).
- **Corrections to bake in.** Do **NOT** add `branch_label` to `CacheKey.partition_field()` (CFG branches share
embeddings; partitioning by branch is a semantic bug). Do **NOT** add a new by-ref type — reuse
`InProcKVConnector` + `TransferManifest.cache_key`. Wiring `cache_blocks` admission is greenfield ⇒ effort **L**.
CPU version proves label/sharing semantics; the real latency win needs a FlashInfer paged kernel (out of scope)
— **merge** with a future "real KVCacheEngine" effort rather than landing isolated.
---
## 4. What v2 already does ≥ M\* — do NOT regress
1. **Required+validated cost model** on every `LoopSpec` (13-kind `WorkUnitKind`) — typed, pre-GPU-validated.
2. **Interleave bit-parity as a hard gate** (`parity.interleave_required=True` on 40+ cards). M\* has no such
gate (its speculative scheduling deliberately wastes steps). Load-bearing invariant; every new primitive
must pass it.
3. **C0–C4 consistency ladder** wired into RL methods, with first-divergence tap reporting. No M\* equivalent.
4. **Integrated training plane** — DiffusionNFT/DMD2/self_forcing, RL→distill flywheel, `WeightSyncController`
hot weight-sync with drain-to-boundary + scoped cache invalidation, driving the **same** serving Loop.
M\* is serving-only. Protect with a toy fixture asserting `rollout_loop` drives the served Loop object.
5. **CPU-toy parity for the whole stack** — loops/CFG/caches/parity/RL run in CI without a GPU. Every new
primitive must ship a toy exercise (this is what makes all PRs above testable without H100s).
6. **Partition-not-flush cache invalidation** + four independent per-class pools.
7. **`extend/` plugin seam** with capability negotiation (a 4-step distilled card *rejects* a residual-skip
interceptor) — M\* describes extensibility; v2 has the mechanism.
8. **Dynamo citizenship** (`deploy/dynamo.py`: one `DeploymentCard`+cost model, two consumers) — beyond M\*'s
self-contained runtime.
---
## 5. Dropped / merged / deferred (and why)
- **DROP declarative `Parallel` as a CFG-execution win.** The runner walks nodes linearly (ignores
`Program.edges`), so `Parallel` lowers to sequential sugar and the CFG 3-pass braid is already one
co-scheduled `WorkPlan.run`; splitting it risks the interleave gate. Salvage only the no-op refactor
extracting `branch_forward` from `WanDenoiseLoop._velocity`. Reassign `Parallel` to the placement workstream.
- **MERGE the full Walk/state-machine layer** into "defer until a re-entrant phase graph needs it" (PR-1 gets the
min-components win with ~20 lines, no new abstraction). If built: the validator must check a walk's node-id
order is a *subsequence* of `program.nodes` (not just membership) or the runner can reorder and break parity.
- **MERGE `StreamBuffer`/`ChunkPolicy` into pipelined-scheduling.** Causal-chunk emit *already ships*
(`wan_causal/loop.py` per-chunk `StepResult.emit` + slab-KV); the gap is the declarative `ChunkPolicy` vocab
+ a concurrent producer/consumer runner. If built: keep all policies pure (per-request `StreamBuffer` history,
not shared edge state) and restrict the bit-identical claim to the token-only handoff.
- **MERGE CFG-fan-out exec + cross-rank transport + PD loop-splitting into a multi-GPU-runtime program.** These
need real collectives (`v2/distributed/` is a stub) and KV-by-reference (KV lives in `CacheManager`, not the
transferable `slots`). **Keep cheaply now:** the *declarative* halves — per-component degree, `(node,Walk)`
placement key with node-only fallback, `ReplicaSet` under `LocalFleet`, and populate `parallel_plan_hash` on
the **serving** cache path (it is already populated in `training/behavior.py:40` — the gap is serving-only).
- **DEFER** speculative deferred-termination, loop-spanning CUDA graphs, N+1 prefetch, attention-plan
double-buffer — all gated on a real GPU executor; benefit unobservable on CPU-toy CI. Keep the cheap
`EngineKind` tag (`STATELESS|KV_CACHE|DIFFUSION`) now. Correct the stale `cudagraph.py:51-52` docstring
(per-step capture ships in 14 cards, not just wan21).
- **RESCOPE per-node TP.** Wan/LTX use `ReplicatedLinear` + **sequence parallelism** (`sp`), not TP; the `sp`
axis already exists in `parallel/plan.py:AXIS_NAMES`. The work is wiring degrees into the runtime, not
inventing vocabulary; a `tp_size=2` "one-line activation" is a no-op for the shipped models.
---
## 6. The first integration test, if/when multi-GPU placement work starts
The **live Qwen-Omni 2-GPU bring-up** (Thinker on rank 0, Talker+Code2Wav on rank 1; see
`v2_debug_videos/vlm.md` Session 4) is the natural first validation target for any `(node,Walk)→rank`
placement work — it is the one place this repo already has real multi-rank composite-model execution.
---
## Anchor files for P0/P1
`v2/program/specs.py`, `v2/runtime/engine.py` (**line 88 fix**), `v2/runtime/disaggregated.py`,
`v2/recipes/omni/ar_loop.py`, `v2/loop/contracts.py`, `v2/card/specs.py`, `v2/cache/{classes.py,keys.py}`,
`v2/parity/interleave_gate.py`, `v2/extend/{interceptors,registry}.py`, `recipes/__init__.py` +
`v2/program/workflow.py` (registry-driven delivery).
@@ -1,23 +1,20 @@
---
name: reseed-performance-baseline
description: Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, environment-caused benchmark shift, or reviewed v2 calibration using one or more reviewed normalized performance JSONs. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline in FastVideo/performance-tracking must be advanced, or when a new v2 exact comparable identity needs its first approved baseline. The workflow backs up existing history under /tmp, validates all source JSONs for the same legacy (model_id, gpu_type) target or the same v2 exact identity, rejects internally inconsistent source batches, uploads one success=true baseline record per accepted source JSON, and offers to clean local temp state after a successful upload.
description: Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, or environment-caused benchmark shift using one or more reviewed normalized performance JSONs. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline in FastVideo/performance-tracking must be advanced from a consistent batch of reviewed source results. The workflow backs up existing history under /tmp, validates all source JSONs for the same (model_id, gpu_type), rejects internally inconsistent source batches, uploads one success=true reseed record per accepted source JSON, and offers to clean local temp state after a successful upload.
---
# Re-seed Performance Baseline
## Purpose
Replace or advance the rolling performance baseline in the HF dataset
`FastVideo/performance-tracking`. Legacy targets are scoped by
`(model_id, gpu_type)`. V2 targets are scoped by exact comparable identity:
`workload_id`, `variant_id`, `benchmark_version`, `hardware_profile_id`,
`software_profile_id`, and `recipe_fingerprint`.
Replace or advance the rolling performance baseline for a single
`(model_id, gpu_type)` pair in the HF dataset
`FastVideo/performance-tracking`.
Performance comparison uses the median of up to the last 5 successful,
baseline-eligible records for the same target. Failed or calibration-only
records are useful audit history, but they do not move the future baseline
because `compare_baseline.py` loads records with `successful_only=True` and
`baseline_eligible_only=True`.
Performance comparison uses the median of up to the last 5 successful records
for the same model and GPU. Failed records are useful audit history, but they
do not move the future baseline because `compare_baseline.py` loads records
with `successful_only=True`.
This skill now reseeds from a reviewed batch of one or more source performance
JSONs. It uploads one new `success=true` record per accepted source JSON; it
@@ -25,13 +22,11 @@ does not blindly replicate one measurement into 3 or 5 records. The effective
reseed size is therefore dynamic and equals the number of provided, validated,
internally consistent source JSONs.
For baseline shifts with existing history, if the operator provides fewer than
3 records, call out that the last-5 rolling median may not move immediately. If
the operator provides 3 consistent shifted records, the rolling median usually
moves immediately. If the operator provides 5 consistent shifted records, the
last-5 window is effectively reset to the new runtime profile. For the first
approved v2 baseline of a new exact identity, one reviewed calibration seed is
enough for the next comparable run to leave `CALIBRATION_NEEDED`.
If the operator provides fewer than 3 records, call out that the last-5 rolling
median may not move immediately. If the operator provides 3 consistent shifted
records, the rolling median usually moves immediately. If the operator provides
5 consistent shifted records, the last-5 window is effectively reset to the new
runtime profile.
These records are intentional operator-approved baseline resets, not ordinary
independent main-branch persistence. Mark them clearly with provenance fields
@@ -71,10 +66,10 @@ approval, then upload reviewed accepted baseline records.
| Parameter | Required | Description |
|-----------|----------|-------------|
| `model_id` | Legacy required; v2 inferred | Benchmark id, e.g. `wan-t2v-1.3b-2gpu`. This maps to the HF subdirectory after `sanitize(model_id)`. For v2 records, use the `model_id` from each source artifact only as the upload directory; comparison is by exact identity. |
| `gpu_type` | Legacy required; v2 inferred | Exact GPU device string from the performance record, e.g. the L40S device name emitted by CI. V2 hardware matching uses `hardware_profile_id`; preserve `gpu_type` as display metadata. |
| `model_id` | Yes | Benchmark id, e.g. `wan-t2v-1.3b-2gpu`. This maps to the HF subdirectory after `sanitize(model_id)`. |
| `gpu_type` | Yes | Exact GPU device string from the performance record, e.g. the L40S device name emitted by CI. Baselines are GPU-specific. |
| `source_results` | Yes | One or more local paths or Buildkite artifact URLs for accepted shifted performance JSONs. Prefer normalized `normalized_perf_*.json` artifacts emitted by `compare_baseline.py`. Accept `source_result` as an alias only for a single JSON. |
| `max_intra_batch_regression` | No | Maximum allowed regression of any source JSON against the source batch median. Default: `0.05` (5%). |
| `max_intra_batch_regression` | No | Maximum allowed regression of any source JSON against the source batch median. Default: `PERF_MAX_REGRESSION` if set, otherwise `0.05` (5%). |
| `intent_rationale` | Yes | One-line explanation for why the baseline shift is legitimate. This is written into provenance and should be reused in the PR. |
Hardcoded defaults:
@@ -83,14 +78,10 @@ Hardcoded defaults:
supported by the code, but use the default unless the user explicitly asks).
- Local sync root: `/tmp/perf-tracking` (`PERFORMANCE_TRACKING_ROOT` override
is supported).
- Prepared-record staging root: `/tmp/performance_reseed_prepared`
(`PERFORMANCE_RESEED_STAGING_ROOT` override is supported). Keep it separate
and non-nested from the sync root.
- Backup root: `/tmp/performance_reseed_backup`.
- Download scratch root for source artifact URLs: `/tmp/performance_reseed_source`.
- Baseline window: last 5 `success=true`, `baseline_eligible=true` records
for the same legacy `(model_id, gpu_type)` target or the same v2 exact
comparable identity.
- Baseline window: last 5 `success=true` records for the same
`(model_id, gpu_type)`.
- Reseed count: dynamic. Upload exactly one accepted seed record per validated
source JSON.
@@ -124,24 +115,12 @@ with open(source_result, encoding="utf-8") as f:
record = json.load(f)
```
Classify the source batch before continuing:
- **Legacy source records** have no v2 exact identity fields. Stop if any
normalized record's `model_id` or `gpu_type` does not match the requested
`model_id` and `gpu_type`.
- **V2 source records** have exact identity fields. Stop unless every source
record has all six comparable identity fields and they are identical across
the batch: `workload_id`, `variant_id`, `benchmark_version`,
`hardware_profile_id`, `software_profile_id`, and `recipe_fingerprint`.
Do not fall back to legacy `(model_id, gpu_type)` matching for v2 records.
Stop if any normalized record's `model_id` or `gpu_type` does not match the
requested `model_id` and `gpu_type`.
The source records may have `success: false` when they came from failed
rolling baseline comparisons. That is expected; only the reviewed reseed
records become new `success: true` baseline records after explicit approval.
For a first v2 baseline seed, the source records must instead be successful
scheduled-main full-suite `CALIBRATION_NEEDED` normalized artifacts. Reject PR,
local, direct-run, non-main-branch, or non-full-suite calibration artifacts as
seed sources.
Sort validated source records by their original `timestamp` ascending before
preparing the seed records. If a source timestamp is missing or unparsable,
@@ -169,7 +148,8 @@ For each metric with at least two non-null source values:
4. Stop if any source record regresses against the batch median by more than
`max_intra_batch_regression`.
Default `max_intra_batch_regression` to `0.05`. Print a table with per-source values, batch median, and
Default `max_intra_batch_regression` to `PERF_MAX_REGRESSION` when set,
otherwise `0.05`. Print a table with per-source values, batch median, and
worst intra-batch regression.
This check prevents uploading a mixed batch where one JSON is materially
@@ -203,7 +183,7 @@ present, that run is not a valid source for baseline reseeding.
### 2. Sync and back up existing HF records under /tmp
Use `fastvideo/performance/hf_store.py` helpers directly. Do **not** use
Use `fastvideo/tests/performance/hf_store.py` helpers directly. Do **not** use
`compare_baseline.py` as a sync shortcut; on full main runs it can persist
records, while this step must only fetch and back up existing history.
@@ -212,16 +192,16 @@ The sync command pattern is:
```bash
export PERFORMANCE_TRACKING_ROOT="${PERFORMANCE_TRACKING_ROOT:-/tmp/perf-tracking}"
export HF_REPO_ID="${HF_REPO_ID:-FastVideo/performance-tracking}"
python -c 'from fastvideo.performance.hf_store import sync_from_hf; import os; sync_from_hf(os.environ["PERFORMANCE_TRACKING_ROOT"], strict=True)'
PYTHONPATH=fastvideo/tests/performance python -c 'from hf_store import sync_from_hf; import os; sync_from_hf(os.environ["PERFORMANCE_TRACKING_ROOT"], strict=True)'
```
For legacy records, back up the sanitized model directory under `/tmp`:
Then back up only the sanitized model directory under `/tmp`:
```bash
SHORT_COMMIT=$(git rev-parse --short=12 HEAD)
TIMESTAMP=$(date -u +%Y%m%d_%H%M%S)
MODEL_SAFE=$(python - <<'PY'
from fastvideo.performance.hf_store import sanitize
MODEL_SAFE=$(PYTHONPATH=fastvideo/tests/performance python - <<'PY'
from hf_store import sanitize
print(sanitize("<model_id>"))
PY
)
@@ -230,16 +210,6 @@ mkdir -p "$BACKUP_DIR"
cp -R "${PERFORMANCE_TRACKING_ROOT}/${MODEL_SAFE}" "$BACKUP_DIR/" 2>/dev/null || true
```
For v2 records, back up the full local tracking root after sync. Exact identity
lookup scans across model directories, so a benchmark rename may have relevant
history outside the current source artifact's `model_id` directory:
```bash
BACKUP_DIR="/tmp/performance_reseed_backup/${TIMESTAMP}_${SHORT_COMMIT}_v2_exact_identity"
mkdir -p "$BACKUP_DIR"
cp -R "${PERFORMANCE_TRACKING_ROOT}" "$BACKUP_DIR/tracking-root"
```
Write provenance next to the backup:
```bash
@@ -262,12 +232,10 @@ first baseline seed. Continue, but report that baseline history was empty.
### 3. Compute old baseline and candidate shift
Load the last 5 successful baseline records for the target.
For legacy targets:
Load the last 5 successful records for the target:
```python
from fastvideo.performance.hf_store import load_records_for_model
from hf_store import load_records_for_model
records = load_records_for_model(
"/tmp/perf-tracking",
@@ -275,28 +243,6 @@ records = load_records_for_model(
"<gpu_type>",
last_n=5,
successful_only=True,
baseline_eligible_only=True,
)
```
For v2 exact-identity targets:
```python
from fastvideo.performance.hf_store import load_records_for_identity
records = load_records_for_identity(
"/tmp/perf-tracking",
{
"workload_id": "<workload_id>",
"variant_id": "<variant_id>",
"benchmark_version": "<benchmark_version>",
"hardware_profile_id": "<hardware_profile_id>",
"software_profile_id": "<software_profile_id>",
"recipe_fingerprint": "<recipe_fingerprint>",
},
last_n=5,
successful_only=True,
baseline_eligible_only=True,
)
```
@@ -312,8 +258,7 @@ medians after appending the proposed seed records, and source batch spread for:
Also print how many successful old records exist. Make clear:
- 1 seed record usually does not move an existing last-5 median by itself, but
it is enough to establish the first v2 baseline for a new exact identity.
- 1 seed record usually does not move a last-5 median by itself.
- 3 consistent seed records usually move the last-5 median immediately.
- 5 consistent seed records effectively reset the last-5 window.
- The records are intentional approved baseline resets and must be labeled
@@ -323,10 +268,10 @@ Also print how many successful old records exist. Make clear:
Require an explicit confirmation phrase before preparing the upload:
> About to RE-SEED performance baseline for `<target description>`.
> About to RE-SEED performance baseline for `<model_id>` on `<gpu_type>`.
> This will upload `<N>` new `success=true` records to
> `FastVideo/performance-tracking/<sanitize(model_id)>/` or the source
> artifact's v2 model directory, one per accepted source JSON.
> `FastVideo/performance-tracking/<sanitize(model_id)>/`, one per accepted
> source JSON.
>
> Reason: `<intent_rationale>`
> Source results: `<source_results>`
@@ -344,63 +289,25 @@ Do not continue unless the user types exactly `confirm performance reseed`.
### 5. Create the accepted seed records
Create one seed record from each normalized source result.
For first v2 baseline seeds, use the scoped utility. It validates exact
identity, requires successful scheduled-main full-suite `CALIBRATION_NEEDED`
source artifacts, preserves the normalized v2 identity and metadata fields,
and writes seed records with `success=true`, `baseline_eligible=true`, and
`comparison_status=PASS`:
```bash
python fastvideo/tests/performance/seed_baseline.py \
--source-result <normalized_perf_1.json> \
--source-result <normalized_perf_2.json> \
--intent-rationale "<intent_rationale>" \
--max-intra-batch-regression 0.05 \
--tracking-root "${PERFORMANCE_TRACKING_ROOT}" \
--staging-root "${PERFORMANCE_RESEED_STAGING_ROOT:-/tmp/performance_reseed_prepared}"
```
The utility is prepare-only and intentionally has no upload option. Upload the
scoped records only after the separate confirmation in step 6.
The utility validates against an isolated fresh HF snapshot and leaves
`PERFORMANCE_TRACKING_ROOT` untouched; that argument only proves the staging
root is separate from the operator's tracking mirror. Before writing, it stops
if the exact identity already has a successful baseline-eligible record or if
the workload/variant/version already trusts another recipe. It atomically
reserves the exact identity and writes a digest-protected upload manifest bound
to the current HF endpoint, repository id, and repository type. Keep the
prepared records, manifest, source files, and reservation unchanged until the
operation is uploaded or explicitly cleaned up.
If the prepared seed records look correct, upload only those scoped records in
step 7. Do not rerun the utility with a different source list after approval.
For legacy reseeds or accepted v2 baseline shifts from regression artifacts,
create one seed record from each normalized source result. Do not copy the
Create one seed record from each normalized source result. Do not copy the
source JSON wholesale.
Infer the baseline field allowlist from all existing HF records for the target
after syncing, including both `success=true` and `success=false` records. For
legacy targets the target is `(model_id, gpu_type)`. For v2 baseline-shift
reseeds the target is the exact comparable identity. Use the union of
non-provenance keys present in those target records, preserving only fields
that also exist in the normalized source record or are explicitly set by the
reseed workflow. Always include `model_id`, `timestamp`, `success`,
`baseline_eligible`, and `comparison_status` because the upload path and
baseline loader depend on them. Always set `timestamp` to a fresh reseed
timestamp, `success` to `true`, `baseline_eligible` to `true`, and
`comparison_status` to `PASS`. Do not include unrelated source-only fields
that are absent from existing HF records.
`(model_id, gpu_type)` after syncing, including both `success=true` and
`success=false` records. Use the union of non-provenance keys present in those
target records, preserving only fields that also exist in the normalized
source record or are explicitly set by the reseed workflow. Always include
`model_id`, `timestamp`, and `success` because the upload path and baseline
loader depend on them. Always set `timestamp` to a fresh reseed timestamp and
`success` to `true`. Do not include unrelated source-only fields that are
absent from existing HF records.
Exclude existing provenance or operator metadata from the inferred baseline
field allowlist. At minimum, exclude keys prefixed with `baseline_reseed` and
any fields known to be local-only audit metadata.
If there are no previous HF records for the target, fall back to this default
baseline field list:
If there are no previous HF records for the target model/GPU, fall back to this
default baseline field list:
- `model_id`
- `timestamp`
@@ -413,22 +320,6 @@ baseline field list:
- `dit_time_s`
- `vae_decode_time_s`
- `success`
- `baseline_eligible`
- `comparison_status`
For v2 baseline-shift reseeds with no previous HF records for the exact
identity, also preserve:
- `workload_id`
- `variant_id`
- `benchmark_version`
- `recipe_fingerprint`
- `hardware_profile_id`
- `software_profile_id`
- `recipe`
- `hardware_profile`
- `software_profile`
- `software_comparison_profile`
Do not upload extra fields from the source artifact.
@@ -444,22 +335,6 @@ Optional provenance fields are allowed and useful:
- `baseline_reseed_operator`
- `baseline_reseed_max_intra_batch_regression`
The v2 calibration seed utility writes analogous first-seed provenance:
- `baseline_seed: true`
- `baseline_seed_reason`
- `baseline_seed_source_result`
- `baseline_seed_source_status`
- `baseline_seed_source_timestamp`
- `baseline_seed_source_success`
- `baseline_seed_source_run_source`
- `baseline_seed_source_branch`
- `baseline_seed_source_test_scope`
- `baseline_seed_source_pr_number`
- `baseline_seed_batch_size`
- `baseline_seed_batch_index`
- `baseline_seed_operator`
Use a fresh reseed timestamp for each seed record, not the original source
result timestamp. This is required because
`load_records_for_model(..., last_n=5)` keeps the last records after loading
@@ -482,8 +357,7 @@ Prefer uploading new accepted seed records so failed history remains visible.
Print:
- Backup directory path under `/tmp`.
- Prepared local record paths under `PERFORMANCE_RESEED_STAGING_ROOT`.
- Prepared upload-manifest path under the identity reservation.
- Prepared local record paths under `PERFORMANCE_TRACKING_ROOT`.
- HF paths that will receive the new records.
- Old rolling medians.
- Source batch medians, source batch spread, reseed count, and candidate
@@ -495,36 +369,22 @@ prepared records plus backup on disk.
### 7. Upload only the scoped records
For a first v2 calibration seed, use the manifest uploader after the user
replies exactly `upload`:
Use the shared storage helper so the path and repo type match CI:
```bash
python -c 'from fastvideo.tests.performance.seed_baseline import upload_prepared_seed_manifest; print(upload_prepared_seed_manifest("<prepared_manifest>"))'
```python
from hf_store import upload_record
upload_record("<local_record_path>", record, strict=True)
```
The uploader verifies the source and prepared-record digests, pins and scans
the current HF revision, rechecks exact-identity and recipe-cohort conflicts,
and writes the entire batch in one commit whose `parent_commit` must still be
current. A concurrent Hub update makes the commit fail. Do not retry
automatically: preserve staging, refresh/review remote state, and request a new
explicit `upload` after the conflict is understood. Each record goes to:
Run it once per prepared record. Each upload goes to:
```text
FastVideo/performance-tracking/<sanitize(model_id)>/<record_filename>.json
```
Never call `upload_record()` once per first-seed record: that can partially
land the batch and has no compare-and-swap guard.
For a legacy reseed or an accepted v2 baseline shift, the first-seed manifest
validator does not apply because an eligible baseline already exists. Upload
only the individually reviewed records prepared in step 5 with the shared
`upload_record(local_path, record, strict=True)` helper. Stop on the first
failure and report exactly which records reached HF; do not silently rerun or
replicate the remainder.
Never bulk upload the tracking or staging root, and never modify another
model's directory in the same operation.
Never bulk upload the whole tracking root. Never modify another model's
directory in the same operation.
### 8. Report outcome and offer cleanup
@@ -546,14 +406,9 @@ distinguish an accepted baseline shift from a hidden regression.
After the upload is verified, ask whether the user wants to clear temporary
local state. Explain what each directory is for:
- `PERFORMANCE_TRACKING_ROOT`, usually `/tmp/perf-tracking`: read-only local
synced mirror used for operator review and reporting. First-v2 preparation
independently proves remote state from a fresh temporary HF snapshot.
- `PERFORMANCE_RESEED_STAGING_ROOT`, usually
`/tmp/performance_reseed_prepared`: prepared local seed records used for the
scoped upload, plus the identity reservation and digest manifest. Keeping
this separate prevents aborted preparations from appearing in later
baseline reads.
- `PERFORMANCE_TRACKING_ROOT`, usually `/tmp/perf-tracking`: local synced
mirror of `FastVideo/performance-tracking` plus the prepared local seed
records used for scoped upload.
- `/tmp/performance_reseed_backup/<...>`: local backup of the target model's
pre-reseed HF history plus `PROVENANCE.txt`, kept so a bad reseed can be
audited or corrected.
@@ -563,18 +418,14 @@ local state. Explain what each directory is for:
Ask:
> Reseed succeeded. Do you want me to delete the local temp tracking mirror,
> this reseed's prepared staging records, source downloads, and reseed backup
> under `/tmp`? These files are local safety/audit artifacts only; HF already
> has the uploaded records.
> source downloads, and reseed backup under `/tmp`? These files are local
> safety/audit artifacts only; HF already has the uploaded records.
>
> Reply `cleanup reseed temp` to delete them, anything else to keep them.
Do not delete anything unless the user replies exactly
`cleanup reseed temp`. If cleanup is requested, remove only the specific
directories and prepared record paths created for this reseed. Do not remove
the shared staging root when it contains other records. Remove this operation's
identity reservation only with its prepared records and manifest, and never
remove unrelated `/tmp` contents.
directories created for this reseed. Never remove unrelated `/tmp` contents.
## Failure modes and handling
@@ -586,34 +437,19 @@ remove unrelated `/tmp` contents.
against the source batch median by more than `max_intra_batch_regression`.
Ask for cleaner sources or a reviewed explanation before continuing.
- **Too few source records to move the median.** Continue only after making
clear that one or two records may not immediately move an existing last-5
median. This warning does not block a first v2 calibration seed for an exact
identity with no eligible baseline yet.
clear that one or two records may not immediately move the last-5 median.
- **The source results are noisy or suspicious.** Stop. Reseeding amplifies
those measurements into the baseline, so they must be reviewed first.
- **HF sync fails.** Stop for destructive reseeds. A stale or empty sync can
make the old baseline look missing.
- **The exact v2 identity already has an eligible baseline.** Stop. The
`CALIBRATION_NEEDED` artifact is stale; use the reviewed baseline-shift path
instead of the first-seed utility.
- **The workload/variant/version trusts another recipe.** Stop. The source is
stale relative to the current recipe cohort and must not bypass
`RECIPE_MISMATCH` by creating a second trusted recipe.
- **The staging root already has a prepared seed for the exact identity.**
Stop and reuse, upload, or explicitly clean that preparation. Do not prepare
another copy of the same measurement.
- **The conditional Hub commit loses its parent race.** Stop without retrying.
Keep the preparation, refresh and review the new remote state, then request
a new explicit `upload` only if the seed is still valid.
- **Candidate still violates fixed thresholds.** Report that this skill only
handles the rolling HF baseline; update benchmark JSON thresholds in code
review if maintainers accept the new absolute limit.
- **The user aborts at either confirmation.** Leave the backup and prepared
records on disk. Nothing should be uploaded.
- **The user declines cleanup.** Keep `/tmp/perf-tracking`, the prepared seed
records under `/tmp/performance_reseed_prepared`, the source download
directory if any, and `/tmp/performance_reseed_backup/<...>` in place for
audit/debugging.
- **The user declines cleanup.** Keep `/tmp/perf-tracking`, the source
download directory if any, and `/tmp/performance_reseed_backup/<...>` in
place for audit/debugging.
- **A bad seed was uploaded.** Use the backup and HF history to identify the
uploaded file, then remove or supersede it with an explicitly reviewed
corrective record. Do not silently rewrite unrelated history.
@@ -624,9 +460,8 @@ remove unrelated `/tmp` contents.
intentional baseline replacement.
- `fastvideo/tests/performance/compare_baseline.py` — normalization, rolling
median comparison, and persistence rules.
- `fastvideo/performance/hf_store.py` — HF sync and record loading helpers.
- `fastvideo/tests/performance/seed_baseline.py` — first-seed preparation,
staging reservation, manifest validation, and conditional batch upload.
- `fastvideo/tests/performance/hf_store.py` — HF sync, record loading,
`sanitize()`, and `upload_record()`.
- `fastvideo/tests/performance/test_inference_performance.py` — source result
JSON schema.
- `.buildkite/performance-benchmarks/tests/*.json` — fixed absolute benchmark
@@ -639,4 +474,3 @@ remove unrelated `/tmp` contents.
| 2026-05-03 | Initial version. Sister workflow to `reseed-ssim-references`, scoped to one performance `(model_id, gpu_type)` baseline seed with backup, confirmation, provenance, and `success=true` upload. |
| 2026-05-03 | Previous policy: replicate one approved shifted source result into 3 success records by default, or 5 only when explicitly requested. Add provenance marker for replicated-source reseeds. Superseded by the 2026-05-08 dynamic multi-source policy. |
| 2026-05-08 | Replace fixed 3/5 replication with dynamic multi-source reseeding: upload one seed record per reviewed source JSON, validate intra-batch consistency, move backup/source scratch under `/tmp`, and ask whether to clean temp state after successful upload. |
| 2026-07-13 | Keep first-v2-seed preparation outside the canonical mirror, reserve staging identities atomically, reject stale or replayed calibration seeds, and upload reviewed manifests with a single parent-guarded Hub commit. |
@@ -1,9 +1,5 @@
{
"benchmark_id": "wan-t2v-1.3b-2gpu",
"config_schema_version": 2,
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp2",
"benchmark_version": 3,
"description": "Wan2.1 T2V 1.3B inference performance",
"model": {
"model_path": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
+3 -43
View File
@@ -1,8 +1,6 @@
env:
IMAGE_VERSION: "py3.12-latest"
BUILDKITE_CLEAN_CHECKOUT: true
# Buildkite only launches Modal; remote jobs initialize their own submodules.
BUILDKITE_GIT_SUBMODULES: false
notify:
- github_commit_status:
@@ -68,17 +66,6 @@ steps:
limit: 2
agents:
queue: "default"
- label: ":vertical_traffic_light: Golden-Gate Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "golden_gate"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Unit Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "unit_test"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
@@ -127,17 +114,6 @@ steps:
limit: 2
agents:
queue: "default"
- label: ":test_tube: LoRA Extraction Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "lora_extraction"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Training Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
@@ -395,27 +371,12 @@ steps:
- TEST_TYPE=inference_lora
agents:
queue: "default"
- path:
- "scripts/lora_extraction/**"
- "fastvideo/tests/lora_extraction/**"
- "fastvideo/models/loader/**"
- "fastvideo/training/training_utils.py"
- "fastvideo/layers/lora/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Extraction Tests"
env:
- TEST_TYPE=lora_extraction
agents:
queue: "default"
- path:
- "fastvideo/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 25m .buildkite/scripts/pr_test.sh"
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Training Tests"
env:
- TEST_TYPE=training
@@ -426,7 +387,7 @@ steps:
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 25m .buildkite/scripts/pr_test.sh"
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Distillation DMD Tests"
env:
- TEST_TYPE=distillation_dmd
@@ -449,7 +410,7 @@ steps:
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 25m .buildkite/scripts/pr_test.sh"
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Training Tests"
env:
- TEST_TYPE=training_lora
@@ -494,7 +455,6 @@ steps:
- "fastvideo/layers/**"
- "fastvideo/worker/**"
- "fastvideo/entrypoints/**"
- "fastvideo/performance/**"
- "fastvideo/tests/performance/**"
- ".buildkite/performance-benchmarks/**"
- "pyproject.toml"
+2 -28
View File
@@ -76,27 +76,10 @@ EFFECTIVE_PR=${BUILDKITE_PULL_REQUEST:-false}
if [ "$EFFECTIVE_PR" = "false" ] && [ -n "${PR_NUMBER:-}" ]; then
EFFECTIVE_PR=$PR_NUMBER
fi
MODAL_ENV="BUILDKITE_REPO=$BUILDKITE_REPO BUILDKITE_COMMIT=$BUILDKITE_COMMIT BUILDKITE_PULL_REQUEST=$EFFECTIVE_PR BUILDKITE_BRANCH=${BUILDKITE_BRANCH:-} BUILDKITE_SOURCE=${BUILDKITE_SOURCE:-} TEST_SCOPE=${TEST_SCOPE:-} BUILDKITE_BUILD_URL=${BUILDKITE_BUILD_URL:-} BUILDKITE_BUILD_ID=${BUILDKITE_BUILD_ID:-} BUILDKITE_JOB_ID=${BUILDKITE_JOB_ID:-} IMAGE_VERSION=$IMAGE_VERSION"
MODAL_ENV="BUILDKITE_REPO=$BUILDKITE_REPO BUILDKITE_COMMIT=$BUILDKITE_COMMIT BUILDKITE_PULL_REQUEST=$EFFECTIVE_PR BUILDKITE_BRANCH=${BUILDKITE_BRANCH:-} TEST_SCOPE=${TEST_SCOPE:-} BUILDKITE_BUILD_URL=${BUILDKITE_BUILD_URL:-} BUILDKITE_BUILD_ID=${BUILDKITE_BUILD_ID:-} BUILDKITE_JOB_ID=${BUILDKITE_JOB_ID:-} IMAGE_VERSION=$IMAGE_VERSION"
POST_RUN_HOOK=""
is_truthy() {
case "${1:-}" in
1|true|TRUE|yes|YES|on|ON) return 0 ;;
*) return 1 ;;
esac
}
ssim_bootstrap_args() {
local title="${PR_TITLE:-}"
local message="${BUILDKITE_MESSAGE:-}"
if is_truthy "${FASTVIDEO_SSIM_BOOTSTRAP_MODE:-}" \
|| [[ "$title" == *"[new-model]"* ]] \
|| [[ "$message" == *"[new-model]"* ]]; then
printf ' --bootstrap-mode'
fi
}
upload_performance_artifacts() {
SHORT_SHA=${BUILDKITE_COMMIT:0:7}
LOCAL_DIR="downloaded_reports"
@@ -187,18 +170,9 @@ case "$TEST_TYPE" in
log "Running transformer tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_transformer_tests"
;;
"golden_gate")
log "Running golden-gate tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_golden_gate_tests"
;;
"ssim")
log "Running SSIM tests..."
SSIM_BOOTSTRAP_ARGS=$(ssim_bootstrap_args)
if [ -n "$SSIM_BOOTSTRAP_ARGS" ]; then
log "SSIM bootstrap mode enabled for new-model reference draft generation"
fi
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run "
MODAL_COMMAND+="$MODAL_SSIM_TEST_FILE::run_ssim_tests$SSIM_BOOTSTRAP_ARGS"
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_SSIM_TEST_FILE::run_ssim_tests"
;;
"training")
log "Running training tests..."
-133
View File
@@ -1,133 +0,0 @@
#!/usr/bin/env bash
# Gate the expensive Buildkite full suite on the cheap GitHub checks.
#
# Polls the workflow runs for the PR head commit and only exits 0 once the
# watched cheap workflows (pre-commit, docs build) have succeeded, so the
# 'ready' label cannot burn ~20 GPU lanes on a head that a cheap check has
# already doomed.
#
# Semantics:
# - watched run completed with a bad conclusion -> exit 1 (fail CLOSED:
# no full suite; the next push re-arms via the 'synchronize' trigger)
# - watched run cancelled -> still pending: the docs
# workflow's repo-global 'pages' concurrency group cancels runs superseded
# by unrelated pushes, so 'cancelled' is not a verdict on this PR
# - watched runs pending -> poll until done
# - docs run absent -> not applicable after a
# short grace period ('Deploy Documentation' is path-filtered on PRs)
# - pre-commit run absent -> keep polling: pre-commit
# is never path-filtered, so its absence is always anomalous
# - 'ready' label removed while waiting -> exit 1 (fail CLOSED:
# un-labeling is a deliberate maintainer action)
# - GitHub API unreachable or timeout -> exit 0 (fail OPEN,
# loud warning: never brick CI on a GitHub outage)
#
# Required env: PR_SHA (PR head commit), PR_NUMBER, GITHUB_REPOSITORY, GH_TOKEN.
set -euo pipefail
: "${PR_SHA:?PR_SHA (PR head commit) is required}"
: "${PR_NUMBER:?PR_NUMBER (pull request number) is required}"
: "${GITHUB_REPOSITORY:?GITHUB_REPOSITORY is required}"
# Workflow-level `name:` values that must be green before the full suite
# may start. "Deploy Documentation" is path-filtered on PRs, so its run may
# legitimately never exist; pre-commit always runs, so it must appear.
WATCHED_NAMES='["pre-commit", "Deploy Documentation"]'
WATCHED_REGEX='^(pre-commit|Deploy Documentation)$'
POLL_SECS="${POLL_SECS:-20}"
GRACE_SECS="${GRACE_SECS:-60}"
MAX_WAIT_SECS="${MAX_WAIT_SECS:-1500}"
# Bound each API call so a hung connection hits the 3-strike fail-open path
# instead of pinning the loop until the job timeout (which would fail closed
# on exactly the GitHub-outage case this script is meant to survive).
if command -v timeout >/dev/null 2>&1; then
gh_api() { timeout 30 gh api "$@"; }
else
gh_api() { gh api "$@"; } # macOS dev boxes; CI always has coreutils timeout
fi
# The workflow checked the label before starting the gate, but the wait can
# last ~25 min: re-check once before any exit 0 and fail closed if 'ready'
# was removed in the meantime. An API error here proceeds (the label was
# present when the gate started; never brick CI on an outage).
recheck_ready_label() {
local pr_json
if pr_json=$(gh_api "repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}" 2>/dev/null); then
if ! jq -e '[.labels[]?.name] | index("ready")' <<<"$pr_json" >/dev/null 2>&1; then
echo "::error::PR #${PR_NUMBER} no longer has the 'ready' label —" \
"NOT triggering the Buildkite full suite. Re-add the label to re-arm."
exit 1
fi
else
echo "::warning::Could not re-check the 'ready' label on PR #${PR_NUMBER}; proceeding (it was present when the gate started)."
fi
}
start=$(date +%s)
api_fails=0
missing=""
while true; do
elapsed=$(( $(date +%s) - start ))
if runs_json=$(gh_api "repos/${GITHUB_REPOSITORY}/actions/runs?head_sha=${PR_SHA}&per_page=100" 2>/dev/null) \
&& state=$(jq --arg re "$WATCHED_REGEX" '
[.workflow_runs[]? | select(.name // "" | test($re))]
| group_by(.name) | map(max_by(.id))
| map({name, status, conclusion})' <<<"$runs_json" 2>/dev/null); then
api_fails=0
echo "t+${elapsed}s watched checks: $(jq -c . <<<"$state")"
failed=$(jq -r '[.[] | select(.status == "completed"
and (.conclusion | IN("success", "skipped", "neutral", "cancelled") | not))]
| map(.name) | join(", ")' <<<"$state")
if [ -n "$failed" ]; then
echo "::error::Cheap check(s) failed on ${PR_SHA}: ${failed}." \
"NOT triggering the Buildkite full suite. Push a fix (the 'ready'" \
"label re-arms on every push), or re-run the failed check and then" \
"re-run this workflow."
exit 1
fi
# 'cancelled' counts as pending: wait for a re-run to reach a real verdict
# (bounded by MAX_WAIT, then the fail-open below).
pending=$(jq '[.[] | select(.status != "completed" or .conclusion == "cancelled")] | length' <<<"$state")
missing=$(jq -r --argjson watched "$WATCHED_NAMES" '($watched - map(.name)) | join(", ")' <<<"$state")
if [ "$pending" -eq 0 ]; then
if [ -z "$missing" ]; then
recheck_ready_label
echo "All watched cheap checks are green — full suite may proceed."
exit 0
fi
case "$missing" in
*pre-commit*)
echo "pre-commit run not found for ${PR_SHA} yet; waiting (pre-commit is never path-filtered, so its absence is anomalous)."
;;
*)
if [ "$elapsed" -ge "$GRACE_SECS" ]; then
recheck_ready_label
echo "::warning::Watched run(s) never appeared for ${PR_SHA}: ${missing} (path-filtered, likely not applicable). Proceeding on the checks that did run."
exit 0
fi
echo "Waiting up to ${GRACE_SECS}s grace for path-filtered run(s) to appear: ${missing}."
;;
esac
fi
else
api_fails=$(( api_fails + 1 ))
echo "::warning::GitHub API error querying workflow runs for ${PR_SHA} (attempt ${api_fails}/3)."
if [ "$api_fails" -ge 3 ]; then
recheck_ready_label
echo "::warning::FAILING OPEN: cannot query GitHub check status — triggering the full suite WITHOUT the cheap-check gate."
exit 0
fi
fi
if [ "$elapsed" -ge "$MAX_WAIT_SECS" ]; then
recheck_ready_label
echo "::warning::FAILING OPEN: watched checks still pending after $(( MAX_WAIT_SECS / 60 )) min${missing:+ (never appeared: ${missing})} — triggering the full suite anyway."
exit 0
fi
sleep "$POLL_SECS"
done
-122
View File
@@ -1,122 +0,0 @@
#!/usr/bin/env bash
# Self-test for gate_full_suite.sh using a mocked `gh`. No network, runs on
# any dev box: bash .github/scripts/test_gate_full_suite.sh
set -u
here=$(cd "$(dirname "$0")" && pwd)
tmp=$(mktemp -d)
trap 'rm -rf "$tmp"' EXIT
# Mock gh. Asserts the exact endpoint (including head_sha) it is called
# with — an endpoint typo in the gate script fails the test rather than
# silently serving canned data. On the runs endpoint it serves
# $MOCK_DIR/response_<call#>.json, sticking on the highest existing file,
# and exits 1 if none exist (simulates a GitHub API outage). On the pulls
# endpoint it serves $MOCK_DIR/pr.json, defaulting to a 'ready'-labeled PR.
cat > "$tmp/gh" <<'EOF'
#!/usr/bin/env bash
if [ "${1:-}" != "api" ]; then
echo "unexpected gh invocation: $*" >> "$MOCK_DIR/endpoint_error"
exit 2
fi
case "${2:-}" in
"repos/o/r/actions/runs?head_sha=deadbeef&per_page=100")
n=$(( $(cat "$MOCK_DIR/count" 2>/dev/null || echo 0) + 1 ))
echo "$n" > "$MOCK_DIR/count"
while [ "$n" -gt 0 ]; do
if [ -f "$MOCK_DIR/response_$n.json" ]; then
cat "$MOCK_DIR/response_$n.json"
exit 0
fi
n=$(( n - 1 ))
done
echo "api outage" >&2
exit 1
;;
"repos/o/r/pulls/42")
if [ -f "$MOCK_DIR/pr.json" ]; then
cat "$MOCK_DIR/pr.json"
else
echo '{"labels": [{"name": "ready"}]}'
fi
;;
*)
echo "unexpected gh endpoint: $2" >> "$MOCK_DIR/endpoint_error"
exit 2
;;
esac
EOF
chmod +x "$tmp/gh"
PC_OK='{"name": "pre-commit", "id": 1, "status": "completed", "conclusion": "success"}'
PC_BAD='{"name": "pre-commit", "id": 1, "status": "completed", "conclusion": "failure"}'
PC_PENDING='{"name": "pre-commit", "id": 1, "status": "in_progress", "conclusion": null}'
DOCS_OK='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "success"}'
DOCS_BAD='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "failure"}'
DOCS_CANCELLED='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "cancelled"}'
OTHER='{"name": "Trigger Full Suite", "id": 3, "status": "in_progress", "conclusion": null}'
NULL_NAME='{"name": null, "id": 4, "status": "completed", "conclusion": "failure"}'
PC_OK_RERUN='{"name": "pre-commit", "id": 5, "status": "completed", "conclusion": "success"}'
fails=0
want_log="" # optional: expect() also greps out.log for this regex, then resets
pr_json="" # optional: served for the pulls (label re-check) endpoint, then resets
raw_body="" # optional: serve responses verbatim instead of wrapping in workflow_runs
expect() { # <name> <expected-exit> <response json>...
local name=$1 want=$2 dir i=1
shift 2
dir=$(mktemp -d "$tmp/test_XXXXXX")
for body in "$@"; do
if [ -n "$raw_body" ]; then
printf '%s' "$body" > "$dir/response_$i.json"
else
printf '{"workflow_runs": [%s]}' "$body" > "$dir/response_$i.json"
fi
i=$(( i + 1 ))
done
[ -n "$pr_json" ] && printf '%s' "$pr_json" > "$dir/pr.json"
( export PATH="$tmp:$PATH" MOCK_DIR="$dir" PR_SHA=deadbeef PR_NUMBER=42 \
GITHUB_REPOSITORY=o/r POLL_SECS=0 GRACE_SECS=1 MAX_WAIT_SECS=3
bash "$here/gate_full_suite.sh" > "$dir/out.log" 2>&1 )
local rc=$?
if [ "$rc" -ne "$want" ]; then
echo "FAIL: $name (exit $rc, want $want)"
cat "$dir/out.log"
fails=1
elif [ -f "$dir/endpoint_error" ]; then
echo "FAIL: $name (mock gh got an unexpected call)"
cat "$dir/endpoint_error"
fails=1
elif [ -n "$want_log" ] && ! grep -Eq "$want_log" "$dir/out.log"; then
echo "FAIL: $name (log does not match: $want_log)"
cat "$dir/out.log"
fails=1
else
echo "ok: $name"
fi
want_log="" pr_json="" raw_body=""
}
expect "both green -> proceed" 0 "$PC_OK, $DOCS_OK, $OTHER, $NULL_NAME"
expect "docs build failed -> blocked" 1 "$PC_OK, $DOCS_BAD"
expect "pre-commit failed -> blocked" 1 "$PC_BAD"
expect "pending then green -> proceed" 0 "$PC_PENDING" "$PC_OK, $DOCS_OK"
want_log="never appeared.*Deploy Documentation"
expect "docs run absent (path-filtered) -> proceed after grace" 0 "$PC_OK"
expect "API outage -> fail open" 0
want_log="FAILING OPEN"
expect "pending past MAX_WAIT -> fail open" 0 "$PC_PENDING"
want_log="FAILING OPEN"
expect "unrelated runs only -> no grace, fail open at MAX_WAIT" 0 "$OTHER"
expect "cancelled docs then green -> proceed" 0 \
"$PC_OK, $DOCS_CANCELLED" "$PC_OK, $DOCS_OK"
want_log="FAILING OPEN"
expect "cancelled docs forever -> fail open at MAX_WAIT" 0 "$PC_OK, $DOCS_CANCELLED"
want_log="FAILING OPEN"
expect "pre-commit absent -> no grace, fail open at MAX_WAIT" 0 "$DOCS_OK"
expect "duplicate run names -> latest wins" 0 "$PC_BAD, $PC_OK_RERUN, $DOCS_OK"
raw_body=1
expect "garbage response body -> fail open" 0 "this is not json"
pr_json='{"labels": [{"name": "other"}]}'
expect "ready label removed mid-gate -> blocked" 1 "$PC_OK, $DOCS_OK"
exit "$fails"
+2 -35
View File
@@ -1,11 +1,7 @@
name: pre-commit
on:
# pull_request_target instead of pull_request: the workflow definition and
# the hook config are always taken from the BASE branch, so fork /
# first-time-contributor PRs run immediately without a maintainer clicking
# "Approve and run". The PR head is checked out as data only.
pull_request_target:
pull_request:
branches: [main]
workflow_call:
inputs:
@@ -19,36 +15,12 @@ permissions:
jobs:
pre-commit:
if: github.event.pull_request.draft != true
if: github.event_name == 'workflow_call' || github.event.pull_request.draft != true
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
ref: ${{ inputs.ref || '' }}
# For PR events, lint the PR head — but keep the hook definitions from
# the base branch so an untrusted PR cannot alter what gets executed.
# The gate scripts are saved too: the self-test step below executes them,
# so it must run the base-branch copies, not the PR head's.
- name: Save trusted hook config and gate scripts
if: github.event_name == 'pull_request_target'
run: |
cp .pre-commit-config.yaml "$RUNNER_TEMP/trusted-pre-commit-config.yaml"
cp -a .github/scripts "$RUNNER_TEMP/trusted-scripts"
echo "GATE_SCRIPTS_DIR=$RUNNER_TEMP/trusted-scripts" >> "$GITHUB_ENV"
# allow-unsafe-pr-checkout acknowledges checkout's pull_request_target
# guard: the head is data for the trusted hooks to lint; nothing from it
# is executed (config and gate scripts are pinned to the base branch
# above) and credentials are not persisted. SHA-pinned to v4.4.0 because
# actionlint's action schema does not know the new input yet.
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
if: github.event_name == 'pull_request_target'
with:
ref: ${{ github.event.pull_request.head.sha }}
persist-credentials: false
allow-unsafe-pr-checkout: true
- name: Restore trusted hook config
if: github.event_name == 'pull_request_target'
run: cp "$RUNNER_TEMP/trusted-pre-commit-config.yaml" .pre-commit-config.yaml
- uses: actions/setup-python@v5
with:
python-version: "3.12"
@@ -58,8 +30,3 @@ jobs:
- uses: pre-commit/action@v3.0.1
with:
extra_args: --all-files --hook-stage manual
# After pre-commit so a self-test failure cannot mask lint failures.
# GATE_SCRIPTS_DIR points at the base-branch copy on fork PRs (set above);
# push / workflow_call runs use the checked-out tree directly.
- name: Full-suite gate self-test
run: bash "${GATE_SCRIPTS_DIR:-.github/scripts}/test_gate_full_suite.sh"
+12 -20
View File
@@ -52,7 +52,6 @@ jobs:
core.setOutput('pr_sha', pr.head.sha);
core.setOutput('pr_branch', pr.head.ref);
core.setOutput('pr_number', String(prNumber));
core.setOutput('pr_title', pr.title);
- name: Trigger Full Suite
if: steps.perm.outputs.has_write == 'true'
@@ -61,7 +60,6 @@ jobs:
PR_SHA: ${{ steps.label.outputs.pr_sha }}
PR_BRANCH: ${{ steps.label.outputs.pr_branch }}
PR_NUMBER: ${{ steps.label.outputs.pr_number }}
PR_TITLE: ${{ steps.label.outputs.pr_title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
@@ -73,7 +71,6 @@ jobs:
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER} (via /merge)" \
--arg pr_title "$PR_TITLE" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
@@ -83,12 +80,11 @@ jobs:
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: "full",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring),
PR_TITLE: $pr_title
}
}')"
TEST_SCOPE: "full",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring)
}
}')"
parse-command:
if: >-
@@ -129,7 +125,7 @@ jobs:
set -euo pipefail
TEST_NAME=$(echo "$COMMENT" | grep -oP '(?<=/test\s)\S+' | head -1 || true)
VALID="encoder vae transformer kernel unit dreamverse ssim golden-gate training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval full fastcheck pre-commit"
VALID="encoder vae transformer kernel unit dreamverse ssim training lora-inference lora-training distillation self-forcing vsa vmoba performance api train-framework eval full fastcheck pre-commit"
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
echo "Unknown test: '$TEST_NAME'. Valid: $VALID"
exit 1
@@ -138,9 +134,8 @@ jobs:
declare -A MAP=(
[encoder]=encoder [vae]=vae [transformer]=transformer
[kernel]=kernel_tests [unit]=unit_test [dreamverse]=dreamverse_app
[ssim]=ssim [golden-gate]=golden_gate [training]=training
[ssim]=ssim [training]=training
[lora-inference]=inference_lora [lora-training]=training_lora
[lora-extraction]=lora_extraction
[distillation]=distillation_dmd [self-forcing]=self_forcing
[vsa]=training_vsa [vmoba]=inference_vmoba
[performance]=performance [api]=api_server
@@ -245,7 +240,6 @@ jobs:
TEST_SCOPE: ${{ needs.parse-command.outputs.test_scope }}
FULL_SUITE: ${{ needs.parse-command.outputs.full_suite }}
TEST_TYPE: ${{ needs.parse-command.outputs.test_type }}
PR_TITLE: ${{ github.event.issue.title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
@@ -262,7 +256,6 @@ jobs:
--arg full_suite "$FULL_SUITE" \
--arg test_type "$TEST_TYPE" \
--arg pr_number "$PR_NUMBER" \
--arg pr_title "$PR_TITLE" \
'{
commit: $commit,
branch: $branch,
@@ -272,9 +265,8 @@ jobs:
pull_request_base_branch: "main",
env: {
TEST_SCOPE: $test_scope,
FULL_SUITE: $full_suite,
TEST_TYPE: $test_type,
PR_NUMBER: $pr_number,
PR_TITLE: $pr_title
}
}')"
FULL_SUITE: $full_suite,
TEST_TYPE: $test_type,
PR_NUMBER: $pr_number
}
}')"
+1 -21
View File
@@ -7,7 +7,6 @@ on:
permissions:
contents: read
pull-requests: read
actions: read
concurrency:
group: full-suite-${{ github.event.pull_request.number }}
@@ -19,8 +18,6 @@ jobs:
(github.event.action == 'labeled' && github.event.label.name == 'ready')
|| github.event.action == 'synchronize'
runs-on: ubuntu-latest
# Gate below may wait for cheap checks (up to MAX_WAIT_SECS = 25 min).
timeout-minutes: 35
steps:
- name: Check ready label
id: check
@@ -52,20 +49,6 @@ jobs:
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds/${build_num}/cancel"
done
# Checks out the BASE branch (default for pull_request_target), so PR
# authors cannot tamper with the gate script.
- name: Checkout gate script
if: steps.check.outputs.has_ready == 'true'
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
- name: Wait for pre-commit and docs build
if: steps.check.outputs.has_ready == 'true'
env:
GH_TOKEN: ${{ github.token }}
PR_SHA: ${{ github.event.pull_request.head.sha }}
PR_NUMBER: ${{ github.event.pull_request.number }}
run: bash .github/scripts/gate_full_suite.sh
- name: Trigger Buildkite Full Suite
if: steps.check.outputs.has_ready == 'true'
env:
@@ -73,7 +56,6 @@ jobs:
PR_SHA: ${{ github.event.pull_request.head.sha }}
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
PR_NUMBER: ${{ github.event.pull_request.number }}
PR_TITLE: ${{ github.event.pull_request.title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
@@ -85,7 +67,6 @@ jobs:
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER}" \
--arg pr_title "$PR_TITLE" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
@@ -97,7 +78,6 @@ jobs:
env: {
TEST_SCOPE: "full",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring),
PR_TITLE: $pr_title
PR_NUMBER: ($pr_id | tostring)
}
}')"
+5 -37
View File
@@ -13,40 +13,12 @@ on:
required: false
default: false
type: boolean
# Auto-rebuild the CUDA images when a repository-controlled image input
# changes on main. This includes the trusted SM89 kernel artifact's source,
# metadata/key helper, ABI dependency metadata, and build orchestration.
# Dreamverse (apps/dreamverse/docker/Dockerfile) and the ROCm Dockerfile stay
# manual-dispatch only.
push:
branches: [main]
paths:
- '.dockerignore'
- '.github/workflows/_template-build-image.yml'
- '.github/workflows/infra-build-image.yml'
- '.gitmodules'
- 'docker/Dockerfile'
- 'docker/uv-excludes'
- 'fastvideo-kernel/**'
- 'fastvideo/tests/modal/kernel_build_cache.py'
- 'pyproject.toml'
permissions:
contents: read
packages: write
# One static group, no cancellation: every run of this workflow writes the same
# mutable registry tags (latest, py3.12-latest, ...), so runs must serialize —
# concurrent push/dispatch runs would race on those tags, and cancelling a run
# mid-publish can strand the cu126/cu130 tag families at different commits. An
# in-flight superseded build wastes its runner time, but its tags are then
# overwritten by the newer queued run. GitHub keeps a single pending run per
# group: the newest queued run replaces any older queued one.
concurrency:
group: infra-build-image
cancel-in-progress: false
jobs:
# CUDA matrix: Python 3.12 x {12.6.3, 13.0.0} x {amd64, arm64}. Each architecture
# builds natively and pushes only by digest; publish-cuda-manifests is the sole
@@ -56,11 +28,7 @@ jobs:
# aliases; 13.0.0/cu130 is published under explicit versioned tags. Flash-attn
# 2.8.3 comes from the architecture-specific prebuilt releases.
build-cuda-images:
# Runs on a manual dispatch when build_cuda_matrix is set, or automatically
# on an in-scope main push (inputs are null on push). The
# repository guard keeps fork syncs from auto-building; manual dispatch
# still works in forks.
if: ${{ (github.event_name == 'push' && github.repository == 'hao-ai-lab/FastVideo') || github.event.inputs.build_cuda_matrix == 'true' }}
if: ${{ github.event.inputs.build_cuda_matrix == 'true' }}
strategy:
fail-fast: false
matrix:
@@ -107,10 +75,10 @@ jobs:
secrets: inherit
publish-cuda-manifests:
# !cancelled(): publish lanes whose digests exist even if a sibling build
# leg failed (the digest-count check fails incomplete lanes); it also
# bypasses skipped-needs propagation, hence the explicit skipped check.
if: ${{ !cancelled() && needs.build-cuda-images.result != 'skipped' }}
# !cancelled(): a failed sibling build leg must not skip the manifests for a
# CUDA lane whose own digests all exist; the digest-count check below fails
# the incomplete lane loudly instead.
if: ${{ !cancelled() && github.event.inputs.build_cuda_matrix == 'true' }}
needs: build-cuda-images
runs-on: ubuntu-latest
permissions:
+2 -6
View File
@@ -6,7 +6,6 @@ results/
wandb/
*.ipynb
*.jpg
!examples/datasets/lingbotworld2/image.jpg
*.safetensors
*.mp4
*.png
@@ -35,7 +34,6 @@ env
*.log
weights/
logs/
/Z-Image/
official_weights/
converted_weights/
@@ -74,8 +72,8 @@ docs/distillation/examples/
# Python pickle files
*.pkl
# Reference videos (negations must come after the catch-all on line below)
!fastvideo/tests/nightly/reference_video_*.mp4
# Reference videos
!fastvideo/tests/ssim/reference_videos/**/*.mp4
# Static images
!docs/assets/images/**/*.png
@@ -129,8 +127,6 @@ apps/dreamverse/web/.env.production.local
.sisyphus/
openspec/
fastvideo/tests/ssim/reference_videos/**
!fastvideo/tests/ssim/reference_videos/**/*.mp4
!fastvideo/tests/ssim/reference_videos/**/*.png
# Editor logs and local Python version pins (accidentally committed)
*.nvimlog
+2
View File
@@ -10,6 +10,8 @@ exclude: |
scripts/.*|
fastvideo/dataset/.*|
fastvideo/models/.*|
v2/(layers|attention|platforms|configs|distributed|models|logging_utils|third_party|hooks|api)/.*|
v2/(envs|logger|utils|version|forward_context|fastvideo_args)\.py|
^apps/dreamverse/web/.*|
examples/.*|
\.agents/.*|
+2 -2
View File
@@ -9,7 +9,7 @@
**FastVideo is a unified post-training and real-time inference framework for accelerated video generation.**
## NEWS
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. See the [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), [Attn-QAT training guide](https://haoailab.com/FastVideo/training/attn_qat/), and [blog](https://haoailab.com/blogs/fastwan-qad/).
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), check out the [Blog](https://haoailab.com/blogs/fastwan-qad/).
- `2026/03/17`: Release demo: Into the Dreamverse: Vibe Directing in FastVideo, check out the [Blog](https://haoailab.com/blogs/dreamverse/).
- `2026/03/13`: Release demo: Create a 5s 1080p Video in 4.5s with FastVideo on a Single GPU, check out the [Blog](https://haoailab.com/blogs/fastvideo_realtime_1080p/).
- `2025/11/19`: Release [CausalWan2.2 I2V A14B Preview](https://huggingface.co/FastVideo/CausalWan2.2-I2V-A14B-Preview-Diffusers) models, [Blog](https://hao-ai-lab.github.io/blogs/fastvideo_causalwan_preview/) and [Inference Code!](https://github.com/hao-ai-lab/FastVideo/blob/main/examples/inference/basic/basic_self_forcing_causal_wan2_2_i2v.py).
@@ -33,7 +33,7 @@ FastVideo has the following features:
- [Sparse distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/) to achieve >50x denoising speedup
- Scalable training with FSDP2, sequence parallelism, and selective activation checkpointing.
- Causal distillation through Self-Forcing
- See this [page](https://hao-ai-lab.github.io/FastVideo/training/overview/) for the supported training workflows, and the [support matrix](https://hao-ai-lab.github.io/FastVideo/inference/support_matrix/) for supported models.
- See this [page](https://hao-ai-lab.github.io/FastVideo/training/overview/) for full list of supported models and recipes.
- State-of-the-art performance optimizations for inference
- Sequence Parallelism for distributed inference
- Multiple state-of-the-art attention backends
-3
View File
@@ -84,12 +84,9 @@ RUN source /opt/venv/bin/activate \
FFMPEG_NATIVE_CXX=/usr/bin/g++ \
bash /opt/FastVideo/apps/dreamverse/scripts/install_native_ffmpeg.sh
# FASTVIDEO_FA4: FA4 (flash_attn.cute) is opt-in; this image installs it via
# the dreamverse extra and is validated with it, so enable it here.
ENV FASTVIDEO_DREAMVERSE_HOME=/var/lib/dreamverse \
STREAM_MODE=av_fmp4 \
FASTVIDEO_ENABLE_PROMPT_SAFETY=0 \
FASTVIDEO_FA4=1 \
HF_HOME=/root/.cache/huggingface
RUN mkdir -p /var/lib/dreamverse
@@ -39,17 +39,6 @@ export FASTVIDEO_GENERATION_SEGMENT_CAP="${FASTVIDEO_GENERATION_SEGMENT_CAP:-6}"
export FASTVIDEO_PROMPT_AUTO_SLEEP_MS="${FASTVIDEO_PROMPT_AUTO_SLEEP_MS:-120}"
export FASTVIDEO_PROMPT_AUTO_TIMEOUT_MS="${FASTVIDEO_PROMPT_AUTO_TIMEOUT_MS:-1800}"
if [[ "${ENABLE_TORCH_COMPILE}" == "1" ]]; then
# Persist Inductor, AOTAutograd, and Triton artifacts across launches.
export DREAMVERSE_TORCH_COMPILE_CACHE_ROOT="${DREAMVERSE_TORCH_COMPILE_CACHE_ROOT:-${HOME}/.cache/dreamverse/torch_compile}"
export TORCHINDUCTOR_CACHE_DIR="${TORCHINDUCTOR_CACHE_DIR:-${DREAMVERSE_TORCH_COMPILE_CACHE_ROOT}/inductor}"
export TRITON_CACHE_DIR="${TRITON_CACHE_DIR:-${DREAMVERSE_TORCH_COMPILE_CACHE_ROOT}/triton}"
export TORCHINDUCTOR_FX_GRAPH_CACHE="${TORCHINDUCTOR_FX_GRAPH_CACHE:-1}"
export TORCHINDUCTOR_AUTOGRAD_CACHE="${TORCHINDUCTOR_AUTOGRAD_CACHE:-1}"
mkdir -p "${TORCHINDUCTOR_CACHE_DIR}" "${TRITON_CACHE_DIR}"
echo "[launch-demo] torch.compile cache: ${DREAMVERSE_TORCH_COMPILE_CACHE_ROOT}"
fi
cd "${DREAMVERSE_ROOT}"
if ! command -v dreamverse-server >/dev/null 2>&1; then
+1 -1
View File
@@ -1 +1 @@
NEXT_PUBLIC_API_BASE_URL=http://localhost:8189/api
PUBLIC_API_BASE_URL=http://localhost:8189/api
+1 -6
View File
@@ -13,12 +13,6 @@
# testing
/coverage
# playwright
/test-results/
/playwright-report/
/blob-report/
/.playwright/
# next.js
/.next/
/out/
@@ -45,3 +39,4 @@ yarn-error.log*
# typescript
*.tsbuildinfo
next-env.d.ts
.svelte-kit
+4
View File
@@ -0,0 +1,4 @@
{
"useTabs": true,
"tabWidth": 4
}
+5 -49
View File
@@ -11,63 +11,19 @@ The UI currently supports:
- Datasets
- Gallery View
## Project Structure
```
apps/fastvideo_studio/
├── server.py / job_runner.py / database.py # FastAPI backend + job lifecycle
├── mock_server.py # In-memory API mock for e2e tests
├── models/ # Pydantic request models (shared with the mock)
├── training_config.py # Studio workloads → fastvideo/train YAML configs
├── tests/ # Backend unit tests (pytest)
├── e2e/ # Playwright specs (run against the mock)
└── src/
├── app/ # Next.js App Router pages (thin routes)
├── components/
│ ├── shell/ # App chrome: header, sidebars, layout
│ ├── jobs/ # Job queue, cards, create-job modal, log sidebar
│ ├── datasets/ # Dataset cards, upload, captions
│ └── ui/ # shadcn-style primitives (shared theme)
├── stores/ # Framework-agnostic state + React bridge (hooks/)
├── lib/ # API client, types, option persistence
└── test/ # Vitest setup + factories
```
The visual theme (slate light/dark palettes, IBM Plex type, `#356cff` accent)
is shared with `apps/dreamverse`; the toggle in the header persists the choice
per browser.
## Testing
```bash
npm run typecheck # tsc
npm test # vitest unit tests
npm run e2e # Playwright against the in-memory mock backend
python -m pytest tests/ # backend unit tests (from apps/fastvideo_studio)
```
## Quick Start
For local development, install dependencies and start the Next.js dev server:
```bash
cd apps/fastvideo_studio
npm install
npm run dev
```
You can then access the app at [http://localhost:3000](http://localhost:3000).
Start the Python API server (default port 8189) in a separate terminal — see below.
For a production build of the web app together with the API server:
The easiest way to run the application is:
```bash
cd apps/fastvideo_studio
npm install
npm run build
npm run start:all
npm run start
```
You can then access the app at [http://localhost:3000](http://localhost:3000).
### Running API and Web Separately
The UI is composed of two separate components:
@@ -75,7 +31,7 @@ The UI is composed of two separate components:
- API Server
- Web Server
To run each component separately, you can use the commands `npm run start:web` (after `npm run build`) and `npm run start:api`.
To run each component separately, you can use the commands `npm run start:web` and `npm run start:api`.
You can also run the API server with this command from the `apps/` directory:
+28 -4
View File
@@ -251,10 +251,11 @@ class Database:
sp_size, negative_prompt,
data_path, max_train_steps, train_batch_size, learning_rate,
num_latent_t, validation_dataset_file, lora_rank,
ltx2_first_frame_conditioning_p,
dmd_use_vsa, dmd_vsa_sparsity, dmd_denoising_steps,
real_score_guidance_scale,
min_timestep_ratio, max_timestep_ratio, real_score_guidance_scale,
generator_update_interval, real_score_model_path, fake_score_model_path
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
job["id"],
@@ -296,9 +297,12 @@ class Database:
job.get("num_latent_t", 20),
job.get("validation_dataset_file", ""),
job.get("lora_rank", 32),
job.get("ltx2_first_frame_conditioning_p"),
1 if job.get("dmd_use_vsa") else 0,
job.get("dmd_vsa_sparsity", 0.8),
job.get("dmd_denoising_steps", "1000,757,522"),
job.get("min_timestep_ratio", 0.02),
job.get("max_timestep_ratio", 0.98),
job.get("real_score_guidance_scale", 3.5),
job.get("generator_update_interval", 5),
job.get("real_score_model_path", ""),
@@ -397,6 +401,24 @@ class Database:
for file_name, caption in captions.items():
self.upsert_dataset_caption(dataset_id, file_name, caption)
def update_dataset(self, dataset_id: str, updates: dict[str, Any]) -> None:
"""Update dataset fields. Only provided keys are updated."""
if not updates:
return
allowed: set[str] = set()
cols = []
vals = []
for k, v in updates.items():
if k in allowed:
cols.append(f"{k} = ?")
vals.append(v)
if not cols:
return
vals.append(dataset_id)
sql = f"UPDATE datasets SET {', '.join(cols)} WHERE id = ?"
self._execute(sql, tuple(vals))
self._commit()
def delete_dataset(self, dataset_id: str) -> bool:
"""Delete a dataset. Returns True if a row was deleted."""
cur = self._execute("DELETE FROM datasets WHERE id = ?", (dataset_id, ))
@@ -421,7 +443,7 @@ class Database:
cur = self._execute("SELECT * FROM settings WHERE id = 1")
row = cur.fetchone()
if not row:
return default_settings_dict()
return _default_settings_dict()
t2v = ((row["default_model_id_t2v"] or row["default_model_id"] or "") if "default_model_id_t2v" in row else
(row["default_model_id"] or ""))
i2v = ((row["default_model_id_i2v"] or "") if "default_model_id_i2v" in row else "")
@@ -535,6 +557,8 @@ def _row_to_job(row: sqlite3.Row) -> dict[str, Any]:
}
float_defaults = {
"dmd_vsa_sparsity": 0.8,
"min_timestep_ratio": 0.02,
"max_timestep_ratio": 0.98,
"real_score_guidance_scale": 3.5,
}
result = {
@@ -601,7 +625,7 @@ def _row_to_job(row: sqlite3.Row) -> dict[str, Any]:
return result
def default_settings_dict() -> dict[str, Any]:
def _default_settings_dict() -> dict[str, Any]:
"""Return default settings as API-style dict."""
return {
"defaultModelId": DEFAULT_SETTINGS["default_model_id"],
@@ -1,42 +0,0 @@
import { expect, test } from '@playwright/test';
import { skipWithoutMock } from './helpers';
/**
* Create-job flow: open the Create Job modal on /inference, fill the prompt
* (the model auto-selects once the mock's /api/models loads), submit, and
* confirm the new job lands in the queue.
*/
test.describe('create inference job', () => {
skipWithoutMock();
test('creates a T2V job and shows it in the queue', async ({ page }) => {
await page.goto('/inference');
// The trigger opens a real menu on click, so this path works for touch,
// mouse, and keyboard users.
await page.getByRole('button', { name: /create job/i }).click();
const t2vItem = page.getByRole('menuitem', { name: /T2V/i });
await expect(t2vItem).toBeVisible();
await t2vItem.click();
const dialog = page.getByRole('dialog');
await expect(dialog).toBeVisible();
// Wait for the mock's model catalogue to populate the dropdown (more than
// just the disabled placeholder), then pick one explicitly — the app's
// auto-selection is racy.
const modelSelect = dialog.getByLabel('Model', { exact: true });
await expect(modelSelect.locator('option')).not.toHaveCount(1);
await modelSelect.selectOption({ index: 1 });
const prompt = `e2e raccoon in sunflowers ${Date.now()}`;
await dialog.getByLabel('Prompt', { exact: true }).fill(prompt);
await dialog.getByRole('button', { name: 'Create Job' }).click();
// Modal closes and the queue refreshes with the newly created job.
await expect(dialog).toBeHidden();
await expect(page.getByText(prompt)).toBeVisible();
});
});
@@ -1,37 +0,0 @@
import { expect, test } from '@playwright/test';
import { API_BASE, skipWithoutMock } from './helpers';
/**
* Datasets page: the seeded datasets render as cards, and the Create Dataset
* modal opens from the header action.
*/
test.describe('datasets', () => {
skipWithoutMock();
test('lists the seeded datasets', async ({ page, request }) => {
// Read the seeded names from the mock so the assertion tracks the fixture.
const res = await request.get(`${API_BASE}/datasets`);
const datasets = (await res.json()) as Array<{ name: string }>;
expect(datasets.length).toBeGreaterThan(0);
await page.goto('/datasets');
for (const ds of datasets) {
await expect(page.getByText(ds.name, { exact: true })).toBeVisible();
}
});
test('opens the Create Dataset modal', async ({ page }) => {
await page.goto('/datasets');
await page.getByRole('button', { name: 'Add Dataset' }).click();
const dialog = page.getByRole('dialog');
await expect(dialog).toBeVisible();
await expect(
dialog.getByRole('heading', { name: /Add Dataset/i }),
).toBeVisible();
await expect(dialog.getByLabel('Name', { exact: true })).toBeVisible();
});
});
-44
View File
@@ -1,44 +0,0 @@
import { expect, test } from '@playwright/test';
import { API_BASE, skipWithoutMock } from './helpers';
/**
* Gallery page: the seeded completed inference job surfaces as a media tile
* with playback controls or an explicit media-error fallback.
*/
test.describe('gallery', () => {
skipWithoutMock();
test('shows a media tile for the seeded completed job', async ({
page,
request,
}) => {
const res = await request.get(`${API_BASE}/jobs?job_type=inference`);
const jobs = (await res.json()) as Array<{
status: string;
output_path: string | null;
prompt: string;
}>;
const completed = jobs.find(
(j) => j.status === 'completed' && j.output_path,
);
expect(completed, 'mock should seed a completed inference job').toBeTruthy();
await page.goto('/gallery');
await expect(
page.getByRole('heading', { level: 1, name: 'Gallery' }),
).toBeVisible();
const tile = page.locator('article').filter({ hasText: completed!.prompt });
await expect(tile).toBeVisible();
await expect(
tile.locator('video').or(tile.getByText('Preview unavailable')),
).toBeVisible();
const video = tile.locator('video');
if (await video.isVisible()) {
await expect(video).toHaveAttribute('controls', '');
}
});
});
@@ -1,39 +0,0 @@
import { expect, test } from '@playwright/test';
import { skipWithoutMock } from './helpers';
/**
* Warm-model slot: load a model through the mock (which flips
* loading -> ready after ~2s, surfaced by the panel's 5s poll), then unload it.
*/
test.describe('generators', () => {
skipWithoutMock();
test('loads and unloads the resident model', async ({ page }) => {
await page.goto('/inference');
const panel = page.getByRole('region', { name: 'Warm models' });
await expect(panel).toBeVisible();
await expect(panel.getByText('No model loaded')).toBeVisible();
await panel
.getByLabel('Model to load')
.selectOption('Wan-AI/Wan2.1-T2V-1.3B-Diffusers');
await panel.getByRole('button', { name: 'Load model' }).click();
await expect(panel.getByText('Wan2.1 T2V 1.3B Diffusers')).toBeVisible();
await expect(panel.getByText('ready')).toBeVisible();
await panel.getByRole('button', { name: 'Unload' }).click();
await expect(panel.getByText('No model loaded')).toBeVisible();
});
test('engine console streams output while open', async ({ page }) => {
await page.goto('/inference');
const engineConsole = page.getByRole('region', { name: 'Engine output' });
await engineConsole.getByRole('button', { name: 'Engine output' }).click();
await expect(engineConsole.getByText(/\[engine\]/).first()).toBeVisible();
});
});
-30
View File
@@ -1,30 +0,0 @@
import { expect, test } from '@playwright/test';
import { skipWithoutMock } from './helpers';
/**
* GPU status page: the mock's fake GPUs render as cards with meters, and the
* page is reachable from the primary sidebar.
*/
test.describe('gpus', () => {
skipWithoutMock();
test('lists the mock GPUs with utilization meters', async ({ page }) => {
await page.goto('/gpus');
await expect(
page.getByRole('heading', { level: 1, name: 'GPUs' }),
).toBeVisible();
await expect(page.getByText('GPU 0')).toBeVisible();
await expect(page.getByText('GPU 1')).toBeVisible();
expect(await page.getByRole('meter').count()).toBe(4);
});
test('is reachable from the sidebar', async ({ page }) => {
await page.goto('/inference');
await page.getByRole('link', { name: 'GPUs' }).click();
await expect(page).toHaveURL(/\/gpus$/);
await expect(page.getByText('GPU 0')).toBeVisible();
});
});
-44
View File
@@ -1,44 +0,0 @@
import { test, type APIRequestContext } from '@playwright/test';
/**
* Port and base URL of the mock API — the single source of truth shared by
* playwright.config.ts (which boots the mock and points `npm run dev` at it
* via NEXT_PUBLIC_API_BASE_URL) and the specs (which probe/read mock state
* directly through the Playwright `request` fixture).
*
* The default port is deliberately NOT the real server's default (8189): the
* e2e specs create and (if autostart is on) launch jobs, so running them
* against a real backend would mutate the developer's database and fire real
* GPU runs. API_BASE is derived solely from MOCK_API_PORT (it does not inherit
* NEXT_PUBLIC_API_BASE_URL) so the specs always talk to the mock.
*/
export const MOCK_API_PORT = process.env.MOCK_API_PORT || '8190';
export const API_BASE = `http://127.0.0.1:${MOCK_API_PORT}/api`;
export const MOCK_SKIP_MESSAGE =
`Mock backend not reachable at ${API_BASE}/__mock__ ` +
`(start fastvideo_studio.mock_server on port ${MOCK_API_PORT}).`;
/**
* True only when the mock server answers on the configured port. Probes the
* mock-only /api/__mock__ sentinel (which the real server does not serve) so
* the suite refuses to run its mutating specs against a real backend.
*/
export async function mockIsUp(request: APIRequestContext): Promise<boolean> {
try {
const res = await request.get(`${API_BASE}/__mock__`);
if (!res.ok()) return false;
const body = await res.json();
return body?.mock === true;
} catch {
return false;
}
}
/** Self-skip every test in the enclosing describe when the mock is down. */
export function skipWithoutMock(): void {
test.beforeEach(async ({ request }) => {
test.skip(!(await mockIsUp(request)), MOCK_SKIP_MESSAGE);
});
}
-115
View File
@@ -1,115 +0,0 @@
import { expect, test } from '@playwright/test';
import { skipWithoutMock } from './helpers';
/**
* App-shell smoke: the Next.js frontend hydrates, renders the FastVideo logo
* and the primary-sidebar navigation, and routing between the top-level
* sections works. Each spec self-skips when the mock backend isn't reachable.
*/
test.describe('app shell', () => {
skipWithoutMock();
test('loads with the logo and primary-sidebar nav', async ({ page }) => {
await page.goto('/');
// The root route redirects to /inference.
await expect(page).toHaveURL(/\/inference$/);
await expect(page.getByRole('img', { name: /fastvideo/i })).toBeVisible();
for (const label of ['Inference', 'Datasets', 'Gallery', 'Settings']) {
await expect(page.getByRole('link', { name: label })).toBeVisible();
}
});
test('navigates between the primary sections', async ({ page }) => {
await page.goto('/inference');
await expect(
page.getByRole('heading', { level: 1, name: 'Jobs' }),
).toBeVisible();
const sections: Array<{ link: string; url: RegExp; title: string }> = [
{ link: 'Datasets', url: /\/datasets$/, title: 'Datasets' },
{ link: 'Gallery', url: /\/gallery$/, title: 'Gallery' },
{ link: 'Settings', url: /\/settings$/, title: 'Settings' },
{ link: 'Inference', url: /\/inference$/, title: 'Jobs' },
];
for (const section of sections) {
await page.getByRole('link', { name: section.link }).click();
await expect(page).toHaveURL(section.url);
// The header <h1> is the only level-1 heading and reflects the route.
await expect(
page.getByRole('heading', { level: 1, name: section.title }),
).toBeVisible();
await expect(page.getByRole('main')).toHaveCount(1);
}
});
test('keeps navigation and content usable at responsive breakpoints', async ({
page,
}) => {
for (const width of [320, 375, 414, 768]) {
await page.setViewportSize({ width, height: 800 });
await page.goto('/inference');
const main = page.getByRole('main');
await expect(main).toBeVisible();
await expect(
page.getByRole('button', { name: /Create Job/i }),
).toBeVisible();
const initialBox = await main.boundingBox();
expect(initialBox?.x).toBe(width < 768 ? 0 : 220);
expect(initialBox?.width).toBe(width < 768 ? width : width - 220);
const navigation = page.getByRole('navigation', {
name: 'Primary navigation',
});
if (width < 768) {
await expect(
page.getByRole('button', { name: 'Open navigation' }),
).toBeVisible();
await page.getByRole('button', { name: 'Open navigation' }).click();
}
await expect(navigation).toBeVisible();
await navigation.getByRole('link', { name: 'Datasets' }).click();
await expect(page).toHaveURL(/\/datasets$/);
expect(
await page.evaluate(
() => document.documentElement.scrollWidth <= window.innerWidth,
),
).toBe(true);
}
});
test('uses full-width detail drawers on mobile', async ({ page }) => {
await page.setViewportSize({ width: 320, height: 800 });
await page.goto('/inference');
await page
.locator('article button[aria-pressed="false"]')
.first()
.click();
const jobDrawer = page.getByRole('dialog', { name: 'Job details' });
await expect(jobDrawer).toBeVisible();
expect(await jobDrawer.boundingBox()).toMatchObject({ x: 0, width: 320 });
await jobDrawer.getByRole('button', { name: 'Close' }).click();
await page.goto('/datasets');
await page
.locator('article button[aria-pressed="false"]')
.first()
.click();
const datasetDrawer = page.getByRole('dialog', {
name: /dataset details$/,
});
await expect(datasetDrawer).toBeVisible();
expect(await datasetDrawer.boundingBox()).toMatchObject({
x: 0,
width: 320,
});
});
});
+2
View File
@@ -0,0 +1,2 @@
// Minimal ESLint config for SvelteKit (add typescript-eslint / eslint-plugin-svelte as needed)
export default [];
-161
View File
@@ -1,161 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
"""GPU telemetry for the studio status page, via NVML (nvidia-ml-py)."""
from __future__ import annotations
import contextlib
import logging
from typing import Any
logger = logging.getLogger("fastvideo.studio.gpu")
_nvml_initialized = False
def _ensure_nvml() -> Any:
"""Import and initialize NVML once; raises on machines without it."""
global _nvml_initialized # noqa: PLW0603
import pynvml
if not _nvml_initialized:
pynvml.nvmlInit()
_nvml_initialized = True
return pynvml
def _device_snapshot(pynvml: Any, index: int) -> dict[str, Any]:
handle = pynvml.nvmlDeviceGetHandleByIndex(index)
name = pynvml.nvmlDeviceGetName(handle)
if isinstance(name, bytes):
name = name.decode()
mem = pynvml.nvmlDeviceGetMemoryInfo(handle)
util = pynvml.nvmlDeviceGetUtilizationRates(handle)
# Optional sensors: not every GPU/driver exposes them.
temperature: int | None = None
power_watts: float | None = None
power_limit_watts: float | None = None
with contextlib.suppress(pynvml.NVMLError):
temperature = int(pynvml.nvmlDeviceGetTemperature(handle, pynvml.NVML_TEMPERATURE_GPU))
with contextlib.suppress(pynvml.NVMLError):
power_watts = pynvml.nvmlDeviceGetPowerUsage(handle) / 1000.0
power_limit_watts = (pynvml.nvmlDeviceGetEnforcedPowerLimit(handle) / 1000.0)
return {
"index": index,
"name": name,
"utilization": int(util.gpu),
"memory_used_mib": int(mem.used / (1024 * 1024)),
"memory_total_mib": int(mem.total / (1024 * 1024)),
"temperature_c": temperature,
"power_watts": power_watts,
"power_limit_watts": power_limit_watts,
}
def get_gpu_snapshot() -> dict[str, Any]:
"""Return {"available", "gpus", "error"} — never raises.
``available: False`` covers both "no NVIDIA driver/library on this host"
and transient NVML failures; the frontend shows ``error`` as-is.
"""
try:
pynvml = _ensure_nvml()
count = pynvml.nvmlDeviceGetCount()
gpus = [_device_snapshot(pynvml, i) for i in range(count)]
return {"available": True, "gpus": gpus, "error": None}
except ImportError:
return {
"available": False,
"gpus": [],
"error": "nvidia-ml-py is not installed on the API server host.",
}
except Exception as exc: # NVMLError, driver issues, …
logger.warning("GPU snapshot failed: %s", exc)
return {"available": False, "gpus": [], "error": str(exc)}
def _remote_gpu_probe() -> dict[str, Any]:
"""Self-contained per-node NVML probe (runs as a ray task on each node;
no fastvideo_studio import — worker environments don't have apps/ on
their path, so cloudpickle must carry this by value)."""
import contextlib as _ctx
import socket as _socket
out: dict[str, Any] = {"hostname": _socket.gethostname(), "available": False, "gpus": [], "error": None}
try:
import pynvml
pynvml.nvmlInit()
for i in range(pynvml.nvmlDeviceGetCount()):
h = pynvml.nvmlDeviceGetHandleByIndex(i)
name = pynvml.nvmlDeviceGetName(h)
if isinstance(name, bytes):
name = name.decode()
mem = pynvml.nvmlDeviceGetMemoryInfo(h)
util = pynvml.nvmlDeviceGetUtilizationRates(h)
temp = power = plimit = None
with _ctx.suppress(pynvml.NVMLError):
temp = int(pynvml.nvmlDeviceGetTemperature(h, pynvml.NVML_TEMPERATURE_GPU))
with _ctx.suppress(pynvml.NVMLError):
power = pynvml.nvmlDeviceGetPowerUsage(h) / 1000.0
plimit = pynvml.nvmlDeviceGetEnforcedPowerLimit(h) / 1000.0
out["gpus"].append({
"index": i,
"name": name,
"utilization": int(util.gpu),
"memory_used_mib": int(mem.used / (1024 * 1024)),
"memory_total_mib": int(mem.total / (1024 * 1024)),
"temperature_c": temp,
"power_watts": power,
"power_limit_watts": plimit,
})
out["available"] = True
except Exception as exc: # noqa: BLE001 -- reported per node
out["error"] = str(exc)
return out
def get_cluster_snapshot() -> dict[str, Any]:
"""Cluster-wide GPU/host telemetry.
When this process is connected to a ray cluster (a model has been
loaded), probe every alive node via per-node ray tasks. Otherwise fall
back to this host's NVML snapshot. Never raises.
"""
import socket
local = get_gpu_snapshot()
local_node = {"hostname": socket.gethostname(), "ip": None, "is_this_host": True,
"cpus": None, "ray_gpus": None, **local}
out: dict[str, Any] = {"mode": "local", "nodes": [local_node],
"resources": None, "error": None}
try:
import ray
if not ray.is_initialized():
out["error"] = "not connected to a ray cluster yet (load a model first); showing the API host only"
return out
from ray.util.scheduling_strategies import NodeAffinitySchedulingStrategy
alive = [n for n in ray.nodes() if n.get("Alive")]
probe = ray.remote(num_cpus=0)(_remote_gpu_probe)
refs = [probe.options(scheduling_strategy=NodeAffinitySchedulingStrategy(
node_id=n["NodeID"], soft=True)).remote() for n in alive]
snaps = ray.get(refs, timeout=15)
nodes = []
for n, snap in zip(alive, snaps, strict=True):
nodes.append({
"ip": n.get("NodeManagerAddress"),
"is_this_host": snap.get("hostname") == socket.gethostname(),
"cpus": n.get("Resources", {}).get("CPU"),
"ray_gpus": n.get("Resources", {}).get("GPU"),
**snap,
})
out["mode"] = "ray"
out["nodes"] = nodes
out["resources"] = {
"gpus_total": ray.cluster_resources().get("GPU", 0.0),
"gpus_available": ray.available_resources().get("GPU", 0.0),
}
except Exception as exc: # noqa: BLE001 -- degrade to the local view
logger.warning("cluster snapshot failed: %s", exc)
out["mode"] = "local"
out["nodes"] = [local_node]
out["error"] = f"cluster probe failed: {exc}"
return out
+119 -315
View File
@@ -23,13 +23,12 @@ import time
from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
import yaml
from fastvideo.utils import get_mp_context
from fastvideo_studio.database import Database
from fastvideo_studio.training_config import (
build_training_config,
build_training_args,
get_training_env,
get_training_module_info,
)
logger = logging.getLogger("fastvideo.studio.job_runner")
@@ -41,11 +40,6 @@ _TQDM_FRAC_RE = re.compile(r"\b(\d+)/(\d+)\b")
_MAX_LOG_LINES = 2000 # ring-buffer cap per job
# ray's log relay prefixes worker lines like "(RayWorkerWrapper pid=123, ip=…)"
# — usually wrapped in ANSI color codes, which must be stripped before matching.
_RAY_RELAY_RE = re.compile(r"^\(\w+ pid=")
_ANSI_RE = re.compile(r"\x1b\[[0-9;]*m")
class JobStatus(str, enum.Enum):
PENDING = "pending"
@@ -166,10 +160,13 @@ class Job:
num_latent_t: int = 20
validation_dataset_file: str = ""
lora_rank: int = 32
ltx2_first_frame_conditioning_p: float | None = None
# DMD options
dmd_use_vsa: bool = False
dmd_vsa_sparsity: float = 0.8
dmd_denoising_steps: str = "1000,757,522"
min_timestep_ratio: float = 0.02
max_timestep_ratio: float = 0.98
real_score_guidance_scale: float = 3.5
generator_update_interval: int = 5
real_score_model_path: str = ""
@@ -224,9 +221,12 @@ class Job:
"num_width": self.width,
"validation_dataset_file": self.validation_dataset_file,
"lora_rank": self.lora_rank,
"ltx2_first_frame_conditioning_p": self.ltx2_first_frame_conditioning_p,
"dmd_use_vsa": self.dmd_use_vsa,
"dmd_vsa_sparsity": self.dmd_vsa_sparsity,
"dmd_denoising_steps": self.dmd_denoising_steps,
"min_timestep_ratio": self.min_timestep_ratio,
"max_timestep_ratio": self.max_timestep_ratio,
"real_score_guidance_scale": self.real_score_guidance_scale,
"generator_update_interval": self.generator_update_interval,
"real_score_model_path": self.real_score_model_path or "",
@@ -264,36 +264,13 @@ class JobRunner:
self._jobs_lock = threading.Lock()
self._load_jobs()
# Exactly one generator lives in memory at a time. Loading a new
# config always releases the old instance first (shutdown + placement
# group teardown); unload deletes it outright.
self._generator: Any | None = None
self._generator_config: dict[str, Any] | None = None
self._generator_state: str = "empty" # empty | loading | ready | failed
self._generator_error: str | None = None
self._generator_lock = threading.Lock() # guards the fields above
# Serializes every slot transition (preload, job-triggered replace,
# unload). Held for the full duration of a load.
self._load_lock = threading.Lock()
# All generator creations run on this ONE persistent thread: the mp
# executor arms prctl(PR_SET_PDEATHSIG, SIGKILL) in its workers, and
# on Linux that fires when the CREATING THREAD exits — a generator
# spawned from a short-lived thread loses all its workers (silent
# SIGKILL, zombies) the moment that thread finishes.
import queue as _queue
self._loader_queue: _queue.Queue = _queue.Queue()
threading.Thread(target=self._loader_loop, daemon=True,
name="generator-loader").start()
# The inference job currently generating, fed by the engine log tee
# (ray relays worker output to the driver; tqdm lines land there).
self._active_inference_job: Job | None = None
# Cache of loaded generators keyed by model config so that we only pay
# the model-loading cost once per model configuration.
self._generators: dict[tuple, Any] = {}
self._generators_lock = threading.Lock()
# Shared Manager for log queues (avoids spawning a new process per job)
self._mp_manager = get_mp_context().Manager()
# One queue for the generator's whole life: mp workers get it at spawn
# (creation-time attach). Sending a Manager proxy over the executor's
# worker pipes post-hoc (set_log_queue RPC) breaks the pipe.
self._worker_log_queue = self._mp_manager.Queue()
atexit.register(self._shutdown)
# Ensure directories exist
@@ -353,9 +330,12 @@ class JobRunner:
num_latent_t=row.get("num_latent_t", 20),
validation_dataset_file=row.get("validation_dataset_file", "") or "",
lora_rank=row.get("lora_rank", 32),
ltx2_first_frame_conditioning_p=row.get("ltx2_first_frame_conditioning_p"),
dmd_use_vsa=row.get("dmd_use_vsa", False),
dmd_vsa_sparsity=float(row.get("dmd_vsa_sparsity", 0.8)),
dmd_denoising_steps=row.get("dmd_denoising_steps", "1000,757,522") or "1000,757,522",
min_timestep_ratio=float(row.get("min_timestep_ratio", 0.02)),
max_timestep_ratio=float(row.get("max_timestep_ratio", 0.98)),
real_score_guidance_scale=float(row.get("real_score_guidance_scale", 3.5)),
generator_update_interval=int(row.get("generator_update_interval", 5)),
real_score_model_path=row.get("real_score_model_path", "") or "",
@@ -432,9 +412,12 @@ class JobRunner:
num_latent_t: int = 20,
validation_dataset_file: str = "",
lora_rank: int = 32,
ltx2_first_frame_conditioning_p: float | None = None,
dmd_use_vsa: bool = False,
dmd_vsa_sparsity: float = 0.8,
dmd_denoising_steps: str = "1000,757,522",
min_timestep_ratio: float = 0.02,
max_timestep_ratio: float = 0.98,
real_score_guidance_scale: float = 3.5,
generator_update_interval: int = 5,
real_score_model_path: str = "",
@@ -474,9 +457,12 @@ class JobRunner:
num_latent_t=num_latent_t,
validation_dataset_file=validation_dataset_file or "",
lora_rank=lora_rank,
ltx2_first_frame_conditioning_p=ltx2_first_frame_conditioning_p,
dmd_use_vsa=dmd_use_vsa,
dmd_vsa_sparsity=dmd_vsa_sparsity,
dmd_denoising_steps=dmd_denoising_steps,
min_timestep_ratio=min_timestep_ratio,
max_timestep_ratio=max_timestep_ratio,
real_score_guidance_scale=real_score_guidance_scale,
generator_update_interval=generator_update_interval,
real_score_model_path=real_score_model_path or "",
@@ -645,211 +631,6 @@ class JobRunner:
"phase": job._log_buf.phase,
}
@staticmethod
def _generator_config_dict(
model_id: str,
workload_type: str,
num_gpus: int,
dit_cpu_offload: bool = False,
text_encoder_cpu_offload: bool = False,
vae_cpu_offload: bool = False,
image_encoder_cpu_offload: bool = False,
use_fsdp_inference: bool = False,
enable_torch_compile: bool = False,
vsa_sparsity: float = 0.0,
tp_size: int = -1,
sp_size: int = -1,
) -> dict[str, Any]:
"""Canonical engine-config dict; equality here == same generator."""
return {
"model_id": model_id,
"workload_type": workload_type,
"num_gpus": num_gpus,
"dit_cpu_offload": dit_cpu_offload,
"text_encoder_cpu_offload": text_encoder_cpu_offload,
"vae_cpu_offload": vae_cpu_offload,
"image_encoder_cpu_offload": image_encoder_cpu_offload,
"use_fsdp_inference": use_fsdp_inference,
"enable_torch_compile": enable_torch_compile,
"vsa_sparsity": vsa_sparsity,
"tp_size": tp_size,
"sp_size": sp_size,
}
def _slot_entry(self) -> dict[str, Any]:
return {
"state": self._generator_state,
"error": self._generator_error,
**(self._generator_config or {}),
}
def _running_inference_jobs(self) -> list[str]:
with self._jobs_lock:
return [j.id for j in self._jobs.values()
if j.status == JobStatus.RUNNING and j.job_type == "inference"]
def _loader_loop(self) -> None:
while True:
fn = self._loader_queue.get()
try:
fn()
except BaseException: # noqa: BLE001 -- surfaced via the caller's box
pass
finally:
self._loader_queue.task_done()
def _run_on_loader(self, fn: Any) -> Any:
"""Run ``fn`` on the persistent loader thread and return its result."""
box: dict[str, Any] = {}
done = threading.Event()
def wrapped() -> None:
try:
box["r"] = fn()
except BaseException as exc: # noqa: BLE001 -- re-raised below
box["e"] = exc
finally:
done.set()
self._loader_queue.put(wrapped)
done.wait()
if "e" in box:
raise box["e"]
return box["r"]
def preload_generator(self, **params: Any) -> dict[str, Any]:
"""Load a model into memory ahead of time. One load at a time; loading
a different config always releases the current instance first."""
config = self._generator_config_dict(**params)
with self._generator_lock:
if self._generator_state == "ready" and self._generator_config == config:
return self._slot_entry()
if not self._load_lock.acquire(blocking=False):
raise RuntimeError("a model load is already in progress")
try:
if self._running_inference_jobs():
raise RuntimeError("cannot swap models while inference jobs are running")
with self._generator_lock:
self._generator_state = "loading"
self._generator_config = config
self._generator_error = None
entry = self._slot_entry()
except BaseException:
self._load_lock.release()
raise
def _load() -> None:
try: # the preload owns _load_lock until the load resolves
self._run_on_loader(lambda: self._load_into_slot_locked(config))
except Exception: # noqa: BLE001 -- state already set to failed
pass
finally:
self._load_lock.release()
threading.Thread(target=_load, daemon=True, name="generator-preload").start()
return entry
def _load_into_slot_locked(self, config: dict[str, Any]) -> Any:
"""Release whatever is resident and load ``config``. Caller MUST hold
``_load_lock``. State is 'loading' on entry or set here."""
with self._generator_lock:
gen = self._generator
self._generator = None
self._generator_state = "loading"
self._generator_config = config
self._generator_error = None
if gen is not None:
logger.info("Releasing resident generator before loading a new one")
gen.shutdown()
del gen
# Import lazily so starting the server is fast even without a GPU.
from fastvideo import VideoGenerator
# Deployment-level knobs (set where the API server is launched):
# FASTVIDEO_STUDIO_MODEL_PATHS="id=/local/dir,..." serves a registered
# model id from local weights instead of the HF hub;
# FASTVIDEO_STUDIO_EXECUTOR_BACKEND=ray runs workers on an existing
# Ray cluster (the multi-node path — "mp" spawns local processes only).
model_path = config["model_id"]
for pair in os.environ.get("FASTVIDEO_STUDIO_MODEL_PATHS", "").split(","):
mid, sep, path = pair.partition("=")
if sep and mid.strip() == config["model_id"]:
model_path = path.strip()
executor_kwargs: dict[str, Any] = {}
backend = os.environ.get("FASTVIDEO_STUDIO_EXECUTOR_BACKEND")
if backend:
executor_kwargs["distributed_executor_backend"] = backend
logger.info("Loading model %s (%s)", config["model_id"],
", ".join(f"{k}={v}" for k, v in config.items() if k != "model_id"))
try:
new_gen = VideoGenerator.from_pretrained(
model_path,
workload_type=config["workload_type"],
num_gpus=config["num_gpus"],
dit_cpu_offload=config["dit_cpu_offload"],
# FastVideoArgs defaults this True, which disables FSDP and
# parks a full DiT copy in host RAM per worker — the mp
# executor's 4 workers OOM-killed the node silently. The UI's
# offload toggles are the studio's contract; layerwise off.
dit_layerwise_offload=False,
text_encoder_cpu_offload=config["text_encoder_cpu_offload"],
vae_cpu_offload=config["vae_cpu_offload"],
image_encoder_cpu_offload=config["image_encoder_cpu_offload"],
use_fsdp_inference=config["use_fsdp_inference"],
enable_torch_compile=config["enable_torch_compile"],
VSA_sparsity=config["vsa_sparsity"],
tp_size=config["tp_size"],
sp_size=config["sp_size"],
log_queue=self._worker_log_queue,
**executor_kwargs,
)
except Exception as exc:
logger.exception("Model load failed for %s", config["model_id"])
with self._generator_lock:
self._generator_state = "failed"
self._generator_error = str(exc)
raise
with self._generator_lock:
self._generator = new_gen
self._generator_config = config
self._generator_state = "ready"
self._generator_error = None
return new_gen
def list_generators(self) -> list[dict[str, Any]]:
"""The resident slot, or empty when nothing is loaded."""
with self._generator_lock:
if self._generator_state == "empty":
return []
return [self._slot_entry()]
def unload_generator(self, **_ignored: Any) -> bool:
"""Shut down and delete the resident generator, freeing GPU memory."""
if not self._load_lock.acquire(blocking=False):
raise RuntimeError("cannot unload while a model load is in progress")
try:
running = self._running_inference_jobs()
if running:
raise RuntimeError(f"cannot unload while inference jobs are running: {running}")
with self._generator_lock:
gen = self._generator
empty = self._generator_state == "empty"
self._generator = None
self._generator_state = "empty"
self._generator_config = None
self._generator_error = None
if empty:
return False
if gen is not None:
logger.info("Releasing resident generator")
gen.shutdown()
del gen
return True
finally:
self._load_lock.release()
def _get_or_create_generator(
self,
model_id: str,
@@ -866,53 +647,69 @@ class JobRunner:
sp_size: int = -1,
log_queue: mp.Queue | None = None,
) -> Any:
"""Return the resident generator if the config matches; otherwise
replace the slot. Blocks behind any in-flight load — every slot
transition happens under ``_load_lock``, so a job can never observe a
half-replaced slot."""
del log_queue # single-slot generators log via the engine tee
config = self._generator_config_dict(
model_id=model_id,
cache_key = (
model_id,
workload_type,
num_gpus,
dit_cpu_offload,
text_encoder_cpu_offload,
vae_cpu_offload,
image_encoder_cpu_offload,
use_fsdp_inference,
enable_torch_compile,
vsa_sparsity,
tp_size,
sp_size,
)
# Generators are cached by model_id and configuration parameters
with self._generators_lock:
if cache_key in self._generators:
return self._generators[cache_key]
# Import lazily so starting the server is fast even without a GPU.
from fastvideo import VideoGenerator
logger.info(
"Loading model %s (workload=%s, num_gpus=%d, offloads: "
"dit=%s text_encoder=%s vae=%s image_encoder=%s, fsdp=%s, "
"torch_compile=%s, vsa_sparsity=%.2f, tp=%d sp=%d) …",
model_id,
workload_type,
num_gpus,
dit_cpu_offload,
text_encoder_cpu_offload,
vae_cpu_offload,
image_encoder_cpu_offload,
use_fsdp_inference,
enable_torch_compile,
vsa_sparsity,
tp_size,
sp_size,
)
gen = VideoGenerator.from_pretrained(
model_id,
workload_type=workload_type,
num_gpus=num_gpus,
dit_cpu_offload=dit_cpu_offload,
text_encoder_cpu_offload=text_encoder_cpu_offload,
vae_cpu_offload=vae_cpu_offload,
image_encoder_cpu_offload=image_encoder_cpu_offload,
use_fsdp_inference=use_fsdp_inference,
enable_torch_compile=enable_torch_compile,
vsa_sparsity=vsa_sparsity,
VSA_sparsity=vsa_sparsity,
tp_size=tp_size,
sp_size=sp_size,
log_queue=log_queue,
)
with self._load_lock: # waits out preloads / other jobs' replaces
with self._generator_lock:
if (self._generator_state == "ready"
and self._generator_config == config
and self._generator is not None):
return self._generator
return self._run_on_loader(lambda: self._load_into_slot_locked(config))
def feed_engine_line(self, line: str) -> None:
"""Bridge ray-relayed worker output into the running job's log buffer.
On the ray backend worker logs cannot cross nodes via the mp queue,
but ray already relays them to the driver's stdout — which the engine
tee captures. Lines with ray's actor prefix are attributed to the one
running inference job, whose buffer parses tqdm into UI progress.
Driver-side logging is excluded (it reaches the buffer via the
logging handlers already).
"""
job = self._active_inference_job
if job is None:
return
line = _ANSI_RE.sub("", line)
if not _RAY_RELAY_RE.match(line):
return
try:
job._log_buf.write(line)
except Exception: # noqa: BLE001 -- never break the tee
pass
with self._generators_lock:
if cache_key not in self._generators:
self._generators[cache_key] = gen
else: # Another thread may have created it while we were loading.
gen.shutdown()
gen = self._generators[cache_key]
return gen
def _run_job(self, job: Job):
if job.job_type == "inference":
@@ -936,23 +733,46 @@ class JobRunner:
self._save_job(job)
return
env = os.environ.copy()
env.update(get_training_env())
try:
# Job.to_dict() carries every key the config builder reads
# (extra keys are ignored by its .get() lookups).
train_config = build_training_config(job.to_dict(), job_output_dir)
except ValueError as exc:
module_info = get_training_module_info(job.workload_type, job.model_id)
if not module_info:
job.status = JobStatus.FAILED
job.error = str(exc)
job.error = f"Unknown workload type: {job.workload_type}"
job.finished_at = time.time()
self._save_job(job)
return
config_path = os.path.join(job_output_dir, "train_config.yaml")
with open(config_path, "w", encoding="utf-8") as f:
yaml.safe_dump(train_config, f, sort_keys=False)
module_path, _pipeline_workload, use_vsa, _is_lora = module_info
dmd_use_vsa = (job.workload_type.startswith("dmd_") and getattr(job, "dmd_use_vsa", False))
env = os.environ.copy()
env.update(get_training_env(use_vsa or dmd_use_vsa))
job_dict = {
"model_id": job.model_id,
"data_path": job.data_path,
"workload_type": job.workload_type,
"num_gpus": job.num_gpus,
"max_train_steps": job.max_train_steps,
"train_batch_size": job.train_batch_size,
"learning_rate": job.learning_rate,
"num_latent_t": job.num_latent_t,
"num_height": job.height,
"num_width": job.width,
"num_frames": job.num_frames,
"validation_dataset_file": job.validation_dataset_file,
"lora_rank": job.lora_rank,
"ltx2_first_frame_conditioning_p": job.ltx2_first_frame_conditioning_p,
}
if job.workload_type.startswith("dmd_") or job.workload_type.startswith("self_forcing_"):
job_dict["dmd_use_vsa"] = getattr(job, "dmd_use_vsa", False)
job_dict["dmd_vsa_sparsity"] = getattr(job, "dmd_vsa_sparsity", 0.8)
job_dict["dmd_denoising_steps"] = getattr(job, "dmd_denoising_steps", "1000,757,522")
job_dict["min_timestep_ratio"] = getattr(job, "min_timestep_ratio", 0.02)
job_dict["max_timestep_ratio"] = getattr(job, "max_timestep_ratio", 0.98)
job_dict["real_score_guidance_scale"] = getattr(job, "real_score_guidance_scale", 3.5)
job_dict["generator_update_interval"] = getattr(job, "generator_update_interval", 5)
job_dict["real_score_model_path"] = (getattr(job, "real_score_model_path", "") or job.model_id)
job_dict["fake_score_model_path"] = (getattr(job, "fake_score_model_path", "") or job.model_id)
train_args = build_training_args(job_dict, job_output_dir)
repo_root = Path(__file__).resolve().parent.parent
torchrun_cmd = [
@@ -963,12 +783,9 @@ class JobRunner:
str(job.num_gpus),
"--nnodes",
"1",
"-m",
"fastvideo.train.entrypoint.train",
"--config",
config_path,
]
buf.write(f"Starting training: {' '.join(torchrun_cmd)}")
str(repo_root / module_path),
] + train_args
buf.write(f"Starting training: {' '.join(torchrun_cmd[:12])}...")
buf.phase = "starting"
try:
@@ -1008,21 +825,12 @@ class JobRunner:
job._process.wait()
exit_code = job._process.returncode or 0
if job._stop_event.is_set():
# Terminated by stop_job() without the reader loop observing the
# flag (e.g. the process died between log lines).
job.status = JobStatus.STOPPED
buf.phase = "stopped"
elif exit_code == 0:
if exit_code == 0:
job.status = JobStatus.COMPLETED
buf.progress = 100.0
buf.phase = "done"
# Training outputs checkpoints, not video. Sort by step number,
# not lexically ("checkpoint-1000" < "checkpoint-500" as strings).
ckpt_dirs = sorted(
Path(job_output_dir).glob("checkpoint-*"),
key=lambda p: int(m.group(1)) if (m := re.fullmatch(r"checkpoint-(\d+)", p.name)) else -1,
)
# Training outputs checkpoints, not video
ckpt_dirs = sorted(Path(job_output_dir).glob("checkpoint-*"))
if ckpt_dirs:
job.output_path = str(ckpt_dirs[-1])
else:
@@ -1058,12 +866,10 @@ class JobRunner:
fastvideo_logger.addHandler(buffer_handler)
fastvideo_logger.addHandler(file_handler)
# Worker logs flow through the runner-wide queue the generator was
# created with; drain anything stale, then listen for this job.
log_queue = self._worker_log_queue
with contextlib.suppress(Exception):
while True:
log_queue.get_nowait()
# Queue for worker process logs (fsdp_load, cuda, etc.)
# Use Manager().Queue() so it can be shared with spawned workers (spawn
# does not inherit memory; mp.Queue only works through inheritance).
log_queue = self._mp_manager.Queue()
queue_listener = logging.handlers.QueueListener(log_queue,
buffer_handler,
file_handler,
@@ -1145,7 +951,6 @@ class JobRunner:
generator = _gen_result[0]
buf.phase = "generating"
self._active_inference_job = job # engine tee feeds tqdm from here
logger.info("Starting generation for job %s (model=%s)", job.id, job.model_id)
gen_kwargs: dict[str, Any] = {
@@ -1161,6 +966,7 @@ class JobRunner:
"fps": job.fps,
"seed": job.seed,
"negative_prompt": job.negative_prompt or "",
"log_queue": log_queue,
}
if job.image_path:
gen_kwargs["image_path"] = job.image_path
@@ -1204,8 +1010,6 @@ class JobRunner:
buf.phase = "failed"
finally:
if self._active_inference_job is job:
self._active_inference_job = None
queue_listener.stop()
# Remove handlers and close file
fastvideo_logger.removeHandler(buffer_handler)
-845
View File
@@ -1,845 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
"""
In-memory mock of the FastVideo Studio API (``server.py``).
This mock implements the same ``/api`` routes and response shapes as the real
FastAPI server but keeps everything in memory and never touches the FastVideo
library, a GPU, or a database. It is meant to back the Playwright e2e suite so
the Next.js frontend can be exercised end-to-end without a real backend.
Job lifecycle is simulated by recording a start timestamp and *computing* the
status on read: a started job reports ``running`` for a few seconds and then
flips to ``completed`` with an ``output_path``, so polling the job list / logs
shows progression. Generated media is a tiny 1s ``testsrc`` MP4 built lazily
with ffmpeg and cached.
Usage (from the ``apps/`` directory)::
PYTHONPATH=.. python -m fastvideo_studio.mock_server --port 8190
"""
from __future__ import annotations
import contextlib
import os
import random
import shutil
import subprocess
import tempfile
import threading
import time
import uuid
from typing import Annotated, Any
from fastapi import FastAPI, File, HTTPException, UploadFile
from fastapi.middleware.cors import CORSMiddleware
from fastapi.responses import FileResponse, PlainTextResponse
from fastvideo_studio.database import default_settings_dict
from fastvideo_studio.models import (CreateDatasetRequest, CreateJobRequest, GeneratorRequest, SettingsUpdate,
UpdateCaptionRequest, model_label)
# --- Config -----------------------------------------------------------------
# How long a started job stays "running" before it flips to "completed".
COMPLETE_AFTER_SECONDS = 3.0
# How long a preloading generator stays "loading" before it flips to "ready".
GENERATOR_READY_AFTER_SECONDS = 2.0
FFMPEG_BIN = shutil.which(os.getenv("FASTVIDEO_FFMPEG_BIN", "ffmpeg"))
# A small catalogue of fake models keyed by workload type. Mirrors the real
# server's {id, label} shape (label derived from the HF-style path).
_MODELS_BY_WORKLOAD: dict[str, list[str]] = {
"t2v": [
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
"FastVideo/FastHunyuan-diffusers",
],
"i2v": [
"Wan-AI/Wan2.1-I2V-14B-480P-Diffusers",
],
"t2i": [
"black-forest-labs/FLUX.1-schnell",
],
}
def _models_for(workload_type: str | None) -> list[dict[str, str]]:
if workload_type:
paths = _MODELS_BY_WORKLOAD.get(workload_type, [])
else:
seen: dict[str, None] = {}
for paths_for_workload in _MODELS_BY_WORKLOAD.values():
for path in paths_for_workload:
seen.setdefault(path, None)
paths = list(seen)
return [{"id": path, "label": model_label(path)} for path in paths]
# The real settings catalogue lives in database.py; the mock only pre-fills
# the default model ids so the Create Job modal auto-selects one.
_DEFAULT_SETTINGS: dict[str, Any] = {
**default_settings_dict(),
"defaultModelId": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
"defaultModelIdT2v": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
"defaultModelIdI2v": "Wan-AI/Wan2.1-I2V-14B-480P-Diffusers",
"defaultModelIdT2i": "black-forest-labs/FLUX.1-schnell",
}
# --- In-memory state --------------------------------------------------------
_state_lock = threading.Lock()
_settings: dict[str, Any] = dict(_DEFAULT_SETTINGS)
_jobs: dict[str, dict[str, Any]] = {}
# The engine's single model slot: None when empty, else the state dict.
_generator_slot: dict[str, Any] | None = None
# Fake engine stdout/stderr tail; grows a little on every poll.
_engine_log_lines: list[str] = ["[engine] FastVideo studio mock engine started"]
_datasets: dict[str, dict[str, Any]] = {}
# dataset_id -> {"file_names": [...], "captions": {file_name: caption}}
_dataset_files: dict[str, dict[str, Any]] = {}
# Lazily-built, cached media clips (shared by every job/dataset media response),
# keyed by extension: "mp4" (1s testsrc video) or "png" (single testsrc frame).
_mock_media_cache: dict[str, str] = {}
_mock_media_lock = threading.Lock()
def _build_mock_media(kind: str) -> str:
"""Build (once) a tiny testsrc clip of the given ``kind`` with ffmpeg; cache the path."""
with _mock_media_lock:
cached = _mock_media_cache.get(kind)
if cached and os.path.isfile(cached):
return cached
if not FFMPEG_BIN:
raise HTTPException(status_code=500, detail="ffmpeg is required to build mock media")
fd, path = tempfile.mkstemp(prefix="fvstudio_mock_", suffix=f".{kind}")
os.close(fd)
if kind == "png":
command = [
FFMPEG_BIN,
"-hide_banner",
"-loglevel",
"error",
"-y",
"-f",
"lavfi",
"-i",
"testsrc=size=320x240:rate=1",
"-frames:v",
"1",
"-f",
"image2",
path,
]
else:
command = [
FFMPEG_BIN,
"-hide_banner",
"-loglevel",
"error",
"-y",
"-f",
"lavfi",
"-i",
"testsrc=size=320x240:rate=24",
"-t",
"1",
"-c:v",
"libx264",
"-preset",
"ultrafast",
"-pix_fmt",
"yuv420p",
"-movflags",
"+faststart",
"-f",
"mp4",
path,
]
try:
subprocess.run(command, check=True, capture_output=True)
except (subprocess.CalledProcessError, OSError) as exc:
with contextlib.suppress(OSError):
os.remove(path)
raise HTTPException(status_code=500, detail=f"ffmpeg failed to build mock {kind}: {exc}") from exc
if not os.path.isfile(path) or os.path.getsize(path) == 0:
raise HTTPException(status_code=500, detail=f"ffmpeg produced no mock {kind} bytes")
_mock_media_cache[kind] = path
return path
# --- Job helpers ------------------------------------------------------------
def _new_job_dict(req: CreateJobRequest) -> dict[str, Any]:
"""Build a job dict (mirrors job_runner.Job.to_dict()) in the pending state."""
job_id = str(uuid.uuid4())
return {
"id": job_id,
"model_id": req.model_id,
"prompt": req.prompt,
"workload_type": req.workload_type or "t2v",
"job_type": req.job_type or "inference",
"image_path": req.image_path or "",
"status": "pending",
"created_at": time.time(),
"started_at": None,
"finished_at": None,
"error": None,
"output_path": None,
"log_file_path": None,
"num_inference_steps": req.num_inference_steps,
"num_frames": req.num_frames,
"height": req.height,
"width": req.width,
"guidance_scale": req.guidance_scale,
"guidance_rescale": req.guidance_rescale,
"fps": req.fps,
"seed": req.seed,
"negative_prompt": req.negative_prompt or "",
"num_gpus": req.num_gpus,
"data_path": req.data_path or "",
"progress": 0.0,
"progress_msg": "",
"phase": "pending",
}
def _advance_job(job: dict[str, Any]) -> None:
"""Flip a running job to completed once enough wall-clock time has passed.
Status is *computed on read* from the recorded start timestamp, so polling
the job list / logs naturally shows pending -> running -> completed.
"""
if job["status"] != "running" or not job.get("started_at"):
return
elapsed = time.time() - job["started_at"]
if elapsed >= COMPLETE_AFTER_SECONDS:
job["status"] = "completed"
job["finished_at"] = job["started_at"] + COMPLETE_AFTER_SECONDS
job["progress"] = 100.0
job["progress_msg"] = "50/50 steps"
job["phase"] = "done"
ext = "png" if job.get("workload_type") == "t2i" else "mp4"
job["output_path"] = f"/mock/outputs/{job['id']}/output.{ext}"
def _public_job(job: dict[str, Any]) -> dict[str, Any]:
"""Advance + return a copy safe to serialize."""
_advance_job(job)
return dict(job)
_LOG_TAIL = [
"Loading model...",
"Model loaded.",
"Starting generation...",
"Denoising step 10/50",
"Denoising step 20/50",
"Denoising step 30/50",
"Denoising step 40/50",
"Denoising step 50/50",
"Generation complete. Saving output...",
"Saved output file.",
"Job completed successfully.",
]
def _log_sequence(job: dict[str, Any]) -> list[str]:
return [
f"Job {job['id']} started",
f"Model: {job['model_id']}",
f"Prompt: {job['prompt']}",
*_LOG_TAIL,
]
def _compute_logs(job: dict[str, Any]) -> dict[str, Any]:
"""Return JobLogs-shaped data, growing the visible lines as time passes."""
seq = _log_sequence(job)
status = job["status"]
if status == "pending":
return {"lines": [], "progress": 0.0, "progress_msg": "", "phase": "pending"}
if status == "completed":
return {"lines": seq, "progress": 100.0, "progress_msg": "50/50 steps", "phase": "done"}
if status in ("failed", "stopped"):
# Keep the total line count monotonic across running -> terminal (the
# `after` cursor relies on it): show every non-final line plus a terminal
# notice, which is always >= any prefix a running job revealed.
lines = seq[:-1] + [f"Job {status}."]
return {"lines": lines, "progress": job.get("progress", 0.0), "progress_msg": "", "phase": status}
# running: reveal a prefix proportional to elapsed time, but never the final
# "completed" line — that appears only once the job is actually completed.
elapsed = time.time() - (job.get("started_at") or time.time())
frac = max(0.0, min(elapsed / COMPLETE_AFTER_SECONDS, 0.99))
reveal = min(max(4, round(frac * len(seq))), len(seq) - 1)
lines = seq[:reveal]
if frac < 0.2:
phase = "loading model"
elif frac < 0.85:
phase = "denoising"
else:
phase = "saving"
return {
"lines": lines,
"progress": round(frac * 100.0, 1),
"progress_msg": f"{int(frac * 50)}/50 steps",
"phase": phase,
}
# --- Dataset helpers --------------------------------------------------------
def _dataset_stats(dataset_id: str) -> tuple[int, int]:
files = _dataset_files.get(dataset_id, {}).get("file_names", [])
count = len(files)
# Fake but stable per-file size so the UI shows a non-zero footprint.
return count, count * 1_048_576
def _public_dataset(dataset: dict[str, Any]) -> dict[str, Any]:
count, size = _dataset_stats(dataset["id"])
return {**dataset, "file_count": count, "size_bytes": size}
def _seed() -> None:
"""Seed a couple of datasets and one completed inference job."""
now = time.time()
seeds = [
("Sunset Clips", ["sunset_01.mp4", "sunset_02.mp4"], {
"sunset_01.mp4": "A sunset over the ocean",
"sunset_02.mp4": "A sunset over the mountains"
}),
("City Timelapse", ["city_01.mp4"], {
"city_01.mp4": "A busy city intersection at night"
}),
]
for idx, (name, file_names, captions) in enumerate(seeds):
dataset_id = str(uuid.uuid4())
_datasets[dataset_id] = {"id": dataset_id, "name": name, "created_at": now - 600 + idx}
_dataset_files[dataset_id] = {"file_names": list(file_names), "captions": dict(captions)}
job_id = str(uuid.uuid4())
_jobs[job_id] = {
"id": job_id,
"model_id": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
"prompt": "A curious raccoon peers through a field of yellow sunflowers",
"workload_type": "t2v",
"job_type": "inference",
"image_path": "",
"status": "completed",
"created_at": now - 300,
"started_at": now - 280,
"finished_at": now - 250,
"error": None,
"output_path": f"/mock/outputs/{job_id}/output.mp4",
"log_file_path": f"/mock/logs/{job_id}.log",
"num_inference_steps": 50,
"num_frames": 81,
"height": 480,
"width": 832,
"guidance_scale": 5.0,
"guidance_rescale": 0.0,
"fps": 24,
"seed": 1024,
"negative_prompt": "",
"num_gpus": 1,
"data_path": "",
"progress": 100.0,
"progress_msg": "50/50 steps",
"phase": "done",
}
# --- App --------------------------------------------------------------------
app = FastAPI(title="FastVideo Studio Mock API", version="0.1.0")
app.add_middleware(
CORSMiddleware,
allow_origins=["*"],
allow_methods=["*"],
allow_headers=["*"],
)
_seed()
@app.get("/api/__mock__")
def mock_sentinel() -> dict[str, bool]:
"""Mock-only marker so the e2e suite can prove it isn't hitting a real
backend before running its mutating specs (the real server has no such
route)."""
return {"mock": True}
# --- Settings ---------------------------------------------------------------
@app.get("/api/settings")
def get_settings() -> dict[str, Any]:
with _state_lock:
return dict(_settings)
@app.put("/api/settings")
def update_settings(settings: SettingsUpdate) -> dict[str, Any]:
updates = settings.model_dump(exclude_unset=True)
with _state_lock:
_settings.update({k: v for k, v in updates.items() if v is not None})
return dict(_settings)
# --- Models -----------------------------------------------------------------
@app.get("/api/models")
def list_models(workload_type: str | None = None) -> list[dict[str, Any]]:
return _models_for(workload_type)
# Per-model sampling presets, mirroring the real /api/models/presets shape
# (keys the model has no recommendation for are simply absent).
_MODEL_PRESETS: dict[str, dict[str, Any]] = {
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers": {
"height": 480,
"width": 832,
"num_frames": 81,
"fps": 16,
"num_inference_steps": 50,
"guidance_scale": 3.0,
"guidance_rescale": 0.0,
"negative_prompt": "Bright tones, overexposed, static, blurred details",
"seed": 1024,
},
"Wan-AI/Wan2.1-I2V-14B-480P-Diffusers": {
"height": 480,
"width": 832,
"num_frames": 81,
"fps": 16,
"num_inference_steps": 40,
"guidance_scale": 5.0,
"seed": 1024,
},
"black-forest-labs/FLUX.1-schnell": {
"height": 1024,
"width": 1024,
"num_frames": 1,
"num_inference_steps": 4,
"guidance_scale": 0.0,
"seed": 42,
},
}
_GENERIC_PRESETS: dict[str, Any] = {
"height": 720,
"width": 1280,
"num_frames": 81,
"fps": 24,
"num_inference_steps": 50,
"guidance_scale": 5.0,
"guidance_rescale": 0.0,
"seed": 1024,
}
@app.get("/api/models/presets")
def model_presets(model_id: str) -> dict[str, Any]:
"""Recommended sampling settings; unknown models get generic defaults."""
return _MODEL_PRESETS.get(model_id, _GENERIC_PRESETS)
# --- GPUs -------------------------------------------------------------------
@app.get("/api/gpus")
def list_gpus() -> dict[str, Any]:
"""Two fake GPUs with slight per-request jitter so the page looks live."""
gpus = []
for index, (base_util, used_mib) in enumerate([(62, 61_440), (7, 4_096)]):
gpus.append({
"index": index,
"name": "NVIDIA Mock GPU 80GB",
"utilization": max(0, min(100, base_util + random.randint(-5, 5))),
"memory_used_mib": used_mib + random.randint(-256, 256),
"memory_total_mib": 81_920,
"temperature_c": 55 + random.randint(-3, 3),
"power_watts": 310.0 + random.randint(-20, 20),
"power_limit_watts": 700.0,
})
return {"available": True, "gpus": gpus, "error": None}
@app.get("/api/cluster")
def cluster_status() -> dict[str, Any]:
"""Two fake ray nodes x 4 GPUs with per-request jitter, mirroring the
real server's /api/cluster shape."""
nodes = []
for host_idx, (hostname, ip, is_this_host) in enumerate([
("mock-node-0", "10.0.0.10", True),
("mock-node-1", "10.0.0.11", False),
]):
gpus = []
for index in range(4):
base_util = (13 + 29 * index + 41 * host_idx) % 100
gpus.append({
"index": index,
"name": "NVIDIA Mock GPU 80GB",
"utilization": max(0, min(100, base_util + random.randint(-5, 5))),
"memory_used_mib": 6_144 + 17_408 * index + random.randint(-256, 256),
"memory_total_mib": 81_920,
"temperature_c": 42 + 6 * index + random.randint(-3, 3),
"power_watts": 110.0 + 140.0 * index + random.randint(-20, 20),
"power_limit_watts": 700.0,
})
nodes.append({
"hostname": hostname,
"ip": ip,
"is_this_host": is_this_host,
"cpus": 64.0,
"ray_gpus": 4.0,
"available": True,
"error": None,
"gpus": gpus,
})
return {
"mode": "ray",
"nodes": nodes,
"resources": {
"gpus_total": 8.0,
"gpus_available": 5.0
},
"error": None,
}
# --- Uploads ----------------------------------------------------------------
@app.post("/api/upload-image")
async def upload_image(file: Annotated[UploadFile, File()]) -> dict[str, str]:
name = file.filename or "image.png"
return {"path": f"/mock/uploads/{uuid.uuid4().hex}_{os.path.basename(name)}"}
@app.post("/api/upload-raw-dataset")
async def upload_raw_dataset(files: Annotated[list[UploadFile], File()]) -> dict[str, Any]:
video_exts = {".mp4", ".webm", ".avi", ".mov", ".mkv"}
file_names: list[str] = []
for uf in files:
name = os.path.basename(uf.filename or f"{uuid.uuid4().hex}.mp4")
if os.path.splitext(name)[1].lower() in video_exts:
file_names.append(name)
if not file_names:
raise HTTPException(status_code=400, detail="No video files found.")
upload_id = uuid.uuid4().hex
path = f"/mock/uploads/{upload_id}"
return {"path": path, "upload_id": upload_id, "file_names": file_names}
# --- Jobs -------------------------------------------------------------------
@app.get("/api/jobs")
def list_jobs(job_type: str | None = None) -> list[dict[str, Any]]:
# _public_job mutates the shared job dict via _advance_job, so it must run
# under the lock (matching every other job route) to avoid racing start/stop.
with _state_lock:
jobs = list(_jobs.values())
if job_type:
jobs = [j for j in jobs if j.get("job_type") == job_type]
jobs.sort(key=lambda j: j["created_at"], reverse=True)
return [_public_job(j) for j in jobs]
@app.get("/api/jobs/{job_id}")
def get_job(job_id: str) -> dict[str, Any]:
with _state_lock:
job = _jobs.get(job_id)
if job is None:
raise HTTPException(status_code=404, detail="Job not found")
return _public_job(job)
@app.post("/api/jobs", status_code=201)
def create_job(req: CreateJobRequest) -> dict[str, Any]:
job = _new_job_dict(req)
with _state_lock:
_jobs[job["id"]] = job
if _settings.get("autoStartJob"):
_start(job)
return _public_job(job)
def _start(job: dict[str, Any]) -> None:
job["status"] = "running"
job["started_at"] = time.time()
job["finished_at"] = None
job["error"] = None
job["output_path"] = None
# The real server assigns the log path once the job starts running; mirror
# that so the Job Details "Download Log" button (gated on log_file_path) works.
job["log_file_path"] = f"/mock/logs/{job['id']}.log"
job["progress"] = 0.0
job["progress_msg"] = ""
job["phase"] = "starting"
@app.post("/api/jobs/{job_id}/start")
def start_job(job_id: str) -> dict[str, Any]:
with _state_lock:
job = _jobs.get(job_id)
if job is None:
raise HTTPException(status_code=404, detail=f"Job {job_id} not found")
_advance_job(job)
if job["status"] == "running":
raise HTTPException(status_code=409, detail="Job is already running")
if job["status"] == "completed":
raise HTTPException(status_code=409, detail="Job already completed. Delete and re-create to run again.")
_start(job)
return _public_job(job)
@app.post("/api/jobs/{job_id}/stop")
def stop_job(job_id: str) -> dict[str, Any]:
with _state_lock:
job = _jobs.get(job_id)
if job is None:
raise HTTPException(status_code=404, detail=f"Job {job_id} not found")
_advance_job(job)
if job["status"] != "running":
raise HTTPException(status_code=409, detail=f"Job is not running (status={job['status']})")
job["status"] = "stopped"
job["finished_at"] = time.time()
job["phase"] = "stopped"
return dict(job)
@app.delete("/api/jobs/{job_id}")
def delete_job(job_id: str) -> dict[str, str]:
with _state_lock:
if _jobs.pop(job_id, None) is None:
raise HTTPException(status_code=404, detail="Job not found")
return {"detail": f"Job {job_id} deleted"}
@app.get("/api/jobs/{job_id}/logs")
def get_job_logs(job_id: str, after: int = 0) -> dict[str, Any]:
with _state_lock:
job = _jobs.get(job_id)
if job is None:
raise HTTPException(status_code=404, detail=f"Job {job_id} not found")
_advance_job(job)
result = _compute_logs(job)
all_lines = result["lines"]
return {
"lines": all_lines[after:],
"total": len(all_lines),
"progress": result["progress"],
"progress_msg": result["progress_msg"],
"phase": result["phase"],
}
@app.get("/api/jobs/{job_id}/video")
def get_video(job_id: str) -> FileResponse:
with _state_lock:
job = _jobs.get(job_id)
if job is None:
raise HTTPException(status_code=404, detail="Job not found")
_advance_job(job)
output_path = job.get("output_path")
if job["status"] != "completed" or not output_path:
raise HTTPException(status_code=404, detail="No output available for this job")
# Mirror the real server: image workloads (t2i) output a .png served as an
# image; everything else is a video.
if output_path.endswith(".png"):
return FileResponse(_build_mock_media("png"), media_type="image/png", filename=f"job_{job_id}.png")
return FileResponse(_build_mock_media("mp4"), media_type="video/mp4", filename=f"job_{job_id}.mp4")
@app.get("/api/jobs/{job_id}/download_log")
def download_log(job_id: str) -> PlainTextResponse:
with _state_lock:
job = _jobs.get(job_id)
if job is None:
raise HTTPException(status_code=404, detail="Job not found")
_advance_job(job)
lines = _compute_logs(job)["lines"]
return PlainTextResponse("\n".join(lines) + "\n", media_type="text/plain")
# --- Generators (warm models) -------------------------------------------------
def _advance_generator(entry: dict[str, Any]) -> None:
"""Flip a loading generator to ready once enough wall-clock time has passed.
Like job status, generator state is *computed on read*, so polling the
generators list naturally shows loading -> ready.
"""
if entry["state"] == "loading" and time.time() - entry["started_at"] >= GENERATOR_READY_AFTER_SECONDS:
entry["state"] = "ready"
def _running_inference_ids() -> list[str]:
return [
j["id"] for j in _jobs.values() if j.get("job_type") == "inference" and _public_job(j)["status"] == "running"
]
@app.get("/api/generators")
def list_generators() -> list[dict[str, Any]]:
with _state_lock:
if _generator_slot is None:
return []
_advance_generator(_generator_slot)
return [dict(_generator_slot)]
@app.post("/api/generators/preload", status_code=202)
def preload_generator(req: GeneratorRequest) -> dict[str, Any]:
global _generator_slot
valid_ids = {m["id"] for m in _models_for(None)}
if req.model_id not in valid_ids:
raise HTTPException(
status_code=400,
detail=f"Unknown model_id '{req.model_id}'. Valid options: {sorted(valid_ids)}",
)
with _state_lock:
if _generator_slot is not None:
_advance_generator(_generator_slot)
if _generator_slot["state"] == "loading":
raise HTTPException(status_code=409, detail="a model load is already in progress")
if _generator_slot["state"] == "ready" and all(
_generator_slot.get(k) == v for k, v in req.model_dump().items()):
return dict(_generator_slot)
running = _running_inference_ids()
if running:
raise HTTPException(status_code=409,
detail=f"cannot swap models while inference jobs are running: {running}")
# Loading a new model always replaces (releases) the resident one.
_generator_slot = {"state": "loading", "started_at": time.time(), "error": None, **req.model_dump()}
return dict(_generator_slot)
@app.post("/api/generators/unload")
def unload_generator() -> dict[str, Any]:
global _generator_slot
with _state_lock:
if _generator_slot is not None:
_advance_generator(_generator_slot)
if _generator_slot["state"] == "loading":
raise HTTPException(status_code=409, detail="cannot unload while a model load is in progress")
if _generator_slot is None:
raise HTTPException(status_code=404, detail="No model is loaded")
running = _running_inference_ids()
if running:
raise HTTPException(status_code=409, detail=f"cannot unload while inference jobs are running: {running}")
_generator_slot = None
return {"unloaded": True}
# --- Engine logs --------------------------------------------------------------
@app.get("/api/engine/logs")
def engine_logs(after: int = 0) -> dict[str, Any]:
"""Growing fake tail of the engine's stdout/stderr: every poll appends a
couple of lines so the console visibly streams."""
with _state_lock:
n = len(_engine_log_lines)
_engine_log_lines.append(f"[engine] step {n}: worker heartbeat ok")
_engine_log_lines.append(f"[engine] step {n + 1}: gpu mem {random.randint(20, 80)}% used")
total = len(_engine_log_lines)
return {"lines": _engine_log_lines[max(0, after):], "total": total}
# --- Datasets ---------------------------------------------------------------
@app.get("/api/datasets")
def list_datasets() -> list[dict[str, Any]]:
with _state_lock:
datasets = sorted(_datasets.values(), key=lambda d: d["created_at"], reverse=True)
return [_public_dataset(d) for d in datasets]
@app.get("/api/datasets/{dataset_id}")
def get_dataset(dataset_id: str) -> dict[str, Any]:
with _state_lock:
dataset = _datasets.get(dataset_id)
if dataset is None:
raise HTTPException(status_code=404, detail="Dataset not found")
return _public_dataset(dataset)
@app.post("/api/datasets", status_code=201)
def create_dataset(req: CreateDatasetRequest) -> dict[str, Any]:
if not req.upload_path:
raise HTTPException(status_code=400, detail="upload_path is required. Upload media files first.")
if not req.file_names:
raise HTTPException(status_code=400, detail="No media files found.")
dataset_id = str(uuid.uuid4())
dataset = {"id": dataset_id, "name": req.name, "created_at": time.time()}
captions = {fn: (req.captions.get(fn, "") if req.captions else "") for fn in req.file_names}
with _state_lock:
_datasets[dataset_id] = dataset
_dataset_files[dataset_id] = {"file_names": list(req.file_names), "captions": captions}
return _public_dataset(dataset)
@app.get("/api/datasets/{dataset_id}/files")
def get_dataset_files(dataset_id: str) -> dict[str, Any]:
with _state_lock:
if dataset_id not in _datasets:
raise HTTPException(status_code=404, detail="Dataset not found")
files = _dataset_files.get(dataset_id, {"file_names": [], "captions": {}})
return {"file_names": list(files["file_names"]), "captions": dict(files["captions"])}
@app.put("/api/datasets/{dataset_id}/captions")
def update_dataset_caption(dataset_id: str, req: UpdateCaptionRequest) -> dict[str, str]:
with _state_lock:
if dataset_id not in _datasets:
raise HTTPException(status_code=404, detail="Dataset not found")
files = _dataset_files.setdefault(dataset_id, {"file_names": [], "captions": {}})
files["captions"][req.file_name] = req.caption
return {"detail": "Caption updated"}
@app.get("/api/datasets/{dataset_id}/media/{file_name:path}")
def serve_dataset_media(dataset_id: str, file_name: str) -> FileResponse:
with _state_lock:
if dataset_id not in _datasets:
raise HTTPException(status_code=404, detail="Dataset not found")
if file_name not in _dataset_files.get(dataset_id, {}).get("file_names", []):
raise HTTPException(status_code=404, detail="File not found")
return FileResponse(_build_mock_media("mp4"), media_type="video/mp4")
@app.delete("/api/datasets/{dataset_id}")
def delete_dataset(dataset_id: str) -> dict[str, str]:
with _state_lock:
if _datasets.pop(dataset_id, None) is None:
raise HTTPException(status_code=404, detail="Dataset not found")
_dataset_files.pop(dataset_id, None)
return {"detail": f"Dataset {dataset_id} deleted"}
def main() -> None:
import argparse
import uvicorn
parser = argparse.ArgumentParser(description="FastVideo Studio mock API server")
parser.add_argument("--host", default="127.0.0.1", help="Bind address (default: 127.0.0.1)")
# Default off the real server's 8189 so the mock never shadows a real API.
parser.add_argument("--port", type=int, default=8190, help="Port number (default: 8190)")
args = parser.parse_args()
uvicorn.run(app, host=args.host, port=args.port, log_level="info")
if __name__ == "__main__":
main()
+1 -11
View File
@@ -1,24 +1,14 @@
# SPDX-License-Identifier: Apache-2.0
"""Pydantic request/response models for the API, plus tiny shared helpers
usable by both the real server and the dependency-light mock server."""
"""Pydantic request/response models for the API."""
from fastvideo_studio.models.create_job_request import CreateJobRequest
from fastvideo_studio.models.generator_request import GeneratorRequest
from fastvideo_studio.models.settings_update import SettingsUpdate
from fastvideo_studio.models.create_dataset_request import CreateDatasetRequest
from fastvideo_studio.models.update_caption_request import UpdateCaptionRequest
def model_label(model_path: str) -> str:
"""Derive a readable label from an HF-style model path."""
return model_path.split("/")[-1].replace("-", " ").replace("_", " ")
__all__ = [
"CreateJobRequest",
"GeneratorRequest",
"SettingsUpdate",
"CreateDatasetRequest",
"UpdateCaptionRequest",
"model_label",
]
@@ -17,6 +17,7 @@ class CreateJobRequest(BaseModel):
num_latent_t: int = 20
validation_dataset_file: str = ""
lora_rank: int = 32
ltx2_first_frame_conditioning_p: float | None = None
negative_prompt: str = ""
num_inference_steps: int = 50
num_frames: int = 81
@@ -40,6 +41,8 @@ class CreateJobRequest(BaseModel):
dmd_use_vsa: bool = False
dmd_vsa_sparsity: float = 0.8
dmd_denoising_steps: str = "1000,757,522"
min_timestep_ratio: float = 0.02
max_timestep_ratio: float = 0.98
real_score_guidance_scale: float = 3.5
generator_update_interval: int = 5
real_score_model_path: str = ""
@@ -1,24 +0,0 @@
# SPDX-License-Identifier: Apache-2.0
"""Request model for preloading/unloading a resident generator.
Field names and defaults mirror the engine subset of ``CreateJobRequest`` so
the UI can send exactly the values it would put on a job — guaranteeing the
job's generator lookup hits this cache entry.
"""
from pydantic import BaseModel
class GeneratorRequest(BaseModel):
model_id: str
workload_type: str = "t2v"
num_gpus: int = 1
dit_cpu_offload: bool = False
text_encoder_cpu_offload: bool = False
vae_cpu_offload: bool = False
image_encoder_cpu_offload: bool = False
use_fsdp_inference: bool = False
enable_torch_compile: bool = False
vsa_sparsity: float = 0.0
tp_size: int = -1
sp_size: int = -1
-13
View File
@@ -1,13 +0,0 @@
import path from 'node:path';
import { fileURLToPath } from 'node:url';
import type { NextConfig } from 'next';
const configDir = path.dirname(fileURLToPath(import.meta.url));
const nextConfig: NextConfig = {
// Point tracing at the monorepo root (apps/fastvideo_studio -> apps -> repo
// root) so Next stops warning about multiple lockfiles in the workspace.
outputFileTracingRoot: path.join(configDir, '..', '..'),
};
export default nextConfig;
+1082 -5343
View File
File diff suppressed because it is too large Load Diff
+22 -44
View File
@@ -4,55 +4,33 @@
"private": true,
"type": "module",
"scripts": {
"dev": "next dev --port 3000",
"build": "next build",
"start": "next start --port 3000",
"typecheck": "tsc --noEmit",
"test": "vitest run",
"test:watch": "vitest",
"test:coverage": "vitest run --coverage",
"e2e": "playwright test",
"dev": "vite dev",
"build": "vite build",
"preview": "vite preview",
"start": "concurrently --kill-others-on-fail \"npm:start:api\" \"npm:start:web\"",
"start:api": "cd .. && python -m fastvideo_studio.server",
"start:web": "next start --port 3000",
"start:all": "concurrently --kill-others-on-fail \"npm:start:api\" \"npm:start:web\""
"start:web": "node build",
"lint": "eslint ."
},
"dependencies": {
"@radix-ui/react-dialog": "^1.1.0",
"@radix-ui/react-dropdown-menu": "^2.1.24",
"@radix-ui/react-label": "^2.1.8",
"@radix-ui/react-scroll-area": "^1.2.10",
"@radix-ui/react-select": "^2.2.6",
"@radix-ui/react-separator": "^1.1.8",
"@radix-ui/react-slider": "^1.2.0",
"@radix-ui/react-slot": "^1.2.4",
"@radix-ui/react-switch": "^1.1.0",
"@radix-ui/react-tabs": "^1.1.0",
"class-variance-authority": "^0.7.1",
"clsx": "^2.1.1",
"lucide-react": "^0.577.0",
"next": "15.5.18",
"react": "^19.1.0",
"react-dom": "^19.1.0",
"sonner": "^2.0.7",
"tailwind-merge": "^3.5.0"
"@sveltejs/kit": "2.57.1",
"svelte": "5.55.7"
},
"devDependencies": {
"@playwright/test": "^1.59.1",
"@tailwindcss/postcss": "^4.2.1",
"@testing-library/dom": "^10.4.1",
"@testing-library/jest-dom": "^6.9.1",
"@testing-library/react": "^16.3.0",
"@testing-library/user-event": "^14.6.1",
"@types/node": "^22",
"@types/react": "^19.1.0",
"@types/react-dom": "^19.1.0",
"@vitejs/plugin-react": "^4.5.2",
"@vitest/coverage-v8": "^3.2.4",
"@sveltejs/adapter-node": "^5.5.4",
"@sveltejs/vite-plugin-svelte": "^5.0.0",
"@types/node": "^20",
"concurrently": "^9",
"jsdom": "^26.1.0",
"postcss": "8.5.10",
"tailwindcss": "^4.2.1",
"typescript": "^5.8.3",
"vitest": "^3.2.4"
"eslint": "^9",
"typescript": "^5",
"vite": "6.4.2"
},
"overrides": {
"cookie": "0.7.0",
"devalue": "5.8.1",
"flatted": "3.4.2",
"minimatch": "3.1.4",
"picomatch": "4.0.4",
"postcss": "8.5.10"
}
}
@@ -1,54 +0,0 @@
import { defineConfig, devices } from '@playwright/test';
import { API_BASE, MOCK_API_PORT } from './e2e/helpers';
/**
* Playwright config for FastVideo Studio end-to-end tests.
*
* The Next.js frontend runs on port 3000 and talks to the in-memory mock API
* (fastvideo_studio.mock_server) on port 8189. The webServer block boots both
* the mock backend and `npm run dev` (pointed at the mock via
* NEXT_PUBLIC_API_BASE_URL) if nothing is already listening, so the suite works
* both locally and in CI. Set PLAYWRIGHT_SKIP_WEBSERVER to reuse externally
* managed servers.
*/
export default defineConfig({
testDir: './e2e',
fullyParallel: false,
workers: 1,
forbidOnly: !!process.env.CI,
retries: process.env.CI ? 2 : 0,
reporter: process.env.CI ? 'github' : 'list',
timeout: 120_000,
expect: { timeout: 30_000 },
use: {
baseURL: process.env.PLAYWRIGHT_BASE_URL ?? 'http://127.0.0.1:3000',
headless: true,
viewport: { width: 1280, height: 720 },
screenshot: 'only-on-failure',
trace: 'retain-on-failure',
},
projects: [{ name: 'chromium', use: { ...devices['Desktop Chrome'] } }],
webServer: process.env.PLAYWRIGHT_SKIP_WEBSERVER
? undefined
: [
{
command: `PYTHONPATH=.. python -m fastvideo_studio.mock_server --port ${MOCK_API_PORT}`,
// Probe the mock-only sentinel, not /api/models (which a real server
// also serves), so a real backend on this port is never mistaken for
// the mock. Always boot a fresh mock in CI.
url: `http://127.0.0.1:${MOCK_API_PORT}/api/__mock__`,
reuseExistingServer: !process.env.CI,
timeout: 120_000,
},
{
command: 'npm run dev',
url: 'http://127.0.0.1:3000',
reuseExistingServer: true,
timeout: 120_000,
env: {
NEXT_PUBLIC_API_BASE_URL: API_BASE,
},
},
],
});
-5
View File
@@ -1,5 +0,0 @@
export default {
plugins: {
'@tailwindcss/postcss': {},
},
};
+142
View File
@@ -0,0 +1,142 @@
# SPDX-License-Identifier: Apache-2.0
"""
Runs preprocessing subprocess for datasets.
"""
from __future__ import annotations
import logging
import os
import subprocess
import sys
import threading
from pathlib import Path
from collections.abc import Callable
logger = logging.getLogger("fastvideo.studio.preprocess_runner")
def build_preprocess_args(
dataset_id: str,
raw_path: str,
output_dir: str,
workload_type: str,
model_path: str,
dataset_type: str = "merged",
num_gpus: int = 1,
) -> list[str]:
"""Build CLI args for v1_preprocessing_new."""
return [
"--model-path",
model_path,
"--mode",
"preprocess",
"--workload-type",
workload_type,
"--preprocess.video_loader_type",
"torchvision",
"--preprocess.dataset_type",
dataset_type,
"--preprocess.dataset_path",
raw_path,
"--preprocess.dataset_output_dir",
output_dir,
"--preprocess.preprocess_video_batch_size",
"2",
"--preprocess.dataloader_num_workers",
"0",
"--preprocess.max_height",
"480",
"--preprocess.max_width",
"832",
"--preprocess.num_frames",
"77",
"--preprocess.train_fps",
"16",
"--preprocess.samples_per_file",
"8",
"--preprocess.flush_frequency",
"8",
"--preprocess.video_length_tolerance_range",
"5",
]
def run_preprocess(
dataset_id: str,
raw_path: str,
output_dir: str,
workload_type: str,
model_path: str,
dataset_type: str,
num_gpus: int,
log_file_path: str,
on_status_change: Callable[[str, str | None], None],
stop_event: threading.Event,
) -> None:
"""Run preprocessing subprocess. Calls on_status_change(status, error)."""
repo_root = Path(__file__).resolve().parent.parent
preprocess_module = "fastvideo.pipelines.preprocess.v1_preprocessing_new"
args = build_preprocess_args(
dataset_id=dataset_id,
raw_path=raw_path,
output_dir=output_dir,
workload_type=workload_type,
model_path=model_path,
dataset_type=dataset_type,
num_gpus=num_gpus,
)
cmd = [
sys.executable,
"-m",
"torch.distributed.run",
"--nproc_per_node",
str(num_gpus),
"--nnodes",
"1",
"-m",
preprocess_module,
] + args
env = os.environ.copy()
env["TOKENIZERS_PARALLELISM"] = "false"
try:
on_status_change("preprocessing", None)
with open(log_file_path, "w", encoding="utf-8") as log_file:
proc = subprocess.Popen(
cmd,
cwd=str(repo_root),
env=env,
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
text=True,
bufsize=1,
)
assert proc.stdout is not None
for line in iter(proc.stdout.readline, ""):
if stop_event.is_set():
proc.terminate()
try:
proc.wait(timeout=30)
except subprocess.TimeoutExpired:
proc.kill()
on_status_change("stopped", "Preprocessing stopped by user")
return
line = line.rstrip()
if line:
log_file.write(line + "\n")
log_file.flush()
proc.wait()
exit_code = proc.returncode or 0
if exit_code == 0:
on_status_change("ready", None)
else:
on_status_change("failed", f"Preprocessing exited with code {exit_code}")
except Exception as exc:
on_status_change("failed", f"{type(exc).__name__}: {exc}")
logger.exception("Preprocessing failed for dataset %s", dataset_id)
+18 -199
View File
@@ -32,10 +32,8 @@ from fastapi.responses import FileResponse
from fastvideo.registry import (get_registered_model_paths, get_registered_models_with_workloads)
from fastvideo_studio.database import Database, _get_db_path
from fastvideo_studio.gpu import get_cluster_snapshot, get_gpu_snapshot
from fastvideo_studio.job_runner import JobRunner, JobStatus
from fastvideo_studio.models import (CreateDatasetRequest, CreateJobRequest, GeneratorRequest, SettingsUpdate,
UpdateCaptionRequest, model_label)
from fastvideo_studio.models import (CreateDatasetRequest, CreateJobRequest, SettingsUpdate, UpdateCaptionRequest)
logging.basicConfig(
level=logging.INFO,
@@ -46,69 +44,14 @@ logger = logging.getLogger("fastvideo.studio.api")
DEFAULT_OUTPUT_DIR = os.path.join(os.path.dirname(__file__), "..", "outputs", "ui_jobs")
class _EngineLogBuffer:
"""Thread-safe ring buffer over the server process's stdout/stderr.
def _get_model_label(model_path: str) -> str:
"""Derive a readable label from a HF model path."""
return model_path.split("/")[-1].replace("-", " ").replace("_", " ")
Because ray relays worker output to the driver (log_to_driver), teeing the
server's own streams captures engine output from every rank on every node;
under the mp executor, worker logs arrive via the logging handlers which
also write to stderr.
"""
def __init__(self, maxlen: int = 5000) -> None:
import collections
import threading
self._lines: collections.deque[str] = collections.deque(maxlen=maxlen)
self._dropped = 0
self._lock = threading.Lock()
self._partial = ""
on_line = None # optional callable(str), set once at startup
def write(self, text: str) -> None:
with self._lock:
buf = self._partial + text
*complete, self._partial = buf.split("\n")
for line in complete:
if len(self._lines) == self._lines.maxlen:
self._dropped += 1
self._lines.append(line)
if self.on_line is not None:
for line in complete:
# never break stdout on a bad feed
with contextlib.suppress(Exception):
self.on_line(line)
def get_lines(self, after: int = 0) -> tuple[list[str], int]:
with self._lock:
total = self._dropped + len(self._lines)
start = max(0, after - self._dropped)
return list(self._lines)[start:], total
class _Tee:
"""File-like that forwards to the original stream and the ring buffer."""
def __init__(self, orig: Any, buffer: _EngineLogBuffer) -> None:
self._orig = orig
self._buffer = buffer
def write(self, text: str) -> int:
self._buffer.write(text)
return self._orig.write(text)
def flush(self) -> None:
self._orig.flush()
def __getattr__(self, name: str) -> Any:
return getattr(self._orig, name)
engine_log = _EngineLogBuffer()
_available_models: list[dict[str, str]] = [{
"id": path,
"label": model_label(path)
"label": _get_model_label(path)
} for path in get_registered_model_paths()]
job_runner: JobRunner
@@ -157,19 +100,6 @@ def update_settings(settings: SettingsUpdate) -> dict[str, Any]:
return database.get_settings()
@app.get("/api/gpus")
def list_gpus() -> dict[str, Any]:
"""Return an NVML snapshot of every GPU on the API server host."""
return get_gpu_snapshot()
@app.get("/api/cluster")
def cluster_status() -> dict[str, Any]:
"""Cluster-wide GPU/host telemetry (per-node NVML via ray when connected,
the local host otherwise)."""
return get_cluster_snapshot()
@app.get("/api/models")
def list_models(workload_type: str | None = None) -> list[dict[str, Any]]:
"""Return the catalogue of available video-generation models.
@@ -184,25 +114,6 @@ def list_models(workload_type: str | None = None) -> list[dict[str, Any]]:
return _available_models
_PRESET_FIELDS = ("height", "width", "num_frames", "fps", "num_inference_steps", "guidance_scale", "guidance_rescale",
"negative_prompt", "seed")
_preset_cache: dict[str, dict[str, Any]] = {}
@app.get("/api/models/presets")
def model_presets(model_id: str) -> dict[str, Any]:
"""The model's recommended sampling settings (config-only — never loads
weights). The UI populates the job form from these on model selection."""
if model_id not in _preset_cache:
from fastvideo.api.sampling_param import SamplingParam
try:
sp = SamplingParam.from_pretrained(model_id)
except Exception as exc: # noqa: BLE001 -- unknown/unresolvable model
raise HTTPException(status_code=404, detail=f"No presets for '{model_id}': {exc}") from exc
_preset_cache[model_id] = {f: getattr(sp, f) for f in _PRESET_FIELDS if getattr(sp, f, None) is not None}
return _preset_cache[model_id]
ALLOWED_IMAGE_EXTENSIONS = {".png", ".jpg", ".jpeg", ".webp", ".bmp"}
@@ -237,7 +148,7 @@ async def upload_image(file: Annotated[UploadFile, File()], ) -> dict[str, str]:
return {"path": os.path.abspath(dest_path)}
ALLOWED_VIDEO_EXTENSIONS = {".mp4", ".webm", ".avi", ".mov", ".mkv"}
ALLOWED_VIDEO_EXTENSIONS = {".mp4", ".webm", ".avi", ".mov"}
def _filter_video_files(files: list[UploadFile]) -> list[UploadFile]:
@@ -245,23 +156,6 @@ def _filter_video_files(files: list[UploadFile]) -> list[UploadFile]:
return [f for f in files if Path(f.filename or "").suffix.lower() in ALLOWED_VIDEO_EXTENSIONS]
def _path_is_within(child: str, parent: str) -> bool:
"""True if ``child`` resolves to a location inside ``parent``."""
try:
parent_real = os.path.realpath(parent)
return os.path.commonpath([os.path.realpath(child), parent_real]) == parent_real
except (ValueError, OSError):
return False
def _staging_base_path() -> str:
"""Root under which raw uploads are staged (settings override or default)."""
settings = database.get_settings() if database is not None else {}
raw_path = (settings.get("datasetUploadPath") or settings.get("dataset_upload_path") or "")
base_path = (raw_path.strip() if raw_path and isinstance(raw_path, str) else "")
return datasets_upload_dir if not base_path else os.path.abspath(base_path)
@app.post("/api/upload-raw-dataset")
async def upload_raw_dataset(files: Annotated[list[UploadFile], File()], ) -> dict[str, Any]:
"""
@@ -273,7 +167,10 @@ async def upload_raw_dataset(files: Annotated[list[UploadFile], File()], ) -> di
status_code=503,
detail="Database not initialized",
)
base_path = _staging_base_path()
settings = database.get_settings()
raw_path = (settings.get("datasetUploadPath") or settings.get("dataset_upload_path") or "")
base_path = (raw_path.strip() if raw_path and isinstance(raw_path, str) else "")
base_path = (datasets_upload_dir if not base_path else os.path.abspath(base_path))
if not base_path:
raise HTTPException(
status_code=503,
@@ -340,77 +237,19 @@ def get_job(job_id: str) -> dict[str, Any]:
return job.to_dict()
@app.get("/api/generators")
def list_generators() -> list[dict[str, Any]]:
"""Generators resident in memory plus preloads in flight or failed."""
return job_runner.list_generators()
@app.post("/api/generators/preload", status_code=202)
def preload_generator(req: GeneratorRequest) -> dict[str, Any]:
"""Load a model into memory ahead of time (replacing whatever is
resident). One load at a time — 409 while another load is in flight."""
valid_ids = {m["id"] for m in _available_models}
if req.model_id not in valid_ids and not os.path.isdir(req.model_id):
raise HTTPException(
status_code=400,
detail=(f"Unknown model_id '{req.model_id}'. "
f"Valid options: {sorted(valid_ids)}"),
)
try:
return job_runner.preload_generator(**req.model_dump())
except RuntimeError as exc:
raise HTTPException(status_code=409, detail=str(exc)) from exc
@app.post("/api/generators/unload")
def unload_generator() -> dict[str, Any]:
"""Shut down and delete the resident generator, freeing GPU memory."""
try:
unloaded = job_runner.unload_generator()
except RuntimeError as exc:
raise HTTPException(status_code=409, detail=str(exc)) from exc
if not unloaded:
raise HTTPException(status_code=404, detail="No model is loaded")
return {"unloaded": True}
@app.get("/api/engine/logs")
def engine_logs(after: int = 0) -> dict[str, Any]:
"""Incremental tail of the engine's stdout/stderr (driver + relayed
worker output). Poll with ?after=<total from the previous response>."""
lines, total = engine_log.get_lines(after=after)
return {"lines": lines, "total": total}
@app.post("/api/jobs", status_code=201)
def create_job(req: CreateJobRequest) -> dict[str, Any]:
"""Create a new job (does **not** start it automatically)."""
job_type = req.job_type or "inference"
if job_type == "inference":
valid_ids = {m["id"] for m in _available_models}
# A local weights directory is as valid as a registered hub id —
# the registry resolves the pipeline from its model_index.
if req.model_id not in valid_ids and not os.path.isdir(req.model_id):
if req.model_id not in valid_ids:
raise HTTPException(
status_code=400,
detail=(f"Unknown model_id '{req.model_id}'. "
f"Valid options: {sorted(valid_ids)}"),
)
# Training jobs reference a dataset by id; resolve it to the on-disk media
# directory the trainer reads (the UI has no free-text path field). Falls
# through unchanged for inference and for anything already a real path.
def _resolve_dataset_path(value: str) -> str:
media_dir = _dataset_media_dir(value) if value else ""
return media_dir if media_dir and os.path.isdir(media_dir) else value
data_path = req.data_path or ""
validation_dataset_file = req.validation_dataset_file or ""
if job_type != "inference":
data_path = _resolve_dataset_path(data_path)
validation_dataset_file = _resolve_dataset_path(validation_dataset_file)
job = job_runner.create_job(
job_id=str(uuid.uuid4()),
model_id=req.model_id,
@@ -418,13 +257,14 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
workload_type=req.workload_type or "t2v",
job_type=job_type,
image_path=req.image_path or "",
data_path=data_path,
data_path=req.data_path or "",
max_train_steps=req.max_train_steps,
train_batch_size=req.train_batch_size,
learning_rate=req.learning_rate,
num_latent_t=req.num_latent_t,
validation_dataset_file=validation_dataset_file,
validation_dataset_file=req.validation_dataset_file or "",
lora_rank=req.lora_rank,
ltx2_first_frame_conditioning_p=req.ltx2_first_frame_conditioning_p,
negative_prompt=req.negative_prompt,
num_inference_steps=req.num_inference_steps,
num_frames=req.num_frames,
@@ -447,6 +287,8 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
dmd_use_vsa=req.dmd_use_vsa,
dmd_vsa_sparsity=req.dmd_vsa_sparsity,
dmd_denoising_steps=req.dmd_denoising_steps or "1000,757,522",
min_timestep_ratio=req.min_timestep_ratio,
max_timestep_ratio=req.max_timestep_ratio,
real_score_guidance_scale=req.real_score_guidance_scale,
generator_update_interval=req.generator_update_interval,
real_score_model_path=req.real_score_model_path or "",
@@ -576,13 +418,6 @@ def create_dataset(req: CreateDatasetRequest) -> dict[str, Any]:
status_code=400,
detail="No media files found. Ensure at least one image or video.",
)
# upload_path must be a staging dir produced by /api/upload-raw-dataset,
# not an arbitrary client path — the code below copies then rmtrees it.
if not _path_is_within(req.upload_path, _staging_base_path()):
raise HTTPException(
status_code=400,
detail="upload_path must be a staged upload directory.",
)
dataset_id = str(uuid.uuid4())
created_at = time.time()
dest_dir = os.path.join(datasets_upload_dir, dataset_id)
@@ -656,9 +491,8 @@ def serve_dataset_media(dataset_id: str, file_name: str) -> FileResponse:
ds = database.get_dataset(dataset_id)
if ds is None:
raise HTTPException(status_code=404, detail="Dataset not found")
media_dir = _dataset_media_dir(dataset_id)
media_path = os.path.realpath(os.path.join(media_dir, file_name))
if not (_path_is_within(media_path, media_dir) and os.path.isfile(media_path)):
media_path = os.path.join(_dataset_media_dir(dataset_id), file_name)
if not os.path.isfile(media_path):
raise HTTPException(status_code=404, detail="File not found")
import mimetypes
mime, _ = mimetypes.guess_type(media_path)
@@ -754,14 +588,6 @@ def create_local_env(host: str, port: int) -> None:
def main() -> None:
import sys
sys.stdout = _Tee(sys.stdout, engine_log)
sys.stderr = _Tee(sys.stderr, engine_log)
# handlers created before the tee (module-level basicConfig) hold the
# original stream objects — re-point them or their output bypasses the buffer
for h in logging.getLogger().handlers:
if isinstance(h, logging.StreamHandler) and h.stream in (sys.__stderr__, sys.__stdout__):
h.setStream(sys.stderr) # type: ignore[arg-type] # duck-typed file-like
global job_runner, database, upload_dir, verbose, datasets_upload_dir # noqa: PLW0603
# Set up signal handlers to prevent worker crashes from killing the server
@@ -826,9 +652,6 @@ def main() -> None:
verbose=args.verbose,
database=database,
)
# ray relays worker output (incl. denoising tqdm) to the driver's stdout;
# feed those lines to the running job so the UI progress bar moves.
engine_log.on_line = job_runner.feed_engine_line
logger.info("Output directory: %s", output_dir)
logger.info("Log directory: %s", log_dir)
@@ -838,10 +661,6 @@ def main() -> None:
host=args.host,
port=args.port,
log_level="info",
# The engine-output console tails this process's stdout/stderr; the
# frontend polls several endpoints every few seconds, so access-log
# lines are pure self-noise there. App/job logging is unaffected.
access_log=False,
)
+66
View File
@@ -0,0 +1,66 @@
/* CSS Variables */
:root {
--header-height: 74px;
--bg: #0f1117;
--surface: #1a1d27;
--border: #2a2d3a;
--text: #e4e4e7;
--text-dim: #9ca3af;
--accent: #356cff;
--accent-h: #818cf8;
--green: #22c55e;
--red: #ef4444;
--yellow: #eab308;
--blue: #3b82f6;
--radius: 8px;
--shadow: 0 2px 8px rgba(0, 0, 0, 0.35);
}
/* Reset */
*,
*::before,
*::after {
box-sizing: border-box;
margin: 0;
padding: 0;
}
/* Base body styles */
body {
font-family:
"Inter",
system-ui,
-apple-system,
sans-serif;
background: var(--bg);
color: var(--text);
line-height: 1.6;
min-height: 100vh;
display: flex;
flex-direction: column;
}
/* Base form input styles (used globally) */
select,
input,
textarea {
background: var(--bg);
color: var(--text);
border: 1px solid var(--border);
border-radius: var(--radius);
padding: 0.55rem 0.75rem;
font-size: 0.9rem;
font-family: inherit;
outline: none;
transition: border-color 0.15s;
}
select:focus,
input:focus,
textarea:focus {
border-color: var(--accent);
}
textarea {
resize: vertical;
}
+13
View File
@@ -0,0 +1,13 @@
<!doctype html>
<html lang="en">
<head>
<meta charset="utf-8" />
<link rel="icon" href="%sveltekit.assets%/fastvideo.ico" />
<meta name="viewport" content="width=device-width, initial-scale=1.0" />
<title>FastVideo</title>
%sveltekit.head%
</head>
<body class="antialiased" data-sveltekit-preload-data="hover">
<div style="display: contents">%sveltekit.body%</div>
</body>
</html>
@@ -1,64 +0,0 @@
import { act, fireEvent, render, screen } from '@testing-library/react';
import { describe, expect, it, vi } from 'vitest';
import { HeaderActionsProvider } from '@/components/shell/HeaderActionsContext';
import { getDatasets, type Dataset } from '@/lib/api';
import DatasetsPage from './page';
vi.mock('@/lib/api', () => ({
getDatasets: vi.fn(),
}));
vi.mock('@/components/datasets/AddDatasetButton', () => ({
default: () => null,
}));
vi.mock('@/components/datasets/CreateDatasetModal', () => ({
default: () => null,
}));
vi.mock('@/components/datasets/DatasetCard', () => ({
default: ({ dataset }: { dataset: Dataset }) => <div>{dataset.name}</div>,
}));
function renderPage() {
return render(
<HeaderActionsProvider>
<DatasetsPage />
</HeaderActionsProvider>,
);
}
describe('DatasetsPage', () => {
it('shows loading content before the initial request settles', async () => {
let resolveDatasets: (datasets: Dataset[]) => void = () => {};
vi.mocked(getDatasets).mockReturnValue(
new Promise<Dataset[]>((resolve) => {
resolveDatasets = resolve;
}),
);
renderPage();
expect(screen.getByLabelText('Loading datasets')).toBeInTheDocument();
expect(screen.queryByText('No datasets yet.')).not.toBeInTheDocument();
act(() => resolveDatasets([]));
expect(await screen.findByText('No datasets yet.')).toBeInTheDocument();
});
it('shows API failures separately from an empty list and retries', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
vi.mocked(getDatasets).mockRejectedValueOnce(new Error('network down'));
renderPage();
expect(
await screen.findByText(/Could not load datasets from the Studio API/),
).toBeInTheDocument();
expect(screen.queryByText('No datasets yet.')).not.toBeInTheDocument();
vi.mocked(getDatasets).mockResolvedValueOnce([]);
fireEvent.click(screen.getByRole('button', { name: 'Try Again' }));
expect(await screen.findByText('No datasets yet.')).toBeInTheDocument();
});
});
@@ -1,136 +0,0 @@
'use client';
import * as React from 'react';
import { AlertTriangle } from 'lucide-react';
import AddDatasetButton from '@/components/datasets/AddDatasetButton';
import CreateDatasetModal from '@/components/datasets/CreateDatasetModal';
import DatasetCard from '@/components/datasets/DatasetCard';
import { HeaderActions } from '@/components/shell/HeaderActionsContext';
import { Card } from '@/components/ui/card';
import { Button } from '@/components/ui/button';
import { useStore } from '@/hooks/useStore';
import { getDatasets } from '@/lib/api';
import type { Dataset } from '@/lib/api';
import {
setActiveDataset,
setActiveDatasetId,
} from '@/stores/activeDataset';
import {
createDatasetModalStore,
setCreateDatasetModalOpen,
} from '@/stores/createDatasetModalOpen';
export default function DatasetsPage() {
const [datasets, setDatasets] = React.useState<Dataset[]>([]);
const [isInitialLoading, setIsInitialLoading] = React.useState(true);
const [error, setError] = React.useState<string | null>(null);
const { open } = useStore(createDatasetModalStore);
const fetchSequence = React.useRef(0);
const fetchDatasets = React.useCallback(async () => {
const sequence = ++fetchSequence.current;
try {
const next = await getDatasets();
if (sequence === fetchSequence.current) {
setDatasets(next);
setError(null);
}
} catch (err) {
console.error('Failed to fetch datasets:', err);
if (sequence === fetchSequence.current) {
setError(
'Could not load datasets from the Studio API. Check the server and try again.',
);
}
} finally {
if (sequence === fetchSequence.current) setIsInitialLoading(false);
}
}, []);
React.useEffect(() => {
fetchDatasets();
}, [fetchDatasets]);
function handleSelectDataset(ds: Dataset) {
setActiveDataset(ds);
setActiveDatasetId(ds.id);
}
return (
<>
<HeaderActions>
<AddDatasetButton />
</HeaderActions>
<div className="mx-auto flex w-full max-w-[850px] flex-col gap-6 px-4 pb-12">
<Card className="p-6">
<div aria-busy={isInitialLoading}>
{isInitialLoading ? (
<div
aria-label="Loading datasets"
className="flex flex-col gap-3 py-2"
>
{[0, 1, 2].map((item) => (
<div
key={item}
className="h-24 animate-pulse rounded-lg border border-border bg-muted/50"
/>
))}
</div>
) : error && datasets.length === 0 ? (
<div
role="alert"
className="flex flex-col items-center gap-3 py-8 text-center"
>
<AlertTriangle
className="size-6 text-destructive"
aria-hidden
/>
<p className="max-w-md text-sm text-muted-foreground">
{error}
</p>
<Button type="button" variant="outline" onClick={fetchDatasets}>
Try Again
</Button>
</div>
) : (
<>
{error && (
<p
role="status"
className="mb-3 rounded-lg border border-amber-500/50 bg-amber-500/10 px-3 py-2 text-sm text-foreground"
>
Dataset updates are temporarily unavailable. Showing the
most recent results.
</p>
)}
{datasets.length === 0 ? (
<p className="py-8 text-center text-muted-foreground">
No datasets yet.
</p>
) : (
datasets.map((ds) => (
<DatasetCard
key={ds.id}
dataset={ds}
onUpdated={fetchDatasets}
onSelect={() => handleSelectDataset(ds)}
/>
))
)}
</>
)}
</div>
</Card>
</div>
<CreateDatasetModal
isOpen={open}
onClose={() => setCreateDatasetModalOpen(false)}
onSuccess={() => {
fetchDatasets();
setCreateDatasetModalOpen(false);
}}
/>
</>
);
}
@@ -1,16 +0,0 @@
'use client';
import CreateJobButton from '@/components/jobs/CreateJobButton';
import { HeaderActions } from '@/components/shell/HeaderActionsContext';
import JobQueue from '@/components/jobs/JobQueue';
export default function DistillationPage() {
return (
<>
<HeaderActions>
<CreateJobButton jobType="distillation" />
</HeaderActions>
<JobQueue jobType="distillation" />
</>
);
}
@@ -1,21 +0,0 @@
'use client';
import CreateJobButton from '@/components/jobs/CreateJobButton';
import { HeaderActions } from '@/components/shell/HeaderActionsContext';
import JobQueue from '@/components/jobs/JobQueue';
import type { JobType } from '@/lib/types';
// 'lora' is a backend job_type (LoRA finetunes) that isn't part of the
// JobType union; the finetuning queue lists both alongside full finetunes.
const FINETUNING_LIST = ['finetuning', 'lora'] as JobType[];
export default function FinetuningPage() {
return (
<>
<HeaderActions>
<CreateJobButton jobType="finetuning" />
</HeaderActions>
<JobQueue jobType="finetuning" jobTypesForList={FINETUNING_LIST} />
</>
);
}
@@ -1,76 +0,0 @@
import { fireEvent, render, screen } from '@testing-library/react';
import { describe, expect, it, vi } from 'vitest';
import GalleryPage from './page';
import { HeaderActionsProvider } from '@/components/shell/HeaderActionsContext';
import { getJobsList } from '@/lib/api';
import type { Job } from '@/lib/types';
import { makeJob as makeBaseJob } from '@/test/factories';
vi.mock('@/lib/api', () => ({
getJobsList: vi.fn(),
getJobVideoUrl: (id: string) => `http://test.local/api/jobs/${id}/video`,
}));
const makeJob = (overrides: Partial<Job> = {}): Job =>
makeBaseJob({
model_id: 'wan',
prompt: 'a cat surfing a wave',
status: 'completed',
created_at: 1,
finished_at: 2,
output_path: '/out/clip.mp4',
...overrides,
});
function renderGallery() {
return render(
<HeaderActionsProvider>
<GalleryPage />
</HeaderActionsProvider>,
);
}
describe('GalleryPage', () => {
it('renders a grid item for a completed inference job', async () => {
vi.mocked(getJobsList).mockResolvedValue([
makeJob({ prompt: 'a cat surfing a wave' }),
]);
renderGallery();
expect(await screen.findByText('a cat surfing a wave')).toBeInTheDocument();
expect(getJobsList).toHaveBeenCalledWith('inference');
});
it('provides video controls and a visible fallback when media fails', async () => {
vi.mocked(getJobsList).mockResolvedValue([makeJob()]);
renderGallery();
const video = await screen.findByLabelText(
'Generated video: a cat surfing a wave',
);
expect(video).toHaveAttribute('controls');
fireEvent.error(video);
expect(screen.getByText('Preview unavailable')).toBeInTheDocument();
expect(
screen.getByText('The generated file could not be loaded.'),
).toBeInTheDocument();
});
it('shows the empty state when no completed videos exist', async () => {
vi.mocked(getJobsList).mockResolvedValue([
makeJob({ status: 'running', output_path: null }),
]);
renderGallery();
expect(
await screen.findByText('No completed videos yet'),
).toBeInTheDocument();
expect(
screen.queryByText('a cat surfing a wave'),
).not.toBeInTheDocument();
});
});
@@ -1,161 +0,0 @@
'use client';
import { AlertTriangle, ImageOff, Loader2 } from 'lucide-react';
import { useEffect, useState } from 'react';
import { Button } from '@/components/ui/button';
import { Card } from '@/components/ui/card';
import { getJobVideoUrl, getJobsList } from '@/lib/api';
import type { Job } from '@/lib/types';
function isImage(job: Job): boolean {
return job.output_path?.toLowerCase().endsWith('.png') ?? false;
}
function GalleryMedia({ job }: { job: Job }) {
const [failed, setFailed] = useState(false);
if (failed) {
return (
<div
role="status"
className="flex h-full flex-col items-center justify-center gap-2 px-4 text-center text-muted-foreground"
>
<ImageOff className="size-7" aria-hidden />
<span className="text-sm font-medium">Preview unavailable</span>
<span className="text-xs">
The generated file could not be loaded.
</span>
</div>
);
}
if (isImage(job)) {
return (
// eslint-disable-next-line @next/next/no-img-element
<img
src={getJobVideoUrl(job.id)}
alt={job.prompt}
className="block h-full w-full object-contain"
loading="lazy"
onError={() => setFailed(true)}
/>
);
}
return (
<video
src={getJobVideoUrl(job.id)}
aria-label={
job.prompt ? `Generated video: ${job.prompt}` : 'Generated video'
}
className="block h-full w-full object-contain"
controls
muted
loop
playsInline
preload="metadata"
onError={() => setFailed(true)}
/>
);
}
export default function GalleryPage() {
const [jobs, setJobs] = useState<Job[]>([]);
const [isLoading, setIsLoading] = useState(true);
const [error, setError] = useState<string | null>(null);
const [reloadKey, setReloadKey] = useState(0);
useEffect(() => {
let cancelled = false;
async function load() {
try {
const list = await getJobsList('inference');
if (cancelled) return;
const sorted = [...list].sort(
(a, b) =>
(b.finished_at ?? b.created_at ?? 0) -
(a.finished_at ?? a.created_at ?? 0),
);
setJobs(sorted);
} catch (e) {
if (!cancelled) {
setError(e instanceof Error ? e.message : 'Failed to load jobs');
}
} finally {
if (!cancelled) setIsLoading(false);
}
}
void load();
return () => {
cancelled = true;
};
}, [reloadKey]);
function retry() {
setError(null);
setIsLoading(true);
setReloadKey((k) => k + 1);
}
const galleryJobs = jobs.filter(
(j) =>
j.status === 'completed' &&
j.output_path &&
(j.job_type === 'inference' || !j.job_type),
);
return (
<div className="mx-auto w-full max-w-[1200px] px-4 pb-12 pt-4">
<Card className="p-6">
<h2 className="mb-1 text-2xl font-semibold text-foreground">Gallery</h2>
<p className="mb-6 text-sm text-muted-foreground">
Generated videos from completed inference jobs. Captions show the
prompt used for each generation.
</p>
{isLoading ? (
<div className="flex items-center gap-3 p-8 text-muted-foreground">
<Loader2 className="h-6 w-6 animate-spin text-primary" />
<span>Loading gallery…</span>
</div>
) : error ? (
<div
role="alert"
className="flex flex-col items-center gap-3 py-8 text-center"
>
<AlertTriangle className="size-6 text-destructive" aria-hidden />
<p className="max-w-md text-sm text-muted-foreground">{error}</p>
<Button type="button" variant="outline" onClick={retry}>
Try Again
</Button>
</div>
) : galleryJobs.length === 0 ? (
<p className="py-8 text-center text-muted-foreground">
No completed videos yet
</p>
) : (
<div className="grid grid-cols-[repeat(auto-fill,minmax(280px,1fr))] gap-5">
{galleryJobs.map((job) => (
<article
key={job.id}
className="flex flex-col overflow-hidden rounded-lg border border-border bg-background"
>
<div className="relative aspect-video overflow-hidden bg-muted">
<GalleryMedia job={job} />
</div>
<p
className="line-clamp-3 border-t border-border px-4 py-3 text-sm text-muted-foreground"
title={job.prompt}
>
{job.prompt || '—'}
</p>
</article>
))}
</div>
)}
</Card>
</div>
);
}
-238
View File
@@ -1,238 +0,0 @@
@import "tailwindcss";
@custom-variant dark (&:is(.dark *));
/*
* FastVideo Studio theme — shared with the Dreamverse design system:
* slate light/dark palettes, #356cff accent blue, IBM Plex type. The studio
* defaults to dark (see the theme init script in layout.tsx); the toggle in
* the header persists the choice to localStorage.
*/
:root {
color-scheme: light;
--header-height: 74px;
--accent-blue: #356cff;
--background: #f5f4f4;
--foreground: #0f172a;
--card: #ffffff;
--card-foreground: #0f172a;
--popover: #ffffff;
--popover-foreground: #0f172a;
--primary: #0f172a;
--primary-foreground: #f8fafc;
--secondary: #f1f5f9;
--secondary-foreground: #1e293b;
--muted: #f1f5f9;
--muted-foreground: #64748b;
--accent: #e2e8f0;
--accent-foreground: #0f172a;
--destructive: #ef4444;
--destructive-foreground: #ffffff;
--border: #e2e8f0;
--input: #cbd5e1;
--ring: #1d4ed8;
--radius: 0.5rem;
}
.dark {
color-scheme: dark;
--accent-blue: #356cff;
--background: #0f172a;
--foreground: #e2e8f0;
--card: #0f172a;
--card-foreground: #f1f5f9;
--popover: #0f172a;
--popover-foreground: #f1f5f9;
--primary: #f1f5f9;
--primary-foreground: #0f172a;
--secondary: #1e293b;
--secondary-foreground: #f1f5f9;
--muted: #1e293b;
--muted-foreground: #94a3b8;
--accent: #1e293b;
--accent-foreground: #f1f5f9;
--destructive: #991b1b;
--destructive-foreground: #fecaca;
--border: #334155;
--input: #334155;
--ring: #7dd3fc;
}
@theme inline {
--font-sans: var(--font-plex-sans), ui-sans-serif, system-ui, sans-serif;
--font-mono: var(--font-plex-mono), ui-monospace, monospace;
--color-accent-blue: var(--accent-blue);
--color-background: var(--background);
--color-foreground: var(--foreground);
--color-card: var(--card);
--color-card-foreground: var(--card-foreground);
--color-popover: var(--popover);
--color-popover-foreground: var(--popover-foreground);
--color-primary: var(--primary);
--color-primary-foreground: var(--primary-foreground);
--color-secondary: var(--secondary);
--color-secondary-foreground: var(--secondary-foreground);
--color-muted: var(--muted);
--color-muted-foreground: var(--muted-foreground);
--color-accent: var(--accent);
--color-accent-foreground: var(--accent-foreground);
--color-destructive: var(--destructive);
--color-destructive-foreground: var(--destructive-foreground);
--color-border: var(--border);
--color-input: var(--input);
--color-ring: var(--ring);
--radius-sm: calc(var(--radius) - 4px);
--radius-md: calc(var(--radius) - 2px);
--radius-lg: var(--radius);
--radius-xl: calc(var(--radius) + 4px);
}
/* ——— Resets ——— */
html,
body {
min-height: 100%;
overflow-x: clip;
}
html {
background: var(--background);
}
body {
margin: 0;
line-height: 1.6;
color: var(--foreground);
font-family: var(--font-plex-sans), ui-sans-serif, system-ui, sans-serif;
background: var(--background);
}
/* Dreamverse's dark-mode backdrop: subtle radial glows over a deep fade. */
body::after {
content: "";
position: fixed;
inset: 0;
z-index: -1;
opacity: 0;
pointer-events: none;
transition: opacity 300ms ease;
background:
radial-gradient(circle at top, rgba(56, 189, 248, 0.14), transparent 36%),
radial-gradient(circle at right top, rgba(129, 140, 248, 0.12), transparent 28%),
linear-gradient(180deg, #020617 0%, #000000 100%);
background-attachment: fixed;
}
.dark body::after {
opacity: 1;
}
a {
color: inherit;
}
:where(
a,
button,
input,
textarea,
select,
summary,
[role="button"],
[role="menuitem"],
[role="slider"],
[tabindex]
):focus-visible {
outline: 3px solid var(--ring) !important;
outline-offset: 2px !important;
}
summary {
list-style: none;
}
summary::-webkit-details-marker {
display: none;
}
::selection {
background: rgba(56, 189, 248, 0.25);
}
.dark ::selection {
background: rgba(56, 189, 248, 0.35);
color: #ffffff;
}
/* Smooth the light/dark switch (class applied briefly by the theme toggle). */
html.theme-transition,
html.theme-transition *,
html.theme-transition *::before,
html.theme-transition *::after {
transition:
background-color 300ms ease,
color 300ms ease,
border-color 300ms ease,
box-shadow 300ms ease,
fill 300ms ease,
stroke 300ms ease !important;
}
/* ——— Base layer ——— */
@layer base {
button,
input,
textarea,
select {
font: inherit;
}
img,
svg,
video,
canvas {
display: block;
max-width: 100%;
}
* {
@apply border-border;
}
body {
@apply bg-background text-foreground;
}
}
@@ -1,55 +0,0 @@
import { readFileSync } from 'node:fs';
import { join } from 'node:path';
import { describe, expect, it } from 'vitest';
const css = readFileSync(join(process.cwd(), 'src/app/globals.css'), 'utf8');
function token(block: string, name: string): string {
const match = block.match(new RegExp(`--${name}:\\s*(#[0-9a-fA-F]{6})`));
if (!match) throw new Error(`Missing --${name} token`);
return match[1];
}
function luminance(hex: string): number {
const channels = hex
.slice(1)
.match(/.{2}/g)!
.map((channel) => parseInt(channel, 16) / 255)
.map((channel) =>
channel <= 0.04045
? channel / 12.92
: ((channel + 0.055) / 1.055) ** 2.4,
);
return (
0.2126 * channels[0] + 0.7152 * channels[1] + 0.0722 * channels[2]
);
}
function contrast(first: string, second: string): number {
const firstLuminance = luminance(first);
const secondLuminance = luminance(second);
return (
(Math.max(firstLuminance, secondLuminance) + 0.05) /
(Math.min(firstLuminance, secondLuminance) + 0.05)
);
}
describe('global focus styles', () => {
it('keeps focus tokens above 3:1 against both page themes', () => {
const light = css.match(/:root\s*{([\s\S]*?)\n}/)?.[1] ?? '';
const dark = css.match(/\.dark\s*{([\s\S]*?)\n}/)?.[1] ?? '';
expect(contrast(token(light, 'ring'), token(light, 'background'))).toBeGreaterThanOrEqual(
3,
);
expect(contrast(token(dark, 'ring'), token(dark, 'background'))).toBeGreaterThanOrEqual(
3,
);
});
it('applies a non-animated three-pixel outline to focus-visible controls', () => {
expect(css).toContain('):focus-visible {');
expect(css).toContain('outline: 3px solid var(--ring) !important;');
expect(css).toContain('outline-offset: 2px !important;');
});
});
@@ -1,182 +0,0 @@
import { render, screen } from '@testing-library/react';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import GpusPage from './page';
import { getClusterStatus } from '@/lib/api';
import type { ClusterSnapshot } from '@/lib/api';
vi.mock('@/lib/api', () => ({
getClusterStatus: vi.fn(),
}));
const RAY_SNAPSHOT: ClusterSnapshot = {
mode: 'ray',
error: null,
resources: { gpus_total: 8, gpus_available: 5 },
nodes: [
{
hostname: 'node-a',
ip: '10.0.0.10',
is_this_host: true,
cpus: 64,
ray_gpus: 4,
available: true,
error: null,
gpus: [
{
index: 0,
name: 'NVIDIA B200',
utilization: 62,
memory_used_mib: 40_960,
memory_total_mib: 81_920,
temperature_c: 41,
power_watts: 312.4,
power_limit_watts: 1000,
},
{
index: 1,
name: 'NVIDIA B200',
utilization: 0,
memory_used_mib: 1_024,
memory_total_mib: 81_920,
temperature_c: null,
power_watts: null,
power_limit_watts: null,
},
],
},
{
hostname: 'node-b',
ip: '10.0.0.11',
is_this_host: false,
cpus: 32,
ray_gpus: 2,
available: true,
error: null,
gpus: [
{
index: 0,
name: 'NVIDIA B200',
utilization: 90,
memory_used_mib: 20_480,
memory_total_mib: 81_920,
temperature_c: 70,
power_watts: 900,
power_limit_watts: 1000,
},
],
},
],
};
const LOCAL_SNAPSHOT: ClusterSnapshot = {
mode: 'local',
error:
'not connected to a ray cluster yet (load a model first); showing the API host only',
resources: null,
nodes: [
{
hostname: 'localhost',
ip: null,
is_this_host: true,
cpus: null,
ray_gpus: null,
available: true,
error: null,
gpus: [
{
index: 0,
name: 'NVIDIA RTX 5090',
utilization: 12,
memory_used_mib: 2_048,
memory_total_mib: 32_768,
temperature_c: 38,
power_watts: 80,
power_limit_watts: 575,
},
],
},
],
};
beforeEach(() => {
vi.mocked(getClusterStatus).mockResolvedValue(RAY_SNAPSHOT);
});
describe('GpusPage', () => {
it('renders the header with mode and GPU totals', async () => {
render(<GpusPage />);
expect(await screen.findByText('ray cluster')).toBeInTheDocument();
expect(screen.getByText(/5 \/\s*8 GPUs available/)).toBeInTheDocument();
});
it('renders a section per node with host details and GPU rows', async () => {
render(<GpusPage />);
// Each hostname appears twice: once in the strip, once as a section.
expect(await screen.findAllByText('node-a')).toHaveLength(2);
expect(screen.getAllByText('node-b')).toHaveLength(2);
expect(screen.getByText('10.0.0.10')).toBeInTheDocument();
// Only node-a is the API host.
expect(screen.getAllByText('API host')).toHaveLength(1);
expect(screen.getByText(/64 CPUs · 4 ray GPUs/)).toBeInTheDocument();
expect(screen.getAllByText('NVIDIA B200')).toHaveLength(3);
expect(screen.getByText('GPU 1')).toBeInTheDocument();
expect(screen.getByText('62%')).toBeInTheDocument();
expect(
screen.getByText('40960 / 81920 MiB (40.0 GiB / 80.0 GiB)'),
).toBeInTheDocument();
// Optional sensors render only when present.
expect(screen.getByText('41°C')).toBeInTheDocument();
expect(screen.getByText('312 W / 1000 W')).toBeInTheDocument();
});
it('bars reflect utilization and VRAM values', async () => {
render(<GpusPage />);
await screen.findAllByText('node-a');
const utilMeters = screen
.getAllByRole('meter', { name: 'Utilization' })
.map((m) => m.getAttribute('aria-valuenow'));
expect(utilMeters).toEqual(['62', '0', '90']);
const vramMeters = screen
.getAllByRole('meter', { name: 'VRAM' })
.map((m) => m.getAttribute('aria-valuenow'));
// 40960/81920 = 50%, 1024/81920 ≈ 1%, 20480/81920 = 25%
expect(vramMeters).toEqual(['50', '1', '25']);
});
it('renders the compact strip with per-GPU segments', async () => {
render(<GpusPage />);
await screen.findAllByText('node-a');
const segments = screen.getAllByRole('img');
expect(segments).toHaveLength(3);
expect(segments[0]).toHaveAccessibleName(
'GPU 0: 62% utilization, 40.0 GiB / 80.0 GiB VRAM',
);
});
it('shows the informational banner and local mode', async () => {
vi.mocked(getClusterStatus).mockResolvedValue(LOCAL_SNAPSHOT);
render(<GpusPage />);
expect(await screen.findByText('local host only')).toBeInTheDocument();
expect(
screen.getByText(/not connected to a ray cluster yet/),
).toBeInTheDocument();
// No resources in local mode.
expect(screen.queryByText(/GPUs available/)).not.toBeInTheDocument();
});
it('explains when the API server is unreachable', async () => {
vi.mocked(getClusterStatus).mockRejectedValue(new Error('network down'));
render(<GpusPage />);
expect(
await screen.findByText(/Could not reach the API server/),
).toBeInTheDocument();
});
});
-258
View File
@@ -1,258 +0,0 @@
'use client';
import * as React from 'react';
import { AlertTriangle, Info } from 'lucide-react';
import ClusterStrip, {
clampPercent,
formatGib,
utilizationColor,
} from '@/components/cluster/ClusterStrip';
import { Badge } from '@/components/ui/badge';
import { Button } from '@/components/ui/button';
import { Card, CardContent } from '@/components/ui/card';
import {
getClusterStatus,
type ClusterGpu,
type ClusterNode,
type ClusterSnapshot,
} from '@/lib/api';
import { cn } from '@/lib/utils';
const POLL_INTERVAL_MS = 5000;
function Meter({
label,
percent,
detail,
fillClass,
}: {
label: string;
percent: number;
detail: string;
fillClass: string;
}) {
const clamped = clampPercent(percent);
return (
<div className="flex flex-col gap-1">
<div className="flex items-baseline justify-between gap-2 text-xs">
<span className="text-muted-foreground">{label}</span>
<span className="font-medium tabular-nums text-foreground">
{detail}
</span>
</div>
<div
role="meter"
aria-label={label}
aria-valuenow={Math.round(clamped)}
aria-valuemin={0}
aria-valuemax={100}
className="h-1.5 overflow-hidden rounded-full bg-muted"
>
<div
className={cn(
'h-full rounded-full transition-[width] duration-500',
fillClass,
)}
style={{ width: `${clamped}%` }}
/>
</div>
</div>
);
}
function GpuRow({ gpu }: { gpu: ClusterGpu }) {
const memPercent =
gpu.memory_total_mib > 0
? (gpu.memory_used_mib / gpu.memory_total_mib) * 100
: 0;
return (
<div className="grid items-center gap-x-6 gap-y-2 border-t border-border pt-3 first:border-t-0 first:pt-0 md:grid-cols-[minmax(0,1fr)_minmax(0,1.2fr)_minmax(0,1.6fr)_auto]">
<div className="flex min-w-0 items-baseline gap-2">
<span className="min-w-0 truncate text-sm font-semibold">
{gpu.name}
</span>
<span className="shrink-0 text-xs font-medium uppercase tracking-wider text-muted-foreground">
GPU {gpu.index}
</span>
</div>
<Meter
label="Utilization"
percent={gpu.utilization}
detail={`${gpu.utilization}%`}
fillClass={utilizationColor(gpu.utilization)}
/>
<Meter
label="VRAM"
percent={memPercent}
detail={`${gpu.memory_used_mib} / ${gpu.memory_total_mib} MiB (${formatGib(gpu.memory_used_mib)} / ${formatGib(gpu.memory_total_mib)})`}
fillClass={memPercent >= 90 ? 'bg-rose-500' : 'bg-accent-blue'}
/>
<div className="flex flex-wrap gap-x-4 gap-y-1 text-xs tabular-nums text-muted-foreground md:w-28 md:justify-end">
{gpu.temperature_c != null && <span>{gpu.temperature_c}°C</span>}
{gpu.power_watts != null && (
<span>
{Math.round(gpu.power_watts)} W
{gpu.power_limit_watts != null &&
` / ${Math.round(gpu.power_limit_watts)} W`}
</span>
)}
</div>
</div>
);
}
function NodeSection({ node }: { node: ClusterNode }) {
return (
<Card>
<CardContent className="flex flex-col gap-3 p-5">
<div className="flex flex-wrap items-center gap-x-3 gap-y-1">
<span className="min-w-0 truncate text-sm font-semibold">
{node.hostname}
</span>
{node.ip && (
<span className="text-xs tabular-nums text-muted-foreground">
{node.ip}
</span>
)}
{node.is_this_host && <Badge variant="secondary">API host</Badge>}
<span className="ml-auto text-xs tabular-nums text-muted-foreground">
{node.cpus != null && `${Math.round(node.cpus)} CPUs`}
{node.cpus != null && node.ray_gpus != null && ' · '}
{node.ray_gpus != null && `${Math.round(node.ray_gpus)} ray GPUs`}
</span>
</div>
{!node.available && (
<p className="text-sm text-muted-foreground">
GPU telemetry unavailable
{node.error ? `: ${node.error}` : '.'}
</p>
)}
{node.gpus.map((gpu) => (
<GpuRow key={gpu.index} gpu={gpu} />
))}
</CardContent>
</Card>
);
}
export default function GpusPage() {
const [snapshot, setSnapshot] = React.useState<ClusterSnapshot | null>(null);
const [fetchError, setFetchError] = React.useState<string | null>(null);
const [retryToken, setRetryToken] = React.useState(0);
React.useEffect(() => {
let mounted = true;
let inFlight = false;
async function poll() {
if (inFlight || document.hidden) return;
inFlight = true;
try {
const next = await getClusterStatus();
if (mounted) {
setSnapshot(next);
setFetchError(null);
}
} catch {
if (mounted) {
setFetchError(
'Cluster status could not be refreshed. The values below may be stale.',
);
}
} finally {
inFlight = false;
}
}
poll();
const interval = setInterval(poll, POLL_INTERVAL_MS);
// Refresh immediately when the tab becomes visible again (polls are
// skipped while hidden).
document.addEventListener('visibilitychange', poll);
return () => {
mounted = false;
clearInterval(interval);
document.removeEventListener('visibilitychange', poll);
};
}, [retryToken]);
let body: React.ReactNode;
if (fetchError && !snapshot) {
body = (
<div
role="alert"
className="flex flex-col items-center gap-3 py-8 text-center"
>
<AlertTriangle className="size-6 text-destructive" aria-hidden />
<p className="text-muted-foreground">
Could not reach the API server. Cluster status needs the Studio API
server running.
</p>
<Button
type="button"
variant="outline"
onClick={() => setRetryToken((token) => token + 1)}
>
Try Again
</Button>
</div>
);
} else if (!snapshot) {
body = <p className="py-8 text-center text-muted-foreground">Loading…</p>;
} else {
body = (
<div className="flex flex-col gap-4">
<header className="flex flex-wrap items-center gap-3">
<h1 className="text-lg font-semibold">Cluster</h1>
<Badge variant="outline">
{snapshot.mode === 'ray' ? 'ray cluster' : 'local host only'}
</Badge>
{snapshot.resources && (
<span className="text-sm tabular-nums text-muted-foreground">
{Math.round(snapshot.resources.gpus_available)} /{' '}
{Math.round(snapshot.resources.gpus_total)} GPUs available
</span>
)}
</header>
{snapshot.error && (
<div
role="status"
className="flex flex-wrap items-center gap-3 rounded-lg border border-blue-400/40 bg-blue-500/10 px-3 py-2 text-sm"
>
<Info className="size-4 shrink-0 text-blue-600" aria-hidden />
<span className="min-w-0 flex-1">{snapshot.error}</span>
</div>
)}
{fetchError && (
<div
role="status"
aria-live="polite"
className="flex flex-wrap items-center gap-3 rounded-lg border border-amber-500/50 bg-amber-500/10 px-3 py-2 text-sm"
>
<AlertTriangle className="size-4 text-amber-600" aria-hidden />
<span className="min-w-0 flex-1">{fetchError}</span>
<Button
type="button"
variant="outline"
size="sm"
onClick={() => setRetryToken((token) => token + 1)}
>
Refresh Now
</Button>
</div>
)}
<ClusterStrip nodes={snapshot.nodes} />
{snapshot.nodes.map((node, i) => (
<NodeSection key={`${node.hostname}-${i}`} node={node} />
))}
</div>
);
}
return (
<div className="mx-auto flex w-full max-w-[1100px] flex-col gap-6 px-4 pb-12 pt-6">
{body}
</div>
);
}
@@ -1,20 +0,0 @@
'use client';
import CreateJobButton from '@/components/jobs/CreateJobButton';
import EngineConsole from '@/components/jobs/EngineConsole';
import { HeaderActions } from '@/components/shell/HeaderActionsContext';
import JobQueue from '@/components/jobs/JobQueue';
import WarmModelsPanel from '@/components/jobs/WarmModelsPanel';
export default function InferencePage() {
return (
<>
<HeaderActions>
<CreateJobButton jobType="inference" />
</HeaderActions>
<WarmModelsPanel />
<EngineConsole />
<JobQueue jobType="inference" />
</>
);
}
-52
View File
@@ -1,52 +0,0 @@
import type { Metadata } from 'next';
import { IBM_Plex_Mono, IBM_Plex_Sans } from 'next/font/google';
import { AppShell } from '@/components/shell/AppShell';
import './globals.css';
const plexSans = IBM_Plex_Sans({
subsets: ['latin'],
weight: ['400', '500', '600', '700'],
variable: '--font-plex-sans',
display: 'swap',
});
const plexMono = IBM_Plex_Mono({
subsets: ['latin'],
weight: ['400', '500', '600'],
variable: '--font-plex-mono',
display: 'swap',
});
export const metadata: Metadata = {
title: 'FastVideo Studio',
icons: { icon: '/fastvideo.ico' },
};
export default function RootLayout({
children,
}: {
children: React.ReactNode;
}) {
return (
<html
lang="en"
suppressHydrationWarning
className={`${plexSans.variable} ${plexMono.variable}`}
>
<head>
{/* The studio defaults to dark; apply the stored choice before paint. */}
<script
id="theme-init-script"
suppressHydrationWarning
dangerouslySetInnerHTML={{
__html: `(function(){var dark=true;try{dark=localStorage.getItem('theme')!=='light'}catch(e){}if(dark)document.documentElement.classList.add('dark')})()`,
}}
/>
</head>
<body className="antialiased">
<AppShell>{children}</AppShell>
</body>
</html>
);
}
@@ -1,5 +0,0 @@
import { redirect } from 'next/navigation';
export default function Page() {
redirect('/finetuning');
}
-5
View File
@@ -1,5 +0,0 @@
import { redirect } from 'next/navigation';
export default function Page() {
redirect('/inference');
}
@@ -1,74 +0,0 @@
import { fireEvent, render, screen } from '@testing-library/react';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import { HeaderActionsProvider } from '@/components/shell/HeaderActionsContext';
import { resetToDefaults, updateOption } from '@/stores/defaultOptions';
import SettingsPage from './page';
vi.mock('@/lib/api', () => ({
getModels: vi.fn().mockResolvedValue([]),
getSettings: vi.fn().mockResolvedValue({}),
updateSettings: vi.fn().mockResolvedValue({}),
}));
// Keep the real store (so `useStore` resolves) but spy on the mutators.
vi.mock('@/stores/defaultOptions', async (importOriginal) => {
const actual =
await importOriginal<typeof import('@/stores/defaultOptions')>();
return {
...actual,
updateOption: vi.fn(),
resetToDefaults: vi.fn(),
};
});
function renderPage() {
return render(
<HeaderActionsProvider>
<SettingsPage />
</HeaderActionsProvider>,
);
}
describe('Settings page', () => {
beforeEach(() => {
vi.clearAllMocks();
});
it('calls updateOption when a toggle changes', () => {
renderPage();
fireEvent.click(
screen.getByRole('switch', { name: 'Auto Start Job on Create' }),
);
expect(updateOption).toHaveBeenCalledWith('autoStartJob', true);
});
it('calls updateOption when a slider changes', () => {
renderPage();
// First slider in DOM order is "Frames".
const slider = screen.getAllByRole('slider')[0];
fireEvent.keyDown(slider, { key: 'ArrowRight' });
expect(updateOption).toHaveBeenCalledWith('numFrames', expect.any(Number));
});
it('gives every slider an accessible name', () => {
renderPage();
const sliders = screen.getAllByRole('slider');
expect(sliders).toHaveLength(11);
for (const slider of sliders) {
expect(slider).toHaveAccessibleName();
}
expect(screen.getByRole('slider', { name: 'Frames' })).toBeInTheDocument();
expect(
screen.getByRole('slider', { name: 'Guidance Scale' }),
).toBeInTheDocument();
});
it('calls resetToDefaults when Reset to Defaults is clicked', () => {
renderPage();
fireEvent.click(screen.getByRole('button', { name: 'Reset to Defaults' }));
expect(resetToDefaults).toHaveBeenCalledTimes(1);
});
});
@@ -1,358 +0,0 @@
'use client';
import * as React from 'react';
import {
FieldRow,
NumberRow,
SliderRow,
ToggleRow,
} from '@/components/form-rows';
import { Button } from '@/components/ui/button';
import { Card, CardContent } from '@/components/ui/card';
import { Input } from '@/components/ui/input';
import { Label } from '@/components/ui/label';
import { NativeSelect } from '@/components/ui/native-select';
import { Separator } from '@/components/ui/separator';
import { Switch } from '@/components/ui/switch';
import { useStore } from '@/hooks/useStore';
import { getModels, type Model } from '@/lib/api';
import {
defaultOptionsStore,
resetToDefaults,
updateOption,
} from '@/stores/defaultOptions';
function TextSettingRow({
id,
label,
value,
placeholder,
onCommit,
}: {
id: string;
label: string;
value: string;
placeholder?: string;
onCommit: (v: string) => void;
}) {
// Buffer keystrokes locally and persist on blur/Enter so each character
// doesn't fire a settings PUT (or a localStorage write).
const [draft, setDraft] = React.useState<string | null>(null);
return (
<FieldRow htmlFor={id} label={label}>
<Input
id={id}
type="text"
className="font-mono text-sm"
value={draft ?? value}
onChange={(e) => setDraft(e.target.value)}
onBlur={() => {
if (draft !== null && draft !== value) onCommit(draft);
setDraft(null);
}}
onKeyDown={(e) => {
if (e.key === 'Enter') e.currentTarget.blur();
}}
placeholder={placeholder}
/>
</FieldRow>
);
}
function SelectRow({
id,
label,
value,
onChange,
children,
}: {
id: string;
label: string;
value: string;
onChange: (v: string) => void;
children: React.ReactNode;
}) {
return (
<FieldRow htmlFor={id} label={label}>
<NativeSelect
id={id}
value={value}
onChange={(e) => onChange(e.target.value)}
>
{children}
</NativeSelect>
</FieldRow>
);
}
export default function SettingsPage() {
const { options } = useStore(defaultOptionsStore);
const [models, setModels] = React.useState<
Record<'t2v' | 'i2v' | 't2i', Model[]>
>({ t2v: [], i2v: [], t2i: [] });
React.useEffect(() => {
Promise.all([getModels('t2v'), getModels('i2v'), getModels('t2i')])
.then(([t2v, i2v, t2i]) => setModels({ t2v, i2v, t2i }))
.catch((e) => console.error('Failed to load models:', e));
}, []);
return (
<div className="mx-auto flex w-full max-w-[850px] flex-col gap-6 px-4 pb-12 pt-6">
<Card>
<CardContent className="space-y-4 p-6">
<h2 className="text-lg font-semibold">Behavior</h2>
<div className="flex items-center justify-between gap-4">
<Label
htmlFor="settings-auto-start-job"
className="pl-0.5 text-xs font-normal tracking-wide text-muted-foreground"
>
Auto Start Job on Create
</Label>
<Switch
id="settings-auto-start-job"
checked={options.autoStartJob}
onCheckedChange={(v) => updateOption('autoStartJob', v)}
/>
</div>
<Separator />
<h2 className="text-lg font-semibold">Paths</h2>
<div className="space-y-4">
<TextSettingRow
id="settings-api-server-base-url"
label="API Server Base URL"
value={options.apiServerBaseUrl ?? ''}
onCommit={(v) => updateOption('apiServerBaseUrl', v)}
placeholder="http://localhost:8189/api"
/>
<TextSettingRow
id="settings-dataset-upload-path"
label="Dataset Upload Path"
value={options.datasetUploadPath ?? ''}
onCommit={(v) => updateOption('datasetUploadPath', v)}
placeholder="outputs/ui_data/uploads/datasets"
/>
</div>
<Separator />
<div className="flex items-center justify-between">
<h2 className="text-lg font-semibold">Default Options</h2>
<Button
type="button"
variant="outline"
size="sm"
onClick={resetToDefaults}
>
Reset to Defaults
</Button>
</div>
<p className="text-sm text-muted-foreground">
These values are used as defaults when creating new jobs.
</p>
<div className="grid gap-x-3 gap-y-2 [grid-template-columns:repeat(auto-fill,minmax(160px,1fr))]">
<SelectRow
id="settings-default-model-t2v"
label="Default Model (T2V)"
value={options.defaultModelIdT2v}
onChange={(v) => updateOption('defaultModelIdT2v', v)}
>
<option value="">None (select when creating job)</option>
{models.t2v.map((model) => (
<option key={model.id} value={model.id}>
{model.label} ({model.id})
</option>
))}
</SelectRow>
<SelectRow
id="settings-default-model-i2v"
label="Default Model (I2V)"
value={options.defaultModelIdI2v}
onChange={(v) => updateOption('defaultModelIdI2v', v)}
>
<option value="">None</option>
{models.i2v.map((model) => (
<option key={model.id} value={model.id}>
{model.label} ({model.id})
</option>
))}
</SelectRow>
<SelectRow
id="settings-default-model-t2i"
label="Default Model (T2I)"
value={options.defaultModelIdT2i}
onChange={(v) => updateOption('defaultModelIdT2i', v)}
>
<option value="">None</option>
{models.t2i.map((model) => (
<option key={model.id} value={model.id}>
{model.label} ({model.id})
</option>
))}
</SelectRow>
<SliderRow
id="settings-num-frames"
label="Frames"
min={1}
max={500}
step={1}
value={options.numFrames}
onChange={(v) => updateOption('numFrames', v)}
/>
<SliderRow
id="settings-height"
label="Height"
min={64}
max={1080}
step={16}
value={options.height}
onChange={(v) => updateOption('height', v)}
/>
<SliderRow
id="settings-width"
label="Width"
min={64}
max={1920}
step={16}
value={options.width}
onChange={(v) => updateOption('width', v)}
/>
<SliderRow
id="settings-num-steps"
label="Inference Steps"
min={1}
max={200}
step={1}
value={options.numInferenceSteps}
onChange={(v) => updateOption('numInferenceSteps', v)}
/>
<SliderRow
id="settings-vsa-sparsity"
label="VSA Sparsity"
title="VSA sparsity (0–1)"
min={0}
max={1}
step={0.05}
value={options.vsaSparsity}
onChange={(v) => updateOption('vsaSparsity', v)}
format={(v) => v.toFixed(2)}
/>
<SliderRow
id="settings-guidance"
label="Guidance Scale"
min={0}
max={20}
step={0.1}
value={options.guidanceScale}
onChange={(v) => updateOption('guidanceScale', v)}
format={(v) => v.toFixed(1)}
/>
<SliderRow
id="settings-guidance-rescale"
label="Guidance Rescale"
title="0 = disabled"
min={0}
max={1}
step={0.05}
value={options.guidanceRescale ?? 0}
onChange={(v) => updateOption('guidanceRescale', v)}
format={(v) => v.toFixed(2)}
/>
<SliderRow
id="settings-tp-size"
label="TP Size"
title="-1 = auto"
min={-1}
max={8}
step={1}
value={options.tpSize}
onChange={(v) => updateOption('tpSize', v)}
format={(v) => (v === -1 ? 'Auto' : String(v))}
/>
<SliderRow
id="settings-sp-size"
label="SP Size"
title="-1 = auto"
min={-1}
max={8}
step={1}
value={options.spSize}
onChange={(v) => updateOption('spSize', v)}
format={(v) => (v === -1 ? 'Auto' : String(v))}
/>
<SliderRow
id="settings-fps"
label="FPS"
min={1}
max={60}
step={1}
value={options.fps ?? 24}
onChange={(v) => updateOption('fps', v)}
/>
<ToggleRow
id="settings-dit-cpu-offload"
label="DiT CPU Offload"
checked={options.ditCpuOffload}
onChange={(v) => updateOption('ditCpuOffload', v)}
/>
<ToggleRow
id="settings-text-encoder-cpu-offload"
label="Text Encoder CPU Offload"
checked={options.textEncoderCpuOffload}
onChange={(v) => updateOption('textEncoderCpuOffload', v)}
/>
<ToggleRow
id="settings-use-fsdp-inference"
label="Use FSDP Inference"
checked={options.useFsdpInference}
onChange={(v) => updateOption('useFsdpInference', v)}
/>
<ToggleRow
id="settings-vae-cpu-offload"
label="VAE CPU Offload"
checked={options.vaeCpuOffload}
onChange={(v) => updateOption('vaeCpuOffload', v)}
/>
<ToggleRow
id="settings-image-encoder-cpu-offload"
label="Image Encoder CPU Offload"
checked={options.imageEncoderCpuOffload}
onChange={(v) => updateOption('imageEncoderCpuOffload', v)}
/>
<ToggleRow
id="settings-enable-torch-compile"
label="Torch Compile"
checked={options.enableTorchCompile}
onChange={(v) => updateOption('enableTorchCompile', v)}
/>
<SliderRow
id="settings-num-gpus"
label="GPUs"
min={1}
max={8}
step={1}
value={options.numGpus}
onChange={(v) => updateOption('numGpus', v)}
/>
<NumberRow
id="settings-seed"
label="Seed"
min={0}
value={options.seed}
onChange={(v) => updateOption('seed', v)}
/>
</div>
</CardContent>
</Card>
</div>
);
}
@@ -1,93 +0,0 @@
'use client';
import type { ClusterNode } from '@/lib/api';
import { cn } from '@/lib/utils';
export function formatGib(mib: number): string {
return `${(mib / 1024).toFixed(1)} GiB`;
}
/** Bar fill class by load: blue when idle, amber under pressure, rose hot. */
export function utilizationColor(percent: number): string {
if (percent >= 85) return 'bg-rose-500';
if (percent >= 50) return 'bg-amber-500';
return 'bg-accent-blue';
}
export function clampPercent(percent: number): number {
return Math.max(0, Math.min(100, percent));
}
function GpuSegment({
index,
utilization,
memUsedMib,
memTotalMib,
}: {
index: number;
utilization: number;
memUsedMib: number;
memTotalMib: number;
}) {
const memPercent =
memTotalMib > 0 ? clampPercent((memUsedMib / memTotalMib) * 100) : 0;
const label =
`GPU ${index}: ${utilization}% utilization, ` +
`${formatGib(memUsedMib)} / ${formatGib(memTotalMib)} VRAM`;
return (
<div
role="img"
aria-label={label}
title={label}
className="flex w-10 shrink-0 flex-col gap-0.5"
>
<div className="h-1.5 overflow-hidden rounded-full bg-muted">
<div
className={cn('h-full rounded-full', utilizationColor(utilization))}
style={{ width: `${clampPercent(utilization)}%` }}
/>
</div>
<div className="h-1.5 overflow-hidden rounded-full bg-muted">
<div
className={cn(
'h-full rounded-full',
memPercent >= 90 ? 'bg-rose-500' : 'bg-accent-blue',
)}
style={{ width: `${memPercent}%` }}
/>
</div>
</div>
);
}
/** One compact line per node: hostname + tiny util/VRAM bars per GPU. */
export default function ClusterStrip({ nodes }: { nodes: ClusterNode[] }) {
return (
<div className="flex flex-col gap-2">
{nodes.map((node, i) => (
<div
key={`${node.hostname}-${i}`}
className="flex items-center gap-3"
>
<span className="w-40 shrink-0 truncate text-xs font-medium">
{node.hostname}
</span>
<div className="flex min-w-0 flex-wrap items-center gap-1.5">
{node.gpus.map((gpu) => (
<GpuSegment
key={gpu.index}
index={gpu.index}
utilization={gpu.utilization}
memUsedMib={gpu.memory_used_mib}
memTotalMib={gpu.memory_total_mib}
/>
))}
{node.gpus.length === 0 && (
<span className="text-xs text-muted-foreground">no GPUs</span>
)}
</div>
</div>
))}
</div>
);
}
@@ -1,12 +0,0 @@
'use client';
import { Button } from '@/components/ui/button';
import { setCreateDatasetModalOpen } from '@/stores/createDatasetModalOpen';
export default function AddDatasetButton() {
return (
<Button type="button" onClick={() => setCreateDatasetModalOpen(true)}>
Add Dataset
</Button>
);
}
@@ -1,43 +0,0 @@
import { describe, expect, it, vi } from 'vitest';
import { fireEvent, render, screen } from '@testing-library/react';
import CreateDatasetModal from '@/components/datasets/CreateDatasetModal';
import { createDataset } from '@/lib/api';
vi.mock('@/lib/api', () => ({
createDataset: vi.fn(),
uploadRawDataset: vi.fn(),
}));
const mockedCreateDataset = vi.mocked(createDataset);
describe('CreateDatasetModal', () => {
it('renders nothing when closed', () => {
render(
<CreateDatasetModal isOpen={false} onClose={() => {}} onSuccess={() => {}} />,
);
expect(screen.queryByText('Add Dataset — Raw')).not.toBeInTheDocument();
});
it('renders the form and the default JSON caption upload when open', () => {
render(
<CreateDatasetModal isOpen onClose={() => {}} onSuccess={() => {}} />,
);
expect(screen.getByText('Add Dataset — Raw')).toBeInTheDocument();
expect(screen.getByText('Upload video files')).toBeInTheDocument();
expect(screen.getByText('Upload videos2caption.json')).toBeInTheDocument();
expect(
screen.getByRole('button', { name: 'Create Dataset' }),
).toBeInTheDocument();
});
it('does not create a dataset when the name is empty', () => {
const onSuccess = vi.fn();
render(
<CreateDatasetModal isOpen onClose={() => {}} onSuccess={onSuccess} />,
);
fireEvent.click(screen.getByRole('button', { name: 'Create Dataset' }));
expect(mockedCreateDataset).not.toHaveBeenCalled();
expect(onSuccess).not.toHaveBeenCalled();
});
});
@@ -1,465 +0,0 @@
'use client';
import * as React from 'react';
import UploadZone from '@/components/datasets/UploadZone';
import { Button } from '@/components/ui/button';
import {
Dialog,
DialogContent,
DialogHeader,
DialogTitle,
} from '@/components/ui/dialog';
import { Input } from '@/components/ui/input';
import { Label } from '@/components/ui/label';
import { Tabs, TabsContent, TabsList, TabsTrigger } from '@/components/ui/tabs';
import { createDataset, uploadRawDataset } from '@/lib/api';
import {
parseCaptionCsv,
parseVideos2Caption,
parseVideosCaptionsTxt,
} from '@/lib/captionParsing';
const ALLOWED_VIDEO_EXT = '.mp4,.webm,.avi,.mov,.mkv';
type CaptionFormat = 'json' | 'txt' | 'csv';
export interface CreateDatasetModalProps {
isOpen: boolean;
onClose: () => void;
onSuccess: () => void;
}
export default function CreateDatasetModal({
isOpen,
onClose,
onSuccess,
}: CreateDatasetModalProps) {
const [name, setName] = React.useState('');
const [isSubmitting, setIsSubmitting] = React.useState(false);
const [rawPath, setRawPath] = React.useState('');
const [fileNames, setFileNames] = React.useState<string[]>([]);
const [isUploading, setIsUploading] = React.useState(false);
const [validationError, setValidationError] = React.useState<string | null>(
null,
);
const [captionFormat, setCaptionFormat] = React.useState<CaptionFormat>('json');
const [captionMap, setCaptionMap] = React.useState<Record<
string,
string
> | null>(null);
const [captionFileName, setCaptionFileName] = React.useState<string | null>(
null,
);
const [videosTxtLines, setVideosTxtLines] = React.useState<string[] | null>(
null,
);
const [videosTxtFileName, setVideosTxtFileName] = React.useState<
string | null
>(null);
const [captionsTxtLines, setCaptionsTxtLines] = React.useState<
string[] | null
>(null);
const [captionsTxtFileName, setCaptionsTxtFileName] = React.useState<
string | null
>(null);
// Bumped on every new media selection and on reset/clear; an in-flight
// upload whose generation no longer matches is discarded, so a superseded or
// abandoned upload can't repopulate/overwrite state.
const uploadGeneration = React.useRef(0);
const txtCaptionMap = React.useMemo(() => {
if (!captionsTxtLines || fileNames.length === 0) return null;
const { captions, error } = parseVideosCaptionsTxt(
videosTxtLines ?? null,
captionsTxtLines,
fileNames,
);
if (error) return null;
return Object.keys(captions).length > 0 ? captions : null;
}, [captionsTxtLines, fileNames, videosTxtLines]);
const effectiveCaptionMap =
captionFormat === 'txt' ? txtCaptionMap : captionMap;
// Keep the validation message in sync with the TXT inputs.
React.useEffect(() => {
if (captionFormat === 'txt' && captionsTxtLines && fileNames.length > 0) {
const hasVideosTxt =
!!videosTxtLines &&
videosTxtLines.length > 0 &&
videosTxtLines.some((s) => s.trim());
if (hasVideosTxt && videosTxtLines) {
const { error } = parseVideosCaptionsTxt(
videosTxtLines,
captionsTxtLines,
fileNames,
);
setValidationError(error);
} else {
setValidationError(null);
}
}
}, [captionFormat, captionsTxtLines, fileNames, videosTxtLines]);
function resetState() {
uploadGeneration.current += 1;
setName('');
setRawPath('');
setFileNames([]);
setValidationError(null);
setCaptionFormat('json');
setCaptionMap(null);
setCaptionFileName(null);
setVideosTxtLines(null);
setVideosTxtFileName(null);
setCaptionsTxtLines(null);
setCaptionsTxtFileName(null);
}
function handleClose() {
if (isSubmitting) return;
resetState();
onClose();
}
async function handleMediaChange(files: File[]) {
const gen = (uploadGeneration.current += 1);
setValidationError(null);
if (files.length === 0) {
setRawPath('');
setFileNames([]);
return;
}
setIsUploading(true);
try {
const res = await uploadRawDataset(files);
if (gen !== uploadGeneration.current) return; // superseded or abandoned
setRawPath(res.path);
setFileNames(res.file_names);
if (res.file_names.length === 0) {
setValidationError(`No video files found. Allowed: ${ALLOWED_VIDEO_EXT}`);
}
} catch (err) {
if (gen !== uploadGeneration.current) return;
setRawPath('');
setFileNames([]);
setValidationError(err instanceof Error ? err.message : 'Upload failed');
} finally {
if (gen === uploadGeneration.current) setIsUploading(false);
}
}
async function handleCaptionJsonChange(files: File[]) {
setValidationError(null);
setCaptionMap(null);
setCaptionFileName(null);
if (files.length === 0) return;
const file = files[0];
try {
const text = await file.text();
const { captions, error } = parseVideos2Caption(text, fileNames);
if (error) {
setValidationError(error);
return;
}
setCaptionMap(captions);
setCaptionFileName(file.name);
} catch {
setValidationError('Could not read the file.');
}
}
async function handleCaptionCsvChange(files: File[]) {
setValidationError(null);
setCaptionMap(null);
setCaptionFileName(null);
if (files.length === 0) return;
const file = files[0];
try {
const text = await file.text();
const { captions, error } = parseCaptionCsv(text, fileNames);
if (error) {
setValidationError(error);
return;
}
setCaptionMap(captions);
setCaptionFileName(file.name);
} catch {
setValidationError('Could not read the file.');
}
}
async function handleVideosTxtChange(files: File[]) {
setValidationError(null);
setVideosTxtLines(null);
setVideosTxtFileName(null);
if (files.length === 0) return;
try {
const text = await files[0].text();
setVideosTxtLines(text.split(/\r?\n/).map((s) => s.trim()));
setVideosTxtFileName(files[0].name);
} catch {
setValidationError('Could not read videos.txt.');
}
}
async function handleCaptionsTxtChange(files: File[]) {
setValidationError(null);
setCaptionsTxtLines(null);
setCaptionsTxtFileName(null);
if (files.length === 0) return;
try {
const text = await files[0].text();
setCaptionsTxtLines(text.split(/\r?\n/).map((s) => s.trim()));
setCaptionsTxtFileName(files[0].name);
} catch {
setValidationError('Could not read captions.txt.');
}
}
function handleCaptionFormatChange(format: CaptionFormat) {
setCaptionFormat(format);
setValidationError(null);
setCaptionMap(null);
setCaptionFileName(null);
setVideosTxtLines(null);
setVideosTxtFileName(null);
setCaptionsTxtLines(null);
setCaptionsTxtFileName(null);
}
async function handleSubmit(e: React.FormEvent) {
e.preventDefault();
setValidationError(null);
if (!name.trim()) return;
if (!rawPath || fileNames.length === 0) {
setValidationError('No data was found. Upload at least one video.');
return;
}
if (captionFormat === 'json' || captionFormat === 'csv') {
if (captionFileName && !captionMap) {
setValidationError(
'Caption file has errors. Fix or remove it before creating the dataset.',
);
return;
}
} else if (captionFormat === 'txt') {
if (videosTxtFileName || captionsTxtFileName) {
if (!captionsTxtFileName) {
setValidationError('Upload captions.txt to use TXT captions.');
return;
}
// Re-validate against current state rather than the stale render-time
// `validationError`, so a videos.txt mismatch both blocks submission
// and keeps its message visible.
const hasVideosTxt =
!!videosTxtLines &&
videosTxtLines.length > 0 &&
videosTxtLines.some((s) => s.trim());
if (hasVideosTxt && videosTxtLines) {
const { error } = parseVideosCaptionsTxt(
videosTxtLines,
captionsTxtLines ?? [],
fileNames,
);
if (error) {
setValidationError(error);
return;
}
}
}
}
const finalCaptionMap =
effectiveCaptionMap && Object.keys(effectiveCaptionMap).length > 0
? effectiveCaptionMap
: null;
if (finalCaptionMap) {
const missing = fileNames.filter((fn) => !(fn in finalCaptionMap));
if (missing.length > 0) {
const list =
missing.length <= 5
? missing.join(', ')
: `${missing.slice(0, 5).join(', ')} and ${missing.length - 5} more`;
const ok = window.confirm(
`The caption file does not include captions for ${missing.length} video(s): ${list}. They will get empty captions. Continue?`,
);
if (!ok) return;
}
}
setIsSubmitting(true);
try {
await createDataset({
name: name.trim(),
upload_path: rawPath,
file_names: fileNames,
...(finalCaptionMap ? { captions: finalCaptionMap } : {}),
});
onSuccess();
resetState();
onClose();
} catch (err) {
setValidationError(
err instanceof Error ? err.message : 'Failed to create dataset',
);
} finally {
setIsSubmitting(false);
}
}
return (
<Dialog
open={isOpen}
onOpenChange={(open) => {
if (!open) handleClose();
}}
>
<DialogContent
aria-describedby={undefined}
className="max-h-[90vh] w-[90vw] max-w-[850px] overflow-y-auto"
>
<DialogHeader>
<DialogTitle>Add Dataset — Raw</DialogTitle>
</DialogHeader>
<form onSubmit={handleSubmit} autoComplete="off">
<div className="mb-3.5 flex flex-col gap-1.5">
<Label htmlFor="add-dataset-name">Name</Label>
<Input
id="add-dataset-name"
type="text"
value={name}
onChange={(e) => setName(e.target.value)}
placeholder="My dataset"
required
disabled={isSubmitting}
/>
</div>
<div className="mb-3.5 flex flex-col gap-1.5">
<Label>Videos</Label>
<UploadZone
label="Upload video files"
hint="Select files or a folder (.mp4, .webm, .avi, .mov, .mkv)"
accept={ALLOWED_VIDEO_EXT}
multiple
directory
allowBothFileAndDirectory
value={rawPath}
fileName={
fileNames.length > 0 ? `${fileNames.length} file(s)` : undefined
}
onFiles={handleMediaChange}
onClear={() => {
uploadGeneration.current += 1;
setRawPath('');
setFileNames([]);
setCaptionMap(null);
setCaptionFileName(null);
setValidationError(null);
}}
disabled={isSubmitting}
uploading={isUploading}
/>
</div>
<div className="mb-3.5 flex flex-col gap-1.5">
<Tabs
value={captionFormat}
onValueChange={(value) =>
handleCaptionFormatChange(value as CaptionFormat)
}
>
<div className="mb-1.5 flex flex-wrap items-center gap-4">
<Label>Captions (optional)</Label>
<TabsList>
<TabsTrigger value="json" disabled={isSubmitting}>
JSON
</TabsTrigger>
<TabsTrigger value="txt" disabled={isSubmitting}>
TXT
</TabsTrigger>
<TabsTrigger value="csv" disabled={isSubmitting}>
CSV
</TabsTrigger>
</TabsList>
</div>
<TabsContent value="json">
<UploadZone
label="Upload videos2caption.json"
hint="Array of { path, cap } or object mapping file names to captions"
accept=".json,application/json"
value={captionFileName ? '1' : ''}
fileName={captionFileName ?? undefined}
onFiles={handleCaptionJsonChange}
onClear={() => {
setCaptionMap(null);
setCaptionFileName(null);
setValidationError(null);
}}
disabled={isSubmitting}
/>
</TabsContent>
<TabsContent value="txt">
<div className="flex flex-wrap gap-[15px] [&>*]:min-w-[200px] [&>*]:flex-1">
<UploadZone
label="Upload videos.txt (optional)"
hint="One video path per line, or leave empty to match captions to videos in alphabetical order"
accept=".txt,text/plain"
value={videosTxtFileName ? '1' : ''}
fileName={videosTxtFileName ?? undefined}
onFiles={handleVideosTxtChange}
onClear={() => {
setVideosTxtLines(null);
setVideosTxtFileName(null);
setValidationError(null);
}}
disabled={isSubmitting}
/>
<UploadZone
label="Upload captions.txt"
hint="One caption per line (same order as videos.txt or alphabetical)"
accept=".txt,text/plain"
value={captionsTxtFileName ? '1' : ''}
fileName={captionsTxtFileName ?? undefined}
onFiles={handleCaptionsTxtChange}
onClear={() => {
setCaptionsTxtLines(null);
setCaptionsTxtFileName(null);
setValidationError(null);
}}
disabled={isSubmitting}
/>
</div>
</TabsContent>
<TabsContent value="csv">
<UploadZone
label="Upload captions CSV"
hint="Header: video_name, caption"
accept=".csv,text/csv"
value={captionFileName ? '1' : ''}
fileName={captionFileName ?? undefined}
onFiles={handleCaptionCsvChange}
onClear={() => {
setCaptionMap(null);
setCaptionFileName(null);
setValidationError(null);
}}
disabled={isSubmitting}
/>
</TabsContent>
</Tabs>
</div>
{validationError && (
<p className="mb-2 text-sm text-destructive">{validationError}</p>
)}
<Button type="submit" disabled={isSubmitting}>
{isSubmitting ? 'Creating…' : 'Create Dataset'}
</Button>
</form>
</DialogContent>
</Dialog>
);
}
@@ -1,119 +0,0 @@
import { beforeEach, describe, expect, it, vi } from 'vitest';
import { fireEvent, render, screen, waitFor } from '@testing-library/react';
import DatasetCard from '@/components/datasets/DatasetCard';
import { deleteDataset } from '@/lib/api';
import type { Dataset } from '@/lib/api';
import { setActiveDatasetId } from '@/stores/activeDataset';
vi.mock('@/lib/api', () => ({
deleteDataset: vi.fn(),
}));
const mockedDeleteDataset = vi.mocked(deleteDataset);
const dataset: Dataset = {
id: 'ds-1',
name: 'My Dataset',
created_at: 0,
file_count: 3,
size_bytes: 2048,
};
beforeEach(() => {
setActiveDatasetId(null);
});
describe('DatasetCard', () => {
it('keeps selection and delete buttons as semantic siblings', () => {
render(<DatasetCard dataset={dataset} onUpdated={() => {}} />);
const selectButton = screen.getByRole('button', { pressed: false });
const deleteButton = screen.getByRole('button', { name: 'Delete' });
expect(selectButton).toHaveTextContent('My Dataset');
expect(selectButton).not.toContainElement(deleteButton);
});
it('renders the name, file count and human-readable size', () => {
render(<DatasetCard dataset={dataset} onUpdated={() => {}} />);
expect(screen.getByText('My Dataset')).toBeInTheDocument();
expect(screen.getByText('3 files · 2.0 KB')).toBeInTheDocument();
});
it('uses the singular "file" label and byte units for a small dataset', () => {
render(
<DatasetCard
dataset={{ ...dataset, file_count: 1, size_bytes: 512 }}
onUpdated={() => {}}
/>,
);
expect(screen.getByText('1 file · 512 B')).toBeInTheDocument();
});
it('calls onSelect when the card body is clicked', () => {
const onSelect = vi.fn();
render(
<DatasetCard dataset={dataset} onUpdated={() => {}} onSelect={onSelect} />,
);
fireEvent.click(screen.getByText('My Dataset'));
expect(onSelect).toHaveBeenCalledTimes(1);
});
it('deletes after confirmation and notifies the parent without selecting', async () => {
const onUpdated = vi.fn();
const onSelect = vi.fn();
const confirmSpy = vi.spyOn(window, 'confirm').mockReturnValue(true);
mockedDeleteDataset.mockResolvedValue(undefined);
render(
<DatasetCard
dataset={dataset}
onUpdated={onUpdated}
onSelect={onSelect}
/>,
);
fireEvent.click(screen.getByRole('button', { name: 'Delete' }));
expect(confirmSpy).toHaveBeenCalledWith('Delete dataset "My Dataset"?');
expect(mockedDeleteDataset).toHaveBeenCalledWith('ds-1');
await waitFor(() => expect(onUpdated).toHaveBeenCalledTimes(1));
expect(onSelect).not.toHaveBeenCalled();
});
it('keeps the selection and delete actions separate', () => {
const onSelect = vi.fn();
render(
<DatasetCard dataset={dataset} onUpdated={() => {}} onSelect={onSelect} />,
);
// Activating the Delete button must not bubble into a card selection.
fireEvent.keyDown(screen.getByRole('button', { name: 'Delete' }), {
key: 'Enter',
});
expect(onSelect).not.toHaveBeenCalled();
// Activating the dedicated selection button selects the dataset.
fireEvent.click(
screen.getByRole('button', {
name: /My Dataset.*3 files.*2.0 KB/,
}),
);
expect(onSelect).toHaveBeenCalledTimes(1);
});
it('does not delete when the user cancels the confirm dialog', () => {
vi.spyOn(window, 'confirm').mockReturnValue(false);
render(<DatasetCard dataset={dataset} onUpdated={() => {}} />);
fireEvent.click(screen.getByRole('button', { name: 'Delete' }));
expect(mockedDeleteDataset).not.toHaveBeenCalled();
});
it('applies selected styling when it is the active dataset', () => {
setActiveDatasetId('ds-1');
const { container } = render(
<DatasetCard dataset={dataset} onUpdated={() => {}} />,
);
expect(container.firstChild).toHaveClass('border-accent-blue', 'bg-accent-blue/5');
});
});
@@ -1,87 +0,0 @@
'use client';
import * as React from 'react';
import { Button } from '@/components/ui/button';
import { deleteDataset } from '@/lib/api';
import type { Dataset } from '@/lib/api';
import { cn } from '@/lib/utils';
import { activeDatasetStore, setActiveDatasetId } from '@/stores/activeDataset';
import { useStore } from '@/hooks/useStore';
function formatSize(sizeBytes: number): string {
if (sizeBytes < 1024) return `${sizeBytes} B`;
if (sizeBytes < 1024 * 1024) return `${(sizeBytes / 1024).toFixed(1)} KB`;
if (sizeBytes < 1024 * 1024 * 1024) {
return `${(sizeBytes / (1024 * 1024)).toFixed(1)} MB`;
}
return `${(sizeBytes / (1024 * 1024 * 1024)).toFixed(1)} GB`;
}
export default function DatasetCard({
dataset,
onUpdated,
onSelect = () => {},
}: {
dataset: Dataset;
onUpdated: () => void;
onSelect?: () => void;
}) {
const { activeDatasetId } = useStore(activeDatasetStore);
const [isLoading, setIsLoading] = React.useState(false);
const isSelected = activeDatasetId === dataset.id;
const fileCount = dataset.file_count ?? 0;
const sizeLabel = formatSize(dataset.size_bytes ?? 0);
async function handleDelete(e: React.MouseEvent) {
e.preventDefault();
e.stopPropagation();
if (isLoading) return;
if (!window.confirm(`Delete dataset "${dataset.name}"?`)) return;
setIsLoading(true);
try {
await deleteDataset(dataset.id);
if (activeDatasetStore.get().activeDatasetId === dataset.id) {
setActiveDatasetId(null);
}
onUpdated();
} catch (err) {
window.alert(
err instanceof Error ? err.message : 'Failed to delete dataset',
);
} finally {
setIsLoading(false);
}
}
return (
<article
className={cn(
'mb-3 flex items-start gap-3 rounded-lg border border-border bg-background px-[1.15rem] py-4',
isSelected && 'border-accent-blue bg-accent-blue/5',
)}
>
<button
type="button"
aria-pressed={isSelected}
onClick={onSelect}
className="flex min-w-0 flex-1 cursor-pointer flex-col gap-[0.6rem] rounded-md text-left"
>
<span className="text-[0.95rem] font-semibold">{dataset.name}</span>
<span className="text-sm text-muted-foreground">
{fileCount} {fileCount === 1 ? 'file' : 'files'} · {sizeLabel}
</span>
</button>
<Button
type="button"
variant="destructive"
size="sm"
onClick={handleDelete}
disabled={isLoading}
>
Delete
</Button>
</article>
);
}
@@ -1,223 +0,0 @@
import { describe, it, expect, vi, beforeEach } from 'vitest';
import { render, screen, fireEvent, act } from '@testing-library/react';
import DatasetSidebar from '@/components/datasets/DatasetSidebar';
import * as api from '@/lib/api';
import type { Dataset } from '@/lib/api';
vi.mock('@/lib/api');
vi.mock('sonner', () => ({
toast: { error: vi.fn() },
}));
const mockedApi = vi.mocked(api);
const dataset: Dataset = {
id: 'ds-1',
name: 'My Dataset',
created_at: 0,
};
beforeEach(() => {
mockedApi.getDatasetFiles.mockResolvedValue({
file_names: ['a.mp4', 'b.mp4'],
captions: { 'a.mp4': 'cap a', 'b.mp4': '' },
});
mockedApi.getDatasetMediaUrl.mockImplementation(
(id, fileName) => `http://test/${id}/${fileName}`,
);
mockedApi.updateDatasetCaption.mockResolvedValue(undefined);
});
describe('DatasetSidebar', () => {
it('fills the mobile viewport without reserving main-content width', async () => {
const onWidthChange = vi.fn();
render(
<DatasetSidebar
dataset={dataset}
isMobile
onClose={() => {}}
onWidthChange={onWidthChange}
/>,
);
const drawer = screen.getByRole('dialog', {
name: 'My Dataset dataset details',
});
expect(drawer).toHaveStyle({ width: '100%', maxWidth: 'none' });
expect(drawer).toHaveAttribute('aria-modal', 'true');
expect(drawer).toHaveFocus();
expect(onWidthChange).toHaveBeenCalledWith(0);
});
it('lists dataset files after loading', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
// The first file's caption is rendered once loading resolves.
expect(await screen.findByDisplayValue('cap a')).toBeInTheDocument();
expect(mockedApi.getDatasetFiles).toHaveBeenCalledWith('ds-1');
// One caption editor per returned file.
const captionFields = screen.getAllByPlaceholderText('Caption');
expect(captionFields).toHaveLength(2);
// Media URLs are requested per visible file.
expect(mockedApi.getDatasetMediaUrl).toHaveBeenCalledWith('ds-1', 'a.mp4');
expect(mockedApi.getDatasetMediaUrl).toHaveBeenCalledWith('ds-1', 'b.mp4');
});
it('shows a fallback when a dataset preview cannot load', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
const preview = await screen.findByLabelText('Preview of a.mp4');
fireEvent.error(preview);
expect(screen.getByText('Preview unavailable')).toBeInTheDocument();
expect(screen.queryByLabelText('Preview of a.mp4')).not.toBeInTheDocument();
});
it('debounces caption save by 500ms', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
const textarea = await screen.findByDisplayValue('cap a');
vi.useFakeTimers();
try {
fireEvent.change(textarea, { target: { value: 'updated caption' } });
// Nothing saved immediately or just before the debounce window closes.
expect(mockedApi.updateDatasetCaption).not.toHaveBeenCalled();
act(() => {
vi.advanceTimersByTime(499);
});
expect(mockedApi.updateDatasetCaption).not.toHaveBeenCalled();
// Saved exactly once after the full 500ms.
act(() => {
vi.advanceTimersByTime(1);
});
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledTimes(1);
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledWith(
'ds-1',
'a.mp4',
'updated caption',
);
} finally {
vi.useRealTimers();
}
});
it('coalesces rapid edits into a single debounced save', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
const textarea = await screen.findByDisplayValue('cap a');
vi.useFakeTimers();
try {
fireEvent.change(textarea, { target: { value: 'one' } });
act(() => {
vi.advanceTimersByTime(300);
});
fireEvent.change(textarea, { target: { value: 'two' } });
act(() => {
vi.advanceTimersByTime(500);
});
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledTimes(1);
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledWith(
'ds-1',
'a.mp4',
'two',
);
} finally {
vi.useRealTimers();
}
});
it('shows a failed save and lets the user retry it', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
mockedApi.updateDatasetCaption
.mockRejectedValueOnce(new Error('network down'))
.mockResolvedValueOnce(undefined);
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
const textarea = await screen.findByDisplayValue('cap a');
vi.useFakeTimers();
try {
fireEvent.change(textarea, { target: { value: 'needs retry' } });
await act(async () => {
await vi.advanceTimersByTimeAsync(500);
});
expect(screen.getByText(/Not saved/)).toBeInTheDocument();
fireEvent.click(screen.getByRole('button', { name: 'Retry' }));
await act(async () => {
await Promise.resolve();
});
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledTimes(2);
expect(mockedApi.updateDatasetCaption).toHaveBeenLastCalledWith(
'ds-1',
'a.mp4',
'needs retry',
);
expect(screen.getByText('Saved')).toBeInTheDocument();
} finally {
vi.useRealTimers();
}
});
it('debounces per file: editing another caption does not cancel a pending save', async () => {
render(<DatasetSidebar dataset={dataset} onClose={() => {}} />);
await screen.findByDisplayValue('cap a');
const [textareaA, textareaB] = screen.getAllByPlaceholderText('Caption');
vi.useFakeTimers();
try {
fireEvent.change(textareaA, { target: { value: 'new cap a' } });
act(() => {
vi.advanceTimersByTime(300);
});
// Editing b.mp4 inside a.mp4's debounce window must not drop a's save.
fireEvent.change(textareaB, { target: { value: 'new cap b' } });
act(() => {
vi.advanceTimersByTime(500);
});
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledTimes(2);
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledWith(
'ds-1',
'a.mp4',
'new cap a',
);
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledWith(
'ds-1',
'b.mp4',
'new cap b',
);
} finally {
vi.useRealTimers();
}
});
it('flushes a pending save on unmount instead of dropping it', async () => {
const { unmount } = render(
<DatasetSidebar dataset={dataset} onClose={() => {}} />,
);
const textarea = await screen.findByDisplayValue('cap a');
vi.useFakeTimers();
try {
fireEvent.change(textarea, { target: { value: 'edited just before close' } });
unmount();
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledTimes(1);
expect(mockedApi.updateDatasetCaption).toHaveBeenCalledWith(
'ds-1',
'a.mp4',
'edited just before close',
);
} finally {
vi.useRealTimers();
}
});
});
@@ -1,347 +0,0 @@
'use client';
import * as React from 'react';
import { ImageOff, X } from 'lucide-react';
import { toast } from 'sonner';
import DownloadCaptions from '@/components/datasets/DownloadCaptions';
import { Textarea } from '@/components/ui/textarea';
import { useResizable } from '@/hooks/useResizable';
import {
getDatasetFiles,
getDatasetMediaUrl,
updateDatasetCaption,
type Dataset,
} from '@/lib/api';
import { cn } from '@/lib/utils';
import { useDrawerFocus } from '@/hooks/useDrawerFocus';
const SIDEBAR_MIN_WIDTH = 320;
const SIDEBAR_MAX_WIDTH = 900;
const INITIAL_PAGE_SIZE = 24;
const PAGE_SIZE = 24;
const SCROLL_THRESHOLD = 200;
type CaptionSaveState = 'idle' | 'saving' | 'saved' | 'error';
// Memoized so a caption keystroke re-renders only the edited card, not every
// visible <video> in the grid (visibleCount grows unbounded with scrolling).
const DatasetFileCard = React.memo(function DatasetFileCard({
fileName,
mediaUrl,
caption,
thumbLoaded,
saveState,
onCaptionChange,
onCaptionRetry,
onThumbLoaded,
}: {
fileName: string;
mediaUrl: string;
caption: string;
thumbLoaded: boolean;
saveState: CaptionSaveState;
onCaptionChange: (fileName: string, value: string) => void;
onCaptionRetry: (fileName: string, value: string) => void;
onThumbLoaded: (fileName: string) => void;
}) {
const [mediaFailed, setMediaFailed] = React.useState(false);
React.useEffect(() => {
setMediaFailed(false);
}, [mediaUrl]);
return (
<div className="relative flex flex-col overflow-hidden rounded-lg border border-border bg-background">
{!thumbLoaded && !mediaFailed && (
<div className="pointer-events-none absolute inset-0 flex items-center justify-center bg-background/70">
<div className="h-6 w-6 animate-spin rounded-full border-2 border-muted-foreground/40 border-t-accent-blue" />
</div>
)}
{mediaFailed ? (
<div
role="status"
className="flex aspect-video w-full flex-col items-center justify-center gap-1 bg-muted px-2 text-center text-muted-foreground"
>
<ImageOff className="size-5" aria-hidden />
<span className="text-xs">Preview unavailable</span>
</div>
) : (
// eslint-disable-next-line jsx-a11y/media-has-caption
<video
src={mediaUrl}
aria-label={`Preview of ${fileName}`}
className="aspect-video w-full bg-border object-cover"
muted
autoPlay
loop
playsInline
onLoadedData={() => onThumbLoaded(fileName)}
onError={() => {
setMediaFailed(true);
onThumbLoaded(fileName);
}}
/>
)}
<Textarea
aria-label={`Caption for ${fileName}`}
value={caption}
onChange={(e) => onCaptionChange(fileName, e.target.value)}
placeholder="Caption"
rows={2}
className="min-h-[2.5rem] resize-y rounded-none border-0 bg-transparent p-1.5 text-xs shadow-none focus-visible:border-transparent focus-visible:ring-0"
/>
<div
aria-live="polite"
className="flex min-h-6 items-center px-1.5 pb-1 text-[0.7rem] text-muted-foreground"
>
{saveState === 'saving' && <span>Saving…</span>}
{saveState === 'saved' && <span>Saved</span>}
{saveState === 'error' && (
<span role="alert" className="text-destructive">
Not saved.{' '}
<button
type="button"
onClick={() => onCaptionRetry(fileName, caption)}
className="inline-flex min-h-11 items-center font-medium underline underline-offset-2"
>
Retry
</button>
</span>
)}
</div>
</div>
);
});
export default function DatasetSidebar({
dataset,
isMobile = false,
onClose,
onWidthChange,
}: {
dataset: Dataset;
isMobile?: boolean;
onClose: () => void;
onWidthChange?: (w: number) => void;
}) {
const drawerRef = useDrawerFocus<HTMLElement>(isMobile);
const [width, setWidth] = React.useState(400);
const [isDragging, setIsDragging] = React.useState(false);
const [fileNames, setFileNames] = React.useState<string[]>([]);
const [captions, setCaptions] = React.useState<Record<string, string>>({});
const [visibleCount, setVisibleCount] = React.useState(INITIAL_PAGE_SIZE);
const [isLoading, setIsLoading] = React.useState(true);
const [thumbLoaded, setThumbLoaded] = React.useState<
Record<string, boolean>
>({});
const [captionSaveStates, setCaptionSaveStates] = React.useState<
Record<string, CaptionSaveState>
>({});
// Pending debounced caption saves, keyed per file so editing one caption
// can't cancel another file's pending save.
const pendingSaves = React.useRef(
new Map<string, { timer: ReturnType<typeof setTimeout>; save: () => void }>(),
);
const captionVersions = React.useRef(new Map<string, number>());
const scrollRef = React.useRef<HTMLDivElement>(null);
React.useEffect(() => {
onWidthChange?.(isMobile ? 0 : width);
}, [isMobile, width, onWidthChange]);
React.useEffect(() => {
let cancelled = false;
setIsLoading(true);
getDatasetFiles(dataset.id)
.then((data) => {
if (cancelled) return;
setFileNames(data.file_names);
setCaptions(data.captions);
setVisibleCount(INITIAL_PAGE_SIZE);
setThumbLoaded({});
setCaptionSaveStates({});
captionVersions.current.clear();
})
.catch((err) => console.error('Failed to load dataset files:', err))
.finally(() => {
if (!cancelled) setIsLoading(false);
});
return () => {
cancelled = true;
};
}, [dataset.id]);
const flushPendingSaves = React.useCallback(() => {
for (const { timer, save } of pendingSaves.current.values()) {
clearTimeout(timer);
save();
}
pendingSaves.current.clear();
}, []);
// Flush (not drop) pending saves when the dataset changes or on unmount, so
// an edit made within the debounce window of closing isn't lost.
React.useEffect(() => flushPendingSaves, [dataset.id, flushPendingSaves]);
const { onMouseDown } = useResizable({
edge: 'right',
minWidth: SIDEBAR_MIN_WIDTH,
maxWidth: SIDEBAR_MAX_WIDTH,
getWidth: () => width,
onWidth: setWidth,
onDragChange: setIsDragging,
});
const datasetId = dataset.id;
const persistCaption = React.useCallback(
(fileName: string, value: string, version: number) => {
setCaptionSaveStates((prev) => ({ ...prev, [fileName]: 'saving' }));
void updateDatasetCaption(datasetId, fileName, value)
.then(() => {
if (captionVersions.current.get(fileName) !== version) return;
setCaptionSaveStates((prev) => ({ ...prev, [fileName]: 'saved' }));
})
.catch((error) => {
if (captionVersions.current.get(fileName) !== version) return;
console.error('Failed to save caption:', error);
setCaptionSaveStates((prev) => ({ ...prev, [fileName]: 'error' }));
toast.error('Caption was not saved', {
description: `${fileName}: check the Studio API, then retry.`,
});
});
},
[datasetId],
);
const handleCaptionChange = React.useCallback(
(fileName: string, value: string) => {
setCaptions((prev) => ({ ...prev, [fileName]: value }));
setCaptionSaveStates((prev) => ({ ...prev, [fileName]: 'idle' }));
const pending = pendingSaves.current.get(fileName);
if (pending) clearTimeout(pending.timer);
const version = (captionVersions.current.get(fileName) ?? 0) + 1;
captionVersions.current.set(fileName, version);
const save = () => persistCaption(fileName, value, version);
const timer = setTimeout(() => {
pendingSaves.current.delete(fileName);
save();
}, 500);
pendingSaves.current.set(fileName, { timer, save });
},
[persistCaption],
);
const handleCaptionRetry = React.useCallback(
(fileName: string, value: string) => {
const version = (captionVersions.current.get(fileName) ?? 0) + 1;
captionVersions.current.set(fileName, version);
persistCaption(fileName, value, version);
},
[persistCaption],
);
function handleScroll() {
const el = scrollRef.current;
if (!el || isLoading || visibleCount >= fileNames.length) return;
const { scrollTop, scrollHeight, clientHeight } = el;
const distanceFromBottom = scrollHeight - (scrollTop + clientHeight);
if (distanceFromBottom < SCROLL_THRESHOLD) {
setVisibleCount((c) => Math.min(c + PAGE_SIZE, fileNames.length));
}
}
// When the visible grid doesn't overflow (a wide/tall sidebar can fit the
// first page without a scrollbar), no scroll event ever fires — so top up
// visibleCount until it overflows or every file is shown. Re-runs on resize
// (width) and after each page grows, otherwise files 25..N are unreachable.
React.useEffect(() => {
if (isLoading || visibleCount >= fileNames.length) return;
const el = scrollRef.current;
if (el && el.scrollHeight <= el.clientHeight) {
setVisibleCount((c) => Math.min(c + PAGE_SIZE, fileNames.length));
}
}, [visibleCount, fileNames.length, isLoading, width]);
const markThumbLoaded = React.useCallback((fileName: string) => {
setThumbLoaded((prev) =>
prev[fileName] ? prev : { ...prev, [fileName]: true },
);
}, []);
const visibleFiles = fileNames.slice(0, visibleCount);
return (
<aside
ref={drawerRef}
tabIndex={-1}
role="dialog"
aria-label={`${dataset.name} dataset details`}
aria-modal={isMobile || undefined}
className="fixed bottom-0 right-0 top-[var(--header-height)] z-50 flex max-h-[calc(100dvh-var(--header-height))] min-w-0 shrink-0 flex-col border-l border-border bg-card md:min-w-[320px]"
style={{
width: isMobile ? '100%' : width,
maxWidth: isMobile ? 'none' : SIDEBAR_MAX_WIDTH,
}}
>
<div className="flex shrink-0 items-center justify-between border-b border-border px-5 py-4">
<h2 className="m-0 min-w-0 truncate text-base font-semibold text-foreground">
{dataset.name}
</h2>
<div className="flex shrink-0 items-center gap-2">
<DownloadCaptions fileNames={fileNames} captions={captions} />
<button
type="button"
onClick={onClose}
title="Close"
aria-label="Close"
className="flex size-11 items-center justify-center rounded-lg text-muted-foreground transition-colors hover:bg-accent hover:text-foreground"
>
<X className="h-[18px] w-[18px]" />
</button>
</div>
</div>
<div className="flex min-h-0 flex-1 flex-col overflow-hidden">
<div
ref={scrollRef}
onScroll={handleScroll}
className="flex-1 overflow-y-auto p-4"
>
{isLoading ? (
<p className="p-8 text-center text-muted-foreground">Loading…</p>
) : fileNames.length === 0 ? (
<p className="p-8 text-center text-muted-foreground">
No media files
</p>
) : (
<div className="grid gap-4 [grid-template-columns:repeat(auto-fill,minmax(140px,1fr))]">
{visibleFiles.map((fileName) => (
<DatasetFileCard
key={fileName}
fileName={fileName}
mediaUrl={getDatasetMediaUrl(dataset.id, fileName)}
caption={captions[fileName] ?? ''}
thumbLoaded={!!thumbLoaded[fileName]}
saveState={captionSaveStates[fileName] ?? 'idle'}
onCaptionChange={handleCaptionChange}
onCaptionRetry={handleCaptionRetry}
onThumbLoaded={markThumbLoaded}
/>
))}
</div>
)}
</div>
</div>
{!isMobile && <div
role="presentation"
onMouseDown={onMouseDown}
className={cn(
'absolute bottom-0 left-0 top-0 z-[1] w-1.5 cursor-col-resize hover:bg-accent-blue/25',
isDragging && 'bg-accent-blue/25',
)}
/>}
</aside>
);
}
@@ -1,113 +0,0 @@
'use client';
import { ChevronDown } from 'lucide-react';
import { Button } from '@/components/ui/button';
import { downloadBlob } from '@/lib/utils';
const MENU_ITEM =
'block min-h-11 w-full cursor-pointer px-4 py-2 text-left text-sm font-medium text-foreground transition-colors hover:bg-muted disabled:cursor-not-allowed disabled:opacity-50';
export default function DownloadCaptions({
fileNames,
captions,
}: {
fileNames: string[];
captions: Record<string, string>;
}) {
const disabled = fileNames.length === 0;
// Sorted lazily in the click handlers: this component re-renders with the
// sidebar on every caption keystroke, and the sort is only needed on click.
const sortedFileNames = () => [...fileNames].sort();
function handleDownloadJson() {
const data = sortedFileNames().map((path) => ({
path,
cap: captions[path] ?? '',
}));
const blob = new Blob([JSON.stringify(data, null, 2)], {
type: 'application/json',
});
downloadBlob(blob, 'videos2caption.json');
}
function handleDownloadTxt() {
const sortedNames = sortedFileNames();
const videosContent = sortedNames.join('\n');
const promptContent = sortedNames
.map((fn) => captions[fn] ?? '')
.join('\n');
downloadBlob(
new Blob([videosContent], { type: 'text/plain' }),
'videos.txt',
);
setTimeout(() => {
downloadBlob(
new Blob([promptContent], { type: 'text/plain' }),
'captions.txt',
);
}, 100);
}
function handleDownloadCsv() {
const escape = (s: string) =>
s.includes('"') || s.includes(',') || s.includes('\n')
? `"${s.replace(/"/g, '""')}"`
: s;
const rows = sortedFileNames().map(
(fn) => `${escape(fn)},${escape(captions[fn] ?? '')}`,
);
const csv = ['video_name,caption', ...rows].join('\n');
downloadBlob(new Blob([csv], { type: 'text/csv' }), 'captions.csv');
}
return (
<div className="group relative inline-block">
<Button
type="button"
variant="outline"
size="sm"
disabled={disabled}
aria-haspopup="menu"
className="gap-1.5"
>
Download Captions
<ChevronDown className="h-3.5 w-3.5 opacity-85" />
</Button>
{!disabled && (
<div
role="menu"
className="invisible absolute right-0 top-full z-[200] min-w-full -translate-y-1 pt-1 opacity-0 transition-all group-focus-within:visible group-focus-within:translate-y-0 group-focus-within:opacity-100 group-hover:visible group-hover:translate-y-0 group-hover:opacity-100"
>
<div className="overflow-hidden rounded-lg border border-border bg-popover py-1 shadow-lg">
<button
type="button"
role="menuitem"
className={MENU_ITEM}
onClick={handleDownloadJson}
>
JSON
</button>
<button
type="button"
role="menuitem"
className={MENU_ITEM}
onClick={handleDownloadTxt}
>
TXT
</button>
<button
type="button"
role="menuitem"
className={MENU_ITEM}
onClick={handleDownloadCsv}
>
CSV
</button>
</div>
</div>
)}
</div>
);
}
@@ -1,205 +0,0 @@
'use client';
import * as React from 'react';
import { cn } from '@/lib/utils';
export interface UploadZoneProps {
label: string;
hint?: string;
accept?: string;
multiple?: boolean;
directory?: boolean;
allowBothFileAndDirectory?: boolean;
value?: string;
fileName?: string;
onFiles?: (files: File[]) => void;
onClear?: () => void;
disabled?: boolean;
uploading?: boolean;
}
const linkButtonClass =
'cursor-pointer bg-transparent p-0 text-accent-blue hover:underline disabled:cursor-not-allowed disabled:no-underline disabled:opacity-50';
export default function UploadZone({
label,
hint,
accept,
multiple = false,
directory = false,
allowBothFileAndDirectory = false,
value = '',
fileName,
onFiles,
onClear,
disabled = false,
uploading = false,
}: UploadZoneProps) {
const fileInputRef = React.useRef<HTMLInputElement>(null);
const directoryInputRef = React.useRef<HTMLInputElement>(null);
const useBoth = directory && allowBothFileAndDirectory;
const clickable = !useBoth;
const hasContent = !!(value || fileName);
// `webkitdirectory` is a DOM property without a typed JSX prop, so set it
// imperatively. The primary input only selects folders when it is the sole
// input; the dedicated directory input always does.
React.useEffect(() => {
if (fileInputRef.current) {
fileInputRef.current.webkitdirectory = directory && !allowBothFileAndDirectory;
}
if (directoryInputRef.current) {
directoryInputRef.current.webkitdirectory = true;
}
}, [directory, allowBothFileAndDirectory]);
function handleChange(e: React.ChangeEvent<HTMLInputElement>) {
const files = e.target.files;
if (files && files.length > 0) {
onFiles?.(Array.from(files));
}
e.target.value = '';
}
function handleClick() {
if (!disabled) {
fileInputRef.current?.click();
}
}
function handleKeyActivate(e: React.KeyboardEvent) {
if ((e.target as HTMLElement).closest('button')) return;
if (e.key === 'Enter' || e.key === ' ') {
e.preventDefault();
handleClick();
}
}
function handleDrop(e: React.DragEvent) {
if (disabled) return;
e.preventDefault();
const files = e.dataTransfer.files;
if (files && files.length > 0) {
onFiles?.(Array.from(files));
}
}
function handleDragOver(e: React.DragEvent) {
if (disabled) return;
e.preventDefault();
}
function clearInputs() {
if (fileInputRef.current) fileInputRef.current.value = '';
if (directoryInputRef.current) directoryInputRef.current.value = '';
}
return (
<div
className={cn(
'flex min-h-[150px] grow flex-col items-center justify-center rounded-lg border-2 border-dashed border-border bg-muted/40 px-5 py-6 text-center transition-colors hover:border-accent-blue hover:bg-muted/60',
clickable ? 'cursor-pointer' : 'cursor-default',
hasContent && 'border-solid border-accent-blue',
)}
onClick={clickable ? handleClick : undefined}
onKeyDown={clickable ? handleKeyActivate : undefined}
onDragOver={handleDragOver}
onDrop={handleDrop}
role={clickable ? 'button' : undefined}
tabIndex={clickable ? 0 : undefined}
>
<input
ref={fileInputRef}
type="file"
className="hidden"
accept={accept}
multiple={multiple}
onChange={handleChange}
disabled={disabled}
/>
{useBoth && (
<input
ref={directoryInputRef}
type="file"
className="hidden"
multiple
onChange={handleChange}
disabled={disabled}
/>
)}
<div className="mb-2 text-sm text-muted-foreground">{label}</div>
{!hasContent && (
<span className="mt-1.5 text-xs text-muted-foreground">
{uploading ? (
'Uploading…'
) : useBoth ? (
<>
<span
role="button"
tabIndex={0}
className="cursor-pointer text-accent-blue hover:underline"
onClick={(e) => {
e.stopPropagation();
handleClick();
}}
onKeyDown={(e) => {
if (e.key === 'Enter' || e.key === ' ') {
e.preventDefault();
handleClick();
}
}}
>
Select files
</span>
{' · '}
<button
type="button"
className={linkButtonClass}
onClick={(e) => {
e.preventDefault();
e.stopPropagation();
if (!disabled) directoryInputRef.current?.click();
}}
disabled={disabled}
>
Select folder
</button>
</>
) : directory ? (
'Click or drop folder'
) : (
'Click or drop file(s)'
)}
</span>
)}
{fileName && (
<div className="mt-2 text-sm text-foreground">
{fileName}
{onClear && (
<>
{' · '}
<button
type="button"
className={linkButtonClass}
onClick={(e) => {
e.stopPropagation();
onClear();
clearInputs();
}}
disabled={disabled || uploading}
>
Clear
</button>
</>
)}
</div>
)}
{hint && <div className="mt-1.5 text-xs text-muted-foreground">{hint}</div>}
</div>
);
}
@@ -1,174 +0,0 @@
'use client';
import * as React from 'react';
import { Input } from '@/components/ui/input';
import { Label } from '@/components/ui/label';
import { Slider } from '@/components/ui/slider';
import { Switch } from '@/components/ui/switch';
import { cn } from '@/lib/utils';
/**
* Labeled form-field rows shared by the Settings page and the Create Job
* modal. They mirror the Svelte `Toggle`/`Slider` UX on top of the shadcn
* `Switch`/`Slider`/`Input` primitives.
*/
export function FieldRow({
htmlFor,
label,
title,
className,
children,
}: {
htmlFor: string;
label: string;
title?: string;
className?: string;
children: React.ReactNode;
}) {
return (
<div className={cn('flex flex-col gap-1.5', className)}>
<Label
htmlFor={htmlFor}
title={title}
className="pl-0.5 text-xs font-normal tracking-wide text-muted-foreground"
>
{label}
</Label>
{children}
</div>
);
}
export function SliderRow({
id,
label,
title,
min,
max,
step,
value,
onChange,
disabled,
format = (v) => String(v),
}: {
id: string;
label: string;
title?: string;
min: number;
max: number;
step: number;
value: number;
/** Called once per gesture (pointer release / key press), not per drag tick. */
onChange: (v: number) => void;
disabled?: boolean;
format?: (v: number) => string;
}) {
// Track the in-progress drag locally so `onChange` only fires on commit;
// any external `value` change (commit landing, reset button) takes over.
const [dragValue, setDragValue] = React.useState<number | null>(null);
React.useEffect(() => {
setDragValue(null);
}, [value]);
const shown = dragValue ?? value;
return (
<FieldRow htmlFor={id} label={label} title={title}>
<div className="flex items-center gap-2">
<Slider
id={id}
min={min}
max={max}
step={step}
value={[shown]}
onValueChange={(v) => setDragValue(v[0])}
onValueCommit={(v) => onChange(v[0])}
disabled={disabled}
aria-label={label}
className="min-w-0 flex-1"
/>
<span
aria-hidden="true"
className="min-w-10 shrink-0 text-right text-sm tabular-nums text-muted-foreground"
>
{format(shown)}
</span>
</div>
</FieldRow>
);
}
export function ToggleRow({
id,
label,
title,
checked,
onChange,
disabled,
}: {
id: string;
label: string;
title?: string;
checked: boolean;
onChange: (v: boolean) => void;
disabled?: boolean;
}) {
return (
<FieldRow htmlFor={id} label={label} title={title}>
<Switch
id={id}
checked={checked}
onCheckedChange={onChange}
disabled={disabled}
/>
</FieldRow>
);
}
export function NumberRow({
id,
label,
title,
min,
max,
step,
value,
onChange,
disabled,
}: {
id: string;
label: string;
title?: string;
min?: number;
max?: number;
step?: number | string;
value: number;
onChange: (v: number) => void;
disabled?: boolean;
}) {
// Buffer the raw text so the field can be emptied while retyping; only
// valid numbers are committed, and blur restores the last committed value.
const [draft, setDraft] = React.useState<string | null>(null);
React.useEffect(() => {
setDraft(null);
}, [value]);
return (
<FieldRow htmlFor={id} label={label} title={title}>
<Input
id={id}
type="number"
min={min}
max={max}
step={step}
value={draft ?? value}
onChange={(e) => {
setDraft(e.target.value);
const v = e.target.valueAsNumber;
if (!Number.isNaN(v)) onChange(v);
}}
onBlur={() => setDraft(null)}
disabled={disabled}
/>
</FieldRow>
);
}
@@ -1,53 +0,0 @@
import { render, screen } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { describe, expect, it, vi } from 'vitest';
import CreateJobButton from './CreateJobButton';
vi.mock('./CreateJobModal', () => ({
default: ({
isOpen,
workloadType,
}: {
isOpen: boolean;
workloadType: string;
}) =>
isOpen ? (
<div role="dialog" data-workload-type={workloadType}>
Create job form
</div>
) : null,
}));
describe('CreateJobButton', () => {
it('opens the workload menu on click and selects an item', async () => {
const user = userEvent.setup();
render(<CreateJobButton jobType="inference" />);
await user.click(screen.getByRole('button', { name: 'Create Job' }));
await user.click(screen.getByRole('menuitem', { name: /I2V/i }));
expect(screen.getByRole('dialog')).toHaveAttribute(
'data-workload-type',
'i2v',
);
});
it('opens and operates the workload menu from the keyboard', async () => {
const user = userEvent.setup();
render(<CreateJobButton jobType="inference" />);
const trigger = screen.getByRole('button', { name: 'Create Job' });
trigger.focus();
await user.keyboard('{Enter}');
const firstItem = await screen.findByRole('menuitem', { name: /T2V/i });
expect(firstItem).toHaveFocus();
await user.keyboard('{Enter}');
expect(screen.getByRole('dialog')).toHaveAttribute(
'data-workload-type',
't2v',
);
});
});
@@ -1,75 +0,0 @@
'use client';
import * as React from 'react';
import { ChevronDown } from 'lucide-react';
import * as DropdownMenu from '@radix-ui/react-dropdown-menu';
import CreateJobModal from '@/components/jobs/CreateJobModal';
import { Button } from '@/components/ui/button';
import { WORKLOAD_OPTIONS } from '@/lib/jobConfig';
import type { JobType } from '@/lib/types';
import { triggerRefresh } from '@/stores/jobsRefresh';
interface CreateJobButtonProps {
jobType: JobType;
}
export default function CreateJobButton({ jobType }: CreateJobButtonProps) {
const options = WORKLOAD_OPTIONS[jobType] ?? [];
const [modalOpen, setModalOpen] = React.useState(false);
const [workloadType, setWorkloadType] = React.useState(
options[0]?.type ?? 't2v',
);
function openModal(type: string) {
setWorkloadType(type);
setModalOpen(true);
}
function handleSuccess() {
triggerRefresh();
setModalOpen(false);
}
return (
<>
<DropdownMenu.Root>
<DropdownMenu.Trigger asChild>
<Button type="button" className="gap-1.5">
Create Job
<ChevronDown className="size-3.5 opacity-85" aria-hidden />
</Button>
</DropdownMenu.Trigger>
<DropdownMenu.Portal>
<DropdownMenu.Content
align="end"
sideOffset={4}
collisionPadding={8}
className="z-[200] min-w-48 overflow-hidden rounded-lg border border-border bg-popover py-1 text-popover-foreground shadow-lg"
>
{options.map((opt) => (
<DropdownMenu.Item
key={opt.type}
onSelect={() => openModal(opt.type)}
className="flex min-h-11 cursor-pointer select-none flex-col justify-center px-4 py-2 text-left text-sm font-medium outline-none data-[highlighted]:bg-secondary"
>
{opt.label}
<span className="mt-0.5 block text-xs font-normal text-muted-foreground">
{opt.desc}
</span>
</DropdownMenu.Item>
))}
</DropdownMenu.Content>
</DropdownMenu.Portal>
</DropdownMenu.Root>
<CreateJobModal
isOpen={modalOpen}
onClose={() => setModalOpen(false)}
onSuccess={handleSuccess}
jobType={jobType}
workloadType={workloadType}
/>
</>
);
}
@@ -1,409 +0,0 @@
import * as React from 'react';
import { render, screen, waitFor, within } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import CreateJobModal from './CreateJobModal';
import {
createJob,
getDatasets,
getModelPresets,
getModels,
listGenerators,
uploadImage,
type GeneratorInfo,
} from '@/lib/api';
import { defaultOptionsStore } from '@/stores/defaultOptions';
import { DEFAULT_OPTIONS } from '@/lib/defaultOptions';
vi.mock('@/lib/api', () => ({
createJob: vi.fn(),
getModels: vi.fn(),
getModelPresets: vi.fn(),
getDatasets: vi.fn(),
listGenerators: vi.fn(),
uploadImage: vi.fn(),
getSettings: vi.fn(),
updateSettings: vi.fn(),
}));
const MODELS = [
{ id: 'wan/t2v-1.3b', label: 'Wan T2V', type: 't2v' },
{ id: 'wan/t2v-14b', label: 'Wan T2V Large', type: 't2v' },
];
// A resident engine slot whose engine config differs from the persisted
// defaults on every field the modal adopts.
const WARM_SLOT: GeneratorInfo = {
state: 'ready',
model_id: 'wan/t2v-14b',
workload_type: 't2v',
num_gpus: 8,
dit_cpu_offload: true,
text_encoder_cpu_offload: true,
vae_cpu_offload: true,
image_encoder_cpu_offload: false,
use_fsdp_inference: true,
enable_torch_compile: true,
vsa_sparsity: 0.5,
tp_size: 1,
sp_size: 8,
error: null,
};
beforeEach(() => {
// Reset the shared options store to a known baseline for test isolation.
defaultOptionsStore.set({ options: DEFAULT_OPTIONS });
vi.mocked(getModels).mockResolvedValue(MODELS);
vi.mocked(getModelPresets).mockResolvedValue({});
vi.mocked(getDatasets).mockResolvedValue([]);
vi.mocked(listGenerators).mockResolvedValue([]);
vi.mocked(uploadImage).mockResolvedValue({ path: '/uploads/x.png' });
vi.mocked(createJob).mockResolvedValue({ id: 'job-1' } as never);
});
function renderModal(
overrides: Partial<React.ComponentProps<typeof CreateJobModal>> = {},
) {
const onClose = vi.fn();
const onSuccess = vi.fn();
render(
<CreateJobModal
isOpen
onClose={onClose}
onSuccess={onSuccess}
jobType="inference"
workloadType="t2v"
{...overrides}
/>,
);
return { onClose, onSuccess };
}
describe('CreateJobModal', () => {
it('shows a model loading error instead of an empty model list', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
vi.mocked(getModels).mockRejectedValueOnce(new Error('network down'));
renderModal();
expect(
await screen.findByText(/Models could not be loaded/),
).toBeInTheDocument();
expect(screen.getByLabelText('Model')).toHaveAttribute(
'aria-invalid',
'true',
);
});
it('keeps the form open and reports job creation failures', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
vi.mocked(createJob).mockRejectedValueOnce(new Error('API rejected job'));
const user = userEvent.setup();
const { onClose, onSuccess } = renderModal();
await screen.findByRole('option', { name: 'Wan T2V (wan/t2v-1.3b)' });
await user.type(screen.getByLabelText('Prompt'), 'a careful test prompt');
await user.click(screen.getByRole('button', { name: 'Create Job' }));
expect(
await screen.findByText(/API rejected job.*then try again/),
).toBeInTheDocument();
expect(onSuccess).not.toHaveBeenCalled();
expect(onClose).not.toHaveBeenCalled();
});
it('reports image upload failures next to the file input', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
vi.mocked(uploadImage).mockRejectedValueOnce(new Error('Upload failed'));
const user = userEvent.setup();
renderModal({ workloadType: 'i2v' });
await screen.findByRole('option', { name: 'Wan T2V (wan/t2v-1.3b)' });
const input = screen.getByLabelText('Image');
await user.upload(
input,
new File(['image'], 'input.png', { type: 'image/png' }),
);
expect(
await screen.findByText(/Upload failed.*Choose the image again/),
).toBeInTheDocument();
expect(input).toHaveAttribute('aria-invalid', 'true');
});
it('renders the form fields for an inference job', async () => {
renderModal();
expect(
await screen.findByText('New Inference Job (T2V)'),
).toBeInTheDocument();
expect(screen.getByLabelText('Model')).toBeInTheDocument();
expect(screen.getByLabelText('Prompt')).toBeInTheDocument();
expect(screen.getByLabelText('Negative Prompt')).toBeInTheDocument();
expect(
screen.getByRole('button', { name: 'Create Job' }),
).toBeInTheDocument();
// The model dropdown is populated once getModels resolves.
expect(
await screen.findByRole('option', {
name: 'Wan T2V (wan/t2v-1.3b)',
}),
).toBeInTheDocument();
});
it('seeds fields from the options store and submits an inference payload', async () => {
// Non-default store values prove the open-time seeding effect ran (the
// useState defaults are 50 / 480).
defaultOptionsStore.set({
options: { ...DEFAULT_OPTIONS, numInferenceSteps: 25, height: 720 },
});
const user = userEvent.setup();
const { onClose, onSuccess } = renderModal();
// Wait for models to load so the default model is selected.
await screen.findByRole('option', { name: 'Wan T2V (wan/t2v-1.3b)' });
await user.type(
screen.getByLabelText('Prompt'),
'a raccoon in sunflowers',
);
await user.click(screen.getByRole('button', { name: 'Create Job' }));
await waitFor(() => expect(createJob).toHaveBeenCalledTimes(1));
const payload = vi.mocked(createJob).mock.calls[0][0];
expect(payload).toMatchObject({
model_id: 'wan/t2v-1.3b',
prompt: 'a raccoon in sunflowers',
workload_type: 't2v',
job_type: 'inference',
num_inference_steps: 25,
height: 720,
num_frames: 81,
width: 832,
guidance_scale: 5,
seed: 1024,
});
await waitFor(() => expect(onSuccess).toHaveBeenCalledTimes(1));
expect(onClose).toHaveBeenCalledTimes(1);
});
it('submits a dmd_t2v distillation payload including the DMD fields', async () => {
vi.mocked(getDatasets).mockResolvedValue([
{ id: 'ds1', name: 'My Dataset', created_at: 0 },
]);
const user = userEvent.setup();
const { onSuccess } = renderModal({
jobType: 'distillation',
workloadType: 'dmd_t2v',
});
// Models + datasets load asynchronously on open. The model/dataset option
// labels also appear in the Real/Fake Score Model and Validation Dataset
// selects, so scope each wait to the relevant select.
await within(screen.getByLabelText('Model')).findByRole('option', {
name: 'Wan T2V (wan/t2v-1.3b)',
});
const datasetSelect = screen.getByLabelText('Dataset *');
await within(datasetSelect).findByRole('option', { name: 'My Dataset' });
await user.type(screen.getByLabelText('Description'), 'distill run');
await user.selectOptions(datasetSelect, 'ds1');
await user.click(screen.getByRole('button', { name: 'Create Job' }));
await waitFor(() => expect(createJob).toHaveBeenCalledTimes(1));
const payload = vi.mocked(createJob).mock.calls[0][0];
expect(payload).toMatchObject({
workload_type: 'dmd_t2v',
job_type: 'distillation',
// The dataset id is sent; the backend resolves it to the on-disk dir.
data_path: 'ds1',
lora_rank: 32,
// DMD-specific fields added to CreateJobRequest for this modal.
dmd_use_vsa: false,
dmd_vsa_sparsity: 0.8,
dmd_denoising_steps: '1000,757,522',
real_score_guidance_scale: 3.5,
generator_update_interval: 5,
real_score_model_path: 'wan/t2v-1.3b',
fake_score_model_path: 'wan/t2v-1.3b',
});
// Inference-only keys must be absent for a training job.
expect(payload).not.toHaveProperty('num_inference_steps');
await waitFor(() => expect(onSuccess).toHaveBeenCalledTimes(1));
});
it('defaults to the warm resident model and adopts its engine config', async () => {
vi.mocked(listGenerators).mockResolvedValue([WARM_SLOT]);
const user = userEvent.setup();
renderModal();
// Without a warm model the default logic picks the first model
// (wan/t2v-1.3b, covered by the seeding test above); the warm slot wins.
await waitFor(() =>
expect(screen.getByLabelText('Model')).toHaveValue('wan/t2v-14b'),
);
// The warm selection also triggers its presets fetch.
await waitFor(() =>
expect(getModelPresets).toHaveBeenCalledWith('wan/t2v-14b'),
);
await user.type(screen.getByLabelText('Prompt'), 'warm run');
await user.click(screen.getByRole('button', { name: 'Create Job' }));
await waitFor(() => expect(createJob).toHaveBeenCalledTimes(1));
// The engine fields mirror the resident slot (not the persisted defaults:
// num_gpus 1, all offloads/fsdp/compile false) so the job reuses the warm
// instance instead of replacing it.
expect(vi.mocked(createJob).mock.calls[0][0]).toMatchObject({
model_id: 'wan/t2v-14b',
num_gpus: 8,
tp_size: 1,
sp_size: 8,
dit_cpu_offload: true,
text_encoder_cpu_offload: true,
vae_cpu_offload: true,
image_encoder_cpu_offload: false,
use_fsdp_inference: true,
enable_torch_compile: true,
vsa_sparsity: 0.5,
});
});
it('restores persisted-default engine fields when switching away from the warm model', async () => {
defaultOptionsStore.set({
options: { ...DEFAULT_OPTIONS, numGpus: 2, tpSize: 2 },
});
vi.mocked(listGenerators).mockResolvedValue([WARM_SLOT]);
const user = userEvent.setup();
renderModal();
await waitFor(() =>
expect(screen.getByLabelText('Model')).toHaveValue('wan/t2v-14b'),
);
await user.selectOptions(screen.getByLabelText('Model'), 'wan/t2v-1.3b');
await user.type(screen.getByLabelText('Prompt'), 'cold run');
await user.click(screen.getByRole('button', { name: 'Create Job' }));
await waitFor(() => expect(createJob).toHaveBeenCalledTimes(1));
expect(vi.mocked(createJob).mock.calls[0][0]).toMatchObject({
model_id: 'wan/t2v-1.3b',
num_gpus: 2,
tp_size: 2,
use_fsdp_inference: false,
enable_torch_compile: false,
});
});
it('applies resolution preset chips and the orientation toggle to the payload', async () => {
const user = userEvent.setup();
renderModal();
await screen.findByRole('option', { name: 'Wan T2V (wan/t2v-1.3b)' });
await user.click(screen.getByText('Options'));
// 720p chip sets the /32-rounded dims; the orientation toggle swaps them.
await user.click(screen.getByRole('button', { name: '720p' }));
await user.click(screen.getByRole('button', { name: 'Landscape' }));
expect(
screen.getByRole('button', { name: 'Portrait' }),
).toBeInTheDocument();
await user.type(screen.getByLabelText('Prompt'), 'portrait 720p');
await user.click(screen.getByRole('button', { name: 'Create Job' }));
await waitFor(() => expect(createJob).toHaveBeenCalledTimes(1));
expect(vi.mocked(createJob).mock.calls[0][0]).toMatchObject({
height: 1280,
width: 704,
});
});
it('restores the model preset resolution via the Native chip', async () => {
vi.mocked(getModelPresets).mockResolvedValue({ height: 720, width: 1280 });
const user = userEvent.setup();
renderModal();
await screen.findByRole('option', { name: 'Wan T2V (wan/t2v-1.3b)' });
// Native is disabled until the selected model's presets have loaded.
await waitFor(() =>
expect(screen.getByRole('button', { name: 'Native' })).toBeEnabled(),
);
await user.click(screen.getByText('Options'));
await user.click(screen.getByRole('button', { name: '1080p' }));
await user.click(screen.getByRole('button', { name: 'Native' }));
await user.type(screen.getByLabelText('Prompt'), 'native dims');
await user.click(screen.getByRole('button', { name: 'Create Job' }));
await waitFor(() => expect(createJob).toHaveBeenCalledTimes(1));
expect(vi.mocked(createJob).mock.calls[0][0]).toMatchObject({
height: 720,
width: 1280,
});
});
it('populates sampling fields from the selected model presets; engine fields stay from defaults', async () => {
defaultOptionsStore.set({
options: { ...DEFAULT_OPTIONS, numGpus: 4, tpSize: 2, seed: 999 },
});
vi.mocked(getModelPresets).mockImplementation(async (id) =>
id === 'wan/t2v-14b'
? {
height: 720,
width: 1280,
num_frames: 121,
fps: 30,
num_inference_steps: 40,
guidance_scale: 6,
guidance_rescale: 0.5,
negative_prompt: 'blurry, low quality',
seed: 7,
}
: {},
);
const user = userEvent.setup();
renderModal();
await screen.findByRole('option', { name: 'Wan T2V Large (wan/t2v-14b)' });
await user.selectOptions(screen.getByLabelText('Model'), 'wan/t2v-14b');
// The negative prompt is the easiest preset-populated field to observe.
await waitFor(() =>
expect(screen.getByLabelText('Negative Prompt')).toHaveValue(
'blurry, low quality',
),
);
await user.type(screen.getByLabelText('Prompt'), 'preset test');
await user.click(screen.getByRole('button', { name: 'Create Job' }));
await waitFor(() => expect(createJob).toHaveBeenCalledTimes(1));
const payload = vi.mocked(createJob).mock.calls[0][0];
expect(payload).toMatchObject({
model_id: 'wan/t2v-14b',
// Sampling fields come from the model presets…
height: 720,
width: 1280,
num_frames: 121,
fps: 30,
num_inference_steps: 40,
guidance_scale: 6,
guidance_rescale: 0.5,
negative_prompt: 'blurry, low quality',
seed: 7,
// …while engine fields still come from the persisted defaults.
num_gpus: 4,
tp_size: 2,
});
});
});
File diff suppressed because it is too large Load Diff
@@ -1,71 +0,0 @@
import { act, fireEvent, render, screen } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import EngineConsole from './EngineConsole';
import { getEngineLogs } from '@/lib/api';
vi.mock('@/lib/api', () => ({
getEngineLogs: vi.fn(),
}));
beforeEach(() => {
vi.mocked(getEngineLogs).mockResolvedValue({
lines: ['[engine] booted', '[engine] worker heartbeat ok'],
total: 2,
});
});
describe('EngineConsole', () => {
it('is collapsed by default and does not fetch', () => {
render(<EngineConsole />);
expect(
screen.getByRole('button', { name: 'Engine output' }),
).toHaveAttribute('aria-expanded', 'false');
expect(getEngineLogs).not.toHaveBeenCalled();
});
it('shows log lines and polls with the cursor while open', async () => {
vi.useFakeTimers();
try {
render(<EngineConsole />);
fireEvent.click(screen.getByRole('button', { name: 'Engine output' }));
// Flush the immediate poll fired on expand.
await act(async () => {
await vi.advanceTimersByTimeAsync(0);
});
expect(getEngineLogs).toHaveBeenCalledWith(0);
expect(screen.getByText(/\[engine\] booted/)).toBeInTheDocument();
// The 2s interval polls again, from the previous total.
await act(async () => {
await vi.advanceTimersByTimeAsync(2000);
});
expect(getEngineLogs).toHaveBeenLastCalledWith(2);
// Collapsing stops the polling.
fireEvent.click(screen.getByRole('button', { name: 'Engine output' }));
const calls = vi.mocked(getEngineLogs).mock.calls.length;
await act(async () => {
await vi.advanceTimersByTimeAsync(10000);
});
expect(getEngineLogs).toHaveBeenCalledTimes(calls);
} finally {
vi.useRealTimers();
}
});
it('clear view empties the scrollback locally', async () => {
const user = userEvent.setup();
render(<EngineConsole />);
await user.click(screen.getByRole('button', { name: 'Engine output' }));
expect(await screen.findByText(/\[engine\] booted/)).toBeInTheDocument();
await user.click(screen.getByRole('button', { name: 'Clear view' }));
expect(screen.queryByText(/\[engine\] booted/)).not.toBeInTheDocument();
expect(screen.getByText('Waiting for engine output…')).toBeInTheDocument();
});
});
@@ -1,123 +0,0 @@
'use client';
import * as React from 'react';
import { ChevronDown, ChevronRight } from 'lucide-react';
import { Button } from '@/components/ui/button';
import { getEngineLogs } from '@/lib/api';
const POLL_INTERVAL_MS = 2000;
// Cap the DOM at the last ~500 lines; the server keeps its own ring buffer.
const MAX_LINES = 500;
// The engine tail includes uvicorn's access log (until a server restart picks
// up access_log=False); the frontend's own polling would otherwise flood the
// console with "GET /api/... 200 OK" lines. Keep non-GET and non-2xx lines.
const ACCESS_LOG_NOISE = /^INFO:\s+[\d.:]+\s+- "(?:GET|HEAD) \S+ HTTP\/[\d.]+" 2\d\d/;
/**
* Collapsible tail of the engine's stdout/stderr (driver + relayed worker
* output). Polls only while open; sticks to the bottom unless the user has
* scrolled up.
*/
export default function EngineConsole() {
const [open, setOpen] = React.useState(false);
const [lines, setLines] = React.useState<string[]>([]);
// Poll cursor + stick-to-bottom flag live in refs: they must update
// synchronously from async polls / scroll events, outside React's cycle.
const afterRef = React.useRef(0);
const stickRef = React.useRef(true);
const consoleRef = React.useRef<HTMLPreElement | null>(null);
React.useEffect(() => {
if (!open) return;
let mounted = true;
let locked = false;
async function poll() {
if (!mounted || locked) return;
locked = true;
try {
const data = await getEngineLogs(afterRef.current);
afterRef.current = data.total;
const fresh = data.lines.filter((l) => !ACCESS_LOG_NOISE.test(l));
if (mounted && fresh.length > 0) {
setLines((prev) => [...prev, ...fresh].slice(-MAX_LINES));
}
} catch (e) {
console.error('Failed to fetch engine logs:', e);
} finally {
locked = false;
}
}
poll();
const interval = setInterval(poll, POLL_INTERVAL_MS);
return () => {
mounted = false;
clearInterval(interval);
};
}, [open]);
// Follow the tail after new lines land, unless the user scrolled up.
React.useEffect(() => {
const el = consoleRef.current;
if (el && stickRef.current) el.scrollTop = el.scrollHeight;
}, [lines]);
function handleScroll() {
const el = consoleRef.current;
if (!el) return;
stickRef.current = el.scrollHeight - el.scrollTop - el.clientHeight < 40;
}
return (
<section
aria-label="Engine output"
className="mx-auto w-full max-w-[850px] px-10 pt-3"
>
<div className="rounded-lg border border-border bg-background">
<div className="flex items-center gap-2 px-2 py-1.5">
<button
type="button"
onClick={() => setOpen((o) => !o)}
aria-expanded={open}
className="flex flex-1 items-center gap-2 rounded-md px-2 py-1 text-sm font-semibold text-foreground hover:bg-accent"
>
{open ? (
<ChevronDown className="h-4 w-4" />
) : (
<ChevronRight className="h-4 w-4" />
)}
Engine output
</button>
{open && (
<Button
size="sm"
variant="ghost"
// Resets the local view only; the server buffer is untouched.
onClick={() => setLines([])}
>
Clear view
</Button>
)}
</div>
{open && (
<pre
ref={consoleRef}
onScroll={handleScroll}
className="m-0 h-64 overflow-auto whitespace-pre-wrap break-words rounded-b-lg border-t border-border bg-zinc-950 p-3 font-mono text-xs leading-normal text-zinc-200"
>
{lines.length === 0 ? (
<span className="italic text-zinc-500">
Waiting for engine output…
</span>
) : (
lines.join('\n')
)}
</pre>
)}
</div>
</section>
);
}
@@ -1,135 +0,0 @@
import { render, screen, waitFor } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import JobCard from '@/components/jobs/JobCard';
import {
deleteJob,
downloadJobVideo,
startJob,
stopJob,
} from '@/lib/api';
import type { Job } from '@/lib/types';
import { activeJobStore, setActiveJobId } from '@/stores/activeJob';
import { makeJob as makeBaseJob } from '@/test/factories';
vi.mock('@/lib/api', () => ({
startJob: vi.fn(),
stopJob: vi.fn(),
deleteJob: vi.fn(),
downloadJobVideo: vi.fn(),
}));
vi.mock('@/lib/utils', async (importOriginal) => {
const actual = await importOriginal<typeof import('@/lib/utils')>();
return {
...actual,
downloadBlob: vi.fn(),
};
});
const makeJob = (overrides: Partial<Job> = {}): Job =>
makeBaseJob({
model_id: 'Wan2.1-T2V',
prompt: 'a cat surfing a wave',
created_at: 1_700_000_000,
num_inference_steps: 50,
num_frames: 81,
height: 480,
width: 832,
guidance_scale: 5,
seed: 42,
num_gpus: 1,
...overrides,
});
beforeEach(() => {
setActiveJobId(null);
vi.mocked(startJob).mockResolvedValue({} as Job);
vi.mocked(stopJob).mockResolvedValue({} as Job);
vi.mocked(deleteJob).mockResolvedValue(undefined);
vi.mocked(downloadJobVideo).mockResolvedValue(new Blob());
vi.spyOn(window, 'confirm').mockReturnValue(true);
vi.spyOn(window, 'alert').mockImplementation(() => {});
});
describe('JobCard', () => {
it('keeps selection and job action buttons as semantic siblings', () => {
render(<JobCard job={makeJob()} />);
const selectButton = screen.getByRole('button', { pressed: false });
const deleteButton = screen.getByRole('button', { name: 'Delete' });
expect(selectButton).toHaveTextContent('Wan2.1-T2V');
expect(selectButton).not.toContainElement(deleteButton);
});
it('renders the model, prompt, status and inference meta', () => {
render(<JobCard job={makeJob()} />);
expect(screen.getByText('Wan2.1-T2V')).toBeInTheDocument();
expect(screen.getByText('a cat surfing a wave')).toBeInTheDocument();
expect(screen.getByText('pending')).toBeInTheDocument();
expect(screen.getByText('81 frames')).toBeInTheDocument();
expect(screen.getByText('480×832')).toBeInTheDocument();
});
it('shows the workload type (not frames) for non-inference jobs', () => {
render(
<JobCard
job={makeJob({ job_type: 'finetuning', workload_type: 'lora_t2v' })}
/>,
);
expect(screen.getByText('lora t2v')).toBeInTheDocument();
expect(screen.queryByText('81 frames')).not.toBeInTheDocument();
});
it('starts a pending job and notifies the parent', async () => {
const onJobUpdated = vi.fn();
render(<JobCard job={makeJob({ status: 'pending' })} onJobUpdated={onJobUpdated} />);
await userEvent.click(screen.getByRole('button', { name: 'Start' }));
await waitFor(() => expect(startJob).toHaveBeenCalledWith('job-1'));
expect(onJobUpdated).toHaveBeenCalled();
});
it('stops a running job', async () => {
render(<JobCard job={makeJob({ status: 'running', started_at: Date.now() })} />);
await userEvent.click(screen.getByRole('button', { name: 'Stop' }));
await waitFor(() => expect(stopJob).toHaveBeenCalledWith('job-1'));
});
it('restarts a failed job via startJob', async () => {
render(<JobCard job={makeJob({ status: 'failed' })} />);
await userEvent.click(screen.getByRole('button', { name: 'Restart' }));
await waitFor(() => expect(startJob).toHaveBeenCalledWith('job-1'));
});
it('deletes when confirmed and skips when cancelled', async () => {
const { rerender } = render(<JobCard job={makeJob()} />);
await userEvent.click(screen.getByRole('button', { name: 'Delete' }));
await waitFor(() => expect(deleteJob).toHaveBeenCalledWith('job-1'));
vi.mocked(deleteJob).mockClear();
vi.mocked(window.confirm).mockReturnValue(false);
rerender(<JobCard job={makeJob()} />);
await userEvent.click(screen.getByRole('button', { name: 'Delete' }));
expect(deleteJob).not.toHaveBeenCalled();
});
it('downloads the video for a completed inference job', async () => {
render(
<JobCard
job={makeJob({ status: 'completed', output_path: '/out/video.mp4' })}
/>,
);
await userEvent.click(
screen.getByRole('button', { name: 'Download Video' }),
);
await waitFor(() => expect(downloadJobVideo).toHaveBeenCalledWith('job-1'));
});
it('selects the job when the card body is clicked', async () => {
render(<JobCard job={makeJob()} />);
await userEvent.click(screen.getByText('Wan2.1-T2V'));
expect(activeJobStore.get().activeJobId).toBe('job-1');
});
});
@@ -1,242 +0,0 @@
'use client';
import * as React from 'react';
import { Timer } from 'lucide-react';
import { Badge, type BadgeProps } from '@/components/ui/badge';
import { Button } from '@/components/ui/button';
import { useStore } from '@/hooks/useStore';
import {
deleteJob,
downloadJobVideo,
startJob,
stopJob,
} from '@/lib/api';
import type { Job } from '@/lib/types';
import { cn, downloadBlob } from '@/lib/utils';
import { activeJobStore, setActiveJobId } from '@/stores/activeJob';
interface JobCardProps {
job: Job;
onJobUpdated?: () => void;
}
function formatDuration(seconds: number): string {
const roundedSeconds = Math.round(seconds);
if (roundedSeconds < 60) return `${roundedSeconds}s`;
if (roundedSeconds < 3600) {
const mins = Math.floor(roundedSeconds / 60);
const secs = roundedSeconds % 60;
return secs > 0 ? `${mins}m ${secs}s` : `${mins}m`;
}
const hours = Math.floor(roundedSeconds / 3600);
const mins = Math.floor((roundedSeconds % 3600) / 60);
return mins > 0 ? `${hours}h ${mins}m` : `${hours}h`;
}
function computeElapsed(job: Job, currentTime: number): string | null {
if (!job.started_at) return null;
const endTime =
job.status === 'running' ? currentTime : (job.finished_at ?? 0);
if (!endTime && job.status !== 'running') return null;
const startedAtMs =
job.started_at < 1e12 ? job.started_at * 1000 : job.started_at;
const endTimeMs = endTime < 1e12 ? endTime * 1000 : endTime;
const elapsedSeconds = (endTimeMs - startedAtMs) / 1000;
if (elapsedSeconds <= 0) return null;
return formatDuration(elapsedSeconds);
}
const BADGE_VARIANTS: Record<string, BadgeProps['variant']> = {
pending: 'secondary',
running: 'warning',
completed: 'success',
ready: 'success',
failed: 'destructive',
stopped: 'secondary',
preprocessing: 'default',
};
export default function JobCard({ job, onJobUpdated }: JobCardProps) {
const { activeJobId } = useStore(activeJobStore);
const isSelected = activeJobId === job.id;
const [isLoading, setIsLoading] = React.useState(false);
const [currentTime, setCurrentTime] = React.useState(() => Date.now());
const elapsedTime = computeElapsed(job, currentTime);
React.useEffect(() => {
if (job.status !== 'running' || !job.started_at) return;
const interval = setInterval(() => setCurrentTime(Date.now()), 1000);
return () => clearInterval(interval);
}, [job.status, job.started_at]);
async function handleStart(e: React.MouseEvent) {
e.preventDefault();
e.stopPropagation();
if (isLoading || job.status === 'running' || job.status === 'completed')
return;
setIsLoading(true);
try {
await startJob(job.id);
onJobUpdated?.();
} catch (err) {
alert(err instanceof Error ? err.message : 'Failed to start job');
} finally {
setIsLoading(false);
}
}
async function handleStop(e: React.MouseEvent) {
e.preventDefault();
e.stopPropagation();
if (isLoading || job.status !== 'running') return;
setIsLoading(true);
try {
await stopJob(job.id);
onJobUpdated?.();
} catch (err) {
alert(err instanceof Error ? err.message : 'Failed to stop job');
} finally {
setIsLoading(false);
}
}
async function handleDelete(e: React.MouseEvent) {
e.preventDefault();
e.stopPropagation();
if (isLoading) return;
if (!confirm('Delete this job?')) return;
setIsLoading(true);
try {
await deleteJob(job.id);
onJobUpdated?.();
} catch (err) {
alert(err instanceof Error ? err.message : 'Failed to delete job');
} finally {
setIsLoading(false);
}
}
function handleSelectJob() {
setActiveJobId(isSelected ? null : job.id);
}
async function handleDownloadVideo(e: React.MouseEvent) {
e.preventDefault();
e.stopPropagation();
if (isLoading || !job.output_path) return;
setIsLoading(true);
try {
const blob = await downloadJobVideo(job.id);
const ext = job.output_path.endsWith('.png') ? 'png' : 'mp4';
downloadBlob(blob, `job_${job.id}.${ext}`);
} catch (err) {
alert(err instanceof Error ? err.message : 'Failed to download video');
} finally {
setIsLoading(false);
}
}
return (
<article
className={cn(
'mb-1.5 flex cursor-pointer flex-col gap-1 rounded-lg border bg-background px-3 py-1.5 transition-colors last:mb-0',
isSelected
? 'border-accent-blue bg-accent-blue/5'
: 'border-border hover:border-muted-foreground/40',
)}
>
<button
type="button"
aria-pressed={isSelected}
onClick={handleSelectJob}
className="flex w-full min-w-0 flex-col gap-1 rounded-md text-left"
>
<span className="flex w-full min-w-0 items-center gap-2">
<span className="shrink-0 text-sm font-semibold text-foreground">
{job.model_id}
</span>
<Badge variant={BADGE_VARIANTS[job.status] ?? 'secondary'}>
{job.status}
</Badge>
<span className="ml-auto flex shrink-0 items-center gap-3 text-xs text-muted-foreground">
{job.job_type === 'inference' ? (
<>
<span>{job.num_frames} frames</span>
<span>
{job.height}×{job.width}
</span>
</>
) : (
<span>{job.workload_type?.replace(/_/g, ' ') ?? job.job_type}</span>
)}
{elapsedTime && (
<span className="inline-flex items-center gap-1">
<Timer className="size-3.5" aria-hidden />
{elapsedTime}
</span>
)}
</span>
</span>
<span className="w-full whitespace-pre-wrap break-words text-xs text-muted-foreground">
{job.prompt}
</span>
</button>
<div className="flex flex-wrap items-center gap-1.5">
{job.status === 'running' ? (
<Button
size="sm"
onClick={handleStop}
disabled={isLoading}
className="h-6 px-2 text-xs border-transparent bg-amber-500 text-black shadow-md hover:bg-amber-400"
>
Stop
</Button>
) : job.status === 'failed' ? (
<Button
size="sm"
onClick={handleStart}
disabled={isLoading}
className="h-6 px-2 text-xs border-transparent bg-emerald-600 text-white shadow-md hover:bg-emerald-500"
>
Restart
</Button>
) : job.status === 'pending' || job.status === 'stopped' ? (
<Button
size="sm"
onClick={handleStart}
disabled={isLoading}
className="h-6 px-2 text-xs border-transparent bg-emerald-600 text-white shadow-md hover:bg-emerald-500"
>
Start
</Button>
) : null}
{job.status === 'completed' &&
job.output_path &&
job.job_type === 'inference' && (
<Button
size="sm"
variant="outline"
onClick={handleDownloadVideo}
disabled={isLoading}
title="Download video"
className="h-6 px-2 text-xs"
>
Download Video
</Button>
)}
<Button
size="sm"
variant="destructive"
onClick={handleDelete}
disabled={isLoading}
className="h-6 px-2 text-xs"
>
Delete
</Button>
</div>
</article>
);
}
@@ -1,174 +0,0 @@
import { act, render, screen } from '@testing-library/react';
import { describe, expect, it, vi } from 'vitest';
import JobDetailsSidebar from './JobDetailsSidebar';
import { getJobLogs } from '@/lib/api';
import type { Job } from '@/lib/types';
import { makeJob as makeBaseJob } from '@/test/factories';
vi.mock('@/lib/api', () => ({
getJobLogs: vi.fn(),
downloadJobLog: vi.fn(),
getJobVideoUrl: (id: string) => `http://test.local/api/jobs/${id}/video`,
}));
const makeJob = (overrides: Partial<Job> = {}): Job =>
makeBaseJob({
status: 'running',
log_file_path: '/logs/job-1.log',
...overrides,
});
describe('JobDetailsSidebar', () => {
it('fills the mobile viewport without reserving main-content width', async () => {
vi.mocked(getJobLogs).mockResolvedValue({
lines: [],
total: 0,
progress: 0,
progress_msg: '',
phase: '',
});
const onWidthChange = vi.fn();
render(
<JobDetailsSidebar
job={makeJob({ status: 'completed' })}
isMobile
onClose={vi.fn()}
onWidthChange={onWidthChange}
/>,
);
const drawer = screen.getByRole('dialog', { name: 'Job details' });
expect(drawer).toHaveStyle({ width: '100%', maxWidth: 'none' });
expect(drawer).toHaveAttribute('aria-modal', 'true');
expect(drawer).toHaveFocus();
expect(onWidthChange).toHaveBeenCalledWith(0);
});
it('plays completed inference output inline; running jobs get no player', async () => {
vi.mocked(getJobLogs).mockResolvedValue({
lines: [],
total: 0,
progress: 0,
progress_msg: '',
phase: '',
});
const { rerender } = render(
<JobDetailsSidebar
job={makeJob({
status: 'completed',
output_path: '/outputs/job-1.mp4',
prompt: 'a cat surfing a wave',
})}
onClose={vi.fn()}
/>,
);
const video = screen.getByLabelText('Generated video: a cat surfing a wave');
expect(video.tagName).toBe('VIDEO');
expect(video).toHaveAttribute('controls');
expect(video).toHaveAttribute(
'src',
'http://test.local/api/jobs/job-1/video',
);
rerender(
<JobDetailsSidebar
job={makeJob({ status: 'running', output_path: null })}
onClose={vi.fn()}
/>,
);
expect(screen.queryByLabelText(/Generated video/)).not.toBeInTheDocument();
});
it('renders log lines streamed from the job log poll', async () => {
vi.mocked(getJobLogs).mockResolvedValue({
lines: ['boot sequence started', 'loading model weights'],
total: 2,
progress: 0,
progress_msg: '',
phase: '',
});
render(
<JobDetailsSidebar job={makeJob({ status: 'running' })} onClose={vi.fn()} />,
);
expect(await screen.findByText(/boot sequence started/)).toBeInTheDocument();
expect(screen.getByText(/loading model weights/)).toBeInTheDocument();
expect(getJobLogs).toHaveBeenCalledWith('job-1', 0);
});
it('keeps polling while the job is running', async () => {
vi.useFakeTimers();
try {
vi.mocked(getJobLogs).mockResolvedValue({
lines: [],
total: 0,
progress: 0,
progress_msg: '',
phase: '',
});
render(
<JobDetailsSidebar
job={makeJob({ status: 'running' })}
onClose={vi.fn()}
/>,
);
// Flush the immediate poll fired on mount.
await act(async () => {
await vi.advanceTimersByTimeAsync(0);
});
const initialCalls = vi.mocked(getJobLogs).mock.calls.length;
// Two 2s interval ticks should fire while the job is running.
await act(async () => {
await vi.advanceTimersByTimeAsync(4000);
});
expect(vi.mocked(getJobLogs).mock.calls.length).toBeGreaterThan(
initialCalls,
);
} finally {
vi.useRealTimers();
}
});
it('stops polling once the job is completed', async () => {
vi.useFakeTimers();
try {
vi.mocked(getJobLogs).mockResolvedValue({
lines: ['final line'],
total: 1,
progress: 0,
progress_msg: '',
phase: '',
});
render(
<JobDetailsSidebar
job={makeJob({ status: 'completed' })}
onClose={vi.fn()}
/>,
);
// The component fetches once on mount even for terminal jobs.
await act(async () => {
await vi.advanceTimersByTimeAsync(0);
});
expect(getJobLogs).toHaveBeenCalledTimes(1);
// No interval is registered, so advancing the clock must not re-poll.
await act(async () => {
await vi.advanceTimersByTimeAsync(10000);
});
expect(getJobLogs).toHaveBeenCalledTimes(1);
} finally {
vi.useRealTimers();
}
});
});
@@ -1,255 +0,0 @@
'use client';
import * as React from 'react';
import { X } from 'lucide-react';
import { Button } from '@/components/ui/button';
import { useDrawerFocus } from '@/hooks/useDrawerFocus';
import { useResizable } from '@/hooks/useResizable';
import { downloadJobLog, getJobLogs, getJobVideoUrl } from '@/lib/api';
import type { Job } from '@/lib/types';
import { cn, downloadBlob } from '@/lib/utils';
const SIDEBAR_MIN_WIDTH = 280;
const SIDEBAR_MAX_WIDTH = 750;
const POLL_INTERVAL_MS = 2000;
export default function JobDetailsSidebar({
job,
isMobile = false,
onClose,
onWidthChange,
}: {
job: Job;
isMobile?: boolean;
onClose: () => void;
onWidthChange?: (w: number) => void;
}) {
const drawerRef = useDrawerFocus<HTMLElement>(isMobile);
const [width, setWidth] = React.useState(360);
const [isDragging, setIsDragging] = React.useState(false);
const [isLoading, setIsLoading] = React.useState(false);
const [logs, setLogs] = React.useState('');
// Race-guard ref (mirrors the Svelte original): the cursor + accumulated
// text live here so in-flight polls don't read stale React state. Do NOT
// swap these for effect deps — they must update synchronously, outside
// React's render cycle. The log is one growing string (amortized O(new)
// appends), not an array re-copied and re-joined on every poll tick.
const stateRef = React.useRef<{ text: string; logAfter: number }>({
text: '',
logAfter: 0,
});
const previousJobId = React.useRef<string | null>(null);
const previousStatus = React.useRef<string | null>(null);
const consoleRef = React.useRef<HTMLPreElement | null>(null);
const { onMouseDown } = useResizable({
// Right-docked panel with the drag handle on its left edge: dragging the
// handle left must grow the panel, which is `edge: 'right'` (matches the
// Svelte original). `edge: 'left'` would invert the drag.
edge: 'right',
minWidth: SIDEBAR_MIN_WIDTH,
maxWidth: SIDEBAR_MAX_WIDTH,
getWidth: () => width,
onWidth: setWidth,
onDragChange: setIsDragging,
});
React.useEffect(() => {
onWidthChange?.(isMobile ? 0 : width);
}, [isMobile, width, onWidthChange]);
// Auto-scroll the console to the bottom whenever new lines land. Runs after
// commit so scrollHeight reflects the freshly-rendered output.
React.useEffect(() => {
const el = consoleRef.current;
if (el) el.scrollTop = el.scrollHeight;
}, [logs]);
// Reset on job switch or restart. Declared before the polling effect so the
// cursor is cleared before the next poll reads it.
React.useEffect(() => {
const wasTerminal =
previousStatus.current === 'failed' ||
previousStatus.current === 'stopped' ||
previousStatus.current === 'completed';
const isRestarting =
previousJobId.current === job.id &&
wasTerminal &&
(job.status === 'pending' || job.status === 'running');
if (previousJobId.current !== job.id || isRestarting) {
stateRef.current.text = '';
stateRef.current.logAfter = 0;
setLogs('');
}
previousJobId.current = job.id;
previousStatus.current = job.status;
}, [job.id, job.status]);
React.useEffect(() => {
const shouldPoll = job.status === 'running' || job.status === 'pending';
let pollInterval: ReturnType<typeof setInterval> | null = null;
let mounted = true;
// Effect-local lock so this job's first fetch is never blocked by a
// previous job's in-flight poll (stale writes are dropped via `mounted`).
let locked = false;
async function pollLogs() {
if (!mounted || locked) return;
locked = true;
try {
const logData = await getJobLogs(job.id, stateRef.current.logAfter);
if (mounted && logData.lines.length > 0) {
const chunk = logData.lines.join('\n');
stateRef.current.text = stateRef.current.text
? `${stateRef.current.text}\n${chunk}`
: chunk;
stateRef.current.logAfter = logData.total;
setLogs(stateRef.current.text);
}
} catch (e) {
console.error('Failed to fetch logs:', e);
} finally {
locked = false;
}
}
pollLogs();
if (shouldPoll) pollInterval = setInterval(pollLogs, POLL_INTERVAL_MS);
return () => {
mounted = false;
if (pollInterval) clearInterval(pollInterval);
};
}, [job.id, job.status]);
async function handleDownloadLog() {
if (isLoading) return;
setIsLoading(true);
try {
const blob = await downloadJobLog(job.id);
downloadBlob(blob, `job_${job.id}.log`);
} catch (err) {
console.error('Failed to download log:', err);
alert(err instanceof Error ? err.message : 'Failed to download log');
} finally {
setIsLoading(false);
}
}
return (
<aside
ref={drawerRef}
tabIndex={-1}
role="dialog"
aria-label="Job details"
aria-modal={isMobile || undefined}
className="fixed bottom-0 right-0 top-[var(--header-height)] z-50 flex max-h-[calc(100dvh-var(--header-height))] min-w-0 shrink-0 flex-col border-l border-border bg-card md:min-w-[280px]"
style={{
width: isMobile ? '100%' : width,
maxWidth: isMobile ? 'none' : SIDEBAR_MAX_WIDTH,
}}
>
<div className="flex items-center justify-between border-b border-border px-5 py-4">
<h2 className="m-0 text-base font-semibold text-foreground">
Job Details
</h2>
<div className="flex items-center gap-2">
<Button
type="button"
variant="outline"
size="sm"
onClick={handleDownloadLog}
disabled={isLoading || !job.log_file_path}
title="Download log file"
>
Download Log
</Button>
<Button
type="button"
variant="ghost"
size="icon-sm"
onClick={onClose}
title="Close"
aria-label="Close"
>
<X className="h-[18px] w-[18px]" />
</Button>
</div>
</div>
{job.status === 'completed' &&
job.output_path &&
(job.job_type === 'inference' || !job.job_type) && (
<div className="border-b border-border px-5 py-4">
<span className="mb-2 block text-xs font-semibold uppercase tracking-wider text-muted-foreground">
Output
</span>
{job.output_path.toLowerCase().endsWith('.png') ? (
// eslint-disable-next-line @next/next/no-img-element
<img
src={getJobVideoUrl(job.id)}
alt={
job.prompt
? `Generated image: ${job.prompt}`
: 'Generated image'
}
className="block w-full rounded-lg border border-border bg-background object-contain"
/>
) : (
<video
src={getJobVideoUrl(job.id)}
aria-label={
job.prompt
? `Generated video: ${job.prompt}`
: 'Generated video'
}
controls
playsInline
preload="metadata"
className="block max-h-80 w-full rounded-lg border border-border bg-background object-contain"
/>
)}
</div>
)}
<div className="flex min-h-0 flex-1 flex-col px-5 py-4">
<div className="mb-2 flex items-center justify-between">
<span className="text-xs font-semibold uppercase tracking-wider text-muted-foreground">
Console Output
</span>
{job.status === 'running' && (
<span className="text-[0.7rem] font-medium text-emerald-600 dark:text-emerald-400">
● Live
</span>
)}
</div>
<pre
ref={consoleRef}
className="m-0 min-h-0 flex-1 overflow-auto whitespace-pre-wrap break-words rounded-lg border border-border bg-background p-3 font-mono text-xs leading-normal text-foreground"
>
{logs === '' ? (
<span className="italic text-muted-foreground">
{job.status === 'running'
? 'Waiting for logs...'
: 'No logs available'}
</span>
) : (
logs
)}
</pre>
</div>
{!isMobile && <div
role="presentation"
onMouseDown={onMouseDown}
className={cn(
'absolute bottom-0 left-0 top-0 z-[1] w-1.5 cursor-col-resize hover:bg-accent-blue/25',
isDragging && 'bg-accent-blue/25',
)}
/>}
</aside>
);
}
@@ -1,122 +0,0 @@
import { act, fireEvent, render, screen, waitFor } from '@testing-library/react';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import JobQueue from '@/components/jobs/JobQueue';
import { getJobsList } from '@/lib/api';
import type { Job, JobType } from '@/lib/types';
import { setActiveJobId } from '@/stores/activeJob';
import { triggerRefresh } from '@/stores/jobsRefresh';
import { makeJob as makeBaseJob } from '@/test/factories';
vi.mock('@/lib/api', () => ({
getJobsList: vi.fn(),
startJob: vi.fn(),
stopJob: vi.fn(),
deleteJob: vi.fn(),
downloadJobVideo: vi.fn(),
}));
const makeJob = (overrides: Partial<Job> = {}): Job =>
makeBaseJob({
model_id: 'Wan2.1-T2V',
status: 'completed',
created_at: 1_700_000_000,
num_inference_steps: 50,
num_frames: 81,
height: 480,
width: 832,
guidance_scale: 5,
seed: 42,
num_gpus: 1,
...overrides,
});
beforeEach(() => {
setActiveJobId(null);
vi.mocked(getJobsList).mockResolvedValue([]);
});
describe('JobQueue', () => {
it('shows a loading placeholder before the initial request settles', async () => {
let resolveJobs: (jobs: Job[]) => void = () => {};
vi.mocked(getJobsList).mockReturnValue(
new Promise<Job[]>((resolve) => {
resolveJobs = resolve;
}),
);
render(<JobQueue jobType="inference" />);
expect(screen.getByLabelText('Loading jobs')).toBeInTheDocument();
expect(
screen.queryByText('No inference jobs yet. Create one above.'),
).not.toBeInTheDocument();
act(() => resolveJobs([]));
expect(
await screen.findByText('No inference jobs yet. Create one above.'),
).toBeInTheDocument();
});
it('shows request failures separately from an empty queue and retries', async () => {
vi.spyOn(console, 'error').mockImplementation(() => {});
vi.mocked(getJobsList).mockRejectedValueOnce(new Error('network down'));
render(<JobQueue jobType="inference" />);
expect(
await screen.findByText(/Could not load jobs from the Studio API/),
).toBeInTheDocument();
expect(
screen.queryByText('No inference jobs yet. Create one above.'),
).not.toBeInTheDocument();
vi.mocked(getJobsList).mockResolvedValueOnce([]);
fireEvent.click(screen.getByRole('button', { name: 'Try Again' }));
expect(
await screen.findByText('No inference jobs yet. Create one above.'),
).toBeInTheDocument();
});
it('shows an empty placeholder and fetches for the single job type', async () => {
render(<JobQueue jobType="inference" />);
expect(
await screen.findByText('No inference jobs yet. Create one above.'),
).toBeInTheDocument();
expect(getJobsList).toHaveBeenCalledWith('inference');
});
it('renders a JobCard for each fetched job', async () => {
vi.mocked(getJobsList).mockResolvedValue([
makeJob({ id: 'a', model_id: 'Model-A' }),
]);
render(<JobQueue jobType="distillation" />);
expect(await screen.findByText('Model-A')).toBeInTheDocument();
});
it('merges all job types when jobTypesForList is provided', async () => {
vi.mocked(getJobsList).mockImplementation((t?: JobType) =>
Promise.resolve(
t === ('lora' as JobType)
? [makeJob({ id: 'l', model_id: 'Lora-Model', created_at: 2 })]
: [makeJob({ id: 'f', model_id: 'Full-Model', created_at: 1 })],
),
);
render(
<JobQueue
jobType="finetuning"
jobTypesForList={['finetuning', 'lora'] as JobType[]}
/>,
);
expect(await screen.findByText('Lora-Model')).toBeInTheDocument();
expect(screen.getByText('Full-Model')).toBeInTheDocument();
expect(getJobsList).toHaveBeenCalledWith('finetuning');
expect(getJobsList).toHaveBeenCalledWith('lora');
});
it('refetches when a refresh is triggered', async () => {
render(<JobQueue jobType="inference" />);
await waitFor(() => expect(getJobsList).toHaveBeenCalledTimes(1));
act(() => triggerRefresh());
await waitFor(() => expect(getJobsList).toHaveBeenCalledTimes(2));
});
});
@@ -1,195 +0,0 @@
'use client';
import * as React from 'react';
import { AlertTriangle } from 'lucide-react';
import JobCard from '@/components/jobs/JobCard';
import { Button } from '@/components/ui/button';
import { useStore } from '@/hooks/useStore';
import { getJobsList } from '@/lib/api';
import type { Job, JobType } from '@/lib/types';
import {
activeJobStore,
setActiveJob,
setActiveJobId,
} from '@/stores/activeJob';
import { jobsRefreshStore } from '@/stores/jobsRefresh';
interface JobQueueProps {
jobType: JobType;
jobTypesForList?: JobType[];
}
// Jobs are flat objects of primitives, so a shallow compare detects "nothing
// changed" across poll responses (which are referentially fresh every fetch).
function jobsShallowEqual(a: Job | null, b: Job | null): boolean {
if (a === b) return true;
if (!a || !b) return false;
const aKeys = Object.keys(a) as (keyof Job)[];
return (
aKeys.length === Object.keys(b).length &&
aKeys.every((k) => a[k] === b[k])
);
}
export default function JobQueue({ jobType, jobTypesForList }: JobQueueProps) {
const [jobs, setJobs] = React.useState<Job[]>([]);
const [isInitialLoading, setIsInitialLoading] = React.useState(true);
const [error, setError] = React.useState<string | null>(null);
const { nonce } = useStore(jobsRefreshStore);
const { activeJobId } = useStore(activeJobStore);
// Stable primitive key so the memoized list below keeps its identity when an
// inline array prop (e.g. ['finetuning', 'lora']) gets a fresh reference each
// render, which would otherwise re-run the fetch effects needlessly.
const typesKey = (jobTypesForList ?? [jobType]).join(',');
const typesToFetch = React.useMemo<JobType[]>(
() => jobTypesForList ?? [jobType],
// eslint-disable-next-line react-hooks/exhaustive-deps
[typesKey],
);
// Guard the poll against slow responses: `inFlight` lets the interval skip
// a tick instead of stacking requests, and the sequence counter drops
// out-of-order responses so a delayed older payload can't overwrite a newer
// job list (direct refetches always run and supersede in-flight polls).
const fetchSeq = React.useRef(0);
const inFlight = React.useRef(false);
const fetchJobs = React.useCallback(async () => {
const seq = ++fetchSeq.current;
inFlight.current = true;
try {
let next: Job[];
if (typesToFetch.length === 1) {
next = await getJobsList(typesToFetch[0]);
} else {
const results = await Promise.all(
typesToFetch.map((t) => getJobsList(t)),
);
next = results
.flat()
.sort(
(a, b) =>
new Date(b.created_at ?? 0).getTime() -
new Date(a.created_at ?? 0).getTime(),
);
}
if (seq === fetchSeq.current) {
setJobs(next);
setError(null);
}
} catch (e) {
console.error('Failed to fetch jobs:', e);
if (seq === fetchSeq.current) {
setError(
'Could not load jobs from the Studio API. Check the server and try again.',
);
}
} finally {
if (seq === fetchSeq.current) {
inFlight.current = false;
setIsInitialLoading(false);
}
}
}, [typesKey]);
// Keep the latest fetchJobs reachable from the polling interval without
// tearing it down/recreating it on every fetchJobs identity change.
const fetchJobsRef = React.useRef(fetchJobs);
React.useEffect(() => {
fetchJobsRef.current = fetchJobs;
}, [fetchJobs]);
// Initial fetch + refetch whenever an external refresh is triggered.
React.useEffect(() => {
fetchJobs();
}, [fetchJobs, nonce]);
// Poll every second while any job is running/pending; stop otherwise.
const hasActive = jobs.some(
(j) => j.status === 'running' || j.status === 'pending',
);
React.useEffect(() => {
if (!hasActive) return;
const interval = setInterval(() => {
if (!inFlight.current) fetchJobsRef.current();
}, 1000);
return () => clearInterval(interval);
}, [hasActive]);
// Keep the active job in sync with the selected id, and void the selection
// if its job disappeared (e.g. deleted). Skip the store write when nothing
// changed — each poll returns fresh objects, and an unconditional write
// would re-render every subscriber (shell, sidebars, cards) once a second.
React.useEffect(() => {
if (!activeJobId) {
if (activeJobStore.get().activeJob) setActiveJob(null);
return;
}
const activeJob = jobs.find((j) => j.id === activeJobId) ?? null;
if (!jobsShallowEqual(activeJob, activeJobStore.get().activeJob)) {
setActiveJob(activeJob);
}
if (!activeJob) setActiveJobId(null);
}, [activeJobId, jobs]);
const multiType = typesToFetch.length > 1;
return (
<div className="mx-auto flex w-full max-w-[850px] flex-col gap-6 px-4 pb-12">
<section className="p-6">
<div aria-busy={isInitialLoading}>
{isInitialLoading ? (
<div
aria-label="Loading jobs"
className="flex flex-col gap-3 py-2"
>
{[0, 1, 2].map((item) => (
<div
key={item}
className="h-32 animate-pulse rounded-lg border border-border bg-muted/50"
/>
))}
</div>
) : error && jobs.length === 0 ? (
<div
role="alert"
className="flex flex-col items-center gap-3 py-8 text-center"
>
<AlertTriangle
className="size-6 text-destructive"
aria-hidden
/>
<p className="max-w-md text-sm text-muted-foreground">{error}</p>
<Button type="button" variant="outline" onClick={fetchJobs}>
Try Again
</Button>
</div>
) : (
<>
{error && (
<p
role="status"
className="mb-3 rounded-lg border border-amber-500/50 bg-amber-500/10 px-3 py-2 text-sm text-foreground"
>
Job updates are temporarily unavailable. Showing the most
recent results.
</p>
)}
{jobs.length === 0 ? (
<p className="py-8 text-center text-muted-foreground">
No {multiType ? 'jobs' : `${jobType} jobs`} yet. Create one above.
</p>
) : (
jobs.map((job) => (
<JobCard key={job.id} job={job} onJobUpdated={fetchJobs} />
))
)}
</>
)}
</div>
</section>
</div>
);
}
@@ -1,188 +0,0 @@
import { render, screen, waitFor } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import WarmModelsPanel from './WarmModelsPanel';
import {
getModels,
listGenerators,
preloadGenerator,
unloadGenerator,
type GeneratorInfo,
} from '@/lib/api';
import { DEFAULT_OPTIONS } from '@/lib/defaultOptions';
import { defaultOptionsStore } from '@/stores/defaultOptions';
import { toast } from 'sonner';
vi.mock('@/lib/api', () => ({
getModels: vi.fn(),
listGenerators: vi.fn(),
preloadGenerator: vi.fn(),
unloadGenerator: vi.fn(),
getSettings: vi.fn(),
updateSettings: vi.fn(),
}));
vi.mock('sonner', () => ({
toast: { error: vi.fn() },
}));
const makeGenerator = (
overrides: Partial<GeneratorInfo> = {},
): GeneratorInfo => ({
state: 'ready',
model_id: 'Wan-AI/Wan2.1-T2V-1.3B-Diffusers',
workload_type: 't2v',
num_gpus: 1,
dit_cpu_offload: false,
text_encoder_cpu_offload: false,
vae_cpu_offload: false,
image_encoder_cpu_offload: false,
use_fsdp_inference: false,
enable_torch_compile: false,
vsa_sparsity: 0,
tp_size: -1,
sp_size: -1,
error: null,
...overrides,
});
beforeEach(() => {
// Reset the shared options store to a known baseline for test isolation.
defaultOptionsStore.set({ options: DEFAULT_OPTIONS });
vi.mocked(getModels).mockResolvedValue([
{ id: 'Wan-AI/Wan2.1-T2V-1.3B-Diffusers', label: 'Wan2.1 T2V 1.3B' },
]);
vi.mocked(listGenerators).mockResolvedValue([]);
vi.mocked(preloadGenerator).mockResolvedValue(
makeGenerator({ state: 'loading' }),
);
vi.mocked(unloadGenerator).mockResolvedValue(undefined);
});
describe('WarmModelsPanel', () => {
it('shows the empty slot when no model is loaded', async () => {
render(<WarmModelsPanel />);
expect(await screen.findByText('No model loaded')).toBeInTheDocument();
expect(
screen.queryByRole('button', { name: 'Unload' }),
).not.toBeInTheDocument();
});
it('renders the resident slot with its state and config summary', async () => {
vi.mocked(listGenerators).mockResolvedValue([
makeGenerator({ num_gpus: 8, enable_torch_compile: true }),
]);
render(<WarmModelsPanel />);
// Model label: last path segment with dashes/underscores as spaces.
expect(
await screen.findByText('Wan2.1 T2V 1.3B Diffusers'),
).toBeInTheDocument();
expect(screen.getByText('ready')).toBeInTheDocument();
expect(screen.getByText('8 GPU · compile')).toBeInTheDocument();
expect(screen.getByRole('button', { name: 'Unload' })).toBeInTheDocument();
});
it('disables loading a new model while a load is in flight', async () => {
vi.mocked(listGenerators).mockResolvedValue([
makeGenerator({ state: 'loading' }),
]);
render(<WarmModelsPanel />);
expect(await screen.findByText('loading')).toBeInTheDocument();
expect(screen.getByRole('button', { name: 'Load model' })).toBeDisabled();
expect(
screen.queryByRole('button', { name: 'Unload' }),
).not.toBeInTheDocument();
});
it('shows the error on a failed slot and keeps retry enabled', async () => {
vi.mocked(listGenerators).mockResolvedValue([
makeGenerator({ state: 'failed', error: 'CUDA out of memory' }),
]);
render(<WarmModelsPanel />);
expect(await screen.findByText('failed')).toHaveAttribute(
'title',
'CUDA out of memory',
);
await waitFor(() =>
expect(screen.getByRole('button', { name: 'Load model' })).toBeEnabled(),
);
});
it('labels the load button as a swap when a different model is resident', async () => {
vi.mocked(listGenerators).mockResolvedValue([
makeGenerator({ model_id: 'FastVideo/FastHunyuan-diffusers' }),
]);
render(<WarmModelsPanel />);
expect(
await screen.findByRole('button', { name: 'Load (replaces current)' }),
).toBeInTheDocument();
});
it('loads the selected model with the persisted default options', async () => {
defaultOptionsStore.set({
options: { ...DEFAULT_OPTIONS, numGpus: 4, enableTorchCompile: true },
});
const user = userEvent.setup();
render(<WarmModelsPanel />);
const button = await screen.findByRole('button', { name: 'Load model' });
await waitFor(() => expect(button).toBeEnabled());
await user.click(button);
await waitFor(() =>
expect(preloadGenerator).toHaveBeenCalledWith({
model_id: 'Wan-AI/Wan2.1-T2V-1.3B-Diffusers',
workload_type: 't2v',
num_gpus: 4,
dit_cpu_offload: false,
text_encoder_cpu_offload: false,
vae_cpu_offload: false,
image_encoder_cpu_offload: false,
use_fsdp_inference: false,
enable_torch_compile: true,
vsa_sparsity: 0,
tp_size: -1,
sp_size: -1,
}),
);
// The panel refetches so the new "loading" slot appears promptly.
await waitFor(() => expect(listGenerators).toHaveBeenCalledTimes(2));
});
it('surfaces the backend detail when a load is rejected', async () => {
vi.mocked(preloadGenerator).mockRejectedValue(
new Error('a model load is already in progress'),
);
const user = userEvent.setup();
render(<WarmModelsPanel />);
const button = await screen.findByRole('button', { name: 'Load model' });
await waitFor(() => expect(button).toBeEnabled());
await user.click(button);
await waitFor(() =>
expect(toast.error).toHaveBeenCalledWith('Model was not loaded', {
description: 'a model load is already in progress',
}),
);
});
it('unloads the resident model with no payload', async () => {
vi.mocked(listGenerators).mockResolvedValue([makeGenerator()]);
const user = userEvent.setup();
render(<WarmModelsPanel />);
await user.click(await screen.findByRole('button', { name: 'Unload' }));
await waitFor(() => expect(unloadGenerator).toHaveBeenCalledTimes(1));
expect(vi.mocked(unloadGenerator).mock.calls[0]).toEqual([]);
// The slot refreshes after the unload succeeds.
await waitFor(() => expect(listGenerators).toHaveBeenCalledTimes(2));
});
});
@@ -1,237 +0,0 @@
'use client';
import * as React from 'react';
import { toast } from 'sonner';
import { Badge, type BadgeProps } from '@/components/ui/badge';
import { Button } from '@/components/ui/button';
import { NativeSelect } from '@/components/ui/native-select';
import { useStore } from '@/hooks/useStore';
import {
getModels,
listGenerators,
preloadGenerator,
unloadGenerator,
type GeneratorInfo,
type Model,
} from '@/lib/api';
import { getDefaultModelForWorkload } from '@/lib/defaultOptions';
import { defaultOptionsStore } from '@/stores/defaultOptions';
// Mirrors the backend's model_label(): readable label from an HF-style path.
function modelLabel(modelId: string): string {
return (modelId.split('/').pop() ?? modelId).replace(/[-_]/g, ' ');
}
// Compact summary of only the non-default engine bits, e.g. "8 GPU · compile".
function configSummary(gen: GeneratorInfo): string {
const parts: string[] = [];
if (gen.num_gpus !== 1) parts.push(`${gen.num_gpus} GPU`);
if (gen.sp_size !== -1) parts.push(`SP ${gen.sp_size}`);
if (gen.tp_size !== -1) parts.push(`TP ${gen.tp_size}`);
if (gen.dit_cpu_offload) parts.push('DiT offload');
if (gen.text_encoder_cpu_offload) parts.push('TE offload');
if (gen.vae_cpu_offload) parts.push('VAE offload');
if (gen.image_encoder_cpu_offload) parts.push('image enc offload');
if (gen.use_fsdp_inference) parts.push('FSDP');
if (gen.enable_torch_compile) parts.push('compile');
if (gen.vsa_sparsity > 0) parts.push(`VSA ${gen.vsa_sparsity.toFixed(2)}`);
return parts.join(' · ');
}
const STATE_VARIANTS: Record<GeneratorInfo['state'], BadgeProps['variant']> = {
ready: 'success',
loading: 'warning',
failed: 'destructive',
};
/**
* Utility strip for the engine's single model slot: shows the resident model
* (ready/loading/failed), loads the selected model using the persisted
* default job options (replacing whatever is resident), and unloads it.
*/
export default function WarmModelsPanel() {
const { options } = useStore(defaultOptionsStore);
const [slot, setSlot] = React.useState<GeneratorInfo | null>(null);
const [models, setModels] = React.useState<Model[]>([]);
const [modelId, setModelId] = React.useState('');
const [isBusy, setIsBusy] = React.useState(false);
const fetchSlot = React.useCallback(async () => {
try {
const list = await listGenerators();
setSlot(list[0] ?? null);
} catch (e) {
console.error('Failed to fetch generators:', e);
}
}, []);
React.useEffect(() => {
fetchSlot();
}, [fetchSlot]);
// Poll every 5s while a load is in flight; stop otherwise.
const isLoading = slot?.state === 'loading';
React.useEffect(() => {
if (!isLoading) return;
const interval = setInterval(fetchSlot, 5000);
return () => clearInterval(interval);
}, [isLoading, fetchSlot]);
// Same model catalogue (and default selection) as the create-job modal.
React.useEffect(() => {
getModels('t2v')
.then((list) => {
setModels(list);
const defaultId = getDefaultModelForWorkload(
defaultOptionsStore.get().options,
't2v',
);
setModelId(
list.some((m) => m.id === defaultId)
? defaultId
: (list[0]?.id ?? ''),
);
})
.catch((e) => console.error('Failed to load models:', e));
}, []);
async function handleLoad() {
if (!modelId || isBusy || isLoading) return;
setIsBusy(true);
try {
await preloadGenerator({
model_id: modelId,
workload_type: 't2v',
num_gpus: options.numGpus,
dit_cpu_offload: options.ditCpuOffload,
text_encoder_cpu_offload: options.textEncoderCpuOffload,
vae_cpu_offload: options.vaeCpuOffload,
image_encoder_cpu_offload: options.imageEncoderCpuOffload,
use_fsdp_inference: options.useFsdpInference,
enable_torch_compile: options.enableTorchCompile,
vsa_sparsity: options.vsaSparsity,
tp_size: options.tpSize,
sp_size: options.spSize,
});
await fetchSlot();
} catch (err) {
console.error('Failed to load model:', err);
toast.error('Model was not loaded', {
description:
err instanceof Error
? err.message
: 'Check the Studio API, then retry.',
});
} finally {
setIsBusy(false);
}
}
async function handleUnload() {
if (isBusy) return;
setIsBusy(true);
try {
await unloadGenerator();
await fetchSlot();
} catch (err) {
console.error('Failed to unload model:', err);
toast.error('Model was not unloaded', {
description:
err instanceof Error
? err.message
: 'Check the Studio API, then retry.',
});
} finally {
setIsBusy(false);
}
}
// Loading a model always replaces the resident one — say so on the button.
const replaces = slot !== null && !!modelId && slot.model_id !== modelId;
return (
<section
aria-label="Warm models"
className="mx-auto w-full max-w-[850px] px-10 pt-6"
>
<div className="flex flex-col gap-3 rounded-lg border border-border bg-background p-4">
<div className="flex flex-wrap items-center gap-2">
<h2 className="mr-auto text-sm font-semibold text-foreground">
Warm Model
</h2>
<label htmlFor="warm-model-select" className="sr-only">
Model to load
</label>
<NativeSelect
id="warm-model-select"
value={modelId}
onChange={(e) => setModelId(e.target.value)}
disabled={isBusy || models.length === 0}
className="h-9 w-auto max-w-64 rounded-lg"
>
<option value="" disabled>
{models.length === 0 ? 'Loading models…' : 'Select a model…'}
</option>
{models.map((model) => (
<option key={model.id} value={model.id}>
{model.label}
</option>
))}
</NativeSelect>
<Button
size="sm"
onClick={handleLoad}
disabled={isBusy || !modelId || isLoading}
>
{replaces ? 'Load (replaces current)' : 'Load model'}
</Button>
</div>
<p className="text-xs text-muted-foreground">
One model at a time stays resident in GPU memory so jobs skip the
load wait; loading a new one replaces it (uses your default job
options).
</p>
<div className="flex min-h-8 flex-wrap items-center gap-2">
{slot ? (
<>
<Badge
variant={STATE_VARIANTS[slot.state]}
className={
slot.state === 'loading' ? 'animate-pulse' : undefined
}
title={
slot.state === 'failed' ? (slot.error ?? undefined) : undefined
}
>
{slot.state}
</Badge>
<span className="text-sm font-medium text-foreground">
{modelLabel(slot.model_id)}
</span>
<span className="text-xs text-muted-foreground">
{configSummary(slot)}
</span>
{slot.state === 'ready' && (
<Button
size="sm"
variant="outline"
className="ml-auto"
onClick={handleUnload}
disabled={isBusy}
>
Unload
</Button>
)}
</>
) : (
<span className="text-sm text-muted-foreground">
No model loaded
</span>
)}
</div>
</div>
</section>
);
}
@@ -1,125 +0,0 @@
'use client';
import * as React from 'react';
import { usePathname } from 'next/navigation';
import DatasetSidebar from '@/components/datasets/DatasetSidebar';
import Header from '@/components/shell/Header';
import { HeaderActionsProvider } from '@/components/shell/HeaderActionsContext';
import PrimarySidebar from '@/components/shell/PrimarySidebar';
import JobDetailsSidebar from '@/components/jobs/JobDetailsSidebar';
import { Toaster } from '@/components/ui/sonner';
import { useMediaQuery } from '@/hooks/useMediaQuery';
import { useStore } from '@/hooks/useStore';
import {
activeDatasetStore,
setActiveDatasetId,
} from '@/stores/activeDataset';
import { activeJobStore, setActiveJobId } from '@/stores/activeJob';
import { initDefaultOptions } from '@/stores/defaultOptions';
const JOB_ROUTES = ['/inference', '/finetuning', '/distillation'];
export function AppShell({ children }: { children: React.ReactNode }) {
const pathname = usePathname();
const { activeJob } = useStore(activeJobStore);
const { activeDataset } = useStore(activeDatasetStore);
const isMobile = useMediaQuery('(max-width: 767px)');
const [primaryWidth, setPrimaryWidth] = React.useState(220);
const [secondaryWidth, setSecondaryWidth] = React.useState(0);
const [primaryOpen, setPrimaryOpen] = React.useState(false);
const jobSidebarOpen = JOB_ROUTES.includes(pathname) && activeJob != null;
const datasetSidebarOpen =
pathname === '/datasets' && activeDataset != null;
const secondaryOpen = jobSidebarOpen || datasetSidebarOpen;
// Mobile detail drawers claim aria-modal, so everything behind them must
// actually be inert — the platform enforces what the ARIA claims.
const drawerModal = isMobile && secondaryOpen;
React.useEffect(() => {
initDefaultOptions();
}, []);
React.useEffect(() => {
setPrimaryOpen(false);
}, [pathname]);
React.useEffect(() => {
function handleKeyDown(e: KeyboardEvent) {
if (e.key !== 'Escape' || document.querySelector('[data-modal]')) return;
if (primaryOpen) {
setPrimaryOpen(false);
return;
}
if (activeJobStore.get().activeJob) setActiveJobId(null);
if (activeDatasetStore.get().activeDataset) setActiveDatasetId(null);
}
document.addEventListener('keydown', handleKeyDown);
return () => document.removeEventListener('keydown', handleKeyDown);
}, [primaryOpen]);
return (
<HeaderActionsProvider>
<div
style={{ display: 'contents' }}
inert={drawerModal ? true : undefined}
>
<Header
navigationOpen={primaryOpen}
onNavigationToggle={() => setPrimaryOpen((open) => !open)}
/>
</div>
<div
className="flex overflow-hidden"
style={{
marginTop: 'var(--header-height)',
height: 'calc(100dvh - var(--header-height))',
}}
>
<PrimarySidebar
isMobile={isMobile}
mobileOpen={primaryOpen}
onMobileClose={() => setPrimaryOpen(false)}
onWidthChange={setPrimaryWidth}
/>
{primaryOpen && (
<button
type="button"
aria-label="Close navigation"
onClick={() => setPrimaryOpen(false)}
className="fixed inset-x-0 bottom-0 top-[var(--header-height)] z-40 bg-black/55 md:hidden"
/>
)}
<main
className="flex min-w-0 flex-1 flex-col overflow-auto"
inert={drawerModal ? true : undefined}
style={{
marginLeft: isMobile ? 0 : primaryWidth,
marginRight: isMobile || !secondaryOpen ? 0 : secondaryWidth,
}}
>
{children}
</main>
{jobSidebarOpen && activeJob && (
<JobDetailsSidebar
job={activeJob}
isMobile={isMobile}
onClose={() => setActiveJobId(null)}
onWidthChange={setSecondaryWidth}
/>
)}
{datasetSidebarOpen && activeDataset && (
<DatasetSidebar
dataset={activeDataset}
isMobile={isMobile}
onClose={() => setActiveDatasetId(null)}
onWidthChange={setSecondaryWidth}
/>
)}
</div>
<Toaster />
</HeaderActionsProvider>
);
}
@@ -1,66 +0,0 @@
'use client';
import { Menu, X } from 'lucide-react';
import { usePathname } from 'next/navigation';
import { useHeaderActions } from '@/components/shell/HeaderActionsContext';
import { Button } from '@/components/ui/button';
import { ThemeToggle } from '@/components/ui/theme-toggle';
const TAB_TITLES: Record<string, string> = {
'/inference': 'Studio',
'/finetuning': 'Studio',
'/distillation': 'Studio',
'/datasets': 'Datasets',
'/gallery': 'Gallery',
'/gpus': 'GPUs',
'/settings': 'Settings',
};
export default function Header({
navigationOpen,
onNavigationToggle,
}: {
navigationOpen: boolean;
onNavigationToggle: () => void;
}) {
const pathname = usePathname();
const { actions } = useHeaderActions();
const title = TAB_TITLES[pathname] ?? 'FastVideo';
return (
<header className="fixed inset-x-0 top-0 z-[100] flex h-[var(--header-height)] items-center gap-2 border-b border-border bg-background/80 px-2 backdrop-blur sm:px-4 md:gap-6 md:px-6">
<Button
type="button"
variant="outline"
size="icon"
aria-label={navigationOpen ? 'Close navigation' : 'Open navigation'}
aria-controls="primary-navigation"
aria-expanded={navigationOpen}
onClick={onNavigationToggle}
className="shrink-0 md:hidden"
>
{navigationOpen ? (
<X className="size-5" aria-hidden />
) : (
<Menu className="size-5" aria-hidden />
)}
</Button>
{/* eslint-disable-next-line @next/next/no-img-element */}
<img
src="/logo.svg"
alt="FastVideo Logo"
width={100}
height={42}
className="hidden h-[42px] w-[78px] shrink-0 object-contain min-[361px]:block md:w-[100px]"
/>
<h1 className="sr-only m-0 flex-1 text-xl font-semibold tracking-tight md:not-sr-only">
{title}
</h1>
<div className="ml-auto flex min-w-0 items-center gap-2 md:gap-3">
{actions}
<ThemeToggle />
</div>
</header>
);
}
@@ -1,52 +0,0 @@
'use client';
import * as React from 'react';
interface HeaderActionsContextValue {
actions: React.ReactNode;
setActions: (node: React.ReactNode) => void;
}
const HeaderActionsContext =
React.createContext<HeaderActionsContextValue | null>(null);
export function HeaderActionsProvider({
children,
}: {
children: React.ReactNode;
}) {
const [actions, setActions] = React.useState<React.ReactNode>(null);
const value = React.useMemo(() => ({ actions, setActions }), [actions]);
return (
<HeaderActionsContext.Provider value={value}>
{children}
</HeaderActionsContext.Provider>
);
}
export function useHeaderActions(): HeaderActionsContextValue {
const ctx = React.useContext(HeaderActionsContext);
if (!ctx) {
throw new Error(
'useHeaderActions must be used within a HeaderActionsProvider',
);
}
return ctx;
}
/**
* Declaratively publish the current page's header actions: renders nothing,
* registers `children` in the header on mount and clears them on unmount.
* Pages without actions simply don't render it.
*/
export function HeaderActions({ children }: { children: React.ReactNode }) {
const { setActions } = useHeaderActions();
// Mount-only on purpose: pages pass inline JSX, which is referentially new
// every render and would re-register per render if it were a dependency.
React.useEffect(() => {
setActions(children);
return () => setActions(null);
// eslint-disable-next-line react-hooks/exhaustive-deps
}, [setActions]);
return null;
}
@@ -1,218 +0,0 @@
'use client';
import * as React from 'react';
import { X } from 'lucide-react';
import Link from 'next/link';
import { usePathname } from 'next/navigation';
import { useResizable } from '@/hooks/useResizable';
import { cn } from '@/lib/utils';
const SIDEBAR_MIN_WIDTH = 100;
const SIDEBAR_MAX_WIDTH = 300;
const SIDEBAR_COLLAPSED_WIDTH = 0;
const SIDEBAR_COLLAPSED_VISIBLE_WIDTH = 60;
const JOB_ROUTES = [
{ href: '/inference', label: 'Inference' },
{ href: '/finetuning', label: 'Finetuning' },
{ href: '/distillation', label: 'Distillation' },
] as const;
const TAB_BASE =
'block min-h-11 px-5 py-[0.65rem] text-left text-sm text-muted-foreground transition-colors hover:bg-accent/60 hover:text-foreground';
const TAB_ACTIVE = 'bg-accent-blue/10 font-medium text-accent-blue';
export default function PrimarySidebar({
isMobile,
mobileOpen,
onMobileClose,
onWidthChange,
}: {
isMobile: boolean;
mobileOpen: boolean;
onMobileClose: () => void;
onWidthChange?: (w: number) => void;
}) {
const pathname = usePathname();
const [width, setWidth] = React.useState(220);
const [isCollapsed, setIsCollapsed] = React.useState(false);
const [isDragging, setIsDragging] = React.useState(false);
const [jobsOpen, setJobsOpen] = React.useState(true);
const effectiveWidth = isCollapsed ? SIDEBAR_COLLAPSED_WIDTH : width;
const layoutWidth = isCollapsed ? SIDEBAR_COLLAPSED_VISIBLE_WIDTH : width;
const isJobsActive = JOB_ROUTES.some((r) => pathname === r.href);
React.useEffect(() => {
onWidthChange?.(isMobile ? 0 : layoutWidth);
}, [isMobile, layoutWidth, onWidthChange]);
React.useEffect(() => {
if (JOB_ROUTES.some((r) => pathname === r.href)) {
setJobsOpen(true);
}
}, [pathname]);
const { onMouseDown } = useResizable({
edge: 'left',
minWidth: SIDEBAR_MIN_WIDTH,
maxWidth: SIDEBAR_MAX_WIDTH,
getWidth: () => width,
onWidth: setWidth,
onDragChange: setIsDragging,
});
return (
<aside
id="primary-navigation"
aria-hidden={isMobile && !mobileOpen}
inert={isMobile && !mobileOpen ? true : undefined}
className={cn(
'fixed bottom-0 left-0 top-[var(--header-height)] z-50 flex max-h-[calc(100dvh-var(--header-height))] shrink-0 flex-col border-r border-border bg-card transition-transform duration-200 md:translate-x-0',
mobileOpen ? 'translate-x-0' : '-translate-x-full',
)}
style={{
width: isMobile
? 'min(18rem, calc(100vw - 3rem))'
: effectiveWidth,
}}
>
{isMobile && (
<div className="flex h-14 items-center justify-between border-b border-border px-4">
<span className="text-sm font-semibold">Navigation</span>
<button
type="button"
onClick={onMobileClose}
aria-label="Close navigation"
className="flex size-11 items-center justify-center rounded-lg text-muted-foreground hover:bg-accent hover:text-foreground"
>
<X className="size-5" aria-hidden />
</button>
</div>
)}
{!isCollapsed && (
<nav
aria-label="Primary navigation"
className="flex flex-col overflow-y-auto py-2"
>
<div className="flex flex-col">
<button
type="button"
onClick={() => setJobsOpen((v) => !v)}
aria-expanded={jobsOpen}
aria-haspopup="true"
className={cn(
TAB_BASE,
'flex w-full cursor-pointer items-center justify-between',
isJobsActive && TAB_ACTIVE,
)}
>
<span>Studio</span>
<svg
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
strokeWidth={2}
className={cn(
'h-4 w-4 shrink-0 opacity-85 transition-transform',
jobsOpen && 'rotate-180',
)}
>
<path d="M6 9l6 6 6-6" />
</svg>
</button>
{jobsOpen && (
<div className="mb-1 ml-4 flex flex-col border-l-2 border-border pl-2">
{JOB_ROUTES.map((route) => (
<Link
key={route.href}
href={route.href}
aria-current={pathname === route.href ? 'page' : undefined}
onClick={onMobileClose}
className={cn(
TAB_BASE,
'px-4 py-2 text-[0.85rem]',
pathname === route.href && TAB_ACTIVE,
)}
>
{route.label}
</Link>
))}
</div>
)}
</div>
<Link
href="/datasets"
aria-current={pathname === '/datasets' ? 'page' : undefined}
onClick={onMobileClose}
className={cn(TAB_BASE, pathname === '/datasets' && TAB_ACTIVE)}
>
Datasets
</Link>
<Link
href="/gallery"
aria-current={pathname === '/gallery' ? 'page' : undefined}
onClick={onMobileClose}
className={cn(TAB_BASE, pathname === '/gallery' && TAB_ACTIVE)}
>
Gallery
</Link>
<Link
href="/gpus"
aria-current={pathname === '/gpus' ? 'page' : undefined}
onClick={onMobileClose}
className={cn(TAB_BASE, pathname === '/gpus' && TAB_ACTIVE)}
>
GPUs
</Link>
<Link
href="/settings"
aria-current={pathname === '/settings' ? 'page' : undefined}
onClick={onMobileClose}
className={cn(TAB_BASE, pathname === '/settings' && TAB_ACTIVE)}
>
Settings
</Link>
</nav>
)}
{!isMobile && <div
className={cn(
'absolute bottom-0 p-2',
isCollapsed ? '-right-[60px] top-0' : 'right-0',
)}
>
<button
type="button"
onClick={() => setIsCollapsed((v) => !v)}
title={isCollapsed ? 'Expand sidebar' : 'Collapse sidebar'}
className={cn(
'flex size-11 items-center justify-center rounded-lg text-muted-foreground transition-colors hover:bg-accent hover:text-foreground',
)}
>
<svg
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
strokeWidth={2}
className="h-[18px] w-[18px]"
>
<path d={isCollapsed ? 'M9 18l6-6-6-6' : 'M15 18l-6-6 6-6'} />
</svg>
</button>
</div>}
{!isMobile && !isCollapsed && (
<div
role="presentation"
onMouseDown={onMouseDown}
className={cn(
'absolute bottom-0 right-0 top-0 z-[1] w-1.5 cursor-col-resize hover:bg-accent-blue/25',
isDragging && 'bg-accent-blue/25',
)}
/>
)}
</aside>
);
}
@@ -1,102 +0,0 @@
import { act, render, screen } from '@testing-library/react';
import { beforeEach, describe, expect, it, vi } from 'vitest';
import GpuGrid from './GpuGrid';
import { getGpus } from '@/lib/api';
import type { GpuSnapshot } from '@/lib/api';
vi.mock('@/lib/api', () => ({
getGpus: vi.fn(),
}));
const SNAPSHOT: GpuSnapshot = {
available: true,
error: null,
gpus: [
{
index: 0,
name: 'NVIDIA B200',
utilization: 62,
memory_used_mib: 61_440,
memory_total_mib: 183_359,
temperature_c: 41,
power_watts: 312.4,
power_limit_watts: 1000,
},
{
index: 1,
name: 'NVIDIA B200',
utilization: 0,
memory_used_mib: 1_024,
memory_total_mib: 183_359,
temperature_c: null,
power_watts: null,
power_limit_watts: null,
},
],
};
beforeEach(() => {
vi.mocked(getGpus).mockResolvedValue(SNAPSHOT);
});
describe('GpuGrid', () => {
it('renders a card per GPU with utilization and memory', async () => {
render(<GpuGrid />);
expect(await screen.findAllByText('NVIDIA B200')).toHaveLength(2);
expect(screen.getByText('GPU 0')).toBeInTheDocument();
expect(screen.getByText('GPU 1')).toBeInTheDocument();
expect(screen.getByText('62%')).toBeInTheDocument();
expect(screen.getByText('60.0 GiB / 179.1 GiB')).toBeInTheDocument();
// Optional sensors render only when present.
expect(screen.getByText('41°C')).toBeInTheDocument();
expect(screen.getByText('312 W / 1000 W')).toBeInTheDocument();
});
it('shows the backend-reported reason when telemetry is unavailable', async () => {
vi.mocked(getGpus).mockResolvedValue({
available: false,
gpus: [],
error: 'NVML Shared Library Not Found',
});
render(<GpuGrid />);
expect(
await screen.findByText(/GPU telemetry unavailable: NVML/),
).toBeInTheDocument();
});
it('explains when the API server is unreachable', async () => {
vi.mocked(getGpus).mockRejectedValue(new Error('network down'));
render(<GpuGrid />);
expect(
await screen.findByText(/Could not reach the API server/),
).toBeInTheDocument();
});
it('keeps the last snapshot visible and warns when a refresh fails', async () => {
vi.useFakeTimers();
try {
vi.mocked(getGpus)
.mockResolvedValueOnce(SNAPSHOT)
.mockRejectedValueOnce(new Error('network down'));
render(<GpuGrid />);
await act(async () => {
await vi.advanceTimersByTimeAsync(0);
});
expect(screen.getAllByText('NVIDIA B200')).toHaveLength(2);
await act(async () => {
await vi.advanceTimersByTimeAsync(3000);
});
expect(
screen.getByText(/values below may be stale/),
).toBeInTheDocument();
expect(screen.getAllByText('NVIDIA B200')).toHaveLength(2);
} finally {
vi.useRealTimers();
}
});
});
@@ -1,197 +0,0 @@
'use client';
import * as React from 'react';
import { AlertTriangle } from 'lucide-react';
import { Button } from '@/components/ui/button';
import { Card, CardContent } from '@/components/ui/card';
import { getGpus, type GpuInfo, type GpuSnapshot } from '@/lib/api';
import { cn } from '@/lib/utils';
const POLL_INTERVAL_MS = 3000;
function formatGib(mib: number): string {
return `${(mib / 1024).toFixed(1)} GiB`;
}
function Meter({
label,
percent,
detail,
warnAt,
}: {
label: string;
percent: number;
detail: string;
/** Turn the fill rose at this percentage (e.g. VRAM pressure). */
warnAt?: number;
}) {
const clamped = Math.max(0, Math.min(100, percent));
return (
<div className="flex flex-col gap-1">
<div className="flex items-baseline justify-between gap-2 text-xs">
<span className="text-muted-foreground">{label}</span>
<span className="font-medium tabular-nums text-foreground">
{detail}
</span>
</div>
<div
role="meter"
aria-label={label}
aria-valuenow={Math.round(clamped)}
aria-valuemin={0}
aria-valuemax={100}
className="h-1.5 overflow-hidden rounded-full bg-muted"
>
<div
className={cn(
'h-full rounded-full bg-accent-blue transition-[width] duration-500',
warnAt !== undefined && clamped >= warnAt && 'bg-rose-500',
)}
style={{ width: `${clamped}%` }}
/>
</div>
</div>
);
}
function GpuCard({ gpu }: { gpu: GpuInfo }) {
const memPercent =
gpu.memory_total_mib > 0
? (gpu.memory_used_mib / gpu.memory_total_mib) * 100
: 0;
return (
<Card>
<CardContent className="flex flex-col gap-4 p-5">
<div className="flex items-baseline justify-between gap-2">
<span className="min-w-0 truncate text-sm font-semibold">
{gpu.name}
</span>
<span className="shrink-0 text-xs font-medium uppercase tracking-wider text-muted-foreground">
GPU {gpu.index}
</span>
</div>
<Meter
label="Utilization"
percent={gpu.utilization}
detail={`${gpu.utilization}%`}
/>
<Meter
label="Memory"
percent={memPercent}
warnAt={90}
detail={`${formatGib(gpu.memory_used_mib)} / ${formatGib(gpu.memory_total_mib)}`}
/>
<div className="flex flex-wrap gap-x-5 gap-y-1 text-xs tabular-nums text-muted-foreground">
{gpu.temperature_c != null && <span>{gpu.temperature_c}°C</span>}
{gpu.power_watts != null && (
<span>
{Math.round(gpu.power_watts)} W
{gpu.power_limit_watts != null &&
` / ${Math.round(gpu.power_limit_watts)} W`}
</span>
)}
</div>
</CardContent>
</Card>
);
}
export default function GpuGrid() {
const [snapshot, setSnapshot] = React.useState<GpuSnapshot | null>(null);
const [fetchError, setFetchError] = React.useState<string | null>(null);
const [retryToken, setRetryToken] = React.useState(0);
React.useEffect(() => {
let mounted = true;
let inFlight = false;
async function poll() {
if (inFlight) return;
inFlight = true;
try {
const next = await getGpus();
if (mounted) {
setSnapshot(next);
setFetchError(null);
}
} catch {
if (mounted) {
setFetchError(
'GPU status could not be refreshed. The values below may be stale.',
);
}
} finally {
inFlight = false;
}
}
poll();
const interval = setInterval(poll, POLL_INTERVAL_MS);
return () => {
mounted = false;
clearInterval(interval);
};
}, [retryToken]);
if (fetchError && !snapshot) {
return (
<div
role="alert"
className="flex flex-col items-center gap-3 py-8 text-center"
>
<AlertTriangle className="size-6 text-destructive" aria-hidden />
<p className="text-muted-foreground">
Could not reach the API server. GPU status needs the Studio API server
running.
</p>
<Button
type="button"
variant="outline"
onClick={() => setRetryToken((token) => token + 1)}
>
Try Again
</Button>
</div>
);
}
if (!snapshot) {
return <p className="py-8 text-center text-muted-foreground">Loading…</p>;
}
if (!snapshot.available) {
return (
<p className="py-8 text-center text-muted-foreground">
GPU telemetry unavailable
{snapshot.error ? `: ${snapshot.error}` : '.'}
</p>
);
}
return (
<div className="flex flex-col gap-4">
{fetchError && (
<div
role="status"
aria-live="polite"
className="flex flex-wrap items-center gap-3 rounded-lg border border-amber-500/50 bg-amber-500/10 px-3 py-2 text-sm"
>
<AlertTriangle className="size-4 text-amber-600" aria-hidden />
<span className="min-w-0 flex-1">{fetchError}</span>
<Button
type="button"
variant="outline"
size="sm"
onClick={() => setRetryToken((token) => token + 1)}
>
Refresh Now
</Button>
</div>
)}
<div className="grid gap-4 [grid-template-columns:repeat(auto-fill,minmax(280px,1fr))]">
{snapshot.gpus.map((gpu) => (
<GpuCard key={gpu.index} gpu={gpu} />
))}
</div>
</div>
);
}
@@ -1,48 +0,0 @@
import { render, screen } from '@testing-library/react';
import { describe, expect, it } from 'vitest';
import { Button } from './button';
import { Input } from './input';
import { NativeSelect } from './native-select';
import { Slider } from './slider';
import { Switch } from './switch';
describe('shared control accessibility', () => {
it('keeps button, input, and select targets at least 44px tall', () => {
render(
<>
<Button size="sm">Small action</Button>
<Input aria-label="Text value" />
<NativeSelect aria-label="Choice" defaultValue="one">
<option value="one">One</option>
</NativeSelect>
</>,
);
expect(screen.getByRole('button', { name: 'Small action' })).toHaveClass(
'h-11',
);
expect(screen.getByRole('textbox', { name: 'Text value' })).toHaveClass(
'h-11',
);
expect(screen.getByRole('combobox', { name: 'Choice' })).toHaveClass(
'h-11',
);
});
it('uses 44px switch and slider interaction surfaces', () => {
render(
<>
<Switch aria-label="Enabled" />
<Slider aria-label="Amount" defaultValue={[50]} />
</>,
);
expect(screen.getByRole('switch', { name: 'Enabled' })).toHaveClass(
'size-11',
);
expect(screen.getByRole('slider', { name: 'Amount' })).toHaveClass(
'size-11',
);
});
});
@@ -1,40 +0,0 @@
'use client';
import * as React from 'react';
import { cva, type VariantProps } from 'class-variance-authority';
import { cn } from '@/lib/utils';
const badgeVariants = cva(
'inline-flex items-center rounded-full border px-2.5 py-1 text-[11px] font-semibold uppercase tracking-[0.16em] transition-colors',
{
variants: {
variant: {
default:
'border-blue-400/35 bg-blue-500/15 text-blue-900 dark:text-blue-100',
secondary: 'border-border bg-secondary text-secondary-foreground',
outline: 'border-border bg-transparent text-muted-foreground',
success:
'border-emerald-400/30 bg-emerald-500/15 text-emerald-800 dark:text-emerald-100',
warning:
'border-amber-400/30 bg-amber-500/15 text-amber-800 dark:text-amber-100',
destructive:
'border-rose-400/30 bg-rose-500/15 text-rose-800 dark:text-rose-100',
},
},
defaultVariants: {
variant: 'default',
},
},
);
export interface BadgeProps
extends React.HTMLAttributes<HTMLSpanElement>,
VariantProps<typeof badgeVariants> {}
// A span (phrasing content), so badges stay valid inside buttons and links.
function Badge({ className, variant, ...props }: BadgeProps) {
return <span className={cn(badgeVariants({ variant }), className)} {...props} />;
}
export { Badge, badgeVariants };
@@ -1,45 +0,0 @@
"use client";
import * as React from "react";
import { Slot } from "@radix-ui/react-slot";
import { cva, type VariantProps } from "class-variance-authority";
import { cn } from "@/lib/utils";
const buttonVariants = cva(
"inline-flex items-center justify-center gap-2 whitespace-nowrap rounded-xl border !text-sm !font-semibold transition-colors duration-150 disabled:pointer-events-none disabled:cursor-not-allowed disabled:opacity-50",
{
variants: {
variant: {
default: "border-slate-700 bg-slate-800 text-white shadow-md hover:bg-slate-700 dark:border-slate-300 dark:bg-slate-300 dark:text-slate-900 dark:hover:bg-slate-200",
secondary: "border-input bg-card text-card-foreground hover:bg-accent",
outline: "border-input bg-secondary text-secondary-foreground hover:bg-accent",
ghost: "border-transparent bg-transparent text-foreground hover:bg-accent",
destructive: "border-rose-500/60 bg-rose-600/90 text-white hover:bg-rose-500",
},
size: {
default: "h-11 px-4 py-2",
sm: "h-11 rounded-lg px-3 !text-xs",
lg: "h-12 px-5 !text-sm",
icon: "size-11",
"icon-sm": "size-11",
},
},
defaultVariants: {
variant: "default",
size: "default",
},
},
);
export interface ButtonProps extends React.ButtonHTMLAttributes<HTMLButtonElement>, VariantProps<typeof buttonVariants> {
asChild?: boolean;
}
const Button = React.forwardRef<HTMLButtonElement, ButtonProps>(({ className, variant, size, asChild = false, ...props }, ref) => {
const Comp = asChild ? Slot : "button";
return <Comp className={cn(buttonVariants({ variant, size, className }))} ref={ref} {...props} />;
});
Button.displayName = "Button";
export { Button, buttonVariants };
@@ -1,84 +0,0 @@
'use client';
import * as React from 'react';
import { cn } from '@/lib/utils';
const Card = React.forwardRef<HTMLDivElement, React.HTMLAttributes<HTMLDivElement>>(
({ className, ...props }, ref) => (
<div
ref={ref}
className={cn(
'rounded-2xl border border-border bg-card text-card-foreground shadow-sm',
className,
)}
{...props}
/>
),
);
Card.displayName = 'Card';
const CardHeader = React.forwardRef<
HTMLDivElement,
React.HTMLAttributes<HTMLDivElement>
>(({ className, ...props }, ref) => (
<div
ref={ref}
className={cn('flex flex-col gap-1.5 p-6', className)}
{...props}
/>
));
CardHeader.displayName = 'CardHeader';
const CardTitle = React.forwardRef<
HTMLParagraphElement,
React.HTMLAttributes<HTMLHeadingElement>
>(({ className, ...props }, ref) => (
<h3
ref={ref}
className={cn('text-lg font-semibold leading-none tracking-tight', className)}
{...props}
/>
));
CardTitle.displayName = 'CardTitle';
const CardDescription = React.forwardRef<
HTMLParagraphElement,
React.HTMLAttributes<HTMLParagraphElement>
>(({ className, ...props }, ref) => (
<p
ref={ref}
className={cn('text-sm text-muted-foreground', className)}
{...props}
/>
));
CardDescription.displayName = 'CardDescription';
const CardContent = React.forwardRef<
HTMLDivElement,
React.HTMLAttributes<HTMLDivElement>
>(({ className, ...props }, ref) => (
<div ref={ref} className={cn('p-6 pt-0', className)} {...props} />
));
CardContent.displayName = 'CardContent';
const CardFooter = React.forwardRef<
HTMLDivElement,
React.HTMLAttributes<HTMLDivElement>
>(({ className, ...props }, ref) => (
<div
ref={ref}
className={cn('flex items-center p-6 pt-0', className)}
{...props}
/>
));
CardFooter.displayName = 'CardFooter';
export {
Card,
CardHeader,
CardFooter,
CardTitle,
CardDescription,
CardContent,
};
@@ -1,117 +0,0 @@
'use client';
import * as React from 'react';
import * as DialogPrimitive from '@radix-ui/react-dialog';
import { X } from 'lucide-react';
import { cn } from '@/lib/utils';
const Dialog = DialogPrimitive.Root;
const DialogTrigger = DialogPrimitive.Trigger;
const DialogPortal = DialogPrimitive.Portal;
const DialogClose = DialogPrimitive.Close;
const DialogOverlay = React.forwardRef<
React.ElementRef<typeof DialogPrimitive.Overlay>,
React.ComponentPropsWithoutRef<typeof DialogPrimitive.Overlay>
>(({ className, ...props }, ref) => (
<DialogPrimitive.Overlay
ref={ref}
className={cn(
'fixed inset-0 z-50 bg-black/70 backdrop-blur-sm data-[state=open]:animate-in data-[state=closed]:animate-out data-[state=closed]:fade-out-0 data-[state=open]:fade-in-0',
className,
)}
{...props}
/>
));
DialogOverlay.displayName = DialogPrimitive.Overlay.displayName;
const DialogContent = React.forwardRef<
React.ElementRef<typeof DialogPrimitive.Content>,
React.ComponentPropsWithoutRef<typeof DialogPrimitive.Content>
>(({ className, children, ...props }, ref) => (
<DialogPortal>
<DialogOverlay />
<DialogPrimitive.Content
ref={ref}
data-modal=""
className={cn(
'fixed left-1/2 top-1/2 z-50 grid w-full max-w-lg -translate-x-1/2 -translate-y-1/2 gap-4 rounded-2xl border border-border bg-card p-6 text-card-foreground shadow-lg duration-200 data-[state=open]:animate-in data-[state=closed]:animate-out data-[state=closed]:fade-out-0 data-[state=open]:fade-in-0',
className,
)}
{...props}
>
{children}
<DialogPrimitive.Close className="absolute right-2 top-2 flex size-11 items-center justify-center rounded-lg text-muted-foreground opacity-70 transition-opacity hover:bg-secondary hover:opacity-100 disabled:pointer-events-none sm:right-4 sm:top-4">
<X className="h-4 w-4" />
<span className="sr-only">Close</span>
</DialogPrimitive.Close>
</DialogPrimitive.Content>
</DialogPortal>
));
DialogContent.displayName = DialogPrimitive.Content.displayName;
const DialogHeader = ({
className,
...props
}: React.HTMLAttributes<HTMLDivElement>) => (
<div
className={cn('flex flex-col gap-1.5 text-left', className)}
{...props}
/>
);
DialogHeader.displayName = 'DialogHeader';
const DialogFooter = ({
className,
...props
}: React.HTMLAttributes<HTMLDivElement>) => (
<div
className={cn(
'flex flex-col-reverse gap-2 sm:flex-row sm:justify-end',
className,
)}
{...props}
/>
);
DialogFooter.displayName = 'DialogFooter';
const DialogTitle = React.forwardRef<
React.ElementRef<typeof DialogPrimitive.Title>,
React.ComponentPropsWithoutRef<typeof DialogPrimitive.Title>
>(({ className, ...props }, ref) => (
<DialogPrimitive.Title
ref={ref}
className={cn(
'text-lg font-semibold leading-none tracking-tight',
className,
)}
{...props}
/>
));
DialogTitle.displayName = DialogPrimitive.Title.displayName;
const DialogDescription = React.forwardRef<
React.ElementRef<typeof DialogPrimitive.Description>,
React.ComponentPropsWithoutRef<typeof DialogPrimitive.Description>
>(({ className, ...props }, ref) => (
<DialogPrimitive.Description
ref={ref}
className={cn('text-sm text-muted-foreground', className)}
{...props}
/>
));
DialogDescription.displayName = DialogPrimitive.Description.displayName;
export {
Dialog,
DialogPortal,
DialogOverlay,
DialogTrigger,
DialogClose,
DialogContent,
DialogHeader,
DialogFooter,
DialogTitle,
DialogDescription,
};
@@ -1,22 +0,0 @@
"use client";
import * as React from "react";
import { cn } from "@/lib/utils";
const Input = React.forwardRef<HTMLInputElement, React.ComponentProps<"input">>(({ className, type, ...props }, ref) => {
return (
<input
type={type}
className={cn(
"flex h-11 w-full rounded-xl border border-input bg-card/60 px-3 py-2 text-sm text-foreground shadow-sm backdrop-blur-md transition-colors placeholder:text-muted-foreground focus-visible:border-ring disabled:cursor-not-allowed disabled:opacity-50",
className,
)}
ref={ref}
{...props}
/>
);
});
Input.displayName = "Input";
export { Input };
@@ -1,26 +0,0 @@
'use client';
import * as React from 'react';
import * as LabelPrimitive from '@radix-ui/react-label';
import { cva, type VariantProps } from 'class-variance-authority';
import { cn } from '@/lib/utils';
const labelVariants = cva(
'text-sm font-medium leading-none text-foreground peer-disabled:cursor-not-allowed peer-disabled:opacity-70',
);
const Label = React.forwardRef<
React.ElementRef<typeof LabelPrimitive.Root>,
React.ComponentPropsWithoutRef<typeof LabelPrimitive.Root> &
VariantProps<typeof labelVariants>
>(({ className, ...props }, ref) => (
<LabelPrimitive.Root
ref={ref}
className={cn(labelVariants(), className)}
{...props}
/>
));
Label.displayName = LabelPrimitive.Root.displayName;
export { Label };
@@ -1,24 +0,0 @@
'use client';
import * as React from 'react';
import { cn } from '@/lib/utils';
const NativeSelect = React.forwardRef<
HTMLSelectElement,
React.ComponentProps<'select'>
>(({ className, children, ...props }, ref) => (
<select
ref={ref}
className={cn(
'flex h-11 w-full appearance-none rounded-xl border border-input bg-card px-3 py-2 text-sm text-foreground shadow-sm transition-colors focus-visible:border-ring disabled:cursor-not-allowed disabled:opacity-50',
className,
)}
{...props}
>
{children}
</select>
));
NativeSelect.displayName = 'NativeSelect';
export { NativeSelect };
@@ -1,45 +0,0 @@
'use client';
import * as React from 'react';
import * as ScrollAreaPrimitive from '@radix-ui/react-scroll-area';
import { cn } from '@/lib/utils';
const ScrollArea = React.forwardRef<
React.ElementRef<typeof ScrollAreaPrimitive.Root>,
React.ComponentPropsWithoutRef<typeof ScrollAreaPrimitive.Root>
>(({ className, children, ...props }, ref) => (
<ScrollAreaPrimitive.Root
ref={ref}
className={cn('relative overflow-hidden', className)}
{...props}
>
<ScrollAreaPrimitive.Viewport className="h-full w-full rounded-[inherit]">
{children}
</ScrollAreaPrimitive.Viewport>
<ScrollBar />
<ScrollAreaPrimitive.Corner />
</ScrollAreaPrimitive.Root>
));
ScrollArea.displayName = ScrollAreaPrimitive.Root.displayName;
const ScrollBar = React.forwardRef<
React.ElementRef<typeof ScrollAreaPrimitive.ScrollAreaScrollbar>,
React.ComponentPropsWithoutRef<typeof ScrollAreaPrimitive.ScrollAreaScrollbar>
>(({ className, orientation = 'vertical', ...props }, ref) => (
<ScrollAreaPrimitive.ScrollAreaScrollbar
ref={ref}
orientation={orientation}
className={cn(
'flex touch-none select-none p-0.5 transition-colors',
orientation === 'vertical' ? 'h-full w-2.5' : 'h-2.5 flex-col',
className,
)}
{...props}
>
<ScrollAreaPrimitive.ScrollAreaThumb className="relative flex-1 rounded-full bg-border" />
</ScrollAreaPrimitive.ScrollAreaScrollbar>
));
ScrollBar.displayName = ScrollAreaPrimitive.ScrollAreaScrollbar.displayName;
export { ScrollArea, ScrollBar };

Some files were not shown because too many files have changed in this diff Show More