Compare commits

...
Author SHA1 Message Date
SolitaryThinker 4a33540c08 [perf] v2: on-device denoise loop — kill the per-step numpy<->torch round-trip (Wan2.1)
The torch adapter boundary marshalled the latent host<->device on EVERY denoise
step: _t uploaded the latent (and re-uploaded the text embeds) and _n downloaded
the velocity with a forced CUDA sync — 2*N PCIe copies + N syncs per generation,
buying nothing, since the latent could stay resident on the GPU the whole loop.

Root cause was a numpy loop surface. But the loop MATH is already array-agnostic
(CFG combine + flow-match Euler are pure arithmetic; the solver kernel already
passes torch through). So introduce a per-platform array namespace (v2/platform/
array_ns): numpy on CPU (torch-free — the parity mini is unchanged), torch-on-
device on cuda. The latent is seeded with numpy and uploaded ONCE; it then stays
resident through forward -> CFG combine -> solver -> next step; a single host
marshal happens at the request/output boundary (engine._to_artifact).

Opt-in per recipe via ModelCard.device_io (set on the Wan cards). When set on a
GPU box, build_component flips the components' TorchComponent.device_io so _out
keeps tensors on-device (in fp32, matching the old _to_numpy cast so the combine
dtype is unchanged). Un-migrated families and the CPU toy keep numpy in/out.
Also: PrecisionPolicy.cast is array-preserving; _t accepts resident tensors.

Verified BIT-IDENTICAL on real Wan2.1-1.3B / H100: the on-device latent equals
the pre-change numpy-path latent exactly (max_abs_diff 0.0, np.array_equal True).
CPU mini holds 237 passed / 5 pre-existing; pre-commit clean.

Other WanDenoiseLoop families can flip device_io next (per-family GPU re-verify);
non-Wan loops migrate to the xp namespace later.
2026-06-20 00:08:16 +00:00
SolitaryThinker ba83e9bef7 [refactor] v2: co-locate per-model torch adapters into their recipe packages
platform/backends/ had become a flat dump of 15 per-model torch_<model>.py
adapters next to the genuinely-shared infra. Each adapter is referenced from
its card by a plain 'module:Class' string loaded via importlib, so there was
no real coupling forcing it into platform/ — the Cosmos/Flux/etc adapter
belongs WITH its recipe (card/loop/program).

Move each torch_<model>.py -> v2/recipes/<model>/adapter.py and flip the card
strings to v2.recipes.<model>.adapter:<Class>. backends/ now holds only the
shared substrate (torch_backend base, torch_cuda registration, torch_kernels,
toy/cpu/accel). Each recipe is now a self-contained package.

Cross-refs updated: gen3c/adapter imports CosmosT5Encoder from cosmos2/adapter;
sd35/program + stable_audio/card import from their own package. Recipes still
import torch-free (adapters pulled only via the string on a GPU box) — CPU mini
holds 237 passed / 5 pre-existing (bucket_c torch-absent). pre-commit clean.
2026-06-19 23:48:59 +00:00
SolitaryThinker 21e59268d8 [misc] v2: full pre-commit clean (ruff UP038/SIM/UP031 + mypy annotations)
Sweep all of v2/ through pre-commit (was previously only run on changed
files). Fixes surfaced across untouched modules:

- ruff: isinstance-tuple -> X | Y (UP038), try/except/pass ->
  contextlib.suppress (SIM105), negated-return (SIM103), %-format ->
  f-string (UP031).
- mypy: add annotations for no-untyped-call + var-annotated across recipes,
  training methods, torch/toy backends, and serving.
- yapf reflow of the SF-Wan KV-cache call sites (semantics unchanged).

yapf/ruff/codespell/mypy all pass; v2 tests 237 passed / 5 pre-existing
(bucket_c torch-absent on CPU venv).
2026-06-19 22:16:28 +00:00
SolitaryThinker 139577639f [bugfix] v2: SF-Wan cross-chunk KV cache — condition each chunk on prior clean chunks
The causal adapter never passed a kv_cache, so CausalWanTransformer3DModel.forward routed to
_forward_train (no cross-chunk KV) on every chunk instead of _forward_inference (the CausVid
Alg-2 KV-cache path). Each chunk denoised blind to the previous ones; the loop's cross-chunk
"context" was a toy mean(prior_latents) the adapter ignored. Result: hard discontinuities at
every chunk boundary (frame-to-frame absdiff spikes 43-55 every ~12 frames).

Fix (cuda path only; toy/CPU path and the 237-test suite untouched):
- WanDiT.alloc_causal_caches(): allocate the persistent per-block KV + cross-attn caches sized
  from the model config (mirrors CausalDenoisingStage._initialize_kv_cache).
- WanDiT.__call__: thread kv_cache/crossattn_cache/current_start/cache_start/start_frame/
  frame_seqlen so the model runs _forward_inference.
- wan_causal/loop.py: own the caches in LoopState (per-request -> interleave-safe); pass
  current_start = chunk_idx*chunk_size*frame_seqlen per chunk; do the clean-KV write
  (timestep ~0) after each chunk so the next attends to it.

Verified on H100: frame-to-frame absdiff mean 12.8->4.4, max 55.4->8.3; chunk-boundary spikes
eliminated; coherent across all 7 chunks. CPU causal toy tests unchanged (24 passed).
2026-06-19 22:01:24 +00:00
SolitaryThinker db866dfb04 [misc] v2: simplify docstrings/comments + drop deleted-design-doc citations
Sweep all 296 v2 modules: simplify verbose docstrings/comments and remove 376 dangling
"(design_vN §X)" citations to the now-deleted design docs (v2/README.md is the source of
truth). Comment/docstring-only — AST-verified code-identical; the CPU suite holds at 237
passed / 5 pre-existing. Also applies yapf + ruff --fix auto-fixes (import ordering,
forward-ref annotation de-quoting under `from __future__ import annotations`; behavior-
neutral, suite-confirmed) and adds the legitimate domain terms mot/clen/te to the codespell
ignore-list. Remaining ruff (24) + mypy (68 no-untyped-call) findings are pre-existing v2
debt, untouched here.
2026-06-19 20:54:59 +00:00
SolitaryThinker 53d63afa76 [docs] v2: make v2/README.md the design source of truth + one-page philosophy + M* roadmap
Unify the four design docs (design.md, designv2.md, design_v3.md, designv4.md) into a single
authoritative v2/README.md: the (recipe, runtime) thesis, driven loops, planes, one-WorkUnit
scheduler, the parity ladder + interleave gate, training-on-shared-loops, the weight-sharing
topology catalog, the current GPU status (20+ models + the BAGEL/Qwen-Omni/Cosmos3 trio
verified), and a prioritized roadmap. Recast design_summary.md as a one-page design philosophy
pointing to it. Add .agents/exploration/mstar-v2-roadmap.md (the adversarially-verified M*
Walk-Graph gap analysis driving the roadmap). Delete the four superseded design docs.
2026-06-19 20:54:44 +00:00
SolitaryThinker 685702e330 [bugfix] v2: MatrixGame2/3 causal-loop progress counter + MG3 patch alignment
MatrixGame2 (causal DMD loop): bump st.step_idx on every executed work unit (each
DMD step and each clean-context pass). The loop drives its own control flow off
block_idx/dmd_idx/phase, but the runtime's no-progress watchdog keys on
st.step_idx, so a multi-block causal rollout was seen as stalled. Mirrors what
every other recipe loop does.

MatrixGame3 (5B WanModel): patch-align the latent H/W (patch_size (1,2,2)) before
denoise. The model folds (H/2, W/2) tokens, so an odd latent dim made the
unpatchified velocity come back one row/col short of the noise latent. Crop to
(latent // patch) * patch, faithful to MatrixGame3DenoisingStage.

v2 mini: 240 passed.
2026-06-19 11:58:17 +00:00
SolitaryThinker 9e8d515842 [bugfix] v2: correct SF-Wan + LTX2 2-stage SR sampling defaults (GPU frame-verified)
Two distilled few-step video models rendered incorrectly on GPU; root-caused via
dense frame sampling (contact sheets) and fixed in the recipe cards.

SF-Wan2.1 (self-forcing causal, wan_causal card): was oversaturated/overcooked.
The distilled student is CFG-FREE (guidance 1.0, single forward/step) and denoises
with the 4-step DMD schedule [1000,750,500,250] (warped by FlowShiftPolicy(5.0)),
at a native causal block of 3 latent frames. Defaults were ClassicCFG@6.0 + 2 steps
+ block 2 -> overcooked AND under-denoised. Fixed: num_chunks=7, chunk_size=3,
steps_per_chunk=4; SamplingDefaults num_steps=4, guidance_scale=1.0 (7x3=21 latent
-> 81 frames). Renders a clean raccoon-in-sunflowers across all 81 frames.

LTX2-Distilled 2-stage SR (ltx2 card + LTX2VAE): was temporally blocky. Root cause
was an OOM-forced 57-frame reduction (only 8 latent temporal frames); the model is
designed for 121 (16 latent frames). Enable VAE tiling in LTX2VAE so the 121-frame
full-res decode fits the 80 GiB GPU; keep base cfg_scale=3.0 (drives brightness;
cfg=1 washed out) with stg_scale=0.0 (v2's drop-text perturbation is not real
skip-layer STG). Renders the on-prompt backyard shot, bright + temporally coherent.

torch_backend.py: enable LTX2VAE tiling; clear pre-existing mypy no-untyped-call /
yapf debt on the file (surfaced once per-file linting bypassed the duplicate-module
flakiness) by annotating the helper/constructor/maker signatures.

Tests: update the 3 affected CPU defaults/chunk-count tests. v2 mini: 240 passed.
2026-06-19 11:57:48 +00:00
SolitaryThinker 1302a27470 [docs] v2: GPU bring-up results — 20 models verified on H100
V2_PORTING_STATUS.md now records the GPU bring-up outcome: 20 models generate
real video/audio on H100 (the 7 prior + 13 newly-ported), with the remaining
split into fastvideo/env-blocked (SLA/VSA kernels, transformers incompat, fastvideo
registry/flash_attn gaps) and HF-access-blocked (gated cosmos2/flux2/sd35) — none
a v2 recipe bug.
2026-06-18 19:39:53 +00:00
SolitaryThinker e1a2c9ede8 [feat] v2 GPU bring-up: 5 more models verified on H100 (huge MoE/world + DMD)
Second GPU pass (distilled + huge dense, 2-wide). 5 verified end-to-end with real
weights; CPU toy path kept green (240 passed, 2 skipped).

VERIFIED:
  * matrixgame3   — mp4 (9,256,256,3), zero fixes. 6.47B, standard Wan attn (NOT
                    sparse-attn-blocked, like its mg2 sibling); degenerate single-clip.
  * fastwan       — FastWan2.2-TI2V-5B-FullAttn DMD 3-step, mp4 (17,256,256,3), zero
                    fixes. The FULL-ATTENTION variant has no VSA params -> the generic
                    Wan loader maps it cleanly (reuses WanDiT via load_id, no adapter).
  * longcat       — LongCat-Video-T2V 13.58B, mp4 (17,256,256,3), zero fixes, CPU
                    offload (~40GB peak).
  * sfwan22       — Self-Forcing Wan2.2-A14B causal+MoE (2x14B), mp4 (29,288,288,3),
                    CPU expert offload (~80GB peak).
  * lingbotworld  — Wan2.2-class 2x14B camera world model, mp4 (9,256,256,3), offload
                    (~98GB peak transient).

BLOCKED: turbowan-i2v-a14b — SLA sparse-attn params (attn1.attn_impl.proj_l on all
40 layers of both experts) cannot load into the dense Wan build + needs the
fastvideo-kernel SLA Triton kernels (no nvcc here). Confirmed via the safetensors
header (no 60GB download). Same SLA family as turbowan-1.3b.

Fixes (own-port only): sfwan22/loop.py, lingbotworld/{card,program}.py +
torch_lingbotworld.py. pre-commit clean per-package.
2026-06-18 19:38:50 +00:00
SolitaryThinker 8ee7270368 [bugfix] v2 VideoGenerator: modality-aware result path (audio/image, not only video)
VideoGenerator._result hardcoded out.artifacts['video'].frames, so an audio-only
(Stable Audio) or image-only (SD3.5 / FLUX.2 T2I) generation crashed with
KeyError 'video' even though the engine had correctly produced the AudioArtifact /
image TensorArtifact (surfaced during stable_audio GPU bring-up). _result now
guards on the artifact present: video -> mp4 (unchanged), else image -> png
([C,H,W]/[B,..] normalized to [H,W,C]), plus the existing audio -> sibling .wav;
image_path recorded in result.extra. The video path is byte-for-byte unchanged.

v2 mini 240 passed, 2 skipped; pre-commit clean.
2026-06-18 18:52:24 +00:00
SolitaryThinker 26c46b901a [feat] v2 GPU bring-up: 8 ported models verified end-to-end on H100 (+CPU-safe fixes)
Ran each dense public port through the real VideoGenerator on H100 (2-wide across
both GPUs). 8 produce real finite output end-to-end; per-model adapter/loop fixes
landed in each port's OWN files (no shared/fastvideo edits). CPU toy path kept
green (240 passed, 2 skipped) — GPU-only conditioning gated to the cuda backend.

VERIFIED (real GPU output):
  * stable_audio    — stereo audio (2, 441000) @44.1kHz. Fixes: dedicated
                      'conditioner' component kind (SA owns its T5, empty
                      text_encoder_configs) + ConditionerLoader from conditioner/
                      + VDenoiser c_noise = atan(sigma)/(pi/2).
  * matrixgame2     — mp4 (9,256,256,3). Loads CLEAN (not sparse-attn-blocked).
                      Fixes: 20ch cond_concat (4ch mask + 16ch img), mandatory i2v
                      (synth blank first frame), pre-sized kv_cache/crossattn_cache
                      (SDPA inference path, avoids flex_attention compile), bf16
                      autocast, per-request reset_caches.
  * gen3c           — video (9,256,256,3), zero fixes (worked first try).
  * wan_fun_control — video (9,256,256,3).
  * lucy_edit       — video (17,256,256,3).
  * hunyuangamecraft— video (9,256,256,3).
  * hunyuan_video   — video (3,9,256,256), dual LLaMA+CLIP + Hunyuan VAE.
  * hunyuan_video15 — video (9,256,256,3), dual Qwen+ByT5 (gated to cuda; CPU passes
                      single embed).

BLOCKED (not v2 recipe bugs — load/run reached, then a fastvideo/env wall):
  * cosmos25 — DiT + VAE ran finite on GPU; the Qwen2.5-VL Reason1 encoder hits a
               transformers 5.12.1 incompat in fastvideo shared code
               (Qwen2_5_VLConfig.pad_token_id).
  * kandinsky5 — fastvideo registry.py registers it with a bare PipelineConfig (no
                 Kandinsky5 config) -> load fails fastvideo-side. (latent z=16 +
                 visual_cond adapter corrected; toy decoupled to its own channels.)
  * hyworld — fastvideo's hyworld DiT hardcodes flash_attn (no SDPA fallback);
              flash_attn kernel not built here.

CPU-safety fixes (mine): kandinsky5 toy ToyDiT/ToyVAE use the toy LATENT_CHANNELS
(not the real z=16); hunyuan_video15 dual-encoder packing gated to cuda.
Sparse-attn distilled (turbowan/SLA, fastwan/VSA) + gated (cosmos2/flux2/sd35) +
huge (>80GB) handled separately. pre-commit clean per-package.
2026-06-18 18:46:32 +00:00
SolitaryThinker fcc5a26b46 [docs] v2: porting status — ALL fastvideo models ported (63/64 by-id; VSA env-blocked)
Rewrites V2_PORTING_STATUS.md to reflect completion: the scope is now ALL
fastvideo models (not Wan+LTX-2 only). Documents the self-contained recipe-package
porting mechanism (ComponentSpec.adapter), the 15 net-new architectures + 5
Wan-family variants newly ported (CPU-verified end-to-end; GPU=BRINGUP), the 7
GPU-verified models, and the single env-blocked id (VSA-14B, needs nvcc).
2026-06-18 16:24:43 +00:00
SolitaryThinker 6450f9fa26 [feat] v2 registry: LTX-2/2.3 repo aliases -> by-id resolution 63/64
Adds explicit ModelEntry aliases for the LTX-2 (FastVideo/LTX2-Diffusers,
LTX2-base, Lightricks/LTX-2 -> single-stage base) and LTX-2.3 (LTX2.3-Diffusers,
LTX2.3-Distilled-Diffusers, LTX2.3-base, Lightricks/LTX-2.3, lightricks/ltx-2.3
-> the distilled joint-A/V card) naming variants of already-ported LTX
checkpoints (the arch fallback also resolves LTX2Transformer3DModel from a root).

v2 now resolves 63/64 fastvideo registry ids by exact id; the only remaining id,
FastVideo/Wan2.1-VSA-T2V-14B-720P-Diffusers, is ENV-BLOCKED (VSA Sparse-Linear
Attention kernels require nvcc, not built in this bring-up; it arch-resolves to
the base Wan card but needs the VSA kernel build to run faithfully).

v2 mini 240 passed, 2 skipped; pre-commit clean.
2026-06-18 16:23:30 +00:00
SolitaryThinker 4d079c65ed [feat] v2: port the residual Wan-family variants (rCM/DMD/v2v/control/causal-MoE)
Closes the bucket-B sampler/conditioning gap — each reuses the Wan/Causal ARCH
(no new torch adapter) with a new in-package loop/sampler/conditioning, declared
in _BUCKET_C as explicit-HF-id-only (transformer_cls="" so the generic Wan/Causal
arch fallback is NOT hijacked — only the exact id distinguishes the capability
variant from a base Wan of the same class).

  * turbowan      — TurboWan rCM (Reparameterized Consistency Model) few-step: a
                    faithful in-package RCMScheduler port (TrigFlow->RectifiedFlow
                    schedule + stochastic consistency SDE step), 1.3B/14B T2V +
                    TurboWan2.2-I2V-A14B (MoE i2v, boundary 0.9 in raw-sigma space)
  * lucy_edit     — Lucy-Edit v2v editor: a video_vae_encode node (the input video
                    -> 48ch cond latent) threaded via the shared i2v_cond hook ->
                    96ch Lucy DiT input (faithful to denoising.py is_lucy_edit)
  * wan_fun_control — Wan2.1-Fun-Control: control-video conditioning (reuses the
                    i2v [mask|cond] concat pattern)
  * sfwan22       — Self-Forcing Wan2.2-A14B: causal chunk_rollout + Wan2.2 MoE
                    boundary routing, i2v (boundary 0.9) + t2v (boundary 0.875)
  * fastwan       — FastWan DMD 3-step: TI2V-5B-FullAttn loadable; the VSA-trained
                    variants + non-strict to_gate_compress load are BRINGUP

All 12 residual ids resolve+build; base Wan/Causal resolution unchanged (arch
fallback not hijacked). The _BUCKET_C regression test auto-extended -> v2 mini
240 passed, 2 skipped; pre-commit clean (per-package + registry).
2026-06-18 16:21:19 +00:00
SolitaryThinker 55fcf27759 [test] v2: end-to-end CPU regression guard for the bucket-C ports
Data-driven from registry._BUCKET_C (+ cosmos2): each net-new ported arch
resolves through the registry (exact id + arch fallback) AND runs end-to-end on
the CPU toy backend via the public Engine path (resolve -> build card+program ->
load_card -> Engine.run), emitting exactly one modality-correct artifact
(video / image / audio) + latents. Auto-covers future _BUCKET_C rows.

21 tests pass; full v2 mini 232 passed, 2 skipped.
2026-06-18 16:14:15 +00:00
SolitaryThinker 3149d0ec89 [feat] v2: port the 14 remaining bucket-C archs as self-contained recipe packages
Completes the bucket-C porting backlog. Each arch is a self-contained recipe
package (card-declared torch adapter via ComponentSpec.adapter + a new/forked
loop + program, NO edit to the shared torch_backend dispatch), following the
cosmos2 reference pattern. One _BUCKET_C table in v2/registry.py drives both the
explicit HF-id registry (PRIMARY) and the select_by_architecture fallback.

Ported (CPU-verified: import + card/program build + registry resolve + denoise
loop runs end-to-end on the CPU toy backend; GPU load/run is BRINGUP):
  * cosmos25       — Cosmos-Predict2.5 (flow-match, per-frame plain-sigma timestep;
                     reuse FLOW_MATCH_STEP; Reason1/Qwen2.5-VL encoder adapter)
  * hunyuan_video  — HunyuanVideo (reuses WanDenoiseLoop; dual LLaMA+CLIP encoders;
                     Hunyuan VAE scaling_factor) + FastHunyuan variant
  * hunyuan_video15 — HunyuanVideo 1.5 (480p/720p cards)
  * longcat        — LongCat-Video T2V/I2V/VC
  * sd35           — SD3.5 MMDiT (flow-match, image; triple-encoder joint embed +
                     pooled_projections)
  * gen3c          — GEN3C (EDM; 82ch pose-buffer DiT; camera/depth -> BRINGUP)
  * kandinsky5     — Kandinsky 5.0 T2V Lite
  * flux2          — FLUX.2 dev/klein (MMDiT, image; gated weights -> BRINGUP)
  * stable_audio   — Stable Audio Open (audio modality)
  * hunyuangamecraft, hyworld, lingbotworld, matrixgame2, matrixgame3 — interactive
                     world models; t2v/degenerate path CPU-verified, action/camera/
                     memory conditioning is BRINGUP (needs request-API extension)

Adapters declared via ComponentSpec.adapter (the ac29750b enabler) so each port
adds only NEW files (recipe package + per-arch torch_<arch>.py + optional facade
stub) — zero shared-file edits. Registry resolves all 31 bucket-C HF ids by exact
id + 14 architecture fallbacks; no regression (cosmos2/wan/ltx2 unchanged).
v2 mini green (211 passed, 2 skipped); pre-commit clean (per-file/registry).
2026-06-18 16:04:36 +00:00
SolitaryThinker 1f4b729c4a [feat] v2: Cosmos-Predict2-2B-Video2World port (EDM-Karras) — reference bucket-C recipe
First net-new architecture ported via the self-contained recipe-package pattern
(card-declared adapter + new loop, no shared-dispatch edit):

* CosmosDenoiseLoop (v2/recipes/cosmos2/loop.py): EDM preconditioning folded into
  a flow-match Euler integrator. Faithful port of CosmosDenoisingStage — Karras
  sigma schedule (rho=7, sigma_max=80 -> sigma_min=0.002, terminal clamp), latent
  init randn*sigma_max, per-step c_in/c_skip/c_out (sigma_data=1) -> x0, CFG in x0
  space, x0 -> velocity (x-x0)/sigma, FLOW_MATCH_STEP. video2world frame-replace
  conditioning threaded but inert for the t2v preset.
* build_karras_sigmas helper added to v2/loop/sampler.py.
* CosmosDiT + CosmosT5Encoder adapters (v2/platform/backends/torch_cosmos.py),
  declared on the card via ComponentSpec.adapter (the ac29750b enabler) — DiT
  returns raw EDM output + builds the mandatory zero condition/padding masks + fps;
  T5 uses the raw last_hidden_state (no Wan zero-pad). Reuses the WanVAE adapter.
* card/program/registry (HF id nvidia/Cosmos-Predict2-2B-Video2World + arch
  fallback on CosmosTransformer3DModel) + COSMOS_NEG prompt + SamplingDefaults
  (35 steps, gs 7, 704x1280, 93f, 16fps).

CPU-verified: Karras schedule, card/program build, registry resolve (id + arch),
EDM loop runs end-to-end on the CPU toy backend. GPU load/run is BRINGUP.
v2 mini green (211 passed, 2 skipped); pre-commit clean.
2026-06-18 15:38:46 +00:00
SolitaryThinker ac29750b55 [feat] v2 torch backend: ComponentSpec.adapter — card-declared per-arch TorchComponent
A new architecture can declare its own torch adapter on the card
(ComponentSpec.adapter="module:Class") instead of editing the shared _make_dit/
_make_vae/_make_text_encoder dispatch. _explicit_adapter() constructs it as
cls(module, *extra, device=, dtype=) and short-circuits the built-in Wan/LTX2
class-name dispatch when set. This makes each bucket-C port a self-contained
recipe package (card + adapter module + loop + program) with no shared-file edit
-> conflict-free parallel porting. Unset -> unchanged built-in dispatch.

CPU mini green (211 passed, 2 skipped).
2026-06-18 15:27:07 +00:00
Will Lin 50fd379f3a [feat] v2: Wan2.2-I2V-A14B (MoE i2v) — combine boundary-routed experts + i2v conditioning
Reuses everything: 2 WanTransformer3DModel experts + BoundaryTimestepRouting (from the A14B MoE pattern),
the CLIP image encoder + first-frame [mask|cond] conditioning + the i2v program (from the Fun-InP i2v
port), and the shared WanDenoiseLoop (i2v hooks + the boundary expert). No new adapter. CPU-verified: the
toy MoE i2v runs end-to-end (2 experts + boundary + conditioning -> finite video); resolves with i2v caps.
Structural (GPU-pending: 2x14B, like the A14B T2V). Wan family now largely covered (T2V 1.3B/14B/TI2V-5B/
A14B, causal SF, i2v 1.3B/14B/A14B). CPU mini 211/2.
2026-06-18 11:56:52 +00:00
Will Lin fa9d8040e1 [feat] v2: register Wan2.1-I2V-14B 480P/720P (reuse the GPU-verified i2v card)
The 14B i2v variants reuse the Wan2.1 i2v card/path proven on Fun-1.3B-InP — just per-variant params
(480P flow_shift 3.0 / 480x832, 720P flow_shift 5.0 / 720x1280). Registry resolution + caps verified;
specific 14B weights GPU-pending (same generic Wan i2v loader path that Fun-InP validated). i2v cluster
now supported; roadmap updated (11 models ported).
2026-06-18 11:53:25 +00:00
Will Lin 9794be7202 [feat] v2: Wan2.1 i2v port (Wan2.1-Fun-1.3B-InP) — CLIP encoder + first-frame conditioning, GPU-verified
Real image-to-video, unlocking the i2v cluster. v2/recipes/wan21/i2v.py: CLIP image-encode -> the DiT's
encoder_hidden_states_image; first-frame VAE conditioning + a 4-channel mask -> the 20ch [mask|cond] that
the Wan adapter concatenates with the 16ch noise -> the 36ch i2v DiT input (mirrors fastvideo's
ImageEncodingStage + ImageVAEEncodingStage; v2's WanVAE.encode already applies the matching (z-mean)/std).
Reuses the shared WanDenoiseLoop (its None-default i2v hooks) + the Wan torch adapter unchanged. Adds
ToyImageEncoder + the image_encoder checkpoint subfolder stamp. Registered Wan2.1-Fun-1.3B-InP.

GPU-verified: loads via the generic Wan loader (1.56B, no param-mapping issue), runs the full i2v
conditioning, produces real video (3,9,256,384, std 0.44, finite, motion 0.041). CPU mini 211/2 (T2V
unregressed). BRINGUP: visual confirmation that the output follows the conditioning image is
human-in-the-loop.
2026-06-18 11:52:00 +00:00
Will Lin e4b14e124d [feat] v2 Wan loop+adapter: optional i2v conditioning hooks (cond concat + CLIP context); T2V unchanged
Threads i2v conditioning through the SHARED WanDenoiseLoop with zero T2V risk: init() reads optional
slots i2v_cond (the [mask|cond] latent) + i2v_img_embeds (CLIP) into scratch; _velocity passes them to
the dit (context=, cond=); WanDiT concats cond (16->36ch) and uses the embeds as
encoder_hidden_states_image; capture is disabled only when i2v conditioning is present. For T2V both are
None -> the dit call, CFG, and cudagraph capture are byte-identical (CPU mini 211 pass, 2 skip — no
regression). ToyDiT accepts+ignores cond (image-conditioning is a GPU-path concern). Completes the i2v
backend seam; the program's mask+cond construction + the Wan i2v card/registry + GPU verify follow.
2026-06-18 11:42:35 +00:00
Will Lin 60ca94796c [feat] v2 torch backend: CLIP image-encoder adapter + image_encoder component kind (i2v groundwork)
Adds CLIPImageEncoder (encode_image -> the DiT's encoder_hidden_states_image) + the generic builder's
image_encoder maker (ImageEncoderLoader + ImageProcessorLoader), registered as the cuda 'image_encoder'
kind. Mirrors fastvideo's ImageEncodingStage. CPU-verified (component-kinds + lazy invariant, 211/2);
GPU path marked BRINGUP/written-not-run (processor subfolder + dtype to confirm on a real i2v checkpoint),
matching how the rest of the torch backend was originally landed. Reusable by the Wan i2v cluster + many
bucket-C models (Hunyuan/Cosmos i2v). Next i2v increments: the mask+cond latent construction + the
concat-into-DiT-input loop, then the card + registry + GPU-verify with Wan2.1-Fun-1.3B-InP.
2026-06-18 11:35:18 +00:00
Will Lin 060bdd85b6 [revert] v2: drop FastWan/VSA registry entries — generic Wan loader can't map their gated-attn params
GPU verification (Wan2.1-Fun... no: FastWan2.1-T2V-1.3B) failed at load: 'Parameter blocks.0.to_gate_compress.bias
not found in custom model state dict' — the FastVideo/* DMD-distilled checkpoints carry gated-attention
params (to_gate_compress) that the generic WanTransformer3DModel loader can't map. v2/registry.py's
select_by_architecture ALREADY rejects WanDMDPipeline for exactly this reason; my explicit ModelEntry
wrongly bypassed it. Reverted FastWan (1.3B + 14B-480P) and the unverified VSA-14B alias (same FastVideo/*
risk). Kept the official Wan2.1-T2V-14B (standard weights, same loader path as the GPU-verified 1.3B).

Lesson recorded in V2_PORTING_STATUS.md: FastWan/Turbo/VSA need a param-mapping fix (like LTX-2.3 did),
not just a schedule — bucket-C-effort. GPU-verify every port before claiming support. CPU mini 211/2.
2026-06-18 11:31:31 +00:00
Will Lin 0a18b6a031 [feat] v2: alias FastVideo/Wan2.1-VSA-T2V-14B-720P to the Wan-14B card (bucket B)
Same WanTransformer3DModel arch (VSA is an attention-backend choice, not a weight/arch difference); v2
runs dense TORCH_SDPA, so it resolves to build_wan_t2v_14b_card. Registry resolution verified; the GPU
forward path is the 1.3B-proven Wan adapter (specific 14B/VSA weights not separately GPU-run).
2026-06-18 11:19:17 +00:00
Will Lin 2e5bad1d6c [feat] v2: port Wan2.1-T2V-14B (bucket B) + document the all-models backlog
First bucket-B port toward 'support every fastvideo model': Wan2.1-T2V-14B reuses the Wan recipe +
torch adapter unchanged (same WanTransformer3DModel/AutoencoderKLWan/UMT5) — only a registry entry +
build_wan_t2v_14b_card (720p, flow_shift 5.0) + SamplingDefaults differ. Without the entry the arch
fallback would give it the 1.3B 480p defaults; the explicit entry gives 50 steps / 720x1280.

Also recorded the full backlog in V2_PORTING_STATUS.md: 63 fastvideo models = 8 ported / 21 bucket-B
(reuse Wan/Causal/LTX2 arch — registry+recipe+defaults, no new adapter) / 34 bucket-C (13 new
architectures needing a TorchComponent adapter). Updated the stale 'how to add a model' steps to the
post-redesign structure (v2/recipes/, torch_backend.py, SamplingDefaults). CPU mini 211 pass, 2 skip.
2026-06-18 11:10:05 +00:00
Will Lin 705f580542 [refactor] v2 torch backend: TorchComponent base + one generic builder + v2.* facade (Phase 1b/1c)
Addresses the adapter-setup pains: collapses the torch_cuda(trampolines)/torch_adapters/torch_ltx2 split
+ 11 near-identical adapter classes + 6 build_torch_* builders into:
- v2/platform/backends/torch_backend.py: a TorchComponent base centralizing .to/.eval, the numpy<->torch
  marshalling (ONE place), the set_forward_context wrap, and the weight surface; thin per-model subclasses
  (WanDiT/LTX2DiT/WanVAE/LTX2VAE/T5Encoder/Gemma/LTX2Upsampler/LTX2AudioVAE/LTX2Vocoder) carrying only
  forward semantics; and ONE build_component(spec) dispatching by spec.kind via _MAKERS.
- torch_cuda.py: registers that single generic builder for all 6 cuda kinds (no per-kind trampolines).
- v2 owns its namespace via re-export STUBS (facade, marked '# STUB'): v2/forward_context, v2/fastvideo_args,
  v2/distributed, v2/loader (the load_component seam), v2/api, v2/models/{dits,audio,upsamplers}/*. All v2
  code imports v2.*; 'from fastvideo' now lives ONLY in those 8 stub files -> a future per-module vendored
  cutover swaps a stub body, no caller changes. No divergence (stubs run fastvideo's live code).
- Deleted torch_adapters.py + torch_ltx2.py.

Verified: CPU mini 210 pass/2 skip; lazy invariant (platform load imports no torch); GPU bit-parity LTX-2.3
T2VS (audio std 0.04304, identical to pre-redesign) + Wan2.1 (video std 31.94, motion 5.997).
2026-06-18 11:01:45 +00:00
Will Lin cbe67337c3 [feat] v2: per-model sampling defaults on ModelCard (Phase 1a)
v2 had no per-model defaults — generate_video hardcoded 30 steps/25 frames/480x832/cfg5/16fps for
every model, badly wrong for e.g. LTX-2 distilled (wants 8 steps @1024x1536) or Wan2.2-TI2V (704x1280@24fps).

- New SamplingDefaults dataclass + ModelCard.sampling_defaults (v2/card/specs.py), exported from v2.card.
- Populated all 7 supported cards from fastvideo's InferencePreset defaults (steps/guidance/HxW/frames/fps
  + per-modality guidance for LTX-2.3 A/V). Negative prompts copied verbatim into v2/recipes/_prompts.py
  (recipe DATA, not model code -> v2-owned, no fastvideo import).
- VideoGenerator stores the resolved card; generate_video applies card defaults with precedence
  kwargs > SamplingParam > card > generic fallback (pure _resolve_default helper, unit-tested).
- test_sampling_defaults.py: per-card values + precedence (incl. empty-neg-prompt edge). CPU mini 210 pass, 2 skip.
2026-06-18 10:47:28 +00:00
Will Lin 113478fd74 [refactor] v2: rename models/ (recipe layer) -> recipes/; move toy backend -> platform/backends/toy.py
Frees v2/models/ to become the vendored-architecture namespace that mirrors fastvideo/models
(part of making v2 self-contained / able to replace fastvideo). The v2 recipe layer (per-family
card.py/loop.py/program.py + common.py + the build_*_engine re-exports) is the recipe, not the
architectures, so it moves to v2/recipes/. The pure-numpy toy/parity implementations (ToyDiT etc.)
move from v2/models/backend.py to v2/platform/backends/toy.py (alongside cpu.py/accel.py/torch_*).

Mechanical: all imports are absolute, so v2.models.<x> -> v2.recipes.<x> and v2.models.backend ->
v2.platform.backends.toy across v2/ + examples/ (89 files, 178 refs). No behavior change.
CPU mini green (202 passed, 2 skipped); all v2 files compile.
2026-06-18 10:35:41 +00:00
SolitaryThinker 9b2007c058 [refactor] v2: use absolute imports (v2.*) everywhere instead of relative
Mechanical conversion of every relative import under v2/ to an absolute v2.* path
(from .x / ..x / ...x -> from v2.<pkg>.x) so imports are unambiguous, grep-able, and
stable when code is copied/moved between entrypoints (VideoGenerator, CLI, server).

Surgical prefix-only rewrite: only the 'from <dots><module>' prefix changed — import
names, parentheses, multi-line formatting, comments, and ordering are byte-for-byte
preserved (no collapsing, no reorder, no unrelated reformatting).

- 471 imports across 125 files; v2/tests/ was already absolute (untouched).
- Validated: all 125 files compile, every 'from v2.* import' target resolves to a real
  module/package, zero relative imports remain (full sweep), CPU mini suite green
  (202 passed, 2 skipped).
2026-06-18 09:41:19 +00:00
SolitaryThinker 1542b7a03d [feat] v2 VideoGenerator: A/V convenience path (generate_video -> T2VS -> mp4 + 24kHz wav)
Makes the 'Full A/V' LTX-2.3 deliverable reachable from the user-facing entrypoint, not just the
engine. A model advertising TEXT_TO_VIDEO_SOUND (LTX-2.3) now auto-issues a T2VS request, so generate()
/ generate_video() return BOTH modalities in one joint pass:
- VideoGenerator stores the resident instance + a supports_av flag (from card.capabilities); generate()
  gains want_audio (None=auto-by-capability, True/False to force) and routes T2V vs T2VS+{video,audio}.
- _result saves the stereo waveform as a sibling .wav at the vocoder's REAL rate (24000) — read off the
  built audio_vae adapter (TorchLTX2AudioVAE.sample_rate = Vocoder.output_sample_rate), since the
  AudioArtifact default rate is a placeholder. Populates GenerationResult.audio/.audio_sample_rate and
  extra['audio_path']. scipy IEEE-float WAV; [channels,samples] auto-transposed.
- Rewrote v2_basic_ltx2_3_distilled.py: registry routes to build_ltx2_3_card (its own joint T2VS A/V
  card, not the LTX-2 base/2-stage card); the example prints both the mp4 and the wav.
- GPU-verified via the convenience API: ev.mp4 + ev.wav (24000 Hz, stereo 61920x2, nonzero, std 0.043).
  CPU mini green (202 passed, 2 skipped); engine/program/toy paths untouched (test_ltx2_av pins 44100).
2026-06-18 09:41:18 +00:00
SolitaryThinker e4b7261b66 [feat] v2 LTX-2.3 T2VS GPU-verified: audio VAE/vocoder wiring + dual-connector audio fix
The full joint text->video+audio LTX-2.3 path now generates on the real 18.99B model:
- GPU audio components: build_torch_audio_vae (AudioDecoderLoader -> LTX2AudioDecoder, chains the
  vocoder) + build_torch_vocoder (VocoderLoader -> LTX2Vocoder); registered the 'audio_vae'/'vocoder'
  cuda component kinds; stamped their checkpoint subfolders (_WAN21_SUBFOLDERS).
- Fix: TorchGemma.encode_av must pass output_hidden_states=True — the 2.3 connector's SEPARATE audio
  projection lives in hidden_states[0] only then (gemma.py:703); without it the audio text fell back to
  the video embedding (4096 vs 2048 -> audio cross-attn shape mismatch).
- GPU-verified T2VS: video (3,33,256,384, std 0.68) + audio (stereo 2x61920 @24kHz, nonzero, std 0.059).
- test_torch_backend: cuda component kinds now include audio_vae + vocoder. CPU suite green (202+2).
2026-06-18 09:41:18 +00:00
SolitaryThinker b39b68c6f5 [feat] v2 LTX-2.3 T2VS: single-stage joint audio+video card/loop/program (CPU-verified)
Makes LTX-2.3 a first-class, faithful card (was wrongly merged into the single-stage base):
- LTX23DenoiseLoop (loop.py): single-pass joint A/V denoise — one DiT forward per step cross-attends
  video<->audio via the adapter's (v_vel,a_vel) return; full-res video latent + a [8,T,16] audio latent;
  distilled few-step schedule (BASE_SIGMAS). Video-only when no audio requested.
- build_ltx2_3_card (model_id 'ltx2.3-distilled' — the name now correctly names the REAL 2.3): 5
  components incl. audio_vae (AudioDecoder) + vocoder (required_for t2vs, optional_for t2v); caps
  T2V + T2VS. build_ltx2_3_program: dual-connector text-encode -> joint denoise -> video + audio decode.
- registry.py routes FastVideo/LTX-2.3-Distilled-Diffusers -> this card (split from the base entry).
- Toy support: ToyTextEncoder.encode_av (separate video/audio text), ToyDiT joint A/V (audio now a
  keyword-only arg so positional  callers like the talker are unaffected), channel-agnostic
  ToyAudioVAE (np.resize identity for the existing 2-stage T2VS).
- CPU-verified: toy T2VS -> video+audio (8 steps), T2V -> video-only; CPU suite green (202+2).
GPU audio-VAE/vocoder loaders + the real-T2VS GPU verify are the next step.
2026-06-18 09:41:17 +00:00
SolitaryThinker 5f698f03f5 [feat] v2 LTX-2 adapters: A/V foundation (joint DiT forward + dual text connector + audio decode)
Foundation for the LTX-2.3 T2VS port (card/loop/program wiring + GPU verify to follow):
- TorchLTX2DiT.__call__ gains an optional joint audio path: pass audio_latent[8,T,16] + audio_text and it
  feeds audio_hidden_states/audio_encoder_hidden_states/audio_timestep/audio_sigma in ONE forward
  (LTX-2.3 cross-attends video<->audio) and returns (video_velocity, audio_velocity). Video-only call is
  byte-for-byte unchanged (audio_latent=None).
- TorchGemma.encode_av returns the SEPARATE (video_text, audio_text) projections from the 2.3 connector
  (video=last_hidden_state, audio=hidden_states[0]); 2.0 returns them equal.
- TorchLTX2AudioVAE (AudioDecoder -> Vocoder -> waveform@24kHz) + TorchLTX2Vocoder wrapper.
Additive + backward-compatible; CPU suite unaffected (adapters are GPU-lazy).
2026-06-18 09:41:17 +00:00
SolitaryThinker 799a9b69a2 [refactor] v2: shared model registry (HF-id primary + arch fallback) for all entrypoints
Per review: dispatch should be a directly-mapped HF-string -> card registry (like fastvideo), shared
by every entrypoint (VideoGenerator + a future CLI / server), not buried in the generator.

- New v2/registry.py mirrors fastvideo's fastvideo/registry.py hybrid resolution: (1) exact HF repo id
  in an explicit ModelEntry registry [PRIMARY — correct per-model card/capabilities, and the only way to
  split same-architecture capability variants like Wan2.1 T2V vs the i2v 'InP' 1.3B], (2) short repo-name
  match, (3) architecture inference [FALLBACK — local paths / unregistered repos]. resolve(model_path[,
  root]) is the single shared entry point.
- video_generator.py: moved _read_arch_signature/_select_builders into the registry; from_config now
  calls resolve() — registered ids resolve with no config read, else arch inference on a cheap *.json
  snapshot. Reconciles the earlier 'no brittle table' refactor with the 'map hf string -> card' ask: one
  clean registry + a fallback, not three coupled structures.
- Verified: registry resolves exact-id / short-name / arch-fallback / unregistered correctly; CPU suite
  green (202+2); wan21 GPU smoke generates via the new resolve path.
2026-06-18 09:41:17 +00:00
SolitaryThinker 27b320afd8 [fix] v2: name LTX-2 cards by architecture (2-stage vs single-stage) + Wan2.1 is T2V-only
Addresses the 'how is ltx2 separate from ltx2.3' confusion + a wrong capability:

- LTX-2 cards renamed by ARCHITECTURE (the version labels did not map to it): build_ltx2_card model_id
  'ltx2.3-distilled' -> 'ltx2-2stage-distilled' (two-stage base->upsample->refine; serves the
  upsampler-having FastVideo/LTX2-Distilled-Diffusers); build_ltx2_base_card 'ltx2.base' ->
  'ltx2-single-stage' (one loop; serves Davids048 base + the single-stage FastVideo/LTX-2.3-Distilled,
  which has NO spatial_upsampler). Dispatch already splits on has_spatial_upsampler. Updated the
  model-id refs in the mini's tests/examples.
- Wan2.1 base is T2V-only: dropped the wrong Capability.IMAGE_TO_VIDEO + narrowed components'
  required_for to {t2v} (i2v is the separate InP variant; v2 has no i2v path yet). build_wan21_card is
  shared by wan21 + wan2.2-ti2v; the A14B card was already T2V-only.
- CPU suite green (202 passed, 2 skipped).
2026-06-18 09:41:17 +00:00
SolitaryThinker c74e9cb4fe [docs] v2: LTX-2 base/2.3 GPU-verified on rebuilt x86 stack + remaining-port mechanisms
- LTX-2 base (Davids048) and LTX-2.3-Distilled both generate real video (inter-frame motion 4.5 / 6.6)
  via the single-stage base card — moved to Working (7 models now verified).
- Environment: the aarch64 venv was rebuilt for x86 (torch 2.11.0+cu128) + re-validated (CPU suite green,
  wan21 + LTX-2 base/2.3 generate).
- Remaining Wan+LTX-2 ports documented with concrete mechanisms: Wan2.2-i2v (SigLIP image_encoder +
  VAE-encode first-frame + concat-mask -> larger-in_channels i2v DiT), TurboWan (RCMScheduler consistency
  loop), Lucy-Edit (Wan v2v via VideoVAEEncodingStage), FastWan (VSA, env-blocked: no nvcc).
2026-06-18 09:41:17 +00:00
SolitaryThinker 0a2c985b56 [docs] v2: LTX-2.3 example + roadmap (A14B offload working, base/2.3 ported, env status)
- v2_basic_ltx2_3_distilled.py: LTX-2.3-Distilled routes to the single-stage base card (no
  spatial_upsampler) via the arch dispatch; pass few steps for the distilled schedule.
- V2_PORTING_STATUS.md: A14B moved to Working (CPU expert offload, 60GB peak); LTX-2 base + 2.3 added
  (code-complete, GPU re-verify pending); Environment-status note on the mid-session aarch64->x86 host
  reschedule that blocks GPU re-verify.
2026-06-18 09:41:17 +00:00
SolitaryThinker 8b343db604 [feat] v2: Wan2.2-A14B MoE CPU offload (fits 1 GPU) + LTX-2 base/2.3 single-stage port
Within the bounded Wan+LTX-2 scope:

- Wan2.2-A14B MoE now GENERATES on a single 80GB GPU via CPU offload (TorchWanDiT offload_group):
  the two 14B experts live on CPU and only the active one is swapped onto the GPU at the boundary-
  timestep transition (a single swap, not per-step). GPU-verified: 60GB peak (vs 79GB OOM), produced
  wan22_a14b_lion.mp4 (17x480x832, std 54.7, motion 8.36). Single-expert Wan stays resident.

- LTX-2 base (single-stage) port: build_ltx2_base_card + build_ltx2_base_program reuse the LTX-2
  adapters at FULL latent res with a request-driven many-step flow-match (LTX2DenoiseLoop full_res/
  request_steps/base_flow_sigmas; distilled base/refine path preserved via False defaults). The SAME
  single-stage card serves LTX-2.3-Distilled (also single-stage: no spatial_upsampler) — dispatched by
  the new has_spatial_upsampler discriminator in _select_builders. v2_basic_ltx2.py added; VideoGenerator
  gains shutdown() for API parity.

- Fixes: from_config 'os' scoping (shadowed module import); LTX-2 upsampler per_channel_statistics
  source (the AE's .decoder, not the top-level module).

Verification status: A14B offload, the upsampler, and the arch-dispatch refactor were GPU-verified
earlier this session. LTX-2 base/2.3 are CPU-verified (cards/programs build, dispatch routes, schedule
correct); their GPU smoke tests were pending when the box was rescheduled aarch64->x86 mid-session,
which broke the aarch64 venv (numpy/torch unrunnable on x86) — GPU re-verify blocked on the env.
2026-06-18 09:41:17 +00:00
SolitaryThinker 24b76d372c [feat] v2: architecture-driven dispatch + real LTX-2 upsampler + Wan2.2-A14B MoE card
Two reviewer asks + the next Wan port, all within the bounded Wan+LTX-2 scope:

1. Architecture-driven dispatch (replaces the HF-id table + substring fallback): from_config reads the
   checkpoint's pipeline/transformer/VAE class names (+ z_dim, transformer_2) and picks the v2 card via
   _select_builders — mirroring fastvideo's get_pipeline_config_cls_from_name. Resolves local paths /
   renamed repos / new distilled variants of a known arch with no table edits, and cleanly REJECTS
   FastWan (detected by WanDMDPipeline) with a precise message instead of a confusing load crash.

2. Real LTX-2 spatial upsampler (was a nearest-neighbor np.repeat stand-in): new 'upsampler' component
   kind -> TorchLTX2Upsampler wraps the real LTX2LatentUpsampler and applies the repo's upsample_video
   (un_normalize via the VAE decoder's per_channel_statistics -> learned 2x upsample -> normalize). CPU
   keeps ToyUpsampler (np.repeat) via the factory terminal, so the program calls
   component('spatial_upsampler').upsample(...) on both backends with no device branch. GPU-verified:
   9x512x768, std 70.5, motion 6.51.

3. Wan2.2-T2V-A14B MoE card (build_wan22_a14b_card): two WanTransformer3DModel experts +
   BoundaryTimestepRouting @0.875. GPU-verified that both experts denoise; OOMs in VAE decode on one
   80GB GPU (~70GB resident) — upstream offloads the DiT for MoE; documented as offload-blocked.

CPU suite green (202 passed, 2 skipped). FastWan root-caused (non-strict load of VSA gate_compress +
VSA not built); roadmap (V2_PORTING_STATUS.md) updated with the bounded scope + per-model status.
2026-06-18 09:41:16 +00:00
SolitaryThinker f7256ab197 [feat] v2 port: Wan2.2-TI2V-5B (T2V) — 4th GPU-verified model
- Wan2.2-TI2V-5B reuses the Wan adapters (WanTransformer3DModel / AutoencoderKLWan / UMT5); deltas
  are the higher-compression VAE geometry (z_dim=48, 16x spatial, 4x temporal) and 480p flow-shift 5.0.
  The DiT forward accepts a scalar timestep (1D path), so no per-frame expand_timesteps for pure t2v.
- WanDenoiseLoop / build_wan21_card gain optional geometry params (latent_channels/spatial_ratio/
  temporal_ratio) defaulting to Wan2.1 (16/8/4) -> wan21 path unchanged; build_wan22_ti2v_card sets
  48/16/4. Registered as family 'wan2.2-ti2v' in VideoGenerator; v2_basic_wan2_2_ti2v.py added (T2V).
- Verified: 25x448x768 mp4, std 62.3, inter-frame motion 4.89 (coherent). CPU suite green (202+2skip).
- Corrected the now-disproven FastWan='wan21 reuse' mapping (its DMD checkpoint to_gate_compress param
  mapping differs); roadmap updated (TI2V-5B working; A14B MoE + I2V remain).
2026-06-18 09:41:16 +00:00
SolitaryThinker 9de0e2bcd1 [feat] v2 VideoGenerator: convenience API (from_pretrained/generate_video) + porting roadmap
- VideoGenerator gains the convenience surface most basic examples use: from_pretrained(model,
  num_gpus/*_cpu_offload/...) and generate_video(prompt, sampling_param=, **kwargs) -> result, on top
  of the typed from_config/generate. Accepts SamplingParam.
- v2 examples matching the upstream convenience-API examples for the verified models: v2_basic.py
  (Wan2.1) and v2_basic_self_forcing_causal.py (SF-causal).
- V2_PORTING_STATUS.md: honest per-family roadmap. Working: wan21, wan_causal, ltx2-distilled. Each
  further model needs per-model work (FastWan: WanDMD to_gate_compress param mapping; TurboWan: RCM
  consistency sampler; Wan2.2: MoE card; LTX2 base/i2v; new families: new cards/adapters; gated Flux2 /
  local GEN3C / audio StableAudio / interactive MatrixGame blocked in this env).
2026-06-18 09:41:16 +00:00
SolitaryThinker 5f33f80d6c [feat] v2 VideoGenerator: typed fastvideo.api entrypoint over the v2 engine
Mirrors fastvideo.entrypoints.VideoGenerator (from_config(GeneratorConfig) -> generate(
GenerationRequest) -> GenerationResult.video_path), reusing the OFFICIAL fastvideo.api config classes
so a basic_dmd_new_api.py-style script differs only by importing VideoGenerator from v2.
- model_path -> v2 card registry (Wan2.1 / FastWan -> wan21; SFWan -> wan_causal; LTX2 -> ltx2);
  snapshot_download + stamp_wan21_checkpoints + Engine(cuda) + program; generate maps SamplingConfig
  -> DiffusionParams -> make_request -> eng.run, saves the [C,T,H,W] decode as an mp4.
- Lazy v2.__getattr__ keeps 'import v2' torch-free (verified) so the CPU mini stays green (202+2).
- examples/inference/basic/v2_basic_new_api.py runs all three GPU models through this API.
- Single-GPU, resident, TORCH_SDPA (EngineConfig offload/num_gpus>1/VSA accepted for parity, not applied).
verified: wan21 from_config->generate->mp4 (frames (5,256,256,3) uint8, video_path written).
2026-06-18 09:41:16 +00:00
SolitaryThinker e7d4aaeef4 [feat] v2 ltx2: two-stage distilled GPU bring-up (LTX2Transformer3DModel 18.88B)
Official FastVideo/LTX2-Distilled-Diffusers. New torch_ltx2.py adapters (build_torch_* dispatch on
class):
- TorchLTX2DiT: patchify-internal; per-token timestep ones(B,tok,1)*sigma (sigma direct) + per-sample
  video_sigma; DiT predicts x0 so the adapter returns velocity=(x_t-x0)/sigma for the v2 flow-match step.
- TorchLTX2VAE: CausalVideoAutoencoder decode (internal per-channel un_normalize).
- TorchGemma: LTX2GemmaTextEncoderModel (Gemma + feature-extractor + connectors) -> last_hidden_state.
- ltx2 loop: real 128-ch latent geometry on cuda (32x spatial / 8x temporal; half-res base, 2x upsample).
e2e two-stage (8+3 steps) -> coherent, high-quality video (surfers at sunset), (3,9,512,768). NOTE:
still uses the v2 program's np.repeat upsampler between stages (the refine regenerates from noise so
output is faithful-quality); real LTX2LatentUpsampler swap-in is a follow-up.
2026-06-18 09:41:16 +00:00
SolitaryThinker 2a085a5c39 [feat] v2 wan_causal: causal DiT (CausalWanTransformer3DModel) GPU bring-up
Official SF checkpoint wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers (reuses TorchWanVAE + TorchT5Encoder).
- TorchWanDiT detects the causal transformer: ignores the chunk_rollout loop's latent `context` (the
  real model conditions across chunks via an internal kv_cache, not a forward arg) -> dispatches to
  full-attention _forward_train, and passes a per-latent-frame timestep [B, num_frames] (the causal
  block asserts a per-frame temb), uniform per chunk.
- chunk_rollout real geometry on cuda (16ch; chunk_size latent frames; 8x spatial).
e2e: chunk_rollout over the SF student -> coherent video (cat in a garden), (3,21,480,832). Fidelity
gap (artifacts): the v2 loop's per-chunk few-step sampling != the official kv-cache streaming + SF
schedule (a follow-up).
2026-06-18 09:41:16 +00:00
SolitaryThinker 261d3007a9 [fix] v2 tests: make no-torch-import guards GPU-aware (skipif torch installed)
The cuda-availability probe imports torch by design; the no-torch-import invariant is only
verifiable when torch is absent. skipif torch installed -> green on GPU box (202 passed, 2 skipped),
still enforced in torchless CI.
2026-06-18 09:41:16 +00:00
SolitaryThinker b2da3b2b4d [feat] v2 wan21: backend-aware latent geometry + checkpoint stamping
- latent_shape(req, model): real Wan geometry (16ch; (T-1)//4+1, H/8, W/8) on the cuda backend;
  toy stand-in stays on accel/cpu.
- stamp_wan21_checkpoints(card, model_root): map a root (local dir or HF id) onto the 3 components'
  ComponentSpec.checkpoint; build_wan21_card(checkpoint_root=...) optional.
2026-06-18 09:41:16 +00:00
SolitaryThinker 7ec4d01afc [feat] v2 cuda backend: real fastvideo construction + risk A-E fixes (Wan2.1 verified on H100)
Take the written-not-run torch adapters to runs-and-generates on 1x H100 (aarch64):
- A: FastVideoArgs.from_kwargs(model_path=root) builds the real pipeline_config; single-GPU dist
  init; load each component from its subfolder; tokenizer from the sibling <root>/tokenizer.
- B/C: DiT forward wrapped in set_forward_context(attn_metadata=None) (SDPA dense path);
  timestep=sigma*1000 + bare-velocity output confirmed.
- D: VAE decode denormalizes z*std+mean; removed the double-mean (it re-added
  shift_factor==latents_mean) that washed out the video.
- E: UMT5 from config; text embeds zero-padded to text_len (Wan t5_postprocess_text) - the fix
  that took output from a dark blur to a coherent prompt-matching scene.
- Components run at native precision (DiT bf16, VAE/text fp32); checkpoint check before dist init.
2026-06-18 09:41:15 +00:00
SolitaryThinker 8b0e2bbb4b [docs] add v2/HANDOFF.md for GPU-side bring-up of the torch backend
Orientation + process doc for an agent on a GPU branch: the 6 commits already
landed, the files to touch, the gating tasks (Risk A FastVideoArgs/checkpoint),
the verification bar (CPU suite stays 204; GPU generation matches a reference via
the SSIM harness), commit/push rules (no Claude co-author; don't rewrite history;
wandb token referenced not embedded), and the gotchas. Points to
GPU_BRINGUP.md for the detailed checklist + risk table.
2026-06-18 09:41:15 +00:00
SolitaryThinker 7fd58f26ea [fix] correct GPU adapters against real fastvideo API (cross-check findings)
Adversarial cross-check of the written-not-run torch adapters against the real
fastvideo source confirmed the interface contracts (DiT returns bare velocity
tensor; timestep=sigma*1000; encode().mode() + bare decode; .last_hidden_state;
no fused solver kernel) but caught a wrong construction layer. Fixed in code:

- Construction: WanTransformer3DModel / AutoencoderKLWan have NO from_pretrained.
  Replace it with the real FastVideo loaders (TransformerLoader / VAELoader /
  TextEncoderLoader + TokenizerLoader, each load(model_path, fastvideo_args)).
  The loader resolves the class from the checkpoint config — so UMT5-vs-T5 is
  chosen correctly instead of hardcoded (was BLOCKER #1/#4/#5).
- Text encoder: wrap the forward in set_forward_context(...) — the (U)MT5
  attention reads global state via get_forward_context(); a bare call mis-encodes
  (was BLOCKER #2). Drop the wrong padding="max_length".
- VAE: apply latent normalization the DiT expects — (z-mean)*inv_std on encode,
  inverse on decode, with latents_std stored as its reciprocal; shift_factor
  before decode (was BLOCKER #3). Skipping it yields washed-out video, not error.

The remaining unknowns are genuinely box-dependent (FastVideoArgs fields,
shift_factor placement, exact tokenizer kwargs, FSDP) — GPU_BRINGUP.md reconciled
to mark what's now fixed-in-code vs what still needs the box. 204 CPU tests pass.
2026-06-18 09:41:15 +00:00
SolitaryThinker 15b6278e9f [feat] real torch/CUDA backend (written-not-run) behind the cuda cells
Implement the GPU backend the substrate was built for: Platform.detect() ->
cuda resolves real torch adapters + torch solver ops instead of the numpy
rungs, with the existing loops/policies/scheduler/training unchanged.

WRITTEN-NOT-RUN: this environment has no GPU/torch, so the torch code is
grounded in the verbatim real fastvideo APIs (DiT forward signature confirmed
from source) but cannot be executed/verified here. It is gated available=False
(CPU mini stays green; importing the backends never imports torch), with every
on-box confirm point marked `# BRINGUP` and an ordered checklist in
platform/backends/GPU_BRINGUP.md.

- torch_adapters.py: TorchWanDiT / TorchWanVAE / TorchT5Encoder wrap the real
  module named by each card's load_id and bridge it to the mini's duck-typed
  surface (numpy<->torch at the boundary; loop math stays numpy fp32). DiT
  weight-surface (copy_from/blend_from/clone) for serving sync; mse_grad_step
  raises (GPU training is a separate workstream).
- torch_kernels.py: flow_match_step / flow_sde_step as plain torch elementwise.
  Grounded conclusion from the kernel audit: fastvideo-kernel ships NO fused
  solver kernel (only attention/norm/quant primitives), so the cuda solver is
  torch, registered at arch generic with an honest source string.
- torch_cuda.py: rewritten as lazy trampolines (torch imported only inside
  builder/kernel bodies). Adds the missing vae + text_encoder cuda components
  (they'd otherwise silently fall back to the toy) and corrects the dishonest
  "fastvideo-kernel:flow_*" labels.
- ComponentSpec.checkpoint: the weights source for the torch adapter (risk A;
  the one field the cards didn't carry). Empty on the CPU toys.
- 7 CPU-verifiable wiring tests: honest registration/sources, torch-free import,
  cuda-resolves-real-cells-not-toy, build-fails-loudly-without-torch.

204 tests pass (CPU). The torch path needs a GPU box to verify (GPU_BRINGUP.md).
2026-06-18 09:41:15 +00:00
SolitaryThinker cd0fbbd3a0 [feat] static-buffer capture form for the cudagraph step body (Path A)
Close the loudest deferred gap from the cudagraph audit: the capturable step
now binds its I/O to address-stable static buffers (modeling real CUDA static
I/O buffers), instead of allocating fresh arrays per call.

- StaticWorkspace: address-stable buffers allocated once per capture key; bind()
  copies the current step's inputs in place via np.copyto, which RAISES on a
  shape/dtype mismatch — turning the weak peak_activation_bytes proxy into a real
  key-soundness backstop (a step whose shape doesn't fit can't replay an
  incompatible graph). Output written into a static buffer too.
- WorkPlan.graph_fn / graph_inputs: a capturable step exposes its deterministic
  op-structure as graph_fn(model, workspace) reading EVERY per-step input (latent,
  sigmas, conditioning, scale) from the workspace — never from closure over
  per-step data — plus the dict of current values. Loops without both stay on the
  eager path. wan21 factors a shared _velocity() so graph_fn and the eager run
  stay bit-identical.
- Capturer dispatch captures/replays via graph_fn against the keyed workspace;
  the workspace is shared per key on the instance. Correct under the engine's
  synchronous step execution (bind+graph_fn atomic per dispatch, output returned
  as a copy) — proven by the batch-of-N interleave gate running two same-key
  requests through the shared workspace bit-identically. A concurrent/multi-stream
  executor would need a per-stream pool (documented).
- 2 new tests (no-static-form eager-break, static-buffer shape-mismatch raises);
  workspace collision test split into bytes-proxy vs shape-backstop.

197 tests pass.
2026-06-18 09:41:15 +00:00
SolitaryThinker 02151584cd [feat] piecewise CUDA-graph capture/replay at the step boundary (Path A)
Wire the capture/replay lifecycle into the driven-loop step boundary — the
other half of Path A (hand-fused kernels behind the registry + piecewise
cudagraphs, no compiler). Models and tests the correctness-critical control
logic; replay re-runs the current step thunk (CPU models the lifecycle, not the
GPU speedup).

- GraphCapturer on the instance (cross-request cache), wired into
  RuntimeLoopContext.execute and gated by LoopSpec.graph_capture ==
  "breakable_cudagraph". Capture key = (device, arch, loop, shape_sig,
  resident-weight-versions, graph_key). Eager-break for non-capturable (SDE) and
  interceptor-overridden steps. Never executes a stored thunk, so interleaved
  requests can't smear state.
- WorkPlan.capturable / graph_key. wan21 sets capturable=not sde and folds
  compute dtype (shape_sig.dtype) + CFG branch set + expert + scheduler-precision
  into the key — closing a cross-precision key-collision corruption path a
  non-fp32 build would otherwise hit on a real GPU (audit finding).
- Version-in-key auto-invalidation + real eviction: set_weights_version evicts
  the synced component's graphs (duck-typed, so card/ imports no runtime),
  preventing a GPU graph leak across FlowGRPO's per-iteration syncs.
- register_kernel gains the workspace_bytes capture-safety contract (declared in
  the matrix; cuda cells declare real scratch, numpy reference is 0).
- 11 tests: capture-once/replay-many, eager-break (SDE + override), capture ≡
  pure-eager bit-identical, recapture on shape + weight-version change with
  eviction, eager-loop gating, accel-backend capture, and capturer unit tests
  (key discrimination, eager-break, workspace-collision safety net, eviction).

Honestly deferred (bite a real GPU, not the CPU tests): the static-buffer
refactor of the step body, admission budgeting of capture cost (GRAPH_CAPTURE),
and per-card opt-in beyond wan21 — all documented in cudagraph.py + README.

195 tests pass.
2026-06-18 09:41:15 +00:00
SolitaryThinker 51bb5e349a [feat] route all diffusion loops + RL recompute through the kernel table
Finish the kernel seam across the board so the platform's KernelTable is the
universal solver-dispatch path, not just wan21.

- Loops: ltx2, wan_causal, adapters, adaptive now resolve the flow-match solver
  via model.platform.kernels.get(FLOW_MATCH_STEP) instead of importing the numpy
  sampler directly (wan21 already did). On CPU this is bit-identical (the cpu
  kernel IS the old function); a GPU/accel backend now overrides every loop.
- RL: the FlowGRPO log-prob recompute in unified_rl / joint_multi_rl /
  workflow_rl dispatches FLOW_SDE_STEP through the platform, pinned to the SAME
  kernel the rollout used (C2 kernel-pinning — otherwise the PPO ratio biases on
  a real GPU where rollout and recompute kernels could differ).
- accel backend: add a kind-generic AccelComponent wrapper and override the vae
  component too (text_encoder left unregistered to keep the device->cpu fallback
  demonstrated), closing the "only dit is overridden" gap.
- tests: vae override assertion; a second-loop (wan-causal chunk rollout) parity
  oracle proving accel == cpu bit-identical beyond wan21. README scope updated.

184 tests pass.
2026-06-18 09:41:15 +00:00
SolitaryThinker 1bdbf1d5b3 [feat] multi-backend dispatch substrate (device/arch/kernel registries)
Add v2/platform/: the (recipe, runtime) backend membrane that lets CPU, GPU,
and other devices coexist behind one dispatch substrate.

- Two tuple-keyed registries: COMPONENTS(kind, device, variant) for
  weight-bearing components and KERNELS(op, device, arch, variant) for
  stateless primitives, each with an availability predicate and an enumerable
  manifest (declared-but-unavailable cells listed without importing torch).
- Platform: detected (device, arch) owning the device + arch fallback chains
  and a per-platform cached KernelTable; detect() is honest (CPU/numpy unless
  torch+CUDA are actually present). Arch fallback is monotonic — only degrades
  to older/portable archs, never a newer binary-incompatible one.
- Seams wired: ModelInstance.component() -> platform.build_component(spec,self)
  with spec.factory as the numpy terminal rung (existing cards untouched); the
  wan21 denoise thunk dispatches solver ops through model.platform.kernels.
- Three backends: cpu (numpy terminal + parity oracle), accel (pure-python
  stand-in proving cross-device resolution, arch fallback, and the oracle),
  torch_cuda (declared-but-unavailable; no faked GPU).
- test_platform.py: 16 tests — detection, terminal rung, arch-fallback walk +
  monotonicity, device precedence, variant fallback, component override +
  per-kind device fallback, the parity oracle (accel == cpu, bit_identical via
  the C1 ladder), and matrix enumeration without importing torch.

Scope is honestly bounded in the README/docstrings: only wan21-denoise routes
through the kernel table and only the dit kind is overridden today (each a
one-line adoption); the torch/CUDA path and a cudagraph workspace-safety
contract are declared/deferred, not implemented. 183 tests pass.
2026-06-18 09:41:14 +00:00
SolitaryThinker de295403c7 [feat] Adapter plane, non-linear workflows, RL→distill flywheel
Three more capabilities on distinct untested surfaces (167 tests pass, no new runtime primitive).

A — Adapter plane (§9.19): one base + swappable LoRA/ControlNet adapters, selected per request
(DiffusionParams.adapters); AdapterDenoiseLoop applies each active adapter's velocity delta. Per-request
selection changes output, multi-LoRA composes, ControlNet conditions on a control image, mixed-adapter
requests interleave without smearing, hot-swap changes generation, cache key partitions by adapter stack
(the adapter_versions field, previously declared-only). ToyLoRA/ToyControlNet; models/adapters/. 6 tests.

B — Non-linear workflows (§9.17): ParallelWorkflow (fan-out: one input → N models → merged) and
BestOfNWorkflow (generate N → score with the served reward card → return best; inference-time scaling).
The shapes a linear chain can't express. 4 tests.

D — RL→distill flywheel (§9.18): run_flywheel RL-improves the base (NFT), then distills FROM the RL'd model
(DMD2 teacher = RL'd policy) into a faster card, recording the base→rl→distilled provenance chain in
RecipeSpec.parents. The distilled student is measurably closer to the RL'd teacher than the base; the
distilled card serves few-step. training/flywheel.py. 4 tests.

designv4 §9.17–§9.19 + layout/closing/counts (167 tests, 29 files).
2026-06-18 09:41:14 +00:00
SolitaryThinker 34834eb4f4 [feat] #8b speculative (draft-verify) decoding — exact + lower-latency AR
The last audit stress test. A cheap draft model proposes K tokens, the target verifies
them in one batched step, and SpeculativeARLoop accepts the matching prefix + one target
correction — a variable accepted-length per round (a ragged AR loop the model owns).

- Exactness: the emitted sequence equals the target's OWN greedy decode for any draft
  quality (every accepted token is one the target would produce; the correction is the
  target's token) — the speedup is free.
- Speedup scales with accept rate: draft-agree 0.3→1x, 0.7→3x, 1.0→4x=K tokens/round
  (fewer verify_rounds, the expensive model's latency steps, for the same output).
- Two components (draft + target) co-scheduled on one resident instance; each round an
  AR_TOKEN WorkUnit.

models/speculative/ (loop+card+program); backend ToyTargetModel/ToyDraftModel with a
shared length-dependent target formula (no degenerate fixed point). 5 tests; full suite
153 passed. designv4 §9.16 + counts (153 tests, 26 files).
2026-06-18 09:41:14 +00:00
SolitaryThinker bd36d69680 [feat] Five more stress tests: LTX-2 A/V, weight-sync, served reward, cache-dit, nested workflows
The remaining design_v3 probes (all except 8b speculative decoding). All fit with no new runtime
primitive (148 tests pass).

#6 LTX-2 joint audio+video (§9.11) — LTX-2 declared an audio_vae required_for t2vs but never used
   it; now a single 2-stage denoise carries a synchronized audio latent (conditioned on video),
   applies per-modality CFG (guidance_per_modality), and decodes via video VAE + audio VAE → video +
   audio. Gated on requesting audio, so the T2V path is byte-identical (existing tests untouched).
   ToyAudioVAE; build_ltx2_av_program. 5 tests.

#4 Live weight-sync under in-flight serving (§9.14) — WeightSyncController makes the freeze → drain →
   transfer → bump version + invalidate → resume lifecycle explicit. Tests: a mid-flight swap corrupts
   (the hazard); draining first leaves the in-flight request bit-identical to baseline while a
   post-sync request reflects new weights; transformer-only sync, so the frozen text-encoder cache
   survives. The RL flywheel's hardest correctness. 3 tests.

#5 Reward-model-as-a-served-card (§9.15) — a reward model is a card (scorer + a score loop emitting
   REWARD_BATCH units); ServedRewardScorer drop-in-replaces the numpy scorer so any RL method becomes
   RLHF/RLAIF with no method change. ToyRewardModel; models/reward/. 4 tests.

#7 Content-adaptive control flow (§9.12) — CacheDiTDenoiseLoop (isolated WanDenoiseLoop subclass)
   reuses the cached velocity when predictions barely change (cache-dit skip) and early-exits on
   convergence — variable step count; interleave parity holds across ragged loops. models/adaptive/. 4 tests.

#8a Nested workflows (§9.13) — a workflow stage can invoke another workflow (engine.run routes ids);
   requires/validate recurse; cycles caught at registration + a run-time guard (engine._wf_running).
   build_t2i_i2v_extend_workflow. 5 tests.

designv4 §9.11–§9.15 + falsifier/layout/closing updates (148 tests, 25 files). Also removed a
pre-existing unused import in ltx2/loop.py.
2026-06-18 09:41:14 +00:00
SolitaryThinker 94b84b03b9 [examples] Add v2_examples/{training,omni,workflows}/ — runnable examples
Three more example folders alongside inference/, all CPU/numpy, self-contained
(sys.path bootstrap), public API only, every script verified to run green.

training/ (7) — one per method, all via the uniform method.train_step seam:
  01 finetune · 02 dmd2 distillation · 03 diffusion_nft (likelihood-free RL,
  samples from the old policy, feature-cache reuse) · 04 self-forcing (causal
  chunk_rollout) · 05 joint LM+generator RL (UniRL; joint + prompt-only) ·
  06 N-way joint RL (per_expert vs shared credit) · 07 end-to-end workflow RL
  (T2I+I2V from one final-video reward).

omni/ (4) — 01 Cosmos3 (reason→joint denoise, shared MoT) · 02 BAGEL
  (text→image, shared MoT; scheduler prices both WorkUnit kinds) · 03 Qwen-Omni
  (thinker→talker→vocoder, three separate experts, text+audio) · 04 interleave
  parity across AR + diffusion loop types.

workflows/ (2) — 01 cross-model T2I→I2V workflow (image provably conditions the
  video) · 02 workflow as a first-class servable (requires/validate, address by
  id, register_workflows catalog, WorkflowRegistry).

Each folder has a README indexing its scripts.
2026-06-18 09:41:14 +00:00
SolitaryThinker a16ce1a07f [examples] Add v2_examples/inference/ — runnable Wan2.1 inference examples
Five self-contained, runnable scripts (CPU/numpy) for the Wan2.1-1.3B card on the
v2 runtime, each bootstrapping the repo onto sys.path so they run from anywhere:

- 01_basic_t2v.py                  minimal path: build engine → T2V request → run → video
- 02_params_and_reproducibility.py DiffusionParams knobs + seeded bit-identical reproducibility
- 03_streaming.py                  per-denoise-step preview chunks (OutputSpec stream)
- 04_concurrent_interleaved.py     step-interleaved batching + interleave parity gate + cache reuse
- 05_async_serving.py              AsyncEngine: concurrent generate, event stream, step-boundary cancel

+ README.md indexing them. All five run green; use only the public API.
2026-06-18 09:41:14 +00:00
SolitaryThinker 37544c68a5 [docs] Update v2/README to designv4 + current scope (127 tests)
- v2/README.md: point to designv4.md as the unified design (design_v3 as
  north star); refresh the scope table (joint/N-way/workflow RL, Qwen-Omni
  cascade, cross-model Workflow, tiled VAE co-scheduling, WorldModelSession),
  package layout (program/Workflow, runtime/session, the 7 methods, new model
  dirs), the demonstrated-stress-tests list (§9.3–§9.10), and counts (49→127,
  20 files). Sessions moved out of "out of scope"; WebRTC wire stays out.
- designv4.md: drop two intermediate absolute suite totals (milestone "91/97
  passed") in favor of "zero regressions" so the only absolute count is the
  current 127 (intro + layout) — no stale numbers.
2026-06-18 09:41:13 +00:00
SolitaryThinker 88e4a4355c [feat] Three stress tests: interactive sessions, workflow RL, heterogeneous co-scheduling
Targets the three design_v3 claims that were most load-bearing AND least
exercised (sessions/realtime, training-plane boundary, the WorkUnit-generality
falsifier). All fit with no new runtime primitive (127 tests pass).

1. Interactive world-model session (runtime/session.py, §9.8) — the Session
   plane had ZERO coverage. WorldModelSession drives the causal chunk_rollout
   loop as a long-lived session: persistent cross-request world state on the
   Session.kv_handle, frame streaming, transactional step-boundary cancellation
   (a cancelled act leaves the world resumable), no cross-session smearing.
   Added only a continuation seam to the chunk loop (init seeds context from a
   world_context slot; default empty = unchanged one-shot path). 5 tests.

2. End-to-end RL over a cross-model workflow (training/methods/workflow_rl.py,
   §9.9) — trains BOTH flux-t2i and wan-i2v from ONE final-video reward. Rolls
   out the whole workflow with SDE capture in both instances; the same final
   advantage drives FlowGRPO PPO on each stage's transformer; two WeightSyncPlans
   on two instances. The earlier model (T2I) is trained by a reward on the final
   video — end-to-end credit across a model boundary — proven causal by a control
   (constant reward => zero advantage => nothing moves). 4 tests.

3. Heterogeneous WorkUnit co-scheduling (models/tiled/, §9.10) — the §17
   falsifier. VAETileLoop makes VAE decode a loop of VAE_TILE units; tiling is
   exact (== one-shot, C0), and VAE_TILE + DIFFUSION_STEP pipelines interleave
   bit-identically and co-run in one batch. Validates the mechanism; the
   economic half (does it pay) stays a port-time measurement. 4 tests.

designv4 §9.8–§9.10 + falsifier/layout/closing updates (127 tests, 20 files).
2026-06-18 09:41:13 +00:00
SolitaryThinker f1dead1361 [feat] Register cross-model workflows as first-class named servables
Answers "what's the right way to name/register custom pipelines like T2I→I2V":
treat a Workflow like a card — a stable namespaced id in the same servable
namespace, declared dependencies, and a two-level registry. No new concepts;
mirrors how cards are registered (and vllm-omni's pipeline_registry).

- program/workflow.py: Workflow gains `requires` (the cards it composes, derived
  from stages) and `validate(engine)` (fail-fast if a required card is absent,
  P7). New WorkflowRegistry: declarative workflow_id -> builder catalog for
  out-of-tree/ad hoc use.
- runtime/engine.py: `_workflows` registry + register_workflow (validates deps,
  rejects id collision with a model_id) + serves(); engine.run routes a request
  whose model_id is a workflow to workflow.run — addressable exactly like a model.
  Single-model hot path untouched.
- runtime/async_engine.py + serving/server.py: serves() and /models include
  workflows (discoverable as servables).
- models/__init__.py: declarative `_WORKFLOWS` catalog (the cross-model analog of
  _BUILDERS) + register_workflows() helper; build_image_video_engine now registers
  the workflow too. Adding a custom pipeline = one catalog line.
- Naming convention: dotted/namespaced workflow_id (`image_video.t2i_i2v`),
  distinct from kebab model ids, collision-checked. Renamed from `t2i_then_i2v`.
- tests (+5, 12 total in the file): addressable by id, requires/validate, id
  collision, registry catalog, register_workflows helper. Full suite 114 passed.
- designv4 §9.6: the naming & registration convention documented.
2026-06-18 09:41:13 +00:00
SolitaryThinker cfc2ad7487 [feat] Cross-model T2I→I2V workflow + N-way joint RL over arbitrary experts
Two more pipelines stress-testing the design, plus a BAGEL-placement note in
designv4. Both fit with no new runtime primitive (109 tests pass).

Pipeline 1 — cross-model T2I→I2V (program/workflow.py, models/image_video/):
- Realizes ProgramKind.WORKFLOW as a thin multi-instance orchestrator ABOVE the
  engine (the hot path stays single-instance). A Program composes one model's
  loops; a Workflow chains full engine.run calls across distinct cards, threading
  artifacts. (LTX-2 already covers same-card multi-stage; cross-model — FLUX→Wan
  — is the new capability the single-instance runner can't express.)
- flux-t2i (text→image) and wan-i2v (text+image→video) cards; the I2V program
  folds the conditioning image into text_embeds so WanDenoiseLoop is unchanged.
- Each model keeps its own interleave-parity guarantee (crossing instances is a
  Workflow boundary, not a loop step). 7 tests incl. video-depends-on-image.

Pipeline 2 — N-way joint RL (training/methods/joint_multi_rl.py, models/multi_expert/):
- JointMultiExpertRL generalizes UnifiedRLMethod (N=2) to N refiner LMs + a
  generator: one reward → one group advantage → N token-PG updates + 1 FlowGRPO
  PPO update, N+1 independent WeightSyncPlans. Proves the substrate was already
  N-ready (card holds N components/loops; per-component weight-sync; dict grad
  targets) — only the method body looped over two; now it loops over a list.
- credit="per_expert" learns all N cleanly; credit="shared" (faithful to UniRL)
  works but is noisier — the honest multi-agent credit-assignment result, a
  reward-shaping choice, not a substrate limit. 6 tests (N=1,3,4; prompt-only).

Fix — flow_sde_ml_velocity (loop/sampler.py): the toy FlowGRPO generator update
targeted the velocity the model already produced (a no-op once guidance_scale=1
was set for the C2 identity; the unified generator moved only on ~1e-7 noise).
The correct PG surrogate targets the max-likelihood velocity of the realized
sample — nonzero at ratio==1. Both UniRL and N-way generators now learn for real;
the C2 ratio==1 identity still holds (measured before the update).

BAGEL: MoT/shared-weight (one transformer on both loops), same row as Cosmos3;
real BAGEL's co-resident experts are expressible via the expert-routing policy
(partial sharing) — captured in designv4 §2.3.
2026-06-18 09:41:13 +00:00
SolitaryThinker 97fedaf432 [feat] Add Qwen-Omni thinker→talker→vocoder model (3 experts, 3 loops)
Ports vllm-omni's canonical qwen2_5_omni omni-speech cascade as a v2 card:
a third weight-sharing topology — three disjoint experts (thinker, talker,
vocoder) on three loop types (ar_decode → ar_decode → audio_decode) in one
request, with chained cross-stage conditioning and streaming codec→waveform.
vllm-omni runs these as three opaque request-scheduled stages; v2 makes every
thinker token, talker token, and vocoder chunk a runtime-visible WorkUnit.

- models/backend.py: ToyTalker (a genuinely distinct AR expert, weight-salted)
  + ToyVocoder (streaming code2wav: codec tokens → waveform chunks).
- models/omni/vocoder_loop.py: VocoderLoop filling the pre-declared
  LoopKind.AUDIO_DECODE / WorkUnitKind.AUDIO_CHUNK slot.
- models/omni/ar_loop.py: ARDecodeLoop gains a configurable prompt_slot so two
  chained AR loops don't collide on the prefill slot (thinker vs talker).
- models/qwen_omni/: card (3 experts/3 loops) + program (tokenize → thinker →
  emit_text → thinker→talker full-payload hand-off → talker → talker→vocoder →
  vocoder → emit_audio). Cross-stage hand-offs are explicit Program nodes, the
  model-native form of vllm-omni's custom_process_input_func.
- _enums.py: Capability.TEXT_TO_SPEECH.
- tests/test_thinker_talker.py: 6 tests incl. three-loop interleave parity,
  cascade conditioning, AUDIO_CHUNK streaming. Full suite 97 passed.

designv4.md: §2.3 topology table extended to four topologies; new §9.5 on the
cascade; reference-synthesis + package layout updated.
2026-06-18 09:41:13 +00:00
SolitaryThinker 3ec511419f [feat] UniRL/PromptRL joint LM+generator RL stress test + designv4
Stress-tests the v2 Card/Loop/Program design with a UniRL/PromptRL-style
joint RL recipe: a prompt-refiner LM expert and a flow generator expert,
two separate experts driven by two loop types in one request, both updated
simultaneously from a single RL reward.

- loop/sampler.py: flow_sde_step_with_logprob — FlowGRPO SDE rollout sampler
  (per-step Gaussian log-prob), distinct from the deterministic ODE serve step.
- request/params.py: gated sde_rollout/sde_noise_scale on DiffusionParams so
  the serve path stays byte-identical (default ODE).
- models/wan21/loop.py: gated SDE-rollout capture in WanDenoiseLoop.advance.
- models/backend.py: ToyPromptRefiner — a real REINFORCE categorical policy
  (the Qwen role), separate weights from the generator.
- models/unified/: the unified card+program — two disjoint experts (llm +
  transformer) on ar_decode + diffusion_denoise; the topological opposite of
  the Cosmos3 MoT card, same vocabulary.
- training/methods/unified_rl.py: joint GRPO — one reward -> group advantage
  -> LM token policy gradient + DiT FlowGRPO PPO; two LRs; prompt-only/joint
  flag; reuses the shared diffusion loop for rollout.
- training/weight_sync.py: WeightSyncPlan gains a component scope so the two
  experts version + cache-invalidate independently (LM sync never flushes the
  frozen text-encoder feature cache).
- tests/test_unified_rl.py: 9 tests incl. likelihood-based C2 identity, the
  two-loop interleave parity gate, joint vs prompt-only. Full suite 91 passed.

designv4.md: unified design doc reflecting v2 as built+tested, with the joint
RL stress test as the validating case study (the design held — new card +
new method, no new runtime primitive).
2026-06-18 09:41:13 +00:00
SolitaryThinker 145082cf74 [refactor] rename package mini_fastvideo → v2
Directory rename (git mv, history preserved) plus rewrite of all references: absolute imports in
tests, the zero-dep runner, docstrings, comments, and the README. No behavior change.

Run: python3 -m pytest v2/tests/ -q ; python3 v2/run_tests.py ; python3 -m v2.examples
2026-06-18 09:41:12 +00:00
SolitaryThinker 71caa58ba9 [fix] mini-fastvideo serving: address adversarial-review findings (capacity/credit leaks, robustness)
Review confirmed the core bets (concurrent disaggregation is bit-identical, design conformance holds,
Dynamo genuinely optional, engine stays step-scheduled). Fixes for the untested failure paths:

- HIGH: pool capacity (RolePool.in_flight) no longer leaks when a disaggregated request is cancelled
  or errors mid-occupancy — DisaggregatedRunner.close() releases the occupied pool and AsyncEngine._run
  calls it in a finally (a cancel on a capacity-1 denoiser no longer bricks the pool).
- credit flow-control: cross-pool transfer wraps acquire/release in try/finally (no credit leak on a
  failing transfer); slot is re-homed only on a successful fetch.
- AsyncEngine: duplicate in-flight request_id is rejected (was a deadlock); submit() cancels the driver
  task when the consumer abandons the stream (client disconnect → no orphaned compute); bounded
  per-request history (no unbounded _events/_states/_results/_runners growth).
- cancellation is common-path on the OFFLINE path too (cancel check at the top of every runner.tick()).
- HTTP server: read timeout (slowloris guard → 408), body-size cap (→ 413), invalid Content-Length
  (→ 400), explicit StreamReader limit, and aclose() of the SSE generator on client disconnect.
- build_deployment_card no longer aliases one mutable CostModel across replica cards (dataclasses.replace),
  so online calibration of one worker's cost doesn't mutate another's.
- video-job tasks tracked (not fire-and-forget); server.close() cancels/drains them; jobs dict bounded;
  fleet affinity map bounded.

6 regression tests added for these paths. 82 tests pass (pytest + zero-dep runner).
2026-06-18 09:41:12 +00:00
SolitaryThinker 8240970ea6 [feat] mini-fastvideo serving + fleet (our own version, Dynamo-optional)
Builds the full serving layer the design files specify, instead of deferring it to Dynamo:

- transport/ (§7.3): pluggable Connectors (in-proc zero-copy / SHM-fake copy) with chunk_ready
  readiness (vllm-omni) AND credit-based flow control (sglang-omni Relay); KVConnector protocol shape;
  TransferManifest.
- runtime/ (§6, §13; plan M3/M4): AsyncEngine — request queue, lifecycle state machine
  (waiting→running→completed/cancelled/failed), live AsyncIterator[OmniEvent] streaming, common-path
  cancellation, step-level concurrency. RolePool + DisaggregatedRunner (encoder→denoiser→decoder,
  capacity-aware dispatch, cross-pool transfers via connectors); disaggregated output is bit-identical
  to inline. No-progress detection (no busy-spin).
- deploy/ (§14, §6.3.5-6): DeploymentCard; OUR OWN LocalFleet (discovery, health/drain, least-loaded
  / cost-model / sticky-affinity routing) so we never rely on Dynamo; DynamoWorkerAdapter +
  FakeDynamoRuntime export the SAME card + cost model so Dynamo CAN front us — one object, two consumers.
- serving/ (§6.3.5, §12): framework-free stdlib-asyncio OpenAI server (our own version of the
  vllm-omni pattern): /v1/chat/completions (SSE), /v1/images/generations, /v1/videos (async job+poll)
  + /v1/videos/sync, /v1/models, /health, /metrics. A thin shim over the STEP-scheduled engine — the
  runtime-visible loop scheduler vllm-omni's request-scheduled opaque DIFFUSION stage lacks.

15 serving tests (real-socket HTTP+SSE via stdlib asyncio, disagg==inline, fleet routing, Dynamo
contract, cancellation); 76 tests pass total (pytest + zero-dep runner). ~7900 LOC.
2026-06-18 09:41:12 +00:00
SolitaryThinker 44d10b5c79 [feat] mini-fastvideo phase 2: omni/MoT — Cosmos3 + canonical vllm-omni (BAGEL/lance)
One resident MoT instance runs BOTH an ar_decode loop and a diffusion_denoise loop on shared weights
(the §16 claim no DAG-of-engines can express), with both loops runtime-visible: the scheduler prices
ar_token AND diffusion_step WorkUnits — unlike vllm-omni's opaque DIFFUSION stage the scheduler never
sees inside.

- ARDecodeLoop: token decode until EOS/max_tokens, paged text-KV — the omni AR pathway (loop/§5).
- ToyMoTDiT: one module exposing an und pathway (ar_forward) AND a gen pathway (denoise __call__);
  ToyTokenizer. Binding both loops to one instance = shared weights, no duplication.
- models/cosmos3/: tokenize → reason(ar_decode) → pack(tokens→conditioning) → diffusion_denoise →
  vae_decode; sound_vae declared optional_for non-t2vs (the lazy-component P8 fix, not an env-var hack).
- models/bagel/: the canonical vllm-omni model — generate_text(ar_decode) → generate_image(diffusion),
  text+image outputs, both loops step-scheduled.
- The diffusion loop is WanDenoiseLoop reused (one loop definition bound to the MoT module).
- build_omni_engine() + an omni worked example; 7 omni tests (shared-instance, both-kinds-scheduled,
  interleave parity across loop types, lazy sound_vae). 61 tests pass (pytest + zero-dep runner).
2026-06-18 09:41:12 +00:00
SolitaryThinker 9d874aa609 [fix] mini-fastvideo: address adversarial-review findings (admission liveness + §7.1 cache key)
- Admission fails fast with AdmissionInfeasible on infeasible/deadlocked reservations instead of a
  10M-iteration busy-spin: no-progress detection in run_to_completion/run_interleaved via a real
  progress token, plus feasibility pre-checks (need > pool capacity).
- Compute budget is now a refundable concurrency gate (release() refunds spent), not a
  never-refunded lifetime cap that silently deadlocks.
- §7.1: the text-encoder feature CacheKey carries adapter_versions + precision (no stale serve across
  te-LoRA stacks); per-component weight versions mean a transformer-only RL weight sync no longer
  flushes the frozen text-encoder cache (component-scoped invalidation, not wholesale).
- Interleave gate flags symmetric-empty output instead of passing it vacuously.
- skipped_steps counted only when the override is actually consumed; BatchScheduler wired for round
  batch-accounting (metric renamed stepped_units); ResidualCache.get cleanup; stream chunks carry a
  latent preview payload; dead progress-vars removed.
- 5 regression tests added for the previously-untested paths. 54 tests pass (pytest + zero-dep runner).
2026-06-18 09:41:11 +00:00
SolitaryThinker 8e828680c6 [feat] mini-fastvideo: model-native runtime per design_v3 (Wan2.1/LTX2 + 4 training methods)
A scoped, CPU-testable realization of design_v3.md — the architecture where the atomic unit
is a typed (recipe, runtime) ModelCard, the model owns loop semantics while the runtime owns
loop lifecycle, one resident instance runs many loops, and training records behavior on the
same loops it serves.

Implements:
- card/ loop/ runtime/ cache/ memory/ parallel/ parity/ extend/ program/ request/ training/
  spanning design_v3 §4-§13: ModelCard + validate(); driven loops (init/next/advance/finalize);
  step-interleaving Engine with reservation-before-admission + per-class caches keyed by CacheKey;
  the C0-C4 consistency ladder + the non-negotiable batch-of-N interleave parity gate.
- Inference: Wan2.1-1.3B (T2V), LTX2.3 (two-stage distilled, shared transformer), Wan-causal
  (chunk rollout + slab-KV streaming).
- Training (Wan2.1-1.3B), each driving the SAME loops the engine serves: finetune (flow-match),
  DMD2 (teacher/critic distribution matching), DiffusionNFT (likelihood-free C2, samples from the
  decay-blended old policy, group-relative advantages, shared-prompt cache reuse), self-forcing
  (causal chunk loop). The engine never imports training (the §10 dependency rule, grep-verified).

numpy-only core (no torch/GPU here); heavy Wan/LTX forwards are deterministic toy stand-ins with
lazy torch-adapter seams (ComponentSpec.load_id/factory) for a GPU box. 49 tests pass via pytest
and a zero-dependency runner; interleave parity verified (and a buggy module-global interceptor
provably breaks it). Omni-ready spine (ar_decode/chunk_step loop kinds, multi-loop instances,
LoopState.extension) for the phase-2 Cosmos3 + vllm-omni omni ports.

Run: python3 -m pytest mini_fastvideo/tests/ -q ; python3 -m mini_fastvideo.examples
2026-06-18 09:41:11 +00:00
SolitaryThinker 735ee6aa56 update 2026-06-18 09:41:11 +00:00
SolitaryThinker 068a532b31 design 2026-06-18 09:41:11 +00:00
sumyyyyy 87f98c9b8b [kernel] Add varlen support for block-sparse attention (#1319) 2026-06-18 01:06:00 +00:00
e60601df7f [feat] QAD 5090: QAT training recipe — finetune + DMD distillation (12/12) (#1462)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 15:18:52 -07:00
eed9c4bfbf [kernel] QAD 5090: Add Attn-QAT training Triton kernels (11/12) (#1460)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
2026-06-17 13:35:12 -07:00
1dee77f4a4 [feat] QAD 5090: Wire the Attn-QAT training attention backend (10/12) (#1459)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: alexzms <26690162+alexzms@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 13:32:56 -07:00
Satyam Srivastava b80148819c [ci] Add performance dashboard visualisation scripts (#1469) 2026-06-16 23:29:29 -07:00
c3b971488e [docs] QAD 5090: Add NVFP4 + Attn-QAT inference example and how-to (9/12) (#1458)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: alexzms <26690162+alexzms@users.noreply.github.com>
2026-06-16 14:36:50 -07:00
88e753f281 [feat] QAD 5090: Wire the Attn-QAT inference attention backend (8/12) (#1457)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: alexzms <26690162+alexzms@users.noreply.github.com>
2026-06-16 13:12:33 -07:00
77832059cc [kernel] QAD 5090: Add modified SageAttention3 FP4 inference kernels (7/12) (#1455)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: Edenzzzz <wtan45@wisc.edu>
2026-06-16 12:02:30 -07:00
alexzmsandmergify[bot] 633d393568 [ci] layer-0 grad-norm regression for per-method training tests (5a-ii) (#1396)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 04:45:07 +00:00
Junda SuandPeiyuan Zhang 5854aec2ce [feat] Add Wan RL DiffusionNFT training (#1450)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2026-06-11 21:18:59 -07:00
Mook 30e45c2411 [bugfix] Classify new config/sampling fields in schema parity inventory (#1446) 2026-06-10 13:38:36 -07:00
alexzms 2a4fe697a6 [docs] LTX-2.3 distilled i2v: typed-API example (from_config + generate) (#1448) 2026-06-10 10:58:14 -07:00
alexzms 921db7479d [perf] LTX-2.3 distilled i2v: drop max-autotune from compile kwargs (#1445) 2026-06-10 10:06:41 -07:00
Aryan KumarandAryan Kumar 7f539424cb [feat]: add Lucy Edit inference scaffold (#1363)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-06-09 15:21:43 -07:00
19a838f54f [bugfix]: release VSA tile cache during training (#1434)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-09 15:01:22 -07:00
d922ab2cbc [model] Flux2 Klein Port (#1349)
Co-authored-by: Gnav3852 <63612880+Gnav3852@users.noreply.github.com>
Co-authored-by: Mac Lee <macthecadillac@gmail.com>
2026-06-09 14:55:55 -07:00
Kaiqin Kongandmergify[bot] 9ea77d37f3 [bugfix] EMA shadow on resume and EMA under MoE path (#1441)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-09 00:55:43 +00:00
2e35b0c6bd [refactor]: linear/mlp FP4 path additions for Wan-2.1 (Attn-QAT 6/12) (#1390)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
Co-authored-by: Matthew Noto <notomatthew31@gmail.com>
2026-06-08 14:51:57 -07:00
William Lin 1c627a3f98 [bugfix]: build fastvideo-kernel on GPU-less Docker runners (#1437) 2026-06-06 17:52:45 -07:00
Kaiqin Kong a931efe33a [bugfix] EMA in distillation pipeline (#1440) 2026-06-06 17:38:03 -07:00
William Lin 041e5e9029 [bugfix]: unblock PyPI publish (flash-attn-cute direct dep) (#1436) 2026-06-05 10:07:12 -07:00
Kaiqin Kongandmergify[bot] efcc245c2e [feat] VLM as judge for WM (#1429)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-05 04:25:54 +00:00
alexzmsandmergify[bot] 922e7e0813 [docs] LTX-2.3 distilled i2v example with compile + timing breakdown (#1430)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-05 01:26:50 +00:00
William Lin c62a8514b0 [chore]: release v0.2.0 (#1432) 2026-06-04 14:22:49 -07:00
William Lin 3eb8081801 [chore]: unpin runtime deps in pyproject.toml (#1431) 2026-06-04 13:46:49 -07:00
Raghav K 3505d09564 [bugfix] tests: include ltx2_3_base in expected LTX2 preset set (#1427) (#1428) 2026-06-04 12:17:18 -07:00
Shao DuanandSolitaryThinker 570607945c [feat] dreamverse: sequence parallelism for serving (#1424)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-01 23:13:49 -07:00
Kaiqin KongandSolitaryThinker d3a821cdcf [feat] LoRA controls and integration for Dreamverse (#1420)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-01 18:54:11 -07:00
Kaiqin Kong 3f24578139 [bugfix] LTX2: honor video_position_offset_sec in the DiT (#1422) 2026-06-01 16:59:15 -07:00
Kevin Lin 89fcf08378 [bugfix] Fix STFT dtype mismatch (#1419) 2026-05-31 20:56:07 -07:00
c2930b2aa1 [ci] Add additional Dreamverse UI tests (#1417)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-31 19:53:29 -07:00
KUAN-HAO HUANGandSolitaryThinker 019239690b [perf] Add Adaptive Guidance (CFG gating) for stale-uncond reuse (#1372)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-30 13:58:55 -07:00
alexzms d6119c1f82 [bugfix]: dreamverse modal bypasses ENTRYPOINT — set ffmpeg env + key check (#1413) 2026-05-29 16:50:01 -07:00
alexzms 84214c80bb [feat] LTX-2.3 audio: BWE vocoder path (#1398) 2026-05-29 16:43:26 -07:00
alexzmsandmergify[bot] c2d7143c72 [feat] LTX-2.3 transformer support (config-gated extension of LTX-2) (#1397)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-29 16:42:37 -07:00
Kaiqin KongandSolitaryThinker afdb6fbfa5 [feat] Add MatrixGame3.0 (#1201)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-27 21:07:01 +00:00
alexzms ba4c02d883 [docs]: highlight Dreamverse deployment paths + add Server B200 (SSH) guide (#1409) 2026-05-27 09:30:58 -07:00
Junda Su 2c137931f3 [bugfix] Fix Dreamverse Modal compile warmup latency (#1394) 2026-05-26 17:58:38 -07:00
Shao Duan 36682797a0 [feat] eval: input ergonomics + Evaluator features + bug fixes (#1392) 2026-05-26 13:59:45 -07:00
William Lin 0ef1357a77 [docs]: surface activation-trace utility in add-model skills (#1399) 2026-05-26 13:45:12 -07:00
William Lin a75d19786a [docs]: Wire activation trace into mkdocs nav + perf/troubleshooting (#1304) 2026-05-26 13:14:36 -07:00
alexzms be548a78ea [feat] VSA-256 fastpath on Blackwell via FA4 CuTe block-sparse attention (#1354) 2026-05-26 12:58:29 -07:00
alexzms 6a610e2bc9 [ci] add per-method single-step training tests for fastvideo.train (#1343) 2026-05-25 17:09:54 -07:00
Shao Duanandabaghyangor 321d5112b4 [refactor] eval: consolidate FVD into common.fvd, remove benchmarks/fvd (#1380)
Co-authored-by: abaghyangor <abaghyangor@gmail.com>
2026-05-24 12:09:47 -07:00
William Lin ba75ad82db [refactor]: shared attention infra additions for QAT-compat (Attn-QAT 5/12) (#1383) 2026-05-23 16:19:24 -07:00
Junda Su 58caa5109f [ci] Add DreamVerse app CI tests (#1386) 2026-05-23 14:28:26 -07:00
2f3ca8aaad [perf]: register FA2/FA3 default flash_attn_func as a torch.library custom op (#1373)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-23 00:39:06 -07:00
Junda Su 3a67319cb6 [infra] Use npm for Dreamverse web builds (#1385) 2026-05-22 18:50:10 -07:00
Junda Su 266fa044b3 [infra] Add Dreamverse Modal UI image build (#1381) 2026-05-22 11:36:53 -07:00
William Lin f3398db868 chore: pin dreamverse npm deps to address Dependabot alerts (#1359) 2026-05-22 11:29:33 -07:00
fda02036bc [feat]: Attn-QAT inference + training backends (deadcode) (Attn-QAT 4/12) (#1358)
Co-authored-by: jzhang38 <42993249+jzhang38@users.noreply.github.com>
Co-authored-by: RandNMR73 <99706358+RandNMR73@users.noreply.github.com>
2026-05-22 02:59:25 -07:00
Raghav K 68179cd752 [docs] Document enable_torch_compile (+ A/B example) (#1366) 2026-05-22 00:25:42 -07:00
Satyam SrivastavaandSolitaryThinker 2dd5760291 [docs] Document performance benchmark workflow (#1376)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-21 20:25:58 -07:00
Wenxuan TanandSolitaryThinker af2ee9c78a [feat] Optimize distributed weight loading in multi-node training (#572)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-21 16:59:48 +00:00
Satyam SrivastavaandSatyam Srivastava 1c80371b27 [ci] Component time performance + reseed hf baseline skill (#1292)
Co-authored-by: Satyam Srivastava <satyam53@Satyams-MacBook-Air.local>
2026-05-20 13:51:11 -07:00
Junda Su eef473225d [ci] Add Dreamverse Docker image workflow (#1369) 2026-05-20 13:50:02 -07:00
Junda Su 44fb84ef6a [bugfix]: shrink Dreamverse Docker context (#1368) 2026-05-20 13:38:52 -07:00
e8597b7448 [Bugfix] FP4 FA4 installation fix (#1367)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-19 22:45:24 -07:00
Raghav Kandmergify[bot] e2252c0a5e [perf] Mark LayerwiseOffloadHook entry points torch.compiler.disable (remove per-layer graph break) (#1365)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-20 03:09:23 +00:00
Junda Su 72cb427cd9 [feat]: add FastLTX-2.3 Gradio demo package (draft) (#1247) 2026-05-17 23:03:16 -07:00
Mingjia Huo 63030cf6ec [fix] Fix causal self-forcing attention settings (#1355) 2026-05-17 22:59:47 -07:00
Kaiqin Kongandmergify[bot] 773d44b875 [misc] Rename MatrixGame to MatrixGame2 (#1357)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-17 22:57:03 -07:00
1df513922f [feat] Add minimal LoRA finetuning support to the YAML training stack (#1242)
Co-authored-by: alexzms <3036648523@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-17 14:49:13 -07:00
William LinandDavids048 30c45620a2 [infra] [dreamverse]: add instruction to install nasm and update ffmpeg installer to work in plain venv (#1361)
Co-authored-by: Davids048 <jundasu@ucsd.edu>
2026-05-17 01:10:55 +00:00
Shao Duanandklhhhhh 6b2c731596 [feat] eval: add audio metrics (#1352)
Co-authored-by: klhhhhh <1412841649@qq.com>
2026-05-16 14:47:37 -07:00
William Lin e6022c20b2 [misc]: demote ROCm-unavailable startup message to DEBUG (#1360) 2026-05-15 22:12:16 -07:00
460f6e398e [feat]: Add NVFP4QAT linear layer (Attn-QAT 3/12) (#1350)
Co-authored-by: jzhang38 <42993249+jzhang38@users.noreply.github.com>
Co-authored-by: RandNMR73 <99706358+RandNMR73@users.noreply.github.com>
2026-05-15 18:42:02 -07:00
d2ffec5cce [perf] Dreamverse 14/14: Add LTX2 profile speedups (#1337)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-15 18:10:47 -07:00
alexzmsandmergify[bot] cb12e88713 [perf] shallow-copy VSA attn_metadata in train model plugins (#1342)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-15 16:10:12 -07:00
0958c344b8 [feat] Dreamverse 13/14: Activate LTX2 integration (#1336)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-15 15:35:05 -07:00
Junda Su 71dc27ea7b [misc]: Add Dreamverse deploy skill frontmatter (#1353) 2026-05-15 15:23:17 -07:00
Junda SuandSolitaryThinker 1263449d2a [docs] Add copy page action (#1351)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-15 14:59:04 -07:00
b4de5a9f1b [feat] Dreamverse 12/14: Add LTX2 refine and upsampler support (#1335)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-15 12:21:40 -07:00
8acd8e21f9 [feat]: Add NVFP4QAT quantization config (Attn-QAT 2/12) (#1348)
Co-authored-by: jzhang38 <42993249+jzhang38@users.noreply.github.com>
Co-authored-by: RandNMR73 <99706358+RandNMR73@users.noreply.github.com>
2026-05-14 16:53:48 -07:00
d45d82334f [infra] Dreamverse 11/14: Add NVFP4 quantization support (#1334)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-14 16:26:39 -07:00
6392bd40a9 [feat] Dreamverse 10/14: Add serving API contracts (#1333)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-14 10:07:01 -07:00
Shao Duan 17f07bc313 [feat] eval: async VideoPool + metric streamlines (#1320) 2026-05-13 15:57:37 -07:00
William Lin 325861fb99 [misc]: PR-1225 sync — housekeeping (1/12) (#1347) 2026-05-13 15:21:52 -07:00
William Lin c5088670c8 [misc]: empty __init__.py files with no logic (#1346) 2026-05-13 13:09:40 -07:00
900 changed files with 103104 additions and 20599 deletions
+207
View File
@@ -0,0 +1,207 @@
# v2 ← M\*: Architecture Gap-Analysis & Improvement Roadmap
**Status:** exploration, flagged for review. **Date:** 2026-06-19.
**Source paper:** *M\*: A Modular, Extensible, Serving System for Multimodal Models* (arXiv 2606.12688,
Stanford/UW/CMU; Jha, Sagan, Kamahori, …, Kasikci, S. Wang). It is a universal serving runtime for composite
multimodal models built on the **Walk Graph** abstraction (a model is a dataflow graph `G`; a request is a
*Walk* — a labeled subgraph — and the runtime executes walks). It beats vLLM-Omni (~20% lower T2I latency on
**BAGEL**, up to 2.64× on I2I), SGLang-Omni (2.7× TTS throughput on **Qwen3-Omni**), and native V-JEPA2
rollout (12.5×). It explicitly names **FastVideo's own** sparse/sliding-tile attention, xDiT/PipeFusion/USP,
Inferix, and FlashDrive as techniques integratable into the graph runtime.
**Method:** a 28-agent workflow — 6 parallel v2-subsystem maps → 10 M\*-dimension analyses, each
*adversarially verified against the actual v2 code* → synthesis + a completeness critic. The critic's
corrections and three P0 claims were then **spot-verified by hand** (file:line below). This doc folds those
corrections in; it is the corrected, authoritative synthesis.
---
## 1. Executive summary
v2 already implements the **harder half** of M\*'s thesis and in several axes **exceeds** it:
- v2's `Program` *is* M\*'s graph `G` (typed `ComponentNode`/`ModelLoopNode` + edges).
- v2's `shared_weight_components` *is* M\*'s cross-Walk node sharing — BAGEL/Cosmos3/LTX2 each bind two
`ModelLoopNode`s to **one resident transformer** (`instance.component()` returns the same live object). This
is the exact MoT serving property the omni cards in this repo already express.
- v2 adds three things M\* (serving-only) has **no equivalent for**: a required+validated per-loop **cost
model**, a non-negotiable **interleave bit-parity gate**, and an **integrated training plane** (RL→distill
flywheel driving the *same* serving Loop).
- The `extend/` plugin seam (interceptors/observers/registry with capability negotiation) is precisely the
hook M\*'s "extensible / integrate FastVideo-STA, xDiT, Inferix, FlashDrive" call-out asks for — **v2
already has the seam M\* only gestures at.**
What v2 lacks is M\*'s **declarative authoring layer above the substrate**, and — the key insight — *much of
that substrate is already authored but inert*: v2 has declared the metadata for "minimum components per
request" (`required_for`/`optional_for` on every omni card) and "branch as a cache axis" (`guidance_sig`,
`CacheKey`) but **never wired it to an executor**. The substrate is ~80% built and switched off.
**Highest-leverage cluster:** three small, parity-safe wires that turn on inert substrate and unblock the
BAGEL/Qwen-Omni/Cosmos3 latency wins M\* measured **on the exact models this repo already runs** — plus one
P1 that aligns v2 with the paper's headline "extensible" claim using a seam v2 already has.
### Verified P0 correctness findings (spot-checked by hand)
1. **Runner divergence (real bug).** `v2/runtime/engine.py:88` → `nodes = self.program.nodes`;
`v2/runtime/disaggregated.py:96` → `nodes = self.program.active_nodes(self.request)`. The inline and
disaggregated runners execute *different node sets*. ✅ confirmed.
2. **EOS is faked.** `v2/recipes/omni/ar_loop.py` docstring says "done on EOS/max_tokens"; `next()` (`:46-48`)
checks **only** `max_tokens`. M\*'s marquee `DynamicLoop` use case (EOS) is unimplemented in the loop that
serves the Qwen-Omni Thinker/Talker and Cosmos3 reasoner. ✅ confirmed.
3. **`required_for`/`optional_for` have zero runtime consumers** (grep outside `specs.py`/recipes/tests is
empty). The min-components metadata is declared on every card and never read. ✅ confirmed.
---
## 2. Dimension table (corrected)
| # | Dimension | v2 status | Gap | Priority | Effort | Payoff | Action |
|---|---|---|---|---|---|---|---|
| 1 | Min-components per request (`required_for` + `when_task`) | substrate built, **inert** | real, cheap | **P0** | S | Consume `required_for` in `active_nodes`; unify `engine.py:88` onto `active_nodes`; deliver via registry/card builder so all ~40 cards inherit it |
| 2 | Real EOS + declarative `DynamicLoop` | early-exit emergent; **EOS faked** | real | **P0** | S | `ARDecodeLoop` honors `eos_id` + `req.sampling.stop`; add `LoopSpec.dynamic_stop` + `register_loop_stop`. **Training-enabling** (world-model rollout horizon) |
| 3 | CFG/branch as label over one paged KV pool | absent (`PagedKVCache` is a counter) | real | **P1** | L | `(namespace,label)` paged store w/ one budget; reuse `guidance_sig` for hash (NOT `partition_field`); by-ref via existing `InProcKVConnector`. AR path only (diffusion has no KV) |
| 4 | `extend/` plugin seam → integrate FastVideo-STA / Inferix | **seam exists, unused for attn** | real (paper headline) | **P1** | M | Expose FastVideo sparse/sliding-tile attention + Inferix block-diffusion as `Interceptor`/`EngineKind` plugins — the paper's named integration targets, on this repo's own code |
| 5 | `ParitySpec.output_determinism` (C3 distributional) | C3 rung defined, **0 users** | real, dormant | **P1** | S | Add field; `compare_outputs` consults it. **Training-enabling** (SDE/FlowGRPO stochastic rollouts) |
| 6 | Registry-driven delivery of #1 | present, not leveraged | integration | **P1** | S | Express `when_task`/min-components through `WorkflowRegistry`/card builders, not 3 bespoke recipe patches |
| 7 | Serving conductor + pluggable data plane | conductor exists (`serving/http.py`); **single-process transport** | real | **P2** | L | v2 already has the step-scheduled worker surface; gap is ZeroMQ/Mooncake + direct worker→worker tensor routing (today `InProcKVConnector` only) |
| 8 | Fleet/Dynamo placement + replicas | **live** (`deploy/fleet.py`,`dynamo.py`) | partial | **P2** | M | Fleet-level placement/affinity/replica is real & ≥M\*; missing piece is only the intra-engine `(node,Walk)→rank` map decoupled from model code |
| 9 | Per-node TP / SP + cross-rank transport | axis vocab **exists** (`sp` incl.); not wired to runtime | partial | **P2** | XL | Wire declarative degrees into runtime; Wan/LTX are **SP-native** (TP is a no-op there); populate `parallel_plan_hash` on the serving cache path |
| 10 | Named Walks + per-model state machine | `Program`=G, sharing real; no Walk/SM | real | **P2** | M | Defer until a *re-entrant* phase graph (Thinker↔Talker, rollout) needs it; #1 captures the min-components win without it |
| 11 | Declarative `Parallel/Sequential/Loop` IR | imperative loop classes | real (authoring) | **P2** | M | Thin Section IR lowering to flat `Program`; scope to one AR recipe |
| 12 | Streaming `ChunkPolicy` + `StreamBuffer` | causal-chunk emit **already ships** (`wan_causal`); `EdgeKind.STREAM` inert | real | **P2** | L | Declarative `ChunkPolicy` vocab over the existing chunk mechanism; needs concurrent producer/consumer runner (= pipelined scheduling). Inferix integration point |
| 13 | Speculative deferred-termination; loop-spanning CUDA graphs; N+1 prefetch; attn double-buffer | absent / per-step capture (14 cards) | real | **P3** | L | Gate behind a real GPU executor; unobservable on CPU-toy CI; loop-span needs an `allows_interleaving=False` carve-out |
| — | Cost model + interleave/consistency parity | **exceeds M\*** | none | **guard** | — | Do not regress; keep `step_cost_model` mandatory + `bit_identical` default |
| — | Integrated training plane (flywheel, weight-sync) | **exceeds M\*** | none | **guard** | — | Protect train==serve loop identity with a toy fixture |
---
## 3. P0/P1 deep-dives (sequenced)
```
PR-1 (P0) min-components ──┐
PR-2 (P0) real EOS ─┼─► prereqs for honest "DynamicLoop" + min-component claims; both training-enabling
PR-3 (P1) output_determinism (independent)
PR-5 (P1) extend/ plugin: FastVideo-STA / Inferix as Interceptors (independent; highest paper-alignment)
PR-4 (P1) CFG-as-label paged pool ──► depends on PR-2 (AR loop is the only KV consumer)
```
PR-1, PR-2, PR-3, PR-5 are mutually independent; PR-4 depends on PR-2.
### PR-1 (P0) — Turn on the inert min-components substrate + fix runner divergence
- **Change.** Extend `Program.active_nodes(request)` (`v2/program/specs.py`) to also drop any node whose bound
`ComponentSpec.required_for` (`v2/card/specs.py:144`) excludes `request.task` (and isn't in `optional_for`).
**Fix the bug:** change `v2/runtime/engine.py:88` to `nodes = self.program.active_nodes(self.request)` so the
inline `ProgramRunner` matches `DisaggregatedRunner` (`disaggregated.py:96`). Deliver the `when_task` gating
through the **registry/card builder** (`recipes/__init__.py`, `program/workflow.py:WorkflowRegistry`) so all
~40 cards inherit it uniformly — not three bespoke `program.py` patches.
- **Why (this repo's models).** BAGEL T2I currently steps the AR-text loop and Cosmos3 t2v materializes the
reasoner even though the cards declare `transformer required_for={'reason','t2i'}`, `vae required_for={'t2i'}`.
On the GPU backend that is wasted resident-weight load + wasted steps on every single-modality request —
exactly M\*'s "execute the MINIMUM components per request," delivered by consuming existing metadata.
- **Risk/invariant.** Validate in `ModelCard.validate()` that every active node's `reads` are produced by an
active node for each declared `TaskType` (avoid dropping a producer). Pure node-id filtering ⇒ serial and
interleaved still walk the same filtered list ⇒ §9.3 interleave bit-parity holds by construction. CPU-toy clean.
### PR-2 (P0) — Real EOS + declarative `dynamic_stop` *(also training-enabling)*
- **Change.** In `v2/recipes/omni/ar_loop.py`, `advance()` reads the emitted token; if it equals the model
`eos_id` (toy backend exposes `EOS=0`) or matches `req.sampling.stop` (`params.py:21`, currently dead),
register termination; `next()` returns `Done()` on stop OR `max_tokens`. Add `StopRegistry` to `LoopState` +
`register_loop_stop(name)` to the `LoopContext` protocol (`contracts.py:204`) and to
`DisaggregatedRunner`'s `RuntimeLoopContext`. Add `LoopSpec.dynamic_stop: bool=False`, opt the AR cards in.
- **Why.** The docstring-vs-code lie sits in the loop serving Qwen-Omni Thinker/Talker and the Cosmos3 reasoner;
M\*'s second named `DynamicLoop` use case (world-model **rollout horizon**) is exactly what `self_forcing` RL
needs — so this is both a serving-credibility fix and a training enabler (raise its payoff accordingly).
- **Risk/invariant.** `dynamic_stop=False` is byte-identical back-compat. Must pass **all three** parity gates:
serial==interleaved AND disaggregated==inline. **Not** in this PR: speculative deferred-termination (unobservable
on CPU-toy, fights the interleave invariant — P3, gated on GPU executor).
### PR-3 (P1) — `ParitySpec.output_determinism` (close the dormant C3 hole) *(training-enabling)*
- **Change.** Add `output_determinism: str = "bit_identical"` to `ParitySpec` (`card/specs.py:88`); make
`compare_outputs` (`parity/interleave_gate.py:54`) consult it (`bit_identical` → today's exact check;
`distributional` → a moment/tolerance check — land a simple moment match first; a real KS test is new code).
- **Why.** `ConsistencyLevel.C3` is defined and used by zero recipes; an SDE/FlowGRPO stochastic rollout cannot
honestly declare its parity contract and would falsely fail the bit-identical gate. Additive; default unchanged.
### PR-5 (P1) — Expose FastVideo's own attention + Inferix as `extend/` plugins *(highest paper-alignment)*
- **Change.** Use the existing `extend/{interceptors,observers,registry}.py` seam (capability-negotiated, with
per-(request,branch) `plugin_state` that already passes the interleave gate) to register FastVideo's
sparse/sliding-tile attention and Inferix-style block-diffusion as `Interceptor`s / an `EngineKind` plugin.
- **Why.** M\*'s title is "Modular, **Extensible**" and it explicitly lists FastVideo-STA, xDiT/PipeFusion/USP,
Inferix, FlashDrive as integratable. v2 already has the seam M\* only describes — this is where v2 most
directly answers the paper, using this repo's own attention code. Low risk (the seam + capability negotiation
already exist and are tested).
### PR-4 (P1) — CFG/branch as a LABEL over one paged KV pool
- **Change.** Rewrite `PagedKVCache` (`cache/classes.py:155-172`) from a block *counter* into a real
`(namespace,label)->[block-handle]` store with **one shared `total_blocks` budget** (M\*'s single-pool
property). Reuse the existing-but-unpopulated `CacheKey.guidance_sig` (`keys.py:53`) for the hash. Thread the
label through `ar_loop.py` (alloc/append/get per `(request_id, branch)`; prefill once per shared-prefix label;
combine via `CFGPolicy.combine`). Wire `ResourceRequest.cache_blocks` (`contracts.py:64`, zero consumers) into
admission per (class,label).
- **Why.** The dossier-identified driver of M\*'s BAGEL win (3 CFG contexts as 3 labels over ONE pool vs dense
per-context). Targets AR_DECODE (BAGEL `generate_text`, omni Thinker); **correctly excludes diffusion**
(Wan/LTX are bidirectional, no KV — their CFG stays dense-but-batched).
- **Corrections to bake in.** Do **NOT** add `branch_label` to `CacheKey.partition_field()` (CFG branches share
embeddings; partitioning by branch is a semantic bug). Do **NOT** add a new by-ref type — reuse
`InProcKVConnector` + `TransferManifest.cache_key`. Wiring `cache_blocks` admission is greenfield ⇒ effort **L**.
CPU version proves label/sharing semantics; the real latency win needs a FlashInfer paged kernel (out of scope)
— **merge** with a future "real KVCacheEngine" effort rather than landing isolated.
---
## 4. What v2 already does ≥ M\* — do NOT regress
1. **Required+validated cost model** on every `LoopSpec` (13-kind `WorkUnitKind`) — typed, pre-GPU-validated.
2. **Interleave bit-parity as a hard gate** (`parity.interleave_required=True` on 40+ cards). M\* has no such
gate (its speculative scheduling deliberately wastes steps). Load-bearing invariant; every new primitive
must pass it.
3. **C0–C4 consistency ladder** wired into RL methods, with first-divergence tap reporting. No M\* equivalent.
4. **Integrated training plane** — DiffusionNFT/DMD2/self_forcing, RL→distill flywheel, `WeightSyncController`
hot weight-sync with drain-to-boundary + scoped cache invalidation, driving the **same** serving Loop.
M\* is serving-only. Protect with a toy fixture asserting `rollout_loop` drives the served Loop object.
5. **CPU-toy parity for the whole stack** — loops/CFG/caches/parity/RL run in CI without a GPU. Every new
primitive must ship a toy exercise (this is what makes all PRs above testable without H100s).
6. **Partition-not-flush cache invalidation** + four independent per-class pools.
7. **`extend/` plugin seam** with capability negotiation (a 4-step distilled card *rejects* a residual-skip
interceptor) — M\* describes extensibility; v2 has the mechanism.
8. **Dynamo citizenship** (`deploy/dynamo.py`: one `DeploymentCard`+cost model, two consumers) — beyond M\*'s
self-contained runtime.
---
## 5. Dropped / merged / deferred (and why)
- **DROP declarative `Parallel` as a CFG-execution win.** The runner walks nodes linearly (ignores
`Program.edges`), so `Parallel` lowers to sequential sugar and the CFG 3-pass braid is already one
co-scheduled `WorkPlan.run`; splitting it risks the interleave gate. Salvage only the no-op refactor
extracting `branch_forward` from `WanDenoiseLoop._velocity`. Reassign `Parallel` to the placement workstream.
- **MERGE the full Walk/state-machine layer** into "defer until a re-entrant phase graph needs it" (PR-1 gets the
min-components win with ~20 lines, no new abstraction). If built: the validator must check a walk's node-id
order is a *subsequence* of `program.nodes` (not just membership) or the runner can reorder and break parity.
- **MERGE `StreamBuffer`/`ChunkPolicy` into pipelined-scheduling.** Causal-chunk emit *already ships*
(`wan_causal/loop.py` per-chunk `StepResult.emit` + slab-KV); the gap is the declarative `ChunkPolicy` vocab
+ a concurrent producer/consumer runner. If built: keep all policies pure (per-request `StreamBuffer` history,
not shared edge state) and restrict the bit-identical claim to the token-only handoff.
- **MERGE CFG-fan-out exec + cross-rank transport + PD loop-splitting into a multi-GPU-runtime program.** These
need real collectives (`v2/distributed/` is a stub) and KV-by-reference (KV lives in `CacheManager`, not the
transferable `slots`). **Keep cheaply now:** the *declarative* halves — per-component degree, `(node,Walk)`
placement key with node-only fallback, `ReplicaSet` under `LocalFleet`, and populate `parallel_plan_hash` on
the **serving** cache path (it is already populated in `training/behavior.py:40` — the gap is serving-only).
- **DEFER** speculative deferred-termination, loop-spanning CUDA graphs, N+1 prefetch, attention-plan
double-buffer — all gated on a real GPU executor; benefit unobservable on CPU-toy CI. Keep the cheap
`EngineKind` tag (`STATELESS|KV_CACHE|DIFFUSION`) now. Correct the stale `cudagraph.py:51-52` docstring
(per-step capture ships in 14 cards, not just wan21).
- **RESCOPE per-node TP.** Wan/LTX use `ReplicatedLinear` + **sequence parallelism** (`sp`), not TP; the `sp`
axis already exists in `parallel/plan.py:AXIS_NAMES`. The work is wiring degrees into the runtime, not
inventing vocabulary; a `tp_size=2` "one-line activation" is a no-op for the shipped models.
---
## 6. The first integration test, if/when multi-GPU placement work starts
The **live Qwen-Omni 2-GPU bring-up** (Thinker on rank 0, Talker+Code2Wav on rank 1; see
`v2_debug_videos/vlm.md` Session 4) is the natural first validation target for any `(node,Walk)→rank`
placement work — it is the one place this repo already has real multi-rank composite-model execution.
---
## Anchor files for P0/P1
`v2/program/specs.py`, `v2/runtime/engine.py` (**line 88 fix**), `v2/runtime/disaggregated.py`,
`v2/recipes/omni/ar_loop.py`, `v2/loop/contracts.py`, `v2/card/specs.py`, `v2/cache/{classes.py,keys.py}`,
`v2/parity/interleave_gate.py`, `v2/extend/{interceptors,registry}.py`, `recipes/__init__.py` +
`v2/program/workflow.py` (registry-driven delivery).
@@ -0,0 +1,41 @@
---
date: 2026-05-22
experiment: PR #1386 DreamVerse app CI backend tests
category: infrastructure
severity: important
---
# DreamVerse App CI Streaming Imports Need GPU
## What Happened
DreamVerse app CI backend pytest collection imports FastVideo streaming surfaces.
When those tests run in a CPU-only Modal environment, collection can fail before
any app assertions run with Triton reporting:
```text
RuntimeError: 0 active drivers
```
## Root Cause
Some streaming import paths can import `fastvideo_kernel` at module import time.
Triton then probes for an active GPU driver during pytest collection. A CPU-only
Modal container has no active driver, so the failure appears as an import-time
collection error rather than a DreamVerse app behavior failure.
## Fix / Workaround
For PR #1386, use a surgical CI fix: allocate a GPU to
`run_dreamverse_app_tests`. Do not refactor core streaming/kernel imports just to
unstick this app CI path.
Keep `build_kernel=False` for this job. The DreamVerse app backend test imports
streaming surfaces but does not need to rebuild or exercise custom kernels.
## Prevention
When adding or modifying DreamVerse app CI jobs that import FastVideo streaming
modules, make the GPU requirement explicit if the import graph may touch
`fastvideo_kernel`. Prefer small CI resource fixes for app test collection issues
unless the product code genuinely requires lazy import cleanup.
+3 -3
View File
@@ -31,7 +31,7 @@ FastVideo-WorldModel/
│ │ │ └── distribution_matching/ # DMD2Method, SelfForcingMethod
│ │ ├── models/ # Per-role model wrappers (ModelBase, CausalModelBase)
│ │ │ ├── wan/ # WanModel, WanCausalModel
│ │ │ └── matrixgame/ # MatrixGameModel, MatrixGameCausalModel
│ │ │ └── matrixgame2/ # MatrixGame2Model, MatrixGame2CausalModel
│ │ ├── callbacks/ # Composable hooks (grad_clip, ema, validation)
│ │ └── utils/ # Config, builder, checkpoint, optimizer, tracking
│ ├── training/ # Legacy training infrastructure (being phased out)
@@ -44,7 +44,7 @@ FastVideo-WorldModel/
│ │ ├── wan_distillation_pipeline.py # Wan distillation
│ │ ├── self_forcing_distillation_pipeline.py # Self-forcing distill
│ │ ├── ltx2_training_pipeline.py # LTX-2 training
│ │ └── matrixgame_training_pipeline.py # MatrixGame training
│ │ └── matrixgame2_training_pipeline.py # Matrix-Game 2.0 training
│ ├── attention/ # Attention backends
│ ├── distributed/ # Sequence/tensor parallel utilities
│ ├── layers/ # Tensor-parallel layers
@@ -95,7 +95,7 @@ FastVideo-WorldModel/
| Wan distillation (DMD) | `fastvideo/training/wan_distillation_pipeline.py` | `torchrun --nproc_per_node N` |
| Self-forcing distill | `fastvideo/training/wan_self_forcing_distillation_pipeline.py` | `torchrun --nproc_per_node N` |
| LTX-2 finetune | `fastvideo/training/ltx2_training_pipeline.py` | `torchrun --nproc_per_node N` |
| MatrixGame | `fastvideo/training/matrixgame_training_pipeline.py` | `torchrun --nproc_per_node N` |
| Matrix-Game 2.0 | `fastvideo/training/matrixgame2_training_pipeline.py` | `torchrun --nproc_per_node N` |
## W&B Integration
@@ -8,6 +8,8 @@ follow-up actions see [open-threads.md](open-threads.md).
**Last updated:** 2026-05-06 (added D-21 — chunk-stutter root cause is software libx264 encoding consuming ~22% of segment wall-time, NOT a migration regression; verified by 3 parallel explore agents that NVFP4 + torch.compile coverage matches FastVideo-internal exactly; landed opt-in NVENC build path in install_native_ffmpeg.sh + `--nvenc`/`--no-nvenc` flag in dreamverse-deploy.sh + `apps/dreamverse/server/benchmarks/benchmark_av_streaming.py` regression test + memory dir update; default codec stays `libx264` for backward compat, opt-in via `--nvenc`. Earlier: added D-12 — GpuPool layer separation, Oracle review post-#1257-merge; added D-13 — prompt enhancer / LLMProvider abstraction shape, Oracle review pre-#1258-merge; added D-14 — streaming auxiliaries cohesion, Oracle review during #1284 review cycle; added D-15 — streaming router placement + sticky/active-active deferral, Oracle review during #1286 review cycle; added D-16 — streaming router polish round 2, second-pass review on top of D-15 covering bridge cancellation hygiene, registry state machine, httpx hard-fail, replica YAML parsing, and `websockets` dep; added D-17 — strategy reversal: abandon 6-PR split in favor of single mega-PR #1288 on `will/ltx2_sr_port`; added D-18 — Option B+ chosen: Dreamverse FE+product-server move into FastVideo as `apps/dreamverse/` subfolder while generic backend stays at `fastvideo.entrypoints.streaming.*`; integration-review.md deprecated, integration-plan.md is the executable migration plan; added D-19 — D-18 executed: 5 commits land on `will/dreamverse-monorepo`, fix-up commits corrected the integration-plan's invalid "delete generic-merged, import public substitutes" assumption — generic-merged files carried product-local instead, e2e passes against migrated code with /proc-verified evidence; added D-20 — segment-2 BrokenPipe root cause was a TWO-direction silent drop of LTX-2 audio kwargs in public `VideoGenerator`).
**Update 2026-05:** Dreamverse frontend tooling migrated from standalone pnpm to standalone npm. `apps/dreamverse/web/package-lock.json` is authoritative; see PR #1385.
## Status legend
- ✅ **Resolved** — decision made and implementation complete (or no implementation needed)
@@ -297,14 +299,14 @@ Playwright (8/8 PASS in 5.1s):
- **Python ML library** stays at root: `fastvideo/`, `fastvideo-kernel/`.
- **Generic backend** stays at `fastvideo.entrypoints.streaming.*` (already there per #1257/#1258/#1284/#1286/#1288).
- **Dreamverse product** moves into `apps/dreamverse/{server,web,prompts,serve_configs,scripts}/`.
- **Tooling**: uv workspace for Python (`[tool.uv.workspace] members = ["apps/dreamverse/server"]`), standalone pnpm for the FE (no root `package.json`), split CI workflows with path-filter triggers.
- **Tooling**: uv workspace for Python (`[tool.uv.workspace] members = ["apps/dreamverse/server"]`), standalone npm for the FE (no root `package.json`), split CI workflows with path-filter triggers.
**Rationale:**
- Drops the cross-repo coordination overhead identified in the post-#1286 rebase cycle (D-17 handled by consolidating into mega-PR; D-18 prevents the next round of cross-repo coordination from happening).
- Keeps the architectural separation Option D recommended (FastVideo owns reusable runtime; product owns product). The boundary is now `apps/dreamverse/` directory rather than two repos.
- Single repo means atomic cross-cutting refactors (e.g. GpuPool API change + Dreamverse adoption) ship as one PR.
- OSS precedents support the shape (chainlit uv-workspace + pnpm; open-webui Python + Svelte with paths-ignore CI). The librarian explicitly noted no precedent for "Python ML library + Next.js product merged into library namespace" — but this isn't that pattern. Dreamverse goes into a sibling directory, NOT into `fastvideo.entrypoints.dreamverse.*`. Library namespace stays clean.
- OSS precedents support the shape (chainlit uv-workspace + frontend package manager; open-webui Python + Svelte with paths-ignore CI). The librarian explicitly noted no precedent for "Python ML library + Next.js product merged into library namespace" — but this isn't that pattern. Dreamverse goes into a sibling directory, NOT into `fastvideo.entrypoints.dreamverse.*`. Library namespace stays clean.
**Why not Option D (separate repos):**
@@ -311,7 +311,7 @@ Notes:
development and CI.
- Product server release is Docker/deploy workflow, not PyPI.
### Frontend build: standalone pnpm
### Frontend build: standalone npm
Do not add a root `package.json`.
@@ -319,14 +319,14 @@ Keep all frontend tooling under:
```text
apps/dreamverse/web/package.json
apps/dreamverse/web/pnpm-lock.yaml
apps/dreamverse/web/package-lock.json
apps/dreamverse/web/playwright.config.*
```
Rationale:
- FastVideo remains primarily a Python ML library.
- Python contributors should not need Node or pnpm for normal work.
- Python contributors should not need Node or npm for normal work.
- This intentionally diverges from chainlit's root JS workspace pattern and
follows the simpler open-webui-style split.
@@ -455,31 +455,24 @@ jobs:
steps:
- uses: actions/checkout@v4
# IMPORTANT: pnpm/action-setup MUST run BEFORE setup-node when using
# cache: pnpm — setup-node otherwise can't find pnpm to populate cache.
- name: Setup pnpm
uses: pnpm/action-setup@v4
with:
version: 9
- name: Setup Node
uses: actions/setup-node@v4
with:
node-version: '22'
cache: pnpm
cache-dependency-path: apps/dreamverse/web/pnpm-lock.yaml
cache: npm
cache-dependency-path: apps/dreamverse/web/package-lock.json
- name: Install dependencies
run: pnpm install --frozen-lockfile
run: npm ci
- name: Typecheck
run: pnpm run typecheck --if-present
run: npm run typecheck --if-present
- name: Unit tests
run: pnpm run test --if-present
run: npm run test --if-present
- name: Build
run: pnpm run build
run: npm run build
# NOTE: Playwright tests require a running backend. Until Phase 4 lands
# `/healthz`/`/readyz`/`/status`/`/prompt-system-config`/`/curated-presets`
@@ -489,13 +482,13 @@ jobs:
# `fastvideo.entrypoints.streaming.build_app` and add product routes).
#
# - name: Install Playwright browsers
# run: pnpm exec playwright install --with-deps chromium
# run: npm exec -- playwright install --with-deps chromium
#
# - name: Playwright (re-enable in Phase 4)
# run: pnpm exec playwright test
# run: npm exec -- playwright test
```
**Playwright config update needed** when moving FE: `Dreamverse/apps/web/playwright.config.ts` line 39 currently uses `npm run dev`; change to `pnpm run dev` post-move ([source](file:///home/william5lin/Dreamverse/apps/web/playwright.config.ts#L39)).
**Playwright config update needed** when moving FE: keep `Dreamverse/apps/web/playwright.config.ts` line 39 on `npm run dev` after the move ([source](file:///home/william5lin/Dreamverse/apps/web/playwright.config.ts#L39)).
#### `.github/workflows/ci-dreamverse-backend.yml`
@@ -604,7 +597,7 @@ monorepo CI path.
| uv workspaces | https://docs.astral.sh/uv/concepts/workspaces/ | Authoritative Python workspace model. |
| Hatch monorepo | https://hatch.pypa.io/latest/how-to/environment/workspace/ | Alternative workspace model; not selected. |
| chainlit | https://github.com/Chainlit/chainlit | uv workspace precedent plus **per-language CI split** (separate `check-frontend.yaml` / `check-backend.yaml` workflows path-filtered by directory). Not a PR-level split — independent of D-17 single-mega-PR decision. |
| open-webui | https://github.com/open-webui/open-webui | Frontend path filtering (`paths-ignore` on backend-only changes) and separate release tracks. **Note:** open-webui has a root `package.json`; we are choosing standalone-pnpm despite the precedent, to avoid forcing Python-only contributors to install Node. |
| open-webui | https://github.com/open-webui/open-webui | Frontend path filtering (`paths-ignore` on backend-only changes) and separate release tracks. **Note:** open-webui has a root `package.json`; we are choosing standalone npm under `apps/dreamverse/web/` despite the precedent, to avoid forcing Python-only contributors to install Node. |
| streamlit | https://github.com/streamlit/streamlit | Split Python and JS testing in one repo. |
| gradio | https://github.com/gradio-app/gradio | Python package plus JS workspace precedent. |
| full-stack-fastapi-template-nextjs | https://github.com/nemanjam/full-stack-fastapi-template-nextjs | Separate frontend/backend build and deploy workflows. |
@@ -623,8 +616,8 @@ class runs where, because not all tests can run on `ubuntu-latest` CI.
| **Unit** | (none / `unit`) | `ci-dreamverse-backend.yml` (ubuntu-latest CI) + locally | Pure logic, no GPU, no live service. Mocked FastVideo backends, schema validation, helper functions. | `test_config.py`, `test_rewrite_prompt_payload.py`, `test_session_init_image.py`, the new `test_import_contract.py` |
| **Integration (fakes)** | `integration` | `ci-dreamverse-backend.yml` + locally | FastAPI test client + in-process fakes/mocks for GPU pool. Validates routes, request/response shapes, session state machine. | `test_health_endpoints.py`, `test_mock_server.py`, `test_entrypoints.py`, `test_prompt_safety.py`, `test_batching.py` (deleted) |
| **Live-service GPU** | `gpu` (skip-by-default in CI) | **Local GPU4 manual QA** + Buildkite-Modal (when added) | Real `fastvideo serve` process + real model weights + real WebSocket round-trips. Validates LTX-2 streaming, NVFP4 wiring, continuation state, frame emission. | `test_realtime_stress.py` (947 LOC), `test_session_logging.py` (1278 LOC) — these spin up real workers per their current shape |
| **Frontend unit / build** | (n/a — pnpm) | `ci-dreamverse-frontend.yml` (ubuntu-latest, no GPU) | Vitest + tsc + Next.js build. No backend needed. | `apps/dreamverse/web/src/**/*.test.ts(x)` |
| **Frontend Playwright E2E** | (n/a — pnpm) | **Local GPU4 manual QA** until Phase 4 lands public health routes; then `ci-dreamverse-frontend.yml` against a mock backend OR a deployed staging | Real browser → real backend WebSocket flow. Requires `/healthz`, `/readyz`, `/status`, `/prompt-system-config`, `/curated-presets`, `/v1/stream`. | `apps/dreamverse/web/e2e/{backend-health,frontend-shell,preset-prompt-generation}.spec.ts` |
| **Frontend unit / build** | (n/a — npm) | `ci-dreamverse-frontend.yml` (ubuntu-latest, no GPU) | Vitest + tsc + Next.js build. No backend needed. | `apps/dreamverse/web/src/**/*.test.ts(x)` |
| **Frontend Playwright E2E** | (n/a — npm) | **Local GPU4 manual QA** until Phase 4 lands public health routes; then `ci-dreamverse-frontend.yml` against a mock backend OR a deployed staging | Real browser → real backend WebSocket flow. Requires `/healthz`, `/readyz`, `/status`, `/prompt-system-config`, `/curated-presets`, `/v1/stream`. | `apps/dreamverse/web/e2e/{backend-health,frontend-shell,preset-prompt-generation}.spec.ts` |
| **FastVideo public contract** | (none) | Existing FastVideo CI (`ci-precommit` + Buildkite for GPU) | Schema/shape guards that this migration must not break. | `fastvideo/tests/contract/test_dreamverse_shape.py`, `test_dynamo_shape.py`, `test_generate_async.py` |
| **FastVideo SSIM regression** | (Buildkite path-filter) | Buildkite-Modal | Inference-quality gates for ported models. | `fastvideo/tests/ssim/test_*.py` |
@@ -709,7 +702,7 @@ CUDA_VISIBLE_DEVICES=4 uv run --locked --package dreamverse-server \
# Smoke-test from another terminal
curl -s http://localhost:8009/health | jq .
curl -s http://localhost:8009/readyz | jq . # Phase 4+ only
# Drive the FE against it: cd apps/dreamverse/web && pnpm run dev
# Drive the FE against it: cd apps/dreamverse/web && npm run dev
```
This is the **manual QA gate** for any phase that touches the live-service
@@ -907,7 +900,7 @@ mass move. This should be a small PR.
6. Add `apps/dreamverse/web/.*` to pre-commit global exclude.
7. Add `docs/contributing/dreamverse-development.md` with local dev commands:
backend `uv run --locked --package dreamverse-server --extra test pytest ...`;
frontend `cd apps/dreamverse/web && pnpm install && pnpm run build`.
frontend `cd apps/dreamverse/web && npm ci && npm run build`.
8. Run `uv lock` to regenerate `uv.lock` with the new workspace member; commit
the lock change in the same PR.
9. Land D-12-A docstring caveat: mark `GpuPool` experimental/server-internal
@@ -990,7 +983,7 @@ These are **explicit shims** — Phase 4 is responsible for their promotion to p
- [`config.py:13`](file:///home/william5lin/Dreamverse/server/config.py#L13): `_APP_ROOT / "apps" / "web"` → `_APP_ROOT / "web"` (since `_APP_ROOT` will resolve to `apps/dreamverse/` in the new layout).
- [`apps/web/next.config.ts:11`](file:///home/william5lin/Dreamverse/apps/web/next.config.ts#L11): `outputFileTracingRoot: path.resolve(__dirname, "../..")` → `path.resolve(__dirname, "../../..")` (one extra `..` since the FE is one level deeper in the monorepo).
- [`playwright.config.ts:39`](file:///home/william5lin/Dreamverse/apps/web/playwright.config.ts#L39): `command: "npm run dev"` → `command: "pnpm run dev"` (matches Phase 1 tooling decision).
- [`playwright.config.ts:39`](file:///home/william5lin/Dreamverse/apps/web/playwright.config.ts#L39): keep `command: "npm run dev"` (matches Phase 1 tooling decision).
- Any hardcoded `../FastVideo` paths in scripts/configs → make repo-root-relative since they now share a repo.
**Steps:**
@@ -1084,9 +1077,9 @@ scripts, and product docs after backend tests are green.
**Verification gate:**
- `cd apps/dreamverse/web && pnpm install --frozen-lockfile` succeeds.
- `cd apps/dreamverse/web && pnpm run build` succeeds.
- `cd apps/dreamverse/web && pnpm run test --if-present` succeeds (Vitest + tsc).
- `cd apps/dreamverse/web && npm ci` succeeds.
- `cd apps/dreamverse/web && npm run build` succeeds.
- `cd apps/dreamverse/web && npm run test --if-present` succeeds (Vitest + tsc).
- **Frontend CI Playwright is intentionally DEFERRED to Phase 4** — at Phase 3
the public `build_app` does not yet expose `/healthz`+`/readyz`+`/status`+
`/prompt-system-config`+`/curated-presets`. The `ci-dreamverse-frontend.yml`
@@ -1094,7 +1087,7 @@ scripts, and product docs after backend tests are green.
reactivates them. No PR note required.
- **Manual GPU4 Playwright smoke** (recommended): on this dev node, run
`apps/dreamverse/web/e2e/frontend-shell.spec.ts` against the GPU4-deployed
backend from Phase 2 manual QA + a `pnpm run dev` frontend at port 5274
backend from Phase 2 manual QA + a `npm run dev` frontend at port 5274
to confirm shell hydration. `backend-health.spec.ts` and
`preset-prompt-generation.spec.ts` will fail until Phase 4 — that is
expected; document as deferred.
@@ -1249,7 +1242,7 @@ product-specific prompt orchestration.
| CI cost increase | Medium | Medium | Add Dreamverse path-specific workflows; add `paths-ignore` to broad workflows; rely on Buildkite monorepo diff watch lists. | CI owner |
| FastVideo PyPI release accidentally includes Dreamverse app | High | Low | Add `apps*` to `[tool.setuptools.packages.find]` and `[tool.wheel]` excludes; verify built wheel contents. | Release owner |
| Release cadence coupling | Medium | Medium | Keep FastVideo PyPI version release unchanged; Dreamverse uses Docker/Vercel deploy from app paths. | Release owner |
| Frontend tooling drift | Medium | Medium | Pin pnpm lockfile in `apps/dreamverse/web/`; no root JS workspace. | Frontend owner |
| Frontend tooling drift | Medium | Medium | Pin npm lockfile in `apps/dreamverse/web/`; no root JS workspace. | Frontend owner |
| Security surface enlargement | Medium | Medium | Product routes stay in `apps/dreamverse/server`; only generic health/streaming routes go into FastVideo. | Backend owner |
| Product-specific API leakage into `fastvideo.*` | High | Medium | Enforce import/module boundary; keep curated presets and prompt UX product-local. | Architecture owner |
| Migration regression | High | Medium | Phase gates; backend before frontend; can stop after any phase with Dreamverse repo still usable until Phase 7. | Migration owner |
@@ -6,7 +6,10 @@ recommended next action.
For why each item is open see [decisions-log.md](decisions-log.md). For
PR-level context see [pr-roadmap.md](pr-roadmap.md).
**Last updated:** 2026-05-12 (added DR-4 follow-up for PR #1330's skipped
**Last updated:** 2026-05-14 (added DR-7 follow-up to remove LTX2 debug
logging env vars and replace them with per-pipeline/model state. Earlier: 2026-05-14 added DR-6 follow-up for PR #1335's Gemma
lazy-load/device-placement behavior outside compiled `forward`. Earlier: 2026-05-13 added DR-5 follow-up for PR #1333's LTX2
distilled SSIM reference refresh. Earlier: 2026-05-12 added DR-4 follow-up for PR #1330's skipped
`App websocket integration` suite. Earlier: 2026-05-05 D-20 broken-pipe root cause + fix landed on
`will/dreamverse-monorepo` @ `5eaf0a13`; added new thread D-20-CP for
cherry-picking the public-API audio routing fix to `will/ltx2_sr_port`
@@ -34,6 +37,9 @@ different vehicle.).
| **DR-2** | Med | Decide `cerebras_ifm` provider path: (a) public Literal + `CerebrasIFMProvider` shipped, OR (b) Dreamverse-side custom provider via `enhancer.register_provider(...)` | S (decision) + S-M (impl) | Resolves the cerebras_ifm gap left by PR #1258. Same item as legacy #3 below; DR-2 is the Dreamverse-side framing. |
| **DR-3** | Low | Replace Dreamverse `PromptEnhancer._run_blocking_request` manual thread/queue polling with `asyncio.to_thread` after the PR #1327 prompt-enhancer compatibility surface is retired or isolated | S | Review comment #1327 (`prompt_enhancer.py`) is valid, but deferred to avoid patching the local fork in this PR. |
| **DR-4** | Low | Investigate unskipping PR #1330's skipped public `App websocket integration` suite | S-M | Public PR #1330 has `describe.skip(...)` around 27 websocket tests while the internal equivalent suite is active with 25 tests. The 2 public-only tests cover backend unreachable / GPU workers not ready. Not blocking while skipped, but stale assertions may need safe refresh before unskip. |
| **DR-5** | Low | Regenerate LTX2-Distilled latent SSIM references under the intended neutral/distilled defaults, then remove the PR #1333 historical full-guidance pins | S-M | PR #1333 changed public LTX2 distilled defaults to neutral/distilled values, but existing LTX2 latent SSIM references appear to have been generated with historical full-guidance defaults. The current PR pins the SSIM test to old values to keep CI compatible until references are refreshed. |
| **DR-6** | Low | Decide whether `LTX2GemmaTextEncoderModel` needs a non-forward device-placement hook after lazy Gemma load | S | PR #1335 should remove the `model.device` / `model.to(...)` guard from `forward` for Dynamo/fullgraph compatibility. Non-compiled runs probably do not need it because `gemma_model` moves Gemma at first load, but a later wrapper `.to(...)` after lazy load could leave Gemma on the old device unless lifecycle placement handles it. |
| **DR-7** | Low | Remove LTX2 debug logging env-var plumbing and replace it with per-pipeline/model debug state | S-M | PR #1335 review flagged that `initialize_pipeline()` mutates process-global LTX2 debug env vars, which can leak/race across pipeline instances. Deleting only the mutation is low-risk but loses config-driven debug logging; deleting all reads without replacement would remove useful SSIM/latent drift diagnostics. |
| **3** | Med | Add `cerebras_ifm` to `PromptEnhancerConfig.provider` Literal + provider | S-M | Public-side resolution if DR-2 picks (a) |
| **4** | Med | Expose `layer_profile` on typed `engine.quantization` | M | Removes Dreamverse's `experimental["pipeline_config"]` dodge for stage profiles |
| **5** | Med | Design typed `dit_config.quant_config` carrier | L design + L impl | Removes broader `experimental["pipeline_config"]` escape hatch |
@@ -421,6 +427,65 @@ current PR #1330 review.
stack end-to-end and compare the public assertions against the internal active
suite.
### Item DR-6: Gemma lazy-load device placement outside `forward`
**Why:** PR #1335 review flagged this pattern in
`fastvideo/models/encoders/gemma.py::LTX2GemmaTextEncoderModel.forward`:
```py
if model.device != target_device:
model.to(device=target_device)
```
The immediate concern is the compiled/Dynamo path: `model.device` and
`model.to(...)` inside `forward` can introduce non-tensor/device parsing work
that fullgraph tracing should not see. Removing the guard from `forward` is the
right PR #1335 review fix.
For non-compiled execution, the guard is mostly defensive rather than required:
`gemma_model` already moves the lazily loaded HF Gemma model to the wrapper's
current parameter device on first load. The remaining edge case is a lifecycle
sequence where Gemma is loaded, then the parent wrapper is later moved to a
different device; in that case Gemma could stay behind unless placement is
handled outside `forward`.
**Action:** After PR #1335 review is unblocked, decide whether FastVideo needs a
small lifecycle hook/helper for this class so lazy Gemma is moved whenever the
wrapper/device placement changes. If yes, implement it outside `forward`; if no,
document that Gemma must be loaded after final device placement.
**Effort:** Small.
**Dependencies:** Not blocking PR #1335 if the forward-path guard is removed and
the normal load path keeps placing Gemma on the wrapper's current device.
### Item DR-7: Remove LTX2 debug logging env-var plumbing
**Why:** PR #1335 review flagged that
`fastvideo/pipelines/basic/ltx2/ltx2_pipeline.py::initialize_pipeline()` sets
and pops LTX2 debug env vars such as `LTX2_PIPELINE_DEBUG_LOG`,
`LTX2_PIPELINE_DEBUG_PATH`, `LTX2_DEBUG_DETAIL`, and
`LTX2_PIPELINE_DEBUG_DETAIL_PATH`. Those env vars are process-global, so one
pipeline instance can enable, overwrite, or clear debug behavior for another
pipeline instance running in the same process.
Deleting only the `initialize_pipeline()` env mutation is low risk for normal
generation, but it would stop config-driven debug logging unless replaced.
Deleting all env-var reads without replacement is riskier because these logs are
useful for SSIM/latent drift diagnosis and may be used by local debug scripts.
**Action:** Replace LTX2 debug env-var plumbing with per-pipeline/model debug
state. Use pipeline/model config for construction-time hooks and a
`ForwardContext`/`ForwardBatch`-style carrier for forward-time logging. Keep
external env-var compatibility only if there is a documented operator workflow
that still needs it.
**Effort:** Small-Medium.
**Dependencies:** Not blocking PR #1335 if the immediate fix is limited to
removing process-global mutation from pipeline initialization while preserving
existing externally supplied env-var reads.
### Item D-12-C: Avoid locking `PoolAssignment.gpu_id: int` as public
**Why:** Today `PoolAssignment` exposes `gpu_id: int`, assuming
+29 -20
View File
@@ -12,7 +12,7 @@ _Last updated: 2026-03-02_
| Metric | Category | Status | Location | Trust |
|--------|----------|--------|----------|-------|
| **FVD** | Distribution | ✅ Implemented | `benchmarks/fvd/` | High |
| **FVD** | Distribution | ✅ Implemented | `fastvideo/eval/metrics/common/fvd/` | High |
| **SSIM** | Reference | ✅ Implemented | `fastvideo/tests/ssim/` | High |
| **LPIPS** | Perceptual | ✅ Implemented | `scripts/lora_extraction/` | Medium |
| **Loss trajectory** | Training signal | ✅ Implemented | W&B `train_loss` | Medium |
@@ -27,8 +27,8 @@ _Last updated: 2026-03-02_
### FVD — Fréchet Video Distance
**Category**: Distribution-level quality metric
**Status**: ✅ Fully implemented in `benchmarks/fvd/`
**Trust**: High — standard protocol, I3D feature extractor
**Status**: ✅ Registered as the `common.fvd` eval metric in `fastvideo/eval/metrics/common/fvd/`
**Trust**: High — standard protocol, I3D feature extractor (CLIP / VideoMAE backbones also available, research-grade)
#### What It Measures
FVD measures the distance between the **distribution** of generated videos and
@@ -59,30 +59,39 @@ Lower FVD = generated videos are more statistically similar to real videos.
#### How to Use
```python
# Programmatic
from benchmarks.fvd import compute_fvd_with_config, FVDConfig
# Programmatic — drive the metric directly for custom kwargs
from fastvideo.eval import get_metric
config = FVDConfig.fvd2048_16f() # Standard: 2048 videos, 16 frames
results = compute_fvd_with_config('data/real/', 'outputs/gen/', config)
print(f"FVD: {results['fvd']:.2f}")
metric = get_metric("common.fvd", extractor="i3d") # or "clip" / "videomae"
metric.to("cuda")
metric.setup()
metric.reset()
# First sample carries the reference set; later samples reuse the cache.
metric.accumulate({"video": gen_tensors[0], "reference": real_tensors})
for gen in gen_tensors[1:]:
metric.accumulate({"video": gen})
result = metric.finalize()
print(f"FVD: {result.score:.2f}")
```
```bash
# CLI
python -m benchmarks.fvd.cli \
--real-path data/real/ \
--gen-path outputs/gen/ \
--protocol fvd2048_16f
# CLI — folder of generated mp4s vs a reference folder
python examples/inference/eval/eval_fvd.py \
--gen-dir outputs/gen/ \
--reference-dir data/real/ \
--extractor i3d \
--output fvd_scores.json
```
**Preset protocols**:
| Protocol | Videos | Frames | Use Case |
|----------|--------|--------|----------|
| `fvd2048_16f` | 2048 | 16 | Standard benchmark (papers) |
| `fvd2048_128f` | 2048 | 128 | Long video evaluation |
| `quick_test` | 100 | 16 | Fast dev iteration |
**Feature extractors**: `i3d` (default, standard FVD spec used in papers),
`clip`, `videomae` (research-grade; not directly comparable to published
FVD numbers).
**Feature extractors**: `i3d` (default, standard), `clip`, `videomae`
**Protocol**: standard FVD uses 2048 generated + 2048 reference videos at
16 frames each. A warning fires below 256 — the score becomes
statistically unreliable.
#### Interpretation
| FVD Range | Interpretation |
+331
View File
@@ -0,0 +1,331 @@
# Performance Dashboard Memory
Date: 2026-06-16
Branch: `ci/dashboard`
## Purpose
This branch adds a local live dashboard for FastVideo performance benchmark
history. It is intended for maintainer/operator use: inspect latest benchmark
status, compare current values with recent baseline context, and view trends
from the Hugging Face performance-tracking dataset.
The dashboard is a FastAPI + React app. It is separate from the existing
Svelte `ui/` app.
## Main Files Added Or Changed
Backend:
- `fastvideo/performance_dashboard/__init__.py`
- `fastvideo/performance_dashboard/__main__.py`
- `fastvideo/performance_dashboard/api.py`
- `fastvideo/performance_dashboard/metrics.py`
- `fastvideo/performance_dashboard/service.py`
Frontend:
- `performance_dashboard/frontend/package.json`
- `performance_dashboard/frontend/package-lock.json`
- `performance_dashboard/frontend/tsconfig.json`
- `performance_dashboard/frontend/vite.config.ts`
- `performance_dashboard/frontend/index.html`
- `performance_dashboard/frontend/scripts/build.mjs`
- `performance_dashboard/frontend/src/main.tsx`
- `performance_dashboard/frontend/src/api.ts`
- `performance_dashboard/frontend/src/App.tsx`
- `performance_dashboard/frontend/src/styles.css`
Docs/tests:
- `performance_dashboard/README.md`
- `docs/contributing/performance_benchmarks.md`
- `fastvideo/tests/performance/test_dashboard_service.py`
- `fastvideo/tests/performance/test_dashboard_api.py`
Shared HF utility change:
- `fastvideo/tests/performance/hf_store.py`
## Data Source
The source of truth remains the Hugging Face dataset repo used by existing
performance CI:
```text
HF_REPO_ID=FastVideo/performance-tracking
```
The dataset stores normalized JSON records emitted by
`fastvideo/tests/performance/compare_baseline.py`. The current v1 normalized
schema includes:
- `model_id`
- `timestamp`
- `commit_sha`
- `gpu_type`
- `latency`
- `throughput`
- `memory`
- `text_encoder_time_s`
- `dit_time_s`
- `vae_decode_time_s`
- `success`
Records are grouped by `(model_id, gpu_type)` for v1 dashboard behavior.
## Local Cache
The backend syncs the HF dataset to a local cache directory:
```text
PERFORMANCE_TRACKING_ROOT=/tmp/fastvideo-perf-dashboard
```
If `PERFORMANCE_TRACKING_ROOT` is not set, the dashboard defaults to:
```text
/tmp/fastvideo-perf-dashboard
```
The sync is performed through the existing helper:
```python
fastvideo.tests.performance.hf_store.sync_from_hf(...)
```
The dashboard then loads JSON files from the local cache through:
```python
fastvideo.tests.performance.hf_store.load_records(...)
```
## Authentication
Originally `hf_store.py` only read `HF_API_KEY`. This caused local dashboard
runs to fail when users had standard Hugging Face token variables set.
`hf_store.py` now resolves tokens from the first available variable in:
```text
HF_API_KEY
HUGGINGFACE_HUB_TOKEN
HF_TOKEN
```
For local use:
```bash
export HF_TOKEN=hf_...
```
If the HF repo is private or gated, the token must have dataset read access.
## Backend API
The FastAPI app is created by:
```python
fastvideo.performance_dashboard.api:create_app
```
The module-level app is:
```python
fastvideo.performance_dashboard.api:app
```
Endpoints:
- `GET /api/performance/health`
- `POST /api/performance/refresh`
- `GET /api/performance/records?days=90`
- `GET /api/performance/summary?days=90`
- `GET /api/performance/trends?days=90`
`POST /api/performance/refresh` forces a fresh HF sync.
## Status Semantics
Important: the dashboard intentionally separates stored CI status from
recomputed context.
Stored status:
- Comes directly from the latest JSON record's `success` field.
- This is what the dashboard displays as `Stored Status`.
- This is the primary latest status.
Recomputed status:
- Calculated locally from cached records for explanatory context.
- Uses the latest record's metric values compared to the median of the latest
five previous successful records in the same `(model_id, gpu_type)` group.
- Displayed separately as `Recomputed`.
- Does not override the stored JSON `success` status.
This distinction was added after observing that recomputing pass/fail from the
local cache can disagree with the status originally uploaded by CI.
## Time Window Behavior
The default dashboard time window is 90 days.
The selected `days` value affects:
- trend charts
- record browsing/filtering
The selected `days` value does not affect:
- latest stored status
- latest summary baseline context
Reason: latest status should not change when users widen or narrow the trend
window. The API keeps `days` on `/summary` only for shared frontend filter
state, but summary loading uses all cached records.
This fixed a bug where changing from roughly 35 days to 42 days could change
the latest status from pass to fail because older records entered the local
baseline window.
## Metric Logic
Dashboard metric definitions live in:
```text
fastvideo/performance_dashboard/metrics.py
```
Tracked metrics:
- `latency` lower is better
- `throughput` higher is better
- `memory` lower is better
- `text_encoder_time_s` lower is better
- `dit_time_s` lower is better
- `vae_decode_time_s` lower is better
Baseline context uses the median of up to five previous successful records for
the same `(model_id, gpu_type)`.
## Frontend Behavior
The React app:
- fetches `/api/performance/summary`
- fetches `/api/performance/trends`
- displays summary cards
- displays latest rows by model/GPU
- displays native SVG trend charts
- has model/GPU/day filters
- includes a refresh button
- auto-refreshes every five minutes
The UI is implemented without a charting library. Trend charts are native SVG
in `performance_dashboard/frontend/src/App.tsx`.
The production frontend build uses `esbuild` through
`performance_dashboard/frontend/scripts/build.mjs`. Vite is still used for the
dev server and `/api` proxy.
Why esbuild for production build:
- Vite/Rollup hit a local macOS native optional dependency code-signing issue
in this environment.
- Direct esbuild worked reliably and is sufficient for this small dashboard.
## Static Serving
After frontend build, the FastAPI server serves:
- static JS/CSS from `performance_dashboard/frontend/dist/assets`
- `performance_dashboard/frontend/dist/index.html` for the dashboard page
This allows a single local port to serve both the API and UI.
## Local Run Workflow
Build frontend:
```bash
cd performance_dashboard/frontend
conda run -n fastvideo env PATH=/Applications/Codex.app/Contents/Resources/cua_node/bin:/usr/local/bin:/usr/bin:/bin \
/Applications/Codex.app/Contents/Resources/cua_node/bin/npm install
conda run -n fastvideo env PATH=/Applications/Codex.app/Contents/Resources/cua_node/bin:/usr/local/bin:/usr/bin:/bin \
/Applications/Codex.app/Contents/Resources/cua_node/bin/npm run build
```
Run dashboard:
```bash
export HF_TOKEN=hf_...
python -m fastvideo.performance_dashboard --host 0.0.0.0 --port 8000
```
Open locally:
```text
http://127.0.0.1:8000
```
## ngrok Workflow
`python -m fastvideo.performance_dashboard --host 0.0.0.0 --port 8000`
starts the actual local dashboard server.
`ngrok http 8000` does not start the dashboard. It exposes the already-running
local server through a temporary public URL.
Typical flow:
```bash
python -m fastvideo.performance_dashboard --host 0.0.0.0 --port 8000
ngrok http 8000
```
Use the HTTPS URL printed by ngrok to view the dashboard remotely.
## Verification Commands
Backend tests:
```bash
conda run -n fastvideo python -m pytest \
fastvideo/tests/performance/test_dashboard_service.py \
fastvideo/tests/performance/test_dashboard_api.py \
-q
```
Expected after latest changes:
```text
8 passed
```
Frontend build:
```bash
cd performance_dashboard/frontend
conda run -n fastvideo env PATH=/Applications/Codex.app/Contents/Resources/cua_node/bin:/usr/local/bin:/usr/bin:/bin \
/Applications/Codex.app/Contents/Resources/cua_node/bin/npm run build
```
Expected:
```text
tsc && node scripts/build.mjs
```
with exit code 0.
## Known Notes
- `performance_dashboard/frontend/node_modules/` and
`performance_dashboard/frontend/dist/` are ignored by git.
- `npm install` reported two high-severity audit findings in dependency tree.
`npm audit fix --force` was not run because it can introduce breaking
dependency upgrades.
- Existing `fastvideo` package imports may emit platform warnings such as NPU
or macOS torch distributed messages. These are not dashboard-specific errors.
@@ -14,7 +14,7 @@ based on **Wan2.1** (SkyReels-V2) DiT models with causal attention for
auto-regressive streaming generation.
**Key techniques you will work with:**
- Full finetuning and LoRA on Wan / LTX-2 / MatrixGame models
- Full finetuning and LoRA on Wan / LTX-2 / Matrix-Game 2.0 models
- DMD-based distillation (few-step generation)
- Self-Forcing distillation (causal streaming)
- Diffusion-Forcing SFT (DFSFT) for causal models
@@ -225,10 +225,10 @@ callbacks:
|----------|-----------|----------|
| Wan T2V finetune | `fastvideo/training/wan_training_pipeline.py` | Standard text-to-video finetune / LoRA |
| Wan I2V finetune | `fastvideo/training/wan_i2v_training_pipeline.py` | Image-to-video (first frame conditioned) |
| MatrixGame finetune | `fastvideo/training/matrixgame_training_pipeline.py` | Action-conditioned world model |
| MatrixGame AR diffusion | `fastvideo/training/matrixgame_ar_diffusion_pipeline.py` | AR diffusion-forcing training |
| MatrixGame ODE-init | `fastvideo/training/matrixgame_ode_causal_pipeline.py` | ODE-trajectory init |
| MatrixGame self-forcing distill | `fastvideo/training/matrixgame_self_forcing_distillation_pipeline.py` | Self-forcing distillation |
| Matrix-Game 2.0 finetune | `fastvideo/training/matrixgame2_training_pipeline.py` | Action-conditioned world model |
| Matrix-Game 2.0 AR diffusion | `fastvideo/training/matrixgame2_ar_diffusion_pipeline.py` | AR diffusion-forcing training |
| Matrix-Game 2.0 ODE-init | `fastvideo/training/matrixgame2_ode_causal_pipeline.py` | ODE-trajectory init |
| Matrix-Game 2.0 self-forcing distill | `fastvideo/training/matrixgame2_self_forcing_distillation_pipeline.py` | Self-forcing distillation |
| LTX-2 finetune | `fastvideo/training/ltx2_training_pipeline.py` | LTX-2 architecture finetuning |
| Wan DMD distillation | `fastvideo/training/wan_distillation_pipeline.py` | Few-step distillation via DMD |
| Self-Forcing distill | `fastvideo/training/wan_self_forcing_distillation_pipeline.py` | Causal streaming distillation |
@@ -258,7 +258,7 @@ Read `.agents/memory/evaluation-registry/README.md` for the full metric catalog.
|--------|-------------|-------|
| **Loss trajectory** | Every run, real-time from W&B | Medium |
| **SSIM** | When comparing against reference outputs | High |
| **FVD** | For benchmarking model quality (`benchmarks/fvd/`) | High |
| **FVD** | For benchmarking model quality (`common.fvd` eval metric; example: `examples/inference/eval/eval_fvd.py`) | High |
| **LPIPS** | LoRA merge validation | Medium |
| **Human preference** | Major checkpoints | Highest |
@@ -278,8 +278,8 @@ Read `.agents/memory/evaluation-registry/README.md` for the full metric catalog.
## World Model–Specific Concepts
### Action Injection (MatrixGame)
The MatrixGame pipeline adds **action modules** to each DiT block, enabling
### Action Injection (Matrix-Game 2.0)
The Matrix-Game 2.0 pipeline adds **action modules** to each DiT block, enabling
frame-level mouse/keyboard input conditioning. The action sequence is injected
per-frame alongside the latent video tokens.
@@ -197,6 +197,12 @@ Tolerance guide:
Element-wise `assert_close` alone is not enough for deep full-DiT parity. Also
log global abs-mean drift and per-modality summaries.
When a non-skip component parity run is numerically red after weight/input
checks, invoke `../add-model-08-trace/SKILL.md` before adding bespoke forward
hooks. Use `docs/contributing/activation_trace.md` to keep
`FASTVIDEO_TRACE_LAYERS`, `FASTVIDEO_TRACE_STATS`, and `FASTVIDEO_TRACE_STEPS`
identical across FastVideo and upstream traces.
Useful local commands:
```bash
@@ -94,8 +94,12 @@ Run the shared parity-debug loop. The component test command is:
pytest <parity_test> -v -s
```
For numerical drift, narrow the first divergent block with per-block hooks or
intermediate tensor comparisons before changing layers.
For numerical drift, use `../add-model-08-trace/SKILL.md` before writing bespoke
hooks. Start with FastVideo's activation trace (`fastvideo/hooks/activation_trace.py`;
`docs/contributing/activation_trace.md`) and a block-level regex such as
`FASTVIDEO_TRACE_LAYERS="^block\.layers\.[0-9]+$"`. Only fall back to custom
per-block hooks if the needed boundary or statistic is not exposed by
`FASTVIDEO_TRACE_STATS`.
## Escape Hatches
+82 -90
View File
@@ -1,6 +1,6 @@
---
name: add-model-08-trace
description: Use during /add-model Phase 6 when component parity has failed and root cause requires layer-by-layer divergence analysis. Instruments both the official reference and FastVideo port with forward hooks to find the first numerical divergence point.
description: Use during /add-model Phase 6 when component parity has failed and root cause requires layer-by-layer divergence analysis. Uses FastVideo activation trace first, falling back to custom hooks only for boundaries or stats the utility cannot observe.
---
# Add-Model Trace
@@ -38,104 +38,90 @@ Required inputs before starting:
- Shared deterministic test inputs (same tensors on both sides).
- The component parity test file path and its current failure output.
## Hard Rules: Instrumentation Hierarchy
## Primary Path: FastVideo Activation Trace
Apply these in priority order. Use the highest-priority method that works for
the target site.
Use FastVideo's first-class activation trace before writing custom hooks:
`fastvideo/hooks/activation_trace.py`, documented in
`docs/contributing/activation_trace.md`.
### (1) Forward hooks (PREFERRED)
Pipeline runs attach trace to the transformer during pipeline initialization.
Component-only parity harnesses may call `attach_activation_trace(model)` from
local test/debug code; do not add trace calls to production model code.
Prefix the failing parity command with a tight layer regex:
```bash
FASTVIDEO_TRACE_ACTIVATIONS=1 \
FASTVIDEO_TRACE_LAYERS="^block\.layers\.[0-9]+$" \
FASTVIDEO_TRACE_STATS="abs_mean,sum,max,shape" \
FASTVIDEO_TRACE_STEPS="0" \
FASTVIDEO_TRACE_OUTPUT="/tmp/opencode/fv_trace.jsonl" \
pytest tests/local_tests -k "parity" -v -s
```
Match the layer regex to the actual `model.named_modules()` names. Empty or
broad regexes are expensive; prefer block-level names first, then narrow to
submodules after the first divergent block is known.
## Trace Compare Contract
One JSONL file per side. FastVideo output should use `FASTVIDEO_TRACE_OUTPUT`;
the upstream harness should emit the same JSONL shape:
```json
{"module":"block.layers.0","tensor":"out","step":0,"abs_mean":0.0123,"sum":1.0,"max":0.5,"shape":[1,16,32]}
```
Compare rows by `(module, step, tensor)`. The first row whose `shape`,
`abs_mean`, or `max` diverges beyond the component tolerance is the first broken
boundary. Keep `FASTVIDEO_TRACE_LAYERS`, `FASTVIDEO_TRACE_STATS`, and
`FASTVIDEO_TRACE_STEPS` identical between sides; if row order differs, sort or
normalize before diffing.
## Drill-Down Loop
**Initial run:** trace every top-level block (`^block\.layers\.[0-9]+$` or the
family's equivalent). Identify the first block index where `abs_mean` or `max`
drifts beyond tolerance while earlier blocks match.
**Drill run:** tighten `FASTVIDEO_TRACE_LAYERS` to submodules inside the first
divergent block: attention output, MLP projections, norm outputs, modality
adapters, or other named boundaries exposed by `named_modules()`.
**Iterate:** if the first divergent operation is a free function or tensor op not
visible as an `nn.Module`, use the fallback instrumentation hierarchy below.
The loop ends when the first divergent submodule or operation is identified with
a file:line citation in the official source.
## Fallback Instrumentation Hierarchy
Use these only when activation trace cannot observe the needed boundary or
statistic.
### (1) Custom forward hooks
`module.register_forward_hook(...)` and `register_forward_pre_hook(...)`.
Always within `try/finally` with `handle.remove()`. Zero source residue.
```python
handle = module.register_forward_hook(fn)
try:
output = model(inputs)
finally:
handle.remove()
```
### (2) Runtime monkey-patch (PREFERRED over source edits)
### (2) Runtime monkey-patch
`module.attr = wrapped_func` or `cls.method = wrapped_method`, restored via
`try/finally` (save original first). Use for free functions and non-Module
sites such as activation functions (`swiglu`, `apply_rotary_emb`) that cannot
be hooked as `nn.Module` submodules.
```python
original = cls.method
cls.method = wrapped
try:
output = model(inputs)
finally:
cls.method = original
```
`try/finally` (save original first). Use for free functions and non-Module sites
such as activation functions (`swiglu`, `apply_rotary_emb`).
### (3) Source edits in FastVideo's own code
Only when (1) and (2) are insufficient. Track all edits within a single named
`git stash` boundary OR a temporary branch. Run `git diff` before closing the
investigation to confirm the stash or branch is clean. The cleanup gate
enforces this.
investigation to confirm cleanup.
### (4) Source edits in official repo source
Allowed if EITHER:
- (a) The official repo is a git-tracked clone (e.g. `daVinci-MagiHuman/` at
the repo root): use `git diff` in the clone path to verify cleanup.
- (b) It's installed editable (`pip install -e .`): use `git diff` in the
editable source path to verify cleanup.
If the official repo is installed non-editable in site-packages: back up the
target file (`cp original.py original.py.trace-backup`) before editing, then
restore from backup at the end (or `pip install --force-reinstall <pkg>`).
The cleanup gate verifies via diff-against-backup or zero-diff-in-clone.
## Logging Contract
One log file per side. Paths:
```
/tmp/opencode/<family>_<component>_up_layers.log
/tmp/opencode/<family>_<component>_fv_layers.log
```
Format: one line per captured tensor, space-separated:
```
<name> <shape> <abs_mean> <sum> <min> <max>
```
Example:
```
block[00] (1,512,1024) 0.012345 6.3210 -0.4321 0.4321
```
Keep the format diff-friendly. Running `diff /tmp/opencode/x_up.log
/tmp/opencode/x_fv.log` should highlight the first divergent line directly.
Retain side-by-side stdout output alongside the per-side files for human
review.
## Drill-Down Loop
**Initial run:** attach hooks to every top-level block (`model.block.layers[i]`
or equivalent). Identify the first block index `NN` where abs_mean relative
drift exceeds 0.5% compared to the previous block.
**Drill run:** set `<FAMILY>_DEBUG_DRILL_LAYER=NN` and re-run. The script
attaches submodule hooks inside block `NN`: attention output, mlp.pre_norm,
mlp.up_gate_proj, mlp.down_proj input (via pre-hook) and output, mlp output,
attn_post_norm (if present), mlp_post_norm (if present).
**Iterate:** if the drill run points to a free function (e.g. an activation
not wrapped in an `nn.Module`), switch to a monkey-patch (method 2) to
intercept its output via the next module's pre-hook.
The loop ends when the first divergent submodule is identified with a
file:line citation in the official source.
Allowed only when hook and monkey-patch approaches cannot capture the site.
For git-tracked or editable official clones, use `git diff` in the clone path to
verify cleanup. For non-editable site-packages, back up the target file before
editing and restore it before handoff.
## Hypothesis Toggles
@@ -194,11 +180,13 @@ Escalate to the calling bucket skill when:
Return to the calling subagent with:
- File paths to per-side logs (`/tmp/opencode/<family>_<component>_{up,fv}_layers.log`).
- The identified first divergent layer or submodule name.
- FastVideo trace JSONL path and upstream trace JSONL path.
- Trace settings used: `FASTVIDEO_TRACE_LAYERS`, `FASTVIDEO_TRACE_STATS`, and
`FASTVIDEO_TRACE_STEPS`.
- The first divergent `(module, step, tensor)` row and observed drift.
- The upstream file:line citation where the divergence originates.
- Hypothesis verdict if an A/B toggle was used (e.g. "PATCH_LINEAR=1 closes
the gap, confirming dtype-cast difference in PackedExpertLinear").
- Fallback hook/patch verdict if activation trace could not observe the boundary.
- Hypothesis verdict if an A/B toggle was used, for example `PATCH_LINEAR=1`.
- Cleanup-gate status: `[cleanup-gate] PASS` or a list of unresolved items.
The calling agent uses this to scope the production fix in the FastVideo
@@ -206,10 +194,14 @@ component file.
## References
- `templates/block_trace_debug.py` in this skill directory: the canonical
template this skill generalizes.
- `docs/contributing/activation_trace.md` for canonical activation-trace env vars,
JSONL output, cost model, and troubleshooting.
- `fastvideo/hooks/activation_trace.py` for the implementation and
`attach_activation_trace(model)` entry point.
- `templates/block_trace_debug.py` in this skill directory: fallback custom-hook
template when activation trace cannot observe the needed boundary or stat.
- `tests/local_tests/transformers/_debug_magi_human_block_parity.py` in the
FastVideo3 repo: the worked magi-human example this skill was extracted from.
FastVideo3 repo: historical worked example for custom hook/patch debugging.
- `add-model/SKILL.md` Phase 6: the calling context for this skill.
- `add-model-03-port-dit/SKILL.md`, `add-model-04-port-vae/SKILL.md`,
`add-model-05-port-encoder/SKILL.md`, `add-model-06-port-generic/SKILL.md`:
@@ -157,6 +157,13 @@ Debug pipeline drift in this order:
channel order, sample rate, FPS, and final slicing.
7. Add targeted stage-level diagnostics to identify the first divergent stage.
If stage diagnostics show the first bad stage is transformer/denoising or a
mid-DiT block, enable activation trace before adding ad hoc pipeline prints; see
`docs/contributing/activation_trace.md` and `../add-model-08-trace/SKILL.md`.
Keep `FASTVIDEO_TRACE_LAYERS`, `FASTVIDEO_TRACE_STATS`, and
`FASTVIDEO_TRACE_STEPS` identical across reruns so pipeline parity traces diff
one-to-one.
If the first divergence belongs to component implementation, strict loading, or
conversion mapping, stop pipeline edits and return `next_step=return_to_phase_6`
with the exact failing evidence. Do not patch conversion from this skill.
+5
View File
@@ -267,6 +267,11 @@ conversion, route it through `../add-model-07-conversion/SKILL.md` with a retry
request matching `contracts/conversion_request.md`, then resume the component
skill with the updated conversion handoff.
When a component failure narrows to layer-by-layer numerical drift, load
`../add-model-08-trace/SKILL.md` before writing custom hooks. It uses
`fastvideo/hooks/activation_trace.py`; canonical env vars and JSONL format are
documented in `docs/contributing/activation_trace.md`.
Phase 6 ends only when every required component handoff reports
`parity_status=non_skip_pass`, or when a precise blocker or escape hatch is
recorded in `port_state_file`.
+32
View File
@@ -0,0 +1,32 @@
---
name: add-reward-model
description: Use when adding reusable reward models under fastvideo/train/methods/rl/rewards for RLHF or online RL training.
---
# Add Reward Model
Use for reward models consumed by RL methods.
## Placement
- Put reusable reward code under `fastvideo/train/methods/rl/rewards/`.
- Expose public builders from `fastvideo/train/methods/rl/rewards/__init__.py`.
- Keep method-specific aggregation or advantage logic out of reward classes.
## Media Inputs
- Reward callables receive decoded media tensors.
- Accept single-frame tensors as `[B, C, H, W]` and multi-frame tensors as `[B, C, T, H, W]` when practical.
- Frame selection is reward-specific. Frame scorers such as PickScore and CLIPScore should explicitly select frame `0`; temporal rewards should inspect whichever frames they need.
- Return one scalar reward per prompt/sample.
## Attribution
- If code is ported or closely adapted from another repo, add a short comment or docstring naming the source file/function.
- Preserve SPDX headers used by FastVideo files.
## Tests
- Unit-test tensor layout handling without loading large reward checkpoints.
- Allow fake scorer injection for multi-reward tests.
- Test weighted reward aggregation and metric keys.
+38
View File
@@ -0,0 +1,38 @@
---
name: add-rl-method
description: Use when adding or modifying an RL/RLHF method under fastvideo/train/methods/rl, including DiffusionNFT-like methods.
---
# Add RL Method
Use for new RL methods in the modular `fastvideo/train` stack.
## Required Shape
- Add the method under `fastvideo/train/methods/rl/`.
- Subclass `TrainingMethod`.
- Keep model-family logic in `ModelBase` wrappers.
- Decode generated latents through `ModelBase.decode_latents`; add that hook to the new model wrapper instead of decoding inside the RL method.
- Use `fastvideo/train/methods/rl/common/sampling.py` for generation unless the method has a documented reason to avoid sampling.
- Use `fastvideo/train/methods/rl/common/prompt_sampling.py` for reusable grouped prompt sampling patterns such as DiffusionNFT K-repeat.
- Use `fastvideo/train/methods/rl/rewards/` for reward models.
## Optimization
- Return `manages_optimization() == True` only when the method must own a nonstandard outer/inner loop.
- If using managed optimization, implement `managed_train_step(data_stream, iteration)`.
- Existing trainer callbacks, checkpointing, tracking, and validation should still work.
## Config
- Put method knobs under `method`.
- Put sampler knobs under `method.sampling`.
- Do not put scheduler or trajectory policy into model configs.
- Do not split a diffusers-style scheduler from its built-in `step()` solver in YAML; use `trajectory` only for higher-level ODE vs re-noise behavior.
- Avoid fixed timestep lists in examples unless reproducing a known baseline; prefer scheduler-generated defaults.
## Tests
- Add fake-model tests for sampler/method behavior.
- Add config parse tests for the public YAML.
- Confirm existing train methods stay on the default Trainer path.
+9 -4
View File
@@ -1,3 +1,8 @@
---
name: dreamverse-deploy
description: Use when redeploying the migrated Dreamverse app backend and frontend on a chosen local GPU; tears down existing ports, launches services, and waits for readiness checks.
---
# dreamverse-deploy — redeploy migrated Dreamverse on a chosen GPU
**Scope:** project (lives in this repo at `.agents/skills/dreamverse-deploy/`)
@@ -17,7 +22,7 @@ both `/readyz` and the FE root to return 200.
`cerebras-cloud-sdk`, `openai` installed (override the default path with
`DREAMVERSE_PYTHON=/path/to/python`)
- `~/.env` exporting `CEREBRAS_API_KEY`, `GROQ_API_KEY`, etc.
- pnpm installed at `/home/william5lin/.local/share/pnpm/pnpm` (or in `$PATH`)
- npm available in `$PATH` (or set `NPM=/path/to/npm`)
- `gcc-13` + `g++-13` at `/usr/bin/` (workaround for nvcc gcc-15 rejection)
- **Recommended:** native ffmpeg env file at `apps/dreamverse/scripts/ffmpeg-env.sh`
(built once via `bash apps/dreamverse/scripts/install_native_ffmpeg.sh`).
@@ -105,7 +110,7 @@ Flags can appear in any position relative to the positional args. Explicit flag
5. Launches the backend via `apps/dreamverse/scripts/dreamverse-server` in a
detached `setsid` session, captures PID.
6. Polls `/readyz` until 200 (max 5 min).
7. Launches the frontend via `pnpm run dev:devtools` in a detached session,
7. Launches the frontend via `npm run dev:devtools` in a detached session,
captures PID.
8. Polls FE `/` until 200 (max 60s).
9. Prints URLs, PIDs, and log paths.
@@ -120,7 +125,7 @@ Flags can appear in any position relative to the positional args. Explicit flag
PLAYWRIGHT_SKIP_WEBSERVER=1 BACKEND_URL=http://127.0.0.1:8009 \
PLAYWRIGHT_BASE_URL=http://127.0.0.1:5274 \
NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
pnpm exec playwright test
npm exec -- playwright test
```
The fast suite (8 specs, ~5s) runs by default; the long-running
two-segment audio-continuation spec is gated behind
@@ -146,7 +151,7 @@ PLAYWRIGHT_SKIP_WEBSERVER=1 \
PLAYWRIGHT_BASE_URL=http://127.0.0.1:5274 \
NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
PLAYWRIGHT_LONG_RUNNING=1 \
pnpm exec playwright test e2e/long-running-segments.spec.ts
npm exec -- playwright test e2e/long-running-segments.spec.ts
```
Expected runtime: ~7-9 minutes on a B200 (torch.compile max-autotune
@@ -185,17 +185,10 @@ CONDA_ENV_PYTHON="${DREAMVERSE_PYTHON:-${HOME}/miniconda3/envs/fv-main/bin/pytho
"${CONDA_ENV_PYTHON}" -c 'import flashinfer' 2>/dev/null \
|| bail "flashinfer-python not installed in ${CONDA_ENV_PYTHON} (run: ${CONDA_ENV_PYTHON} -m pip install flashinfer-python --no-build-isolation)"
PNPM="${PNPM:-}"
if [[ -n "${PNPM}" ]]; then
PNPM_REQUESTED="${PNPM}"
PNPM="$(command -v "${PNPM}" 2>/dev/null || true)"
[[ -n "${PNPM}" ]] || bail "pnpm not executable or not in PATH: ${PNPM_REQUESTED} (set PNPM to override)"
elif [[ -x "${HOME}/.local/share/pnpm/pnpm" ]]; then
PNPM="${HOME}/.local/share/pnpm/pnpm"
else
PNPM="$(command -v pnpm 2>/dev/null || true)"
fi
[[ -n "${PNPM}" ]] && [[ -x "${PNPM}" ]] || bail "pnpm not found. Set PNPM, install at ${HOME}/.local/share/pnpm/pnpm, or add pnpm to PATH"
NPM="${NPM:-npm}"
NPM_REQUESTED="${NPM}"
NPM="$(command -v "${NPM}" 2>/dev/null || true)"
[[ -n "${NPM}" ]] && [[ -x "${NPM}" ]] || bail "npm not executable or not in PATH: ${NPM_REQUESTED} (set NPM to override)"
GCC13="$(command -v "${GCC13:-gcc-13}" 2>/dev/null || true)"
GPP13="$(command -v "${GPP13:-g++-13}" 2>/dev/null || true)"
@@ -386,7 +379,7 @@ frontend_log="${LOG_DIR}/frontend-port${FRONTEND_PORT}.log"
# requested port differs, run `next dev --port` directly with devtools env.
fe_cmd="run dev:devtools"
if [[ "${FRONTEND_PORT}" != "5274" ]]; then
fe_cmd="exec next dev --port ${FRONTEND_PORT}"
fe_cmd="exec -- next dev --port ${FRONTEND_PORT}"
fi
setsid bash -c "
@@ -395,7 +388,7 @@ setsid bash -c "
export BACKEND_URL=http://127.0.0.1:${BACKEND_PORT}
export BACKEND_HOST=127.0.0.1
export BACKEND_PORT=${BACKEND_PORT}
exec '${PNPM}' ${fe_cmd}
exec '${NPM}' ${fe_cmd}
" > "${frontend_log}" 2>&1 < /dev/null &
disown
@@ -462,5 +455,5 @@ cat <<SUMMARY
BACKEND_URL=http://127.0.0.1:${BACKEND_PORT} \\
PLAYWRIGHT_BASE_URL=http://127.0.0.1:${FRONTEND_PORT} \\
NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \\
pnpm exec playwright test
npm exec -- playwright test
SUMMARY
+4
View File
@@ -6,4 +6,8 @@
{"name": "seed-ssim-references", "description": "Run a new or updated fastvideo/tests/ssim/ test on Modal, pull generated videos, and upload them to FastVideo/ssim-reference-videos so the test has a regression baseline", "path": "seed-ssim-references/SKILL.md", "status": "draft", "trust": "low"}
{"name": "reseed-ssim-references", "description": "Re-seed (overwrite) HF reference videos for an existing fastvideo/tests/ssim/ test and a single model id on Modal L40S. Always backs up current refs first, regenerates on Modal, pauses for the user to eyeball before-vs-after, then uploads with --force scoped to --model-id. Sister skill to seed-ssim-references; use when intentional code change has invalidated existing refs", "path": "reseed-ssim-references/SKILL.md", "status": "draft", "trust": "low"}
{"name": "decompose-pipeline-pr", "description": "Decompose an oversized FastVideo pipeline PR into a stack of independently-reviewable PRs. Tiers the diff by blast radius (invisible / dead code / cross-cutting infra / activation), produces a branch graph and worktree bootstrap, drafts the AGENTS.md manifest, flags missing tests on cross-cutting infra changes, and extracts lessons from the PR body. Worked example: PR #1280 daVinci-MagiHuman (9.8k LOC) decomposed into 10 stacked PRs.", "path": "decompose-pipeline-pr/SKILL.md", "status": "tested", "trust": "medium"}
{"name": "reseed-performance-baseline", "description": "Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, or environment-caused benchmark shift. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline must be advanced by replicating one reviewed shifted source result into three success=true records, or five records when explicitly requested", "path": "reseed-performance-baseline/SKILL.md", "status": "draft", "trust": "low"}
{"name": "add-model", "description": "Add a new model (or variant) to FastVideo: DiT + configs + pipeline + presets + registry + tests. Walks through FastVideo's single stage-based pipeline architecture with exact file paths and registration hooks.", "path": "add-model/SKILL.md", "status": "draft", "trust": "low"}
{"name": "rlhf-training-abstractions", "description": "Use when changing FastVideo RLHF/RL training infrastructure, especially sampler, reward, scheduler trajectory, or method boundaries under fastvideo/train.", "path": "rlhf-training-abstractions/SKILL.md", "status": "draft", "trust": "low"}
{"name": "add-rl-method", "description": "Use when adding or modifying an RL/RLHF method under fastvideo/train/methods/rl, including DiffusionNFT-like methods.", "path": "add-rl-method/SKILL.md", "status": "draft", "trust": "low"}
{"name": "add-reward-model", "description": "Use when adding reusable reward models under fastvideo/train/methods/rl/rewards for RLHF or online RL training.", "path": "add-reward-model/SKILL.md", "status": "draft", "trust": "low"}
+1 -1
View File
@@ -38,7 +38,7 @@ entrypoint, and applying defaults from the closest example script.
| `finetune` (Wan T2V) | `fastvideo/training/wan_training_pipeline.py` |
| `finetune` (Wan I2V) | `fastvideo/training/wan_i2v_training_pipeline.py` |
| `finetune` (LTX-2) | `fastvideo/training/ltx2_training_pipeline.py` |
| `finetune` (MatrixGame) | `fastvideo/training/matrixgame_training_pipeline.py` |
| `finetune` (Matrix-Game 2.0) | `fastvideo/training/matrixgame2_training_pipeline.py` |
| `distill-dmd` | `fastvideo/training/wan_distillation_pipeline.py` |
| `self-forcing` | `fastvideo/training/wan_self_forcing_distillation_pipeline.py` |
@@ -0,0 +1,476 @@
---
name: reseed-performance-baseline
description: Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, or environment-caused benchmark shift using one or more reviewed normalized performance JSONs. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline in FastVideo/performance-tracking must be advanced from a consistent batch of reviewed source results. The workflow backs up existing history under /tmp, validates all source JSONs for the same (model_id, gpu_type), rejects internally inconsistent source batches, uploads one success=true reseed record per accepted source JSON, and offers to clean local temp state after a successful upload.
---
# Re-seed Performance Baseline
## Purpose
Replace or advance the rolling performance baseline for a single
`(model_id, gpu_type)` pair in the HF dataset
`FastVideo/performance-tracking`.
Performance comparison uses the median of up to the last 5 successful records
for the same model and GPU. Failed records are useful audit history, but they
do not move the future baseline because `compare_baseline.py` loads records
with `successful_only=True`.
This skill now reseeds from a reviewed batch of one or more source performance
JSONs. It uploads one new `success=true` record per accepted source JSON; it
does not blindly replicate one measurement into 3 or 5 records. The effective
reseed size is therefore dynamic and equals the number of provided, validated,
internally consistent source JSONs.
If the operator provides fewer than 3 records, call out that the last-5 rolling
median may not move immediately. If the operator provides 3 consistent shifted
records, the rolling median usually moves immediately. If the operator provides
5 consistent shifted records, the last-5 window is effectively reset to the new
runtime profile.
These records are intentional operator-approved baseline resets, not ordinary
independent main-branch persistence. Mark them clearly with provenance fields
so the HF history remains auditable.
Use this skill when a performance test fails for an intentional and reviewed
reason, such as a torch/runtime/container upgrade that legitimately increases
peak memory or changes timings. This is the performance equivalent of
`reseed-ssim-references`: backup first, scope tightly, require explicit human
approval, then upload reviewed accepted baseline records.
## When to use
- A PR or main run failed the rolling performance comparison by more than the
allowed regression threshold, and maintainers agree the shift is caused by
an intentional runtime, dependency, hardware image, or benchmark environment
change rather than a FastVideo logic regression.
- One or more shifted source result JSONs have been reviewed and accepted, and
the operator wants to use those exact reviewed results to advance the rolling
baseline.
- The source batch is internally consistent: no provided source JSON regresses
against the batch median by more than the configured tolerance.
## When not to use
- The benchmark failure might be a real code regression. Fix or investigate
the code path first.
- The fixed benchmark thresholds in
`.buildkite/performance-benchmarks/tests/*.json` are too low. Those are a
separate gate from the rolling HF baseline and may need a code review change.
- There is no clear source run, commit, and rationale. Baseline history is a
production signal; do not edit it without provenance.
- The provided source JSONs disagree materially with each other. Rerun or
investigate instead of uploading a noisy reseed batch.
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| `model_id` | Yes | Benchmark id, e.g. `wan-t2v-1.3b-2gpu`. This maps to the HF subdirectory after `sanitize(model_id)`. |
| `gpu_type` | Yes | Exact GPU device string from the performance record, e.g. the L40S device name emitted by CI. Baselines are GPU-specific. |
| `source_results` | Yes | One or more local paths or Buildkite artifact URLs for accepted shifted performance JSONs. Prefer normalized `normalized_perf_*.json` artifacts emitted by `compare_baseline.py`. Accept `source_result` as an alias only for a single JSON. |
| `max_intra_batch_regression` | No | Maximum allowed regression of any source JSON against the source batch median. Default: `PERF_MAX_REGRESSION` if set, otherwise `0.05` (5%). |
| `intent_rationale` | Yes | One-line explanation for why the baseline shift is legitimate. This is written into provenance and should be reused in the PR. |
Hardcoded defaults:
- HF repo: `FastVideo/performance-tracking` (`HF_REPO_ID` override is
supported by the code, but use the default unless the user explicitly asks).
- Local sync root: `/tmp/perf-tracking` (`PERFORMANCE_TRACKING_ROOT` override
is supported).
- Backup root: `/tmp/performance_reseed_backup`.
- Download scratch root for source artifact URLs: `/tmp/performance_reseed_source`.
- Baseline window: last 5 `success=true` records for the same
`(model_id, gpu_type)`.
- Reseed count: dynamic. Upload exactly one accepted seed record per validated
source JSON.
## Steps
### 1. Validate the target and source results
Normalize `source_results` to a list. If the user passes a single
`source_result`, treat it as a one-element `source_results` list and report
that a single record may not move the last-5 median immediately.
If any source result is a Buildkite artifact URL, download it first into a
local scratch directory under `/tmp/performance_reseed_source/` and use the
downloaded JSON path for the rest of the workflow. If the agent cannot access
the artifact because Buildkite authentication is missing, ask the user to
download the artifact manually and provide the local path.
Prefer the normalized Buildkite artifact emitted by `compare_baseline.py`:
```text
perf_reports/results/normalized_perf_*.json
```
That file is already in the HF tracking schema. Load each normalized JSON
directly:
```python
import json
with open(source_result, encoding="utf-8") as f:
record = json.load(f)
```
Stop if any normalized record's `model_id` or `gpu_type` does not match the
requested `model_id` and `gpu_type`.
The source records may have `success: false` when they came from failed
rolling baseline comparisons. That is expected; only the reviewed reseed
records become new `success: true` baseline records after explicit approval.
Sort validated source records by their original `timestamp` ascending before
preparing the seed records. If a source timestamp is missing or unparsable,
preserve input order for those records and print a warning. This makes the
fresh reseed timestamps deterministic and makes it clear which records enter
the last-5 window when more than 5 source JSONs are provided.
Check that `HF_API_KEY` is exported. The sync path may be public, but the
upload path requires write access.
### 1a. Check source batch consistency
Before syncing or preparing uploads, reject source batches that are internally
inconsistent. Use the same metric direction as `compare_baseline.py`:
- Lower is better: `latency`, `memory`, `text_encoder_time_s`, `dit_time_s`,
`vae_decode_time_s`.
- Higher is better: `throughput`.
For each metric with at least two non-null source values:
1. Compute the source batch median.
2. For lower-is-better metrics, compute `(source_value - batch_median) / batch_median`.
3. For `throughput`, compute `(batch_median - source_value) / batch_median`.
4. Stop if any source record regresses against the batch median by more than
`max_intra_batch_regression`.
Default `max_intra_batch_regression` to `PERF_MAX_REGRESSION` when set,
otherwise `0.05`. Print a table with per-source values, batch median, and
worst intra-batch regression.
This check prevents uploading a mixed batch where one JSON is materially
slower or faster than the others. If the batch fails this check, ask the user
to provide a cleaner batch or explicitly investigate the variance. Do not
silently drop outliers unless the user gives a concrete reviewed reason and a
new source list.
### 1b. How to obtain source results from CI
The performance CI exports normalized source results for failed rolling
baseline comparisons when `compare_baseline.py` ran. The preferred artifacts
come from:
```text
perf_reports/results/normalized_perf_*.json
```
The normal operator flow is:
1. Open the failed Buildkite performance job or several reruns of the same
benchmark after the accepted environment shift.
2. Download the `normalized_perf_*.json` artifacts for the target benchmark.
3. Pass all reviewed local paths or artifact URLs as `source_results`.
Do not scrape the Markdown performance summary to reconstruct JSON. The
normalized JSON artifacts are the only supported source of truth for reseed
metrics and provenance. Raw `fastvideo/tests/performance/results/perf_*.json`
artifacts are not accepted by this skill. If no normalized JSON artifact is
present, that run is not a valid source for baseline reseeding.
### 2. Sync and back up existing HF records under /tmp
Use `fastvideo/tests/performance/hf_store.py` helpers directly. Do **not** use
`compare_baseline.py` as a sync shortcut; on full main runs it can persist
records, while this step must only fetch and back up existing history.
The sync command pattern is:
```bash
export PERFORMANCE_TRACKING_ROOT="${PERFORMANCE_TRACKING_ROOT:-/tmp/perf-tracking}"
export HF_REPO_ID="${HF_REPO_ID:-FastVideo/performance-tracking}"
PYTHONPATH=fastvideo/tests/performance python -c 'from hf_store import sync_from_hf; import os; sync_from_hf(os.environ["PERFORMANCE_TRACKING_ROOT"], strict=True)'
```
Then back up only the sanitized model directory under `/tmp`:
```bash
SHORT_COMMIT=$(git rev-parse --short=12 HEAD)
TIMESTAMP=$(date -u +%Y%m%d_%H%M%S)
MODEL_SAFE=$(PYTHONPATH=fastvideo/tests/performance python - <<'PY'
from hf_store import sanitize
print(sanitize("<model_id>"))
PY
)
BACKUP_DIR="/tmp/performance_reseed_backup/${TIMESTAMP}_${SHORT_COMMIT}_${MODEL_SAFE}"
mkdir -p "$BACKUP_DIR"
cp -R "${PERFORMANCE_TRACKING_ROOT}/${MODEL_SAFE}" "$BACKUP_DIR/" 2>/dev/null || true
```
Write provenance next to the backup:
```bash
cat > "$BACKUP_DIR/PROVENANCE.txt" <<EOF
model_id: <model_id>
gpu_type: <gpu_type>
source_results:
- <source_result_1>
- <source_result_2>
reseed_record_count: <len(source_results)>
max_intra_batch_regression: <threshold>
head_commit: $(git rev-parse HEAD)
timestamp_utc: $(date -u +%FT%TZ)
reason: <intent_rationale>
EOF
```
If the backup has no prior records, this is not a destructive reseed; it is a
first baseline seed. Continue, but report that baseline history was empty.
### 3. Compute old baseline and candidate shift
Load the last 5 successful records for the target:
```python
from hf_store import load_records_for_model
records = load_records_for_model(
"/tmp/perf-tracking",
"<model_id>",
"<gpu_type>",
last_n=5,
successful_only=True,
)
```
Print a small table showing old medians, source batch medians, candidate
medians after appending the proposed seed records, and source batch spread for:
- `latency`
- `throughput`
- `memory`
- `text_encoder_time_s`
- `dit_time_s`
- `vae_decode_time_s`
Also print how many successful old records exist. Make clear:
- 1 seed record usually does not move a last-5 median by itself.
- 3 consistent seed records usually move the last-5 median immediately.
- 5 consistent seed records effectively reset the last-5 window.
- The records are intentional approved baseline resets and must be labeled
that way.
### 4. Confirm intent
Require an explicit confirmation phrase before preparing the upload:
> About to RE-SEED performance baseline for `<model_id>` on `<gpu_type>`.
> This will upload `<N>` new `success=true` records to
> `FastVideo/performance-tracking/<sanitize(model_id)>/`, one per accepted
> source JSON.
>
> Reason: `<intent_rationale>`
> Source results: `<source_results>`
> Reseed record count: `<N>`
> Max intra-batch regression: `<threshold>`
> Note: these records come from a reviewed source batch and are intended to
> move the rolling median to the accepted runtime profile. They are not
> ordinary main-branch persistence.
> HEAD: `<git rev-parse --short=12 HEAD>`
> Backup: `<BACKUP_DIR>`
>
> Reply `confirm performance reseed` to proceed, anything else to abort.
Do not continue unless the user types exactly `confirm performance reseed`.
### 5. Create the accepted seed records
Create one seed record from each normalized source result. Do not copy the
source JSON wholesale.
Infer the baseline field allowlist from all existing HF records for the target
`(model_id, gpu_type)` after syncing, including both `success=true` and
`success=false` records. Use the union of non-provenance keys present in those
target records, preserving only fields that also exist in the normalized
source record or are explicitly set by the reseed workflow. Always include
`model_id`, `timestamp`, and `success` because the upload path and baseline
loader depend on them. Always set `timestamp` to a fresh reseed timestamp and
`success` to `true`. Do not include unrelated source-only fields that are
absent from existing HF records.
Exclude existing provenance or operator metadata from the inferred baseline
field allowlist. At minimum, exclude keys prefixed with `baseline_reseed` and
any fields known to be local-only audit metadata.
If there are no previous HF records for the target model/GPU, fall back to this
default baseline field list:
- `model_id`
- `timestamp`
- `commit_sha`
- `gpu_type`
- `latency`
- `throughput`
- `memory`
- `text_encoder_time_s`
- `dit_time_s`
- `vae_decode_time_s`
- `success`
Do not upload extra fields from the source artifact.
Optional provenance fields are allowed and useful:
- `baseline_reseed: true`
- `baseline_reseed_reason`
- `baseline_reseed_source_result`
- `baseline_reseed_source_timestamp`
- `baseline_reseed_source_success`
- `baseline_reseed_batch_size`
- `baseline_reseed_batch_index`
- `baseline_reseed_operator`
- `baseline_reseed_max_intra_batch_regression`
Use a fresh reseed timestamp for each seed record, not the original source
result timestamp. This is required because
`load_records_for_model(..., last_n=5)` keeps the last records after loading
the model directory; stale filenames/timestamps may not enter the last-5
window and therefore may not move the median. Preserve the original source
timestamp in `baseline_reseed_source_timestamp`.
Use the existing filename convention from `_write_tracking_record()`:
`<sanitize(timestamp)>_<sanitize(commit_sha)>.json` under the sanitized model
directory, but include a deterministic suffix such as `_reseed_01`,
`_reseed_02`, and so on before `.json` so multiple records from the same
batch do not overwrite each other.
If a source record already exists on HF with `success=false`, do not edit it
in place unless the user explicitly asked for an audit-preserving correction.
Prefer uploading new accepted seed records so failed history remains visible.
### 6. Pause before upload
Print:
- Backup directory path under `/tmp`.
- Prepared local record paths under `PERFORMANCE_TRACKING_ROOT`.
- HF paths that will receive the new records.
- Old rolling medians.
- Source batch medians, source batch spread, reseed count, and candidate
medians.
- Rationale.
Ask the user to reply exactly `upload`. Anything else aborts and leaves the
prepared records plus backup on disk.
### 7. Upload only the scoped records
Use the shared storage helper so the path and repo type match CI:
```python
from hf_store import upload_record
upload_record("<local_record_path>", record, strict=True)
```
Run it once per prepared record. Each upload goes to:
```text
FastVideo/performance-tracking/<sanitize(model_id)>/<record_filename>.json
```
Never bulk upload the whole tracking root. Never modify another model's
directory in the same operation.
### 8. Report outcome and offer cleanup
Report:
- Uploaded HF paths.
- Backup directory under `/tmp`.
- Local tracking root, usually `/tmp/perf-tracking`.
- Old baseline window count and medians.
- Source batch medians, source batch spread, reseed count, and candidate
medians.
- Expected effect based on reseed count.
- Any separate threshold changes still needed in
`.buildkite/performance-benchmarks/tests/*.json`.
Include the `intent_rationale` in the PR or follow-up comment so reviewers can
distinguish an accepted baseline shift from a hidden regression.
After the upload is verified, ask whether the user wants to clear temporary
local state. Explain what each directory is for:
- `PERFORMANCE_TRACKING_ROOT`, usually `/tmp/perf-tracking`: local synced
mirror of `FastVideo/performance-tracking` plus the prepared local seed
records used for scoped upload.
- `/tmp/performance_reseed_backup/<...>`: local backup of the target model's
pre-reseed HF history plus `PROVENANCE.txt`, kept so a bad reseed can be
audited or corrected.
- `/tmp/performance_reseed_source/<...>` when used: downloaded source JSON
artifacts from Buildkite URLs.
Ask:
> Reseed succeeded. Do you want me to delete the local temp tracking mirror,
> source downloads, and reseed backup under `/tmp`? These files are local
> safety/audit artifacts only; HF already has the uploaded records.
>
> Reply `cleanup reseed temp` to delete them, anything else to keep them.
Do not delete anything unless the user replies exactly
`cleanup reseed temp`. If cleanup is requested, remove only the specific
directories created for this reseed. Never remove unrelated `/tmp` contents.
## Failure modes and handling
- **`HF_API_KEY` unset.** Stop before upload. Do not create an untracked
process that appears to have reseeded but never reached HF.
- **Source result does not match target.** Stop. The wrong benchmark or GPU
would poison a separate baseline.
- **Source batch is internally inconsistent.** Stop if any source regresses
against the source batch median by more than `max_intra_batch_regression`.
Ask for cleaner sources or a reviewed explanation before continuing.
- **Too few source records to move the median.** Continue only after making
clear that one or two records may not immediately move the last-5 median.
- **The source results are noisy or suspicious.** Stop. Reseeding amplifies
those measurements into the baseline, so they must be reviewed first.
- **HF sync fails.** Stop for destructive reseeds. A stale or empty sync can
make the old baseline look missing.
- **Candidate still violates fixed thresholds.** Report that this skill only
handles the rolling HF baseline; update benchmark JSON thresholds in code
review if maintainers accept the new absolute limit.
- **The user aborts at either confirmation.** Leave the backup and prepared
records on disk. Nothing should be uploaded.
- **The user declines cleanup.** Keep `/tmp/perf-tracking`, the source
download directory if any, and `/tmp/performance_reseed_backup/<...>` in
place for audit/debugging.
- **A bad seed was uploaded.** Use the backup and HF history to identify the
uploaded file, then remove or supersede it with an explicitly reviewed
corrective record. Do not silently rewrite unrelated history.
## References
- `.agents/skills/reseed-ssim-references/SKILL.md` — safety pattern for
intentional baseline replacement.
- `fastvideo/tests/performance/compare_baseline.py` — normalization, rolling
median comparison, and persistence rules.
- `fastvideo/tests/performance/hf_store.py` — HF sync, record loading,
`sanitize()`, and `upload_record()`.
- `fastvideo/tests/performance/test_inference_performance.py` — source result
JSON schema.
- `.buildkite/performance-benchmarks/tests/*.json` — fixed absolute benchmark
thresholds, separate from rolling baseline comparisons.
## Changelog
| Date | Change |
|------|--------|
| 2026-05-03 | Initial version. Sister workflow to `reseed-ssim-references`, scoped to one performance `(model_id, gpu_type)` baseline seed with backup, confirmation, provenance, and `success=true` upload. |
| 2026-05-03 | Previous policy: replicate one approved shifted source result into 3 success records by default, or 5 only when explicitly requested. Add provenance marker for replicated-source reseeds. Superseded by the 2026-05-08 dynamic multi-source policy. |
| 2026-05-08 | Replace fixed 3/5 replication with dynamic multi-source reseeding: upload one seed record per reviewed source JSON, validate intra-batch consistency, move backup/source scratch under `/tmp`, and ask whether to clean temp state after successful upload. |
@@ -0,0 +1,41 @@
---
name: rlhf-training-abstractions
description: Use when changing FastVideo RLHF/RL training infrastructure, especially sampler, reward, scheduler trajectory, or method boundaries under fastvideo/train.
---
# RLHF Training Abstractions
Use this skill before editing RLHF-style training code in `fastvideo/train`.
## Boundaries
- RL methods live under `fastvideo/train/methods/rl/` and own algorithm logic: reward collection, advantage computation, policy loss, KL/reference terms, and optimizer cadence.
- Rewards live under `fastvideo/train/methods/rl/rewards/` and must be reusable across RL methods.
- RL methods pass decoded media to rewards; each reward decides whether to use the first frame, sampled frames, or the full video.
- Sampling lives under `fastvideo/train/methods/rl/common/` and must use `ModelBase` primitives plus scheduler math, not model-family inference pipelines.
- Model wrappers under `fastvideo/train/models/` own model-specific forward details.
- Model wrappers also own model-specific latent decoding via `ModelBase.decode_latents`; RL methods should not reach into VAE normalization internals.
- Shared RL helpers such as K-repeat prompt sampling belong under `fastvideo/train/methods/rl/common/` when they are reusable across RL methods.
## Anti-Patterns
- Do not bind RL methods to inference pipeline classes such as `WanDMDPipeline`.
- Do not hardcode timestep lists in a method when the scheduler can generate them.
- Do not put reward-model code inside one RL method.
- Do not make existing non-RL methods use method-managed optimization unless explicitly requested.
## Sampling Policy
- Prefer YAML-configured `method.sampling` with `scheduler`, `trajectory`, `num_steps`, `timesteps`, and `sigmas`.
- Treat diffusers-style scheduler classes as owning both the timestep schedule and their `step()` update rule; avoid a separate `solver` field unless a new sampler truly implements solver math outside the scheduler object.
- Missing `timesteps` means “ask the scheduler”; explicit `timesteps` or `sigmas` are overrides.
- ODE-style trajectories should not re-noise between denoising steps.
- SDE/re-noise behavior must be explicit in config.
## Validation
- Run focused local tests for sampler config and Trainer opt-in behavior.
- Verify existing train methods still report `manages_optimization() == False`.
- Keep fixed-prompt validation helpers in `fastvideo/train/methods/rl/common/validation.py` so new RL methods can reuse sharding and captions.
- Test distributed prompt grouping helpers separately from heavyweight model loading.
- Run `pre-commit run --files <changed paths>`; respect configured excludes.
@@ -89,5 +89,4 @@ The following land in follow-up PRs:
- **MIND** metrics (depends on a separate `vipe` submodule).
- **VBench-2.0** sibling package.
- Native conversion of **FVD** under `fastvideo/eval/metrics/fvd/`.
- The training-time `EvalCallback`.
@@ -36,7 +36,10 @@
"thresholds": {
"L40S": {
"max_generation_time_s": 34.0,
"max_peak_memory_mb": 11000.0
"max_peak_memory_mb": 11000.0,
"max_text_encoder_time_s": 5.0,
"max_dit_time_s": 10.0,
"max_vae_decode_time_s": 10.0
},
"default": {
"max_generation_time_s": 120.0,
+21
View File
@@ -77,6 +77,17 @@ steps:
limit: 2
agents:
queue: "default"
- label: ":microscope: DreamVerse App Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "dreamverse_app"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
# --- Full-suite-scope direct tests ---
- label: ":bar_chart: SSIM Tests"
@@ -289,6 +300,16 @@ steps:
- TEST_TYPE=unit_test
agents:
queue: "default"
- path:
- "apps/dreamverse/**"
- "pyproject.toml"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":microscope: DreamVerse App Tests"
env:
- TEST_TYPE=dreamverse_app
agents:
queue: "default"
# ============================================================
# Full Suite: Runs when TEST_SCOPE=full
+18
View File
@@ -121,6 +121,19 @@ upload_performance_artifacts() {
fi
}
_upload_normalized_perf_results() {
local found=0
while IFS= read -r -d '' target; do
found=1
log "Found normalized performance result: $target. Uploading to Buildkite..."
buildkite-agent artifact upload "$target"
done < <(find "$LOCAL_DIR" -path "*/results/normalized_perf_*.json" -print0)
if [ "$found" -eq 0 ]; then
log "No normalized performance result artifacts found. This is expected when the rolling performance comparison did not run."
fi
}
_cleanup_modal_volume() {
log "Cleaning up perf_reports/ from Modal Volume..."
if modal volume rm hf-model-weights "perf_reports/" --recursive; then
@@ -139,6 +152,7 @@ upload_performance_artifacts() {
_download_reports || { _cleanup_local; return 1; }
_upload_dashboard
_upload_perf_summary
_upload_normalized_perf_results
_cleanup_modal_volume
_cleanup_local
}
@@ -197,6 +211,10 @@ case "$TEST_TYPE" in
log "Running unit tests..."
MODAL_COMMAND="$MODAL_ENV python3 -m modal run $MODAL_TEST_FILE::run_unit_test"
;;
"dreamverse_app")
log "Running DreamVerse app tests..."
MODAL_COMMAND="$MODAL_ENV python3 -m modal run $MODAL_TEST_FILE::run_dreamverse_app_tests"
;;
"train_framework")
log "Running fastvideo.train framework tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_train_framework_tests"
+23 -7
View File
@@ -12,6 +12,18 @@ on:
tag_suffix:
required: true
type: string
image_name:
required: false
type: string
default: fastvideo-dev
build_args:
required: false
type: string
default: ''
include_latest_tags:
required: false
type: boolean
default: true
jobs:
build-and-push:
@@ -67,11 +79,14 @@ jobs:
run: |
SHORT_SHA=$(echo ${{ github.sha }} | cut -c1-7)
TAGS="type=raw,value=${{ inputs.tag_suffix }}-latest"
TAGS="${TAGS}\ntype=raw,value=${{ inputs.tag_suffix }}-sha-${SHORT_SHA}"
TAGS="type=raw,value=${{ inputs.tag_suffix }}-sha-${SHORT_SHA}"
if [[ "${{ inputs.include_latest_tags }}" == "true" ]]; then
TAGS="type=raw,value=${{ inputs.tag_suffix }}-latest\n${TAGS}"
fi
# Set Python 3.10 as the default image
if [[ "${{ inputs.python_version }}" == "3.10" ]]; then
if [[ "${{ inputs.include_latest_tags }}" == "true" && "${{ inputs.python_version }}" == "3.10" ]]; then
TAGS="${TAGS}\ntype=raw,value=latest"
fi
@@ -85,7 +100,7 @@ jobs:
id: meta
uses: docker/metadata-action@v5
with:
images: ghcr.io/${{ github.repository }}/fastvideo-dev
images: ghcr.io/${{ github.repository }}/${{ inputs.image_name }}
tags: ${{ steps.prepare-tags.outputs.tags }}
- name: Build and push Docker image
@@ -97,10 +112,11 @@ jobs:
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
build-args: ${{ inputs.build_args }}
cache-from: type=gha
cache-to: type=gha,mode=max
- name: Success message
run: |
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ghcr.io/${{ github.repository }}/fastvideo-dev:${{ inputs.tag_suffix }}-latest"
echo "To run tests with this image, manually trigger the 'Run Tests' workflow."
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ghcr.io/${{ github.repository }}/${{ inputs.image_name }}:${{ inputs.tag_suffix }}-sha-${GITHUB_SHA::7}"
echo "To run tests with this image, manually trigger the 'Run Tests' workflow."
+2 -2
View File
@@ -125,7 +125,7 @@ jobs:
set -euo pipefail
TEST_NAME=$(echo "$COMMENT" | grep -oP '(?<=/test\s)\S+' | head -1 || true)
VALID="encoder vae transformer kernel unit ssim training lora-inference lora-training distillation self-forcing vsa vmoba performance api train-framework full fastcheck pre-commit"
VALID="encoder vae transformer kernel unit dreamverse ssim training lora-inference lora-training distillation self-forcing vsa vmoba performance api train-framework full fastcheck pre-commit"
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
echo "Unknown test: '$TEST_NAME'. Valid: $VALID"
exit 1
@@ -133,7 +133,7 @@ jobs:
declare -A MAP=(
[encoder]=encoder [vae]=vae [transformer]=transformer
[kernel]=kernel_tests [unit]=unit_test
[kernel]=kernel_tests [unit]=unit_test [dreamverse]=dreamverse_app
[ssim]=ssim [training]=training
[lora-inference]=inference_lora [lora-training]=training_lora
[distillation]=distillation_dmd [self-forcing]=self_forcing
+29
View File
@@ -23,6 +23,11 @@ on:
required: false
default: false
type: boolean
dreamverse_cuda_12_9:
description: 'Build Dreamverse CUDA 12.9 backend-only and UI images'
required: false
default: false
type: boolean
permissions:
@@ -65,3 +70,27 @@ jobs:
dockerfile_path: docker/Dockerfile.python3.12.cuda12.9.1
tag_suffix: py3.12-cuda12.9.1
secrets: inherit
build-dreamverse-backend-cuda-12-9:
if: ${{ github.event.inputs.dreamverse_cuda_12_9 == 'true' }}
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.12'
dockerfile_path: apps/dreamverse/docker/Dockerfile
tag_suffix: dreamverse-backend-cuda12.9.1
image_name: dreamverse
build_args: BUILD_DREAMVERSE_UI=0
include_latest_tags: false
secrets: inherit
build-dreamverse-ui-cuda-12-9:
if: ${{ github.event.inputs.dreamverse_cuda_12_9 == 'true' }}
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.12'
dockerfile_path: apps/dreamverse/docker/Dockerfile
tag_suffix: dreamverse-ui-cuda12.9.1
image_name: dreamverse
build_args: BUILD_DREAMVERSE_UI=1
include_latest_tags: false
secrets: inherit
+2
View File
@@ -34,6 +34,8 @@ env
*.log
weights/
logs/
official_weights/
converted_weights/
# SSIM test outputs
fastvideo/tests/ssim/generated_videos/
+15
View File
@@ -42,6 +42,8 @@ FastVideo has the following features:
- Support H100, A100, 4090
- Support Linux, Windows, MacOS
- See this [page](https://hao-ai-lab.github.io/FastVideo/inference/support_matrix/) for full list of supported models, hardware assumptions, and optimization compatibility.
- Realtime video generation & editing
- [Dreamverse](apps/dreamverse/README.md): stream and "vibe direct" video in realtime ([live demo](https://dreamverse.fastvideo.org/)), deployable on local GPU, a self-hosted B200 server, Docker, or serverless Modal
## Getting Started
@@ -69,6 +71,19 @@ See below for recipes and datasets:
| [FastWan2.1-T2V-1.3B](https://huggingface.co/FastVideo/FastWan2.1-T2V-1.3B-Diffusers) | [Recipe](https://github.com/hao-ai-lab/FastVideo/tree/main/examples/distill/Wan2.1-T2V/Wan-Syn-Data-480P) | [FastVideo Synthetic Wan2.1 480P](https://huggingface.co/datasets/FastVideo/Wan-Syn_77x448x832_600k) |
| [FastWan2.2-TI2V-5B](https://huggingface.co/FastVideo/FastWan2.2-TI2V-5B-Diffusers) | [Recipe](https://github.com/hao-ai-lab/FastVideo/tree/main/examples/distill/Wan2.2-TI2V-5B-Diffusers/Data-free) | [FastVideo Synthetic Wan2.2 720P](https://huggingface.co/datasets/FastVideo/Wan2.2-Syn-121x704x1280_32k) |
## Dreamverse — Realtime Video Generation & Editing
[Dreamverse](apps/dreamverse/README.md) is FastVideo's realtime video generation
and editing platform — "vibe directing" a video as it streams. It lives in the
monorepo under [`apps/dreamverse/`](apps/dreamverse/) and ships its own backend
(`dreamverse-server`) plus a web UI.
Try the [live demo](https://dreamverse.fastvideo.org/), read the
[blog](https://haoailab.com/blogs/dreamverse/), or run it yourself. Dreamverse
deploys on a local GPU, a self-hosted B200 server over SSH, Docker, or
serverless [Modal](apps/dreamverse/scripts/modal/README.md) — see the
[Dreamverse README](apps/dreamverse/README.md).
## Inference
### Generating Your First Video
+68 -11
View File
@@ -2,6 +2,8 @@
Dreamverse is the FastVideo realtime video generation & editing platform. It lives in this monorepo under `apps/dreamverse/`.
**Deploy on:** [local GPU](#quick-start-local-gpu) · [self-hosted B200 (SSH)](#server-b200-deployment-ssh) · [Docker](docker/README.md) · [Modal](scripts/modal/README.md)
## Install Dreamverse
You can install Dreamverse using one of the methods below.
@@ -43,13 +45,36 @@ See `apps/dreamverse/docker/README.md` for Docker build and run option details.
## Optional: Building FFmpeg For Better Performance
For full streaming performance in a non-Docker install, build a custom FFmpeg
binary:
binary from a FastVideo source checkout. The command below is repo-relative,
so run it from the repository root:
```bash
bash apps/dreamverse/scripts/install_native_ffmpeg.sh
```
This builds and installs into `~/opt/ffmpeg-native/` and writes
The installer supports Linux `x86_64` and `aarch64`. It prefers conda-forge
triplet compilers when those commands are on `PATH`, otherwise it falls back to
system `gcc`/`g++` (plain venv). On `x86_64`, x264's hand-tuned SIMD also
requires `nasm`; install via whichever path fits your host:
```bash
sudo apt install nasm # Debian/Ubuntu
conda install -c conda-forge nasm # inside an active conda env
```
No sudo and no conda? Build `nasm` from source (~30s, installs into `$HOME`):
```bash
(
mkdir -p "$HOME/src" "$HOME/opt" && cd "$HOME/src"
curl -fsSL -O https://www.nasm.us/pub/nasm/releasebuilds/2.16.03/nasm-2.16.03.tar.gz
tar -xf nasm-2.16.03.tar.gz && cd nasm-2.16.03
./configure --prefix="$HOME/opt/nasm" && make -j"$(nproc)" && make install
)
export PATH="$HOME/opt/nasm/bin:$PATH" # add to ~/.bashrc to persist
```
The installer writes to `~/opt/ffmpeg-native/` and emits
`apps/dreamverse/scripts/ffmpeg-env.sh`. Source it before starting the backend
so Dreamverse uses the custom FFmpeg binary:
@@ -59,8 +84,7 @@ dreamverse-server
```
Docker images already run this FFmpeg build during image creation and source the
generated environment file at container startup. The installer supports Linux
`x86_64` and `aarch64`.
generated environment file at container startup.
## Launch Dreamverse
@@ -71,17 +95,24 @@ dreamverse-server --port 8009
dreamverse-mock-server --port 8009
```
> **Expect a slow first boot.** With `torch.compile` and startup warmup enabled
> (the default), the backend compiles the segment 1 and segment 2 inference
> paths before it reports ready — this can take **tens of minutes on a cold
> cache**, regardless of how you deploy (local, server, Docker, or Modal).
> `/healthz` responds as soon as the process is up; `/readyz` stays `503` until
> warmup finishes. For a faster, uncompiled startup while testing, set
> `FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend.
## Frontend Setup
Install the web dependencies once from the FastVideo checkout:
```bash
cd apps/dreamverse/web
pnpm install --frozen-lockfile
npm ci
```
The frontend package also has an npm lockfile, but the bundled launch scripts
use `pnpm`.
The frontend package uses `package-lock.json`; use npm for installs and scripts.
## Quick Start: Local GPU
@@ -137,11 +168,37 @@ Start the frontend:
```bash
cd apps/dreamverse/web
BACKEND_HOST=localhost BACKEND_PORT=8009 pnpm run dev
BACKEND_HOST=localhost BACKEND_PORT=8009 npm run dev
```
Open `http://localhost:5299`.
## Server B200 deployment (SSH)
Deploying on a remote GPU host (for example a B200 box) is a local install run
over SSH, plus a few server-specific concerns. Two paths:
### Option A: Native (source install)
SSH in, then follow [Install → From source](#method-2-from-source) and
(recommended) [Building FFmpeg](#optional-building-ffmpeg-for-better-performance),
then start the backend as in [Quick Start: Local GPU](#quick-start-local-gpu).
For a remote host, a few things differ from localhost:
- Bind all interfaces: `dreamverse-server --host 0.0.0.0 --port 8009`.
- Point the frontend/client at the host: `BACKEND_HOST=<b200-host> BACKEND_PORT=8009 npm run dev`.
- Keep the backend alive across SSH sessions (`tmux` / `systemd` / `nohup`).
- Expose / firewall port `8009`, or front it with a reverse proxy + auth.
### Option B: Docker (on the server)
SSH in, then follow [Install → Using Docker](#method-3-using-docker) and the run
steps in [`docker/README.md`](docker/README.md):
```bash
CEREBRAS_API_KEY="<key>" GROQ_API_KEY="<key>" apps/dreamverse/docker/docker_run.sh
```
## Quick Start: Mock Backend (For UI development)
The mock server emulates the Dreamverse backend protocol and streams a
@@ -173,14 +230,14 @@ Run the frontend tests:
```bash
cd apps/dreamverse/web
pnpm test
npm test
```
Run the frontend e2e tests:
```bash
cd apps/dreamverse/web
pnpm run e2e
npm run e2e
```
## Troubleshooting
@@ -216,7 +273,7 @@ Only one GPU is used
Frontend cannot connect to backend
- confirm the backend is running on `8009`; if not, point the frontend at it
with `BACKEND_HOST=<host> BACKEND_PORT=<port> pnpm run dev`
with `BACKEND_HOST=<host> BACKEND_PORT=<port> npm run dev`
- confirm `http://localhost:8009/healthz` responds before starting the frontend
- confirm `http://localhost:8009/readyz` returns `200` before clicking Generate
- use `apps/dreamverse/scripts/smoke_local.sh` for a repeatable local startup
-40
View File
@@ -1,40 +0,0 @@
.git/
**/.git/
.venv/
**/.venv/
**/__pycache__/
**/*.pyc
**/*.pyo
**/*.egg-info/
.pytest_cache/
**/.pytest_cache/
.mypy_cache/
.ruff_cache/
.cache/
apps/dreamverse/web/node_modules/
apps/dreamverse/web/.next/
apps/dreamverse/web/out/
apps/dreamverse/web/dist/
apps/dreamverse/web/test-results/
apps/dreamverse/web/playwright-report/
apps/dreamverse/outputs/
apps/dreamverse/dreamverse/outputs/
apps/dreamverse/dreamverse/prompts.local/
apps/dreamverse/logs/
outputs/
logs/
slurm-logs/
wandb/
.env
.env.*
**/prompts.local/
.codex/
.agents/exploration/
.vscode/
.idea/
*.log
*.tmp
*.pdf
+13
View File
@@ -3,6 +3,7 @@ ARG CUDA_TAG=12.9.1-cudnn-devel-ubuntu22.04
FROM nvidia/cuda:${CUDA_TAG}
ARG BUILD_FASTVIDEO_KERNEL_FROM_SOURCE=0
ARG BUILD_DREAMVERSE_UI=0
ENV DEBIAN_FRONTEND=noninteractive \
PYTHONUNBUFFERED=1 \
@@ -50,6 +51,18 @@ RUN if [[ "${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE}" == "1" ]]; then \
echo "Skipping source fastvideo-kernel build; using installed fastvideo-kernel package."; \
fi
RUN if [[ "${BUILD_DREAMVERSE_UI}" == "1" ]]; then \
curl -fsSL https://deb.nodesource.com/setup_22.x | bash - \
&& apt-get install -y --no-install-recommends nodejs \
&& rm -rf /var/lib/apt/lists/* \
&& cd /opt/FastVideo/apps/dreamverse/web \
&& npm ci --ignore-scripts \
&& NEXT_OUTPUT_EXPORT=1 npm run build \
&& rm -rf node_modules .next; \
else \
echo "Skipping Dreamverse UI build; Building backend-only image."; \
fi
# The monorepo ffmpeg installer force-selects conda compiler triplets for
# local dev shells. Inside this image we explicitly opt into the system
# gcc/g++ toolchain.
@@ -2,6 +2,7 @@
**/.git/
.venv/
**/.venv/
.*dreamverse*/
**/__pycache__/
**/*.pyc
**/*.pyo
@@ -24,6 +25,9 @@ apps/dreamverse/dreamverse/outputs/
apps/dreamverse/dreamverse/prompts.local/
apps/dreamverse/logs/
outputs/
outputs_video/
quality_check_outputs/
data/
logs/
slurm-logs/
wandb/
@@ -32,7 +36,9 @@ wandb/
.env.*
**/prompts.local/
.codex/
.agents/exploration/
.agents/
.opencode/
.agent_tmp/
.vscode/
.idea/
*.log
+14 -4
View File
@@ -1,9 +1,9 @@
# Dreamverse Docker Image
This folder contains the backend-only Docker image for Dreamverse inside the
FastVideo monorepo. Build commands use the FastVideo repository root as the
Docker context, so run the helper scripts from this folder or from any path in
the checkout.
This folder contains the Docker image for Dreamverse inside the FastVideo
monorepo. Build commands use the FastVideo repository root as the Docker
context, so run the helper scripts from this folder or from any path in the
checkout.
## Build
@@ -17,6 +17,16 @@ The image defaults to `dreamverse:dev`. Override it with:
DREAMVERSE_IMAGE=dreamverse:local apps/dreamverse/docker/docker_build.sh
```
Backend-only remains the default image. To include the static Dreamverse UI
served by the backend, set `BUILD_DREAMVERSE_UI=1` and choose a specific image
tag:
```bash
BUILD_DREAMVERSE_UI=1 DREAMVERSE_IMAGE=<image-tag> apps/dreamverse/docker/docker_build.sh
```
Prefer SHA-specific tags for deployable images; avoid `latest`.
The Dockerfile builds a CUDA 12.9.1 image, installs FastVideo from this
checkout with the `dreamverse` extra, installs the FA4
flash-attention fork, builds native FFmpeg, and installs FlashInfer for NVFP4
+32 -1
View File
@@ -4,13 +4,44 @@ set -euo pipefail
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd -- "${SCRIPT_DIR}/../../.." && pwd)"
IMAGE="${DREAMVERSE_IMAGE:-dreamverse:dev}"
ROOT_DOCKERIGNORE="${REPO_ROOT}/.dockerignore"
DOCKERFILE_DOCKERIGNORE="${SCRIPT_DIR}/Dockerfile.dockerignore"
CREATED_ROOT_DOCKERIGNORE_SYMLINK=0
cleanup_root_dockerignore_symlink() {
if [[ "${CREATED_ROOT_DOCKERIGNORE_SYMLINK}" == "1" && -L "${ROOT_DOCKERIGNORE}" ]] && \
[[ "$(readlink "${ROOT_DOCKERIGNORE}")" == "${DOCKERFILE_DOCKERIGNORE}" ]]; then
rm -- "${ROOT_DOCKERIGNORE}"
fi
}
trap cleanup_root_dockerignore_symlink EXIT INT TERM
if [[ -z "${DOCKER_BUILDKIT:-}" ]] && docker buildx version >/dev/null 2>&1; then
export DOCKER_BUILDKIT=1
fi
build_args=()
[[ -n "${CUDA_TAG:-}" ]] && build_args+=(--build-arg "CUDA_TAG=${CUDA_TAG}")
[[ -n "${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE:-}" ]] && \
build_args+=(--build-arg "BUILD_FASTVIDEO_KERNEL_FROM_SOURCE=${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE}")
build_args+=(--build-arg "BUILD_DREAMVERSE_UI=${BUILD_DREAMVERSE_UI:-0}")
exec docker build \
if [[ "${DOCKER_BUILDKIT:-}" == "1" ]]; then
if [[ -L "${ROOT_DOCKERIGNORE}" ]] && \
[[ "$(readlink "${ROOT_DOCKERIGNORE}")" == "${DOCKERFILE_DOCKERIGNORE}" ]]; then
rm -- "${ROOT_DOCKERIGNORE}"
printf 'Removed stale temporary root .dockerignore symlink: %s\n' "${ROOT_DOCKERIGNORE}"
fi
elif [[ -e "${ROOT_DOCKERIGNORE}" || -L "${ROOT_DOCKERIGNORE}" ]]; then
printf 'Using existing root .dockerignore: %s\n' "${ROOT_DOCKERIGNORE}"
else
ln -s -- "${DOCKERFILE_DOCKERIGNORE}" "${ROOT_DOCKERIGNORE}"
CREATED_ROOT_DOCKERIGNORE_SYMLINK=1
printf 'Created temporary root .dockerignore symlink for legacy Docker builder: %s -> %s\n' \
"${ROOT_DOCKERIGNORE}" "${DOCKERFILE_DOCKERIGNORE}"
fi
docker build \
-f "${SCRIPT_DIR}/Dockerfile" \
-t "${IMAGE}" \
"${build_args[@]}" \
+101 -3
View File
@@ -4,6 +4,7 @@ from pathlib import Path
_REPO_ROOT = Path(__file__).resolve().parents[1]
_SERVER_ROOT = Path(__file__).resolve().parent
_FASTVIDEO_DREAMVERSE_HOME = os.environ.get("FASTVIDEO_DREAMVERSE_HOME")
_FASTVIDEO_DREAMVERSE_FRONTEND_ROOT = os.environ.get("FASTVIDEO_DREAMVERSE_FRONTEND_ROOT")
_XDG_STATE_HOME = os.environ.get("XDG_STATE_HOME")
_DEFAULT_STATE_ROOT = (Path(_FASTVIDEO_DREAMVERSE_HOME) if _FASTVIDEO_DREAMVERSE_HOME else
(Path(_XDG_STATE_HOME) if _XDG_STATE_HOME else Path.home() / ".local/state") /
@@ -16,18 +17,39 @@ _APP_ROOT = _REPO_ROOT
def _resolve_frontend_root() -> Path:
for candidate in (
Path(_FASTVIDEO_DREAMVERSE_FRONTEND_ROOT) if _FASTVIDEO_DREAMVERSE_FRONTEND_ROOT else None,
_APP_ROOT / "web",
_APP_ROOT / "prod-ui",
):
if candidate.is_dir():
if candidate is not None and candidate.is_dir():
return candidate
if _FASTVIDEO_DREAMVERSE_FRONTEND_ROOT:
return Path(_FASTVIDEO_DREAMVERSE_FRONTEND_ROOT)
return _APP_ROOT / "web"
def _resolve_frontend_static_dir_candidates() -> tuple[str, ...]:
roots = (
FRONTEND_ROOT,
Path.cwd() / "apps/dreamverse/web",
Path.cwd() / "web",
Path("/opt/FastVideo/apps/dreamverse/web"),
)
candidates: list[str] = []
seen: set[Path] = set()
for root in roots:
resolved_root = root.resolve(strict=False)
if resolved_root in seen:
continue
seen.add(resolved_root)
candidates.extend(str(resolved_root / dirname) for dirname in ("out", "dist"))
return tuple(candidates)
FRONTEND_ROOT = _resolve_frontend_root()
_CLIENT_PROMPTS_ROOT = FRONTEND_ROOT / "prompts"
_CLIENT_PROMPTS_LOCAL_ROOT = FRONTEND_ROOT / "prompts.local"
FRONTEND_STATIC_DIR_CANDIDATES = tuple(str(FRONTEND_ROOT / dirname) for dirname in ("out", "dist"))
FRONTEND_STATIC_DIR_CANDIDATES = _resolve_frontend_static_dir_candidates()
# Model registry
MODEL_REGISTRY = {
@@ -35,18 +57,24 @@ MODEL_REGISTRY = {
"name": "FastLTX2",
"model_path": "FastVideo/LTX2-Distilled-Diffusers",
"config_model_path": "FastVideo/LTX2-Distilled-Diffusers",
"lora_repo": "FastVideo/LTX2-OmniNFT-LoRA",
},
"fast-ltx23": {
"name": "FastLTX23",
"model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
"config_model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
"lora_repo": "FastVideo/LTX-2.3-OmniNFT-LoRA",
},
}
DEFAULT_MODEL_ID = "fast-ltx2"
ACTIVE_MODEL_ID = (os.getenv("DREAMVERSE_MODEL_ID", "").strip() or DEFAULT_MODEL_ID)
if ACTIVE_MODEL_ID not in MODEL_REGISTRY:
ACTIVE_MODEL_ID = DEFAULT_MODEL_ID
# Active model configuration
MODEL_CONFIG = MODEL_REGISTRY[DEFAULT_MODEL_ID]
MODEL_CONFIG = MODEL_REGISTRY[ACTIVE_MODEL_ID]
# Generation limits
SESSION_TIMEOUT_SECONDS = 300
@@ -142,6 +170,76 @@ def _optional_env(*names: str) -> str | None:
DEVTOOLS_ENABLED = _env_bool("FASTVIDEO_ENABLE_DEVTOOLS", False)
PROMPT_SAFETY_ENABLED = _env_bool("FASTVIDEO_ENABLE_PROMPT_SAFETY", False)
DREAMVERSE_MAX_AUTOTUNE = _env_bool("DREAMVERSE_MAX_AUTOTUNE", True)
DREAMVERSE_SP_SIZE = max(1, _env_int("DREAMVERSE_SP_SIZE", 1))
DREAMVERSE_MODEL_PATH = (os.getenv("DREAMVERSE_MODEL_PATH", "").strip() or None)
if DREAMVERSE_MODEL_PATH:
MODEL_CONFIG = {
**MODEL_CONFIG,
"model_path": DREAMVERSE_MODEL_PATH,
"config_model_path": DREAMVERSE_MODEL_PATH,
}
AVAILABLE_LORAS = {
"pixar": {
"repo": "vrgamedevgirl84/LTX_2.3_Pixar_Toon_Style_LoRa",
"trigger": "P1x4r",
"model": "fast-ltx23",
"position": "prepend",
"label": "Pixar Toon",
},
"transition": {
"repo": "valiantcat/LTX-2.3-Transition-LORA",
"trigger": "zhuanchang",
"model": "fast-ltx23",
"position": "append",
"label": "Transition",
},
}
def _active_model_key() -> str:
return ACTIVE_MODEL_ID
def _available_styles_for_active_model() -> list[str]:
active = _active_model_key()
return [k for k, v in AVAILABLE_LORAS.items() if v.get("model") == active]
def _resolve_lora_spec(spec: str) -> str | None:
spec = (spec or "").strip()
if not spec:
return None
if spec.lower() == "omninft":
return MODEL_CONFIG.get("lora_repo")
if spec.lower() in AVAILABLE_LORAS:
return AVAILABLE_LORAS[spec.lower()]["repo"]
return spec
def _parse_lora_stack(raw: str) -> list[tuple[str, float]]:
out: list[tuple[str, float]] = []
for item in (raw or "").split(","):
item = item.strip()
if not item:
continue
spec, _, stren = item.partition("@")
resolved = _resolve_lora_spec(spec)
if resolved:
try:
strength = float(stren) if stren.strip() else 1.0
except ValueError:
strength = 1.0
out.append((resolved, strength))
return out
DREAMVERSE_LORA_PATH = _resolve_lora_spec(os.getenv("DREAMVERSE_LORA_PATH", ""))
DREAMVERSE_LORA_NICKNAME = (os.getenv("DREAMVERSE_LORA_NICKNAME", "omninft").strip() or "omninft")
DREAMVERSE_LORA_STRENGTH = _env_float("DREAMVERSE_LORA_STRENGTH", 1.0)
DREAMVERSE_LORA_STACK = _parse_lora_stack(os.getenv("DREAMVERSE_LORA_STACK", ""))
def _resolve_devtools_paths(
+82 -2
View File
@@ -13,6 +13,7 @@ from multiprocessing import Process, Queue
from dreamverse.config import (
DEFAULT_MODEL_ID,
DREAMVERSE_SP_SIZE,
MODEL_REGISTRY,
STARTUP_WARMUP_ENABLED,
STARTUP_WARMUP_PROMPT,
@@ -33,6 +34,8 @@ from dreamverse.worker_ipc import (
InitAck,
JoinAck,
LeaveAck,
LoraAck,
LoraStackPayload,
MediaChunk,
MediaComplete,
MediaInit,
@@ -80,6 +83,7 @@ class CommandType(Enum):
USER_STEP = "user_step"
USER_LEAVE = "user_leave"
RELOAD_MODEL = "reload_model"
APPLY_LORA = "apply_lora"
@dataclass
@@ -282,6 +286,25 @@ def gpu_worker_process(
message=str(e),
))
elif cmd.type == CommandType.APPLY_LORA:
try:
assert isinstance(cmd.payload, LoraStackPayload), (f"APPLY_LORA requires LoraStackPayload, "
f"got {type(cmd.payload).__name__}")
trigger, position = worker.apply_lora_stack(cmd.payload.stack)
response_queue.put(
LoraAck(
user_id=cmd.user_id,
style_trigger=trigger,
style_trigger_position=position,
))
except Exception as e:
print(f"[GPU {gpu_id}] Apply LoRA error: {e}")
traceback.print_exc()
response_queue.put(WorkerError(
user_id=cmd.user_id,
message=str(e),
))
return True # Continue loop
if first_cmd is not None:
@@ -359,6 +382,25 @@ def gpu_worker_process(
message=str(e),
))
elif cmd.type == CommandType.APPLY_LORA:
try:
assert isinstance(cmd.payload, LoraStackPayload), (f"APPLY_LORA requires LoraStackPayload, "
f"got {type(cmd.payload).__name__}")
trigger, position = worker.apply_lora_stack(cmd.payload.stack)
response_queue.put(
LoraAck(
user_id=cmd.user_id,
style_trigger=trigger,
style_trigger_position=position,
))
except Exception as e:
print(f"[GPU {gpu_id}] Apply LoRA error: {e}")
traceback.print_exc()
response_queue.put(WorkerError(
user_id=cmd.user_id,
message=str(e),
))
elif cmd.type in (CommandType.USER_JOIN, CommandType.USER_STEP, CommandType.USER_LEAVE):
event_loop(first_cmd=cmd)
break
@@ -729,6 +771,25 @@ class GPUSlot:
raise RuntimeError(f"Unexpected step response for {user_id[:8]}: "
f"{type(response).__name__}")
async def apply_lora_stack(
self,
stack: list[tuple[str, float]],
) -> LoraAck:
"""Re-apply a runtime LoRA stack on this GPU's worker."""
self._active = True
user_id = "__lora__"
payload = LoraStackPayload(stack=stack)
response = await self._send_command_tagged(Command(CommandType.APPLY_LORA, payload=payload, user_id=user_id),
timeout=120.0)
match response:
case LoraAck() as ack:
return ack
case WorkerError(message=msg):
raise RuntimeError(f"Apply LoRA failed on GPU {self.gpu_id}: {msg}")
case _:
raise RuntimeError(f"Unexpected LoRA response on GPU {self.gpu_id}: "
f"{type(response).__name__}")
async def leave_user(self, user_id: str) -> None:
"""Remove a user from this GPU."""
try:
@@ -785,8 +846,16 @@ class GPUPool:
"""Manages multiple GPU worker subprocesses."""
def __init__(self, gpu_ids: list[int]):
self.gpu_ids = gpu_ids
self.slots: dict[int, GPUSlot] = {gpu_id: GPUSlot(gpu_id, str(gpu_id)) for gpu_id in gpu_ids}
sp_size = DREAMVERSE_SP_SIZE
groups = [gpu_ids[i:i + sp_size] for i in range(0, len(gpu_ids), sp_size)]
groups = [g for g in groups if len(g) == sp_size]
if not groups:
raise RuntimeError(f"Not enough GPUs for DREAMVERSE_SP_SIZE={sp_size}: available={gpu_ids}")
if sp_size > 1:
print(f"[INFO] Sequence-parallel slots (sp_size={sp_size}): " + ", ".join("{" + ",".join(map(str, g)) + "}"
for g in groups))
self.gpu_ids = [g[0] for g in groups]
self.slots: dict[int, GPUSlot] = {g[0]: GPUSlot(g[0], ",".join(str(x) for x in g)) for g in groups}
self.waiting_list: list[tuple[str, asyncio.Event, WebSocket]] = []
self.client_gpu_map: dict[str, int] = {}
self._pool_lock = asyncio.Lock()
@@ -903,6 +972,17 @@ class GPUPool:
except Exception:
pass
async def apply_lora_stack(
self,
stack: list[tuple[str, float]],
) -> dict[int, str | None]:
"""Re-apply a runtime LoRA stack across all ready GPU workers."""
ready_slots = [slot for slot in self.slots.values() if slot.ready]
if not ready_slots:
raise RuntimeError("No ready GPU workers to apply LoRA stack.")
results = await asyncio.gather(*(slot.apply_lora_stack(stack) for slot in ready_slots))
return {slot.gpu_id: ack.style_trigger for slot, ack in zip(ready_slots, results, strict=False)}
async def shutdown(self):
"""Shutdown all GPU workers."""
print("Shutting down GPU pool...")
+82 -1
View File
@@ -6,18 +6,22 @@ import os
from contextlib import asynccontextmanager
from pathlib import Path
from fastapi import FastAPI, WebSocket
from fastapi import FastAPI, HTTPException, WebSocket
from fastapi.middleware.cors import CORSMiddleware
from fastapi.staticfiles import StaticFiles
from pydantic import BaseModel
from fastvideo.entrypoints.streaming import build_health_router
from dreamverse.gpu_pool import GPUPool, get_available_gpus
from dreamverse.session_logger import SessionEventLogger
from dreamverse.config import (
AVAILABLE_LORAS,
DEVTOOLS_ENABLED,
FRONTEND_STATIC_DIR_CANDIDATES,
PROMPT_SAFETY_ENABLED,
SESSION_LOG_ROOT,
_available_styles_for_active_model,
_resolve_lora_spec,
)
from dreamverse.prompt_enhancer import PromptEnhancer
from dreamverse.prompt_safety import PromptSafetyFilter
@@ -104,6 +108,83 @@ async def websocket_endpoint(websocket: WebSocket):
await controller.run()
class LoraRequest(BaseModel):
strength: float = 1.0
styles: dict[str, float] = {}
style: str = ""
@app.get("/lora/options")
async def lora_options() -> dict:
if not DEVTOOLS_ENABLED:
raise HTTPException(status_code=404, detail="Not found")
styles = _available_styles_for_active_model()
has_base_lora = _resolve_lora_spec("omninft") is not None
labels = {"none": "None"}
labels.update({k: AVAILABLE_LORAS[k].get("label", k) for k in styles})
return {
"styles": ["none", *styles],
"labels": labels,
"has_base_lora": has_base_lora,
}
@app.post("/lora")
async def apply_lora(request: LoraRequest) -> dict:
if not DEVTOOLS_ENABLED:
raise HTTPException(status_code=404, detail="Not found")
if runtime.gpu_pool is None:
raise HTTPException(status_code=503, detail="GPU pool not ready")
strength = max(0.0, min(1.0, float(request.strength)))
allowed_styles = _available_styles_for_active_model()
requested = dict(request.styles) if request.styles else {}
if not requested and request.style and request.style.strip().lower() != "none":
requested = {request.style: 1.0}
styles_map: dict[str, float] = {}
for name, intensity in requested.items():
key = str(name).strip().lower()
if key in ("", "none"):
continue
if key not in allowed_styles:
raise HTTPException(status_code=400, detail=f"Unknown style for active model: {key}")
value = max(0.0, min(1.0, float(intensity)))
if value <= 0.0:
continue
styles_map[key] = value
stack: list[tuple[str, float]] = []
if _resolve_lora_spec("omninft") is not None:
stack.append(("omninft", strength))
for key, intensity in styles_map.items():
stack.append((key, intensity))
if not stack:
raise HTTPException(status_code=400,
detail="Active model has no OmniNFT LoRA and no style selected; nothing to apply.")
try:
triggers = await runtime.gpu_pool.apply_lora_stack(stack)
except Exception as exc:
raise HTTPException(status_code=500, detail=str(exc)) from exc
return {
"applied": True,
"strength": strength,
"styles": styles_map,
"triggers": {
key: AVAILABLE_LORAS[key]["trigger"]
for key in styles_map
},
"gpus": {
str(gpu_id): trigger
for gpu_id, trigger in triggers.items()
},
}
# Serve an exported frontend bundle when present.
for static_dir in FRONTEND_STATIC_DIR_CANDIDATES:
if os.path.isdir(static_dir):
+62 -2
View File
@@ -41,7 +41,7 @@ MOCK_FPS = 24
MOCK_DURATION_SECONDS = 5.0
MOCK_STREAM_CHUNK_SIZE_BYTES = 256 * 1024
MOCK_STREAM_CHUNK_DELAY_MS = 15
MOCK_AV_MIME = 'video/mp4; codecs="avc1.42E01E"'
MOCK_AV_MIME = 'video/mp4; codecs="avc1.42E01E,mp4a.40.2"'
FFMPEG_BIN = shutil.which(os.getenv("FASTVIDEO_FFMPEG_BIN", "ffmpeg"))
MOCK_SEGMENT_BYTES: bytes | None = None
@@ -75,9 +75,14 @@ def _build_mock_segment_bytes() -> bytes:
"lavfi",
"-i",
f"testsrc2=size={MOCK_FRAME_WIDTH}x{MOCK_FRAME_HEIGHT}:rate={MOCK_FPS}",
# Silent stereo AAC track. We don't need audible content — the FE's
# audio-decode coverage only needs a real AAC track present
"-f",
"lavfi",
"-i",
"anullsrc=channel_layout=stereo:sample_rate=48000",
"-t",
f"{MOCK_DURATION_SECONDS}",
"-an",
"-c:v",
"libx264",
"-preset",
@@ -90,6 +95,12 @@ def _build_mock_segment_bytes() -> bytes:
"baseline",
"-level",
"3.0",
"-c:a",
"aac",
"-b:a",
"128k",
"-ar",
"48000",
"-movflags",
"+empty_moov+default_base_moof+frag_keyframe",
"-frag_duration",
@@ -183,6 +194,55 @@ async def status():
"queue_size": 0,
"available_gpus": 1,
"total_gpus": 1,
"warmup_enabled": False,
"warmup_successful_gpus": 1,
"warmup_failed_gpus": 0,
"gpu_status": {
0: {
"ready": True,
"available": True,
"client_count": 0,
"current_model_id": "mock-ltx2",
"process_alive": True,
"warmup_enabled": False,
"warmup_success": True,
"warmup_error": None,
"warmup_timings": {},
},
},
}
@app.get("/prompt-system-config")
async def prompt_system_config():
return {
"next_segment_system_prompt": "Mock next segment prompt.",
"auto_extension_system_prompt": "Mock auto extension prompt.",
"rewrite_window_system_prompt": "Mock rewrite window prompt.",
"rewrite_user_system_prompt": "Mock rewrite user prompt.",
"rewrite_model": "mock-rewrite-model",
"rewrite_temperature": 0.0,
}
@app.get("/curated-presets")
async def curated_presets():
presets = [
{
"id": "mock_story",
"label": "Mock Story",
"description": "CI-safe mock story preset.",
"segment_prompts": [
"A tiny robot opens a glowing door.",
"The robot waves at a friendly moon.",
],
},
]
return {
"presets": presets,
"count": len(presets),
"file_path": None,
"fallback_file_path": None,
}
@@ -124,12 +124,21 @@ def test_config_uses_local_overlay_paths_when_devtools_enabled(monkeypatch, tmp_
assert module.CURATED_PRESETS_FALLBACK_FILE_PATH.endswith(
"apps/dreamverse/web/prompts/selected_ltx2_continuation_story_presets.json"
)
assert module.FRONTEND_STATIC_DIR_CANDIDATES == (
assert module.FRONTEND_STATIC_DIR_CANDIDATES[:2] == (
str(module.FRONTEND_ROOT / "out"),
str(module.FRONTEND_ROOT / "dist"),
)
def test_config_includes_docker_checkout_static_frontend_candidates(monkeypatch):
_set_required_prompt_keys(monkeypatch)
module = _load_config_module()
assert "/opt/FastVideo/apps/dreamverse/web/out" in module.FRONTEND_STATIC_DIR_CANDIDATES
assert "/opt/FastVideo/apps/dreamverse/web/dist" in module.FRONTEND_STATIC_DIR_CANDIDATES
def test_config_enables_prompt_safety_when_requested(monkeypatch):
_set_required_prompt_keys(monkeypatch)
monkeypatch.setenv("FASTVIDEO_ENABLE_PROMPT_SAFETY", "true")
@@ -171,3 +171,69 @@ def test_mock_server_cli_updates_latency(monkeypatch):
assert mock_server.LATENCY_MS == 321
finally:
mock_server.LATENCY_MS = old_latency_ms
def _lora_client(monkeypatch, captured):
server_main = _import_server_main(monkeypatch)
monkeypatch.setattr(server_main, "DEVTOOLS_ENABLED", True)
monkeypatch.setattr(server_main, "_available_styles_for_active_model", lambda: ["pixar", "transition"])
monkeypatch.setattr(
server_main,
"_resolve_lora_spec",
lambda spec: "FastVideo/OmniNFT" if str(spec).strip().lower() == "omninft" else spec,
)
class _FakePool:
async def apply_lora_stack(self, stack):
captured["stack"] = list(stack)
return {0: "trigger"}
server_main.runtime.gpu_pool = _FakePool()
return TestClient(server_main.app)
def test_apply_lora_stacks_multiple_styles_with_intensities(monkeypatch):
captured: dict = {}
client = _lora_client(monkeypatch, captured)
response = client.post("/lora", json={"strength": 0.8, "styles": {"pixar": 0.45, "transition": 0.5}})
assert response.status_code == 200
body = response.json()
assert body["applied"] is True
assert body["styles"] == {"pixar": 0.45, "transition": 0.5}
assert captured["stack"] == [("omninft", 0.8), ("pixar", 0.45), ("transition", 0.5)]
def test_apply_lora_rejects_unknown_style(monkeypatch):
captured: dict = {}
client = _lora_client(monkeypatch, captured)
response = client.post("/lora", json={"strength": 0.5, "styles": {"watercolor": 0.5}})
assert response.status_code == 400
assert "stack" not in captured
def test_apply_lora_accepts_legacy_single_style(monkeypatch):
captured: dict = {}
client = _lora_client(monkeypatch, captured)
response = client.post("/lora", json={"strength": 0.6, "style": "pixar"})
assert response.status_code == 200
assert captured["stack"] == [("omninft", 0.6), ("pixar", 1.0)]
def test_apply_lora_clamps_and_drops_zero_intensity(monkeypatch):
captured: dict = {}
client = _lora_client(monkeypatch, captured)
response = client.post("/lora", json={"strength": 1.5, "styles": {"pixar": 2.0, "transition": 0.0}})
assert response.status_code == 200
body = response.json()
assert body["strength"] == 1.0
assert body["styles"] == {"pixar": 1.0}
assert captured["stack"] == [("omninft", 1.0), ("pixar", 1.0)]
def test_lora_endpoints_hidden_without_devtools(monkeypatch):
server_main = _import_server_main(monkeypatch)
monkeypatch.setattr(server_main, "DEVTOOLS_ENABLED", False)
client = TestClient(server_main.app)
assert client.get("/lora/options").status_code == 404
assert client.post("/lora", json={"strength": 0.5}).status_code == 404
@@ -20,9 +20,11 @@ os.environ.setdefault("GROQ_API_KEY", "dummy")
import dreamverse.main as server_main # noqa: E402
import dreamverse.runtime as runtime # noqa: E402
from dreamverse.config import PROMPT_TIMEOUT_MS # noqa: E402
from dreamverse.session_logger import SessionEventLogger # noqa: E402
from dreamverse.worker_ipc import MediaChunk, MediaComplete, MediaInit # noqa: E402
from dreamverse.session import controller as session_controller # noqa: E402
from dreamverse.utils import PROMPT_EXTENSION_FAILURE_USER_MESSAGE # noqa: E402
pytestmark = pytest.mark.gpu
@@ -338,7 +340,7 @@ def test_session_event_logger_initializes_hostname_folder_and_utc_filename(
logger = SessionEventLogger(tmp_path)
assert logger.directory == tmp_path / socket.gethostname()
assert logger.directory.is_dir()
assert re.match(r"^\d{6}_\d{6}\.jsonl$", logger.path.name)
assert re.match(r"^\d{6}_\d{6}_\d{6}\.jsonl$", logger.path.name)
assert logger.path.is_file()
@@ -833,7 +835,7 @@ def test_initial_custom_rollout_prompt_generates_seed_window_before_streaming():
"rewrite_instruction": "A moonbase corridor thriller with flooding",
"rewrite_model": "gpt-4.1-mini",
"rewrite_temperature": 0.4,
"timeout_ms": server_main.PROMPT_TIMEOUT_MS,
"timeout_ms": PROMPT_TIMEOUT_MS,
"system_prompt_override": "Session-specific rewrite prompt",
}
assert fake_pool._slot.calls
@@ -1256,7 +1258,7 @@ def test_single5s_enhancement_fallback_does_not_start_generation():
assert fallback_payloads[0]["prompt_id"] == "simple_custom_prompt"
assert (
fallback_payloads[0]["error"]
== server_main.PROMPT_EXTENSION_FAILURE_USER_MESSAGE
== PROMPT_EXTENSION_FAILURE_USER_MESSAGE
)
assert fallback_payloads[0]["source"] == "user_enhancement_failed"
+105 -2
View File
@@ -12,6 +12,7 @@ module import time.
# mypy: ignore-errors
import gc
import os
import re
import time
from dataclasses import dataclass
from typing import Any
@@ -20,11 +21,19 @@ import numpy as np
import torch
from dreamverse.config import (
AVAILABLE_LORAS,
FRAME_HEIGHT,
FRAME_WIDTH,
MODEL_CONFIG,
NUM_FRAMES,
NUM_INFERENCE_STEPS,
DREAMVERSE_MAX_AUTOTUNE,
DREAMVERSE_SP_SIZE,
DREAMVERSE_LORA_PATH,
DREAMVERSE_LORA_NICKNAME,
DREAMVERSE_LORA_STRENGTH,
DREAMVERSE_LORA_STACK,
_resolve_lora_spec,
)
# Multi-frame decoded continuation defaults from
@@ -62,6 +71,15 @@ DEFAULT_LTX2_AUDIO_HOP_LENGTH = 160
DEFAULT_LTX2_AUDIO_DOWNSAMPLE = 4
def _reset_lora_registry(worker) -> dict:
pipeline = getattr(worker, "pipeline", None)
if pipeline is None:
return {"status": "no_pipeline"}
pipeline.cur_adapter_name = ""
pipeline.cur_adapter_strength = 1.0
return {"status": "lora_registry_reset"}
@dataclass
class StepResult:
"""Output of one generation step.
@@ -199,6 +217,8 @@ class VideoGenerationWorker:
self.continuation = ContinuationState()
self.audio_encoder_module = None
self.audio_processor_module = None
self.active_style_prepend: list[str] = []
self.active_style_append: list[str] = []
def _gpu_mem(self) -> str:
a = torch.cuda.memory_allocated() / 1024**3
@@ -257,6 +277,7 @@ class VideoGenerationWorker:
or self.current_model_config["model_path"])
enable_compile = os.getenv("ENABLE_TORCH_COMPILE", "1") == "1"
compile_mode = "max-autotune-no-cudagraphs" if DREAMVERSE_MAX_AUTOTUNE else None
components = ComponentConfig(
config_root=config_model_path,
@@ -269,7 +290,7 @@ class VideoGenerationWorker:
generator_config = GeneratorConfig(
model_path=model_root,
engine=EngineConfig(
num_gpus=1,
num_gpus=DREAMVERSE_SP_SIZE,
offload=OffloadConfig(
dit=False,
dit_layerwise=False,
@@ -280,9 +301,10 @@ class VideoGenerationWorker:
compile=CompileConfig(
enabled=enable_compile,
text_encoder_enabled=enable_compile,
vae_enabled=enable_compile,
backend="inductor",
fullgraph=True,
mode="max-autotune-no-cudagraphs",
mode=compile_mode,
dynamic=False,
),
use_fsdp_inference=False,
@@ -304,9 +326,88 @@ class VideoGenerationWorker:
self.generator = VideoGenerator.from_pretrained(config=generator_config)
print(f"[GPU {self.gpu_id}] After model load: {self._gpu_mem()}")
lora_stack = DREAMVERSE_LORA_STACK or ([(DREAMVERSE_LORA_PATH,
DREAMVERSE_LORA_STRENGTH)] if DREAMVERSE_LORA_PATH else [])
for i, (lora_path, lora_strength) in enumerate(lora_stack):
nickname = DREAMVERSE_LORA_NICKNAME if i == 0 else f"{DREAMVERSE_LORA_NICKNAME}_{i}"
print(f"[GPU {self.gpu_id}] Applying LoRA '{nickname}' from {lora_path} "
f"@{lora_strength} accumulate={i > 0}")
self.generator.set_lora_adapter(
lora_nickname=nickname,
lora_path=lora_path,
strength=lora_strength,
accumulate=(i > 0),
)
if lora_stack:
print(f"[GPU {self.gpu_id}] LoRA stack applied ({len(lora_stack)})")
self._load_audio_encoder(model_root)
print(f"[GPU {self.gpu_id}] LTX2 model loaded (warmup pending)")
def apply_lora_stack(
self,
stack: list[tuple[str, float]],
) -> tuple[str | None, str | None]:
"""Re-apply a runtime LoRA stack and update the active style trigger."""
if self.generator is None:
raise RuntimeError("Generator not initialized; cannot apply LoRA stack.")
try:
self.generator.unmerge_lora_weights()
except Exception as exc:
print(f"[GPU {self.gpu_id}] unmerge_lora_weights skipped: {exc}")
try:
self.generator.executor.collective_rpc(_reset_lora_registry)
except Exception as exc:
print(f"[GPU {self.gpu_id}] lora registry reset skipped: {exc}")
resolved_stack: list[tuple[str, float]] = []
for spec, strength in stack:
resolved = _resolve_lora_spec(spec)
if resolved:
resolved_stack.append((resolved, float(strength)))
for i, (lora_path, lora_strength) in enumerate(resolved_stack):
nickname = DREAMVERSE_LORA_NICKNAME if i == 0 else f"{DREAMVERSE_LORA_NICKNAME}_{i}"
print(f"[GPU {self.gpu_id}] Re-applying LoRA '{nickname}' from {lora_path} "
f"@{lora_strength} accumulate={i > 0}")
self.generator.set_lora_adapter(
lora_nickname=nickname,
lora_path=lora_path,
strength=lora_strength,
accumulate=(i > 0),
)
print(f"[GPU {self.gpu_id}] Runtime LoRA stack applied ({len(resolved_stack)})")
prepend: list[str] = []
append: list[str] = []
for spec, _ in stack:
key = (spec or "").strip().lower()
if key in AVAILABLE_LORAS:
trigger = AVAILABLE_LORAS[key]["trigger"]
if AVAILABLE_LORAS[key].get("position", "append") == "prepend":
prepend.append(trigger)
else:
append.append(trigger)
self.active_style_prepend = prepend
self.active_style_append = append
combined = " ".join(prepend + append) or None
return combined, None
def _inject_style_trigger(self, prompt: str) -> str:
result = prompt
for trigger in self.active_style_prepend:
trigger = (trigger or "").strip()
if trigger and not re.search(rf"\b{re.escape(trigger)}\b", result, re.IGNORECASE):
result = f"{trigger} {result}".strip()
for trigger in self.active_style_append:
trigger = (trigger or "").strip()
if trigger and not re.search(rf"\b{re.escape(trigger)}\b", result, re.IGNORECASE):
result = f"{result} {trigger}".strip()
return result
def _load_audio_encoder(self, model_root: str) -> None:
if not ENABLE_AUDIO_RE_ENCODE:
return
@@ -375,6 +476,8 @@ class VideoGenerationWorker:
"""Execute one generation step; snapshot state for the next segment."""
timings: dict = {}
prompt = self._inject_style_trigger(prompt)
request_kwargs = dict(
prompt=prompt,
negative_prompt="",
+14 -1
View File
@@ -48,6 +48,13 @@ class ReloadAck:
user_id: str
@dataclass(frozen=True)
class LoraAck:
user_id: str | None
style_trigger: str | None
style_trigger_position: str | None
@dataclass(frozen=True)
class WarmupComplete:
user_id: str | None
@@ -117,6 +124,7 @@ WorkerEvent = (StepComplete
| JoinAck
| LeaveAck
| ReloadAck
| LoraAck
| WarmupComplete
| MediaInit
| MediaChunk
@@ -154,4 +162,9 @@ class ReloadModelPayload:
model_config: dict
CommandPayload = UserStepPayload | WarmupPayload | ReloadModelPayload
@dataclass(frozen=True)
class LoraStackPayload:
stack: list[tuple[str, float]]
CommandPayload = (UserStepPayload | WarmupPayload | ReloadModelPayload | LoraStackPayload)
@@ -32,14 +32,21 @@
# FFMPEG_NATIVE_CC explicit C compiler command for native builds
# FFMPEG_NATIVE_CXX explicit C++ compiler command for native builds
#
# Toolchain selection: CC / CXX / AS are pinned to the conda-forge
# triplet matching `uname -m`. Inherited values are intentionally
# ignored — conda envs that have BOTH `gcc_linux-64` and
# `gcc_linux-aarch64` installed export the cross-compiler triplet on
# every `conda activate` (the `aarch64` activation script sorts later
# and wins), which silently breaks x264's compiler probe on the
# opposite host. If you genuinely need a non-host toolchain, set
# FFMPEG_NATIVE_CC and/or FFMPEG_NATIVE_CXX explicitly.
# Toolchain selection (precedence, highest first):
# 1. FFMPEG_NATIVE_CC / FFMPEG_NATIVE_CXX, if set — explicit override.
# 2. The conda-forge triplet matching `uname -m`
# (`x86_64-conda-linux-gnu-cc` or `aarch64-conda-linux-gnu-cc`),
# if present on PATH. Pinning the triplet sidesteps a bug where
# conda envs with BOTH `gcc_linux-64` and `gcc_linux-aarch64`
# installed export the cross-compiler triplet on every
# `conda activate` (the `aarch64` activation script sorts later
# and wins), silently breaking x264's compiler probe on the
# opposite host.
# 3. System `gcc` / `g++` — fallback so the script also works in a
# plain venv or bare shell with no conda toolchain installed.
# Inherited bare CC / CXX from the caller environment are ignored in
# cases (2) and (3); use FFMPEG_NATIVE_CC/CXX to inject a non-host
# toolchain.
set -euo pipefail
# ─── Defaults (override via env) ──────────────────────────────────────────
@@ -61,8 +68,8 @@ MAKE_JOBS="${MAKE_JOBS:-$(( NPROC < 16 ? NPROC : 16 ))}"
ARCH="$(uname -m)"
case "$ARCH" in
x86_64)
DEFAULT_CC=x86_64-conda-linux-gnu-cc
DEFAULT_CXX=x86_64-conda-linux-gnu-c++
CONDA_CC=x86_64-conda-linux-gnu-cc
CONDA_CXX=x86_64-conda-linux-gnu-c++
AS=nasm
X264_CFLAGS="-O3 -march=native -mtune=native -fPIC -flto"
X264_LDFLAGS="-flto -fuse-linker-plugin"
@@ -70,8 +77,8 @@ case "$ARCH" in
FFMPEG_LDFLAGS="-flto -Wl,-rpath,$INSTALL_PREFIX/lib"
;;
aarch64)
DEFAULT_CC=aarch64-conda-linux-gnu-cc
DEFAULT_CXX=aarch64-conda-linux-gnu-c++
CONDA_CC=aarch64-conda-linux-gnu-cc
CONDA_CXX=aarch64-conda-linux-gnu-c++
unset AS # GNU as on ARM
X264_CFLAGS="-O3 -mcpu=native -fPIC -flto"
X264_LDFLAGS="-flto -fuse-linker-plugin"
@@ -84,8 +91,27 @@ case "$ARCH" in
;;
esac
# Prefer the conda-forge triplet when it's actually on PATH; otherwise fall
# back to system gcc/g++ so a plain venv works too. FFMPEG_NATIVE_CC/CXX
# overrides both.
if command -v -- "$CONDA_CC" >/dev/null 2>&1 \
&& command -v -- "$CONDA_CXX" >/dev/null 2>&1; then
DEFAULT_CC="$CONDA_CC"
DEFAULT_CXX="$CONDA_CXX"
default_source="conda-forge ($ARCH triplet)"
else
DEFAULT_CC=gcc
DEFAULT_CXX=g++
default_source="system gcc/g++"
fi
CC="${FFMPEG_NATIVE_CC:-$DEFAULT_CC}"
CXX="${FFMPEG_NATIVE_CXX:-$DEFAULT_CXX}"
if [[ -n "${FFMPEG_NATIVE_CC:-}" || -n "${FFMPEG_NATIVE_CXX:-}" ]]; then
toolchain_source="FFMPEG_NATIVE_CC/CXX override"
else
toolchain_source="$default_source"
fi
require_compiler() {
local name="$1" compiler="$2"
@@ -95,6 +121,8 @@ require_compiler() {
fi
if ! command -v -- "$compiler" >/dev/null 2>&1; then
echo "[install_native_ffmpeg] $name is unavailable: $compiler" >&2
echo "[install_native_ffmpeg] install a compiler, or set FFMPEG_NATIVE_CC/FFMPEG_NATIVE_CXX" \
"to compiler commands on PATH." >&2
exit 1
fi
if ! "$compiler" --version >/dev/null 2>&1; then
@@ -107,7 +135,7 @@ require_compiler CC "$CC"
require_compiler CXX "$CXX"
export CC CXX
[[ -n "${AS:-}" ]] && export AS
echo "[install_native_ffmpeg] toolchain: CC=$CC CXX=$CXX AS=${AS:-<gnu-as>} (uname -m=$ARCH)"
echo "[install_native_ffmpeg] toolchain: CC=$CC CXX=$CXX AS=${AS:-<gnu-as>} (uname -m=$ARCH, source: $toolchain_source)"
# ─── Step 0: probe required tools ─────────────────────────────────────────
required=("$CC" "$CXX" make pkg-config git)
@@ -118,7 +146,12 @@ for cmd in "${required[@]}"; do
done
if (( ${#missing[@]} > 0 )); then
echo "[install_native_ffmpeg] missing required tools: ${missing[*]}" >&2
echo "[install_native_ffmpeg] install them, or set FFMPEG_NATIVE_CC/FFMPEG_NATIVE_CXX explicitly." >&2
echo "[install_native_ffmpeg] install the missing tools. FFMPEG_NATIVE_CC/FFMPEG_NATIVE_CXX" \
"only override compiler selection; they do not provide make, pkg-config, git, or nasm." >&2
if [[ " ${missing[*]} " == *" nasm "* ]]; then
echo "[install_native_ffmpeg] nasm is required on x86_64 for x264 SIMD; install nasm" \
"(for example via apt or conda-forge) and retry." >&2
fi
exit 1
fi
+1 -1
View File
@@ -28,7 +28,7 @@ NO_BROWSER=1 apps/dreamverse/scripts/launch/launch_demo.sh
`launch_backend_dreamverse.sh` starts the full Dreamverse backend path used by
the web app.
`launch_frontend.sh` starts the Next.js frontend and installs `pnpm`
`launch_frontend.sh` starts the Next.js frontend and installs npm
dependencies when `node_modules/` is missing.
`launch_backend_fastvideo.sh` starts the typed `fastvideo serve --config` path
@@ -7,8 +7,8 @@
# FRONTEND_MODE=dev bash launch_frontend.sh # plain dev (5299, no devtools)
# FRONTEND_MODE=single5s bash launch_frontend.sh # single-5s product mode
#
# The script ``cd``'s into ``web`` and shells out to pnpm. It runs
# ``pnpm install --frozen-lockfile`` only when ``node_modules/`` is missing so repeat
# The script ``cd``'s into ``web`` and shells out to npm. It runs
# ``npm ci`` only when ``node_modules/`` is missing so repeat
# launches are fast.
set -euo pipefail
@@ -34,21 +34,21 @@ fi
cd "${WEB_ROOT}"
if [[ ! -d node_modules ]]; then
echo "[launch-demo] node_modules missing — running pnpm install --frozen-lockfile"
pnpm install --frozen-lockfile
echo "[launch-demo] node_modules missing — running npm ci"
npm ci
fi
case "${FRONTEND_MODE}" in
devtools)
echo "[launch-demo] starting Next.js dev:devtools (port 5274)"
exec pnpm run dev:devtools -- "$@"
exec npm run dev:devtools -- "$@"
;;
dev)
echo "[launch-demo] starting Next.js dev (port 5299)"
exec pnpm run dev -- "$@"
exec npm run dev -- "$@"
;;
single5s)
echo "[launch-demo] starting Next.js dev:single5s (port 5274)"
exec pnpm run dev:single5s -- "$@"
exec npm run dev:single5s -- "$@"
;;
esac
+127
View File
@@ -0,0 +1,127 @@
# Dreamverse Modal deployment
This directory contains the Modal wrapper for running the registry-built
Dreamverse Docker image on a Modal B200 container. The wrapper pulls a prebuilt Docker image and exposes the dreamverse application running inside.
## 1. Install Modal CLI
Install and configure the Modal CLI for the target workspace/profile:
```bash
pip install modal
modal token set --token-id <token-id> --token-secret <token-secret> --profile=<profile-name>
modal profile activate <profile-name>
```
## 2. Create the API key secret
Dreamverse needs several API keys for prompt-rewriter LLMs and model access in the Modal secret named
`dreamverse-api-keys`. Create or replace it with placeholders like this:
```bash
modal secret create dreamverse-api-keys \
CEREBRAS_API_KEY=<your-cerebras-key> \
GROQ_API_KEY=<your-groq-key> \
HF_TOKEN=<your-huggingface-token> \
--force
```
## 3. Deploy with Docker image
`DREAMVERSE_IMAGE` is required at deploy time:
```bash
DREAMVERSE_IMAGE=ghcr.io/<org>/<repo>/dreamverse:<tag> \
modal deploy apps/dreamverse/scripts/modal/modal_app.py
```
Use a SHA-specific tag, not `latest`. Use a `dreamverse-backend-cuda12.9.1-sha-*`
tag for backend-only deploys, or a `dreamverse-ui-cuda12.9.1-sha-*` tag for an
image that includes the static UI served by the backend.
For local image build details, see
`apps/dreamverse/docker/README.md` and `apps/dreamverse/docker/docker_build.sh`.
## 4. Validate the deployment
Set the URL returned by `modal deploy`, then probe the backend:
```bash
URL=https://<workspace>--dreamverse-b200-serve.modal.run
curl -fsS --max-time 300 "$URL/healthz"
curl -fsS --max-time 300 "$URL/status"
curl -fsS --max-time 300 "$URL/readyz"
```
Check logs with function and container IDs so you can confirm only one container
is active:
```bash
modal app logs dreamverse-b200 \
--tail 200 \
--timestamps \
--show-function-id \
--show-container-id
```
Useful Modal commands:
```bash
modal app list
modal app list --json
modal billing report --for today --resolution h
modal app stop dreamverse-b200
```
## Runtime details
### Volumes and assumed locations
`modal_app.py` creates the named volumes if they do not exist and mounts them at
the locations expected by the image:
| Modal volume | Container path | Environment variable | Purpose |
| --- | --- | --- | --- |
| `dreamverse-hf-cache` | `/root/.cache/huggingface` | `HF_HOME` | Hugging Face model/cache data |
| `dreamverse-state` | `/var/lib/dreamverse` | `FASTVIDEO_DREAMVERSE_HOME` | Dreamverse outputs, session logs, and runtime state |
### Useful dev env vars
Most deploys only need the Modal secret and required `DREAMVERSE_IMAGE`.
`DREAMVERSE_IMAGE` is read from your local shell at deploy time; the runtime
knobs below live in the image or `modal_app.py` unless you intentionally change them:
- `DREAMVERSE_IMAGE`: deploy a prebuilt SHA-specific image tag without editing `modal_app.py`.
- `FASTVIDEO_ENABLE_DEVTOOLS`: enable Dreamverse devtools behavior.
- `FASTVIDEO_PROMPT_PROVIDER` / `FASTVIDEO_PROMPT_*_MODEL`: try prompt-rewriter provider or model choices.
- `FASTVIDEO_ENABLE_STARTUP_WARMUP`: trade slower startup for a warmer first request.
- `DREAMVERSE_MAX_AUTOTUNE`: enable or disable PyTorch Inductor max-autotune for the compiled Dreamverse runtime.
The Modal wrapper defaults to torch compile with Inductor max-autotune enabled
for the fastest generation after startup warmup. If you need shorter
compile/warmup time, disable max-autotune via:
```bash
DREAMVERSE_IMAGE=ghcr.io/<org>/<repo>/dreamverse:<tag> \
DREAMVERSE_MAX_AUTOTUNE=0 \
modal deploy apps/dreamverse/scripts/modal/modal_app.py
```
### Autoscaling and cost safety
The current deployed script keeps exactly one B200 container warm by setting
`min_containers=1` and `max_containers=1`. `min_containers=1` prevents Modal
from scaling the deployment down to zero after idle periods, which avoids paying
the expensive torch-compile/startup-warmup cost again on the next request.
`max_containers=1` caps concurrency so concurrent requests queue onto the warm
container instead of spawning additional B200 containers and multiplying cost.
### Common gotchas
- Modal profile/workspace matters: create the secret and deploy with the same active profile.
- `modal secret create ... --force` replaces the existing `dreamverse-api-keys` secret.
- Use the exact URL printed by `modal deploy`; the placeholder URL above is only an example shape.
- B200 cold starts and first model downloads can make `/readyz` slow. Check logs before redeploying.
- `modal app stop dreamverse-b200` stops the app; it is not a read-only inspection command.
@@ -0,0 +1,77 @@
# pyright: reportAttributeAccessIssue=false
"""Modal deployment entrypoint for the Dreamverse B200 backend."""
import os
import subprocess
import modal
IMAGE = os.environ.get("DREAMVERSE_IMAGE")
if not IMAGE:
raise RuntimeError(
"DREAMVERSE_IMAGE is required. Set it to a published SHA-specific Dreamverse image, "
"for example a dreamverse-backend-cuda12.9.1-sha-* tag or a "
"dreamverse-ui-cuda12.9.1-sha-* tag if serving the static UI."
)
# ``@modal.web_server`` invokes ``serve()`` directly and bypasses the image
# ENTRYPOINT (``docker/docker_entrypoint.sh``). That entrypoint normally
# ``:?``-validates these secret entries and ``source``-s ``ffmpeg-env.sh``;
# neither runs on Modal, so we replicate both here.
_REQUIRED_SECRET_KEYS: tuple[str, ...] = ("CEREBRAS_API_KEY", "GROQ_API_KEY")
image = modal.Image.from_registry(IMAGE)
image = image.env({
"DREAMVERSE_IMAGE": IMAGE,
"HF_HOME": "/root/.cache/huggingface",
"FASTVIDEO_DREAMVERSE_HOME": "/var/lib/dreamverse",
"FASTVIDEO_ENABLE_STARTUP_WARMUP": "1",
"FASTVIDEO_GPU_COUNT": "1",
"ENABLE_TORCH_COMPILE": "1",
"DREAMVERSE_MAX_AUTOTUNE": os.environ.get("DREAMVERSE_MAX_AUTOTUNE", "1"),
"STREAM_MODE": "av_fmp4",
# Native ffmpeg is built into ``/opt/ffmpeg-native`` by the Dockerfile but
# is not on the image's ``PATH``. Without these, ``av_streaming``'s
# ``shutil.which("ffmpeg")`` returns ``None`` and ``av_fmp4`` muxing has
# no encoder. ``ffmpeg-env.sh`` would normally set them.
"FASTVIDEO_FFMPEG_BIN": "/opt/ffmpeg-native/bin/ffmpeg",
"FASTVIDEO_VIDEO_CODEC": "libx264",
})
app = modal.App("dreamverse-b200")
hf_cache = modal.Volume.from_name("dreamverse-hf-cache", create_if_missing=True)
dreamverse_state = modal.Volume.from_name("dreamverse-state", create_if_missing=True)
@app.function(
image=image,
gpu="B200",
cpu=16,
memory=65536,
timeout=7200,
startup_timeout=4800,
min_containers=1,
max_containers=1,
secrets=[modal.Secret.from_name("dreamverse-api-keys")],
volumes={
"/root/.cache/huggingface": hf_cache,
"/var/lib/dreamverse": dreamverse_state,
},
)
@modal.web_server(8009, startup_timeout=4800)
def serve():
# ``or ""`` collapses ``None`` (unset) into an empty string, ``.strip()``
# collapses whitespace-only values (e.g. ``" "``) — both should be
# treated as missing.
missing = [
k for k in _REQUIRED_SECRET_KEYS
if not (os.environ.get(k) or "").strip()
]
if missing:
raise RuntimeError(
"dreamverse-api-keys secret is missing required entries: "
f"{', '.join(missing)}. Add them with `modal secret create "
"dreamverse-api-keys ... --force` and redeploy "
"(see apps/dreamverse/scripts/modal/README.md).")
subprocess.Popen(["dreamverse-server", "--host", "0.0.0.0", "--port", "8009"])
+2 -2
View File
@@ -87,9 +87,9 @@ if [[ "${started_backend}" == "1" ]]; then
echo "Backend log: ${BACKEND_LOG_PATH}"
echo "Backend is still running so you can launch the frontend:"
echo " cd ${ROOT_DIR}/web"
echo " BACKEND_HOST=${HOST} BACKEND_PORT=${PORT} pnpm run dev"
echo " BACKEND_HOST=${HOST} BACKEND_PORT=${PORT} npm run dev"
else
echo "You can now launch the frontend:"
echo " cd ${ROOT_DIR}/web"
echo " BACKEND_HOST=${HOST} BACKEND_PORT=${PORT} pnpm run dev"
echo " BACKEND_HOST=${HOST} BACKEND_PORT=${PORT} npm run dev"
fi
@@ -11,7 +11,7 @@ test.describe('backend health', () => {
expect(response.ok()).toBeTruthy();
const body = await response.json();
expect(body.status).toBe('ok');
expect(body.service).toBe('ltx2-streaming-backend');
expect(['ltx2-streaming-backend', 'ltx2-streaming-mock-server']).toContain(body.service);
});
test('readyz reports gpu pool state', async ({ request }) => {
@@ -18,7 +18,7 @@ import { test, expect, type WebSocket as PWWebSocket } from '@playwright/test';
* PLAYWRIGHT_BASE_URL=http://127.0.0.1:5274 \
* NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
* PLAYWRIGHT_LONG_RUNNING=1 \
* pnpm exec playwright test e2e/long-running-segments.spec.ts
* npm exec -- playwright test e2e/long-running-segments.spec.ts
*
* The full run takes ~7-9 minutes on a B200: ~3-4min for torch.compile
* max-autotune to warm both DiT + text-encoder graphs, then ~30s for
@@ -0,0 +1,311 @@
import { stat } from 'node:fs/promises';
import { test, expect, type WebSocket as PWWebSocket } from '@playwright/test';
test.describe('mock-backed generation smoke', () => {
test('accepts a custom prompt and enters the streaming state', async ({ page, request }) => {
const health = await request.get('/healthz');
const body = health.ok() ? await health.json() : {};
test.skip(
body.service !== 'ltx2-streaming-mock-server',
'Mock-backed smoke requires dreamverse.mock_server on BACKEND_PORT (default 8009).',
);
const wsEventTypes: string[] = [];
page.on('websocket', (ws: PWWebSocket) => {
ws.on('framereceived', ({ payload }) => {
if (typeof payload !== 'string') return;
try {
const parsed = JSON.parse(payload);
if (typeof parsed?.type === 'string') wsEventTypes.push(parsed.type);
} catch {
}
});
});
await page.goto('/');
await expect(page.getByRole('img', { name: 'FastVideo' })).toBeVisible();
const continuation = page.getByLabel('Continuation prompt');
await expect(continuation).toBeVisible();
await continuation.fill('A glass whale swims above a neon forest');
await page.getByRole('button', { name: /^generate$/i }).click();
await expect(continuation).toBeDisabled({ timeout: 30_000 });
await expect(continuation).toHaveAttribute('placeholder', /generating video/i);
await expect(page.getByRole('button', { name: /^leave$/i })).toBeVisible({ timeout: 30_000 });
await expect(page.locator('video').first()).toHaveCount(1);
await expect.poll(() => wsEventTypes).toEqual(
expect.arrayContaining(['ltx2_stream_start', 'ltx2_segment_start', 'media_init']),
);
});
test('streams, plays, and surfaces a downloadable clip', async ({ page, request }) => {
const health = await request.get('/healthz');
const body = health.ok() ? await health.json() : {};
test.skip(
body.service !== 'ltx2-streaming-mock-server',
'Mock-backed playback assertions require dreamverse.mock_server on BACKEND_PORT (default 8009).',
);
await page.addInitScript(() => {
(window as unknown as { __sharedFiles: unknown }).__sharedFiles = null;
const stub = async (data: { files?: File[] }) => {
const files = Array.isArray(data?.files) ? data.files : [];
(window as unknown as { __sharedFiles: unknown }).__sharedFiles = files.map((f) => ({
name: f.name,
type: f.type,
size: f.size,
}));
};
Object.defineProperty(Navigator.prototype, 'share', {
value: stub,
configurable: true,
writable: true,
});
Object.defineProperty(Navigator.prototype, 'canShare', {
value: (data: { files?: File[] }) => Array.isArray(data?.files),
configurable: true,
writable: true,
});
});
await page.goto('/');
const continuation = page.getByLabel('Continuation prompt');
await expect(continuation).toBeVisible();
await continuation.fill('A glass whale swims above a neon forest');
await page.getByRole('button', { name: /^generate$/i }).click();
const liveVideo = page.locator('video:not(.hidden)');
await expect(liveVideo).toHaveCount(1);
await test.step('MSE pipeline attaches the live <video>', async () => {
// Chromium/Firefox attach via blob: URL; Safari uses srcObject with ManagedMediaSource.
await expect
.poll(
async () =>
liveVideo.evaluate(
(v: HTMLVideoElement) =>
v.src.startsWith('blob:') || v.srcObject !== null,
),
{ timeout: 30_000 },
)
.toBe(true);
});
await test.step('SourceBuffer accepts fMP4 and decoder produces frames', async () => {
await expect
.poll(
async () =>
liveVideo.evaluate((v: HTMLVideoElement) =>
v.buffered.length > 0 ? v.buffered.end(0) : 0,
),
{ timeout: 60_000 },
)
.toBeGreaterThan(0);
// readyState >= 3 = HAVE_FUTURE_DATA.
await expect
.poll(
async () => liveVideo.evaluate((v: HTMLVideoElement) => v.readyState),
{ timeout: 60_000 },
)
.toBeGreaterThanOrEqual(3);
});
await test.step('<video> playback advances past the first second', async () => {
await liveVideo.evaluate(async (v: HTMLVideoElement) => {
if (v.paused) {
try {
await v.play();
} catch {
}
}
});
await expect
.poll(
async () => liveVideo.evaluate((v: HTMLVideoElement) => v.currentTime),
{ timeout: 30_000 },
)
.toBeGreaterThan(0);
await expect
.poll(
async () =>
liveVideo.evaluate((v: HTMLVideoElement) =>
v.ended ? v.duration : v.currentTime,
),
{ timeout: 10_000 },
)
.toBeGreaterThanOrEqual(1.0);
});
await test.step('the AAC audio track is demuxed and decoded', async () => {
// The mock emits AAC (mp4a.40.2), matching the real backend
// (av_streaming.py). Proves the FE decoded the audio track.
await expect
.poll(
async () =>
liveVideo.evaluate((el: HTMLVideoElement) => {
const v = el as HTMLVideoElement & {
audioTracks?: { length: number };
mozHasAudio?: boolean;
webkitAudioDecodedByteCount?: number;
};
return (
(v.audioTracks?.length ?? 0) > 0 ||
v.mozHasAudio === true ||
(v.webkitAudioDecodedByteCount ?? 0) > 0
);
}),
{ timeout: 10_000 },
)
.toBe(true);
});
await test.step('completed clip surfaces a working download', async () => {
const downloadButton = page.getByRole('button', { name: /download video|share video/i });
await expect(downloadButton).toBeVisible({ timeout: 30_000 });
const buttonLabel = (await downloadButton.getAttribute('aria-label')) ?? '';
const isShareFlow = /share video/i.test(buttonLabel);
if (isShareFlow) {
await downloadButton.click();
await expect
.poll(
async () => page.evaluate(() => (window as { __sharedFiles?: unknown }).__sharedFiles),
{ timeout: 10_000 },
)
.not.toBeNull();
const shared = (await page.evaluate(
() => (window as { __sharedFiles?: Array<{ name: string; type: string; size: number }> }).__sharedFiles,
)) ?? [];
expect(shared).toHaveLength(1);
expect(shared[0].name).toMatch(/\.(mp4|webm)$/);
expect(shared[0].size).toBeGreaterThan(0);
} else {
const downloadPromise = page.waitForEvent('download');
await downloadButton.click();
const download = await downloadPromise;
expect(download.suggestedFilename()).toMatch(/\.(mp4|webm)$/);
const savedPath = await download.path();
expect(savedPath).not.toBeNull();
const { size } = await stat(savedPath!);
expect(size).toBeGreaterThan(0);
}
});
await test.step('project history sidebar lists the current session', async () => {
const sidebar = page.getByRole('complementary', { name: 'Project history' });
await page.getByRole('button', { name: 'Toggle sidebar' }).click();
await expect(sidebar).toBeInViewport();
await expect(sidebar.getByText('Current', { exact: true })).toBeVisible();
await expect(sidebar.getByText('Active', { exact: true })).toBeVisible();
});
const lastError = await liveVideo.evaluate((v: HTMLVideoElement) =>
v.error ? `${v.error.code}: ${v.error.message ?? ''}` : null,
);
expect(lastError).toBeNull();
});
test('starts a new project and switches back to the prior session', async ({ page, request }) => {
const health = await request.get('/healthz');
const body = health.ok() ? await health.json() : {};
test.skip(
body.service !== 'ltx2-streaming-mock-server',
'Mock-backed project lifecycle assertions require dreamverse.mock_server on BACKEND_PORT (default 8009).',
);
await page.goto('/');
await test.step('complete a generation so there is a project to save', async () => {
const continuation = page.getByLabel('Continuation prompt');
await expect(continuation).toBeVisible();
await continuation.fill('Aurora over a frozen lake');
await page.getByRole('button', { name: /^generate$/i }).click();
await expect(
page.getByRole('button', { name: /download video|share video/i }),
).toBeVisible({ timeout: 60_000 });
});
const sidebar = page.getByRole('complementary', { name: 'Project history' });
await test.step('"New project" closes the sidebar and resets the composer', async () => {
await page.getByRole('button', { name: 'Toggle sidebar' }).click();
await expect(sidebar).toBeInViewport();
await sidebar.getByRole('button', { name: /^new project$/i }).click();
await expect(sidebar).not.toBeInViewport();
await expect(page.getByRole('button', { name: /^generate$/i })).toBeVisible({ timeout: 30_000 });
const continuation = page.getByLabel('Continuation prompt');
await expect(continuation).toBeEnabled();
await expect(continuation).toHaveValue('');
});
await test.step('the prior session appears under timeline', async () => {
await page.getByRole('button', { name: 'Toggle sidebar' }).click();
await expect(sidebar).toBeInViewport();
await expect(sidebar.getByText('Previous', { exact: true })).toBeVisible({ timeout: 30_000 });
await expect(sidebar.getByText(/^(just now|\d+m ago)$/).first()).toBeVisible();
});
await test.step('clicking the prior session enters viewing mode', async () => {
const priorRow = sidebar.locator('div[role="button"]').filter({ hasText: /just now|\d+m ago/ }).first();
await priorRow.click();
await expect(sidebar).not.toBeInViewport();
await expect(page.locator('video[autoplay][loop]')).toBeVisible({ timeout: 30_000 });
});
});
test('saved projects persist across a page reload', async ({ page, request }) => {
const health = await request.get('/healthz');
const body = health.ok() ? await health.json() : {};
test.skip(
body.service !== 'ltx2-streaming-mock-server',
'Mock-backed persistence assertions require dreamverse.mock_server on BACKEND_PORT (default 8009).',
);
await page.goto('/');
await test.step('complete a generation', async () => {
const continuation = page.getByLabel('Continuation prompt');
await expect(continuation).toBeVisible();
await continuation.fill('Aurora over a frozen lake');
await page.getByRole('button', { name: /^generate$/i }).click();
await expect(
page.getByRole('button', { name: /download video|share video/i }),
).toBeVisible({ timeout: 60_000 });
});
const sidebar = page.getByRole('complementary', { name: 'Project history' });
await test.step('persist via "New project" and confirm it lands under Previous', async () => {
await page.getByRole('button', { name: 'Toggle sidebar' }).click();
await expect(sidebar).toBeInViewport();
await sidebar.getByRole('button', { name: /^new project$/i }).click();
await expect(page.getByRole('button', { name: /^generate$/i })).toBeVisible({ timeout: 30_000 });
await page.getByRole('button', { name: 'Toggle sidebar' }).click();
await expect(sidebar).toBeInViewport();
await expect(sidebar.getByText('Previous', { exact: true })).toBeVisible({ timeout: 30_000 });
await expect(sidebar.getByText(/^(just now|\d+m ago)$/).first()).toBeVisible();
});
await test.step('after page reload, the prior project is still in Previous', async () => {
await page.reload();
await page.getByRole('button', { name: 'Toggle sidebar' }).click();
await expect(sidebar).toBeInViewport();
await expect(sidebar.getByText('Previous', { exact: true })).toBeVisible({ timeout: 30_000 });
await expect(sidebar.getByText(/^(just now|\d+m ago)$/).first()).toBeVisible();
});
});
});
+13 -2
View File
@@ -6,10 +6,13 @@ const backendHost = process.env.BACKEND_HOST || '127.0.0.1';
const backendPort = Number(process.env.BACKEND_PORT) || 8009;
const backendUrl = `http://${backendHost}:${backendPort}`;
const configDir = path.dirname(fileURLToPath(import.meta.url));
const staticExport = process.env.NEXT_OUTPUT_EXPORT === '1';
const nextConfig: NextConfig = {
...(staticExport ? { output: 'export' as const } : {}),
...(staticExport ? { images: { unoptimized: true } } : {}),
outputFileTracingRoot: path.join(configDir, '..', '..', '..'),
async rewrites() {
...(staticExport ? {} : { async rewrites() {
return [
{
source: '/ws',
@@ -47,8 +50,16 @@ const nextConfig: NextConfig = {
source: '/curated-presets/:path*',
destination: `${backendUrl}/curated-presets/:path*`,
},
{
source: '/lora',
destination: `${backendUrl}/lora`,
},
{
source: '/lora/:path*',
destination: `${backendUrl}/lora/:path*`,
},
];
},
} }),
webpack: (config) => {
config.module.rules.push({
test: /\.jsonl$/,
+11669 -11443
View File
File diff suppressed because it is too large Load Diff
+9 -2
View File
@@ -38,7 +38,7 @@
"geist": "^1.7.0",
"lucide-react": "^0.577.0",
"mp4box": "^2.3.0",
"next": "^15.3.3",
"next": "15.5.18",
"radix-ui": "^1.4.3",
"react": "^19.1.0",
"react-dom": "^19.1.0",
@@ -59,9 +59,16 @@
"@vitest/coverage-v8": "^3.2.4",
"jsdom": "^26.1.0",
"mock-socket": "^9.3.1",
"postcss": "^8.5.8",
"postcss": "8.5.10",
"tailwindcss": "^4.2.1",
"typescript": "^5.8.3",
"vitest": "^3.2.4"
},
"overrides": {
"esbuild": "0.25.12",
"picomatch": "4.0.4",
"postcss": "8.5.10",
"rollup": "4.59.0",
"vite": "7.3.2"
}
}
+30 -15
View File
@@ -1,12 +1,16 @@
import { defineConfig, devices } from '@playwright/test';
const backendHost = process.env.BACKEND_HOST || '127.0.0.1';
const backendPort = process.env.BACKEND_PORT || '8009';
/**
* Playwright config for Dreamverse end-to-end tests.
*
* Tests assume the Dreamverse Python server is reachable at
* BACKEND_HOST:BACKEND_PORT (default 127.0.0.1:8009) and the Next.js frontend
* runs on port 5299. The webServer block boots `pnpm run dev` if no
* server is already listening so tests work both locally and in CI.
* runs on port 5299. The webServer block boots the mock backend and
* `npm run dev` if no server is already listening so tests work both
* locally and in CI.
*/
export default defineConfig({
testDir: './e2e',
@@ -26,23 +30,34 @@ export default defineConfig({
viewport: { width: 1280, height: 720 },
screenshot: 'only-on-failure',
trace: 'retain-on-failure',
video: process.env.CI ? 'retain-on-failure' : 'on',
},
projects: [
{
name: 'chromium',
use: { ...devices['Desktop Chrome'] },
},
{ name: 'chromium', use: { ...devices['Desktop Chrome'] } },
{ name: 'webkit', use: { ...devices['Desktop Safari'] } },
{ name: 'firefox', use: { ...devices['Desktop Firefox'] } },
{ name: 'msedge', use: { ...devices['Desktop Edge'], channel: 'msedge' } },
{ name: 'mobile-safari', use: { ...devices['iPhone 14'] } },
{ name: 'mobile-chromium', use: { ...devices['Pixel 7'] } },
],
webServer: process.env.PLAYWRIGHT_SKIP_WEBSERVER
? undefined
: {
command: 'pnpm run dev',
url: 'http://127.0.0.1:5299',
reuseExistingServer: true,
timeout: 120_000,
env: {
BACKEND_HOST: process.env.BACKEND_HOST || '127.0.0.1',
BACKEND_PORT: process.env.BACKEND_PORT || '8009',
: [
{
command: `PYTHONPATH=.. python3 -m uvicorn dreamverse.mock_server:app --host ${backendHost} --port ${backendPort}`,
url: `http://${backendHost}:${backendPort}/healthz`,
reuseExistingServer: true,
timeout: 120_000,
},
},
{
command: 'npm run dev',
url: 'http://127.0.0.1:5299',
reuseExistingServer: true,
timeout: 120_000,
env: {
BACKEND_HOST: backendHost,
BACKEND_PORT: backendPort,
},
},
],
});
-5199
View File
File diff suppressed because it is too large Load Diff
+2
View File
@@ -2489,6 +2489,8 @@ export default function Page() {
if (devtoolsMode) {
return (
<DevtoolsShell
videoRef={videoRefCallback}
archivedPlaybackRef={archivedPlaybackRefCallback}
connected={connected as boolean}
gpuAssigned={gpuAssigned as boolean}
sessionStarted={sessionStarted as boolean}
@@ -161,6 +161,7 @@ export default function VideoPlayer({
}}
size="icon"
variant="outline"
aria-label={canShare ? "Share video" : "Download video"}
className="absolute top-3 left-3 z-10 cursor-pointer bg-slate-800/50 text-white/90 shadow-md backdrop-blur-sm transition-all border-white/30 hover:bg-slate-800/85 hover:border-white/50 hover:text-white hover:scale-105"
>
{canShare ? <Share className="size-5" /> : <Download className="size-5" />}
@@ -9,6 +9,7 @@ import { Label } from '@/components/ui/label';
import { ScrollArea } from '@/components/ui/scroll-area';
import { Textarea } from '@/components/ui/textarea';
import { cn } from '@/lib/utils';
import LoraControls from '@/components/devtools/LoraControls';
interface DevtoolsDrawerProps {
editableMode?: boolean;
@@ -126,6 +127,8 @@ export default function DevtoolsDrawer({
</div>
<div className="grid gap-4 xl:grid-cols-2">
<LoraControls />
{/* Prompt window memory */}
<details className="group overflow-hidden rounded-2xl border border-border bg-card shadow-sm" open>
<summary className="flex cursor-pointer list-none items-start justify-between gap-4 px-5 py-4">
@@ -113,6 +113,8 @@ interface DevtoolsShellProps {
formatTime?: (seconds: number) => string;
formatDurationMs?: (durationMs: number) => string;
onPlaying?: () => void;
videoRef?: React.RefCallback<HTMLVideoElement>;
archivedPlaybackRef?: React.RefCallback<HTMLVideoElement>;
}
export default function DevtoolsShell({
@@ -209,6 +211,8 @@ export default function DevtoolsShell({
formatTime = (seconds) => `${seconds}`,
formatDurationMs = (durationMs) => `${durationMs}`,
onPlaying = () => {},
videoRef,
archivedPlaybackRef,
}: DevtoolsShellProps) {
return (
<WorkspaceShell
@@ -235,6 +239,8 @@ export default function DevtoolsShell({
workspace={
<div className="flex flex-col gap-4">
<VideoPlayer
videoRef={videoRef}
archivedPlaybackRef={archivedPlaybackRef}
activeClip={activeClip}
sessionStarted={sessionStarted}
avPlaybackStarted={avPlaybackStarted}
@@ -0,0 +1,181 @@
'use client';
import React, { useEffect, useRef, useState } from 'react';
import { Badge } from '@/components/ui/badge';
import { Label } from '@/components/ui/label';
const STYLE_LABELS: Record<string, string> = {
none: 'None',
pixar: 'Pixar Toon',
transition: 'Transition',
};
export default function LoraControls() {
const [styleKeys, setStyleKeys] = useState<string[]>([]);
const [labels, setLabels] = useState<Record<string, string>>(STYLE_LABELS);
const [strength, setStrength] = useState(0.8);
const [enabled, setEnabled] = useState<Record<string, boolean>>({});
const [intensity, setIntensity] = useState<Record<string, number>>({});
const [status, setStatus] = useState('idle');
const strengthRef = useRef(0.8);
const enabledRef = useRef<Record<string, boolean>>({});
const intensityRef = useRef<Record<string, number>>({});
const debounceRef = useRef<ReturnType<typeof setTimeout> | null>(null);
const requestIdRef = useRef(0);
useEffect(() => {
fetch('/lora/options')
.then((response) => (response.ok ? response.json() : null))
.then((data) => {
if (data && Array.isArray(data.styles)) {
const keys = data.styles.filter((s: string) => s !== 'none');
setStyleKeys(keys);
const initIntensity: Record<string, number> = {};
keys.forEach((k: string) => {
initIntensity[k] = 1.0;
});
setIntensity(initIntensity);
intensityRef.current = initIntensity;
}
if (data && data.labels && typeof data.labels === 'object') {
setLabels((prev) => ({ ...prev, ...data.labels }));
}
})
.catch(() => setStatus('options unavailable'));
return () => {
if (debounceRef.current) {
clearTimeout(debounceRef.current);
}
};
}, []);
const applyNow = () => {
const reqId = ++requestIdRef.current;
const stylesPayload: Record<string, number> = {};
for (const key of Object.keys(enabledRef.current)) {
if (enabledRef.current[key]) {
stylesPayload[key] = intensityRef.current[key] ?? 1.0;
}
}
setStatus('applying…');
fetch('/lora', {
method: 'POST',
headers: { 'Content-Type': 'application/json' },
body: JSON.stringify({ strength: strengthRef.current, styles: stylesPayload }),
})
.then(async (response) => {
if (!response.ok) {
throw new Error(await response.text());
}
return response.json();
})
.then((data) => {
if (reqId !== requestIdRef.current) return;
const active = Object.keys(data?.styles ?? {});
const desc = active.length
? active.map((k) => `${labels[k] ?? k}@${data.styles[k]}`).join(' + ')
: 'no style';
setStatus(`applied OmniNFT@${data.strength} + ${desc}`);
})
.catch((error) => {
if (reqId !== requestIdRef.current) return;
setStatus(`error: ${String(error).slice(0, 120)}`);
});
};
const scheduleApply = () => {
if (debounceRef.current) {
clearTimeout(debounceRef.current);
}
debounceRef.current = setTimeout(applyNow, 250);
};
return (
<details
className="group overflow-hidden rounded-2xl border border-border bg-card shadow-sm xl:col-span-2"
open
>
<summary className="flex cursor-pointer list-none items-start justify-between gap-4 px-5 py-4">
<div className="space-y-1">
<span className="block text-lg font-semibold text-foreground">
LoRA stack
</span>
<span className="block text-sm leading-6 text-muted-foreground">
OmniNFT strength + stackable style adapters, each with its own intensity (live).
</span>
</div>
<Badge variant="secondary">{status}</Badge>
</summary>
<div className="space-y-5 border-t border-border px-5 py-4">
<div className="space-y-2">
<Label htmlFor="lora-strength">
OmniNFT strength · {strength.toFixed(2)}
</Label>
<input
id="lora-strength"
type="range"
min={0}
max={1}
step={0.05}
value={strength}
onChange={(event) => {
const next = Number(event.target.value);
setStrength(next);
strengthRef.current = next;
scheduleApply();
}}
className="w-full accent-sky-400"
/>
</div>
{styleKeys.length === 0 ? (
<p className="text-sm text-muted-foreground">
No style adapters for the active model.
</p>
) : (
styleKeys.map((key) => (
<div key={key} className="space-y-2">
<div className="flex items-center justify-between">
<label className="flex cursor-pointer items-center gap-2 text-sm font-medium text-foreground">
<input
type="checkbox"
checked={!!enabled[key]}
onChange={(event) => {
const next = { ...enabledRef.current, [key]: event.target.checked };
enabledRef.current = next;
setEnabled(next);
scheduleApply();
}}
className="accent-sky-400"
/>
{labels[key] ?? STYLE_LABELS[key] ?? key}
</label>
<span className="text-xs text-muted-foreground">
{(intensity[key] ?? 1).toFixed(2)}
</span>
</div>
<input
type="range"
min={0}
max={1}
step={0.05}
value={intensity[key] ?? 1}
disabled={!enabled[key]}
onChange={(event) => {
const next = { ...intensityRef.current, [key]: Number(event.target.value) };
intensityRef.current = next;
setIntensity(next);
scheduleApply();
}}
className="w-full accent-sky-400 disabled:opacity-40"
/>
</div>
))
)}
</div>
</details>
);
}
-106
View File
@@ -1,106 +0,0 @@
# FVD (Fréchet Video Distance) Benchmark
Evaluate generated video quality using FVD with the I3D feature extractor.
## Quick Start
**Run the benchmark:**
```bash
bash benchmarks/scripts/run.sh
```
That's it! The script auto-installs dependencies and runs the benchmark.
**To customize:** Edit `benchmarks/fvd/run_fvd.py` to change:
- Video paths (`real_dir`, `gen_dir`)
- Number of videos, frames, sampling strategy
- Device, batch size, caching, etc.
## Advanced Usage (CLI)
For more control without editing Python files, use the CLI.
**First-time setup** (one-time per pod/environment):
```bash
bash benchmarks/scripts/setup_fvd.sh
```
Then run any configuration you want:
```bash
# Custom configuration
python -m benchmarks.fvd.cli \
--real-path data/real/ \
--gen-path outputs/gen/ \
--num-videos 1024 \
--num-frames 32 \
--clip-strategy random \
--batch-size 32 \
--seed 42 \
--extractor clip
```
**Standard protocols:**
```bash
# Use predefined protocols
python -m benchmarks.fvd.cli \
--real-path data/real/ \
--gen-path outputs/gen/ \
--protocol fvd2048_16f # or fvd2048_128f, quick_test, etc.
```
This would use i3d model by default as the feature extractor
**Feature caching** (speed up repeated evaluations):
```bash
python -m benchmarks.fvd.cli \
--real-path data/real/ \
--gen-path outputs/gen/ \
--protocol fvd2048_16f \
--cache-real-features fvd-cache/extractor_name # Directory path (will save/load fvd-cache/extractor_name/extractor-name_real_features.pkl)
```
Run `python -m benchmarks.fvd.cli --help` for all options.
## Available Protocols
- `fvd2048_16f` - Standard (2048 videos, 16 frames)
- `fvd2048_128f` - Long videos (128 frames)
- `fvd2048_128f_subsample8` - Subsampled long videos
- `quick_test` - Fast testing (10 videos)
## Configuration Options
Key options in `FVDConfig`:
```python
num_videos=2048, # Videos to evaluate
num_frames_per_clip=16, # Frames per clip
clip_strategy='beginning', # beginning|random|uniform|middle|sliding
frame_stride=1, # Frame subsampling
batch_size=32, # GPU batch size
device='cuda', # cuda|cpu
cache_real_features=None, # Cache path for speed
seed=42, # Reproducibility
extractor='i3d', # i3d|clip|videomae
```
## Programmatic Usage
```python
from benchmarks.fvd import compute_fvd_with_config, FVDConfig
config = FVDConfig.fvd2048_16f() # or custom config
results = compute_fvd_with_config('data/real/', 'outputs/gen/', config)
print(f"FVD: {results['fvd']:.2f}")
```
## Notes
- Requires minimum 10 frames per clip
- Supports both video files (.mp4, .avi, etc.) and frame directories
- `--cache-real-features` expects a **directory path** (e.g., `cache/real`), it will automatically create/load `real_features.pkl` inside that directory
-37
View File
@@ -1,37 +0,0 @@
"""
FastVideo Frechet Video Distance (FVD) Benchmark Module.
>>> from fastvideo.benchmarks.fvd import compute_fvd_with_config, FVDConfig
>>> config = FVDConfig.fvd2048_16f() # Standard protocol
>>> results = compute_fvd_with_config('data/real/', 'outputs/gen/', config)
>>> print(f"FVD: {results['fvd']:.2f}")
"""
from .fvd import (
compute_fvd,
compute_fvd_with_config,
compute_frechet_distance,
compute_statistics,
FVDConfig,
)
from .feature_extractors import (BaseFeatureExtractor, I3DFeatureExtractor, load_extractor)
from .video_utils import (
load_video_auto,
sample_clips_from_video,
load_video_clips_streaming,
ClipSamplingStrategy,
)
__all__ = [
'compute_fvd',
'compute_fvd_with_config',
'compute_frechet_distance',
'compute_statistics',
'FVDConfig',
'BaseFeatureExtractor',
'I3DFeatureExtractor',
'load_extractor',
'load_video_auto',
'sample_clips_from_video',
'load_video_clips_streaming',
'ClipSamplingStrategy',
]
-77
View File
@@ -1,77 +0,0 @@
import argparse
import sys
import traceback
from .fvd import compute_fvd_with_config, FVDConfig
def main() -> int:
parser = argparse.ArgumentParser(description='Compute Fréchet Video Distance (FVD)')
# Required arguments
parser.add_argument('--real-path', type=str, required=True, help='Path to real videos')
parser.add_argument('--gen-path', type=str, required=True, help='Path to generated videos')
# Extractor selection
parser.add_argument('--extractor',
type=str,
default='i3d',
choices=['i3d', 'clip', 'videomae'],
help='Feature extractor model to use (default: i3d)')
# Standard args
parser.add_argument('--seed', type=int, default=None, help='Random seed for reproducibility')
parser.add_argument('--protocol',
type=str,
default=None,
choices=['fvd2048_16f', 'fvd2048_128f', 'quick_test'],
help='Use standard protocol (overrides other settings)')
parser.add_argument('--num-videos', type=int, default=2048, help='Number of videos to use')
parser.add_argument('--num-frames', type=int, default=16, help='Number of frames per clip')
parser.add_argument('--clip-strategy', type=str, default='beginning', help='Clip sampling strategy')
parser.add_argument('--batch-size', type=int, default=32, help='Batch size for feature extraction')
parser.add_argument('--device', type=str, default='cuda', help='Device to use (cuda or cpu)')
parser.add_argument('--cache-real-features', type=str, default=None, help='Path to cache real video features')
parser.add_argument('--quiet', action='store_true', help='Suppress progress output')
args = parser.parse_args()
# Create config
if args.protocol:
protocol_map = {
'fvd2048_16f': FVDConfig.fvd2048_16f,
'fvd2048_128f': FVDConfig.fvd2048_128f,
'quick_test': FVDConfig.quick_test,
}
config = protocol_map[args.protocol]()
# Apply overrides
config.device = args.device
config.cache_real_features = args.cache_real_features
config.extractor_model = args.extractor # Apply extractor arg
else:
config = FVDConfig(
num_videos=args.num_videos,
num_frames_per_clip=args.num_frames,
extractor_model=args.extractor, # Apply extractor arg
clip_strategy=args.clip_strategy,
batch_size=args.batch_size,
device=args.device,
cache_real_features=args.cache_real_features,
seed=args.seed)
try:
_ = compute_fvd_with_config(
args.real_path, # noqa: F841
args.gen_path,
config,
verbose=not args.quiet)
return 0
except Exception as e:
print(f"Error: {e}", file=sys.stderr)
traceback.print_exc(file=sys.stderr)
return 1
if __name__ == '__main__':
sys.exit(main())
-232
View File
@@ -1,232 +0,0 @@
"""
Pluggable Feature Extractors for FVD Computation.
Supports I3D (standard), CLIP, and VideoMAE via a common interface.
"""
import torch
import torch.nn as nn
import torch.nn.functional as F
from abc import ABC, abstractmethod
from huggingface_hub import hf_hub_download
from tqdm import tqdm
try:
from transformers import CLIPModel, CLIPProcessor, VideoMAEModel
TRANSFORMERS_AVAILABLE = True
except ImportError:
TRANSFORMERS_AVAILABLE = False
class BaseFeatureExtractor(ABC, nn.Module):
"""Abstract base class for all video feature extractors."""
def __init__(self, device: str = 'cuda'):
super().__init__()
self.device = torch.device(device if torch.cuda.is_available() else 'cpu')
@property
@abstractmethod
def feature_dim(self) -> int:
"""Dimension of the output feature vector."""
pass
@abstractmethod
def preprocess(self, videos: torch.Tensor) -> torch.Tensor:
"""
Args:
videos: [B, T, C, H, W] in [0, 255] range.
Returns:
Preprocessed tensor ready for the model.
"""
pass
@abstractmethod
def extract_features_batch(self, videos: torch.Tensor) -> torch.Tensor:
"""
Extract features for a single batch.
Args:
videos: [B, T, C, H, W] (raw input)
Returns:
Features: [B, feature_dim]
"""
pass
@torch.no_grad()
def extract_features(self, videos: torch.Tensor, batch_size: int = 32, verbose: bool = True) -> torch.Tensor:
"""
Extract features for a large tensor of videos by batching.
"""
N = len(videos)
all_features = []
iterator = range(0, N, batch_size)
if verbose:
iterator = tqdm(iterator, desc=f"Extracting features ({self.__class__.__name__})")
for i in iterator:
batch = videos[i:i + batch_size].to(self.device)
features = self.extract_features_batch(batch)
all_features.append(features.cpu())
return torch.cat(all_features, dim=0)
# 1. I3D Extractor (The Standard FVD Metric)
class I3DFeatureExtractor(BaseFeatureExtractor):
REPO_ID = 'flateon/FVD-I3D-torchscript'
MODEL_FILENAME = 'i3d_torchscript.pt'
def __init__(self, device: str = 'cuda', cache_dir: str | None = None):
super().__init__(device)
self.cache_dir = cache_dir
self.model = self._load_model()
self.model.eval()
self.model.to(self.device)
@property
def feature_dim(self) -> int:
return 400
def _load_model(self) -> torch.nn.Module:
try:
model_path = hf_hub_download(repo_id=self.REPO_ID, filename=self.MODEL_FILENAME, cache_dir=self.cache_dir)
return torch.jit.load(model_path, map_location=self.device)
except Exception as e:
raise RuntimeError(f"Failed to load I3D model: {e}") from e
def preprocess(self, videos: torch.Tensor) -> torch.Tensor:
"""Standard I3D preprocessing: Resize to 224, Norm to [-1, 1]."""
B, T, C, H, W = videos.shape
if T < 10:
raise ValueError(f"I3D requires at least 10 frames, got {T}")
# Normalize to [0, 1]
if videos.max() > 1.0:
videos = videos / 255.0
# Scale to [-1, 1]
videos = videos * 2.0 - 1.0
# Resize to 224x224
if H != 224 or W != 224:
videos = videos.reshape(B * T, C, H, W)
videos = F.interpolate(videos, size=(224, 224), mode='bilinear', align_corners=False)
videos = videos.reshape(B, T, C, 224, 224)
# [B, T, C, H, W] -> [B, C, T, H, W]
return videos.permute(0, 2, 1, 3, 4).contiguous()
def extract_features_batch(self, videos: torch.Tensor) -> torch.Tensor:
batch = self.preprocess(videos)
# TorchScript I3D returns raw logits when return_features=True
return self.model(batch, rescale=False, resize=False, return_features=True)
# 2. CLIP Extractor (Semantic/Content Quality)
class CLIPFeatureExtractor(BaseFeatureExtractor):
def __init__(self, device: str = 'cuda', model_name: str = "openai/clip-vit-base-patch32"):
if not TRANSFORMERS_AVAILABLE:
raise ImportError("Please install transformers: uv pip install transformers")
super().__init__(device)
self.processor = CLIPProcessor.from_pretrained(model_name)
self.model = CLIPModel.from_pretrained(model_name).to(self.device)
self.model.eval()
self._feature_dim = self.model.config.projection_dim
@property
def feature_dim(self) -> int:
return self._feature_dim
def preprocess(self, videos: torch.Tensor) -> torch.Tensor:
# Ensure values are [0, 255]
if videos.max() <= 1.0:
videos = videos * 255.0
return videos.to(torch.uint8)
def extract_features_batch(self, videos: torch.Tensor) -> torch.Tensor:
# Input: [B, T, C, H, W]
B, T, C, H, W = videos.shape
videos = self.preprocess(videos)
# Flatten B*T to treat frames as images
images = videos.view(B * T, C, H, W)
# HF Processor
inputs = self.processor(images=images, return_tensors="pt", padding=True)
inputs = {k: v.to(self.device) for k, v in inputs.items()}
# Extract features [B*T, Dim]
outputs = self.model.get_image_features(**inputs)
# Reshape [B, T, Dim] and Average Pooling over time
outputs = outputs.view(B, T, -1)
return outputs.mean(dim=1)
# 3. VideoMAE Extractor (Structure/Motion Quality)
class VideoMAEFeatureExtractor(BaseFeatureExtractor):
def __init__(self, device: str = 'cuda', model_name: str = "MCG-NJU/videomae-base"):
if not TRANSFORMERS_AVAILABLE:
raise ImportError("Please install transformers: uv pip install transformers")
super().__init__(device)
self.model = VideoMAEModel.from_pretrained(model_name).to(self.device)
self.model.eval()
self.register_buffer('mean', torch.tensor([0.485, 0.456, 0.406], device=self.device).view(1, 1, 3, 1, 1))
self.register_buffer('std', torch.tensor([0.229, 0.224, 0.225], device=self.device).view(1, 1, 3, 1, 1))
@property
def feature_dim(self) -> int:
return self.model.config.hidden_size
def preprocess(self, videos: torch.Tensor) -> torch.Tensor:
"""
Efficient GPU-based preprocessing.
Input: [B, T, C, H, W] in range [0, 255]
"""
B, T, C, H, W = videos.shape
# 1. Resize to 224x224
if H != 224 or W != 224:
videos = videos.view(B * T, C, H, W)
videos = F.interpolate(videos, size=(224, 224), mode='bilinear', align_corners=False)
videos = videos.view(B, T, C, 224, 224)
# 2. Normalize to [0, 1]
if videos.dtype != torch.float32:
videos = videos.float()
if videos.max() > 1.0:
videos = videos / 255.0
# 3. Apply ImageNet Mean/Std
return (videos - self.mean) / self.std
def extract_features_batch(self, videos: torch.Tensor) -> torch.Tensor:
# Input: [B, T, C, H, W]
# Fast GPU Preprocessing
pixel_values = self.preprocess(videos)
# Forward pass
outputs = self.model(pixel_values)
# Global Average Pooling of last hidden state [B, T_patches, 768] -> [B, 768]
return outputs.last_hidden_state.mean(dim=1)
# Factory
def load_extractor(name: str, device: str = 'cuda') -> BaseFeatureExtractor:
name = name.lower()
if name == 'i3d':
return I3DFeatureExtractor(device)
elif name == 'clip':
return CLIPFeatureExtractor(device)
elif name == 'videomae':
return VideoMAEFeatureExtractor(device)
else:
raise ValueError(f"Unknown extractor: {name}. Options: i3d, clip, videomae")
-384
View File
@@ -1,384 +0,0 @@
import numpy as np
import scipy.linalg
import torch
from pathlib import Path
from collections.abc import Iterator
import pickle
from dataclasses import dataclass, field
from .feature_extractors import BaseFeatureExtractor, load_extractor
from .video_utils import ClipSamplingStrategy, load_video_clips_streaming
def compute_statistics(features: np.ndarray) -> tuple[np.ndarray, np.ndarray]:
"""Compute mean and covariance."""
mu = np.mean(features, axis=0)
sigma = np.cov(features, rowvar=False)
return mu, sigma
def compute_frechet_distance(mu1: np.ndarray,
sigma1: np.ndarray,
mu2: np.ndarray,
sigma2: np.ndarray,
eps: float = 1e-6) -> float:
"""
Compute Fréchet distance between two Gaussians.
"""
sigma1 = sigma1 + eps * np.eye(sigma1.shape[0])
sigma2 = sigma2 + eps * np.eye(sigma2.shape[0])
diff = mu1 - mu2
mean_distance = np.sum(diff**2)
trace_sum = np.trace(sigma1 + sigma2)
covmean = scipy.linalg.sqrtm(sigma1 @ sigma2)
if np.iscomplexobj(covmean):
if not np.allclose(np.diagonal(covmean).imag, 0, atol=1e-3):
print(f"Warning: Imaginary component: {np.max(np.abs(covmean.imag))}")
covmean = covmean.real
trace_product = np.trace(covmean)
fvd = mean_distance + trace_sum - 2 * trace_product
return float(fvd)
@dataclass
class FVDConfig:
# default configuration for FVD computation:
# Video selection
num_videos: int = 2048
# Feature Extractor Selection
extractor_model: str = 'i3d' # Options: 'i3d', 'clip', 'videomae'
# Clip sampling
num_frames_per_clip: int = 16
num_clips_per_video: int = 1
clip_strategy: str | ClipSamplingStrategy = 'beginning'
# Temporal subsampling
frame_stride: int = 1 # 1=no subsampling, 2=every 2nd, 8=every 8th
temporal_stride: int = 1 # For sliding window clips
# Data processing
video_extensions: list[str] = field(default_factory=lambda: ['.mp4', '.avi', '.mov', '.mkv'])
support_frame_dirs: bool = True
# Computation
batch_size: int = 32
device: str = 'cuda'
use_streaming: bool = True
resize_before_extraction: bool = True
# Caching
cache_real_features: str | None = None
i3d_model_path: str | None = None
# Reproducibility
seed: int | None = None
@classmethod
def fvd2048_16f(cls) -> 'FVDConfig':
"""Standard FVD protocol: 2048 videos, 16 frames, beginning clip."""
return cls(num_videos=2048, num_frames_per_clip=16, clip_strategy='beginning', use_streaming=True)
@classmethod
def fvd2048_128f(cls) -> 'FVDConfig':
"""Long video protocol: 2048 videos, 128 frames."""
return cls(num_videos=2048, num_frames_per_clip=128, clip_strategy='beginning', use_streaming=True)
@classmethod
def quick_test(cls) -> 'FVDConfig':
"""Quick test config: 100 videos, 16 frames."""
return cls(num_videos=100, num_frames_per_clip=16, clip_strategy='beginning')
def to_dict(self) -> dict:
"""Export config to dict for logging"""
d = self.__dict__.copy()
d['clip_strategy'] = str(self.clip_strategy)
return d
def __str__(self) -> str:
"""Human-readable protocol name"""
desc = f"FVD_{self.extractor_model.upper()}_{self.num_videos}_{self.num_frames_per_clip}f"
if self.frame_stride > 1:
desc += f"_subsample{self.frame_stride}"
if self.num_clips_per_video > 1:
desc += f"_{self.num_clips_per_video}clips"
if self.clip_strategy != 'beginning':
desc += f"_{self.clip_strategy}"
return desc
def extract_features_streaming(video_generator: Iterator[torch.Tensor],
extractor: BaseFeatureExtractor,
batch_size: int = 32,
max_clips: int | None = None,
verbose: bool = True) -> np.ndarray:
"""
Extract features from a video clip generator using streaming.
"""
all_features = []
batch = []
if verbose:
print(f"Extracting features with batch_size={batch_size}...")
with torch.no_grad():
for clip_count, clip in enumerate(video_generator):
batch.append(clip)
# Process batch when full
if len(batch) == batch_size:
batch_tensor = torch.stack(batch).to(extractor.device)
features = extractor.extract_features_batch(batch_tensor)
all_features.append(features.detach().cpu().numpy())
batch = []
if verbose and clip_count % (batch_size * 10) == 0:
print(f"Processed {clip_count} clips...")
if max_clips is not None and clip_count >= max_clips:
break
# Process remaining clips
if len(batch) > 0:
batch_tensor = torch.stack(batch).to(extractor.device)
features = extractor.extract_features_batch(batch_tensor)
all_features.append(features.detach().cpu().numpy())
if len(all_features) == 0:
raise RuntimeError("No features extracted - check video loading")
features = np.concatenate(all_features, axis=0)
if verbose:
print(f"Extracted {len(features)} feature vectors")
return features
def load_or_compute_features(videos: str | Path | torch.Tensor,
extractor: BaseFeatureExtractor,
config: FVDConfig,
cache_path: str | None = None,
cache_name: str = "real_features") -> np.ndarray:
"""Load features from cache or compute (with streaming support)"""
if cache_path is not None:
script_dir = Path(__file__).parent
cache_dir = script_dir / cache_path
cache_file = cache_dir / f"{config.extractor_model}_{cache_name}.pkl"
if cache_file.exists():
print(f"Loading cached features from {cache_file}")
with open(cache_file, 'rb') as f:
features = pickle.load(f)
# Validate and limit based on config
max_features = config.num_videos * config.num_clips_per_video
if len(features) < max_features:
print(f"WARNING: Cache has {len(features)} features but need {max_features}")
print("Cached features insufficient - will recompute...")
elif len(features) > max_features:
print(f"Using {max_features} features from cache (truncated from {len(features)})")
features = features[:max_features]
return features
else:
print(f"Using all {len(features)} cached features")
return features
print("Computing features from scratch...")
if isinstance(videos, (str | Path)):
target_size = (224, 224) if config.resize_before_extraction else None
video_generator = load_video_clips_streaming(videos,
num_frames=config.num_frames_per_clip,
max_videos=config.num_videos,
clip_strategy=config.clip_strategy,
frame_stride=config.frame_stride,
num_clips_per_video=config.num_clips_per_video,
video_extensions=config.video_extensions,
support_frame_dirs=config.support_frame_dirs,
target_size=target_size,
verbose=True)
max_clips = config.num_videos * config.num_clips_per_video
features = extract_features_streaming(video_generator,
extractor,
batch_size=config.batch_size,
max_clips=max_clips,
verbose=True)
else:
print(f"Extracting features from {len(videos)} video tensors...")
features = extractor.extract_features(videos, batch_size=config.batch_size, verbose=True)
features = features.numpy()
# Validate feature count
expected_count = config.num_videos * config.num_clips_per_video
if len(features) < expected_count:
raise ValueError(f"ERROR: Only extracted {len(features)} features, but need {expected_count}!\n"
f"Found fewer videos than expected. Check your video directory.")
elif len(features) > expected_count:
print(f"Truncating {len(features)} features to {expected_count}")
features = features[:expected_count]
# Cache features if requested
if cache_path is not None:
script_dir = Path(__file__).parent
cache_dir = script_dir / cache_path
cache_dir.mkdir(parents=True, exist_ok=True)
cache_file = cache_dir / f"{config.extractor_model}_{cache_name}.pkl"
print(f"Caching features to {cache_file}")
with open(cache_file, 'wb') as f:
pickle.dump(features, f)
return features
def compute_fvd_with_config(real_videos: str | Path | torch.Tensor,
gen_videos: str | Path | torch.Tensor,
config: FVDConfig,
verbose: bool = True) -> dict:
"""
Compute FVD using a standardized configuration.
This is the recommended way to compute FVD for reproducibility.
Args:
real_videos: Path or tensors
gen_videos: Path or tensors
config: FVDConfig specifying protocol
verbose: Print progress
Returns:
results: Dictionary with:
- 'fvd': FVD score (float)
- 'protocol': Protocol name (str)
- 'model': Feature extractor model name (str)
- 'config': Configuration dict
Example:
>>> config = FVDConfig.fvd2048_16f()
>>> results = compute_fvd_with_config('data/real/', 'outputs/gen/', config)
>>> print(f"FVD: {results['fvd']:.2f}")
"""
# Seed for reproducibility
if config.seed is not None:
import random as _rnd
_rnd.seed(config.seed)
np.random.seed(config.seed)
torch.manual_seed(config.seed)
if torch.cuda.is_available():
torch.cuda.manual_seed_all(config.seed)
if verbose:
print("=" * 70)
print(f"Computing FVD with protocol: {config}")
print(f"Model: {config.extractor_model.upper()}")
print("=" * 70)
print("\nConfiguration:")
for key, value in config.to_dict().items():
print(f" {key}: {value}")
print()
# Initialize Extractor using Factory
if verbose:
print(f"\nInitializing {config.extractor_model.upper()} model on {config.device}...")
extractor = load_extractor(config.extractor_model, device=config.device)
# Extract features
if verbose:
print(f"\n{'='*70}")
print("Extracting REAL video features...")
print(f"{'='*70}")
real_features = load_or_compute_features(videos=real_videos,
extractor=extractor,
config=config,
cache_path=config.cache_real_features,
cache_name="real_features")
if verbose:
print(f"\n{'='*70}")
print("Extracting GENERATED video features...")
print(f"{'='*70}")
gen_features = load_or_compute_features(videos=gen_videos,
extractor=extractor,
config=config,
cache_path=None,
cache_name="gen_features")
if verbose:
print(f"\nReal videos/clips: {len(real_features)}")
print(f"Generated videos/clips: {len(gen_features)}")
print(f"\n{'='*70}")
print("Computing statistics...")
print(f"{'='*70}")
mu_real, sigma_real = compute_statistics(real_features)
mu_gen, sigma_gen = compute_statistics(gen_features)
if verbose:
print(f"\n{'='*70}")
print("Computing Fréchet distance...")
print(f"{'='*70}")
fvd = compute_frechet_distance(mu_real, sigma_real, mu_gen, sigma_gen)
if verbose:
print(f"\n{'='*70}")
print(f"FVD Score ({config.extractor_model.upper()}): {fvd:.4f}")
print(f"Protocol: {config}")
print(f"{'='*70}\n")
results = {
'fvd': fvd,
'protocol': str(config),
'model': config.extractor_model,
'config': config.to_dict(),
}
return results
def compute_fvd(real_videos: str | Path | torch.Tensor,
gen_videos: str | Path | torch.Tensor,
num_frames: int = 16,
batch_size: int = 32,
device: str = 'cuda',
num_videos: int | None = 2048,
cache_real_features: str | None = None,
i3d_model_path: str | None = None,
seed: int | None = None,
verbose: bool = True) -> float:
"""
Backward compatibility wrapper for computing FVD (defaults to I3D).
"""
num_videos = num_videos if num_videos is not None else 2048
config = FVDConfig(
num_videos=num_videos,
num_frames_per_clip=num_frames,
extractor_model='i3d', # Default to I3D
batch_size=batch_size,
device=device,
cache_real_features=cache_real_features,
i3d_model_path=i3d_model_path,
seed=seed,
)
result = compute_fvd_with_config(real_videos, gen_videos, config, verbose)
return result['fvd']
-124
View File
@@ -1,124 +0,0 @@
"""I3D Feature Extractor for FVD Computation"""
import torch
import torch.nn as nn
import torch.nn.functional as F
from pathlib import Path
from huggingface_hub import hf_hub_download
from tqdm import tqdm
from contextlib import suppress
class I3DFeatureExtractor(nn.Module):
"""
I3D feature extractor for FVD computation.
Extracts 400-dimensional features from videos using I3D model
trained on Kinetics-400.
"""
REPO_ID = 'flateon/FVD-I3D-torchscript'
MODEL_FILENAME = 'i3d_torchscript.pt'
def __init__(self, device: str = 'cuda', cache_dir: str | Path | None = None):
super().__init__()
self.device_str = device
if device == 'cuda' and not torch.cuda.is_available():
print("Warning: CUDA requested but not available – falling back to CPU")
self.device = torch.device('cpu')
else:
self.device = torch.device(device)
self.cache_dir: str | None
if cache_dir is not None:
self.cache_dir = str(Path(cache_dir).resolve())
else:
self.cache_dir = None # Use HF default cache
self.model = self._load_model()
self.model.eval()
with suppress(Exception):
self.model.to(self.device)
def _load_model(self) -> torch.nn.Module:
"""Download and load I3D TorchScript model from Hugging Face Hub."""
print(f"Loading I3D model from Hugging Face Hub ({self.REPO_ID})...")
try:
# Download model from Hugging Face Hub
model_path = hf_hub_download(repo_id=self.REPO_ID, filename=self.MODEL_FILENAME, cache_dir=self.cache_dir)
# Load directly to chosen device
model = torch.jit.load(model_path, map_location=self.device)
print("I3D model loaded successfully")
return model
except Exception as e:
raise RuntimeError(f"Failed to load I3D model from Hugging Face Hub. Error: {e}\n"
f"Ensure you have internet connection and huggingface_hub installed:\n"
f"uv pip install huggingface_hub") from e
def preprocess(self, videos: torch.Tensor) -> torch.Tensor:
"""
Preprocess videos for I3D.
Args:
videos: [B, T, C, H, W], values in [0, 255]
Returns:
Preprocessed videos [B, C, T, 224, 224] (normalized and resized)
"""
B, T, C, H, W = videos.shape
if T < 10:
raise ValueError(f"I3D requires at least 10 frames, got {T}")
# Normalize to [0, 1] if needed
if videos.max() > 1.0:
videos = videos / 255.0
# Resize to 224x224 if needed
if H != 224 or W != 224:
videos = videos.reshape(B * T, C, H, W)
videos = F.interpolate(videos, size=(224, 224), mode='bilinear', align_corners=False)
videos = videos.reshape(B, T, C, 224, 224)
# Convert to [B, C, T, H, W] format
videos = videos.permute(0, 2, 1, 3, 4).contiguous()
return videos
@torch.no_grad()
def extract_features(self, videos: torch.Tensor, batch_size: int = 32, verbose: bool = True) -> torch.Tensor:
"""
Extract I3D features
Args:
videos: [N, T, C, H, W], values in [0, 255]
batch_size: Batch size for processing
verbose: Show progress bar
Returns:
Features [N, 400]
"""
N = len(videos)
all_features = []
iterator = range(0, N, batch_size)
if verbose:
iterator = tqdm(iterator, desc="Extracting I3D features")
for i in iterator:
batch = videos[i:i + batch_size].to(self.device)
batch = self.preprocess(batch) # Now returns [B, C, T, H, W]
# Use the HF model without rescale/resize (we handle it in preprocess)
features = self.model(batch, rescale=False, resize=False, return_features=True)
all_features.append(features.cpu())
return torch.cat(all_features, dim=0)
def __call__(self, videos: torch.Tensor, batch_size: int = 32) -> torch.Tensor:
return self.extract_features(videos, batch_size=batch_size)
-51
View File
@@ -1,51 +0,0 @@
import sys
from pathlib import Path
root_dir = Path(__file__).parent.parent.parent
sys.path.insert(0, str(root_dir))
from benchmarks.fvd.fvd import FVDConfig, compute_fvd_with_config # noqa: E402
def main() -> None:
script_dir = Path(__file__).parent.resolve()
# Define directories
real_dir = "benchmarks/data/real_videos"
gen_dir = "benchmarks/data/generated_videos"
# Compare all 3 models
models_to_test = ['i3d', 'clip', 'videomae']
print(f"\n{'='*60}")
print("STARTING COMPARISON BENCHMARK")
print(f"{'='*60}")
for model_name in models_to_test:
print(f"\n>>> Running evaluation with {model_name.upper()}...")
try:
cfg = FVDConfig(
num_videos=650,
num_frames_per_clip=16,
extractor_model=model_name,
clip_strategy='beginning',
device='cuda',
seed=42,
# Use separate cache folders for each model to avoid conflicts
cache_real_features=str(script_dir / f'fvd-cache/{model_name}'),
)
results = compute_fvd_with_config(real_dir, gen_dir, cfg, verbose=False)
print(f"FVD: {results['fvd']}\nModel: {results['model']}")
except Exception as e:
print(f"{model_name.upper()} Failed: {e}")
print(f"\n{'='*60}")
print("BENCHMARK COMPLETE")
print(f"{'='*60}")
if __name__ == '__main__':
main()
-89
View File
@@ -1,89 +0,0 @@
#!/usr/bin/env python3
import sys
from pathlib import Path
import shutil
import random
from fvd import compute_fvd_with_config, FVDConfig
script_path = Path(__file__).resolve()
fastvideo_root = script_path.parent.parent.parent
sys.path.insert(0, str(fastvideo_root))
def split_videos(video_dir: Path, n_per_subset: int = 128, seed: int = 42):
subset_a = video_dir.parent / 'bair_full_subset_A'
subset_b = video_dir.parent / 'bair_full_subset_B'
if subset_a.exists():
shutil.rmtree(subset_a)
if subset_b.exists():
shutil.rmtree(subset_b)
subset_a.mkdir(parents=True)
subset_b.mkdir(parents=True)
videos = sorted(video_dir.glob('*.mp4'))
random.seed(seed)
shuffled = list(videos)
random.shuffle(shuffled)
needed = n_per_subset * 2
if len(shuffled) > needed:
shuffled = shuffled[:needed]
mid = len(shuffled) // 2
print(f"\nSplitting {len(shuffled)} BAIR FULL videos:")
print(f" Subset A: {mid} videos")
print(f" Subset B: {len(shuffled) - mid} videos")
for v in shuffled[:mid]:
shutil.copy2(v, subset_a / v.name)
for v in shuffled[mid:]:
shutil.copy2(v, subset_b / v.name)
return subset_a, subset_b, mid
def validate_fvd(subset_a: Path, subset_b: Path, num_videos: int):
config = FVDConfig(num_videos=num_videos,
num_frames_per_clip=16,
clip_strategy='beginning',
batch_size=8,
device='cuda',
seed=42)
print("\n" + "=" * 70)
print("TEST 1: Identity Test")
print("=" * 70)
result1 = compute_fvd_with_config(real_videos=str(subset_a), gen_videos=str(subset_a), config=config, verbose=False)
fvd_identity = result1['fvd']
print(f"\nIdentity FVD: {fvd_identity:.2f}")
print("\n" + "=" * 70)
print("TEST 2: Real vs Real")
print("=" * 70)
result2 = compute_fvd_with_config(real_videos=str(subset_a), gen_videos=str(subset_b), config=config, verbose=False)
fvd_real = result2['fvd']
print(f"\nReal vs Real FVD: {fvd_real:.2f}")
print("\n" + "=" * 70)
print("RESULTS")
print("=" * 70)
print(f"Identity: {fvd_identity:.2f}")
print(f"Real vs Real: {fvd_real:.2f}")
def main() -> None:
bair_dir = Path('benchmarks/data/bair_full_videos')
subset_a, subset_b, count = split_videos(bair_dir, n_per_subset=128, seed=42)
validate_fvd(subset_a, subset_b, count)
if __name__ == '__main__':
main()
-460
View File
@@ -1,460 +0,0 @@
import torch
import cv2
import numpy as np
from pathlib import Path
from collections.abc import Iterator
from tqdm import tqdm
from enum import Enum
class ClipSamplingStrategy(Enum):
"""Clip sampling strategies for FVD evaluation."""
BEGINNING = 'beginning' # Take first N frames (most common)
RANDOM = 'random' # Random N consecutive frames
UNIFORM = 'uniform' # Uniformly spaced frames across video
MIDDLE = 'middle' # Middle N frames
SLIDING = 'sliding' # Multiple sliding windows
ALL = 'all' # All possible clips
def _load_video_cv2(video_path: str | Path,
num_frames: int | None = 16,
sample_strategy: str = 'uniform') -> torch.Tensor:
"""
Load video from video file using OpenCV.
Args:
video_path: Path to video file (MP4, AVI, MOV, MKV)
num_frames: Number of frames to extract
sample_strategy: 'uniform' or 'random'
Returns:
video: [T, C, H, W]
"""
video_path = str(video_path)
cap = cv2.VideoCapture(video_path)
if not cap.isOpened():
raise RuntimeError(f"Cannot open video: {video_path}")
frames = []
total_frames = int(cap.get(cv2.CAP_PROP_FRAME_COUNT))
if num_frames is None:
# Read all available frames
while True:
ret, frame = cap.read()
if not ret:
break
frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
frames.append(frame)
cap.release()
if len(frames) == 0:
raise RuntimeError(f"Video has 0 frames: {video_path}")
frames = np.stack(frames) # [T, H, W, C]
frames = torch.from_numpy(frames).permute(0, 3, 1, 2).float() # [T, C, H, W]
return frames
if total_frames == 0:
raise RuntimeError(f"Video has 0 frames: {video_path}")
# Determine frame indices for sampling
if total_frames < num_frames:
frame_indices = list(range(total_frames)) + [total_frames - 1] * (num_frames - total_frames)
elif sample_strategy == 'uniform':
frame_indices = np.linspace(0, total_frames - 1, num_frames, dtype=int).tolist()
elif sample_strategy == 'random':
frame_indices = sorted(np.random.choice(total_frames, num_frames, replace=False))
else:
raise ValueError(f"Unknown sample_strategy: {sample_strategy}")
# Extract frames
for idx in frame_indices:
cap.set(cv2.CAP_PROP_POS_FRAMES, idx)
ret, frame = cap.read()
if not ret:
if len(frames) > 0:
frames.append(frames[-1].copy())
else:
h = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))
w = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))
frames.append(np.zeros((h, w, 3), dtype=np.uint8))
continue
frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
frames.append(frame)
cap.release()
frames = np.stack(frames) # [T, H, W, C]
frames = torch.from_numpy(frames).permute(0, 3, 1, 2).float() # [T, C, H, W]
return frames
def _load_video_from_frames(frame_dir: str | Path,
num_frames: int | None = 16,
sample_strategy: str = 'uniform',
frame_extensions: list[str] | None = None) -> torch.Tensor:
"""
Load video from directory of frame images.
Args:
frame_dir: Directory containing frames
num_frames: Number of frames to sample
sample_strategy: 'uniform' or 'random'
frame_extensions: Image file extensions to look for
Returns:
video: [T, C, H, W]
"""
if frame_extensions is None:
frame_extensions = ['.jpg', '.png', '.jpeg', '.bmp']
frame_dir = Path(frame_dir)
if not frame_dir.exists():
raise FileNotFoundError(f"Frame directory not found: {frame_dir}")
# Find all frames
frame_files: list[Path] = []
for ext in frame_extensions:
frame_files.extend(frame_dir.glob(f"*{ext}"))
if len(frame_files) == 0:
raise ValueError(f"No frames found in {frame_dir} with extensions {frame_extensions}")
frame_files = sorted(frame_files, key=lambda x: x.name)
total_frames = len(frame_files)
# Determine frame indices
if num_frames is None:
frame_indices = list(range(total_frames))
else:
if total_frames < num_frames:
frame_indices = list(range(total_frames)) + [total_frames - 1] * (num_frames - total_frames)
elif sample_strategy == 'uniform':
frame_indices = np.linspace(0, total_frames - 1, num_frames, dtype=int).tolist()
elif sample_strategy == 'random':
frame_indices = sorted(np.random.choice(total_frames, num_frames, replace=False))
else:
raise ValueError(f"Unknown sample_strategy: {sample_strategy}")
# Load frames
frames = []
for idx in frame_indices:
frame_path = frame_files[idx]
frame = cv2.imread(str(frame_path))
if frame is None:
raise RuntimeError(f"Failed to load frame: {frame_path}")
frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
frames.append(frame)
# Stack and convert to tensor
frames = np.stack(frames) # [T, H, W, C]
frames = torch.from_numpy(frames).permute(0, 3, 1, 2).float() # [T, C, H, W]
return frames
def _detect_video_format(path: str | Path) -> str:
"""
Detect if path is a video file or frame directory.
Returns:
'video_file', 'frame_directory', or 'unknown'
"""
path = Path(path)
if path.is_file():
return 'video_file'
elif path.is_dir():
# Check if contains image files
image_extensions = ['.jpg', '.jpeg', '.png', '.bmp']
for ext in image_extensions:
if list(path.glob(f"*{ext}")):
return 'frame_directory'
return 'unknown'
else:
raise ValueError(f"Path does not exist: {path}")
def load_video_auto(video_path: str | Path,
num_frames: int | None = 16,
sample_strategy: str = 'uniform') -> torch.Tensor:
"""
Automatically detect format and load video.
Supports:
- Video files (MP4, AVI, MOV, MKV)
- Frame directories (JPG, PNG)
Args:
video_path: Path to video file or frame directory
num_frames: Number of frames to extract
sample_strategy: 'uniform' or 'random'
Returns:
video: [T, C, H, W]
"""
format_type = _detect_video_format(video_path)
if format_type == 'video_file':
return _load_video_cv2(video_path, num_frames, sample_strategy)
elif format_type == 'frame_directory':
return _load_video_from_frames(video_path, num_frames, sample_strategy)
else:
raise ValueError(f"Unknown video format at {video_path}")
def sample_clips_from_video(video: torch.Tensor,
num_frames_per_clip: int = 16,
num_clips: int = 1,
strategy: str | ClipSamplingStrategy = ClipSamplingStrategy.BEGINNING,
frame_stride: int = 1,
temporal_stride: int = 1) -> list[torch.Tensor]:
"""
Sample clips from a video with various strategies.
Args:
video: [T, C, H, W] full video
num_frames_per_clip: Frames per clip
num_clips: Number of clips to extract
strategy: ClipSamplingStrategy or string ('beginning', 'random', etc.)
frame_stride: Skip frames (FPS control: 1=all, 2=every 2nd, 8=every 8th)
temporal_stride: Stride between clips for sliding window
Returns:
List of clips, each [num_frames_per_clip, C, H, W]
Examples:
>>> # Beginning clip (most common for FVD)
>>> clips = sample_clips_from_video(video, 16, strategy='beginning')
>>> # Multiple random clips
>>> clips = sample_clips_from_video(video, 16, num_clips=4, strategy='random')
>>> # Subsample FPS by 2x (every 2nd frame)
>>> clips = sample_clips_from_video(video, 16, frame_stride=2)
>>> # Sliding window with overlap
>>> clips = sample_clips_from_video(video, 16, strategy='sliding', temporal_stride=8)
"""
# Convert string to enum if needed
if isinstance(strategy, str):
strategy = ClipSamplingStrategy(strategy)
T, C, H, W = video.shape
# Apply frame stride (FPS subsampling)
if frame_stride > 1:
video = video[::frame_stride]
T = len(video)
effective_clip_length = num_frames_per_clip
# Handle videos shorter than clip length
if effective_clip_length > T:
pad_length = effective_clip_length - T
last_frame = video[-1:].repeat(pad_length, 1, 1, 1)
video = torch.cat([video, last_frame], dim=0)
T = len(video)
clips = []
if strategy == ClipSamplingStrategy.BEGINNING:
# Take first clip (most common for FVD evaluation)
clip = video[:effective_clip_length]
clips.append(clip)
elif strategy == ClipSamplingStrategy.MIDDLE:
# Take middle clip
start = (T - effective_clip_length) // 2
clip = video[start:start + effective_clip_length]
clips.append(clip)
elif strategy == ClipSamplingStrategy.RANDOM:
# Sample N random clips
for _ in range(num_clips):
start = 0 if effective_clip_length == T else np.random.randint(0, T - effective_clip_length + 1)
clip = video[start:start + effective_clip_length]
clips.append(clip)
elif strategy == ClipSamplingStrategy.UNIFORM:
# Uniformly spaced clips
if num_clips == 1:
# Single clip from middle
start = (T - effective_clip_length) // 2
clip = video[start:start + effective_clip_length]
clips.append(clip)
else:
# Multiple uniformly spaced clips
step = (T - effective_clip_length) / (num_clips - 1) if num_clips > 1 else 0
for i in range(num_clips):
start = int(i * step)
start = min(start, T - effective_clip_length)
clip = video[start:start + effective_clip_length]
clips.append(clip)
elif strategy == ClipSamplingStrategy.SLIDING:
# Sliding window with stride
for start in range(0, T - effective_clip_length + 1, temporal_stride):
clip = video[start:start + effective_clip_length]
clips.append(clip)
if len(clips) >= num_clips:
break
elif strategy == ClipSamplingStrategy.ALL:
# All possible clips (overlapping)
for start in range(T - effective_clip_length + 1):
clip = video[start:start + effective_clip_length]
clips.append(clip)
else:
raise ValueError(f"Unknown strategy: {strategy}")
return clips
def load_video_clips_streaming(directory: str | Path,
num_frames: int = 16,
max_videos: int | None = None,
clip_strategy: str
| ClipSamplingStrategy = 'beginning',
frame_stride: int = 1,
num_clips_per_video: int = 1,
video_extensions: list[str] | None = None,
support_frame_dirs: bool = True,
target_size: tuple[int, int] | None = (224, 224),
verbose: bool = True) -> Iterator[torch.Tensor]:
"""
This generator yields clips one-by-one instead of loading all videos into RAM.
Perfect for large datasets where memory is limited.
Args:
directory: Path to directory with videos
num_frames: Frames per clip
max_videos: Max videos to load
clip_strategy: 'beginning', 'random', 'uniform', etc.
frame_stride: Frame skip (1=all, 2=every 2nd, 8=every 8th)
num_clips_per_video: Number of clips per video
video_extensions: Video file extensions
support_frame_dirs: Also load frame directories
target_size: Resize clips to (H, W). If None, keep original size.
verbose: Show progress
Yields:
clip: [T, C, H, W] individual clips
Example:
>>> for clip in load_video_clips_streaming('data/videos/', num_frames=16):
>>> features = model.extract_features(clip.unsqueeze(0))
>>> # Process one clip at a time - low memory usage!
"""
if video_extensions is None:
video_extensions = ['.mp4', '.avi', '.mov', '.mkv']
directory = Path(directory)
if not directory.exists():
raise FileNotFoundError(f"Directory not found: {directory}")
# Find video paths
video_paths: list[Path] = []
# Find video files
for ext in video_extensions:
video_paths.extend(directory.glob(f"**/*{ext}"))
# Find frame directories if enabled
if support_frame_dirs:
for subdir in directory.iterdir():
if subdir.is_dir():
# Check if it contains frames
image_extensions = ['.jpg', '.jpeg', '.png', '.bmp']
for ext in image_extensions:
if list(subdir.glob(f"*{ext}")):
video_paths.append(subdir)
break
if len(video_paths) == 0:
raise ValueError(f"No videos found in {directory}")
video_paths = sorted(video_paths)
if max_videos is not None:
video_paths = video_paths[:max_videos]
if verbose:
print(f"Found {len(video_paths)} videos in {directory}")
if num_clips_per_video > 1:
print(f"Extracting {num_clips_per_video} clips per video...")
if frame_stride > 1:
print(f"Subsampling frames with stride {frame_stride}...")
if target_size:
print(f"Resizing clips to {target_size}...")
# Track statistics
failed_count = 0
total_clips = 0
iterator = tqdm(video_paths, desc="Loading videos") if verbose else video_paths
for video_path in iterator:
try:
# Load full video
video = load_video_auto(video_path, num_frames=None, sample_strategy='uniform')
# Sample clips from video
clips = sample_clips_from_video(video,
num_frames_per_clip=num_frames,
num_clips=num_clips_per_video,
strategy=clip_strategy,
frame_stride=frame_stride)
if target_size is not None:
resized_clips = []
for clip in clips:
T, C, H, W = clip.shape
if target_size != (H, W):
# Resize to target size
clip = clip.contiguous() # Fix non-contiguous tensors first
clip_flat = clip.view(T * C, H, W).unsqueeze(0) # [1, T*C, H, W]
clip_resized = torch.nn.functional.interpolate(clip_flat,
size=target_size,
mode='bilinear',
align_corners=False)
clip = clip_resized.squeeze(0).view(T, C, target_size[0],
target_size[1]) # Back to [T, C, H, W]
resized_clips.append(clip)
clips = resized_clips
# Yield clips one by one
for clip in clips:
yield clip
total_clips += 1
# Free memory
del video, clips
except Exception as e:
failed_count += 1
if verbose:
print(f"\nWarning: Failed to load {video_path}: {e}")
continue
# Validate
if total_clips == 0:
raise RuntimeError(f"Failed to load any videos from {directory}")
failure_rate = failed_count / len(video_paths)
if failure_rate > 0.1: # More than 10% failed
print(f"\nWARNING: {failure_rate:.1%} of videos failed to load ({failed_count}/{len(video_paths)})")
if verbose:
print(f"\nSuccessfully loaded {total_clips} clips from {len(video_paths) - failed_count} videos")
-7
View File
@@ -1,7 +0,0 @@
#!/bin/bash
# 1. Install missing dependency
uv pip install -q opencv-python-headless transformers huggingface_hub
# 2. Run FVD script
python benchmarks/fvd/run_fvd.py
-4
View File
@@ -1,4 +0,0 @@
#!/bin/bash
# 1. Install missing dependency
uv pip install -q opencv-python-headless
+56
View File
@@ -0,0 +1,56 @@
# FastVideo — Design Philosophy
One page on *why* FastVideo is built the way it is. The full architecture, the as-built status, and the
forward roadmap live in **[`v2/README.md`](v2/README.md)** — this is the philosophy beneath it.
---
**A deployable model is a post-training artifact.** Unlike an LLM — where inference optimizes frozen weights
after the fact — a *usable* video/omni model is *created* by training: step distillation for latency, QAT for
precision, distillation + self-forcing for causal/world models. So every inference capability is a
**(recipe, runtime) pair**: the weights and the loop that produced-and-assumes them are one versioned object.
This is the source of the moat — whoever owns *both* sides of the pair owns the optimization frontier — and it
is why training and serving cannot be two systems.
**The work is loops, not `forward()`.** Denoise timesteps, AR decode, chunked rollout, VAE tiles, encoder
chunks, audio tokens, reward batches, optimizer steps, media chunks — video and omni inference is iteration. A
runtime that collapses everything to a single `forward` can't schedule, batch, cancel, stream, reserve memory
for, or capture the behavior of what actually runs. So loops are first-class, and they are **driven**: the
model describes the next step it needs, the runtime decides when and with whom it runs, the model folds the
result back. The model keeps content-adaptive control flow; the runtime keeps admission, batching, streaming,
and behavior capture. Per-request state lives in typed `LoopState`, never in module globals — so interleaving
requests through one model instance cannot smear state, by construction.
**The model is the center; everything else is a view over it.** A typed `ModelCard` owns components, loops,
the recipe, and the parity contract. Programs compose a card's loops into a task; Workflows compose cards into
pipelines; the scheduler runs the *steps* of all loops as `WorkUnit`s under one currency (predicted GPU-time,
because a bidirectional denoise step and an AR token are ~1000× apart and incommensurable in counts);
deployment places and routes; products stream artifacts. None of them define model semantics — they reference
the Model Plane. One resident instance can run many loop types on shared weights, which is what makes omni/MoT
native rather than a DAG that doubles weights.
**Correctness is a typed contract, not a hope.** Caches are correct by *key* — if a field can change output
semantics it is in the key, so reuse is partitioned, never blindly flushed. Parity between the train-forward
and the serve-forward is *measured* on a declared ladder (component → loop → behavioral → distribution →
artifact-quality), never assumed. And the non-negotiable gate is **interleave bit-parity**: N requests
interleaved at step granularity must be bit-identical to running them serially — the test the whole
loop-inversion bet lives or dies on.
**One substrate for inference, training, and RL.** The rollout forward *is* the serve forward plus capture —
same loop, same caches, same batcher, same numerics — so every serving optimization is automatically a rollout
optimization, and there is one numerics surface the ladder measures rather than a correction layer papering
over it. The engine doubles as the RL rollout engine under a strict rule: `training` consumes the engine; the
**engine never imports `training`**.
**Borrow aggressively; copy nothing as the core.** vLLM/SGLang scheduling, vLLM-Omni/SGLang-Omni omni serving,
Dynamo fleet orchestration, diffusers components, xDiT parallelism, TorchTitan mesh discipline,
verl-omni/miles RL lessons, ComfyUI workflows, Dreamverse/LiveKit sessions — each contributes a take, none is
the center. Deployment orchestration (Dynamo) sits *above* the engine, never inside it. Extensions are
versioned hook points, never monkeypatching. New frontier capabilities arrive as a card, a method, a loop, a
workflow, or a controller — **not a rewrite**.
> A model card is a (recipe, runtime) pair with a parity obligation. The model owns loop semantics; the runtime
> owns loop lifecycle. One resident instance runs many loops; one scheduler runs their steps in one currency.
> Caches are correct by key; parity is correct by test; the interleave gate is non-negotiable. Training records
> behavior on the same loops it serves. Deployment places and routes; products stream artifacts; neither defines
> the model.
+4 -2
View File
@@ -55,12 +55,14 @@ RUN source $HOME/.local/bin/env && \
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
# Install FastVideo Unified Kernel
# Install FastVideo Unified Kernel.
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
# of probing a live device for the arch (matches the released kernel wheel).
RUN source $HOME/.local/bin/env && \
source /opt/venv/bin/activate && \
cd fastvideo-kernel && \
git submodule update --init --recursive && \
./build.sh
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
EXPOSE 22
+4 -2
View File
@@ -55,12 +55,14 @@ RUN source $HOME/.local/bin/env && \
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
# Install FastVideo Unified Kernel
# Install FastVideo Unified Kernel.
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
# of probing a live device for the arch (matches the released kernel wheel).
RUN source $HOME/.local/bin/env && \
source /opt/venv/bin/activate && \
cd fastvideo-kernel && \
git submodule update --init --recursive && \
./build.sh
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
EXPOSE 22
+4 -2
View File
@@ -55,11 +55,13 @@ RUN source $HOME/.local/bin/env && \
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
# Install FastVideo Unified Kernel
# Install FastVideo Unified Kernel.
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
# of probing a live device for the arch (matches the released kernel wheel).
RUN source $HOME/.local/bin/env && \
source /opt/venv/bin/activate && \
cd fastvideo-kernel && \
git submodule update --init --recursive && \
./build.sh
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
EXPOSE 22
+4 -2
View File
@@ -55,12 +55,14 @@ RUN source $HOME/.local/bin/env && \
echo 'source /opt/venv/bin/activate' >> /root/.bashrc && \
echo 'if [ -n "$ZSH_VERSION" ] && [ -f ~/.zshrc ]; then . ~/.zshrc; elif [ -f ~/.bashrc ]; then . ~/.bashrc; fi' > /root/.profile
# Install FastVideo Unified Kernel
# Install FastVideo Unified Kernel.
# This build machine has no GPU, so target Hopper (sm_90a) explicitly instead
# of probing a live device for the arch (matches the released kernel wheel).
RUN source $HOME/.local/bin/env && \
source /opt/venv/bin/activate && \
cd fastvideo-kernel && \
git submodule update --init --recursive && \
./build.sh
TORCH_CUDA_ARCH_LIST=9.0a ./build.sh
EXPOSE 22
+205
View File
@@ -0,0 +1,205 @@
// Add a one-click copy button for the rendered MkDocs article.
(function () {
const BUTTON_ID = "copy-page-button";
function text(value) {
return (value || "").replace(/\s+/g, " ").trim();
}
function codeLanguage(code) {
const classes = Array.from(code.classList || []);
const language = classes.find((name) => name.startsWith("language-"));
return language ? language.replace("language-", "") : "";
}
function serializeInline(node) {
if (node.nodeType === Node.TEXT_NODE) {
return node.textContent || "";
}
if (node.nodeType !== Node.ELEMENT_NODE) {
return "";
}
const tagName = node.tagName.toLowerCase();
if (tagName === "code" && node.parentElement && node.parentElement.tagName.toLowerCase() !== "pre") {
return "`" + (node.textContent || "").trim() + "`";
}
if (tagName === "a") {
if (node.classList.contains("headerlink")) {
return "";
}
const label = text(Array.from(node.childNodes).map(serializeInline).join(""));
const href = node.href;
return href && label ? `${label} (${href})` : label;
}
if (tagName === "img") {
const alt = node.getAttribute("alt") || "image";
const src = node.src || "";
return src ? `[${alt}](${src})` : `[${alt}]`;
}
if (tagName === "br") {
return "\n";
}
return Array.from(node.childNodes).map(serializeInline).join("");
}
function serializeTable(table) {
const rows = Array.from(table.rows);
if (rows.length === 0) return "";
const mdRows = rows.map(
(row) => "| " + Array.from(row.children).map((cell) => text(serializeInline(cell))).join(" | ") + " |"
);
const separator = "| " + Array.from(rows[0].children)
.map(() => "---")
.join(" | ") + " |";
mdRows.splice(1, 0, separator);
return mdRows.join("\n");
}
function serializeBlock(node, listDepth = 0) {
if (node.nodeType === Node.TEXT_NODE) {
return text(node.textContent);
}
if (node.nodeType !== Node.ELEMENT_NODE) {
return "";
}
const tagName = node.tagName.toLowerCase();
if (["script", "style", "nav", "button"].includes(tagName) || node.id === BUTTON_ID) {
return "";
}
if (/^h[1-6]$/.test(tagName)) {
const level = Number(tagName.slice(1));
return `${"#".repeat(level)} ${text(serializeInline(node))}`;
}
if (tagName === "pre") {
const code = node.querySelector("code");
const content = code ? code.textContent || "" : node.textContent || "";
return `\`\`\`${code ? codeLanguage(code) : ""}\n${content.replace(/\n$/, "")}\n\`\`\``;
}
if (["p", "figcaption"].includes(tagName)) {
return text(serializeInline(node));
}
if (tagName === "blockquote") {
return serializeChildren(node, listDepth)
.split("\n")
.map((line) => (line ? `> ${line}` : ">"))
.join("\n");
}
if (tagName === "ul" || tagName === "ol") {
return Array.from(node.children)
.filter((child) => child.tagName && child.tagName.toLowerCase() === "li")
.map((item, index) => serializeListItem(item, tagName === "ol", index, listDepth))
.join("\n");
}
if (tagName === "table") {
return serializeTable(node);
}
if (["hr"].includes(tagName)) {
return "---";
}
return serializeChildren(node, listDepth);
}
function serializeListItem(item, ordered, index, listDepth) {
const marker = ordered ? `${index + 1}. ` : "- ";
const indent = " ".repeat(listDepth);
const childBlocks = [];
const inlineParts = [];
Array.from(item.childNodes).forEach((child) => {
if (child.nodeType === Node.ELEMENT_NODE && ["ul", "ol"].includes(child.tagName.toLowerCase())) {
childBlocks.push(serializeBlock(child, listDepth + 1));
} else {
const content = serializeInline(child);
if (content) inlineParts.push(content);
}
});
const firstLine = `${indent}${marker}${text(inlineParts.join(" "))}`.trimEnd();
return [firstLine, ...childBlocks.filter(Boolean)].join("\n");
}
function serializeChildren(node, listDepth = 0) {
return Array.from(node.childNodes)
.map((child) => serializeBlock(child, listDepth))
.map((value) => value.trim())
.filter(Boolean)
.join("\n\n");
}
function articleText(article) {
const clone = article.cloneNode(true);
clone.querySelectorAll("script, style, .headerlink, .md-clipboard, #copy-page-button").forEach((node) => node.remove());
const content = serializeChildren(clone).trim();
const title = document.querySelector("h1") || document.querySelector("title");
const pageTitle = title ? text(title.textContent) : "";
if (pageTitle && !content.startsWith("# ")) {
return `# ${pageTitle}\n\n${content}`.trim();
}
return content;
}
async function copyArticle(button, article) {
const originalLabel = button.textContent;
try {
await navigator.clipboard.writeText(articleText(article));
button.textContent = "Copied!";
button.classList.add("copy-page-button--copied");
} catch (error) {
button.textContent = "Copy failed";
button.classList.add("copy-page-button--error");
console.error("Failed to copy page", error);
}
window.setTimeout(() => {
button.textContent = originalLabel;
button.classList.remove("copy-page-button--copied", "copy-page-button--error");
}, 2000);
}
function addCopyButton() {
const article = document.querySelector("article.md-content__inner");
if (!article || document.getElementById(BUTTON_ID)) {
return;
}
const button = document.createElement("button");
button.id = BUTTON_ID;
button.type = "button";
button.className = "copy-page-button md-button md-button--primary";
button.textContent = "Copy page";
button.setAttribute("aria-label", "Copy this page as plain text");
button.addEventListener("click", () => copyArticle(button, article));
article.insertBefore(button, article.firstChild);
}
if (typeof document$ !== "undefined") {
document$.subscribe(addCopyButton);
}
document.addEventListener("DOMContentLoaded", addCopyButton);
window.addEventListener("load", addCopyButton);
})();
+17
View File
@@ -41,3 +41,20 @@ img {
display: block;
margin: 0 auto;
}
.md-typeset .copy-page-button.md-button {
float: right;
margin: 0 0 1rem 1rem;
padding: 0.2em 0.6em;
font-size: 1em;
font-weight: 400;
line-height: 1.2;
}
.md-typeset .copy-page-button.md-button.copy-page-button--copied {
background-color: var(--md-accent-fg-color);
}
.md-typeset .copy-page-button.md-button.copy-page-button--error {
background-color: var(--md-code-hl-number-color);
}
+118
View File
@@ -85,6 +85,124 @@ Each line in `/tmp/fv_trace.jsonl` is a JSON record:
5. The first divergent line identifies the first layer where FastVideo and the
upstream produce different outputs. Start debugging there.
## Performance impact
When the master toggle is off, overhead is nil. When on, the cost is
proportional to how broadly the layer regex matches.
### What runs when tracing is on
For every match against `model.named_modules()`, FastVideo registers an
`ActivationStatHook` that runs after the module's forward returns:
1. Walks the (possibly nested) output `tuple` / `list` / `dict` and extracts
every `torch.Tensor` leaf.
2. Computes each enabled stat on the tensor's `.detach().float()` view — the
cast is required for bf16 inputs because some reductions are not stable
in bf16.
3. Writes one JSON record per tensor to the line-buffered JSONL sink.
Cost is roughly `O(num_matched_modules × num_output_tensors × num_stats × tensor_numel)`
per forward, dominated by `tensor_numel` for value-stats (`abs_mean`, `sum`,
`min`, `max`, `mean`, `std`). The `shape` and `dtype` stats are O(1).
### Cost-shaping knobs
The default config (empty `FASTVIDEO_TRACE_LAYERS` is treated as `.*`, default
two stats, all steps) is intentionally blunt — useful only for a one-shot
smoke run. For real debugging, scope down:
| Knob | Effect |
|---|---|
| Tighten `FASTVIDEO_TRACE_LAYERS` to a regex matching <50 modules | Linear reduction in hook count |
| Drop unused stats from `FASTVIDEO_TRACE_STATS` (`mean`, `std`, `min`, `max`) if you only need divergence detection | One full reduction per stat saved per matched module per forward |
| Set `FASTVIDEO_TRACE_STEPS="0,15,31"` for a 32-step run | ~10x reduction vs all-steps (the hook still fires but exits early when the step doesn't match) |
| Use `shape` + `dtype` only on layers where you only care about layout | Skips tensor reductions entirely on those layers |
### Disk
Output is line-buffered (`open(..., buffering=1)`), so every record flushes
on write. A typical 32-step run that traces 40 DiT blocks with 4 stats
writes roughly 5K records — about 1 MB of JSONL. Point
`FASTVIDEO_TRACE_OUTPUT` at a fast local disk for parity runs; slow network
mounts will dominate runtime once tracing is on.
## Troubleshooting
### "I set the env var but no JSONL file appears"
Three things to check, in order:
1. **Toggle semantics.** `FASTVIDEO_TRACE_ACTIVATIONS` uses a strict
not-equal-to-`"0"` test. Setting it to `1`, `true`, or even `""` all
enable tracing. Only an unset variable or `FASTVIDEO_TRACE_ACTIVATIONS=0`
disables it.
2. **Module exposure.** The hook attaches to `pipeline.modules.get("transformer")`
at the end of `post_init`. If your pipeline does not expose a module
under that key (e.g. a non-standard custom pipeline whose DiT is
reachable only via `pipeline.modules["sr_transformer"]`), the trace is
silently a no-op. Check the
`Activation trace attached to N modules` log line at startup; if
`N=0`, either the regex didn't match anything or the expected module
isn't exposed.
3. **Output path failure.** The parent of `FASTVIDEO_TRACE_OUTPUT` is
auto-created. If creation fails (permissions, read-only mount),
`JsonlSink.__init__` raises at startup — look for an `OSError` early
in the log.
### "My regex isn't filtering the way I expect"
`FASTVIDEO_TRACE_LAYERS` is compiled with `re.compile(spec)` and matched
with `pattern.search(name)`. Two consequences:
- `search`, not `fullmatch`. `block.layers` matches
`transformer.block.layers.0.attn`. Anchor with `^...$` if you want
exact matches.
- Module names use Python dot notation (`transformer.block.layers.0`),
not slashes. The `.` in your regex is a metacharacter — escape it as
`\.` if you want a literal dot.
Run once with `FASTVIDEO_LOGGING_LEVEL=DEBUG` to see the
`Activation trace attached to N modules (pattern=...)` line. `N` is the
ground truth for how many modules survived your regex.
### "Stats are NaN or `<error: ...>` for some layers"
Some outputs hit numerical edge cases:
- `mean` / `std` on a 0-dim tensor returns NaN.
- `abs_mean` on an empty tensor or one full of `inf` returns NaN.
- Non-tensor outputs (e.g. a Python `bool` from a verification gate) are
silently skipped — the hook only walks `torch.Tensor` leaves.
Dropping `mean`/`std` and using `abs_mean`/`max` is more robust. For
layout-only debugging, `shape`+`dtype` never fail.
### "Tracing slows the run by 5x"
You're probably matching too broadly. An empty `FASTVIDEO_TRACE_LAYERS` is
treated as `.*` and matches every named module — for a 15B-param DiT
that's hundreds of submodules, each running stat reductions on every
forward. Tighten to a single block depth: `^transformer\.blocks\.\d+$`
typically matches a few dozen modules, which is a manageable trace.
### "Trace records are missing tensors I expect"
The hook walks `tuple` / `list` / `dict` outputs recursively but does
**not** unpack custom dataclasses or named tuples — those are silently
skipped. If your module returns
`BlockOutput(hidden_states=..., attn_logits=...)`, no records are
emitted. Workaround: either return a plain `dict` (`{"hidden_states": ..., "attn_logits": ...}`)
or attach the hook to a deeper module that already returns a raw tensor.
### "I want tracing inside a `torch.compile`'d region"
Module forward hooks run on the eager wrapper. If the entire module is
compiled, the hook sees only the wrapped op's output, not internal FX
nodes. This is by design — Extension 1 (FX backend rewrite) in
[Future extensions](#future-extensions-design-only-not-yet-implemented)
is the planned path when inside-graph granularity is required.
## Architecture (Extension 0: module forward hooks)
At pipeline initialization, `attach_activation_trace()` reads the env vars once.
+3 -1
View File
@@ -73,6 +73,7 @@ replicate CI results before pushing.
| Transformer Tests | `transformer` | `fastvideo/models/dits/**`, `fastvideo/models/loader/**`, `fastvideo/tests/transformers/**`, `fastvideo/layers/**`, `fastvideo/attention/**`, `pyproject.toml`, `docker/Dockerfile.python3.12` |
| Kernel Tests | `kernel_tests` | `fastvideo-kernel/**`, `pyproject.toml`, `docker/Dockerfile.python3.12` |
| Unit Tests | `unit_test` | `fastvideo/**`, `.buildkite/**`, `.github/**`, `pyproject.toml`, `docker/Dockerfile.python3.12` |
| DreamVerse App Tests | `dreamverse_app` | `apps/dreamverse/**`, `pyproject.toml` |
A Fastcheck failure means a component-level regression. Check the Buildkite build log for the
failing test's output.
@@ -99,7 +100,7 @@ failing test's output.
| LoRA Training Tests | `training_lora` | 15 min |
| Training Tests VSA | `training_vsa` | 15 min |
| Inference Tests VMoBA | `inference_vmoba` | 15 min |
| Performance Tests | `performance` | 30 min |
| [Performance Tests](performance_benchmarks.md) | `performance` | 30 min |
| API Server Tests | `api_server` | 30 min |
| Train Framework Tests | `train_framework` | 30 min |
@@ -268,6 +269,7 @@ Triggers a specific Buildkite test or suite on the current PR branch.
| `/test transformer` | Transformer Tests (Fastcheck) | `transformer` |
| `/test kernel` | Kernel Tests (Fastcheck) | `kernel_tests` |
| `/test unit` | Unit Tests (Fastcheck) | `unit_test` |
| `/test dreamverse` | DreamVerse App Tests (Fastcheck) | `dreamverse_app` |
| `/test ssim` | SSIM regression tests | `ssim` |
| `/test training` | Training pipeline tests | `training` |
| `/test lora-inference` | LoRA inference tests | `inference_lora` |
+100
View File
@@ -0,0 +1,100 @@
# Dreamverse Development
Dreamverse lives under `apps/dreamverse/` as a product app inside the
FastVideo monorepo. Backend code uses the local FastVideo workspace package;
frontend tooling remains standalone under `apps/dreamverse/web/`.
## Backend tests
Run CPU-safe backend tests from the FastVideo repository root:
```bash
uv run --locked --package dreamverse --extra test pytest apps/dreamverse/server/tests/ -m 'not gpu' -q
```
## Backend launch
Launch the migrated backend through the installed console commands:
```bash
dreamverse-server --port 8009
dreamverse-mock-server --port 8009
```
If `dreamverse-server` is missing, install FastVideo with the `dreamverse`
extra from the checkout:
```bash
uv pip install -e ".[dreamverse]"
```
## Frontend build and tests
Run frontend commands from the standalone web app:
```bash
cd apps/dreamverse/web
npm ci
npm run build
npm test
```
Playwright is intentionally run against a live backend as part of the GPU4
manual verification flow, not in the Phase 3 migration gate.
## Local GPU4 verification hook
Use physical GPU 4 for migration smoke tests. `CUDA_VISIBLE_DEVICES=4` makes
that GPU appear as logical GPU 0 inside the process, preserving the previous
Dreamverse deployment behavior.
```bash
CUDA_VISIBLE_DEVICES=4 dreamverse-server --host 0.0.0.0 --port 8009
```
In another shell, verify the service:
```bash
curl -s http://localhost:8009/healthz
```
Phase 4 adds the public `/healthz`, `/readyz`, `/status`,
`/prompt-system-config`, and `/curated-presets` route coverage needed for the
full Playwright suite.
## Phase 0 production-equivalent prerequisites
For the production-equivalent NVFP4 path, install these dependencies
in the FastVideo `.venv` before GPU smoke tests:
```bash
uv pip install --python .venv/bin/python \
flashinfer-python flash-attn cerebras-cloud-sdk openai \
--no-build-isolation
```
| Package | Why |
|---|---|
| `flashinfer-python` | Required for NVFP4 quantization. Without it, model load fails with `ImportError: NVFP4 quantization requires flashinfer`. |
| `flash-attn` | Optional but recommended; without it attention falls back to Torch SDPA (functional but slower). |
| `cerebras-cloud-sdk` | Required by the migrated prompt enhancer for the default `cerebras` provider. |
| `openai` | Required by the prompt enhancer's OpenAI-compatible providers + downstream rewrites. |
### B200 / sm_100a + gcc-15 conda toolchain (flashinfer JIT workaround)
On hosts where the conda toolchain ships gcc-15 (which nvcc rejects with
`#error -- unsupported GNU version! gcc versions later than 14 are not
supported!`), set these env vars before launching anything that triggers
flashinfer's JIT kernel build:
```bash
export CC=/usr/bin/gcc-13
export CXX=/usr/bin/g++-13
export CUDAHOSTCXX=/usr/bin/g++-13
export NVCC_PREPEND_FLAGS="-ccbin /usr/bin/gcc-13 -allow-unsupported-compiler"
```
`dreamverse-server` does NOT set these — they need to come from the launching
shell. The `dreamverse-deploy` skill
([`.agents/skills/dreamverse-deploy/`](../../.agents/skills/dreamverse-deploy/SKILL.md))
sets them for you and is the recommended local-deploy path.
+17 -13
View File
@@ -11,14 +11,14 @@ Use this guide when you are:
- Adding a new metric (native or wrapping a third-party library).
- Porting a benchmark (e.g. VBench, MIND, EvalCrafter) whose Python
code needs to be importable from a pinned upstream.
- Adding a new metric group (audio, vlm, etc.).
- Adding a new metric group (audio, videoscore2, etc.).
## TL;DR
Metrics are auto-discovered from
`fastvideo/eval/metrics/<group>/<name>/metric.py`. Each declares itself
with `@register("<group>.<name>")` and subclasses `BaseMetric`. Three
recipes:
with `@register("<group>.<name>")` and subclasses `BaseMetric`. Five
recipes, depending on how the metric ships and what its licence allows:
1. **Native metric** (pure-PyTorch, no submodule). Drop a file,
declare deps, implement `compute(sample)`.
@@ -26,11 +26,18 @@ recipes:
Same as above, plus route the library's cache through
`get_cache_dir()` if it has a `download_root=` / `cache_dir=`
kwarg.
3. **Upstream-submodule-wrapped metric** (vbench-style). Pin upstream
as a git submodule under `fastvideo/third_party/eval/<bench>/`. The
adapter `__init__.py` does the `sys.path` insert and any runtime
compat shims for modern dep versions. Patches live as Python in
that file rather than as on-disk patches to the submodule.
3. **Submodule-wrapped metric** (vbench-style). Pin upstream as a git
submodule under `fastvideo/third_party/eval/<bench>/`. The adapter
`__init__.py` does the `sys.path` insert and any runtime compat
shims. Best for large research packages with stable layouts.
4. **Vendored upstream** (synchformer / glmasr-style). Copy a small,
surgical piece of upstream into `fastvideo/third_party/eval/<name>/`
with its `LICENSE` alongside. Best for permissive-licensed
(MIT / Apache-2.0) source you need a few files from.
5. **Git-source dep via `[tool.uv.sources]`** (ImageBind-style). For
license-restricted upstream (e.g. CC BY-NC-SA) that cannot be
redistributed in the FastVideo tree. uv pulls the source at install
time pinned to a SHA in `pyproject.toml`.
The full recipes are below.
@@ -43,7 +50,8 @@ fastvideo/eval/metrics/
├── base.py # BaseMetric + lifecycle contract
├── common/ # group: SSIM, PSNR, LPIPS
├── optical_flow/ # group: gt_optical_flow, synthetic_optical_flow
├── vlm/ # group: VideoScore-2
├── audio/ # group: CLAP, AudioBox, KL, FAD, WER, DeSync, ImageBind
├── videoscore2/ # VideoScore-2 (single metric at group level)
├── physics_iq/ # group + sub-metrics
└── vbench/ # group: 16 sub-metrics
├── __init__.py # sys.path bootstrap + runtime compat shims
@@ -546,10 +554,6 @@ scores ± tolerance, and add a calibration test under
## 10) When not to add a metric
- **Set-vs-set distribution metrics** (FVD, FID-style) do not fit
`BaseMetric.compute(sample)` cleanly; they need a population.
Adding them requires a stateful accumulator interface that does
not exist yet. Open an issue first.
- **Metrics requiring a single-GPU model larger than available
memory.** Eval is not the place for tensor-parallel sharding;
metrics are expected to fit on one GPU.
+319
View File
@@ -0,0 +1,319 @@
# Performance Benchmarks
FastVideo's performance benchmark suite measures end-to-end inference latency,
throughput, peak GPU memory, and component-level pipeline timings for
representative pipeline configurations. It tracks those metrics over time
against a rolling baseline stored on the Hugging Face Hub.
It serves three audiences:
* **CI** — gates pull requests against a per-GPU static threshold and a
rolling-median regression check.
* **Maintainers** — surfaces regressions in a Markdown summary on every
performance build and a long-form Plotly dashboard.
* **Local developers** — lets you run the same benchmark on your own machine,
then compare against the historical baseline for the same model and GPU.
## Quick start (local)
```bash
# Run all benchmarks; writes raw perf_*.json under
# fastvideo/tests/performance/results/
pytest fastvideo/tests/performance/ -vs
# Optional: compare against the rolling HF baseline (read-only outside CI).
# PERF_REPORTS_DIR defaults to /root/data/perf_reports for Modal/CI, so
# override it when running outside the container.
PERF_REPORTS_DIR=/tmp/fastvideo_perf_reports \
python fastvideo/tests/performance/compare_baseline.py
# Optional: build the Plotly dashboard locally.
PERF_REPORTS_DIR=/tmp/fastvideo_perf_reports \
python fastvideo/tests/performance/dashboard.py
```
The pytest run never uploads anything. `compare_baseline.py` only writes to
the HF dataset when `TEST_SCOPE=full` *and* `BUILDKITE_BRANCH=main`, so local
runs are always read-only. The report directory default is container-oriented;
set `PERF_REPORTS_DIR` to a writable local path when generating dashboards or
when you want local Markdown/normalized-result artifacts from the comparator.
`compare_baseline.py` reads every `perf_*.json` currently present in
`fastvideo/tests/performance/results/`; remove stale result files if you only
want to compare the latest local run.
## Local live dashboard
For an app-style local dashboard backed by the same HF performance-tracking
records, see `performance_dashboard/README.md`. The dashboard provides a
FastAPI API plus a React UI and can be exposed with `ngrok` after building the
frontend.
## Architecture
```
.buildkite/performance-benchmarks/tests/*.json
└── per-benchmark configs: model, gen kwargs, per-GPU thresholds
fastvideo/tests/performance/
├── test_inference_performance.py
│ └── pytest test that runs each config, writes perf_*.json
│ with latency, memory, throughput, and component timings
├── compare_baseline.py
│ └── normalizes raw results, compares against HF rolling baseline,
│ writes Markdown summary + (optionally) uploads new records
├── dashboard.py
│ └── builds time-series Plotly HTML from HF history
└── hf_store.py # shared HF I/O + DataFrame helpers
```
The HF dataset (`FastVideo/performance-tracking` by default) holds one
normalized JSON per `(model_id, gpu_type, run)` tuple. The rolling baseline is
the median of the last 5 successful records for that model+GPU.
## Planned Coverage
The current rollout tracks a small set of representative inference workloads.
Broader coverage is planned for additional models, GPU types, attention
backends, workload shapes, and inference recipes. As that coverage lands, the
performance tracking system will also add environment-specific considerations
so comparisons remain meaningful across hardware, runtime, attention backend,
and recipe changes instead of treating all records for a model as equivalent.
## Metrics
Each benchmark records six metrics:
| Metric | Raw key | Normalized key | Direction |
|---|---|---|---|
| End-to-end generation latency | `avg_generation_time_s` | `latency` | Lower is better |
| Video throughput | `throughput_fps` | `throughput` | Higher is better |
| Peak GPU memory | `max_peak_memory_mb` | `memory` | Lower is better |
| Text encoder time | `text_encoder_time_s` | `text_encoder_time_s` | Lower is better |
| DiT denoising time | `dit_time_s` | `dit_time_s` | Lower is better |
| VAE decode time | `vae_decode_time_s` | `vae_decode_time_s` | Lower is better |
`test_inference_performance.py` temporarily sets `FASTVIDEO_STAGE_LOGGING=1`
while it runs so pipeline stage execution times are available in
`generate_video(...).logging_info`. It maps `TextEncodingStage` to
`text_encoder_time_s`, `DenoisingStage` and `DmdDenoisingStage` to
`dit_time_s`, and `DecodingStage` to `vae_decode_time_s`. If a pipeline does
not report one of those stages, that component metric is stored as `null` and
is skipped by the static threshold and rolling baseline checks.
## The two gates
There are **two independent regression gates** — they protect against
different failure modes and are not redundant.
### Static thresholds (per-GPU)
Defined in `.buildkite/performance-benchmarks/tests/<benchmark>.json` under
`thresholds`. Example:
```json
"thresholds": {
"L40S": {
"max_generation_time_s": 34.0,
"max_peak_memory_mb": 11000.0,
"max_text_encoder_time_s": 5.0,
"max_dit_time_s": 10.0,
"max_vae_decode_time_s": 10.0
},
"default": { "max_generation_time_s": 120.0, "max_peak_memory_mb": 30000.0 }
}
```
Selection: `_get_thresholds(cfg)` matches the current GPU name (substring
match) against keys; falls back to `default` if no GPU matches.
`max_generation_time_s` and `max_peak_memory_mb` are required for every
selected threshold block. Component limits are optional: if
`max_text_encoder_time_s`, `max_dit_time_s`, or `max_vae_decode_time_s` is
absent, the pytest static-threshold gate skips that component.
These are **fail-safes** — they catch order-of-magnitude regressions,
unrealistic memory growth, and optionally large component-specific slowdowns
even when the rolling baseline is empty. They are hand-set with generous
headroom and almost never need touching.
### Rolling baseline (per `(model_id, gpu_type)`)
`compare_baseline.py` loads the last 5 successful records for the same
`(model_id, gpu_type)` from the HF dataset, computes the median for each
available metric, and fails if the current run regresses by more than
`PERF_MAX_REGRESSION` (default 5%). For latency, memory, and component times,
higher values are regressions. For throughput, lower values are regressions.
This is the **drift detector** — it catches sub-threshold regressions that
slowly add up. It only persists new records when running the full suite on
`main`. Local and pull-request runs can compare against the HF baseline, but
they do not update it.
When the baseline shifts for a legitimate reason (torch upgrade, kernel
change, etc.) and CI starts failing, use the
[`reseed-performance-baseline`](https://github.com/hao-ai-lab/FastVideo/blob/main/.agents/skills/reseed-performance-baseline/SKILL.md)
agent skill to advance the rolling median.
## Schemas
### Raw record (`results/perf_*.json`)
Written by `test_inference_performance.py`. One file per benchmark run.
```jsonc
{
"benchmark_id": "wan-t2v-1.3b-2gpu",
"model_short_name": "Wan2.1-T2V-1.3B-Diffusers",
"device": "NVIDIA L40S",
"num_gpus": 2,
"num_warmup_runs": 1,
"num_measurement_runs": 3,
"avg_generation_time_s": 28.4,
"individual_times_s": [28.5, 28.3, 28.4],
"throughput_fps": 1.58,
"max_peak_memory_mb": 10840.0,
"individual_peak_memories_mb": [10840.0, 10822.0, 10833.0],
"thresholds": {
"max_generation_time_s": 34.0,
"max_peak_memory_mb": 11000.0,
"max_text_encoder_time_s": 5.0,
"max_dit_time_s": 10.0,
"max_vae_decode_time_s": 10.0
},
"commit": "<full sha>",
"pr_number": "1234",
"timestamp": "2026-05-08T22:00:00+00:00",
"text_encoder_time_s": 2.141,
"dit_time_s": 8.437,
"vae_decode_time_s": 3.208
}
```
### Normalized record (HF dataset, also dumped as `normalized_perf_*.json`)
Written by `compare_baseline.py:_normalize_record`. One file per benchmark
result, used as the rolling-baseline source of truth.
```jsonc
{
"model_id": "wan-t2v-1.3b-2gpu",
"timestamp": "2026-05-08T22:00:00+00:00",
"commit_sha": "<full sha>",
"gpu_type": "NVIDIA L40S",
"latency": 28.4,
"throughput": 1.58,
"memory": 10840.0,
"text_encoder_time_s": 2.141,
"dit_time_s": 8.437,
"vae_decode_time_s": 3.208,
"success": true
}
```
### Compatibility with legacy records
Older records in the HF dataset may not have component timing fields. The
comparator ignores missing or `null` metrics when computing a median, and the
dashboard lists skipped plots for metric series that have no non-null values.
## Environment variable reference
| Variable | Default | Used by | Purpose |
|---|---|---|---|
| `PERF_MAX_REGRESSION` | `0.05` | `compare_baseline.py` | Per-metric regression fraction that fails the build. |
| `PERFORMANCE_TRACKING_ROOT` | `/tmp/perf-tracking` | `compare_baseline.py`, `dashboard.py` | Local directory the HF dataset is synced to. |
| `PERF_REPORTS_DIR` | `/root/data/perf_reports` | `compare_baseline.py`, `dashboard.py` | Where the Markdown summary and Plotly HTML get written for Buildkite to pick up. |
| `HF_REPO_ID` | `FastVideo/performance-tracking` | `hf_store.py` | HF dataset repo holding rolling-baseline records. |
| `HF_API_KEY` | unset | `hf_store.py` | Required for upload (main-branch full-suite only); reads work without it. |
| `TEST_SCOPE` | unset | `compare_baseline.py` | Set to `full` together with `BUILDKITE_BRANCH=main` to enable HF persistence. |
| `BUILDKITE_BRANCH`, `BUILDKITE_COMMIT`, `BUILDKITE_PULL_REQUEST` | unset | `compare_baseline.py`, `test_inference_performance.py` | CI metadata stamped into records. |
| `DASHBOARD_DAYS` | `30` | `dashboard.py` | Lookback window for the Plotly trend pages. |
| `PERFORMANCE_TRACKING_SYNC_REUSE_TTL_SECONDS` | `3600` | `hf_store.py` | Freshness window for reusing an existing HF sync when requested by dashboard consumers. |
| `FASTVIDEO_STAGE_LOGGING` | set by the pytest test | `test_inference_performance.py` | Enables pipeline stage timing capture for component metrics during benchmark runs. |
## CI integration
The performance step can run on demand with `/test performance` and as part of
the Full Suite (see [CI Architecture](ci_architecture.md)). The Modal entry
point is `fastvideo/tests/modal/pr_test.py:run_performance_tests` and the
Buildkite artifact upload is in
`.buildkite/scripts/pr_test.sh:upload_performance_artifacts`.
Each performance build runs pytest first. If that fixed-threshold phase fails,
`compare_baseline.py` is skipped, so Markdown summaries and normalized JSON
artifacts are not emitted. The dashboard still runs best-effort for
observability. When pytest passes, the rolling-baseline phase emits:
* **Markdown summary** — appended to `$GITHUB_STEP_SUMMARY` when that variable
is set, and written as `perf_<sha>_<ts>.md` for Buildkite upload. Contains a
per-benchmark row with current vs. baseline values for latency, throughput,
memory, text encoder time, DiT time, and VAE decode time.
* **Plotly dashboard** — `dashboard_<sha>_<ts>.html` showing time-series for
each metric grouped by `(model_id, gpu_type)`.
* **Normalized records** — `normalized_perf_*.json`, one per benchmark.
Useful as input to the
[`reseed-performance-baseline`](https://github.com/hao-ai-lab/FastVideo/blob/main/.agents/skills/reseed-performance-baseline/SKILL.md)
skill.
## Adding a new benchmark
1. Drop a new JSON config into
`.buildkite/performance-benchmarks/tests/<name>.json`. Required keys:
```json
{
"benchmark_id": "<unique-id>",
"model": { "model_path": "...", "model_short_name": "..." },
"init_kwargs": { "num_gpus": 1, ... },
"generation_kwargs": { "num_frames": 45, ... },
"test_prompts": ["..."],
"run_config": { "required_gpus": 1,
"num_warmup_runs": 1, "num_measurement_runs": 3 },
"thresholds": {
"L40S": {
"max_generation_time_s": 34.0,
"max_peak_memory_mb": 11000.0,
"max_text_encoder_time_s": 5.0,
"max_dit_time_s": 10.0,
"max_vae_decode_time_s": 10.0
},
"default": { "max_generation_time_s": 120.0, "max_peak_memory_mb": 30000.0 }
}
}
```
2. The pytest test auto-discovers all configs — no test code needed. CI
picks it up on the next `/test performance` run.
3. The first persisted main-branch run with no HF history initializes the
baseline (passes automatically). Subsequent runs compare against it. Local
and pull-request runs with no HF history also pass, but they do not seed the
shared baseline.
4. If the benchmark targets a GPU not currently in `thresholds`, either add
that GPU as a key or rely on the `default` block. Note that `default` is
intended for slower fallback GPUs, so its values should be relaxed
relative to the fastest entry.
5. Add component thresholds only when the stage timing is stable enough to be
a useful fixed gate. The rolling baseline will still track component times
when static component thresholds are omitted.
## Troubleshooting
**"No baseline for ... Initializing"** — first run for this `(model_id,
gpu_type)`. Run will pass and (if persisting) seed the first record.
**Persistent failure right after a torch / kernel / image upgrade** —
genuine regression *or* baseline drift. Compare the failing normalized record
with recent successful records in the HF dataset. If the shift is expected and
reviewed, use the `reseed-performance-baseline` skill.
**Dashboard reports skipped metric plots** — the loaded records do not have
non-null values for that metric. This is expected for older records or for
pipelines that did not report a mapped component stage.
**Component timing is `null`** — the generated result did not include a mapped
stage in `logging_info.stages`. Check that the pipeline emits stage logging
and that the stage name is listed in `STAGE_METRIC_MAP` in
`test_inference_performance.py`.
+8
View File
@@ -51,3 +51,11 @@ Traces can be visualized using <https://ui.perfetto.dev/>.
- Keep the profiled step count small; traces can be large and slow down job shutdown while the profiler flushes data.
- After profiling, clean up trace directories to avoid filling disk storage.
- When adding new regions, register them in `fastvideo.profiler` and wrap the corresponding code block with `with self.profiler_controller.region("your_region"):` or the `@profile_region` decorator.
## Related: Activation Trace Mode
For per-layer **numerical-divergence** debugging (parity bring-up against an
upstream reference), use the env-gated [Activation Trace Mode](activation_trace.md)
instead of the torch profiler. Activation trace dumps per-tensor stats to
JSONL for offline `diff`-ing; the torch profiler captures kernel timing.
They solve different problems and can run together if needed.
+4 -1
View File
@@ -6,10 +6,12 @@ This guide explains how to add and run tests in FastVideo. The testing suite is
* **Unit Tests**: Located in `fastvideo/tests/api`, `fastvideo/tests/dataset`, `fastvideo/tests/entrypoints`, `fastvideo/tests/workflow`, and the CPU-only subset of `fastvideo/tests/train` (callbacks, utils). These test individual functions and classes.
* **Component Tests**: Located in `fastvideo/tests/encoders`, `fastvideo/tests/transformers`, and `fastvideo/tests/vaes`. These verify the loading and basic functionality of model components.
* **Train Framework Tests** (GPU): Located in `fastvideo/tests/train/models`. Cover model loading + forward smoke for the new `fastvideo/train/` framework. Triggered via `/test train-framework` or as part of the Full Suite.
* **Train Framework Tests** (GPU): Located in `fastvideo/tests/train/models` (model loading + forward smoke) and `fastvideo/tests/train/methods` (per-method single training step). Cover the new `fastvideo/train/` framework end-to-end on real checkpoints with tiny synthetic batches. Triggered via `/test train-framework` or as part of the Full Suite.
* **SSIM Tests**: Located in `fastvideo/tests/ssim`. These are regression tests that compare generated videos against reference videos using the Structural Similarity Index Measure (SSIM) to detect quality degradation.
* **Training Tests**: Located in `fastvideo/tests/training`. These validate training loops, loss calculations, and specific training techniques like LoRA, Distillation, and VSA.
* **Inference Tests**: Located in `fastvideo/tests/inference`. These test specialized inference pipelines and optimizations (e.g., VSA, V-MoBA).
* **Performance Tests**: Located in `fastvideo/tests/performance`. End-to-end latency, throughput, and peak-memory benchmarks gated against per-GPU static thresholds and a rolling Hugging Face baseline. See [Performance Benchmarks](performance_benchmarks.md) for the full workflow, schema and future environment-specific considerations.
* **DreamVerse App Tests**: Located in `apps/dreamverse`. These validate the DreamVerse app frontend and related app code.
For now, we will focus on **SSIM Tests**.
@@ -212,6 +214,7 @@ from a PR comment. The workflow reacts with a 🚀 emoji to confirm the command
/test transformer # Transformer / DiT tests (Fastcheck)
/test kernel # CUDA kernel tests (Fastcheck)
/test unit # Unit tests (Fastcheck)
/test dreamverse # DreamVerse app tests (Fastcheck)
/test full # Entire Full Suite
/test fastcheck # Entire Fastcheck suite
```
@@ -29,24 +29,46 @@ surfaces:
vae_cpu_offload: generator.engine.offload.vae
pin_cpu_memory: generator.engine.offload.pin_cpu_memory
enable_torch_compile: generator.engine.compile.enabled
enable_torch_compile_text_encoder: generator.engine.compile.text_encoder_enabled
enable_torch_compile_vae: generator.engine.compile.vae_enabled
enable_torch_compile_audio_vae: generator.engine.compile.audio_vae_enabled
torch_compile_kwargs: generator.engine.compile.backend,fullgraph,mode,dynamic,extras
torch_compile_kwargs_dit: generator.engine.compile.dit_kwargs
torch_compile_kwargs_text_encoder: generator.engine.compile.text_encoder_kwargs
torch_compile_kwargs_vae: generator.engine.compile.vae_kwargs
torch_compile_kwargs_audio_vae: generator.engine.compile.audio_vae_kwargs
transformer_quant: generator.engine.quantization.transformer_quant
disable_autocast: generator.engine.disable_autocast
enable_stage_verification: generator.engine.enable_stage_verification
prompt_txt: request.inputs.prompt_path
override_text_encoder_safetensors: generator.pipeline.components.text_encoder_weights
override_text_encoder_quant: generator.engine.quantization.text_encoder_quant
transformer_quant: generator.engine.quantization.transformer_quant
override_transformer_cls_name: generator.pipeline.components.override_transformer_cls_name
init_weights_from_safetensors: generator.pipeline.components.transformer_weights
init_weights_from_safetensors_2: generator.pipeline.components.transformer_2_weights
override_pipeline_cls_name: generator.pipeline.components.override_pipeline_cls_name
boundary_ratio: request.sampling.boundary_ratio
ltx2_vae_tiling: generator.pipeline.vae_tiling
refine_enabled: generator.pipeline.preset_overrides.refine.enabled
refine_upsampler_path: generator.pipeline.components.upsampler_weights
refine_lora_path: generator.pipeline.components.lora_path
refine_num_inference_steps: request.stage_overrides.refine.num_inference_steps
refine_guidance_scale: request.stage_overrides.refine.guidance_scale
refine_add_noise: generator.pipeline.preset_overrides.refine.add_noise
ltx2_refine_enabled: generator.pipeline.preset_overrides.refine.enabled
ltx2_refine_upsampler_path: generator.pipeline.components.upsampler_weights
ltx2_refine_lora_path: generator.pipeline.components.lora_path
ltx2_refine_num_inference_steps: request.stage_overrides.refine.num_inference_steps
ltx2_refine_guidance_scale: request.stage_overrides.refine.guidance_scale
ltx2_refine_add_noise: generator.pipeline.preset_overrides.refine.add_noise
preset_owned:
ltx2_vae_spatial_tile_size_in_pixels: generator.pipeline.preset_overrides.ltx2.vae.spatial_tile_size_in_pixels
ltx2_vae_spatial_tile_overlap_in_pixels: generator.pipeline.preset_overrides.ltx2.vae.spatial_tile_overlap_in_pixels
ltx2_vae_temporal_tile_size_in_frames: generator.pipeline.preset_overrides.ltx2.vae.temporal_tile_size_in_frames
ltx2_vae_temporal_tile_overlap_in_frames: generator.pipeline.preset_overrides.ltx2.vae.temporal_tile_overlap_in_frames
ltx2_initial_latent_path: request.extensions.ltx2.initial_latent_path
ltx2_audio_latent_path: request.extensions.ltx2.audio_latent_path
compatibility_only:
mode: "Legacy multi-mode FastVideoArgs switch; typed inference config should not expose execution mode."
inference_mode: "Legacy boolean mirror of mode; kept only through adapters while FastVideoArgs remains."
@@ -56,6 +78,14 @@ surfaces:
VSA_sparsity: "Model-specific inference optimization not yet represented in the typed public schema."
moba_config_path: "Model-specific MoBA optimization surface not yet represented in the typed public schema."
master_port: "Executor/bootstrap compatibility field; not part of the canonical inference schema."
refine_transformer_path: "Generic stage-2 refine transformer override; no typed equivalent yet."
refine_noise_path: "Generic stage-2 refine noise override; no typed equivalent yet."
refine_audio_noise_path: "Generic stage-2 refine audio noise override; no typed equivalent yet."
ltx2_refine_transformer_path: "LTX-2 refine transformer carrier; no typed equivalent yet."
ltx2_refine_noise_path: "LTX-2 refine noise carrier; no typed equivalent yet."
ltx2_refine_audio_noise_path: "LTX-2 refine audio noise carrier; no typed equivalent yet."
ltx2_legacy_native_noise_order: "LTX-2 SSIM compatibility knob preserving legacy native latent noise ordering."
ltx2_use_distilled_sigmas: "LTX-2 compatibility knob gating use of distilled sigma schedule."
private_only:
ray_placement_group: "Ray deployment-only field."
ray_runtime_env: "Ray deployment-only field."
@@ -78,6 +108,7 @@ surfaces:
vae_sp: generator.pipeline.preset_overrides.vae_sp
dmd_denoising_steps: generator.pipeline.preset_overrides.dmd_denoising_steps
ti2v_task: generator.pipeline.preset_overrides.ti2v_task
lucy_edit_task: generator.pipeline.preset_overrides.lucy_edit_task
boundary_ratio: generator.pipeline.preset_overrides.boundary_ratio
compatibility_only:
model_path: "Redundant with generator.model_path."
@@ -95,9 +126,18 @@ surfaces:
text_encoder_configs: "Legacy internal component config object."
preprocess_text_funcs: "Internal text preprocessing hooks."
postprocess_text_funcs: "Internal text postprocessing hooks."
scheduler_step_in_fp32: "Runtime scheduler precision toggle; not part of the public typed inference API."
pipeline_config_extensions:
preset_owned:
flux2_text_encoder_type:
sources:
- fastvideo.configs.pipelines.flux_2.Flux2PipelineConfig
- fastvideo.configs.pipelines.flux_2.Flux2KleinPipelineConfig
text_encoder_out_layers:
sources:
- fastvideo.configs.pipelines.flux_2.Flux2PipelineConfig
- fastvideo.configs.pipelines.flux_2.Flux2KleinPipelineConfig
conditioning_strategy:
sources:
- fastvideo.configs.pipelines.cosmos.CosmosConfig
@@ -231,8 +271,8 @@ surfaces:
- fastvideo.configs.pipelines.turbodiffusion.TurboDiffusionT2V_1_3B_Config
- fastvideo.configs.pipelines.wan.FastWan2_1_T2V_480P_Config
- fastvideo.configs.pipelines.wan.FastWan2_2_TI2V_5B_Config
- fastvideo.configs.pipelines.wan.MatrixGameBaseI2V480PConfig
- fastvideo.configs.pipelines.wan.MatrixGameI2V480PConfig
- fastvideo.configs.pipelines.matrixgame2.MatrixGame2BaseI2V480PConfig
- fastvideo.configs.pipelines.matrixgame2.MatrixGame2I2V480PConfig
- fastvideo.configs.pipelines.wan.SelfForcingWan2_2_T2V480PConfig
- fastvideo.configs.pipelines.wan.SelfForcingWanT2V480PConfig
- fastvideo.configs.pipelines.wan.WANV2VConfig
@@ -254,8 +294,8 @@ surfaces:
- fastvideo.configs.pipelines.turbodiffusion.TurboDiffusionT2V_1_3B_Config
- fastvideo.configs.pipelines.wan.FastWan2_1_T2V_480P_Config
- fastvideo.configs.pipelines.wan.FastWan2_2_TI2V_5B_Config
- fastvideo.configs.pipelines.wan.MatrixGameBaseI2V480PConfig
- fastvideo.configs.pipelines.wan.MatrixGameI2V480PConfig
- fastvideo.configs.pipelines.matrixgame2.MatrixGame2BaseI2V480PConfig
- fastvideo.configs.pipelines.matrixgame2.MatrixGame2I2V480PConfig
- fastvideo.configs.pipelines.wan.SelfForcingWan2_2_T2V480PConfig
- fastvideo.configs.pipelines.wan.SelfForcingWanT2V480PConfig
- fastvideo.configs.pipelines.wan.WANV2VConfig
@@ -303,9 +343,9 @@ surfaces:
- fastvideo.configs.pipelines.wan.FastWan2_2_TI2V_5B_Config
- fastvideo.configs.pipelines.wan.Wan2_2_TI2V_5B_Config
context_noise:
sources: [fastvideo.configs.pipelines.wan.MatrixGameI2V480PConfig]
sources: [fastvideo.configs.pipelines.matrixgame2.MatrixGame2I2V480PConfig]
num_frames_per_block:
sources: [fastvideo.configs.pipelines.wan.MatrixGameI2V480PConfig]
sources: [fastvideo.configs.pipelines.matrixgame2.MatrixGame2I2V480PConfig]
audio_channels:
sources:
- fastvideo.configs.pipelines.stable_audio.StableAudioT2AConfig
@@ -330,6 +370,40 @@ surfaces:
sources:
- fastvideo.configs.pipelines.stable_audio.StableAudioT2AConfig
- fastvideo.configs.pipelines.stable_audio.StableAudioOpenSmallConfig
audio_txt_guidance_scale:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanBaseConfig]
cfg_number:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanBaseConfig]
cfg_trick_start_frame:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
cfg_trick_value:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
noise_value:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
sr_audio_noise_scale:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
sr_height:
sources:
- fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig
- fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR1080pConfig
sr_num_inference_steps:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
sr_video_txt_guidance_scale:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
sr_width:
sources:
- fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig
- fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR1080pConfig
t5_gemma_target_length:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanBaseConfig]
use_cfg_trick:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanSR540pConfig]
video_guidance_high_t_threshold:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanBaseConfig]
video_guidance_low_t_value:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanBaseConfig]
video_txt_guidance_scale:
sources: [fastvideo.pipelines.basic.magi_human.pipeline_configs.MagiHumanBaseConfig]
compatibility_only:
batch_size: "Gen3C inference-only tuning field pending typed batching design."
gradient_checkpointing: "Gen3C inference-only compatibility field pending typed batching design."
@@ -338,6 +412,15 @@ surfaces:
internal_only:
audio_decoder_config: "Legacy internal component config object."
audio_decoder_precision: "Precision override pending dedicated component precision design."
audio_vae_config: "MagiHuman internal audio VAE component config object."
coords_style: "MagiHuman internal data-proxy coordinate convention."
frame_receptive_field: "MagiHuman internal data-proxy receptive-field setting."
image_conditioning: "MagiHuman preset variant marker for reference-image conditioning."
ref_audio_offset: "MagiHuman internal data-proxy audio alignment offset."
sr_local_attn_layers: "MagiHuman SR internal sparse-attention layer selection."
text_offset: "MagiHuman internal data-proxy text alignment offset."
vae_stride: "MagiHuman internal VAE/data-proxy stride setting."
z_dim: "MagiHuman internal VAE latent channel setting."
vocoder_config: "Legacy internal component config object."
vocoder_precision: "Precision override pending dedicated component precision design."
@@ -404,6 +487,11 @@ surfaces:
ltx2_stg_scale_audio: request.extensions.ltx2.stg_scale_audio
ltx2_stg_blocks_video: request.extensions.ltx2.stg_blocks_video
ltx2_stg_blocks_audio: request.extensions.ltx2.stg_blocks_audio
ltx2_images: request.extensions.ltx2.images
ltx2_image_crf: request.stage_overrides.refine.image_crf
ltx2_conditioning_latent_stage1: request.extensions.ltx2.conditioning_latent_stage1
ltx2_conditioning_latent_stage2: request.extensions.ltx2.conditioning_latent_stage2
ltx2_video_conditions: request.extensions.ltx2.video_conditions
audio_start_in_s: request.extensions.stable_audio.audio_start_in_s
audio_end_in_s: request.extensions.stable_audio.audio_end_in_s
init_audio: request.extensions.stable_audio.init_audio
@@ -413,6 +501,8 @@ surfaces:
inpaint_mask: request.extensions.stable_audio.inpaint_mask
internal_only:
data_type: "Derived from the request shape and not a public input."
latents: "Pre-generated diffusion latents supplied by parity/debug harnesses; not a public input."
max_sequence_length: "Model-specific text-encoder sequence cap; not part of the public typed inference API."
sampling_param_extensions: {}
+5
View File
@@ -168,6 +168,11 @@ How this maps to FastVideo:
- Attention backends live in `fastvideo/attention/` and can be selected via
`FASTVIDEO_ATTENTION_BACKEND`.
- SageAttention3 is split into two selectable backends:
`SAGE_ATTN_THREE` for the regular upstream package and
`ATTN_QAT_INFER` for the FastVideoKernel-backed inference variant.
- `ATTN_QAT_TRAIN` is a separate FastVideoKernel Triton backend for the QAT attention
path.
- `LocalAttention` is used for cross-attention and most attention layers.
- `DistributedAttention` is used for full-sequence self-attention in the DiT.
- Tensor-parallel layers live in `fastvideo/layers/`.
+354
View File
@@ -0,0 +1,354 @@
# Dynamo Native Backend Integration
FastVideo exposes a stable Python API that the
[ai-dynamo/dynamo](https://github.com/ai-dynamo/dynamo) project consumes
as a pure-Python import, same tier as `vllm`, `sglang`, `trtllm`.
**FastVideo hosts no Dynamo code.** The backend subpackage
(`components/src/dynamo/fastvideo/`) lives in the Dynamo repo. This doc
is the reference integrators copy when standing up that package — it
mirrors the structure used by `dynamo/components/src/dynamo/sglang/`
and is known to satisfy the (closed) draft
[ai-dynamo/dynamo#7544](https://github.com/ai-dynamo/dynamo/pull/7544)
pattern.
## What FastVideo provides
The public surface Dynamo imports is intentionally small:
```python
from fastvideo import VideoGenerator
from fastvideo.api import (
ContinuationState,
GenerationRequest,
InputConfig,
OutputConfig,
SamplingConfig,
# Post-PR 7.10:
VideoEvent, VideoProgressEvent, VideoPartialEvent, VideoFinalEvent,
VideoResult,
)
```
| Surface | Availability | Notes |
| --- | --- | --- |
| `VideoGenerator.from_pretrained(model_path, **typed_kwargs)` | Today | `typed_kwargs` is a stable subset from `GeneratorConfig` — no flat legacy LTX-2 kwargs (guaranteed after PR 6) |
| `VideoGenerator.generate(request: GenerationRequest) -> GenerationResult` | Today | Aggregated; Dynamo wraps in `asyncio.to_thread` under `asyncio.Lock` |
| `VideoGenerator.generate_async(request) -> AsyncGenerator[VideoEvent, None]` | **PR 7.10** | Canonical execution substrate; sync wrapper reroutes through this |
| `VideoGenerator.default_health_check_request() -> GenerationRequest` | **PR 7.10** | 256x256 / 8 frames / 1 step; lets Dynamo build its health payload without knowing any FastVideo internals |
| `fastvideo.api.GenerationRequest` / `SamplingConfig` / `InputConfig` | Today | Stable public dataclasses |
| `fastvideo.api.ContinuationState` | Today (PR 7) | JSON-safe envelope; kind-versioned payloads |
| `fastvideo.api.VideoResult` | Today | `frames`, `video_path`, `state`, `metadata` |
| `config_to_dict(cfg)` | Today | Used by Dynamo's `dump_config(path, config)` |
## Backend package layout
Modeled on `components/src/dynamo/sglang/`:
```
components/src/dynamo/fastvideo/
├── __init__.py
├── __main__.py # Entry: python -m dynamo.fastvideo
├── main.py # worker() dispatch — mirrors sglang/main.py
├── args.py # FastVideoArgGroup — CLI → GeneratorConfig
├── backend_args.py # Dynamo runtime flags (namespace, fs_url, ...)
├── init_video_generation.py # init_video_generation(runtime, config)
├── register.py # register_video_generation_model() for Dynamo
├── backend.py # VideoGenerationWorkerHandler
├── health_check.py # FastVideoHealthCheckPayload
├── protocol.py # NvCreateVideoRequest ↔ GenerationRequest adapter
├── request_handlers/
│ └── video_generation/
│ └── video_generation_handler.py # async generate(req, ctx)
├── README.md
└── CLAUDE.md # per-backend guidance
```
None of these files live in FastVideo.
## Request/response mapping
Dynamo's `NvCreateVideoRequest` / `VideoNvExt` / `NvVideosResponse` map
one-to-one onto FastVideo's typed schema:
```
NvCreateVideoRequest -> fastvideo.api.GenerationRequest
prompt -> request.prompt
size="WxH" -> request.sampling.width, height
seconds -> seconds * nvext.fps -> request.sampling.num_frames
input_reference -> request.inputs.image_path / video_path
nvext.fps -> request.sampling.fps
nvext.num_frames -> request.sampling.num_frames (overrides seconds*fps)
nvext.num_inference_steps -> request.sampling.num_inference_steps
nvext.guidance_scale -> request.sampling.guidance_scale
nvext.seed -> request.sampling.seed
nvext.negative_prompt -> request.sampling.negative_prompt
nvext.continuation_state -> request.state (opaque ContinuationState)
response_format -> (handled by adapter at output)
VideoFinalEvent -> NvVideosResponse
video_bytes -> data[0].b64_json (if response_format=b64_json)
uploaded URL -> data[0].url (if response_format=url)
metadata.inference_time_s -> inference_time_s
continuation_state -> nvext.continuation_state (reserved for disagg)
```
## Example: aggregated handler (sync wrap)
Satisfies the PR #7544 shape; works today against
`VideoGenerator.generate`, upgrades cleanly to `generate_async` after
PR 7.10.
```python
# components/src/dynamo/fastvideo/request_handlers/video_generation/
# video_generation_handler.py
from __future__ import annotations
import asyncio
import base64
import time
from typing import Any, AsyncGenerator
from fastvideo import VideoGenerator
from fastvideo.api import GenerationRequest, InputConfig, OutputConfig, SamplingConfig
class VideoGenerationWorkerHandler:
def __init__(self, generator: VideoGenerator, config, fs=None):
self.generator = generator
self.config = config
self.fs = fs
self._lock = asyncio.Lock() # aggregated = one-in-flight
async def generate(
self,
request: dict[str, Any],
context,
) -> AsyncGenerator[dict[str, Any], None]:
req = _to_fastvideo_request(request)
t0 = time.perf_counter()
async with self._lock:
result = await asyncio.to_thread(self.generator.generate, req)
elapsed = time.perf_counter() - t0
video_bytes = _materialize(result, self.fs, request.get("response_format"))
yield {
"data": [video_bytes],
"inference_time_s": elapsed,
"model": request.get("model"),
}
def _to_fastvideo_request(request: dict[str, Any]) -> GenerationRequest:
nvext = request.get("nvext") or {}
fps = nvext.get("fps", 24)
num_frames = nvext.get("num_frames") or (request.get("seconds") or 4) * fps
width, height = _parse_size(request.get("size"))
return GenerationRequest(
prompt=request["prompt"],
negative_prompt=nvext.get("negative_prompt"),
inputs=InputConfig(
image_path=request.get("input_reference"),
),
sampling=SamplingConfig(
width=width, height=height,
num_frames=num_frames, fps=fps,
num_inference_steps=nvext.get("num_inference_steps", 50),
guidance_scale=nvext.get("guidance_scale", 1.0),
seed=nvext.get("seed", 1024),
),
output=OutputConfig(save_video=False, return_frames=False),
state=nvext.get("continuation_state"), # public ContinuationState
)
```
`_parse_size` and `_materialize` are small adapter helpers owned by the
Dynamo backend package; they never appear in FastVideo.
## Example: streaming handler (post-PR 7.10)
```python
async def generate(self, request, context):
req = _to_fastvideo_request(request)
async for event in self.generator.generate_async(req):
if event.__class__.__name__ == "VideoProgressEvent":
yield {"status": "generating", "progress": event.step / event.total_steps}
elif event.__class__.__name__ == "VideoFinalEvent":
yield {
"data": [{"b64_json": base64.b64encode(event.video_bytes).decode()}],
"inference_time_s": event.metadata.get("inference_time_s"),
"nvext": {"continuation_state": _serialize_state(event.continuation_state)},
}
```
Aggregated and streaming differ only in which events the handler
forwards; both share one `generate_async` substrate.
## Example: health check
```python
# components/src/dynamo/fastvideo/health_check.py
from dynamo.health_check import HealthCheckPayload
from fastvideo import VideoGenerator
class FastVideoHealthCheckPayload(HealthCheckPayload):
def __init__(self, generator: VideoGenerator) -> None:
# Post-PR 7.10: generator.default_health_check_request() returns a
# typed GenerationRequest; dump it into the same dict shape that
# FastVideo's adapter accepts.
req = generator.default_health_check_request()
self.default_payload = {
"prompt": req.prompt or "test",
"size": f"{req.sampling.width}x{req.sampling.height}",
"response_format": "b64_json",
"nvext": {
"fps": req.sampling.fps,
"num_frames": req.sampling.num_frames,
"num_inference_steps": req.sampling.num_inference_steps,
"guidance_scale": req.sampling.guidance_scale,
},
}
super().__init__()
```
Fallback (pre-PR 7.10) — hardcoded 256×256 / 8 frames / 1 step, matching
[`VideoGenerationHealthCheckPayload`](https://github.com/ai-dynamo/dynamo/blob/main/components/src/dynamo/sglang/health_check.py#L198-L226).
## Init function sketch
```python
# components/src/dynamo/fastvideo/init_video_generation.py
async def init_video_generation(runtime, config, shutdown_endpoints):
from fastvideo import VideoGenerator
from fastvideo.api import config_to_dict
server_args, dynamo_args = config.server_args, config.dynamo_args
generator = VideoGenerator.from_pretrained(**config.fastvideo_kwargs())
dump_config(dynamo_args.dump_config_to, config)
endpoint = runtime.endpoint(
f"{dynamo_args.namespace}.{dynamo_args.component}.{dynamo_args.endpoint}"
)
shutdown_endpoints[:] = [endpoint]
handler = VideoGenerationWorkerHandler(
generator, config, fs=get_fs(dynamo_args.media_output_fs_url)
)
payload = FastVideoHealthCheckPayload(generator).to_dict()
await asyncio.gather(
endpoint.serve_endpoint(
handler.generate,
graceful_shutdown=True,
health_check_payload=payload,
),
register_video_generation_model(
generator, endpoint, server_args,
),
)
```
## Args adapter
`FastVideoArgGroup` (Dynamo-side) converts CLI flags into a typed
`GeneratorConfig` — **never** into legacy flat kwargs. Because PR 6
added typed homes for every kwarg the internal `gpu_pool.py` used,
this adapter can build the config purely from the public typed schema:
```python
def build_generator_config(args) -> "GeneratorConfig":
from fastvideo.api import (
CompileConfig, ComponentConfig, EngineConfig, GeneratorConfig,
OffloadConfig, ParallelismConfig, PipelineSelection,
)
return GeneratorConfig(
model_path=args.model_path,
engine=EngineConfig(
num_gpus=args.num_gpus,
parallelism=ParallelismConfig(tp_size=args.tp_size, sp_size=args.sp_size),
offload=OffloadConfig(dit=args.dit_offload, text_encoder=args.te_offload),
compile=CompileConfig(enabled=args.compile, mode=args.compile_mode),
),
pipeline=PipelineSelection(
workload_type=args.workload or "t2v",
preset=args.preset, # e.g. "ltx2_two_stage"
components=ComponentConfig(
upsampler_weights=args.refine_upsampler,
lora_path=args.refine_lora,
),
),
)
```
## Registration
Dynamo's Rust side skips HuggingFace `config.json` downloads for
`ModelType::Videos`, same fast path used by image diffusion. The
Python-side registration:
```python
# components/src/dynamo/fastvideo/register.py
from dynamo.llm import ModelDeploymentCard, ModelType, register_model
async def register_video_generation_model(generator, endpoint, server_args):
mdc = ModelDeploymentCard.with_name_only(server_args.model_name or server_args.model_path)
await register_model(endpoint, mdc, ModelType.Videos, readiness_gate=asyncio.Event())
```
## Contract guarantees
These guardrails let the Dynamo backend be written once and not
re-chase FastVideo drift:
1. `GenerationRequest` field paths are stable across PR 6 onward. Any
breaking rename triggers a major bump and appears in
[`inference_schema_parity_inventory.yaml`](../inference_schema_parity_inventory.yaml).
2. `ContinuationState.payload` is JSON-serializable or references
opaque blob ids. Dynamo can round-trip it through RPC without
special-casing torch tensors.
3. `VideoGenerator.from_pretrained` accepts a typed `GeneratorConfig`;
legacy flat kwargs are compatibility-only and deprecate in PR 13.
4. `generate_async` (PR 7.10+) emits events in order
`Progress* → Partial* → Final`; the final event always has exactly
one occurrence per request.
5. `default_health_check_request()` (PR 7.10+) returns a request that
passes `parse_config` and produces a non-zero-latency but bounded
workload (256×256 / 8 frames / 1 step).
FastVideo's contract tests (`fastvideo/tests/contract/`) assert these
with mocked Dynamo-style handlers that import only the public surface.
If a change to FastVideo breaks the adapter pattern, those tests fail
at FastVideo's CI — before the Dynamo-side integration even knows.
## What the Dynamo adapter MUST NOT import
* Anything under `fastvideo.pipelines.*` directly (pipelines are
internal; presets identify them by name on
`PipelineSelection.preset`).
* `fastvideo.fastvideo_args.FastVideoArgs` (legacy compat type).
* `fastvideo.api.compat.*` private helpers
(`_validate_continuation_state` etc.) — the public boundary is
`VideoGenerator` + `fastvideo.api`.
* Any flat legacy LTX-2 kwarg (`ltx2_refine_upsampler_path`,
`torch_compile_kwargs`, etc.) — all have typed homes in
`GeneratorConfig`.
## Future: disaggregated prefill/decode
PR 7's continuation state was designed to survive RPC transport, so a
future Dynamo split where prefill yields state and decode hydrates it
is expressible without changing the contract. The streaming server
(PR 7.6) already uses the `SessionStore` pattern; Dynamo's disagg
could wire a distributed `SessionStore` backend by the same interface.
## See also
* [OpenAI HTTP contract](openai.md)
* [Streaming WebSocket protocol](streaming.md)
* Draft PR reference: [ai-dynamo/dynamo#7544](https://github.com/ai-dynamo/dynamo/pull/7544)
* Dynamo SGLang backend (template this doc is modeled on):
[ai-dynamo/dynamo](https://github.com/ai-dynamo/dynamo/tree/main/components/src/dynamo/sglang)
+41
View File
@@ -0,0 +1,41 @@
# Server Contracts
FastVideo's typed public API (`fastvideo.api`) is consumed by three
server-class integrations that must share one execution substrate so we
don't grow three near-duplicate progress loops:
| Consumer | Transport | Request shape | State model |
| --- | --- | --- | --- |
| [Stateless OpenAI](openai.md) | HTTP POST `/v1/videos` | `VideoGenerationsRequest` → `GenerationRequest` merged onto `ServeConfig.default_request` | Stateless; optional `ContinuationState` round-trip |
| [Streaming WebSocket](streaming.md) | WebSocket JSON + binary fMP4 | `GenerationRequest` per segment | Server-held `SessionStore`, snapshot-on-demand |
| [Dynamo native backend](dynamo.md) | Dynamo RPC | `NvCreateVideoRequest` → adapter → `GenerationRequest` | Aggregated today; disaggregated via `ContinuationState` later |
All three consume the same underlying surface:
```python
from fastvideo import VideoGenerator
from fastvideo.api import (
ContinuationState,
GenerationRequest,
InputConfig,
OutputConfig,
SamplingConfig,
ServeConfig,
)
# Sync today, async after PR 7.10 lands VideoGenerator.generate_async.
result = generator.generate(request)
```
These docs lock down the request/response shapes so drift between
FastVideo, the internal UI, and Dynamo can be caught at review time.
PR 8 does not ship runtime code; it ships the contract reference and
the contract tests that guard it.
## Related
- [API refactor design](../overview.md)
- Parity inventory: [`inference_schema_parity_inventory.yaml`](../inference_schema_parity_inventory.yaml)
- [Streaming server upstream plan](../../../.agents/memory/dreamverse-integration/source-archive/streaming-server-upstream-plan.md)
- Draft PR (closed) that establishes the Dynamo shape:
https://github.com/ai-dynamo/dynamo/pull/7544
+116
View File
@@ -0,0 +1,116 @@
# OpenAI-compatible HTTP Contract
The stateless FastVideo HTTP server lives at
[`fastvideo/entrypoints/openai/`](https://github.com/hao-ai-lab/FastVideo/tree/main/fastvideo/entrypoints/openai).
Launch: `fastvideo serve --config serve.yaml`.
## Endpoints
| Method | Path | Description |
| --- | --- | --- |
| `POST` | `/v1/videos/generations` | Synchronous video generation |
| `GET` | `/v1/videos` | List prior jobs held in the in-memory store |
| `GET` | `/v1/videos/{id}` | Job status / result |
| `GET` | `/v1/videos/{id}/content` | Download the MP4 once ready |
| `POST` | `/v1/images/generations` | Synchronous image generation |
| `GET` | `/v1/models` | Enumerate registered models |
| `GET` | `/health` | Liveness probe |
## `VideoGenerationsRequest` shape
Mirrors the OpenAI `POST /v1/videos/generations` shape:
```json
{
"prompt": "a fox running through snow",
"size": "1024x1536",
"seconds": 5,
"fps": 24,
"num_frames": 121,
"seed": 42,
"num_inference_steps": 8,
"guidance_scale": 1.0,
"negative_prompt": "blurry, low quality",
"input_reference": "/path/to/init.png"
}
```
SGLang-compatible extensions carried today:
`num_inference_steps`, `guidance_scale`, `guidance_scale_2`,
`true_cfg_scale`, `negative_prompt`, `enable_teacache`, `output_path`.
## Merge precedence
The server builds a `GenerationRequest` each call using three layers,
highest first:
1. **Request body (client-explicit)** — only fields carried in
`request.model_fields_set` (Pydantic v2). Unset fields do not count,
even if the Pydantic model has a schema default for them.
2. **`ServeConfig.default_request` (operator-explicit)** — projected via
[`explicit_request_updates()`](../../../fastvideo/api/compat.py);
only fields the operator actually wrote into the YAML count as
defaults. Every other field inherits the schema default rather than
being pinned.
3. **Hardcoded fallback** — e.g. `fps = 24`.
The gate matters: both surfaces carry schema defaults. Without
`model_fields_set` / explicit-path tracking, schema defaults would
masquerade as intent and silently shadow the other side.
See [`video_api.py::_build_generation_kwargs`](../../../fastvideo/entrypoints/openai/video_api.py)
for the canonical implementation; the per-request assembly lives there,
not in pipeline code.
## Continuation state
The stateless surface accepts an opaque `ContinuationState` round-trip.
Clients that want continuation pass the prior `state` blob back on the
next request, and receive a new one on the response when
`request.output.return_state = true`.
Shape:
```json
{
"state": {
"kind": "ltx2.v1",
"payload": { "schema_version": 1, "segment_index": 3, ... }
}
}
```
Payload is always JSON-serializable. Large tensors may live in an
opaque blob-store reference the client simply round-trips; see
[`LTX2ContinuationState`](../../../fastvideo/pipelines/basic/ltx2/continuation.py).
Continuation is not yet wired all the way through to
`generator.generate_video(...)` — PR 7.6 (GPU pool upstream) is the
pipeline-level consumer. PR 7 locked the envelope so this surface is
stable ahead of that plumbing.
## Error codes
| HTTP | Condition |
| --- | --- |
| `400 Bad Request` | Parse/validation failure (unknown field, type mismatch, incompatible preset/state) |
| `404 Not Found` | `GET /v1/videos/{id}` for an unknown job |
| `409 Conflict` | Job id already exists |
| `500 Internal Server Error` | Pipeline raised; body mirrors upstream OpenAI error envelope |
| `503 Service Unavailable` | No generator loaded, or shutdown in progress |
Errors include a JSON body with
`{"error": {"type": "...", "message": "..."}}` matching the OpenAI
Python SDK's expectation.
## What does not cross this boundary
* Flat legacy kwargs (`ltx2_refine_enabled`, `torch_compile_kwargs`,
etc.) — these are init-time, configured via `ServeConfig.generator`,
never per-request.
* Private Dreamverse-only fields — those live in a private adapter on
the Dreamverse side; the public FastVideo surface never promises
backward compatibility for them.
* Raw tensor payloads (`ltx2_audio_clean_latent` et al.) — these are
derived by the pipeline from `ContinuationState`, never shipped as
request fields.
+2 -2
View File
@@ -226,7 +226,7 @@ Dataclass carrying all pipeline state between stages. Key field groups:
`width_latents`.
- **Scheduler**: `timesteps`, `num_inference_steps`, `guidance_scale`,
`sigmas`.
- **Task-specific**: `mouse_cond`/`keyboard_cond` (MatrixGame), `pose`
- **Task-specific**: `mouse_cond`/`keyboard_cond` (Matrix-Game 2.0), `pose`
(HYWorld), `camera_states` (GameCraft), `c2ws_plucker_emb`
(LingBotWorld).
- **Output**: `output: Tensor | None`.
@@ -252,7 +252,7 @@ Standard stages (typical execution order):
Specialized variants: `CausalDenoisingStage`, `LTX2DenoisingStage`,
`LongCatDenoisingStage`, `GameCraftDenoisingStage`,
`HYWorldDenoisingStage`, `MatrixGameDenoisingStage`,
`HYWorldDenoisingStage`, `MatrixGame2CausalDenoisingStage`,
`SRDenoisingStage`, `LTX2AudioDecodingStage`, `SD35ConditioningStage`,
`LTX2TextEncodingStage`, `LTX2LatentPreparationStage`.
+2
View File
@@ -107,6 +107,8 @@ If you encounter CUDA out of memory errors:
(single GPU) or `use_fsdp_inference=True` (multi-GPU)
- Try a smaller model or use distilled versions
- Use `num_gpus` > 1 if multiple GPUs are available
- Try enabling FSDP inference with `use_fsdp_inference=True` (may slow down generation)
- Try enabling DiT layerwise offload with `dit_layerwise_offload=True` (now only a few models support this, but may introduce less overhead than FSDP)
### Slow Generation
+187 -12
View File
@@ -11,6 +11,9 @@ This page describes the various options for speeding up generation times in Fast
- [Sliding Tile Attention (Archived)](#sliding-tile-attention-archived)
- [Sage Attention](#sage-attention)
- [Sage Attention 3](#sage-attention-3)
- [Adaptive Guidance (CFG gating)](#adaptive-guidance-cfg-gating)
- [torch.compile](#torch-compile)
## Attention Backends
@@ -21,6 +24,7 @@ This page describes the various options for speeding up generation times in Fast
- Video Sparse Attention: `FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN`
- Sage Attention: `FASTVIDEO_ATTENTION_BACKEND=SAGE_ATTN`
- Sage Attention 3: `FASTVIDEO_ATTENTION_BACKEND=SAGE_ATTN_THREE`
- Attn-QAT inference (modified SageAttention3 FP4, sm_120/RTX 5090): `FASTVIDEO_ATTENTION_BACKEND=ATTN_QAT_INFER`
- Video MoBA Attention: `FASTVIDEO_ATTENTION_BACKEND=VMOBA_ATTN`
- Sparse Linear Attention: `FASTVIDEO_ATTENTION_BACKEND=SLA_ATTN`
- SageSLA Attention: `FASTVIDEO_ATTENTION_BACKEND=SAGE_SLA_ATTN`
@@ -69,9 +73,9 @@ python setup.py install
### FP4 Flash Attention 4 (Blackwell only)
**`FLASH_ATTN`** with **`FASTVIDEO_NVFP4_FA4=1`**
**`FLASH_ATTN`** with **`--nvfp4_fa4`**
On Blackwell GPUs (B200/B300), you can enable FP4 quantized Q/K attention for up to **1.39x kernel speedup** over BF16 FA4, peaking at **1801 TFLOPS**. This quantizes Q and K to NVFP4 E2M1 with per-block E4M3 scale factors while keeping V in BF16.
On Blackwell GPUs (B200/B300), you can enable FP4 quantized Q/K attention for up to **1.31x kernel speedup** over BF16 FA4, peaking at **2018 TFLOPS**. This quantizes Q and K to NVFP4 E2M1 with per-block E4M3 scale factors while keeping V in BF16 or FP8.
See the [Attn-QAT paper](https://arxiv.org/abs/2603.00040) and [flash-attention-fp4 benchmark results](https://github.com/hao-ai-lab/flash-attention-fp4/blob/fp4/flash_attn/cute/README.md) for details.
@@ -83,32 +87,30 @@ See the [Attn-QAT paper](https://arxiv.org/abs/2603.00040) and [flash-attention-
#### Installation
Install the FP4 flash attention kernel and its dependencies:
Install the FP4 flash attention kernel (without upgrading your existing torch):
```bash
pip install "git+ssh://git@github.com/hao-ai-lab/flash-attention-fp4.git@fp4#subdirectory=flash_attn/cute"
pip install --no-deps "git+ssh://git@github.com/hao-ai-lab/flash-attention-fp4.git@fp4#subdirectory=flash_attn/cute"
pip install "nvidia-cutlass-dsl>=4.4.2" apache-tvm-ffi flashinfer-python
```
This installs the FP4 kernel and all dependencies (nvidia-cutlass-dsl, flashinfer-python, apache-tvm-ffi).
The `--no-deps` flag prevents upgrading torch/torchvision. The kernel requires torch >= 2.4 with CUDA 12.8+ support (already present in FastVideo's environment).
#### Usage
Enable FP4 attention via environment variables:
Enable FP4 attention via the `--nvfp4_fa4` flag:
```bash
FASTVIDEO_NVFP4_FA4=1 CUTE_DSL_ENABLE_TVM_FFI=1 python examples/inference/optimizations/fp4_attn_wan2_1_1_3b.py --nvfp4_fa4
python examples/inference/optimizations/fp4_attn_wan2_1_1_3b.py --nvfp4_fa4
```
Or in Python:
Or in Python via the `nvfp4_fa4` kwarg (sets env vars automatically):
```python
import os
os.environ["FASTVIDEO_NVFP4_FA4"] = "1"
os.environ["CUTE_DSL_ENABLE_TVM_FFI"] = "1"
from fastvideo import VideoGenerator
gen = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
nvfp4_fa4=True,
num_gpus=1,
use_fsdp_inference=False, # FSDP is incompatible with FP4 pointer path
)
@@ -121,6 +123,45 @@ gen.generate_video(prompt="A raccoon in sunflowers", save_video=True)
- Per-call cosine similarity vs BF16: ~0.99 (slight quantization error accumulates over denoising steps)
- Only supports `headdim >= 128`
### NVFP4 + Attn-QAT (modified SageAttention3, Blackwell sm_120)
**`ATTN_QAT_INFER`** with **`transformer_quant=nvfp4_qat`**
Runs the DiT fully in 4-bit: NVFP4 linear layers (activations quantized on the
fly) plus the modified SageAttention3 FP4 attention backend. This is the
inference half of the Quantization-Aware Distillation (QAD) recipe and the path
used for the RTX 5090 release.
The `attn_qat_infer` kernel hard-gates on **sm_120 (consumer Blackwell / RTX
5090)**; on other GPUs the backend logs a notice and falls back to Flash
Attention. See the [Attn-QAT paper](https://arxiv.org/abs/2603.00040).
Enable both halves — attention via the env var, linear via `transformer_quant`:
```python
import os
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "ATTN_QAT_INFER"
from fastvideo import VideoGenerator
from fastvideo.layers.quantization import get_quantization_config
gen = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
num_gpus=1,
# Wan-2.1 uses the nvfp4_qat config (NVFP4 is LTX2-specific). Pass an
# instance — the bare string is not resolved on the from_pretrained path.
transformer_quant=get_quantization_config("nvfp4_qat")(),
use_fsdp_inference=False, # FSDP shards invalidate the FP4 tensor pointers
)
gen.generate(request={"prompt": "A raccoon in sunflowers", "output": {"save_video": True}})
```
Or run the example script:
```bash
python examples/inference/optimizations/nvfp4_qat_wan2_1_1_3b.py
python examples/inference/optimizations/nvfp4_qat_wan2_1_1_3b.py --bf16 # baseline
```
### Sliding Tile Attention (Archived)
**`SLIDING_TILE_ATTN`**
@@ -177,6 +218,106 @@ These backends are model-specific and require the corresponding kernels and
dependencies. Use the support matrix and model examples to confirm compatibility
before enabling them.
<a id="torch-compile"></a>
## torch.compile
FastVideo can `torch.compile` the DiT (transformer) for a substantial
end-to-end speedup. It is **off by default** and enabled per-run.
### Enabling
```python
generator = VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
enable_torch_compile=True,
)
```
A complete A/B example (eager vs compiled, warmup excluded) is in
[`examples/inference/optimizations/torch_compile_example.py`](https://github.com/hao-ai-lab/FastVideo/blob/main/examples/inference/optimizations/torch_compile_example.py).
`fastvideo generate` is config-file driven; to enable `torch.compile`
from the CLI, set the relevant field in your run config and pass it via
`fastvideo generate --config run.yaml`. There is no top-level
`--enable-torch-compile` flag on the subcommand.
Only DiT submodules that declare `_compile_conditions` are compiled
(most shipped models). The text encoder and VAE are not compiled by this
flag.
### What to expect
| Config | Effect |
|---|---|
| Wan2.1-T2V-1.3B, A100-80GB, 480×832×81f, 50 steps | end-to-end **259.7s → 198.1s (−23.7%)**; per-step **4.91 → 3.78 s/it** |
The speedup is **configuration-dependent** — it varies with model,
resolution, step count, and GPU. Treat the number above as one measured
data point, not a guarantee; benchmark your own config (recipe below).
There is a **one-time graph-build cost** on the first generation (tens of
seconds to minutes, model-dependent). It amortizes over subsequent
generations with the same input shapes. Always exclude the first
(warmup) generation when measuring steady-state latency — measuring the
warmup is the most common way to wrongly conclude "compile is slower".
**Numerics.** Inductor's lowering is designed to preserve eager
semantics within floating-point tolerance, but per-model equivalence is
not asserted by any standing SSIM regression here — the SSIM tests in
[`fastvideo/tests/ssim/`](https://github.com/hao-ai-lab/FastVideo/tree/main/fastvideo/tests/ssim)
run with `enable_torch_compile` disabled. If you depend on compile
output staying close to eager (or your previous compiled run), run an
MS-SSIM gate on *your* config, especially when combining
`enable_torch_compile=True` with other numerics-affecting flags
(quantized attention backends, FP4, layerwise offload edge cases).
### Known interactions
- **Layerwise CPU offload** (`dit_layerwise_offload=True`, the default):
the offload hook previously caused an implicit graph break once per
transformer layer, fragmenting the compiled region. Addressed in
hao-ai-lab/FastVideo#1365 — keep that fix to get a clean compiled
region under the default offload path.
- **`mode="reduce-overhead"` / CUDA graphs**: not yet supported
end-to-end. The attention dispatch is an untraceable custom op and
still breaks the graph, which CUDA-graph trees cannot span. Use the
default inductor mode (shown above) until that is resolved.
Extra `torch.compile` options are passed through `torch_compile_kwargs`
(a dict), accepted by `VideoGenerator.from_pretrained(...)` and by the
CLI as a JSON string via `--torch-compile-kwargs`. Example (currently
**not** recommended — see the CUDA-graphs caveat above):
```python
VideoGenerator.from_pretrained(
"Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
enable_torch_compile=True,
torch_compile_kwargs={"mode": "reduce-overhead"}, # may error today
)
```
### Benchmarking torch.compile
Same discipline as attention backends — same prompt, same seed, same
config; **discard the first generation** (graph build):
```python
import time
from fastvideo import VideoGenerator
gen = VideoGenerator.from_pretrained("your-model-id", enable_torch_compile=True)
req = {"prompt": "Your prompt", "sampling": {"seed": 1024},
"output": {"save_video": False}}
gen.generate(req) # warmup: graph build, discard
t0 = time.perf_counter()
gen.generate(req) # measured: shapes reused
print(f"compiled steady-state: {time.perf_counter() - t0:.2f}s")
```
See `examples/inference/optimizations/torch_compile_example.py` for a
baseline-vs-compile A/B with the warmup correctly excluded.
## Benchmarking different optimizations
To benchmark backend performance, generate the same prompt with the same seed and compare end-to-end generation times:
@@ -198,3 +339,37 @@ for backend in ["TORCH_SDPA", "FLASH_ATTN", "SAGE_ATTN"]:
```
Note: reinstantiate `VideoGenerator` after changing `FASTVIDEO_ATTENTION_BACKEND`.
## Adaptive Guidance (CFG gating)
CFG gating accelerates classifier-free guidance by reusing the cached
`noise_pred_cond - noise_pred_uncond` delta after a configurable fraction of
the denoising schedule, skipping the unconditional model forward for the
remaining steps. The technique is the LinearAG variant of Adaptive Guidance
(Castillo et al. 2023, [arXiv:2312.12487](https://arxiv.org/abs/2312.12487)).
### Enabling
Set the `FASTVIDEO_CFG_GATE_STEP` environment variable to a float in `[0, 1]`:
| Value | Behavior |
|-------|----------|
| `1.0` (default) | Disabled — legacy two-pass CFG every step. |
| `0.5` | Cache the delta after `len(timesteps) * 0.5` steps; reuse for the rest. |
| `0.0` | Cache from the very first step (most aggressive). |
```bash
export FASTVIDEO_CFG_GATE_STEP=0.5
```
### Trade-offs
- **Memory**: one extra model-output-sized tensor per rank held during the
gating window.
- **Quality**: VBench-measured quality is preserved within noise on 4 of 5
dimensions at `FASTVIDEO_CFG_GATE_STEP=0.5` for Wan T2V 1.3B per the PR's
reported numbers (see [#1372](https://github.com/hao-ai-lab/FastVideo/pull/1372)).
- **Speed**: ~22% e2e on 4xL40S and ~24% on 1xH100 at the same settings.
Default behavior is byte-for-byte equivalent to the legacy two-pass CFG path;
the feature is fully opt-in.
+8 -3
View File
@@ -58,6 +58,7 @@ pipeline initialization and sampling.
| FastWan2.1 T2V 1.3B | `FastVideo/FastWan2.1-T2V-1.3B-Diffusers` | 480P | ⭕ | ⭕ | ⭕ | ✅ | ⭕ |
| FastWan2.2 TI2V 5B Full Attn* | `FastVideo/FastWan2.2-TI2V-5B-FullAttn-Diffusers` | 720P | ⭕ | ⭕ | ⭕ | ✅ | ⭕ |
| Wan2.2 TI2V 5B | `Wan-AI/Wan2.2-TI2V-5B-Diffusers` | 720P | ⭕ | ⭕ | ✅ | ⭕ | ⭕ |
| Lucy Edit Dev 5B*** | `decart-ai/Lucy-Edit-Dev` | 480P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| Wan2.2 T2V A14B | `Wan-AI/Wan2.2-T2V-A14B-Diffusers` | 480P<br>720P | ❌ | ❌ | ✅ | ⭕ | ⭕ |
| Wan2.2 I2V A14B | `Wan-AI/Wan2.2-I2V-A14B-Diffusers` | 480P<br>720P | ❌ | ❌ | ✅ | ⭕ | ⭕ |
| HunyuanVideo | `hunyuanvideo-community/HunyuanVideo` | 720px1280p<br>544px960p | ❌ | ✅ | ✅ | ⭕ | ⭕ |
@@ -70,13 +71,17 @@ pipeline initialization and sampling.
| TurboWan2.1 T2V 14B | `loayrashid/TurboWan2.1-T2V-14B-Diffusers` | 480P, 720P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| TurboWan2.2 I2V A14B | `loayrashid/TurboWan2.2-I2V-A14B-Diffusers` | 480P<br>720P | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| LongCat T2V 13.6B | See note** | 480P<br>720P | ❌ | ❌ | ❌ | ⭕ | ✅ |
| Matrix Game 2.0 Base | `FastVideo/Matrix-Game-2.0-Base-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| Matrix Game 2.0 GTA | `FastVideo/Matrix-Game-2.0-GTA-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| Matrix Game 2.0 TempleRun | `FastVideo/Matrix-Game-2.0-TempleRun-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| Matrix Game 2.0 Base Distilled | `FastVideo/Matrix-Game-2.0-Base-Distilled-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| Matrix Game 2.0 GTA Distilled | `FastVideo/Matrix-Game-2.0-GTA-Distilled-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| Matrix Game 2.0 TempleRun Distilled | `FastVideo/Matrix-Game-2.0-TempleRun-Distilled-Diffusers` | 352x640 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| Matrix Game 3.0 Base Distilled | `FastVideo/Matrix-Game-3.0-Base-Distilled-Diffusers` | 720x1280 | ⭕ | ⭕ | ⭕ | ⭕ | ⭕ |
| GEN3C Cosmos 7B | `FastVideo/GEN3C-Cosmos-7B-Diffusers` | 704px1280p | ❌ | ❌ | ❌ | ⭕ | ⭕ |
**Note**: Wan2.2 TI2V 5B has some quality issues when performing I2V generation. We are working on fixing this issue.
***Lucy Edit Dev uses a non-commercial model license. FastVideo support is
focused on inference integration for video editing workflows.
`Sliding Tile Attn (Legacy Branch)` entries refer to the archived
`sta_do_not_delete` branch workflow, not active `main` inference wiring.
+28 -2
View File
@@ -27,7 +27,29 @@ Useful variables:
- `FASTVIDEO_LOGGING_LEVEL`: `DEBUG`, `INFO`, `WARNING`, `ERROR`
- `FASTVIDEO_STAGE_LOGGING`: print per-stage timings during pipeline execution
- `FASTVIDEO_ATTENTION_BACKEND`: force an attention backend (for example
`TORCH_SDPA` or `FLASH_ATTN`)
`TORCH_SDPA`, `FLASH_ATTN`, `SAGE_ATTN_THREE`, or
`ATTN_QAT_INFER`, or `ATTN_QAT_TRAIN`)
## Layer-by-Layer Activation Tracing
For numerical-divergence debugging — typically when porting a new model and
needing to find the first layer where FastVideo and an upstream reference
produce different outputs — use the env-gated activation trace mode:
```bash
FASTVIDEO_TRACE_ACTIVATIONS=1 \
FASTVIDEO_TRACE_LAYERS="^transformer\.blocks\.\d+$" \
FASTVIDEO_TRACE_OUTPUT=/tmp/fv_trace.jsonl \
python your_script.py
```
The trace dumps per-tensor stats (`abs_mean`, `sum`, `shape`, etc.) to a
JSONL file. Run the same workload with tracing on the upstream side, then
`diff` the two files to localize the first divergent layer.
See [Activation Trace Mode](../contributing/activation_trace.md) for the
full guide (env var reference, JSONL output schema, parity-debug workflow,
performance impact, and troubleshooting).
## Common Failure Modes
@@ -52,7 +74,11 @@ If forcing a backend fails, verify optional dependencies are installed:
- `VIDEO_SPARSE_ATTN`: `fastvideo-kernel`
- `SLIDING_TILE_ATTN`: STA legacy workflow in
`sta_do_not_delete` + `fastvideo-kernel`
- `SAGE_ATTN` / `SAGE_ATTN_THREE`: SageAttention packages
- `SAGE_ATTN`: SageAttention package
- `SAGE_ATTN_THREE`: upstream `sageattn3` package
- `ATTN_QAT_INFER`: `fastvideo-kernel` checkout/source install that exposes
`attn_qat_infer`
- `ATTN_QAT_TRAIN`: `fastvideo-kernel` install exposing `fastvideo_kernel`
As a fallback, use:

Some files were not shown because too many files have changed in this diff Show More