Compare commits

..
372 Commits
Author SHA1 Message Date
Peiyuan Zhang 67f5b53595 [fix]: keep v2 examples runnable 2026-07-04 20:22:18 +00:00
Peiyuan Zhang aa6bf5079b [docs]: scope v2 docs to inference runtime 2026-07-04 19:33:05 +00:00
Peiyuan Zhang 4ca56d8951 [fix]: route v2 serving tasks by capabilities 2026-07-04 19:32:46 +00:00
Peiyuan Zhang 1e19650eb3 [misc]: simplify v2 program execution model 2026-07-04 19:32:30 +00:00
Peiyuan Zhang de354f804c [feat]: add FastWan QAD FP8 card 2026-07-04 19:32:18 +00:00
Peiyuan Zhang 6feee4e4fa [feat]: add FP8 vendor support for Wan weights 2026-07-04 19:32:05 +00:00
Peiyuan Zhang 82e00615e5 [feat]: wire v2 VideoGenerator to real torch backend 2026-07-04 19:31:55 +00:00
Peiyuan Zhang 30cfbfc4f6 [refactor]: remove training semantics from v2 inference contracts 2026-07-04 19:31:35 +00:00
Peiyuan Zhang a4d8978c37 [refactor]: remove v2 training package 2026-07-04 19:31:12 +00:00
Will Lin 803a6c99ae [refactor] v2: drop unwired ConditioningInjector policy
ConditioningInjector (ABC) + PassthroughConditioning were defined and exported
but never instantiated or called — zero call sites anywhere. Unlike the policies
the §5 thesis actually names (CFG / flow-shift / precision / expert-routing),
which every recipe wires (cfg=ClassicCFG(), expert=NoRouting(), ...), conditioning
is injected INLINE by the loops: `st.cond["prompt_embeds"] = ctx.slots.get(...)`
(wan21/loop.py, ltx2/loop.py); the qwen_omni cascade conditions via loop wiring.
PassthroughConditioning even described a dataflow (state.scratch["cond"]) that is
not how conditioning actually flows (loops read ctx.slots). So it was a designed-
but-bypassed seam, not forward-design — and conditioning isn't in the README §5
policy list. Removed the ABC + impl + exports; kept the live `cond` field (the
loops fill it directly) and fixed its comment.

Tests: 143 passed, 2 skipped.
2026-07-04 17:15:35 +00:00
Will Lin 8a637c3215 [refactor] v2: remove unused DataRef provenance spec
DataRef (dataset_id/revision/description — "what a recipe trained on") was only
the type of RecipeSpec.data_contract, which no recipe authored and no code read
(0/0). It is absent from the README RecipeSpec contract (§2.1: method, parents,
assumes_loop, assumes_precision, consistency_required) and not among the §18
"wire the inert metadata" roadmap items — i.e. unwired governance metadata, not
deliberate forward-design. RecipeSpec now matches §2.1 exactly.

Kept CheckpointManifest: unlike DataRef it IS wired (ModelCard.checkpoint), the
declarative "explicit components + key maps, no name-detector guessing" load
contract — declared intent, not dead.

Tests: 143 passed, 2 skipped.
2026-07-04 17:15:35 +00:00
Will Lin 51bae33c7b [refactor] v2: drop unused typed-schema slots from Component/LoopSpec
state_schema / step_schema / result_schema (LoopSpec) and config_schema /
io_schema (ComponentSpec), plus valid_parallel_plans and parallel_constraints,
were authored by no recipe and read by no executor (audited: 0 reads / 0 sets
across native + tests, no dynamic dataclasses.fields/asdict/__dict__ access, no
consumer in scripts/_vendor/examples). They duplicated mechanisms that already
exist: the concrete LoopState/WorkPlan/StepResult classes the driver uses
directly, and the card-level ParallelismContract. Removing them slims the core
spec surface with zero behavior change.

Thesis untouched: cards still own components/loops/recipe/parity; kept the
parity, precision/placement, behavior-capture (behavior_schema), wired
extension_schema, and roadmap required_for/optional_for/resident_for fields.
Also drop the stale README step_cost_model mention (that field went with the
cost mechanism in 352c1b28).

Tests: 143 passed, 2 skipped (toy backend).
2026-07-04 17:15:34 +00:00
Will Lin 54b7fe3a55 [refactor] v2: drop dead toy components + unused Karras schedule
ToyLoRA / ToyControlNet / ToyTargetModel / ToyDraftModel / ToyRewardModel (+ the
_spec_target_next helper) in the toy backend, and build_karras_sigmas in the
sampler, were defined but referenced nowhere — no recipe, card, loop, test, or
example used them. They were toy stand-ins for capabilities (adapter plane,
speculative decode, served reward) and an EDM/Karras noise schedule that were
written ahead of being wired. Remove them (-134 LOC); a toy can be re-added when
the capability is actually wired. No thesis impact, no behavior change.

Suite green: 143 passed, 2 skipped (toy backend); import v2 stays torch-free.
2026-07-04 17:15:34 +00:00
Will Lin 1fe50c0092 [refactor] v2: group flat top-level into planes; isolate vendored under _vendor/
The v2 top level had grown to ~27 dirs + 13 loose files — half of them
single-concept abstraction shells (memory/ 86 LOC, transport/ 222, parity/ 178,
extend/ 226) and half vendored fastvideo code sitting as peers to the actual v2
design. Regroup to mirror README section 3 "Planes & dependency order":

  core/     enums+types, card, loop, program, parity, request, parallel
            (the model-native contracts; no kernels)
  runtime/  + folded-in substrate: cache, memory, transport, extend
            (the import graph shows only runtime consumes them; compile/cudagraph
            already lived here)
  serving/  + deploy/ (products / fleet)
  _vendor/  all copied fastvideo: models, layers, attention, configs, distributed,
            platforms, api, hooks, logging_utils, third_party + fastvideo_args/
            utils/logger/envs/forward_context — internal layout unchanged, still
            mirrors upstream for diffing

28 dirs + 13 root files -> 8 dirs + 6 root files. Pure mechanical move: 947
absolute-import paths rewritten (v2.X -> v2.{core,runtime,serving,_vendor}.X),
boundary-anchored so platform/ (native dispatch) and platforms/ (vendored CUDA
detect) no longer collide and neither does hooks/ vs extend/. Deleted the empty
v2/loader/. README section 3 + 16 updated; stale test count corrected
(34 files/216 tests -> 22 files/143 tests).

Validated: `import v2` stays torch-free; `v2/run_tests.py` and `pytest v2/tests/`
-> 143 passed, 2 skipped on the numpy toy backend. (On a GPU box force the toy
backend with CUDA_VISIBLE_DEVICES="" or detect() picks cuda.) No external importer
changed — examples use the re-exported `from v2 import VideoGenerator`.
2026-07-04 17:15:34 +00:00
SolitaryThinker 290795daf8 [refactor] v2: remove interleave; pooled run-to-completion serving (P2)
Second step of the runtime simplification (after cost removal). Drop the
coordinated step-interleave scheduler + the interleave-parity gate; serving is
now pooled run-to-completion.

- Engine: remove run_interleaved + the WorkUnit/BatchScheduler imports; run /
  run_serial drive each request to completion (tick/run_to_completion kept as the
  per-request stepper).
- scheduler.py: remove BatchScheduler + WorkUnit + the batches metric; the
  AdmissionController is now a pure refundable memory/OOM guard.
- AsyncEngine: bound concurrency with a serving pool (asyncio.Semaphore,
  max_concurrent) — each request waits for a slot, then runs to completion.
- parity: remove assert_interleave_parity (the run_serial==run_interleaved gate);
  rename interleave_gate.py -> compare.py (compare_outputs stays — bit-parity
  between execution paths, e.g. disaggregated==inline).
- card specs: drop ParitySpec.interleave_required + LoopSpec.allows_interleaving;
  stripped interleave_required from all cards.
- tests: delete the interleave-gate/parity tests; refocus the ones that exercised
  real behavior (residual-skip, compare_outputs symmetric-empty).
- README: removed the design-doc references at the top; simplified the thesis /
  scheduler (§6) / parity (§9) / package-layout / comparison sections to pooled
  run-to-completion (no cost model, no interleave gate).

CPU mini: 143 passed / 2 skipped. Native omni port (P3) still to come. (pyproject
kernel hack excluded.)
2026-07-04 17:15:34 +00:00
SolitaryThinker 0047d54a1b [refactor] v2: remove the cost mechanism (P1 of runtime simplification)
First step toward lean pooled run-to-completion serving: rip out the GPU-time
cost/budget machinery entirely (it priced nothing useful for the target design).

Removed: CostModel + LoopSpec.step_cost_model; ResourceRequest.compute_seconds;
StepResult.actual_seconds; AdmissionController's compute budget + SchedulerMetrics
.gpu_seconds (the memory/OOM reservation guard stays); the Profiler observer
(cost calibration); per-step timing in RuntimeLoopContext; cost-based fleet/Dynamo
routing (now a coarse step-count load proxy); DeploymentCard.cost_model. Stripped
cost from all 9 recipe cards + their loops.

Tests: dropped the 3 cost-specific tests (cost routing, cost_model aliasing,
compute-budget gate); refocused 2 (loop cache validation, NaNWatch-clean).

CPU mini: 151 passed / 2 skipped. Interleave removal + pooled serving (P2) and the
native omni port (P3) follow. (pyproject kernel hack excluded as always.)
2026-07-04 17:15:34 +00:00
SolitaryThinker b2d55a7ba0 [refactor] v2: full vendor cutover — copy fastvideo modeling + layer code into v2 (zero fastvideo imports)
Replace the re-export stubs with real vendored copies of the fastvideo modeling
+ layer + supporting infra, for the kept diffusion models (wan21, wan_causal,
ltx2, flux2, matrixgame2). v2 now imports ZERO `fastvideo.*` — it is
self-contained. (bagel/qwen_omni load from vllm_omni, an external pkg, not
fastvideo; cosmos3's load_id was already dangling — both out of scope here.)

Vendored (cp + `sed fastvideo. -> v2.`):
- models/  the 5 models' nn.Module dits/vaes/encoders/audio/upsamplers + the
           component loader/ + the lazy class registry (other families' rows are
           dormant/lazy — only the 5 resolve).
- layers/ attention/ platforms/ distributed/ configs/ logging_utils/ hooks/
  third_party/pynvml + top-level forward_context/fastvideo_args/envs/logger/
  utils/version — copied verbatim (layers et al. 'as is').
- api/  slimmed to schema + results (the VideoGenerator's config dataclasses);
  the fastvideo parser/presets/overrides (which pull the pipeline runtime) are
  intentionally NOT vendored — v2 has its own runtime/loop.

Decoupling surgery (cut the loader's coupling to the fastvideo runtime):
- configs/pipeline_registry.py (vendored from fastvideo/registry.py, renamed to
  avoid colliding with v2/registry.py): dropped the _register_presets() auto-call
  and matrixgame3 (removed model); config-class resolution preserved.
- configs/pipelines/__init__.py: dropped the registry back-edge (fixes an import
  cycle) — base.py imports the registry lazily where used.
- torch_backend: load_component now uses v2.models.loader.

The only remaining external 'fastvideo*' refs are `fastvideo_kernel` (the
separate optional CUDA-kernel pkg for sparse/MoBA attention) — guarded; the dense
TORCH_SDPA path v2 uses never imports it.

Vendored subtrees added to the pre-commit exclude (faithful copies, mirroring the
existing fastvideo/models exclusion — not re-linted, to stay re-syncable).

Verified: grep finds zero fastvideo-package imports in v2/; Wan2.1 T2V on H100 is
BIT-IDENTICAL to the fastvideo-backed path (same .npy SHA256, byte-for-byte); CPU
mini 156 (154 passed + 2 env-skipped on x86, torch present). Backup: branch v2_backup.
2026-07-04 17:15:34 +00:00
SolitaryThinker 24cfe281b3 [refactor] v2: prune recipes to 8 models (+ omni shared infra)
Keep: bagel, cosmos3, flux2, matrixgame2, ltx2, qwen_omni, wan21, wan_causal
(plus the shared omni/ package that bagel/cosmos3/qwen_omni depend on). Remove
the other 26 recipe packages.

- Delete 26 recipe dirs (adapters, adaptive, cosmos2, cosmos25, fastwan, gen3c,
  hunyuangamecraft, hunyuan_video(15), hyworld, image_video, kandinsky5,
  lingbotworld, longcat, lucy_edit, matrixgame3, multi_expert, reward, sd35,
  sfwan22, speculative, stable_audio, tiled, turbowan, unified, wan_fun_control).
- recipes/__init__.py: keep-closure builders + build_default_engine /
  build_omni_engine (dropped the workflow/tiled/unified/image_video engine helpers).
- registry.py: _BUCKET_C pruned to flux2 + matrixgame2; removed the cosmos2
  ModelEntry + the CosmosTransformer3DModel arch branch + the TurboWan-14B entry.
- Delete 11 tests for removed recipes/features; patch test_bucket_c_ports to drop
  the cosmos2 reference (it now auto-derives from the pruned _BUCKET_C).
- README: correct the recipes/ roster to the kept families.

Backup of the full pre-prune tree is on branch v2_backup. CPU mini 156 passed / 0
failed (the prior 5 torch-absent bucket_c failures are gone with sd35/stable_audio);
no dangling references to any removed recipe; all kept model ids still resolve.
2026-07-04 17:15:34 +00:00
SolitaryThinker dc086b207f [perf] v2: on-device denoise loop — kill the per-step numpy<->torch round-trip (Wan2.1)
The torch adapter boundary marshalled the latent host<->device on EVERY denoise
step: _t uploaded the latent (and re-uploaded the text embeds) and _n downloaded
the velocity with a forced CUDA sync — 2*N PCIe copies + N syncs per generation,
buying nothing, since the latent could stay resident on the GPU the whole loop.

Root cause was a numpy loop surface. But the loop MATH is already array-agnostic
(CFG combine + flow-match Euler are pure arithmetic; the solver kernel already
passes torch through). So introduce a per-platform array namespace (v2/platform/
array_ns): numpy on CPU (torch-free — the parity mini is unchanged), torch-on-
device on cuda. The latent is seeded with numpy and uploaded ONCE; it then stays
resident through forward -> CFG combine -> solver -> next step; a single host
marshal happens at the request/output boundary (engine._to_artifact).

Opt-in per recipe via ModelCard.device_io (set on the Wan cards). When set on a
GPU box, build_component flips the components' TorchComponent.device_io so _out
keeps tensors on-device (in fp32, matching the old _to_numpy cast so the combine
dtype is unchanged). Un-migrated families and the CPU toy keep numpy in/out.
Also: PrecisionPolicy.cast is array-preserving; _t accepts resident tensors.

Verified BIT-IDENTICAL on real Wan2.1-1.3B / H100: the on-device latent equals
the pre-change numpy-path latent exactly (max_abs_diff 0.0, np.array_equal True).
CPU mini holds 237 passed / 5 pre-existing; pre-commit clean.

Other WanDenoiseLoop families can flip device_io next (per-family GPU re-verify);
non-Wan loops migrate to the xp namespace later.
2026-07-04 17:15:34 +00:00
SolitaryThinker b9db151658 [refactor] v2: co-locate per-model torch adapters into their recipe packages
platform/backends/ had become a flat dump of 15 per-model torch_<model>.py
adapters next to the genuinely-shared infra. Each adapter is referenced from
its card by a plain 'module:Class' string loaded via importlib, so there was
no real coupling forcing it into platform/ — the Cosmos/Flux/etc adapter
belongs WITH its recipe (card/loop/program).

Move each torch_<model>.py -> v2/recipes/<model>/adapter.py and flip the card
strings to v2.recipes.<model>.adapter:<Class>. backends/ now holds only the
shared substrate (torch_backend base, torch_cuda registration, torch_kernels,
toy/cpu/accel). Each recipe is now a self-contained package.

Cross-refs updated: gen3c/adapter imports CosmosT5Encoder from cosmos2/adapter;
sd35/program + stable_audio/card import from their own package. Recipes still
import torch-free (adapters pulled only via the string on a GPU box) — CPU mini
holds 237 passed / 5 pre-existing (bucket_c torch-absent). pre-commit clean.
2026-07-04 17:15:34 +00:00
SolitaryThinker 9124963238 [misc] v2: full pre-commit clean (ruff UP038/SIM/UP031 + mypy annotations)
Sweep all of v2/ through pre-commit (was previously only run on changed
files). Fixes surfaced across untouched modules:

- ruff: isinstance-tuple -> X | Y (UP038), try/except/pass ->
  contextlib.suppress (SIM105), negated-return (SIM103), %-format ->
  f-string (UP031).
- mypy: add annotations for no-untyped-call + var-annotated across recipes,
  training methods, torch/toy backends, and serving.
- yapf reflow of the SF-Wan KV-cache call sites (semantics unchanged).

yapf/ruff/codespell/mypy all pass; v2 tests 237 passed / 5 pre-existing
(bucket_c torch-absent on CPU venv).
2026-07-04 17:15:34 +00:00
SolitaryThinker f90e8f3e76 [bugfix] v2: SF-Wan cross-chunk KV cache — condition each chunk on prior clean chunks
The causal adapter never passed a kv_cache, so CausalWanTransformer3DModel.forward routed to
_forward_train (no cross-chunk KV) on every chunk instead of _forward_inference (the CausVid
Alg-2 KV-cache path). Each chunk denoised blind to the previous ones; the loop's cross-chunk
"context" was a toy mean(prior_latents) the adapter ignored. Result: hard discontinuities at
every chunk boundary (frame-to-frame absdiff spikes 43-55 every ~12 frames).

Fix (cuda path only; toy/CPU path and the 237-test suite untouched):
- WanDiT.alloc_causal_caches(): allocate the persistent per-block KV + cross-attn caches sized
  from the model config (mirrors CausalDenoisingStage._initialize_kv_cache).
- WanDiT.__call__: thread kv_cache/crossattn_cache/current_start/cache_start/start_frame/
  frame_seqlen so the model runs _forward_inference.
- wan_causal/loop.py: own the caches in LoopState (per-request -> interleave-safe); pass
  current_start = chunk_idx*chunk_size*frame_seqlen per chunk; do the clean-KV write
  (timestep ~0) after each chunk so the next attends to it.

Verified on H100: frame-to-frame absdiff mean 12.8->4.4, max 55.4->8.3; chunk-boundary spikes
eliminated; coherent across all 7 chunks. CPU causal toy tests unchanged (24 passed).
2026-07-04 17:15:34 +00:00
SolitaryThinker 4ff33a28b7 [misc] v2: simplify docstrings/comments + drop deleted-design-doc citations
Sweep all 296 v2 modules: simplify verbose docstrings/comments and remove 376 dangling
"(design_vN §X)" citations to the now-deleted design docs (v2/README.md is the source of
truth). Comment/docstring-only — AST-verified code-identical; the CPU suite holds at 237
passed / 5 pre-existing. Also applies yapf + ruff --fix auto-fixes (import ordering,
forward-ref annotation de-quoting under `from __future__ import annotations`; behavior-
neutral, suite-confirmed) and adds the legitimate domain terms mot/clen/te to the codespell
ignore-list. Remaining ruff (24) + mypy (68 no-untyped-call) findings are pre-existing v2
debt, untouched here.
2026-07-04 17:15:34 +00:00
SolitaryThinker 940f94f435 [docs] v2: make v2/README.md the design source of truth + one-page philosophy + M* roadmap
Unify the four design docs (design.md, designv2.md, design_v3.md, designv4.md) into a single
authoritative v2/README.md: the (recipe, runtime) thesis, driven loops, planes, one-WorkUnit
scheduler, the parity ladder + interleave gate, training-on-shared-loops, the weight-sharing
topology catalog, the current GPU status (20+ models + the BAGEL/Qwen-Omni/Cosmos3 trio
verified), and a prioritized roadmap. Recast design_summary.md as a one-page design philosophy
pointing to it. Add .agents/exploration/mstar-v2-roadmap.md (the adversarially-verified M*
Walk-Graph gap analysis driving the roadmap). Delete the four superseded design docs.
2026-07-04 17:15:34 +00:00
SolitaryThinker ac9dcb63ea [bugfix] v2: MatrixGame2/3 causal-loop progress counter + MG3 patch alignment
MatrixGame2 (causal DMD loop): bump st.step_idx on every executed work unit (each
DMD step and each clean-context pass). The loop drives its own control flow off
block_idx/dmd_idx/phase, but the runtime's no-progress watchdog keys on
st.step_idx, so a multi-block causal rollout was seen as stalled. Mirrors what
every other recipe loop does.

MatrixGame3 (5B WanModel): patch-align the latent H/W (patch_size (1,2,2)) before
denoise. The model folds (H/2, W/2) tokens, so an odd latent dim made the
unpatchified velocity come back one row/col short of the noise latent. Crop to
(latent // patch) * patch, faithful to MatrixGame3DenoisingStage.

v2 mini: 240 passed.
2026-07-04 17:15:34 +00:00
SolitaryThinker 861f87e843 [bugfix] v2: correct SF-Wan + LTX2 2-stage SR sampling defaults (GPU frame-verified)
Two distilled few-step video models rendered incorrectly on GPU; root-caused via
dense frame sampling (contact sheets) and fixed in the recipe cards.

SF-Wan2.1 (self-forcing causal, wan_causal card): was oversaturated/overcooked.
The distilled student is CFG-FREE (guidance 1.0, single forward/step) and denoises
with the 4-step DMD schedule [1000,750,500,250] (warped by FlowShiftPolicy(5.0)),
at a native causal block of 3 latent frames. Defaults were ClassicCFG@6.0 + 2 steps
+ block 2 -> overcooked AND under-denoised. Fixed: num_chunks=7, chunk_size=3,
steps_per_chunk=4; SamplingDefaults num_steps=4, guidance_scale=1.0 (7x3=21 latent
-> 81 frames). Renders a clean raccoon-in-sunflowers across all 81 frames.

LTX2-Distilled 2-stage SR (ltx2 card + LTX2VAE): was temporally blocky. Root cause
was an OOM-forced 57-frame reduction (only 8 latent temporal frames); the model is
designed for 121 (16 latent frames). Enable VAE tiling in LTX2VAE so the 121-frame
full-res decode fits the 80 GiB GPU; keep base cfg_scale=3.0 (drives brightness;
cfg=1 washed out) with stg_scale=0.0 (v2's drop-text perturbation is not real
skip-layer STG). Renders the on-prompt backyard shot, bright + temporally coherent.

torch_backend.py: enable LTX2VAE tiling; clear pre-existing mypy no-untyped-call /
yapf debt on the file (surfaced once per-file linting bypassed the duplicate-module
flakiness) by annotating the helper/constructor/maker signatures.

Tests: update the 3 affected CPU defaults/chunk-count tests. v2 mini: 240 passed.
2026-07-04 17:15:34 +00:00
SolitaryThinker 7ef2d9083e [docs] v2: GPU bring-up results — 20 models verified on H100
V2_PORTING_STATUS.md now records the GPU bring-up outcome: 20 models generate
real video/audio on H100 (the 7 prior + 13 newly-ported), with the remaining
split into fastvideo/env-blocked (SLA/VSA kernels, transformers incompat, fastvideo
registry/flash_attn gaps) and HF-access-blocked (gated cosmos2/flux2/sd35) — none
a v2 recipe bug.
2026-07-04 17:15:34 +00:00
SolitaryThinker 0ac5367b54 [feat] v2 GPU bring-up: 5 more models verified on H100 (huge MoE/world + DMD)
Second GPU pass (distilled + huge dense, 2-wide). 5 verified end-to-end with real
weights; CPU toy path kept green (240 passed, 2 skipped).

VERIFIED:
  * matrixgame3   — mp4 (9,256,256,3), zero fixes. 6.47B, standard Wan attn (NOT
                    sparse-attn-blocked, like its mg2 sibling); degenerate single-clip.
  * fastwan       — FastWan2.2-TI2V-5B-FullAttn DMD 3-step, mp4 (17,256,256,3), zero
                    fixes. The FULL-ATTENTION variant has no VSA params -> the generic
                    Wan loader maps it cleanly (reuses WanDiT via load_id, no adapter).
  * longcat       — LongCat-Video-T2V 13.58B, mp4 (17,256,256,3), zero fixes, CPU
                    offload (~40GB peak).
  * sfwan22       — Self-Forcing Wan2.2-A14B causal+MoE (2x14B), mp4 (29,288,288,3),
                    CPU expert offload (~80GB peak).
  * lingbotworld  — Wan2.2-class 2x14B camera world model, mp4 (9,256,256,3), offload
                    (~98GB peak transient).

BLOCKED: turbowan-i2v-a14b — SLA sparse-attn params (attn1.attn_impl.proj_l on all
40 layers of both experts) cannot load into the dense Wan build + needs the
fastvideo-kernel SLA Triton kernels (no nvcc here). Confirmed via the safetensors
header (no 60GB download). Same SLA family as turbowan-1.3b.

Fixes (own-port only): sfwan22/loop.py, lingbotworld/{card,program}.py +
torch_lingbotworld.py. pre-commit clean per-package.
2026-07-04 17:15:34 +00:00
SolitaryThinker 56f84018df [bugfix] v2 VideoGenerator: modality-aware result path (audio/image, not only video)
VideoGenerator._result hardcoded out.artifacts['video'].frames, so an audio-only
(Stable Audio) or image-only (SD3.5 / FLUX.2 T2I) generation crashed with
KeyError 'video' even though the engine had correctly produced the AudioArtifact /
image TensorArtifact (surfaced during stable_audio GPU bring-up). _result now
guards on the artifact present: video -> mp4 (unchanged), else image -> png
([C,H,W]/[B,..] normalized to [H,W,C]), plus the existing audio -> sibling .wav;
image_path recorded in result.extra. The video path is byte-for-byte unchanged.

v2 mini 240 passed, 2 skipped; pre-commit clean.
2026-07-04 17:15:34 +00:00
SolitaryThinker d04127fe6f [feat] v2 GPU bring-up: 8 ported models verified end-to-end on H100 (+CPU-safe fixes)
Ran each dense public port through the real VideoGenerator on H100 (2-wide across
both GPUs). 8 produce real finite output end-to-end; per-model adapter/loop fixes
landed in each port's OWN files (no shared/fastvideo edits). CPU toy path kept
green (240 passed, 2 skipped) — GPU-only conditioning gated to the cuda backend.

VERIFIED (real GPU output):
  * stable_audio    — stereo audio (2, 441000) @44.1kHz. Fixes: dedicated
                      'conditioner' component kind (SA owns its T5, empty
                      text_encoder_configs) + ConditionerLoader from conditioner/
                      + VDenoiser c_noise = atan(sigma)/(pi/2).
  * matrixgame2     — mp4 (9,256,256,3). Loads CLEAN (not sparse-attn-blocked).
                      Fixes: 20ch cond_concat (4ch mask + 16ch img), mandatory i2v
                      (synth blank first frame), pre-sized kv_cache/crossattn_cache
                      (SDPA inference path, avoids flex_attention compile), bf16
                      autocast, per-request reset_caches.
  * gen3c           — video (9,256,256,3), zero fixes (worked first try).
  * wan_fun_control — video (9,256,256,3).
  * lucy_edit       — video (17,256,256,3).
  * hunyuangamecraft— video (9,256,256,3).
  * hunyuan_video   — video (3,9,256,256), dual LLaMA+CLIP + Hunyuan VAE.
  * hunyuan_video15 — video (9,256,256,3), dual Qwen+ByT5 (gated to cuda; CPU passes
                      single embed).

BLOCKED (not v2 recipe bugs — load/run reached, then a fastvideo/env wall):
  * cosmos25 — DiT + VAE ran finite on GPU; the Qwen2.5-VL Reason1 encoder hits a
               transformers 5.12.1 incompat in fastvideo shared code
               (Qwen2_5_VLConfig.pad_token_id).
  * kandinsky5 — fastvideo registry.py registers it with a bare PipelineConfig (no
                 Kandinsky5 config) -> load fails fastvideo-side. (latent z=16 +
                 visual_cond adapter corrected; toy decoupled to its own channels.)
  * hyworld — fastvideo's hyworld DiT hardcodes flash_attn (no SDPA fallback);
              flash_attn kernel not built here.

CPU-safety fixes (mine): kandinsky5 toy ToyDiT/ToyVAE use the toy LATENT_CHANNELS
(not the real z=16); hunyuan_video15 dual-encoder packing gated to cuda.
Sparse-attn distilled (turbowan/SLA, fastwan/VSA) + gated (cosmos2/flux2/sd35) +
huge (>80GB) handled separately. pre-commit clean per-package.
2026-07-04 17:15:34 +00:00
SolitaryThinker ae0cf3e2d1 [docs] v2: porting status — ALL fastvideo models ported (63/64 by-id; VSA env-blocked)
Rewrites V2_PORTING_STATUS.md to reflect completion: the scope is now ALL
fastvideo models (not Wan+LTX-2 only). Documents the self-contained recipe-package
porting mechanism (ComponentSpec.adapter), the 15 net-new architectures + 5
Wan-family variants newly ported (CPU-verified end-to-end; GPU=BRINGUP), the 7
GPU-verified models, and the single env-blocked id (VSA-14B, needs nvcc).
2026-07-04 17:15:34 +00:00
SolitaryThinker f4af3cf886 [feat] v2 registry: LTX-2/2.3 repo aliases -> by-id resolution 63/64
Adds explicit ModelEntry aliases for the LTX-2 (FastVideo/LTX2-Diffusers,
LTX2-base, Lightricks/LTX-2 -> single-stage base) and LTX-2.3 (LTX2.3-Diffusers,
LTX2.3-Distilled-Diffusers, LTX2.3-base, Lightricks/LTX-2.3, lightricks/ltx-2.3
-> the distilled joint-A/V card) naming variants of already-ported LTX
checkpoints (the arch fallback also resolves LTX2Transformer3DModel from a root).

v2 now resolves 63/64 fastvideo registry ids by exact id; the only remaining id,
FastVideo/Wan2.1-VSA-T2V-14B-720P-Diffusers, is ENV-BLOCKED (VSA Sparse-Linear
Attention kernels require nvcc, not built in this bring-up; it arch-resolves to
the base Wan card but needs the VSA kernel build to run faithfully).

v2 mini 240 passed, 2 skipped; pre-commit clean.
2026-07-04 17:15:34 +00:00
SolitaryThinker 39580ad73a [feat] v2: port the residual Wan-family variants (rCM/DMD/v2v/control/causal-MoE)
Closes the bucket-B sampler/conditioning gap — each reuses the Wan/Causal ARCH
(no new torch adapter) with a new in-package loop/sampler/conditioning, declared
in _BUCKET_C as explicit-HF-id-only (transformer_cls="" so the generic Wan/Causal
arch fallback is NOT hijacked — only the exact id distinguishes the capability
variant from a base Wan of the same class).

  * turbowan      — TurboWan rCM (Reparameterized Consistency Model) few-step: a
                    faithful in-package RCMScheduler port (TrigFlow->RectifiedFlow
                    schedule + stochastic consistency SDE step), 1.3B/14B T2V +
                    TurboWan2.2-I2V-A14B (MoE i2v, boundary 0.9 in raw-sigma space)
  * lucy_edit     — Lucy-Edit v2v editor: a video_vae_encode node (the input video
                    -> 48ch cond latent) threaded via the shared i2v_cond hook ->
                    96ch Lucy DiT input (faithful to denoising.py is_lucy_edit)
  * wan_fun_control — Wan2.1-Fun-Control: control-video conditioning (reuses the
                    i2v [mask|cond] concat pattern)
  * sfwan22       — Self-Forcing Wan2.2-A14B: causal chunk_rollout + Wan2.2 MoE
                    boundary routing, i2v (boundary 0.9) + t2v (boundary 0.875)
  * fastwan       — FastWan DMD 3-step: TI2V-5B-FullAttn loadable; the VSA-trained
                    variants + non-strict to_gate_compress load are BRINGUP

All 12 residual ids resolve+build; base Wan/Causal resolution unchanged (arch
fallback not hijacked). The _BUCKET_C regression test auto-extended -> v2 mini
240 passed, 2 skipped; pre-commit clean (per-package + registry).
2026-07-04 17:15:34 +00:00
SolitaryThinker 521c2845e6 [test] v2: end-to-end CPU regression guard for the bucket-C ports
Data-driven from registry._BUCKET_C (+ cosmos2): each net-new ported arch
resolves through the registry (exact id + arch fallback) AND runs end-to-end on
the CPU toy backend via the public Engine path (resolve -> build card+program ->
load_card -> Engine.run), emitting exactly one modality-correct artifact
(video / image / audio) + latents. Auto-covers future _BUCKET_C rows.

21 tests pass; full v2 mini 232 passed, 2 skipped.
2026-07-04 17:15:34 +00:00
SolitaryThinker 0467edbd07 [feat] v2: port the 14 remaining bucket-C archs as self-contained recipe packages
Completes the bucket-C porting backlog. Each arch is a self-contained recipe
package (card-declared torch adapter via ComponentSpec.adapter + a new/forked
loop + program, NO edit to the shared torch_backend dispatch), following the
cosmos2 reference pattern. One _BUCKET_C table in v2/registry.py drives both the
explicit HF-id registry (PRIMARY) and the select_by_architecture fallback.

Ported (CPU-verified: import + card/program build + registry resolve + denoise
loop runs end-to-end on the CPU toy backend; GPU load/run is BRINGUP):
  * cosmos25       — Cosmos-Predict2.5 (flow-match, per-frame plain-sigma timestep;
                     reuse FLOW_MATCH_STEP; Reason1/Qwen2.5-VL encoder adapter)
  * hunyuan_video  — HunyuanVideo (reuses WanDenoiseLoop; dual LLaMA+CLIP encoders;
                     Hunyuan VAE scaling_factor) + FastHunyuan variant
  * hunyuan_video15 — HunyuanVideo 1.5 (480p/720p cards)
  * longcat        — LongCat-Video T2V/I2V/VC
  * sd35           — SD3.5 MMDiT (flow-match, image; triple-encoder joint embed +
                     pooled_projections)
  * gen3c          — GEN3C (EDM; 82ch pose-buffer DiT; camera/depth -> BRINGUP)
  * kandinsky5     — Kandinsky 5.0 T2V Lite
  * flux2          — FLUX.2 dev/klein (MMDiT, image; gated weights -> BRINGUP)
  * stable_audio   — Stable Audio Open (audio modality)
  * hunyuangamecraft, hyworld, lingbotworld, matrixgame2, matrixgame3 — interactive
                     world models; t2v/degenerate path CPU-verified, action/camera/
                     memory conditioning is BRINGUP (needs request-API extension)

Adapters declared via ComponentSpec.adapter (the ac29750b enabler) so each port
adds only NEW files (recipe package + per-arch torch_<arch>.py + optional facade
stub) — zero shared-file edits. Registry resolves all 31 bucket-C HF ids by exact
id + 14 architecture fallbacks; no regression (cosmos2/wan/ltx2 unchanged).
v2 mini green (211 passed, 2 skipped); pre-commit clean (per-file/registry).
2026-07-04 17:15:34 +00:00
SolitaryThinker bfbd90ea3a [feat] v2: Cosmos-Predict2-2B-Video2World port (EDM-Karras) — reference bucket-C recipe
First net-new architecture ported via the self-contained recipe-package pattern
(card-declared adapter + new loop, no shared-dispatch edit):

* CosmosDenoiseLoop (v2/recipes/cosmos2/loop.py): EDM preconditioning folded into
  a flow-match Euler integrator. Faithful port of CosmosDenoisingStage — Karras
  sigma schedule (rho=7, sigma_max=80 -> sigma_min=0.002, terminal clamp), latent
  init randn*sigma_max, per-step c_in/c_skip/c_out (sigma_data=1) -> x0, CFG in x0
  space, x0 -> velocity (x-x0)/sigma, FLOW_MATCH_STEP. video2world frame-replace
  conditioning threaded but inert for the t2v preset.
* build_karras_sigmas helper added to v2/loop/sampler.py.
* CosmosDiT + CosmosT5Encoder adapters (v2/platform/backends/torch_cosmos.py),
  declared on the card via ComponentSpec.adapter (the ac29750b enabler) — DiT
  returns raw EDM output + builds the mandatory zero condition/padding masks + fps;
  T5 uses the raw last_hidden_state (no Wan zero-pad). Reuses the WanVAE adapter.
* card/program/registry (HF id nvidia/Cosmos-Predict2-2B-Video2World + arch
  fallback on CosmosTransformer3DModel) + COSMOS_NEG prompt + SamplingDefaults
  (35 steps, gs 7, 704x1280, 93f, 16fps).

CPU-verified: Karras schedule, card/program build, registry resolve (id + arch),
EDM loop runs end-to-end on the CPU toy backend. GPU load/run is BRINGUP.
v2 mini green (211 passed, 2 skipped); pre-commit clean.
2026-07-04 17:15:34 +00:00
SolitaryThinker 5fd6e23e30 [feat] v2 torch backend: ComponentSpec.adapter — card-declared per-arch TorchComponent
A new architecture can declare its own torch adapter on the card
(ComponentSpec.adapter="module:Class") instead of editing the shared _make_dit/
_make_vae/_make_text_encoder dispatch. _explicit_adapter() constructs it as
cls(module, *extra, device=, dtype=) and short-circuits the built-in Wan/LTX2
class-name dispatch when set. This makes each bucket-C port a self-contained
recipe package (card + adapter module + loop + program) with no shared-file edit
-> conflict-free parallel porting. Unset -> unchanged built-in dispatch.

CPU mini green (211 passed, 2 skipped).
2026-07-04 17:15:34 +00:00
Will Lin fc3332550c [feat] v2: Wan2.2-I2V-A14B (MoE i2v) — combine boundary-routed experts + i2v conditioning
Reuses everything: 2 WanTransformer3DModel experts + BoundaryTimestepRouting (from the A14B MoE pattern),
the CLIP image encoder + first-frame [mask|cond] conditioning + the i2v program (from the Fun-InP i2v
port), and the shared WanDenoiseLoop (i2v hooks + the boundary expert). No new adapter. CPU-verified: the
toy MoE i2v runs end-to-end (2 experts + boundary + conditioning -> finite video); resolves with i2v caps.
Structural (GPU-pending: 2x14B, like the A14B T2V). Wan family now largely covered (T2V 1.3B/14B/TI2V-5B/
A14B, causal SF, i2v 1.3B/14B/A14B). CPU mini 211/2.
2026-07-04 17:15:34 +00:00
Will Lin 7464ef8308 [feat] v2: register Wan2.1-I2V-14B 480P/720P (reuse the GPU-verified i2v card)
The 14B i2v variants reuse the Wan2.1 i2v card/path proven on Fun-1.3B-InP — just per-variant params
(480P flow_shift 3.0 / 480x832, 720P flow_shift 5.0 / 720x1280). Registry resolution + caps verified;
specific 14B weights GPU-pending (same generic Wan i2v loader path that Fun-InP validated). i2v cluster
now supported; roadmap updated (11 models ported).
2026-07-04 17:15:34 +00:00
Will Lin 4a56274d10 [feat] v2: Wan2.1 i2v port (Wan2.1-Fun-1.3B-InP) — CLIP encoder + first-frame conditioning, GPU-verified
Real image-to-video, unlocking the i2v cluster. v2/recipes/wan21/i2v.py: CLIP image-encode -> the DiT's
encoder_hidden_states_image; first-frame VAE conditioning + a 4-channel mask -> the 20ch [mask|cond] that
the Wan adapter concatenates with the 16ch noise -> the 36ch i2v DiT input (mirrors fastvideo's
ImageEncodingStage + ImageVAEEncodingStage; v2's WanVAE.encode already applies the matching (z-mean)/std).
Reuses the shared WanDenoiseLoop (its None-default i2v hooks) + the Wan torch adapter unchanged. Adds
ToyImageEncoder + the image_encoder checkpoint subfolder stamp. Registered Wan2.1-Fun-1.3B-InP.

GPU-verified: loads via the generic Wan loader (1.56B, no param-mapping issue), runs the full i2v
conditioning, produces real video (3,9,256,384, std 0.44, finite, motion 0.041). CPU mini 211/2 (T2V
unregressed). BRINGUP: visual confirmation that the output follows the conditioning image is
human-in-the-loop.
2026-07-04 17:15:34 +00:00
Will Lin 9ede9af123 [feat] v2 Wan loop+adapter: optional i2v conditioning hooks (cond concat + CLIP context); T2V unchanged
Threads i2v conditioning through the SHARED WanDenoiseLoop with zero T2V risk: init() reads optional
slots i2v_cond (the [mask|cond] latent) + i2v_img_embeds (CLIP) into scratch; _velocity passes them to
the dit (context=, cond=); WanDiT concats cond (16->36ch) and uses the embeds as
encoder_hidden_states_image; capture is disabled only when i2v conditioning is present. For T2V both are
None -> the dit call, CFG, and cudagraph capture are byte-identical (CPU mini 211 pass, 2 skip — no
regression). ToyDiT accepts+ignores cond (image-conditioning is a GPU-path concern). Completes the i2v
backend seam; the program's mask+cond construction + the Wan i2v card/registry + GPU verify follow.
2026-07-04 17:15:34 +00:00
Will Lin 3d405f6dfa [feat] v2 torch backend: CLIP image-encoder adapter + image_encoder component kind (i2v groundwork)
Adds CLIPImageEncoder (encode_image -> the DiT's encoder_hidden_states_image) + the generic builder's
image_encoder maker (ImageEncoderLoader + ImageProcessorLoader), registered as the cuda 'image_encoder'
kind. Mirrors fastvideo's ImageEncodingStage. CPU-verified (component-kinds + lazy invariant, 211/2);
GPU path marked BRINGUP/written-not-run (processor subfolder + dtype to confirm on a real i2v checkpoint),
matching how the rest of the torch backend was originally landed. Reusable by the Wan i2v cluster + many
bucket-C models (Hunyuan/Cosmos i2v). Next i2v increments: the mask+cond latent construction + the
concat-into-DiT-input loop, then the card + registry + GPU-verify with Wan2.1-Fun-1.3B-InP.
2026-07-04 17:15:34 +00:00
Will Lin d891771ba3 [revert] v2: drop FastWan/VSA registry entries — generic Wan loader can't map their gated-attn params
GPU verification (Wan2.1-Fun... no: FastWan2.1-T2V-1.3B) failed at load: 'Parameter blocks.0.to_gate_compress.bias
not found in custom model state dict' — the FastVideo/* DMD-distilled checkpoints carry gated-attention
params (to_gate_compress) that the generic WanTransformer3DModel loader can't map. v2/registry.py's
select_by_architecture ALREADY rejects WanDMDPipeline for exactly this reason; my explicit ModelEntry
wrongly bypassed it. Reverted FastWan (1.3B + 14B-480P) and the unverified VSA-14B alias (same FastVideo/*
risk). Kept the official Wan2.1-T2V-14B (standard weights, same loader path as the GPU-verified 1.3B).

Lesson recorded in V2_PORTING_STATUS.md: FastWan/Turbo/VSA need a param-mapping fix (like LTX-2.3 did),
not just a schedule — bucket-C-effort. GPU-verify every port before claiming support. CPU mini 211/2.
2026-07-04 17:15:34 +00:00
Will Lin 67dea39052 [feat] v2: alias FastVideo/Wan2.1-VSA-T2V-14B-720P to the Wan-14B card (bucket B)
Same WanTransformer3DModel arch (VSA is an attention-backend choice, not a weight/arch difference); v2
runs dense TORCH_SDPA, so it resolves to build_wan_t2v_14b_card. Registry resolution verified; the GPU
forward path is the 1.3B-proven Wan adapter (specific 14B/VSA weights not separately GPU-run).
2026-07-04 17:15:34 +00:00
Will Lin 5f1d2ef7d2 [feat] v2: port Wan2.1-T2V-14B (bucket B) + document the all-models backlog
First bucket-B port toward 'support every fastvideo model': Wan2.1-T2V-14B reuses the Wan recipe +
torch adapter unchanged (same WanTransformer3DModel/AutoencoderKLWan/UMT5) — only a registry entry +
build_wan_t2v_14b_card (720p, flow_shift 5.0) + SamplingDefaults differ. Without the entry the arch
fallback would give it the 1.3B 480p defaults; the explicit entry gives 50 steps / 720x1280.

Also recorded the full backlog in V2_PORTING_STATUS.md: 63 fastvideo models = 8 ported / 21 bucket-B
(reuse Wan/Causal/LTX2 arch — registry+recipe+defaults, no new adapter) / 34 bucket-C (13 new
architectures needing a TorchComponent adapter). Updated the stale 'how to add a model' steps to the
post-redesign structure (v2/recipes/, torch_backend.py, SamplingDefaults). CPU mini 211 pass, 2 skip.
2026-07-04 17:15:34 +00:00
Will Lin 440b99523e [refactor] v2 torch backend: TorchComponent base + one generic builder + v2.* facade (Phase 1b/1c)
Addresses the adapter-setup pains: collapses the torch_cuda(trampolines)/torch_adapters/torch_ltx2 split
+ 11 near-identical adapter classes + 6 build_torch_* builders into:
- v2/platform/backends/torch_backend.py: a TorchComponent base centralizing .to/.eval, the numpy<->torch
  marshalling (ONE place), the set_forward_context wrap, and the weight surface; thin per-model subclasses
  (WanDiT/LTX2DiT/WanVAE/LTX2VAE/T5Encoder/Gemma/LTX2Upsampler/LTX2AudioVAE/LTX2Vocoder) carrying only
  forward semantics; and ONE build_component(spec) dispatching by spec.kind via _MAKERS.
- torch_cuda.py: registers that single generic builder for all 6 cuda kinds (no per-kind trampolines).
- v2 owns its namespace via re-export STUBS (facade, marked '# STUB'): v2/forward_context, v2/fastvideo_args,
  v2/distributed, v2/loader (the load_component seam), v2/api, v2/models/{dits,audio,upsamplers}/*. All v2
  code imports v2.*; 'from fastvideo' now lives ONLY in those 8 stub files -> a future per-module vendored
  cutover swaps a stub body, no caller changes. No divergence (stubs run fastvideo's live code).
- Deleted torch_adapters.py + torch_ltx2.py.

Verified: CPU mini 210 pass/2 skip; lazy invariant (platform load imports no torch); GPU bit-parity LTX-2.3
T2VS (audio std 0.04304, identical to pre-redesign) + Wan2.1 (video std 31.94, motion 5.997).
2026-07-04 17:15:34 +00:00
Will Lin 1cb2b4e84c [feat] v2: per-model sampling defaults on ModelCard (Phase 1a)
v2 had no per-model defaults — generate_video hardcoded 30 steps/25 frames/480x832/cfg5/16fps for
every model, badly wrong for e.g. LTX-2 distilled (wants 8 steps @1024x1536) or Wan2.2-TI2V (704x1280@24fps).

- New SamplingDefaults dataclass + ModelCard.sampling_defaults (v2/card/specs.py), exported from v2.card.
- Populated all 7 supported cards from fastvideo's InferencePreset defaults (steps/guidance/HxW/frames/fps
  + per-modality guidance for LTX-2.3 A/V). Negative prompts copied verbatim into v2/recipes/_prompts.py
  (recipe DATA, not model code -> v2-owned, no fastvideo import).
- VideoGenerator stores the resolved card; generate_video applies card defaults with precedence
  kwargs > SamplingParam > card > generic fallback (pure _resolve_default helper, unit-tested).
- test_sampling_defaults.py: per-card values + precedence (incl. empty-neg-prompt edge). CPU mini 210 pass, 2 skip.
2026-07-04 17:15:34 +00:00
Will Lin 51898f48e9 [refactor] v2: rename models/ (recipe layer) -> recipes/; move toy backend -> platform/backends/toy.py
Frees v2/models/ to become the vendored-architecture namespace that mirrors fastvideo/models
(part of making v2 self-contained / able to replace fastvideo). The v2 recipe layer (per-family
card.py/loop.py/program.py + common.py + the build_*_engine re-exports) is the recipe, not the
architectures, so it moves to v2/recipes/. The pure-numpy toy/parity implementations (ToyDiT etc.)
move from v2/models/backend.py to v2/platform/backends/toy.py (alongside cpu.py/accel.py/torch_*).

Mechanical: all imports are absolute, so v2.models.<x> -> v2.recipes.<x> and v2.models.backend ->
v2.platform.backends.toy across v2/ + examples/ (89 files, 178 refs). No behavior change.
CPU mini green (202 passed, 2 skipped); all v2 files compile.
2026-07-04 17:15:34 +00:00
SolitaryThinker 4e331eb7ff [refactor] v2: use absolute imports (v2.*) everywhere instead of relative
Mechanical conversion of every relative import under v2/ to an absolute v2.* path
(from .x / ..x / ...x -> from v2.<pkg>.x) so imports are unambiguous, grep-able, and
stable when code is copied/moved between entrypoints (VideoGenerator, CLI, server).

Surgical prefix-only rewrite: only the 'from <dots><module>' prefix changed — import
names, parentheses, multi-line formatting, comments, and ordering are byte-for-byte
preserved (no collapsing, no reorder, no unrelated reformatting).

- 471 imports across 125 files; v2/tests/ was already absolute (untouched).
- Validated: all 125 files compile, every 'from v2.* import' target resolves to a real
  module/package, zero relative imports remain (full sweep), CPU mini suite green
  (202 passed, 2 skipped).
2026-07-04 17:15:34 +00:00
SolitaryThinker c6d2976fc2 [feat] v2 VideoGenerator: A/V convenience path (generate_video -> T2VS -> mp4 + 24kHz wav)
Makes the 'Full A/V' LTX-2.3 deliverable reachable from the user-facing entrypoint, not just the
engine. A model advertising TEXT_TO_VIDEO_SOUND (LTX-2.3) now auto-issues a T2VS request, so generate()
/ generate_video() return BOTH modalities in one joint pass:
- VideoGenerator stores the resident instance + a supports_av flag (from card.capabilities); generate()
  gains want_audio (None=auto-by-capability, True/False to force) and routes T2V vs T2VS+{video,audio}.
- _result saves the stereo waveform as a sibling .wav at the vocoder's REAL rate (24000) — read off the
  built audio_vae adapter (TorchLTX2AudioVAE.sample_rate = Vocoder.output_sample_rate), since the
  AudioArtifact default rate is a placeholder. Populates GenerationResult.audio/.audio_sample_rate and
  extra['audio_path']. scipy IEEE-float WAV; [channels,samples] auto-transposed.
- Rewrote v2_basic_ltx2_3_distilled.py: registry routes to build_ltx2_3_card (its own joint T2VS A/V
  card, not the LTX-2 base/2-stage card); the example prints both the mp4 and the wav.
- GPU-verified via the convenience API: ev.mp4 + ev.wav (24000 Hz, stereo 61920x2, nonzero, std 0.043).
  CPU mini green (202 passed, 2 skipped); engine/program/toy paths untouched (test_ltx2_av pins 44100).
2026-07-04 17:15:34 +00:00
SolitaryThinker 6096b00aeb [feat] v2 LTX-2.3 T2VS GPU-verified: audio VAE/vocoder wiring + dual-connector audio fix
The full joint text->video+audio LTX-2.3 path now generates on the real 18.99B model:
- GPU audio components: build_torch_audio_vae (AudioDecoderLoader -> LTX2AudioDecoder, chains the
  vocoder) + build_torch_vocoder (VocoderLoader -> LTX2Vocoder); registered the 'audio_vae'/'vocoder'
  cuda component kinds; stamped their checkpoint subfolders (_WAN21_SUBFOLDERS).
- Fix: TorchGemma.encode_av must pass output_hidden_states=True — the 2.3 connector's SEPARATE audio
  projection lives in hidden_states[0] only then (gemma.py:703); without it the audio text fell back to
  the video embedding (4096 vs 2048 -> audio cross-attn shape mismatch).
- GPU-verified T2VS: video (3,33,256,384, std 0.68) + audio (stereo 2x61920 @24kHz, nonzero, std 0.059).
- test_torch_backend: cuda component kinds now include audio_vae + vocoder. CPU suite green (202+2).
2026-07-04 17:15:34 +00:00
SolitaryThinker 1d23399d81 [feat] v2 LTX-2.3 T2VS: single-stage joint audio+video card/loop/program (CPU-verified)
Makes LTX-2.3 a first-class, faithful card (was wrongly merged into the single-stage base):
- LTX23DenoiseLoop (loop.py): single-pass joint A/V denoise — one DiT forward per step cross-attends
  video<->audio via the adapter's (v_vel,a_vel) return; full-res video latent + a [8,T,16] audio latent;
  distilled few-step schedule (BASE_SIGMAS). Video-only when no audio requested.
- build_ltx2_3_card (model_id 'ltx2.3-distilled' — the name now correctly names the REAL 2.3): 5
  components incl. audio_vae (AudioDecoder) + vocoder (required_for t2vs, optional_for t2v); caps
  T2V + T2VS. build_ltx2_3_program: dual-connector text-encode -> joint denoise -> video + audio decode.
- registry.py routes FastVideo/LTX-2.3-Distilled-Diffusers -> this card (split from the base entry).
- Toy support: ToyTextEncoder.encode_av (separate video/audio text), ToyDiT joint A/V (audio now a
  keyword-only arg so positional  callers like the talker are unaffected), channel-agnostic
  ToyAudioVAE (np.resize identity for the existing 2-stage T2VS).
- CPU-verified: toy T2VS -> video+audio (8 steps), T2V -> video-only; CPU suite green (202+2).
GPU audio-VAE/vocoder loaders + the real-T2VS GPU verify are the next step.
2026-07-04 17:15:34 +00:00
SolitaryThinker fefcd415ff [feat] v2 LTX-2 adapters: A/V foundation (joint DiT forward + dual text connector + audio decode)
Foundation for the LTX-2.3 T2VS port (card/loop/program wiring + GPU verify to follow):
- TorchLTX2DiT.__call__ gains an optional joint audio path: pass audio_latent[8,T,16] + audio_text and it
  feeds audio_hidden_states/audio_encoder_hidden_states/audio_timestep/audio_sigma in ONE forward
  (LTX-2.3 cross-attends video<->audio) and returns (video_velocity, audio_velocity). Video-only call is
  byte-for-byte unchanged (audio_latent=None).
- TorchGemma.encode_av returns the SEPARATE (video_text, audio_text) projections from the 2.3 connector
  (video=last_hidden_state, audio=hidden_states[0]); 2.0 returns them equal.
- TorchLTX2AudioVAE (AudioDecoder -> Vocoder -> waveform@24kHz) + TorchLTX2Vocoder wrapper.
Additive + backward-compatible; CPU suite unaffected (adapters are GPU-lazy).
2026-07-04 17:15:34 +00:00
SolitaryThinker 7ac2ff0d1c [refactor] v2: shared model registry (HF-id primary + arch fallback) for all entrypoints
Per review: dispatch should be a directly-mapped HF-string -> card registry (like fastvideo), shared
by every entrypoint (VideoGenerator + a future CLI / server), not buried in the generator.

- New v2/registry.py mirrors fastvideo's fastvideo/registry.py hybrid resolution: (1) exact HF repo id
  in an explicit ModelEntry registry [PRIMARY — correct per-model card/capabilities, and the only way to
  split same-architecture capability variants like Wan2.1 T2V vs the i2v 'InP' 1.3B], (2) short repo-name
  match, (3) architecture inference [FALLBACK — local paths / unregistered repos]. resolve(model_path[,
  root]) is the single shared entry point.
- video_generator.py: moved _read_arch_signature/_select_builders into the registry; from_config now
  calls resolve() — registered ids resolve with no config read, else arch inference on a cheap *.json
  snapshot. Reconciles the earlier 'no brittle table' refactor with the 'map hf string -> card' ask: one
  clean registry + a fallback, not three coupled structures.
- Verified: registry resolves exact-id / short-name / arch-fallback / unregistered correctly; CPU suite
  green (202+2); wan21 GPU smoke generates via the new resolve path.
2026-07-04 17:15:34 +00:00
SolitaryThinker fa9c58b419 [fix] v2: name LTX-2 cards by architecture (2-stage vs single-stage) + Wan2.1 is T2V-only
Addresses the 'how is ltx2 separate from ltx2.3' confusion + a wrong capability:

- LTX-2 cards renamed by ARCHITECTURE (the version labels did not map to it): build_ltx2_card model_id
  'ltx2.3-distilled' -> 'ltx2-2stage-distilled' (two-stage base->upsample->refine; serves the
  upsampler-having FastVideo/LTX2-Distilled-Diffusers); build_ltx2_base_card 'ltx2.base' ->
  'ltx2-single-stage' (one loop; serves Davids048 base + the single-stage FastVideo/LTX-2.3-Distilled,
  which has NO spatial_upsampler). Dispatch already splits on has_spatial_upsampler. Updated the
  model-id refs in the mini's tests/examples.
- Wan2.1 base is T2V-only: dropped the wrong Capability.IMAGE_TO_VIDEO + narrowed components'
  required_for to {t2v} (i2v is the separate InP variant; v2 has no i2v path yet). build_wan21_card is
  shared by wan21 + wan2.2-ti2v; the A14B card was already T2V-only.
- CPU suite green (202 passed, 2 skipped).
2026-07-04 17:15:34 +00:00
SolitaryThinker 0662b42510 [docs] v2: LTX-2 base/2.3 GPU-verified on rebuilt x86 stack + remaining-port mechanisms
- LTX-2 base (Davids048) and LTX-2.3-Distilled both generate real video (inter-frame motion 4.5 / 6.6)
  via the single-stage base card — moved to Working (7 models now verified).
- Environment: the aarch64 venv was rebuilt for x86 (torch 2.11.0+cu128) + re-validated (CPU suite green,
  wan21 + LTX-2 base/2.3 generate).
- Remaining Wan+LTX-2 ports documented with concrete mechanisms: Wan2.2-i2v (SigLIP image_encoder +
  VAE-encode first-frame + concat-mask -> larger-in_channels i2v DiT), TurboWan (RCMScheduler consistency
  loop), Lucy-Edit (Wan v2v via VideoVAEEncodingStage), FastWan (VSA, env-blocked: no nvcc).
2026-07-04 17:15:34 +00:00
SolitaryThinker ae6d8085de [docs] v2: LTX-2.3 example + roadmap (A14B offload working, base/2.3 ported, env status)
- v2_basic_ltx2_3_distilled.py: LTX-2.3-Distilled routes to the single-stage base card (no
  spatial_upsampler) via the arch dispatch; pass few steps for the distilled schedule.
- V2_PORTING_STATUS.md: A14B moved to Working (CPU expert offload, 60GB peak); LTX-2 base + 2.3 added
  (code-complete, GPU re-verify pending); Environment-status note on the mid-session aarch64->x86 host
  reschedule that blocks GPU re-verify.
2026-07-04 17:15:34 +00:00
SolitaryThinker d0648ba8d9 [feat] v2: Wan2.2-A14B MoE CPU offload (fits 1 GPU) + LTX-2 base/2.3 single-stage port
Within the bounded Wan+LTX-2 scope:

- Wan2.2-A14B MoE now GENERATES on a single 80GB GPU via CPU offload (TorchWanDiT offload_group):
  the two 14B experts live on CPU and only the active one is swapped onto the GPU at the boundary-
  timestep transition (a single swap, not per-step). GPU-verified: 60GB peak (vs 79GB OOM), produced
  wan22_a14b_lion.mp4 (17x480x832, std 54.7, motion 8.36). Single-expert Wan stays resident.

- LTX-2 base (single-stage) port: build_ltx2_base_card + build_ltx2_base_program reuse the LTX-2
  adapters at FULL latent res with a request-driven many-step flow-match (LTX2DenoiseLoop full_res/
  request_steps/base_flow_sigmas; distilled base/refine path preserved via False defaults). The SAME
  single-stage card serves LTX-2.3-Distilled (also single-stage: no spatial_upsampler) — dispatched by
  the new has_spatial_upsampler discriminator in _select_builders. v2_basic_ltx2.py added; VideoGenerator
  gains shutdown() for API parity.

- Fixes: from_config 'os' scoping (shadowed module import); LTX-2 upsampler per_channel_statistics
  source (the AE's .decoder, not the top-level module).

Verification status: A14B offload, the upsampler, and the arch-dispatch refactor were GPU-verified
earlier this session. LTX-2 base/2.3 are CPU-verified (cards/programs build, dispatch routes, schedule
correct); their GPU smoke tests were pending when the box was rescheduled aarch64->x86 mid-session,
which broke the aarch64 venv (numpy/torch unrunnable on x86) — GPU re-verify blocked on the env.
2026-07-04 17:15:34 +00:00
SolitaryThinker e10828346f [feat] v2: architecture-driven dispatch + real LTX-2 upsampler + Wan2.2-A14B MoE card
Two reviewer asks + the next Wan port, all within the bounded Wan+LTX-2 scope:

1. Architecture-driven dispatch (replaces the HF-id table + substring fallback): from_config reads the
   checkpoint's pipeline/transformer/VAE class names (+ z_dim, transformer_2) and picks the v2 card via
   _select_builders — mirroring fastvideo's get_pipeline_config_cls_from_name. Resolves local paths /
   renamed repos / new distilled variants of a known arch with no table edits, and cleanly REJECTS
   FastWan (detected by WanDMDPipeline) with a precise message instead of a confusing load crash.

2. Real LTX-2 spatial upsampler (was a nearest-neighbor np.repeat stand-in): new 'upsampler' component
   kind -> TorchLTX2Upsampler wraps the real LTX2LatentUpsampler and applies the repo's upsample_video
   (un_normalize via the VAE decoder's per_channel_statistics -> learned 2x upsample -> normalize). CPU
   keeps ToyUpsampler (np.repeat) via the factory terminal, so the program calls
   component('spatial_upsampler').upsample(...) on both backends with no device branch. GPU-verified:
   9x512x768, std 70.5, motion 6.51.

3. Wan2.2-T2V-A14B MoE card (build_wan22_a14b_card): two WanTransformer3DModel experts +
   BoundaryTimestepRouting @0.875. GPU-verified that both experts denoise; OOMs in VAE decode on one
   80GB GPU (~70GB resident) — upstream offloads the DiT for MoE; documented as offload-blocked.

CPU suite green (202 passed, 2 skipped). FastWan root-caused (non-strict load of VSA gate_compress +
VSA not built); roadmap (V2_PORTING_STATUS.md) updated with the bounded scope + per-model status.
2026-07-04 17:15:34 +00:00
SolitaryThinker 655f362cf4 [feat] v2 port: Wan2.2-TI2V-5B (T2V) — 4th GPU-verified model
- Wan2.2-TI2V-5B reuses the Wan adapters (WanTransformer3DModel / AutoencoderKLWan / UMT5); deltas
  are the higher-compression VAE geometry (z_dim=48, 16x spatial, 4x temporal) and 480p flow-shift 5.0.
  The DiT forward accepts a scalar timestep (1D path), so no per-frame expand_timesteps for pure t2v.
- WanDenoiseLoop / build_wan21_card gain optional geometry params (latent_channels/spatial_ratio/
  temporal_ratio) defaulting to Wan2.1 (16/8/4) -> wan21 path unchanged; build_wan22_ti2v_card sets
  48/16/4. Registered as family 'wan2.2-ti2v' in VideoGenerator; v2_basic_wan2_2_ti2v.py added (T2V).
- Verified: 25x448x768 mp4, std 62.3, inter-frame motion 4.89 (coherent). CPU suite green (202+2skip).
- Corrected the now-disproven FastWan='wan21 reuse' mapping (its DMD checkpoint to_gate_compress param
  mapping differs); roadmap updated (TI2V-5B working; A14B MoE + I2V remain).
2026-07-04 17:15:34 +00:00
SolitaryThinker 3541e81d66 [feat] v2 VideoGenerator: convenience API (from_pretrained/generate_video) + porting roadmap
- VideoGenerator gains the convenience surface most basic examples use: from_pretrained(model,
  num_gpus/*_cpu_offload/...) and generate_video(prompt, sampling_param=, **kwargs) -> result, on top
  of the typed from_config/generate. Accepts SamplingParam.
- v2 examples matching the upstream convenience-API examples for the verified models: v2_basic.py
  (Wan2.1) and v2_basic_self_forcing_causal.py (SF-causal).
- V2_PORTING_STATUS.md: honest per-family roadmap. Working: wan21, wan_causal, ltx2-distilled. Each
  further model needs per-model work (FastWan: WanDMD to_gate_compress param mapping; TurboWan: RCM
  consistency sampler; Wan2.2: MoE card; LTX2 base/i2v; new families: new cards/adapters; gated Flux2 /
  local GEN3C / audio StableAudio / interactive MatrixGame blocked in this env).
2026-07-04 17:15:34 +00:00
SolitaryThinker f51497ee6d [feat] v2 VideoGenerator: typed fastvideo.api entrypoint over the v2 engine
Mirrors fastvideo.entrypoints.VideoGenerator (from_config(GeneratorConfig) -> generate(
GenerationRequest) -> GenerationResult.video_path), reusing the OFFICIAL fastvideo.api config classes
so a basic_dmd_new_api.py-style script differs only by importing VideoGenerator from v2.
- model_path -> v2 card registry (Wan2.1 / FastWan -> wan21; SFWan -> wan_causal; LTX2 -> ltx2);
  snapshot_download + stamp_wan21_checkpoints + Engine(cuda) + program; generate maps SamplingConfig
  -> DiffusionParams -> make_request -> eng.run, saves the [C,T,H,W] decode as an mp4.
- Lazy v2.__getattr__ keeps 'import v2' torch-free (verified) so the CPU mini stays green (202+2).
- examples/inference/basic/v2_basic_new_api.py runs all three GPU models through this API.
- Single-GPU, resident, TORCH_SDPA (EngineConfig offload/num_gpus>1/VSA accepted for parity, not applied).
verified: wan21 from_config->generate->mp4 (frames (5,256,256,3) uint8, video_path written).
2026-07-04 17:15:34 +00:00
SolitaryThinker d8af2e60d2 [feat] v2 ltx2: two-stage distilled GPU bring-up (LTX2Transformer3DModel 18.88B)
Official FastVideo/LTX2-Distilled-Diffusers. New torch_ltx2.py adapters (build_torch_* dispatch on
class):
- TorchLTX2DiT: patchify-internal; per-token timestep ones(B,tok,1)*sigma (sigma direct) + per-sample
  video_sigma; DiT predicts x0 so the adapter returns velocity=(x_t-x0)/sigma for the v2 flow-match step.
- TorchLTX2VAE: CausalVideoAutoencoder decode (internal per-channel un_normalize).
- TorchGemma: LTX2GemmaTextEncoderModel (Gemma + feature-extractor + connectors) -> last_hidden_state.
- ltx2 loop: real 128-ch latent geometry on cuda (32x spatial / 8x temporal; half-res base, 2x upsample).
e2e two-stage (8+3 steps) -> coherent, high-quality video (surfers at sunset), (3,9,512,768). NOTE:
still uses the v2 program's np.repeat upsampler between stages (the refine regenerates from noise so
output is faithful-quality); real LTX2LatentUpsampler swap-in is a follow-up.
2026-07-04 17:15:34 +00:00
SolitaryThinker f79919ba8a [feat] v2 wan_causal: causal DiT (CausalWanTransformer3DModel) GPU bring-up
Official SF checkpoint wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers (reuses TorchWanVAE + TorchT5Encoder).
- TorchWanDiT detects the causal transformer: ignores the chunk_rollout loop's latent `context` (the
  real model conditions across chunks via an internal kv_cache, not a forward arg) -> dispatches to
  full-attention _forward_train, and passes a per-latent-frame timestep [B, num_frames] (the causal
  block asserts a per-frame temb), uniform per chunk.
- chunk_rollout real geometry on cuda (16ch; chunk_size latent frames; 8x spatial).
e2e: chunk_rollout over the SF student -> coherent video (cat in a garden), (3,21,480,832). Fidelity
gap (artifacts): the v2 loop's per-chunk few-step sampling != the official kv-cache streaming + SF
schedule (a follow-up).
2026-07-04 17:15:34 +00:00
SolitaryThinker 1594f8e6be [fix] v2 tests: make no-torch-import guards GPU-aware (skipif torch installed)
The cuda-availability probe imports torch by design; the no-torch-import invariant is only
verifiable when torch is absent. skipif torch installed -> green on GPU box (202 passed, 2 skipped),
still enforced in torchless CI.
2026-07-04 17:15:34 +00:00
SolitaryThinker 64cadaa0bf [feat] v2 wan21: backend-aware latent geometry + checkpoint stamping
- latent_shape(req, model): real Wan geometry (16ch; (T-1)//4+1, H/8, W/8) on the cuda backend;
  toy stand-in stays on accel/cpu.
- stamp_wan21_checkpoints(card, model_root): map a root (local dir or HF id) onto the 3 components'
  ComponentSpec.checkpoint; build_wan21_card(checkpoint_root=...) optional.
2026-07-04 17:15:34 +00:00
SolitaryThinker 750fc1245b [feat] v2 cuda backend: real fastvideo construction + risk A-E fixes (Wan2.1 verified on H100)
Take the written-not-run torch adapters to runs-and-generates on 1x H100 (aarch64):
- A: FastVideoArgs.from_kwargs(model_path=root) builds the real pipeline_config; single-GPU dist
  init; load each component from its subfolder; tokenizer from the sibling <root>/tokenizer.
- B/C: DiT forward wrapped in set_forward_context(attn_metadata=None) (SDPA dense path);
  timestep=sigma*1000 + bare-velocity output confirmed.
- D: VAE decode denormalizes z*std+mean; removed the double-mean (it re-added
  shift_factor==latents_mean) that washed out the video.
- E: UMT5 from config; text embeds zero-padded to text_len (Wan t5_postprocess_text) - the fix
  that took output from a dark blur to a coherent prompt-matching scene.
- Components run at native precision (DiT bf16, VAE/text fp32); checkpoint check before dist init.
2026-07-04 17:15:34 +00:00
SolitaryThinker 7725998b0c [docs] add v2/HANDOFF.md for GPU-side bring-up of the torch backend
Orientation + process doc for an agent on a GPU branch: the 6 commits already
landed, the files to touch, the gating tasks (Risk A FastVideoArgs/checkpoint),
the verification bar (CPU suite stays 204; GPU generation matches a reference via
the SSIM harness), commit/push rules (no Claude co-author; don't rewrite history;
wandb token referenced not embedded), and the gotchas. Points to
GPU_BRINGUP.md for the detailed checklist + risk table.
2026-07-04 17:15:34 +00:00
SolitaryThinker 4b61dedc43 [fix] correct GPU adapters against real fastvideo API (cross-check findings)
Adversarial cross-check of the written-not-run torch adapters against the real
fastvideo source confirmed the interface contracts (DiT returns bare velocity
tensor; timestep=sigma*1000; encode().mode() + bare decode; .last_hidden_state;
no fused solver kernel) but caught a wrong construction layer. Fixed in code:

- Construction: WanTransformer3DModel / AutoencoderKLWan have NO from_pretrained.
  Replace it with the real FastVideo loaders (TransformerLoader / VAELoader /
  TextEncoderLoader + TokenizerLoader, each load(model_path, fastvideo_args)).
  The loader resolves the class from the checkpoint config — so UMT5-vs-T5 is
  chosen correctly instead of hardcoded (was BLOCKER #1/#4/#5).
- Text encoder: wrap the forward in set_forward_context(...) — the (U)MT5
  attention reads global state via get_forward_context(); a bare call mis-encodes
  (was BLOCKER #2). Drop the wrong padding="max_length".
- VAE: apply latent normalization the DiT expects — (z-mean)*inv_std on encode,
  inverse on decode, with latents_std stored as its reciprocal; shift_factor
  before decode (was BLOCKER #3). Skipping it yields washed-out video, not error.

The remaining unknowns are genuinely box-dependent (FastVideoArgs fields,
shift_factor placement, exact tokenizer kwargs, FSDP) — GPU_BRINGUP.md reconciled
to mark what's now fixed-in-code vs what still needs the box. 204 CPU tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 1e58d90d02 [feat] real torch/CUDA backend (written-not-run) behind the cuda cells
Implement the GPU backend the substrate was built for: Platform.detect() ->
cuda resolves real torch adapters + torch solver ops instead of the numpy
rungs, with the existing loops/policies/scheduler/training unchanged.

WRITTEN-NOT-RUN: this environment has no GPU/torch, so the torch code is
grounded in the verbatim real fastvideo APIs (DiT forward signature confirmed
from source) but cannot be executed/verified here. It is gated available=False
(CPU mini stays green; importing the backends never imports torch), with every
on-box confirm point marked `# BRINGUP` and an ordered checklist in
platform/backends/GPU_BRINGUP.md.

- torch_adapters.py: TorchWanDiT / TorchWanVAE / TorchT5Encoder wrap the real
  module named by each card's load_id and bridge it to the mini's duck-typed
  surface (numpy<->torch at the boundary; loop math stays numpy fp32). DiT
  weight-surface (copy_from/blend_from/clone) for serving sync; mse_grad_step
  raises (GPU training is a separate workstream).
- torch_kernels.py: flow_match_step / flow_sde_step as plain torch elementwise.
  Grounded conclusion from the kernel audit: fastvideo-kernel ships NO fused
  solver kernel (only attention/norm/quant primitives), so the cuda solver is
  torch, registered at arch generic with an honest source string.
- torch_cuda.py: rewritten as lazy trampolines (torch imported only inside
  builder/kernel bodies). Adds the missing vae + text_encoder cuda components
  (they'd otherwise silently fall back to the toy) and corrects the dishonest
  "fastvideo-kernel:flow_*" labels.
- ComponentSpec.checkpoint: the weights source for the torch adapter (risk A;
  the one field the cards didn't carry). Empty on the CPU toys.
- 7 CPU-verifiable wiring tests: honest registration/sources, torch-free import,
  cuda-resolves-real-cells-not-toy, build-fails-loudly-without-torch.

204 tests pass (CPU). The torch path needs a GPU box to verify (GPU_BRINGUP.md).
2026-07-04 17:15:34 +00:00
SolitaryThinker 47f7a04e09 [feat] static-buffer capture form for the cudagraph step body (Path A)
Close the loudest deferred gap from the cudagraph audit: the capturable step
now binds its I/O to address-stable static buffers (modeling real CUDA static
I/O buffers), instead of allocating fresh arrays per call.

- StaticWorkspace: address-stable buffers allocated once per capture key; bind()
  copies the current step's inputs in place via np.copyto, which RAISES on a
  shape/dtype mismatch — turning the weak peak_activation_bytes proxy into a real
  key-soundness backstop (a step whose shape doesn't fit can't replay an
  incompatible graph). Output written into a static buffer too.
- WorkPlan.graph_fn / graph_inputs: a capturable step exposes its deterministic
  op-structure as graph_fn(model, workspace) reading EVERY per-step input (latent,
  sigmas, conditioning, scale) from the workspace — never from closure over
  per-step data — plus the dict of current values. Loops without both stay on the
  eager path. wan21 factors a shared _velocity() so graph_fn and the eager run
  stay bit-identical.
- Capturer dispatch captures/replays via graph_fn against the keyed workspace;
  the workspace is shared per key on the instance. Correct under the engine's
  synchronous step execution (bind+graph_fn atomic per dispatch, output returned
  as a copy) — proven by the batch-of-N interleave gate running two same-key
  requests through the shared workspace bit-identically. A concurrent/multi-stream
  executor would need a per-stream pool (documented).
- 2 new tests (no-static-form eager-break, static-buffer shape-mismatch raises);
  workspace collision test split into bytes-proxy vs shape-backstop.

197 tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 9b8838834c [feat] piecewise CUDA-graph capture/replay at the step boundary (Path A)
Wire the capture/replay lifecycle into the driven-loop step boundary — the
other half of Path A (hand-fused kernels behind the registry + piecewise
cudagraphs, no compiler). Models and tests the correctness-critical control
logic; replay re-runs the current step thunk (CPU models the lifecycle, not the
GPU speedup).

- GraphCapturer on the instance (cross-request cache), wired into
  RuntimeLoopContext.execute and gated by LoopSpec.graph_capture ==
  "breakable_cudagraph". Capture key = (device, arch, loop, shape_sig,
  resident-weight-versions, graph_key). Eager-break for non-capturable (SDE) and
  interceptor-overridden steps. Never executes a stored thunk, so interleaved
  requests can't smear state.
- WorkPlan.capturable / graph_key. wan21 sets capturable=not sde and folds
  compute dtype (shape_sig.dtype) + CFG branch set + expert + scheduler-precision
  into the key — closing a cross-precision key-collision corruption path a
  non-fp32 build would otherwise hit on a real GPU (audit finding).
- Version-in-key auto-invalidation + real eviction: set_weights_version evicts
  the synced component's graphs (duck-typed, so card/ imports no runtime),
  preventing a GPU graph leak across FlowGRPO's per-iteration syncs.
- register_kernel gains the workspace_bytes capture-safety contract (declared in
  the matrix; cuda cells declare real scratch, numpy reference is 0).
- 11 tests: capture-once/replay-many, eager-break (SDE + override), capture ≡
  pure-eager bit-identical, recapture on shape + weight-version change with
  eviction, eager-loop gating, accel-backend capture, and capturer unit tests
  (key discrimination, eager-break, workspace-collision safety net, eviction).

Honestly deferred (bite a real GPU, not the CPU tests): the static-buffer
refactor of the step body, admission budgeting of capture cost (GRAPH_CAPTURE),
and per-card opt-in beyond wan21 — all documented in cudagraph.py + README.

195 tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 634f0828a2 [feat] route all diffusion loops + RL recompute through the kernel table
Finish the kernel seam across the board so the platform's KernelTable is the
universal solver-dispatch path, not just wan21.

- Loops: ltx2, wan_causal, adapters, adaptive now resolve the flow-match solver
  via model.platform.kernels.get(FLOW_MATCH_STEP) instead of importing the numpy
  sampler directly (wan21 already did). On CPU this is bit-identical (the cpu
  kernel IS the old function); a GPU/accel backend now overrides every loop.
- RL: the FlowGRPO log-prob recompute in unified_rl / joint_multi_rl /
  workflow_rl dispatches FLOW_SDE_STEP through the platform, pinned to the SAME
  kernel the rollout used (C2 kernel-pinning — otherwise the PPO ratio biases on
  a real GPU where rollout and recompute kernels could differ).
- accel backend: add a kind-generic AccelComponent wrapper and override the vae
  component too (text_encoder left unregistered to keep the device->cpu fallback
  demonstrated), closing the "only dit is overridden" gap.
- tests: vae override assertion; a second-loop (wan-causal chunk rollout) parity
  oracle proving accel == cpu bit-identical beyond wan21. README scope updated.

184 tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 2f044c02dd [feat] multi-backend dispatch substrate (device/arch/kernel registries)
Add v2/platform/: the (recipe, runtime) backend membrane that lets CPU, GPU,
and other devices coexist behind one dispatch substrate.

- Two tuple-keyed registries: COMPONENTS(kind, device, variant) for
  weight-bearing components and KERNELS(op, device, arch, variant) for
  stateless primitives, each with an availability predicate and an enumerable
  manifest (declared-but-unavailable cells listed without importing torch).
- Platform: detected (device, arch) owning the device + arch fallback chains
  and a per-platform cached KernelTable; detect() is honest (CPU/numpy unless
  torch+CUDA are actually present). Arch fallback is monotonic — only degrades
  to older/portable archs, never a newer binary-incompatible one.
- Seams wired: ModelInstance.component() -> platform.build_component(spec,self)
  with spec.factory as the numpy terminal rung (existing cards untouched); the
  wan21 denoise thunk dispatches solver ops through model.platform.kernels.
- Three backends: cpu (numpy terminal + parity oracle), accel (pure-python
  stand-in proving cross-device resolution, arch fallback, and the oracle),
  torch_cuda (declared-but-unavailable; no faked GPU).
- test_platform.py: 16 tests — detection, terminal rung, arch-fallback walk +
  monotonicity, device precedence, variant fallback, component override +
  per-kind device fallback, the parity oracle (accel == cpu, bit_identical via
  the C1 ladder), and matrix enumeration without importing torch.

Scope is honestly bounded in the README/docstrings: only wan21-denoise routes
through the kernel table and only the dit kind is overridden today (each a
one-line adoption); the torch/CUDA path and a cudagraph workspace-safety
contract are declared/deferred, not implemented. 183 tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 431f4daddb [feat] Adapter plane, non-linear workflows, RL→distill flywheel
Three more capabilities on distinct untested surfaces (167 tests pass, no new runtime primitive).

A — Adapter plane (§9.19): one base + swappable LoRA/ControlNet adapters, selected per request
(DiffusionParams.adapters); AdapterDenoiseLoop applies each active adapter's velocity delta. Per-request
selection changes output, multi-LoRA composes, ControlNet conditions on a control image, mixed-adapter
requests interleave without smearing, hot-swap changes generation, cache key partitions by adapter stack
(the adapter_versions field, previously declared-only). ToyLoRA/ToyControlNet; models/adapters/. 6 tests.

B — Non-linear workflows (§9.17): ParallelWorkflow (fan-out: one input → N models → merged) and
BestOfNWorkflow (generate N → score with the served reward card → return best; inference-time scaling).
The shapes a linear chain can't express. 4 tests.

D — RL→distill flywheel (§9.18): run_flywheel RL-improves the base (NFT), then distills FROM the RL'd model
(DMD2 teacher = RL'd policy) into a faster card, recording the base→rl→distilled provenance chain in
RecipeSpec.parents. The distilled student is measurably closer to the RL'd teacher than the base; the
distilled card serves few-step. training/flywheel.py. 4 tests.

designv4 §9.17–§9.19 + layout/closing/counts (167 tests, 29 files).
2026-07-04 17:15:34 +00:00
SolitaryThinker 6220d02746 [feat] #8b speculative (draft-verify) decoding — exact + lower-latency AR
The last audit stress test. A cheap draft model proposes K tokens, the target verifies
them in one batched step, and SpeculativeARLoop accepts the matching prefix + one target
correction — a variable accepted-length per round (a ragged AR loop the model owns).

- Exactness: the emitted sequence equals the target's OWN greedy decode for any draft
  quality (every accepted token is one the target would produce; the correction is the
  target's token) — the speedup is free.
- Speedup scales with accept rate: draft-agree 0.3→1x, 0.7→3x, 1.0→4x=K tokens/round
  (fewer verify_rounds, the expensive model's latency steps, for the same output).
- Two components (draft + target) co-scheduled on one resident instance; each round an
  AR_TOKEN WorkUnit.

models/speculative/ (loop+card+program); backend ToyTargetModel/ToyDraftModel with a
shared length-dependent target formula (no degenerate fixed point). 5 tests; full suite
153 passed. designv4 §9.16 + counts (153 tests, 26 files).
2026-07-04 17:15:34 +00:00
SolitaryThinker 9489c6c1dd [feat] Five more stress tests: LTX-2 A/V, weight-sync, served reward, cache-dit, nested workflows
The remaining design_v3 probes (all except 8b speculative decoding). All fit with no new runtime
primitive (148 tests pass).

#6 LTX-2 joint audio+video (§9.11) — LTX-2 declared an audio_vae required_for t2vs but never used
   it; now a single 2-stage denoise carries a synchronized audio latent (conditioned on video),
   applies per-modality CFG (guidance_per_modality), and decodes via video VAE + audio VAE → video +
   audio. Gated on requesting audio, so the T2V path is byte-identical (existing tests untouched).
   ToyAudioVAE; build_ltx2_av_program. 5 tests.

#4 Live weight-sync under in-flight serving (§9.14) — WeightSyncController makes the freeze → drain →
   transfer → bump version + invalidate → resume lifecycle explicit. Tests: a mid-flight swap corrupts
   (the hazard); draining first leaves the in-flight request bit-identical to baseline while a
   post-sync request reflects new weights; transformer-only sync, so the frozen text-encoder cache
   survives. The RL flywheel's hardest correctness. 3 tests.

#5 Reward-model-as-a-served-card (§9.15) — a reward model is a card (scorer + a score loop emitting
   REWARD_BATCH units); ServedRewardScorer drop-in-replaces the numpy scorer so any RL method becomes
   RLHF/RLAIF with no method change. ToyRewardModel; models/reward/. 4 tests.

#7 Content-adaptive control flow (§9.12) — CacheDiTDenoiseLoop (isolated WanDenoiseLoop subclass)
   reuses the cached velocity when predictions barely change (cache-dit skip) and early-exits on
   convergence — variable step count; interleave parity holds across ragged loops. models/adaptive/. 4 tests.

#8a Nested workflows (§9.13) — a workflow stage can invoke another workflow (engine.run routes ids);
   requires/validate recurse; cycles caught at registration + a run-time guard (engine._wf_running).
   build_t2i_i2v_extend_workflow. 5 tests.

designv4 §9.11–§9.15 + falsifier/layout/closing updates (148 tests, 25 files). Also removed a
pre-existing unused import in ltx2/loop.py.
2026-07-04 17:15:34 +00:00
SolitaryThinker 69c9871154 [examples] Add v2_examples/{training,omni,workflows}/ — runnable examples
Three more example folders alongside inference/, all CPU/numpy, self-contained
(sys.path bootstrap), public API only, every script verified to run green.

training/ (7) — one per method, all via the uniform method.train_step seam:
  01 finetune · 02 dmd2 distillation · 03 diffusion_nft (likelihood-free RL,
  samples from the old policy, feature-cache reuse) · 04 self-forcing (causal
  chunk_rollout) · 05 joint LM+generator RL (UniRL; joint + prompt-only) ·
  06 N-way joint RL (per_expert vs shared credit) · 07 end-to-end workflow RL
  (T2I+I2V from one final-video reward).

omni/ (4) — 01 Cosmos3 (reason→joint denoise, shared MoT) · 02 BAGEL
  (text→image, shared MoT; scheduler prices both WorkUnit kinds) · 03 Qwen-Omni
  (thinker→talker→vocoder, three separate experts, text+audio) · 04 interleave
  parity across AR + diffusion loop types.

workflows/ (2) — 01 cross-model T2I→I2V workflow (image provably conditions the
  video) · 02 workflow as a first-class servable (requires/validate, address by
  id, register_workflows catalog, WorkflowRegistry).

Each folder has a README indexing its scripts.
2026-07-04 17:15:34 +00:00
SolitaryThinker 18dd295e8d [examples] Add v2_examples/inference/ — runnable Wan2.1 inference examples
Five self-contained, runnable scripts (CPU/numpy) for the Wan2.1-1.3B card on the
v2 runtime, each bootstrapping the repo onto sys.path so they run from anywhere:

- 01_basic_t2v.py                  minimal path: build engine → T2V request → run → video
- 02_params_and_reproducibility.py DiffusionParams knobs + seeded bit-identical reproducibility
- 03_streaming.py                  per-denoise-step preview chunks (OutputSpec stream)
- 04_concurrent_interleaved.py     step-interleaved batching + interleave parity gate + cache reuse
- 05_async_serving.py              AsyncEngine: concurrent generate, event stream, step-boundary cancel

+ README.md indexing them. All five run green; use only the public API.
2026-07-04 17:15:34 +00:00
SolitaryThinker 4c333e0509 [docs] Update v2/README to designv4 + current scope (127 tests)
- v2/README.md: point to designv4.md as the unified design (design_v3 as
  north star); refresh the scope table (joint/N-way/workflow RL, Qwen-Omni
  cascade, cross-model Workflow, tiled VAE co-scheduling, WorldModelSession),
  package layout (program/Workflow, runtime/session, the 7 methods, new model
  dirs), the demonstrated-stress-tests list (§9.3–§9.10), and counts (49→127,
  20 files). Sessions moved out of "out of scope"; WebRTC wire stays out.
- designv4.md: drop two intermediate absolute suite totals (milestone "91/97
  passed") in favor of "zero regressions" so the only absolute count is the
  current 127 (intro + layout) — no stale numbers.
2026-07-04 17:15:34 +00:00
SolitaryThinker 32a7a6b87b [feat] Three stress tests: interactive sessions, workflow RL, heterogeneous co-scheduling
Targets the three design_v3 claims that were most load-bearing AND least
exercised (sessions/realtime, training-plane boundary, the WorkUnit-generality
falsifier). All fit with no new runtime primitive (127 tests pass).

1. Interactive world-model session (runtime/session.py, §9.8) — the Session
   plane had ZERO coverage. WorldModelSession drives the causal chunk_rollout
   loop as a long-lived session: persistent cross-request world state on the
   Session.kv_handle, frame streaming, transactional step-boundary cancellation
   (a cancelled act leaves the world resumable), no cross-session smearing.
   Added only a continuation seam to the chunk loop (init seeds context from a
   world_context slot; default empty = unchanged one-shot path). 5 tests.

2. End-to-end RL over a cross-model workflow (training/methods/workflow_rl.py,
   §9.9) — trains BOTH flux-t2i and wan-i2v from ONE final-video reward. Rolls
   out the whole workflow with SDE capture in both instances; the same final
   advantage drives FlowGRPO PPO on each stage's transformer; two WeightSyncPlans
   on two instances. The earlier model (T2I) is trained by a reward on the final
   video — end-to-end credit across a model boundary — proven causal by a control
   (constant reward => zero advantage => nothing moves). 4 tests.

3. Heterogeneous WorkUnit co-scheduling (models/tiled/, §9.10) — the §17
   falsifier. VAETileLoop makes VAE decode a loop of VAE_TILE units; tiling is
   exact (== one-shot, C0), and VAE_TILE + DIFFUSION_STEP pipelines interleave
   bit-identically and co-run in one batch. Validates the mechanism; the
   economic half (does it pay) stays a port-time measurement. 4 tests.

designv4 §9.8–§9.10 + falsifier/layout/closing updates (127 tests, 20 files).
2026-07-04 17:15:34 +00:00
SolitaryThinker 58223c0c41 [feat] Register cross-model workflows as first-class named servables
Answers "what's the right way to name/register custom pipelines like T2I→I2V":
treat a Workflow like a card — a stable namespaced id in the same servable
namespace, declared dependencies, and a two-level registry. No new concepts;
mirrors how cards are registered (and vllm-omni's pipeline_registry).

- program/workflow.py: Workflow gains `requires` (the cards it composes, derived
  from stages) and `validate(engine)` (fail-fast if a required card is absent,
  P7). New WorkflowRegistry: declarative workflow_id -> builder catalog for
  out-of-tree/ad hoc use.
- runtime/engine.py: `_workflows` registry + register_workflow (validates deps,
  rejects id collision with a model_id) + serves(); engine.run routes a request
  whose model_id is a workflow to workflow.run — addressable exactly like a model.
  Single-model hot path untouched.
- runtime/async_engine.py + serving/server.py: serves() and /models include
  workflows (discoverable as servables).
- models/__init__.py: declarative `_WORKFLOWS` catalog (the cross-model analog of
  _BUILDERS) + register_workflows() helper; build_image_video_engine now registers
  the workflow too. Adding a custom pipeline = one catalog line.
- Naming convention: dotted/namespaced workflow_id (`image_video.t2i_i2v`),
  distinct from kebab model ids, collision-checked. Renamed from `t2i_then_i2v`.
- tests (+5, 12 total in the file): addressable by id, requires/validate, id
  collision, registry catalog, register_workflows helper. Full suite 114 passed.
- designv4 §9.6: the naming & registration convention documented.
2026-07-04 17:15:34 +00:00
SolitaryThinker b254d1affe [feat] Cross-model T2I→I2V workflow + N-way joint RL over arbitrary experts
Two more pipelines stress-testing the design, plus a BAGEL-placement note in
designv4. Both fit with no new runtime primitive (109 tests pass).

Pipeline 1 — cross-model T2I→I2V (program/workflow.py, models/image_video/):
- Realizes ProgramKind.WORKFLOW as a thin multi-instance orchestrator ABOVE the
  engine (the hot path stays single-instance). A Program composes one model's
  loops; a Workflow chains full engine.run calls across distinct cards, threading
  artifacts. (LTX-2 already covers same-card multi-stage; cross-model — FLUX→Wan
  — is the new capability the single-instance runner can't express.)
- flux-t2i (text→image) and wan-i2v (text+image→video) cards; the I2V program
  folds the conditioning image into text_embeds so WanDenoiseLoop is unchanged.
- Each model keeps its own interleave-parity guarantee (crossing instances is a
  Workflow boundary, not a loop step). 7 tests incl. video-depends-on-image.

Pipeline 2 — N-way joint RL (training/methods/joint_multi_rl.py, models/multi_expert/):
- JointMultiExpertRL generalizes UnifiedRLMethod (N=2) to N refiner LMs + a
  generator: one reward → one group advantage → N token-PG updates + 1 FlowGRPO
  PPO update, N+1 independent WeightSyncPlans. Proves the substrate was already
  N-ready (card holds N components/loops; per-component weight-sync; dict grad
  targets) — only the method body looped over two; now it loops over a list.
- credit="per_expert" learns all N cleanly; credit="shared" (faithful to UniRL)
  works but is noisier — the honest multi-agent credit-assignment result, a
  reward-shaping choice, not a substrate limit. 6 tests (N=1,3,4; prompt-only).

Fix — flow_sde_ml_velocity (loop/sampler.py): the toy FlowGRPO generator update
targeted the velocity the model already produced (a no-op once guidance_scale=1
was set for the C2 identity; the unified generator moved only on ~1e-7 noise).
The correct PG surrogate targets the max-likelihood velocity of the realized
sample — nonzero at ratio==1. Both UniRL and N-way generators now learn for real;
the C2 ratio==1 identity still holds (measured before the update).

BAGEL: MoT/shared-weight (one transformer on both loops), same row as Cosmos3;
real BAGEL's co-resident experts are expressible via the expert-routing policy
(partial sharing) — captured in designv4 §2.3.
2026-07-04 17:15:34 +00:00
SolitaryThinker c7e0a8e894 [feat] Add Qwen-Omni thinker→talker→vocoder model (3 experts, 3 loops)
Ports vllm-omni's canonical qwen2_5_omni omni-speech cascade as a v2 card:
a third weight-sharing topology — three disjoint experts (thinker, talker,
vocoder) on three loop types (ar_decode → ar_decode → audio_decode) in one
request, with chained cross-stage conditioning and streaming codec→waveform.
vllm-omni runs these as three opaque request-scheduled stages; v2 makes every
thinker token, talker token, and vocoder chunk a runtime-visible WorkUnit.

- models/backend.py: ToyTalker (a genuinely distinct AR expert, weight-salted)
  + ToyVocoder (streaming code2wav: codec tokens → waveform chunks).
- models/omni/vocoder_loop.py: VocoderLoop filling the pre-declared
  LoopKind.AUDIO_DECODE / WorkUnitKind.AUDIO_CHUNK slot.
- models/omni/ar_loop.py: ARDecodeLoop gains a configurable prompt_slot so two
  chained AR loops don't collide on the prefill slot (thinker vs talker).
- models/qwen_omni/: card (3 experts/3 loops) + program (tokenize → thinker →
  emit_text → thinker→talker full-payload hand-off → talker → talker→vocoder →
  vocoder → emit_audio). Cross-stage hand-offs are explicit Program nodes, the
  model-native form of vllm-omni's custom_process_input_func.
- _enums.py: Capability.TEXT_TO_SPEECH.
- tests/test_thinker_talker.py: 6 tests incl. three-loop interleave parity,
  cascade conditioning, AUDIO_CHUNK streaming. Full suite 97 passed.

designv4.md: §2.3 topology table extended to four topologies; new §9.5 on the
cascade; reference-synthesis + package layout updated.
2026-07-04 17:15:34 +00:00
SolitaryThinker b3ddf6014d [feat] UniRL/PromptRL joint LM+generator RL stress test + designv4
Stress-tests the v2 Card/Loop/Program design with a UniRL/PromptRL-style
joint RL recipe: a prompt-refiner LM expert and a flow generator expert,
two separate experts driven by two loop types in one request, both updated
simultaneously from a single RL reward.

- loop/sampler.py: flow_sde_step_with_logprob — FlowGRPO SDE rollout sampler
  (per-step Gaussian log-prob), distinct from the deterministic ODE serve step.
- request/params.py: gated sde_rollout/sde_noise_scale on DiffusionParams so
  the serve path stays byte-identical (default ODE).
- models/wan21/loop.py: gated SDE-rollout capture in WanDenoiseLoop.advance.
- models/backend.py: ToyPromptRefiner — a real REINFORCE categorical policy
  (the Qwen role), separate weights from the generator.
- models/unified/: the unified card+program — two disjoint experts (llm +
  transformer) on ar_decode + diffusion_denoise; the topological opposite of
  the Cosmos3 MoT card, same vocabulary.
- training/methods/unified_rl.py: joint GRPO — one reward -> group advantage
  -> LM token policy gradient + DiT FlowGRPO PPO; two LRs; prompt-only/joint
  flag; reuses the shared diffusion loop for rollout.
- training/weight_sync.py: WeightSyncPlan gains a component scope so the two
  experts version + cache-invalidate independently (LM sync never flushes the
  frozen text-encoder feature cache).
- tests/test_unified_rl.py: 9 tests incl. likelihood-based C2 identity, the
  two-loop interleave parity gate, joint vs prompt-only. Full suite 91 passed.

designv4.md: unified design doc reflecting v2 as built+tested, with the joint
RL stress test as the validating case study (the design held — new card +
new method, no new runtime primitive).
2026-07-04 17:15:34 +00:00
SolitaryThinker a9e5f6ee7a [refactor] rename package mini_fastvideo → v2
Directory rename (git mv, history preserved) plus rewrite of all references: absolute imports in
tests, the zero-dep runner, docstrings, comments, and the README. No behavior change.

Run: python3 -m pytest v2/tests/ -q ; python3 v2/run_tests.py ; python3 -m v2.examples
2026-07-04 17:15:34 +00:00
SolitaryThinker 7467076d72 [fix] mini-fastvideo serving: address adversarial-review findings (capacity/credit leaks, robustness)
Review confirmed the core bets (concurrent disaggregation is bit-identical, design conformance holds,
Dynamo genuinely optional, engine stays step-scheduled). Fixes for the untested failure paths:

- HIGH: pool capacity (RolePool.in_flight) no longer leaks when a disaggregated request is cancelled
  or errors mid-occupancy — DisaggregatedRunner.close() releases the occupied pool and AsyncEngine._run
  calls it in a finally (a cancel on a capacity-1 denoiser no longer bricks the pool).
- credit flow-control: cross-pool transfer wraps acquire/release in try/finally (no credit leak on a
  failing transfer); slot is re-homed only on a successful fetch.
- AsyncEngine: duplicate in-flight request_id is rejected (was a deadlock); submit() cancels the driver
  task when the consumer abandons the stream (client disconnect → no orphaned compute); bounded
  per-request history (no unbounded _events/_states/_results/_runners growth).
- cancellation is common-path on the OFFLINE path too (cancel check at the top of every runner.tick()).
- HTTP server: read timeout (slowloris guard → 408), body-size cap (→ 413), invalid Content-Length
  (→ 400), explicit StreamReader limit, and aclose() of the SSE generator on client disconnect.
- build_deployment_card no longer aliases one mutable CostModel across replica cards (dataclasses.replace),
  so online calibration of one worker's cost doesn't mutate another's.
- video-job tasks tracked (not fire-and-forget); server.close() cancels/drains them; jobs dict bounded;
  fleet affinity map bounded.

6 regression tests added for these paths. 82 tests pass (pytest + zero-dep runner).
2026-07-04 17:15:33 +00:00
SolitaryThinker 01dc0c3377 [feat] mini-fastvideo serving + fleet (our own version, Dynamo-optional)
Builds the full serving layer the design files specify, instead of deferring it to Dynamo:

- transport/ (§7.3): pluggable Connectors (in-proc zero-copy / SHM-fake copy) with chunk_ready
  readiness (vllm-omni) AND credit-based flow control (sglang-omni Relay); KVConnector protocol shape;
  TransferManifest.
- runtime/ (§6, §13; plan M3/M4): AsyncEngine — request queue, lifecycle state machine
  (waiting→running→completed/cancelled/failed), live AsyncIterator[OmniEvent] streaming, common-path
  cancellation, step-level concurrency. RolePool + DisaggregatedRunner (encoder→denoiser→decoder,
  capacity-aware dispatch, cross-pool transfers via connectors); disaggregated output is bit-identical
  to inline. No-progress detection (no busy-spin).
- deploy/ (§14, §6.3.5-6): DeploymentCard; OUR OWN LocalFleet (discovery, health/drain, least-loaded
  / cost-model / sticky-affinity routing) so we never rely on Dynamo; DynamoWorkerAdapter +
  FakeDynamoRuntime export the SAME card + cost model so Dynamo CAN front us — one object, two consumers.
- serving/ (§6.3.5, §12): framework-free stdlib-asyncio OpenAI server (our own version of the
  vllm-omni pattern): /v1/chat/completions (SSE), /v1/images/generations, /v1/videos (async job+poll)
  + /v1/videos/sync, /v1/models, /health, /metrics. A thin shim over the STEP-scheduled engine — the
  runtime-visible loop scheduler vllm-omni's request-scheduled opaque DIFFUSION stage lacks.

15 serving tests (real-socket HTTP+SSE via stdlib asyncio, disagg==inline, fleet routing, Dynamo
contract, cancellation); 76 tests pass total (pytest + zero-dep runner). ~7900 LOC.
2026-07-04 17:15:33 +00:00
SolitaryThinker f1dc587c74 [feat] mini-fastvideo phase 2: omni/MoT — Cosmos3 + canonical vllm-omni (BAGEL/lance)
One resident MoT instance runs BOTH an ar_decode loop and a diffusion_denoise loop on shared weights
(the §16 claim no DAG-of-engines can express), with both loops runtime-visible: the scheduler prices
ar_token AND diffusion_step WorkUnits — unlike vllm-omni's opaque DIFFUSION stage the scheduler never
sees inside.

- ARDecodeLoop: token decode until EOS/max_tokens, paged text-KV — the omni AR pathway (loop/§5).
- ToyMoTDiT: one module exposing an und pathway (ar_forward) AND a gen pathway (denoise __call__);
  ToyTokenizer. Binding both loops to one instance = shared weights, no duplication.
- models/cosmos3/: tokenize → reason(ar_decode) → pack(tokens→conditioning) → diffusion_denoise →
  vae_decode; sound_vae declared optional_for non-t2vs (the lazy-component P8 fix, not an env-var hack).
- models/bagel/: the canonical vllm-omni model — generate_text(ar_decode) → generate_image(diffusion),
  text+image outputs, both loops step-scheduled.
- The diffusion loop is WanDenoiseLoop reused (one loop definition bound to the MoT module).
- build_omni_engine() + an omni worked example; 7 omni tests (shared-instance, both-kinds-scheduled,
  interleave parity across loop types, lazy sound_vae). 61 tests pass (pytest + zero-dep runner).
2026-07-04 17:15:33 +00:00
SolitaryThinker 098bcf014a [fix] mini-fastvideo: address adversarial-review findings (admission liveness + §7.1 cache key)
- Admission fails fast with AdmissionInfeasible on infeasible/deadlocked reservations instead of a
  10M-iteration busy-spin: no-progress detection in run_to_completion/run_interleaved via a real
  progress token, plus feasibility pre-checks (need > pool capacity).
- Compute budget is now a refundable concurrency gate (release() refunds spent), not a
  never-refunded lifetime cap that silently deadlocks.
- §7.1: the text-encoder feature CacheKey carries adapter_versions + precision (no stale serve across
  te-LoRA stacks); per-component weight versions mean a transformer-only RL weight sync no longer
  flushes the frozen text-encoder cache (component-scoped invalidation, not wholesale).
- Interleave gate flags symmetric-empty output instead of passing it vacuously.
- skipped_steps counted only when the override is actually consumed; BatchScheduler wired for round
  batch-accounting (metric renamed stepped_units); ResidualCache.get cleanup; stream chunks carry a
  latent preview payload; dead progress-vars removed.
- 5 regression tests added for the previously-untested paths. 54 tests pass (pytest + zero-dep runner).
2026-07-04 17:15:33 +00:00
SolitaryThinker 270fae959d [feat] mini-fastvideo: model-native runtime per design_v3 (Wan2.1/LTX2 + 4 training methods)
A scoped, CPU-testable realization of design_v3.md — the architecture where the atomic unit
is a typed (recipe, runtime) ModelCard, the model owns loop semantics while the runtime owns
loop lifecycle, one resident instance runs many loops, and training records behavior on the
same loops it serves.

Implements:
- card/ loop/ runtime/ cache/ memory/ parallel/ parity/ extend/ program/ request/ training/
  spanning design_v3 §4-§13: ModelCard + validate(); driven loops (init/next/advance/finalize);
  step-interleaving Engine with reservation-before-admission + per-class caches keyed by CacheKey;
  the C0-C4 consistency ladder + the non-negotiable batch-of-N interleave parity gate.
- Inference: Wan2.1-1.3B (T2V), LTX2.3 (two-stage distilled, shared transformer), Wan-causal
  (chunk rollout + slab-KV streaming).
- Training (Wan2.1-1.3B), each driving the SAME loops the engine serves: finetune (flow-match),
  DMD2 (teacher/critic distribution matching), DiffusionNFT (likelihood-free C2, samples from the
  decay-blended old policy, group-relative advantages, shared-prompt cache reuse), self-forcing
  (causal chunk loop). The engine never imports training (the §10 dependency rule, grep-verified).

numpy-only core (no torch/GPU here); heavy Wan/LTX forwards are deterministic toy stand-ins with
lazy torch-adapter seams (ComponentSpec.load_id/factory) for a GPU box. 49 tests pass via pytest
and a zero-dependency runner; interleave parity verified (and a buggy module-global interceptor
provably breaks it). Omni-ready spine (ar_decode/chunk_step loop kinds, multi-loop instances,
LoopState.extension) for the phase-2 Cosmos3 + vllm-omni omni ports.

Run: python3 -m pytest mini_fastvideo/tests/ -q ; python3 -m mini_fastvideo.examples
2026-07-04 17:15:33 +00:00
SolitaryThinker 4a14c1afa3 update 2026-07-04 17:15:33 +00:00
SolitaryThinker f1c19050c3 design 2026-07-04 17:15:33 +00:00
William Lin 51ed1ea423 [build] Bump fastvideo-kernel pin to 0.3.2 (#1541) 2026-07-03 14:48:15 -07:00
Mac Lee 00ec3e7388 [bugfix]: compute VSA topk from padded blocks (#1517) 2026-07-03 14:16:51 -07:00
William Lin 0d626ef2d1 [kernel] Bump fastvideo-kernel pin to 0.3.1 and version the FA4 tile_mn port as 0.3.2 (#1539) 2026-07-03 14:11:08 -07:00
William Lin b36d0ef085 [infra] Add DGX Spark and multi-architecture CUDA support (#1447) 2026-07-03 13:41:43 -07:00
William Lin 10c8c5df4d [misc] reorg: relocate ui/ and performance_dashboard/ under apps/ (#1537) 2026-07-02 20:29:17 -07:00
William Lin 31e26abec4 [chore] release fastvideo-kernel 0.3.1 (#1520) 2026-06-30 12:25:27 -07:00
sumyyyyyandSolitaryThinker 9e83ba630c [kernel] Extract VSA utility functions into fastvideo_kernel (#1408)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-30 12:19:25 -07:00
William Lin fc02a9ce8e [ci] aarch64 kernel wheel: build for Blackwell (sm_100a/sm_120a), not Hopper (#1516) 2026-06-29 16:10:42 -07:00
William Lin 3d36160fc4 [bugfix] -fsigned-char so ThunderKittens compiles on aarch64 (Grace Hopper) (#1515) 2026-06-29 14:05:38 -07:00
William Lin a11ec43de2 [ci] build + publish aarch64 (Grace Hopper) kernel wheels alongside x86_64 (#1514) 2026-06-29 12:28:34 -07:00
William Lin 2656d6530c [bugfix] compile FP4 (attn_qat_infer) kernels for sm_120a only via per-arch split (#1508) 2026-06-29 11:07:27 -07:00
Kevin Lin e3f54e7169 [bugfix] Fix causal attention mask for Blackwell FP4 MMA column layout (#1506) 2026-06-28 20:36:40 -07:00
William Lin 4ba5681307 [bugfix] guard ThunderKittens Hopper kernels for non-sm_90a device passes (#1507) 2026-06-28 20:07:43 -07:00
alexzms 8c23c86994 [docs]: add raw-video preprocess script for the QAD MixKit recipe (#1487) 2026-06-28 18:01:13 -07:00
William Lin 78d606a3a7 [feat]: make VSA tile cache configurable for training (#1444) 2026-06-28 02:18:31 -07:00
William Lin c4e108de78 [docs]: modernize documentation build (#1503) 2026-06-28 02:02:09 -07:00
Kaiqin Kong 16bf2eaf77 [feat] Relativistic RoPE re-indexing for long causal rollouts (#1454) 2026-06-28 01:43:00 -07:00
William Lin 1d15d974fa [misc]: cleanup outdated or unneeded agent infrastructure (#1504) 2026-06-27 19:37:49 -07:00
alexzms 5e57868b76 [ci] train-framework model coverage: Cosmos/MatrixGame2 finetune grad-norm + Cosmos/LongCat/MatrixGame2 loading smokes (#1497) 2026-06-27 17:16:00 -07:00
William Lin 6be280c914 [bugfix]: fix docs build (#1502) 2026-06-27 13:00:09 -07:00
KyleNeverGivesUpandClaude Opus 4.8 3ccdec9798 [bugfix] Make enable_torch_compile_vae actually compile the Wan VAE (#1498)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 12:30:59 -07:00
Satyam Srivastava 8658f774f3 [docs] Restructure contributing CI/CD and testing docs (#1501) 2026-06-27 11:46:18 -07:00
Raghav K 4f3ad3f6df [perf] Default Wan VAE decode to bf16 (lossless, faster) (#1472) 2026-06-26 12:42:29 -07:00
William Lin b454aa56c3 [misc] fix pre-commit (#1500) 2026-06-26 12:39:28 -07:00
William Lin f9b8e30ff3 Update README.md 2026-06-26 09:23:30 -07:00
William Lin 0356205b84 [bugfix] revert fastvideo kernel version to 0.2.6 (#1495) 2026-06-25 03:52:43 -07:00
Kevin Lin 719a1879bd [bugfix] Remove incorrect V-row permutation in scaled_fp4_quant_trans_kernel (#1493) 2026-06-25 03:05:12 -07:00
Utkarsh RanjanandClaude Opus 4.8 b57180bf97 [bugfix] Warn when a requested attention backend is unsupported by a layer (#1254) (#1486)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 15:56:39 -07:00
William Lin 7cebf5f82c [ci] cap kernel wheel build parallelism to avoid runner OOM (#1483) 2026-06-23 14:17:56 -07:00
William Lin dd0f4b6753 [misc] cleanup misc files (#1484) 2026-06-23 14:17:22 -07:00
William Lin d303b4e03a [docs] update README (#1482) 2026-06-23 12:07:21 -07:00
xsankandmergify[bot] 31b719ae49 [perf] optimize compress & topk kernel (#1421)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 18:51:22 +00:00
William Lin 887aaf3d3e [ci] bump cuda-toolkit action to v0.2.35 to fix kernel cu130 publish (#1481) 2026-06-23 10:33:31 -07:00
William Lin b1d89eba1f [chore] release fastvideo-kernel 0.3.0 (#1478) 2026-06-23 10:02:27 -07:00
Loay RashidandSolitaryThinker 995a5fdf97 [bugfix] fixing denoising time in the fastwan script (#1480)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-23 10:01:23 -07:00
William Lin 70a70b689e [ci] mergify: stop auto-syncing ready PRs (#1477) 2026-06-23 03:23:58 -07:00
Shao Duan 4d6ac89b43 [ci] eval: add metric regression + identity-invariant ci tests (#1451) 2026-06-23 01:27:34 -07:00
Raghav K 4171cacd93 [bugfix] Wire Cosmos-Predict2.5 2B to its sampling preset (#1468) 2026-06-23 01:25:35 -07:00
b2ade71467 [Bugfix] QAD 5090: Torch.compile and other optimizations (15/12) (#1466)
Co-authored-by: alexzms <3036648523@qq.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-23 01:23:57 -07:00
Kevin Lin 82ed9fe58d [feat] QAD 5090: FP8 linear layer inference (#1465) 2026-06-22 18:25:06 -07:00
Satyam Srivastava 3d8cc4f0a0 [bugfix] Fix performance component timing extraction (#1473) 2026-06-22 13:05:35 -07:00
Satyam Srivastava 0557f7a7d9 [ci] Add performance dashboard metadata and visualizations (#1470) 2026-06-19 14:27:01 -07:00
dc66cd97ef [feat] QAD 5090: FP8 QAT linear training (14/12) (#1464)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-19 01:03:51 -07:00
6da206e196 [feat] QAD 5090: FP4 QAT linear STE for training (13/12) (#1463)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 15:55:47 -07:00
sumyyyyy 87f98c9b8b [kernel] Add varlen support for block-sparse attention (#1319) 2026-06-18 01:06:00 +00:00
e60601df7f [feat] QAD 5090: QAT training recipe — finetune + DMD distillation (12/12) (#1462)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 15:18:52 -07:00
eed9c4bfbf [kernel] QAD 5090: Add Attn-QAT training Triton kernels (11/12) (#1460)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
2026-06-17 13:35:12 -07:00
1dee77f4a4 [feat] QAD 5090: Wire the Attn-QAT training attention backend (10/12) (#1459)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: alexzms <26690162+alexzms@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 13:32:56 -07:00
Satyam Srivastava b80148819c [ci] Add performance dashboard visualisation scripts (#1469) 2026-06-16 23:29:29 -07:00
c3b971488e [docs] QAD 5090: Add NVFP4 + Attn-QAT inference example and how-to (9/12) (#1458)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: alexzms <26690162+alexzms@users.noreply.github.com>
2026-06-16 14:36:50 -07:00
88e753f281 [feat] QAD 5090: Wire the Attn-QAT inference attention backend (8/12) (#1457)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: alexzms <26690162+alexzms@users.noreply.github.com>
2026-06-16 13:12:33 -07:00
77832059cc [kernel] QAD 5090: Add modified SageAttention3 FP4 inference kernels (7/12) (#1455)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: Edenzzzz <wtan45@wisc.edu>
2026-06-16 12:02:30 -07:00
alexzmsandmergify[bot] 633d393568 [ci] layer-0 grad-norm regression for per-method training tests (5a-ii) (#1396)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-12 04:45:07 +00:00
Junda SuandPeiyuan Zhang 5854aec2ce [feat] Add Wan RL DiffusionNFT training (#1450)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2026-06-11 21:18:59 -07:00
Mook 30e45c2411 [bugfix] Classify new config/sampling fields in schema parity inventory (#1446) 2026-06-10 13:38:36 -07:00
alexzms 2a4fe697a6 [docs] LTX-2.3 distilled i2v: typed-API example (from_config + generate) (#1448) 2026-06-10 10:58:14 -07:00
alexzms 921db7479d [perf] LTX-2.3 distilled i2v: drop max-autotune from compile kwargs (#1445) 2026-06-10 10:06:41 -07:00
Aryan KumarandAryan Kumar 7f539424cb [feat]: add Lucy Edit inference scaffold (#1363)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-06-09 15:21:43 -07:00
19a838f54f [bugfix]: release VSA tile cache during training (#1434)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-09 15:01:22 -07:00
d922ab2cbc [model] Flux2 Klein Port (#1349)
Co-authored-by: Gnav3852 <63612880+Gnav3852@users.noreply.github.com>
Co-authored-by: Mac Lee <macthecadillac@gmail.com>
2026-06-09 14:55:55 -07:00
Kaiqin Kongandmergify[bot] 9ea77d37f3 [bugfix] EMA shadow on resume and EMA under MoE path (#1441)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-09 00:55:43 +00:00
2e35b0c6bd [refactor]: linear/mlp FP4 path additions for Wan-2.1 (Attn-QAT 6/12) (#1390)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
Co-authored-by: Matthew Noto <notomatthew31@gmail.com>
2026-06-08 14:51:57 -07:00
William Lin 1c627a3f98 [bugfix]: build fastvideo-kernel on GPU-less Docker runners (#1437) 2026-06-06 17:52:45 -07:00
Kaiqin Kong a931efe33a [bugfix] EMA in distillation pipeline (#1440) 2026-06-06 17:38:03 -07:00
William Lin 041e5e9029 [bugfix]: unblock PyPI publish (flash-attn-cute direct dep) (#1436) 2026-06-05 10:07:12 -07:00
Kaiqin Kongandmergify[bot] efcc245c2e [feat] VLM as judge for WM (#1429)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-05 04:25:54 +00:00
alexzmsandmergify[bot] 922e7e0813 [docs] LTX-2.3 distilled i2v example with compile + timing breakdown (#1430)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-05 01:26:50 +00:00
William Lin c62a8514b0 [chore]: release v0.2.0 (#1432) 2026-06-04 14:22:49 -07:00
William Lin 3eb8081801 [chore]: unpin runtime deps in pyproject.toml (#1431) 2026-06-04 13:46:49 -07:00
Raghav K 3505d09564 [bugfix] tests: include ltx2_3_base in expected LTX2 preset set (#1427) (#1428) 2026-06-04 12:17:18 -07:00
Shao DuanandSolitaryThinker 570607945c [feat] dreamverse: sequence parallelism for serving (#1424)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-01 23:13:49 -07:00
Kaiqin KongandSolitaryThinker d3a821cdcf [feat] LoRA controls and integration for Dreamverse (#1420)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-01 18:54:11 -07:00
Kaiqin Kong 3f24578139 [bugfix] LTX2: honor video_position_offset_sec in the DiT (#1422) 2026-06-01 16:59:15 -07:00
Kevin Lin 89fcf08378 [bugfix] Fix STFT dtype mismatch (#1419) 2026-05-31 20:56:07 -07:00
c2930b2aa1 [ci] Add additional Dreamverse UI tests (#1417)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-31 19:53:29 -07:00
KUAN-HAO HUANGandSolitaryThinker 019239690b [perf] Add Adaptive Guidance (CFG gating) for stale-uncond reuse (#1372)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-30 13:58:55 -07:00
alexzms d6119c1f82 [bugfix]: dreamverse modal bypasses ENTRYPOINT — set ffmpeg env + key check (#1413) 2026-05-29 16:50:01 -07:00
alexzms 84214c80bb [feat] LTX-2.3 audio: BWE vocoder path (#1398) 2026-05-29 16:43:26 -07:00
alexzmsandmergify[bot] c2d7143c72 [feat] LTX-2.3 transformer support (config-gated extension of LTX-2) (#1397)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-29 16:42:37 -07:00
Kaiqin KongandSolitaryThinker afdb6fbfa5 [feat] Add MatrixGame3.0 (#1201)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-27 21:07:01 +00:00
alexzms ba4c02d883 [docs]: highlight Dreamverse deployment paths + add Server B200 (SSH) guide (#1409) 2026-05-27 09:30:58 -07:00
Junda Su 2c137931f3 [bugfix] Fix Dreamverse Modal compile warmup latency (#1394) 2026-05-26 17:58:38 -07:00
Shao Duan 36682797a0 [feat] eval: input ergonomics + Evaluator features + bug fixes (#1392) 2026-05-26 13:59:45 -07:00
William Lin 0ef1357a77 [docs]: surface activation-trace utility in add-model skills (#1399) 2026-05-26 13:45:12 -07:00
William Lin a75d19786a [docs]: Wire activation trace into mkdocs nav + perf/troubleshooting (#1304) 2026-05-26 13:14:36 -07:00
alexzms be548a78ea [feat] VSA-256 fastpath on Blackwell via FA4 CuTe block-sparse attention (#1354) 2026-05-26 12:58:29 -07:00
alexzms 6a610e2bc9 [ci] add per-method single-step training tests for fastvideo.train (#1343) 2026-05-25 17:09:54 -07:00
Shao Duanandabaghyangor 321d5112b4 [refactor] eval: consolidate FVD into common.fvd, remove benchmarks/fvd (#1380)
Co-authored-by: abaghyangor <abaghyangor@gmail.com>
2026-05-24 12:09:47 -07:00
William Lin ba75ad82db [refactor]: shared attention infra additions for QAT-compat (Attn-QAT 5/12) (#1383) 2026-05-23 16:19:24 -07:00
Junda Su 58caa5109f [ci] Add DreamVerse app CI tests (#1386) 2026-05-23 14:28:26 -07:00
2f3ca8aaad [perf]: register FA2/FA3 default flash_attn_func as a torch.library custom op (#1373)
Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-23 00:39:06 -07:00
Junda Su 3a67319cb6 [infra] Use npm for Dreamverse web builds (#1385) 2026-05-22 18:50:10 -07:00
Junda Su 266fa044b3 [infra] Add Dreamverse Modal UI image build (#1381) 2026-05-22 11:36:53 -07:00
William Lin f3398db868 chore: pin dreamverse npm deps to address Dependabot alerts (#1359) 2026-05-22 11:29:33 -07:00
fda02036bc [feat]: Attn-QAT inference + training backends (deadcode) (Attn-QAT 4/12) (#1358)
Co-authored-by: jzhang38 <42993249+jzhang38@users.noreply.github.com>
Co-authored-by: RandNMR73 <99706358+RandNMR73@users.noreply.github.com>
2026-05-22 02:59:25 -07:00
Raghav K 68179cd752 [docs] Document enable_torch_compile (+ A/B example) (#1366) 2026-05-22 00:25:42 -07:00
Satyam SrivastavaandSolitaryThinker 2dd5760291 [docs] Document performance benchmark workflow (#1376)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-21 20:25:58 -07:00
Wenxuan TanandSolitaryThinker af2ee9c78a [feat] Optimize distributed weight loading in multi-node training (#572)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-21 16:59:48 +00:00
Satyam SrivastavaandSatyam Srivastava 1c80371b27 [ci] Component time performance + reseed hf baseline skill (#1292)
Co-authored-by: Satyam Srivastava <satyam53@Satyams-MacBook-Air.local>
2026-05-20 13:51:11 -07:00
Junda Su eef473225d [ci] Add Dreamverse Docker image workflow (#1369) 2026-05-20 13:50:02 -07:00
Junda Su 44fb84ef6a [bugfix]: shrink Dreamverse Docker context (#1368) 2026-05-20 13:38:52 -07:00
e8597b7448 [Bugfix] FP4 FA4 installation fix (#1367)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-19 22:45:24 -07:00
Raghav Kandmergify[bot] e2252c0a5e [perf] Mark LayerwiseOffloadHook entry points torch.compiler.disable (remove per-layer graph break) (#1365)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-20 03:09:23 +00:00
Junda Su 72cb427cd9 [feat]: add FastLTX-2.3 Gradio demo package (draft) (#1247) 2026-05-17 23:03:16 -07:00
Mingjia Huo 63030cf6ec [fix] Fix causal self-forcing attention settings (#1355) 2026-05-17 22:59:47 -07:00
Kaiqin Kongandmergify[bot] 773d44b875 [misc] Rename MatrixGame to MatrixGame2 (#1357)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-17 22:57:03 -07:00
1df513922f [feat] Add minimal LoRA finetuning support to the YAML training stack (#1242)
Co-authored-by: alexzms <3036648523@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-17 14:49:13 -07:00
William LinandDavids048 30c45620a2 [infra] [dreamverse]: add instruction to install nasm and update ffmpeg installer to work in plain venv (#1361)
Co-authored-by: Davids048 <jundasu@ucsd.edu>
2026-05-17 01:10:55 +00:00
Shao Duanandklhhhhh 6b2c731596 [feat] eval: add audio metrics (#1352)
Co-authored-by: klhhhhh <1412841649@qq.com>
2026-05-16 14:47:37 -07:00
William Lin e6022c20b2 [misc]: demote ROCm-unavailable startup message to DEBUG (#1360) 2026-05-15 22:12:16 -07:00
460f6e398e [feat]: Add NVFP4QAT linear layer (Attn-QAT 3/12) (#1350)
Co-authored-by: jzhang38 <42993249+jzhang38@users.noreply.github.com>
Co-authored-by: RandNMR73 <99706358+RandNMR73@users.noreply.github.com>
2026-05-15 18:42:02 -07:00
d2ffec5cce [perf] Dreamverse 14/14: Add LTX2 profile speedups (#1337)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-15 18:10:47 -07:00
alexzmsandmergify[bot] cb12e88713 [perf] shallow-copy VSA attn_metadata in train model plugins (#1342)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-15 16:10:12 -07:00
0958c344b8 [feat] Dreamverse 13/14: Activate LTX2 integration (#1336)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-15 15:35:05 -07:00
Junda Su 71dc27ea7b [misc]: Add Dreamverse deploy skill frontmatter (#1353) 2026-05-15 15:23:17 -07:00
Junda SuandSolitaryThinker 1263449d2a [docs] Add copy page action (#1351)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-15 14:59:04 -07:00
b4de5a9f1b [feat] Dreamverse 12/14: Add LTX2 refine and upsampler support (#1335)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-15 12:21:40 -07:00
8acd8e21f9 [feat]: Add NVFP4QAT quantization config (Attn-QAT 2/12) (#1348)
Co-authored-by: jzhang38 <42993249+jzhang38@users.noreply.github.com>
Co-authored-by: RandNMR73 <99706358+RandNMR73@users.noreply.github.com>
2026-05-14 16:53:48 -07:00
d45d82334f [infra] Dreamverse 11/14: Add NVFP4 quantization support (#1334)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-14 16:26:39 -07:00
6392bd40a9 [feat] Dreamverse 10/14: Add serving API contracts (#1333)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-14 10:07:01 -07:00
Shao Duan 17f07bc313 [feat] eval: async VideoPool + metric streamlines (#1320) 2026-05-13 15:57:37 -07:00
William Lin 325861fb99 [misc]: PR-1225 sync — housekeeping (1/12) (#1347) 2026-05-13 15:21:52 -07:00
William Lin c5088670c8 [misc]: empty __init__.py files with no logic (#1346) 2026-05-13 13:09:40 -07:00
0403c6f47e [infra] Dreamverse 09/14: Add Docker and launch scripts (#1332)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 18:01:11 -07:00
a55b1cdde4 [feat] Dreamverse 08/14: Add frontend media and E2E coverage (#1331)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-05-12 17:45:05 -07:00
aec9a20a16 [feat] Dreamverse 07/14: Add frontend session UI (#1330)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 17:22:45 -07:00
b4458e5bae feat: FP4 Flash Attention 4 for Blackwell GPUs (#1221)
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-12 17:04:39 -07:00
Kaiqin Kongandmergify[bot] a790153705 [bugfix] MatrixGame2 SF distillation under gradient checkpointing (#1340)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-12 17:01:04 -07:00
alexzmsandmergify[bot] 0df1445d0d [ci] add GPU model loading tests for fastvideo.train (PR 4/9) (#1274)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-12 16:54:37 -07:00
038da6e02b [feat] Dreamverse 06/14: Add frontend scaffold (#1329)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 16:44:13 -07:00
William Lin 3fb2fbe1a2 [infra]: MagiHuman checkpoint conversion + push scripts (7/8) (#1301) 2026-05-12 15:45:15 -07:00
William Lin acea9d23e1 [docs]: MagiHuman provenance - AGENTS.md, JOURNAL.md, lessons (6/8) (#1300) 2026-05-12 15:40:47 -07:00
William Lin b7a448cf5b [feat]: MagiHuman pipeline orchestrator + 10-test parity battery (5/8) (#1299) 2026-05-12 15:12:54 -07:00
04e32991c6 [feat] Dreamverse 05/14: Add streaming runtime (#1328)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 15:11:18 -07:00
William Lin 424c643b6e [feat]: MagiHuman pipeline stages (4/8) (#1298) 2026-05-12 14:59:21 -07:00
William Lin de803cb250 [feat]: MagiHuman DiT (transformer) port + parity tests (3/8) (#1297) 2026-05-12 14:49:03 -07:00
effc1d3492 [feat] Dreamverse 04/14: Add session and prompt logic (#1327)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 14:45:56 -07:00
William Lin b1ddb1ba33 [feat]: T5-Gemma encoder for MagiHuman pipeline (2/8) (#1296) 2026-05-12 13:59:55 -07:00
9ff65c83b6 [feat] Dreamverse 03/14: Add backend skeleton (#1326)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:58:23 -07:00
15473f3c77 [docs] Dreamverse 02/14: Add app documentation (#1325)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 13:34:42 -07:00
William Lin 490641cf47 [infra]: MagiHuman housekeeping (gitignore, codespell, skills index) (1/8) (#1295) 2026-05-12 12:50:01 -07:00
4ad6880ca6 [docs] Dreamverse 01/14: Add integration provenance (#1324)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-12 11:42:42 -07:00
Raghav K 9d0af28307 [feat] Add Cosmos 2.5 T2W training pipeline (LoRA + full fine-tune) (#1227) 2026-05-11 19:15:00 -07:00
d6dfe95466 [feat] FastVideo World Model Training (#1179)
Co-authored-by: mignonjia <mhuo@ucsd.edu>
Co-authored-by: alexzms <3036648523@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-11 05:13:12 +00:00
alexzmsandmergify[bot] 636d3b743e [misc] attention hot-path cleanup + denoising loop hoists (#1272)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-10 11:51:10 -07:00
William LinandRaghav e3a5c6954f [misc]: import add-model skill stack to .agents/skills/ (#1308)
Co-authored-by: Raghav <ragg04@gmail.com>
2026-05-09 13:06:54 -07:00
f633e30ebb [feat]: add LongCat bidirectional finetuning support (#1244)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: alexzms <3036648523@qq.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-09 12:20:43 -07:00
William Lin 5ce4947ac2 [ci] mergify: accept [skill]/[skills] and [infra] PR title tags (#1309) 2026-05-08 17:14:21 -07:00
William Lin 323d74c0d2 [skills]: New skill - decompose-pipeline-pr (#1303) 2026-05-08 16:05:58 -07:00
William Lin d98aeafc86 [feat]: Loader umbrella-repo support + optional component dirs (#1294) 2026-05-08 15:24:26 -07:00
Shao Duan f6396fb8c6 [feat] Add fastvideo.eval video evaluation suite (#1305) 2026-05-07 20:31:15 -07:00
William Lin 6300329cd5 [infra]: Add activation trace hooks for pipeline debugging (#1293) 2026-05-07 16:38:13 -07:00
MookandSolitaryThinker c17d33bf33 [ci] Replace flaky LTX-2 pixel SSIM with latent-slice cosine regression (#1253)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-05 03:36:58 -07:00
2aaeee2ab8 [feat] Improve API: streaming router (multi-replica load balancer + ws proxy) (#1286)
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 03:00:08 -07:00
eb3a394224 [feat] Improve API: streaming auxiliaries (safety, rewrite, logger, mock) (#1284)
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-05 00:14:34 -07:00
f673423b51 [feat] Improve API: streaming prompt enhancer with LLMProvider abstraction (#1258)
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-04 13:44:40 -07:00
eb0a41528a [feat] Improve API: streaming server GpuPool + worker subprocess (#1257)
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
2026-05-04 12:56:31 -07:00
William Lin 140bd1a6cf [misc]: standardize install instructions on uv pip install (#1279) 2026-05-02 12:45:50 -07:00
William Lin 11f5a8e582 [misc] pin torch to 2.11.0 (#1277) 2026-05-02 11:48:07 -07:00
71b3cb8c34 [ci] Add CI Performance Regression Tracking Changes (#1248)
Co-authored-by: Satyam Srivastava <satyam53@Mac.lan1>
Co-authored-by: Satyam Srivastava <satyam53@Satyams-MacBook-Air.local>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-05-02 03:29:53 -07:00
William Lin c85f6a477f [docs] add hierarchical AGENTS.md per-directory guidance (#1278) 2026-05-02 03:28:18 -07:00
Junda Su 40d4930d73 [bugfix] Update fa import (#1271) 2026-05-02 01:25:22 -07:00
William Lin f9be085243 [ci] pre-commit: drop stale excludes + document agent lint flow (#1276) 2026-05-02 01:19:06 -07:00
William Lin 36b53ff350 [bugfix]: classify stable_audio fields in schema parity inventory (#1275) 2026-05-02 00:12:10 -07:00
William Lin 9801037c3d [refactor] tests/local_tests: organize by model family (#1269) 2026-05-01 01:49:54 -07:00
alexzmsandmergify[bot] 74d09b0efd [misc] cleanup: grad-norm asserts, dead offload file, callback names (#1268)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-05-01 01:16:13 -07:00
alexzms 38dc8820ac [ci] add CPU unit tests for train callback system in fastvideo.train (#1267) 2026-05-01 00:53:29 -07:00
William Lin c77a76c6af [feat] Stable Audio Open 1.0: T2A + A2A + RePaint inpainting (native) (#1260) 2026-05-01 00:07:11 -07:00
alexzms d14d5aadea [feat] Cosmos 2.5 training support in fastvideo.train (#1224) 2026-05-01 01:15:02 +00:00
alexzms 4c915b7742 [ci] add CPU unit tests for train checkpoint utilities in fastvideo.train (#1265) 2026-04-29 18:55:39 +00:00
alexzms 9a8bbe18fa [bugfix]: fix SP deadlock in negative prompt encoding during training (#1178) 2026-04-28 01:06:49 +00:00
alexzms ea25441ef0 [ci] add CPU unit tests for fastvideo.train load_run_config (#1264) 2026-04-28 01:06:18 +00:00
48957fcde1 [bugfix] Fix modal remote functions crash container on sys exit in CI remote functions (#1261)
Co-authored-by: Satyam Srivastava <satyam53@Mac.lan1>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-27 21:50:55 +00:00
Mook 7b872cc41e [Perf] Skip bool-mask round-trip in block-sparse VSA attention (#1243) 2026-04-26 15:14:37 -07:00
alexzms 37418946c8 [docs]: clarify real_score_guidance_scale CFG parameterization (#1256) 2026-04-26 16:38:00 +08:00
William Lin 95fd29e0cb [feat] Streaming WebSocket server skeleton (single generator + fMP4) (#1251) 2026-04-26 00:33:49 -07:00
Junda Suandmergify[bot] e17cd2633c [bugfix]: normalize uint8 pil_image in I2V VAE encoding (#1249)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-24 09:16:01 +00:00
William Lin e0dc5f2b0c [feat] Add typed LTX-2 continuation state and streaming session store (#1250) 2026-04-24 01:28:07 -07:00
William Lin 70ee5d230c [feat] [6/n] Improve API: LTX-2 public preset + asset wiring + gpu_pool translation (#1239) 2026-04-23 11:36:45 -07:00
William Lin 24ced500f5 [test] add LTX-2 distilled T2V SSIM regression test (#1240) 2026-04-21 12:03:38 -07:00
William Lin 4ddcdf541f [feat] [5.5/n] Improve API: streaming server config surface + serve dispatch (#1238) 2026-04-17 15:36:21 -07:00
William Lin 0e3529869c [feat] [5/n] Improve API: wire ServeConfig.default_request into OpenAI serving (#1237) 2026-04-17 13:26:18 -07:00
William Lin e1e0d91c00 [misc] small cleanup for API handling (#1235) 2026-04-16 16:21:21 -07:00
William Lin 145a3f166b [feat] [4/n] Improve API: refactor sampling param and merge with presets (#1234) 2026-04-16 14:10:02 -07:00
William Lin 88a5a933ab [feat] [3/n] Improve API: extend support to cli (#1226) 2026-04-14 15:20:47 -07:00
William Lin c591d6d2a6 [feat] [2/n] Improve API: add initial support in video_generator (#1220) 2026-04-06 10:33:54 -07:00
Kun Linandmergify[bot] 65dff806a8 [bugfix]Fixing Lora distillation training distributed checkpointing bug (#1192)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-06 02:20:26 +00:00
KUAN-HAO HUANGandmergify[bot] b85f0f4c2a [perf]: Eliminate CPU-GPU synchronization bottlenecks in training pipeline (#1217)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-06 02:03:46 +00:00
William Lin 76c62d7a00 [feat] [1/n] API improvements: add intial files for new fastvideo public API (#1218) 2026-04-05 18:13:19 -07:00
f6e65ff668 [Feature] Add BSA (Bidirectional Sparse Attention) inference backend (#1174)
Co-authored-by: Satyam Srivastava <satyam53@Mac.lan1>
Co-authored-by: Satyam Srivastava <satyam53@Satyams-MacBook-Air.local>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-05 05:00:33 +00:00
mergify[bot] c220aa8000 [ci](mergify): upgrade configuration to current format (#1216)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-04-04 23:09:17 +00:00
Jinzhe PanandDarren Sadr 4713fc17ed [feat] Job Runner UI (#1189)
Co-authored-by: Darren Sadr <darrensadr@gmail.com>
2026-04-02 16:07:24 -07:00
vishruthb 5789955bbe [feat] add gen3c (cosmos-7b) model and pipeline support (#1059) 2026-04-01 11:42:02 +00:00
Jinzhe Pan 2ad84a3b78 [ci] Use update instead of rebase for auto branch sync (#1215) 2026-04-01 19:16:59 +08:00
Jinzhe Pan 12d699cd78 [ci] Add direct test retry with check overwrite and aggregate status refresh (#1214) 2026-04-01 17:21:28 +08:00
Jinzhe Pan 34f14ded21 [ci] Use pull_request_target for Full Suite trigger (#1213) 2026-04-01 03:01:07 +08:00
Jinzhe Pan 71d1ab411f [ci] Fix jq crash when Buildkite build env is null (#1212) 2026-04-01 02:35:01 +08:00
Jinzhe Pan 805e487773 [ci] Ignore legacy reference videos when checking for HF download (#1211) 2026-04-01 02:12:09 +08:00
Jinzhe Pan 8803b4547e [ci] Add retry for flaky tests and fix stale SSIM references (#1210) 2026-04-01 01:11:49 +08:00
Jinzhe Pan 3b3806b3f6 [ci] Fix /merge to directly trigger Full Suite + simplify rebase conditions (#1209) 2026-03-31 23:17:09 +08:00
Jinzhe Pan 38d962e89d [ci] Remove Mergify ready-label race condition (#1208) 2026-03-31 20:59:13 +08:00
Jinzhe Pan 3966a365d0 [ci] Add statuses:write permission for /test pre-commit (#1207) 2026-03-31 20:33:18 +08:00
Jinzhe Pan d73fd14af0 [ci] Post pre-commit status to PR commit SHA (#1206) 2026-03-31 20:21:21 +08:00
Jinzhe Pan a87cc89916 [ci] Trigger pre-commit on /test slash commands (#1205) 2026-03-31 20:12:57 +08:00
Jinzhe Pan 81fd80c8ee [ci] Add TEST_SCOPE routing for clean single-test execution (#1203) 2026-03-31 19:40:59 +08:00
Jinzhe Pan ff22439f28 [ci] Fix fork PR checkout for /test and Full Suite triggers (#1202) 2026-03-31 13:42:13 +08:00
Jinzhe Pan de0de04212 [ci] Replace Merge Queue with auto-merge — reduce CI complexity (#1200) 2026-03-31 10:09:33 +08:00
mergify[bot] 7f2c3e1f64 [ci](mergify): upgrade configuration to current format (#1194)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-03-31 00:47:48 +08:00
Jinzhe Pan ab55e57c22 [ci] Fix Merge Queue requeue and draft PR pre-commit skip (#1197) 2026-03-30 22:35:13 +08:00
Jinzhe Pan 46f6b43a53 [ci] Fix Merge Queue immediate dequeue (#1196) 2026-03-30 21:33:45 +08:00
Jinzhe Pan 833a33b663 [ci] CI follow-up: gate checks, issue label unification, draft PR skip (#1193) 2026-03-30 20:19:46 +08:00
Jinzhe Pan 9ea1307cd4 [ci] Add approval and pre-commit checks to merge protections (#1190)
## Summary

Follow-up to #1187. Two small changes:

1. **Merge Protections expanded** — adds `#approved-reviews-by>=1` and `check-success~=pre-commit` to `merge_protections` so the Mergify check shows a unified requirements checklist on every PR (title format + approval + pre-commit), instead of only showing the title format.

2. **Buildkite pipeline comment fix** — updates the outdated Full Suite section comment from "Triggered by adding the 'ready' label via GitHub Actions → Buildkite API" to reflect the new Merge Queue trigger path.
2026-03-30 05:14:49 +00:00
Jinzhe Pan be35003cb1 [ci] Merge Queue, label system overhaul, and slash commands (2/2) (#1187) 2026-03-30 08:22:35 +08:00
Jinzhe Pan 26bd4db253 [ci] CI infrastructure cleanup and workflow reorganization (1/2) (#1186) 2026-03-29 17:01:09 -07:00
Jinzhe PanandWill Lin e294ca011c [feat]: overhaul SSIM test infrastructure — partition scheduling, helper migration, CI fixes (#1185)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-03-29 23:59:25 +00:00
alexzms 2085a4fc4a [bugfix]: fix VAE temporal tiling blend corruption in tiled_encode (#1181) 2026-03-29 23:12:42 +00:00
Jinzhe Pan c0c8e39c04 Revert "[feat] Job Runner UI" (#1188) 2026-03-29 16:46:27 +08:00
Darren f72618dafb [feat] Job Runner UI (#1172) 2026-03-29 16:17:56 +08:00
alexzms b3edfacdd8 [bugfix]: fix I2V preprocessing crash for models without CLIP (Wan2.2 I2V) (#1184) 2026-03-28 10:28:50 +08:00
alexzms 30129a3350 [misc]: reorganize training configs and add documentation (#1177) 2026-03-26 16:38:31 -07:00
jaisurya27 4d49f7b0aa Kandinsky5 lite dit clean (#1088) 2026-03-26 08:07:11 +08:00
alexzms 71bfc13d75 [feat]: add HunyuanVideo model plugin for fastvideo/train framework (#1175) 2026-03-24 16:49:51 -07:00
Kaiqin Kong 74db6e18d1 [misc] update action loading in validation and preprocess (#1143) 2026-03-24 15:10:03 -07:00
Kaiqin Kong 7d263c6a36 [bugfix] self-forcing train/validation step mismatch (#1173) 2026-03-20 00:52:45 -07:00
Zhang Peiyuan 454c32d1d1 Update README.md 2026-03-17 14:26:07 -07:00
Jinzhe Pan d1240b9238 [CI] add contributor interaction automation (#1170) 2026-03-17 12:05:40 +08:00
Hao Zhang 4105094fa5 [docs] Update README with realtime demo announcement (#1169) 2026-03-13 15:52:02 -07:00
alexzms f036469d3d [feat]: Knowledge Distillation training method for ODE-init (KDMethod + KDCausalMethod) (#1166) 2026-03-11 20:43:24 -07:00
alexzms 14261bc98c [feat] pre-commit support 120 col num (#1167) 2026-03-11 20:19:30 -07:00
alexzms d92858659d [feat] Self-Forcing methods in refactored training infra (#1164) 2026-03-09 20:20:59 -07:00
1a383f3f66 [refactor] train v1 clean up
Co-authored-by: alexzms <3036648523@qq.com>
Co-authored-by: Peiyuan Zhang <a1286225768@slurm-h200-204-215.slurm-compute.tenant-slurm.svc.cluster.local>
2026-03-09 18:58:56 -07:00
alexzmsandPeiyuan Zhang bc27a032c5 [feat] Refactor training framework into fastvideo/train (#1159)
Co-authored-by: Peiyuan Zhang <a1286225768@gmail.com>
2026-03-09 15:16:42 -07:00
alexzms 2b13e117f0 [Feat] Add causal Wan pipeline with multi-step denoising (#1161) 2026-03-08 13:32:04 -07:00
Junda Chen 99c166c381 feat: Building agent friendly repo (#1151) 2026-03-07 17:46:29 -08:00
XOR-op 95066245db [misc] FlashAttention 4 support (#1114) 2026-03-07 16:53:43 -08:00
Jinzhe Panandgemini-code-assist[bot] 6dcaac768b [CI] PR template (#1157)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-03-07 11:51:28 -08:00
Ajay Anubolu 02c1c49b75 [CI] Add inference performance regression tests (#1140) 2026-03-07 08:26:54 +08:00
Zhang Peiyuan cd1b7cf139 [Refactor] SP Mask --> original seq len; HunyuanVideo 1.5 does not need mask (#1142) 2026-03-04 11:44:47 +08:00
Ajay Anubolu e63b7d8ac4 [Feat] Added OpenAI-compatible API server and benchmark script (#1109) 2026-03-02 17:12:32 -05:00
Jinzhe Pan 5190c1bb1e [Doc] add doc for inference architecture (#1147) 2026-03-02 13:44:27 -08:00
Darren 2cb3bba658 [bugfix]: fix a bug where collect_env was not running properly... (#1145) 2026-03-02 10:57:29 -08:00
Jinzhe Pan f9e1c46c3c [CI][Feat] launch 2 instance to run ssim (#1137) 2026-03-01 01:49:29 -08:00
Peiyuan Zhang e1eda47589 remove temporal frame adjustment 2026-02-27 20:47:16 +00:00
Zhang Peiyuan d902967208 Py/fix sp (#1138) 2026-02-27 12:14:44 +08:00
Zhang Peiyuan fea556269b [Misc] Fix memory leakage in VideoGenerator (#1132) 2026-02-26 19:51:32 -08:00
William Lin 69dd3c68f6 [bugfix] fix matrix game kv indexing and CI (#1135) 2026-02-26 01:24:42 -08:00
Jinzhe Pan 5433f6e80b [fix] preprocessing issue (#1134) 2026-02-25 21:49:52 -08:00
Junda (David) Su e315657066 [docs] [kernel] Migrate to uv (#1127) 2026-02-25 14:13:50 -08:00
Zhang Peiyuan f8d9a0c57f [misc] fix hunyuan (#1125) 2026-02-25 08:26:29 +08:00
Jinzhe PanandWill Lin fa6d276925 [Feat] Improved CI (#1119)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-02-24 12:53:02 -08:00
Zhang PeiyuanandWill Lin fc80d95d7e [Misc] Remove STA (#1124)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-02-23 15:14:42 -08:00
Shao DuanandSolitaryThinker 37cab18780 [bugfix] Added ltx2 guidance missing modulation term (#1100)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-02-23 14:22:31 -08:00
Zhang Peiyuan 128d0b7fc5 [Misc] Remove Teacache (#1121) 2026-02-22 16:53:07 -08:00
Matthew Noto 8092f02e6d small refactor in post-processing to improve efficiency (#1123) 2026-02-22 16:45:44 -08:00
Zhang Peiyuan 03d9ce2edb [Misc] Remove StepVideo (#1118) 2026-02-21 17:15:42 -08:00
10fc92dba5 Upstream LTX2 Training (#1116)
Co-authored-by: RandNMR73 <notomatthew31@gmail.com>
Co-authored-by: JerryZhou54 <zhouw.jerry2017@outlook.com>
Co-authored-by: Davids048 <jundasu@ucsd.edu>
Co-authored-by: Peiyuan Zhang <a1286225768@slurm-h200-204-239.slurm-compute.tenant-slurm.svc.cluster.local>
2026-02-21 16:06:55 -08:00
6736dc06a5 Improve Docs (#1112)
Co-authored-by: Peiyuan Zhang <a1286225768@slurm-h200-204-227.slurm-compute.tenant-slurm.svc.cluster.local>
Co-authored-by: Peiyuan Zhang <a1286225768@slurm-login-0.slurm-login.tenant-slurm.svc.cluster.local>
2026-02-19 14:23:32 -08:00
William Lin 8c002c62af [misc] add hy-world link to readme (#1113) 2026-02-18 12:01:10 -08:00
Darren 7061313d04 [bugfix] get_torch_device and other device calls were being made on non-cuda platforms (#1107) 2026-02-18 11:43:46 -08:00
Zhang Peiyuanandgemini-code-assist[bot] 76d3ba69e0 [Misc] clean up VSA finetuning examples. (#1111)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-18 11:37:20 -08:00
8e39ce38c9 [Feat] Native dit implementation for SD3.5 (#1093)
Co-authored-by: Ishan Vaish <vaish.ishan@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
2026-02-18 10:19:40 +08:00
Darren d4bd8bf2c0 Update README.md (#1110) 2026-02-17 14:54:34 -08:00
Darren e83d7bc50c [bugfix] fix import PreTrainedModel in stepllm.py (#1108) 2026-02-16 21:45:08 -08:00
Jinzhe Pan ff3d5aff75 [Fix] hunyuan postprecessing issue (#1104) 2026-02-15 12:26:35 -08:00
XOR-op 959dbcc8a2 [perf] causal MatrixGame optimization (#1078) 2026-02-15 09:51:00 +08:00
William Lin 36bf37e9ba [bugfix] Fix failed kernel publish and SFT regressions (#1103) 2026-02-14 16:19:17 -08:00
Mihir Jagtap 7a83e0e6fc [feature] Add Hunyuan-GameCraft model support (#1071) 2026-02-14 08:07:44 +08:00
William Lin 8be1313b86 [kernel] add torch 2.10 to package build matrix (#1099) 2026-02-13 13:03:05 -08:00
alexzms d925ad05f3 [bugfix] fastvideo-kernel: fix VSA Triton padding NaNs and support q/kv length mismatch (#1094) 2026-02-13 12:39:49 -08:00
Shao DuanandWill Lin 7f795600c8 [bugfix] Fixed ltx2 base cfg guidance (#1095)
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-02-13 11:53:11 -08:00
William Lin ec16b6b01d [misc] update wechat group link (#1098) 2026-02-13 01:19:11 -08:00
Kaiqin Kong 31c0f1b341 [feat] Port LingBot-World-Base (Cam) (#1081) 2026-02-10 11:12:33 -08:00
William Lin 4bee0fa199 [misc] cleanup assets/ and demo/ (#1091) 2026-02-10 02:26:09 -08:00
530e6b8363 [Model] LTX 2 Base (#1064)
Co-authored-by: Davids048 <jundasu@ucsd.edu>
Co-authored-by: Will Lin <wlsaidhi@gmail.com>
2026-02-10 01:11:17 -08:00
Jinzhe Pan 9ab2725db1 [ci] CI Transformer Tests (#1089) 2026-02-10 01:08:59 -08:00
IshanandJinzhe Pan 0aff68f51d [Feat] Add Stable Diffusion 3.5 (#1075)
Co-authored-by: Jinzhe Pan <eigensystem1318@gmail.com>
2026-02-10 14:31:36 +08:00
ad58f802f3 [Feat] Port LTX2 trainer (#1074)
Co-authored-by: Davids048 <jundasu@ucsd.edu>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
2026-02-09 17:32:57 -08:00
Wei Zhou 04fa356ee3 [Misc] [Training] Fixed a bunch of bugs in current training pipeline (#1084) 2026-02-09 16:01:05 -08:00
Matthew Notoandgemini-code-assist[bot] f9c076fe2b [misc] add AGENTS.md file (#1085)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-02-09 01:03:49 -08:00
2062 changed files with 297233 additions and 1365935 deletions
+207
View File
@@ -0,0 +1,207 @@
# v2 ← M\*: Architecture Gap-Analysis & Improvement Roadmap
**Status:** exploration, flagged for review. **Date:** 2026-06-19.
**Source paper:** *M\*: A Modular, Extensible, Serving System for Multimodal Models* (arXiv 2606.12688,
Stanford/UW/CMU; Jha, Sagan, Kamahori, …, Kasikci, S. Wang). It is a universal serving runtime for composite
multimodal models built on the **Walk Graph** abstraction (a model is a dataflow graph `G`; a request is a
*Walk* — a labeled subgraph — and the runtime executes walks). It beats vLLM-Omni (~20% lower T2I latency on
**BAGEL**, up to 2.64× on I2I), SGLang-Omni (2.7× TTS throughput on **Qwen3-Omni**), and native V-JEPA2
rollout (12.5×). It explicitly names **FastVideo's own** sparse/sliding-tile attention, xDiT/PipeFusion/USP,
Inferix, and FlashDrive as techniques integratable into the graph runtime.
**Method:** a 28-agent workflow — 6 parallel v2-subsystem maps → 10 M\*-dimension analyses, each
*adversarially verified against the actual v2 code* → synthesis + a completeness critic. The critic's
corrections and three P0 claims were then **spot-verified by hand** (file:line below). This doc folds those
corrections in; it is the corrected, authoritative synthesis.
---
## 1. Executive summary
v2 already implements the **harder half** of M\*'s thesis and in several axes **exceeds** it:
- v2's `Program` *is* M\*'s graph `G` (typed `ComponentNode`/`ModelLoopNode` + edges).
- v2's `shared_weight_components` *is* M\*'s cross-Walk node sharing — BAGEL/Cosmos3/LTX2 each bind two
`ModelLoopNode`s to **one resident transformer** (`instance.component()` returns the same live object). This
is the exact MoT serving property the omni cards in this repo already express.
- v2 adds three things M\* (serving-only) has **no equivalent for**: a required+validated per-loop **cost
model**, a non-negotiable **interleave bit-parity gate**, and an **integrated training plane** (RL→distill
flywheel driving the *same* serving Loop).
- The `extend/` plugin seam (interceptors/observers/registry with capability negotiation) is precisely the
hook M\*'s "extensible / integrate FastVideo-STA, xDiT, Inferix, FlashDrive" call-out asks for — **v2
already has the seam M\* only gestures at.**
What v2 lacks is M\*'s **declarative authoring layer above the substrate**, and — the key insight — *much of
that substrate is already authored but inert*: v2 has declared the metadata for "minimum components per
request" (`required_for`/`optional_for` on every omni card) and "branch as a cache axis" (`guidance_sig`,
`CacheKey`) but **never wired it to an executor**. The substrate is ~80% built and switched off.
**Highest-leverage cluster:** three small, parity-safe wires that turn on inert substrate and unblock the
BAGEL/Qwen-Omni/Cosmos3 latency wins M\* measured **on the exact models this repo already runs** — plus one
P1 that aligns v2 with the paper's headline "extensible" claim using a seam v2 already has.
### Verified P0 correctness findings (spot-checked by hand)
1. **Runner divergence (real bug).** `v2/runtime/engine.py:88` → `nodes = self.program.nodes`;
`v2/runtime/disaggregated.py:96` → `nodes = self.program.active_nodes(self.request)`. The inline and
disaggregated runners execute *different node sets*. ✅ confirmed.
2. **EOS is faked.** `v2/recipes/omni/ar_loop.py` docstring says "done on EOS/max_tokens"; `next()` (`:46-48`)
checks **only** `max_tokens`. M\*'s marquee `DynamicLoop` use case (EOS) is unimplemented in the loop that
serves the Qwen-Omni Thinker/Talker and Cosmos3 reasoner. ✅ confirmed.
3. **`required_for`/`optional_for` have zero runtime consumers** (grep outside `specs.py`/recipes/tests is
empty). The min-components metadata is declared on every card and never read. ✅ confirmed.
---
## 2. Dimension table (corrected)
| # | Dimension | v2 status | Gap | Priority | Effort | Payoff | Action |
|---|---|---|---|---|---|---|---|
| 1 | Min-components per request (`required_for` + `when_task`) | substrate built, **inert** | real, cheap | **P0** | S | Consume `required_for` in `active_nodes`; unify `engine.py:88` onto `active_nodes`; deliver via registry/card builder so all ~40 cards inherit it |
| 2 | Real EOS + declarative `DynamicLoop` | early-exit emergent; **EOS faked** | real | **P0** | S | `ARDecodeLoop` honors `eos_id` + `req.sampling.stop`; add `LoopSpec.dynamic_stop` + `register_loop_stop`. **Training-enabling** (world-model rollout horizon) |
| 3 | CFG/branch as label over one paged KV pool | absent (`PagedKVCache` is a counter) | real | **P1** | L | `(namespace,label)` paged store w/ one budget; reuse `guidance_sig` for hash (NOT `partition_field`); by-ref via existing `InProcKVConnector`. AR path only (diffusion has no KV) |
| 4 | `extend/` plugin seam → integrate FastVideo-STA / Inferix | **seam exists, unused for attn** | real (paper headline) | **P1** | M | Expose FastVideo sparse/sliding-tile attention + Inferix block-diffusion as `Interceptor`/`EngineKind` plugins — the paper's named integration targets, on this repo's own code |
| 5 | `ParitySpec.output_determinism` (C3 distributional) | C3 rung defined, **0 users** | real, dormant | **P1** | S | Add field; `compare_outputs` consults it. **Training-enabling** (SDE/FlowGRPO stochastic rollouts) |
| 6 | Registry-driven delivery of #1 | present, not leveraged | integration | **P1** | S | Express `when_task`/min-components through `WorkflowRegistry`/card builders, not 3 bespoke recipe patches |
| 7 | Serving conductor + pluggable data plane | conductor exists (`serving/http.py`); **single-process transport** | real | **P2** | L | v2 already has the step-scheduled worker surface; gap is ZeroMQ/Mooncake + direct worker→worker tensor routing (today `InProcKVConnector` only) |
| 8 | Fleet/Dynamo placement + replicas | **live** (`deploy/fleet.py`,`dynamo.py`) | partial | **P2** | M | Fleet-level placement/affinity/replica is real & ≥M\*; missing piece is only the intra-engine `(node,Walk)→rank` map decoupled from model code |
| 9 | Per-node TP / SP + cross-rank transport | axis vocab **exists** (`sp` incl.); not wired to runtime | partial | **P2** | XL | Wire declarative degrees into runtime; Wan/LTX are **SP-native** (TP is a no-op there); populate `parallel_plan_hash` on the serving cache path |
| 10 | Named Walks + per-model state machine | `Program`=G, sharing real; no Walk/SM | real | **P2** | M | Defer until a *re-entrant* phase graph (Thinker↔Talker, rollout) needs it; #1 captures the min-components win without it |
| 11 | Declarative `Parallel/Sequential/Loop` IR | imperative loop classes | real (authoring) | **P2** | M | Thin Section IR lowering to flat `Program`; scope to one AR recipe |
| 12 | Streaming `ChunkPolicy` + `StreamBuffer` | causal-chunk emit **already ships** (`wan_causal`); `EdgeKind.STREAM` inert | real | **P2** | L | Declarative `ChunkPolicy` vocab over the existing chunk mechanism; needs concurrent producer/consumer runner (= pipelined scheduling). Inferix integration point |
| 13 | Speculative deferred-termination; loop-spanning CUDA graphs; N+1 prefetch; attn double-buffer | absent / per-step capture (14 cards) | real | **P3** | L | Gate behind a real GPU executor; unobservable on CPU-toy CI; loop-span needs an `allows_interleaving=False` carve-out |
| — | Cost model + interleave/consistency parity | **exceeds M\*** | none | **guard** | — | Do not regress; keep `step_cost_model` mandatory + `bit_identical` default |
| — | Integrated training plane (flywheel, weight-sync) | **exceeds M\*** | none | **guard** | — | Protect train==serve loop identity with a toy fixture |
---
## 3. P0/P1 deep-dives (sequenced)
```
PR-1 (P0) min-components ──┐
PR-2 (P0) real EOS ─┼─► prereqs for honest "DynamicLoop" + min-component claims; both training-enabling
PR-3 (P1) output_determinism (independent)
PR-5 (P1) extend/ plugin: FastVideo-STA / Inferix as Interceptors (independent; highest paper-alignment)
PR-4 (P1) CFG-as-label paged pool ──► depends on PR-2 (AR loop is the only KV consumer)
```
PR-1, PR-2, PR-3, PR-5 are mutually independent; PR-4 depends on PR-2.
### PR-1 (P0) — Turn on the inert min-components substrate + fix runner divergence
- **Change.** Extend `Program.active_nodes(request)` (`v2/program/specs.py`) to also drop any node whose bound
`ComponentSpec.required_for` (`v2/card/specs.py:144`) excludes `request.task` (and isn't in `optional_for`).
**Fix the bug:** change `v2/runtime/engine.py:88` to `nodes = self.program.active_nodes(self.request)` so the
inline `ProgramRunner` matches `DisaggregatedRunner` (`disaggregated.py:96`). Deliver the `when_task` gating
through the **registry/card builder** (`recipes/__init__.py`, `program/workflow.py:WorkflowRegistry`) so all
~40 cards inherit it uniformly — not three bespoke `program.py` patches.
- **Why (this repo's models).** BAGEL T2I currently steps the AR-text loop and Cosmos3 t2v materializes the
reasoner even though the cards declare `transformer required_for={'reason','t2i'}`, `vae required_for={'t2i'}`.
On the GPU backend that is wasted resident-weight load + wasted steps on every single-modality request —
exactly M\*'s "execute the MINIMUM components per request," delivered by consuming existing metadata.
- **Risk/invariant.** Validate in `ModelCard.validate()` that every active node's `reads` are produced by an
active node for each declared `TaskType` (avoid dropping a producer). Pure node-id filtering ⇒ serial and
interleaved still walk the same filtered list ⇒ §9.3 interleave bit-parity holds by construction. CPU-toy clean.
### PR-2 (P0) — Real EOS + declarative `dynamic_stop` *(also training-enabling)*
- **Change.** In `v2/recipes/omni/ar_loop.py`, `advance()` reads the emitted token; if it equals the model
`eos_id` (toy backend exposes `EOS=0`) or matches `req.sampling.stop` (`params.py:21`, currently dead),
register termination; `next()` returns `Done()` on stop OR `max_tokens`. Add `StopRegistry` to `LoopState` +
`register_loop_stop(name)` to the `LoopContext` protocol (`contracts.py:204`) and to
`DisaggregatedRunner`'s `RuntimeLoopContext`. Add `LoopSpec.dynamic_stop: bool=False`, opt the AR cards in.
- **Why.** The docstring-vs-code lie sits in the loop serving Qwen-Omni Thinker/Talker and the Cosmos3 reasoner;
M\*'s second named `DynamicLoop` use case (world-model **rollout horizon**) is exactly what `self_forcing` RL
needs — so this is both a serving-credibility fix and a training enabler (raise its payoff accordingly).
- **Risk/invariant.** `dynamic_stop=False` is byte-identical back-compat. Must pass **all three** parity gates:
serial==interleaved AND disaggregated==inline. **Not** in this PR: speculative deferred-termination (unobservable
on CPU-toy, fights the interleave invariant — P3, gated on GPU executor).
### PR-3 (P1) — `ParitySpec.output_determinism` (close the dormant C3 hole) *(training-enabling)*
- **Change.** Add `output_determinism: str = "bit_identical"` to `ParitySpec` (`card/specs.py:88`); make
`compare_outputs` (`parity/interleave_gate.py:54`) consult it (`bit_identical` → today's exact check;
`distributional` → a moment/tolerance check — land a simple moment match first; a real KS test is new code).
- **Why.** `ConsistencyLevel.C3` is defined and used by zero recipes; an SDE/FlowGRPO stochastic rollout cannot
honestly declare its parity contract and would falsely fail the bit-identical gate. Additive; default unchanged.
### PR-5 (P1) — Expose FastVideo's own attention + Inferix as `extend/` plugins *(highest paper-alignment)*
- **Change.** Use the existing `extend/{interceptors,observers,registry}.py` seam (capability-negotiated, with
per-(request,branch) `plugin_state` that already passes the interleave gate) to register FastVideo's
sparse/sliding-tile attention and Inferix-style block-diffusion as `Interceptor`s / an `EngineKind` plugin.
- **Why.** M\*'s title is "Modular, **Extensible**" and it explicitly lists FastVideo-STA, xDiT/PipeFusion/USP,
Inferix, FlashDrive as integratable. v2 already has the seam M\* only describes — this is where v2 most
directly answers the paper, using this repo's own attention code. Low risk (the seam + capability negotiation
already exist and are tested).
### PR-4 (P1) — CFG/branch as a LABEL over one paged KV pool
- **Change.** Rewrite `PagedKVCache` (`cache/classes.py:155-172`) from a block *counter* into a real
`(namespace,label)->[block-handle]` store with **one shared `total_blocks` budget** (M\*'s single-pool
property). Reuse the existing-but-unpopulated `CacheKey.guidance_sig` (`keys.py:53`) for the hash. Thread the
label through `ar_loop.py` (alloc/append/get per `(request_id, branch)`; prefill once per shared-prefix label;
combine via `CFGPolicy.combine`). Wire `ResourceRequest.cache_blocks` (`contracts.py:64`, zero consumers) into
admission per (class,label).
- **Why.** The dossier-identified driver of M\*'s BAGEL win (3 CFG contexts as 3 labels over ONE pool vs dense
per-context). Targets AR_DECODE (BAGEL `generate_text`, omni Thinker); **correctly excludes diffusion**
(Wan/LTX are bidirectional, no KV — their CFG stays dense-but-batched).
- **Corrections to bake in.** Do **NOT** add `branch_label` to `CacheKey.partition_field()` (CFG branches share
embeddings; partitioning by branch is a semantic bug). Do **NOT** add a new by-ref type — reuse
`InProcKVConnector` + `TransferManifest.cache_key`. Wiring `cache_blocks` admission is greenfield ⇒ effort **L**.
CPU version proves label/sharing semantics; the real latency win needs a FlashInfer paged kernel (out of scope)
— **merge** with a future "real KVCacheEngine" effort rather than landing isolated.
---
## 4. What v2 already does ≥ M\* — do NOT regress
1. **Required+validated cost model** on every `LoopSpec` (13-kind `WorkUnitKind`) — typed, pre-GPU-validated.
2. **Interleave bit-parity as a hard gate** (`parity.interleave_required=True` on 40+ cards). M\* has no such
gate (its speculative scheduling deliberately wastes steps). Load-bearing invariant; every new primitive
must pass it.
3. **C0–C4 consistency ladder** wired into RL methods, with first-divergence tap reporting. No M\* equivalent.
4. **Integrated training plane** — DiffusionNFT/DMD2/self_forcing, RL→distill flywheel, `WeightSyncController`
hot weight-sync with drain-to-boundary + scoped cache invalidation, driving the **same** serving Loop.
M\* is serving-only. Protect with a toy fixture asserting `rollout_loop` drives the served Loop object.
5. **CPU-toy parity for the whole stack** — loops/CFG/caches/parity/RL run in CI without a GPU. Every new
primitive must ship a toy exercise (this is what makes all PRs above testable without H100s).
6. **Partition-not-flush cache invalidation** + four independent per-class pools.
7. **`extend/` plugin seam** with capability negotiation (a 4-step distilled card *rejects* a residual-skip
interceptor) — M\* describes extensibility; v2 has the mechanism.
8. **Dynamo citizenship** (`deploy/dynamo.py`: one `DeploymentCard`+cost model, two consumers) — beyond M\*'s
self-contained runtime.
---
## 5. Dropped / merged / deferred (and why)
- **DROP declarative `Parallel` as a CFG-execution win.** The runner walks nodes linearly (ignores
`Program.edges`), so `Parallel` lowers to sequential sugar and the CFG 3-pass braid is already one
co-scheduled `WorkPlan.run`; splitting it risks the interleave gate. Salvage only the no-op refactor
extracting `branch_forward` from `WanDenoiseLoop._velocity`. Reassign `Parallel` to the placement workstream.
- **MERGE the full Walk/state-machine layer** into "defer until a re-entrant phase graph needs it" (PR-1 gets the
min-components win with ~20 lines, no new abstraction). If built: the validator must check a walk's node-id
order is a *subsequence* of `program.nodes` (not just membership) or the runner can reorder and break parity.
- **MERGE `StreamBuffer`/`ChunkPolicy` into pipelined-scheduling.** Causal-chunk emit *already ships*
(`wan_causal/loop.py` per-chunk `StepResult.emit` + slab-KV); the gap is the declarative `ChunkPolicy` vocab
+ a concurrent producer/consumer runner. If built: keep all policies pure (per-request `StreamBuffer` history,
not shared edge state) and restrict the bit-identical claim to the token-only handoff.
- **MERGE CFG-fan-out exec + cross-rank transport + PD loop-splitting into a multi-GPU-runtime program.** These
need real collectives (`v2/distributed/` is a stub) and KV-by-reference (KV lives in `CacheManager`, not the
transferable `slots`). **Keep cheaply now:** the *declarative* halves — per-component degree, `(node,Walk)`
placement key with node-only fallback, `ReplicaSet` under `LocalFleet`, and populate `parallel_plan_hash` on
the **serving** cache path (it is already populated in `training/behavior.py:40` — the gap is serving-only).
- **DEFER** speculative deferred-termination, loop-spanning CUDA graphs, N+1 prefetch, attention-plan
double-buffer — all gated on a real GPU executor; benefit unobservable on CPU-toy CI. Keep the cheap
`EngineKind` tag (`STATELESS|KV_CACHE|DIFFUSION`) now. Correct the stale `cudagraph.py:51-52` docstring
(per-step capture ships in 14 cards, not just wan21).
- **RESCOPE per-node TP.** Wan/LTX use `ReplicatedLinear` + **sequence parallelism** (`sp`), not TP; the `sp`
axis already exists in `parallel/plan.py:AXIS_NAMES`. The work is wiring degrees into the runtime, not
inventing vocabulary; a `tp_size=2` "one-line activation" is a no-op for the shipped models.
---
## 6. The first integration test, if/when multi-GPU placement work starts
The **live Qwen-Omni 2-GPU bring-up** (Thinker on rank 0, Talker+Code2Wav on rank 1; see
`v2_debug_videos/vlm.md` Session 4) is the natural first validation target for any `(node,Walk)→rank`
placement work — it is the one place this repo already has real multi-rank composite-model execution.
---
## Anchor files for P0/P1
`v2/program/specs.py`, `v2/runtime/engine.py` (**line 88 fix**), `v2/runtime/disaggregated.py`,
`v2/recipes/omni/ar_loop.py`, `v2/loop/contracts.py`, `v2/card/specs.py`, `v2/cache/{classes.py,keys.py}`,
`v2/parity/interleave_gate.py`, `v2/extend/{interceptors,registry}.py`, `recipes/__init__.py` +
`v2/program/workflow.py` (registry-driven delivery).
@@ -0,0 +1,73 @@
---
date: 2026-05-07
experiment: PR #1280 (daVinci-MagiHuman port), distill DiT parity bring-up
category: porting
severity: important
---
# Conversion `--cast-bf16` Needs an FP32-Keep Suffix Allowlist
## What Happened
`scripts/checkpoint_conversion/convert_magi_human_to_diffusers.py --cast-bf16`
produced a converted distill DiT checkpoint that loaded cleanly, ran end-to-
end, and emitted reasonable output — but `test_magi_human_distill_parity`
showed `diff_mean=0.114` against the upstream reference. The base DiT was
bit-exact with the same conversion script. Only the distill variant
regressed.
The error was small enough that visual quality looked normal, but large
enough to fail bit-exact parity. The MagiHuman base + distill DiTs share
most of their architecture, so a difference that affected only distill was
counterintuitive.
## Root Cause
`--cast-bf16` was downcasting **all** fp32 tensors to bf16 indiscriminately.
The base checkpoint and the FastVideo `final_linear` / adapter modules
require eight specific tensors to remain in fp32:
- LayerNorm `gamma` / `beta` weights for the final residual exit
- Adapter projection biases
- A handful of scale parameters in the output projection chain
These tensors participate in chains where bf16 precision causes accumulation
error large enough to drift the parity check. The base DiT happened to not
hit those specific chains in the path the test exercised (different
attention mask shape, different audio interleave); the distill variant did.
## Fix / Workaround
Added `_FP32_KEEP_SUFFIXES` allowlist to
`convert_magi_human_to_diffusers.py` (commit `829f70d3`) and gated `--cast-
bf16` on it. Tensors whose state-dict key ends with any allowlisted suffix
keep their original fp32 dtype regardless of the flag.
Distill DiT parity went from `diff_mean=0.114` (silently wrong) to bit-exact
in one commit.
## Prevention
1. **Treat `--cast-bf16` as opinionated, not blanket.** Any conversion
script that supports a global dtype downcast flag MUST own an explicit
allowlist of fp32-keep tensors, documented at the top of the file.
2. **The `add-model-conversion` skill** should enforce two checks for any
converter that ships a `--cast-bf16`-style flag:
- Run the parity test for **every** variant of the model (base, distill,
SR, etc.), not just the headline variant. Different variants exercise
different code paths.
- Diff the converted checkpoint's dtype map against the upstream
reference and assert the allowlist covers every fp32 tensor in the
reference.
3. **For MagiHuman specifically**: if you add or rename DiT modules that
touch `final_linear`, the adapter, or any LayerNorm in the residual exit
path, **check that any fp32-required tensors are covered by
`_FP32_KEEP_SUFFIXES`** in the conversion script and re-run
`test_magi_human_distill_parity` (it's the canary).
4. The lesson generalizes beyond MagiHuman: any DiT that uses bf16 mixed
precision but keeps specific tensors in fp32 (a common pattern with
flash-attn-style backends) needs this allowlist for any conversion that
downcasts.
@@ -0,0 +1,77 @@
---
date: 2026-05-07
experiment: PR #1280 (daVinci-MagiHuman port), DiT parity bring-up
category: porting
severity: important
---
# DiT Dtype Boundary Alignment with Flash-Attn-Style Backends
## What Happened
DiT bit-exact parity for daVinci-MagiHuman against the upstream reference
sat at `diff_max=0.5` after the architecture port was complete and weight
loading was correct. The error grew with depth (later layers diverged more
than earlier ones), suggesting an accumulating numerical drift rather than
a structural mismatch. None of the obvious culprits (RoPE, GQA expansion,
attention mask handling) accounted for the pattern.
## Root Cause
Four cumulative dtype-boundary mismatches, each individually small but
together pushing parity from `diff_max=0.5` to bit-exact (`diff_max=0.0`):
1. **SDPA inputs were not cast to bf16.** Upstream's `flash_attn_with_cp`
internally casts Q/K/V to bf16 at `dit_module.py:508` before the kernel.
FastVideo was passing fp32 tensors through, getting numerically different
intermediates even though the kernel accepts both.
2. **Post-attention output was kept in bf16 across the per-head gating
multiply.** Upstream upcasts to fp32 before the gating, FastVideo did the
gate in bf16 then upcast.
3. **A residual-stream cast at the block boundary.** FastVideo had a
`.to(bf16)` then `.to(fp32)` at the start of each block. Upstream keeps
the residual stream **continuously in fp32** across all 40 layers; only
the inputs to specific kernels are temporarily downcast.
4. **Parity test scheduler used a double-shift.** A separate per-block fix
(Wave 11 production migration) — single-shift schedule is what upstream
uses; the parity test was double-shifting.
## Fix / Workaround
Four cumulative changes in `fastvideo/models/dits/magi_human.py` (commit
`3a4816cb`), each with a comment at the call site explaining the upstream
parity rationale:
- Cast SDPA inputs to bf16 right before the attention call.
- Upcast attention output to fp32 before the per-head gate multiply.
- Drop the residual-stream `.to(bf16)`/`.to(fp32)` wrapper at the block
boundary; let the residual stay fp32 throughout.
- Single-shift schedule in the parity test fixture (matches upstream Wave 11).
## Prevention
1. **For any DiT port with a flash-attn-style backend**, treat the dtype of
the residual stream as a load-bearing invariant, not a performance knob.
Document it in the model's per-pipeline AGENTS.md. MagiHuman's invariant:
*residual stream stays fp32 across all blocks; only kernel inputs are
temporarily bf16*.
2. **Use layer-by-layer activation hooks** when DiT parity is close-but-not-
bit-exact and the gap grows with depth. The
`fastvideo/hooks/activation_trace.py` infra exists exactly for this case
(`add-model-trace` skill). In MagiHuman's case it would have localized the
first divergence point in one pass.
3. **The `add-model-port-dit` skill** should explicitly call out:
- SDPA input dtype must match the upstream kernel's internal cast.
- Post-attention upcast happens **before** any per-head gate, not after.
- Residual stream dtype across block boundaries is a parity invariant.
These rules apply to any DiT port whose upstream uses a flash-attn-style
backend (`flash_attn_with_cp`, `flex_flash_attn_func`, etc.).
4. **Add an "intermediate-layer parity" test** for new DiT ports — comparing
activations at layer 5, 10, 20, 30 — not just the final output. A growing-
with-depth pattern is otherwise indistinguishable from "almost right".
@@ -0,0 +1,69 @@
---
date: 2026-05-07
experiment: PR #1280 (daVinci-MagiHuman port), Wave 14
category: porting
severity: critical
---
# Silent Channel-Major Token-Packing Bugs
## What Happened
While porting daVinci-MagiHuman (`fastvideo/pipelines/basic/magi_human/`),
the pipeline-parity test passed bit-exactly but the E2E user-visible output
was **pure static noise**. Latent tensors compared identically against the
upstream reference at every checkpointed boundary, yet decoded videos showed
no recognizable content. The discrepancy reproduced on every variant
(base / distill / SR-540p / SR-1080p) with the same noise profile.
## Root Cause
Video tokens were being packed **spatial-major** instead of **channel-major**:
```python
# What we had (spatial-major, WRONG)
einops.rearrange(x, "b c (T pT) (H pH) (W pW) -> b (T H W) (pT pH pW C)", ...)
# What upstream's UnfoldNd produces (channel-major, CORRECT)
einops.rearrange(x, "b c (T pT) (H pH) (W pW) -> b (T H W) (C pT pH pW)", ...)
```
A single-character einops reorder. The pipeline-parity test used FastVideo's
own packer on **both** sides of the comparison, so the bug was invisible there
— both sides agreed on the wrong layout. The DiT consumed those tokens
without complaint because the channel dimension only matters at decode time,
when the VAE's first conv expects channel-major input. By that point the test
boundary was already passed.
The bug was load-bearing for any token-packed format that downstream feeds
into a `UnfoldNd`-shaped consumer. Wave 14 of the port took multiple bug-hunt
iterations and an Oracle consultation to localize.
## Fix / Workaround
Single-character einops change in `stages/latent_preparation.py:_img2tokens`
(commit `6d190693` of the original PR). After the fix, all four variants
produced expected E2E output and the pipeline-parity tests still passed
because both sides of the parity check are now correct.
## Prevention
1. **Never use the FastVideo-side packer on both sides of a parity test.**
At least one parity boundary must compare against an upstream tensor
produced by the upstream packer. For MagiHuman this means a separate
`_img2tokens` parity test that feeds upstream `UnfoldNd` output as the
reference, not FastVideo's reformatted equivalent.
2. **Add an E2E hash check** alongside latent-parity. The mp4 SHA was the
first signal that something was wrong; if it had been part of the standard
parity battery, the bug would have surfaced in Wave 1, not Wave 14. See
`fastvideo/tests/ssim/test_magi_human_similarity.py` for the CI version.
3. **For any new model port that involves explicit tensor reshaping into
tokens**, document the expected packing order (`(C pT pH pW)` vs
`(pT pH pW C)`) at the call site and assert the layout matches the
downstream consumer's expectation.
4. The `add-model-port-dit` skill's parity gate should require an E2E hash
check for any DiT that does video token packing, not just latent
bit-exactness.
@@ -0,0 +1,41 @@
---
date: 2026-05-22
experiment: PR #1386 DreamVerse app CI backend tests
category: infrastructure
severity: important
---
# DreamVerse App CI Streaming Imports Need GPU
## What Happened
DreamVerse app CI backend pytest collection imports FastVideo streaming surfaces.
When those tests run in a CPU-only Modal environment, collection can fail before
any app assertions run with Triton reporting:
```text
RuntimeError: 0 active drivers
```
## Root Cause
Some streaming import paths can import `fastvideo_kernel` at module import time.
Triton then probes for an active GPU driver during pytest collection. A CPU-only
Modal container has no active driver, so the failure appears as an import-time
collection error rather than a DreamVerse app behavior failure.
## Fix / Workaround
For PR #1386, use a surgical CI fix: allocate a GPU to
`run_dreamverse_app_tests`. Do not refactor core streaming/kernel imports just to
unstick this app CI path.
Keep `build_kernel=False` for this job. The DreamVerse app backend test imports
streaming surfaces but does not need to rebuild or exercise custom kernels.
## Prevention
When adding or modifying DreamVerse app CI jobs that import FastVideo streaming
modules, make the GPU requirement explicit if the import graph may touch
`fastvideo_kernel`. Prefer small CI resource fixes for app test collection issues
unless the product code genuinely requires lazy import cleanup.
+48
View File
@@ -0,0 +1,48 @@
# Lessons Learned Database
This directory stores documented mistakes, unexpected behaviors, and their fixes.
Each lesson is a permanent record that helps agents and humans avoid repeating
past errors.
## When to Create a Lesson
- An experiment failed for a non-obvious reason.
- A configuration or hyperparameter choice led to wasted compute.
- A porting, data, or infrastructure issue was discovered and resolved.
- A workaround was needed for a known framework/library bug.
## File Naming
`<YYYY-MM-DD>_<short-slug>.md` — e.g., `2026-03-02_lr-too-high-for-lora.md`
## Template
```markdown
---
date: <ISO-8601>
experiment: <reference to experiment_journal.md entry, if applicable>
category: hyperparameter | data | infrastructure | evaluation | porting | other
severity: critical | important | minor
---
# <Short Descriptive Title>
## What Happened
<Description of the problem and its symptoms.>
## Root Cause
<Analysis of why it happened.>
## Fix / Workaround
<What resolved the issue.>
## Prevention
<How to avoid this in the future — updated skills, SOPs, or checks.>
```
## Usage
- Before starting a task, **search this directory** for relevant lessons.
- After completing or failing a task, **check if a new lesson should be created**.
- Periodically review lessons for **patterns** — recurring themes may warrant
a new skill, SOP, or codebase fix.
+96
View File
@@ -0,0 +1,96 @@
#!/usr/bin/env bash
# Sync .agents/skills/ into .claude/skills/ via per-skill symlinks.
#
# Why: Claude Code only scans .claude/skills/ and ~/.claude/skills/ for
# user-invocable skills (no skillsPath config exists — see
# https://code.claude.com/docs/en/skills.md). This repo's skills live
# in .agents/skills/ so they travel with the repo and stay under git.
# Run this once after cloning (or after adding/removing a skill) to
# expose them to Claude Code without maintaining a parallel tree.
#
# Usage:
# .agents/scripts/sync-skills.sh
#
# Idempotent and safe to re-run. Prunes stale symlinks whose source
# has been removed from .agents/skills/. Leaves hand-written
# .claude/skills/<name>/ directories untouched (only symlinks are
# managed).
set -euo pipefail
REPO_ROOT="$(git -C "$(dirname "$0")" rev-parse --show-toplevel)"
SRC_DIR="$REPO_ROOT/.agents/skills"
DST_DIR="$REPO_ROOT/.claude/skills"
if [[ ! -d "$SRC_DIR" ]]; then
echo "Error: $SRC_DIR does not exist." >&2
exit 1
fi
mkdir -p "$DST_DIR"
linked=0
unchanged=0
skipped=0
pruned=0
link_skill() {
local name="$1"
local src="$SRC_DIR/$name"
local dst="$DST_DIR/$name"
# Relative target keeps symlinks portable across clones.
local rel="../../.agents/skills/$name"
if [[ -L "$dst" ]]; then
if [[ "$(readlink "$dst")" == "$rel" ]]; then
unchanged=$((unchanged + 1))
return
fi
rm "$dst"
elif [[ -e "$dst" ]]; then
echo "Skipped (not a symlink): .claude/skills/$name" >&2
skipped=$((skipped + 1))
return
fi
ln -s "$rel" "$dst"
echo "Linked: .claude/skills/$name -> $rel"
linked=$((linked + 1))
}
prune_stale() {
local link="$1"
local target
target="$(readlink "$link")"
case "$target" in
../../.agents/skills/*) ;;
*) return ;;
esac
local name="${target##*/}"
if [[ ! -d "$SRC_DIR/$name" ]]; then
rm "$link"
echo "Pruned stale: .claude/skills/$(basename "$link")"
pruned=$((pruned + 1))
fi
}
for src in "$SRC_DIR"/*/; do
[[ -d "$src" ]] || continue
name="$(basename "$src")"
# Only treat directories that actually contain a SKILL.md as skills.
[[ -f "$src/SKILL.md" ]] || continue
link_skill "$name"
done
shopt -s nullglob
for link in "$DST_DIR"/*; do
[[ -L "$link" ]] || continue
prune_stale "$link"
done
shopt -u nullglob
printf "\nSummary: %d linked, %d unchanged, %d pruned" "$linked" "$unchanged" "$pruned"
if [[ "$skipped" -gt 0 ]]; then
printf ", %d skipped (non-symlink collision)" "$skipped"
fi
printf "\n"
+55
View File
@@ -0,0 +1,55 @@
---
name: <skill-name>
description: <one-line description — Codex uses this for implicit invocation matching>
---
# <Skill Name>
## Purpose
<Why this skill exists and when to use it.>
## Prerequisites
- <What must be true before using this skill>
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| `param1` | Yes | ... |
## Steps
1. **Step 1 title**
- Detail...
2. **Step 2 title**
- Detail...
## Outputs
- <What this skill produces>
## Example Usage
```
<Example invocation or prompt snippet>
```
## References
- <Links to relevant files in the codebase>
---
## Folder Structure
Each skill lives in its own directory under `.agents/skills/`:
```
.agents/skills/<skill-name>/
├── SKILL.md # Required: instructions + metadata (this file)
├── scripts/ # Optional: executable helper scripts
├── references/ # Optional: documentation, papers
└── assets/ # Optional: templates, resources
```
Skill discovery is directory-based; no hand-maintained registry entry is
required. Run `.agents/scripts/sync-skills.sh` if a local Claude Code checkout
needs refreshed `.claude/skills/` symlinks.
+173
View File
@@ -0,0 +1,173 @@
---
name: add-model-01-prep
description: Use at the start of a FastVideo model port to gather required inputs, inspect/download HF weights, clone and install the official reference repo in the current environment, create a local_tests README skeleton, and produce a handoff before conversion or implementation.
---
# Add Model Prep
## Goal
Prepare external assets and the shared parity-test environment for a FastVideo
model port. Stop before writing conversion scripts, model components, pipeline
code, registry entries, or executable parity tests.
## Ask First
Ask once, then proceed if the HF token is already exported:
```text
Before prep: (1) official reference repo or Diffusers pipeline URL, (2) HF repo
id or local weights path and whether it has a root model_index.json, (3) target
model_family, (4) workload types, (5) which token env var is exported:
HF_TOKEN, HUGGINGFACE_HUB_TOKEN, or HF_API_KEY, (6) may I stage clone and
weights under the FastVideo repo root, and (7) may I install official reference
dependencies into the current FastVideo conda/env for parity tests?
```
Useful optional inputs: `pipeline_class`, `reference_dir`, `hf_revision`,
`official_revision`, `reuse_hints`, `download_scope`.
## Rules
- Follow `../add-model/shared/common_rules.md` for token/auth safety, state files,
escape hatches, and skip/pass semantics.
- Run from the FastVideo repo root.
- Use repo-relative defaults: `<ReferenceDir>/`,
`official_weights/<model_family>/`, `converted_weights/<model_family>/`.
- Install official reference deps into the current FastVideo environment, not a
new venv/conda env, so parity tests run both implementations with one shared
numeric stack.
- If the reference is a Diffusers class/package instead of a cloneable repo,
record import path and version instead of cloning.
- Prep may create only the local-test README and `PORT_STATUS.md` skeletons;
executable `.py` parity tests belong to `../add-model-02-parity/SKILL.md`.
## Escape Hatches
Follow `../add-model/shared/common_rules.md`. Prep-specific ask cases include
overwriting an existing clone or weight directory, installing untrusted/private
deps, choosing between incompatible official references, large downloads outside
the agreed scope, or missing gated-repo auth setup by env var name.
## Workflow
1. Verify the repo:
```bash
git rev-parse --show-toplevel
```
Expected markers: `fastvideo/`, `scripts/checkpoint_conversion/`,
`scripts/huggingface/download_hf.py`, `fastvideo/registry.py`.
2. Inspect HF or local weight layout:
```bash
python ".agents/skills/add-model-01-prep/scripts/inspect_hf_layout.py" \
"Org/Model" \
--revision "<revision>" \
--json
```
For a local path, replace `Org/Model` with `/path/to/weights`. Record
`source_layout`, `needs_conversion`, `model_index_class`, and
`components_seen`.
3. Download HF weights if needed:
```bash
python ".agents/skills/add-model-01-prep/scripts/download_hf_weights.py" \
"Org/Model" \
"official_weights/<model_family>" \
--revision "<revision>"
```
For selected files, repeat `--file-name`. For partial snapshots, repeat
`--allow-pattern` or `--ignore-pattern`. If the user provided a local path,
record it instead of copying large weights by default.
4. Clone the official reference repo if applicable:
```bash
python ".agents/skills/add-model-01-prep/scripts/clone_reference_repo.py" \
"<official_repo_url>" \
"<ReferenceDir>" \
--branch "<tag-or-branch>" \
--commit "<commit-sha>" \
--update-gitignore
```
Omit `--branch`, `--commit`, or `--update-gitignore` when not needed. The
helper refuses to overwrite existing paths and prints remote/HEAD instead.
5. Keep prep assets ignored. Ensure `.gitignore` includes relevant entries:
```gitignore
/<ReferenceDir>/
/official_weights/
/converted_weights/
```
6. Follow the official repo's setup instructions in the current environment.
Inspect dependency files and README install docs before installing anything:
- `README*`, install docs, or model-card instructions.
- `requirements*.txt`, `pyproject.toml`, `setup.py`, `environment.yml`.
Use the current FastVideo conda/env. Do not create a new env even if upstream
docs recommend one; translate the needed install commands into the active env.
Prefer editable/no-deps first so the official source is importable without
changing shared pins:
```bash
uv pip install --no-deps -e ./<ReferenceDir>
```
Then install only missing official deps needed for parity imports. Stop before
installing requirements that would change FastVideo's core stack. If upstream
requires private/non-PyPI deps, record that parity needs a local stub helper
rather than pretending setup is complete.
7. Create the model-family local test skeleton and top-level port state file:
```bash
mkdir -p tests/local_tests/<model_family>
cp ".agents/skills/add-model-01-prep/templates/local_tests_readme.md" \
tests/local_tests/<model_family>/README.md
cp ".agents/skills/add-model-01-prep/templates/port_status.md" \
tests/local_tests/<model_family>/PORT_STATUS.md
```
Edit every placeholder in the README and `PORT_STATUS.md`. The README gives
later review agents enough information to reproduce the shared environment and
run/review parity work:
- official code URL or import path, local clone path, and commit/version;
- HF URL or local weight path, revision, access notes, and token env var name
only;
- commands already run and any blocked official dependency installs;
- shared-env install commands to re-run without changing core pins;
- expected local parity test paths and pytest commands;
- private-dependency stubs or known setup gaps;
- PR/review notes explaining which parity tests are required before handoff.
Do not include raw tokens, absolute cache paths that are not repo-reproducible,
or large generated outputs. If prep is blocked before imports work, still create
the README with `official_env_status=blocked` and the exact blocker.
`PORT_STATUS.md` must follow `../add-model/contracts/port_state.md`. Record open
questions and prep issues immediately, using stable IDs such as `Q001` and
`I001`. Keep resolved questions/issues in the table with a resolution instead of
deleting them.
## Handoff
End with the canonical prep handoff contract from
`../add-model/contracts/prep_handoff.md` and update the shared state files before
handoff.
## Helper Scripts
- `scripts/inspect_hf_layout.py`: classify HF/local layout.
- `scripts/download_hf_weights.py`: download HF snapshot or selected files.
- `scripts/clone_reference_repo.py`: clone reference repo safely.
@@ -0,0 +1,123 @@
#!/usr/bin/env python3
"""Clone an official reference repo without overwriting existing paths."""
from __future__ import annotations
import argparse
import subprocess
import sys
from pathlib import Path
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Clone a reference repo for FastVideo parity tests."
)
parser.add_argument("repo_url", help="Official reference repository URL")
parser.add_argument("target_dir", help="Directory to clone into")
parser.add_argument("--branch", help="Branch or tag to clone")
parser.add_argument("--commit", help="Commit SHA to check out after clone")
parser.add_argument(
"--update-gitignore",
action="store_true",
help="Add the target directory to .gitignore if missing",
)
parser.add_argument(
"--gitignore",
default=".gitignore",
help="Path to gitignore file when --update-gitignore is used",
)
return parser.parse_args()
def run(command: list[str], check: bool = True) -> subprocess.CompletedProcess[str]:
return subprocess.run(
command,
check=check,
text=True,
capture_output=True,
)
def print_existing_repo_info(target: Path) -> int:
print(f"target_exists: {target}")
if not (target / ".git").exists():
print("error: target exists but is not a git repo", file=sys.stderr)
return 1
remote = run(["git", "-C", str(target), "remote", "-v"], check=False)
head = run(["git", "-C", str(target), "rev-parse", "HEAD"], check=False)
if remote.stdout:
print("remote_v:")
print(remote.stdout.rstrip())
if head.stdout:
print(f"head: {head.stdout.strip()}")
print("not_overwritten: true")
return 0
def gitignore_entry_for(target: Path) -> str:
root = Path.cwd().resolve()
resolved = target.resolve()
try:
relative = resolved.relative_to(root)
except ValueError as exc:
raise ValueError(
"--update-gitignore requires target_dir to be under the current directory"
) from exc
text = relative.as_posix().rstrip("/")
return "/" + text + "/"
def update_gitignore(path: Path, target: Path) -> bool:
entry = gitignore_entry_for(target)
existing = path.read_text().splitlines() if path.exists() else []
if entry in existing:
return False
new_text = "\n".join(existing).rstrip("\n")
if new_text:
new_text += "\n"
new_text += entry + "\n"
path.write_text(new_text)
return True
def main() -> int:
args = parse_args()
target = Path(args.target_dir)
if target.exists():
return print_existing_repo_info(target)
command = ["git", "clone", "--depth", "1"]
if args.branch:
command.extend(["--branch", args.branch])
command.extend([args.repo_url, str(target)])
try:
run(command)
if args.commit:
run(["git", "-C", str(target), "fetch", "--depth", "1", "origin", args.commit])
run(["git", "-C", str(target), "checkout", args.commit])
except subprocess.CalledProcessError as exc:
if exc.stdout:
print(exc.stdout, end="")
if exc.stderr:
print(exc.stderr, end="", file=sys.stderr)
return exc.returncode
head = run(["git", "-C", str(target), "rev-parse", "HEAD"])
print(f"cloned: {target}")
print(f"head: {head.stdout.strip()}")
if args.update_gitignore:
changed = update_gitignore(Path(args.gitignore), target)
print(f"gitignore_updated: {str(changed).lower()}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,105 @@
#!/usr/bin/env python3
"""Download HF weights using the standard FastVideo token env vars."""
from __future__ import annotations
import argparse
import os
import sys
from pathlib import Path
HF_TOKEN_ENV_KEYS = ("HF_TOKEN", "HUGGINGFACE_HUB_TOKEN", "HF_API_KEY")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Download a HF model snapshot or selected files into a local directory."
)
parser.add_argument("repo_id", help="HF repo id, for example Org/Model")
parser.add_argument("local_dir", help="Destination directory")
parser.add_argument("--repo-type", default="model", help="HF repo type (default: model)")
parser.add_argument("--revision", help="HF branch, tag, or commit")
parser.add_argument(
"--file-name",
action="append",
default=[],
help="Download one file; may be repeated. If omitted, download full snapshot.",
)
parser.add_argument(
"--allow-pattern",
action="append",
default=[],
help="Snapshot allow pattern; may be repeated. Ignored when --file-name is used.",
)
parser.add_argument(
"--ignore-pattern",
action="append",
default=[],
help="Snapshot ignore pattern; may be repeated. Ignored when --file-name is used.",
)
return parser.parse_args()
def resolve_token() -> tuple[str | None, str | None]:
for key in HF_TOKEN_ENV_KEYS:
value = os.environ.get(key)
if value:
return key, value
return None, None
def main() -> int:
args = parse_args()
token_env, token = resolve_token()
local_dir = Path(args.local_dir).expanduser()
if token_env:
print(f"token_env: {token_env}")
else:
print("token_env: none", file=sys.stderr)
try:
if local_dir.exists() and not local_dir.is_dir():
print(
f"error: destination exists and is not a directory: {local_dir}",
file=sys.stderr,
)
return 1
local_dir.mkdir(parents=True, exist_ok=True)
if args.file_name:
from huggingface_hub import hf_hub_download
for file_name in args.file_name:
path = hf_hub_download(
repo_id=args.repo_id,
filename=file_name,
repo_type=args.repo_type,
revision=args.revision,
local_dir=str(local_dir),
token=token,
)
print(f"downloaded_file: {path}")
else:
from huggingface_hub import snapshot_download
path = snapshot_download(
repo_id=args.repo_id,
repo_type=args.repo_type,
revision=args.revision,
local_dir=str(local_dir),
token=token,
allow_patterns=args.allow_pattern or None,
ignore_patterns=args.ignore_pattern or None,
)
print(f"downloaded_snapshot: {path}")
except Exception as exc: # noqa: BLE001 - CLI should print concise failures.
print(f"error: {exc}", file=sys.stderr)
return 1
print(f"local_dir: {local_dir.resolve()}")
return 0
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,264 @@
#!/usr/bin/env python3
"""Inspect a Hugging Face repo or local weight directory layout."""
from __future__ import annotations
import argparse
import json
import os
import sys
from pathlib import Path
from typing import Any
HF_TOKEN_ENV_KEYS = ("HF_TOKEN", "HUGGINGFACE_HUB_TOKEN", "HF_API_KEY")
RAW_WEIGHT_SUFFIXES = (".safetensors", ".pt", ".pth", ".ckpt", ".bin")
KNOWN_COMPONENTS = {
"audio_vae",
"conditioner",
"feature_extractor",
"image_encoder",
"scheduler",
"text_encoder",
"text_encoder_2",
"tokenizer",
"tokenizer_2",
"transformer",
"transformer_2",
"unet",
"upsampler",
"vae",
"vocoder",
}
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Classify a HF repo or local directory as Diffusers, raw, custom, or unknown."
)
parser.add_argument("source", help="HF repo id or local weights directory")
parser.add_argument("--repo-type", default="model", help="HF repo type (default: model)")
parser.add_argument("--revision", help="HF revision to inspect")
parser.add_argument(
"--max-local-files",
type=int,
default=20000,
help="Maximum local files to scan recursively (default: 20000)",
)
parser.add_argument(
"--sample-limit",
type=int,
default=80,
help="Number of file paths to print in human output (default: 80)",
)
parser.add_argument("--json", action="store_true", help="Emit JSON only")
return parser.parse_args()
def resolve_token() -> tuple[str | None, str | None]:
for key in HF_TOKEN_ENV_KEYS:
value = os.environ.get(key)
if value:
return key, value
return None, None
def load_local_files(root: Path, max_files: int) -> tuple[list[str], bool]:
files: list[str] = []
truncated = False
for path in root.rglob("*"):
if not path.is_file():
continue
files.append(path.relative_to(root).as_posix())
if len(files) >= max_files:
truncated = True
break
return sorted(files), truncated
def load_local_model_index(root: Path) -> tuple[dict[str, Any] | None, str | None]:
index_path = root / "model_index.json"
if not index_path.is_file():
return None, None
try:
return json.loads(index_path.read_text()), None
except Exception as exc: # noqa: BLE001 - surface malformed JSON clearly.
return None, f"failed to parse local model_index.json: {exc}"
def load_remote_files(
repo_id: str,
repo_type: str,
revision: str | None,
token: str | None,
) -> list[str]:
from huggingface_hub import list_repo_files
return sorted(
list_repo_files(
repo_id,
repo_type=repo_type,
revision=revision,
token=token,
)
)
def load_remote_model_index(
repo_id: str,
repo_type: str,
revision: str | None,
token: str | None,
) -> tuple[dict[str, Any] | None, str | None]:
from huggingface_hub import hf_hub_download
try:
path = hf_hub_download(
repo_id=repo_id,
filename="model_index.json",
repo_type=repo_type,
revision=revision,
token=token,
)
except Exception as exc: # noqa: BLE001 - missing/inaccessible file is data.
return None, f"failed to download model_index.json: {exc}"
try:
return json.loads(Path(path).read_text()), None
except Exception as exc: # noqa: BLE001 - surface malformed JSON clearly.
return None, f"failed to parse remote model_index.json: {exc}"
def root_file_names(files: list[str]) -> set[str]:
return {name for name in files if "/" not in name}
def component_names(files: list[str], model_index: dict[str, Any] | None) -> list[str]:
components: set[str] = set()
for name in files:
parts = name.split("/", 1)
if len(parts) != 2:
continue
top, rest = parts
if top in KNOWN_COMPONENTS or rest == "config.json":
components.add(top)
if model_index:
for key, value in model_index.items():
if key.startswith("_"):
continue
if isinstance(value, list) and len(value) == 2:
components.add(key)
return sorted(components)
def classify_layout(
files: list[str],
model_index: dict[str, Any] | None,
components: list[str],
) -> tuple[str, str]:
roots = root_file_names(files)
raw_weight_files = [name for name in roots if name.endswith(RAW_WEIGHT_SUFFIXES)]
has_model_index = "model_index.json" in roots or model_index is not None
if has_model_index and components:
return "diffusers", "no"
if has_model_index:
return "custom", "unknown"
if raw_weight_files:
return "raw_official", "yes"
if any(name.endswith(RAW_WEIGHT_SUFFIXES) for name in files):
return "custom", "yes"
return "unknown", "unknown"
def build_result(args: argparse.Namespace) -> dict[str, Any]:
token_env, token = resolve_token()
source_path = Path(args.source).expanduser()
is_local = source_path.exists()
if is_local:
root = source_path.resolve()
if not root.is_dir():
raise ValueError(f"local source is not a directory: {root}")
files, truncated = load_local_files(root, args.max_local_files)
model_index, model_index_error = load_local_model_index(root)
source_kind = "local"
source = str(root)
else:
files = load_remote_files(args.source, args.repo_type, args.revision, token)
truncated = False
model_index, model_index_error = load_remote_model_index(
args.source,
args.repo_type,
args.revision,
token,
)
source_kind = "hf"
source = args.source
components = component_names(files, model_index)
source_layout, needs_conversion = classify_layout(files, model_index, components)
return {
"source": source,
"source_kind": source_kind,
"repo_type": None if is_local else args.repo_type,
"revision": args.revision,
"token_env": token_env,
"source_layout": source_layout,
"needs_conversion": needs_conversion,
"model_index_class": (model_index or {}).get("_class_name"),
"model_index_diffusers_version": (model_index or {}).get("_diffusers_version"),
"model_index_error": model_index_error,
"components_seen": components,
"file_count": len(files),
"file_scan_truncated": truncated,
"files_sample": files[: args.sample_limit],
}
def print_human(result: dict[str, Any]) -> None:
for key in (
"source",
"source_kind",
"repo_type",
"revision",
"token_env",
"source_layout",
"needs_conversion",
"model_index_class",
"model_index_diffusers_version",
"model_index_error",
"file_count",
"file_scan_truncated",
):
value = result.get(key)
if value is not None:
print(f"{key}: {value}")
components = result["components_seen"]
print("components_seen: " + (", ".join(components) if components else "none"))
print("files_sample:")
for name in result["files_sample"]:
print(f" {name}")
def main() -> int:
args = parse_args()
try:
result = build_result(args)
except Exception as exc: # noqa: BLE001 - CLI should print concise failures.
print(f"error: {exc}", file=sys.stderr)
return 1
if args.json:
print(json.dumps(result, indent=2, sort_keys=True))
else:
print_human(result)
return 0
if __name__ == "__main__":
raise SystemExit(main())
@@ -0,0 +1,122 @@
# <Model Family> Local Tests
Local-only parity and smoke tests for the `<model_family>` FastVideo port. These
tests compare FastVideo against the official reference implementation and are
not expected to run in CI unless explicitly promoted later.
Port progress, open questions, issues, and handoff notes live in
`tests/local_tests/<model_family>/PORT_STATUS.md`.
## Reference Assets
| Field | Value |
|---|---|
| Model family | `<model_family>` |
| Workload types | `<T2V/I2V/V2V/T2I/or compatibility shim with rationale>` |
| Official reference | `<url or import path>` |
| Local reference dir | `<ReferenceDir or none>` |
| Official commit/version | `<sha, tag, package version, or unknown>` |
| HF weights | `<HF repo id/url or local path>` |
| HF revision | `<revision or default>` |
| Local weights dir | `<official_weights/model_family or local path>` |
| Source layout | `<diffusers/raw_official/monolithic/separate_components/mixed/custom/unknown>` |
| Needs conversion | `<yes/no/unknown>` |
Do not write token values in this file. Use only the token env var name:
`<HF_TOKEN or HUGGINGFACE_HUB_TOKEN or HF_API_KEY>`.
## Shared Environment Setup
Run from the FastVideo repo root in the same conda/env used for FastVideo.
Do not create a separate upstream environment for parity tests.
```bash
# Official reference source, if cloneable.
python ".agents/skills/add-model-01-prep/scripts/clone_reference_repo.py" \
"<official_repo_url>" \
"<ReferenceDir>" \
--commit "<commit-sha>" \
--update-gitignore
# Editable install without changing shared core pins.
uv pip install --no-deps -e ./<ReferenceDir>
# Additional official deps installed or required for imports:
# <package list or none>
```
Do not change core dependency versions (`torch`, `diffusers`, `transformers`,
`flash-attn`, `triton`, CUDA packages) without explicit approval.
## Official Environment Status
```text
dependency_changes: <none | installed no-deps editable | installed official deps in current env | blocked on user>
official_env_status: <imports_ok | private_deps_need_stubs | blocked>
private_dep_stubs: <none or tests/local_tests/helpers/<model_family>_upstream.py>
blocked_on: <none or exact blocker>
```
## Weight Setup
```bash
python ".agents/skills/add-model-01-prep/scripts/download_hf_weights.py" \
"<Org/Model>" \
"official_weights/<model_family>" \
--revision "<revision>"
```
If weights are local-only, record the local path and do not copy large files into
the repository.
## Prototype And Conversion Artifacts
State-dict key/shape dumps are generated after FastVideo native prototypes exist
and are used to build the conversion mapping.
```text
official_key_dumps:
<component>: converted_weights/<model_family>/_mapping/<component>_official_keys.json
fastvideo_key_dumps:
<component>: converted_weights/<model_family>/_mapping/<component>_fastvideo_keys.json
conversion_script: scripts/checkpoint_conversion/<model_family>_to_diffusers.py
conversion_source_layout: <diffusers | separate_components | monolithic | mixed | custom>
converted_weights_dir: converted_weights/<model_family>
strict_load_status: <not_run | pass | pass_with_documented_exclusions | blocked>
```
For monolithic official checkpoints, record the component prefix split here. For
example, a single checkpoint may contain transformer, VAE/pretransform,
conditioner, and scheduler/vocoder keys that the conversion script writes into
separate FastVideo component subfolders.
## Expected Parity Tests
Planned local tests for this family:
| Component | Official files / args | Test | Concerns | Status |
|---|---|---|---|---|
| `<component>` | `<definition path; instantiation path + args>` | `tests/local_tests/<bucket>/test_<model_family>_<component>_parity.py` | `<prototype or setup concerns>` | `<planned/scaffold_skip/debug_red/non_skip_pass/blocked>` |
| `pipeline` | `<official pipeline call>` | `tests/local_tests/pipelines/test_<model_family>_pipeline_parity.py` | `<pipeline concerns>` | `<planned/scaffold_skip/debug_red/non_skip_pass/blocked>` |
Include reused components in this table. Reuse is accepted only after the
FastVideo component definition and official instantiation arguments have both
been checked and the component parity test passes non-skip.
Run the relevant tests with:
```bash
pytest tests/local_tests/<bucket>/test_<model_family>_<component>_parity.py -v -s
pytest tests/local_tests/pipelines/test_<model_family>_pipeline_parity.py -v -s
```
## Review Notes
- Required before handoff: non-skip PASS for each required component parity
test, including reused components that own weights or numerical behavior.
- Pipeline parity may start as a scaffold, but final handoff requires non-skip
PASS or an explicit blocker accepted through the escape-hatch process.
- User decisions and pause points are tracked as `E###` rows in
`PORT_STATUS.md`; do not rely on chat history for escape-hatch context.
- Review agents should verify this README's setup commands still match the PR,
then run the listed parity tests or report the exact blocker.
@@ -0,0 +1,69 @@
# <Model Family> Port Status
## Summary
- model_family: `<model_family>`
- workload_types: `<T2V/I2V/V2V/T2I/or compatibility shim with rationale>`
- official_ref: `<url or import path>`
- official_ref_dir: `<ReferenceDir or none>`
- hf_weights_path: `<HF repo id/url or local path>`
- local_weights_dir: `<official_weights/model_family or local path>`
- source_layout: `<diffusers/raw_official/monolithic/separate_components/mixed/custom/unknown>`
- local_tests_readme: `tests/local_tests/<model_family>/README.md`
## Current Phase
- phase: `prep`
- status: `in_progress`
- owner: `prep`
- last_updated: `<YYYY-MM-DD>`
## Component Matrix
| Component | Type | Reuse/Port | Official Definition | Official Instantiation | FastVideo Target | Prototype | Conversion | Parity | Open Issues |
|---|---|---|---|---|---|---|---|---|---|
| `<component>` | `<dit/vae/encoder/generic>` | `<unknown/reuse/port>` | `<path + symbols>` | `<path + args>` | `<target files>` | `<not_started/in_progress/pass/blocked>` | `<not_started/pass/blocked>` | `<not_started/scaffold_skip/debug_red/non_skip_pass/blocked>` | `<none or IDs>` |
## Conversion State
- conversion_script: `scripts/checkpoint_conversion/<model_family>_to_diffusers.py`
- converted_weights_dir: `converted_weights/<model_family>`
- source_layout: `<diffusers/separate_components/monolithic/mixed/custom/unknown>`
- strict_load_status: `not_run`
- passthrough_components: `<none or list>`
- retry_history: `<none>`
## Parity Commands
| Scope | Command | Last Result | Notes |
|---|---|---|---|
| component | `pytest tests/local_tests/<bucket>/test_<model_family>_<component>_parity.py -v -s` | `not_run` | `<notes>` |
| pipeline | `pytest tests/local_tests/pipelines/test_<model_family>_pipeline_parity.py -v -s` | `not_run` | `<notes>` |
## Open Questions
| ID | Question | Owner | Needed By Phase | Status | Resolution |
|---|---|---|---|---|---|
| Q001 | `<question>` | `<owner>` | `<phase>` | `<open/resolved>` | `<resolution or blank>` |
## Issues And Blockers
| ID | Phase | Component | Severity | Issue | Evidence | Owner | Status | Resolution |
|---|---|---|---|---|---|---|---|---|
| I001 | `<phase>` | `<component or all>` | `<low/medium/high/blocker>` | `<issue>` | `<logs/paths/commands>` | `<owner>` | `<open/resolved>` | `<resolution or blank>` |
## Escape Hatches
| ID | Phase | Decision Type | Question | Recommended Option | Status | Resolution |
|---|---|---|---|---|---|---|
| E001 | `<phase>` | `<scope/dependency/auth/cost/destructive/ambiguity/blocker>` | `<one precise question>` | `<safe recommended option>` | `<open/resolved>` | `<resolution or blank>` |
## Decisions
| Date | Decision | Rationale | Impact |
|---|---|---|---|
| `<YYYY-MM-DD>` | `<decision>` | `<why>` | `<affected components/phases>` |
## Handoff Notes
- `<short notes for the next agent>`
+234
View File
@@ -0,0 +1,234 @@
---
name: add-model-02-parity
description: Use during /add-model after reference/architecture study to scaffold and later activate local FastVideo component parity tests. Emphasizes early test creation, official-reference loading, standardized FastVideo loading, and non-skip handoff gates.
---
# Add Model Parity
## Goal
Create parity tests as early as possible in a FastVideo port. The first pass can
land before conversion or component implementation as an executable scaffold;
handoff is blocked until the same tests become non-skip PASS with real weights.
## When To Run
Follow `../add-model/shared/common_rules.md` for token/auth safety, state files,
escape hatches, and skip/pass semantics.
Run immediately after `/add-model` Phase 1 has identified:
- official component classes and call signatures;
- FastVideo target component buckets/classes/configs;
- local reference clone or import path from `add-model-01-prep`;
- local raw or Diffusers weight path;
- `official_env_status=imports_ok`, or private deps that will be stubbed
locally in tests;
- `local_tests_readme` documenting setup and planned review/test commands;
- expected component inputs and output tensors.
Do not wait for all FastVideo components to be implemented. Write the tests
first, then let component-porting subagents make them pass.
## Outputs
- One component parity test per required component, including reused components:
`tests/local_tests/<bucket>/test_<family>_<component>_parity.py`.
- Optional helper for upstream private deps:
`tests/local_tests/helpers/<family>_upstream.py`.
- Pipeline parity is owned later by `../add-model-09-pipeline/SKILL.md` after all
component parity tests pass non-skip.
- A parity status block for the `/add-model` parity verification phase.
## Early Scaffold Rules
- A scaffold may skip while the FastVideo class, converted weights, or official
import is missing.
- A scaffold must already encode the real official load path, FastVideo load
path, deterministic inputs, expected output extraction, and tolerance target.
- Each parity test must declare its coverage scope in the file docstring or a
module constant: `production_loader`, `implementation_subcomponent`, or `both`.
Implementation/subcomponent parity may bypass production loaders deliberately,
but final handoff still needs production-loader coverage somewhere before the
pipeline depends on that component.
- Official reference imports must run in the current FastVideo environment; do
not create or assume a separate upstream venv/conda env.
- A scaffold is not evidence of correctness. It becomes evidence only after a
local non-skip PASS.
- Prefer env-var path overrides with repo-relative defaults.
- Keep tests local-only under `tests/local_tests/`; package/CI quality tests are
added later.
- Update shared state files as described in
`../add-model/shared/common_rules.md` whenever adding or activating parity
tests.
## Component Template
Copy `templates/component_parity_test.py` and fill every `TODO` marker. The
template is distilled from:
- `tests/local_tests/transformers/test_ltx2.py`
- `tests/local_tests/transformers/test_gamecraft_parity.py`
- `tests/local_tests/encoders/test_ltx2_gemma_parity.py`
- `tests/local_tests/vaes/test_oobleck_vae_parity.py`
- `tests/local_tests/sd35/test_sd35_component_parity.py`
The template supports three states:
| State | Meaning |
|---|---|
| Scaffold skip | Test is committed early, but official import, FastVideo class, or weights are not available yet. |
| Debug red | Both sides load and the test fails numerically. This is useful: porting can chase the first drift. |
| Non-skip pass | Required before `/add-model` handoff. |
## Subagent Dispatch Pattern
After Phase 1, dispatch one parity subagent per component before or alongside
component implementation:
```text
Create a local parity test scaffold for <family> <component>.
Use the prep handoff:
- official_ref_dir/import: <...>
- local_weights_dir: <...>
- source_layout: <...>
- needs_conversion: <yes/no>
- official_env_status: <imports_ok | private_deps_need_stubs>
- local_tests_readme: tests/local_tests/<model_family>/README.md
- port_state_file: tests/local_tests/<model_family>/PORT_STATUS.md
- official_definition_files: <paths + classes/functions>
- official_instantiation_files: <paths + factory/pipeline/config call sites + args>
- concerns_or_unknowns: <known ambiguous inputs, outputs, deps, or args>
The complete per-component packet must match
`../add-model/contracts/component_context.md`.
Read the official component call path and the planned FastVideo component API.
Add tests/local_tests/<bucket>/test_<family>_<component>_parity.py based on
add-model-02-parity/templates/component_parity_test.py.
The scaffold must load the official model with real weights when available,
load the FastVideo model through the standardized config/class/loader path when
available, create deterministic inputs, compare concrete outputs, and skip only
when a dependency is genuinely missing. Do not make an unconditional skip or a
shape-only test.
```
## FastVideo Load Patterns
Pick the narrowest load path that matches the component:
| Component | Preferred FastVideo load path |
|---|---|
| DiT / transformer | Bucket config + model class, or `TransformerLoader` when testing converted Diffusers component dirs. |
| VAE | VAE class `from_pretrained(...)` when implemented, or bucket config + class for local converted dirs. |
| Text/image encoder | Bucket config + model class; pass HF subpaths from `local_weights_dir` or converted component dirs. |
| Scheduler/conditioner | Native class/config plus exact official kwargs. |
For early scaffolds, an import of the planned FastVideo class may be inside a
helper that calls `pytest.skip` if the class does not exist yet. Replace that
skip with a real import once the component PR adds the class.
Direct class/config construction is allowed for implementation or subcomponent
parity, such as connector-only encoder checks or official monolithic-checkpoint
mapping tests. Label that scope explicitly and add separate production-loader
coverage when converted component dirs are available.
## Official Load Patterns
- Clone/reference repo path: add its source dir to `sys.path` before imports.
- HF/Diffusers reference: import only inside the test, not production code.
- Private deps: add a helper under `tests/local_tests/helpers/` to install
stubs before importing upstream modules; do not rely on an external upstream
environment.
- Gated HF repos: resolve `HF_TOKEN`, `HUGGINGFACE_HUB_TOKEN`, or `HF_API_KEY`
under the token rules in `../add-model/shared/common_rules.md`.
## Non-Skip Activation Checklist
Before `/add-model` handoff, each scaffolded test must be activated:
```text
[ ] Official side imports and loads real weights.
[ ] FastVideo side imports and loads the converted or original weights.
[ ] Test executes at least one real forward call on both sides.
[ ] Test compares output tensors, not only shapes or state-dict keys.
[ ] Local pytest output contains PASSED, not SKIPPED or XFAIL.
[ ] Tolerance is justified for the component scope and kernel alignment.
```
## Component Parity Details
Reference imports:
- Import from `official_ref_dir` or the recorded package/import path.
- If upstream has private deps, add a helper under
`tests/local_tests/helpers/<family>_upstream.py` that installs minimal stubs
before importing upstream modules.
- Common stubs: identity compile/op-registration decorators, CP world size set to
1, identity scatter/gather, and test-friendly custom-op kernels.
- Stub decorators that register `torch.ops.<ns>.<op>` must preserve the
`torch.library` registration side-effect. Identity decorators alone are not
enough.
- Delete stub helpers and every `install_stubs()` call as soon as the real deps
become required installs. No-op shims are dead code.
Kernel and wrapper pitfalls:
- If parity routes flash-attn GQA through SDPA, expand KV heads manually on the
SDPA side with `repeat_interleave` along the head axis.
- If upstream VAE `decode()` denormalizes internally but FastVideo/Diffusers
expects pre-denormalized latents, apply `z = z * std + mean` only on the
FastVideo side in the parity test.
- Per-channel VAE `latents_mean` / `latents_std` must be reshaped explicitly,
e.g. `.view(1, z_dim, 1, 1, 1)` for 5D video latents.
Tolerance guide:
| Scope | Start `atol` / `rtol` | Notes |
|---|---|---|
| Single block, same kernel | `1e-4` / `1e-4` | Tight default. |
| Full DiT, aligned kernels | `1e-2` / `1e-2` | Cross-layer accumulation. |
| Full DiT, cross-kernel bf16 | `0.1` / `0.1` | Also require abs-mean drift below 5% and per-modality diagnostics. |
| VAE decode fp32 | `5e-2` / `5e-2` | After normalization alignment. |
| Encoder wrapper around same HF class | `1e-3` / `1e-3` | Should be near-zero. |
Element-wise `assert_close` alone is not enough for deep full-DiT parity. Also
log global abs-mean drift and per-modality summaries.
When a non-skip component parity run is numerically red after weight/input
checks, invoke `../add-model-08-trace/SKILL.md` before adding bespoke forward
hooks. Use `docs/contributing/activation_trace.md` to keep
`FASTVIDEO_TRACE_LAYERS`, `FASTVIDEO_TRACE_STATS`, and `FASTVIDEO_TRACE_STEPS`
identical across FastVideo and upstream traces.
Useful local commands:
```bash
pytest tests/local_tests/<bucket>/test_<family>_*parity*.py -v -s
pytest tests/local_tests -k "<family> and parity" -v -s
```
## Escape Hatches
Follow `../add-model/shared/common_rules.md`. Parity-specific ask cases include
private dependency approval, choosing between incompatible official references,
accepting a shape-only substitute, or loosening required tolerances.
## Pipeline Parity
Pipeline parity is later than component parity because it needs stages, presets,
registry wiring, converted weights, and green component parity. Record official
pipeline call notes in `local_tests_readme`, but do not treat pipeline parity as
owned by this skill.
Use `../add-model-09-pipeline/SKILL.md` and its
`templates/pipeline_parity_test.py` for pipeline parity scaffolding and
debugging. Compare denoised latents or decoded media, not just successful
generation.
## Handoff Status Block
Return `../add-model/contracts/parity_status.md` to `/add-model` and update the
shared state files before handoff.
@@ -0,0 +1,201 @@
# SPDX-License-Identifier: Apache-2.0
"""Component parity scaffold for <FAMILY> <COMPONENT>.
This file is intended to be created early in a port. It may skip until the
official reference, FastVideo class, and real weights are available, but it must
never become an unconditional skip or shape-only test.
Fill every TODO before considering this test active.
"""
from __future__ import annotations
import importlib
import os
from pathlib import Path
import sys
import pytest
import torch
from torch.testing import assert_close
os.environ.setdefault("MASTER_ADDR", "localhost")
os.environ.setdefault("MASTER_PORT", "29519")
os.environ.setdefault("DISABLE_SP", "1")
os.environ.setdefault("FASTVIDEO_ATTENTION_BACKEND", "TORCH_SDPA")
REPO_ROOT = Path(__file__).resolve().parents[3]
FAMILY = "<family>" # TODO: snake_case family name.
COMPONENT = "<component>" # TODO: transformer | vae | encoder | conditioner | ...
PARITY_SCOPE = "implementation_subcomponent" # TODO: production_loader | implementation_subcomponent | both
OFFICIAL_MODULE = "<official.module>" # TODO: e.g. "ltx_core.model.transformer".
OFFICIAL_CLASS = "<OfficialClass>" # TODO: official class/factory name.
FASTVIDEO_CONFIG_MODULE = "fastvideo.configs.models.<bucket>" # TODO.
FASTVIDEO_CONFIG_CLASS = "<FastVideoConfig>" # TODO.
FASTVIDEO_MODEL_MODULE = "fastvideo.models.<bucket>.<module>" # TODO.
FASTVIDEO_MODEL_CLASS = "<FastVideoModel>" # TODO.
OFFICIAL_REF_DIR = Path(
os.getenv("<FAMILY_UPPER>_OFFICIAL_REF_DIR", REPO_ROOT / "<ReferenceDir>")
)
LOCAL_WEIGHTS_DIR = Path(
os.getenv("<FAMILY_UPPER>_LOCAL_WEIGHTS_DIR", REPO_ROOT / "official_weights" / FAMILY)
)
CONVERTED_WEIGHTS_DIR = Path(
os.getenv("<FAMILY_UPPER>_CONVERTED_WEIGHTS_DIR", REPO_ROOT / "converted_weights" / FAMILY)
)
def _resolve_hf_token() -> str | None:
for key in ("HF_TOKEN", "HUGGINGFACE_HUB_TOKEN", "HF_API_KEY"):
value = os.environ.get(key)
if value:
return value
return None
def _add_official_to_path() -> None:
"""Add the official source path before importing upstream modules."""
# TODO: adjust for the official repo layout. Common examples:
# OFFICIAL_REF_DIR / "src"
# OFFICIAL_REF_DIR / "packages" / "<pkg>" / "src"
# OFFICIAL_REF_DIR
official_src = OFFICIAL_REF_DIR / "src"
if not official_src.exists():
official_src = OFFICIAL_REF_DIR
if official_src.exists() and str(official_src) not in sys.path:
sys.path.insert(0, str(official_src))
def _import_or_skip(module_name: str, attr_name: str | None = None):
if "<" in module_name or (attr_name is not None and "<" in attr_name):
pytest.skip(f"Template import placeholder not filled: {module_name}.{attr_name}")
try:
module = importlib.import_module(module_name)
except Exception as exc: # noqa: BLE001 - local parity should skip missing refs.
pytest.skip(f"Cannot import {module_name}: {exc}")
if attr_name is None:
return module
try:
return getattr(module, attr_name)
except AttributeError:
pytest.skip(f"{module_name} has no attribute {attr_name}")
def _load_official_model(device: torch.device, dtype: torch.dtype) -> torch.nn.Module:
"""Load the official component with real weights."""
_add_official_to_path()
if not OFFICIAL_REF_DIR.exists():
pytest.skip(f"Official reference missing: {OFFICIAL_REF_DIR}")
if not LOCAL_WEIGHTS_DIR.exists():
pytest.skip(f"Local weights missing: {LOCAL_WEIGHTS_DIR}")
# TODO: import official class/factory and load real weights strictly.
# Examples in-tree:
# - LTX2: SingleGPUModelBuilder(...).build(device=device, dtype=dtype)
# - GameCraft: torch.load(...)["module"] -> official_model.load_state_dict(...)
# - Oobleck: create_model_from_config(config) + ckpt state_dict
OfficialClass = _import_or_skip(OFFICIAL_MODULE, OFFICIAL_CLASS)
model = OfficialClass() # TODO: pass official config kwargs.
state_dict = {} # TODO: load official state dict from LOCAL_WEIGHTS_DIR.
missing, unexpected = model.load_state_dict(state_dict, strict=True)
assert not missing and not unexpected, (
f"official load mismatch missing={missing[:5]} unexpected={unexpected[:5]}"
)
return model.to(device=device, dtype=dtype).eval()
def _load_fastvideo_model(device: torch.device, dtype: torch.dtype) -> torch.nn.Module:
"""Load the FastVideo component with the same tensor content."""
if not CONVERTED_WEIGHTS_DIR.exists() and not LOCAL_WEIGHTS_DIR.exists():
pytest.skip(
f"No FastVideo loadable weights: {CONVERTED_WEIGHTS_DIR} or {LOCAL_WEIGHTS_DIR}"
)
# TODO: replace with the bucket-specific FastVideo config/class/loader.
# DiT examples:
# from fastvideo.configs.models.dits import <Config>
# from fastvideo.models.dits.<module> import <Model>
# VAE examples:
# from fastvideo.models.vaes.<module> import <VAE>
# model = <VAE>.from_pretrained(...)
FastVideoConfig = _import_or_skip(FASTVIDEO_CONFIG_MODULE, FASTVIDEO_CONFIG_CLASS)
FastVideoModel = _import_or_skip(FASTVIDEO_MODEL_MODULE, FASTVIDEO_MODEL_CLASS)
config = FastVideoConfig()
model = FastVideoModel(config=config)
state_dict = {} # TODO: load converted or directly mapped state dict.
missing, unexpected = model.load_state_dict(state_dict, strict=True)
assert not missing and not unexpected, (
f"FastVideo load mismatch missing={missing[:5]} unexpected={unexpected[:5]}"
)
return model.to(device=device, dtype=dtype).eval()
def _make_inputs(device: torch.device, dtype: torch.dtype) -> dict[str, torch.Tensor]:
"""Create deterministic inputs matching the official component call."""
torch.manual_seed(0)
# TODO: replace with component-specific tensors and metadata.
return {
"hidden_states": torch.randn(1, 4, 16, device=device, dtype=dtype),
"timestep": torch.tensor([10], device=device),
}
def _run_official(model: torch.nn.Module, inputs: dict[str, torch.Tensor]) -> torch.Tensor:
"""Run official component and return the tensor to compare."""
with torch.inference_mode():
output = model(**inputs) # TODO: adapt official call signature.
if isinstance(output, dict):
sample = output.get("sample")
output = sample if sample is not None else output.get("x")
elif hasattr(output, "sample"):
output = output.sample
elif isinstance(output, tuple):
output = output[0]
assert torch.is_tensor(output), f"official output is not tensor: {type(output)}"
return output.detach().float().cpu()
def _run_fastvideo(model: torch.nn.Module, inputs: dict[str, torch.Tensor]) -> torch.Tensor:
"""Run FastVideo component and return the tensor to compare."""
with torch.inference_mode():
output = model(**inputs) # TODO: adapt FastVideo call signature.
if isinstance(output, dict):
sample = output.get("sample")
output = sample if sample is not None else output.get("x")
elif hasattr(output, "sample"):
output = output.sample
elif isinstance(output, tuple):
output = output[0]
assert torch.is_tensor(output), f"FastVideo output is not tensor: {type(output)}"
return output.detach().float().cpu()
@pytest.mark.skipif(not torch.cuda.is_available(), reason="CUDA required for this parity test.")
def test_component_parity():
"""Compare official and FastVideo outputs on identical inputs."""
device = torch.device("cuda:0")
dtype = torch.bfloat16
official = _load_official_model(device, dtype)
fastvideo = _load_fastvideo_model(device, dtype)
inputs = _make_inputs(device, dtype)
official_out = _run_official(official, inputs)
fastvideo_out = _run_fastvideo(fastvideo, inputs)
assert official_out.shape == fastvideo_out.shape
diff = (official_out - fastvideo_out).abs()
print(
f"official abs_mean={official_out.abs().mean().item():.6f} "
f"fastvideo abs_mean={fastvideo_out.abs().mean().item():.6f} "
f"diff_max={diff.max().item():.6f} diff_mean={diff.mean().item():.6f}"
)
# TODO: pick tolerance by scope:
# - single block / same kernel: 1e-4
# - full DiT aligned kernels: 1e-2
# - full DiT cross-kernel bf16: 1e-1 + abs_mean drift check
# - VAE decode fp32: 5e-2 after normalization alignment
assert_close(fastvideo_out, official_out, atol=1e-4, rtol=1e-4)
@@ -0,0 +1,114 @@
---
name: add-model-03-port-dit
description: Use during /add-model Phase 4 or Phase 6 to prototype or parity-debug one FastVideo-native DiT/transformer component.
---
# Add Model Port DiT
## Goal
Prototype or parity-debug one diffusion transformer in FastVideo-native code.
This skill is for one component only; do not work on the VAE, encoders,
pipeline, or unrelated conversion code unless the current component cannot load
without a minimal fix there.
## Inputs
Follow `../add-model/shared/component_skill_common.md` and require the complete
packet from `../add-model/contracts/component_context.md`.
DiT-specific packet fields:
- `component`: transformer or DiT name.
- `parity_test`: `tests/local_tests/<bucket>/test_<family>_<component>_parity.py`.
- `weights`: converted transformer dir or local official path.
- `target_files`: `fastvideo/models/dits/<family>.py` and
`fastvideo/configs/models/dits/<family>.py`.
## Modes
Use the common prototype and parity-debug modes from
`../add-model/shared/component_skill_common.md`.
DiT-specific prototype concerns include ambiguous official flags, shape
mismatches, missing FastVideo layer equivalents, and dedicated output heads.
## Reuse Proof
Apply the shared reuse proof. DiT-specific comparison must include attention
algorithm, positional embeddings, RoPE/patching, timestep/guidance embeddings,
scaling constants, dtype casts, state-dict names, and every output head.
## Existing FastVideo Patterns
- Base class: `fastvideo/models/dits/base.py::BaseDiT`.
- Config bases: `DiTConfig` and `DiTArchConfig` in
`fastvideo/configs/models/dits/base.py`.
- Use the matching DiT config bucket. Wrong bucket inheritance can typecheck but
fail during pipeline wiring.
- Config export: add the config to
`fastvideo/configs/models/dits/__init__.py`.
- Registry discovery: set `EntryClass = <ClassName>` in the model file.
- Loader path: `TransformerLoader` reads `transformer/config.json`, calls
`dit_config.update_model_arch(config)`, resolves `_class_name` through
`ModelRegistry`, and constructs the class with `config` and `hf_config`.
- Reference examples: `stable_audio.py`, `wanvideo.py`, `sd3.py`, `longcat.py`,
and `ltx2.py`.
- Layer guidance: `fastvideo/layers/AGENTS.md`.
## Implementation Rules
- Use FastVideo-native layers by default: `ReplicatedLinear` for DiT hot-path
linears, `DistributedAttention` for standard full-sequence attention, and
`LocalAttention` for local/window attention or simple single-GPU parity paths.
- Raw SDPA is acceptable for cross-modality flat streams when no FastVideo
distributed equivalent exists; document the SP gap in the module docstring.
- Mirror official tensor contracts exactly: latent packing, patch ordering,
timestep embedding scale, RoPE/positional embedding, guidance embedding,
cross-attention context order, output head order, and dtype casts.
- Preserve all output heads that the official DiT emits. Do not silently drop
audio, depth, pose, mask, or auxiliary heads.
- Put architecture fields on `DiTArchConfig`; keep inference steps, CFG scales,
FPS, flow shift, and sampling defaults out of the arch config.
- Define `_fsdp_shard_conditions`, `_compile_conditions`,
`param_names_mapping`, and `reverse_param_names_mapping` where needed.
- Follow the production import boundary in
`../add-model/shared/common_rules.md`.
## Prototype Checks
Follow the shared prototype success criteria. A useful one-off check is:
```bash
python - <<'PY'
# Import the target config/class, instantiate with random weights, and print
# state_dict names/shapes for the conversion mapping.
PY
```
## Parity-Debug Loop
Run the shared parity-debug loop. The component test command is:
```bash
pytest <parity_test> -v -s
```
For numerical drift, use `../add-model-08-trace/SKILL.md` before writing bespoke
hooks. Start with FastVideo's activation trace (`fastvideo/hooks/activation_trace.py`;
`docs/contributing/activation_trace.md`) and a block-level regex such as
`FASTVIDEO_TRACE_LAYERS="^block\.layers\.[0-9]+$"`. Only fall back to custom
per-block hooks if the needed boundary or statistic is not exposed by
`FASTVIDEO_TRACE_STATS`.
## Escape Hatches
Follow `../add-model/shared/common_rules.md` and the component-specific guidance
in `../add-model/shared/component_skill_common.md`. DiT-specific ask cases include
dropping an output head/modality, accepting an unsupported kernel/private op, or
choosing between incompatible official transformer definitions.
## Handoff
Return `../add-model/contracts/component_skill_handoff.md` following the common
handoff rules in `../add-model/shared/component_skill_common.md`.
@@ -0,0 +1,100 @@
---
name: add-model-04-port-vae
description: Use during /add-model Phase 4 or Phase 6 to prototype or parity-debug one FastVideo-native VAE component.
---
# Add Model Port VAE
## Goal
Prototype or parity-debug one VAE or autoencoder in FastVideo-native code. This
skill covers video, image, and audio VAEs.
## Inputs
Follow `../add-model/shared/component_skill_common.md` and require the complete
packet from `../add-model/contracts/component_context.md`.
VAE-specific packet fields:
- `component`: VAE or autoencoder name.
- `parity_test`: `tests/local_tests/vaes/test_<family>_<component>_parity.py`.
- `weights`: converted VAE dir, HF subfolder, or local official path.
- `target_files`: `fastvideo/models/vaes/<arch_or_family>.py` and
`fastvideo/configs/models/vaes/<arch_or_family>.py`.
## Modes
Use the common prototype and parity-debug modes from
`../add-model/shared/component_skill_common.md`.
VAE-specific prototype concerns include latent normalization, stochastic
posterior behavior, tiling incompatibility, temporal/spatial/audio layout, and
decode output containers.
## Reuse Proof
Apply the shared reuse proof. VAE-specific comparison must include latent layout,
temporal/spatial/audio compression, scaling factor, mean/std normalization,
posterior behavior, encode/decode output objects, tiling flags, and cropping.
## Existing FastVideo Patterns
- Shared tiling wrapper: `fastvideo/models/vaes/common.py::ParallelTiledVAE`.
- Config bases: `VAEConfig` and `VAEArchConfig` in
`fastvideo/configs/models/vaes/base.py`.
- Use the matching VAE config bucket. Wrong bucket inheritance can typecheck but
fail during pipeline wiring.
- Config export: add the config to
`fastvideo/configs/models/vaes/__init__.py`.
- Registry discovery: set `EntryClass = <ClassName>` in the model file.
- Loader path: VAE loaders resolve `_class_name` through `ModelRegistry` and
load converted component weights from the VAE subdir.
- Reference examples: `oobleck.py`, `autoencoder_kl.py`, `wanvae.py`,
`ltx2vae.py`, and `gamecraftvae.py`.
- Layer guidance: `fastvideo/layers/AGENTS.md`.
## Implementation Rules
- Name reusable VAE architectures by architecture (`oobleck.py`,
`autoencoder_kl.py`); name family-specific VAEs by family.
- Match official encode/decode contracts exactly: input layout, latent layout,
temporal/spatial/audio compression, scaling factor, mean/std normalization,
posterior sampling behavior, decode output object, and frame/sample cropping.
- Compare deterministic outputs in parity: decode outputs, encode mean/mode, or
round-trip tensors. Do not compare stochastic samples unless the RNG path is
explicitly controlled.
- Use FastVideo tiling only when it preserves official numerics for the tested
shape; disable it in config for audio or unsupported dimensions.
- Put architecture constants on `VAEArchConfig`; put `load_encoder`,
`load_decoder`, tiling, dtype, and pretrained path fields on `VAEConfig`.
- Follow the production import boundary in
`../add-model/shared/common_rules.md`.
## Prototype Checks
Follow the shared prototype success criteria.
## Parity-Debug Loop
Run the shared parity-debug loop. The component test command is:
```bash
pytest <parity_test> -v -s
```
For numerical drift, check normalization, latent scaling, posterior mode vs
sample, channel order, and temporal/spatial/audio cropping before changing
layers.
## Escape Hatches
Follow `../add-model/shared/common_rules.md` and the component-specific guidance
in `../add-model/shared/component_skill_common.md`. VAE-specific ask cases include
dropping an encode/decode path, accepting an unsupported private op, or choosing
between incompatible official VAE definitions.
## Handoff
Return `../add-model/contracts/component_skill_handoff.md` following the common
handoff rules in `../add-model/shared/component_skill_common.md`.
@@ -0,0 +1,119 @@
---
name: add-model-05-port-encoder
description: Use during /add-model Phase 4 or Phase 6 to prototype or parity-debug one FastVideo-native text, image, audio, or compound encoder component.
---
# Add Model Port Encoder
## Goal
Prototype or parity-debug one encoder or encoder-like conditioner in
FastVideo-native code. Use this for text encoders, image encoders, audio
encoders, and compound conditioners that fit the encoder config/loader bucket.
## Inputs
Follow `../add-model/shared/component_skill_common.md` and require the complete
packet from `../add-model/contracts/component_context.md`.
Encoder-specific packet fields:
- `component`: encoder or encoder-like conditioner name.
- `parity_test`: `tests/local_tests/encoders/test_<family>_<component>_parity.py`.
- `weights`: converted encoder dir, HF subfolder, or external HF id.
- `target_files`: `fastvideo/models/encoders/<arch_or_family>.py` and
`fastvideo/configs/models/encoders/<arch_or_family>.py`.
## Modes
Use the common prototype and parity-debug modes from
`../add-model/shared/component_skill_common.md`.
Encoder-specific prototype concerns include tokenizer kwargs, hidden-state
extraction, output packing, connector order, and external/passthrough weight
needs.
## Reuse Proof
Apply the shared reuse proof. Encoder-specific comparison must include tokenizer
contracts, hidden-state extraction, masks, positional IDs, output packing,
connector/projection ordering, passthrough paths, and returned dataclass shape.
## Existing FastVideo Patterns
- Base classes: `TextEncoder` and `ImageEncoder` in
`fastvideo/models/encoders/base.py`.
- Output type: `BaseEncoderOutput`.
- Config bases: `TextEncoderConfig`, `ImageEncoderConfig`,
`TextEncoderArchConfig`, and `ImageEncoderArchConfig` in
`fastvideo/configs/models/encoders/base.py`.
- Use the matching encoder config bucket. Wrong bucket inheritance can typecheck
but fail during pipeline wiring.
- Config export: add the config to
`fastvideo/configs/models/encoders/__init__.py`.
- Registry discovery: set `EntryClass = <ClassName>` or a list of class names in
the model file.
- Reference examples: native `t5.py`, `clip.py`, `siglip.py`, `llama.py`,
`qwen2_5.py`, `gemma.py`, and compound `stable_audio_conditioner.py`.
- Layer guidance: `fastvideo/layers/AGENTS.md`.
## Implementation Rules
- Reuse tokenizers and pure data utilities when needed, but do not add runtime
third-party model-class imports as a placeholder for a component that owns
weights or numerical behavior.
- For LLM-style encoders, follow existing tensor-parallel patterns such as
`QKVParallelLinear`, `MergedColumnParallelLinear`, `RowParallelLinear`,
`VocabParallelEmbedding`, and `RMSNorm` when matching native examples.
- Match official hidden-state extraction exactly: layer index, pooled output,
attention mask dtype, padding side, truncation, special tokens, final norm,
output_hidden_states, and returned tuple/dataclass shape.
- For connector or conditioner modules, preserve sub-conditioner order and the
exact packing of cross-attention tokens, masks, and global conditioning.
- Put tokenizer kwargs and architecture constants on the arch config when they
affect numerical behavior.
- If an external HF encoder is explicitly accepted as a lazy wrapper, keep it
isolated, document why it is not a native port, and still require parity for
the wrapper's output contract.
Hybrid external-HF encoder checklist:
- Put external model folders in passthrough subfolders such as
`text_encoder/<external_name>/`, or record a root `model_index.json` path field
that the loader resolves to a local directory.
- Keep external model parameters out of the FastVideo-owned state-dict surface
when the external model is loaded lazily from its own HF files.
- Convert and strict-check only the FastVideo-owned connector/projection weights;
document external model weights as passthrough.
- Add parity for the wrapper's final output contract and, when useful, a narrower
connector-only parity test that labels its scope as
`implementation_subcomponent`.
- Verify the production loader resolves the same external path used by the
pipeline, not just the direct class used in the parity test.
## Prototype Checks
Follow the shared prototype success criteria.
## Parity-Debug Loop
Run the shared parity-debug loop. The component test command is:
```bash
pytest <parity_test> -v -s
```
For numerical drift, check tokenization, masks, hidden-state selection,
positional IDs, dtype/autocast, and output packing before changing layers.
## Escape Hatches
Follow `../add-model/shared/common_rules.md` and the component-specific guidance
in `../add-model/shared/component_skill_common.md`. Encoder-specific ask cases
include accepting private model-code execution, choosing between incompatible
tokenizer/encoder references, or dropping a required conditioning stream.
## Handoff
Return `../add-model/contracts/component_skill_handoff.md` following the common
handoff rules in `../add-model/shared/component_skill_common.md`.
@@ -0,0 +1,111 @@
---
name: add-model-06-port-generic
description: Use during /add-model Phase 4 or Phase 6 to prototype or parity-debug one non-DiT, non-VAE, non-encoder FastVideo component.
---
# Add Model Port Generic
## Goal
Prototype or parity-debug one scheduler, conditioner, upsampler, vocoder,
adapter, preprocessor, or unknown component in FastVideo-native code.
## Inputs
Follow `../add-model/shared/component_skill_common.md` and require the complete
packet from `../add-model/contracts/component_context.md`.
Generic-component packet fields:
- `component`: component name.
- `component_type`: scheduler, conditioner, upsampler, vocoder, adapter,
preprocessor, or unknown.
- `parity_test`: `tests/local_tests/<bucket>/test_<family>_<component>_parity.py`.
- `weights`: converted component dir, HF subfolder, or none.
- `target_files`: matching `fastvideo/models/` and `fastvideo/configs/models/`
bucket files when applicable.
## Modes
Use the common prototype and parity-debug modes from
`../add-model/shared/component_skill_common.md`.
Generic-component prototype concerns include stateless/stateful ambiguity,
missing loader buckets, source prefixes, mutable scheduler state, and output
container shape.
## Reuse Proof
Apply the shared reuse proof. Generic-component comparison must include mutable
state, scaling constants, scheduler/conditioner semantics, output containers, and
whether the component owns state or is stateless.
## Existing FastVideo Patterns
- Schedulers live under `fastvideo/models/schedulers/` and expose `EntryClass`.
- Upsamplers use `fastvideo/models/upsamplers/` plus configs under
`fastvideo/configs/models/upsamplers/`; see `hunyuan15.py`.
- Vocoders and audio-specific modules can live under `fastvideo/models/audio/`
with configs under `fastvideo/configs/models/audio/`; see `ltx2_audio_vae.py`.
- Compound conditioners may fit the encoder bucket when the pipeline loader uses
`ConditionerLoader`; see `stable_audio_conditioner.py`.
- Registry discovery uses `EntryClass`; config bucket exports are required when
pipeline configs import them by bucket.
- Use the narrowest matching config bucket. Wrong bucket inheritance can typecheck
but fail during pipeline wiring.
- Layer guidance: `fastvideo/layers/AGENTS.md`.
## Bucket Decision
- If the component is a transformer/DiT, stop and use `add-model-03-port-dit`.
- If the component is a VAE/autoencoder, stop and use `add-model-04-port-vae`.
- If the component is a text/image/audio encoder or encoder-like conditioner,
stop and use `add-model-05-port-encoder` unless the loader requires a different
bucket.
- Otherwise choose the narrowest existing bucket. Add a new bucket only when no
existing loader/config shape can represent the component without misleading
names or unsafe runtime behavior.
## Implementation Rules
- Match official behavior, not just shapes: constructor args, default values,
runtime flags, RNG use, dtype/autocast, scaling constants, masks, and output
containers all matter.
- Keep the implementation minimal and native. Do not keep a runtime import of
the official implementation as the production component.
- For schedulers, compare timesteps, sigmas/noise levels, step outputs, shift
handling, prediction type, and any mutable internal state.
- For upsamplers, compare resize mode, align_corners, residual branches,
causal padding, normalization, and exact target-shape behavior.
- For vocoders/audio components, compare waveform shape, sample-rate contract,
channel order, hop length, normalization, and dtype.
- If private upstream deps are required only for tests, keep stubs under
`tests/local_tests/helpers/` and do not import them from production code.
## Prototype Checks
Follow the shared prototype success criteria.
## Parity-Debug Loop
Run the shared parity-debug loop. The component test command is:
```bash
pytest <parity_test> -v -s
```
For numerical drift, add targeted intermediate comparisons in the test to
identify the first divergent operation.
## Escape Hatches
Follow `../add-model/shared/common_rules.md` and the component-specific guidance
in `../add-model/shared/component_skill_common.md`. Generic-component ask cases
include creating a new loader bucket, accepting an unsupported private op,
choosing between incompatible official definitions, or dropping a required
component.
## Handoff
Return `../add-model/contracts/component_skill_handoff.md` following the common
handoff rules in `../add-model/shared/component_skill_common.md`.
@@ -0,0 +1,186 @@
---
name: add-model-07-conversion
description: Use during /add-model Phase 5 to write and verify a FastVideo checkpoint conversion script after native component prototypes expose FastVideo state-dict keys/shapes.
---
# Add Model Conversion
## Goal
Convert official weights into a FastVideo-loadable component layout after Phase 4
native prototypes exist. The conversion script owns parameter mapping, component
splitting, passthrough assets, config emission, and strict-load verification.
## Inputs
Follow `../add-model/shared/common_rules.md` for token/auth safety, state files,
escape hatches, production boundaries, and skip/pass semantics.
Require the initial request from
`../add-model/contracts/conversion_request.md`.
If the FastVideo key/shape dump is missing, return to `/add-model` Phase 4. Do
not write a final mapping against an unimplemented component.
For Phase 6 retry requests from component skills, also require the retry shape
from `../add-model/contracts/conversion_request.md`.
## Output
- `scripts/checkpoint_conversion/<family>_to_diffusers.py`.
- `converted_weights/<family>/` with `model_index.json` and per-component
subfolders.
- Updated `tests/local_tests/<model_family>/README.md` with conversion command,
source layout, output path, and strict-load status.
- Updated `tests/local_tests/<model_family>/PORT_STATUS.md` with conversion
state, retry history, open questions, and issues/blockers.
## Reference Scripts
- `scripts/checkpoint_conversion/convert_ltx2_weights.py`: component prefix
splitting, metadata config extraction, passthrough Gemma/tokenizer assets, and
optional component-only output.
- `scripts/checkpoint_conversion/stable_audio_to_diffusers.py`: monolithic
`model.safetensors` split into transformer/VAE/conditioner, plus copied
passthrough subfolders. Use this shape for single-checkpoint official repos.
- `scripts/checkpoint_conversion/convert_gamecraft_full.py`: separate official
sources for transformer, VAE, text encoders, tokenizers, scheduler, and root
`model_index.json`.
- `scripts/checkpoint_conversion/longcat_to_fastvideo.py`: fused QKV/KV split,
renamed native transformer weights, and copied existing Diffusers components.
- `scripts/checkpoint_conversion/pt_to_safetensors.py`: simple `.pt` extraction
helper for nested checkpoint dictionaries.
## Source Layout Decision
Choose exactly one primary layout:
| Layout | Conversion behavior |
|---|---|
| `diffusers` | Usually no tensor remap; verify configs/classes and copy or update `_class_name` only when needed. |
| `raw_official` | Convert a raw official checkpoint file or directory. Choose explicit component ownership before writing output. |
| `separate_components` | Convert/copy each component from its own file or directory. |
| `monolithic` | Load one model checkpoint and split state dict by authoritative prefixes into component buckets. |
| `mixed` | Convert some components and copy passthrough components such as tokenizers, text encoders, schedulers, or already-Diffusers VAE dirs. |
| `custom` | Document why none of the above fits before writing conversion code. |
Monolithic checkpoints need explicit prefix ownership. For example, Stable Audio
uses one `model.safetensors` with DiT, pretransform/VAE, and conditioner keys;
the converter splits those keys into FastVideo component subfolders and writes
per-component configs.
## Script Shape
Start from `templates/family_to_diffusers.py` or the closest reference script.
Keep the script explicit and reviewable:
- `COMPONENT_SPECS` or `COMPONENT_PREFIXES` declares component ownership.
- `PARAM_NAME_MAP` declares key renames.
- `SKIP_PATTERNS` declares intentionally dropped training-only keys.
- tensor split/fuse helpers are named by operation, e.g. `split_qkv`.
- `build_component_configs(...)` writes loader-compatible config files. Most
model components use `config.json`; schedulers use `scheduler_config.json`.
- `build_model_index(...)` writes a root `model_index.json` matching the target
FastVideo pipeline and component classes.
- verification reports missing, unexpected, skipped, unchanged, renamed, and
shape-mismatched keys.
`model_index.json` library tokens must match FastVideo loaders:
- standard native DiT/VAE/audio/vocoder/upsampler components loaded by existing
Diffusers-style loaders usually use `"diffusers"` with a FastVideo
`_class_name` in the component `config.json`;
- text encoders, tokenizers, image encoders, processors, and feature extractors
usually use `"transformers"`;
- `conditioner` currently expects `"fastvideo"`;
- use fully qualified `"fastvideo.<module>"` only when intentionally relying on
the custom fastvideo-library escape path;
- do not write bare `"fastvideo"` for transformer, VAE, or other loaders that
expect `"diffusers"` unless the loader explicitly expects it.
## Mapping Rules
- Use Phase 4 key/shape dumps to derive mappings. Do not guess from official key
names alone.
- Preserve each component's official file paths, parity test path, and prototype
concerns in comments or structured constants near the mapping that uses them.
- Every official inference parameter should be mapped, copied through, or listed
as intentionally skipped with a reason.
- Every FastVideo prototype parameter should receive a tensor or be listed as an
intentional external/passthrough parameter.
- Shape matches are necessary but not sufficient; check semantic pairing for
Q/K/V, gate/up/down, norm scale/bias, LoRA/base, and modality-specific heads.
- If official and FastVideo fuse or split tensors differently, convert tensors in
the script rather than changing production code to match checkpoint quirks.
## Verification
Run conversion locally, then verify before returning to Phase 6. For retry
requests, update the mapping, rerun conversion, and refresh only the implicated
converted component when safe; otherwise rerun the full conversion.
```bash
python scripts/checkpoint_conversion/<family>_to_diffusers.py \
--src <official_weights> \
--revision <hf_revision> \
--dst converted_weights/<model_family>
```
Omit `--revision` for local sources or when prep recorded `default` / `none`.
Minimum output layout:
```text
converted_weights/<family>/
model_index.json
transformer/config.json
transformer/*.safetensors
vae/config.json
vae/*.safetensors
scheduler/scheduler_config.json as needed
text_encoder/... as needed
```
Required checks:
- `model_index.json` exists and lists every required component.
- Each converted component has the config filename its loader expects and
safetensors weights when it owns weights. Scheduler dirs require
`scheduler_config.json`; most other native model dirs use `config.json`.
- Weight filenames may vary by loader: transformer and VAE loaders glob all
`*.safetensors`; text encoders may load `*.safetensors`, `*.bin`, and
sometimes `*.pt`; `conditioner` currently expects
`diffusion_pytorch_model.safetensors`. Use the loader's actual accepted layout
rather than assuming one global filename.
- Passthrough components are copied or referenced deliberately.
- Each emitted component config validates through the same path production
loaders use. Instantiate the relevant config and call `update_model_arch(...)`
or `update_model_config(...)` with the emitted JSON so unknown keys fail during
conversion, not at pipeline load time.
- Record production loader strictness for every stateful component. If the loader
intentionally uses non-strict loading, add explicit missing/unexpected-key
assertions in the parity test and document exactly which keys are allowed.
- Each new FastVideo component strict-loads converted weights where its production
loader is strict. If strict loading is impossible, record the exact allowed
missing/unexpected keys and why they are not inference weights.
- Retry fixes include the original component parity evidence and the new
strict-load result in `local_tests_readme` so the component subagent can resume
without rediscovering context.
- `local_tests_readme` records the command, output directory, and strict-load
result.
Do not chase numerical parity in this skill except to identify a conversion
mapping bug. Long parity-debug loops belong to `/add-model` Phase 6.
## Escape Hatches
Follow `../add-model/shared/common_rules.md`. Conversion-specific ask cases
include selecting between incompatible official checkpoints, publishing/uploading
weights, overwriting an existing converted repo not created by this run,
accepting non-strict missing inference weights, or dropping a component/output
from scope.
## Handoff
Return `../add-model/contracts/conversion_handoff.md` and update the shared state
files before handoff.
@@ -0,0 +1,312 @@
#!/usr/bin/env python3
# SPDX-License-Identifier: Apache-2.0
"""Convert <model_family> official weights to a FastVideo Diffusers-style tree.
This template supports both separate component sources and a monolithic pipeline
checkpoint that must be split by component prefix. Replace every TODO before
using it for a real port.
"""
from __future__ import annotations
import argparse
import json
import os
import re
import shutil
from collections import OrderedDict
from pathlib import Path
from typing import Any
import torch
from safetensors import safe_open
from safetensors.torch import load_file, save_file
try:
from huggingface_hub import snapshot_download
except ImportError: # pragma: no cover - optional local conversion dependency
snapshot_download = None
# TODO: fill with authoritative component prefixes for monolithic checkpoints.
# Example: {"model.model.": "transformer", "pretransform.model.": "vae"}
COMPONENT_PREFIXES: dict[str, str] = {}
# TODO: fill with component-specific source paths for separate-component repos.
# Example: {"transformer": "transformer/model.safetensors", "vae": "vae/"}
SEPARATE_COMPONENT_PATHS: dict[str, str] = {}
# TODO: copy passthrough dirs that are already loadable by FastVideo/Diffusers.
PASSTHROUGH_SUBFOLDERS: tuple[str, ...] = ("tokenizer", "scheduler")
# TODO: add regex renames derived from Phase 4 key/shape dumps.
PARAM_NAME_MAP: dict[str, str] = {}
# TODO: include training-only or dynamically-computed keys that must not load.
SKIP_PATTERNS: tuple[str, ...] = ()
def _hf_token() -> str | None:
return (
os.environ.get("HF_TOKEN") or os.environ.get("HUGGINGFACE_HUB_TOKEN")
or os.environ.get("HF_API_KEY")
)
def resolve_src(src: str, revision: str | None) -> Path:
if os.path.exists(src):
return Path(src)
if snapshot_download is None:
raise RuntimeError("huggingface_hub is required when --src is a repo id")
return Path(snapshot_download(repo_id=src, revision=revision, token=_hf_token()))
def load_checkpoint(path: Path) -> dict[str, torch.Tensor]:
if path.is_dir():
weights: dict[str, torch.Tensor] = {}
for shard in sorted(path.glob("*.safetensors")):
weights.update(load_file(str(shard)))
if weights:
return weights
raise FileNotFoundError(f"No safetensors found in {path}")
if path.suffix == ".safetensors":
return load_file(str(path))
checkpoint = torch.load(path, map_location="cpu", weights_only=True)
if isinstance(checkpoint, dict):
for key in ("state_dict", "model_state_dict", "model", "module", "ema"):
if key in checkpoint and isinstance(checkpoint[key], dict):
return checkpoint[key]
return checkpoint
raise TypeError(f"Unsupported checkpoint type: {type(checkpoint)!r}")
def should_skip_key(key: str) -> bool:
return any(re.search(pattern, key) for pattern in SKIP_PATTERNS)
def apply_mapping(key: str) -> str | None:
if should_skip_key(key):
return None
for pattern, replacement in PARAM_NAME_MAP.items():
if re.match(pattern, key):
return re.sub(pattern, replacement, key)
return key
def split_monolithic(
state: dict[str, torch.Tensor],
) -> dict[str, OrderedDict[str, torch.Tensor]]:
components: dict[str, OrderedDict[str, torch.Tensor]] = {
name: OrderedDict() for name in set(COMPONENT_PREFIXES.values())
}
intentionally_skipped: list[str] = []
unowned: list[str] = []
for key, value in state.items():
if should_skip_key(key):
intentionally_skipped.append(key)
continue
for prefix, component in COMPONENT_PREFIXES.items():
if key.startswith(prefix):
mapped = apply_mapping(key[len(prefix):])
if mapped is not None:
components[component][mapped] = value
break
else:
unowned.append(key)
if unowned:
sample = ", ".join(unowned[:10])
raise ValueError(
f"Unowned monolithic keys: {len(unowned)}. "
f"Add COMPONENT_PREFIXES or SKIP_PATTERNS entries. Sample: {sample}"
)
if intentionally_skipped:
print(f"Intentionally skipped {len(intentionally_skipped)} keys")
return {name: weights for name, weights in components.items() if weights}
def load_separate_components(src_dir: Path) -> dict[str, OrderedDict[str, torch.Tensor]]:
components: dict[str, OrderedDict[str, torch.Tensor]] = {}
for component, rel_path in SEPARATE_COMPONENT_PATHS.items():
state = load_checkpoint(src_dir / rel_path)
converted: OrderedDict[str, torch.Tensor] = OrderedDict()
for key, value in state.items():
mapped = apply_mapping(key)
if mapped is not None:
converted[mapped] = value
components[component] = converted
return components
def build_component_configs(_src_dir: Path) -> dict[str, dict[str, Any]]:
# TODO: emit config content accepted by FastVideo loaders. Most components use
# config.json; schedulers use scheduler_config.json.
return {
"transformer": {"_class_name": "<FastVideoTransformerClass>"},
"vae": {"_class_name": "<FastVideoVAEClass>"},
}
def config_filename(component: str) -> str:
if component == "scheduler":
return "scheduler_config.json"
return "config.json"
def source_label(src: str) -> str:
if os.path.exists(src):
return Path(src).name
return src
def build_model_index(
src: str,
revision: str | None,
available_components: set[str],
) -> dict[str, Any]:
# TODO: match the target pipeline and every required component.
index: dict[str, Any] = {
"_class_name": "<FastVideoPipelineClass>",
"_diffusers_version": "0.30.0",
"_fastvideo_converted_from": source_label(src),
# Existing transformer/VAE loaders expect "diffusers" even when
# _class_name names a FastVideo-native class registered in FastVideo.
"transformer": ["diffusers", "<FastVideoTransformerClass>"],
"vae": ["diffusers", "<FastVideoVAEClass>"],
}
if revision:
index["_fastvideo_converted_revision"] = revision
return {
key: value
for key, value in index.items()
if key.startswith("_") or key in available_components
}
def validate_component_configs(configs: dict[str, dict[str, Any]]) -> None:
# TODO: instantiate each FastVideo config and call update_model_arch(...) or
# update_model_config(...) with this JSON so unknown emitted keys fail here.
placeholder_configs = [
name for name, config in configs.items() if "<" in json.dumps(config)
]
if placeholder_configs:
raise ValueError(f"Replace config placeholders for: {placeholder_configs}")
def verify_conversion(
dst_dir: Path,
components: dict[str, OrderedDict[str, torch.Tensor]],
) -> None:
del dst_dir, components
# TODO: load each emitted stateful component through its production loader and
# assert strict load, or document exact allowed missing/unexpected keys.
raise NotImplementedError(
"Implement production config validation and strict-load checks"
)
def write_component(
dst_dir: Path,
name: str,
state: dict[str, torch.Tensor],
config: dict[str, Any] | None,
) -> None:
component_dir = dst_dir / name
if component_dir.exists() and any(component_dir.iterdir()):
shutil.rmtree(component_dir)
component_dir.mkdir(parents=True, exist_ok=True)
save_file(
dict(state), str(component_dir / "diffusion_pytorch_model.safetensors")
)
if config is not None:
config_path = component_dir / config_filename(name)
with config_path.open("w", encoding="utf-8") as f:
json.dump(config, f, indent=2)
f.write("\n")
print(f"Wrote {name}: {len(state)} tensors")
def copy_passthrough(src_dir: Path, dst_dir: Path) -> list[str]:
copied: list[str] = []
for subfolder in PASSTHROUGH_SUBFOLDERS:
src = src_dir / subfolder
if not src.is_dir():
continue
dst = dst_dir / subfolder
if dst.exists():
shutil.rmtree(dst)
shutil.copytree(src, dst)
copied.append(subfolder)
print(f"Copied {subfolder}/")
return copied
def default_monolithic_checkpoint(src_path: Path) -> Path:
if src_path.is_file():
return src_path
return src_path / "model.safetensors"
def convert(
src: str,
dst: str,
layout: str,
revision: str | None,
) -> None:
src_path = resolve_src(src, revision)
dst_dir = Path(dst)
dst_dir.mkdir(parents=True, exist_ok=True)
model_index_path = dst_dir / "model_index.json"
if layout in {"monolithic", "raw_official"}:
# TODO: replace model.safetensors with the official monolithic file name.
components = split_monolithic(
load_checkpoint(default_monolithic_checkpoint(src_path))
)
elif layout in {"separate_components", "mixed"}:
if not src_path.is_dir():
raise ValueError(f"{layout} layout requires a source directory: {src_path}")
components = load_separate_components(src_path)
else:
raise ValueError(f"Unsupported template layout: {layout}")
copied = (
copy_passthrough(src_path, dst_dir) if src_path.is_dir() else []
)
configs = build_component_configs(src_path if src_path.is_dir() else src_path.parent)
validate_component_configs(configs)
for name, state in components.items():
write_component(dst_dir, name, state, configs.get(name))
available = set(components) | set(copied)
with model_index_path.open("w", encoding="utf-8") as f:
json.dump(build_model_index(src, revision, available), f, indent=2)
f.write("\n")
print(f"Wrote {dst_dir / 'model_index.json'}")
verify_conversion(dst_dir, components)
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--src", required=True, help="HF repo id, local dir, or checkpoint path"
)
parser.add_argument("--revision", help="HF branch, tag, or commit for repo sources")
parser.add_argument(
"--dst",
required=True,
help="Output converted_weights/<model_family> directory",
)
parser.add_argument(
"--layout",
choices=("raw_official", "monolithic", "separate_components", "mixed"),
required=True,
help="Official source layout",
)
args = parser.parse_args()
convert(args.src, args.dst, args.layout, args.revision)
if __name__ == "__main__":
main()
+214
View File
@@ -0,0 +1,214 @@
---
name: add-model-08-trace
description: Use during /add-model Phase 6 when component parity has failed and root cause requires layer-by-layer divergence analysis. Uses FastVideo activation trace first, falling back to custom hooks only for boundaries or stats the utility cannot observe.
---
# Add-Model Trace
## Manual Invocation
Load this skill when `/add-model` Phase 6 component parity has failed and the
root cause requires layer-by-layer divergence analysis. This skill is not
auto-fired. The calling subagent (DiT, VAE, encoder, or generic port skill)
loads it when its standard parity-debug loop hits a wall and cannot isolate
the divergence from end-to-end tensor comparisons alone.
Do not load this skill for first-pass parity failures. Try weight-diff and
end-to-end tensor comparison first. Load this skill only when those do not
isolate the cause.
## Goal
Find the first numerical divergence point between FastVideo's port and the
official reference, layer by layer, by instrumenting both sides at matching
tensor boundaries. The investigation must leave zero source residue in
production code when it closes.
## When To Run
After a component parity test FAILS at a bf16-noise-realistic tolerance AND
the calling subagent's first-pass debug (weight-diff, end-to-end tensor
compare) does not isolate the cause.
Required inputs before starting:
- A working FastVideo loader for the component under investigation.
- A working official loader, typically via
`tests/local_tests/helpers/<family>_upstream.py::load_upstream_<component>`.
- Shared deterministic test inputs (same tensors on both sides).
- The component parity test file path and its current failure output.
## Primary Path: FastVideo Activation Trace
Use FastVideo's first-class activation trace before writing custom hooks:
`fastvideo/hooks/activation_trace.py`, documented in
`docs/contributing/activation_trace.md`.
Pipeline runs attach trace to the transformer during pipeline initialization.
Component-only parity harnesses may call `attach_activation_trace(model)` from
local test/debug code; do not add trace calls to production model code.
Prefix the failing parity command with a tight layer regex:
```bash
FASTVIDEO_TRACE_ACTIVATIONS=1 \
FASTVIDEO_TRACE_LAYERS="^block\.layers\.[0-9]+$" \
FASTVIDEO_TRACE_STATS="abs_mean,sum,max,shape" \
FASTVIDEO_TRACE_STEPS="0" \
FASTVIDEO_TRACE_OUTPUT="/tmp/opencode/fv_trace.jsonl" \
pytest tests/local_tests -k "parity" -v -s
```
Match the layer regex to the actual `model.named_modules()` names. Empty or
broad regexes are expensive; prefer block-level names first, then narrow to
submodules after the first divergent block is known.
## Trace Compare Contract
One JSONL file per side. FastVideo output should use `FASTVIDEO_TRACE_OUTPUT`;
the upstream harness should emit the same JSONL shape:
```json
{"module":"block.layers.0","tensor":"out","step":0,"abs_mean":0.0123,"sum":1.0,"max":0.5,"shape":[1,16,32]}
```
Compare rows by `(module, step, tensor)`. The first row whose `shape`,
`abs_mean`, or `max` diverges beyond the component tolerance is the first broken
boundary. Keep `FASTVIDEO_TRACE_LAYERS`, `FASTVIDEO_TRACE_STATS`, and
`FASTVIDEO_TRACE_STEPS` identical between sides; if row order differs, sort or
normalize before diffing.
## Drill-Down Loop
**Initial run:** trace every top-level block (`^block\.layers\.[0-9]+$` or the
family's equivalent). Identify the first block index where `abs_mean` or `max`
drifts beyond tolerance while earlier blocks match.
**Drill run:** tighten `FASTVIDEO_TRACE_LAYERS` to submodules inside the first
divergent block: attention output, MLP projections, norm outputs, modality
adapters, or other named boundaries exposed by `named_modules()`.
**Iterate:** if the first divergent operation is a free function or tensor op not
visible as an `nn.Module`, use the fallback instrumentation hierarchy below.
The loop ends when the first divergent submodule or operation is identified with
a file:line citation in the official source.
## Fallback Instrumentation Hierarchy
Use these only when activation trace cannot observe the needed boundary or
statistic.
### (1) Custom forward hooks
`module.register_forward_hook(...)` and `register_forward_pre_hook(...)`.
Always within `try/finally` with `handle.remove()`. Zero source residue.
### (2) Runtime monkey-patch
`module.attr = wrapped_func` or `cls.method = wrapped_method`, restored via
`try/finally` (save original first). Use for free functions and non-Module sites
such as activation functions (`swiglu`, `apply_rotary_emb`).
### (3) Source edits in FastVideo's own code
Only when (1) and (2) are insufficient. Track all edits within a single named
`git stash` boundary OR a temporary branch. Run `git diff` before closing the
investigation to confirm cleanup.
### (4) Source edits in official repo source
Allowed only when hook and monkey-patch approaches cannot capture the site.
For git-tracked or editable official clones, use `git diff` in the clone path to
verify cleanup. For non-editable site-packages, back up the target file before
editing and restore it before handoff.
## Hypothesis Toggles
Use env-var-gated monkey-patches to A/B test suspect implementations without
source edits. Pattern: `<FAMILY>_DEBUG_PATCH_<HYPOTHESIS>=1`.
Example from the magi-human investigation:
```
MAGI_DEBUG_PATCH_LINEAR=1
```
This patched `PackedExpertLinear.forward` to mirror upstream's
`_BF16ComputeLinear` explicit-cast pattern, isolating a dtype-cast difference
as the root cause.
Document all toggles in the script docstring. Each toggle must:
- save the original before patching;
- restore the original in a `try/finally` block;
- print a `[debug] Patched <ClassName>.<method>` line to stdout when active.
## Cleanup Gate
The calling agent MUST report `[cleanup-gate] PASS` on all five items before
handoff. Do not hand off with any item unresolved.
1. `git diff` in the FastVideo repo: empty. No stray prints, hooks, or
monkey-patches in production code.
2. `git diff` in the official-repo clone (if used): empty. For non-editable
site-packages installs: `diff original.py original.py.trace-backup` is
empty OR `pip install --force-reinstall <pkg>` succeeded and the installed
file matches the original.
3. `git stash list`: only the named investigation stash (or empty). No
unnamed stashes left from this session.
4. No new untracked files outside `/tmp/opencode/` (logs) and the existing
debug script directory (`tests/local_tests/transformers/` or equivalent).
5. `mypy` clean on any production files touched during the investigation.
## Escape Hatches
Escalate to the calling bucket skill when:
- A forward hook on an official module raises because of a custom `forward`
signature or varlen handler args that the hook closure cannot satisfy. The
bucket skill has component-specific knowledge to work around this.
- The first divergent layer is `block[0]`, meaning the divergence is in the
adapter, modality dispatcher, coordinate embedding, or packing step before
any block runs. Check those sites first; the bug is not in attention or MLP.
- Per-block drift is never zero anywhere across all blocks. This usually means
the inputs are not bit-identical between sides. Verify with a state-dict
compare (weight-diff script) AND confirm the input tensors are the same
object or have identical values before the forward call.
## Handoff
Return to the calling subagent with:
- FastVideo trace JSONL path and upstream trace JSONL path.
- Trace settings used: `FASTVIDEO_TRACE_LAYERS`, `FASTVIDEO_TRACE_STATS`, and
`FASTVIDEO_TRACE_STEPS`.
- The first divergent `(module, step, tensor)` row and observed drift.
- The upstream file:line citation where the divergence originates.
- Fallback hook/patch verdict if activation trace could not observe the boundary.
- Hypothesis verdict if an A/B toggle was used, for example `PATCH_LINEAR=1`.
- Cleanup-gate status: `[cleanup-gate] PASS` or a list of unresolved items.
The calling agent uses this to scope the production fix in the FastVideo
component file.
## References
- `docs/contributing/activation_trace.md` for canonical activation-trace env vars,
JSONL output, cost model, and troubleshooting.
- `fastvideo/hooks/activation_trace.py` for the implementation and
`attach_activation_trace(model)` entry point.
- `templates/block_trace_debug.py` in this skill directory: fallback custom-hook
template when activation trace cannot observe the needed boundary or stat.
- `tests/local_tests/transformers/_debug_magi_human_block_parity.py` in the
FastVideo3 repo: historical worked example for custom hook/patch debugging.
- `add-model/SKILL.md` Phase 6: the calling context for this skill.
- `add-model-03-port-dit/SKILL.md`, `add-model-04-port-vae/SKILL.md`,
`add-model-05-port-encoder/SKILL.md`, `add-model-06-port-generic/SKILL.md`:
bucket-specific debug language and component-specific escape-hatch knowledge.
## Changelog
| Date | Change |
|---|---|
| 2026-05-01 | Initial skill extracted from `_debug_magi_human_block_parity.py` pattern. |
@@ -0,0 +1,356 @@
# SPDX-License-Identifier: Apache-2.0
"""Per-block divergence debugger template for FastVideo model ports.
Run directly (not a pytest test):
python tests/local_tests/transformers/_debug_<family>_<component>_parity.py
Generalizes: tests/local_tests/transformers/_debug_magi_human_block_parity.py
Fill FAMILY, COMPONENT, and the two loader functions. Run once for the initial
drift table, then set <FAMILY>_DEBUG_DRILL_LAYER=NN to drill into submodules.
Add <FAMILY>_DEBUG_PATCH_<HYPOTHESIS>=1 to A/B test a suspect implementation.
CLEANUP: all hooks removed in try/finally; monkey-patches restored in
try/finally; source edits tracked in a named git stash. Zero source residue.
See add-model-08-trace/SKILL.md for the full cleanup gate checklist.
"""
from __future__ import annotations
import gc
import os
import sys
from pathlib import Path
from typing import Any
import torch
FAMILY: str = "<family>" # e.g. "magi_human", "ltx2", "wan"
COMPONENT: str = "<component>" # e.g. "dit", "vae", "encoder"
DRILL_LAYER_ENV: str = "<FAMILY>_DEBUG_DRILL_LAYER"
HYPOTHESIS_ENV: str = "<FAMILY>_DEBUG_PATCH_<HYPOTHESIS>"
REL_THRESHOLD: float = 0.005 # 0.5% abs_mean drift flags a block as divergent
LOG_DIR: Path = Path("/tmp/opencode")
REPO_ROOT = Path(__file__).resolve().parents[3]
sys.path.insert(0, str(REPO_ROOT))
def load_official(device: torch.device) -> torch.nn.Module:
"""Load the official upstream model. TODO: implement for your family.
Example (magi-human):
from tests.local_tests.helpers.magi_human_upstream import install_stubs, load_upstream_dit
install_stubs()
return load_upstream_dit(base_shard_dir, device=device, dtype=None)
"""
raise NotImplementedError(f"Fill load_official() for {FAMILY}/{COMPONENT}.")
def load_fastvideo(device: torch.device) -> torch.nn.Module:
"""Load the FastVideo-native model. TODO: implement for your family.
Example (magi-human):
from fastvideo.configs.models.dits.magi_human import MagiHumanVideoConfig
from fastvideo.models.dits.magi_human import MagiHumanDiT
from safetensors.torch import load_file; import glob
fv = MagiHumanDiT(MagiHumanVideoConfig())
state = {}
for shard in sorted(glob.glob(str(transformer_dir / "*.safetensors"))): state.update(load_file(shard))
fv.load_state_dict(state, strict=False); return fv.to(device).eval()
"""
raise NotImplementedError(f"Fill load_fastvideo() for {FAMILY}/{COMPONENT}.")
def build_inputs(device: torch.device) -> dict[str, Any]:
"""Return deterministic inputs shared by both sides. TODO: replace.
Both sides must receive the SAME tensors (clone before each forward call).
Non-identical inputs cause non-zero drift everywhere.
"""
torch.manual_seed(0)
return {"x": torch.randn(64, 1024, dtype=torch.bfloat16, device=device)}
def _stat(name: str, t: torch.Tensor) -> dict:
f = t.detach().float()
return {
"name": name,
"shape": tuple(t.shape),
"abs_mean": f.abs().mean().item(),
"sum": f.sum().item(),
"min": f.min().item(),
"max": f.max().item(),
}
def _attach_block_hooks(
model: torch.nn.Module,
label: str,
log: list[dict],
tensors: dict[str, torch.Tensor] | None = None,
drill_layer: int | None = None,
) -> list[Any]:
"""Return hook handles. Caller MUST remove them in try/finally."""
handles: list[Any] = []
def _hook(name: str):
def fn(_module, _inputs, outputs):
t = outputs[0] if isinstance(outputs, tuple) else outputs
if not torch.is_tensor(t):
return
log.append({"side": label, **_stat(name, t)})
if tensors is not None:
tensors[name] = t.detach().float().cpu()
return fn
def _pre_hook(name: str):
# Pre-hooks observe a free function's output by intercepting the next
# module's input (useful when the activation is not an nn.Module).
def fn(_module, inputs):
t = inputs[0] if isinstance(inputs, tuple) else inputs
if not torch.is_tensor(t):
return
key = f"{name}<in>"
log.append({"side": label, **_stat(key, t)})
if tensors is not None:
tensors[key] = t.detach().float().cpu()
return fn
# TODO: adapt attribute paths to your model. Remove adapter block if absent.
if hasattr(model, "adapter"):
handles.append(model.adapter.register_forward_hook(_hook("adapter")))
# TODO: adapt model.block.layers to your block container.
# Alternatives: model.transformer.layers, model.blocks, model.layers
block_layers = model.block.layers # type: ignore[attr-defined]
for i, layer in enumerate(block_layers):
handles.append(layer.register_forward_hook(_hook(f"block[{i:02d}]")))
if drill_layer is not None and i == drill_layer:
tag = f"L{i:02d}"
# TODO: adapt submodule names to your layer's attributes.
# magi-human uses: attention, mlp.pre_norm, mlp.up_gate_proj,
# mlp.down_proj (pre+post), mlp, attn_post_norm, mlp_post_norm.
if hasattr(layer, "attention"):
handles.append(
layer.attention.register_forward_hook(_hook(f"{tag}.attention"))
)
if hasattr(layer, "mlp"):
mlp = layer.mlp
if hasattr(mlp, "pre_norm"):
handles.append(
mlp.pre_norm.register_forward_hook(_hook(f"{tag}.mlp.pre_norm"))
)
if hasattr(mlp, "up_gate_proj"):
handles.append(
mlp.up_gate_proj.register_forward_hook(
_hook(f"{tag}.mlp.up_gate_proj")
)
)
if hasattr(mlp, "down_proj"):
handles.append(
mlp.down_proj.register_forward_pre_hook(
_pre_hook(f"{tag}.mlp.down_proj")
)
)
handles.append(
mlp.down_proj.register_forward_hook(_hook(f"{tag}.mlp.down_proj"))
)
handles.append(mlp.register_forward_hook(_hook(f"{tag}.mlp")))
if hasattr(layer, "attn_post_norm"):
handles.append(
layer.attn_post_norm.register_forward_hook(
_hook(f"{tag}.attn_post_norm")
)
)
if hasattr(layer, "mlp_post_norm"):
handles.append(
layer.mlp_post_norm.register_forward_hook(
_hook(f"{tag}.mlp_post_norm")
)
)
return handles
def _apply_hypothesis_patch() -> bool:
"""Apply an optional monkey-patch gated by HYPOTHESIS_ENV. TODO: implement.
Pattern: save original on the class, patch, restore in _restore_hypothesis_patch().
"""
if os.getenv(HYPOTHESIS_ENV) != "1":
return False
# TODO: import FastVideo class, save original, apply patch.
print(f"[debug] Hypothesis patch {HYPOTHESIS_ENV}=1 applied.")
return True
def _restore_hypothesis_patch() -> None:
if os.getenv(HYPOTHESIS_ENV) != "1":
return
# TODO: restore original, e.g.: _mod.TargetClass.method = _mod._ORIGINAL_METHOD
def _write_log(entries: list[dict], path: Path) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
with open(path, "w") as f:
for e in entries:
f.write(
f"{e['name']} {e['shape']} "
f"{e['abs_mean']:.8f} {e['sum']:.4f} "
f"{e['min']:.6f} {e['max']:.6f}\n"
)
def _sort_key(name: str, drill_layer: int) -> tuple:
if name == "adapter":
return (0, "")
if name.startswith(f"L{drill_layer:02d}."):
sub_order = {
"attention": 0, "attn_post_norm": 1, "mlp.pre_norm": 2,
"mlp.up_gate_proj": 3, "mlp.down_proj<in>": 4,
"mlp.down_proj": 5, "mlp": 6, "mlp_post_norm": 7,
}.get(name.split(".", 1)[1], 9)
return (1, f"block[{drill_layer:02d}]", sub_order)
if name.startswith("block["):
return (1, name, 99)
return (2, name, 0)
def _print_table(by_name: dict[str, dict], drill_layer: int) -> int | None:
hdr = (
f"{'name':<18} {'up_shape':<22} {'up_absmean':>12} {'fv_absmean':>12} "
f"{'absmean_diff':>14} {'rel%':>8} {'up_sum':>14} {'fv_sum':>14} {'sum_diff':>12}"
)
print(f"\n{hdr}\n{'-' * len(hdr)}")
first_div: int | None = None
for name in sorted(by_name.keys(), key=lambda n: _sort_key(n, drill_layer)):
d = by_name[name]
up, fv = d.get("up"), d.get("fv")
if up is None or fv is None:
continue
am_diff = abs(up["abs_mean"] - fv["abs_mean"])
am_rel = am_diff / max(up["abs_mean"], 1e-9)
sum_diff = abs(up["sum"] - fv["sum"])
flag = ""
if name.startswith("block[") and am_rel > REL_THRESHOLD:
flag = " <<< DIVERGE"
if first_div is None:
first_div = int(name[len("block["):-1])
print(
f"{name:<18} {str(up['shape']):<22} {up['abs_mean']:>12.6f} "
f"{fv['abs_mean']:>12.6f} {am_diff:>14.6f} {am_rel * 100:>7.3f}% "
f"{up['sum']:>14.4f} {fv['sum']:>14.4f} {sum_diff:>12.4f}{flag}"
)
return first_div
def _print_elementwise(up_t: dict[str, torch.Tensor], fv_t: dict[str, torch.Tensor], drill_layer: int) -> None:
common = set(up_t.keys()) & set(fv_t.keys())
if not common:
return
hdr = f"{'name':<30} {'shape':<22} {'diff_max':>12} {'diff_mean':>12} {'diff_rel%':>10}"
print(f"\nElement-wise diffs for drilled L{drill_layer:02d} submodules:\n{hdr}\n{'-' * len(hdr)}")
for name in sorted(common):
a, b = up_t[name], fv_t[name]
if a.shape != b.shape:
continue
diff = (a - b).abs()
rel = (diff.mean().item() / max(a.abs().mean().item(), 1e-9)) * 100
print(
f"{name:<30} {str(tuple(a.shape)):<22} "
f"{diff.max().item():>12.6f} {diff.mean().item():>12.6f} {rel:>9.4f}%"
)
def main() -> None:
if not torch.cuda.is_available():
print("Need CUDA. Skipping.")
return
# TODO: add precondition checks (official clone present, weights available).
drill_layer = int(os.getenv(DRILL_LAYER_ENV, "0"))
device = torch.device("cuda:0")
patched = _apply_hypothesis_patch()
try:
inputs = build_inputs(device)
print("Loading official model...")
official = load_official(device)
up_log: list[dict] = []
up_t: dict[str, torch.Tensor] = {}
up_handles = _attach_block_hooks(official, "up", up_log, up_t, drill_layer)
print("Running official forward (with hooks)...")
try:
with torch.inference_mode():
# TODO: adapt forward call signature to your component.
ref_out = official(**{k: v.clone() for k, v in inputs.items()})
if isinstance(ref_out, dict):
sample = ref_out.get("sample")
ref_out = sample if sample is not None else ref_out.get("x")
elif hasattr(ref_out, "sample"):
ref_out = ref_out.sample
elif isinstance(ref_out, tuple):
ref_out = ref_out[0]
assert torch.is_tensor(ref_out), f"official output is not tensor: {type(ref_out)}"
ref_out = ref_out.detach().float().cpu()
finally:
for h in up_handles:
h.remove()
del official
gc.collect()
torch.cuda.empty_cache()
print("Loading FastVideo model...")
fv = load_fastvideo(device)
fv_log: list[dict] = []
fv_t: dict[str, torch.Tensor] = {}
fv_handles = _attach_block_hooks(fv, "fv", fv_log, fv_t, drill_layer)
print("Running FastVideo forward (with hooks)...")
try:
with torch.inference_mode():
# TODO: adapt forward call signature to your component.
fv_out = fv(**{k: v.clone() for k, v in inputs.items()})
if isinstance(fv_out, dict):
sample = fv_out.get("sample")
fv_out = sample if sample is not None else fv_out.get("x")
elif hasattr(fv_out, "sample"):
fv_out = fv_out.sample
elif isinstance(fv_out, tuple):
fv_out = fv_out[0]
assert torch.is_tensor(fv_out), f"FastVideo output is not tensor: {type(fv_out)}"
fv_out = fv_out.detach().float().cpu()
finally:
for h in fv_handles:
h.remove()
finally:
_restore_hypothesis_patch()
LOG_DIR.mkdir(parents=True, exist_ok=True)
up_path = LOG_DIR / f"{FAMILY}_{COMPONENT}_up_layers.log"
fv_path = LOG_DIR / f"{FAMILY}_{COMPONENT}_fv_layers.log"
_write_log(up_log, up_path)
_write_log(fv_log, fv_path)
print(f"\nLogs: {up_path} {fv_path}\nDiff: diff {up_path} {fv_path}")
by_name: dict[str, dict] = {}
for entry in up_log + fv_log:
by_name.setdefault(entry["name"], {})[entry["side"]] = entry
first_div = _print_table(by_name, drill_layer)
print()
if first_div is not None:
print(f"First block exceeding {REL_THRESHOLD * 100:.2f}% drift: block[{first_div:02d}]")
print(f"Re-run with {DRILL_LAYER_ENV}={first_div} to drill submodules.")
else:
print(f"No block exceeded {REL_THRESHOLD * 100:.2f}% -- divergence is amortized or pre-block.")
diff = (ref_out - fv_out).abs()
print(f"\nFinal ref_abs={ref_out.abs().mean():.6f} fv_abs={fv_out.abs().mean():.6f} "
f"diff_max={diff.max():.6f} diff_mean={diff.mean():.6f}")
_print_elementwise(up_t, fv_t, drill_layer)
if patched:
print(f"\n[debug] Hypothesis {HYPOTHESIS_ENV}=1 was active this run.")
if __name__ == "__main__":
main()
@@ -0,0 +1,255 @@
---
name: add-model-09-pipeline
description: Use during /add-model Phase 7 after all required component parity tests pass to define FastVideo pipeline wiring, configs, presets, registry entries, examples, smoke tests, and pipeline parity tests.
---
# Add Model Pipeline
## Goal
Implement and verify the end-to-end FastVideo pipeline after the native
components and converted weights have passed non-skip component parity. This
skill owns pipeline class/stage wiring, pipeline configs, presets, registry
entries, examples, smoke tests, and pipeline parity-debug.
FastVideo has one pipeline architecture: stage-based composition through
`ComposedPipelineBase`. Add or specialize stages only when existing stages cannot
represent the official behavior safely.
## Hard Gate
Do not start pipeline work until every required component, including reused
components, has a non-skip local parity PASS.
If any component row is missing, skipped, red, or blocked, return to `/add-model`
Phase 6. Pipeline parity cannot distinguish stage wiring mistakes from broken
component numerics when component parity is still unresolved.
## Inputs
Follow `../add-model/shared/common_rules.md` for token/auth safety, state files,
escape hatches, production boundaries, and skip/pass semantics.
Require a complete packet matching
`../add-model/contracts/pipeline_context.md`.
The packet must include:
- official pipeline files and official call/default sources;
- workload types, input/output modalities, and output contract;
- converted or source `model_index.json` path;
- component parity rows, all `non_skip_pass`;
- target FastVideo pipeline/config/preset/registry/example/test paths;
- `local_tests_readme` and `port_state_file` paths.
## Outputs
- Pipeline package under `fastvideo/pipelines/basic/<family>/`.
- Pipeline config under `fastvideo/configs/pipelines/<family>.py` or a documented
family-local config file when that matches existing project style.
- Presets under `fastvideo/pipelines/basic/<family>/presets.py`.
- Registry updates in `fastvideo/registry.py`.
- Basic example under `examples/inference/basic/basic_<family>*.py`.
- Local smoke and parity tests under `tests/local_tests/pipelines/`.
- Updated `tests/local_tests/<model_family>/README.md`.
- Updated `tests/local_tests/<model_family>/PORT_STATUS.md`.
- Handoff matching `../add-model/contracts/pipeline_handoff.md`.
## Mode: Pipeline Definition
Use this mode first.
1. Read the official pipeline call path before editing FastVideo code.
2. Compare official defaults against the planned FastVideo config and presets:
steps, CFG scales, secondary CFG, flow shift, schedulers, sigmas, seed/RNG,
resolution, frames, FPS, duration, VAE scaling, decode slicing, negative
prompt defaults, and output heads.
3. Create or update the pipeline class with `_required_config_modules` matching
the emitted `model_index.json` and `ComposedPipelineBase.load_modules`.
Runtime pipeline resolution is exact: `model_index.json["_class_name"]` must
match a registered `EntryClass.__name__`, or a wrapper/alias class in
`EntryClass`. Registry detectors do not select the executable pipeline class.
4. Add new public generation kwargs to `fastvideo/api/sampling_param.py` before
examples or presets use them. `SamplingParam.update()` ignores unknown keys
except for logging, and preset defaults apply only to declared fields. Add CLI
args when the option should be available from command-line entrypoints.
5. Put loader-time changes in `load_modules()` or earlier, not
`initialize_pipeline()`. `ComposedPipelineBase.__init__` loads modules before
`post_init()` calls `initialize_pipeline()`, so process-global flags, loader
path rewrites, dtype overrides, and tokenizer path changes needed for loading
cannot be introduced there.
6. Use `self.get_module("transformer_2", None)` and similar optional accessors
for truly optional modules. Do not hard-require optional modules by accident.
7. Avoid mutating class-level `_required_config_modules` in custom code. If a
pipeline needs dynamic modules, copy the list to an instance-owned value or
pass `required_config_modules` explicitly so one pipeline instance cannot leak
module requirements into another.
8. Create the stage chain in official execution order. Prefer existing shared
stages for standard text encoding, timestep preparation, latent preparation,
denoising, and decoding.
9. Add model-specific stages only for family-specific behavior that does not fit
the shared stage contracts.
10. Add pipeline config classes for wiring and runtime defaults. Do not duplicate
component architecture fields unless a loader requires them in the subconfig.
Family-local config files such as
`fastvideo/pipelines/basic/<family>/pipeline_configs.py` are valid only when
`fastvideo/registry.py` imports and registers the classes explicitly.
11. Add `InferencePreset` objects with `model_family`, `name`, `version`,
`defaults`, optional validation-only `stage_schemas`, and an `ALL_PRESETS`
tuple. `stage_schemas` validates user-facing `stage_overrides` names; it does
not drive `create_pipeline_stages()` execution.
12. Register config classes and presets in `fastvideo/registry.py`: add
`register_configs(...)`, import the family's `ALL_PRESETS`, and append it to
`_register_presets()`. Detectors should cover HF paths and `_class_name`
strings for config/preset lookup, but not as a replacement for exact pipeline
class-name resolution.
13. Add a basic example with a user-story docstring and normal file-path inputs
for image, audio, or video references. Keep orchestration glue in the
pipeline or a helper, not in the example.
14. Add a separate smoke test
`tests/local_tests/pipelines/test_<family>_pipeline_smoke.py` that proves
imports, `EntryClass`, registry, presets, config defaults, and at least one
real load/generate path when weights are local. Older local tests sometimes
colocate smoke checks in parity files; new ports should use the separate file
convention.
15. Add or update pipeline parity test scaffolding with
`templates/pipeline_parity_test.py`.
16. Update `local_tests_readme` and `port_state_file` with commands, statuses,
default sources, decisions, and blockers.
Production import boundaries are defined in
`../add-model/shared/common_rules.md`.
## Mode: Pipeline Parity Debug
Run after pipeline definition and after smoke can execute far enough to load the
pipeline. Loop until pipeline parity is a non-skip PASS or a precise blocker is
returned.
Mandatory order:
```bash
pytest tests/local_tests/pipelines/test_<family>_pipeline_smoke.py -v -s
DISABLE_SP=1 pytest tests/local_tests/pipelines/test_<family>_pipeline_parity.py -v -s
python examples/inference/basic/basic_<family>.py
```
Pipeline parity must compare real outputs, not only successful generation:
- denoised latents when decode parity is expensive or nondeterministic;
- decoded videos/images when visual output should be deterministic enough;
- decoded waveform or audio features for audio pipelines;
- separate video and audio targets for joint AV pipelines unless a validated
joint metric exists.
Debug pipeline drift in this order:
1. Confirm both sides use the same component weights and component parity PASS
results are still valid.
2. Align official and FastVideo call arguments, presets, and default values.
3. Align scheduler timesteps, sigmas/noise levels, prediction type, flow shift,
guidance math, and secondary-guidance branches.
4. Align RNG: initial latents/noise, generator device, seed, per-step noise, VAE
sampling, and any official `+1 frame` or crop/slice behavior.
5. Align conditioning: prompt templates, negative prompts, masks, image/audio
preprocessing, modality packing, text truncation, and dtype/autocast.
6. Align decode: latent scaling, per-channel mean/std, tiling flags, output
channel order, sample rate, FPS, and final slicing.
7. Add targeted stage-level diagnostics to identify the first divergent stage.
If stage diagnostics show the first bad stage is transformer/denoising or a
mid-DiT block, enable activation trace before adding ad hoc pipeline prints; see
`docs/contributing/activation_trace.md` and `../add-model-08-trace/SKILL.md`.
Keep `FASTVIDEO_TRACE_LAYERS`, `FASTVIDEO_TRACE_STATS`, and
`FASTVIDEO_TRACE_STEPS` identical across reruns so pipeline parity traces diff
one-to-one.
If the first divergence belongs to component implementation, strict loading, or
conversion mapping, stop pipeline edits and return `next_step=return_to_phase_6`
with the exact failing evidence. Do not patch conversion from this skill.
## Stage And Variant Rules
- Canonical video T2V order: `InputValidationStage`, `TextEncodingStage`,
`ConditioningStage`, `TimestepPreparationStage`, `LatentPreparationStage`,
`DenoisingStage`, `DecodingStage`.
- Canonical video I2V delta adds image loading/encoding and image VAE encoding in
the official order, commonly: `TextEncodingStage`, `ImageEncodingStage`,
`ConditioningStage`, `TimestepPreparationStage`, `LatentPreparationStage`,
`ImageVAEEncodingStage`, `DenoisingStage`, `DecodingStage`.
- Treat `ConditioningStage` as default-present for Wan-style pipelines, but still
follow the reference if another family truly skips or replaces it.
- T2V video pipelines usually use validation, text encoding, conditioning,
timestep preparation, latent preparation, denoising, and decoding.
- I2V adds image loading/encoding and image-latent preparation according to the
official pipeline, not by assuming CLIP or Wan-specific branches.
- Pick image, audio, and video encoders from the reference. Do not assume CLIP or
any other common encoder unless the reference uses it.
- Cross-attention class names are not prescribed; match the family style and
preserve the official tensor contract.
- `WorkloadType` currently has no `T2A`, `A2A`, or `AV` values. Until that enum
is extended, audio-only pipelines may register with `WorkloadType.T2V` and
preset `workload_type="t2v"` as a compatibility shim, but must document the
rationale in code and `PORT_STATUS.md`.
- Audio-only pipelines should not force real video semantics into presets. Use
minimal video-shaped placeholders such as small `height`/`width` and
`num_frames=1` only when shared `VideoGenerator`/validation paths require them,
and document that the real output is audio.
- Record modality-specific shape knobs and output contract in the pipeline
handoff: video uses `height`, `width`, `num_frames`, and `fps`; audio uses
`audio_seconds` and `sampling_rate`; joint AV records both plus whether output
is muxed or paired files.
- Use sibling pipeline classes/configs when required modules, HF repo layout,
stage chains, or inputs differ materially.
- Use one kwargs-driven pipeline class only when variants share weights,
modules, stage chain, and safe call semantics.
- Split later if components diverge, workload tags require separate discovery,
signatures become unsafe, or stage branches become substantial.
- If the DiT branches on `added_kv_proj_dim`, document the T2V/I2V split.
- If the reference uses `transformer_2`, `boundary_ratio`, `guidance_scale_2`, or
DMD step lists, keep those on config, presets, or stages deliberately.
- Support every official output head in scope. If a head is out of scope, record
explicit user approval in `PORT_STATUS.md`.
Pipeline verification order:
```bash
pytest tests/local_tests/pipelines/test_<family>_pipeline_smoke.py -v -s
DISABLE_SP=1 pytest -v -s tests/local_tests/pipelines/test_<family>_pipeline_parity.py
python examples/inference/basic/basic_<family>.py
```
Smoke tests prove loadability only. They are not a substitute for numerical
component or pipeline parity.
## Escape Hatches
Follow `../add-model/shared/common_rules.md`. Pipeline-specific ask cases include
dropping a public mode, modality, or output head; adding a new workload enum;
changing official defaults for user-facing behavior; accepting a known pipeline
parity blocker; running GPU-heavy quality work outside the agreed scope; or
publishing/uploading generated references or converted weights.
## Handoff
Return `../add-model/contracts/pipeline_handoff.md` and update the shared state
files before handoff.
Do not hand back a green pipeline if smoke or parity skipped locally. A skip is a
setup gap, not a pass.
## References
- `fastvideo/pipelines/composed_pipeline_base.py` for module loading and stage
execution.
- `fastvideo/pipelines/basic/wan/` for standard video T2V/I2V/DMD variants.
- `fastvideo/pipelines/basic/stable_audio/` for audio-specific stage composition.
- `fastvideo/configs/pipelines/stable_audio.py` and
`fastvideo/pipelines/basic/stable_audio/presets.py` for config/preset shape.
- `fastvideo/registry.py` for `register_configs(...)` and preset registration.
- `tests/local_tests/pipelines/test_gamecraft_pipeline_parity.py` for latent
parity structure.
- `tests/local_tests/pipelines/test_stable_audio_pipeline_parity.py` for audio
parity structure.
- `tests/local_tests/pipelines/test_stable_audio_pipeline_smoke.py` for no-GPU
import/registry/preset preflight shape.
@@ -0,0 +1,153 @@
# SPDX-License-Identifier: Apache-2.0
"""Pipeline parity scaffold for TODO_MODEL_FAMILY.
Copy this file to
`tests/local_tests/pipelines/test_<family>_pipeline_parity.py` and replace every
TODO before treating it as an executable scaffold.
The filled test should compare denoised latents, decoded media, audio waveform,
or another concrete output from the official pipeline against FastVideo. A
successful generation without tensor/media comparison is not parity.
"""
from __future__ import annotations
import os
import sys
from pathlib import Path
from typing import Any
import pytest
import torch
from torch.testing import assert_close
_REPO_ROOT = Path(__file__).resolve().parents[3]
_MODEL_FAMILY = "TODO_MODEL_FAMILY"
_OFFICIAL_REF_ENV = "TODO_OFFICIAL_REF_PATH"
_OFFICIAL_REF_DEFAULT = _REPO_ROOT / "TODO_OFFICIAL_REF_DIR"
_FASTVIDEO_MODEL_ENV = "TODO_FASTVIDEO_MODEL_PATH"
_FASTVIDEO_MODEL_DEFAULT = _REPO_ROOT / "converted_weights" / _MODEL_FAMILY
def _path_from_env(env_name: str, default: Path) -> Path:
return Path(os.getenv(env_name, str(default))).expanduser()
def _add_official_to_path() -> Path:
official_path = _path_from_env(_OFFICIAL_REF_ENV, _OFFICIAL_REF_DEFAULT)
if not official_path.exists():
pytest.skip(f"Official reference not found at {official_path}")
if str(official_path) not in sys.path:
sys.path.insert(0, str(official_path))
return official_path
def _log_tensor_stats(label: str, tensor: torch.Tensor) -> None:
value = tensor.detach().float()
print(
f"[{_MODEL_FAMILY} PIPELINE] {label}: shape={tuple(tensor.shape)} "
f"dtype={tensor.dtype} device={tensor.device} "
f"min={value.min().item():.6f} max={value.max().item():.6f} "
f"mean={value.mean().item():.6f} std={value.std().item():.6f}"
)
def _extract_tensor(output: Any, key: str) -> torch.Tensor:
if isinstance(output, dict):
value = output.get(key)
else:
value = getattr(output, key, None)
if value is None:
raise AssertionError(f"Pipeline output did not contain {key!r}")
if not torch.is_tensor(value):
try:
import numpy as np
value = torch.from_numpy(np.asarray(value))
except Exception as exc: # pragma: no cover - scaffold guard
raise AssertionError(f"Could not convert {key!r} to tensor") from exc
return value.detach().float().cpu()
def _run_official_pipeline(
official_path: Path,
params: dict[str, Any],
device: torch.device,
) -> Any:
del official_path, params, device
pytest.skip(
"TODO: import the official pipeline/factory, load official weights, "
"run with params, and return the comparison target."
)
def _run_fastvideo_pipeline(model_path: Path, params: dict[str, Any]) -> Any:
from fastvideo import VideoGenerator
generator = VideoGenerator.from_pretrained(
str(model_path),
num_gpus=1,
use_fsdp_inference=False,
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=False,
)
try:
return generator.generate_video(
prompt=params["prompt"],
negative_prompt=params.get("negative_prompt"),
output_path=f"outputs_{_MODEL_FAMILY}/pipeline_parity",
save_video=False,
height=params.get("height"),
width=params.get("width"),
num_frames=params.get("num_frames"),
fps=params.get("fps"),
num_inference_steps=params["num_inference_steps"],
guidance_scale=params.get("guidance_scale"),
seed=params["seed"],
)
finally:
generator.shutdown()
@pytest.mark.skipif(
not torch.cuda.is_available(),
reason="TODO_MODEL_FAMILY pipeline parity requires CUDA.",
)
def test_todo_model_family_pipeline_official_parity() -> None:
official_path = _add_official_to_path()
fastvideo_model_path = _path_from_env(
_FASTVIDEO_MODEL_ENV,
_FASTVIDEO_MODEL_DEFAULT,
)
if not fastvideo_model_path.exists():
pytest.skip(f"FastVideo model path not found at {fastvideo_model_path}")
device = torch.device("cuda:0")
params = {
"prompt": "TODO: stable parity prompt",
"negative_prompt": "",
"height": 64,
"width": 64,
"num_frames": 9,
"fps": 8,
"num_inference_steps": 4,
"guidance_scale": 1.0,
"seed": 0,
}
official_output = _run_official_pipeline(official_path, params, device)
fastvideo_output = _run_fastvideo_pipeline(fastvideo_model_path, params)
comparison_key = "TODO_COMPARISON_KEY"
official_tensor = _extract_tensor(official_output, comparison_key)
fastvideo_tensor = _extract_tensor(fastvideo_output, comparison_key)
_log_tensor_stats("official", official_tensor)
_log_tensor_stats("fastvideo", fastvideo_tensor)
assert official_tensor.shape == fastvideo_tensor.shape
diff = (official_tensor - fastvideo_tensor).abs()
print(
f"diff max={diff.max().item():.6f} "
f"mean={diff.mean().item():.6f} median={diff.median().item():.6f}"
)
assert_close(fastvideo_tensor, official_tensor, atol=1e-2, rtol=1e-2)
@@ -0,0 +1,141 @@
---
name: add-model-10-pr-review
description: Review rubric for FastVideo PRs that add or modify model families, variants, first-class components, checkpoint conversion, pipelines, parity coverage, or generated-media quality baselines. Use when reviewing a PR whose diff touches fastvideo/models/, fastvideo/pipelines/basic/, fastvideo/registry.py, scripts/checkpoint_conversion/, fastvideo/tests/ssim/, or related model-port surfaces. Pairs with review-pr-link as a project-scoped review pass; produces findings, not fixes.
---
# Add-Model PR Review
Use this skill when a reviewed PR appears to add, port, or substantially modify
a FastVideo model family, model variant, first-class model component,
checkpoint conversion, model pipeline, or local parity coverage.
This is a review skill, not an implementation workflow. Do not run `/add-model`
or start writing missing port code during review. Use the add-model skill stack
as a rubric for findings.
## Trigger Paths
Trigger this skill if `git diff --name-only <base>...HEAD` includes any of:
- `fastvideo/models/dits/`, `fastvideo/configs/models/dits/`
- `fastvideo/models/vaes/`, `fastvideo/configs/models/vaes/`
- `fastvideo/models/encoders/`, `fastvideo/configs/models/encoders/`
- `fastvideo/models/schedulers/`, `fastvideo/configs/models/schedulers/`
- `fastvideo/models/upsamplers/`, `fastvideo/configs/models/upsamplers/`
- `fastvideo/models/audio/`, `fastvideo/configs/models/audio/`
- `fastvideo/pipelines/basic/`, `fastvideo/configs/pipelines/`
- `fastvideo/registry.py`, `fastvideo/api/sampling_param.py`
- `scripts/checkpoint_conversion/`
- `examples/inference/basic/`
- `tests/local_tests/`, especially component or pipeline parity tests
- `fastvideo/tests/ssim/` or other quality-regression tests for generated media
Also trigger when the PR title/body claims a new model, model variant, VAE,
encoder, scheduler, conditioner, pipeline, conversion script, or generated-media
quality baseline even if the path list is incomplete.
## Review Inputs
Read these add-model references as review checklists:
- `../add-model/SKILL.md`: phase gates and final handoff requirements.
- `../add-model/shared/common_rules.md`: token/auth safety, production import
boundaries, state files, and skip/pass semantics.
- `../add-model/contracts/final_handoff.md`: final evidence expected from a
complete port.
- `../add-model/contracts/component_context.md` and
`../add-model/contracts/component_skill_handoff.md`: component evidence and
parity-debug expectations.
- `../add-model/contracts/conversion_request.md` and
`../add-model/contracts/conversion_handoff.md`: conversion evidence,
strict-load status, config validation, and retry context.
- `../add-model/contracts/pipeline_context.md` and
`../add-model/contracts/pipeline_handoff.md`: pipeline class/stage/config/
preset/registry/example evidence.
Then read only the satellite skill(s) that match touched areas:
- DiT/transformer changes: `../add-model-03-port-dit/SKILL.md`.
- VAE changes: `../add-model-04-port-vae/SKILL.md`.
- Encoder/conditioner changes: `../add-model-05-port-encoder/SKILL.md`.
- Scheduler/upsampler/vocoder/other components:
`../add-model-06-port-generic/SKILL.md`.
- Component parity tests: `../add-model-02-parity/SKILL.md`.
- Checkpoint conversion: `../add-model-07-conversion/SKILL.md`.
- Pipeline/config/presets/registry/examples:
`../add-model-09-pipeline/SKILL.md`.
- Prep/state docs: `../add-model-01-prep/SKILL.md`.
## Required Review Lanes
For a full model-family or model-variant PR, cover all lanes. For a
component-only PR, cover the component, conversion/parity as applicable, and the
documented downstream consumer.
1. Scope and source-of-truth lane:
Verify the PR clearly identifies the official reference, weights/revision,
supported variants, modalities, output heads, and any approved scope cuts.
2. Component lane:
Verify each required component is FastVideo-native or has a documented and
accepted lazy-wrapper exception. Check bucket/config inheritance, `EntryClass`,
state-dict surface, reused-component evidence, and output heads.
3. Conversion lane:
Verify mappings are derived from prototype key/shape dumps, source layout is
supported, skipped keys are intentional, emitted configs validate through
production paths, component strict-load status is recorded, `model_index.json`
library tokens match loaders, and revisions are pinned when converting from
HF.
4. Component parity lane:
Verify local parity tests exist for every required component, including reused
components. Scaffolds may skip in CI, but the PR must provide local non-skip
PASS evidence or an explicit accepted blocker.
5. Pipeline lane:
Verify stage order, required modules, `_class_name` / `EntryClass.__name__`
resolution, config defaults, presets, `SamplingParam` fields, registry
registration, examples, smoke tests, and pipeline parity.
6. Quality and evidence lane:
Verify media quality regression is added or explicitly deferred, examples run,
generated outputs are non-corrupt, `tests/local_tests/<family>/README.md` and
`PORT_STATUS.md` are current, and final blockers are surfaced in the review.
## Findings To Prioritize
Prioritize review findings in this order:
- Missing or skipped required component parity without accepted blocker.
- Pipeline parity/smoke/example missing or skipped for a pipeline PR.
- Conversion emits unloadable or unvalidated configs/weights.
- Wrong `model_index.json` `_class_name`, component library token, or registry
class resolution.
- Runtime diffusers/transformers model-class imports for components that own
weights or numerical behavior.
- Dropped modalities, output heads, variants, or conditioning streams without
explicit approval.
- Reused FastVideo component lacks exact definition/instantiation proof or
non-skip parity.
- Public generation kwargs/preset defaults missing from `SamplingParam`.
- Tests only check shapes, importability, or successful generation without
numerical/media comparison.
- Tokens, credentials, reference clones, staged weights, or generated bulk assets
committed to the PR.
## Output Format
Write normal code-review findings first, ordered by severity. Include file and
line references from the PR diff when possible.
Use this phrasing for missing add-model evidence:
```text
This PR does not satisfy the add-model <component|conversion|pipeline|final>
gate because <specific required evidence> is missing. The risk is <runtime load,
numerical parity, dropped output, registry resolution, etc.>.
```
Keep the summary short. Mention which lanes were reviewed and which could not be
verified because assets, GPU time, or external credentials were unavailable.
+438
View File
@@ -0,0 +1,438 @@
---
name: add-model
description: Manual /add-model workflow for implementing a FastVideo model or first-class component port after add-model-01-prep has staged reference code and weights. Organizes the port into numbered phases with conversion rules, component policies, parity gates, and handoff checks.
---
# Add Model
## Manual Invocation
This skill is for explicit `/add-model` use only. Do not auto-start it from a
casual model-port mention. The setup-only workflow is
`../add-model-01-prep/SKILL.md`.
## Goal
Port a new FastVideo model family, model variant, or first-class reusable
component so it can be loaded through FastVideo's native model, config, stage,
registry, preset, and test infrastructure.
FastVideo has one pipeline architecture: stage-based composition via
`ComposedPipelineBase`. Vary the stages and modules, not the architecture.
## Scope Shapes
Use this skill for either shape:
| Shape | Required output |
|---|---|
| Full model family or variant | Native components, conversion if needed, pipeline config/class, presets, registry, smoke test, local parity tests, example, quality regression. |
| First-class component contribution | Native component class/config, bucket export, component parity test, and a documented downstream pipeline that will consume it. Skip pipeline/preset/registry rows only when the contribution is intentionally component-only. |
If upstream ships many variants, lock scope before coding. "Base model" means
checkpoint variant, not a modality subset. If the base checkpoint produces
audio, pose, depth, masks, or other output heads, either support those outputs
or get explicit user agreement to drop them.
## Required Input
Start from an `add-model-01-prep` handoff, or equivalent fields matching
`contracts/prep_handoff.md`.
Before Phase 0, read the shared rules and all relevant schemas:
- `shared/common_rules.md`
- `contracts/prep_handoff.md`
- `contracts/port_state.md`
- `contracts/escape_hatch.md`
- `contracts/component_context.md`
- `contracts/parity_status.md`
- `contracts/conversion_request.md`
- `contracts/conversion_handoff.md`
- `contracts/component_skill_handoff.md`
- `contracts/pipeline_context.md`
- `contracts/pipeline_handoff.md`
- `contracts/final_handoff.md`
## Hard Rules
- Follow `shared/common_rules.md` for token/auth safety, state files, escape
hatches, production import boundaries, and skip/pass semantics.
- If the prep handoff is missing or ambiguous, stop and run
`../add-model-01-prep/SKILL.md`.
- If a needed component is not ported, do not ship the pipeline that needs it.
- Wan is grandfathered for missing local parity; do not copy its missing-test
precedent for new work.
## Escape Hatches
Follow `shared/common_rules.md` and `contracts/escape_hatch.md`. The main
orchestrator should ask only when no phase skill can safely continue under the
shared rules.
## Files Map
| Area | Paths |
|---|---|
| DiT | `fastvideo/models/dits/<family>.py`, `fastvideo/configs/models/dits/<family>.py`, bucket `__init__.py`. |
| VAE | `fastvideo/models/vaes/<arch_or_family>.py`, `fastvideo/configs/models/vaes/<arch_or_family>.py`, bucket `__init__.py`. Name by shared arch when reusable (`oobleck.py`, `autoencoder_kl.py`), otherwise by family (`wanvae.py`). |
| Encoder / conditioner / scheduler / upsampler | Native class/config in the matching `fastvideo/models/<bucket>/` and `fastvideo/configs/models/<bucket>/` bucket. |
| Lazy loader wrapper | Optional `fastvideo/models/<bucket>/<family>_loader.py` or similar thin `nn.Module` wrapper when a component is fetched from an external HF repo and should be hidden from host-pipeline state-dict matching. |
| Conversion | `scripts/checkpoint_conversion/<family>_to_diffusers.py` only when `needs_conversion=yes`. |
| Pipeline | `fastvideo/pipelines/basic/<family>/<family>_pipeline.py` plus sibling files for variants whose components or required modules differ. |
| Pipeline config | `fastvideo/configs/pipelines/<family>.py` or `fastvideo/pipelines/basic/<family>/pipeline_configs.py`. |
| Stages | `fastvideo/pipelines/basic/<family>/stages/` only for model-specific stage subclasses. |
| Presets / registry | `fastvideo/pipelines/basic/<family>/presets.py`, `fastvideo/registry.py`. |
| Tests | Component parity under `tests/local_tests/<bucket>/`; pipeline smoke/parity under `tests/local_tests/pipelines/`; CI-backed quality tests under `fastvideo/tests/`. |
| Example | `examples/inference/basic/basic_<family>*.py`, one per public mode/variant. |
## Phase 0: Scope And Handoff Gate
1. Validate every required handoff field.
2. Resolve `needs_conversion=unknown` before component work:
```bash
python ".agents/skills/add-model-01-prep/scripts/inspect_hf_layout.py" \
"<hf-or-local-path>" \
--json
```
3. List first-PR scope across both axes:
- Variant axis: base, distill, SR/refine, causal, DMD, I2V, V2V, etc.
- Modality axis: video, image, audio, pose, depth, masks, text, etc.
4. For component-only work, explicitly name the downstream full-pipeline PR or
planned consumer.
5. Confirm `official_env_status` is `imports_ok` or
`private_deps_need_stubs`. If it is `blocked`, return to
`../add-model-01-prep/SKILL.md` before parity scaffolding.
6. Confirm `local_tests_readme` exists and records official setup, HF weights,
dependency changes, and planned parity commands for reviewers.
7. Confirm `port_state_file` exists, follows `contracts/port_state.md`, and has
rows for open questions/issues found during prep.
8. If there are multiple official implementations, choose the one whose
architecture matches the published weights. A blessed library port can be a
better parity reference than a highly configurable research repo; document
the choice in tests.
## Phase 1: Reference And Architecture Study
Read the official pipeline call path before writing code.
Record:
- Required modules from `model_index.json` or equivalent: transformer, VAE,
text encoders, tokenizers, scheduler, image encoders, audio VAE, vocoder,
conditioners, upsamplers.
- Input/output modalities and every dedicated DiT output head.
- Text/image/audio encoding flow, latent shape, dtype, scaling, packing,
scheduler/timestep math, guidance math, VAE normalization, and decode flow.
- Whether the official code relies on private deps, custom ops, or special
kernels that parity tests must stub.
Arch config rule:
- `ArchConfig` fields must match the emitted per-component config, especially
`transformer/config.json`, one-to-one.
- Pipeline knobs do not belong on the DiT arch config: inference steps, CFG
scales, flow shift, FPS, VAE stride, text target length, data-proxy knobs,
eval defaults, and sampling defaults go on `PipelineConfig`, presets, or
stages.
- If the HF repo is raw or has empty configs, synthesize
`transformer/config.json` from the official Python model-config class, not
from data/eval config classes.
## Phase 2: Early Parity Scaffolding
Create component parity tests before or alongside implementation. Use
`../add-model-02-parity/SKILL.md` and its `templates/component_parity_test.py`.
The official reference must import in the current FastVideo environment, or the
prep handoff must identify private deps that will be stubbed locally for tests.
Use `local_tests_readme` as the reviewer-facing source for setup commands and
update its planned test table as parity scaffolds are added.
This phase is early by design:
- Official loading can be implemented from the reference study.
- FastVideo loading can target planned standardized class/config/loader paths.
- Tests may initially skip because the FastVideo class or converted weights do
not exist yet.
- The scaffold must still contain real official loading, deterministic inputs,
output extraction, and concrete tensor comparisons. No unconditional skips,
no shape-only tests.
Use subagents here: dispatch one parity-test subagent per required component,
including components that may be reused. Their output becomes the red/skip
target that porting or reuse-verification subagents make pass later.
## Phase 3: Reuse Gate And Component Dispatch
Build a component inventory before implementation:
| Field | Meaning |
|---|---|
| Component | transformer, VAE, text encoder, image encoder, scheduler, conditioner, upsampler, vocoder, etc. |
| Official definition | Repo-relative source file, class/function name, and relevant line/range if known. |
| Official instantiation | Repo-relative pipeline/config/factory call site plus constructor args and runtime flags. |
| FastVideo target | Existing class to reuse or new bucket/file/config to add. |
| Parity test | Required local test path, including reused components. |
| Status | `reuse_pending`, `reuse_proven`, `port_pending`, `non_skip_pass`, or `blocked`. |
Reuse is allowed only from the checked-out FastVideo tree. Do not wait for or
depend on an open PR adding a native class; add the native port directly in this
PR if the current tree cannot be reused.
Reuse decision:
1. Record exact official definition and instantiation evidence for every
component.
2. If an existing FastVideo class and config match both definition and
instantiation, pass that reused target to the bucket-specific skill in
`mode=prototype` and require reuse evidence plus key/shape dumps.
3. If either definition or instantiation differs, port the component directly as
FastVideo-native code through the bucket-specific skill.
4. Reused components still require non-skip component parity against the exact
official instantiation used by the target pipeline.
Porting subagent dispatch:
- Dispatch one subagent per component after Phase 2 parity scaffolds exist.
- Use `../add-model-03-port-dit/SKILL.md` for DiTs/transformers.
- Use `../add-model-04-port-vae/SKILL.md` for VAEs.
- Use `../add-model-05-port-encoder/SKILL.md` for text, image, audio, or compound
encoders/conditioners that fit the encoder config bucket.
- Use `../add-model-06-port-generic/SKILL.md` for schedulers, upsamplers,
vocoders, adapters, preprocessors, or unknown components.
- Each subagent owns one component only and must loop on that component's local
parity test until it produces a non-skip PASS or returns a precise blocker.
Every component subagent must receive a complete packet matching
`contracts/component_context.md`. If any required path is unknown, pass `unknown`
plus the exact search already performed. Do not silently omit ambiguous official
files or prototype concerns.
Bucket, layer, and attention rules live in the bucket-specific skills and
`fastvideo/layers/AGENTS.md`.
## Phase 4: Native Component Prototype
Conversion needs a FastVideo state-dict surface. Use the Phase 3
bucket-specific skill in `mode=prototype` for every required component, including
reused components.
Prototype success criteria:
- the FastVideo-native or reused class/config can import and instantiate with the
exact official architecture args;
- official and FastVideo key/shape dumps exist for every stateful component;
- `local_tests_readme` and `port_state_file` record prototype status and concerns;
- the returned handoff matches `contracts/component_skill_handoff.md`.
Do not chase numerical parity in Phase 4. Prototype mode ends when conversion has
the key/shape surface it needs, or when the component skill returns a precise
blocker or escape hatch.
## Phase 5: Param Mapping And Weight Conversion
Use `../add-model-07-conversion/SKILL.md` after Phase 4 prototypes exist.
Send a request matching `contracts/conversion_request.md`; consume the returned
`contracts/conversion_handoff.md` update before Phase 6.
Use the prep handoff's `needs_conversion` value:
- `no`: verify the source already has the component layout FastVideo loaders can
consume, then record any passthrough components.
- `yes`: write `scripts/checkpoint_conversion/<family>_to_diffusers.py` and
output `converted_weights/<family>/`.
- `unknown`: return to Phase 0.
The conversion skill owns source-layout handling, mapping derivation, config and
`model_index.json` emission, passthrough assets, strict-load verification, and
Phase 6 retry requests. Component skills must not patch conversion scripts or
converted weights ad hoc.
## Phase 6: Component Parity Debug
This is the expected expensive loop. Dispatch one subagent per required
component, including reused components, using the bucket-specific skill in
`mode=parity-debug`.
Each subagent gets:
- the complete component context packet from Phase 3/4;
- updated conversion mapping notes and strict-load result from Phase 5;
- any prototype concerns or unknowns that were not resolved before conversion.
The bucket-specific skills own parity-debug tactics. If a failure belongs to
conversion, route it through `../add-model-07-conversion/SKILL.md` with a retry
request matching `contracts/conversion_request.md`, then resume the component
skill with the updated conversion handoff.
When a component failure narrows to layer-by-layer numerical drift, load
`../add-model-08-trace/SKILL.md` before writing custom hooks. It uses
`fastvideo/hooks/activation_trace.py`; canonical env vars and JSONL format are
documented in `docs/contributing/activation_trace.md`.
Phase 6 ends only when every required component handoff reports
`parity_status=non_skip_pass`, or when a precise blocker or escape hatch is
recorded in `port_state_file`.
## Phase 7: Pipeline, Stages, And Variants
Do not start Phase 7 until every required component, reused or ported, has a
non-skip local parity PASS from Phase 6. If any component parity test is still
`scaffold_skip`, `debug_red`, `blocked`, or missing, resume Phase 6 first.
Use `../add-model-09-pipeline/SKILL.md` for pipeline definition and parity-debug.
Send a complete packet matching `contracts/pipeline_context.md`; consume the
returned `contracts/pipeline_handoff.md` before moving to quality regression or
final handoff.
The pipeline skill owns:
- pipeline class, stage chain, and optional model-specific stages;
- pipeline config, presets, registry updates, and examples;
- official args/defaults/presets comparison before setting FastVideo defaults;
- pipeline smoke and parity tests;
- continuous pipeline parity-debug until non-skip PASS or precise blocker;
- updates to `local_tests_readme` and `port_state_file`.
The pipeline handoff must explicitly cover stage order, variants, modality and
output-head handling, config/preset/registry/example status, smoke/parity tests,
and any return-to-Phase-6 evidence.
## Phase 8: PipelineConfig, Presets, Registry, Examples
This phase is implemented through `../add-model-09-pipeline/SKILL.md` after the
Phase 7 component-parity gate passes. Accept the pipeline handoff only if it
covers configs, presets, registry detection/exact class resolution, examples,
new `SamplingParam` fields for public kwargs/defaults, and local smoke/parity
status. Detailed rules live in `../add-model-09-pipeline/SKILL.md`.
## Phase 9: Parity Activation And Local Verification
Local parity is author-run, not CI-enforced. CI may only run package-level
quality tests later. Before handoff, Phase 2 scaffolds must be activated into
non-skip PASS results.
Order is mandatory:
1. Run conversion if needed.
2. Run component parity for every required component, including reused ones.
3. Run pipeline smoke.
4. Run pipeline parity.
5. Run the basic example.
If pipeline smoke or parity points back to component implementation,
strict-load, or conversion mapping, return to Phase 6 or Phase 5 rather than
patching around the issue in the pipeline.
Skip policy:
- Follow `shared/common_rules.md`: a committed local test may skip for absent
clones/weights, but a local skip is not a verified pass.
Use the commands and tolerance guidance from `../add-model-02-parity/SKILL.md` for
component checks and from `../add-model-09-pipeline/SKILL.md` for pipeline smoke,
pipeline parity, and examples. Record exact commands, status, and blockers in
`local_tests_readme` and `port_state_file`.
## Phase 10: Quality Regression
Video outputs:
- Add `fastvideo/tests/ssim/test_<family>_similarity.py` when output video
quality must be preserved.
- Seed references through `seed-ssim-references` after the test exists.
Audio outputs:
- SSIM does not apply. Use an audio-specific regression metric such as
mel-spectrogram L1, multi-resolution STFT, CLAP cosine, or a project-approved
learned metric.
- Document the metric and hardware/runtime assumptions in the test.
Joint AV outputs:
- Keep video and audio regression checks separate unless there is a validated
joint metric.
## Phase 11: Post-Parity Review And Handoff
After parity is green, run a hot-path review before handoff:
- Hoist constant tensor allocations out of sampler/denoising loops.
- Replace per-step `randn_like` churn with preallocated buffers plus
`.normal_()` when safe.
- Move `torch.backends.*` flag changes to one-shot setup/load paths.
- Delete `batch.extra` writes that nothing reads.
- Derive magic constants from configs when possible.
Pre-handoff checklist:
```text
[ ] Prep handoff is complete and committed nowhere with token values.
[ ] Conversion was run if needed and output loads with real weights.
[ ] Every required component, reused or newly ported, has a non-skip local parity PASS.
[ ] `local_tests_readme` lists every component parity test, command, status, and blocker if any.
[ ] `port_state_file` has every open question/issue either resolved or listed as an explicit blocker.
[ ] Any `next_step=ask_user` has a matching `escape_hatch` block and `E###` row.
[ ] Pipeline smoke has a non-skip local PASS.
[ ] Pipeline parity has a non-skip local PASS against the official reference.
[ ] Basic example runs and writes a non-corrupt output.
[ ] Video SSIM or audio-specific quality regression is added or explicitly deferred.
[ ] Runtime production code has no diffusers/transformers model-class imports.
[ ] Production comments are WHY-focused; examples have user-story docstrings.
[ ] Post-parity hot-path pass is complete.
```
Ask before deleting any reference clone or staged weights created by
`add-model-01-prep`. Leave `.gitignore` entries so future parity assets stay
untracked. Never commit the clone, weights, `.env`, credentials, or anything
matching `*secret*`.
## References
- `../add-model-01-prep/SKILL.md` for user-input collection, HF inspection,
weight staging, reference cloning, and setup handoff.
- `contracts/` for canonical handoff schemas used by prep, parity, conversion,
component porting, escape hatches, and final handoff.
- `../add-model-02-parity/SKILL.md` for early component parity scaffolds and
activation templates.
- `../add-model-07-conversion/SKILL.md` for Phase 5 mapping, conversion scripts,
monolithic checkpoint splitting, and strict-load checks.
- `../add-model-03-port-dit/SKILL.md`, `../add-model-04-port-vae/SKILL.md`,
`../add-model-05-port-encoder/SKILL.md`, and
`../add-model-06-port-generic/SKILL.md` for component subagent implementation
and parity-debug loops.
- `../add-model-09-pipeline/SKILL.md` for pipeline definition, config/preset/
registry/example wiring, smoke tests, and pipeline parity-debug.
- `fastvideo/layers/AGENTS.md` for native layer selection and state-dict surface
guidance.
- `docs/contributing/coding_agents.md` for narrative context.
- `docs/design/overview.md` for pipeline/config/registry architecture.
- `fastvideo/pipelines/basic/wan/` for standard T2V/I2V/DMD/Causal variants.
- `fastvideo/pipelines/basic/ltx2/` for non-standard stages and audio/video
patterns.
- `tests/local_tests/pipelines/test_gamecraft_pipeline_parity.py` for pipeline
parity shape.
- `tests/local_tests/transformers/test_ltx2.py`,
`tests/local_tests/vaes/test_ltx2_vae.py`, and
`tests/local_tests/encoders/test_ltx2_gemma_parity.py` for component parity.
- `scripts/checkpoint_conversion/convert_ltx2_weights.py` for modern conversion
script shape.
- `scripts/checkpoint_conversion/wan_to_diffusers.py` for legacy regex mapping
reference only.
## Changelog
| Date | Change |
|---|---|
| 2026-04-24 | Initial FastVideo add-model workflow. |
| 2026-04-30 | Split external setup into `add-model-01-prep`. |
| 2026-04-30 | Rewrote as manual `/add-model` phase workflow and incorporated prior review decisions. |
| 2026-04-30 | Extracted early parity scaffolding into `add-model-02-parity` and moved it before conversion/component implementation. |
| 2026-04-30 | Added component reuse proof gate, bucket-specific porting skills, and parity PASS requirement for reused components. |
| 2026-04-30 | Split prototype, conversion, and parity-debug phases; added conversion skill for monolithic and separate checkpoint layouts. |
| 2026-04-30 | Extracted handoff schemas into `contracts/` for shared use across skills. |
| 2026-04-30 | Added pipeline skill contract and Phase 7 component-parity gate. |
| 2026-04-30 | Added escape-hatch contract for user decisions and `ask_user` handoffs. |
@@ -0,0 +1,29 @@
# Add Model Contracts
Canonical handoff schemas for the `/add-model` workflow. When a skill needs to
send or receive structured context, use these files instead of inventing a local
schema.
| Contract | Use |
|---|---|
| `prep_handoff.md` | `add-model-01-prep` output and `/add-model` Phase 0 input. |
| `port_state.md` | Per-port `PORT_STATUS.md` file tracking progress, open questions, and issues. |
| `escape_hatch.md` | Shared pause-and-ask schema for user decisions the workflow cannot safely choose. |
| `component_context.md` | Per-component packet passed to parity, prototype, conversion, and parity-debug subagents. |
| `parity_status.md` | `add-model-02-parity` scaffold/activation status returned to `/add-model`. |
| `conversion_request.md` | Phase 5 conversion input and Phase 6 conversion retry request. |
| `conversion_handoff.md` | `add-model-07-conversion` output back to `/add-model` and component subagents. |
| `component_skill_handoff.md` | Component porting skill output in prototype or parity-debug mode. |
| `pipeline_context.md` | Phase 7 packet passed to `add-model-09-pipeline` after component parity is green. |
| `pipeline_handoff.md` | `add-model-09-pipeline` output back to `/add-model` after pipeline definition or parity-debug. |
| `final_handoff.md` | Final `/add-model` pre-handoff checklist summary. |
Rules:
- Do not omit required fields. Use `unknown` plus the search already performed
when the value is not known yet.
- Do not include raw token values. Use env var names only.
- Keep model-specific mapping details in conversion scripts and the local tests
README/status notes, not in generic skill docs.
- Use `next_step=ask_user` only with an `escape_hatch` block matching
`escape_hatch.md`.
@@ -0,0 +1,55 @@
# Component Context Contract
Canonical per-component packet passed from `/add-model` to parity, prototype,
conversion, and parity-debug subagents.
```text
component_context:
model_family: <snake_case>
component: <name>
component_type: <dit|vae|encoder|scheduler|conditioner|upsampler|vocoder|generic>
mode: parity-scaffold | prototype | parity-debug
official_ref_dir: <path or import path>
official_definition_files:
- path: <repo-relative or absolute path in official repo>
symbols: <class/function names>
notes: <layer graph, output contract, state-dict owner>
official_instantiation_files:
- path: <repo-relative or absolute path in official repo>
symbols: <factory/pipeline/config names>
args: <constructor args, config values, runtime flags>
official_weight_source: <checkpoint file, subfolder, prefix, or passthrough source>
fastvideo_target_files:
- fastvideo/models/<bucket>/<file>.py
- fastvideo/configs/models/<bucket>/<file>.py
local_tests_readme: tests/local_tests/<model_family>/README.md
port_state_file: tests/local_tests/<model_family>/PORT_STATUS.md
parity_test: tests/local_tests/<bucket>/test_<family>_<component>_parity.py
prototype_key_dumps:
official: converted_weights/<family>/_mapping/<component>_official_keys.json | planned | unknown
fastvideo: converted_weights/<family>/_mapping/<component>_fastvideo_keys.json | planned | unknown
conversion:
script: scripts/checkpoint_conversion/<family>_to_diffusers.py | not_created | not_needed | unknown
converted_component_dir: converted_weights/<family>/<component> | not_created | not_needed | unknown
model_index_library: <diffusers|transformers|fastvideo|fastvideo.*|unknown|none>
config_file: <config.json|scheduler_config.json|none|unknown>
mapping_notes: <key prefixes, split/fuse concerns, skipped keys, not_created, not_needed, or unknown>
production_loader_strictness: <strict|non_strict_with_allowed_keys|stateless|unknown>
strict_load: <not_run | pass | pass_with_documented_exclusions | blocked>
concerns_or_unknowns:
- <prototype mismatch, ambiguous arg, missing op, dtype concern, output head, etc.>
```
Rules:
- If any required path is unknown, pass `unknown` plus the exact search already
performed.
- Do not silently omit ambiguous official files, instantiation args, or prototype
concerns.
- For reused components, still fill every field and set `fastvideo_target_files`
to the reused class/config.
- In `mode=parity-scaffold`, prototype and conversion fields may be `planned`,
`not_created`, `not_needed`, or `unknown`; do not invent paths or statuses that
do not exist yet.
- Update `port_state_file` when concerns, issues, conversion status, or parity
status change.
@@ -0,0 +1,41 @@
# Component Skill Handoff Contract
Returned by `add-model-03-port-dit`, `add-model-04-port-vae`,
`add-model-05-port-encoder`, and `add-model-06-port-generic`.
```text
component: <name>
mode: prototype | parity-debug
files_changed: <model/config/export/test/readme paths>
official_files_used: <definition files, instantiation files>
prototype_key_dumps: <official path, fastvideo path, or none>
port_state_file: tests/local_tests/<model_family>/PORT_STATUS.md
concerns_or_unknowns: <remaining or newly discovered concerns>
parity_test: <path>
parity_status: scaffold_skip | debug_red | non_skip_pass | blocked
production_loader_strictness: strict | non_strict_with_allowed_keys | stateless
strict_load: pass | pass_with_documented_exclusions | blocked | not_run
pytest_output: <command + short result>
blocker: <none or exact missing dependency/weights/numeric mismatch>
conversion_retry_request: <none or failing keys/shapes/prefixes/evidence for add-model-07-conversion>
readme_updated: yes | no
next_step: phase_5_conversion | phase_5_conversion_retry | phase_6_continue | ask_user | blocked
escape_hatch: <none or block matching contracts/escape_hatch.md>
```
Rules:
- In `mode=prototype`, parity may be `scaffold_skip` or `blocked`; key dumps are
the required artifact. Successful prototype handoff should use
`next_step=phase_5_conversion`.
- In `mode=parity-debug`, final success requires `parity_status=non_skip_pass`.
- If conversion is implicated, return `conversion_retry_request` and do not edit
conversion scripts or converted weights directly.
- If production loading is non-strict, list allowed missing/unexpected keys in
the parity test or handoff and mark `strict_load=pass_with_documented_exclusions`.
- Update `port_state_file` before returning: component row, open questions,
issues/blockers, decisions, and handoff notes.
- Return an `escape_hatch` only for user decisions, not for normal component
implementation or parity-debug failures.
- Use `next_step=ask_user` only with an `escape_hatch` block and a matching
`PORT_STATUS.md` row.
@@ -0,0 +1,41 @@
# Conversion Handoff Contract
Returned by `../add-model-07-conversion/SKILL.md` to `/add-model` and component
parity-debug subagents.
```text
conversion_script: scripts/checkpoint_conversion/<family>_to_diffusers.py
source_layout: <diffusers|raw_official|separate_components|monolithic|mixed|custom>
converted_weights_dir: converted_weights/<model_family>
port_state_file: tests/local_tests/<model_family>/PORT_STATUS.md
components_written: <list>
passthrough_components: <list>
strict_load: pass | pass_with_documented_exclusions | blocked
component_context_updates:
- component: <name>
converted_component_dir: <path>
model_index_library: <diffusers|transformers|fastvideo|fastvideo.*>
config_file: <path or none>
config_validation: pass | blocked | not_applicable
mapping_notes: <prefixes, split/fuse ops, skipped keys>
production_loader_strictness: strict | non_strict_with_allowed_keys | stateless
strict_load: pass | pass_with_documented_exclusions | blocked | not_run
retry_resolved: <yes | no | not_a_retry>
concerns_or_unknowns: <remaining list>
blocked_on: <none or exact blocker>
next_step: phase_6_component_parity_debug | ask_user
escape_hatch: <none or block matching contracts/escape_hatch.md>
```
Rules:
- Include strict-load evidence for every stateful converted component.
- If a component intentionally loads non-strictly, list the exact missing or
unexpected keys and why they are safe.
- Include the actual `model_index.json` library token, config filename, and config
validation result for every emitted component.
- Preserve retry evidence so the requesting component subagent can resume with
updated context.
- Keep `port_state_file` synchronized with `component_context_updates`.
- Use `next_step=ask_user` only with an `escape_hatch` block and a matching
`PORT_STATUS.md` row.
@@ -0,0 +1,54 @@
# Conversion Request Contract
Consumed by `../add-model-07-conversion/SKILL.md` in Phase 5 and during Phase 6
conversion retries.
Initial conversion request:
```text
model_family: <snake_case>
source_layout: diffusers | raw_official | monolithic | separate_components | mixed | custom
official_weights: <HF repo, local dir, or checkpoint file>
hf_revision: <revision | default | none>
converted_weights_dir: converted_weights/<model_family>
local_tests_readme: tests/local_tests/<model_family>/README.md
port_state_file: tests/local_tests/<model_family>/PORT_STATUS.md
components:
- name: <transformer|vae|text_encoder|conditioner|scheduler|...>
component_type: <dit|vae|encoder|scheduler|conditioner|upsampler|vocoder|generic>
official_definition_files: <paths + symbols>
official_instantiation_files: <paths + call sites + args>
official_weight_source: <checkpoint file, prefix, subfolder, or passthrough source>
official_keys: <path to official key/shape dump>
fastvideo_keys: <path to FastVideo prototype key/shape dump>
fastvideo_class: <class name>
model_index_library: <diffusers|transformers|fastvideo|fastvideo.*>
config_filename: <config.json|scheduler_config.json|none>
production_loader_strictness: <strict|non_strict_with_allowed_keys|stateless>
source_prefix_or_path: <prefix or path>
parity_test: <component parity test path>
prototype_concerns_or_unknowns: <short list>
```
Retry request from a component skill:
```text
conversion_retry_request:
component: <name>
parity_test: <path>
failing_keys: <official and FastVideo keys, if known>
expected_actual_shapes: <expected vs actual shapes, if known>
source_prefix_or_path: <prefix/path implicated by the failure>
evidence: <strict-load error, first divergent tensor, parity log excerpt>
suspected_fix: <rename | split | fuse | skip | component bucket | config | unknown>
```
Rules:
- Phase 5 conversion requires Phase 4 official/FastVideo key dumps.
- Component skills must use the retry request instead of editing conversion
scripts or converted weights directly.
- Conversion must update `port_state_file` with conversion status, retry history,
strict-load status, new issues, and resolved issues.
- Conversion must validate emitted config keys through the production config
update path and record the config filename expected by each loader.
@@ -0,0 +1,51 @@
# Escape Hatch Contract
Canonical pause-and-ask schema for `/add-model` skills. Use this when the next
action requires user input instead of autonomous debugging.
```text
escape_hatch:
needs_user_input: yes | no
decision_type: scope | dependency | auth | cost | destructive | ambiguity | blocker
question: <one precise question>
recommended_option: <safe recommended choice>
options:
- <option + consequence>
safe_default: <what the agent will do after approval, or none>
blocked_until_answered: yes | no
state_snapshot:
phase: <phase or skill mode>
files_changed:
- <paths>
command_or_test: <last relevant command, or not_run>
evidence: <short logs, paths, error text, or blocker ID>
```
Use `needs_user_input=no` when the handoff is green or the next step is already
specified by the workflow.
Ask the user only for decisions the workflow cannot safely choose:
- product or PR scope changes, including dropping a modality, output head, or
variant;
- core dependency changes, version pin changes, or installing untrusted/private
dependencies;
- auth setup for gated repos, using env var names only and never token values;
- large downloads, publishing weights, SSIM/reference uploads, or GPU-heavy work
where cost/runtime approval is needed;
- destructive file/git operations, overwriting existing clones/weights, or
deleting staged assets;
- ambiguous official sources of truth with incompatible behavior;
- accepting a known blocker, loosening parity/quality tolerances, or shipping
without required non-skip parity.
Do not ask for normal recoverable failures:
- missing imports, missing local paths, skipped tests, failing parity, conversion
mapping errors, strict-load failures, format/lint failures, or implementation
bugs covered by the skill workflow.
Before returning `next_step=ask_user`, update
`tests/local_tests/<model_family>/PORT_STATUS.md` with the blocker/question ID,
include the exact evidence, and provide one recommended option plus at most three
alternatives.
@@ -0,0 +1,38 @@
# Final Handoff Contract
Completed by `/add-model` before handing work back to the user or opening a PR.
```text
final_handoff:
prep_handoff_complete: yes | no
conversion_status: not_needed | pass | blocked
components:
- name: <component>
reuse_or_port: reused | ported
parity_test: <path>
parity_status: non_skip_pass | blocked
concerns_or_unknowns: <none or list>
pipeline_smoke: pass | blocked | not_run
pipeline_parity: pass | blocked | not_run
example_status: pass | blocked | not_run
quality_regression: added | deferred_with_reason | not_applicable
local_tests_readme: tests/local_tests/<model_family>/README.md
port_state_file: tests/local_tests/<model_family>/PORT_STATUS.md
token_values_committed: no
runtime_third_party_model_imports: none | listed_with_rationale
blockers: <none or list>
escape_hatch: <none or block matching contracts/escape_hatch.md>
```
Required before handoff:
- Every required component, reused or ported, has non-skip local parity PASS.
- Pipeline smoke and pipeline parity are non-skip PASS, or a blocker is explicit.
- Basic example runs and writes a non-corrupt output.
- `local_tests_readme` lists every component parity command/status/blocker.
- `port_state_file` has no unresolved blocker that is omitted from the final
response or PR notes.
- No raw HF token values, credentials, `.env`, reference clone, or staged weight
blobs are committed.
- If final handoff is blocked on user input, include an `escape_hatch` block and
matching `PORT_STATUS.md` row.
@@ -0,0 +1,44 @@
# Parity Status Contract
Returned by `../add-model-02-parity/SKILL.md` to `/add-model` and later updated by
component parity-debug subagents.
```text
component_parity:
- component: <name>
test: tests/local_tests/<bucket>/test_<family>_<component>_parity.py
status: scaffold_skip | debug_red | non_skip_pass | blocked
missing: <none | fastvideo_class | converted_weights | official_import | ...>
coverage_scope: production_loader | implementation_subcomponent | both
official_definition_files: <paths>
official_instantiation_files: <paths>
concerns_or_unknowns: <short list>
pipeline_parity:
test: <path or not-created>
status: not_started | scaffold_skip | debug_red | non_skip_pass | blocked
local_tests_readme: tests/local_tests/<model_family>/README.md
port_state_file: tests/local_tests/<model_family>/PORT_STATUS.md
notes: <short list>
escape_hatch: <none or block matching contracts/escape_hatch.md>
```
Status meanings:
- `scaffold_skip`: test is present but skips for a specific missing dependency,
FastVideo class, or weights.
- `debug_red`: both sides load and the test fails numerically.
- `non_skip_pass`: required before final handoff for every required component,
including reused components.
- `blocked`: a precise missing dependency, weight, official call path, or
component/conversion regression prevents local activation.
Coverage meanings:
- `production_loader`: FastVideo side loads through the same loader/path used by
a pipeline.
- `implementation_subcomponent`: FastVideo side constructs classes or remaps
tensors directly to isolate implementation behavior.
- `both`: the test covers both production loading and implementation behavior.
Use `escape_hatch` only when blocked status requires a user decision. Normal
skips or red parity should be debugged by the workflow without asking.
@@ -0,0 +1,82 @@
# Pipeline Context Contract
Canonical packet passed from `/add-model` to `add-model-09-pipeline` for pipeline
definition and pipeline parity-debug work.
```text
pipeline_context:
model_family: <snake_case>
mode: pipeline-definition | pipeline-parity-debug
workload_types:
- <T2V|I2V|V2V|T2I|compatibility-shim-with-rationale>
modalities:
inputs: <text/image/video/audio/pose/depth/mask/etc.>
outputs: <video/image/audio/joint-av/latents/etc.>
official_ref_dir: <path or import path>
official_pipeline_files:
- path: <repo-relative or absolute path in official repo>
symbols: <pipeline/factory/sample functions>
notes: <stage order, mutable state, output contract>
official_call:
command_or_api: <official CLI, Python call, or package entrypoint>
args_and_defaults: <height, width, frames, fps, duration, steps, CFG, scheduler, seeds, etc.>
preset_source: <model card, config file, official script, or unknown>
scheduler_and_rng: <timestep/sigma/noise/generator behavior>
output_contract: <decoded media, denoised latents, waveform, dict keys, etc.>
model_index:
class_name: <FastVideo pipeline class name to emit in model_index.json>
entry_class_names: <registered EntryClass.__name__ values that must include class_name>
required_modules: <text_encoder, tokenizer, vae, transformer, scheduler, etc.>
passthrough_modules: <tokenizer, scheduler, processor, external HF dirs, or none>
sampling_param:
new_fields: <none or list of public kwargs/preset defaults to add to SamplingParam>
cli_fields: <none or list of fields that need CLI args>
placeholder_fields: <none or video-shaped compatibility placeholders with rationale>
components:
- name: <component>
component_type: <dit|vae|encoder|scheduler|conditioner|upsampler|vocoder|generic>
parity_test: tests/local_tests/<bucket>/test_<family>_<component>_parity.py
parity_status: non_skip_pass
fastvideo_target_files: <model/config/export files>
converted_component_dir: converted_weights/<family>/<component>
conversion:
converted_weights_dir: converted_weights/<family>
source_layout: <diffusers|raw_official|monolithic|separate_components|mixed|custom>
model_index_path: converted_weights/<family>/model_index.json
fastvideo_targets:
pipeline_files:
- fastvideo/pipelines/basic/<family>/<family>_pipeline.py
stage_files:
- fastvideo/pipelines/basic/<family>/stages/<stage>.py
pipeline_config_files:
- fastvideo/configs/pipelines/<family>.py
preset_file: fastvideo/pipelines/basic/<family>/presets.py
registry_file: fastvideo/registry.py
example_files:
- examples/inference/basic/basic_<family>.py
smoke_test: tests/local_tests/pipelines/test_<family>_pipeline_smoke.py
parity_test: tests/local_tests/pipelines/test_<family>_pipeline_parity.py
local_tests_readme: tests/local_tests/<model_family>/README.md
port_state_file: tests/local_tests/<model_family>/PORT_STATUS.md
concerns_or_unknowns:
- <pipeline branch, unsupported workload, output head, preset ambiguity, etc.>
```
Rules:
- Start only after every required component, reused or ported, has
`parity_status=non_skip_pass`. If any row is missing or skipped, return to
`/add-model` Phase 6.
- Record official call arguments and default sources before writing FastVideo
presets. Do not invent inference defaults from memory.
- `model_index.class_name` must match a registered pipeline `EntryClass.__name__`;
registry detectors are not sufficient for executable pipeline resolution.
- Every public generation kwarg or preset default must be represented in
`SamplingParam`, or documented as an intentional internal-only field.
- `T2A`, `A2A`, and `AV` may be used only after `WorkloadType` supports them;
otherwise record the compatibility shim and rationale explicitly.
- Keep token values out of the packet. Use only token environment variable names.
- If a target path is unknown, use `unknown` plus the exact search already
performed.
- Update `local_tests_readme` and `port_state_file` whenever pipeline smoke,
parity, presets, registry, examples, or blockers change.
@@ -0,0 +1,71 @@
# Pipeline Handoff Contract
Returned by `add-model-09-pipeline` to `/add-model` after pipeline definition or
pipeline parity-debug work.
```text
pipeline_handoff:
model_family: <snake_case>
mode: pipeline-definition | pipeline-parity-debug
files_changed:
- <pipeline/config/preset/registry/stage/example/test/readme/status paths>
official_files_used:
- <definition/call/default source paths>
required_config_modules:
emitted: <list from pipeline class>
model_index: <list from converted or source model_index.json>
status: match | mismatch | blocked
pipeline_class_resolution:
model_index_class_name: <_class_name>
entry_class_names: <registered EntryClass.__name__ values>
status: exact_match | alias_added | blocked
sampling_param:
fields_added: <none or list>
cli_fields_added: <none or list>
unknown_kwargs_checked: yes | no | blocked
stage_chain:
- <stage names in execution order>
pipeline_config:
file: <path>
classes: <class names>
official_defaults_checked: yes | no | blocked
presets:
file: <path>
names: <preset names>
status: pass | blocked | not_run
registry:
status: pass | blocked | not_run
detectors: <HF paths and model_index _class_name strings covered>
smoke_test:
path: tests/local_tests/pipelines/test_<family>_pipeline_smoke.py
status: non_skip_pass | blocked | not_run
pytest_output: <command + short result>
pipeline_parity:
path: tests/local_tests/pipelines/test_<family>_pipeline_parity.py
status: scaffold_skip | debug_red | non_skip_pass | blocked
pytest_output: <command + short result>
comparison_target: <latents|decoded video|audio|joint outputs>
example:
path: examples/inference/basic/basic_<family>.py
status: pass | blocked | not_run
output: <path or none>
readme_updated: yes | no
port_state_updated: yes | no
blockers: <none or exact blocker list>
next_step: phase_10_quality_regression | return_to_phase_6 | ask_user
escape_hatch: <none or block matching contracts/escape_hatch.md>
```
Rules:
- `pipeline-definition` may return with parity still `scaffold_skip` only if the
exact missing dependency, weight, or call-path blocker is recorded.
- Final `/add-model` handoff requires `smoke_test.status=non_skip_pass` and
`pipeline_parity.status=non_skip_pass`, unless the user explicitly accepts a
documented blocker.
- If parity failure traces to a component, conversion, or strict-load issue,
return `next_step=return_to_phase_6` and include the exact failing evidence.
- Keep `local_tests_readme` and `port_state_file` synchronized with this
handoff before returning.
- Use `next_step=ask_user` only with an `escape_hatch` block and a matching
`PORT_STATUS.md` row.
@@ -0,0 +1,89 @@
# Port State Contract
Canonical per-port state file created during prep and updated by every
`/add-model` phase.
Path:
```text
tests/local_tests/<model_family>/PORT_STATUS.md
```
Purpose:
- Single source of truth for resumable port progress.
- Tracks component status, conversion status, parity status, open questions,
blockers, escape hatches, and issue history.
- Lets review agents run the same setup/tests without reconstructing handoffs
from conversation history.
Required sections:
```text
# <Model Family> Port Status
## Summary
- model_family:
- workload_types:
- official_ref:
- official_ref_dir:
- hf_weights_path:
- local_weights_dir:
- source_layout:
- local_tests_readme:
## Current Phase
- phase:
- status: not_started | in_progress | blocked | complete
- owner: orchestrator | prep | parity | conversion | component:<name> | pipeline
- last_updated:
## Component Matrix
| Component | Type | Reuse/Port | Official Definition | Official Instantiation | FastVideo Target | Prototype | Conversion | Parity | Open Issues |
|---|---|---|---|---|---|---|---|---|---|
## Conversion State
- conversion_script:
- converted_weights_dir:
- source_layout:
- strict_load_status:
- passthrough_components:
- retry_history:
## Parity Commands
| Scope | Command | Last Result | Notes |
|---|---|---|---|
## Open Questions
| ID | Question | Owner | Needed By Phase | Status | Resolution |
|---|---|---|---|---|---|
## Issues And Blockers
| ID | Phase | Component | Severity | Issue | Evidence | Owner | Status | Resolution |
|---|---|---|---|---|---|---|---|---|
## Escape Hatches
| ID | Phase | Decision Type | Question | Recommended Option | Status | Resolution |
|---|---|---|---|---|---|---|
## Decisions
| Date | Decision | Rationale | Impact |
|---|---|---|---|
## Handoff Notes
- <short notes for the next agent>
```
Rules:
- Update this file whenever a phase starts, blocks, resolves an issue, or hands
off to another skill.
- Record open questions and issues immediately. Do not leave blockers only in
chat history or subagent responses.
- Use stable IDs: `Q001`, `Q002`, `I001`, `I002`, etc.
- Use stable escape-hatch IDs: `E001`, `E002`, etc. Link them from handoff
`escape_hatch.state_snapshot.evidence` when returning `next_step=ask_user`.
- Do not include raw token values, machine-local cache internals, or large output
dumps. Use repo-relative paths when possible.
- If a question or issue is resolved, keep the row and fill `Resolution` instead
of deleting it.
@@ -0,0 +1,42 @@
# Prep Handoff Contract
Produced by `../add-model-01-prep/SKILL.md` and consumed by `/add-model` Phase 0.
```text
model_family: <snake_case>
workload_types: <T2V/I2V/V2V/T2I/or compatibility shim with rationale>
official_ref: <url or import path>
official_ref_dir: <ReferenceDir or none>
official_ref_commit: <sha or unknown>
hf_weights_path: <HF id or local path>
hf_revision: <revision or default>
local_weights_dir: official_weights/<model_family> or <local path>
source_layout: diffusers | raw_official | monolithic | separate_components | mixed | custom | unknown
model_index_class: <_class_name or none>
components_seen: <components>
needs_conversion: yes | no | unknown
hf_token_env: <env var name only>
dependency_changes: none | installed no-deps editable | installed official deps in current env | blocked on user
official_env_status: imports_ok | private_deps_need_stubs | blocked
local_tests_readme: tests/local_tests/<model_family>/README.md
port_state_file: tests/local_tests/<model_family>/PORT_STATUS.md
gitignore_entries_added: <list>
next_step: add-model | ask_user
open_questions: <short list>
escape_hatch: <none or block matching contracts/escape_hatch.md>
```
Validation:
- `official_env_status` must be `imports_ok` or `private_deps_need_stubs` before
component parity scaffolding.
- `local_tests_readme` must exist and describe official setup, HF weights,
dependency changes, planned parity commands, and review notes.
- `port_state_file` must exist and follow `contracts/port_state.md`.
- Prep does not go directly to conversion; `/add-model` must run component
prototype/key-dump Phase 4 before Phase 5 conversion.
- `T2A`, `A2A`, and `AV` may be used only after `WorkloadType` supports them;
otherwise record the compatibility shim and rationale explicitly.
- Never include HF token values.
- Use `next_step=ask_user` only with an `escape_hatch` block and a matching
`PORT_STATUS.md` row.
@@ -0,0 +1,96 @@
# Shared Add-Model Rules
These rules apply to every `add-model` related skill: prep, parity, conversion,
component porting, pipeline, and the main `/add-model` orchestrator.
## Token And Auth Safety
- Never accept, print, echo, log, hard-code, or commit raw HF token values.
- Refer only to token environment variable names: `HF_TOKEN`,
`HUGGINGFACE_HUB_TOKEN`, or `HF_API_KEY`.
- Scripts may read those environment variables but must not print their values.
- Ask for auth setup only by env var name. Do not ask the user to paste a token.
- Read scope is needed for gated repos during conversion/load. Write scope is
needed for publishing converted weights or seeding generated references.
## Shared State Files
- `tests/local_tests/<model_family>/README.md` is the reviewer-facing setup and
verification log. Keep it current with setup commands, dependency blockers,
parity commands, conversion commands, and pass/blocker status.
- `tests/local_tests/<model_family>/PORT_STATUS.md` is the per-port state file.
It must follow `../contracts/port_state.md` and keep stable `Q###`, `I###`, and
`E###` IDs.
- Keep resolved questions/issues in `PORT_STATUS.md` with the resolution instead
of deleting them.
- Before returning a handoff, update both state files when the skill changed
setup, tests, conversion, parity status, blockers, or decisions.
- Do not include raw tokens, non-reproducible absolute cache paths, large
generated outputs, `.env`, credentials, or anything matching `*secret*`.
## Escape Hatches
Continue autonomously for recoverable setup, implementation, conversion,
strict-load, smoke, parity-debug, lint, or test failures. Stop and ask the user
only when the next action requires a product, cost, safety, auth, dependency, or
scope decision the workflow cannot safely choose.
Use `../contracts/escape_hatch.md` whenever returning `next_step=ask_user`.
Ask for user input only for:
- scope changes, such as dropping a modality, output head, variant, component, or
public mode;
- core dependency changes, untrusted/private dependency installs, or version pin
changes;
- auth setup for gated repos, using env var names only;
- large downloads, publishing weights, SSIM/reference uploads, or GPU-heavy work
where cost/runtime approval is needed;
- destructive file/git operations, overwriting existing clones/weights, or
deleting staged assets;
- incompatible official sources of truth where no reference can be chosen from
published weights and docs;
- accepting a blocker, loosening tolerances, using shape-only substitutes, or
shipping without required non-skip parity.
Do not ask for normal recoverable failures: missing imports, missing local paths,
skipped tests, failing parity, conversion mapping bugs, strict-load errors,
format/lint failures, smoke failures, registry import issues, example failures,
or implementation bugs covered by the phase workflow.
Before asking:
- update `PORT_STATUS.md` with an `E###` escape-hatch row plus any linked `Q###`
or `I###` row;
- include exact evidence: command, path, short error text, parity or strict-load
excerpt, or blocker ID;
- provide one recommended option and at most three alternatives;
- set the relevant handoff `next_step=ask_user` and include the `escape_hatch`
block.
Skill-specific escape-hatch sections may add extra examples, but they must not
weaken these shared rules.
## Production Boundary
- No runtime `from diffusers import <model class>` or
`from transformers import <model class>` in `fastvideo/` production code.
- Components that own weights or numerical behavior must be FastVideo-native
unless the user explicitly accepts a documented lazy-wrapper exception.
- Allowed third-party runtime exceptions are tokenizers and pure data utilities
when they match existing project patterns.
- Tests may import diffusers/transformers as parity references.
- Production comments explain why, not what or provenance. Avoid narrative
comments like `vendored from`, `matches upstream`, `REVIEW`, or session-history
commentary.
## Verification Semantics
- A committed local test may skip when clones, weights, or private deps are absent
so CI and other contributors are not blocked.
- On the porter's machine, a skip is not a pass. Fix the missing import, weights,
or path before claiming verification.
- New ports require local non-skip parity for required components and pipeline
parity when a pipeline is in scope.
- Smoke tests prove loadability only. They are not a substitute for numerical
component or pipeline parity.
@@ -0,0 +1,117 @@
# Component Skill Common Instructions
These instructions apply to `add-model-03-port-dit`, `add-model-04-port-vae`,
`add-model-05-port-encoder`, and `add-model-06-port-generic`. Bucket-specific skills
add target paths, implementation patterns, drift checks, and scope questions.
## Required Context
Require the complete packet from `../contracts/component_context.md`.
Do not start if the official definition files, official instantiation files, or
parity test path are missing. Ask the `/add-model` orchestrator for the complete
component context packet instead of rediscovering broad scope silently.
If the parity scaffold is missing, create it first with
`../../add-model-02-parity/templates/component_parity_test.py`.
## Prototype Mode
Prototype mode runs before conversion:
- implement or prove reuse for the minimal native component, config, export, and
`EntryClass` surface needed by the relevant loader;
- instantiate with random weights using the exact official architecture args, or
instantiate/document stateless components with no weights;
- dump official and FastVideo `state_dict()` names/shapes for every stateful
component so conversion can derive mappings from real surfaces;
- return concerns discovered during prototype work, such as ambiguous official
flags, shape mismatches, private ops, missing loader buckets, passthrough
weights, or output heads;
- update `local_tests_readme` with prototype status and key-dump paths;
- update `port_state_file` with prototype status, open questions, issues, and
handoff notes;
- do not chase numerical parity and do not block on converted weights.
Prototype mode succeeds when the component imports, instantiates with official
args, and required key/shape dumps exist. Converted weights and parity PASS are
not required yet.
## Parity-Debug Mode
Parity-debug mode runs after conversion:
- strict-load converted weights through the same path the pipeline will use, or
document that the component is stateless or an approved passthrough;
- use `conversion_context` and `concerns_or_unknowns` to decide whether a failure
belongs to mapping, loading, implementation, tokenization, normalization,
scheduler semantics, or the parity test;
- run only the component parity test first with `pytest <parity_test> -v -s`;
- if it skips, fix the missing official import, FastVideo class, tokenizer,
converted weights, or path;
- if it fails numerically, add targeted intermediate comparisons to identify the
first divergent operation or tensor;
- update component implementation only when the failure is a component
layer/config/forward/contract bug;
- update `local_tests_readme` with the command, result, and blocker or PASS;
- update `port_state_file` with parity status, resolved/new issues, open
questions, and handoff notes;
- keep iterating until the test is a non-skip PASS or return a precise blocker.
## Conversion Boundary
Component skills must not patch conversion scripts or converted weights ad hoc.
If the first drift or strict-load failure points to wrong keys, missing tensors,
shape mismatches, component prefixes, split/fuse logic, skipped-key policy, or
config emission, return a conversion retry request for
`../../add-model-07-conversion/SKILL.md` matching
`../contracts/conversion_request.md`.
Resume parity-debug only after conversion returns an updated handoff.
## Reuse Proof
When `fastvideo_target_files` point to existing FastVideo code instead of a new
port:
- compare the official definition against the FastVideo target: graph/operation
structure, parameter or state shapes, normalization, activation, positional or
temporal behavior, scaling constants, dtype behavior, state-dict names, output
containers, and output tensors;
- compare the official instantiation against the FastVideo config and loader
args: constructor args, config values, defaults, variant flags, optional
submodules, checkpoint metadata, tokenizer/media paths, and loader path;
- treat a matching class instantiated with different args as not reusable;
- record reuse evidence in `local_tests_readme` and keep the reused component in
`prototype_key_dumps` when it owns state so conversion and parity-debug use the
same surface;
- still run parity-debug to a non-skip PASS. If mismatch is found, return the
concern so `/add-model` can switch the component to a native port.
## Handoff
Return `../contracts/component_skill_handoff.md`.
Mode-specific expectations:
- In `mode=prototype`, `parity_status` may be `scaffold_skip` or `blocked`; key
dumps are the required artifact and successful prototype handoff should use
`next_step=phase_5_conversion`.
- In `mode=parity-debug`, final success requires
`parity_status=non_skip_pass`.
- If conversion is implicated, return `conversion_retry_request` and leave
conversion edits to `add-model-07-conversion`.
- If production loading is non-strict, list allowed missing/unexpected keys in
the parity test or handoff and mark
`strict_load=pass_with_documented_exclusions`.
## Escape Hatches
Follow `common_rules.md`. Do not ask for normal prototype or parity-debug
failures such as missing imports, tokenizer/path issues, red parity,
strict-load failures, key mismatches, shape mismatches, or implementation bugs.
Return conversion retry requests or precise blockers as directed by the workflow.
Ask only when component work requires a scope or safety decision, such as
dropping a required stream/output/path, changing core dependencies, accepting
private model code or unsupported private ops, choosing between incompatible
official definitions, creating a new loader bucket, or loosening required parity.
@@ -0,0 +1,338 @@
---
name: decompose-pipeline-pr
description: Decompose an oversized FastVideo pipeline PR into a stack of independently-reviewable PRs. Tiers the diff by blast radius (invisible / dead code / cross-cutting infra / activation), produces a branch graph and worktree bootstrap, drafts the AGENTS.md manifest, flags missing tests on cross-cutting infra changes, and extracts lessons from the PR body.
---
# Decompose Pipeline PR
## Purpose
When a PR adds a new pipeline (or first-class component port) and crosses
~3,000 LOC, single-shot review converges to rubber-stamping. This skill
decomposes such a PR into a stack of independently-reviewable PRs without
disturbing `main`.
It is the inverse of `add-model`: where `add-model` walks adding a new
pipeline as a fresh PR, this skill walks decomposing an existing oversized
pipeline PR.
**Worked example:** PR #1280 (daVinci-MagiHuman, 9,812 LOC, 56 files) →
2 prerequisite PRs off main + 8-PR stack:
- #1293 `will/activation-trace` (prerequisite)
- #1294 `will/loader-infra` (prerequisite)
- #1295 (1/8) housekeeping
- #1296 (2/8) t5gemma encoder
- #1297 (3/8) DiT
- #1298 (4/8) pipeline stages
- #1299 (5/8) pipeline orchestrator
- #1300 (6/8) provenance (AGENTS.md, JOURNAL.md, lessons)
- #1301 (7/8) conversion scripts
- #1302 (8/8) registry activation
## Prerequisites
- Open PR number on `hao-ai-lab/FastVideo` (or any FastVideo fork)
- `gh` CLI authenticated against the target remote
- Local git worktree support (`git worktree`)
- Git config `user.name` / `user.email` set
- Pre-commit installed (`pre-commit install --hook-type pre-commit --hook-type commit-msg`)
- The target PR's branch fetched locally as `origin/<feature-branch>`
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| PR number or URL | Yes | E.g. `1280` or `https://github.com/hao-ai-lab/FastVideo/pull/1280` |
| Max desired PR size | No | Defaults to ~2,500 LOC of code per stack PR (excluding generated/journal files) |
| Output directory | No | Defaults to `.agents/tmp/decompose-<pr-number>/` (gitignored) |
## Steps
### 1. Verify ground truth (do not trust `gh pr diff --name-only`)
`gh pr diff <N> --name-only` has been observed to emit phantom file entries.
Always cross-check against the authoritative `git diff`:
```bash
mkdir -p .agents/tmp/decompose-<N>
git fetch origin pull/<N>/head:<feature-branch>
git diff origin/main..origin/<feature-branch> --name-status \
> .agents/tmp/decompose-<N>/files.txt
git diff origin/main..origin/<feature-branch> --stat
```
Use the `--name-status` output as the authoritative file list. If it
disagrees with `gh pr diff --name-only`, trust the git diff.
### 2. Tier the diff by blast radius
Classify every changed file into one of four tiers:
| Tier | Description | Examples |
|---|---|---|
| **Tier 0 — Invisible** | Lint/style/CI configs that don't affect runtime | `.gitignore`, `pyproject.toml` (codespell only), agent documentation |
| **Tier 1 — Dead code** | New files in their own dirs; aggregator one-liners | `fastvideo/models/dits/<new>/`, `fastvideo/pipelines/basic/<new>/`, `examples/inference/basic/basic_<new>*.py`, `tests/local_tests/<new>/`, `__init__.py` exports |
| **Tier 2 — Cross-cutting infra** | Modifications to files used by every pipeline | See protected-paths list below |
| **Tier 3 — Activation switch** | `register_configs(...)` calls + the example scripts that demo them | `fastvideo/registry.py` |
**FastVideo Tier 2 protected paths:**
```
fastvideo/utils.py
fastvideo/pipelines/composed_pipeline_base.py
fastvideo/models/loader/component_loader.py
fastvideo/configs/models/dits/__init__.py
fastvideo/configs/models/encoders/__init__.py
fastvideo/configs/models/vaes/__init__.py
fastvideo/envs.py
fastvideo/fastvideo_args.py
fastvideo/distributed/**
fastvideo/layers/**
fastvideo/attention/**
fastvideo/registry.py # treat as Tier 3 if change is the activation
```
Tier 3 detection (mechanical):
```bash
git diff origin/main..origin/<feature-branch> -- fastvideo/registry.py | \
grep -E "^\+.*register_configs\("
```
If `registry.py` only contains `register_configs` additions, treat it as
Tier 3. If it modifies existing behavior, treat it as Tier 2 (rare).
### 3. Identify reusable Tier-1 components
Within Tier 1, look for sub-trees that are **not** model-specific and could
land separately:
- Encoders matching a known multi-model base (T5/T5-Gemma/Llama/Gemma/CLIP variants)
- New stage classes that subclass shared bases without referencing the new model
- Hook/profiler/debug infra under `fastvideo/hooks/`
- New helpers that have no model-specific dependencies
These get split into their own PRs (e.g. PR 4 `t5gemma-encoder` in the
MagiHuman example).
### 4. Hunt for missing test coverage on Tier 2 changes
For every Tier-2 file modified, check whether the original PR added unit
tests for the new behavior:
```bash
for f in <list-of-tier-2-files>; do
echo "=== Tests for $f ==="
git diff origin/main..origin/<feature-branch> -- \
"$(echo $f | sed 's|fastvideo/|fastvideo/tests/|; s|\.py|*|')"
done
```
If a Tier-2 PR has no accompanying tests, **emit a "must-add tests" list**
with a sketch of the case grid. Tier-2 PRs do not ship without those tests.
The MagiHuman example required this for PR-B (`utils.py`): the original PR
shipped no `test_utils_loader.py`, so the decomposition added 9 unit-test
cases covering the umbrella-detector boundary, the optional-component-dirs
relaxation, and regression coverage on every existing 2-segment HF id.
### 5. Build the dependency DAG and topo-sort
Edges:
- Tier 2 infra → Tier 1 code that imports it
- Reusable Tier 1 components → model-specific Tier 1 code that uses them
(encoder before DiT before pipeline)
- Tier 1 → Tier 3 (activation always last)
- Tier 0 has no dependents (lands first as a freebie)
Topo-sort produces the stack ordering. Pull Tier-2 PRs **out of the stack**
when they have no model-specific dependency — they should land off main
with their own focused review, not buried in a model port.
Render as a tree (markdown):
```
main
├─ <prereq-A>
│ └─ <prereq-B>
│ ├─ <stack-01-housekeeping>
│ │ └─ <stack-02-encoder>
│ │ └─ <stack-03-dit>
│ │ └─ ...
│ │ └─ <stack-N-activate>
│ └─ (parallel) <skill-pr> off main
```
### 6. Detect mis-shelved docs and debug scratch
Two categories to flag:
- **Mis-shelved docs**: Markdown files under `tests/local_tests/` are
journals, not tests. Flag for relocation to the package dir as
`JOURNAL.md`.
- **Debug scratch**: files starting with `_debug_`, `_scratch_`, or
`_explore_`. Flag for drop (do not carry into any output PR).
For MagiHuman: `tests/local_tests/magi-human.md` → relocate. Two
`_debug_magi_human_*.py` files → drop.
### 7. Author the AGENTS.md manifest skeleton
For the new pipeline package, generate a 6-section `AGENTS.md` scaffold
with the file table pre-populated from the diff:
1. **Manifest** — file table by role
2. **Parity invariants** — load-bearing rules with one-paragraph each + lesson refs
3. **Cross-refs** — "If you change X, re-run Y" matrix
4. **Run book** — single pytest command + prereqs (HF tokens, GPU, wall-time)
5. **Open questions** — known issues (e.g. tolerance carve-outs)
6. **Provenance** — PR table with branch names and source SHA
The provenance section is filled incrementally during stack execution and
finalized in the activation PR.
### 8. Extract lessons from the PR body
Scan the PR body for sections titled "Key implementation work", "Bug hunt",
"Lessons", or sentences with patterns like "took N waves to localize",
"silent regression", "investigation revealed". Each becomes a candidate
`.agents/lessons/<YYYY-MM-DD>_<slug>.md` draft.
Lessons MUST follow the existing template in
`.agents/lessons/README.md`:
- YAML frontmatter: `date`, `experiment`, `category`, `severity`
- Sections: What Happened, Root Cause, Fix / Workaround, Prevention
- Filename: `<YYYY-MM-DD>_<short-slug>.md`
Lessons co-locate with the code they concern: a conversion-script lesson
lands in the same PR as the conversion script, not in the docs PR.
### 9. Emit the commit-footer convention
Every commit in the stack ends with:
```
<Feature>-Stack: N/M
```
E.g. `Magi-Stack: 5/8`. Use the package directory name as the feature key.
After all PRs squash-merge, `git log --grep='^<Feature>-Stack:'` reconstructs
the lineage even if PR numbers later get renumbered.
### 10. Produce the worktree bootstrap
Generate a runnable bash script:
```bash
#!/bin/bash
set -euo pipefail
REPO=/home/<user>/FastVideo
WORKTREE=/home/<user>/FastVideoMagi # NB: directory name must be a valid
# Python identifier (no hyphens) so
# mypy doesn't choke
SOURCE_PR=<N>
SOURCE_BRANCH=will/<feature>
SOURCE_SHA=$(git -C "$REPO" rev-parse "origin/$SOURCE_BRANCH")
git -C "$REPO" fetch origin main:main
git -C "$REPO" fetch "origin/$SOURCE_BRANCH"
git -C "$REPO" worktree add "$WORKTREE" origin/main
# Capture baseline for provenance. Everything under .agents/tmp is transient
# and ignored by git.
OUTPUT_DIR="$REPO/.agents/tmp/decompose-$SOURCE_PR"
mkdir -p "$OUTPUT_DIR"
cat > "$OUTPUT_DIR/<feature>-baseline-${SOURCE_SHA:0:8}.txt" <<EOF
Source PR: <repo>#$SOURCE_PR
Source SHA: $SOURCE_SHA
Authoritative file count: $(git -C "$REPO" diff origin/main..origin/$SOURCE_BRANCH --name-only | wc -l)
Date captured: $(date -u +%Y-%m-%dT%H:%M:%SZ)
EOF
```
### 11. Author preserve via `git checkout`, not `cherry-pick`
For each stack PR:
```bash
git -C "$WORKTREE" switch -c <new-branch> <base-branch>
git -C "$WORKTREE" checkout origin/<source-branch> -- <file1> <file2> ...
git -C "$WORKTREE" commit -m "[<scope>]: <subject>
<body>
<Feature>-Stack: N/M"
git -C "$WORKTREE" push -u origin <new-branch>
gh pr create --base <base-branch> --head <new-branch> --title "..." --body "$(cat <<EOF ... EOF)"
```
Notes:
- `git checkout origin/<source> -- <files>` extracts only the named files,
preserving the diff. The original PR's author is **not** preserved on the
new commit (it's authored by whoever runs the script). Reference the
original PR + source SHA in every commit body and PR description for
authorship attribution.
- **Never use `git cherry-pick`** for this workflow — cherry-pick applies
whole commits, which mixes concerns across PR boundaries.
## Outputs
The skill produces all transient planning artifacts under
`.agents/tmp/decompose-<pr>/`:
1. A markdown decomposition plan (`plan.md`)
2. A proposed branch graph
3. A worktree-bootstrap script (`bootstrap.sh`)
4. Per-PR file allocation lists (under `stack/`)
5. AGENTS.md scaffolds for any new pipeline packages
6. Draft lesson files (placed alongside the PR that owns the code they concern)
7. A finalized provenance table for the package AGENTS.md
## Anti-Patterns
The skill should warn against:
- **"Just rebase the megaPR into smaller commits."** Doesn't help review;
reviewer still sees one PR.
- **Co-locating tests under the new package.** FastVideo's convention is
by-kind under `fastvideo/tests/` and `tests/local_tests/<family>/`. Don't
invent a new layout per pipeline.
- **Splitting Tier 2 changes into "one file per PR."** Tier 2 PRs are
about semantic units (e.g., "loader umbrella + optional component dirs"
together because they jointly define the new diffusers-format contract),
not file-count.
- **Landing the activation switch first** ("just register, the code can
be empty"). The skill enforces activation-last so every intermediate
state is dead code, not broken code.
- **Trusting `gh pr diff --name-only`.** Cross-check against
`git diff origin/main..origin/<feature-branch> --name-status` —
`gh`'s output has been observed to include phantom entries.
- **Worktree dir names with hyphens.** mypy interprets them as invalid
Python package names and refuses to run. Use CamelCase or underscores.
- **Skipping the lesson-extraction step.** PR bodies contain the most
expensive learnings of the original implementation. Losing them to a
squash-merge is the silent decay of institutional knowledge.
## Example Usage
```
User: split PR 1280
Agent: [invokes decompose-pipeline-pr]
→ produces .agents/tmp/decompose-1280/plan.md with:
- tiered file table (56 files: 3 tier-0, 35 tier-1, 9 tier-2,
9 tier-3)
- branch graph (PR-A + PR-B + 8-PR stack)
- worktree bootstrap script
- per-PR file lists
- AGENTS.md scaffold for fastvideo/pipelines/basic/magi_human/
- 3 draft lessons extracted from the PR body
→ asks user to confirm before opening branches
```
## References
- The MagiHuman decomposition (worked example):
`fastvideo/pipelines/basic/magi_human/AGENTS.md` (after PR #1302 merges)
- Existing skill: `.agents/skills/add-model/SKILL.md` (the inverse — adding
a new pipeline as a fresh PR)
- Lesson template: `.agents/lessons/README.md`
- Skill template: `.agents/skills/SKILL_TEMPLATE.md`
+196
View File
@@ -0,0 +1,196 @@
---
name: dreamverse-deploy
description: Use when redeploying the migrated Dreamverse app backend and frontend on a chosen local GPU; tears down existing ports, launches services, and waits for readiness checks.
---
# dreamverse-deploy — redeploy migrated Dreamverse on a chosen GPU
**Scope:** project (lives in this repo at `.agents/skills/dreamverse-deploy/`)
**When to use:** you want to (re)launch the migrated `apps/dreamverse/` backend
and frontend on this dev node, pinned to a specific physical GPU. Tears down
any existing deploy on the same ports first, then boots fresh and waits for
both `/readyz` and the FE root to return 200.
## Prerequisites
- Working tree containing `apps/dreamverse/`
- `dreamverse-server` installed from this checkout; if missing, run
`uv pip install -e ".[dreamverse]"`
- Local conda env at `~/miniconda3/envs/fv-main/` with `flashinfer-python`,
`cerebras-cloud-sdk`, `openai` installed (override the default path with
`DREAMVERSE_PYTHON=/path/to/python`)
- `~/.env` exporting `CEREBRAS_API_KEY`, `GROQ_API_KEY`, etc.
- npm available in `$PATH` (or set `NPM=/path/to/npm`)
- `gcc-13` + `g++-13` at `/usr/bin/` (workaround for nvcc gcc-15 rejection)
- **Recommended:** native ffmpeg at `$HOME/opt/ffmpeg-native/bin/ffmpeg`, built
via `bash apps/dreamverse/scripts/install_native_ffmpeg.sh`. The deploy
detects that binary directly and exports it for the backend. The installer's
generated `apps/dreamverse/scripts/ffmpeg-env.sh` is for manual launches.
When the binary is missing, the deploy falls back to system ffmpeg with a
warning. Set
`DREAMVERSE_REQUIRE_NATIVE_FFMPEG=true` to make the missing binary a hard
failure.
If any required prereq is missing, the script fails fast with a clear message.
## Usage
```bash
# Deploy on GPU 4 with the current web port. The legacy helper default remains
# 5274, so pass 5299 explicitly. Torch compile and warmup are both off.
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh 4 8009 5299
# Deploy on GPU 6 with custom ports
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh 6 8089 5275
# Deploy on GPU 0 with warmup enabled
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh --warmup 0 8009 5299
# Deploy with torch.compile enabled (max-autotune; first segment ~3-4min,
# subsequent segments save ~3s — only worth it for benchmarking)
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh --torch-compile 4 8009 5299
# Deploy with both warmup AND torch.compile enabled
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh --warmup --torch-compile 4 8009 5299
# Flags can appear before, between, or after positional args
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh 4 8089 5275 --warmup
```
### Arguments
| Position | Name | Default | Notes |
|---|---|---|---|
| 1 | `GPU` | (required) | Physical GPU index, e.g. `4` |
| 2 | `BACKEND_PORT` | `8009` | TCP port for the FastAPI server |
| 3 | `FRONTEND_PORT` | `5274` | TCP port for the Next.js dev server |
### Flags
| Flag | Default | Notes |
|---|---|---|
| `--warmup` / `--no-warmup` | off | Run GPU warmup at boot (~minutes). Overrides `DREAMVERSE_WARMUP` |
| `--torch-compile` / `--no-torch-compile` | off | Enable max-autotune `torch.compile`. First segment ~3-4min when on, ~45s when off. Overrides `DREAMVERSE_TORCH_COMPILE` |
| `--nvenc` / `--no-nvenc` | off | Use `h264_nvenc` hardware encoder instead of `libx264` software. Eliminates ~1100ms/segment of CPU encoding cost (raises realtime ratio from ~0.78x → ≥1.0x, eliminating inter-segment buffer-drain stutter). Requires native ffmpeg built with `--enable-nvenc` (the install script's default since the NVENC update). Hard-fails up-front if the binary is missing or lacks NVENC. Overrides `DREAMVERSE_NVENC` |
| `-h` / `--help` | — | Show usage |
Flags can appear in any position relative to the positional args. Explicit flag values always win over env-var defaults.
### Environment variables (used when no flag is given)
| Var | Default | Purpose |
|---|---|---|
| `DREAMVERSE_WARMUP` | `false` | Same as `--warmup`/`--no-warmup`. Flag takes precedence |
| `DREAMVERSE_TORCH_COMPILE` | `false` | Same as `--torch-compile`/`--no-torch-compile`. Flag takes precedence |
| `DREAMVERSE_NVENC` | `false` | Same as `--nvenc`/`--no-nvenc`. Flag takes precedence |
| `DREAMVERSE_PYTHON` | `~/miniconda3/envs/fv-main/bin/python` | Conda environment used for the flashinfer prerequisite probe; `dreamverse-server` itself is resolved from `PATH` |
| `DREAMVERSE_REPO_ROOT` | git rev-parse | Repo root override |
| `DREAMVERSE_LOG_DIR` | `/tmp/opencode/dreamverse-deploy` | Directory for the per-GPU backend and per-port frontend logs |
| `DREAMVERSE_REQUIRE_NATIVE_FFMPEG` | `false` | If `true`, fail when `$HOME/opt/ffmpeg-native/bin/ffmpeg` is absent |
## What it does
1. Validates prereqs.
2. Kills any process on the target backend/frontend ports + waits for the
target GPU to release memory (allows up to 30s for cleanup).
3. Sources `~/.env`.
4. Exports the env recipe required for boot:
- `CUDA_VISIBLE_DEVICES=<gpu>`
- `FASTVIDEO_ENABLE_DEVTOOLS=1`
- `FASTVIDEO_ENABLE_STARTUP_WARMUP=<DREAMVERSE_WARMUP>`
- `FASTVIDEO_GPU_COUNT=1`
- `ENABLE_TORCH_COMPILE=<0|1 derived from DREAMVERSE_TORCH_COMPILE>`
- `CC=/usr/bin/gcc-13 CXX=/usr/bin/g++-13 CUDAHOSTCXX=/usr/bin/g++-13`
- `NVCC_PREPEND_FLAGS="-ccbin /usr/bin/gcc-13 -allow-unsupported-compiler"`
- `FASTVIDEO_FFMPEG_BIN=$HOME/opt/ffmpeg-native/bin/ffmpeg` +
`FASTVIDEO_VIDEO_CODEC=<libx264|h264_nvenc>` (when the native binary exists)
5. Launches the installed `dreamverse-server` console command in a detached
`setsid` session and captures its PID.
6. Polls `/readyz` until 200. The budget is 5 minutes by default, 8 minutes
with one startup optimization enabled, and 15 minutes with both warmup and
`torch.compile` enabled.
7. Launches the devtools frontend through npm in a detached session and
captures its PID.
8. Polls FE `/` until 200 (max 60s).
9. Prints URLs, PIDs, and log paths.
## What it does NOT do
- Does not modify `~/.env` or the FastVideo `.venv`.
- Does not push code or commit anything.
- Does not run Playwright. Use the e2e wrapper separately:
```bash
cd apps/dreamverse/web
PLAYWRIGHT_SKIP_WEBSERVER=1 BACKEND_HOST=127.0.0.1 BACKEND_PORT=8009 \
PLAYWRIGHT_BASE_URL=http://127.0.0.1:5299 \
NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
npm exec -- playwright test
```
The standard suite runs by default; the long-running two-segment
audio-continuation spec is gated behind
`PLAYWRIGHT_LONG_RUNNING=1` (see below).
## Long-running e2e (paired with `--warmup --torch-compile`)
[`apps/dreamverse/web/e2e/long-running-segments.spec.ts`](../../../apps/dreamverse/web/e2e/long-running-segments.spec.ts)
drives a real two-segment session through the FE, captures every WS
frame, and asserts segments 1 AND 2 both reach `media_segment_complete`
with at least one binary fMP4 chunk per segment. It guards against the
BrokenPipe regression previously caused by dropped LTX-2 audio continuation
kwargs.
Skipped by default. Enable with:
```bash
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh \
--warmup --torch-compile 4 8009 5299
cd apps/dreamverse/web
PLAYWRIGHT_SKIP_WEBSERVER=1 \
BACKEND_HOST=127.0.0.1 \
BACKEND_PORT=8009 \
PLAYWRIGHT_BASE_URL=http://127.0.0.1:5299 \
NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
PLAYWRIGHT_LONG_RUNNING=1 \
npm exec -- playwright test e2e/long-running-segments.spec.ts
```
Expected runtime: ~7-9 minutes on a B200 (torch.compile max-autotune
warm-up dominates the cold start; per-test timeout is 900s). The spec
hard-fails on any WS `error`/`step_error` frame so the BrokenPipe
regression surfaces with the actual ffmpeg/audio diagnostics rather
than an opaque "test timed out".
## Teardown
Stop both services without redeploying:
```bash
# Stop services on default ports (port-pattern based)
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh --stop
# Stop AND nuke any process holding GPU N
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh --stop 4
```
The redeploy path (`<GPU>` mode) automatically nukes any process holding the
target GPU before launching — including orphan `multiproc_executor` worker
subprocesses left over from a parent backend that was killed without grace.
This was the failure mode of an earlier naive port-only kill: parent dies,
children survive, GPU stays full, next deploy OOMs.
## Notes
- The installed `dreamverse-server` console command enters
`apps/dreamverse/dreamverse/server_entry.py`, which loads the current
Dreamverse runtime from `apps/dreamverse/dreamverse/`.
- The B200 / sm_100a NVCC flags are mandatory on this dev node because the
conda toolchain ships gcc-15, which nvcc rejects. The script requires the
configured gcc-13 and g++-13 binaries during preflight.
## Deployment boundary
This skill is for a local checkout on a directly attached GPU. For a container
image, use `apps/dreamverse/docker/README.md`. For Modal, follow
`apps/dreamverse/scripts/modal/README.md`; do not adapt this process-killing
workflow to a remote deployment.
@@ -0,0 +1,460 @@
#!/usr/bin/env bash
# See ../SKILL.md for full usage.
set -euo pipefail
is_pid_alive() {
kill -0 "$1" 2>/dev/null
}
terminate_pid() {
local pid="$1"
local label="${2:-pid=${pid}}"
[[ -n "${pid}" ]] && [[ "${pid}" != "$$" ]] || return 0
is_pid_alive "${pid}" || return 0
kill "${pid}" 2>/dev/null || true
for _ in $(seq 1 10); do
is_pid_alive "${pid}" || return 0
sleep 0.5
done
if is_pid_alive "${pid}"; then
kill -9 "${pid}" 2>/dev/null && echo " force-killed ${label}" || true
fi
}
terminate_pattern() {
local pattern="$1"
local pid
if ! command -v pgrep >/dev/null 2>&1; then
pkill -TERM -f "${pattern}" 2>/dev/null || true
sleep 2
pkill -KILL -f "${pattern}" 2>/dev/null || true
return 0
fi
for pid in $(pgrep -f -- "${pattern}" 2>/dev/null || true); do
terminate_pid "${pid}" "pattern='${pattern}' pid=${pid}"
done
}
list_port_pids() {
local port="$1"
if command -v lsof >/dev/null 2>&1; then
lsof -t -iTCP:"${port}" -sTCP:LISTEN 2>/dev/null || true
return 0
fi
ss -tlnp 2>/dev/null | awk -v port=":${port}" '
$0 ~ port {
while (match($0, /pid=[0-9]+/)) {
print substr($0, RSTART + 4, RLENGTH - 4)
$0 = substr($0, RSTART + RLENGTH)
}
}
' || true
}
if [[ "${1:-}" == "--stop" ]]; then
for pat in 'apps/dreamverse/dreamverse/main.py' 'dreamverse-server --host 0.0.0.0 --port' 'next dev --port' 'next-server (v'; do
terminate_pattern "${pat}"
done
if [[ -n "${2:-}" ]] && [[ "${2}" =~ ^[0-9]+$ ]]; then
gpu_uuid="$(nvidia-smi --query-gpu=index,uuid --format=csv,noheader 2>/dev/null | awk -F', ' -v g="${2}" '$1==g {print $2}')"
if [[ -n "${gpu_uuid}" ]]; then
for pid in $(nvidia-smi --query-compute-apps=pid,gpu_uuid --format=csv,noheader 2>/dev/null \
| awk -F', ' -v u="${gpu_uuid}" '$2==u {print $1}'); do
terminate_pid "${pid}" "GPU${2} pid=${pid}"
done
fi
fi
sleep 2
echo "stopped: ports may take a few seconds to free"
exit 0
fi
# ---------------------------------------------------------------------------
# Args
# ---------------------------------------------------------------------------
usage() {
cat <<USAGE
Usage: $(basename "$0") [FLAGS] <GPU> [BACKEND_PORT] [FRONTEND_PORT]
$(basename "$0") --stop [GPU]
Positional:
GPU Physical GPU index (required), e.g. 4
BACKEND_PORT default 8009
FRONTEND_PORT default 5274
Flags (override env vars when both set):
--warmup / --no-warmup run GPU warmup at boot (default off)
--torch-compile / --no-torch-compile
enable max-autotune torch.compile
(default off — first segment ~3-4min
when on, ~45s when off)
--nvenc / --no-nvenc use h264_nvenc hardware encoder (default
off — uses libx264 software encoder).
Requires native ffmpeg built with NVENC.
-h, --help show this help
Env overrides:
DREAMVERSE_WARMUP 'true'|'false' (default false)
DREAMVERSE_TORCH_COMPILE 'true'|'false' (default false)
DREAMVERSE_NVENC 'true'|'false' (default false)
DREAMVERSE_REPO_ROOT default: \$(git rev-parse --show-toplevel)
DREAMVERSE_LOG_DIR default: /tmp/opencode/dreamverse-deploy
DREAMVERSE_REQUIRE_NATIVE_FFMPEG 'true'|'false' (default false)
USAGE
}
WARMUP_OVERRIDE=""
TORCH_COMPILE_OVERRIDE=""
NVENC_OVERRIDE=""
POSITIONAL=()
while [[ $# -gt 0 ]]; do
case "$1" in
-h|--help) usage; exit 0 ;;
--warmup) WARMUP_OVERRIDE=true; shift ;;
--no-warmup) WARMUP_OVERRIDE=false; shift ;;
--torch-compile) TORCH_COMPILE_OVERRIDE=true; shift ;;
--no-torch-compile) TORCH_COMPILE_OVERRIDE=false; shift ;;
--nvenc) NVENC_OVERRIDE=true; shift ;;
--no-nvenc) NVENC_OVERRIDE=false; shift ;;
--) shift; while [[ $# -gt 0 ]]; do POSITIONAL+=("$1"); shift; done ;;
-*) echo "error: unknown flag '$1'" >&2; usage >&2; exit 2 ;;
*) POSITIONAL+=("$1"); shift ;;
esac
done
set -- "${POSITIONAL[@]+"${POSITIONAL[@]}"}"
if [[ $# -lt 1 ]]; then
usage >&2
exit 2
fi
GPU="${1}"
BACKEND_PORT="${2:-8009}"
FRONTEND_PORT="${3:-5274}"
if ! [[ "${GPU}" =~ ^[0-9]+$ ]]; then
echo "error: GPU must be a non-negative integer (got '${GPU}')" >&2
exit 2
fi
WARMUP="${WARMUP_OVERRIDE:-${DREAMVERSE_WARMUP:-false}}"
case "${WARMUP}" in
true|false) ;;
*) echo "error: warmup must be 'true' or 'false' (got '${WARMUP}')" >&2; exit 2 ;;
esac
TORCH_COMPILE="${TORCH_COMPILE_OVERRIDE:-${DREAMVERSE_TORCH_COMPILE:-false}}"
case "${TORCH_COMPILE}" in
true|false) ;;
*) echo "error: torch-compile must be 'true' or 'false' (got '${TORCH_COMPILE}')" >&2; exit 2 ;;
esac
TORCH_COMPILE_FLAG=$([[ "${TORCH_COMPILE}" == "true" ]] && echo 1 || echo 0)
NVENC="${NVENC_OVERRIDE:-${DREAMVERSE_NVENC:-false}}"
case "${NVENC}" in
true|false) ;;
*) echo "error: nvenc must be 'true' or 'false' (got '${NVENC}')" >&2; exit 2 ;;
esac
REPO_ROOT="${DREAMVERSE_REPO_ROOT:-$(git rev-parse --show-toplevel 2>/dev/null || pwd)}"
LOG_DIR="${DREAMVERSE_LOG_DIR:-/tmp/opencode/dreamverse-deploy}"
# ---------------------------------------------------------------------------
# Prereq checks
# ---------------------------------------------------------------------------
bail() { echo "error: $*" >&2; exit 3; }
[[ -d "${REPO_ROOT}/apps/dreamverse" ]] \
|| bail "REPO_ROOT '${REPO_ROOT}' does not contain apps/dreamverse/. Are you on a migration branch?"
DREAMVERSE_SERVER="$(command -v dreamverse-server 2>/dev/null || true)"
[[ -n "${DREAMVERSE_SERVER}" ]] && [[ -x "${DREAMVERSE_SERVER}" ]] \
|| bail "dreamverse-server not executable or not in PATH (run: uv pip install -e \".[dreamverse]\")"
CONDA_ENV_PYTHON="${DREAMVERSE_PYTHON:-${HOME}/miniconda3/envs/fv-main/bin/python}"
[[ -x "${CONDA_ENV_PYTHON}" ]] \
|| bail "conda env python missing at ${CONDA_ENV_PYTHON} (set DREAMVERSE_PYTHON to override)"
"${CONDA_ENV_PYTHON}" -c 'import flashinfer' 2>/dev/null \
|| bail "flashinfer-python not installed in ${CONDA_ENV_PYTHON} (run: ${CONDA_ENV_PYTHON} -m pip install flashinfer-python --no-build-isolation)"
NPM="${NPM:-npm}"
NPM_REQUESTED="${NPM}"
NPM="$(command -v "${NPM}" 2>/dev/null || true)"
[[ -n "${NPM}" ]] && [[ -x "${NPM}" ]] || bail "npm not executable or not in PATH: ${NPM_REQUESTED} (set NPM to override)"
GCC13="$(command -v "${GCC13:-gcc-13}" 2>/dev/null || true)"
GPP13="$(command -v "${GPP13:-g++-13}" 2>/dev/null || true)"
[[ -n "${GCC13}" ]] && command -v "${GCC13}" >/dev/null 2>&1 \
|| bail "gcc-13 not found or not executable (needed for nvcc workaround). Set GCC13 or install gcc-13 in PATH"
[[ -n "${GPP13}" ]] && command -v "${GPP13}" >/dev/null 2>&1 \
|| bail "g++-13 not found or not executable (needed for nvcc workaround). Set GPP13 or install g++-13 in PATH"
[[ -f "${HOME}/.env" ]] || echo "warn: ${HOME}/.env missing — provider API keys may be unset" >&2
NATIVE_FFMPEG_BIN="${HOME}/opt/ffmpeg-native/bin/ffmpeg"
if [[ "${NVENC}" == "true" ]]; then
NATIVE_VIDEO_CODEC=h264_nvenc
else
NATIVE_VIDEO_CODEC=libx264
fi
REQUIRE_NATIVE_FFMPEG="${DREAMVERSE_REQUIRE_NATIVE_FFMPEG:-false}"
case "${REQUIRE_NATIVE_FFMPEG}" in
true|false) ;;
*) bail "DREAMVERSE_REQUIRE_NATIVE_FFMPEG must be 'true' or 'false' (got '${REQUIRE_NATIVE_FFMPEG}')" ;;
esac
if [[ -x "${NATIVE_FFMPEG_BIN}" ]]; then
if [[ "${NVENC}" == "true" ]]; then
encoder_list="$("${NATIVE_FFMPEG_BIN}" -hide_banner -encoders 2>/dev/null || true)"
if [[ "${encoder_list}" != *h264_nvenc* ]]; then
bail "--nvenc requested but ${NATIVE_FFMPEG_BIN} was not built with NVENC. Rebuild: bash apps/dreamverse/scripts/install_native_ffmpeg.sh (with ENABLE_NVENC=1, the default)"
fi
if ! "${NATIVE_FFMPEG_BIN}" -hide_banner -loglevel error -y \
-f lavfi -i 'color=red:size=64x64:rate=24:duration=0.2' \
-c:v h264_nvenc -f null - >/dev/null 2>&1; then
bail "--nvenc requested but the GPU on this host has no NVENC silicon (probe failed: 'OpenEncodeSessionEx unsupported device'). Datacenter Blackwell (B200) and some H100 SKUs ship without NVENC; --nvenc only works on hosts with NVENC-capable GPUs (RTX 50-series, T4, A10, etc.)."
fi
fi
echo " native ffmpeg: ${NATIVE_FFMPEG_BIN} (codec=${NATIVE_VIDEO_CODEC})"
elif [[ "${REQUIRE_NATIVE_FFMPEG}" == "true" ]] || [[ "${NVENC}" == "true" ]]; then
bail "${NATIVE_FFMPEG_BIN} missing (required by --nvenc or DREAMVERSE_REQUIRE_NATIVE_FFMPEG=true). Run: bash apps/dreamverse/scripts/install_native_ffmpeg.sh"
else
echo "warn: ${NATIVE_FFMPEG_BIN} missing — backend will fall back to system ffmpeg (\$(command -v ffmpeg))." >&2
echo " Build native ffmpeg with: bash apps/dreamverse/scripts/install_native_ffmpeg.sh" >&2
fi
echo " python: ${CONDA_ENV_PYTHON}"
mkdir -p "${LOG_DIR}"
# ---------------------------------------------------------------------------
# Teardown anything on target ports
# ---------------------------------------------------------------------------
echo "[1/8] killing any existing deploy on ports ${BACKEND_PORT}/${FRONTEND_PORT} and GPU ${GPU}..."
kill_port_pid() {
local port="$1"
local pid
for pid in $(list_port_pids "${port}"); do
terminate_pid "${pid}" "port=${port} pid=${pid}"
done
}
for pat in "dreamverse-server --host 0.0.0.0 --port ${BACKEND_PORT}" "next dev --port ${FRONTEND_PORT}" "NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 next dev --port ${FRONTEND_PORT}"; do
terminate_pattern "${pat}"
done
kill_port_pid "${BACKEND_PORT}"
kill_port_pid "${FRONTEND_PORT}"
gpu_uuid="$(nvidia-smi --query-gpu=index,uuid --format=csv,noheader 2>/dev/null | awk -F', ' -v g="${GPU}" '$1==g {print $2}')"
if [[ -n "${gpu_uuid}" ]]; then
for pid in $(nvidia-smi --query-compute-apps=pid,gpu_uuid --format=csv,noheader 2>/dev/null \
| awk -F', ' -v u="${gpu_uuid}" '$2==u {print $1}'); do
if [[ -n "${pid}" ]] && [[ "${pid}" != "$$" ]]; then
cmd="$(ps -p "${pid}" -o comm= 2>/dev/null || true)"
terminate_pid "${pid}" "GPU${GPU} pid=${pid} (${cmd:-?})"
fi
done
fi
for i in $(seq 1 30); do
free_be=true
free_fe=true
ss -tln 2>/dev/null | grep -qE ":${BACKEND_PORT}\b" && free_be=false
ss -tln 2>/dev/null | grep -qE ":${FRONTEND_PORT}\b" && free_fe=false
gpu_mem="$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits 2>/dev/null | sed -n "$((GPU + 1))p" || echo 99999)"
if "${free_be}" && "${free_fe}" && [[ "${gpu_mem}" -lt 1000 ]]; then
break
fi
sleep 1
done
gpu_mem="$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits 2>/dev/null | sed -n "$((GPU + 1))p" || echo 0)"
echo " ports cleared; GPU${GPU} at ${gpu_mem} MiB"
# ---------------------------------------------------------------------------
# Launch backend
# ---------------------------------------------------------------------------
echo "[2/8] launching backend on GPU ${GPU} port ${BACKEND_PORT} (warmup=${WARMUP} torch_compile=${TORCH_COMPILE} nvenc=${NVENC})..."
backend_log="${LOG_DIR}/backend-gpu${GPU}.log"
: > "${backend_log}"
setsid bash -c "
set -a
if [[ -f \"${HOME}/.env\" ]]; then
source \"${HOME}/.env\"
fi
set +a
if [[ -x \"${NATIVE_FFMPEG_BIN}\" ]]; then
export FASTVIDEO_FFMPEG_BIN=\"${NATIVE_FFMPEG_BIN}\"
export FASTVIDEO_VIDEO_CODEC=\"${NATIVE_VIDEO_CODEC}\"
fi
export DREAMVERSE_PYTHON=\"${CONDA_ENV_PYTHON}\"
export CUDA_VISIBLE_DEVICES=${GPU}
export FASTVIDEO_ENABLE_DEVTOOLS=1
export FASTVIDEO_ENABLE_STARTUP_WARMUP=${WARMUP}
export FASTVIDEO_GPU_COUNT=1
export ENABLE_TORCH_COMPILE=${TORCH_COMPILE_FLAG}
export CC=${GCC13}
export CXX=${GPP13}
export CUDAHOSTCXX=${GPP13}
export NVCC_PREPEND_FLAGS=\"-ccbin ${GCC13} -allow-unsupported-compiler\"
cd \"${REPO_ROOT}\"
exec \"${DREAMVERSE_SERVER}\" --host 0.0.0.0 --port ${BACKEND_PORT}
" > "${backend_log}" 2>&1 < /dev/null &
disown
# Wait briefly, then resolve actual python PID (the inner process, not the
# wrapper bash).
sleep 4
backend_pid="$(pgrep -f "dreamverse-server --host 0.0.0.0 --port ${BACKEND_PORT}" | head -1 || true)"
if [[ -z "${backend_pid}" ]]; then
echo "error: backend failed to spawn. Last 30 lines of log:" >&2
tail -30 "${backend_log}" >&2
exit 4
fi
echo " backend pid=${backend_pid} log=${backend_log}"
# Poll /readyz. Deadline scales with warmup + torch.compile flags
# because warmup runs two synthetic segments before /readyz=200, and
# torch.compile max-autotune adds ~3-4min cold start to the first
# segment. Empirical worst case (warmup=true, torch_compile=true):
# ~7 min on B200; we budget 15 min for safety.
if [[ "${WARMUP}" == "true" ]] && [[ "${TORCH_COMPILE}" == "true" ]]; then
READYZ_BUDGET_SECONDS=900
elif [[ "${WARMUP}" == "true" ]] || [[ "${TORCH_COMPILE}" == "true" ]]; then
READYZ_BUDGET_SECONDS=480
else
READYZ_BUDGET_SECONDS=300
fi
READYZ_POLL_INTERVAL=6
READYZ_MAX_ITERS=$(( READYZ_BUDGET_SECONDS / READYZ_POLL_INTERVAL ))
echo "[3/8] polling http://127.0.0.1:${BACKEND_PORT}/readyz (budget=${READYZ_BUDGET_SECONDS}s) ..."
ready=0
for i in $(seq 1 ${READYZ_MAX_ITERS}); do
code="$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 "http://127.0.0.1:${BACKEND_PORT}/readyz" 2>/dev/null || echo 000)"
if [[ "${code}" == "200" ]]; then
ready=1
break
fi
if ! kill -0 "${backend_pid}" 2>/dev/null; then
echo "error: backend pid ${backend_pid} died. Last 50 lines:" >&2
tail -50 "${backend_log}" >&2
exit 5
fi
sleep ${READYZ_POLL_INTERVAL}
done
if [[ "${ready}" != "1" ]]; then
echo "error: backend did not become /readyz=200 within ${READYZ_BUDGET_SECONDS}s. Last 50 lines:" >&2
tail -50 "${backend_log}" >&2
exit 5
fi
echo "[4/8] backend /readyz OK"
# ---------------------------------------------------------------------------
# Launch frontend
# ---------------------------------------------------------------------------
echo "[5/8] launching frontend on port ${FRONTEND_PORT}..."
frontend_log="${LOG_DIR}/frontend-port${FRONTEND_PORT}.log"
: > "${frontend_log}"
# Resolve dev script: dev:devtools forces port 5274 + devtools env. If the
# requested port differs, run `next dev --port` directly with devtools env.
fe_cmd="run dev:devtools"
if [[ "${FRONTEND_PORT}" != "5274" ]]; then
fe_cmd="exec -- next dev --port ${FRONTEND_PORT}"
fi
setsid bash -c "
cd \"${REPO_ROOT}/apps/dreamverse/web\"
export NEXT_PUBLIC_INCLUDE_DEVTOOLS=1
export BACKEND_URL=http://127.0.0.1:${BACKEND_PORT}
export BACKEND_HOST=127.0.0.1
export BACKEND_PORT=${BACKEND_PORT}
exec '${NPM}' ${fe_cmd}
" > "${frontend_log}" 2>&1 < /dev/null &
disown
sleep 4
frontend_pid="$(pgrep -f "next dev --port ${FRONTEND_PORT}" | head -1 || true)"
if [[ -z "${frontend_pid}" ]]; then
echo "error: frontend failed to spawn. Last 30 lines:" >&2
tail -30 "${frontend_log}" >&2
exit 6
fi
echo " frontend pid=${frontend_pid} log=${frontend_log}"
# Poll FE root
echo "[6/8] polling http://127.0.0.1:${FRONTEND_PORT}/ ..."
fe_ready=0
for i in $(seq 1 30); do
code="$(curl -s -o /dev/null -w '%{http_code}' --max-time 2 "http://127.0.0.1:${FRONTEND_PORT}/" 2>/dev/null || echo 000)"
if [[ "${code}" == "200" ]]; then
fe_ready=1
break
fi
if ! kill -0 "${frontend_pid}" 2>/dev/null; then
echo "error: frontend pid ${frontend_pid} died. Last 30 lines:" >&2
tail -30 "${frontend_log}" >&2
exit 7
fi
sleep 2
done
if [[ "${fe_ready}" != "1" ]]; then
echo "error: frontend did not respond 200 within 60s. Last 30 lines:" >&2
tail -30 "${frontend_log}" >&2
exit 7
fi
echo "[7/8] frontend / OK"
# ---------------------------------------------------------------------------
# Print summary
# ---------------------------------------------------------------------------
cwd="$(readlink "/proc/${backend_pid}/cwd" 2>/dev/null || echo unknown)"
gpu_mem_now="$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits 2>/dev/null | sed -n "$((GPU + 1))p" || echo 0)"
ffmpeg_in_use="$(tr '\0' '\n' < "/proc/${backend_pid}/environ" 2>/dev/null | sed -n 's/^FASTVIDEO_FFMPEG_BIN=//p' | head -1)"
[[ -z "${ffmpeg_in_use}" ]] && ffmpeg_in_use="$(command -v ffmpeg 2>/dev/null || echo '<not found>') (system fallback)"
cat <<SUMMARY
[8/8] redeploy OK
Frontend : http://localhost:${FRONTEND_PORT} (PID ${frontend_pid})
Backend : http://localhost:${BACKEND_PORT} (PID ${backend_pid})
cwd=${cwd}
gpu=${GPU} mem=${gpu_mem_now} MiB
ffmpeg=${ffmpeg_in_use}
Logs : ${backend_log}
${frontend_log}
Stop : ./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh --stop
E2E : cd apps/dreamverse/web && \\
PLAYWRIGHT_SKIP_WEBSERVER=1 \\
BACKEND_URL=http://127.0.0.1:${BACKEND_PORT} \\
PLAYWRIGHT_BASE_URL=http://127.0.0.1:${FRONTEND_PORT} \\
NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \\
npm exec -- playwright test
SUMMARY
@@ -0,0 +1,476 @@
---
name: reseed-performance-baseline
description: Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, or environment-caused benchmark shift using one or more reviewed normalized performance JSONs. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline in FastVideo/performance-tracking must be advanced from a consistent batch of reviewed source results. The workflow backs up existing history under /tmp, validates all source JSONs for the same (model_id, gpu_type), rejects internally inconsistent source batches, uploads one success=true reseed record per accepted source JSON, and offers to clean local temp state after a successful upload.
---
# Re-seed Performance Baseline
## Purpose
Replace or advance the rolling performance baseline for a single
`(model_id, gpu_type)` pair in the HF dataset
`FastVideo/performance-tracking`.
Performance comparison uses the median of up to the last 5 successful records
for the same model and GPU. Failed records are useful audit history, but they
do not move the future baseline because `compare_baseline.py` loads records
with `successful_only=True`.
This skill now reseeds from a reviewed batch of one or more source performance
JSONs. It uploads one new `success=true` record per accepted source JSON; it
does not blindly replicate one measurement into 3 or 5 records. The effective
reseed size is therefore dynamic and equals the number of provided, validated,
internally consistent source JSONs.
If the operator provides fewer than 3 records, call out that the last-5 rolling
median may not move immediately. If the operator provides 3 consistent shifted
records, the rolling median usually moves immediately. If the operator provides
5 consistent shifted records, the last-5 window is effectively reset to the new
runtime profile.
These records are intentional operator-approved baseline resets, not ordinary
independent main-branch persistence. Mark them clearly with provenance fields
so the HF history remains auditable.
Use this skill when a performance test fails for an intentional and reviewed
reason, such as a torch/runtime/container upgrade that legitimately increases
peak memory or changes timings. This is the performance equivalent of
`reseed-ssim-references`: backup first, scope tightly, require explicit human
approval, then upload reviewed accepted baseline records.
## When to use
- A PR or main run failed the rolling performance comparison by more than the
allowed regression threshold, and maintainers agree the shift is caused by
an intentional runtime, dependency, hardware image, or benchmark environment
change rather than a FastVideo logic regression.
- One or more shifted source result JSONs have been reviewed and accepted, and
the operator wants to use those exact reviewed results to advance the rolling
baseline.
- The source batch is internally consistent: no provided source JSON regresses
against the batch median by more than the configured tolerance.
## When not to use
- The benchmark failure might be a real code regression. Fix or investigate
the code path first.
- The fixed benchmark thresholds in
`.buildkite/performance-benchmarks/tests/*.json` are too low. Those are a
separate gate from the rolling HF baseline and may need a code review change.
- There is no clear source run, commit, and rationale. Baseline history is a
production signal; do not edit it without provenance.
- The provided source JSONs disagree materially with each other. Rerun or
investigate instead of uploading a noisy reseed batch.
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| `model_id` | Yes | Benchmark id, e.g. `wan-t2v-1.3b-2gpu`. This maps to the HF subdirectory after `sanitize(model_id)`. |
| `gpu_type` | Yes | Exact GPU device string from the performance record, e.g. the L40S device name emitted by CI. Baselines are GPU-specific. |
| `source_results` | Yes | One or more local paths or Buildkite artifact URLs for accepted shifted performance JSONs. Prefer normalized `normalized_perf_*.json` artifacts emitted by `compare_baseline.py`. Accept `source_result` as an alias only for a single JSON. |
| `max_intra_batch_regression` | No | Maximum allowed regression of any source JSON against the source batch median. Default: `PERF_MAX_REGRESSION` if set, otherwise `0.05` (5%). |
| `intent_rationale` | Yes | One-line explanation for why the baseline shift is legitimate. This is written into provenance and should be reused in the PR. |
Hardcoded defaults:
- HF repo: `FastVideo/performance-tracking` (`HF_REPO_ID` override is
supported by the code, but use the default unless the user explicitly asks).
- Local sync root: `/tmp/perf-tracking` (`PERFORMANCE_TRACKING_ROOT` override
is supported).
- Backup root: `/tmp/performance_reseed_backup`.
- Download scratch root for source artifact URLs: `/tmp/performance_reseed_source`.
- Baseline window: last 5 `success=true` records for the same
`(model_id, gpu_type)`.
- Reseed count: dynamic. Upload exactly one accepted seed record per validated
source JSON.
## Steps
### 1. Validate the target and source results
Normalize `source_results` to a list. If the user passes a single
`source_result`, treat it as a one-element `source_results` list and report
that a single record may not move the last-5 median immediately.
If any source result is a Buildkite artifact URL, download it first into a
local scratch directory under `/tmp/performance_reseed_source/` and use the
downloaded JSON path for the rest of the workflow. If the agent cannot access
the artifact because Buildkite authentication is missing, ask the user to
download the artifact manually and provide the local path.
Prefer the normalized Buildkite artifact emitted by `compare_baseline.py`:
```text
perf_reports/results/normalized_perf_*.json
```
That file is already in the HF tracking schema. Load each normalized JSON
directly:
```python
import json
with open(source_result, encoding="utf-8") as f:
record = json.load(f)
```
Stop if any normalized record's `model_id` or `gpu_type` does not match the
requested `model_id` and `gpu_type`.
The source records may have `success: false` when they came from failed
rolling baseline comparisons. That is expected; only the reviewed reseed
records become new `success: true` baseline records after explicit approval.
Sort validated source records by their original `timestamp` ascending before
preparing the seed records. If a source timestamp is missing or unparsable,
preserve input order for those records and print a warning. This makes the
fresh reseed timestamps deterministic and makes it clear which records enter
the last-5 window when more than 5 source JSONs are provided.
Check that `HF_API_KEY` is exported. The sync path may be public, but the
upload path requires write access.
### 1a. Check source batch consistency
Before syncing or preparing uploads, reject source batches that are internally
inconsistent. Use the same metric direction as `compare_baseline.py`:
- Lower is better: `latency`, `memory`, `text_encoder_time_s`, `dit_time_s`,
`vae_decode_time_s`.
- Higher is better: `throughput`.
For each metric with at least two non-null source values:
1. Compute the source batch median.
2. For lower-is-better metrics, compute `(source_value - batch_median) / batch_median`.
3. For `throughput`, compute `(batch_median - source_value) / batch_median`.
4. Stop if any source record regresses against the batch median by more than
`max_intra_batch_regression`.
Default `max_intra_batch_regression` to `PERF_MAX_REGRESSION` when set,
otherwise `0.05`. Print a table with per-source values, batch median, and
worst intra-batch regression.
This check prevents uploading a mixed batch where one JSON is materially
slower or faster than the others. If the batch fails this check, ask the user
to provide a cleaner batch or explicitly investigate the variance. Do not
silently drop outliers unless the user gives a concrete reviewed reason and a
new source list.
### 1b. How to obtain source results from CI
The performance CI exports normalized source results for failed rolling
baseline comparisons when `compare_baseline.py` ran. The preferred artifacts
come from:
```text
perf_reports/results/normalized_perf_*.json
```
The normal operator flow is:
1. Open the failed Buildkite performance job or several reruns of the same
benchmark after the accepted environment shift.
2. Download the `normalized_perf_*.json` artifacts for the target benchmark.
3. Pass all reviewed local paths or artifact URLs as `source_results`.
Do not scrape the Markdown performance summary to reconstruct JSON. The
normalized JSON artifacts are the only supported source of truth for reseed
metrics and provenance. Raw `fastvideo/tests/performance/results/perf_*.json`
artifacts are not accepted by this skill. If no normalized JSON artifact is
present, that run is not a valid source for baseline reseeding.
### 2. Sync and back up existing HF records under /tmp
Use `fastvideo/tests/performance/hf_store.py` helpers directly. Do **not** use
`compare_baseline.py` as a sync shortcut; on full main runs it can persist
records, while this step must only fetch and back up existing history.
The sync command pattern is:
```bash
export PERFORMANCE_TRACKING_ROOT="${PERFORMANCE_TRACKING_ROOT:-/tmp/perf-tracking}"
export HF_REPO_ID="${HF_REPO_ID:-FastVideo/performance-tracking}"
PYTHONPATH=fastvideo/tests/performance python -c 'from hf_store import sync_from_hf; import os; sync_from_hf(os.environ["PERFORMANCE_TRACKING_ROOT"], strict=True)'
```
Then back up only the sanitized model directory under `/tmp`:
```bash
SHORT_COMMIT=$(git rev-parse --short=12 HEAD)
TIMESTAMP=$(date -u +%Y%m%d_%H%M%S)
MODEL_SAFE=$(PYTHONPATH=fastvideo/tests/performance python - <<'PY'
from hf_store import sanitize
print(sanitize("<model_id>"))
PY
)
BACKUP_DIR="/tmp/performance_reseed_backup/${TIMESTAMP}_${SHORT_COMMIT}_${MODEL_SAFE}"
mkdir -p "$BACKUP_DIR"
cp -R "${PERFORMANCE_TRACKING_ROOT}/${MODEL_SAFE}" "$BACKUP_DIR/" 2>/dev/null || true
```
Write provenance next to the backup:
```bash
cat > "$BACKUP_DIR/PROVENANCE.txt" <<EOF
model_id: <model_id>
gpu_type: <gpu_type>
source_results:
- <source_result_1>
- <source_result_2>
reseed_record_count: <len(source_results)>
max_intra_batch_regression: <threshold>
head_commit: $(git rev-parse HEAD)
timestamp_utc: $(date -u +%FT%TZ)
reason: <intent_rationale>
EOF
```
If the backup has no prior records, this is not a destructive reseed; it is a
first baseline seed. Continue, but report that baseline history was empty.
### 3. Compute old baseline and candidate shift
Load the last 5 successful records for the target:
```python
from hf_store import load_records_for_model
records = load_records_for_model(
"/tmp/perf-tracking",
"<model_id>",
"<gpu_type>",
last_n=5,
successful_only=True,
)
```
Print a small table showing old medians, source batch medians, candidate
medians after appending the proposed seed records, and source batch spread for:
- `latency`
- `throughput`
- `memory`
- `text_encoder_time_s`
- `dit_time_s`
- `vae_decode_time_s`
Also print how many successful old records exist. Make clear:
- 1 seed record usually does not move a last-5 median by itself.
- 3 consistent seed records usually move the last-5 median immediately.
- 5 consistent seed records effectively reset the last-5 window.
- The records are intentional approved baseline resets and must be labeled
that way.
### 4. Confirm intent
Require an explicit confirmation phrase before preparing the upload:
> About to RE-SEED performance baseline for `<model_id>` on `<gpu_type>`.
> This will upload `<N>` new `success=true` records to
> `FastVideo/performance-tracking/<sanitize(model_id)>/`, one per accepted
> source JSON.
>
> Reason: `<intent_rationale>`
> Source results: `<source_results>`
> Reseed record count: `<N>`
> Max intra-batch regression: `<threshold>`
> Note: these records come from a reviewed source batch and are intended to
> move the rolling median to the accepted runtime profile. They are not
> ordinary main-branch persistence.
> HEAD: `<git rev-parse --short=12 HEAD>`
> Backup: `<BACKUP_DIR>`
>
> Reply `confirm performance reseed` to proceed, anything else to abort.
Do not continue unless the user types exactly `confirm performance reseed`.
### 5. Create the accepted seed records
Create one seed record from each normalized source result. Do not copy the
source JSON wholesale.
Infer the baseline field allowlist from all existing HF records for the target
`(model_id, gpu_type)` after syncing, including both `success=true` and
`success=false` records. Use the union of non-provenance keys present in those
target records, preserving only fields that also exist in the normalized
source record or are explicitly set by the reseed workflow. Always include
`model_id`, `timestamp`, and `success` because the upload path and baseline
loader depend on them. Always set `timestamp` to a fresh reseed timestamp and
`success` to `true`. Do not include unrelated source-only fields that are
absent from existing HF records.
Exclude existing provenance or operator metadata from the inferred baseline
field allowlist. At minimum, exclude keys prefixed with `baseline_reseed` and
any fields known to be local-only audit metadata.
If there are no previous HF records for the target model/GPU, fall back to this
default baseline field list:
- `model_id`
- `timestamp`
- `commit_sha`
- `gpu_type`
- `latency`
- `throughput`
- `memory`
- `text_encoder_time_s`
- `dit_time_s`
- `vae_decode_time_s`
- `success`
Do not upload extra fields from the source artifact.
Optional provenance fields are allowed and useful:
- `baseline_reseed: true`
- `baseline_reseed_reason`
- `baseline_reseed_source_result`
- `baseline_reseed_source_timestamp`
- `baseline_reseed_source_success`
- `baseline_reseed_batch_size`
- `baseline_reseed_batch_index`
- `baseline_reseed_operator`
- `baseline_reseed_max_intra_batch_regression`
Use a fresh reseed timestamp for each seed record, not the original source
result timestamp. This is required because
`load_records_for_model(..., last_n=5)` keeps the last records after loading
the model directory; stale filenames/timestamps may not enter the last-5
window and therefore may not move the median. Preserve the original source
timestamp in `baseline_reseed_source_timestamp`.
Use the existing filename convention from `_write_tracking_record()`:
`<sanitize(timestamp)>_<sanitize(commit_sha)>.json` under the sanitized model
directory, but include a deterministic suffix such as `_reseed_01`,
`_reseed_02`, and so on before `.json` so multiple records from the same
batch do not overwrite each other.
If a source record already exists on HF with `success=false`, do not edit it
in place unless the user explicitly asked for an audit-preserving correction.
Prefer uploading new accepted seed records so failed history remains visible.
### 6. Pause before upload
Print:
- Backup directory path under `/tmp`.
- Prepared local record paths under `PERFORMANCE_TRACKING_ROOT`.
- HF paths that will receive the new records.
- Old rolling medians.
- Source batch medians, source batch spread, reseed count, and candidate
medians.
- Rationale.
Ask the user to reply exactly `upload`. Anything else aborts and leaves the
prepared records plus backup on disk.
### 7. Upload only the scoped records
Use the shared storage helper so the path and repo type match CI:
```python
from hf_store import upload_record
upload_record("<local_record_path>", record, strict=True)
```
Run it once per prepared record. Each upload goes to:
```text
FastVideo/performance-tracking/<sanitize(model_id)>/<record_filename>.json
```
Never bulk upload the whole tracking root. Never modify another model's
directory in the same operation.
### 8. Report outcome and offer cleanup
Report:
- Uploaded HF paths.
- Backup directory under `/tmp`.
- Local tracking root, usually `/tmp/perf-tracking`.
- Old baseline window count and medians.
- Source batch medians, source batch spread, reseed count, and candidate
medians.
- Expected effect based on reseed count.
- Any separate threshold changes still needed in
`.buildkite/performance-benchmarks/tests/*.json`.
Include the `intent_rationale` in the PR or follow-up comment so reviewers can
distinguish an accepted baseline shift from a hidden regression.
After the upload is verified, ask whether the user wants to clear temporary
local state. Explain what each directory is for:
- `PERFORMANCE_TRACKING_ROOT`, usually `/tmp/perf-tracking`: local synced
mirror of `FastVideo/performance-tracking` plus the prepared local seed
records used for scoped upload.
- `/tmp/performance_reseed_backup/<...>`: local backup of the target model's
pre-reseed HF history plus `PROVENANCE.txt`, kept so a bad reseed can be
audited or corrected.
- `/tmp/performance_reseed_source/<...>` when used: downloaded source JSON
artifacts from Buildkite URLs.
Ask:
> Reseed succeeded. Do you want me to delete the local temp tracking mirror,
> source downloads, and reseed backup under `/tmp`? These files are local
> safety/audit artifacts only; HF already has the uploaded records.
>
> Reply `cleanup reseed temp` to delete them, anything else to keep them.
Do not delete anything unless the user replies exactly
`cleanup reseed temp`. If cleanup is requested, remove only the specific
directories created for this reseed. Never remove unrelated `/tmp` contents.
## Failure modes and handling
- **`HF_API_KEY` unset.** Stop before upload. Do not create an untracked
process that appears to have reseeded but never reached HF.
- **Source result does not match target.** Stop. The wrong benchmark or GPU
would poison a separate baseline.
- **Source batch is internally inconsistent.** Stop if any source regresses
against the source batch median by more than `max_intra_batch_regression`.
Ask for cleaner sources or a reviewed explanation before continuing.
- **Too few source records to move the median.** Continue only after making
clear that one or two records may not immediately move the last-5 median.
- **The source results are noisy or suspicious.** Stop. Reseeding amplifies
those measurements into the baseline, so they must be reviewed first.
- **HF sync fails.** Stop for destructive reseeds. A stale or empty sync can
make the old baseline look missing.
- **Candidate still violates fixed thresholds.** Report that this skill only
handles the rolling HF baseline; update benchmark JSON thresholds in code
review if maintainers accept the new absolute limit.
- **The user aborts at either confirmation.** Leave the backup and prepared
records on disk. Nothing should be uploaded.
- **The user declines cleanup.** Keep `/tmp/perf-tracking`, the source
download directory if any, and `/tmp/performance_reseed_backup/<...>` in
place for audit/debugging.
- **A bad seed was uploaded.** Use the backup and HF history to identify the
uploaded file, then remove or supersede it with an explicitly reviewed
corrective record. Do not silently rewrite unrelated history.
## References
- `.agents/skills/reseed-ssim-references/SKILL.md` — safety pattern for
intentional baseline replacement.
- `fastvideo/tests/performance/compare_baseline.py` — normalization, rolling
median comparison, and persistence rules.
- `fastvideo/tests/performance/hf_store.py` — HF sync, record loading,
`sanitize()`, and `upload_record()`.
- `fastvideo/tests/performance/test_inference_performance.py` — source result
JSON schema.
- `.buildkite/performance-benchmarks/tests/*.json` — fixed absolute benchmark
thresholds, separate from rolling baseline comparisons.
## Changelog
| Date | Change |
|------|--------|
| 2026-05-03 | Initial version. Sister workflow to `reseed-ssim-references`, scoped to one performance `(model_id, gpu_type)` baseline seed with backup, confirmation, provenance, and `success=true` upload. |
| 2026-05-03 | Previous policy: replicate one approved shifted source result into 3 success records by default, or 5 only when explicitly requested. Add provenance marker for replicated-source reseeds. Superseded by the 2026-05-08 dynamic multi-source policy. |
| 2026-05-08 | Replace fixed 3/5 replication with dynamic multi-source reseeding: upload one seed record per reviewed source JSON, validate intra-batch consistency, move backup/source scratch under `/tmp`, and ask whether to clean temp state after successful upload. |
@@ -0,0 +1,343 @@
---
name: reseed-ssim-references
description: Re-seed HF reference videos for a single existing SSIM test on Modal L40S. Always backs up current refs locally first, regenerates on Modal, pauses for the user to eyeball before-vs-after quality, then overwrites the targeted `<model_id>` subtree on `FastVideo/ssim-reference-videos` with `--force`. Use when an intentional code change (model port fix, attention backend swap, kernel upgrade, hyperparameter change) has invalidated existing refs and they need to be regenerated. Pairs with `seed-ssim-references`, which is for first-time seeding only.
---
# Re-seed SSIM Reference Videos
## Purpose
Replace the existing SSIM reference videos for a single `(test_file, model_id)`
pair on the HF dataset (`FastVideo/ssim-reference-videos`). This is **destructive**
on HF — the old refs are overwritten — so the skill always:
1. Confirms intent with a one-liner the user has to type.
2. Downloads the existing refs as a local, timestamped backup.
3. Regenerates on Modal L40S (same code path that CI uses).
4. Pauses for a side-by-side eyeball of backup vs new mp4s.
5. Uploads with `--force`, scoped to the single `--model-id`.
6. Reminds the user to keep the backup until the PR lands.
Pairs with `seed-ssim-references`, which is the inverse (first-time seeding
only, refuses to overwrite). Re-seeding is intentionally a separate, more
ceremonial operation because mistakenly clobbering production refs is much
harder to recover from than failing closed.
## When to use
- An intentional code change (model port fix, kernel upgrade, attention
backend swap, hyperparameter change in the test itself) has shifted the
expected SSIM output and the existing refs no longer represent the new
ground truth.
- A test is failing in CI **for the right reason** (the new code is correct,
the old refs are stale).
## When not to use
- A test is failing for the **wrong** reason (the port is buggy, not the
refs). Fix the port; re-seeding hides the bug.
- A brand-new test that has no refs on HF yet. Use `seed-ssim-references`.
- "Just to clean up drift" without a concrete code change to point at. The
PR description has to justify *why* refs changed; without a concrete
change, there's nothing to write.
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| `test_file` | Yes | Path to the SSIM test, e.g. `fastvideo/tests/ssim/test_matrixgame_similarity.py`. Validated against `fastvideo/tests/ssim/test_*_similarity.py`. |
| `model_id` | Yes | Single model id from the test's `*_MODEL_TO_PARAMS`, e.g. `Matrix-Game-2.0-Diffusers-Base`. Re-seed runs are **per model**. For multi-model tests, invoke the skill once per model. |
| `intent_rationale` | Yes | One-line explanation of *why* refs are being regenerated (e.g. "Relax FA-2 head_size whitelist to include 80 — matrix_game now uses FLASH_ATTN instead of TORCH_SDPA"). Recorded in the backup directory and reused in the PR description. |
Hardcoded:
- Modal GPU: **L40S** (matches CI; re-seeding from another SKU produces refs
that L40S CI cannot match).
- Quality tier: **`default`**. `full_quality` is a separate, deliberate
operation.
- HF repo: `FastVideo/ssim-reference-videos` (override via
`FASTVIDEO_SSIM_REFERENCE_HF_REPO`).
- Device folder: `L40S_reference_videos`.
## Prerequisites
The user has confirmed:
- `modal` CLI authenticated.
- `hf` CLI authenticated, **and** `HF_API_KEY` (or `HUGGINGFACE_HUB_TOKEN` /
`HF_TOKEN`) exported with **write** access to
`FastVideo/ssim-reference-videos`.
- The current branch's code is the change that motivated the re-seed (i.e.
`git rev-parse HEAD` is the commit that intentionally invalidated refs).
Fail fast if any of these are missing.
## Steps
### 1. Validate inputs and confirm intent
- Verify `test_file` exists and matches `fastvideo/tests/ssim/test_*_similarity.py`.
- Grep the file for `*_MODEL_TO_PARAMS` and assert `model_id` is one of its
keys. If the file has only a single hardcoded model, accept that model id
as the only valid value.
- Print the rationale and ask the user to type **`confirm reseed`** (not just
`y` — make it deliberate):
> About to RE-SEED references for model `<model_id>` from test `<test_file>`.
> This will OVERWRITE existing refs on
> `FastVideo/ssim-reference-videos/reference_videos/default/L40S_reference_videos/<model_id>/`
> after backup + Modal regen + eyeball.
>
> Reason: `<intent_rationale>`
> HEAD: `<git rev-parse --short=12 HEAD>`
>
> Reply `confirm reseed` to proceed, anything else to abort.
Stop until the user types exactly `confirm reseed`. Anything else aborts
with no side effects.
### 2. Back up existing refs
Always required. The backup is the only graceful path back if anything goes
wrong later.
```bash
SHORT_COMMIT=$(git rev-parse --short=12 HEAD)
TIMESTAMP=$(date -u +%Y%m%d_%H%M%S)
MODEL_SAFE=$(echo "<model_id>" | tr '/' '_')
BACKUP_DIR="ssim_reseed_backup/${TIMESTAMP}_${SHORT_COMMIT}_${MODEL_SAFE}"
mkdir -p "$BACKUP_DIR"
hf download \
--repo-type dataset FastVideo/ssim-reference-videos \
--include "reference_videos/default/L40S_reference_videos/<model_id>/**" \
--local-dir "$BACKUP_DIR"
mp4_count=$(find "$BACKUP_DIR" -name "*.mp4" | wc -l)
echo "Backup mp4 count: $mp4_count"
[ "$mp4_count" -gt 0 ] || {
echo "ERROR: backup is empty for <model_id>. Either the model id is wrong"
echo "or there are no existing refs (use seed-ssim-references instead)."
exit 1
}
# Provenance — used in the PR description
cat > "$BACKUP_DIR/PROVENANCE.txt" <<EOF
test_file: <test_file>
model_id: <model_id>
head_commit: $(git rev-parse HEAD)
timestamp_utc: $(date -u +%FT%TZ)
reason: <intent_rationale>
EOF
```
If the `hf download` produces zero mp4s, abort — the user has either picked a
non-existent `model_id` or there are no refs yet (in which case
`seed-ssim-references` is the right tool).
### 3. Regenerate on Modal L40S
Mirror CI's exact env recipe so the regenerated refs are byte-comparable to
what CI will produce on the same commit. Two differences from CI:
1. **Pass the same env prefix CI uses** (`IMAGE_VERSION`, `BUILDKITE_*`) — see
`.buildkite/pipeline.yml:1-3` and `.buildkite/scripts/pr_test.sh:62-83`.
Without this, `ssim_test.py:17-18` resolves a different GHCR image tag
(default is `latest`, CI is `py3.12-latest`), and `ssim_test.py:38-46`
bakes different values into the image's frozen env block. **Mismatched
image or env is the most common source of SSIM drift between reseed and
CI runs.**
2. **Do not pass `--skip-reference-download`**. Letting the test fetch the
existing refs and run the full SSIM compare gives "before" SSIM numbers
for the PR description, and the test still produces the new mp4s
regardless of whether the comparison passes or fails.
```bash
SUBDIR="${TIMESTAMP}_${SHORT_COMMIT}"
IMAGE_VERSION="py3.12-latest" \
BUILDKITE_REPO="$(git config --get remote.origin.url)" \
BUILDKITE_COMMIT="$(git rev-parse HEAD)" \
BUILDKITE_PULL_REQUEST="${BUILDKITE_PULL_REQUEST:-false}" \
modal run fastvideo/tests/modal/ssim_test.py \
--git-repo="$(git config --get remote.origin.url)" \
--git-commit="$(git rev-parse HEAD)" \
--hf-api-key="$HF_API_KEY" \
--test-files="<test_file>" \
--sync-generated-to-volume \
--generated-volume-subdir="$SUBDIR" \
--no-fail-fast
```
Capture the printed `modal volume get ...` hint — its `<SUBDIR>` matches
`$SUBDIR` and is needed for step 4. Capture the SSIM numbers from the test
output (or from the JSON next to the generated mp4) for the PR description.
### 4. Download generated videos
```bash
modal volume get --force hf-model-weights \
ssim_generated_videos/default/"$SUBDIR"/generated_videos \
./generated_videos_modal/default
```
After this, the new mp4s live at:
```
./generated_videos_modal/default/generated_videos/L40S_reference_videos/<model_id>/<backend>/<prompt>.mp4
```
`--force` is required when `./generated_videos_modal/default` already exists
from a prior run; safe on the first run too.
### 5. PAUSE — user reviews quality side-by-side
Print the diff and the comparison:
```bash
echo "=== File list diff (backup vs new) ==="
diff -u \
<(find "$BACKUP_DIR/reference_videos/default/L40S_reference_videos/<model_id>" -name "*.mp4" \
| sed "s|$BACKUP_DIR/reference_videos/default/L40S_reference_videos/||" | sort) \
<(find ./generated_videos_modal/default/generated_videos/L40S_reference_videos/<model_id> -name "*.mp4" \
| sed "s|./generated_videos_modal/default/generated_videos/L40S_reference_videos/||" | sort) \
|| true
echo
echo "=== SSIM numbers from this run (paste into PR) ==="
find ./generated_videos_modal/default/generated_videos/L40S_reference_videos/<model_id> -name "*_ssim.json" -exec cat {} \;
```
Then stop and tell the user:
> Old refs backed up to `$BACKUP_DIR`.
> New videos in `./generated_videos_modal/default/generated_videos/L40S_reference_videos/<model_id>/`.
>
> Open both in a video player. Confirm the new videos:
> 1. Look correct (no obvious artifacts, no black/static frames).
> 2. Are *intentionally* different from the backup in the way described
> in `<intent_rationale>` (e.g. slight numerical drift only, not a
> different scene / different motion / corrupted output).
>
> Reply **`upload`** to overwrite HF, anything else to abort.
> Aborting leaves the backup and new videos on disk for inspection — nothing
> on HF changes.
Do not proceed until the user types exactly `upload`. If they abort, leave
everything on disk and stop here.
### 6. Copy into the local reference layout
Same as `seed-ssim-references` step 5:
```bash
python fastvideo/tests/ssim/reference_videos_cli.py copy-local \
--quality-tier default \
--device-folder L40S_reference_videos \
--generated-dir ./generated_videos_modal/default/generated_videos/L40S_reference_videos
```
Result: `fastvideo/tests/ssim/reference_videos/default/L40S_reference_videos/<model_id>/<backend>/<prompt>.mp4`.
### 7. Upload with `--force`, scoped to `--model-id`
The `--force` flag is what makes this skill different from `seed-ssim-references`.
Always pair it with `--model-id` so a typo cannot accidentally overwrite a
neighboring model's refs.
```bash
python fastvideo/tests/ssim/reference_videos_cli.py upload \
--quality-tier default \
--device-folder L40S_reference_videos \
--model-id "<model_id>" \
--force
```
The CLI's overwrite guard refuses without `--force`; with `--force` it
overwrites only files under
`reference_videos/default/L40S_reference_videos/<model_id>/`.
### 8. Report success and retention guidance
Print:
- The HF path that was overwritten (`<repo>/reference_videos/default/L40S_reference_videos/<model_id>/`).
- The local backup directory path.
- The new SSIM numbers from step 5.
- This restore command, in case the PR review surfaces a problem after
upload:
```bash
python fastvideo/tests/ssim/reference_videos_cli.py upload \
--quality-tier default \
--device-folder L40S_reference_videos \
--model-id "<model_id>" \
--reference-dir "$BACKUP_DIR/reference_videos/default/L40S_reference_videos" \
--force
```
- This PR-description checklist (see `fastvideo/tests/ssim/AGENTS.md` →
*Updating Reference Videos*):
1. Source commit that produced the new refs (HEAD at re-seed time).
2. Test command and GPU SKU (`L40S`).
3. Before/after SSIM numbers.
4. The `<intent_rationale>` from step 1.
5. A note that the backup lives at `$BACKUP_DIR` and should be retained
until CI on the PR is green.
Do **not** auto-rerun the SSIM test — the user does that as part of the PR.
## Failure modes and how to handle them
- **`HF_API_KEY` unset.** Stop before step 2.
- **Backup is empty (zero mp4s).** Stop before step 3 — the model id is
wrong or the refs don't exist yet (use `seed-ssim-references`).
- **Modal run fails before generation.** No mp4s on the volume. Don't
upload. Investigate the failure (test crash, OOM, partition exhaustion),
fix, then retry from step 3. Backup is still intact.
- **Quality regressed (visual or metric).** User aborts at step 5. Backup
retained. New videos retained on disk for inspection. Nothing on HF
changed. Either fix the underlying code change or abandon the re-seed.
- **User confirmed `upload` but later realized the new refs are wrong.**
Run the restore command from step 8 with the backup `--reference-dir`.
This is exactly why the backup exists.
- **Multi-model test, only one model is being re-seeded.** Run the skill
once per model id. The `--model-id` scope on upload guarantees the others
are untouched.
## Design notes (for future skill maintainers)
- Per-`model_id` scope is mandatory. The dataset houses many model subtrees;
re-seeding the wrong one is hard to undo without backup.
- `default` tier only; `full_quality` is a separate, deliberate operation
with different params and ~doubled runtime, and isn't what CI gates on.
- The skill deliberately does **not** pass `--skip-reference-download` to
Modal so we get pre-reseed SSIM numbers for the PR. The `seed`-skill
passes it because no refs exist yet; for re-seed, refs do exist and
exposing the comparison is informative.
- The two-token confirm (`confirm reseed`, then `upload`) is intentional.
Re-seeding is high-blast-radius and should not be one-keystroke.
- The backup directory is plain mp4s + `PROVENANCE.txt`. No HF metadata is
preserved; the restore path uses `reference_videos_cli.py upload
--reference-dir` which doesn't need it.
## References
- `.agents/skills/seed-ssim-references/SKILL.md` — the first-time seed
skill this one parallels. Read it for the Modal flag rationale shared
between the two flows.
- `fastvideo/tests/ssim/AGENTS.md` — directory rules, including the PR
expectations for any reference-video change (rationale, before/after
SSIM, source commit/model/backend).
- `fastvideo/tests/ssim/reference_videos_cli.py` — `copy-local`, `upload`
(with `--model-id`, `--force`), `download`. The overwrite guard at
`upload_reference_videos` is the safety net this skill leans on.
- `fastvideo/tests/modal/ssim_test.py` — Modal orchestrator;
`--sync-generated-to-volume`, `--generated-volume-subdir`,
`--skip-reference-download`, `--no-fail-fast`.
## Changelog
| Date | Change |
|------|--------|
| 2026-05-02 | Initial version. Sister skill to `seed-ssim-references`, scoped to single `(test_file, model_id)` re-seeds, with mandatory backup and two-token confirm. |
@@ -0,0 +1,378 @@
---
name: seed-ssim-references
description: Seed HF reference artefacts for a single newly-added SSIM test (pixel `.mp4` for `run_text_to_video_similarity_test`-style tests, or latent `.pt` for `run_text_to_latent_similarity_test`-style tests). Runs the test on Modal L40S, downloads the generated artefacts via `modal volume get`, pauses for the user to verify (visual eyeball for mp4, numerics dump for pt), then uploads only that test's files to `FastVideo/ssim-reference-videos`. Use when a new `fastvideo/tests/ssim/test_*_similarity.py` has just been added and has no references on HF yet.
---
# Seed SSIM Reference Artefacts (mp4 or pt)
## Purpose
A brand-new SSIM test in `fastvideo/tests/ssim/` fails forever until its
reference artefacts exist on the HF dataset
(`FastVideo/ssim-reference-videos`). The dataset hosts two kinds of artefacts
side-by-side per `(model_id, backend, prompt)`:
- **`.mp4`** — pixel ground-truth for tests that call
`run_text_to_video_similarity_test` / `run_image_to_video_similarity_test`
in `inference_similarity_utils.py`. Compared via SSIM.
- **`.pt`** — pre-VAE latent bundle (fp16 full latent + fp32 slice +
metadata + `slice_spec` + `format_version`) for tests that call
`run_text_to_latent_similarity_test` in `latent_similarity_utils.py`.
Compared via cosine distance on the slice and the full tensor.
This skill:
1. Detects which artefact type the test produces (pixel vs latent).
2. Runs the test on Modal's L40S pool to generate the artefacts.
3. Downloads them to the local repo via `modal volume get`.
4. Pauses so the user can verify quality:
- **mp4**: visual eyeball in a video player.
- **pt**: numerics dump (shape, slice stats, NaN/Inf check, metadata).
5. Uploads only the new test's files to HF, with a guard that refuses to
overwrite anything already present.
The skill is run **manually**, once per new test. Before invoking it, the user
has already sanity-tested the new test locally — it launches `VideoGenerator`
and writes an artefact without crashing (the missing-reference assertion at
the end is expected). The skill does not re-test locally; it goes straight
to Modal L40S (which is what CI uses).
## When to use
- A new `test_*_similarity.py` file has been added in `fastvideo/tests/ssim/`
and the HF dataset has no `reference_videos/default/L40S_reference_videos/<model_id>/`
subtree for it yet.
## When not to use
- Regular CI runs — once refs exist, `pytest fastvideo/tests/ssim/` downloads
them automatically.
- Re-seeding an existing test. That requires `--force` on the upload step, and
is out of scope here; treat as a separate, deliberate operation.
## Inputs
The skill has **one required input**: the path to the new SSIM test file.
Prompt the user for it if they didn't supply it.
| Parameter | Required | Description |
|-----------|----------|-------------|
| `test_file` | Yes | e.g. `fastvideo/tests/ssim/test_ltx2_similarity.py`. The skill's first action is to ask for this if missing. |
Everything else is fixed:
- Modal runner GPU: **L40S** (hardcoded in `fastvideo/tests/modal/ssim_test.py`).
- Device folder: `L40S_reference_videos`.
- Quality tier: `default` (the tier CI runs). The `full_quality` tier is not
seeded by this skill.
- HF repo: `FastVideo/ssim-reference-videos` (dataset).
- Multi-model test files: all model ids in `*_MODEL_TO_PARAMS` are seeded
together; the Modal run produces one mp4 per (model, prompt, backend) and
the upload scopes by `--model-id`, looping if there is more than one.
## Prerequisites
The user has confirmed:
- `modal` CLI authenticated.
- `HF_API_KEY` (or `HUGGINGFACE_HUB_TOKEN` / `HF_TOKEN`) exported with write
access to `FastVideo/ssim-reference-videos`.
- The test file runs locally end-to-end (generates an mp4; SSIM assertion
failure due to missing reference is expected and fine).
Fail fast if the token env var is missing.
## Steps
### 1. Ask for the test file, then detect artefact type
If the user didn't name one, ask: *"Which SSIM test file do you want to seed
references for? (e.g. `fastvideo/tests/ssim/test_ltx2_similarity.py`)"*.
Validate:
- Path exists and matches `fastvideo/tests/ssim/test_*_similarity.py`.
- File defines a `*_MODEL_TO_PARAMS` dict — grep it to extract the set of
model ids. Those ids drive step 5.
Detect artefact type by inspecting the file's imports / helper call:
- **latent** (`.pt`) — file imports `run_text_to_latent_similarity_test`
from `fastvideo.tests.ssim.latent_similarity_utils` (or any other helper
that ends with `_latent_similarity_test`).
- **pixel** (`.mp4`) — file imports
`run_text_to_video_similarity_test` / `run_image_to_video_similarity_test`
from `fastvideo.tests.ssim.inference_similarity_utils`, OR uses the
legacy custom-inline helper pattern (see `test_gamecraft`,
`test_longcat`, etc.). Default to pixel when both heuristics fail.
Record `ARTEFACT_TYPE ∈ {pixel, latent}` for use in step 4. Steps 2, 3, 5,
and 6 are artefact-type-agnostic — `_iter_reference_files`,
`copy_generated_to_reference`, and `upload_reference_videos` already walk
both `.mp4` and `.pt` (see `reference_videos_cli.py`).
If either check fails, stop and tell the user what's wrong.
### 2. Run the test on Modal L40S
Pick a subdir name so repeated runs don't collide:
```bash
SHORT_COMMIT=$(git rev-parse --short=12 HEAD)
TIMESTAMP=$(date -u +%Y%m%d_%H%M%S)
SUBDIR="${TIMESTAMP}_${SHORT_COMMIT}"
```
Then launch the Modal run. The `IMAGE_VERSION` and `BUILDKITE_*` env-prefix
**must** match what CI exports in `.buildkite/scripts/pr_test.sh`, otherwise
`fastvideo/tests/modal/ssim_test.py` resolves a different GHCR image tag
(default is `latest`, CI is `py3.12-latest`) and bakes different values into
the image's frozen env block (`ssim_test.py:17-18, 38-46`). Mismatched image
or env produces SSIM drift that doesn't show up until the same commit runs
in CI.
```bash
IMAGE_VERSION="py3.12-latest" \
BUILDKITE_REPO="$(git config --get remote.origin.url)" \
BUILDKITE_COMMIT="$(git rev-parse HEAD)" \
BUILDKITE_PULL_REQUEST="${BUILDKITE_PULL_REQUEST:-false}" \
modal run fastvideo/tests/modal/ssim_test.py \
--git-repo="$(git config --get remote.origin.url)" \
--git-commit="$(git rev-parse HEAD)" \
--hf-api-key="$HF_API_KEY" \
--test-files="<test_file>" \
--sync-generated-to-volume \
--generated-volume-subdir="$SUBDIR" \
--skip-reference-download \
--no-fail-fast
```
Env prefix rationale (parity with CI; see `.buildkite/pipeline.yml:1-3` and
`.buildkite/scripts/pr_test.sh:62-83`):
- `IMAGE_VERSION=py3.12-latest`: pins the Modal image tag to the same one CI
uses. The published `py3.12-latest` and `latest` tags point at Python 3.12 /
CUDA 12.6.3 / cu126; `py3.12-cuda12.6.3-latest` is the explicit alias for the
same image. CUDA 13 / cu130 is available under the explicit
`py3.12-cuda13.0.0-latest` tag. This tag policy comes from
`infra-build-image.yml`; the unparameterized `docker/Dockerfile` build itself
still defaults to CUDA 13 / cu130.
- `BUILDKITE_REPO`/`BUILDKITE_COMMIT`/`BUILDKITE_PULL_REQUEST`: mirror what
Buildkite exports. `ssim_test.py:38-46` bakes these into the image's
`.env(...)` block; mismatched values can perturb in-container code paths
that branch on PR-vs-non-PR. `false` for `BUILDKITE_PULL_REQUEST` matches
Buildkite's "non-PR build" sentinel.
Flag rationale:
- `--skip-reference-download`: no refs exist yet, so conftest must not try to
pull them.
- `--no-fail-fast`: lets the test finish generation before `_assert_similarity`
raises `FileNotFoundError: Reference video folder does not exist`. The
expected failure is what we want — the mp4 has already been written.
- `--sync-generated-to-volume` + `--generated-volume-subdir`: copies the
generated mp4s to the `hf-model-weights` Modal volume under
`ssim_generated_videos/default/<SUBDIR>/generated_videos/` so we can pull
them locally.
The Modal run will end with a nonzero exit (expected) and print a
`modal volume get hf-model-weights ssim_generated_videos/default/<SUBDIR>/generated_videos ./generated_videos_modal/default`
command. Capture that `<SUBDIR>` — you need it for step 3.
### 3. Download generated videos locally
```bash
modal volume get --force hf-model-weights \
ssim_generated_videos/default/"$SUBDIR"/generated_videos \
./generated_videos_modal/default
```
`--force` is required when the parent `./generated_videos_modal/default`
already exists; without it, `modal volume get` errors with `[Errno 21] Is a
directory`. Safe to pass on the first run too.
After this, the mp4s live at
`./generated_videos_modal/default/generated_videos/L40S_reference_videos/<model_id>/<backend>/<prompt>.mp4`.
The extra `generated_videos/` level comes from the volume layout in
`_sync_generated_videos_to_volume` (`ssim_test.py`) — the command copies
`<repo>/fastvideo/tests/ssim/generated_videos/<tier>` to
`ssim_generated_videos/<tier>/<SUBDIR>/generated_videos/`, and `modal volume
get` preserves that trailing `generated_videos/` segment.
### 4. PAUSE — user reviews quality
Type-aware verification.
**For `ARTEFACT_TYPE = pixel`** — list the downloaded mp4s and ask the user to
open them in a video player:
> "Generated videos downloaded to `./generated_videos_modal/default/generated_videos/L40S_reference_videos/`. Please open them and confirm the quality looks correct. Reply **`upload`** to continue, or anything else to abort."
**For `ARTEFACT_TYPE = latent`** — `.pt` files are not human-watchable. Print
a numerics dump for each `.pt` so the user can sanity-check shape, distribution,
and metadata:
```python
import torch
from pathlib import Path
ROOT = Path("./generated_videos_modal/default/generated_videos/L40S_reference_videos")
for p in sorted(ROOT.rglob("*.pt")):
d = torch.load(p, map_location="cpu", weights_only=False)
s = d["expected_slice"]
L = d["latent"].float()
print(f"=== {p.relative_to(ROOT)} ===")
print(f" format_version: {d['format_version']}")
print(f" shape: {d['shape']}")
print(f" dtype_original: {d['dtype_original']}")
print(f" slice_spec: {d['slice_spec']}")
print(f" slice shape={tuple(s.shape)} mean={s.mean():+.4f} std={s.std():.4f} min={s.min():+.4f} max={s.max():+.4f}")
print(f" latent shape={tuple(L.shape)} mean={L.mean():+.4f} std={L.std():.4f} min={L.min():+.4f} max={L.max():+.4f}")
print(f" finite: latent NaN={torch.isnan(L).any().item()} Inf={torch.isinf(L).any().item()}; "
f"slice NaN={torch.isnan(s).any().item()} Inf={torch.isinf(s).any().item()}")
print(f" metadata: {d['metadata']}\n")
```
Sanity criteria:
- `format_version == 1` (matches `LATENT_REFERENCE_FORMAT_VERSION`).
- `shape` matches what the model produces (e.g. LTX-2 distilled =
`[1, 128, T_lat, H_lat, W_lat]`; Stable Audio Open 1.0 = `[1, 64, 1024]`).
- `slice_spec.kind` matches a registered kind (`corner_3x3_first_frame`
for video, `audio_first_8_timesteps` for audio).
- No `NaN`/`Inf`. `mean ≈ 0`, `std ≈ 1` (denoised latents stay close to
the initial Gaussian distribution; very wide deviations suggest
numerical drift).
- `metadata.prompt` matches the test's prompt.
Then ask:
> "Numerics look right? Reply **`upload`** to continue, or anything else to abort."
Do not proceed until the user explicitly says `upload`. If they abort, leave
everything on disk so they can inspect further — no cleanup.
### 5. Copy into the local reference layout
Scoped copy — only the new test's artefacts. Single command works for both
artefact types because `_iter_reference_files` walks `.mp4` and `.pt`:
```bash
python fastvideo/tests/ssim/reference_videos_cli.py copy-local \
--quality-tier default \
--device-folder L40S_reference_videos \
--generated-dir ./generated_videos_modal/default/generated_videos/L40S_reference_videos
```
(The `--generated-dir` points at the device-folder root inside the
downloaded tree; `copy-local` walks all `<model>/<backend>/*.{mp4,pt}`
underneath it. Since the Modal run was scoped to a single test file via
`--test-files`, only that test's model(s) are present — so the copy is
implicitly per-test.)
Result for pixel: `fastvideo/tests/ssim/reference_videos/default/L40S_reference_videos/<model_id>/<backend>/<prompt>.mp4`.
Result for latent: same path with `.pt` extension.
### 6. Upload to HF — scoped per model_id, with overwrite guard
For each `<model_id>`:
```bash
python fastvideo/tests/ssim/reference_videos_cli.py upload \
--quality-tier default \
--device-folder L40S_reference_videos \
--model-id "<model_id>"
```
The upload command:
- Uploads **only** `reference_videos/default/L40S_reference_videos/<model_id>/`.
- **Refuses** if any file already exists at that path on HF (this is the
guard — seeding a new test should never clobber existing refs). To override,
the user must re-run with `--force`. If the guard fires, stop and report
exactly which files exist; do not silently `--force`.
Reads the HF token from `HF_API_KEY` / `HUGGINGFACE_HUB_TOKEN` / `HF_TOKEN`.
### 7. Report success
List what was uploaded (paths in repo) and remind the user to push any
related code changes. Do **not** auto-verify by re-running Modal — the user
can run `pytest fastvideo/tests/ssim/<test_file>` later to confirm end-to-end;
it will auto-download the refs they just uploaded.
## Failure modes and how to handle them
- **`HF_API_KEY` unset.** Stop before step 2. The Modal run needs it (passed
via `--hf-api-key`), and step 6 needs it for upload. If the user
ran `hf auth login` instead of exporting an env var, read the cached
token via `huggingface_hub.get_token()` and forward it to Modal as
`--hf-api-key="$CACHED_TOKEN"`.
- **Modal run fails before generation.** No artefacts on the volume — nothing
to download. Fix the test locally (`pytest fastvideo/tests/ssim/<test_file>`)
and retry from step 2.
- **`./generated_videos_modal/default/L40S_reference_videos/` missing after
`modal volume get`.** The run didn't produce artefacts (most likely the
test crashed before writing, or `REQUIRED_GPUS` exceeded the partition
capacity — see Modal logs).
- **Latent test crashed with FSDP / inference_mode error
(`RuntimeError: Inference tensors do not track version counter`).** The
test must pass `init_kwargs_override={"use_fsdp_inference": False}` when
`sp_size == 1` — see `test_stable_audio_similarity.py` for the pattern.
Fix in the test, push, retry.
- **Upload guard fires (files already exist).** The test name / model id
collides with something already on HF. Verify the user actually wants to
replace existing refs; if so, re-run the upload with `--force`. If not,
rename the model id in `*_MODEL_TO_PARAMS` and re-seed.
- **Quality looks wrong in step 4.** Abort. The artefacts stay on disk for
inspection. The fix is usually in the test's params (resolution, steps,
seed) — edit the test, then re-run the skill.
- For latent: also check `slice_spec.kind` matches the latent rank
(`corner_3x3_first_frame` requires 5-D, `audio_first_8_timesteps`
requires 3-D); a rank/kind mismatch raises in `_extract_expected_slice`.
## Design notes (for future skill maintainers)
- The skill deliberately runs on Modal, **not** locally, because the CI
runner is L40S. Seeding from a different GPU SKU produces refs that CI's
L40S runs can't match (pixel SSIM drifts across SKUs; latent cosine has
tighter cross-SKU bf16 drift but the configured tolerances assume
same-SKU seed → same-SKU verify).
- The skill is default-tier only. `full_quality` refs are seeded by a
separate, deliberate operation — they double runtime and aren't what CI
gates on.
- The overwrite guard in `reference_videos_cli.py upload` is default-on
specifically because this skill exists. Re-seeding is a distinct operation
that requires explicit `--force`.
- Both artefact types share the same Modal flow: the orchestrator sets
`--skip-reference-download` + `--no-fail-fast`, runs pytest, the test's
helper writes the artefact (`.mp4` via `imageio` for pixel,
`save_latent_reference` → `torch.save` for latent) BEFORE the
missing-reference assertion raises. `_sync_generated_videos_to_volume` in
`ssim_test.py` does a `shutil.copytree` of the whole `generated_videos/`
tree, picking up `.mp4`, `.pt`, and the `*_ssim.json` / `*_latent.json`
metric files alongside.
## References
- `fastvideo/tests/modal/ssim_test.py` — Modal orchestrator; see
`--sync-generated-to-volume`, `--generated-volume-subdir`,
`--skip-reference-download`, `--no-fail-fast`.
- `fastvideo/tests/ssim/reference_videos_cli.py` — `copy-local`, `upload`
(with `--model-id`, `--force`), `download`, `ensure` subcommands.
Extension allowlist is `REFERENCE_EXTENSIONS = VIDEO_EXTENSIONS +
LATENT_EXTENSIONS` (`.pt`).
- `fastvideo/tests/ssim/README.md` — reference layout, HF repo conventions.
- `fastvideo/tests/ssim/inference_similarity_utils.py` — pixel helpers
(`run_text_to_video_similarity_test`,
`run_image_to_video_similarity_test`, `build_init_kwargs`).
- `fastvideo/tests/ssim/latent_similarity_utils.py` — latent helper
(`run_text_to_latent_similarity_test`), slice spec dispatch
(`_extract_expected_slice`), reference schema
(`save_latent_reference` / `load_latent_reference`),
`LATENT_REFERENCE_FORMAT_VERSION`.
## Changelog
| Date | Change |
|------|--------|
| 2026-04-17 | Initial version (Modal sync-to-volume flow). |
| 2026-04-21 | Rewrite: single-test scope, explicit user-review pause, per-`model_id` upload, HF overwrite guard. Dropped `scripts/seed_ssim.sh`. |
| 2026-04-21 | Post-first-run fixes: `modal volume get` needs `--force` when parent exists; download tree has an extra `generated_videos/` level so `--generated-dir` must reflect it. |
| 2026-05-01 | Latent (`*.pt`) artefact support: artefact-type detection in step 1, type-aware verification (visual eyeball for mp4, numerics dump for pt) in step 4, FSDP+inference_mode failure-mode added, design notes for the unified Modal flow. Triggered by PR #1253 (LTX-2 latent migration + Stable Audio latent test). |
@@ -0,0 +1,49 @@
{
"benchmark_id": "wan-t2v-1.3b-2gpu",
"description": "Wan2.1 T2V 1.3B inference performance",
"model": {
"model_path": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
"model_short_name": "Wan2.1-T2V-1.3B"
},
"init_kwargs": {
"num_gpus": 2,
"flow_shift": 7.0,
"sp_size": 2,
"tp_size": 1,
"vae_sp": true,
"vae_tiling": true,
"text_encoder_precisions": ["fp32"]
},
"generation_kwargs": {
"height": 480,
"width": 832,
"num_frames": 45,
"num_inference_steps": 4,
"guidance_scale": 3,
"embedded_cfg_scale": 6,
"seed": 1024,
"fps": 24,
"neg_prompt": "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards"
},
"test_prompts": [
"Will Smith casually eats noodles, his relaxed demeanor contrasting with the energetic background of a bustling street food market. The scene captures a mix of humor and authenticity. Mid-shot framing, vibrant lighting."
],
"run_config": {
"num_warmup_runs": 2,
"num_measurement_runs": 5,
"required_gpus": 2
},
"thresholds": {
"L40S": {
"max_generation_time_s": 34.0,
"max_peak_memory_mb": 11000.0,
"max_text_encoder_time_s": 5.0,
"max_dit_time_s": 10.0,
"max_vae_decode_time_s": 10.0
},
"default": {
"max_generation_time_s": 120.0,
"max_peak_memory_mb": 30000.0
}
}
}
+385 -73
View File
@@ -2,28 +2,259 @@ env:
IMAGE_VERSION: "py3.12-latest"
BUILDKITE_CLEAN_CHECKOUT: true
notify:
- github_commit_status:
context: "fastcheck-passed"
if: build.env("TEST_SCOPE") == "fastcheck" || build.env("TEST_SCOPE") == null
- github_commit_status:
context: "full-suite-passed"
if: build.env("TEST_SCOPE") == "full"
- github_commit_status:
context: "direct-test-completed"
if: build.env("TEST_SCOPE") == "direct"
steps:
- label: "pre-commit"
command: ".buildkite/scripts/pre_commit.sh"
agents:
queue: "default"
# ============================================================
# Direct test: triggered by /test <name> slash command.
# Labels match fastcheck/full-suite counterparts so the GitHub
# check status overwrites the original failed check.
# Only ONE step executes per build (gated by TEST_TYPE).
# ============================================================
- wait
# --- Fastcheck-scope direct tests ---
- label: ":microscope: Encoder Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "encoder"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: VAE Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "vae"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Transformer Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "transformer"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Kernel Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "kernel_tests"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Unit Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "unit_test"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: DreamVerse App Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "dreamverse_app"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: "Trigger Tests"
plugins:
- monorepo-diff#v1.4.0:
diff: 'git fetch origin "$BUILDKITE_PULL_REQUEST_BASE_BRANCH" && git diff --name-only origin/"$BUILDKITE_PULL_REQUEST_BASE_BRANCH"...HEAD'
watch:
- path:
# --- Full-suite-scope direct tests ---
- label: ":bar_chart: SSIM Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "ssim"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "default"
- label: ":test_tube: LoRA Inference Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "inference_lora"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Training Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Distillation DMD Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "distillation_dmd"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Self-Forcing Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "self_forcing"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: LoRA Training Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training_lora"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Training Tests VSA"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training_vsa"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Inference Tests VMoBA"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "inference_vmoba"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Performance Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "performance"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: API Server Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "api_server"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Train Framework Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "train_framework"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Eval Metrics Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "eval"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
# ============================================================
# Fastcheck: Runs on every PR (~10-15 min parallel)
# Core component validation: encoders, VAEs, transformers,
# CUDA kernels, and unit tests.
# ============================================================
- label: "Trigger Fastcheck"
if: build.env("TEST_SCOPE") == "fastcheck" || build.env("TEST_SCOPE") == null
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
plugins:
- monorepo-diff#v1.4.0:
diff: 'git fetch origin "${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}" && git diff --name-only "origin/${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}...HEAD"'
watch:
- path:
- "fastvideo/models/encoders/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/encoders/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: "Encoder Tests"
label: ":microscope: Encoder Tests"
env:
- TEST_TYPE=encoder
agents:
@@ -33,10 +264,10 @@ steps:
- "fastvideo/models/loader/**"
- "fastvideo/tests/vaes/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: "VAE Tests"
label: ":microscope: VAE Tests"
env:
- TEST_TYPE=vae
agents:
@@ -48,23 +279,81 @@ steps:
- "fastvideo/layers/**"
- "fastvideo/attention/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: "Transformer Tests"
label: ":microscope: Transformer Tests"
env:
- TEST_TYPE=transformer
agents:
queue: "default"
- path:
- path:
- "fastvideo-kernel/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Kernel Tests"
env:
- TEST_TYPE=kernel_tests
agents:
queue: "default"
- path:
- "fastvideo/**"
- ".buildkite/**"
- ".github/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Unit Tests"
env:
- TEST_TYPE=unit_test
agents:
queue: "default"
- path:
- "apps/dreamverse/**"
- "pyproject.toml"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":microscope: DreamVerse App Tests"
env:
- TEST_TYPE=dreamverse_app
agents:
queue: "default"
# ============================================================
# Full Suite: Runs when TEST_SCOPE=full
# Triggered by adding the 'ready' label (via ci-trigger-full-suite.yml)
# or on-demand via /test full slash command.
# Includes integration tests, SSIM regression, training pipelines,
# and performance benchmarks.
# ============================================================
- label: "Trigger Full Suite"
if: build.env("TEST_SCOPE") == "full"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
plugins:
- monorepo-diff#v1.4.0:
diff: 'git fetch origin "${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}" && git diff --name-only "origin/${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}...HEAD"'
watch:
- path:
- "fastvideo/**/*.py"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: "SSIM Tests"
label: ":bar_chart: SSIM Tests"
env:
- TEST_TYPE=ssim
retry:
automatic:
- exit_status: 1
limit: 2
agents:
queue: "default"
- path:
@@ -74,10 +363,10 @@ steps:
- "fastvideo/pipelines/**"
- "fastvideo/layers/lora/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: "LoRA Inference Tests"
label: ":test_tube: LoRA Inference Tests"
env:
- TEST_TYPE=inference_lora
agents:
@@ -85,10 +374,10 @@ steps:
- path:
- "fastvideo/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: "Training Tests"
label: ":test_tube: Training Tests"
env:
- TEST_TYPE=training
agents:
@@ -96,10 +385,10 @@ steps:
- path:
- "fastvideo/training/*distillation_pipeline.py"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: "Distillation DMDTests"
label: ":test_tube: Distillation DMD Tests"
env:
- TEST_TYPE=distillation_dmd
agents:
@@ -108,10 +397,10 @@ steps:
- "fastvideo/training/*self_forcing_distillation_pipeline.py"
- "fastvideo/tests/training/self-forcing/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: "Self-Forcing Tests"
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Self-Forcing Tests"
env:
- TEST_TYPE=self_forcing
agents:
@@ -119,78 +408,101 @@ steps:
- path:
- "fastvideo/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: "LoRA Training Tests"
label: ":test_tube: LoRA Training Tests"
env:
- TEST_TYPE=training_lora
retry:
automatic:
- exit_status: 1
limit: 2
agents:
queue: "default"
- path:
- "fastvideo/**"
- "fastvideo-kernel/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: "Training Tests VSA"
label: ":test_tube: Training Tests VSA"
env:
- TEST_TYPE=training_vsa
retry:
automatic:
- exit_status: 1
limit: 2
agents:
queue: "default"
- path:
- "fastvideo/**"
- "fastvideo-kernel/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: "Inference Tests STA"
env:
- TEST_TYPE=inference_sta
agents:
queue: "default"
- path:
- "fastvideo-kernel/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: "Kernel Tests"
env:
- TEST_TYPE=kernel_tests
- path:
- "fastvideo-kernel/**"
- "fastvideo/attention/backends/vmoba.py"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: "Inference Tests VMoBA"
env:
label: ":test_tube: Inference Tests VMoBA"
env:
- TEST_TYPE=inference_vmoba
agents:
queue: "default"
- path:
- "fastvideo/**"
- "fastvideo/models/dits/**"
- "fastvideo/pipelines/**"
- "fastvideo/attention/**"
- "fastvideo/layers/**"
- "fastvideo/worker/**"
- "fastvideo/entrypoints/**"
- "fastvideo/tests/performance/**"
- ".buildkite/performance-benchmarks/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: "Unit Tests"
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Performance Tests"
env:
- TEST_TYPE=unit_test
- TEST_TYPE=performance
agents:
queue: "default"
- path:
- "fastvideo/entrypoints/openai/**"
- "fastvideo/entrypoints/cli/serve.py"
- "fastvideo/tests/entrypoints/test_openai_api_integration.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: API Server Tests"
env:
- TEST_TYPE=api_server
agents:
queue: "default"
- path:
- "fastvideo/train/**"
- "fastvideo/tests/train/models/**"
- "fastvideo/tests/train/fixtures/**"
- "fastvideo/models/dits/**"
- "fastvideo/models/loader/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Train Framework Tests"
env:
- TEST_TYPE=train_framework
agents:
queue: "default"
- path:
- "fastvideo/eval/**"
- "fastvideo/tests/eval/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Eval Metrics Tests"
env:
- TEST_TYPE=eval
agents:
queue: "default"
# - path:
# - "scripts/lora_extraction/**"
# - "pyproject.toml"
# - "docker/Dockerfile.python3.12"
# config:
# command: "timeout 90m .buildkite/scripts/pr_test.sh"
# label: "LoRA Extraction Tests"
# env:
# - TEST_TYPE=lora_extraction
# agents:
# queue: "default"
+127 -8
View File
@@ -15,8 +15,21 @@ log "Project root: $PROJECT_ROOT"
# Install Modal if not available
if ! python3 -m modal --version &> /dev/null; then
log "Modal not found, installing..."
python3 -m pip install modal
if ! command -v uv &> /dev/null; then
log "uv not found, bootstrapping..."
if ! curl -LsSf https://astral.sh/uv/install.sh | sh; then
log "Error: Failed to bootstrap uv via astral.sh installer."
exit 1
fi
export PATH="$HOME/.local/bin:$PATH"
if ! command -v uv &> /dev/null; then
log "Error: uv still not on PATH after bootstrap."
exit 1
fi
fi
# --break-system-packages preserves prior `pip install --user` semantics on PEP 668 agents.
uv pip install --system --break-system-packages modal
# Verify installation
if ! python3 -m modal --version &> /dev/null; then
log "Error: Failed to install modal. Please install it manually."
@@ -51,6 +64,7 @@ else
fi
MODAL_TEST_FILE="fastvideo/tests/modal/pr_test.py"
MODAL_SSIM_TEST_FILE="fastvideo/tests/modal/ssim_test.py"
if [ -z "${TEST_TYPE:-}" ]; then
log "Error: TEST_TYPE environment variable is not set"
@@ -58,7 +72,90 @@ if [ -z "${TEST_TYPE:-}" ]; then
fi
log "Test type: $TEST_TYPE"
MODAL_ENV="BUILDKITE_REPO=$BUILDKITE_REPO BUILDKITE_COMMIT=$BUILDKITE_COMMIT BUILDKITE_PULL_REQUEST=$BUILDKITE_PULL_REQUEST IMAGE_VERSION=$IMAGE_VERSION"
EFFECTIVE_PR=${BUILDKITE_PULL_REQUEST:-false}
if [ "$EFFECTIVE_PR" = "false" ] && [ -n "${PR_NUMBER:-}" ]; then
EFFECTIVE_PR=$PR_NUMBER
fi
MODAL_ENV="BUILDKITE_REPO=$BUILDKITE_REPO BUILDKITE_COMMIT=$BUILDKITE_COMMIT BUILDKITE_PULL_REQUEST=$EFFECTIVE_PR BUILDKITE_BRANCH=${BUILDKITE_BRANCH:-} TEST_SCOPE=${TEST_SCOPE:-} BUILDKITE_BUILD_URL=${BUILDKITE_BUILD_URL:-} BUILDKITE_BUILD_ID=${BUILDKITE_BUILD_ID:-} BUILDKITE_JOB_ID=${BUILDKITE_JOB_ID:-} IMAGE_VERSION=$IMAGE_VERSION"
POST_RUN_HOOK=""
upload_performance_artifacts() {
SHORT_SHA=${BUILDKITE_COMMIT:0:7}
LOCAL_DIR="downloaded_reports"
_download_reports() {
log "Downloading perf_reports/ from Modal Volume..."
mkdir -p "$LOCAL_DIR"
if ! modal volume get hf-model-weights "perf_reports/" "$LOCAL_DIR"; then
log "Error: Failed to download perf_reports/ from Modal Volume."
return 1
fi
}
_upload_dashboard() {
local target
target=$(find "$LOCAL_DIR" -name "dashboard_${SHORT_SHA}_*" | head -n 1)
log "TARGET dashboard: '$target'"
if [ -n "$target" ]; then
log "Found dashboard: $target. Uploading to Buildkite..."
buildkite-agent artifact upload "$target"
buildkite-agent annotate --style info --context "perf-dashboard" < "$target"
else
log "Warning: Could not find a dashboard file matching $SHORT_SHA"
fi
}
_upload_perf_summary() {
local target
target=$(find "$LOCAL_DIR" -name "perf_${SHORT_SHA}_*" | head -n 1)
log "TARGET perf summary: '$target'"
if [ -n "$target" ]; then
log "Found perf summary: $target. Uploading to Buildkite..."
buildkite-agent artifact upload "$target"
buildkite-agent annotate --style info --context "perf-summary" < "$target"
else
log "Warning: Could not find a perf summary file matching $SHORT_SHA"
fi
}
_upload_normalized_perf_results() {
local found=0
while IFS= read -r -d '' target; do
found=1
log "Found normalized performance result: $target. Uploading to Buildkite..."
buildkite-agent artifact upload "$target"
done < <(find "$LOCAL_DIR" -path "*/results/normalized_perf_*.json" -print0)
if [ "$found" -eq 0 ]; then
log "No normalized performance result artifacts found. This is expected when the rolling performance comparison did not run."
fi
}
_cleanup_modal_volume() {
log "Cleaning up perf_reports/ from Modal Volume..."
if modal volume rm hf-model-weights "perf_reports/" --recursive; then
log "Successfully deleted perf_reports/ from Modal Volume."
else
log "Warning: Failed to delete perf_reports/ from Modal Volume. Manual cleanup may be required."
fi
}
_cleanup_local() {
log "Cleaning up local download directory..."
rm -rf "$LOCAL_DIR"
}
# --- Main flow ---
_download_reports || { _cleanup_local; return 1; }
_upload_dashboard
_upload_perf_summary
_upload_normalized_perf_results
_cleanup_modal_volume
_cleanup_local
}
case "$TEST_TYPE" in
"encoder")
@@ -75,7 +172,7 @@ case "$TEST_TYPE" in
;;
"ssim")
log "Running SSIM tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_ssim_tests"
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_SSIM_TEST_FILE::run_ssim_tests"
;;
"training")
log "Running training tests..."
@@ -89,10 +186,6 @@ case "$TEST_TYPE" in
log "Running training VSA tests..."
MODAL_COMMAND="$MODAL_ENV WANDB_API_KEY=$WANDB_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_training_tests_VSA"
;;
"inference_sta")
log "Running inference STA tests..."
MODAL_COMMAND="$MODAL_ENV python3 -m modal run $MODAL_TEST_FILE::run_inference_tests_STA"
;;
"kernel_tests")
log "Running kernel tests..."
MODAL_COMMAND="$MODAL_ENV python3 -m modal run $MODAL_TEST_FILE::run_kernel_tests"
@@ -118,10 +211,31 @@ case "$TEST_TYPE" in
log "Running unit tests..."
MODAL_COMMAND="$MODAL_ENV python3 -m modal run $MODAL_TEST_FILE::run_unit_test"
;;
"dreamverse_app")
log "Running DreamVerse app tests..."
MODAL_COMMAND="$MODAL_ENV python3 -m modal run $MODAL_TEST_FILE::run_dreamverse_app_tests"
;;
"train_framework")
log "Running fastvideo.train framework tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_train_framework_tests"
;;
"eval")
log "Running eval metric tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_eval_tests"
;;
"lora_extraction")
log "Running LoRA extraction tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_lora_extraction_tests"
;;
"performance")
log "Running performance tests on Modal..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_performance_tests"
POST_RUN_HOOK="upload_performance_artifacts"
;;
"api_server")
log "Running API server integration tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_api_server_tests"
;;
*)
log "Error: Unknown test type: $TEST_TYPE"
exit 1
@@ -138,5 +252,10 @@ else
log "Error: Modal test failed with exit code: $TEST_EXIT_CODE"
fi
if [ -n "$POST_RUN_HOOK" ]; then
log "Executing post-run hook: $POST_RUN_HOOK"
"$POST_RUN_HOOK"
fi
log "=== Test execution completed with exit code: $TEST_EXIT_CODE ==="
exit $TEST_EXIT_CODE
+15 -2
View File
@@ -13,8 +13,21 @@ log "Project root: $PROJECT_ROOT"
if ! python3 -m pre_commit --version &> /dev/null; then
log "pre-commit not found, installing..."
python3 -m pip install --user pre-commit==4.0.1
if ! command -v uv &> /dev/null; then
log "uv not found, bootstrapping..."
if ! curl -LsSf https://astral.sh/uv/install.sh | sh; then
log "Error: Failed to bootstrap uv via astral.sh installer."
exit 1
fi
export PATH="$HOME/.local/bin:$PATH"
if ! command -v uv &> /dev/null; then
log "Error: uv still not on PATH after bootstrap."
exit 1
fi
fi
# --break-system-packages preserves prior `pip install --user` semantics on PEP 668 agents.
uv pip install --system --break-system-packages pre-commit==4.0.1
if ! python3 -m pre_commit --version &> /dev/null; then
log "Error: Failed to install pre-commit."
exit 1
+31
View File
@@ -0,0 +1,31 @@
# Build-context excludes: keep the context small and the `COPY . .` layer cache
# stable. Docker uploads everything here to the daemon and bakes it into a layer;
# without this, the 6.5 GB host .venv alone is shipped + cached on every build.
#
# IMPORTANT: do NOT ignore .git — fastvideo-kernel's build runs
# `git submodule update --init --recursive`, which needs the repo metadata.
# Virtualenvs — the image builds its own /opt/venv
.venv/
venv/
env/
# Python caches & build/test/lint artifacts
**/__pycache__/
*.py[cod]
*.egg-info/
.eggs/
.pytest_cache/
.mypy_cache/
.ruff_cache/
.cache/
# Local run outputs / logs (not needed in the image)
outputs/
wandb/
*.log
# Editor / OS cruft
.DS_Store
.idea/
.vscode/
+63
View File
@@ -0,0 +1,63 @@
<!--
PR TITLE: Must start with a type tag, e.g.:
[feat] Add new model [bugfix] Fix VAE tiling [refactor] Restructure pipeline
[perf] Optimize kernel [ci] Update tests [docs] Add guide
[misc] Cleanup configs [new-model] Port Flux2 [infra] Add trace hooks
[skill] Add agent skill
MERGE WORKFLOW:
1. Ensure pre-commit passes and you have at least 1 approval
2. Comment /merge (or add the "ready" label) to enter the Merge Queue
3. Full Test Suite runs automatically on a staging branch → auto-merge on success
ON-DEMAND TESTING (write access required):
/test full — Full Test Suite /test ssim — SSIM regression
/test training — Training pipeline /test encoder — Encoder tests
/test transformer — Transformer tests /test vae — VAE tests
/test kernel — CUDA kernel tests /test unit — Unit tests
See docs/contributing/pull_requests.md for all 17 test commands
-->
## Purpose
<!-- What does this PR do? Link the related issue if applicable. -->
Fixes #
## Changes
<!-- Describe your changes concisely. What approach did you take? -->
-
## Test Plan
<!-- How did you verify your changes? Paste exact commands and output. -->
```bash
# Commands you ran
```
## Test Results
<!-- Paste test output, before/after comparisons, or SSIM scores for model changes. -->
<details>
<summary>Test output</summary>
```
# Paste output here
```
</details>
## Checklist
- [ ] I ran `pre-commit run --all-files` and fixed all issues
- [ ] I added or updated tests for my changes
- [ ] I updated documentation if needed
- [ ] I considered GPU memory impact of my changes
**For model/pipeline changes, also check:**
- [ ] I verified SSIM regression tests pass
- [ ] I updated the support matrix if adding a new model
+324
View File
@@ -0,0 +1,324 @@
merge_protections:
- name: PR merge requirements
if:
- base = main
success_conditions:
- "title~=(?i)^\\[(feat|feature|bugfix|fix|refactor|perf|ci|doc|docs|misc|chore|kernel|new.?model|skill|skills|infra)\\]"
- "#approved-reviews-by>=1"
- check-success~=pre-commit
- check-success=fastcheck-passed
- check-success=full-suite-passed
pull_request_rules:
# ============================================================
# Type labels (from PR title prefix)
# ============================================================
- name: "label type: feat"
conditions:
- "title~=(?i)^\\[(feat|feature)\\]"
- -closed
actions:
label:
add: ["type: feat"]
- name: "label type: bugfix"
conditions:
- "title~=(?i)^\\[(bug)?fix\\]"
- -closed
actions:
label:
add: ["type: bugfix"]
- name: "label type: refactor"
conditions:
- "title~=(?i)^\\[refactor\\]"
- -closed
actions:
label:
add: ["type: refactor"]
- name: "label type: perf"
conditions:
- "title~=(?i)^\\[perf\\]"
- -closed
actions:
label:
add: ["type: perf"]
- name: "label type: ci"
conditions:
- "title~=(?i)^\\[ci\\]"
- -closed
actions:
label:
add: ["type: ci"]
- name: "label type: docs"
conditions:
- "title~=(?i)^\\[(doc|docs)\\]"
- -closed
actions:
label:
add: ["type: docs"]
- name: "label type: misc"
conditions:
- "title~=(?i)^\\[(misc|chore)\\]"
- -closed
actions:
label:
add: ["type: misc"]
- name: "label type: new-model"
conditions:
- "title~=(?i)^\\[new.?model\\]"
- -closed
actions:
label:
add: ["type: new-model"]
- name: "label type: infra"
conditions:
- "title~=(?i)^\\[infra\\]"
- -closed
actions:
label:
add: ["type: infra"]
- name: "label type: skill"
conditions:
- "title~=(?i)^\\[skills?\\]"
- -closed
actions:
label:
add: ["type: skill"]
# ============================================================
# Scope labels (from changed files)
# ============================================================
- name: "label scope: training"
conditions:
- or:
- files~=^fastvideo/train/
- files~=^fastvideo/training/
- files~=^fastvideo/distillation/
- files~=^examples/train/
- files~=^examples/training/
- files~=^examples/distill/
- -closed
actions:
label:
add: ["scope: training"]
- name: "label scope: inference"
conditions:
- or:
- files~=^fastvideo/pipelines/basic/
- files~=^fastvideo/pipelines/stages/
- files~=^fastvideo/pipelines/samplers/
- files~=^fastvideo/entrypoints/
- files~=^fastvideo/worker/
- files~=^fastvideo/api/sampling_param
- files~=^fastvideo/configs/pipelines/
- files~=^examples/inference/
- -closed
actions:
label:
add: ["scope: inference"]
- name: "label scope: attention"
conditions:
- files~=^fastvideo/attention/
- -closed
actions:
label:
add: ["scope: attention"]
- name: "label scope: kernel"
conditions:
- or:
- files~=^fastvideo-kernel/
- files~=^csrc/
- -closed
actions:
label:
add: ["scope: kernel"]
- name: "label scope: data"
conditions:
- or:
- files~=^fastvideo/dataset/
- files~=^fastvideo/pipelines/preprocess/
- files~=^examples/preprocessing/
- -closed
actions:
label:
add: ["scope: data"]
- name: "label scope: infra"
conditions:
- or:
- files~=^\.github/
- files~=^\.buildkite/
- files~=^fastvideo/tests/
- files~=^docker/
- -closed
actions:
label:
add: ["scope: infra"]
- name: "label scope: distributed"
conditions:
- files~=^fastvideo/distributed/
- -closed
actions:
label:
add: ["scope: distributed"]
- name: "label scope: docs"
conditions:
- files~=^docs/
- -closed
actions:
label:
add: ["scope: docs"]
- name: "label scope: studio"
conditions:
- files~=^apps/fastvideo_studio/
- -closed
actions:
label:
add: ["scope: studio"]
- name: "label scope: model"
conditions:
- or:
- files~=^fastvideo/models/
- files~=^fastvideo/layers/
- files~=^fastvideo/configs/models/
- -closed
actions:
label:
add: ["scope: model"]
# ============================================================
# Pre-commit failure help comment
# ============================================================
- name: comment on pre-commit failure
conditions:
- check-failure~=pre-commit
- -closed
actions:
comment:
message: |
## Pre-commit checks failed
Hi @{{author}}, the pre-commit checks have failed. To fix them locally:
```bash
# Install pre-commit if you haven't already
uv pip install pre-commit
pre-commit install
# Run all checks and auto-fix what's possible
pre-commit run --all-files
```
Common fixes:
- **yapf**: `yapf -i <file>` (formatting)
- **ruff**: `ruff check --fix <file>` (linting)
- **codespell**: `codespell --write-changes <file>` (spelling)
After fixing, commit and push the changes. The checks will re-run automatically.
For future commits, `pre-commit` will run automatically on changed files before each commit.
# ============================================================
# Merge conflict detection
# ============================================================
- name: label conflicting PRs
conditions:
- conflict
- -closed
- label!=stale
actions:
label:
add: [needs-rebase]
comment:
message: |
This PR has merge conflicts with the base branch. Please rebase:
```bash
git fetch origin main
git rebase origin/main
# Resolve any conflicts, then:
git push --force-with-lease
```
- name: remove conflict label when resolved
conditions:
- -conflict
- -closed
- label=needs-rebase
actions:
label:
remove: [needs-rebase]
# ============================================================
# Auto-merge
# ============================================================
- name: auto-merge when ready and all checks pass
conditions:
- label=ready
- "title~=(?i)^\\[(feat|feature|bugfix|fix|refactor|perf|ci|doc|docs|misc|chore|kernel|new.?model|skill|skills|infra)\\]"
- "#approved-reviews-by>=1"
- check-success~=pre-commit
- check-success=fastcheck-passed
- check-success=full-suite-passed
- -conflict
- -closed
- -draft
actions:
merge:
method: squash
# ============================================================
# PR title format help
# ============================================================
- name: comment on invalid PR title format
conditions:
- -closed
- -draft
- "-title~=(?i)^\\[(feat|feature|bugfix|fix|refactor|perf|ci|doc|docs|misc|chore|kernel|new.?model|skill|skills|infra)\\]"
actions:
comment:
message: |
## ⚠️ PR title format required
Your PR title must start with a type tag in brackets. Examples:
- `[feat] Add new model support`
- `[bugfix] Fix VAE tiling corruption`
- `[refactor] Restructure training pipeline`
- `[perf] Optimize attention kernel`
- `[ci] Update test infrastructure`
- `[infra] Add activation trace hooks`
- `[docs] Add inference guide`
- `[misc] Clean up configs`
- `[new-model] Port Flux2 to FastVideo`
- `[skill] Add add-model agent skill`
Valid tags: `feat`, `feature`, `bugfix`, `fix`, `refactor`, `perf`, `ci`, `infra`, `doc`, `docs`, `misc`, `chore`, `kernel`, `new-model`, `skill`, `skills`
Please update your PR title and the merge protection check will pass automatically.
merge_protections_settings:
reporting_method: check-runs
-249
View File
@@ -1,249 +0,0 @@
import argparse
import json
import os
import subprocess
import sys
import time
import requests
def parse_arguments():
"""Parse command line arguments"""
parser = argparse.ArgumentParser(description='Run tests on RunPod GPU')
parser.add_argument('--gpu-type', type=str, help='GPU type to use')
parser.add_argument('--gpu-count',
type=int,
help='Number of GPUs to use',
default=1)
parser.add_argument('--test-command', type=str, help='Test command to run')
parser.add_argument('--disk-size',
type=int,
default=20,
help='Container disk size in GB (default: 20)')
parser.add_argument('--volume-size',
type=int,
default=20,
help='Persistent volume size in GB (default: 20)')
parser.add_argument(
'--image',
type=str,
required=True,
help='Docker image to use')
return parser.parse_args()
args = parse_arguments()
API_KEY = os.environ['RUNPOD_API_KEY']
RUN_ID = os.environ['GITHUB_RUN_ID']
JOB_ID = os.environ['JOB_ID']
PODS_API = "https://rest.runpod.io/v1/pods"
HEADERS = {
"Content-Type": "application/json",
"Authorization": f"Bearer {API_KEY}"
}
def create_pod():
"""Create a RunPod instance"""
# Ensure image name is lowercase (Docker requirement)
image_name = args.image.lower()
print(f"Using specified image: {image_name}")
docker_start_cmd = [
"bash",
"-c",
"apt update;DEBIAN_FRONTEND=noninteractive apt-get install openssh-server -y;mkdir -p ~/.ssh;cd $_;chmod 700 ~/.ssh;echo \"$PUBLIC_KEY\" >> authorized_keys;chmod 700 authorized_keys;service ssh start;sleep infinity"
]
print(f"Creating RunPod instance with GPU: {args.gpu_type}...")
payload = {
"name": f"fastvideo-{JOB_ID}-{RUN_ID}",
"containerDiskInGb": args.disk_size,
"volumeInGb": args.volume_size,
"gpuTypeIds": [args.gpu_type],
"gpuCount": args.gpu_count,
"imageName": image_name,
"allowedCudaVersions": ["12.4"],
"dockerStartCmd": docker_start_cmd
}
response = requests.post(PODS_API, headers=HEADERS, json=payload)
response_data = response.json()
print(f"Response: {json.dumps(response_data, indent=2)}")
return response_data["id"]
def wait_for_pod(pod_id):
"""Wait for pod to be in RUNNING state and fully ready with SSH access"""
print("Waiting for RunPod to be ready...")
# First wait for RUNNING status
max_attempts = 10
attempts = 0
while attempts < max_attempts:
response = requests.get(f"{PODS_API}/{pod_id}", headers=HEADERS)
pod_data = response.json()
status = pod_data["desiredStatus"]
if status == "RUNNING":
print("RunPod is running! Now waiting for ports to be assigned...")
break
print(
f"Current status: {status}, waiting... (attempt {attempts+1}/{max_attempts})"
)
time.sleep(2)
attempts += 1
if attempts >= max_attempts:
raise TimeoutError(
"Timed out waiting for RunPod to reach RUNNING state")
# Wait for ports to be assigned
max_attempts = 50
attempts = 0
while attempts < max_attempts:
response = requests.get(f"{PODS_API}/{pod_id}", headers=HEADERS)
pod_data = response.json()
port_mappings = pod_data.get("portMappings")
if (port_mappings is not None and "22" in port_mappings
and pod_data.get("publicIp", "") != ""):
print("RunPod is ready with SSH access!")
print(f"SSH IP: {pod_data['publicIp']}")
print(f"SSH Port: {port_mappings['22']}")
break
print(
f"Waiting for SSH port and public IP to be available... (attempt {attempts+1}/{max_attempts})"
)
time.sleep(20)
attempts += 1
if attempts >= max_attempts:
raise TimeoutError("Timed out waiting for RunPod SSH access")
def execute_command(pod_id):
"""Execute command on the pod via SSH using system SSH client"""
print(f"Running command: {args.test_command}")
response = requests.get(f"{PODS_API}/{pod_id}", headers=HEADERS)
pod_data = response.json()
ssh_ip = pod_data["publicIp"]
ssh_port = pod_data["portMappings"]["22"]
# Copy the repository to the pod using scp
repo_dir = os.path.abspath(os.getcwd())
repo_name = os.path.basename(repo_dir)
print(f"Copying repository from {repo_dir} to RunPod...")
tar_command = [
"tar", "-czf", "/tmp/repo.tar.gz", "-C",
os.path.dirname(repo_dir), repo_name
]
subprocess.run(tar_command, check=True)
# Copy the tarball to the pod
scp_command = [
"scp", "-o", "StrictHostKeyChecking=no", "-o",
"UserKnownHostsFile=/dev/null", "-o", "ServerAliveInterval=60", "-o",
"ServerAliveCountMax=10", "-P",
str(ssh_port), "/tmp/repo.tar.gz", f"root@{ssh_ip}:/tmp/"
]
subprocess.run(scp_command, check=True)
# For custom image, we can use the pre-configured environment
setup_steps = [
"tar -xzf /tmp/repo.tar.gz --no-same-owner -C /workspace/",
f"cd /workspace/{repo_name}",
"source $HOME/.local/bin/env && source /opt/venv/bin/activate",
args.test_command
]
remote_command = " && ".join(setup_steps)
ssh_command = [
"ssh", "-o", "StrictHostKeyChecking=no", "-o",
"UserKnownHostsFile=/dev/null", "-o", "ServerAliveInterval=60", "-o",
"ServerAliveCountMax=10", "-p",
str(ssh_port), f"root@{ssh_ip}", remote_command
]
print(f"Connecting to {ssh_ip}:{ssh_port}...")
try:
process = subprocess.Popen(ssh_command,
stdout=subprocess.PIPE,
stderr=subprocess.STDOUT,
universal_newlines=True,
bufsize=0)
stdout_lines = []
print("Command output:")
for line in iter(process.stdout.readline, ''):
print(line.strip())
stdout_lines.append(line)
process.wait()
return_code = process.returncode
success = return_code == 0
stdout_str = "".join(stdout_lines)
if success:
print("Command executed successfully")
else:
print(f"Command failed with exit code {return_code}")
result = {
"success": success,
"return_code": return_code,
"stdout": stdout_str,
"stderr": ""
}
return result
except Exception as e:
print(f"Error executing SSH command: {str(e)}")
result = {"success": False, "error": str(e), "stdout": "", "stderr": ""}
return result
def terminate_pod(pod_id):
"""Terminate the pod"""
print("Terminating RunPod...")
requests.delete(f"{PODS_API}/{pod_id}", headers=HEADERS)
print(f"Terminated pod {pod_id}")
def main():
pod_id = None
try:
pod_id = create_pod()
wait_for_pod(pod_id)
result = execute_command(pod_id)
if result.get("error") is not None:
print(f"Error executing command: {result['error']}")
sys.exit(1)
if not result.get("success", False):
print(
"Tests failed - check the output above for details on which tests failed"
)
sys.exit(1)
finally:
if pod_id:
terminate_pod(pod_id)
if __name__ == "__main__":
main()
-90
View File
@@ -1,90 +0,0 @@
import json
import os
import sys
import uuid
import requests
API_KEY = os.environ['RUNPOD_API_KEY']
RUN_ID = os.environ.get('GITHUB_RUN_ID', str(uuid.uuid4()))
PODS_API = "https://rest.runpod.io/v1/pods"
HEADERS = {
"Content-Type": "application/json",
"Authorization": f"Bearer {API_KEY}"
}
def get_job_ids():
"""Parse job IDs from environment variable"""
job_ids_str = os.environ.get('JOB_IDS')
try:
job_ids = json.loads(job_ids_str)
if not isinstance(job_ids, list):
print("Error: JOB_IDS is not a list.")
sys.exit(1)
return job_ids
except json.JSONDecodeError as e:
print(f"Error parsing JOB_IDS: {e}")
sys.exit(1)
def cleanup_pods():
"""Find and terminate RunPod instances"""
print(f"Run ID: {RUN_ID}")
single_job_id = os.environ.get('JOB_ID')
if single_job_id:
job_ids = [single_job_id]
print(f"Job ID: {single_job_id}")
else:
job_ids = get_job_ids()
print(f"Job IDs: {job_ids}")
# Get all pods associated with RunPod API_KEY
try:
response = requests.get(PODS_API, headers=HEADERS)
response.raise_for_status()
pods = response.json()
except requests.exceptions.RequestException as e:
print(f"Error getting pods: {e}")
sys.exit(1)
# Find and terminate pods created by this workflow run
terminated_pods = []
for pod in pods:
pod_name = pod.get("name", "")
pod_id = pod.get("id")
# Check if this pod was created by one of our jobs
if any(f"{job_id}-{RUN_ID}" in pod_name for job_id in job_ids):
print(f"Found pod: {pod_id} ({pod_name})")
try:
print(f"Terminating pod {pod_id}...")
term_response = requests.delete(f"{PODS_API}/{pod_id}",
headers=HEADERS)
term_response.raise_for_status()
terminated_pods.append(pod_id)
print(f"Successfully terminated pod {pod_id}")
except requests.exceptions.RequestException as e:
print(f"Error terminating pod {pod_id}: {e}")
sys.exit(1)
if terminated_pods:
if single_job_id:
print(f"Terminated pod: {terminated_pods[0]}")
else:
print(f"Terminated {len(terminated_pods)} pods: {terminated_pods}")
else:
if single_job_id:
print(f"No pod found matching pattern: {single_job_id}-{RUN_ID}")
else:
print("No pods found to terminate.")
def main():
cleanup_pods()
if __name__ == "__main__":
main()
+200
View File
@@ -0,0 +1,200 @@
name: Build Image Template
on:
workflow_call:
inputs:
python_version:
required: true
type: string
dockerfile_path:
required: true
type: string
tag_suffix:
required: true
type: string
image_name:
required: false
type: string
default: fastvideo-dev
build_args:
required: false
type: string
default: ''
include_latest_tags:
required: false
type: boolean
default: true
mark_as_latest:
required: false
type: boolean
default: false
runner:
required: false
type: string
default: ubuntu-latest
architecture:
required: false
type: string
default: amd64
push_by_digest:
required: false
type: boolean
default: false
digest_artifact_name:
required: false
type: string
default: ''
jobs:
build-and-push:
runs-on: ${{ inputs.runner }}
permissions:
contents: read
packages: write
steps:
- name: Checkout code
uses: actions/checkout@v4
# The Docker context intentionally includes .git so the kernel build can
# initialize its pinned submodules. Do not copy the checkout token with it.
with:
persist-credentials: false
- name: Free up disk space
run: |
# Display initial space
echo "Initial disk space:"
df -h
# Remove large directories directly
sudo rm -rf /usr/share/dotnet
sudo rm -rf /usr/local/lib/android
sudo rm -rf /opt/ghc
sudo rm -rf /usr/local/share/boost
sudo rm -rf /usr/share/swift
sudo rm -rf /usr/local/lib/node_modules
sudo rm -rf /usr/local/share/powershell
sudo rm -rf /usr/share/rust
sudo rm -rf /usr/local/.ghcup
# Remove cached files
sudo rm -rf /var/lib/apt/lists/*
sudo rm -rf /var/cache/apt/archives/*
# Clean Docker
docker system prune -af --volumes
# Display available space after cleanup
echo "Disk space after cleanup:"
df -h
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Login to GitHub Container Registry
uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Normalize image reference
id: image
env:
IMAGE: ghcr.io/${{ github.repository }}/${{ inputs.image_name }}
run: echo "name=${IMAGE,,}" >> "${GITHUB_OUTPUT}"
- name: Prepare tags
id: prepare-tags
run: |
SHORT_SHA=$(echo ${{ github.sha }} | cut -c1-7)
TAGS="type=raw,value=${{ inputs.tag_suffix }}-sha-${SHORT_SHA}"
if [[ "${{ inputs.include_latest_tags }}" == "true" ]]; then
TAGS="type=raw,value=${{ inputs.tag_suffix }}-latest\n${TAGS}"
fi
# Tag the designated default variant as the global `latest` image
if [[ "${{ inputs.include_latest_tags }}" == "true" && "${{ inputs.mark_as_latest }}" == "true" ]]; then
TAGS="${TAGS}\ntype=raw,value=latest"
fi
{
echo "tags<<EOF"
echo -e "$TAGS"
echo "EOF"
} >> $GITHUB_OUTPUT
- name: Extract metadata for Docker
id: meta
uses: docker/metadata-action@v5
with:
images: ${{ steps.image.outputs.name }}
tags: ${{ steps.prepare-tags.outputs.tags }}
- name: Build and push Docker image
if: ${{ !inputs.push_by_digest }}
id: build-push
uses: docker/build-push-action@v6
with:
context: .
file: ${{ inputs.dockerfile_path }}
platforms: linux/${{ inputs.architecture }}
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
build-args: ${{ inputs.build_args }}
cache-from: type=gha,scope=${{ inputs.image_name }}-${{ inputs.tag_suffix }}-${{ inputs.architecture }}
cache-to: type=gha,mode=max,scope=${{ inputs.image_name }}-${{ inputs.tag_suffix }}-${{ inputs.architecture }}
# Multi-architecture callers publish immutable manifests by digest here,
# then create the shared user-facing tags in a single downstream job.
# This prevents architecture jobs from racing to replace the same tag.
- name: Build and push Docker image by digest
if: ${{ inputs.push_by_digest }}
id: build-push-digest
uses: docker/build-push-action@v6
with:
context: .
file: ${{ inputs.dockerfile_path }}
platforms: linux/${{ inputs.architecture }}
labels: ${{ steps.meta.outputs.labels }}
build-args: ${{ inputs.build_args }}
outputs: type=image,name=${{ steps.image.outputs.name }},push-by-digest=true,name-canonical=true,push=true
cache-from: type=gha,scope=${{ inputs.image_name }}-${{ inputs.tag_suffix }}-${{ inputs.architecture }}
cache-to: type=gha,mode=max,scope=${{ inputs.image_name }}-${{ inputs.tag_suffix }}-${{ inputs.architecture }}
- name: Export digest
if: ${{ inputs.push_by_digest }}
env:
DIGEST: ${{ steps.build-push-digest.outputs.digest }}
run: |
if [[ -z "${{ inputs.digest_artifact_name }}" ]]; then
echo "digest_artifact_name is required when push_by_digest is true" >&2
exit 1
fi
mkdir -p /tmp/digests
touch "/tmp/digests/${DIGEST#sha256:}"
- name: Upload digest
if: ${{ inputs.push_by_digest }}
uses: actions/upload-artifact@v4
with:
name: ${{ inputs.digest_artifact_name }}
path: /tmp/digests/*
if-no-files-found: error
retention-days: 1
- name: Success message
if: ${{ !inputs.push_by_digest }}
run: |
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ${{ steps.image.outputs.name }}:${{ inputs.tag_suffix }}-sha-${GITHUB_SHA::7}"
echo "To run tests with this image, manually trigger the 'Run Tests' workflow."
- name: Digest success message
if: ${{ inputs.push_by_digest }}
env:
DIGEST: ${{ steps.build-push-digest.outputs.digest }}
run: |
echo "✅ Python ${{ inputs.python_version }} linux/${{ inputs.architecture }} image pushed as ${DIGEST}"
-106
View File
@@ -1,106 +0,0 @@
name: Build Image Template
on:
workflow_call:
inputs:
python_version:
required: true
type: string
dockerfile_path:
required: true
type: string
tag_suffix:
required: true
type: string
jobs:
build-and-push:
runs-on: ubuntu-latest
permissions:
contents: read
packages: write
steps:
- name: Checkout code
uses: actions/checkout@v4
- name: Free up disk space
run: |
# Display initial space
echo "Initial disk space:"
df -h
# Remove large directories directly
sudo rm -rf /usr/share/dotnet
sudo rm -rf /usr/local/lib/android
sudo rm -rf /opt/ghc
sudo rm -rf /usr/local/share/boost
sudo rm -rf /usr/share/swift
sudo rm -rf /usr/local/lib/node_modules
sudo rm -rf /usr/local/share/powershell
sudo rm -rf /usr/share/rust
sudo rm -rf /usr/local/.ghcup
# Remove cached files
sudo rm -rf /var/lib/apt/lists/*
sudo rm -rf /var/cache/apt/archives/*
# Clean Docker
docker system prune -af --volumes
# Display available space after cleanup
echo "Disk space after cleanup:"
df -h
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Login to GitHub Container Registry
uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Prepare tags
id: prepare-tags
run: |
SHORT_SHA=$(echo ${{ github.sha }} | cut -c1-7)
TAGS="type=raw,value=${{ inputs.tag_suffix }}-latest"
TAGS="${TAGS}\ntype=raw,value=${{ inputs.tag_suffix }}-sha-${SHORT_SHA}"
# Set Python 3.10 as the default image
if [[ "${{ inputs.python_version }}" == "3.10" ]]; then
TAGS="${TAGS}\ntype=raw,value=latest"
fi
{
echo "tags<<EOF"
echo -e "$TAGS"
echo "EOF"
} >> $GITHUB_OUTPUT
- name: Extract metadata for Docker
id: meta
uses: docker/metadata-action@v5
with:
images: ghcr.io/${{ github.repository }}/fastvideo-dev
tags: ${{ steps.prepare-tags.outputs.tags }}
- name: Build and push Docker image
id: build-push
uses: docker/build-push-action@v6
with:
context: .
file: ${{ inputs.dockerfile_path }}
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
cache-from: type=gha
cache-to: type=gha,mode=max
- name: Success message
run: |
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ghcr.io/${{ github.repository }}/fastvideo-dev:${{ inputs.tag_suffix }}-latest"
echo "To run tests with this image, manually trigger the 'Run Tests' workflow."
-67
View File
@@ -1,67 +0,0 @@
name: Build and Push Docker Images
on:
workflow_dispatch:
inputs:
python_3_10:
description: 'Build Python 3.10 image'
required: false
default: false
type: boolean
python_3_11:
description: 'Build Python 3.11 image'
required: false
default: false
type: boolean
python_3_12:
description: 'Build Python 3.12 image'
required: false
default: false
type: boolean
python_3_12_cuda_12_9:
description: 'Build Python 3.12 image Cuda 12.9'
required: false
default: false
type: boolean
permissions:
contents: read
packages: write
jobs:
build-python-3-10:
if: ${{ github.event.inputs.python_3_10 == 'true' }}
uses: ./.github/workflows/build-image-template.yml
with:
python_version: '3.10'
dockerfile_path: docker/Dockerfile.python3.10
tag_suffix: py3.10
secrets: inherit
build-python-3-11:
if: ${{ github.event.inputs.python_3_11 == 'true' }}
uses: ./.github/workflows/build-image-template.yml
with:
python_version: '3.11'
dockerfile_path: docker/Dockerfile.python3.11
tag_suffix: py3.11
secrets: inherit
build-python-3-12:
if: ${{ github.event.inputs.python_3_12 == 'true' }}
uses: ./.github/workflows/build-image-template.yml
with:
python_version: '3.12'
dockerfile_path: docker/Dockerfile.python3.12
tag_suffix: py3.12
secrets: inherit
build-python-3-12-cuda-12-9:
if: ${{ github.event.inputs.python_3_12_cuda_12_9 == 'true' }}
uses: ./.github/workflows/build-image-template.yml
with:
python_version: '3.12'
dockerfile_path: docker/Dockerfile.python3.12.cuda12.9.1
tag_suffix: py3.12-cuda12.9.1
secrets: inherit
+80
View File
@@ -0,0 +1,80 @@
name: Aggregate Test Status
on:
status:
permissions:
statuses: write
jobs:
aggregate:
if: >-
github.event.context == 'direct-test-completed'
&& github.event.state == 'success'
runs-on: ubuntu-latest
steps:
- name: Check and update aggregate status
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
script: |
const sha = context.payload.sha;
const { data } = await github.rest.repos.getCombinedStatusForRef({
owner: context.repo.owner,
repo: context.repo.repo,
ref: sha,
per_page: 100,
});
const bkStatuses = data.statuses.filter(
s => s.context.startsWith('buildkite/ci/')
);
const FASTCHECK_PREFIX = 'buildkite/ci/microscope-';
const FULL_SUITE_PREFIXES = [
'buildkite/ci/test-tube-',
'buildkite/ci/bar-chart-',
];
const fastcheck = bkStatuses.filter(
s => s.context.startsWith(FASTCHECK_PREFIX)
);
const fullSuite = bkStatuses.filter(
s => FULL_SUITE_PREFIXES.some(p => s.context.startsWith(p))
);
if (
fastcheck.length > 0
&& fastcheck.every(s => s.state === 'success')
) {
core.info(
`All ${fastcheck.length} fastcheck tests passed — updating fastcheck-passed`
);
await github.rest.repos.createCommitStatus({
owner: context.repo.owner,
repo: context.repo.repo,
sha,
state: 'success',
context: 'fastcheck-passed',
description:
`All ${fastcheck.length} fastcheck tests passed`,
});
}
if (
fullSuite.length > 0
&& fullSuite.every(s => s.state === 'success')
) {
core.info(
`All ${fullSuite.length} full suite tests passed — updating full-suite-passed`
);
await github.rest.repos.createCommitStatus({
owner: context.repo.owner,
repo: context.repo.repo,
sha,
state: 'success',
context: 'full-suite-passed',
description:
`All ${fullSuite.length} full suite tests passed`,
});
}
+32
View File
@@ -0,0 +1,32 @@
name: pre-commit
on:
pull_request:
branches: [main]
workflow_call:
inputs:
ref:
description: 'Git ref to checkout (defaults to github.ref)'
required: false
type: string
permissions:
contents: read
jobs:
pre-commit:
if: github.event_name == 'workflow_call' || github.event.pull_request.draft != true
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
with:
ref: ${{ inputs.ref || '' }}
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: echo "::add-matcher::.github/workflows/matchers/actionlint.json"
- run: echo "::add-matcher::.github/workflows/matchers/mypy.json"
- run: echo "::add-matcher::.github/workflows/matchers/ruff.json"
- uses: pre-commit/action@v3.0.1
with:
extra_args: --all-files --hook-stage manual
+272
View File
@@ -0,0 +1,272 @@
name: Slash Commands
on:
issue_comment:
types: [created]
permissions:
contents: read
pull-requests: write
statuses: write
jobs:
handle-merge:
if: >-
github.event.issue.pull_request != null
&& startsWith(github.event.comment.body, '/merge')
runs-on: ubuntu-latest
steps:
- name: Check write permission
id: perm
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
script: |
const { data: perm } = await github.rest.repos.getCollaboratorPermissionLevel({
owner: context.repo.owner,
repo: context.repo.repo,
username: context.payload.comment.user.login,
});
const hasWrite = ['admin', 'write'].includes(perm.permission);
if (!hasWrite) {
core.setFailed(`User ${context.payload.comment.user.login} lacks write permission (has: ${perm.permission}).`);
}
core.setOutput('has_write', String(hasWrite));
- name: Add ready label and react
id: label
if: steps.perm.outputs.has_write == 'true'
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
script: |
const owner = context.repo.owner;
const repo = context.repo.repo;
const prNumber = context.payload.issue.number;
try { await github.rest.issues.removeLabel({ owner, repo, issue_number: prNumber, name: 'ready' }); } catch {}
await github.rest.issues.addLabels({ owner, repo, issue_number: prNumber, labels: ['ready'] });
await github.rest.reactions.createForIssueComment({
owner, repo,
comment_id: context.payload.comment.id,
content: 'rocket',
});
const { data: pr } = await github.rest.pulls.get({ owner, repo, pull_number: prNumber });
core.setOutput('pr_sha', pr.head.sha);
core.setOutput('pr_branch', pr.head.ref);
core.setOutput('pr_number', String(prNumber));
- name: Trigger Full Suite
if: steps.perm.outputs.has_write == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_SHA: ${{ steps.label.outputs.pr_sha }}
PR_BRANCH: ${{ steps.label.outputs.pr_branch }}
PR_NUMBER: ${{ steps.label.outputs.pr_number }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
-H "Content-Type: application/json" \
--data-raw "$(jq -n \
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER} (via /merge)" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
branch: $branch,
message: $message,
ignore_pipeline_branch_filters: true,
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: "full",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring)
}
}')"
parse-command:
if: >-
github.event.issue.pull_request != null
&& startsWith(github.event.comment.body, '/test')
runs-on: ubuntu-latest
outputs:
test_type: ${{ steps.parse.outputs.test_type }}
test_scope: ${{ steps.parse.outputs.test_scope }}
full_suite: ${{ steps.parse.outputs.full_suite }}
pr_sha: ${{ steps.pr.outputs.sha }}
pr_branch: ${{ steps.pr.outputs.branch }}
has_write: ${{ steps.perm.outputs.has_write }}
steps:
- name: Check write permission
id: perm
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
script: |
const { data: perm } = await github.rest.repos.getCollaboratorPermissionLevel({
owner: context.repo.owner,
repo: context.repo.repo,
username: context.payload.comment.user.login,
});
const hasWrite = ['admin', 'write'].includes(perm.permission);
core.setOutput('has_write', String(hasWrite));
if (!hasWrite) {
core.info(`User ${context.payload.comment.user.login} lacks write permission — ignoring.`);
}
- name: Parse /test command
id: parse
if: steps.perm.outputs.has_write == 'true'
shell: bash
env:
COMMENT: ${{ github.event.comment.body }}
run: |
set -euo pipefail
TEST_NAME=$(echo "$COMMENT" | grep -oP '(?<=/test\s)\S+' | head -1 || true)
VALID="encoder vae transformer kernel unit dreamverse ssim training lora-inference lora-training distillation self-forcing vsa vmoba performance api train-framework eval full fastcheck pre-commit"
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
echo "Unknown test: '$TEST_NAME'. Valid: $VALID"
exit 1
fi
declare -A MAP=(
[encoder]=encoder [vae]=vae [transformer]=transformer
[kernel]=kernel_tests [unit]=unit_test [dreamverse]=dreamverse_app
[ssim]=ssim [training]=training
[lora-inference]=inference_lora [lora-training]=training_lora
[distillation]=distillation_dmd [self-forcing]=self_forcing
[vsa]=training_vsa [vmoba]=inference_vmoba
[performance]=performance [api]=api_server
[train-framework]=train_framework [eval]=eval
)
if [ "$TEST_NAME" = "full" ]; then
{
echo "test_type=all"
echo "test_scope=full"
echo "full_suite=true"
} >> "$GITHUB_OUTPUT"
elif [ "$TEST_NAME" = "fastcheck" ]; then
{
echo "test_type=fastcheck"
echo "test_scope=fastcheck"
echo "full_suite=false"
} >> "$GITHUB_OUTPUT"
elif [ "$TEST_NAME" = "pre-commit" ]; then
{
echo "test_type="
echo "test_scope=precommit"
echo "full_suite=false"
} >> "$GITHUB_OUTPUT"
else
{
echo "test_type=${MAP[$TEST_NAME]}"
echo "test_scope=direct"
echo "full_suite=false"
} >> "$GITHUB_OUTPUT"
fi
- name: Get PR details
id: pr
if: steps.perm.outputs.has_write == 'true'
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
script: |
const { data: pr } = await github.rest.pulls.get({
owner: context.repo.owner,
repo: context.repo.repo,
pull_number: context.payload.issue.number,
});
core.setOutput('sha', pr.head.sha);
core.setOutput('branch', pr.head.ref);
- name: React to comment
if: steps.perm.outputs.has_write == 'true'
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
script: |
await github.rest.reactions.createForIssueComment({
owner: context.repo.owner,
repo: context.repo.repo,
comment_id: context.payload.comment.id,
content: 'rocket',
});
pre-commit:
needs: parse-command
if: >-
needs.parse-command.outputs.has_write == 'true'
&& needs.parse-command.outputs.test_scope == 'precommit'
uses: ./.github/workflows/ci-precommit.yml
with:
ref: refs/pull/${{ github.event.issue.number }}/merge
post-precommit-status:
needs: [parse-command, pre-commit]
if: always() && needs.parse-command.outputs.test_scope == 'precommit'
runs-on: ubuntu-latest
steps:
- uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
env:
PR_SHA: ${{ needs.parse-command.outputs.pr_sha }}
RESULT: ${{ needs.pre-commit.result }}
with:
script: |
const state = process.env.RESULT === 'success' ? 'success' : 'failure';
await github.rest.repos.createCommitStatus({
owner: context.repo.owner,
repo: context.repo.repo,
sha: process.env.PR_SHA,
state,
context: 'pre-commit',
description: `Triggered via /test pre-commit (${state})`,
});
trigger-buildkite:
needs: parse-command
if: >-
needs.parse-command.outputs.has_write == 'true'
&& needs.parse-command.outputs.test_type != ''
runs-on: ubuntu-latest
steps:
- name: Trigger Buildkite
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_SHA: ${{ needs.parse-command.outputs.pr_sha }}
PR_BRANCH: ${{ needs.parse-command.outputs.pr_branch }}
PR_NUMBER: ${{ github.event.issue.number }}
TEST_SCOPE: ${{ needs.parse-command.outputs.test_scope }}
FULL_SUITE: ${{ needs.parse-command.outputs.full_suite }}
TEST_TYPE: ${{ needs.parse-command.outputs.test_type }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
-H "Content-Type: application/json" \
--data-raw "$(jq -n \
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "/test ${TEST_TYPE} on PR #${PR_NUMBER}" \
--argjson pr_id "$PR_NUMBER" \
--arg test_scope "$TEST_SCOPE" \
--arg full_suite "$FULL_SUITE" \
--arg test_type "$TEST_TYPE" \
--arg pr_number "$PR_NUMBER" \
'{
commit: $commit,
branch: $branch,
message: $message,
ignore_pipeline_branch_filters: true,
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: $test_scope,
FULL_SUITE: $full_suite,
TEST_TYPE: $test_type,
PR_NUMBER: $pr_number
}
}')"
@@ -0,0 +1,83 @@
name: Trigger Full Suite
on:
pull_request_target:
types: [labeled, synchronize]
permissions:
contents: read
pull-requests: read
concurrency:
group: full-suite-${{ github.event.pull_request.number }}
cancel-in-progress: false
jobs:
trigger:
if: >-
(github.event.action == 'labeled' && github.event.label.name == 'ready')
|| github.event.action == 'synchronize'
runs-on: ubuntu-latest
steps:
- name: Check ready label
id: check
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
script: |
const { data: pr } = await github.rest.pulls.get({
owner: context.repo.owner,
repo: context.repo.repo,
pull_number: context.payload.pull_request.number,
});
const hasReady = pr.labels.some(l => l.name === 'ready');
core.setOutput('has_ready', String(hasReady));
if (!hasReady) core.info('No ready label — skipping Full Suite trigger.');
- name: Cancel previous Buildkite builds
if: steps.check.outputs.has_ready == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
run: |
# Find running builds for this branch with TEST_SCOPE=full and cancel them
builds=$(curl -sS -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds?branch=${PR_BRANCH}&state=running,scheduled" \
| jq -r '.[] | select(try (.env.TEST_SCOPE == "full") catch false) | .number')
for build_num in $builds; do
echo "Cancelling Buildkite build #$build_num"
curl -sS -X PUT -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds/${build_num}/cancel"
done
- name: Trigger Buildkite Full Suite
if: steps.check.outputs.has_ready == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_SHA: ${{ github.event.pull_request.head.sha }}
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
PR_NUMBER: ${{ github.event.pull_request.number }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
-H "Content-Type: application/json" \
--data-raw "$(jq -n \
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER}" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
branch: $branch,
message: $message,
ignore_pipeline_branch_filters: true,
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: "full",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring)
}
}')"
@@ -0,0 +1,65 @@
name: Auto-Label Issues
on:
issues:
types: [opened, edited]
permissions:
issues: write
jobs:
label-issues:
if: github.repository == 'hao-ai-lab/FastVideo'
runs-on: ubuntu-latest
steps:
- name: Label by keywords
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
script: |
const title = context.payload.issue.title.toLowerCase();
const body = (context.payload.issue.body || '').toLowerCase();
const text = title + ' ' + body;
const labels = [];
const rules = [
// scope labels (shared with PR labeling via Mergify)
// Mapping: label → repo directories
// scope: training → fastvideo/train/, fastvideo/training/, fastvideo/distillation/
// scope: inference → fastvideo/pipelines/, fastvideo/entrypoints/, fastvideo/worker/
// scope: attention → fastvideo/attention/
// scope: kernel → fastvideo-kernel/, csrc/
// scope: model → fastvideo/models/, fastvideo/layers/, fastvideo/configs/models/
// scope: data → fastvideo/dataset/, fastvideo/pipelines/preprocess/
// scope: distributed → fastvideo/distributed/
// scope: docs → docs/
{ keywords: ['training', 'finetune', 'fine-tune', 'lora', 'fsdp', 'distill'], label: 'scope: training' },
{ keywords: ['inference', 'generate', 'pipeline', 'slow', 'latency'], label: 'scope: inference' },
{ keywords: ['attention', 'vsa', 'flash', 'sta', 'vmoba', 'sparse attn'], label: 'scope: attention' },
{ keywords: ['kernel', 'csrc', 'cuda kernel', 'thunderkittens'], label: 'scope: kernel' },
{ keywords: ['wan', 'hunyuan', 'mochi', 'ltx', 'cogvideo', 'flux', 'sd3', 'cosmos'], label: 'scope: model' },
{ keywords: ['dataset', 'dataloader', 'preprocessing', 'preprocess'], label: 'scope: data' },
{ keywords: ['distributed', 'sequence parallel', 'fsdp', 'tensor parallel', 'multi-node', 'multi-gpu'], label: 'scope: distributed' },
{ keywords: ['docs', 'documentation', 'tutorial', 'example'], label: 'scope: docs' },
// issue-only labels (cross-module, no single repo directory)
{ keywords: ['install', 'setup', 'pip', 'cuda', 'uv ', 'import error', 'modulenotfound'], label: 'installation' },
{ keywords: ['memory', 'oom', 'out of memory', 'gpu memory', 'vram'], label: 'performance' },
{ keywords: ['windows', 'macos', 'mac os', 'apple', 'mps', 'rocm', 'amd', 'npu'], label: 'platform' },
];
for (const rule of rules) {
if (rule.keywords.some(kw => text.includes(kw))) {
labels.push(rule.label);
}
}
if (labels.length > 0) {
await github.rest.issues.addLabels({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: context.payload.issue.number,
labels: labels,
});
console.log(`Added labels: ${labels.join(', ')}`);
} else {
console.log('No keyword matches found');
}
+51
View File
@@ -0,0 +1,51 @@
name: Close Stale Issues and PRs
on:
schedule:
# Daily at 1:30 AM UTC
- cron: '30 1 * * *'
jobs:
stale:
if: github.repository == 'hao-ai-lab/FastVideo'
permissions:
issues: write
pull-requests: write
actions: write
runs-on: ubuntu-latest
steps:
- uses: actions/stale@997185467fa4f803885201cee163a9f38240193d # v10.1.1
with:
operations-per-run: 500
exempt-draft-pr: true
exempt-issue-labels: 'keep-open,pinned,security,Bug,RFC'
exempt-pr-labels: 'keep-open,pinned'
labels-to-add-when-unstale: 'unstale'
labels-to-remove-when-stale: 'unstale'
days-before-issue-stale: 90
days-before-issue-close: 30
stale-issue-label: 'stale'
stale-issue-message: >
This issue has been automatically marked as stale because it has not
had any activity within 90 days. It will be automatically closed if
no further activity occurs within 30 days. Leave a comment if you
feel this issue should remain open. Thank you!
close-issue-message: >
This issue has been automatically closed due to inactivity. Please
feel free to reopen if you feel it is still relevant. Thank you!
days-before-pr-stale: 60
days-before-pr-close: 14
stale-pr-label: 'stale'
stale-pr-message: >
This pull request has been automatically marked as stale because it
has not had any activity within 60 days. It will be automatically
closed if no further activity occurs within 14 days. Leave a comment
if you feel this pull request should remain open. Thank you!
close-pr-message: >
This pull request has been automatically closed due to inactivity.
Please feel free to reopen if you intend to continue working on it.
Thank you!
+56
View File
@@ -0,0 +1,56 @@
name: Welcome First-Time Contributors
on:
issues:
types: [opened]
pull_request_target:
types: [opened]
permissions:
issues: write
pull-requests: write
jobs:
welcome:
if: github.repository == 'hao-ai-lab/FastVideo'
runs-on: ubuntu-latest
steps:
- uses: actions/first-interaction@34f15e814fe48ac9312ccf29db4e74fa767cbab7 # v1.3.0
with:
repo-token: ${{ secrets.GITHUB_TOKEN }}
issue-message: |
Welcome to FastVideo! Thanks for opening your first issue.
To help us investigate, please include:
- **FastVideo version**: `pip show fastvideo`
- **GPU**: `nvidia-smi` output (GPU model, driver, CUDA version)
- **Python version**: `python --version`
- **OS**: e.g., Ubuntu 22.04
If this is a bug, a minimal reproduction script helps us fix it faster.
Useful links:
- [Documentation](https://hao-ai-lab.github.io/FastVideo)
- [Contributing Guide](https://hao-ai-lab.github.io/FastVideo/contributing/overview/)
- [Slack](https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ)
pr-message: |
Welcome to FastVideo! Thanks for your first pull request.
**How our CI works:**
PRs run a two-tier CI system:
1. **Pre-commit** — formatting (yapf), linting (ruff), type checking (mypy). Runs immediately on every PR.
2. **Fastcheck** — core GPU tests (encoders, VAEs, transformers, kernels, unit tests). Runs automatically via Buildkite on relevant file changes (~10-15 min).
3. **Full Suite** — integration tests, training pipelines, SSIM regression. Runs only when a reviewer adds the `ready` label.
**Before your PR is reviewed:**
- [ ] `pre-commit run --all-files` passes locally
- [ ] You've added or updated tests for your changes
- [ ] The PR description explains what and why
If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and Full Suite results appear in the Checks section below.
**Useful links:**
- [Contributing Guide](https://hao-ai-lab.github.io/FastVideo/contributing/overview/)
- [Development Roadmap](https://github.com/hao-ai-lab/FastVideo/issues/899)
- [Slack](https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ)
+198
View File
@@ -0,0 +1,198 @@
name: Build and Push Docker Images
on:
workflow_dispatch:
inputs:
build_cuda_matrix:
description: 'Build multi-arch CUDA images (12.6.3 default; 13.0.0 explicitly tagged)'
required: false
default: true
type: boolean
build_dreamverse_matrix:
description: 'Build the amd64 Dreamverse matrix (backend + UI x CUDA 12.6.3/13.0.0)'
required: false
default: false
type: boolean
permissions:
contents: read
packages: write
jobs:
# CUDA matrix: Python 3.12 x {12.6.3, 13.0.0} x {amd64, arm64}. Each architecture
# builds natively and pushes only by digest; publish-cuda-manifests is the sole
# owner of the shared tags. CUDA 12.6 arm64 targets Hopper (sm_90a / GH200 class),
# because CUDA 12.6 cannot compile sm_121. CUDA 13 arm64 targets GB10 (sm_121).
# 12.6.3/cu126 owns the default `py3.12`/`latest` tags and keeps its versioned
# aliases; 13.0.0/cu130 is published under explicit versioned tags. Flash-attn
# 2.8.3 comes from the architecture-specific prebuilt releases.
build-cuda-images:
if: ${{ github.event.inputs.build_cuda_matrix == 'true' }}
strategy:
fail-fast: false
matrix:
cuda:
- version: '12.6.3'
suffix: '-cuda12.6.3'
torch_backend: 'cu126'
fa_tag: 'cu126torch2.12'
cmake_build_parallel_level: '4'
torch_cuda_arch_list:
amd64: '9.0a'
arm64: '9.0a'
- version: '13.0.0'
suffix: '-cuda13.0.0'
torch_backend: 'cu130'
fa_tag: 'cu130torch2.12'
cmake_build_parallel_level: '1'
torch_cuda_arch_list:
amd64: '9.0a'
arm64: '12.1'
architecture:
- name: amd64
runner: ubuntu-latest
- name: arm64
runner: ubuntu-24.04-arm
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.12'
dockerfile_path: docker/Dockerfile
tag_suffix: py3.12${{ matrix.cuda.suffix }}
runner: ${{ matrix.architecture.runner }}
architecture: ${{ matrix.architecture.name }}
push_by_digest: true
digest_artifact_name: fastvideo-dev-cuda${{ matrix.cuda.version }}-${{ matrix.architecture.name }}
build_args: |
PYTHON_VERSION=3.12
CUDA_VERSION=${{ matrix.cuda.version }}
UV_TORCH_BACKEND=${{ matrix.cuda.torch_backend }}
TORCH_CUDA_ARCH_LIST=${{ matrix.cuda.torch_cuda_arch_list[matrix.architecture.name] }}
CMAKE_BUILD_PARALLEL_LEVEL=${{ matrix.cuda.cmake_build_parallel_level }}
FLASH_ATTN_WHEEL_TAG=${{ matrix.cuda.fa_tag }}
FLASH_ATTN_WHEEL_RELEASE=https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.9.17
FLASH_ATTN_WHEEL_RELEASE_ARM64=https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.9.22
secrets: inherit
publish-cuda-manifests:
# !cancelled(): a failed sibling build leg must not skip the manifests for a
# CUDA lane whose own digests all exist; the digest-count check below fails
# the incomplete lane loudly instead.
if: ${{ !cancelled() && github.event.inputs.build_cuda_matrix == 'true' }}
needs: build-cuda-images
runs-on: ubuntu-latest
permissions:
contents: read
packages: write
strategy:
fail-fast: false
matrix:
cuda:
- version: '12.6.3'
suffix: '-cuda12.6.3'
is_default: true
- version: '13.0.0'
suffix: '-cuda13.0.0'
is_default: false
steps:
- name: Download architecture digests
uses: actions/download-artifact@v4
with:
path: /tmp/digests
pattern: fastvideo-dev-cuda${{ matrix.cuda.version }}-*
merge-multiple: true
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Login to GitHub Container Registry
uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Normalize image reference
id: image
env:
IMAGE: ghcr.io/${{ github.repository }}/fastvideo-dev
run: echo "name=${IMAGE,,}" >> "${GITHUB_OUTPUT}"
- name: Prepare tags
id: prepare-tags
env:
VERSIONED_TAG_PREFIX: py3.12${{ matrix.cuda.suffix }}
PUBLISH_DEFAULT_TAGS: ${{ matrix.cuda.is_default }}
run: |
SHORT_SHA="${GITHUB_SHA::7}"
TAGS="type=raw,value=${VERSIONED_TAG_PREFIX}-latest\ntype=raw,value=${VERSIONED_TAG_PREFIX}-sha-${SHORT_SHA}"
if [[ "${PUBLISH_DEFAULT_TAGS}" == "true" ]]; then
TAGS="type=raw,value=latest\ntype=raw,value=py3.12-latest\ntype=raw,value=py3.12-sha-${SHORT_SHA}\n${TAGS}"
fi
{
echo "tags<<EOF"
echo -e "${TAGS}"
echo "EOF"
} >> "${GITHUB_OUTPUT}"
- name: Extract metadata for Docker
id: meta
uses: docker/metadata-action@v5
with:
images: ${{ steps.image.outputs.name }}
tags: ${{ steps.prepare-tags.outputs.tags }}
- name: Create multi-architecture manifests
env:
IMAGE: ${{ steps.image.outputs.name }}
run: |
mapfile -t DIGESTS < <(find /tmp/digests -maxdepth 1 -type f -printf '%f\n' | sort)
if [[ "${#DIGESTS[@]}" -ne 2 ]]; then
echo "Expected exactly two architecture digests, found ${#DIGESTS[@]}" >&2
exit 1
fi
mapfile -t TAGS < <(jq -r '.tags[]' <<< "${DOCKER_METADATA_OUTPUT_JSON}")
TAG_ARGS=()
for TAG in "${TAGS[@]}"; do
TAG_ARGS+=(--tag "${TAG}")
done
IMAGE_REFS=()
for DIGEST in "${DIGESTS[@]}"; do
IMAGE_REFS+=("${IMAGE}@sha256:${DIGEST}")
done
docker buildx imagetools create "${TAG_ARGS[@]}" "${IMAGE_REFS[@]}"
docker buildx imagetools inspect "${TAGS[0]}"
# Dreamverse matrix: {backend, UI} x {12.6.3, 13.0.0}, Python 3.12. Torch backend
# matches the base CUDA (cu126 / cu130). Keep these images amd64-only until the
# required FA4 dependency stack is available and validated on arm64.
build-dreamverse:
if: ${{ github.event.inputs.build_dreamverse_matrix == 'true' }}
strategy:
fail-fast: false
matrix:
cuda:
- version: '12.6.3'
torch_backend: 'cu126'
- version: '13.0.0'
torch_backend: 'cu130'
variant:
- name: backend
ui: '0'
- name: ui
ui: '1'
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.12'
dockerfile_path: apps/dreamverse/docker/Dockerfile
tag_suffix: dreamverse-${{ matrix.variant.name }}-cuda${{ matrix.cuda.version }}
image_name: dreamverse
build_args: |
CUDA_VERSION=${{ matrix.cuda.version }}
UV_TORCH_BACKEND=${{ matrix.cuda.torch_backend }}
BUILD_DREAMVERSE_UI=${{ matrix.variant.ui }}
include_latest_tags: false
secrets: inherit
@@ -5,16 +5,22 @@ on:
branches: [ main ]
paths:
- 'docs/**'
- 'examples/**'
- 'mkdocs.yml'
- 'requirements-mkdocs.in'
- 'requirements-mkdocs.txt'
- '.github/workflows/docs.yml'
- 'scripts/check_docs_links.py'
- '.github/workflows/infra-docs.yml'
pull_request:
branches: [ main ]
paths:
- 'docs/**'
- 'examples/**'
- 'mkdocs.yml'
- 'requirements-mkdocs.in'
- 'requirements-mkdocs.txt'
- '.github/workflows/docs.yml'
- 'scripts/check_docs_links.py'
- '.github/workflows/infra-docs.yml'
permissions:
contents: read
@@ -31,16 +37,19 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Setup Python
uses: actions/setup-python@v5
with:
python-version: '3.12'
- name: Install uv
uses: astral-sh/setup-uv@v3
- name: Install dependencies
run: |
python -m pip install --upgrade pip
pip install -r requirements-mkdocs.txt
run: uv pip install --system -r requirements-mkdocs.txt
- name: Setup Pages
uses: actions/configure-pages@v4
@@ -48,6 +57,9 @@ jobs:
- name: Build documentation
run: mkdocs build
- name: Check docs links
run: python scripts/check_docs_links.py
- name: Upload artifact
uses: actions/upload-pages-artifact@v3
with:
@@ -63,4 +75,4 @@ jobs:
steps:
- name: Deploy to GitHub Pages
id: deployment
uses: actions/deploy-pages@v4
uses: actions/deploy-pages@v4
+17
View File
@@ -0,0 +1,17 @@
{
"problemMatcher": [
{
"owner": "ruff",
"pattern": [
{
"regexp": "^(.+):(\\d+):(\\d+): (\\w+) (.+)$",
"file": 1,
"line": 2,
"column": 3,
"code": 4,
"message": 5
}
]
}
]
}
-401
View File
@@ -1,401 +0,0 @@
name: PR Test
on:
push:
branches: [main]
paths:
- "fastvideo/**/*.py"
- ".github/workflows/pr-test.yml"
pull_request:
branches: [main]
types: [opened, ready_for_review, synchronize, reopened]
paths:
- "fastvideo/**/*.py"
- ".github/workflows/pr-test.yml"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "csrc/**"
workflow_dispatch:
inputs:
run_encoder_test:
description: "Run encoder-test"
required: false
default: false
type: boolean
run_vae_test:
description: "Run vae-test"
required: false
default: false
type: boolean
run_transformer_test:
description: "Run transformer-test"
required: false
default: false
type: boolean
run_ssim_test:
description: "Run ssim-test"
required: false
default: false
type: boolean
run_training_test:
description: "Run training-test"
required: false
default: false
type: boolean
run_training_test_VSA:
description: "Run training-test-VSA"
required: false
default: false
type: boolean
run_inference_test_STA:
description: "Run inference-test-STA"
required: false
default: false
type: boolean
run_precision_test_STA:
description: "Run precision-test-STA"
required: false
default: false
type: boolean
run_precision_test_VSA:
description: "Run precision-test-VSA"
required: false
default: false
type: boolean
run_unit_test:
description: "Run unit-test"
required: false
default: false
type: boolean
env:
PYTHONUNBUFFERED: "1"
concurrency:
group: pr-test-${{ github.ref }}
cancel-in-progress: true
jobs:
pre-commit:
uses: ./.github/workflows/pre-commit.yml
change-filter:
runs-on: ubuntu-latest
needs: pre-commit
if: ${{ github.event.pull_request.draft == false || github.event_name == 'workflow_dispatch' }}
outputs:
encoder-test: ${{ steps.filter.outputs.encoder-test }}
vae-test: ${{ steps.filter.outputs.vae-test }}
transformer-test: ${{ steps.filter.outputs.transformer-test }}
training-test: ${{ steps.filter.outputs.training-test }}
training-test-VSA: ${{ steps.filter.outputs.training-test-VSA }}
inference-test-STA: ${{ steps.filter.outputs.inference-test-STA }}
precision-test-STA: ${{ steps.filter.outputs.precision-test-STA }}
precision-test-VSA: ${{ steps.filter.outputs.precision-test-VSA }}
unit-test: ${{ steps.filter.outputs.unit-test }}
steps:
- uses: actions/checkout@v4
- uses: dorny/paths-filter@v3
id: filter
with:
filters: |
# Define reusable path patterns
common-paths: &common-paths
- 'pyproject.toml'
- 'docker/Dockerfile.python3.10'
- 'docker/Dockerfile.python3.11'
- 'docker/Dockerfile.python3.12'
sta-kernel-paths: &sta-kernel-paths
- 'csrc/attn/sliding_tile_attn/**'
- 'csrc/attn/sliding_tile_attn/tk/**'
- 'csrc/attn/sliding_tile_attn/setup.py'
- 'csrc/attn/sliding_tile_attn/config_sta.py'
- 'csrc/attn/sliding_tile_attn/st_attn.cpp'
vsa-kernel-paths: &vsa-kernel-paths
- 'csrc/attn/video_sparse_attn/**'
- 'csrc/attn/video_sparse_attn/tk/**'
- 'csrc/attn/video_sparse_attn/setup.py'
- 'csrc/attn/video_sparse_attn/config_vsa.py'
- 'csrc/attn/video_sparse_attn/vsa.cpp'
vsa-paths: &vsa-paths
- 'fastvideo/**'
- *common-paths
- *vsa-kernel-paths
# Actual tests
encoder-test:
- 'fastvideo/models/encoders/**'
- 'fastvideo/models/loader/**'
- 'fastvideo/tests/encoders/**'
- *common-paths
vae-test:
- 'fastvideo/models/vaes/**'
- 'fastvideo/models/loader/**'
- 'fastvideo/tests/vaes/**'
- *common-paths
transformer-test:
- 'fastvideo/models/dits/**'
- 'fastvideo/models/loader/**'
- 'fastvideo/tests/transformers/**'
- 'fastvideo/layers/**'
- 'fastvideo/attention/**'
- *common-paths
training-test:
- 'fastvideo/**'
- *common-paths
training-test-VSA:
- 'fastvideo/**'
- *common-paths
- *vsa-kernel-paths
inference-test-STA:
- 'fastvideo/**'
- *common-paths
- *sta-kernel-paths
precision-test-STA:
- *common-paths
- *sta-kernel-paths
precision-test-VSA:
- *common-paths
- *vsa-kernel-paths
unit-test:
- 'fastvideo/**'
- *common-paths
encoder-test:
needs: change-filter
if: >-
(github.event_name != 'workflow_dispatch' && needs.change-filter.outputs.encoder-test == 'true') ||
(github.event_name == 'workflow_dispatch' && github.event.inputs.run_encoder_test == 'true')
uses: ./.github/workflows/runpod-test.yml
with:
job_id: "encoder-test"
gpu_type: "NVIDIA A40"
gpu_count: 1
volume_size: 100
image: "ghcr.io/${{ github.repository }}/fastvideo-dev:py3.12-latest"
test_command: "uv pip install -e .[test] && pytest ./fastvideo/tests/encoders -s"
timeout_minutes: 30
secrets:
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
vae-test:
needs: change-filter
if: >-
(github.event_name != 'workflow_dispatch' && needs.change-filter.outputs.vae-test == 'true') ||
(github.event_name == 'workflow_dispatch' && github.event.inputs.run_vae_test == 'true')
uses: ./.github/workflows/runpod-test.yml
with:
job_id: "vae-test"
gpu_type: "NVIDIA A40"
gpu_count: 1
volume_size: 100
image: "ghcr.io/${{ github.repository }}/fastvideo-dev:py3.12-latest"
test_command: "uv pip install -e .[test] && pytest ./fastvideo/tests/vaes -s"
timeout_minutes: 30
secrets:
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
transformer-test:
needs: change-filter
if: >-
(github.event_name != 'workflow_dispatch' && needs.change-filter.outputs.transformer-test == 'true') ||
(github.event_name == 'workflow_dispatch' && github.event.inputs.run_transformer_test == 'true')
uses: ./.github/workflows/runpod-test.yml
with:
job_id: "transformer-test"
gpu_type: "NVIDIA L40S"
gpu_count: 1
volume_size: 100
image: "ghcr.io/${{ github.repository }}/fastvideo-dev:py3.12-latest"
test_command: "uv pip install -e .[test] && pytest ./fastvideo/tests/transformers -s"
timeout_minutes: 30
secrets:
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
ssim-test:
needs: change-filter
if: >-
github.event_name != 'workflow_dispatch' || (github.event_name == 'workflow_dispatch' && github.event.inputs.run_ssim_test == 'true')
strategy:
fail-fast: false
matrix:
python-version: [
# {version: "3.10", tag: "latest"},
# {version: "3.11", tag: "py3.11-latest"},
{version: "3.12", tag: "py3.12-latest"}
]
uses: ./.github/workflows/runpod-test.yml
with:
job_id: "ssim-test-py${{ matrix.python-version.version }}"
gpu_type: "NVIDIA A40"
gpu_count: 2
volume_size: 200
disk_size: 200
image: "ghcr.io/${{ github.repository }}/fastvideo-dev:${{ matrix.python-version.tag }}"
test_command: "uv pip install -e .[test] && pytest ./fastvideo/tests/ssim -vs"
timeout_minutes: 60
secrets:
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
training-test:
needs: change-filter
if: >-
(github.event_name != 'workflow_dispatch' && needs.change-filter.outputs.training-test == 'true') ||
(github.event_name == 'workflow_dispatch' && github.event.inputs.run_training_test == 'true')
uses: ./.github/workflows/runpod-test.yml
with:
job_id: "training-test"
gpu_type: "NVIDIA A40"
gpu_count: 4
volume_size: 100
disk_size: 100
image: "ghcr.io/${{ github.repository }}/fastvideo-dev:py3.12-latest"
test_command: "wandb login $WANDB_API_KEY && uv pip install -e .[test] && pytest ./fastvideo/tests/training/Vanilla -srP"
timeout_minutes: 30
secrets:
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
WANDB_API_KEY: ${{ secrets.WANDB_API_KEY }}
training-test-VSA:
needs: change-filter
if: >-
(github.event_name != 'workflow_dispatch' && needs.change-filter.outputs.training-test-VSA == 'true') ||
(github.event_name == 'workflow_dispatch' && github.event.inputs.run_training_test_VSA == 'true')
uses: ./.github/workflows/runpod-test.yml
with:
job_id: "training-test-VSA"
gpu_type: "NVIDIA H100 NVL"
gpu_count: 2
volume_size: 100
disk_size: 100
image: "ghcr.io/${{ github.repository }}/fastvideo-dev:py3.12-latest"
test_command: "wandb login $WANDB_API_KEY && uv pip install -e .[test] && pytest ./fastvideo/tests/training/VSA -srP"
timeout_minutes: 30
secrets:
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
WANDB_API_KEY: ${{ secrets.WANDB_API_KEY }}
inference-test-STA:
needs: change-filter
if: >-
(github.event_name != 'workflow_dispatch' && needs.change-filter.outputs.inference-test-STA == 'true') ||
(github.event_name == 'workflow_dispatch' && github.event.inputs.run_inference_test_STA == 'true')
uses: ./.github/workflows/runpod-test.yml
with:
job_id: "inference-test-STA"
gpu_type: "NVIDIA H100 NVL"
gpu_count: 2
volume_size: 100
disk_size: 100
image: "ghcr.io/${{ github.repository }}/fastvideo-dev:py3.12-latest"
test_command: "uv pip install -e .[test] && pytest ./fastvideo/tests/inference/STA -srP"
timeout_minutes: 30
secrets:
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
precision-test-STA:
needs: change-filter
if: >-
(github.event_name != 'workflow_dispatch' && needs.change-filter.outputs.precision-test-STA == 'true') ||
(github.event_name == 'workflow_dispatch' && github.event.inputs.run_precision_test_STA == 'true')
uses: ./.github/workflows/runpod-test.yml
with:
job_id: "precision-test-STA"
gpu_type: "NVIDIA H100 NVL"
gpu_count: 1
volume_size: 100
disk_size: 100
image: "ghcr.io/${{ github.repository }}/fastvideo-dev:py3.12-latest"
test_command: "uv pip install -e .[test] && python csrc/attn/tests/test_sta.py"
timeout_minutes: 30
secrets:
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
precision-test-VSA:
needs: change-filter
if: >-
(github.event_name != 'workflow_dispatch' && needs.change-filter.outputs.precision-test-VSA == 'true') ||
(github.event_name == 'workflow_dispatch' && github.event.inputs.run_precision_test_VSA == 'true')
uses: ./.github/workflows/runpod-test.yml
with:
job_id: "precision-test-VSA"
gpu_type: "NVIDIA H100 NVL"
gpu_count: 1
volume_size: 100
disk_size: 100
image: "ghcr.io/${{ github.repository }}/fastvideo-dev:py3.12-latest"
test_command: "uv pip install -e .[test] && python csrc/attn/tests/test_vsa.py"
timeout_minutes: 30
secrets:
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
unit-test:
needs: change-filter
if: >-
(github.event_name != 'workflow_dispatch' && needs.change-filter.outputs.unit-test == 'true') ||
(github.event_name == 'workflow_dispatch' && github.event.inputs.run_unit_test == 'true')
uses: ./.github/workflows/runpod-test.yml
with:
job_id: "unit-test"
gpu_type: "NVIDIA L40S"
gpu_count: 1
volume_size: 100
disk_size: 100
image: "ghcr.io/${{ github.repository }}/fastvideo-dev:py3.12-latest"
test_command: "uv pip install -e .[test] && pytest ./fastvideo/dataset/ -vs && pytest ./fastvideo/workflow/ -vs && pytest ./fastvideo/entrypoints/ -vs"
timeout_minutes: 30
secrets:
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
# nightly-test:
# if: >-
# (github.event_name == 'workflow_dispatch' && github.event.inputs.run_nightly_test == 'true')
# uses: ./.github/workflows/runpod-test.yml
# with:
# job_id: "nightly-test"
# gpu_type: "NVIDIA A40"
# gpu_count: 4
# volume_size: 100
# disk_size: 100
# image: "ghcr.io/${{ github.repository }}/fastvideo-dev:py3.12-latest"
# test_command: "wandb login $WANDB_API_KEY && uv pip install -e .[test] && pytest ./fastvideo/tests/nightly/test_e2e_overfit_single_sample.py -vs"
# timeout_minutes: 30
# secrets:
# RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
# RUNPOD_PRIVATE_KEY: ${{ secrets.RUNPOD_PRIVATE_KEY }}
# WANDB_API_KEY: ${{ secrets.WANDB_API_KEY }}
runpod-cleanup:
# Add other jobs to this list as you create them
needs: [encoder-test, vae-test, transformer-test, ssim-test, training-test, training-test-VSA, inference-test-STA, precision-test-STA, precision-test-VSA]
if: ${{ always() && ((github.event_name != 'workflow_dispatch' && github.event.pull_request.draft == false) || github.event_name == 'workflow_dispatch') }}
runs-on: ubuntu-latest
steps:
- name: Checkout code
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.10"
- name: Install dependencies
run: pip install requests
- name: Cleanup all RunPod instances
env:
JOB_IDS: '["encoder-test", "vae-test", "transformer-test", "ssim-test-py3.10", "ssim-test-py3.11", "ssim-test-py3.12", "training-test", "training-test-VSA", "inference-test-STA", "precision-test-STA", "precision-test-VSA"]'
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
GITHUB_RUN_ID: ${{ github.run_id }}
run: python .github/scripts/runpod_cleanup.py
-18
View File
@@ -1,18 +0,0 @@
name: pre-commit
on:
workflow_call:
jobs:
pre-commit:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: echo "::add-matcher::.github/workflows/matchers/actionlint.json"
- run: echo "::add-matcher::.github/workflows/matchers/mypy.json"
- uses: pre-commit/action@v3.0.1
with:
extra_args: --all-files --hook-stage manual
@@ -56,10 +56,11 @@ jobs:
with:
python-version: '3.10'
- name: Install uv
uses: astral-sh/setup-uv@v3
- name: Install build dependencies
run: |
python -m pip install --upgrade pip
pip install build twine wheel
run: uv pip install --system build twine wheel
- name: Build package
run: |
@@ -36,37 +36,50 @@ jobs:
if [ "$NEW_VERSION" != "$OLD_VERSION" ]; then
echo "Version changed from $OLD_VERSION to $NEW_VERSION"
echo "changed=true" >> $GITHUB_OUTPUT
echo "new-version=$NEW_VERSION" >> $GITHUB_OUTPUT
echo "changed=true" >> "$GITHUB_OUTPUT"
echo "new-version=$NEW_VERSION" >> "$GITHUB_OUTPUT"
else
echo "Version did not change"
echo "changed=false" >> $GITHUB_OUTPUT
echo "changed=false" >> "$GITHUB_OUTPUT"
fi
build_wheels:
name: Build Wheel
needs: check-version-change
if: ${{ needs.check-version-change.outputs.version-changed == 'true' || github.event_name == 'workflow_dispatch' }}
runs-on: ${{ matrix.os }}
runs-on: ${{ matrix.platform.os }}
strategy:
fail-fast: false
matrix:
os: [ubuntu-22.04]
python-version: ['3.10', '3.11', '3.12']
python-version: ['3.12']
torch-cuda:
# - torch-version: '2.5.1'
# cuda-version: '12.4.1'
# torch-cuda-short: 'cu124'
# - torch-version: '2.6.0'
# cuda-version: '12.6.3'
# torch-cuda-short: 'cu126'
# - torch-version: '2.7.1'
# cuda-version: '12.8.0'
# torch-cuda-short: 'cu128'
- torch-version: '2.9.1'
cuda-version: '12.8.0'
torch-cuda-short: 'cu128'
# torch 2.12 dropped cu128; cu130 (CUDA 13) is the default wheel, cu126 covers older drivers.
- torch-version: '2.12.0'
cuda-version: '12.6.3'
torch-cuda-short: 'cu126'
- torch-version: '2.12.0'
cuda-version: '13.0.0'
torch-cuda-short: 'cu130'
platform:
# x86_64 builds the full cu126 + cu130 set (cu130 ships the consumer
# Blackwell sm_120a FP4 kernels).
- os: ubuntu-22.04
arch: x86_64
wheel-plat: manylinux_2_35_x86_64
# aarch64 is Blackwell (GB200 sm_100a + DGX Spark / consumer sm_120a), not
# Hopper, and Blackwell needs CUDA >= 12.8 — so only the cu130 leg applies.
# Added via include so x86 keeps cu126 + cu130 while aarch64 stays cu130-only.
include:
- python-version: '3.12'
torch-cuda:
torch-version: '2.12.0'
cuda-version: '13.0.0'
torch-cuda-short: 'cu130'
platform:
os: ubuntu-22.04-arm
arch: aarch64
wheel-plat: manylinux_2_35_aarch64
steps:
- name: Free up disk space
@@ -101,7 +114,7 @@ jobs:
python-version: ${{ matrix.python-version }}
- name: Install CUDA ${{ matrix.torch-cuda.cuda-version }}
uses: Jimver/cuda-toolkit@v0.2.21
uses: Jimver/cuda-toolkit@v0.2.35
id: cuda-toolkit
with:
cuda: ${{ matrix.torch-cuda.cuda-version }}
@@ -128,11 +141,13 @@ jobs:
clang-11 --version
nvcc --version
- name: Install uv
uses: astral-sh/setup-uv@v3
- name: Install PyTorch ${{ matrix.torch-cuda.torch-version }}+cu${{ matrix.torch-cuda.cuda-version }}
run: |
pip install --upgrade pip
pip install typing-extensions==4.12.2
pip install --no-cache-dir torch==${{ matrix.torch-cuda.torch-version }} --index-url https://download.pytorch.org/whl/${{matrix.torch-cuda.torch-cuda-short}}
uv pip install --system typing-extensions==4.12.2
uv pip install --system --no-cache-dir torch==${{ matrix.torch-cuda.torch-version }} --index-url https://download.pytorch.org/whl/${{matrix.torch-cuda.torch-cuda-short}}
nvcc --version
python --version
python -c "import torch; print('PyTorch:', torch.__version__)"
@@ -142,20 +157,44 @@ jobs:
- name: Build wheel
run: |
export PYTHONPATH=$GITHUB_WORKSPACE:$PYTHONPATH
pip install setuptools ninja packaging wheel triton scikit-build-core cmake build
uv pip install --system setuptools ninja packaging wheel triton scikit-build-core cmake build
cd fastvideo-kernel
git submodule update --init --recursive # Ensure ThunderKittens submodule is initialized
# Release builds are produced on GPU-less runners, so force-enable TK and target Hopper.
export TORCH_CUDA_ARCH_LIST="9.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=ON -DCMAKE_CUDA_ARCHITECTURES=90a"
# Release builds run on GPU-less runners, so set kernels + arch explicitly:
# * aarch64 = Blackwell (GB200 sm_100a + DGX Spark/consumer sm_120a), NOT
# Hopper, so TK (sm_90a wgmma) is OFF. The C++ FP4 (attn_qat_infer, SM120)
# covers sm_120a; turbodiffusion covers sm_100a+sm_120a. The sm_100 FP4
# forward is the FA4 CuTe DSL path in the fastvideo package (PR #1221),
# JIT-compiled at runtime — not built into this wheel.
# * x86_64 cu130 = Hopper TK + consumer Blackwell sm_120a FP4.
# * x86_64 cu126 = Hopper TK only (older drivers; CUDA < 12.8 has no FP4).
# The per-arch split in CMakeLists pins the FP4 targets to sm_120a and builds
# the main extension for the full arch list. CMAKE_BUILD_PARALLEL_LEVEL caps
# Ninja so heavy CUTLASS/TK template TUs don't OOM the 16 GB runner (exit 143).
if [ "${{ matrix.platform.arch }}" = "aarch64" ]; then
export TORCH_CUDA_ARCH_LIST="10.0a;12.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=OFF -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON"
export CMAKE_BUILD_PARALLEL_LEVEL=1
elif [ "${{ matrix.torch-cuda.torch-cuda-short }}" = "cu130" ]; then
export TORCH_CUDA_ARCH_LIST="9.0a;12.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=ON -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON -DCMAKE_CUDA_ARCHITECTURES=90a"
# A single FP4 TU (attn_qat_infer) can use ~8-12 GB on its own, so serialize.
export CMAKE_BUILD_PARALLEL_LEVEL=1
else
export TORCH_CUDA_ARCH_LIST="9.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=ON -DCMAKE_CUDA_ARCHITECTURES=90a"
# Hopper-only: the two TK TUs are the heavy ones (single-arch) and fit side
# by side; -j4 overlaps the light TUs without OOMing (~12-14 GB peak).
export CMAKE_BUILD_PARALLEL_LEVEL=4
fi
# Build standard wheel (no local version suffix) for PyPI
python -m build --wheel --outdir dist
# Fix the wheel to be manylinux compliant
pip install auditwheel
uv pip install --system auditwheel
# Point auditwheel at torch libs, but do not vendor them into the wheel.
TORCH_LIB_DIR=$(python - <<'PY'
import os
@@ -165,8 +204,8 @@ jobs:
PY
)
export LD_LIBRARY_PATH="${TORCH_LIB_DIR}:${LD_LIBRARY_PATH}"
# Target manylinux_2_35 (Ubuntu 22.04 native)
auditwheel repair dist/*.whl --plat manylinux_2_35_x86_64 -w fixed_dist \
# Target manylinux_2_35 (Ubuntu 22.04 native), per-arch plat tag.
auditwheel repair dist/*.whl --plat ${{ matrix.platform.wheel-plat }} -w fixed_dist \
--exclude libtorch_cuda.so \
--exclude libtorch_cpu.so \
--exclude libtorch.so \
@@ -178,11 +217,11 @@ jobs:
mv fixed_dist/*.whl dist/
- name: Upload wheel artifact
# Only upload if it's the "main" CUDA version we want on PyPI
# We upload all to artifacts for inspection/GH releases, but give them distinct artifact names
# Upload every matrix leg as a distinct artifact (for inspection / GitHub releases).
# The publish job below selects which CUDA build is pushed to PyPI.
uses: actions/upload-artifact@v4
with:
name: fastvideo_kernel-py${{ matrix.python-version }}-${{ matrix.torch-cuda.torch-cuda-short }}-torch${{ matrix.torch-cuda.torch-version }}
name: fastvideo_kernel-py${{ matrix.python-version }}-${{ matrix.torch-cuda.torch-cuda-short }}-${{ matrix.platform.arch }}-torch${{ matrix.torch-cuda.torch-version }}
path: fastvideo-kernel/dist/*.whl
retention-days: 90
@@ -199,19 +238,28 @@ jobs:
- uses: actions/setup-python@v5
with:
python-version: '3.10'
python-version: '3.12'
- name: Download PyPI wheels
# Publish the cu130 (CUDA 13) wheels to PyPI for both architectures:
# x86_64 — Hopper sm_90a TK + consumer Blackwell sm_120a FP4
# aarch64 — Blackwell: turbodiffusion (sm_100a/sm_120a) + C++ FP4 (sm_120a);
# no TK (Hopper). sm_100 FP4 forward is the FA4 CuTe DSL path in the
# fastvideo package (#1221), shipped/JIT separately.
# The x86_64 cu126 wheel stays available as a build artifact / GitHub-release asset.
uses: actions/download-artifact@v4
with:
path: fastvideo-kernel/dist/
pattern: 'fastvideo_kernel-py*'
pattern: 'fastvideo_kernel-py*-cu130-*'
merge-multiple: true
- name: Install uv
uses: astral-sh/setup-uv@v3
- name: Build source distribution
run: |
pip install build scikit-build-core cmake ninja
uv pip install --system build scikit-build-core cmake ninja
cd fastvideo-kernel
# We don't need full CUDA/Torch to just package the source (sdist)
python -m build --sdist --outdir dist
-94
View File
@@ -1,94 +0,0 @@
name: RunPod Test
on:
workflow_call:
inputs:
job_id:
required: true
type: string
description: "Unique identifier for this test job"
gpu_type:
required: true
type: string
description: "GPU type to use (e.g. NVIDIA A40, NVIDIA L40S)"
gpu_count:
required: true
type: number
description: "Number of GPUs to use"
volume_size:
required: false
type: number
default: 20
description: "Volume size in GB"
disk_size:
required: false
type: number
default: 20
description: "Disk size in GB"
image:
required: true
type: string
description: "Docker image to use"
test_command:
required: true
type: string
description: "Command to run tests"
timeout_minutes:
required: false
type: number
default: 30
description: "Timeout in minutes"
secrets:
RUNPOD_API_KEY:
required: true
RUNPOD_PRIVATE_KEY:
required: true
WANDB_API_KEY:
required: false
jobs:
run-test:
runs-on: ubuntu-latest
environment: runpod-runners
steps:
- name: Checkout code
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Set up SSH key
run: |
mkdir -p ~/.ssh
echo "${{ secrets.RUNPOD_PRIVATE_KEY }}" > ~/.ssh/id_rsa
chmod 600 ~/.ssh/id_rsa
ssh-keygen -y -f ~/.ssh/id_rsa > ~/.ssh/id_rsa.pub
- name: Install dependencies
run: pip install requests
- name: Run tests on RunPod
env:
JOB_ID: ${{ inputs.job_id }}
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
GITHUB_RUN_ID: ${{ github.run_id }}
WANDB_API_KEY: ${{ secrets.WANDB_API_KEY }}
timeout-minutes: ${{ inputs.timeout_minutes }}
run: >-
python .github/scripts/runpod_api.py
--gpu-type "${{ inputs.gpu_type }}"
--gpu-count ${{ inputs.gpu_count }}
--volume-size ${{ inputs.volume_size }}
--disk-size ${{ inputs.disk_size }}
--image "${{ inputs.image }}"
--test-command "${{ inputs.test_command }}"
- name: Terminate RunPod Instances
if: ${{ always() }}
env:
RUNPOD_API_KEY: ${{ secrets.RUNPOD_API_KEY }}
GITHUB_RUN_ID: ${{ github.run_id }}
JOB_ID: ${{ inputs.job_id }}
run: python .github/scripts/runpod_cleanup.py
-249
View File
@@ -1,249 +0,0 @@
name: Publish Sliding Tile Attention Kernel to PyPI on Version Change
on:
push:
branches:
- main
paths:
- "csrc/attn/sliding_tile_attn/setup.py"
workflow_dispatch:
jobs:
check-version-change:
runs-on: ubuntu-latest
outputs:
version-changed: ${{ steps.check-version.outputs.changed }}
new-version: ${{ steps.check-version.outputs.new-version }}
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
fetch-depth: 2
- name: Check if version changed
id: check-version
run: |
cd csrc/attn/sliding_tile_attn
# Get current commit's version
NEW_VERSION=$(grep -oP 'VERSION\s*=\s*"\K[^"]+' setup.py)
echo "New version: $NEW_VERSION"
# Get previous version from git history
OLD_VERSION=$(git show HEAD~1:./setup.py | grep -oP 'VERSION\s*=\s*"\K[^"]+' || echo "0.0.0")
echo "Old version: $OLD_VERSION"
if [ "$NEW_VERSION" != "$OLD_VERSION" ]; then
echo "Version changed from $OLD_VERSION to $NEW_VERSION"
echo "changed=true" >> $GITHUB_OUTPUT
echo "new-version=$NEW_VERSION" >> $GITHUB_OUTPUT
else
echo "Version did not change"
echo "changed=false" >> $GITHUB_OUTPUT
fi
build_wheels:
name: Build Wheel
needs: check-version-change
if: ${{ needs.check-version-change.outputs.version-changed == 'true' || github.event_name == 'workflow_dispatch' }}
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
# Using ubuntu-20.04 instead of 22.04 for more compatibility (glibc). Ideally we'd use the
# manylinux docker image, but I haven't figured out how to install CUDA on manylinux.
os: [ubuntu-22.04]
python-version: ['3.10', '3.11', '3.12', '3.13']
torch-version: ['2.5.1', '2.6.0']
cuda-version: ['12.4.1', '12.5.1', '12.6.3']
steps:
- name: Free up disk space
run: |
echo "Initial disk space:"
df -h
# Remove large directories
sudo rm -rf /usr/share/dotnet
sudo rm -rf /usr/local/lib/android
sudo rm -rf /opt/ghc
sudo rm -rf /usr/local/share/boost
sudo rm -rf /usr/share/swift
sudo rm -rf /usr/local/lib/node_modules
sudo rm -rf /usr/local/share/powershell
sudo rm -rf /usr/share/rust
sudo rm -rf /usr/local/.ghcup
# Remove cached files
sudo rm -rf /var/lib/apt/lists/*
sudo rm -rf /var/cache/apt/archives/*
echo "Disk space after cleanup:"
df -h
- name: Checkout
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
- name: Install CUDA ${{ matrix.cuda-version }}
uses: Jimver/cuda-toolkit@v0.2.21
id: cuda-toolkit
with:
cuda: ${{ matrix.cuda-version }}
linux-local-args: '["--toolkit"]'
method: 'network'
- name: Install dependencies (GCC, Clang, CUDA Paths, Git)
run: |
sudo apt update
sudo apt install -y git patchelf gcc-11 g++-11 clang-11
sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
# Allow Git to Access Safe Directory
git config --global --add safe.directory /__w/FastVideo/FastVideo
# Set CUDA environment variables
export CUDA_HOME=/usr/local/cuda-${{ matrix.cuda-version }}
export PATH=${CUDA_HOME}/bin:${PATH}
export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:$LD_LIBRARY_PATH
# Verify installation
gcc --version
g++ --version
clang-11 --version
nvcc --version
- name: Install PyTorch ${{ matrix.torch-version }}+cu${{ matrix.cuda-version }}
run: |
pip install --upgrade pip
# With python 3.13 and torch 2.5.1, unless we update typing-extensions, we get error
# AttributeError: attribute '__default__' of 'typing.ParamSpec' objects is not writable
pip install typing-extensions==4.12.2
# We want to figure out the CUDA version to download pytorch
# e.g. we can have system CUDA version being 11.7 but if torch==1.12 then we need to download the wheel from cu116
# see https://github.com/pytorch/pytorch/blob/main/RELEASE.md#release-compatibility-matrix
export TORCH_CUDA_VERSION=124
pip install --no-cache-dir torch==${{ matrix.torch-version }} --index-url https://download.pytorch.org/whl/cu${TORCH_CUDA_VERSION}
nvcc --version
python --version
python -c "import torch; print('PyTorch:', torch.__version__)"
python -c "import torch; print('CUDA:', torch.version.cuda)"
python -c "from torch.utils import cpp_extension; print (cpp_extension.CUDA_HOME)"
- name: Build wheel
run: |
export PYTHONPATH=$GITHUB_WORKSPACE:$PYTHONPATH
# We want setuptools >= 49.6.0 otherwise we can't compile the extension if system CUDA version is 11.7 and pytorch cuda version is 11.6
# https://github.com/pytorch/pytorch/blob/664058fa83f1d8eede5d66418abff6e20bd76ca8/torch/utils/cpp_extension.py#L810
# However this still fails so I'm using a newer version of setuptools
pip install setuptools
pip install ninja packaging wheel
cd csrc/attn/sliding_tile_attn # Move into the correct folder
git submodule update --init --recursive # Ensure ThunderKittens submodule is initialized
python setup.py bdist_wheel --dist-dir=dist
- name: Rename wheel file
run: |
cd csrc/attn/sliding_tile_attn
CUDA_SHORT_VERSION=$(echo ${{ matrix.cuda-version }} | cut -d. -f1,2 | sed 's/\.//g')
TORCH_SHORT_VERSION=$(echo ${{ matrix.torch-version }} | cut -d. -f1,2)
# Get the correct version format
tmpname=cu${CUDA_SHORT_VERSION}torch${TORCH_SHORT_VERSION}
wheel_name=$(ls dist/*whl | xargs -n 1 basename | sed "s/-/+$tmpname-/2")
# Rename with version information
ls dist/*whl |xargs -I {} mv {} dist/${wheel_name}
echo "wheel_name=${wheel_name}" >> $GITHUB_ENV
- name: Upload wheel artifact
uses: actions/upload-artifact@v4
with:
name: ${{ env.wheel_name }}
path: csrc/attn/sliding_tile_attn/dist/*.whl
retention-days: 90
publish_package:
name: Publish package
needs: [build_wheels, check-version-change]
if: ${{ needs.check-version-change.outputs.version-changed == 'true' || github.event_name == 'workflow_dispatch' }}
runs-on: ubuntu-22.04
permissions:
id-token: write # Needed for OIDC Trusted Publishing
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.10'
- name: Install CUDA 12.4.1
uses: Jimver/cuda-toolkit@v0.2.21
id: cuda-toolkit
with:
cuda: 12.4.1
linux-local-args: '["--toolkit"]'
method: 'network'
sub-packages: '["nvcc"]'
- name: Install dependencies (GCC, Clang, CUDA Paths, Git)
run: |
sudo apt update
sudo apt install -y git patchelf gcc-11 g++-11 clang-11
sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
# Allow Git to Access Safe Directory
git config --global --add safe.directory /__w/FastVideo/FastVideo
# Set CUDA environment variables
export CUDA_HOME=/usr/local/cuda-12.4.1
export PATH=${CUDA_HOME}/bin:${PATH}
export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:$LD_LIBRARY_PATH
# Verify installation
gcc --version
g++ --version
clang-11 --version
nvcc --version
- name: Install PyTorch 2.5.1+cu12.4.1
run: |
pip install --upgrade pip
# With python 3.13 and torch 2.5.1, unless we update typing-extensions, we get error
# AttributeError: attribute '__default__' of 'typing.ParamSpec' objects is not writable
pip install typing-extensions==4.12.2
# We want to figure out the CUDA version to download pytorch
# e.g. we can have system CUDA version being 11.7 but if torch==1.12 then we need to download the wheel from cu116
# see https://github.com/pytorch/pytorch/blob/main/RELEASE.md#release-compatibility-matrix
export TORCH_CUDA_VERSION=124
pip install --no-cache-dir torch==2.5.1 --index-url https://download.pytorch.org/whl/cu${TORCH_CUDA_VERSION}
nvcc --version
python --version
python -c "import torch; print('PyTorch:', torch.__version__)"
python -c "import torch; print('CUDA:', torch.version.cuda)"
python -c "from torch.utils import cpp_extension; print (cpp_extension.CUDA_HOME)"
- name: Build source distribution
run: |
export PYTHONPATH=$GITHUB_WORKSPACE:$PYTHONPATH
# We want setuptools >= 49.6.0 otherwise we can't compile the extension if system CUDA version is 11.7 and pytorch cuda version is 11.6
# https://github.com/pytorch/pytorch/blob/664058fa83f1d8eede5d66418abff6e20bd76ca8/torch/utils/cpp_extension.py#L810
# However this still fails so I'm using a newer version of setuptools
pip install setuptools
pip install ninja packaging wheel
cd csrc/attn/sliding_tile_attn # Move into the correct folder
git submodule update --init --recursive # Ensure ThunderKittens submodule is initialized
python setup.py sdist --dist-dir=dist
- name: Publish release distributions to PyPI
uses: pypa/gh-action-pypi-publish@release/v1
with:
packages-dir: csrc/attn/sliding_tile_attn/dist/
-31
View File
@@ -1,31 +0,0 @@
name: Run Tests
on:
push:
branches: [ main ]
pull_request:
branches: [ main ]
jobs:
test:
runs-on: ubuntu-latest
steps:
- name: Check out repository
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.12' # or any version you need
- name: Install dependencies
run: |
python -m pip install --upgrade pip setuptools wheel
pip install torch
pip install packaging ninja
pip install -e .
pip install pytest
- name: Run Pytest
run: |
pytest --ignore csrc/attn/test
-257
View File
@@ -1,257 +0,0 @@
name: Publish Video Sparse Attention Kernel to PyPI on Version Change
on:
push:
branches:
- main
paths:
- "csrc/attn/video_sparse_attn/setup.py"
workflow_dispatch:
jobs:
check-version-change:
runs-on: ubuntu-latest
outputs:
version-changed: ${{ steps.check-version.outputs.changed }}
new-version: ${{ steps.check-version.outputs.new-version }}
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
fetch-depth: 2
- name: Check if version changed
id: check-version
run: |
cd csrc/attn/video_sparse_attn
# Get current commit's version
NEW_VERSION=$(grep -oP 'VERSION\s*=\s*"\K[^"]+' setup.py)
echo "New version: $NEW_VERSION"
# Get previous version from git history
OLD_VERSION=$(git show HEAD~1:./setup.py | grep -oP 'VERSION\s*=\s*"\K[^"]+' || echo "0.0.0")
echo "Old version: $OLD_VERSION"
if [ "$NEW_VERSION" != "$OLD_VERSION" ]; then
echo "Version changed from $OLD_VERSION to $NEW_VERSION"
echo "changed=true" >> $GITHUB_OUTPUT
echo "new-version=$NEW_VERSION" >> $GITHUB_OUTPUT
else
echo "Version did not change"
echo "changed=false" >> $GITHUB_OUTPUT
fi
build_wheels:
name: Build Wheel
needs: check-version-change
if: ${{ needs.check-version-change.outputs.version-changed == 'true' || github.event_name == 'workflow_dispatch' }}
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
# Using ubuntu-20.04 instead of 22.04 for more compatibility (glibc). Ideally we'd use the
# manylinux docker image, but I haven't figured out how to install CUDA on manylinux.
os: [ubuntu-22.04]
python-version: ['3.10', '3.11', '3.12', '3.13']
# For version reference https://pytorch.org/get-started/previous-versions/
torch-cuda:
- torch-version: '2.5.1'
cuda-version: '12.4.1'
torch-cuda-short: 'cu124'
- torch-version: '2.6.0'
cuda-version: '12.6.3'
torch-cuda-short: 'cu126'
- torch-version: '2.7.1'
cuda-version: '12.8.0'
torch-cuda-short: 'cu128'
steps:
- name: Free up disk space
run: |
echo "Initial disk space:"
df -h
# Remove large directories
sudo rm -rf /usr/share/dotnet
sudo rm -rf /usr/local/lib/android
sudo rm -rf /opt/ghc
sudo rm -rf /usr/local/share/boost
sudo rm -rf /usr/share/swift
sudo rm -rf /usr/local/lib/node_modules
sudo rm -rf /usr/local/share/powershell
sudo rm -rf /usr/share/rust
sudo rm -rf /usr/local/.ghcup
# Remove cached files
sudo rm -rf /var/lib/apt/lists/*
sudo rm -rf /var/cache/apt/archives/*
echo "Disk space after cleanup:"
df -h
- name: Checkout
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
- name: Install CUDA ${{ matrix.torch-cuda.cuda-version }}
uses: Jimver/cuda-toolkit@v0.2.21
id: cuda-toolkit
with:
cuda: ${{ matrix.torch-cuda.cuda-version }}
linux-local-args: '["--toolkit"]'
method: 'network'
- name: Install dependencies (GCC, Clang, CUDA Paths, Git)
run: |
sudo apt update
sudo apt install -y git patchelf gcc-11 g++-11 clang-11
sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
# Allow Git to Access Safe Directory
git config --global --add safe.directory /__w/FastVideo/FastVideo
# Set CUDA environment variables
export CUDA_HOME=/usr/local/cuda-${{ matrix.torch-cuda.cuda-version }}
export PATH=${CUDA_HOME}/bin:${PATH}
export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:$LD_LIBRARY_PATH
# Verify installation
gcc --version
g++ --version
clang-11 --version
nvcc --version
- name: Install PyTorch ${{ matrix.torch-cuda.torch-version }}+cu${{ matrix.torch-cuda.cuda-version }}
run: |
pip install --upgrade pip
# With python 3.13 and torch 2.5.1, unless we update typing-extensions, we get error
# AttributeError: attribute '__default__' of 'typing.ParamSpec' objects is not writable
pip install typing-extensions==4.12.2
# We want to figure out the CUDA version to download pytorch
# e.g. we can have system CUDA version being 11.7 but if torch==1.12 then we need to download the wheel from cu116
# see https://github.com/pytorch/pytorch/blob/main/RELEASE.md#release-compatibility-matrix
pip install --no-cache-dir torch==${{ matrix.torch-cuda.torch-version }} --index-url https://download.pytorch.org/whl/${{matrix.torch-cuda.torch-cuda-short}}
nvcc --version
python --version
python -c "import torch; print('PyTorch:', torch.__version__)"
python -c "import torch; print('CUDA:', torch.version.cuda)"
python -c "from torch.utils import cpp_extension; print (cpp_extension.CUDA_HOME)"
- name: Build wheel
run: |
export PYTHONPATH=$GITHUB_WORKSPACE:$PYTHONPATH
# We want setuptools >= 49.6.0 otherwise we can't compile the extension if system CUDA version is 11.7 and pytorch cuda version is 11.6
# https://github.com/pytorch/pytorch/blob/664058fa83f1d8eede5d66418abff6e20bd76ca8/torch/utils/cpp_extension.py#L810
# However this still fails so I'm using a newer version of setuptools
pip install setuptools
pip install ninja packaging wheel
cd csrc/attn/video_sparse_attn # Move into the correct folder
git submodule update --init --recursive # Ensure ThunderKittens submodule is initialized
python setup.py bdist_wheel --dist-dir=dist
- name: Rename wheel file
run: |
cd csrc/attn/video_sparse_attn
CUDA_SHORT_VERSION=$(echo ${{ matrix.torch-cuda.cuda-version }} | cut -d. -f1,2 | sed 's/\.//g')
TORCH_SHORT_VERSION=$(echo ${{ matrix.torch-cuda.torch-version }} | cut -d. -f1,2)
# Get the correct version format
tmpname=cu${CUDA_SHORT_VERSION}torch${TORCH_SHORT_VERSION}
wheel_name=$(ls dist/*whl | xargs -n 1 basename | sed "s/-/+$tmpname-/2")
# Rename with version information
ls dist/*whl |xargs -I {} mv {} dist/${wheel_name}
echo "wheel_name=${wheel_name}" >> $GITHUB_ENV
- name: Upload wheel artifact
uses: actions/upload-artifact@v4
with:
name: ${{ env.wheel_name }}
path: csrc/attn/video_sparse_attn/dist/*.whl
retention-days: 90
publish_package:
name: Publish package
needs: [build_wheels, check-version-change]
if: ${{ needs.check-version-change.outputs.version-changed == 'true' || github.event_name == 'workflow_dispatch' }}
runs-on: ubuntu-22.04
permissions:
id-token: write # Needed for OIDC Trusted Publishing
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.10'
- name: Install CUDA 12.4.1
uses: Jimver/cuda-toolkit@v0.2.21
id: cuda-toolkit
with:
cuda: 12.4.1
linux-local-args: '["--toolkit"]'
method: 'network'
sub-packages: '["nvcc"]'
- name: Install dependencies (GCC, Clang, CUDA Paths, Git)
run: |
sudo apt update
sudo apt install -y git patchelf gcc-11 g++-11 clang-11
sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
# Allow Git to Access Safe Directory
git config --global --add safe.directory /__w/FastVideo/FastVideo
# Set CUDA environment variables
export CUDA_HOME=/usr/local/cuda-12.4.1
export PATH=${CUDA_HOME}/bin:${PATH}
export LD_LIBRARY_PATH=${CUDA_HOME}/lib64:$LD_LIBRARY_PATH
# Verify installation
gcc --version
g++ --version
clang-11 --version
nvcc --version
- name: Install PyTorch 2.5.1+cu12.4.1
run: |
pip install --upgrade pip
# With python 3.13 and torch 2.5.1, unless we update typing-extensions, we get error
# AttributeError: attribute '__default__' of 'typing.ParamSpec' objects is not writable
pip install typing-extensions==4.12.2
# We want to figure out the CUDA version to download pytorch
# e.g. we can have system CUDA version being 11.7 but if torch==1.12 then we need to download the wheel from cu116
# see https://github.com/pytorch/pytorch/blob/main/RELEASE.md#release-compatibility-matrix
export TORCH_CUDA_VERSION=124
pip install --no-cache-dir torch==2.5.1 --index-url https://download.pytorch.org/whl/cu${TORCH_CUDA_VERSION}
nvcc --version
python --version
python -c "import torch; print('PyTorch:', torch.__version__)"
python -c "import torch; print('CUDA:', torch.version.cuda)"
python -c "from torch.utils import cpp_extension; print (cpp_extension.CUDA_HOME)"
- name: Build source distribution
run: |
export PYTHONPATH=$GITHUB_WORKSPACE:$PYTHONPATH
# We want setuptools >= 49.6.0 otherwise we can't compile the extension if system CUDA version is 11.7 and pytorch cuda version is 11.6
# https://github.com/pytorch/pytorch/blob/664058fa83f1d8eede5d66418abff6e20bd76ca8/torch/utils/cpp_extension.py#L810
# However this still fails so I'm using a newer version of setuptools
pip install setuptools
pip install ninja packaging wheel
cd csrc/attn/video_sparse_attn # Move into the correct folder
git submodule update --init --recursive # Ensure ThunderKittens submodule is initialized
python setup.py sdist --dist-dir=dist
- name: Publish release distributions to PyPI
uses: pypa/gh-action-pypi-publish@release/v1
with:
packages-dir: csrc/attn/video_sparse_attn/dist/
+60
View File
@@ -18,6 +18,7 @@ venv/
.venv/
runs/
samples/
Miniconda3-latest-Linux-x86_64.sh
*validation/
data/
outputs/
@@ -32,6 +33,14 @@ env
**.txt
*.log
weights/
logs/
official_weights/
converted_weights/
# SSIM test outputs
fastvideo/tests/ssim/generated_videos/
**/.cache/**
# Distribution / packaging
build/
@@ -44,6 +53,7 @@ eggs/
# MkDocs documentation
site/
docs/getting_started/examples/
docs/examples/
docs/inference/examples/
docs/training/examples/
docs/distillation/examples/
@@ -69,6 +79,56 @@ docs/distillation/examples/
!docs/assets/images/**/*.png
!comfyui/assets/**/*.png
!comfyui/assets/**/*.gif
!assets/images/**/*.png
!assets/images/**/*.jpg
!assets/images/**/*.jpeg
!assets/images/**/*.gif
!assets/videos/**/*.mp4
dmd_t2v_output/
preprocess_output_text/
# SvelteKit / Node artifacts under apps/fastvideo_studio/: see apps/fastvideo_studio/.gitignore
# Next.js / Node artifacts under apps/dreamverse/web/
apps/dreamverse/web/node_modules/
apps/dreamverse/web/.next/
apps/dreamverse/web/out/
apps/dreamverse/web/coverage/
apps/dreamverse/web/test-results/
apps/dreamverse/web/playwright-report/
apps/dreamverse/web/.env.local
apps/dreamverse/web/.env.development.local
apps/dreamverse/web/.env.test.local
# Generated by apps/dreamverse/scripts/install_native_ffmpeg.sh — host-specific
apps/dreamverse/scripts/ffmpeg-env.sh
apps/dreamverse/web/.env.production.local
# Unignore migrated Dreamverse product assets — root .gitignore globally
# ignores *.png/*.jpg/*.mp4/*.gif, but apps/dreamverse/web/public/ MUST
# be tracked (logo, icons, k2.png, etc.).
!apps/dreamverse/web/public/**/*.png
!apps/dreamverse/web/public/**/*.jpg
!apps/dreamverse/web/public/**/*.jpeg
!apps/dreamverse/web/public/**/*.mp4
!apps/dreamverse/web/public/**/*.gif
!apps/dreamverse/web/prompts/**/*.png
!apps/dreamverse/web/prompts/**/*.jpg
!apps/dreamverse/web/prompts/**/*.jpeg
!apps/dreamverse/web/prompts/**/*.mp4
!apps/dreamverse/web/prompts/**/*.gif
!apps/dreamverse/gpu-pool.svg
!apps/dreamverse/gpu-pool.drawio
.claude/
.codex/
.agents/tmp/
.sisyphus/
openspec/
fastvideo/tests/ssim/reference_videos/**
# Editor logs and local Python version pins (accidentally committed)
*.nvimlog
.nvimlog
.python-version
+3
View File
@@ -4,3 +4,6 @@
[submodule "fastvideo-kernel/include/cutlass"]
path = fastvideo-kernel/include/cutlass
url = https://github.com/NVIDIA/cutlass.git
[submodule "fastvideo/third_party/eval/vbench"]
path = fastvideo/third_party/eval/vbench
url = https://github.com/Vchitect/VBench.git
+7 -13
View File
@@ -7,22 +7,16 @@ exclude: |
fastvideo-kernel/.*|
assets/.*|
tests/.*|
demo/.*|
predict\.py|
scripts/.*|
prompts/.*|
fastvideo/data_preprocess/.*|
fastvideo/dataset/.*|
fastvideo/models/.*|
fastvideo/sample/.*|
fastvideo/train\.py|
fastvideo/utils/.*|
v2/(layers|attention|platforms|configs|distributed|models|logging_utils|third_party|hooks|api)/.*|
v2/(envs|logger|utils|version|forward_context|fastvideo_args)\.py|
^apps/dreamverse/web/.*|
examples/.*|
.github/workflows/fastvideo-publish.yml|
.github/workflows/sta-publish.yml|
.github/workflows/vsa-publish.yml|
.github/workflows/build-image-template.yml|
docs/source/inference/support_matrix.md
\.agents/.*|
.github/workflows/publish-fastvideo.yml|
.github/workflows/_template-build-image.yml
)
repos:
- repo: https://github.com/google/yapf
@@ -60,7 +54,7 @@ repos:
hooks:
- id: mypy
args: [--python-version, '3.10', --follow-imports, "skip", "--disable-error-code", "union-attr", "--disable-error-code", "override" ]
additional_dependencies: [types-cachetools, types-setuptools, types-PyYAML, types-requests]
additional_dependencies: [types-aiofiles, types-cachetools, types-setuptools, types-PyYAML, types-requests]
- repo: local
hooks:
- id: check-filenames
+47 -3
View File
@@ -8,10 +8,10 @@
- `tests/local_tests/` for additional local/component checks.
- Docs and guides: `docs/` (MkDocs source), with contributor docs in `docs/contributing/`.
- Runnable examples and scripts: `examples/` and `scripts/`.
- Static assets: `assets/`, `images/`, `videos/`, and `comfyui/assets/`.
- Static assets: `assets/` (including `assets/images/`, `assets/videos/`, and `assets/prompts/`) and `comfyui/assets/`.
## Build, Test, and Development Commands
- `uv pip install -e .[dev]`: editable install with lint/test extras.
- `UV_TORCH_BACKEND=cu126 uv pip install -e ".[dev]"`: editable CUDA 12 install with lint/test extras (`cu130` on CUDA 13).
- `pre-commit install --hook-type pre-commit --hook-type commit-msg`: enable local hooks.
- `pre-commit run --all-files`: run formatter/lint/type/spelling checks.
- `pytest tests/`: run top-level test suite.
@@ -23,7 +23,8 @@
- Python 3.10+; 4-space indentation; keep code and imports readable and explicit.
- Style tools are configured in `pyproject.toml` and `.pre-commit-config.yaml`:
- `yapf` (format), `ruff` (lint, auto-fix), `mypy` (typing), `codespell`.
- Target line length is 80.
- Lint via `pre-commit run --files <changed paths>` (or `pre-commit run --all-files` for a full sweep) before committing. Do not shell out to `yapf`/`ruff`/`codespell`/`mypy` directly — pre-commit chains them with the project's config and respects the `.pre-commit-config.yaml` excludes (e.g. `fastvideo/tests/` is intentionally skipped). If pre-commit reports `(no files to check)` for your paths, that exclude is deliberate — don't bypass it.
- Target line length is 120 (configured in `pyproject.toml` for ruff, yapf, and isort).
- Naming: `snake_case` for functions/files, `PascalCase` for classes, `UPPER_SNAKE_CASE` for constants.
## Testing Guidelines
@@ -40,3 +41,46 @@
- test evidence (`pytest`/SSIM outputs or rationale if skipped),
- linked issue/PR context,
- screenshots or sample outputs for UI/demo/docs changes.
## Agent Infrastructure
This repository is agent-friendly. Before doing any work:
1. Read the nearest in-scope `AGENTS.md` for every directory you may edit.
2. Read the relevant user-facing design or contributor guide.
3. Check `.agents/skills/*/SKILL.md` for a task-specific workflow.
4. Search `.agents/lessons/` for relevant pitfalls.
Architecture, commands, and operational guidance belong beside the code or
under `docs/`. Do not maintain static codebase maps, experiment journals,
branch-state snapshots, or duplicate user-facing documentation under
`.agents/`. Put a reusable procedure directly in a skill or contributor guide
and capture durable failures in `.agents/lessons/`.
## Per-Directory AGENTS.md
Local guidance lives next to the code. Read the in-scope file before editing:
| Directory | What it covers |
|-----------|----------------|
| `fastvideo/AGENTS.md` | Core package map, public API, registry-driven model dispatch |
| `fastvideo/configs/AGENTS.md` | Arch + pipeline config dataclasses, `param_names_mapping` |
| `fastvideo/models/AGENTS.md` | DiT / VAE / encoder / scheduler / loader layout (pre-commit excluded) |
| `fastvideo/layers/AGENTS.md` | Tensor-parallel linear/attention layer rules for ports |
| `fastvideo/attention/AGENTS.md` | Backend registry + env-var override |
| `fastvideo/pipelines/AGENTS.md` | Stage ABC, `basic/<model>/`, `preprocess/`, presets |
| `fastvideo/training/AGENTS.md` | Legacy monolithic pipelines (frozen for existing models) |
| `fastvideo/train/AGENTS.md` | New modular trainer (methods × models × callbacks, YAML) |
| `fastvideo/tests/AGENTS.md` | Test taxonomy, conftest, pre-commit-excluded path |
| `fastvideo/tests/ssim/AGENTS.md` | GPU SSIM regression authoring + reference video sync |
| `scripts/checkpoint_conversion/AGENTS.md` | Adding a converter for a new HF/official checkpoint |
## Critical: Two Training Stacks Coexist
- `fastvideo/training/` — legacy, monolithic per-model `*_training_pipeline.py` and
`*_distillation_pipeline.py`. Still authoritative for shipped models.
- `fastvideo/train/` — new modular framework (composable methods × models × callbacks
driven by YAML). Preferred for new training work.
Pick the matching stack before editing. Do not migrate a pipeline between them
without an explicit ask — the conventions and config surfaces differ.
+1 -1
View File
@@ -1 +1 @@
@AGENTS.md
@AGENTS.md
+56 -14
View File
@@ -3,19 +3,21 @@
</div>
<p align="center">
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://ibb.co/sv3MMKyv" target="_blank"> <b> WeChat </b> </a> |
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://github.com/hao-ai-lab/FastVideo/discussions/1097" target="_blank"> <b> WeChat </b> </a> |
</p>
**FastVideo is a unified post-training and inference framework for accelerated video generation.**
**FastVideo is a unified post-training and real-time inference framework for accelerated video generation.**
## NEWS
- `2025/11/19`: Release [CausalWan2.2 I2V A14B Preview](https://huggingface.co/FastVideo/CausalWan2.2-I2V-A14B-Preview-Diffusers) models, [Blog](https://hao-ai-lab.github.io/blogs/fastvideo_causalwan_preview/) and [Inference Code!](https://github.com/hao-ai-lab/FastVideo/blob/main/examples/inference/basic/basic_self_forcing_causal_wan2_2_i2v.py)
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), check out the [Blog](https://haoailab.com/blogs/fastwan-qad/).
- `2026/03/17`: Release demo: Into the Dreamverse: Vibe Directing in FastVideo, check out the [Blog](https://haoailab.com/blogs/dreamverse/).
- `2026/03/13`: Release demo: Create a 5s 1080p Video in 4.5s with FastVideo on a Single GPU, check out the [Blog](https://haoailab.com/blogs/fastvideo_realtime_1080p/).
- `2025/11/19`: Release [CausalWan2.2 I2V A14B Preview](https://huggingface.co/FastVideo/CausalWan2.2-I2V-A14B-Preview-Diffusers) models, [Blog](https://hao-ai-lab.github.io/blogs/fastvideo_causalwan_preview/) and [Inference Code!](https://github.com/hao-ai-lab/FastVideo/blob/main/examples/inference/basic/basic_self_forcing_causal_wan2_2_i2v.py).
- `2025/08/04`: Release [FastWan](https://hao-ai-lab.github.io/FastVideo/distillation/dmd) models and [Sparse-Distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/).
### More News
- `2025/06/14`: Release finetuning and inference code for [VSA](https://arxiv.org/pdf/2505.13389)
- `2025/06/14`: Release finetuning and inference code for [VSA](https://arxiv.org/pdf/2505.13389).
- `2025/04/24`: [FastVideo V1](https://hao-ai-lab.github.io/blogs/fastvideo/) is released!
- `2025/02/18`: Release the inference code for [Sliding Tile Attention](https://hao-ai-lab.github.io/blogs/sta/).
@@ -40,23 +42,50 @@ FastVideo has the following features:
- Diverse hardware and OS support
- Support H100, A100, 4090
- Support Linux, Windows, MacOS
- See this [page](https://hao-ai-lab.github.io/FastVideo/inference/hardware_support/) for full list of supported hardware and OS.
- See this [page](https://hao-ai-lab.github.io/FastVideo/inference/support_matrix/) for full list of supported models, hardware assumptions, and optimization compatibility.
- Realtime video generation & editing
- [Dreamverse](apps/dreamverse/README.md): stream and "vibe direct" video in realtime ([live demo](https://dreamverse.fastvideo.org/)), deployable on local GPU, a self-hosted B200 server, Docker, or serverless Modal
## Getting Started
We recommend using an environment manager such as `Conda` to create a clean environment:
We recommend using [uv](https://docs.astral.sh/uv/) to create a clean environment. If you previously used Conda, switching to uv generally gives faster and more stable installs.
```bash
# Create and activate a new conda environment
conda create -n fastvideo python=3.12
conda activate fastvideo
# Create and activate a new uv environment
uv venv --python 3.12 --seed
source .venv/bin/activate
# Install FastVideo
pip install fastvideo
# Install FastVideo on NVIDIA CUDA 12
UV_TORCH_BACKEND=cu126 uv pip install fastvideo
```
Use `UV_TORCH_BACKEND=cu130` on CUDA 13. Apple silicon users should follow the
[MPS installation guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
Please see our [docs](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/) for more detailed installation instructions.
> **On an NVIDIA DGX Spark (GB10 / ARM64 + CUDA 13)?** There's no prebuilt ARM wheel for the FastVideo CUDA kernel, so it's an editable from-source install (`UV_TORCH_BACKEND=cu130 uv pip install -e .`, which compiles that kernel for you) rather than `UV_TORCH_BACKEND=cu130 uv pip install fastvideo`. A compatible prebuilt ARM64 FlashAttention wheel is available separately. Follow the [DGX Spark install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/spark/).
### Install with an AI coding agent
FastVideo is a monorepo with rich agent guidance (see [`AGENTS.md`](AGENTS.md)). If you use Claude Code, Cursor, or another coding agent, paste the prompt below — it detects your platform and follows the matching guide:
```text
Install FastVideo (https://github.com/hao-ai-lab/FastVideo) into a fresh uv virtual environment.
1. Detect the platform: run `uname -m`, `nvidia-smi`, and `nvcc --version`.
2. Read and follow the matching install guide exactly (in this repo, or at
https://hao-ai-lab.github.io/FastVideo/getting_started/installation/):
- NVIDIA GPU, x86_64 -> docs/getting_started/installation/gpu.md
- NVIDIA DGX Spark / GB10, aarch64, CUDA 13 -> docs/getting_started/installation/spark.md
- Apple Silicon, macOS -> docs/getting_started/installation/mps.md
3. Use uv for every step. If a command fails, debug it and tell me what you changed.
4. Verify the result:
python -c "import fastvideo, torch; print('cuda', torch.cuda.is_available())"
fastvideo --help
5. Report which platform you detected and any deviations you had to make.
```
## Sparse Distillation
For our sparse distillation techniques, please see our [distillation docs](https://hao-ai-lab.github.io/FastVideo/distillation/dmd/) and check out our [blog](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/).
@@ -68,11 +97,24 @@ See below for recipes and datasets:
| [FastWan2.1-T2V-1.3B](https://huggingface.co/FastVideo/FastWan2.1-T2V-1.3B-Diffusers) | [Recipe](https://github.com/hao-ai-lab/FastVideo/tree/main/examples/distill/Wan2.1-T2V/Wan-Syn-Data-480P) | [FastVideo Synthetic Wan2.1 480P](https://huggingface.co/datasets/FastVideo/Wan-Syn_77x448x832_600k) |
| [FastWan2.2-TI2V-5B](https://huggingface.co/FastVideo/FastWan2.2-TI2V-5B-Diffusers) | [Recipe](https://github.com/hao-ai-lab/FastVideo/tree/main/examples/distill/Wan2.2-TI2V-5B-Diffusers/Data-free) | [FastVideo Synthetic Wan2.2 720P](https://huggingface.co/datasets/FastVideo/Wan2.2-Syn-121x704x1280_32k) |
## Dreamverse — Realtime Video Generation & Editing
[Dreamverse](apps/dreamverse/README.md) is FastVideo's realtime video generation
and editing platform — "vibe directing" a video as it streams. It lives in the
monorepo under [`apps/dreamverse/`](apps/dreamverse/) and ships its own backend
(`dreamverse-server`) plus a web UI.
Try the [live demo](https://dreamverse.fastvideo.org/), read the
[blog](https://haoailab.com/blogs/dreamverse/), or run it yourself. Dreamverse
deploys on a local GPU, a self-hosted B200 server over SSH, Docker, or
serverless [Modal](apps/dreamverse/scripts/modal/README.md) — see the
[Dreamverse README](apps/dreamverse/README.md).
## Inference
### Generating Your First Video
Here's a minimal example to generate a video using the default settings. Make sure VSA kernels are [installed](https://hao-ai-lab.github.io/FastVideo/video_sparse_attention/installation/). Create a file called `example.py` with the following code:
Here's a minimal example to generate a video using the default settings. Make sure VSA kernels are [installed](https://hao-ai-lab.github.io/FastVideo/attention/vsa/#installation). Create a file called `example.py` with the following code:
```python
import os
@@ -93,7 +135,6 @@ def main():
# Generate the video
video = generator.generate_video(
prompt,
return_frames=True, # Also return frames from this call (defaults to False)
output_path="my_videos/", # Controls where videos are saved
save_video=True
)
@@ -122,6 +163,7 @@ For a more detailed guide, please see our [inference quick start](https://hao-ai
- [DanceGRPO](https://github.com/XueZeyue/DanceGRPO): A unified framework to adapt Group Relative Policy Optimization (GRPO) to visual generation paradigms. Code based on FastVideo.
- [SRPO](https://github.com/Tencent-Hunyuan/SRPO): A method to directly align the full diffusion trajectory with fine-grained human preference. Code based on FastVideo.
- [DCM](https://github.com/Vchitect/DCM): Dual-expert consistency model for efficient and high-quality video generation. Code based on FastVideo.
- [HY-WorldPlay](https://github.com/Tencent-Hunyuan/HY-WorldPlay): An action-conditioned world model model trained using FastVideo framework.
- [Hunyuan Video 1.5](https://github.com/Tencent-Hunyuan/HunyuanVideo-1.5): A leading lightweight video generation model, where they proposed SSTA based on Sliding Tile Attention.
- [Kandinsky-5.0](https://github.com/kandinskylab/kandinsky-5): A family of diffusion models for video & image generation, where their NABLA attention includes a Sliding Tile Attention branch.
- [LongCat Video](https://github.com/meituan-longcat/LongCat-Video): A foundational video generation model with 13.6B parameters with block-sparse attention similar to Video Sparse Attention.
+3 -8
View File
@@ -1,15 +1,10 @@
try:
from .comfyui.video_generator.nodes import (NODE_CLASS_MAPPINGS,
NODE_DISPLAY_NAME_MAPPINGS)
from .comfyui.video_generator.nodes import (NODE_CLASS_MAPPINGS, NODE_DISPLAY_NAME_MAPPINGS)
WEB_DIRECTORY = "./web"
__all__ = [
'NODE_CLASS_MAPPINGS', 'NODE_DISPLAY_NAME_MAPPINGS', 'WEB_DIRECTORY'
]
__all__ = ['NODE_CLASS_MAPPINGS', 'NODE_DISPLAY_NAME_MAPPINGS', 'WEB_DIRECTORY']
except ImportError:
# ComfyUI environment not available, skip comfyui imports
NODE_CLASS_MAPPINGS = {}
NODE_DISPLAY_NAME_MAPPINGS = {}
WEB_DIRECTORY = "./web"
__all__ = [
'NODE_CLASS_MAPPINGS', 'NODE_DISPLAY_NAME_MAPPINGS', 'WEB_DIRECTORY'
]
__all__ = ['NODE_CLASS_MAPPINGS', 'NODE_DISPLAY_NAME_MAPPINGS', 'WEB_DIRECTORY']
+215
View File
@@ -0,0 +1,215 @@
# Dreamverse Agent Notes
## Repo-local workflows
- Use `.agents/skills/dreamverse-deploy/` to redeploy or stop the local
backend/frontend stack on a chosen physical GPU.
- Use `apps/dreamverse/scripts/modal/README.md` for Modal deployments and
`apps/dreamverse/docker/README.md` for image builds. The local deploy skill
intentionally does not manage remote deployments.
## Repo layout
Current paths:
- `apps/dreamverse/web/`: Next.js frontend, client-side stores, websocket
event reduction, prompt-window editing, devtools UI.
- `apps/dreamverse/dreamverse/`: current Python FastAPI runtime, websocket
protocol, prompt enhancement, prompt rewrite orchestration, GPU worker
lifecycle.
- `apps/dreamverse/dreamverse/tests/`: backend unit and integration-oriented
tests.
- `apps/dreamverse/dreamverse/benchmarks/`: prompt-provider latency/token
benchmarking scripts.
Planned paths during the OSS reorg:
- `controller/`: local control plane for provider credentials, compute
lifecycle, and proxying.
- `runtime/`: eventual rename of `apps/dreamverse/dreamverse/` once the
controller/runtime split is stable.
- `providers/`: provider adapters for local, Runpod, and Modal.
Important rule:
- Until the split lands, treat `apps/dreamverse/dreamverse/` as the
authoritative runtime and keep provider orchestration out of
`apps/dreamverse/web`.
## System split
Dreamverse is moving toward a three-part local-first architecture.
- Frontend owns local UI state, drafts, inspection tools, and user-triggered
actions.
- Controller will own local-only credentials, compute provisioning, runtime
lifecycle, and HTTP/websocket proxying.
- Runtime owns generation sessions, prompt rewrite, prompt safety, and
websocket semantics.
The browser should only talk to the local Dreamverse process, never directly to
Modal or Runpod.
## Current runtime responsibilities
The runtime in `apps/dreamverse/dreamverse/` is responsible for:
- websocket session lifecycle on `/ws`
- queueing, GPU assignment, worker startup, and stream chunk emission
- seed prompt memory and the active prompt window used for generation
- prompt enhancement and prompt rewrite execution
- prompt safety checks
- persistence and reload of prompt system prompt files
- curated preset append/read routes in devtools mode
- health and readiness endpoints
Relevant files:
- `apps/dreamverse/dreamverse/main.py`: websocket protocol, session state
machine, REST routes
- `apps/dreamverse/dreamverse/gpu_pool.py`: FastVideo-backed generation
workers
- `apps/dreamverse/dreamverse/prompt_enhancer.py`: provider clients, prompt
enhancement, rewrite execution
- `apps/dreamverse/dreamverse/rewrite_prompt_payload.py`: canonical rewrite
request body format
- `apps/dreamverse/dreamverse/config.py`: prompt file paths, provider
configuration, runtime flags
## Frontend responsibilities
The frontend in `apps/dreamverse/web/` is responsible for:
- collecting user input and deciding whether to send raw prompts or rewrite
requests
- maintaining client-side stores for session, prompt-window, stream, rewrite,
and UI state
- rendering prompt history, playback state, devtools controls, and rewrite
inspection
- building the prompt-window snapshot sent with rewrite requests
- reducing websocket events into UI state
- showing compute status and controller-driven errors once the controller lands
Relevant files:
- `apps/dreamverse/web/src/app/page.tsx`: main orchestration, websocket
connect/send paths
- `apps/dreamverse/web/src/lib/ws/reducer.ts`: applies normalized websocket
events to stores
- `apps/dreamverse/web/src/stores/promptWindow.ts`: prompt window and
preset/editor state
- `apps/dreamverse/web/src/stores/rewrite.ts`: rewrite activity timeline and
flags
- `apps/dreamverse/web/src/lib/prompts/promptWindowSnapshot.ts`: rewrite snapshot
normalization and padding
## Planned controller responsibilities
The future local controller should own:
- local-only provider credential loading and storage
- provider selection
- runtime provisioning, reuse, shutdown, and health checks
- proxying frontend HTTP and websocket traffic to the active runtime
- surfacing provisioning, ready, failed, and idle states to the frontend
- durable local settings that should survive ephemeral remote runtimes
The controller should not own:
- prompt rewrite logic
- seed prompt memory
- generation queue semantics
- websocket event schemas
## Prompt rewrite contract
Prompt rewrite is a shared flow with a strict ownership split.
Frontend responsibilities:
- decide when a user action should trigger `rewrite_seed_prompts` instead of
`append_prompt`
- send `rewrite_instruction` and a snapshot of the current prompt window
- pad the rewrite snapshot to the runtime-expected segment count using
`buildRewritePromptWindowSnapshotFromPrompts(...)`
- show rewrite activity and raw LLM output in local inspection UI
Runtime responsibilities:
- validate and normalize `prompt_window_prompts`
- choose the rewrite system prompt and provider/model/temperature
- build the canonical LLM request body in
`apps/dreamverse/dreamverse/rewrite_prompt_payload.py`
- run the rewrite through `PromptEnhancer.rewrite_prompt_sequence(...)`
- apply safety filtering to rewritten prompts
- replace the authoritative seed prompt memory when rewrite succeeds
- emit `seed_prompts_updated` and `rewrite_seed_prompts_complete`
Controller responsibilities:
- proxy the request and response
- surface runtime availability and provider lifecycle failures
Important rule:
- The frontend may suggest the prompt window to rewrite, but the runtime owns
the actual rewritten rollout and the authoritative prompt window after
acceptance.
## Rewrite modes
There are two runtime rewrite modes:
- edit existing rollout: when `prompt_window_prompts` is non-empty, rewrite the
current rollout while preserving segment count and ordering
- new rollout: when the prompt window is empty but there is a
`rewrite_instruction`, generate a fresh rollout
The frontend should not emulate runtime rewrite behavior locally. It should
prepare the snapshot, send it, and display the result.
## Prompt window ownership
- Frontend owns editable drafts, selected preset UI, and prompt-window
inspection state.
- Runtime owns the active seed prompt memory used for actual generation.
- After any runtime event with reason `rewrite`, the frontend must replace its
prompt window from the server payload instead of keeping a locally-derived
version.
## Devtools and persistence ownership
Prompt config editing is runtime-owned persistence with frontend-owned forms
today.
- Frontend loads and edits drafts through `/prompt-system-config`.
- Runtime reads and writes prompt files and reloads runtime prompt config.
Curated presets follow the same pattern:
- frontend submits append requests and may update local UI optimistically from
the response
- runtime persists the JSON file and resolves overlay vs fallback file paths
During the controller reorg, avoid moving durable user settings into ephemeral
remote runtimes. Controller-owned local persistence is preferred for anything
that must survive provider restarts.
## Editing guidance
- Do not move rewrite logic into the frontend or controller.
- Do not make the frontend the source of truth for the generated prompt window
after rewrite.
- If you change websocket message types or payload fields in
`apps/dreamverse/dreamverse/main.py`, update the reducer in
`apps/dreamverse/web/src/lib/ws/reducer.ts` in the same change.
- If you change rewrite request shape, update both
`apps/dreamverse/web/src/lib/prompts/promptWindowSnapshot.ts` and
`apps/dreamverse/dreamverse/rewrite_prompt_payload.py`.
- If you add controller-managed status or error payloads, keep them separate
from runtime websocket events unless there is a strong reason to merge them.
- If you change prompt file paths or devtools persistence, update
`apps/dreamverse/dreamverse/`, the frontend devtools UI, and any
controller-owned local persistence logic together.
- Keep provider adapters focused on runtime lifecycle and reachability, not on
prompt or session semantics.
+294
View File
@@ -0,0 +1,294 @@
# Dreamverse
Dreamverse is the FastVideo realtime video generation & editing platform. It lives in this monorepo under `apps/dreamverse/`.
**Deploy on:** [local GPU](#quick-start-local-gpu) · [self-hosted B200 (SSH)](#server-b200-deployment-ssh) · [Docker](docker/README.md) · [Modal](scripts/modal/README.md)
## Install Dreamverse
You can install Dreamverse using one of the methods below.
### Method 1: With uv pip
```bash
pip install --upgrade pip
pip install uv
uv venv .venv --python 3.12
source .venv/bin/activate
UV_TORCH_BACKEND=cu126 uv pip install "fastvideo[dreamverse]"
```
### Method 2: From source
```bash
git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo
pip install --upgrade pip
pip install uv
uv venv .venv --python 3.12
source .venv/bin/activate
UV_TORCH_BACKEND=cu126 uv pip install -e ".[dreamverse]"
```
Use `UV_TORCH_BACKEND=cu130` instead on CUDA 13.
### Method 3: Using Docker
```bash
git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo
apps/dreamverse/docker/docker_build.sh
```
See `apps/dreamverse/docker/README.md` for Docker build and run option details.
## Optional: Building FFmpeg For Better Performance
For full streaming performance in a non-Docker install, build a custom FFmpeg
binary from a FastVideo source checkout. The command below is repo-relative,
so run it from the repository root:
```bash
bash apps/dreamverse/scripts/install_native_ffmpeg.sh
```
The installer supports Linux `x86_64` and `aarch64`. It prefers conda-forge
triplet compilers when those commands are on `PATH`, otherwise it falls back to
system `gcc`/`g++` (plain venv). On `x86_64`, x264's hand-tuned SIMD also
requires `nasm`; install via whichever path fits your host:
```bash
sudo apt install nasm # Debian/Ubuntu
conda install -c conda-forge nasm # inside an active conda env
```
No sudo and no conda? Build `nasm` from source (~30s, installs into `$HOME`):
```bash
(
mkdir -p "$HOME/src" "$HOME/opt" && cd "$HOME/src"
curl -fsSL -O https://www.nasm.us/pub/nasm/releasebuilds/2.16.03/nasm-2.16.03.tar.gz
tar -xf nasm-2.16.03.tar.gz && cd nasm-2.16.03
./configure --prefix="$HOME/opt/nasm" && make -j"$(nproc)" && make install
)
export PATH="$HOME/opt/nasm/bin:$PATH" # add to ~/.bashrc to persist
```
The installer writes to `~/opt/ffmpeg-native/` and emits
`apps/dreamverse/scripts/ffmpeg-env.sh`. Source it before starting the backend
so Dreamverse uses the custom FFmpeg binary:
```bash
source apps/dreamverse/scripts/ffmpeg-env.sh
dreamverse-server
```
Docker images already run this FFmpeg build during image creation and source the
generated environment file at container startup.
## Launch Dreamverse
Start the backend with the installed Dreamverse commands:
```bash
dreamverse-server --port 8009
dreamverse-mock-server --port 8009
```
> **Expect a slow first boot.** With `torch.compile` and startup warmup enabled
> (the default), the backend compiles the segment 1 and segment 2 inference
> paths before it reports ready — this can take **tens of minutes on a cold
> cache**, regardless of how you deploy (local, server, Docker, or Modal).
> `/healthz` responds as soon as the process is up; `/readyz` stays `503` until
> warmup finishes. For a faster, uncompiled startup while testing, set
> `FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend.
## Frontend Setup
Install the web dependencies once from the FastVideo checkout:
```bash
cd apps/dreamverse/web
npm ci
```
The frontend package uses `package-lock.json`; use npm for installs and scripts.
## Quick Start: Local GPU
### Start Backend
Export the API keys used for prompt rewrite and prompt enhancement:
```bash
export CEREBRAS_API_KEY=...
export GROQ_API_KEY=...
```
If you built the optional native FFmpeg binary above, source its environment
file in the same shell before starting the backend:
```bash
source apps/dreamverse/scripts/ffmpeg-env.sh
dreamverse-server --host 0.0.0.0 --port 8009
```
The Dreamverse backend defaults to `0.0.0.0:8009` and starts one GPU worker on
the first visible GPU by default.
### Check Readiness
In another shell, verify that the backend process is alive:
```bash
curl http://localhost:8009/healthz
```
Then wait for GPU workers and startup warmup to finish:
```bash
curl http://localhost:8009/readyz
```
You can also run the same readiness path with:
```bash
BACKEND_HOST=localhost BACKEND_PORT=8009 apps/dreamverse/scripts/smoke_local.sh
```
If a backend is already running and you only want the script to probe it:
```bash
DREAMVERSE_SMOKE_START_BACKEND=0 apps/dreamverse/scripts/smoke_local.sh
```
### Start Frontend
Start the frontend:
```bash
cd apps/dreamverse/web
BACKEND_HOST=localhost BACKEND_PORT=8009 npm run dev
```
Open `http://localhost:5299`.
## Server B200 deployment (SSH)
Deploying on a remote GPU host (for example a B200 box) is a local install run
over SSH, plus a few server-specific concerns. Two paths:
### Option A: Native (source install)
SSH in, then follow [Install → From source](#method-2-from-source) and
(recommended) [Building FFmpeg](#optional-building-ffmpeg-for-better-performance),
then start the backend as in [Quick Start: Local GPU](#quick-start-local-gpu).
For a remote host, a few things differ from localhost:
- Bind all interfaces: `dreamverse-server --host 0.0.0.0 --port 8009`.
- Point the frontend/client at the host: `BACKEND_HOST=<b200-host> BACKEND_PORT=8009 npm run dev`.
- Keep the backend alive across SSH sessions (`tmux` / `systemd` / `nohup`).
- Expose / firewall port `8009`, or front it with a reverse proxy + auth.
### Option B: Docker (on the server)
SSH in, then follow [Install → Using Docker](#method-3-using-docker) and the run
steps in [`docker/README.md`](docker/README.md):
```bash
CEREBRAS_API_KEY="<key>" GROQ_API_KEY="<key>" apps/dreamverse/docker/docker_run.sh
```
## Quick Start: Mock Backend (For UI development)
The mock server emulates the Dreamverse backend protocol and streams a
synthetic FFmpeg-generated fMP4 clip, so the frontend can run without a GPU.
```bash
dreamverse-mock-server --latency 200 --port 8009
```
## Tests
Run the focused backend tests that validate local startup wiring, config, GPU
selection, and mock-server behavior:
```bash
pytest apps/dreamverse/dreamverse/tests/test_config.py \
apps/dreamverse/dreamverse/tests/test_entrypoints.py \
apps/dreamverse/dreamverse/tests/test_gpu_pool.py \
apps/dreamverse/dreamverse/tests/test_mock_server.py -q
```
Run the broader Dreamverse backend suite:
```bash
pytest apps/dreamverse/dreamverse/tests -q
```
Run the frontend tests:
```bash
cd apps/dreamverse/web
npm test
```
Run the frontend e2e tests:
```bash
cd apps/dreamverse/web
npm run e2e
```
## Troubleshooting
`dreamverse-server` exits with an install hint
- install the Dreamverse extra with `UV_TORCH_BACKEND=cu126 uv pip install -e ".[dreamverse]"` from a
source checkout, or `UV_TORCH_BACKEND=cu126 uv pip install "fastvideo[dreamverse]"` from PyPI
(`cu130` on CUDA 13).
Prompt-provider environment variable errors
- set `CEREBRAS_API_KEY`
- set `GROQ_API_KEY`
- direct `dreamverse-server` launches do not source `~/.env`; export the keys
in the shell or use the bundled launch scripts, which source `~/.env`.
`/readyz` stays at `503`
- wait for model loading and startup warmup to finish
- confirm a compatible CUDA GPU is visible to the process
- check backend logs for worker startup or warmup failures
- for a startup/debug pass without warmup, set
`FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend
Only one GPU is used
- this is the default local behavior
- set `FASTVIDEO_GPU_COUNT=<N>` to start N GPU worker subprocesses inside one
backend instance
- set `FASTVIDEO_GPU_COUNT=all` to start one worker for every visible GPU
- use `CUDA_VISIBLE_DEVICES` first if you need to pin the visible GPU set
Frontend cannot connect to backend
- confirm the backend is running on `8009`; if not, point the frontend at it
with `BACKEND_HOST=<host> BACKEND_PORT=<port> npm run dev`
- confirm `http://localhost:8009/healthz` responds before starting the frontend
- confirm `http://localhost:8009/readyz` returns `200` before clicking Generate
- use `apps/dreamverse/scripts/smoke_local.sh` for a repeatable local startup
check
Mock backend fails during startup
- install FFmpeg or set `FASTVIDEO_FFMPEG_BIN` to an FFmpeg binary
- for non-mock local GPU streaming performance, use the native FFmpeg installer
described above
## Notes
Dreamverse owns its backend app under `apps/dreamverse/dreamverse/`. It expects
`dreamverse-server`, not `fastvideo serve`.
+354
View File
@@ -0,0 +1,354 @@
# Dreamverse Architecture
## Overview
Dreamverse currently has two main runtime pieces:
- `apps/dreamverse/web/`: Next.js frontend
- `apps/dreamverse/dreamverse/`: Python FastAPI runtime
Today, the browser talks directly to the Dreamverse runtime over HTTP and a
single websocket on `/ws`. The frontend owns UI state and interaction flow. The
server owns generation state, prompt rewrite, prompt safety, websocket session
semantics, and GPU-backed execution.
Near-term OSS note:
- `apps/dreamverse/dreamverse/` is the current runtime implementation.
- A future `controller/` layer is planned for local-only compute management and
provider orchestration, but it does not exist yet.
## Repo Map
### Frontend
- `apps/dreamverse/web/src/app/page.tsx`: main client orchestration,
websocket connect, init payloads, send paths, and top-level app behavior
- `apps/dreamverse/web/src/lib/ws/reducer.ts`: reduces normalized websocket
events into client stores
- `apps/dreamverse/web/src/stores/session.ts`: connection, mode, and top-level
session UI
- `apps/dreamverse/web/src/stores/promptWindow.ts`: editable prompt window and
seed prompt UI state
- `apps/dreamverse/web/src/stores/rewrite.ts`: rewrite activity timeline and
inspection state
- `apps/dreamverse/web/src/stores/stream.ts`: playback and stream-related
client state
- `apps/dreamverse/web/src/lib/prompts/promptWindowSnapshot.ts`: prompt-window snapshot
building for rewrite requests
### Server
- `apps/dreamverse/dreamverse/main.py`: websocket endpoint, request handling,
session state machine, rewrite orchestration, REST routes, and stream relay
- `apps/dreamverse/dreamverse/gpu_pool.py`: GPU worker processes, warmup, model
loading, and `generate_video()` calls through FastVideo
- `apps/dreamverse/dreamverse/prompt_enhancer.py`: prompt enhancement, rollout
rewrite execution, provider selection, and timeout/fallback behavior
- `apps/dreamverse/dreamverse/rewrite_prompt_payload.py`: canonical rewrite request payload
building
- `apps/dreamverse/dreamverse/config.py`: runtime flags, prompt file paths, provider settings, and
warmup config
- `apps/dreamverse/dreamverse/session_init_image.py`: validates and persists uploaded initial
images for segment 1
## Current Split Of Responsibility
### Frontend owns
- local UI state and client-side stores
- prompt drafts and prompt window editing
- websocket connection management
- deciding which user action to send:
- `session_init_v2`
- `project_init_v1`
- `append_prompt`
- `rewrite_seed_prompts`
- `simple_generate`
- showing rewrite progress, stream status, prompt history, and devtools views
### Server owns
- websocket session lifecycle and protocol
- GPU assignment and worker lifecycle
- the authoritative seed prompt memory used for generation
- prompt rewrite execution and prompt safety
- actual generation queue semantics
- stream chunk emission and segment lifecycle events
- prompt config and preset persistence routes
- health and readiness endpoints
Important rule:
- The frontend may propose prompt-window state for rewrite, but the server is
the source of truth for the rewritten rollout and the active prompt memory
used for generation.
## End-To-End Flow
1. The frontend opens `/ws`.
2. The frontend sends `session_init_v2` with the initial prompt-window state,
preset metadata, and current toggles.
3. The server validates init data, persists an optional initial image, acquires
a GPU slot, and emits session status such as `gpu_assigned`.
4. The server starts or resumes project generation and emits events like
`ltx2_stream_start`, `ltx2_segment_start`, media init/chunks, and completion
events.
5. The frontend reduces those websocket events into its stores and updates the
UI.
6. User actions such as appending prompts, rewriting seed prompts, or starting
a single custom clip go back to the server over the same websocket.
## Frontend Architecture
The frontend is store-driven.
- `page.tsx` wires together websocket setup, send helpers, reducer
application, and top-level interaction flows.
- Store modules separate concerns like session state, rewrite state, prompt
window state, and stream state.
- The websocket reducer is responsible for turning normalized runtime events
into store updates. If the server event schema changes, the reducer must
change with it.
The frontend is intentionally not responsible for:
- generating rewritten prompts locally
- deciding final prompt safety outcomes
- reconstructing server session state from scratch
- inventing its own generation semantics independent of the runtime
## Server Architecture
The current runtime is a FastAPI app with a single long-lived websocket per
session.
`apps/dreamverse/dreamverse/main.py` manages:
- websocket connect/init
- prompt queues
- project and segment state
- prompt enhancement/rewrite triggers
- stream relay from GPU workers to the browser
- session logging and REST endpoints
`apps/dreamverse/dreamverse/gpu_pool.py` manages:
- model loading through FastVideo
- one or more worker processes
- startup warmup
- user join/leave commands
- `USER_STEP` execution for each segment
- continuation state between segments
`apps/dreamverse/dreamverse/prompt_enhancer.py` manages:
- prompt enhancement for user-submitted prompts
- rollout rewrite requests for the prompt window
- provider selection and fallback across configured prompt providers
- response normalization and safety-aware failure handling
## FastAPI Surface
The server is a single FastAPI application created in
`apps/dreamverse/dreamverse/main.py`.
Current built-in FastAPI docs are enabled:
- `/docs`: Swagger UI
- `/redoc`: ReDoc
- `/openapi.json`: OpenAPI schema
The runtime also mounts the frontend static build at `/` when one of the
configured frontend static directories exists. It does not expose the backend
Python package as static content.
## HTTP API
The current HTTP API is small. Most realtime behavior still goes through the
websocket.
### Core health and status routes
- `GET /healthz`
- process liveness probe
- returns a small payload with `status`, `service`, and timestamp
- `GET /readyz`
- readiness probe
- returns `503` until prompt services are initialized and at least one GPU
worker is ready
- returns readiness and GPU pool summary fields such as ready workers, total
GPUs, warmup counts, and queue size
- `GET /status`
- returns the current GPU pool status payload from `gpu_pool`
- `GET /internal/monitor/sessions`
- internal monitoring payload for session dashboards
- includes pending session count, max available sessions, prompt provider
success counts, and timestamp
### Prompt config routes
- `GET /prompt-system-config`
- returns the editable prompt-system configuration currently loaded by
`PromptEnhancer`
- `POST /prompt-system-config`
- saves prompt-system configuration to disk and reloads prompt config in the
runtime
- current editable fields include:
- next-segment system prompt
- auto-extension system prompt
- rewrite-window system prompt
- rewrite-user system prompt
- rewrite model
- rewrite temperature
### Devtools-only preset routes
These exist only when `DEVTOOLS_ENABLED` is true in
`apps/dreamverse/dreamverse/config.py`.
- `GET /curated-presets`
- returns merged curated presets, applying the local overlay file on top of
the fallback file when both exist
- `POST /curated-presets/append`
- appends a new curated preset to the overlay presets file
- validates non-empty label, normalized id, and at least two non-empty
segment prompts
## Websocket API
`WS /ws` is the main runtime API.
The websocket owns:
- session init
- project init and reset
- prompt append
- prompt rewrite
- generation toggles
- segment lifecycle events
- media stream delivery
- runtime error delivery
The websocket is the authoritative API for realtime Dreamverse behavior. The
HTTP routes mainly support health checks, devtools persistence, and monitoring.
## Prompt Rewrite Architecture
Prompt rewrite is a shared flow with strict ownership boundaries.
### Frontend responsibilities
- collect the rewrite instruction
- build the prompt-window snapshot from current client state
- send `rewrite_seed_prompts`
- show rewrite progress, raw output, fallback state, and resulting prompt list
### Server responsibilities
- validate and normalize the prompt-window payload
- choose rewrite model, system prompt, timeout, and temperature
- build the canonical prompt payload in
`apps/dreamverse/dreamverse/rewrite_prompt_payload.py`
- execute rewrite through `PromptEnhancer`
- apply safety filtering to rewritten prompts
- replace the authoritative seed prompt memory when rewrite succeeds
- emit `seed_prompts_updated` and `rewrite_seed_prompts_complete`
Important rule:
- The frontend owns editable drafts.
- The server owns the accepted rollout.
After a successful rewrite, the frontend should replace its prompt-window view
from the server payload instead of preserving a locally-derived version.
## Prompt Modes
There are three related prompt paths in the current system:
### Initial rollout
- The frontend sends seed prompts during `session_init_v2`.
- The server uses those prompts as the initial seed prompt memory.
- If the rollout starts from an empty prompt window plus an initial rewrite
instruction, the server can pause generation until rewrite completes.
### Live append
- The frontend sends `append_prompt`.
- The server may enhance that prompt, safety-check it, enqueue it, and use it
as the next generated segment.
### Rewrite
- The frontend sends `rewrite_seed_prompts`.
- The server rewrites the entire seed prompt window or generates a new rollout,
depending on the payload and current state.
## Initial Image And Segment Handling
The frontend currently sends `initial_image` as part of session init or
`simple_generate`.
The server:
- validates and persists the image
- uses it only for segment 1 when present
- keeps continuation state for later segments in the GPU worker
This means the runtime, not the frontend, decides how segment 1 image
conditioning and later continuation conditioning are applied.
## Websocket Contract
The websocket is the main integration surface between UI and runtime.
Typical incoming messages from the frontend:
- `session_init_v2`
- `project_init_v1`
- `append_prompt`
- `rewrite_seed_prompts`
- `simple_generate`
- `set_enhancement`
- `set_auto_extension`
- `set_loop_generation`
Typical outgoing messages from the server:
- `gpu_assigned`
- `ltx2_stream_start`
- `ltx2_segment_start`
- `segment_prompt_source`
- `prompt_received`
- `prompt_ready`
- `prompt_enhancing`
- `seed_prompts_updated`
- `rewrite_seed_prompts_complete`
- `media_init`
- `media_segment_complete`
- `project_idle`
- `error`
Binary websocket frames carry media chunks for playback.
## Current And Planned Architecture
Current architecture:
- browser -> `apps/dreamverse/web`
- `apps/dreamverse/web` -> `apps/dreamverse/dreamverse/main.py`
- `apps/dreamverse/dreamverse/main.py` ->
`apps/dreamverse/dreamverse/gpu_pool.py`
- `gpu_pool.py` -> FastVideo runtime
Planned architecture:
- browser -> `apps/dreamverse/web`
- `apps/dreamverse/web` -> local `controller/`
- `controller/` -> local or remote Dreamverse runtime
- runtime -> FastVideo runtime
That future controller split should not move prompt rewrite, session state, or
generation semantics out of the runtime.
+600
View File
@@ -0,0 +1,600 @@
# Dreamverse OSS Design
## Overview
Dreamverse should ship as a local-first open source application.
- The browser talks only to a local Dreamverse control plane on the user's
machine.
- Provider credentials stay local to that machine.
- Dreamverse may provision compute on the user's behalf, but Dreamverse does
not host that control path as a service.
This keeps the UX simple without turning Dreamverse into a credential-holding
hosted platform.
## Goals
- Support three compute modes behind one product surface:
- local GPU
- managed remote GPU via Runpod
- managed remote GPU via Modal
- Keep prompt rewrite, websocket session state, and generation behavior
consistent across providers.
- Keep provider API keys out of browser state and out of any hosted service.
- Make the existing runtime reusable as the common serving contract.
- Minimize provider-specific code and isolate it behind a narrow interface.
## Non-goals
- Do not make the frontend call provider APIs directly.
- Do not unify providers at the level of SSH, VM, serverless, or pod
semantics.
- Do not move prompt rewrite logic into the frontend or controller.
- Do not require remote compute for the basic product path.
## Current State
Today the repo contains two major pieces:
- `apps/dreamverse/web/`: Next.js frontend
- `apps/dreamverse/dreamverse/`: FastAPI runtime that owns websocket state,
prompt rewrite, prompt safety, and GPU-backed generation
The current runtime already exposes useful health and streaming surfaces such
as `/healthz`, `/readyz`, `/status`, and `/ws`.
## Target Architecture
The target open source structure should be:
```text
Dreamverse/
├── apps/dreamverse/
│ ├── web/ # browser UI
│ ├── dreamverse/ # current FastAPI websocket/generation runtime
│ ├── controller/ # local control plane and provider lifecycle
│ ├── providers/ # provider adapters
│ ├── tests/
│ │ ├── contract/
│ │ ├── controller/
│ │ └── smoke/
│ └── design.md
└── ...
```
Near-term note:
- `apps/dreamverse/dreamverse/` is the current runtime implementation.
- We can keep the code there initially and rename it to `runtime/` only after
the controller lands.
## Trust Model
Dreamverse is local-only for control and secrets.
- The user launches Dreamverse on their own machine.
- Provider API keys are entered into the local app or local CLI.
- The controller uses those credentials to provision or connect to compute.
- The browser never talks to Modal or Runpod directly.
- Dreamverse-hosted infrastructure is not involved.
This is the key reason the provider-based path is acceptable for OSS.
## Responsibility Split
### `apps/dreamverse/web`
The frontend should own:
- UI state, drafts, and local interaction state
- websocket event reduction into client stores
- selection of compute mode and display of cost/health/status
- local forms for provider configuration
- sending prompt requests and rewrite requests to the local controller
The frontend should not own:
- provider credentials after submission
- provider API calls
- runtime lifecycle
- authoritative prompt window after rewrite
- prompt safety or generation policy
### `controller`
The local controller should own:
- provider credential loading and local-only storage
- compute mode selection
- provisioning, reuse, shutdown, and health monitoring of runtimes
- reverse proxying HTTP and websocket traffic from the frontend to the active
runtime
- user-visible status such as provisioning, ready, failed, and idle shutdown
- local persistence for user settings that must survive ephemeral runtimes
The controller should not own:
- prompt rewrite logic
- seed prompt memory semantics
- generation queue behavior
- provider-specific UI state
### `runtime`
The runtime should remain the authoritative owner of:
- `/ws` session state
- prompt rewrite execution
- prompt safety
- seed prompt memory and prompt-window state used for generation
- generation orchestration and GPU worker lifecycle
- websocket event schemas
This preserves the current model and avoids splitting state across layers.
## Runtime Contract
Provider abstraction should happen around a stable Dreamverse runtime contract,
not around infrastructure details.
Minimum runtime surface:
- `GET /healthz`
- `GET /readyz`
- `GET /status`
- `GET/POST /prompt-system-config` if devtools persists config through the
runtime
- curated preset routes if those remain runtime-backed
- `WS /ws`
Important rule:
- The controller only needs to know how to reach a healthy runtime.
- The runtime remains provider-agnostic.
## Provider Abstraction
Use a narrow provider interface:
```python
class ComputeProvider(Protocol):
async def ensure_runtime(self, spec: RuntimeSpec) -> RuntimeHandle: ...
async def wait_until_ready(self, handle: RuntimeHandle) -> None: ...
async def stop_runtime(self, handle: RuntimeHandle) -> None: ...
```
`RuntimeHandle` should include:
- `provider`
- `runtime_id`
- `base_url`
- `ws_url`
- runtime auth headers or tokens if needed
- lifecycle metadata
- cost or hardware metadata for UI display
The controller should work only with `RuntimeHandle`, never with raw SSH hosts
or provider-specific payloads after resolution.
## Provider Notes
### Local
Local mode should be the reference implementation.
- Start the runtime as a local subprocess or connect to an already-running
local runtime URL.
- Reuse the same runtime contract as remote providers.
- Make this the first supported path and the main smoke-test target.
### Runpod
Runpod should be treated as pod lifecycle plus runtime reachability.
- Prefer prepared images or templates that auto-start the Dreamverse runtime.
- Prefer exposed HTTP/TCP ports for steady-state traffic.
- Use SSH only for bootstrap fallback, diagnostics, or repair.
- Avoid a design where the controller shells into the pod for every action.
### Modal
Modal should be treated as deployment-based runtime hosting.
- Wrap the Dreamverse runtime in a thin Modal entrypoint if needed.
- Reuse the same runtime behavior behind that wrapper.
- Do not model Modal as a machine that Dreamverse logs into.
- Do not force the websocket runtime into a per-request serverless handler
shape.
## Config and Persistence
Remote compute may be ephemeral, so mutable user configuration should not live
only inside remote runtimes.
Keep durable state local to the user's machine unless there is a strong reason
otherwise:
- provider selection
- provider credentials or credential references
- default hardware preferences
- editable prompt presets
- prompt system prompt overrides
- idle shutdown policy
Runtime-local state should be treated as disposable unless explicitly synced.
## Prompt Rewrite Ownership
Prompt rewrite remains runtime-owned even after the controller is added.
Frontend responsibilities:
- collect the rewrite instruction
- build the prompt-window snapshot
- display rewrite activity and results
Runtime responsibilities:
- validate and normalize the prompt window
- choose the rewrite system prompt and model settings
- execute rewrite
- apply safety filtering
- replace authoritative seed prompt memory
- emit the canonical completion events
Controller responsibilities:
- proxy the request and response
- surface runtime availability and failure state
This boundary should not move.
## Recommended Rollout
1. Finish the path reorg so docs and code agree on `apps/dreamverse/web`.
2. Introduce `controller/` as a local-only API/proxy process.
3. Keep `apps/dreamverse/dreamverse/` as the runtime and adapt it behind the
controller.
4. Add `local` provider first.
5. Add "bring your own runtime URL" as an escape hatch.
6. Add automated Runpod provisioning.
7. Add Modal deployment support.
8. Rename `apps/dreamverse/dreamverse/` to `runtime/` once the split is stable.
## Implementation Plan
The implementation should start with the smallest milestone that gives users a
working local GPU setup without forcing the controller/provider architecture
into the first patch series.
### Milestone 0: Make local GPU the official baseline
Goal:
- A user with a working `fastvideo` install can run the Dreamverse backend on a
local GPU and connect to it from `apps/dreamverse/web`.
Non-goals for this milestone:
- no controller process yet
- no provider abstraction yet
- no Runpod or Modal support yet
- no secret-management UI yet
Reasoning:
- `apps/dreamverse/dreamverse/` already is the real local GPU runtime.
- `apps/dreamverse/web` already knows how to talk to a backend over `/ws` and
REST rewrites.
- The shortest path is to make the existing local path explicit, reliable, and
tested before adding another layer.
### Milestone 0 work items
#### 0.1 Fix repo path assumptions after the frontend move
Current issue:
- Some paths still assume `prod-ui/`, but the frontend now lives at
`apps/dreamverse/web/`.
Required changes:
- update prompt/preset path resolution in
`apps/dreamverse/dreamverse/config.py`
- update docs that still mention `prod-ui`
- audit any frontend build settings that assume the old repo root
This is prerequisite cleanup. Local GPU mode should not depend on stale
monorepo paths.
#### 0.2 Make local runtime startup the primary supported entrypoint
Required outcome:
- one documented backend command
- one documented frontend command
- one clear env contract for local development
Expected shape:
```bash
uv pip install -e ".[dreamverse]"
dreamverse-server --host 0.0.0.0 --port 8009
cd apps/dreamverse/web
npm ci
BACKEND_HOST=localhost BACKEND_PORT=8009 npm run dev
```
Optional but useful:
- add a small root helper script or Make target for local startup
- add a `dreamverse-doctor` or lightweight startup check later
#### 0.3 Define the minimum local runtime contract
For Milestone 0, the frontend should rely only on the current runtime surface:
- `/ws`
- `/status`
- `/healthz`
- `/readyz`
- existing prompt/devtools routes
Do not add a second local API layer yet unless the current runtime surface is
proven insufficient.
#### 0.4 Make failure states explicit in the UI
Local GPU mode fails in a few predictable ways:
- backend not reachable
- backend reachable but not ready
- `fastvideo` or model runtime missing
- no compatible GPU available
Minimum implementation:
- show a clear connection error when `/ws` or `/status` fails
- surface readiness failures in a human-readable way
- avoid silent retry loops that hide backend startup failures
This is a small UI pass, not a controller project.
#### 0.5 Add a minimal local smoke test path
At this milestone, local GPU support is "done" only if there is a repeatable
test path for the local runtime contract.
Minimum test additions:
- backend tests for `/healthz`, `/readyz`, and `/status`
- a frontend integration test that assumes a reachable backend URL and verifies
connection lifecycle behavior
- one local smoke script that starts the backend and verifies readiness before
the frontend is launched
### Milestone 1: Introduce a thin local controller
Goal:
- Preserve the same local GPU behavior, but place a stable local control-plane
API in front of the runtime.
This should happen only after Milestone 0 is stable.
Scope:
- add `controller/`
- proxy `/ws` and the needed REST routes to `apps/dreamverse/dreamverse/`
- expose controller-owned status for "backend starting", "runtime ready", and
"runtime failed"
- optionally spawn the local runtime as a subprocess
Non-goal:
- do not add remote provider logic yet
Reasoning:
- the controller earns its complexity only once it stabilizes the local
contract that future providers will share
### Milestone 2: Provider abstraction on top of the controller
Goal:
- Keep the same frontend contract while allowing the controller to resolve a
runtime via `local`, then later `runpod` and `modal`.
At this point:
- define `ComputeProvider`
- implement `providers/local.py`
- move local-runtime subprocess management behind the provider interface
The first provider should be `local`, because it is cheapest to debug and
matches the runtime most closely.
## Minimal Code Change Order
If we want the shortest path to a working local GPU milestone, the change order
should be:
1. Fix `apps/dreamverse/dreamverse/config.py` and any remaining path
assumptions from `prod-ui` to `apps/dreamverse/web`.
2. Update `README.md` to document the real local GPU startup flow.
3. Confirm `apps/dreamverse/web` connects cleanly to the local wrapper-backed
backend.
4. Improve frontend error handling for backend-not-ready and backend-missing
cases.
5. Add a local smoke test and keep existing backend/frontend tests green.
6. Only then introduce `controller/`.
## Test Plan for the Local GPU Milestone
### Backend
Keep the current Python test suite as the base:
- `apps/dreamverse/dreamverse/tests/test_health_endpoints.py`
- `apps/dreamverse/dreamverse/tests/test_mock_server.py`
- `apps/dreamverse/dreamverse/tests/test_prompt_enhancer.py`
- `apps/dreamverse/dreamverse/tests/test_rewrite_prompt_payload.py`
- related config and logging tests
Add or tighten tests for:
- path resolution in `apps/dreamverse/dreamverse/config.py`
- readiness behavior when GPU pool initialization fails
- startup error messaging when `fastvideo` is unavailable
### Frontend
Keep the current Vitest suite as the base:
- websocket reducer tests
- prompt-window snapshot tests
- integration tests under `apps/dreamverse/web/src/app/`
Add or tighten tests for:
- connection failure UX when backend is down
- readiness failure UX when backend returns non-ready status
- backend routing configuration through `BACKEND_HOST` and `BACKEND_PORT`
### Manual smoke path
The first manual smoke checklist should be:
1. start `dreamverse-server`
2. confirm `GET /healthz` returns 200
3. confirm `GET /readyz` returns 200 after warmup
4. start `apps/dreamverse/web`
5. confirm the UI opens and the websocket connects
6. submit a prompt and verify the first generation starts
This checklist should be written down in the README once the milestone is
implemented.
## Test Strategy
The test suite should preserve one rule: provider changes must not be able to
break prompt rewrite, websocket semantics, or runtime behavior silently.
### 1. Runtime unit and integration tests
Keep and expand the current `pytest` coverage in
`apps/dreamverse/dreamverse/tests/test_*.py`.
Focus areas:
- config loading
- prompt rewrite payload normalization
- prompt enhancement and prompt safety
- health/readiness endpoints
- websocket session behavior
- session logging
- mock runtime behavior
These tests should remain provider-agnostic.
### 2. Controller unit tests
Add a Python test suite for the controller state machine.
Key cases:
- provider selection and validation
- credential loading from local config or env
- runtime lifecycle transitions:
- idle
- provisioning
- ready
- failed
- stopping
- idle timeout and cleanup behavior
- retry and backoff behavior
- HTTP and websocket proxy routing
These tests should use fake providers and fake runtimes by default.
### 3. Provider contract tests
Each provider should pass the same contract tests.
Examples:
- `ensure_runtime()` returns a usable `RuntimeHandle`
- `wait_until_ready()` surfaces timeout vs readiness correctly
- `stop_runtime()` is safe to call twice
- provider errors are mapped into stable controller error types
Use recorded fixtures or fakes wherever possible to avoid spend in CI.
### 4. Web contract tests
The frontend already has useful Vitest coverage under
`apps/dreamverse/web/src`.
Preserve that and expand around the controller split.
Priority areas:
- websocket event reduction
- prompt-window snapshot construction
- rewrite request shaping
- compute-status UI
- failure and reconnect UX
The frontend should mock the local controller API, not provider APIs.
### 5. Cross-layer protocol tests
Add contract fixtures that validate shared payloads across layers.
Important fixtures:
- websocket event payloads
- rewrite request payloads
- runtime status payloads
- controller status payloads
These can be simple JSON fixtures validated by both Python and TypeScript
tests. They will catch drift earlier than end-to-end tests.
### 6. Smoke tests
Add a small number of high-signal smoke tests:
- local provider + mock runtime
- local provider + real runtime when GPU is available
- controller startup + frontend health path
These should be cheap enough for routine local use.
### 7. Provider-backed manual or nightly tests
Real Modal and Runpod tests should be opt-in.
- Do not run them in default CI.
- Gate them behind explicit credentials and flags.
- Keep them focused on provisioning and reachability, not full product
regression.
This avoids flaky and expensive CI while still validating real provider flows.
## Testing Recommendations for the Next Step
The next practical additions should be:
1. A controller test suite in Python using fake providers.
2. Shared contract fixtures for websocket and rewrite payloads.
3. A smoke test that starts the local controller against the existing mock
runtime.
I would not add Playwright yet. The current web stack already has Vitest and
integration-style component tests, which are cheaper and better aligned with
the immediate reorg. Add browser automation only after the controller path is
stable.
+100
View File
@@ -0,0 +1,100 @@
# syntax=docker/dockerfile:1.7
# CUDA base. CUDA_VERSION/UBUNTU_VERSION feed the default tag (matches the
# unified docker/Dockerfile); override BUILD_BASE_IMAGE wholesale for a mirror.
ARG CUDA_VERSION=13.0.0
ARG UBUNTU_VERSION=22.04
ARG BUILD_BASE_IMAGE=nvidia/cuda:${CUDA_VERSION}-cudnn-devel-ubuntu${UBUNTU_VERSION}
FROM ${BUILD_BASE_IMAGE}
ARG BUILD_FASTVIDEO_KERNEL_FROM_SOURCE=0
ARG BUILD_DREAMVERSE_UI=0
ENV DEBIAN_FRONTEND=noninteractive \
PYTHONUNBUFFERED=1 \
UV_LINK_MODE=copy \
UV_CACHE_DIR=/opt/uv/cache
# pyproject no longer pins a PyTorch index; GPU-less build host -> pin explicitly
# (auto would fall back to CPU). Matches the base CUDA: cu130 (13.0) / cu126 (12.6).
ARG UV_TORCH_BACKEND=cu130
ENV UV_TORCH_BACKEND=${UV_TORCH_BACKEND}
SHELL ["/bin/bash", "-c"]
RUN apt-get update && apt-get install -y --no-install-recommends \
gcc-11 g++-11 clang-11 \
make cmake ninja-build pkg-config nasm \
git curl wget ca-certificates \
libssl-dev zlib1g-dev \
&& rm -rf /var/lib/apt/lists/* \
&& update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 \
--slave /usr/bin/g++ g++ /usr/bin/g++-11
# Version-agnostic /usr/local/cuda symlink so this works for any CUDA_VERSION.
ENV CUDA_HOME=/usr/local/cuda
ENV PATH=/root/.local/bin:/opt/venv/bin:${CUDA_HOME}/bin:${PATH}
ENV LD_LIBRARY_PATH=${CUDA_HOME}/lib64:${LD_LIBRARY_PATH}
ENV VIRTUAL_ENV=/opt/venv
RUN curl -LsSf https://astral.sh/uv/install.sh | sh
RUN uv venv --python 3.12 --seed /opt/venv \
&& echo 'source /opt/venv/bin/activate' >> /root/.bashrc
WORKDIR /opt/FastVideo
COPY . /opt/FastVideo
RUN --mount=type=cache,target=/opt/uv/cache \
source /opt/venv/bin/activate \
&& uv pip install "/opt/FastVideo[dreamverse]"
# Standard docker build does not expose GPUs, while fastvideo-kernel/build.sh
# detects the CUDA architecture with torch at build time. The FastVideo package
# install above brings in the pinned fastvideo-kernel package; rebuild from the
# copied source only on hosts configured for build-time GPU access.
RUN --mount=type=cache,target=/opt/uv/cache \
if [[ "${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE}" == "1" ]]; then \
source /opt/venv/bin/activate \
&& cd /opt/FastVideo/fastvideo-kernel \
&& ./build.sh; \
else \
echo "Skipping source fastvideo-kernel build; using installed fastvideo-kernel package."; \
fi
RUN if [[ "${BUILD_DREAMVERSE_UI}" == "1" ]]; then \
curl -fsSL https://deb.nodesource.com/setup_22.x | bash - \
&& apt-get install -y --no-install-recommends nodejs \
&& rm -rf /var/lib/apt/lists/* \
&& cd /opt/FastVideo/apps/dreamverse/web \
&& npm ci --ignore-scripts \
&& NEXT_OUTPUT_EXPORT=1 npm run build \
&& rm -rf node_modules .next; \
else \
echo "Skipping Dreamverse UI build; Building backend-only image."; \
fi
# The monorepo ffmpeg installer force-selects conda compiler triplets for
# local dev shells. Inside this image we explicitly opt into the system
# gcc/g++ toolchain.
RUN source /opt/venv/bin/activate \
&& INSTALL_PREFIX=/opt/ffmpeg-native \
SOURCE_DIR=/tmp/ffmpeg-native-src \
FFMPEG_NATIVE_CC=/usr/bin/gcc \
FFMPEG_NATIVE_CXX=/usr/bin/g++ \
bash /opt/FastVideo/apps/dreamverse/scripts/install_native_ffmpeg.sh
ENV FASTVIDEO_DREAMVERSE_HOME=/var/lib/dreamverse \
STREAM_MODE=av_fmp4 \
FASTVIDEO_ENABLE_PROMPT_SAFETY=0 \
HF_HOME=/root/.cache/huggingface
RUN mkdir -p /var/lib/dreamverse
EXPOSE 8009
HEALTHCHECK --interval=30s --timeout=5s --start-period=2700s --retries=3 \
CMD curl -fsS http://127.0.0.1:8009/healthz || exit 1
ENTRYPOINT ["/opt/FastVideo/apps/dreamverse/docker/docker_entrypoint.sh"]
CMD ["dreamverse-server", "--host", "0.0.0.0", "--port", "8009"]
@@ -0,0 +1,46 @@
.git/
**/.git/
.venv/
**/.venv/
.*dreamverse*/
**/__pycache__/
**/*.pyc
**/*.pyo
**/*.egg-info/
.pytest_cache/
**/.pytest_cache/
.mypy_cache/
.ruff_cache/
.cache/
apps/dreamverse/web/node_modules/
apps/dreamverse/web/.next/
apps/dreamverse/web/out/
apps/dreamverse/web/dist/
apps/dreamverse/web/test-results/
apps/dreamverse/web/playwright-report/
apps/dreamverse/outputs/
apps/dreamverse/dreamverse/outputs/
apps/dreamverse/dreamverse/prompts.local/
apps/dreamverse/logs/
outputs/
outputs_video/
quality_check_outputs/
data/
logs/
slurm-logs/
wandb/
.env
.env.*
**/prompts.local/
.codex/
.agents/
.opencode/
.agent_tmp/
.vscode/
.idea/
*.log
*.tmp
*.pdf
+97
View File
@@ -0,0 +1,97 @@
# Dreamverse Docker Image
This folder contains the Docker image for Dreamverse inside the FastVideo
monorepo. Build commands use the FastVideo repository root as the Docker
context, so run the helper scripts from this folder or from any path in the
checkout.
## Build
```bash
apps/dreamverse/docker/docker_build.sh
```
The image defaults to `dreamverse:dev`. Override it with:
```bash
DREAMVERSE_IMAGE=dreamverse:local apps/dreamverse/docker/docker_build.sh
```
Backend-only remains the default image. To include the static Dreamverse UI
served by the backend, set `BUILD_DREAMVERSE_UI=1` and choose a specific image
tag:
```bash
BUILD_DREAMVERSE_UI=1 DREAMVERSE_IMAGE=<image-tag> apps/dreamverse/docker/docker_build.sh
```
Prefer SHA-specific tags for deployable images; avoid `latest`.
The Dockerfile defaults to CUDA 13.0.0 with the cu130 PyTorch backend. CI also
builds a CUDA 12.6.3 / cu126 image. Select that local build explicitly with:
```bash
CUDA_VERSION=12.6.3 apps/dreamverse/docker/docker_build.sh
```
The helper derives `UV_TORCH_BACKEND=cu126` for CUDA 12.x and `cu130` for CUDA
13.x. Set `UV_TORCH_BACKEND` explicitly for a custom CUDA version. The legacy
complete-image-tag override remains supported and selects the same matching
backend:
```bash
CUDA_TAG=12.6.3-cudnn-devel-ubuntu22.04 apps/dreamverse/docker/docker_build.sh
```
Do not set `CUDA_TAG` and `CUDA_VERSION` together. The image installs FastVideo
from this checkout with the `dreamverse` extra, including FA4 flash-attention
and FlashInfer for NVFP4 quantization, and builds native FFmpeg.
FastVideo's pinned `fastvideo-kernel==0.3.2` package is installed by default.
To rebuild `fastvideo-kernel` from this checkout during the image build, set:
```bash
BUILD_FASTVIDEO_KERNEL_FROM_SOURCE=1 apps/dreamverse/docker/docker_build.sh
```
That source build detects the GPU architecture with torch during `docker
build`. On hosts where Docker does not expose GPUs during build, leave the
default package install path enabled.
## Run
```bash
CEREBRAS_API_KEY="<your-key>" \
GROQ_API_KEY="<your-key>" \
apps/dreamverse/docker/docker_run.sh
```
The container serves Dreamverse on host port `8009` by default and mounts:
```text
$HOME/.cache/huggingface -> /root/.cache/huggingface
apps/dreamverse/outputs -> /var/lib/dreamverse/outputs
```
Override the host port and output directory with `BACKEND_PORT` and
`DREAMVERSE_OUTPUTS_DIR`.
To pin the container to a specific host GPU, pass Docker's GPU request syntax:
```bash
DREAMVERSE_DOCKER_GPUS=device=4 FASTVIDEO_GPU_COUNT=1 \
CEREBRAS_API_KEY="<your-key>" \
GROQ_API_KEY="<your-key>" \
apps/dreamverse/docker/docker_run.sh
```
## Smoke
```bash
CEREBRAS_API_KEY=placeholder \
GROQ_API_KEY=placeholder \
apps/dreamverse/docker/docker_smoke.sh
```
The smoke script starts the container, polls `/healthz`, then polls `/readyz`.
It removes the container on exit unless `DREAMVERSE_KEEP_CONTAINER=1` is set.
+88
View File
@@ -0,0 +1,88 @@
#!/usr/bin/env bash
set -euo pipefail
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
REPO_ROOT="$(cd -- "${SCRIPT_DIR}/../../.." && pwd)"
IMAGE="${DREAMVERSE_IMAGE:-dreamverse:dev}"
ROOT_DOCKERIGNORE="${REPO_ROOT}/.dockerignore"
DOCKERFILE_DOCKERIGNORE="${SCRIPT_DIR}/Dockerfile.dockerignore"
CREATED_ROOT_DOCKERIGNORE_SYMLINK=0
cleanup_root_dockerignore_symlink() {
if [[ "${CREATED_ROOT_DOCKERIGNORE_SYMLINK}" == "1" && -L "${ROOT_DOCKERIGNORE}" ]] && \
[[ "$(readlink "${ROOT_DOCKERIGNORE}")" == "${DOCKERFILE_DOCKERIGNORE}" ]]; then
rm -- "${ROOT_DOCKERIGNORE}"
fi
}
trap cleanup_root_dockerignore_symlink EXIT INT TERM
if [[ -z "${DOCKER_BUILDKIT:-}" ]] && docker buildx version >/dev/null 2>&1; then
export DOCKER_BUILDKIT=1
fi
build_args=()
cuda_version="${CUDA_VERSION:-}"
torch_backend="${UV_TORCH_BACKEND:-}"
# CUDA_TAG was the Dockerfile's original override and contains the complete
# nvidia/cuda tag (for example, 12.6.3-cudnn-devel-ubuntu22.04). Keep accepting
# it while translating it to the parameterized Dockerfile inputs.
if [[ -n "${CUDA_TAG:-}" ]]; then
if [[ -n "${CUDA_VERSION:-}" ]]; then
printf 'CUDA_TAG and CUDA_VERSION cannot both be set. Use CUDA_VERSION for new builds.\n' >&2
exit 2
fi
cuda_version="${CUDA_TAG%%-*}"
if [[ ! "${cuda_version}" =~ ^[0-9]+\.[0-9]+(\.[0-9]+)?$ ]]; then
printf 'Cannot infer CUDA_VERSION from CUDA_TAG=%s. Use CUDA_VERSION and UV_TORCH_BACKEND instead.\n' \
"${CUDA_TAG}" >&2
exit 2
fi
build_args+=(--build-arg "BUILD_BASE_IMAGE=nvidia/cuda:${CUDA_TAG}")
fi
if [[ -n "${cuda_version}" ]]; then
build_args+=(--build-arg "CUDA_VERSION=${cuda_version}")
fi
if [[ -z "${torch_backend}" && -n "${cuda_version}" ]]; then
case "${cuda_version}" in
12.*) torch_backend=cu126 ;;
13.*) torch_backend=cu130 ;;
*)
printf 'No default UV_TORCH_BACKEND for CUDA_VERSION=%s. Set UV_TORCH_BACKEND explicitly.\n' \
"${cuda_version}" >&2
exit 2
;;
esac
fi
if [[ -n "${torch_backend}" ]]; then
build_args+=(--build-arg "UV_TORCH_BACKEND=${torch_backend}")
fi
[[ -n "${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE:-}" ]] && \
build_args+=(--build-arg "BUILD_FASTVIDEO_KERNEL_FROM_SOURCE=${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE}")
build_args+=(--build-arg "BUILD_DREAMVERSE_UI=${BUILD_DREAMVERSE_UI:-0}")
if [[ "${DOCKER_BUILDKIT:-}" == "1" ]]; then
if [[ -L "${ROOT_DOCKERIGNORE}" ]] && \
[[ "$(readlink "${ROOT_DOCKERIGNORE}")" == "${DOCKERFILE_DOCKERIGNORE}" ]]; then
rm -- "${ROOT_DOCKERIGNORE}"
printf 'Removed stale temporary root .dockerignore symlink: %s\n' "${ROOT_DOCKERIGNORE}"
fi
elif [[ -e "${ROOT_DOCKERIGNORE}" || -L "${ROOT_DOCKERIGNORE}" ]]; then
printf 'Using existing root .dockerignore: %s\n' "${ROOT_DOCKERIGNORE}"
else
ln -s -- "${DOCKERFILE_DOCKERIGNORE}" "${ROOT_DOCKERIGNORE}"
CREATED_ROOT_DOCKERIGNORE_SYMLINK=1
printf 'Created temporary root .dockerignore symlink for legacy Docker builder: %s -> %s\n' \
"${ROOT_DOCKERIGNORE}" "${DOCKERFILE_DOCKERIGNORE}"
fi
docker build \
-f "${SCRIPT_DIR}/Dockerfile" \
-t "${IMAGE}" \
"${build_args[@]}" \
"${REPO_ROOT}"
+13
View File
@@ -0,0 +1,13 @@
#!/usr/bin/env bash
set -euo pipefail
source /opt/venv/bin/activate
if [[ -f /opt/FastVideo/apps/dreamverse/scripts/ffmpeg-env.sh ]]; then
source /opt/FastVideo/apps/dreamverse/scripts/ffmpeg-env.sh
fi
: "${CEREBRAS_API_KEY:?CEREBRAS_API_KEY must be set (pass with -e CEREBRAS_API_KEY=...)}"
: "${GROQ_API_KEY:?GROQ_API_KEY must be set (pass with -e GROQ_API_KEY=...)}"
exec "$@"
+30
View File
@@ -0,0 +1,30 @@
#!/usr/bin/env bash
set -euo pipefail
: "${CEREBRAS_API_KEY:?CEREBRAS_API_KEY not set on host}"
: "${GROQ_API_KEY:?GROQ_API_KEY not set on host}"
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
DREAMVERSE_ROOT="$(cd -- "${SCRIPT_DIR}/.." && pwd)"
IMAGE="${DREAMVERSE_IMAGE:-dreamverse:dev}"
PORT="${BACKEND_PORT:-8009}"
HF_CACHE="${HF_HOME:-$HOME/.cache/huggingface}"
OUTPUTS_DIR="${DREAMVERSE_OUTPUTS_DIR:-${DREAMVERSE_ROOT}/outputs}"
GPU_REQUEST="${DREAMVERSE_DOCKER_GPUS:-all}"
mkdir -p "${HF_CACHE}" "${OUTPUTS_DIR}"
env_args=(
-e "CEREBRAS_API_KEY=${CEREBRAS_API_KEY}"
-e "GROQ_API_KEY=${GROQ_API_KEY}"
)
[[ -n "${ENABLE_TORCH_COMPILE:-}" ]] && env_args+=(-e "ENABLE_TORCH_COMPILE=${ENABLE_TORCH_COMPILE}")
[[ -n "${FASTVIDEO_GPU_COUNT:-}" ]] && env_args+=(-e "FASTVIDEO_GPU_COUNT=${FASTVIDEO_GPU_COUNT}")
exec docker run --rm --gpus "${GPU_REQUEST}" --init \
-p "${PORT}:8009" \
"${env_args[@]}" \
-v "${HF_CACHE}:/root/.cache/huggingface" \
-v "${OUTPUTS_DIR}:/var/lib/dreamverse/outputs" \
"${IMAGE}"
+75
View File
@@ -0,0 +1,75 @@
#!/usr/bin/env bash
set -euo pipefail
: "${CEREBRAS_API_KEY:?CEREBRAS_API_KEY not set on host}"
: "${GROQ_API_KEY:?GROQ_API_KEY not set on host}"
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
DREAMVERSE_ROOT="$(cd -- "${SCRIPT_DIR}/.." && pwd)"
IMAGE="${DREAMVERSE_IMAGE:-dreamverse:dev}"
PORT="${BACKEND_PORT:-8009}"
HF_CACHE="${HF_HOME:-$HOME/.cache/huggingface}"
OUTPUTS_DIR="${DREAMVERSE_OUTPUTS_DIR:-${DREAMVERSE_ROOT}/outputs}"
NAME="${DREAMVERSE_NAME:-dreamverse}"
TIMEOUT_SECONDS="${DREAMVERSE_SMOKE_TIMEOUT_SECONDS:-1200}"
POLL_SECONDS="${DREAMVERSE_SMOKE_POLL_SECONDS:-5}"
GPU_REQUEST="${DREAMVERSE_DOCKER_GPUS:-all}"
mkdir -p "${HF_CACHE}" "${OUTPUTS_DIR}"
docker rm -f "${NAME}" >/dev/null 2>&1 || true
env_args=(
-e "CEREBRAS_API_KEY=${CEREBRAS_API_KEY}"
-e "GROQ_API_KEY=${GROQ_API_KEY}"
-e "ENABLE_TORCH_COMPILE=${ENABLE_TORCH_COMPILE:-0}"
)
[[ -n "${FASTVIDEO_GPU_COUNT:-}" ]] && env_args+=(-e "FASTVIDEO_GPU_COUNT=${FASTVIDEO_GPU_COUNT}")
container_id="$(
docker run -d --rm --gpus "${GPU_REQUEST}" --init \
-p "${PORT}:8009" \
"${env_args[@]}" \
-v "${HF_CACHE}:/root/.cache/huggingface" \
-v "${OUTPUTS_DIR}:/var/lib/dreamverse/outputs" \
--name "${NAME}" \
"${IMAGE}"
)"
cleanup() {
if [[ "${DREAMVERSE_KEEP_CONTAINER:-0}" != "1" ]]; then
docker rm -f "${NAME}" >/dev/null 2>&1 || true
fi
}
trap cleanup EXIT
wait_for_endpoint() {
local path="$1"
local label="$2"
local deadline=$((SECONDS + TIMEOUT_SECONDS))
local url="http://127.0.0.1:${PORT}${path}"
echo "Waiting for ${label} at ${url}"
while (( SECONDS < deadline )); do
if curl -fsS "${url}" >/dev/null 2>&1; then
echo "${label} ok"
return 0
fi
if ! docker ps --format '{{.Names}}' | grep -qx "${NAME}"; then
echo "Container exited before ${label} became healthy." >&2
docker logs "${container_id}" >&2 || true
return 1
fi
sleep "${POLL_SECONDS}"
done
echo "Timed out waiting for ${label}." >&2
docker logs "${container_id}" >&2 || true
return 1
}
wait_for_endpoint "/healthz" "healthz"
wait_for_endpoint "/readyz" "readyz"
echo "Dreamverse Docker smoke passed for ${IMAGE} on host port ${PORT}."
+15
View File
@@ -0,0 +1,15 @@
from __future__ import annotations
DREAMVERSE_RUNTIME_DEPS_MESSAGE = (
"Dreamverse runtime deps missing — install with pip install 'fastvideo[dreamverse]'.")
def require_dreamverse_runtime_deps() -> None:
try:
import cerebras.cloud.sdk # noqa: F401
import openai # noqa: F401
except ModuleNotFoundError as exc:
missing_root = (exc.name or "").split(".", 1)[0]
if missing_root in {"cerebras", "openai"}:
raise SystemExit(DREAMVERSE_RUNTIME_DEPS_MESSAGE) from exc
raise
+445
View File
@@ -0,0 +1,445 @@
# pyright: reportMissingTypeArgument=false, reportArgumentType=false, reportOptionalSubscript=false, reportOptionalMemberAccess=false, reportConstantRedefinition=false, reportCallIssue=false
# ruff: noqa: UP007, SIM108, SIM105
# mypy: ignore-errors
"""ffmpeg fMP4 muxing with chunk-level event emission.
Self-contained: spawns ffmpeg as a subprocess, pipes raw frames into
its stdin, reads fragmented-MP4 chunks from stdout, and publishes each
chunk as a ``StreamEvent`` via the caller-supplied ``publish``
callback. Knows nothing about multiprocessing queues, the GPU pool,
or individual users — the caller decides what "publish" means.
"""
import fcntl
import os
import shutil
import subprocess
import tempfile
import threading
import time
import uuid
import wave
from dataclasses import dataclass
from typing import Union
from collections.abc import Callable
import numpy as np
import torch
FFMPEG_BIN = shutil.which(os.getenv("FASTVIDEO_FFMPEG_BIN", "ffmpeg"))
AV_MEDIA_MIME = os.getenv(
"STREAM_MIME_TYPE",
'video/mp4; codecs="avc1.42E01E,mp4a.40.2"',
)
AV_CHUNK_SIZE_BYTES = 1048576
TARGET_FPS = 24
AV_FRAGMENT_DURATION_US = int(os.getenv("FASTVIDEO_FRAG_US", "250000"))
X264_GOP_FRAMES = int(os.getenv("FASTVIDEO_X264_GOP", "12"))
X264_PROFILE = os.getenv("FASTVIDEO_X264_PROFILE", "baseline").strip().lower()
if X264_PROFILE not in {"baseline", "main", "high", "main10", "high10"}:
print(f"[WARN] Unsupported FASTVIDEO_X264_PROFILE={X264_PROFILE}; using baseline")
X264_PROFILE = "baseline"
USE_SHARED_STREAM_BUFFER = (os.getenv("FASTVIDEO_USE_SHARED_STREAM_BUFFER", "1").strip().lower()
not in {"0", "false", "no"})
SHARED_STREAM_BUFFER_BYTES = int(os.getenv("FASTVIDEO_SHARED_STREAM_BUFFER_BYTES", str(256 * 1024 * 1024)))
@dataclass
class StreamInit:
"""First event emitted — tells the consumer the stream is starting."""
stream_id: str
mime: str
uses_shared_buffer: bool
@dataclass
class StreamChunk:
"""One fMP4 chunk. Either ``chunk`` (raw bytes) or
``chunk_offset``+``chunk_length`` (read from the shared buffer)
will be populated, never both."""
stream_id: str
chunk: bytes | None = None
chunk_offset: int | None = None
chunk_length: int | None = None
uses_shared_buffer: bool = False
@dataclass
class StreamComplete:
"""Final event emitted — muxing finished successfully."""
stream_id: str
chunks: int
StreamEvent = Union[StreamInit, StreamChunk, StreamComplete]
def generate_stream_id(segment_idx: int) -> str:
"""Convenience: build a stream id of the form ``seg007-abcd1234``."""
return f"seg{segment_idx:03d}-{uuid.uuid4().hex[:8]}"
def _normalize_audio_tensor(audio: object) -> tuple[np.ndarray, int] | None:
"""Convert audio tensor/array into int16 ndarray [samples, channels]."""
if audio is None:
return None
if torch.is_tensor(audio):
audio_np = audio.detach().cpu().float().numpy()
else:
audio_np = np.asarray(audio, dtype=np.float32)
if audio_np.ndim == 1:
audio_np = audio_np[:, None]
elif audio_np.ndim == 2:
if audio_np.shape[0] <= 8 and audio_np.shape[1] > audio_np.shape[0]:
audio_np = audio_np.T
else:
return None
audio_np = np.clip(audio_np, -1.0, 1.0)
audio_int16 = (audio_np * 32767.0).astype(np.int16)
num_channels = audio_int16.shape[1]
return audio_int16, num_channels
def _write_audio_wav(
audio_int16: np.ndarray,
num_channels: int,
sample_rate: int,
) -> str:
"""Write normalized int16 audio to a temporary WAV file."""
with tempfile.NamedTemporaryFile(suffix=".wav", delete=False) as f:
wav_path = f.name
try:
with wave.open(wav_path, "wb") as wav_file:
wav_file.setnchannels(num_channels)
wav_file.setsampwidth(2)
wav_file.setframerate(sample_rate)
wav_file.writeframes(audio_int16.tobytes())
except Exception:
try:
os.unlink(wav_path)
except FileNotFoundError:
pass
raise
return wav_path
def stream_fmp4(
*,
frames: list[np.ndarray],
audio: object,
audio_sample_rate: int | None,
stream_id: str,
timings: dict,
head_trim_frames: int = 0,
head_trim_audio_frames: int | None = None,
shared_buffer=None,
shared_buffer_bytes: int = 0,
publish: Callable[[StreamEvent], None],
log_prefix: str = "",
) -> tuple[bool, str | None]:
"""Encode frames+audio with ffmpeg, publish each fMP4 chunk as an event.
Args:
frames: RGB24 video frames as HxWx3 uint8 arrays.
audio: 1D/2D tensor or ndarray, float values in [-1, 1].
audio_sample_rate: sample rate of ``audio``.
stream_id: caller-supplied identifier carried on every event.
timings: dict mutated in place with ffmpeg/stream timing metrics.
head_trim_frames: video frames to drop from the start
(conditioning overlap).
head_trim_audio_frames: video-frame-equivalent audio to drop.
Defaults to ``head_trim_frames``.
shared_buffer: optional ``mp.RawArray``-compatible object; when
provided, chunks are written into it and emitted by offset
rather than by bytes (avoids IPC copies).
shared_buffer_bytes: size of ``shared_buffer`` in bytes.
publish: callback invoked once per stream event.
log_prefix: prepended to warning prints (e.g. ``"[GPU 0]"``).
Returns:
``(True, None)`` on success, ``(False, error_message)`` on
failure. On mid-stream failure, a ``StreamInit`` may have
already been published — the caller is responsible for
handling that.
"""
if not frames:
return False, "no frames returned"
if audio is None:
return False, "audio is None"
if audio_sample_rate is None:
return False, "audio_sample_rate is None"
if FFMPEG_BIN is None:
return False, "ffmpeg not found"
if head_trim_audio_frames is None:
head_trim_audio_frames = head_trim_frames
normalized_audio = _normalize_audio_tensor(audio)
if normalized_audio is None:
shape_hint = getattr(audio, "shape", None)
return False, f"unsupported audio shape={shape_hint}"
audio_int16, num_channels = normalized_audio
if head_trim_frames < 0:
return False, (f"head_trim_frames must be >= 0, "
f"got {head_trim_frames}")
if head_trim_frames >= len(frames):
return False, (f"head_trim_frames={head_trim_frames} removes "
f"all {len(frames)} frames in segment")
out_frames = (frames[head_trim_frames:] if head_trim_frames > 0 else frames)
sample_rate = int(audio_sample_rate)
if head_trim_audio_frames > 0:
trim_start_samples = int(round((head_trim_audio_frames / float(TARGET_FPS)) * sample_rate))
if trim_start_samples >= audio_int16.shape[0]:
return False, ("audio too short after overlap trim: "
f"trim_start_samples={trim_start_samples}"
f", audio_samples={audio_int16.shape[0]}")
keep_samples = int(round((len(out_frames) / float(TARGET_FPS)) * sample_rate))
trim_end_samples = min(
audio_int16.shape[0],
trim_start_samples + keep_samples,
)
if trim_end_samples <= trim_start_samples:
return False, ("invalid audio trim range: "
f"start={trim_start_samples}, "
f"end={trim_end_samples}")
audio_int16 = audio_int16[trim_start_samples:trim_end_samples]
height = int(out_frames[0].shape[0])
width = int(out_frames[0].shape[1])
codec = os.getenv("FASTVIDEO_VIDEO_CODEC", "libx264")
t_wav_start = time.perf_counter()
wav_path = _write_audio_wav(audio_int16, num_channels, sample_rate)
wav_write_ms = (time.perf_counter() - t_wav_start) * 1000
cmd = [
FFMPEG_BIN,
"-hide_banner",
"-loglevel",
"error",
"-y",
"-f",
"rawvideo",
"-pix_fmt",
"rgb24",
"-s:v",
f"{width}x{height}",
"-r",
str(TARGET_FPS),
"-i",
"pipe:0",
"-i",
wav_path,
"-c:v",
codec,
]
if codec.endswith("_nvenc"):
cmd += [
"-preset",
os.getenv("FASTVIDEO_NVENC_PRESET", "p1"),
"-tune",
os.getenv("FASTVIDEO_NVENC_TUNE", "ull"),
"-rc",
os.getenv("FASTVIDEO_NVENC_RC", "constqp"),
"-qp",
os.getenv("FASTVIDEO_NVENC_QP", "28"),
"-bf",
os.getenv("FASTVIDEO_NVENC_BF", "0"),
]
else:
cmd += [
"-preset",
os.getenv("FASTVIDEO_X264_PRESET", "ultrafast"),
"-tune",
"zerolatency",
"-profile:v",
X264_PROFILE,
# Emit frequent keyframes so fragments are independently playable.
"-g",
str(X264_GOP_FRAMES),
"-keyint_min",
str(X264_GOP_FRAMES),
"-x264-params",
"scenecut=0",
]
cmd += [
"-c:a",
"aac",
"-pix_fmt",
os.getenv("FASTVIDEO_OUTPUT_PIX_FMT", "yuv420p"),
"-shortest",
"-movflags",
"+frag_keyframe+empty_moov+default_base_moof",
"-frag_duration",
str(AV_FRAGMENT_DURATION_US),
"-flush_packets",
"1",
"-muxdelay",
"0",
"-muxpreload",
"0",
"-f",
"mp4",
"pipe:1",
]
proc: subprocess.Popen | None = None
stderr_chunks: list[bytes] = []
writer_error: list[Exception | None] = [None]
t_stream_start = time.perf_counter()
use_shared_buffer = (USE_SHARED_STREAM_BUFFER and shared_buffer is not None and shared_buffer_bytes > 0)
shared_write_offset = 0
shared_buffer_fallback = False
shared_np = (np.frombuffer(
shared_buffer,
dtype=np.uint8,
count=shared_buffer_bytes,
) if use_shared_buffer else None)
chunk_intervals_ms: list[float] = []
chunk_publish_ms: list[float] = []
chunk_read_ms: list[float] = []
try:
t_proc_spawn_start = time.perf_counter()
proc = subprocess.Popen(
cmd,
stdin=subprocess.PIPE,
stdout=subprocess.PIPE,
stderr=subprocess.PIPE,
bufsize=0,
)
assert proc.stdin is not None
assert proc.stdout is not None
assert proc.stderr is not None
if hasattr(fcntl, "F_SETPIPE_SZ"):
fcntl.fcntl(proc.stdin.fileno(), fcntl.F_SETPIPE_SZ, 1048576)
fcntl.fcntl(proc.stdout.fileno(), fcntl.F_SETPIPE_SZ, 1048576)
ffmpeg_spawn_ms = (time.perf_counter() - t_proc_spawn_start) * 1000
def _write_frames():
try:
for frame in out_frames:
proc.stdin.write(np.ascontiguousarray(frame).tobytes())
proc.stdin.close()
except Exception as exc:
writer_error[0] = exc
try:
proc.stdin.close()
except Exception:
pass
def _read_stderr():
try:
while True:
data = proc.stderr.read(4096)
if not data:
break
stderr_chunks.append(data)
except Exception:
pass
writer_thread = threading.Thread(target=_write_frames, daemon=True)
stderr_thread = threading.Thread(target=_read_stderr, daemon=True)
writer_thread.start()
stderr_thread.start()
publish(StreamInit(
stream_id=stream_id,
mime=AV_MEDIA_MIME,
uses_shared_buffer=use_shared_buffer,
))
chunk_count = 0
total_bytes = 0
first_chunk_ms: float | None = None
last_chunk_emit_t = time.perf_counter()
while True:
t_read_start = time.perf_counter()
chunk = proc.stdout.read(AV_CHUNK_SIZE_BYTES)
t_read_end = time.perf_counter()
if not chunk:
break
chunk_read_ms.append((t_read_end - t_read_start) * 1000)
chunk_count += 1
total_bytes += len(chunk)
if first_chunk_ms is None:
first_chunk_ms = (t_read_end - t_stream_start) * 1000
chunk_intervals_ms.append((t_read_end - last_chunk_emit_t) * 1000)
t_publish_start = time.perf_counter()
if use_shared_buffer and not shared_buffer_fallback:
chunk_len = len(chunk)
write_end = shared_write_offset + chunk_len
if write_end <= shared_buffer_bytes:
shared_np[shared_write_offset:write_end] = np.frombuffer(chunk, dtype=np.uint8)
publish(
StreamChunk(
stream_id=stream_id,
chunk_offset=shared_write_offset,
chunk_length=chunk_len,
uses_shared_buffer=True,
))
shared_write_offset = write_end
chunk_publish_ms.append((time.perf_counter() - t_publish_start) * 1000)
last_chunk_emit_t = time.perf_counter()
continue
shared_buffer_fallback = True
print(f"{log_prefix} Shared stream buffer exhausted at "
f"{shared_write_offset / (1024 * 1024):.1f}MB; "
"falling back to queue chunk bytes")
publish(StreamChunk(
stream_id=stream_id,
chunk=chunk,
))
chunk_publish_ms.append((time.perf_counter() - t_publish_start) * 1000)
last_chunk_emit_t = time.perf_counter()
writer_thread.join(timeout=5.0)
rc = proc.wait()
stderr_thread.join(timeout=1.0)
if rc != 0:
stderr_tail = b"".join(stderr_chunks).decode(errors="ignore")[-1200:]
return False, f"ffmpeg av_fmp4 stream failed (rc={rc}): {stderr_tail}"
if writer_error[0] is not None:
return False, f"ffmpeg frame writer failed: {writer_error[0]}"
timings["av_encode_stream_ms"] = (time.perf_counter() - t_stream_start) * 1000
timings["av_stream_bytes"] = total_bytes
timings["av_trim_head_frames"] = head_trim_frames
timings["av_trim_head_audio_frames"] = head_trim_audio_frames
timings["av_frames_encoded"] = len(out_frames)
timings["av_shared_buffer_used"] = (bool(use_shared_buffer and not shared_buffer_fallback))
timings["av_wav_write_ms"] = wav_write_ms
timings["av_ffmpeg_spawn_ms"] = ffmpeg_spawn_ms
timings["av_first_chunk_ms"] = first_chunk_ms or 0.0
if chunk_intervals_ms:
timings["av_chunk_interval_ms_min"] = min(chunk_intervals_ms)
timings["av_chunk_interval_ms_median"] = float(np.median(chunk_intervals_ms))
timings["av_chunk_interval_ms_p95"] = (float(np.percentile(chunk_intervals_ms, 95)))
timings["av_chunk_interval_ms_max"] = max(chunk_intervals_ms)
if chunk_publish_ms:
timings["av_chunk_publish_ms_median"] = float(np.median(chunk_publish_ms))
timings["av_chunk_publish_ms_p95"] = float(np.percentile(chunk_publish_ms, 95))
if chunk_read_ms:
timings["av_chunk_read_ms_median"] = float(np.median(chunk_read_ms))
timings["av_chunk_read_ms_p95"] = float(np.percentile(chunk_read_ms, 95))
publish(StreamComplete(
stream_id=stream_id,
chunks=chunk_count,
))
return True, None
except Exception as exc:
return False, str(exc)
finally:
if proc is not None and proc.poll() is None:
try:
proc.kill()
except Exception:
pass
try:
os.remove(wav_path)
except OSError:
pass

Some files were not shown because too many files have changed in this diff Show More