Compare commits

..
145 Commits
Author SHA1 Message Date
Peiyuan Zhang 67f5b53595 [fix]: keep v2 examples runnable 2026-07-04 20:22:18 +00:00
Peiyuan Zhang aa6bf5079b [docs]: scope v2 docs to inference runtime 2026-07-04 19:33:05 +00:00
Peiyuan Zhang 4ca56d8951 [fix]: route v2 serving tasks by capabilities 2026-07-04 19:32:46 +00:00
Peiyuan Zhang 1e19650eb3 [misc]: simplify v2 program execution model 2026-07-04 19:32:30 +00:00
Peiyuan Zhang de354f804c [feat]: add FastWan QAD FP8 card 2026-07-04 19:32:18 +00:00
Peiyuan Zhang 6feee4e4fa [feat]: add FP8 vendor support for Wan weights 2026-07-04 19:32:05 +00:00
Peiyuan Zhang 82e00615e5 [feat]: wire v2 VideoGenerator to real torch backend 2026-07-04 19:31:55 +00:00
Peiyuan Zhang 30cfbfc4f6 [refactor]: remove training semantics from v2 inference contracts 2026-07-04 19:31:35 +00:00
Peiyuan Zhang a4d8978c37 [refactor]: remove v2 training package 2026-07-04 19:31:12 +00:00
Will Lin 803a6c99ae [refactor] v2: drop unwired ConditioningInjector policy
ConditioningInjector (ABC) + PassthroughConditioning were defined and exported
but never instantiated or called — zero call sites anywhere. Unlike the policies
the §5 thesis actually names (CFG / flow-shift / precision / expert-routing),
which every recipe wires (cfg=ClassicCFG(), expert=NoRouting(), ...), conditioning
is injected INLINE by the loops: `st.cond["prompt_embeds"] = ctx.slots.get(...)`
(wan21/loop.py, ltx2/loop.py); the qwen_omni cascade conditions via loop wiring.
PassthroughConditioning even described a dataflow (state.scratch["cond"]) that is
not how conditioning actually flows (loops read ctx.slots). So it was a designed-
but-bypassed seam, not forward-design — and conditioning isn't in the README §5
policy list. Removed the ABC + impl + exports; kept the live `cond` field (the
loops fill it directly) and fixed its comment.

Tests: 143 passed, 2 skipped.
2026-07-04 17:15:35 +00:00
Will Lin 8a637c3215 [refactor] v2: remove unused DataRef provenance spec
DataRef (dataset_id/revision/description — "what a recipe trained on") was only
the type of RecipeSpec.data_contract, which no recipe authored and no code read
(0/0). It is absent from the README RecipeSpec contract (§2.1: method, parents,
assumes_loop, assumes_precision, consistency_required) and not among the §18
"wire the inert metadata" roadmap items — i.e. unwired governance metadata, not
deliberate forward-design. RecipeSpec now matches §2.1 exactly.

Kept CheckpointManifest: unlike DataRef it IS wired (ModelCard.checkpoint), the
declarative "explicit components + key maps, no name-detector guessing" load
contract — declared intent, not dead.

Tests: 143 passed, 2 skipped.
2026-07-04 17:15:35 +00:00
Will Lin 51bae33c7b [refactor] v2: drop unused typed-schema slots from Component/LoopSpec
state_schema / step_schema / result_schema (LoopSpec) and config_schema /
io_schema (ComponentSpec), plus valid_parallel_plans and parallel_constraints,
were authored by no recipe and read by no executor (audited: 0 reads / 0 sets
across native + tests, no dynamic dataclasses.fields/asdict/__dict__ access, no
consumer in scripts/_vendor/examples). They duplicated mechanisms that already
exist: the concrete LoopState/WorkPlan/StepResult classes the driver uses
directly, and the card-level ParallelismContract. Removing them slims the core
spec surface with zero behavior change.

Thesis untouched: cards still own components/loops/recipe/parity; kept the
parity, precision/placement, behavior-capture (behavior_schema), wired
extension_schema, and roadmap required_for/optional_for/resident_for fields.
Also drop the stale README step_cost_model mention (that field went with the
cost mechanism in 352c1b28).

Tests: 143 passed, 2 skipped (toy backend).
2026-07-04 17:15:34 +00:00
Will Lin 54b7fe3a55 [refactor] v2: drop dead toy components + unused Karras schedule
ToyLoRA / ToyControlNet / ToyTargetModel / ToyDraftModel / ToyRewardModel (+ the
_spec_target_next helper) in the toy backend, and build_karras_sigmas in the
sampler, were defined but referenced nowhere — no recipe, card, loop, test, or
example used them. They were toy stand-ins for capabilities (adapter plane,
speculative decode, served reward) and an EDM/Karras noise schedule that were
written ahead of being wired. Remove them (-134 LOC); a toy can be re-added when
the capability is actually wired. No thesis impact, no behavior change.

Suite green: 143 passed, 2 skipped (toy backend); import v2 stays torch-free.
2026-07-04 17:15:34 +00:00
Will Lin 1fe50c0092 [refactor] v2: group flat top-level into planes; isolate vendored under _vendor/
The v2 top level had grown to ~27 dirs + 13 loose files — half of them
single-concept abstraction shells (memory/ 86 LOC, transport/ 222, parity/ 178,
extend/ 226) and half vendored fastvideo code sitting as peers to the actual v2
design. Regroup to mirror README section 3 "Planes & dependency order":

  core/     enums+types, card, loop, program, parity, request, parallel
            (the model-native contracts; no kernels)
  runtime/  + folded-in substrate: cache, memory, transport, extend
            (the import graph shows only runtime consumes them; compile/cudagraph
            already lived here)
  serving/  + deploy/ (products / fleet)
  _vendor/  all copied fastvideo: models, layers, attention, configs, distributed,
            platforms, api, hooks, logging_utils, third_party + fastvideo_args/
            utils/logger/envs/forward_context — internal layout unchanged, still
            mirrors upstream for diffing

28 dirs + 13 root files -> 8 dirs + 6 root files. Pure mechanical move: 947
absolute-import paths rewritten (v2.X -> v2.{core,runtime,serving,_vendor}.X),
boundary-anchored so platform/ (native dispatch) and platforms/ (vendored CUDA
detect) no longer collide and neither does hooks/ vs extend/. Deleted the empty
v2/loader/. README section 3 + 16 updated; stale test count corrected
(34 files/216 tests -> 22 files/143 tests).

Validated: `import v2` stays torch-free; `v2/run_tests.py` and `pytest v2/tests/`
-> 143 passed, 2 skipped on the numpy toy backend. (On a GPU box force the toy
backend with CUDA_VISIBLE_DEVICES="" or detect() picks cuda.) No external importer
changed — examples use the re-exported `from v2 import VideoGenerator`.
2026-07-04 17:15:34 +00:00
SolitaryThinker 290795daf8 [refactor] v2: remove interleave; pooled run-to-completion serving (P2)
Second step of the runtime simplification (after cost removal). Drop the
coordinated step-interleave scheduler + the interleave-parity gate; serving is
now pooled run-to-completion.

- Engine: remove run_interleaved + the WorkUnit/BatchScheduler imports; run /
  run_serial drive each request to completion (tick/run_to_completion kept as the
  per-request stepper).
- scheduler.py: remove BatchScheduler + WorkUnit + the batches metric; the
  AdmissionController is now a pure refundable memory/OOM guard.
- AsyncEngine: bound concurrency with a serving pool (asyncio.Semaphore,
  max_concurrent) — each request waits for a slot, then runs to completion.
- parity: remove assert_interleave_parity (the run_serial==run_interleaved gate);
  rename interleave_gate.py -> compare.py (compare_outputs stays — bit-parity
  between execution paths, e.g. disaggregated==inline).
- card specs: drop ParitySpec.interleave_required + LoopSpec.allows_interleaving;
  stripped interleave_required from all cards.
- tests: delete the interleave-gate/parity tests; refocus the ones that exercised
  real behavior (residual-skip, compare_outputs symmetric-empty).
- README: removed the design-doc references at the top; simplified the thesis /
  scheduler (§6) / parity (§9) / package-layout / comparison sections to pooled
  run-to-completion (no cost model, no interleave gate).

CPU mini: 143 passed / 2 skipped. Native omni port (P3) still to come. (pyproject
kernel hack excluded.)
2026-07-04 17:15:34 +00:00
SolitaryThinker 0047d54a1b [refactor] v2: remove the cost mechanism (P1 of runtime simplification)
First step toward lean pooled run-to-completion serving: rip out the GPU-time
cost/budget machinery entirely (it priced nothing useful for the target design).

Removed: CostModel + LoopSpec.step_cost_model; ResourceRequest.compute_seconds;
StepResult.actual_seconds; AdmissionController's compute budget + SchedulerMetrics
.gpu_seconds (the memory/OOM reservation guard stays); the Profiler observer
(cost calibration); per-step timing in RuntimeLoopContext; cost-based fleet/Dynamo
routing (now a coarse step-count load proxy); DeploymentCard.cost_model. Stripped
cost from all 9 recipe cards + their loops.

Tests: dropped the 3 cost-specific tests (cost routing, cost_model aliasing,
compute-budget gate); refocused 2 (loop cache validation, NaNWatch-clean).

CPU mini: 151 passed / 2 skipped. Interleave removal + pooled serving (P2) and the
native omni port (P3) follow. (pyproject kernel hack excluded as always.)
2026-07-04 17:15:34 +00:00
SolitaryThinker b2d55a7ba0 [refactor] v2: full vendor cutover — copy fastvideo modeling + layer code into v2 (zero fastvideo imports)
Replace the re-export stubs with real vendored copies of the fastvideo modeling
+ layer + supporting infra, for the kept diffusion models (wan21, wan_causal,
ltx2, flux2, matrixgame2). v2 now imports ZERO `fastvideo.*` — it is
self-contained. (bagel/qwen_omni load from vllm_omni, an external pkg, not
fastvideo; cosmos3's load_id was already dangling — both out of scope here.)

Vendored (cp + `sed fastvideo. -> v2.`):
- models/  the 5 models' nn.Module dits/vaes/encoders/audio/upsamplers + the
           component loader/ + the lazy class registry (other families' rows are
           dormant/lazy — only the 5 resolve).
- layers/ attention/ platforms/ distributed/ configs/ logging_utils/ hooks/
  third_party/pynvml + top-level forward_context/fastvideo_args/envs/logger/
  utils/version — copied verbatim (layers et al. 'as is').
- api/  slimmed to schema + results (the VideoGenerator's config dataclasses);
  the fastvideo parser/presets/overrides (which pull the pipeline runtime) are
  intentionally NOT vendored — v2 has its own runtime/loop.

Decoupling surgery (cut the loader's coupling to the fastvideo runtime):
- configs/pipeline_registry.py (vendored from fastvideo/registry.py, renamed to
  avoid colliding with v2/registry.py): dropped the _register_presets() auto-call
  and matrixgame3 (removed model); config-class resolution preserved.
- configs/pipelines/__init__.py: dropped the registry back-edge (fixes an import
  cycle) — base.py imports the registry lazily where used.
- torch_backend: load_component now uses v2.models.loader.

The only remaining external 'fastvideo*' refs are `fastvideo_kernel` (the
separate optional CUDA-kernel pkg for sparse/MoBA attention) — guarded; the dense
TORCH_SDPA path v2 uses never imports it.

Vendored subtrees added to the pre-commit exclude (faithful copies, mirroring the
existing fastvideo/models exclusion — not re-linted, to stay re-syncable).

Verified: grep finds zero fastvideo-package imports in v2/; Wan2.1 T2V on H100 is
BIT-IDENTICAL to the fastvideo-backed path (same .npy SHA256, byte-for-byte); CPU
mini 156 (154 passed + 2 env-skipped on x86, torch present). Backup: branch v2_backup.
2026-07-04 17:15:34 +00:00
SolitaryThinker 24cfe281b3 [refactor] v2: prune recipes to 8 models (+ omni shared infra)
Keep: bagel, cosmos3, flux2, matrixgame2, ltx2, qwen_omni, wan21, wan_causal
(plus the shared omni/ package that bagel/cosmos3/qwen_omni depend on). Remove
the other 26 recipe packages.

- Delete 26 recipe dirs (adapters, adaptive, cosmos2, cosmos25, fastwan, gen3c,
  hunyuangamecraft, hunyuan_video(15), hyworld, image_video, kandinsky5,
  lingbotworld, longcat, lucy_edit, matrixgame3, multi_expert, reward, sd35,
  sfwan22, speculative, stable_audio, tiled, turbowan, unified, wan_fun_control).
- recipes/__init__.py: keep-closure builders + build_default_engine /
  build_omni_engine (dropped the workflow/tiled/unified/image_video engine helpers).
- registry.py: _BUCKET_C pruned to flux2 + matrixgame2; removed the cosmos2
  ModelEntry + the CosmosTransformer3DModel arch branch + the TurboWan-14B entry.
- Delete 11 tests for removed recipes/features; patch test_bucket_c_ports to drop
  the cosmos2 reference (it now auto-derives from the pruned _BUCKET_C).
- README: correct the recipes/ roster to the kept families.

Backup of the full pre-prune tree is on branch v2_backup. CPU mini 156 passed / 0
failed (the prior 5 torch-absent bucket_c failures are gone with sd35/stable_audio);
no dangling references to any removed recipe; all kept model ids still resolve.
2026-07-04 17:15:34 +00:00
SolitaryThinker dc086b207f [perf] v2: on-device denoise loop — kill the per-step numpy<->torch round-trip (Wan2.1)
The torch adapter boundary marshalled the latent host<->device on EVERY denoise
step: _t uploaded the latent (and re-uploaded the text embeds) and _n downloaded
the velocity with a forced CUDA sync — 2*N PCIe copies + N syncs per generation,
buying nothing, since the latent could stay resident on the GPU the whole loop.

Root cause was a numpy loop surface. But the loop MATH is already array-agnostic
(CFG combine + flow-match Euler are pure arithmetic; the solver kernel already
passes torch through). So introduce a per-platform array namespace (v2/platform/
array_ns): numpy on CPU (torch-free — the parity mini is unchanged), torch-on-
device on cuda. The latent is seeded with numpy and uploaded ONCE; it then stays
resident through forward -> CFG combine -> solver -> next step; a single host
marshal happens at the request/output boundary (engine._to_artifact).

Opt-in per recipe via ModelCard.device_io (set on the Wan cards). When set on a
GPU box, build_component flips the components' TorchComponent.device_io so _out
keeps tensors on-device (in fp32, matching the old _to_numpy cast so the combine
dtype is unchanged). Un-migrated families and the CPU toy keep numpy in/out.
Also: PrecisionPolicy.cast is array-preserving; _t accepts resident tensors.

Verified BIT-IDENTICAL on real Wan2.1-1.3B / H100: the on-device latent equals
the pre-change numpy-path latent exactly (max_abs_diff 0.0, np.array_equal True).
CPU mini holds 237 passed / 5 pre-existing; pre-commit clean.

Other WanDenoiseLoop families can flip device_io next (per-family GPU re-verify);
non-Wan loops migrate to the xp namespace later.
2026-07-04 17:15:34 +00:00
SolitaryThinker b9db151658 [refactor] v2: co-locate per-model torch adapters into their recipe packages
platform/backends/ had become a flat dump of 15 per-model torch_<model>.py
adapters next to the genuinely-shared infra. Each adapter is referenced from
its card by a plain 'module:Class' string loaded via importlib, so there was
no real coupling forcing it into platform/ — the Cosmos/Flux/etc adapter
belongs WITH its recipe (card/loop/program).

Move each torch_<model>.py -> v2/recipes/<model>/adapter.py and flip the card
strings to v2.recipes.<model>.adapter:<Class>. backends/ now holds only the
shared substrate (torch_backend base, torch_cuda registration, torch_kernels,
toy/cpu/accel). Each recipe is now a self-contained package.

Cross-refs updated: gen3c/adapter imports CosmosT5Encoder from cosmos2/adapter;
sd35/program + stable_audio/card import from their own package. Recipes still
import torch-free (adapters pulled only via the string on a GPU box) — CPU mini
holds 237 passed / 5 pre-existing (bucket_c torch-absent). pre-commit clean.
2026-07-04 17:15:34 +00:00
SolitaryThinker 9124963238 [misc] v2: full pre-commit clean (ruff UP038/SIM/UP031 + mypy annotations)
Sweep all of v2/ through pre-commit (was previously only run on changed
files). Fixes surfaced across untouched modules:

- ruff: isinstance-tuple -> X | Y (UP038), try/except/pass ->
  contextlib.suppress (SIM105), negated-return (SIM103), %-format ->
  f-string (UP031).
- mypy: add annotations for no-untyped-call + var-annotated across recipes,
  training methods, torch/toy backends, and serving.
- yapf reflow of the SF-Wan KV-cache call sites (semantics unchanged).

yapf/ruff/codespell/mypy all pass; v2 tests 237 passed / 5 pre-existing
(bucket_c torch-absent on CPU venv).
2026-07-04 17:15:34 +00:00
SolitaryThinker f90e8f3e76 [bugfix] v2: SF-Wan cross-chunk KV cache — condition each chunk on prior clean chunks
The causal adapter never passed a kv_cache, so CausalWanTransformer3DModel.forward routed to
_forward_train (no cross-chunk KV) on every chunk instead of _forward_inference (the CausVid
Alg-2 KV-cache path). Each chunk denoised blind to the previous ones; the loop's cross-chunk
"context" was a toy mean(prior_latents) the adapter ignored. Result: hard discontinuities at
every chunk boundary (frame-to-frame absdiff spikes 43-55 every ~12 frames).

Fix (cuda path only; toy/CPU path and the 237-test suite untouched):
- WanDiT.alloc_causal_caches(): allocate the persistent per-block KV + cross-attn caches sized
  from the model config (mirrors CausalDenoisingStage._initialize_kv_cache).
- WanDiT.__call__: thread kv_cache/crossattn_cache/current_start/cache_start/start_frame/
  frame_seqlen so the model runs _forward_inference.
- wan_causal/loop.py: own the caches in LoopState (per-request -> interleave-safe); pass
  current_start = chunk_idx*chunk_size*frame_seqlen per chunk; do the clean-KV write
  (timestep ~0) after each chunk so the next attends to it.

Verified on H100: frame-to-frame absdiff mean 12.8->4.4, max 55.4->8.3; chunk-boundary spikes
eliminated; coherent across all 7 chunks. CPU causal toy tests unchanged (24 passed).
2026-07-04 17:15:34 +00:00
SolitaryThinker 4ff33a28b7 [misc] v2: simplify docstrings/comments + drop deleted-design-doc citations
Sweep all 296 v2 modules: simplify verbose docstrings/comments and remove 376 dangling
"(design_vN §X)" citations to the now-deleted design docs (v2/README.md is the source of
truth). Comment/docstring-only — AST-verified code-identical; the CPU suite holds at 237
passed / 5 pre-existing. Also applies yapf + ruff --fix auto-fixes (import ordering,
forward-ref annotation de-quoting under `from __future__ import annotations`; behavior-
neutral, suite-confirmed) and adds the legitimate domain terms mot/clen/te to the codespell
ignore-list. Remaining ruff (24) + mypy (68 no-untyped-call) findings are pre-existing v2
debt, untouched here.
2026-07-04 17:15:34 +00:00
SolitaryThinker 940f94f435 [docs] v2: make v2/README.md the design source of truth + one-page philosophy + M* roadmap
Unify the four design docs (design.md, designv2.md, design_v3.md, designv4.md) into a single
authoritative v2/README.md: the (recipe, runtime) thesis, driven loops, planes, one-WorkUnit
scheduler, the parity ladder + interleave gate, training-on-shared-loops, the weight-sharing
topology catalog, the current GPU status (20+ models + the BAGEL/Qwen-Omni/Cosmos3 trio
verified), and a prioritized roadmap. Recast design_summary.md as a one-page design philosophy
pointing to it. Add .agents/exploration/mstar-v2-roadmap.md (the adversarially-verified M*
Walk-Graph gap analysis driving the roadmap). Delete the four superseded design docs.
2026-07-04 17:15:34 +00:00
SolitaryThinker ac9dcb63ea [bugfix] v2: MatrixGame2/3 causal-loop progress counter + MG3 patch alignment
MatrixGame2 (causal DMD loop): bump st.step_idx on every executed work unit (each
DMD step and each clean-context pass). The loop drives its own control flow off
block_idx/dmd_idx/phase, but the runtime's no-progress watchdog keys on
st.step_idx, so a multi-block causal rollout was seen as stalled. Mirrors what
every other recipe loop does.

MatrixGame3 (5B WanModel): patch-align the latent H/W (patch_size (1,2,2)) before
denoise. The model folds (H/2, W/2) tokens, so an odd latent dim made the
unpatchified velocity come back one row/col short of the noise latent. Crop to
(latent // patch) * patch, faithful to MatrixGame3DenoisingStage.

v2 mini: 240 passed.
2026-07-04 17:15:34 +00:00
SolitaryThinker 861f87e843 [bugfix] v2: correct SF-Wan + LTX2 2-stage SR sampling defaults (GPU frame-verified)
Two distilled few-step video models rendered incorrectly on GPU; root-caused via
dense frame sampling (contact sheets) and fixed in the recipe cards.

SF-Wan2.1 (self-forcing causal, wan_causal card): was oversaturated/overcooked.
The distilled student is CFG-FREE (guidance 1.0, single forward/step) and denoises
with the 4-step DMD schedule [1000,750,500,250] (warped by FlowShiftPolicy(5.0)),
at a native causal block of 3 latent frames. Defaults were ClassicCFG@6.0 + 2 steps
+ block 2 -> overcooked AND under-denoised. Fixed: num_chunks=7, chunk_size=3,
steps_per_chunk=4; SamplingDefaults num_steps=4, guidance_scale=1.0 (7x3=21 latent
-> 81 frames). Renders a clean raccoon-in-sunflowers across all 81 frames.

LTX2-Distilled 2-stage SR (ltx2 card + LTX2VAE): was temporally blocky. Root cause
was an OOM-forced 57-frame reduction (only 8 latent temporal frames); the model is
designed for 121 (16 latent frames). Enable VAE tiling in LTX2VAE so the 121-frame
full-res decode fits the 80 GiB GPU; keep base cfg_scale=3.0 (drives brightness;
cfg=1 washed out) with stg_scale=0.0 (v2's drop-text perturbation is not real
skip-layer STG). Renders the on-prompt backyard shot, bright + temporally coherent.

torch_backend.py: enable LTX2VAE tiling; clear pre-existing mypy no-untyped-call /
yapf debt on the file (surfaced once per-file linting bypassed the duplicate-module
flakiness) by annotating the helper/constructor/maker signatures.

Tests: update the 3 affected CPU defaults/chunk-count tests. v2 mini: 240 passed.
2026-07-04 17:15:34 +00:00
SolitaryThinker 7ef2d9083e [docs] v2: GPU bring-up results — 20 models verified on H100
V2_PORTING_STATUS.md now records the GPU bring-up outcome: 20 models generate
real video/audio on H100 (the 7 prior + 13 newly-ported), with the remaining
split into fastvideo/env-blocked (SLA/VSA kernels, transformers incompat, fastvideo
registry/flash_attn gaps) and HF-access-blocked (gated cosmos2/flux2/sd35) — none
a v2 recipe bug.
2026-07-04 17:15:34 +00:00
SolitaryThinker 0ac5367b54 [feat] v2 GPU bring-up: 5 more models verified on H100 (huge MoE/world + DMD)
Second GPU pass (distilled + huge dense, 2-wide). 5 verified end-to-end with real
weights; CPU toy path kept green (240 passed, 2 skipped).

VERIFIED:
  * matrixgame3   — mp4 (9,256,256,3), zero fixes. 6.47B, standard Wan attn (NOT
                    sparse-attn-blocked, like its mg2 sibling); degenerate single-clip.
  * fastwan       — FastWan2.2-TI2V-5B-FullAttn DMD 3-step, mp4 (17,256,256,3), zero
                    fixes. The FULL-ATTENTION variant has no VSA params -> the generic
                    Wan loader maps it cleanly (reuses WanDiT via load_id, no adapter).
  * longcat       — LongCat-Video-T2V 13.58B, mp4 (17,256,256,3), zero fixes, CPU
                    offload (~40GB peak).
  * sfwan22       — Self-Forcing Wan2.2-A14B causal+MoE (2x14B), mp4 (29,288,288,3),
                    CPU expert offload (~80GB peak).
  * lingbotworld  — Wan2.2-class 2x14B camera world model, mp4 (9,256,256,3), offload
                    (~98GB peak transient).

BLOCKED: turbowan-i2v-a14b — SLA sparse-attn params (attn1.attn_impl.proj_l on all
40 layers of both experts) cannot load into the dense Wan build + needs the
fastvideo-kernel SLA Triton kernels (no nvcc here). Confirmed via the safetensors
header (no 60GB download). Same SLA family as turbowan-1.3b.

Fixes (own-port only): sfwan22/loop.py, lingbotworld/{card,program}.py +
torch_lingbotworld.py. pre-commit clean per-package.
2026-07-04 17:15:34 +00:00
SolitaryThinker 56f84018df [bugfix] v2 VideoGenerator: modality-aware result path (audio/image, not only video)
VideoGenerator._result hardcoded out.artifacts['video'].frames, so an audio-only
(Stable Audio) or image-only (SD3.5 / FLUX.2 T2I) generation crashed with
KeyError 'video' even though the engine had correctly produced the AudioArtifact /
image TensorArtifact (surfaced during stable_audio GPU bring-up). _result now
guards on the artifact present: video -> mp4 (unchanged), else image -> png
([C,H,W]/[B,..] normalized to [H,W,C]), plus the existing audio -> sibling .wav;
image_path recorded in result.extra. The video path is byte-for-byte unchanged.

v2 mini 240 passed, 2 skipped; pre-commit clean.
2026-07-04 17:15:34 +00:00
SolitaryThinker d04127fe6f [feat] v2 GPU bring-up: 8 ported models verified end-to-end on H100 (+CPU-safe fixes)
Ran each dense public port through the real VideoGenerator on H100 (2-wide across
both GPUs). 8 produce real finite output end-to-end; per-model adapter/loop fixes
landed in each port's OWN files (no shared/fastvideo edits). CPU toy path kept
green (240 passed, 2 skipped) — GPU-only conditioning gated to the cuda backend.

VERIFIED (real GPU output):
  * stable_audio    — stereo audio (2, 441000) @44.1kHz. Fixes: dedicated
                      'conditioner' component kind (SA owns its T5, empty
                      text_encoder_configs) + ConditionerLoader from conditioner/
                      + VDenoiser c_noise = atan(sigma)/(pi/2).
  * matrixgame2     — mp4 (9,256,256,3). Loads CLEAN (not sparse-attn-blocked).
                      Fixes: 20ch cond_concat (4ch mask + 16ch img), mandatory i2v
                      (synth blank first frame), pre-sized kv_cache/crossattn_cache
                      (SDPA inference path, avoids flex_attention compile), bf16
                      autocast, per-request reset_caches.
  * gen3c           — video (9,256,256,3), zero fixes (worked first try).
  * wan_fun_control — video (9,256,256,3).
  * lucy_edit       — video (17,256,256,3).
  * hunyuangamecraft— video (9,256,256,3).
  * hunyuan_video   — video (3,9,256,256), dual LLaMA+CLIP + Hunyuan VAE.
  * hunyuan_video15 — video (9,256,256,3), dual Qwen+ByT5 (gated to cuda; CPU passes
                      single embed).

BLOCKED (not v2 recipe bugs — load/run reached, then a fastvideo/env wall):
  * cosmos25 — DiT + VAE ran finite on GPU; the Qwen2.5-VL Reason1 encoder hits a
               transformers 5.12.1 incompat in fastvideo shared code
               (Qwen2_5_VLConfig.pad_token_id).
  * kandinsky5 — fastvideo registry.py registers it with a bare PipelineConfig (no
                 Kandinsky5 config) -> load fails fastvideo-side. (latent z=16 +
                 visual_cond adapter corrected; toy decoupled to its own channels.)
  * hyworld — fastvideo's hyworld DiT hardcodes flash_attn (no SDPA fallback);
              flash_attn kernel not built here.

CPU-safety fixes (mine): kandinsky5 toy ToyDiT/ToyVAE use the toy LATENT_CHANNELS
(not the real z=16); hunyuan_video15 dual-encoder packing gated to cuda.
Sparse-attn distilled (turbowan/SLA, fastwan/VSA) + gated (cosmos2/flux2/sd35) +
huge (>80GB) handled separately. pre-commit clean per-package.
2026-07-04 17:15:34 +00:00
SolitaryThinker ae0cf3e2d1 [docs] v2: porting status — ALL fastvideo models ported (63/64 by-id; VSA env-blocked)
Rewrites V2_PORTING_STATUS.md to reflect completion: the scope is now ALL
fastvideo models (not Wan+LTX-2 only). Documents the self-contained recipe-package
porting mechanism (ComponentSpec.adapter), the 15 net-new architectures + 5
Wan-family variants newly ported (CPU-verified end-to-end; GPU=BRINGUP), the 7
GPU-verified models, and the single env-blocked id (VSA-14B, needs nvcc).
2026-07-04 17:15:34 +00:00
SolitaryThinker f4af3cf886 [feat] v2 registry: LTX-2/2.3 repo aliases -> by-id resolution 63/64
Adds explicit ModelEntry aliases for the LTX-2 (FastVideo/LTX2-Diffusers,
LTX2-base, Lightricks/LTX-2 -> single-stage base) and LTX-2.3 (LTX2.3-Diffusers,
LTX2.3-Distilled-Diffusers, LTX2.3-base, Lightricks/LTX-2.3, lightricks/ltx-2.3
-> the distilled joint-A/V card) naming variants of already-ported LTX
checkpoints (the arch fallback also resolves LTX2Transformer3DModel from a root).

v2 now resolves 63/64 fastvideo registry ids by exact id; the only remaining id,
FastVideo/Wan2.1-VSA-T2V-14B-720P-Diffusers, is ENV-BLOCKED (VSA Sparse-Linear
Attention kernels require nvcc, not built in this bring-up; it arch-resolves to
the base Wan card but needs the VSA kernel build to run faithfully).

v2 mini 240 passed, 2 skipped; pre-commit clean.
2026-07-04 17:15:34 +00:00
SolitaryThinker 39580ad73a [feat] v2: port the residual Wan-family variants (rCM/DMD/v2v/control/causal-MoE)
Closes the bucket-B sampler/conditioning gap — each reuses the Wan/Causal ARCH
(no new torch adapter) with a new in-package loop/sampler/conditioning, declared
in _BUCKET_C as explicit-HF-id-only (transformer_cls="" so the generic Wan/Causal
arch fallback is NOT hijacked — only the exact id distinguishes the capability
variant from a base Wan of the same class).

  * turbowan      — TurboWan rCM (Reparameterized Consistency Model) few-step: a
                    faithful in-package RCMScheduler port (TrigFlow->RectifiedFlow
                    schedule + stochastic consistency SDE step), 1.3B/14B T2V +
                    TurboWan2.2-I2V-A14B (MoE i2v, boundary 0.9 in raw-sigma space)
  * lucy_edit     — Lucy-Edit v2v editor: a video_vae_encode node (the input video
                    -> 48ch cond latent) threaded via the shared i2v_cond hook ->
                    96ch Lucy DiT input (faithful to denoising.py is_lucy_edit)
  * wan_fun_control — Wan2.1-Fun-Control: control-video conditioning (reuses the
                    i2v [mask|cond] concat pattern)
  * sfwan22       — Self-Forcing Wan2.2-A14B: causal chunk_rollout + Wan2.2 MoE
                    boundary routing, i2v (boundary 0.9) + t2v (boundary 0.875)
  * fastwan       — FastWan DMD 3-step: TI2V-5B-FullAttn loadable; the VSA-trained
                    variants + non-strict to_gate_compress load are BRINGUP

All 12 residual ids resolve+build; base Wan/Causal resolution unchanged (arch
fallback not hijacked). The _BUCKET_C regression test auto-extended -> v2 mini
240 passed, 2 skipped; pre-commit clean (per-package + registry).
2026-07-04 17:15:34 +00:00
SolitaryThinker 521c2845e6 [test] v2: end-to-end CPU regression guard for the bucket-C ports
Data-driven from registry._BUCKET_C (+ cosmos2): each net-new ported arch
resolves through the registry (exact id + arch fallback) AND runs end-to-end on
the CPU toy backend via the public Engine path (resolve -> build card+program ->
load_card -> Engine.run), emitting exactly one modality-correct artifact
(video / image / audio) + latents. Auto-covers future _BUCKET_C rows.

21 tests pass; full v2 mini 232 passed, 2 skipped.
2026-07-04 17:15:34 +00:00
SolitaryThinker 0467edbd07 [feat] v2: port the 14 remaining bucket-C archs as self-contained recipe packages
Completes the bucket-C porting backlog. Each arch is a self-contained recipe
package (card-declared torch adapter via ComponentSpec.adapter + a new/forked
loop + program, NO edit to the shared torch_backend dispatch), following the
cosmos2 reference pattern. One _BUCKET_C table in v2/registry.py drives both the
explicit HF-id registry (PRIMARY) and the select_by_architecture fallback.

Ported (CPU-verified: import + card/program build + registry resolve + denoise
loop runs end-to-end on the CPU toy backend; GPU load/run is BRINGUP):
  * cosmos25       — Cosmos-Predict2.5 (flow-match, per-frame plain-sigma timestep;
                     reuse FLOW_MATCH_STEP; Reason1/Qwen2.5-VL encoder adapter)
  * hunyuan_video  — HunyuanVideo (reuses WanDenoiseLoop; dual LLaMA+CLIP encoders;
                     Hunyuan VAE scaling_factor) + FastHunyuan variant
  * hunyuan_video15 — HunyuanVideo 1.5 (480p/720p cards)
  * longcat        — LongCat-Video T2V/I2V/VC
  * sd35           — SD3.5 MMDiT (flow-match, image; triple-encoder joint embed +
                     pooled_projections)
  * gen3c          — GEN3C (EDM; 82ch pose-buffer DiT; camera/depth -> BRINGUP)
  * kandinsky5     — Kandinsky 5.0 T2V Lite
  * flux2          — FLUX.2 dev/klein (MMDiT, image; gated weights -> BRINGUP)
  * stable_audio   — Stable Audio Open (audio modality)
  * hunyuangamecraft, hyworld, lingbotworld, matrixgame2, matrixgame3 — interactive
                     world models; t2v/degenerate path CPU-verified, action/camera/
                     memory conditioning is BRINGUP (needs request-API extension)

Adapters declared via ComponentSpec.adapter (the ac29750b enabler) so each port
adds only NEW files (recipe package + per-arch torch_<arch>.py + optional facade
stub) — zero shared-file edits. Registry resolves all 31 bucket-C HF ids by exact
id + 14 architecture fallbacks; no regression (cosmos2/wan/ltx2 unchanged).
v2 mini green (211 passed, 2 skipped); pre-commit clean (per-file/registry).
2026-07-04 17:15:34 +00:00
SolitaryThinker bfbd90ea3a [feat] v2: Cosmos-Predict2-2B-Video2World port (EDM-Karras) — reference bucket-C recipe
First net-new architecture ported via the self-contained recipe-package pattern
(card-declared adapter + new loop, no shared-dispatch edit):

* CosmosDenoiseLoop (v2/recipes/cosmos2/loop.py): EDM preconditioning folded into
  a flow-match Euler integrator. Faithful port of CosmosDenoisingStage — Karras
  sigma schedule (rho=7, sigma_max=80 -> sigma_min=0.002, terminal clamp), latent
  init randn*sigma_max, per-step c_in/c_skip/c_out (sigma_data=1) -> x0, CFG in x0
  space, x0 -> velocity (x-x0)/sigma, FLOW_MATCH_STEP. video2world frame-replace
  conditioning threaded but inert for the t2v preset.
* build_karras_sigmas helper added to v2/loop/sampler.py.
* CosmosDiT + CosmosT5Encoder adapters (v2/platform/backends/torch_cosmos.py),
  declared on the card via ComponentSpec.adapter (the ac29750b enabler) — DiT
  returns raw EDM output + builds the mandatory zero condition/padding masks + fps;
  T5 uses the raw last_hidden_state (no Wan zero-pad). Reuses the WanVAE adapter.
* card/program/registry (HF id nvidia/Cosmos-Predict2-2B-Video2World + arch
  fallback on CosmosTransformer3DModel) + COSMOS_NEG prompt + SamplingDefaults
  (35 steps, gs 7, 704x1280, 93f, 16fps).

CPU-verified: Karras schedule, card/program build, registry resolve (id + arch),
EDM loop runs end-to-end on the CPU toy backend. GPU load/run is BRINGUP.
v2 mini green (211 passed, 2 skipped); pre-commit clean.
2026-07-04 17:15:34 +00:00
SolitaryThinker 5fd6e23e30 [feat] v2 torch backend: ComponentSpec.adapter — card-declared per-arch TorchComponent
A new architecture can declare its own torch adapter on the card
(ComponentSpec.adapter="module:Class") instead of editing the shared _make_dit/
_make_vae/_make_text_encoder dispatch. _explicit_adapter() constructs it as
cls(module, *extra, device=, dtype=) and short-circuits the built-in Wan/LTX2
class-name dispatch when set. This makes each bucket-C port a self-contained
recipe package (card + adapter module + loop + program) with no shared-file edit
-> conflict-free parallel porting. Unset -> unchanged built-in dispatch.

CPU mini green (211 passed, 2 skipped).
2026-07-04 17:15:34 +00:00
Will Lin fc3332550c [feat] v2: Wan2.2-I2V-A14B (MoE i2v) — combine boundary-routed experts + i2v conditioning
Reuses everything: 2 WanTransformer3DModel experts + BoundaryTimestepRouting (from the A14B MoE pattern),
the CLIP image encoder + first-frame [mask|cond] conditioning + the i2v program (from the Fun-InP i2v
port), and the shared WanDenoiseLoop (i2v hooks + the boundary expert). No new adapter. CPU-verified: the
toy MoE i2v runs end-to-end (2 experts + boundary + conditioning -> finite video); resolves with i2v caps.
Structural (GPU-pending: 2x14B, like the A14B T2V). Wan family now largely covered (T2V 1.3B/14B/TI2V-5B/
A14B, causal SF, i2v 1.3B/14B/A14B). CPU mini 211/2.
2026-07-04 17:15:34 +00:00
Will Lin 7464ef8308 [feat] v2: register Wan2.1-I2V-14B 480P/720P (reuse the GPU-verified i2v card)
The 14B i2v variants reuse the Wan2.1 i2v card/path proven on Fun-1.3B-InP — just per-variant params
(480P flow_shift 3.0 / 480x832, 720P flow_shift 5.0 / 720x1280). Registry resolution + caps verified;
specific 14B weights GPU-pending (same generic Wan i2v loader path that Fun-InP validated). i2v cluster
now supported; roadmap updated (11 models ported).
2026-07-04 17:15:34 +00:00
Will Lin 4a56274d10 [feat] v2: Wan2.1 i2v port (Wan2.1-Fun-1.3B-InP) — CLIP encoder + first-frame conditioning, GPU-verified
Real image-to-video, unlocking the i2v cluster. v2/recipes/wan21/i2v.py: CLIP image-encode -> the DiT's
encoder_hidden_states_image; first-frame VAE conditioning + a 4-channel mask -> the 20ch [mask|cond] that
the Wan adapter concatenates with the 16ch noise -> the 36ch i2v DiT input (mirrors fastvideo's
ImageEncodingStage + ImageVAEEncodingStage; v2's WanVAE.encode already applies the matching (z-mean)/std).
Reuses the shared WanDenoiseLoop (its None-default i2v hooks) + the Wan torch adapter unchanged. Adds
ToyImageEncoder + the image_encoder checkpoint subfolder stamp. Registered Wan2.1-Fun-1.3B-InP.

GPU-verified: loads via the generic Wan loader (1.56B, no param-mapping issue), runs the full i2v
conditioning, produces real video (3,9,256,384, std 0.44, finite, motion 0.041). CPU mini 211/2 (T2V
unregressed). BRINGUP: visual confirmation that the output follows the conditioning image is
human-in-the-loop.
2026-07-04 17:15:34 +00:00
Will Lin 9ede9af123 [feat] v2 Wan loop+adapter: optional i2v conditioning hooks (cond concat + CLIP context); T2V unchanged
Threads i2v conditioning through the SHARED WanDenoiseLoop with zero T2V risk: init() reads optional
slots i2v_cond (the [mask|cond] latent) + i2v_img_embeds (CLIP) into scratch; _velocity passes them to
the dit (context=, cond=); WanDiT concats cond (16->36ch) and uses the embeds as
encoder_hidden_states_image; capture is disabled only when i2v conditioning is present. For T2V both are
None -> the dit call, CFG, and cudagraph capture are byte-identical (CPU mini 211 pass, 2 skip — no
regression). ToyDiT accepts+ignores cond (image-conditioning is a GPU-path concern). Completes the i2v
backend seam; the program's mask+cond construction + the Wan i2v card/registry + GPU verify follow.
2026-07-04 17:15:34 +00:00
Will Lin 3d405f6dfa [feat] v2 torch backend: CLIP image-encoder adapter + image_encoder component kind (i2v groundwork)
Adds CLIPImageEncoder (encode_image -> the DiT's encoder_hidden_states_image) + the generic builder's
image_encoder maker (ImageEncoderLoader + ImageProcessorLoader), registered as the cuda 'image_encoder'
kind. Mirrors fastvideo's ImageEncodingStage. CPU-verified (component-kinds + lazy invariant, 211/2);
GPU path marked BRINGUP/written-not-run (processor subfolder + dtype to confirm on a real i2v checkpoint),
matching how the rest of the torch backend was originally landed. Reusable by the Wan i2v cluster + many
bucket-C models (Hunyuan/Cosmos i2v). Next i2v increments: the mask+cond latent construction + the
concat-into-DiT-input loop, then the card + registry + GPU-verify with Wan2.1-Fun-1.3B-InP.
2026-07-04 17:15:34 +00:00
Will Lin d891771ba3 [revert] v2: drop FastWan/VSA registry entries — generic Wan loader can't map their gated-attn params
GPU verification (Wan2.1-Fun... no: FastWan2.1-T2V-1.3B) failed at load: 'Parameter blocks.0.to_gate_compress.bias
not found in custom model state dict' — the FastVideo/* DMD-distilled checkpoints carry gated-attention
params (to_gate_compress) that the generic WanTransformer3DModel loader can't map. v2/registry.py's
select_by_architecture ALREADY rejects WanDMDPipeline for exactly this reason; my explicit ModelEntry
wrongly bypassed it. Reverted FastWan (1.3B + 14B-480P) and the unverified VSA-14B alias (same FastVideo/*
risk). Kept the official Wan2.1-T2V-14B (standard weights, same loader path as the GPU-verified 1.3B).

Lesson recorded in V2_PORTING_STATUS.md: FastWan/Turbo/VSA need a param-mapping fix (like LTX-2.3 did),
not just a schedule — bucket-C-effort. GPU-verify every port before claiming support. CPU mini 211/2.
2026-07-04 17:15:34 +00:00
Will Lin 67dea39052 [feat] v2: alias FastVideo/Wan2.1-VSA-T2V-14B-720P to the Wan-14B card (bucket B)
Same WanTransformer3DModel arch (VSA is an attention-backend choice, not a weight/arch difference); v2
runs dense TORCH_SDPA, so it resolves to build_wan_t2v_14b_card. Registry resolution verified; the GPU
forward path is the 1.3B-proven Wan adapter (specific 14B/VSA weights not separately GPU-run).
2026-07-04 17:15:34 +00:00
Will Lin 5f1d2ef7d2 [feat] v2: port Wan2.1-T2V-14B (bucket B) + document the all-models backlog
First bucket-B port toward 'support every fastvideo model': Wan2.1-T2V-14B reuses the Wan recipe +
torch adapter unchanged (same WanTransformer3DModel/AutoencoderKLWan/UMT5) — only a registry entry +
build_wan_t2v_14b_card (720p, flow_shift 5.0) + SamplingDefaults differ. Without the entry the arch
fallback would give it the 1.3B 480p defaults; the explicit entry gives 50 steps / 720x1280.

Also recorded the full backlog in V2_PORTING_STATUS.md: 63 fastvideo models = 8 ported / 21 bucket-B
(reuse Wan/Causal/LTX2 arch — registry+recipe+defaults, no new adapter) / 34 bucket-C (13 new
architectures needing a TorchComponent adapter). Updated the stale 'how to add a model' steps to the
post-redesign structure (v2/recipes/, torch_backend.py, SamplingDefaults). CPU mini 211 pass, 2 skip.
2026-07-04 17:15:34 +00:00
Will Lin 440b99523e [refactor] v2 torch backend: TorchComponent base + one generic builder + v2.* facade (Phase 1b/1c)
Addresses the adapter-setup pains: collapses the torch_cuda(trampolines)/torch_adapters/torch_ltx2 split
+ 11 near-identical adapter classes + 6 build_torch_* builders into:
- v2/platform/backends/torch_backend.py: a TorchComponent base centralizing .to/.eval, the numpy<->torch
  marshalling (ONE place), the set_forward_context wrap, and the weight surface; thin per-model subclasses
  (WanDiT/LTX2DiT/WanVAE/LTX2VAE/T5Encoder/Gemma/LTX2Upsampler/LTX2AudioVAE/LTX2Vocoder) carrying only
  forward semantics; and ONE build_component(spec) dispatching by spec.kind via _MAKERS.
- torch_cuda.py: registers that single generic builder for all 6 cuda kinds (no per-kind trampolines).
- v2 owns its namespace via re-export STUBS (facade, marked '# STUB'): v2/forward_context, v2/fastvideo_args,
  v2/distributed, v2/loader (the load_component seam), v2/api, v2/models/{dits,audio,upsamplers}/*. All v2
  code imports v2.*; 'from fastvideo' now lives ONLY in those 8 stub files -> a future per-module vendored
  cutover swaps a stub body, no caller changes. No divergence (stubs run fastvideo's live code).
- Deleted torch_adapters.py + torch_ltx2.py.

Verified: CPU mini 210 pass/2 skip; lazy invariant (platform load imports no torch); GPU bit-parity LTX-2.3
T2VS (audio std 0.04304, identical to pre-redesign) + Wan2.1 (video std 31.94, motion 5.997).
2026-07-04 17:15:34 +00:00
Will Lin 1cb2b4e84c [feat] v2: per-model sampling defaults on ModelCard (Phase 1a)
v2 had no per-model defaults — generate_video hardcoded 30 steps/25 frames/480x832/cfg5/16fps for
every model, badly wrong for e.g. LTX-2 distilled (wants 8 steps @1024x1536) or Wan2.2-TI2V (704x1280@24fps).

- New SamplingDefaults dataclass + ModelCard.sampling_defaults (v2/card/specs.py), exported from v2.card.
- Populated all 7 supported cards from fastvideo's InferencePreset defaults (steps/guidance/HxW/frames/fps
  + per-modality guidance for LTX-2.3 A/V). Negative prompts copied verbatim into v2/recipes/_prompts.py
  (recipe DATA, not model code -> v2-owned, no fastvideo import).
- VideoGenerator stores the resolved card; generate_video applies card defaults with precedence
  kwargs > SamplingParam > card > generic fallback (pure _resolve_default helper, unit-tested).
- test_sampling_defaults.py: per-card values + precedence (incl. empty-neg-prompt edge). CPU mini 210 pass, 2 skip.
2026-07-04 17:15:34 +00:00
Will Lin 51898f48e9 [refactor] v2: rename models/ (recipe layer) -> recipes/; move toy backend -> platform/backends/toy.py
Frees v2/models/ to become the vendored-architecture namespace that mirrors fastvideo/models
(part of making v2 self-contained / able to replace fastvideo). The v2 recipe layer (per-family
card.py/loop.py/program.py + common.py + the build_*_engine re-exports) is the recipe, not the
architectures, so it moves to v2/recipes/. The pure-numpy toy/parity implementations (ToyDiT etc.)
move from v2/models/backend.py to v2/platform/backends/toy.py (alongside cpu.py/accel.py/torch_*).

Mechanical: all imports are absolute, so v2.models.<x> -> v2.recipes.<x> and v2.models.backend ->
v2.platform.backends.toy across v2/ + examples/ (89 files, 178 refs). No behavior change.
CPU mini green (202 passed, 2 skipped); all v2 files compile.
2026-07-04 17:15:34 +00:00
SolitaryThinker 4e331eb7ff [refactor] v2: use absolute imports (v2.*) everywhere instead of relative
Mechanical conversion of every relative import under v2/ to an absolute v2.* path
(from .x / ..x / ...x -> from v2.<pkg>.x) so imports are unambiguous, grep-able, and
stable when code is copied/moved between entrypoints (VideoGenerator, CLI, server).

Surgical prefix-only rewrite: only the 'from <dots><module>' prefix changed — import
names, parentheses, multi-line formatting, comments, and ordering are byte-for-byte
preserved (no collapsing, no reorder, no unrelated reformatting).

- 471 imports across 125 files; v2/tests/ was already absolute (untouched).
- Validated: all 125 files compile, every 'from v2.* import' target resolves to a real
  module/package, zero relative imports remain (full sweep), CPU mini suite green
  (202 passed, 2 skipped).
2026-07-04 17:15:34 +00:00
SolitaryThinker c6d2976fc2 [feat] v2 VideoGenerator: A/V convenience path (generate_video -> T2VS -> mp4 + 24kHz wav)
Makes the 'Full A/V' LTX-2.3 deliverable reachable from the user-facing entrypoint, not just the
engine. A model advertising TEXT_TO_VIDEO_SOUND (LTX-2.3) now auto-issues a T2VS request, so generate()
/ generate_video() return BOTH modalities in one joint pass:
- VideoGenerator stores the resident instance + a supports_av flag (from card.capabilities); generate()
  gains want_audio (None=auto-by-capability, True/False to force) and routes T2V vs T2VS+{video,audio}.
- _result saves the stereo waveform as a sibling .wav at the vocoder's REAL rate (24000) — read off the
  built audio_vae adapter (TorchLTX2AudioVAE.sample_rate = Vocoder.output_sample_rate), since the
  AudioArtifact default rate is a placeholder. Populates GenerationResult.audio/.audio_sample_rate and
  extra['audio_path']. scipy IEEE-float WAV; [channels,samples] auto-transposed.
- Rewrote v2_basic_ltx2_3_distilled.py: registry routes to build_ltx2_3_card (its own joint T2VS A/V
  card, not the LTX-2 base/2-stage card); the example prints both the mp4 and the wav.
- GPU-verified via the convenience API: ev.mp4 + ev.wav (24000 Hz, stereo 61920x2, nonzero, std 0.043).
  CPU mini green (202 passed, 2 skipped); engine/program/toy paths untouched (test_ltx2_av pins 44100).
2026-07-04 17:15:34 +00:00
SolitaryThinker 6096b00aeb [feat] v2 LTX-2.3 T2VS GPU-verified: audio VAE/vocoder wiring + dual-connector audio fix
The full joint text->video+audio LTX-2.3 path now generates on the real 18.99B model:
- GPU audio components: build_torch_audio_vae (AudioDecoderLoader -> LTX2AudioDecoder, chains the
  vocoder) + build_torch_vocoder (VocoderLoader -> LTX2Vocoder); registered the 'audio_vae'/'vocoder'
  cuda component kinds; stamped their checkpoint subfolders (_WAN21_SUBFOLDERS).
- Fix: TorchGemma.encode_av must pass output_hidden_states=True — the 2.3 connector's SEPARATE audio
  projection lives in hidden_states[0] only then (gemma.py:703); without it the audio text fell back to
  the video embedding (4096 vs 2048 -> audio cross-attn shape mismatch).
- GPU-verified T2VS: video (3,33,256,384, std 0.68) + audio (stereo 2x61920 @24kHz, nonzero, std 0.059).
- test_torch_backend: cuda component kinds now include audio_vae + vocoder. CPU suite green (202+2).
2026-07-04 17:15:34 +00:00
SolitaryThinker 1d23399d81 [feat] v2 LTX-2.3 T2VS: single-stage joint audio+video card/loop/program (CPU-verified)
Makes LTX-2.3 a first-class, faithful card (was wrongly merged into the single-stage base):
- LTX23DenoiseLoop (loop.py): single-pass joint A/V denoise — one DiT forward per step cross-attends
  video<->audio via the adapter's (v_vel,a_vel) return; full-res video latent + a [8,T,16] audio latent;
  distilled few-step schedule (BASE_SIGMAS). Video-only when no audio requested.
- build_ltx2_3_card (model_id 'ltx2.3-distilled' — the name now correctly names the REAL 2.3): 5
  components incl. audio_vae (AudioDecoder) + vocoder (required_for t2vs, optional_for t2v); caps
  T2V + T2VS. build_ltx2_3_program: dual-connector text-encode -> joint denoise -> video + audio decode.
- registry.py routes FastVideo/LTX-2.3-Distilled-Diffusers -> this card (split from the base entry).
- Toy support: ToyTextEncoder.encode_av (separate video/audio text), ToyDiT joint A/V (audio now a
  keyword-only arg so positional  callers like the talker are unaffected), channel-agnostic
  ToyAudioVAE (np.resize identity for the existing 2-stage T2VS).
- CPU-verified: toy T2VS -> video+audio (8 steps), T2V -> video-only; CPU suite green (202+2).
GPU audio-VAE/vocoder loaders + the real-T2VS GPU verify are the next step.
2026-07-04 17:15:34 +00:00
SolitaryThinker fefcd415ff [feat] v2 LTX-2 adapters: A/V foundation (joint DiT forward + dual text connector + audio decode)
Foundation for the LTX-2.3 T2VS port (card/loop/program wiring + GPU verify to follow):
- TorchLTX2DiT.__call__ gains an optional joint audio path: pass audio_latent[8,T,16] + audio_text and it
  feeds audio_hidden_states/audio_encoder_hidden_states/audio_timestep/audio_sigma in ONE forward
  (LTX-2.3 cross-attends video<->audio) and returns (video_velocity, audio_velocity). Video-only call is
  byte-for-byte unchanged (audio_latent=None).
- TorchGemma.encode_av returns the SEPARATE (video_text, audio_text) projections from the 2.3 connector
  (video=last_hidden_state, audio=hidden_states[0]); 2.0 returns them equal.
- TorchLTX2AudioVAE (AudioDecoder -> Vocoder -> waveform@24kHz) + TorchLTX2Vocoder wrapper.
Additive + backward-compatible; CPU suite unaffected (adapters are GPU-lazy).
2026-07-04 17:15:34 +00:00
SolitaryThinker 7ac2ff0d1c [refactor] v2: shared model registry (HF-id primary + arch fallback) for all entrypoints
Per review: dispatch should be a directly-mapped HF-string -> card registry (like fastvideo), shared
by every entrypoint (VideoGenerator + a future CLI / server), not buried in the generator.

- New v2/registry.py mirrors fastvideo's fastvideo/registry.py hybrid resolution: (1) exact HF repo id
  in an explicit ModelEntry registry [PRIMARY — correct per-model card/capabilities, and the only way to
  split same-architecture capability variants like Wan2.1 T2V vs the i2v 'InP' 1.3B], (2) short repo-name
  match, (3) architecture inference [FALLBACK — local paths / unregistered repos]. resolve(model_path[,
  root]) is the single shared entry point.
- video_generator.py: moved _read_arch_signature/_select_builders into the registry; from_config now
  calls resolve() — registered ids resolve with no config read, else arch inference on a cheap *.json
  snapshot. Reconciles the earlier 'no brittle table' refactor with the 'map hf string -> card' ask: one
  clean registry + a fallback, not three coupled structures.
- Verified: registry resolves exact-id / short-name / arch-fallback / unregistered correctly; CPU suite
  green (202+2); wan21 GPU smoke generates via the new resolve path.
2026-07-04 17:15:34 +00:00
SolitaryThinker fa9c58b419 [fix] v2: name LTX-2 cards by architecture (2-stage vs single-stage) + Wan2.1 is T2V-only
Addresses the 'how is ltx2 separate from ltx2.3' confusion + a wrong capability:

- LTX-2 cards renamed by ARCHITECTURE (the version labels did not map to it): build_ltx2_card model_id
  'ltx2.3-distilled' -> 'ltx2-2stage-distilled' (two-stage base->upsample->refine; serves the
  upsampler-having FastVideo/LTX2-Distilled-Diffusers); build_ltx2_base_card 'ltx2.base' ->
  'ltx2-single-stage' (one loop; serves Davids048 base + the single-stage FastVideo/LTX-2.3-Distilled,
  which has NO spatial_upsampler). Dispatch already splits on has_spatial_upsampler. Updated the
  model-id refs in the mini's tests/examples.
- Wan2.1 base is T2V-only: dropped the wrong Capability.IMAGE_TO_VIDEO + narrowed components'
  required_for to {t2v} (i2v is the separate InP variant; v2 has no i2v path yet). build_wan21_card is
  shared by wan21 + wan2.2-ti2v; the A14B card was already T2V-only.
- CPU suite green (202 passed, 2 skipped).
2026-07-04 17:15:34 +00:00
SolitaryThinker 0662b42510 [docs] v2: LTX-2 base/2.3 GPU-verified on rebuilt x86 stack + remaining-port mechanisms
- LTX-2 base (Davids048) and LTX-2.3-Distilled both generate real video (inter-frame motion 4.5 / 6.6)
  via the single-stage base card — moved to Working (7 models now verified).
- Environment: the aarch64 venv was rebuilt for x86 (torch 2.11.0+cu128) + re-validated (CPU suite green,
  wan21 + LTX-2 base/2.3 generate).
- Remaining Wan+LTX-2 ports documented with concrete mechanisms: Wan2.2-i2v (SigLIP image_encoder +
  VAE-encode first-frame + concat-mask -> larger-in_channels i2v DiT), TurboWan (RCMScheduler consistency
  loop), Lucy-Edit (Wan v2v via VideoVAEEncodingStage), FastWan (VSA, env-blocked: no nvcc).
2026-07-04 17:15:34 +00:00
SolitaryThinker ae6d8085de [docs] v2: LTX-2.3 example + roadmap (A14B offload working, base/2.3 ported, env status)
- v2_basic_ltx2_3_distilled.py: LTX-2.3-Distilled routes to the single-stage base card (no
  spatial_upsampler) via the arch dispatch; pass few steps for the distilled schedule.
- V2_PORTING_STATUS.md: A14B moved to Working (CPU expert offload, 60GB peak); LTX-2 base + 2.3 added
  (code-complete, GPU re-verify pending); Environment-status note on the mid-session aarch64->x86 host
  reschedule that blocks GPU re-verify.
2026-07-04 17:15:34 +00:00
SolitaryThinker d0648ba8d9 [feat] v2: Wan2.2-A14B MoE CPU offload (fits 1 GPU) + LTX-2 base/2.3 single-stage port
Within the bounded Wan+LTX-2 scope:

- Wan2.2-A14B MoE now GENERATES on a single 80GB GPU via CPU offload (TorchWanDiT offload_group):
  the two 14B experts live on CPU and only the active one is swapped onto the GPU at the boundary-
  timestep transition (a single swap, not per-step). GPU-verified: 60GB peak (vs 79GB OOM), produced
  wan22_a14b_lion.mp4 (17x480x832, std 54.7, motion 8.36). Single-expert Wan stays resident.

- LTX-2 base (single-stage) port: build_ltx2_base_card + build_ltx2_base_program reuse the LTX-2
  adapters at FULL latent res with a request-driven many-step flow-match (LTX2DenoiseLoop full_res/
  request_steps/base_flow_sigmas; distilled base/refine path preserved via False defaults). The SAME
  single-stage card serves LTX-2.3-Distilled (also single-stage: no spatial_upsampler) — dispatched by
  the new has_spatial_upsampler discriminator in _select_builders. v2_basic_ltx2.py added; VideoGenerator
  gains shutdown() for API parity.

- Fixes: from_config 'os' scoping (shadowed module import); LTX-2 upsampler per_channel_statistics
  source (the AE's .decoder, not the top-level module).

Verification status: A14B offload, the upsampler, and the arch-dispatch refactor were GPU-verified
earlier this session. LTX-2 base/2.3 are CPU-verified (cards/programs build, dispatch routes, schedule
correct); their GPU smoke tests were pending when the box was rescheduled aarch64->x86 mid-session,
which broke the aarch64 venv (numpy/torch unrunnable on x86) — GPU re-verify blocked on the env.
2026-07-04 17:15:34 +00:00
SolitaryThinker e10828346f [feat] v2: architecture-driven dispatch + real LTX-2 upsampler + Wan2.2-A14B MoE card
Two reviewer asks + the next Wan port, all within the bounded Wan+LTX-2 scope:

1. Architecture-driven dispatch (replaces the HF-id table + substring fallback): from_config reads the
   checkpoint's pipeline/transformer/VAE class names (+ z_dim, transformer_2) and picks the v2 card via
   _select_builders — mirroring fastvideo's get_pipeline_config_cls_from_name. Resolves local paths /
   renamed repos / new distilled variants of a known arch with no table edits, and cleanly REJECTS
   FastWan (detected by WanDMDPipeline) with a precise message instead of a confusing load crash.

2. Real LTX-2 spatial upsampler (was a nearest-neighbor np.repeat stand-in): new 'upsampler' component
   kind -> TorchLTX2Upsampler wraps the real LTX2LatentUpsampler and applies the repo's upsample_video
   (un_normalize via the VAE decoder's per_channel_statistics -> learned 2x upsample -> normalize). CPU
   keeps ToyUpsampler (np.repeat) via the factory terminal, so the program calls
   component('spatial_upsampler').upsample(...) on both backends with no device branch. GPU-verified:
   9x512x768, std 70.5, motion 6.51.

3. Wan2.2-T2V-A14B MoE card (build_wan22_a14b_card): two WanTransformer3DModel experts +
   BoundaryTimestepRouting @0.875. GPU-verified that both experts denoise; OOMs in VAE decode on one
   80GB GPU (~70GB resident) — upstream offloads the DiT for MoE; documented as offload-blocked.

CPU suite green (202 passed, 2 skipped). FastWan root-caused (non-strict load of VSA gate_compress +
VSA not built); roadmap (V2_PORTING_STATUS.md) updated with the bounded scope + per-model status.
2026-07-04 17:15:34 +00:00
SolitaryThinker 655f362cf4 [feat] v2 port: Wan2.2-TI2V-5B (T2V) — 4th GPU-verified model
- Wan2.2-TI2V-5B reuses the Wan adapters (WanTransformer3DModel / AutoencoderKLWan / UMT5); deltas
  are the higher-compression VAE geometry (z_dim=48, 16x spatial, 4x temporal) and 480p flow-shift 5.0.
  The DiT forward accepts a scalar timestep (1D path), so no per-frame expand_timesteps for pure t2v.
- WanDenoiseLoop / build_wan21_card gain optional geometry params (latent_channels/spatial_ratio/
  temporal_ratio) defaulting to Wan2.1 (16/8/4) -> wan21 path unchanged; build_wan22_ti2v_card sets
  48/16/4. Registered as family 'wan2.2-ti2v' in VideoGenerator; v2_basic_wan2_2_ti2v.py added (T2V).
- Verified: 25x448x768 mp4, std 62.3, inter-frame motion 4.89 (coherent). CPU suite green (202+2skip).
- Corrected the now-disproven FastWan='wan21 reuse' mapping (its DMD checkpoint to_gate_compress param
  mapping differs); roadmap updated (TI2V-5B working; A14B MoE + I2V remain).
2026-07-04 17:15:34 +00:00
SolitaryThinker 3541e81d66 [feat] v2 VideoGenerator: convenience API (from_pretrained/generate_video) + porting roadmap
- VideoGenerator gains the convenience surface most basic examples use: from_pretrained(model,
  num_gpus/*_cpu_offload/...) and generate_video(prompt, sampling_param=, **kwargs) -> result, on top
  of the typed from_config/generate. Accepts SamplingParam.
- v2 examples matching the upstream convenience-API examples for the verified models: v2_basic.py
  (Wan2.1) and v2_basic_self_forcing_causal.py (SF-causal).
- V2_PORTING_STATUS.md: honest per-family roadmap. Working: wan21, wan_causal, ltx2-distilled. Each
  further model needs per-model work (FastWan: WanDMD to_gate_compress param mapping; TurboWan: RCM
  consistency sampler; Wan2.2: MoE card; LTX2 base/i2v; new families: new cards/adapters; gated Flux2 /
  local GEN3C / audio StableAudio / interactive MatrixGame blocked in this env).
2026-07-04 17:15:34 +00:00
SolitaryThinker f51497ee6d [feat] v2 VideoGenerator: typed fastvideo.api entrypoint over the v2 engine
Mirrors fastvideo.entrypoints.VideoGenerator (from_config(GeneratorConfig) -> generate(
GenerationRequest) -> GenerationResult.video_path), reusing the OFFICIAL fastvideo.api config classes
so a basic_dmd_new_api.py-style script differs only by importing VideoGenerator from v2.
- model_path -> v2 card registry (Wan2.1 / FastWan -> wan21; SFWan -> wan_causal; LTX2 -> ltx2);
  snapshot_download + stamp_wan21_checkpoints + Engine(cuda) + program; generate maps SamplingConfig
  -> DiffusionParams -> make_request -> eng.run, saves the [C,T,H,W] decode as an mp4.
- Lazy v2.__getattr__ keeps 'import v2' torch-free (verified) so the CPU mini stays green (202+2).
- examples/inference/basic/v2_basic_new_api.py runs all three GPU models through this API.
- Single-GPU, resident, TORCH_SDPA (EngineConfig offload/num_gpus>1/VSA accepted for parity, not applied).
verified: wan21 from_config->generate->mp4 (frames (5,256,256,3) uint8, video_path written).
2026-07-04 17:15:34 +00:00
SolitaryThinker d8af2e60d2 [feat] v2 ltx2: two-stage distilled GPU bring-up (LTX2Transformer3DModel 18.88B)
Official FastVideo/LTX2-Distilled-Diffusers. New torch_ltx2.py adapters (build_torch_* dispatch on
class):
- TorchLTX2DiT: patchify-internal; per-token timestep ones(B,tok,1)*sigma (sigma direct) + per-sample
  video_sigma; DiT predicts x0 so the adapter returns velocity=(x_t-x0)/sigma for the v2 flow-match step.
- TorchLTX2VAE: CausalVideoAutoencoder decode (internal per-channel un_normalize).
- TorchGemma: LTX2GemmaTextEncoderModel (Gemma + feature-extractor + connectors) -> last_hidden_state.
- ltx2 loop: real 128-ch latent geometry on cuda (32x spatial / 8x temporal; half-res base, 2x upsample).
e2e two-stage (8+3 steps) -> coherent, high-quality video (surfers at sunset), (3,9,512,768). NOTE:
still uses the v2 program's np.repeat upsampler between stages (the refine regenerates from noise so
output is faithful-quality); real LTX2LatentUpsampler swap-in is a follow-up.
2026-07-04 17:15:34 +00:00
SolitaryThinker f79919ba8a [feat] v2 wan_causal: causal DiT (CausalWanTransformer3DModel) GPU bring-up
Official SF checkpoint wlsaidhi/SFWan2.1-T2V-1.3B-Diffusers (reuses TorchWanVAE + TorchT5Encoder).
- TorchWanDiT detects the causal transformer: ignores the chunk_rollout loop's latent `context` (the
  real model conditions across chunks via an internal kv_cache, not a forward arg) -> dispatches to
  full-attention _forward_train, and passes a per-latent-frame timestep [B, num_frames] (the causal
  block asserts a per-frame temb), uniform per chunk.
- chunk_rollout real geometry on cuda (16ch; chunk_size latent frames; 8x spatial).
e2e: chunk_rollout over the SF student -> coherent video (cat in a garden), (3,21,480,832). Fidelity
gap (artifacts): the v2 loop's per-chunk few-step sampling != the official kv-cache streaming + SF
schedule (a follow-up).
2026-07-04 17:15:34 +00:00
SolitaryThinker 1594f8e6be [fix] v2 tests: make no-torch-import guards GPU-aware (skipif torch installed)
The cuda-availability probe imports torch by design; the no-torch-import invariant is only
verifiable when torch is absent. skipif torch installed -> green on GPU box (202 passed, 2 skipped),
still enforced in torchless CI.
2026-07-04 17:15:34 +00:00
SolitaryThinker 64cadaa0bf [feat] v2 wan21: backend-aware latent geometry + checkpoint stamping
- latent_shape(req, model): real Wan geometry (16ch; (T-1)//4+1, H/8, W/8) on the cuda backend;
  toy stand-in stays on accel/cpu.
- stamp_wan21_checkpoints(card, model_root): map a root (local dir or HF id) onto the 3 components'
  ComponentSpec.checkpoint; build_wan21_card(checkpoint_root=...) optional.
2026-07-04 17:15:34 +00:00
SolitaryThinker 750fc1245b [feat] v2 cuda backend: real fastvideo construction + risk A-E fixes (Wan2.1 verified on H100)
Take the written-not-run torch adapters to runs-and-generates on 1x H100 (aarch64):
- A: FastVideoArgs.from_kwargs(model_path=root) builds the real pipeline_config; single-GPU dist
  init; load each component from its subfolder; tokenizer from the sibling <root>/tokenizer.
- B/C: DiT forward wrapped in set_forward_context(attn_metadata=None) (SDPA dense path);
  timestep=sigma*1000 + bare-velocity output confirmed.
- D: VAE decode denormalizes z*std+mean; removed the double-mean (it re-added
  shift_factor==latents_mean) that washed out the video.
- E: UMT5 from config; text embeds zero-padded to text_len (Wan t5_postprocess_text) - the fix
  that took output from a dark blur to a coherent prompt-matching scene.
- Components run at native precision (DiT bf16, VAE/text fp32); checkpoint check before dist init.
2026-07-04 17:15:34 +00:00
SolitaryThinker 7725998b0c [docs] add v2/HANDOFF.md for GPU-side bring-up of the torch backend
Orientation + process doc for an agent on a GPU branch: the 6 commits already
landed, the files to touch, the gating tasks (Risk A FastVideoArgs/checkpoint),
the verification bar (CPU suite stays 204; GPU generation matches a reference via
the SSIM harness), commit/push rules (no Claude co-author; don't rewrite history;
wandb token referenced not embedded), and the gotchas. Points to
GPU_BRINGUP.md for the detailed checklist + risk table.
2026-07-04 17:15:34 +00:00
SolitaryThinker 4b61dedc43 [fix] correct GPU adapters against real fastvideo API (cross-check findings)
Adversarial cross-check of the written-not-run torch adapters against the real
fastvideo source confirmed the interface contracts (DiT returns bare velocity
tensor; timestep=sigma*1000; encode().mode() + bare decode; .last_hidden_state;
no fused solver kernel) but caught a wrong construction layer. Fixed in code:

- Construction: WanTransformer3DModel / AutoencoderKLWan have NO from_pretrained.
  Replace it with the real FastVideo loaders (TransformerLoader / VAELoader /
  TextEncoderLoader + TokenizerLoader, each load(model_path, fastvideo_args)).
  The loader resolves the class from the checkpoint config — so UMT5-vs-T5 is
  chosen correctly instead of hardcoded (was BLOCKER #1/#4/#5).
- Text encoder: wrap the forward in set_forward_context(...) — the (U)MT5
  attention reads global state via get_forward_context(); a bare call mis-encodes
  (was BLOCKER #2). Drop the wrong padding="max_length".
- VAE: apply latent normalization the DiT expects — (z-mean)*inv_std on encode,
  inverse on decode, with latents_std stored as its reciprocal; shift_factor
  before decode (was BLOCKER #3). Skipping it yields washed-out video, not error.

The remaining unknowns are genuinely box-dependent (FastVideoArgs fields,
shift_factor placement, exact tokenizer kwargs, FSDP) — GPU_BRINGUP.md reconciled
to mark what's now fixed-in-code vs what still needs the box. 204 CPU tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 1e58d90d02 [feat] real torch/CUDA backend (written-not-run) behind the cuda cells
Implement the GPU backend the substrate was built for: Platform.detect() ->
cuda resolves real torch adapters + torch solver ops instead of the numpy
rungs, with the existing loops/policies/scheduler/training unchanged.

WRITTEN-NOT-RUN: this environment has no GPU/torch, so the torch code is
grounded in the verbatim real fastvideo APIs (DiT forward signature confirmed
from source) but cannot be executed/verified here. It is gated available=False
(CPU mini stays green; importing the backends never imports torch), with every
on-box confirm point marked `# BRINGUP` and an ordered checklist in
platform/backends/GPU_BRINGUP.md.

- torch_adapters.py: TorchWanDiT / TorchWanVAE / TorchT5Encoder wrap the real
  module named by each card's load_id and bridge it to the mini's duck-typed
  surface (numpy<->torch at the boundary; loop math stays numpy fp32). DiT
  weight-surface (copy_from/blend_from/clone) for serving sync; mse_grad_step
  raises (GPU training is a separate workstream).
- torch_kernels.py: flow_match_step / flow_sde_step as plain torch elementwise.
  Grounded conclusion from the kernel audit: fastvideo-kernel ships NO fused
  solver kernel (only attention/norm/quant primitives), so the cuda solver is
  torch, registered at arch generic with an honest source string.
- torch_cuda.py: rewritten as lazy trampolines (torch imported only inside
  builder/kernel bodies). Adds the missing vae + text_encoder cuda components
  (they'd otherwise silently fall back to the toy) and corrects the dishonest
  "fastvideo-kernel:flow_*" labels.
- ComponentSpec.checkpoint: the weights source for the torch adapter (risk A;
  the one field the cards didn't carry). Empty on the CPU toys.
- 7 CPU-verifiable wiring tests: honest registration/sources, torch-free import,
  cuda-resolves-real-cells-not-toy, build-fails-loudly-without-torch.

204 tests pass (CPU). The torch path needs a GPU box to verify (GPU_BRINGUP.md).
2026-07-04 17:15:34 +00:00
SolitaryThinker 47f7a04e09 [feat] static-buffer capture form for the cudagraph step body (Path A)
Close the loudest deferred gap from the cudagraph audit: the capturable step
now binds its I/O to address-stable static buffers (modeling real CUDA static
I/O buffers), instead of allocating fresh arrays per call.

- StaticWorkspace: address-stable buffers allocated once per capture key; bind()
  copies the current step's inputs in place via np.copyto, which RAISES on a
  shape/dtype mismatch — turning the weak peak_activation_bytes proxy into a real
  key-soundness backstop (a step whose shape doesn't fit can't replay an
  incompatible graph). Output written into a static buffer too.
- WorkPlan.graph_fn / graph_inputs: a capturable step exposes its deterministic
  op-structure as graph_fn(model, workspace) reading EVERY per-step input (latent,
  sigmas, conditioning, scale) from the workspace — never from closure over
  per-step data — plus the dict of current values. Loops without both stay on the
  eager path. wan21 factors a shared _velocity() so graph_fn and the eager run
  stay bit-identical.
- Capturer dispatch captures/replays via graph_fn against the keyed workspace;
  the workspace is shared per key on the instance. Correct under the engine's
  synchronous step execution (bind+graph_fn atomic per dispatch, output returned
  as a copy) — proven by the batch-of-N interleave gate running two same-key
  requests through the shared workspace bit-identically. A concurrent/multi-stream
  executor would need a per-stream pool (documented).
- 2 new tests (no-static-form eager-break, static-buffer shape-mismatch raises);
  workspace collision test split into bytes-proxy vs shape-backstop.

197 tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 9b8838834c [feat] piecewise CUDA-graph capture/replay at the step boundary (Path A)
Wire the capture/replay lifecycle into the driven-loop step boundary — the
other half of Path A (hand-fused kernels behind the registry + piecewise
cudagraphs, no compiler). Models and tests the correctness-critical control
logic; replay re-runs the current step thunk (CPU models the lifecycle, not the
GPU speedup).

- GraphCapturer on the instance (cross-request cache), wired into
  RuntimeLoopContext.execute and gated by LoopSpec.graph_capture ==
  "breakable_cudagraph". Capture key = (device, arch, loop, shape_sig,
  resident-weight-versions, graph_key). Eager-break for non-capturable (SDE) and
  interceptor-overridden steps. Never executes a stored thunk, so interleaved
  requests can't smear state.
- WorkPlan.capturable / graph_key. wan21 sets capturable=not sde and folds
  compute dtype (shape_sig.dtype) + CFG branch set + expert + scheduler-precision
  into the key — closing a cross-precision key-collision corruption path a
  non-fp32 build would otherwise hit on a real GPU (audit finding).
- Version-in-key auto-invalidation + real eviction: set_weights_version evicts
  the synced component's graphs (duck-typed, so card/ imports no runtime),
  preventing a GPU graph leak across FlowGRPO's per-iteration syncs.
- register_kernel gains the workspace_bytes capture-safety contract (declared in
  the matrix; cuda cells declare real scratch, numpy reference is 0).
- 11 tests: capture-once/replay-many, eager-break (SDE + override), capture ≡
  pure-eager bit-identical, recapture on shape + weight-version change with
  eviction, eager-loop gating, accel-backend capture, and capturer unit tests
  (key discrimination, eager-break, workspace-collision safety net, eviction).

Honestly deferred (bite a real GPU, not the CPU tests): the static-buffer
refactor of the step body, admission budgeting of capture cost (GRAPH_CAPTURE),
and per-card opt-in beyond wan21 — all documented in cudagraph.py + README.

195 tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 634f0828a2 [feat] route all diffusion loops + RL recompute through the kernel table
Finish the kernel seam across the board so the platform's KernelTable is the
universal solver-dispatch path, not just wan21.

- Loops: ltx2, wan_causal, adapters, adaptive now resolve the flow-match solver
  via model.platform.kernels.get(FLOW_MATCH_STEP) instead of importing the numpy
  sampler directly (wan21 already did). On CPU this is bit-identical (the cpu
  kernel IS the old function); a GPU/accel backend now overrides every loop.
- RL: the FlowGRPO log-prob recompute in unified_rl / joint_multi_rl /
  workflow_rl dispatches FLOW_SDE_STEP through the platform, pinned to the SAME
  kernel the rollout used (C2 kernel-pinning — otherwise the PPO ratio biases on
  a real GPU where rollout and recompute kernels could differ).
- accel backend: add a kind-generic AccelComponent wrapper and override the vae
  component too (text_encoder left unregistered to keep the device->cpu fallback
  demonstrated), closing the "only dit is overridden" gap.
- tests: vae override assertion; a second-loop (wan-causal chunk rollout) parity
  oracle proving accel == cpu bit-identical beyond wan21. README scope updated.

184 tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 2f044c02dd [feat] multi-backend dispatch substrate (device/arch/kernel registries)
Add v2/platform/: the (recipe, runtime) backend membrane that lets CPU, GPU,
and other devices coexist behind one dispatch substrate.

- Two tuple-keyed registries: COMPONENTS(kind, device, variant) for
  weight-bearing components and KERNELS(op, device, arch, variant) for
  stateless primitives, each with an availability predicate and an enumerable
  manifest (declared-but-unavailable cells listed without importing torch).
- Platform: detected (device, arch) owning the device + arch fallback chains
  and a per-platform cached KernelTable; detect() is honest (CPU/numpy unless
  torch+CUDA are actually present). Arch fallback is monotonic — only degrades
  to older/portable archs, never a newer binary-incompatible one.
- Seams wired: ModelInstance.component() -> platform.build_component(spec,self)
  with spec.factory as the numpy terminal rung (existing cards untouched); the
  wan21 denoise thunk dispatches solver ops through model.platform.kernels.
- Three backends: cpu (numpy terminal + parity oracle), accel (pure-python
  stand-in proving cross-device resolution, arch fallback, and the oracle),
  torch_cuda (declared-but-unavailable; no faked GPU).
- test_platform.py: 16 tests — detection, terminal rung, arch-fallback walk +
  monotonicity, device precedence, variant fallback, component override +
  per-kind device fallback, the parity oracle (accel == cpu, bit_identical via
  the C1 ladder), and matrix enumeration without importing torch.

Scope is honestly bounded in the README/docstrings: only wan21-denoise routes
through the kernel table and only the dit kind is overridden today (each a
one-line adoption); the torch/CUDA path and a cudagraph workspace-safety
contract are declared/deferred, not implemented. 183 tests pass.
2026-07-04 17:15:34 +00:00
SolitaryThinker 431f4daddb [feat] Adapter plane, non-linear workflows, RL→distill flywheel
Three more capabilities on distinct untested surfaces (167 tests pass, no new runtime primitive).

A — Adapter plane (§9.19): one base + swappable LoRA/ControlNet adapters, selected per request
(DiffusionParams.adapters); AdapterDenoiseLoop applies each active adapter's velocity delta. Per-request
selection changes output, multi-LoRA composes, ControlNet conditions on a control image, mixed-adapter
requests interleave without smearing, hot-swap changes generation, cache key partitions by adapter stack
(the adapter_versions field, previously declared-only). ToyLoRA/ToyControlNet; models/adapters/. 6 tests.

B — Non-linear workflows (§9.17): ParallelWorkflow (fan-out: one input → N models → merged) and
BestOfNWorkflow (generate N → score with the served reward card → return best; inference-time scaling).
The shapes a linear chain can't express. 4 tests.

D — RL→distill flywheel (§9.18): run_flywheel RL-improves the base (NFT), then distills FROM the RL'd model
(DMD2 teacher = RL'd policy) into a faster card, recording the base→rl→distilled provenance chain in
RecipeSpec.parents. The distilled student is measurably closer to the RL'd teacher than the base; the
distilled card serves few-step. training/flywheel.py. 4 tests.

designv4 §9.17–§9.19 + layout/closing/counts (167 tests, 29 files).
2026-07-04 17:15:34 +00:00
SolitaryThinker 6220d02746 [feat] #8b speculative (draft-verify) decoding — exact + lower-latency AR
The last audit stress test. A cheap draft model proposes K tokens, the target verifies
them in one batched step, and SpeculativeARLoop accepts the matching prefix + one target
correction — a variable accepted-length per round (a ragged AR loop the model owns).

- Exactness: the emitted sequence equals the target's OWN greedy decode for any draft
  quality (every accepted token is one the target would produce; the correction is the
  target's token) — the speedup is free.
- Speedup scales with accept rate: draft-agree 0.3→1x, 0.7→3x, 1.0→4x=K tokens/round
  (fewer verify_rounds, the expensive model's latency steps, for the same output).
- Two components (draft + target) co-scheduled on one resident instance; each round an
  AR_TOKEN WorkUnit.

models/speculative/ (loop+card+program); backend ToyTargetModel/ToyDraftModel with a
shared length-dependent target formula (no degenerate fixed point). 5 tests; full suite
153 passed. designv4 §9.16 + counts (153 tests, 26 files).
2026-07-04 17:15:34 +00:00
SolitaryThinker 9489c6c1dd [feat] Five more stress tests: LTX-2 A/V, weight-sync, served reward, cache-dit, nested workflows
The remaining design_v3 probes (all except 8b speculative decoding). All fit with no new runtime
primitive (148 tests pass).

#6 LTX-2 joint audio+video (§9.11) — LTX-2 declared an audio_vae required_for t2vs but never used
   it; now a single 2-stage denoise carries a synchronized audio latent (conditioned on video),
   applies per-modality CFG (guidance_per_modality), and decodes via video VAE + audio VAE → video +
   audio. Gated on requesting audio, so the T2V path is byte-identical (existing tests untouched).
   ToyAudioVAE; build_ltx2_av_program. 5 tests.

#4 Live weight-sync under in-flight serving (§9.14) — WeightSyncController makes the freeze → drain →
   transfer → bump version + invalidate → resume lifecycle explicit. Tests: a mid-flight swap corrupts
   (the hazard); draining first leaves the in-flight request bit-identical to baseline while a
   post-sync request reflects new weights; transformer-only sync, so the frozen text-encoder cache
   survives. The RL flywheel's hardest correctness. 3 tests.

#5 Reward-model-as-a-served-card (§9.15) — a reward model is a card (scorer + a score loop emitting
   REWARD_BATCH units); ServedRewardScorer drop-in-replaces the numpy scorer so any RL method becomes
   RLHF/RLAIF with no method change. ToyRewardModel; models/reward/. 4 tests.

#7 Content-adaptive control flow (§9.12) — CacheDiTDenoiseLoop (isolated WanDenoiseLoop subclass)
   reuses the cached velocity when predictions barely change (cache-dit skip) and early-exits on
   convergence — variable step count; interleave parity holds across ragged loops. models/adaptive/. 4 tests.

#8a Nested workflows (§9.13) — a workflow stage can invoke another workflow (engine.run routes ids);
   requires/validate recurse; cycles caught at registration + a run-time guard (engine._wf_running).
   build_t2i_i2v_extend_workflow. 5 tests.

designv4 §9.11–§9.15 + falsifier/layout/closing updates (148 tests, 25 files). Also removed a
pre-existing unused import in ltx2/loop.py.
2026-07-04 17:15:34 +00:00
SolitaryThinker 69c9871154 [examples] Add v2_examples/{training,omni,workflows}/ — runnable examples
Three more example folders alongside inference/, all CPU/numpy, self-contained
(sys.path bootstrap), public API only, every script verified to run green.

training/ (7) — one per method, all via the uniform method.train_step seam:
  01 finetune · 02 dmd2 distillation · 03 diffusion_nft (likelihood-free RL,
  samples from the old policy, feature-cache reuse) · 04 self-forcing (causal
  chunk_rollout) · 05 joint LM+generator RL (UniRL; joint + prompt-only) ·
  06 N-way joint RL (per_expert vs shared credit) · 07 end-to-end workflow RL
  (T2I+I2V from one final-video reward).

omni/ (4) — 01 Cosmos3 (reason→joint denoise, shared MoT) · 02 BAGEL
  (text→image, shared MoT; scheduler prices both WorkUnit kinds) · 03 Qwen-Omni
  (thinker→talker→vocoder, three separate experts, text+audio) · 04 interleave
  parity across AR + diffusion loop types.

workflows/ (2) — 01 cross-model T2I→I2V workflow (image provably conditions the
  video) · 02 workflow as a first-class servable (requires/validate, address by
  id, register_workflows catalog, WorkflowRegistry).

Each folder has a README indexing its scripts.
2026-07-04 17:15:34 +00:00
SolitaryThinker 18dd295e8d [examples] Add v2_examples/inference/ — runnable Wan2.1 inference examples
Five self-contained, runnable scripts (CPU/numpy) for the Wan2.1-1.3B card on the
v2 runtime, each bootstrapping the repo onto sys.path so they run from anywhere:

- 01_basic_t2v.py                  minimal path: build engine → T2V request → run → video
- 02_params_and_reproducibility.py DiffusionParams knobs + seeded bit-identical reproducibility
- 03_streaming.py                  per-denoise-step preview chunks (OutputSpec stream)
- 04_concurrent_interleaved.py     step-interleaved batching + interleave parity gate + cache reuse
- 05_async_serving.py              AsyncEngine: concurrent generate, event stream, step-boundary cancel

+ README.md indexing them. All five run green; use only the public API.
2026-07-04 17:15:34 +00:00
SolitaryThinker 4c333e0509 [docs] Update v2/README to designv4 + current scope (127 tests)
- v2/README.md: point to designv4.md as the unified design (design_v3 as
  north star); refresh the scope table (joint/N-way/workflow RL, Qwen-Omni
  cascade, cross-model Workflow, tiled VAE co-scheduling, WorldModelSession),
  package layout (program/Workflow, runtime/session, the 7 methods, new model
  dirs), the demonstrated-stress-tests list (§9.3–§9.10), and counts (49→127,
  20 files). Sessions moved out of "out of scope"; WebRTC wire stays out.
- designv4.md: drop two intermediate absolute suite totals (milestone "91/97
  passed") in favor of "zero regressions" so the only absolute count is the
  current 127 (intro + layout) — no stale numbers.
2026-07-04 17:15:34 +00:00
SolitaryThinker 32a7a6b87b [feat] Three stress tests: interactive sessions, workflow RL, heterogeneous co-scheduling
Targets the three design_v3 claims that were most load-bearing AND least
exercised (sessions/realtime, training-plane boundary, the WorkUnit-generality
falsifier). All fit with no new runtime primitive (127 tests pass).

1. Interactive world-model session (runtime/session.py, §9.8) — the Session
   plane had ZERO coverage. WorldModelSession drives the causal chunk_rollout
   loop as a long-lived session: persistent cross-request world state on the
   Session.kv_handle, frame streaming, transactional step-boundary cancellation
   (a cancelled act leaves the world resumable), no cross-session smearing.
   Added only a continuation seam to the chunk loop (init seeds context from a
   world_context slot; default empty = unchanged one-shot path). 5 tests.

2. End-to-end RL over a cross-model workflow (training/methods/workflow_rl.py,
   §9.9) — trains BOTH flux-t2i and wan-i2v from ONE final-video reward. Rolls
   out the whole workflow with SDE capture in both instances; the same final
   advantage drives FlowGRPO PPO on each stage's transformer; two WeightSyncPlans
   on two instances. The earlier model (T2I) is trained by a reward on the final
   video — end-to-end credit across a model boundary — proven causal by a control
   (constant reward => zero advantage => nothing moves). 4 tests.

3. Heterogeneous WorkUnit co-scheduling (models/tiled/, §9.10) — the §17
   falsifier. VAETileLoop makes VAE decode a loop of VAE_TILE units; tiling is
   exact (== one-shot, C0), and VAE_TILE + DIFFUSION_STEP pipelines interleave
   bit-identically and co-run in one batch. Validates the mechanism; the
   economic half (does it pay) stays a port-time measurement. 4 tests.

designv4 §9.8–§9.10 + falsifier/layout/closing updates (127 tests, 20 files).
2026-07-04 17:15:34 +00:00
SolitaryThinker 58223c0c41 [feat] Register cross-model workflows as first-class named servables
Answers "what's the right way to name/register custom pipelines like T2I→I2V":
treat a Workflow like a card — a stable namespaced id in the same servable
namespace, declared dependencies, and a two-level registry. No new concepts;
mirrors how cards are registered (and vllm-omni's pipeline_registry).

- program/workflow.py: Workflow gains `requires` (the cards it composes, derived
  from stages) and `validate(engine)` (fail-fast if a required card is absent,
  P7). New WorkflowRegistry: declarative workflow_id -> builder catalog for
  out-of-tree/ad hoc use.
- runtime/engine.py: `_workflows` registry + register_workflow (validates deps,
  rejects id collision with a model_id) + serves(); engine.run routes a request
  whose model_id is a workflow to workflow.run — addressable exactly like a model.
  Single-model hot path untouched.
- runtime/async_engine.py + serving/server.py: serves() and /models include
  workflows (discoverable as servables).
- models/__init__.py: declarative `_WORKFLOWS` catalog (the cross-model analog of
  _BUILDERS) + register_workflows() helper; build_image_video_engine now registers
  the workflow too. Adding a custom pipeline = one catalog line.
- Naming convention: dotted/namespaced workflow_id (`image_video.t2i_i2v`),
  distinct from kebab model ids, collision-checked. Renamed from `t2i_then_i2v`.
- tests (+5, 12 total in the file): addressable by id, requires/validate, id
  collision, registry catalog, register_workflows helper. Full suite 114 passed.
- designv4 §9.6: the naming & registration convention documented.
2026-07-04 17:15:34 +00:00
SolitaryThinker b254d1affe [feat] Cross-model T2I→I2V workflow + N-way joint RL over arbitrary experts
Two more pipelines stress-testing the design, plus a BAGEL-placement note in
designv4. Both fit with no new runtime primitive (109 tests pass).

Pipeline 1 — cross-model T2I→I2V (program/workflow.py, models/image_video/):
- Realizes ProgramKind.WORKFLOW as a thin multi-instance orchestrator ABOVE the
  engine (the hot path stays single-instance). A Program composes one model's
  loops; a Workflow chains full engine.run calls across distinct cards, threading
  artifacts. (LTX-2 already covers same-card multi-stage; cross-model — FLUX→Wan
  — is the new capability the single-instance runner can't express.)
- flux-t2i (text→image) and wan-i2v (text+image→video) cards; the I2V program
  folds the conditioning image into text_embeds so WanDenoiseLoop is unchanged.
- Each model keeps its own interleave-parity guarantee (crossing instances is a
  Workflow boundary, not a loop step). 7 tests incl. video-depends-on-image.

Pipeline 2 — N-way joint RL (training/methods/joint_multi_rl.py, models/multi_expert/):
- JointMultiExpertRL generalizes UnifiedRLMethod (N=2) to N refiner LMs + a
  generator: one reward → one group advantage → N token-PG updates + 1 FlowGRPO
  PPO update, N+1 independent WeightSyncPlans. Proves the substrate was already
  N-ready (card holds N components/loops; per-component weight-sync; dict grad
  targets) — only the method body looped over two; now it loops over a list.
- credit="per_expert" learns all N cleanly; credit="shared" (faithful to UniRL)
  works but is noisier — the honest multi-agent credit-assignment result, a
  reward-shaping choice, not a substrate limit. 6 tests (N=1,3,4; prompt-only).

Fix — flow_sde_ml_velocity (loop/sampler.py): the toy FlowGRPO generator update
targeted the velocity the model already produced (a no-op once guidance_scale=1
was set for the C2 identity; the unified generator moved only on ~1e-7 noise).
The correct PG surrogate targets the max-likelihood velocity of the realized
sample — nonzero at ratio==1. Both UniRL and N-way generators now learn for real;
the C2 ratio==1 identity still holds (measured before the update).

BAGEL: MoT/shared-weight (one transformer on both loops), same row as Cosmos3;
real BAGEL's co-resident experts are expressible via the expert-routing policy
(partial sharing) — captured in designv4 §2.3.
2026-07-04 17:15:34 +00:00
SolitaryThinker c7e0a8e894 [feat] Add Qwen-Omni thinker→talker→vocoder model (3 experts, 3 loops)
Ports vllm-omni's canonical qwen2_5_omni omni-speech cascade as a v2 card:
a third weight-sharing topology — three disjoint experts (thinker, talker,
vocoder) on three loop types (ar_decode → ar_decode → audio_decode) in one
request, with chained cross-stage conditioning and streaming codec→waveform.
vllm-omni runs these as three opaque request-scheduled stages; v2 makes every
thinker token, talker token, and vocoder chunk a runtime-visible WorkUnit.

- models/backend.py: ToyTalker (a genuinely distinct AR expert, weight-salted)
  + ToyVocoder (streaming code2wav: codec tokens → waveform chunks).
- models/omni/vocoder_loop.py: VocoderLoop filling the pre-declared
  LoopKind.AUDIO_DECODE / WorkUnitKind.AUDIO_CHUNK slot.
- models/omni/ar_loop.py: ARDecodeLoop gains a configurable prompt_slot so two
  chained AR loops don't collide on the prefill slot (thinker vs talker).
- models/qwen_omni/: card (3 experts/3 loops) + program (tokenize → thinker →
  emit_text → thinker→talker full-payload hand-off → talker → talker→vocoder →
  vocoder → emit_audio). Cross-stage hand-offs are explicit Program nodes, the
  model-native form of vllm-omni's custom_process_input_func.
- _enums.py: Capability.TEXT_TO_SPEECH.
- tests/test_thinker_talker.py: 6 tests incl. three-loop interleave parity,
  cascade conditioning, AUDIO_CHUNK streaming. Full suite 97 passed.

designv4.md: §2.3 topology table extended to four topologies; new §9.5 on the
cascade; reference-synthesis + package layout updated.
2026-07-04 17:15:34 +00:00
SolitaryThinker b3ddf6014d [feat] UniRL/PromptRL joint LM+generator RL stress test + designv4
Stress-tests the v2 Card/Loop/Program design with a UniRL/PromptRL-style
joint RL recipe: a prompt-refiner LM expert and a flow generator expert,
two separate experts driven by two loop types in one request, both updated
simultaneously from a single RL reward.

- loop/sampler.py: flow_sde_step_with_logprob — FlowGRPO SDE rollout sampler
  (per-step Gaussian log-prob), distinct from the deterministic ODE serve step.
- request/params.py: gated sde_rollout/sde_noise_scale on DiffusionParams so
  the serve path stays byte-identical (default ODE).
- models/wan21/loop.py: gated SDE-rollout capture in WanDenoiseLoop.advance.
- models/backend.py: ToyPromptRefiner — a real REINFORCE categorical policy
  (the Qwen role), separate weights from the generator.
- models/unified/: the unified card+program — two disjoint experts (llm +
  transformer) on ar_decode + diffusion_denoise; the topological opposite of
  the Cosmos3 MoT card, same vocabulary.
- training/methods/unified_rl.py: joint GRPO — one reward -> group advantage
  -> LM token policy gradient + DiT FlowGRPO PPO; two LRs; prompt-only/joint
  flag; reuses the shared diffusion loop for rollout.
- training/weight_sync.py: WeightSyncPlan gains a component scope so the two
  experts version + cache-invalidate independently (LM sync never flushes the
  frozen text-encoder feature cache).
- tests/test_unified_rl.py: 9 tests incl. likelihood-based C2 identity, the
  two-loop interleave parity gate, joint vs prompt-only. Full suite 91 passed.

designv4.md: unified design doc reflecting v2 as built+tested, with the joint
RL stress test as the validating case study (the design held — new card +
new method, no new runtime primitive).
2026-07-04 17:15:34 +00:00
SolitaryThinker a9e5f6ee7a [refactor] rename package mini_fastvideo → v2
Directory rename (git mv, history preserved) plus rewrite of all references: absolute imports in
tests, the zero-dep runner, docstrings, comments, and the README. No behavior change.

Run: python3 -m pytest v2/tests/ -q ; python3 v2/run_tests.py ; python3 -m v2.examples
2026-07-04 17:15:34 +00:00
SolitaryThinker 7467076d72 [fix] mini-fastvideo serving: address adversarial-review findings (capacity/credit leaks, robustness)
Review confirmed the core bets (concurrent disaggregation is bit-identical, design conformance holds,
Dynamo genuinely optional, engine stays step-scheduled). Fixes for the untested failure paths:

- HIGH: pool capacity (RolePool.in_flight) no longer leaks when a disaggregated request is cancelled
  or errors mid-occupancy — DisaggregatedRunner.close() releases the occupied pool and AsyncEngine._run
  calls it in a finally (a cancel on a capacity-1 denoiser no longer bricks the pool).
- credit flow-control: cross-pool transfer wraps acquire/release in try/finally (no credit leak on a
  failing transfer); slot is re-homed only on a successful fetch.
- AsyncEngine: duplicate in-flight request_id is rejected (was a deadlock); submit() cancels the driver
  task when the consumer abandons the stream (client disconnect → no orphaned compute); bounded
  per-request history (no unbounded _events/_states/_results/_runners growth).
- cancellation is common-path on the OFFLINE path too (cancel check at the top of every runner.tick()).
- HTTP server: read timeout (slowloris guard → 408), body-size cap (→ 413), invalid Content-Length
  (→ 400), explicit StreamReader limit, and aclose() of the SSE generator on client disconnect.
- build_deployment_card no longer aliases one mutable CostModel across replica cards (dataclasses.replace),
  so online calibration of one worker's cost doesn't mutate another's.
- video-job tasks tracked (not fire-and-forget); server.close() cancels/drains them; jobs dict bounded;
  fleet affinity map bounded.

6 regression tests added for these paths. 82 tests pass (pytest + zero-dep runner).
2026-07-04 17:15:33 +00:00
SolitaryThinker 01dc0c3377 [feat] mini-fastvideo serving + fleet (our own version, Dynamo-optional)
Builds the full serving layer the design files specify, instead of deferring it to Dynamo:

- transport/ (§7.3): pluggable Connectors (in-proc zero-copy / SHM-fake copy) with chunk_ready
  readiness (vllm-omni) AND credit-based flow control (sglang-omni Relay); KVConnector protocol shape;
  TransferManifest.
- runtime/ (§6, §13; plan M3/M4): AsyncEngine — request queue, lifecycle state machine
  (waiting→running→completed/cancelled/failed), live AsyncIterator[OmniEvent] streaming, common-path
  cancellation, step-level concurrency. RolePool + DisaggregatedRunner (encoder→denoiser→decoder,
  capacity-aware dispatch, cross-pool transfers via connectors); disaggregated output is bit-identical
  to inline. No-progress detection (no busy-spin).
- deploy/ (§14, §6.3.5-6): DeploymentCard; OUR OWN LocalFleet (discovery, health/drain, least-loaded
  / cost-model / sticky-affinity routing) so we never rely on Dynamo; DynamoWorkerAdapter +
  FakeDynamoRuntime export the SAME card + cost model so Dynamo CAN front us — one object, two consumers.
- serving/ (§6.3.5, §12): framework-free stdlib-asyncio OpenAI server (our own version of the
  vllm-omni pattern): /v1/chat/completions (SSE), /v1/images/generations, /v1/videos (async job+poll)
  + /v1/videos/sync, /v1/models, /health, /metrics. A thin shim over the STEP-scheduled engine — the
  runtime-visible loop scheduler vllm-omni's request-scheduled opaque DIFFUSION stage lacks.

15 serving tests (real-socket HTTP+SSE via stdlib asyncio, disagg==inline, fleet routing, Dynamo
contract, cancellation); 76 tests pass total (pytest + zero-dep runner). ~7900 LOC.
2026-07-04 17:15:33 +00:00
SolitaryThinker f1dc587c74 [feat] mini-fastvideo phase 2: omni/MoT — Cosmos3 + canonical vllm-omni (BAGEL/lance)
One resident MoT instance runs BOTH an ar_decode loop and a diffusion_denoise loop on shared weights
(the §16 claim no DAG-of-engines can express), with both loops runtime-visible: the scheduler prices
ar_token AND diffusion_step WorkUnits — unlike vllm-omni's opaque DIFFUSION stage the scheduler never
sees inside.

- ARDecodeLoop: token decode until EOS/max_tokens, paged text-KV — the omni AR pathway (loop/§5).
- ToyMoTDiT: one module exposing an und pathway (ar_forward) AND a gen pathway (denoise __call__);
  ToyTokenizer. Binding both loops to one instance = shared weights, no duplication.
- models/cosmos3/: tokenize → reason(ar_decode) → pack(tokens→conditioning) → diffusion_denoise →
  vae_decode; sound_vae declared optional_for non-t2vs (the lazy-component P8 fix, not an env-var hack).
- models/bagel/: the canonical vllm-omni model — generate_text(ar_decode) → generate_image(diffusion),
  text+image outputs, both loops step-scheduled.
- The diffusion loop is WanDenoiseLoop reused (one loop definition bound to the MoT module).
- build_omni_engine() + an omni worked example; 7 omni tests (shared-instance, both-kinds-scheduled,
  interleave parity across loop types, lazy sound_vae). 61 tests pass (pytest + zero-dep runner).
2026-07-04 17:15:33 +00:00
SolitaryThinker 098bcf014a [fix] mini-fastvideo: address adversarial-review findings (admission liveness + §7.1 cache key)
- Admission fails fast with AdmissionInfeasible on infeasible/deadlocked reservations instead of a
  10M-iteration busy-spin: no-progress detection in run_to_completion/run_interleaved via a real
  progress token, plus feasibility pre-checks (need > pool capacity).
- Compute budget is now a refundable concurrency gate (release() refunds spent), not a
  never-refunded lifetime cap that silently deadlocks.
- §7.1: the text-encoder feature CacheKey carries adapter_versions + precision (no stale serve across
  te-LoRA stacks); per-component weight versions mean a transformer-only RL weight sync no longer
  flushes the frozen text-encoder cache (component-scoped invalidation, not wholesale).
- Interleave gate flags symmetric-empty output instead of passing it vacuously.
- skipped_steps counted only when the override is actually consumed; BatchScheduler wired for round
  batch-accounting (metric renamed stepped_units); ResidualCache.get cleanup; stream chunks carry a
  latent preview payload; dead progress-vars removed.
- 5 regression tests added for the previously-untested paths. 54 tests pass (pytest + zero-dep runner).
2026-07-04 17:15:33 +00:00
SolitaryThinker 270fae959d [feat] mini-fastvideo: model-native runtime per design_v3 (Wan2.1/LTX2 + 4 training methods)
A scoped, CPU-testable realization of design_v3.md — the architecture where the atomic unit
is a typed (recipe, runtime) ModelCard, the model owns loop semantics while the runtime owns
loop lifecycle, one resident instance runs many loops, and training records behavior on the
same loops it serves.

Implements:
- card/ loop/ runtime/ cache/ memory/ parallel/ parity/ extend/ program/ request/ training/
  spanning design_v3 §4-§13: ModelCard + validate(); driven loops (init/next/advance/finalize);
  step-interleaving Engine with reservation-before-admission + per-class caches keyed by CacheKey;
  the C0-C4 consistency ladder + the non-negotiable batch-of-N interleave parity gate.
- Inference: Wan2.1-1.3B (T2V), LTX2.3 (two-stage distilled, shared transformer), Wan-causal
  (chunk rollout + slab-KV streaming).
- Training (Wan2.1-1.3B), each driving the SAME loops the engine serves: finetune (flow-match),
  DMD2 (teacher/critic distribution matching), DiffusionNFT (likelihood-free C2, samples from the
  decay-blended old policy, group-relative advantages, shared-prompt cache reuse), self-forcing
  (causal chunk loop). The engine never imports training (the §10 dependency rule, grep-verified).

numpy-only core (no torch/GPU here); heavy Wan/LTX forwards are deterministic toy stand-ins with
lazy torch-adapter seams (ComponentSpec.load_id/factory) for a GPU box. 49 tests pass via pytest
and a zero-dependency runner; interleave parity verified (and a buggy module-global interceptor
provably breaks it). Omni-ready spine (ar_decode/chunk_step loop kinds, multi-loop instances,
LoopState.extension) for the phase-2 Cosmos3 + vllm-omni omni ports.

Run: python3 -m pytest mini_fastvideo/tests/ -q ; python3 -m mini_fastvideo.examples
2026-07-04 17:15:33 +00:00
SolitaryThinker 4a14c1afa3 update 2026-07-04 17:15:33 +00:00
SolitaryThinker f1c19050c3 design 2026-07-04 17:15:33 +00:00
William Lin 51ed1ea423 [build] Bump fastvideo-kernel pin to 0.3.2 (#1541) 2026-07-03 14:48:15 -07:00
Mac Lee 00ec3e7388 [bugfix]: compute VSA topk from padded blocks (#1517) 2026-07-03 14:16:51 -07:00
William Lin 0d626ef2d1 [kernel] Bump fastvideo-kernel pin to 0.3.1 and version the FA4 tile_mn port as 0.3.2 (#1539) 2026-07-03 14:11:08 -07:00
William Lin b36d0ef085 [infra] Add DGX Spark and multi-architecture CUDA support (#1447) 2026-07-03 13:41:43 -07:00
William Lin 10c8c5df4d [misc] reorg: relocate ui/ and performance_dashboard/ under apps/ (#1537) 2026-07-02 20:29:17 -07:00
William Lin 31e26abec4 [chore] release fastvideo-kernel 0.3.1 (#1520) 2026-06-30 12:25:27 -07:00
sumyyyyyandSolitaryThinker 9e83ba630c [kernel] Extract VSA utility functions into fastvideo_kernel (#1408)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-30 12:19:25 -07:00
William Lin fc02a9ce8e [ci] aarch64 kernel wheel: build for Blackwell (sm_100a/sm_120a), not Hopper (#1516) 2026-06-29 16:10:42 -07:00
William Lin 3d36160fc4 [bugfix] -fsigned-char so ThunderKittens compiles on aarch64 (Grace Hopper) (#1515) 2026-06-29 14:05:38 -07:00
William Lin a11ec43de2 [ci] build + publish aarch64 (Grace Hopper) kernel wheels alongside x86_64 (#1514) 2026-06-29 12:28:34 -07:00
William Lin 2656d6530c [bugfix] compile FP4 (attn_qat_infer) kernels for sm_120a only via per-arch split (#1508) 2026-06-29 11:07:27 -07:00
Kevin Lin e3f54e7169 [bugfix] Fix causal attention mask for Blackwell FP4 MMA column layout (#1506) 2026-06-28 20:36:40 -07:00
William Lin 4ba5681307 [bugfix] guard ThunderKittens Hopper kernels for non-sm_90a device passes (#1507) 2026-06-28 20:07:43 -07:00
alexzms 8c23c86994 [docs]: add raw-video preprocess script for the QAD MixKit recipe (#1487) 2026-06-28 18:01:13 -07:00
William Lin 78d606a3a7 [feat]: make VSA tile cache configurable for training (#1444) 2026-06-28 02:18:31 -07:00
William Lin c4e108de78 [docs]: modernize documentation build (#1503) 2026-06-28 02:02:09 -07:00
Kaiqin Kong 16bf2eaf77 [feat] Relativistic RoPE re-indexing for long causal rollouts (#1454) 2026-06-28 01:43:00 -07:00
William Lin 1d15d974fa [misc]: cleanup outdated or unneeded agent infrastructure (#1504) 2026-06-27 19:37:49 -07:00
alexzms 5e57868b76 [ci] train-framework model coverage: Cosmos/MatrixGame2 finetune grad-norm + Cosmos/LongCat/MatrixGame2 loading smokes (#1497) 2026-06-27 17:16:00 -07:00
William Lin 6be280c914 [bugfix]: fix docs build (#1502) 2026-06-27 13:00:09 -07:00
KyleNeverGivesUpandClaude Opus 4.8 3ccdec9798 [bugfix] Make enable_torch_compile_vae actually compile the Wan VAE (#1498)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-27 12:30:59 -07:00
Satyam Srivastava 8658f774f3 [docs] Restructure contributing CI/CD and testing docs (#1501) 2026-06-27 11:46:18 -07:00
Raghav K 4f3ad3f6df [perf] Default Wan VAE decode to bf16 (lossless, faster) (#1472) 2026-06-26 12:42:29 -07:00
William Lin b454aa56c3 [misc] fix pre-commit (#1500) 2026-06-26 12:39:28 -07:00
William Lin f9b8e30ff3 Update README.md 2026-06-26 09:23:30 -07:00
William Lin 0356205b84 [bugfix] revert fastvideo kernel version to 0.2.6 (#1495) 2026-06-25 03:52:43 -07:00
Kevin Lin 719a1879bd [bugfix] Remove incorrect V-row permutation in scaled_fp4_quant_trans_kernel (#1493) 2026-06-25 03:05:12 -07:00
Utkarsh RanjanandClaude Opus 4.8 b57180bf97 [bugfix] Warn when a requested attention backend is unsupported by a layer (#1254) (#1486)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-24 15:56:39 -07:00
William Lin 7cebf5f82c [ci] cap kernel wheel build parallelism to avoid runner OOM (#1483) 2026-06-23 14:17:56 -07:00
William Lin dd0f4b6753 [misc] cleanup misc files (#1484) 2026-06-23 14:17:22 -07:00
William Lin d303b4e03a [docs] update README (#1482) 2026-06-23 12:07:21 -07:00
xsankandmergify[bot] 31b719ae49 [perf] optimize compress & topk kernel (#1421)
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-23 18:51:22 +00:00
William Lin 887aaf3d3e [ci] bump cuda-toolkit action to v0.2.35 to fix kernel cu130 publish (#1481) 2026-06-23 10:33:31 -07:00
William Lin b1d89eba1f [chore] release fastvideo-kernel 0.3.0 (#1478) 2026-06-23 10:02:27 -07:00
Loay RashidandSolitaryThinker 995a5fdf97 [bugfix] fixing denoising time in the fastwan script (#1480)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-23 10:01:23 -07:00
William Lin 70a70b689e [ci] mergify: stop auto-syncing ready PRs (#1477) 2026-06-23 03:23:58 -07:00
Shao Duan 4d6ac89b43 [ci] eval: add metric regression + identity-invariant ci tests (#1451) 2026-06-23 01:27:34 -07:00
Raghav K 4171cacd93 [bugfix] Wire Cosmos-Predict2.5 2B to its sampling preset (#1468) 2026-06-23 01:25:35 -07:00
b2ade71467 [Bugfix] QAD 5090: Torch.compile and other optimizations (15/12) (#1466)
Co-authored-by: alexzms <3036648523@qq.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-23 01:23:57 -07:00
Kevin Lin 82ed9fe58d [feat] QAD 5090: FP8 linear layer inference (#1465) 2026-06-22 18:25:06 -07:00
Satyam Srivastava 3d8cc4f0a0 [bugfix] Fix performance component timing extraction (#1473) 2026-06-22 13:05:35 -07:00
Satyam Srivastava 0557f7a7d9 [ci] Add performance dashboard metadata and visualizations (#1470) 2026-06-19 14:27:01 -07:00
dc66cd97ef [feat] QAD 5090: FP8 QAT linear training (14/12) (#1464)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-06-19 01:03:51 -07:00
6da206e196 [feat] QAD 5090: FP4 QAT linear STE for training (13/12) (#1463)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-18 15:55:47 -07:00
sumyyyyy 87f98c9b8b [kernel] Add varlen support for block-sparse attention (#1319) 2026-06-18 01:06:00 +00:00
e60601df7f [feat] QAD 5090: QAT training recipe — finetune + DMD distillation (12/12) (#1462)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 15:18:52 -07:00
eed9c4bfbf [kernel] QAD 5090: Add Attn-QAT training Triton kernels (11/12) (#1460)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
2026-06-17 13:35:12 -07:00
1dee77f4a4 [feat] QAD 5090: Wire the Attn-QAT training attention backend (10/12) (#1459)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: alexzms <26690162+alexzms@users.noreply.github.com>
Co-authored-by: mergify[bot] <37929162+mergify[bot]@users.noreply.github.com>
2026-06-17 13:32:56 -07:00
Satyam Srivastava b80148819c [ci] Add performance dashboard visualisation scripts (#1469) 2026-06-16 23:29:29 -07:00
c3b971488e [docs] QAD 5090: Add NVFP4 + Attn-QAT inference example and how-to (9/12) (#1458)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: alexzms <26690162+alexzms@users.noreply.github.com>
2026-06-16 14:36:50 -07:00
88e753f281 [feat] QAD 5090: Wire the Attn-QAT inference attention backend (8/12) (#1457)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: alexzms <26690162+alexzms@users.noreply.github.com>
2026-06-16 13:12:33 -07:00
77832059cc [kernel] QAD 5090: Add modified SageAttention3 FP4 inference kernels (7/12) (#1455)
Co-authored-by: William Lin <SolitaryThinker@users.noreply.github.com>
Co-authored-by: Loay Rashid <42599591+loaydatrain@users.noreply.github.com>
Co-authored-by: Kaiqin Kong <k1kong@ucsd.edu>
Co-authored-by: Edenzzzz <wtan45@wisc.edu>
2026-06-16 12:02:30 -07:00
743 changed files with 91846 additions and 19874 deletions
-46
View File
@@ -1,46 +0,0 @@
# Exploration Logs
This directory holds draft procedures and investigation notes for tasks that
don't yet have a standardized skill or SOP. Each exploration should follow this
template.
## When to Create an Exploration Log
- You are working on a task with no existing skill or workflow.
- You are experimenting with a new metric, training technique, or tool.
- You want to document findings before they are promoted to a standard.
## File Naming
`<topic-slug>.md` — e.g., `fvd-metric-investigation.md`
## Template
```markdown
# Exploration Log: <Topic>
## Status: draft | under_review | promoted | abandoned
## Context
<Why this exploration is needed — link to experiment or task if applicable.>
## Progress
- [ ] Step 1: ...
- [ ] Step 2: ...
## Findings
<What you have learned so far.>
## Mistakes / Dead Ends
<What didn't work and why — these become lessons.>
## Proposed Standardization
<If this works, describe the skill/SOP/workflow to create.>
```
## Lifecycle
1. **Create** during exploration mode.
2. **Update** as you make progress.
3. **Promote**: If findings are solid, create a skill in `.agents/skills/` or an SOP in `.agents/workflows/`.
4. **Archive mistakes**: Move failures into `.agents/lessons/`.
+207
View File
@@ -0,0 +1,207 @@
# v2 ← M\*: Architecture Gap-Analysis & Improvement Roadmap
**Status:** exploration, flagged for review. **Date:** 2026-06-19.
**Source paper:** *M\*: A Modular, Extensible, Serving System for Multimodal Models* (arXiv 2606.12688,
Stanford/UW/CMU; Jha, Sagan, Kamahori, …, Kasikci, S. Wang). It is a universal serving runtime for composite
multimodal models built on the **Walk Graph** abstraction (a model is a dataflow graph `G`; a request is a
*Walk* — a labeled subgraph — and the runtime executes walks). It beats vLLM-Omni (~20% lower T2I latency on
**BAGEL**, up to 2.64× on I2I), SGLang-Omni (2.7× TTS throughput on **Qwen3-Omni**), and native V-JEPA2
rollout (12.5×). It explicitly names **FastVideo's own** sparse/sliding-tile attention, xDiT/PipeFusion/USP,
Inferix, and FlashDrive as techniques integratable into the graph runtime.
**Method:** a 28-agent workflow — 6 parallel v2-subsystem maps → 10 M\*-dimension analyses, each
*adversarially verified against the actual v2 code* → synthesis + a completeness critic. The critic's
corrections and three P0 claims were then **spot-verified by hand** (file:line below). This doc folds those
corrections in; it is the corrected, authoritative synthesis.
---
## 1. Executive summary
v2 already implements the **harder half** of M\*'s thesis and in several axes **exceeds** it:
- v2's `Program` *is* M\*'s graph `G` (typed `ComponentNode`/`ModelLoopNode` + edges).
- v2's `shared_weight_components` *is* M\*'s cross-Walk node sharing — BAGEL/Cosmos3/LTX2 each bind two
`ModelLoopNode`s to **one resident transformer** (`instance.component()` returns the same live object). This
is the exact MoT serving property the omni cards in this repo already express.
- v2 adds three things M\* (serving-only) has **no equivalent for**: a required+validated per-loop **cost
model**, a non-negotiable **interleave bit-parity gate**, and an **integrated training plane** (RL→distill
flywheel driving the *same* serving Loop).
- The `extend/` plugin seam (interceptors/observers/registry with capability negotiation) is precisely the
hook M\*'s "extensible / integrate FastVideo-STA, xDiT, Inferix, FlashDrive" call-out asks for — **v2
already has the seam M\* only gestures at.**
What v2 lacks is M\*'s **declarative authoring layer above the substrate**, and — the key insight — *much of
that substrate is already authored but inert*: v2 has declared the metadata for "minimum components per
request" (`required_for`/`optional_for` on every omni card) and "branch as a cache axis" (`guidance_sig`,
`CacheKey`) but **never wired it to an executor**. The substrate is ~80% built and switched off.
**Highest-leverage cluster:** three small, parity-safe wires that turn on inert substrate and unblock the
BAGEL/Qwen-Omni/Cosmos3 latency wins M\* measured **on the exact models this repo already runs** — plus one
P1 that aligns v2 with the paper's headline "extensible" claim using a seam v2 already has.
### Verified P0 correctness findings (spot-checked by hand)
1. **Runner divergence (real bug).** `v2/runtime/engine.py:88` → `nodes = self.program.nodes`;
`v2/runtime/disaggregated.py:96` → `nodes = self.program.active_nodes(self.request)`. The inline and
disaggregated runners execute *different node sets*. ✅ confirmed.
2. **EOS is faked.** `v2/recipes/omni/ar_loop.py` docstring says "done on EOS/max_tokens"; `next()` (`:46-48`)
checks **only** `max_tokens`. M\*'s marquee `DynamicLoop` use case (EOS) is unimplemented in the loop that
serves the Qwen-Omni Thinker/Talker and Cosmos3 reasoner. ✅ confirmed.
3. **`required_for`/`optional_for` have zero runtime consumers** (grep outside `specs.py`/recipes/tests is
empty). The min-components metadata is declared on every card and never read. ✅ confirmed.
---
## 2. Dimension table (corrected)
| # | Dimension | v2 status | Gap | Priority | Effort | Payoff | Action |
|---|---|---|---|---|---|---|---|
| 1 | Min-components per request (`required_for` + `when_task`) | substrate built, **inert** | real, cheap | **P0** | S | Consume `required_for` in `active_nodes`; unify `engine.py:88` onto `active_nodes`; deliver via registry/card builder so all ~40 cards inherit it |
| 2 | Real EOS + declarative `DynamicLoop` | early-exit emergent; **EOS faked** | real | **P0** | S | `ARDecodeLoop` honors `eos_id` + `req.sampling.stop`; add `LoopSpec.dynamic_stop` + `register_loop_stop`. **Training-enabling** (world-model rollout horizon) |
| 3 | CFG/branch as label over one paged KV pool | absent (`PagedKVCache` is a counter) | real | **P1** | L | `(namespace,label)` paged store w/ one budget; reuse `guidance_sig` for hash (NOT `partition_field`); by-ref via existing `InProcKVConnector`. AR path only (diffusion has no KV) |
| 4 | `extend/` plugin seam → integrate FastVideo-STA / Inferix | **seam exists, unused for attn** | real (paper headline) | **P1** | M | Expose FastVideo sparse/sliding-tile attention + Inferix block-diffusion as `Interceptor`/`EngineKind` plugins — the paper's named integration targets, on this repo's own code |
| 5 | `ParitySpec.output_determinism` (C3 distributional) | C3 rung defined, **0 users** | real, dormant | **P1** | S | Add field; `compare_outputs` consults it. **Training-enabling** (SDE/FlowGRPO stochastic rollouts) |
| 6 | Registry-driven delivery of #1 | present, not leveraged | integration | **P1** | S | Express `when_task`/min-components through `WorkflowRegistry`/card builders, not 3 bespoke recipe patches |
| 7 | Serving conductor + pluggable data plane | conductor exists (`serving/http.py`); **single-process transport** | real | **P2** | L | v2 already has the step-scheduled worker surface; gap is ZeroMQ/Mooncake + direct worker→worker tensor routing (today `InProcKVConnector` only) |
| 8 | Fleet/Dynamo placement + replicas | **live** (`deploy/fleet.py`,`dynamo.py`) | partial | **P2** | M | Fleet-level placement/affinity/replica is real & ≥M\*; missing piece is only the intra-engine `(node,Walk)→rank` map decoupled from model code |
| 9 | Per-node TP / SP + cross-rank transport | axis vocab **exists** (`sp` incl.); not wired to runtime | partial | **P2** | XL | Wire declarative degrees into runtime; Wan/LTX are **SP-native** (TP is a no-op there); populate `parallel_plan_hash` on the serving cache path |
| 10 | Named Walks + per-model state machine | `Program`=G, sharing real; no Walk/SM | real | **P2** | M | Defer until a *re-entrant* phase graph (Thinker↔Talker, rollout) needs it; #1 captures the min-components win without it |
| 11 | Declarative `Parallel/Sequential/Loop` IR | imperative loop classes | real (authoring) | **P2** | M | Thin Section IR lowering to flat `Program`; scope to one AR recipe |
| 12 | Streaming `ChunkPolicy` + `StreamBuffer` | causal-chunk emit **already ships** (`wan_causal`); `EdgeKind.STREAM` inert | real | **P2** | L | Declarative `ChunkPolicy` vocab over the existing chunk mechanism; needs concurrent producer/consumer runner (= pipelined scheduling). Inferix integration point |
| 13 | Speculative deferred-termination; loop-spanning CUDA graphs; N+1 prefetch; attn double-buffer | absent / per-step capture (14 cards) | real | **P3** | L | Gate behind a real GPU executor; unobservable on CPU-toy CI; loop-span needs an `allows_interleaving=False` carve-out |
| — | Cost model + interleave/consistency parity | **exceeds M\*** | none | **guard** | — | Do not regress; keep `step_cost_model` mandatory + `bit_identical` default |
| — | Integrated training plane (flywheel, weight-sync) | **exceeds M\*** | none | **guard** | — | Protect train==serve loop identity with a toy fixture |
---
## 3. P0/P1 deep-dives (sequenced)
```
PR-1 (P0) min-components ──┐
PR-2 (P0) real EOS ─┼─► prereqs for honest "DynamicLoop" + min-component claims; both training-enabling
PR-3 (P1) output_determinism (independent)
PR-5 (P1) extend/ plugin: FastVideo-STA / Inferix as Interceptors (independent; highest paper-alignment)
PR-4 (P1) CFG-as-label paged pool ──► depends on PR-2 (AR loop is the only KV consumer)
```
PR-1, PR-2, PR-3, PR-5 are mutually independent; PR-4 depends on PR-2.
### PR-1 (P0) — Turn on the inert min-components substrate + fix runner divergence
- **Change.** Extend `Program.active_nodes(request)` (`v2/program/specs.py`) to also drop any node whose bound
`ComponentSpec.required_for` (`v2/card/specs.py:144`) excludes `request.task` (and isn't in `optional_for`).
**Fix the bug:** change `v2/runtime/engine.py:88` to `nodes = self.program.active_nodes(self.request)` so the
inline `ProgramRunner` matches `DisaggregatedRunner` (`disaggregated.py:96`). Deliver the `when_task` gating
through the **registry/card builder** (`recipes/__init__.py`, `program/workflow.py:WorkflowRegistry`) so all
~40 cards inherit it uniformly — not three bespoke `program.py` patches.
- **Why (this repo's models).** BAGEL T2I currently steps the AR-text loop and Cosmos3 t2v materializes the
reasoner even though the cards declare `transformer required_for={'reason','t2i'}`, `vae required_for={'t2i'}`.
On the GPU backend that is wasted resident-weight load + wasted steps on every single-modality request —
exactly M\*'s "execute the MINIMUM components per request," delivered by consuming existing metadata.
- **Risk/invariant.** Validate in `ModelCard.validate()` that every active node's `reads` are produced by an
active node for each declared `TaskType` (avoid dropping a producer). Pure node-id filtering ⇒ serial and
interleaved still walk the same filtered list ⇒ §9.3 interleave bit-parity holds by construction. CPU-toy clean.
### PR-2 (P0) — Real EOS + declarative `dynamic_stop` *(also training-enabling)*
- **Change.** In `v2/recipes/omni/ar_loop.py`, `advance()` reads the emitted token; if it equals the model
`eos_id` (toy backend exposes `EOS=0`) or matches `req.sampling.stop` (`params.py:21`, currently dead),
register termination; `next()` returns `Done()` on stop OR `max_tokens`. Add `StopRegistry` to `LoopState` +
`register_loop_stop(name)` to the `LoopContext` protocol (`contracts.py:204`) and to
`DisaggregatedRunner`'s `RuntimeLoopContext`. Add `LoopSpec.dynamic_stop: bool=False`, opt the AR cards in.
- **Why.** The docstring-vs-code lie sits in the loop serving Qwen-Omni Thinker/Talker and the Cosmos3 reasoner;
M\*'s second named `DynamicLoop` use case (world-model **rollout horizon**) is exactly what `self_forcing` RL
needs — so this is both a serving-credibility fix and a training enabler (raise its payoff accordingly).
- **Risk/invariant.** `dynamic_stop=False` is byte-identical back-compat. Must pass **all three** parity gates:
serial==interleaved AND disaggregated==inline. **Not** in this PR: speculative deferred-termination (unobservable
on CPU-toy, fights the interleave invariant — P3, gated on GPU executor).
### PR-3 (P1) — `ParitySpec.output_determinism` (close the dormant C3 hole) *(training-enabling)*
- **Change.** Add `output_determinism: str = "bit_identical"` to `ParitySpec` (`card/specs.py:88`); make
`compare_outputs` (`parity/interleave_gate.py:54`) consult it (`bit_identical` → today's exact check;
`distributional` → a moment/tolerance check — land a simple moment match first; a real KS test is new code).
- **Why.** `ConsistencyLevel.C3` is defined and used by zero recipes; an SDE/FlowGRPO stochastic rollout cannot
honestly declare its parity contract and would falsely fail the bit-identical gate. Additive; default unchanged.
### PR-5 (P1) — Expose FastVideo's own attention + Inferix as `extend/` plugins *(highest paper-alignment)*
- **Change.** Use the existing `extend/{interceptors,observers,registry}.py` seam (capability-negotiated, with
per-(request,branch) `plugin_state` that already passes the interleave gate) to register FastVideo's
sparse/sliding-tile attention and Inferix-style block-diffusion as `Interceptor`s / an `EngineKind` plugin.
- **Why.** M\*'s title is "Modular, **Extensible**" and it explicitly lists FastVideo-STA, xDiT/PipeFusion/USP,
Inferix, FlashDrive as integratable. v2 already has the seam M\* only describes — this is where v2 most
directly answers the paper, using this repo's own attention code. Low risk (the seam + capability negotiation
already exist and are tested).
### PR-4 (P1) — CFG/branch as a LABEL over one paged KV pool
- **Change.** Rewrite `PagedKVCache` (`cache/classes.py:155-172`) from a block *counter* into a real
`(namespace,label)->[block-handle]` store with **one shared `total_blocks` budget** (M\*'s single-pool
property). Reuse the existing-but-unpopulated `CacheKey.guidance_sig` (`keys.py:53`) for the hash. Thread the
label through `ar_loop.py` (alloc/append/get per `(request_id, branch)`; prefill once per shared-prefix label;
combine via `CFGPolicy.combine`). Wire `ResourceRequest.cache_blocks` (`contracts.py:64`, zero consumers) into
admission per (class,label).
- **Why.** The dossier-identified driver of M\*'s BAGEL win (3 CFG contexts as 3 labels over ONE pool vs dense
per-context). Targets AR_DECODE (BAGEL `generate_text`, omni Thinker); **correctly excludes diffusion**
(Wan/LTX are bidirectional, no KV — their CFG stays dense-but-batched).
- **Corrections to bake in.** Do **NOT** add `branch_label` to `CacheKey.partition_field()` (CFG branches share
embeddings; partitioning by branch is a semantic bug). Do **NOT** add a new by-ref type — reuse
`InProcKVConnector` + `TransferManifest.cache_key`. Wiring `cache_blocks` admission is greenfield ⇒ effort **L**.
CPU version proves label/sharing semantics; the real latency win needs a FlashInfer paged kernel (out of scope)
— **merge** with a future "real KVCacheEngine" effort rather than landing isolated.
---
## 4. What v2 already does ≥ M\* — do NOT regress
1. **Required+validated cost model** on every `LoopSpec` (13-kind `WorkUnitKind`) — typed, pre-GPU-validated.
2. **Interleave bit-parity as a hard gate** (`parity.interleave_required=True` on 40+ cards). M\* has no such
gate (its speculative scheduling deliberately wastes steps). Load-bearing invariant; every new primitive
must pass it.
3. **C0–C4 consistency ladder** wired into RL methods, with first-divergence tap reporting. No M\* equivalent.
4. **Integrated training plane** — DiffusionNFT/DMD2/self_forcing, RL→distill flywheel, `WeightSyncController`
hot weight-sync with drain-to-boundary + scoped cache invalidation, driving the **same** serving Loop.
M\* is serving-only. Protect with a toy fixture asserting `rollout_loop` drives the served Loop object.
5. **CPU-toy parity for the whole stack** — loops/CFG/caches/parity/RL run in CI without a GPU. Every new
primitive must ship a toy exercise (this is what makes all PRs above testable without H100s).
6. **Partition-not-flush cache invalidation** + four independent per-class pools.
7. **`extend/` plugin seam** with capability negotiation (a 4-step distilled card *rejects* a residual-skip
interceptor) — M\* describes extensibility; v2 has the mechanism.
8. **Dynamo citizenship** (`deploy/dynamo.py`: one `DeploymentCard`+cost model, two consumers) — beyond M\*'s
self-contained runtime.
---
## 5. Dropped / merged / deferred (and why)
- **DROP declarative `Parallel` as a CFG-execution win.** The runner walks nodes linearly (ignores
`Program.edges`), so `Parallel` lowers to sequential sugar and the CFG 3-pass braid is already one
co-scheduled `WorkPlan.run`; splitting it risks the interleave gate. Salvage only the no-op refactor
extracting `branch_forward` from `WanDenoiseLoop._velocity`. Reassign `Parallel` to the placement workstream.
- **MERGE the full Walk/state-machine layer** into "defer until a re-entrant phase graph needs it" (PR-1 gets the
min-components win with ~20 lines, no new abstraction). If built: the validator must check a walk's node-id
order is a *subsequence* of `program.nodes` (not just membership) or the runner can reorder and break parity.
- **MERGE `StreamBuffer`/`ChunkPolicy` into pipelined-scheduling.** Causal-chunk emit *already ships*
(`wan_causal/loop.py` per-chunk `StepResult.emit` + slab-KV); the gap is the declarative `ChunkPolicy` vocab
+ a concurrent producer/consumer runner. If built: keep all policies pure (per-request `StreamBuffer` history,
not shared edge state) and restrict the bit-identical claim to the token-only handoff.
- **MERGE CFG-fan-out exec + cross-rank transport + PD loop-splitting into a multi-GPU-runtime program.** These
need real collectives (`v2/distributed/` is a stub) and KV-by-reference (KV lives in `CacheManager`, not the
transferable `slots`). **Keep cheaply now:** the *declarative* halves — per-component degree, `(node,Walk)`
placement key with node-only fallback, `ReplicaSet` under `LocalFleet`, and populate `parallel_plan_hash` on
the **serving** cache path (it is already populated in `training/behavior.py:40` — the gap is serving-only).
- **DEFER** speculative deferred-termination, loop-spanning CUDA graphs, N+1 prefetch, attention-plan
double-buffer — all gated on a real GPU executor; benefit unobservable on CPU-toy CI. Keep the cheap
`EngineKind` tag (`STATELESS|KV_CACHE|DIFFUSION`) now. Correct the stale `cudagraph.py:51-52` docstring
(per-step capture ships in 14 cards, not just wan21).
- **RESCOPE per-node TP.** Wan/LTX use `ReplicatedLinear` + **sequence parallelism** (`sp`), not TP; the `sp`
axis already exists in `parallel/plan.py:AXIS_NAMES`. The work is wiring degrees into the runtime, not
inventing vocabulary; a `tp_size=2` "one-line activation" is a no-op for the shipped models.
---
## 6. The first integration test, if/when multi-GPU placement work starts
The **live Qwen-Omni 2-GPU bring-up** (Thinker on rank 0, Talker+Code2Wav on rank 1; see
`v2_debug_videos/vlm.md` Session 4) is the natural first validation target for any `(node,Walk)→rank`
placement work — it is the one place this repo already has real multi-rank composite-model execution.
---
## Anchor files for P0/P1
`v2/program/specs.py`, `v2/runtime/engine.py` (**line 88 fix**), `v2/runtime/disaggregated.py`,
`v2/recipes/omni/ar_loop.py`, `v2/loop/contracts.py`, `v2/card/specs.py`, `v2/cache/{classes.py,keys.py}`,
`v2/parity/interleave_gate.py`, `v2/extend/{interceptors,registry}.py`, `recipes/__init__.py` +
`v2/program/workflow.py` (registry-driven delivery).
-130
View File
@@ -1,130 +0,0 @@
# FastVideo-WorldModel — Codebase Map
High-level structural index for agent orientation. Updated 2026-03-08.
## Repository Layout
```
FastVideo-WorldModel/
├── fastvideo/ # Core Python package
│ ├── models/ # Model implementations
│ │ ├── dits/ # DiT transformers (wanvideo, ltx2, ...)
│ │ ├── vaes/ # VAE models
│ │ ├── encoders/ # Text/image encoders (T5, CLIP)
│ │ ├── schedulers/ # Noise schedulers
│ │ ├── upsamplers/ # Super-resolution models
│ │ ├── audio/ # Audio models
│ │ └── loader/ # Component loaders for HF repos
│ ├── configs/ # Configuration system
│ │ ├── models/ # Arch configs + param_names_mapping
│ │ ├── pipelines/ # Pipeline wiring
│ │ └── sample/ # Default sampling parameters
│ ├── pipelines/ # End-to-end pipelines
│ │ ├── basic/ # Per-model pipelines (wan/, ltx2/, ...)
│ │ └── stages/ # Reusable pipeline stages
│ ├── train/ # Refactored training framework (YAML-driven, preferred)
│ │ ├── trainer.py # Main training loop coordinator
│ │ ├── entrypoint/ # Training entrypoint (train.py) + checkpoint conversion
│ │ ├── methods/ # Training algorithms (FineTune, DFSFT, DMD2, SelfForcing)
│ │ │ ├── base.py # TrainingMethod ABC
│ │ │ ├── fine_tuning/ # FineTuneMethod, DiffusionForcingSFTMethod
│ │ │ └── distribution_matching/ # DMD2Method, SelfForcingMethod
│ │ ├── models/ # Per-role model wrappers (ModelBase, CausalModelBase)
│ │ │ ├── wan/ # WanModel, WanCausalModel
│ │ │ └── matrixgame2/ # MatrixGame2Model, MatrixGame2CausalModel
│ │ ├── callbacks/ # Composable hooks (grad_clip, ema, validation)
│ │ └── utils/ # Config, builder, checkpoint, optimizer, tracking
│ ├── training/ # Legacy training infrastructure (being phased out)
│ │ ├── trackers.py # W&B tracker (BaseTracker → WandbTracker)
│ │ ├── training_utils.py # Checkpointing, grad clipping, state dicts
│ │ ├── training_pipeline.py # Base training pipeline
│ │ ├── wan_training_pipeline.py # Wan T2V training
│ │ ├── wan_i2v_training_pipeline.py # Wan I2V training
│ │ ├── distillation_pipeline.py # Distillation base
│ │ ├── wan_distillation_pipeline.py # Wan distillation
│ │ ├── self_forcing_distillation_pipeline.py # Self-forcing distill
│ │ ├── ltx2_training_pipeline.py # LTX-2 training
│ │ └── matrixgame2_training_pipeline.py # Matrix-Game 2.0 training
│ ├── attention/ # Attention backends
│ ├── distributed/ # Sequence/tensor parallel utilities
│ ├── layers/ # Tensor-parallel layers
│ ├── tests/ # Package-level tests
│ │ ├── training/ # Training regression tests (W&B summary comparison)
│ │ ├── ssim/ # SSIM visual regression tests
│ │ ├── encoders/ # Encoder parity tests
│ │ └── modal/ # Modal CI test runner
│ └── registry.py # Unified config registry
├── fastvideo-kernel/ # CUDA/custom kernels (separate build: ./build.sh)
├── scripts/ # Utility scripts
│ ├── distill/ # Distillation launch scripts
│ ├── inference/ # Inference scripts
│ ├── checkpoint_conversion/ # Weight conversion tools
│ ├── finetune/ # Finetune scripts
│ └── preprocess/ # Data preprocessing
├── examples/ # Ready-to-run examples
│ ├── training/ # Training examples (finetune/, consistency_finetune/)
│ ├── distill/ # Distillation examples
│ ├── inference/ # Inference examples
│ └── dataset/ # Dataset examples
├── docs/ # MkDocs documentation source
│ ├── design/overview.md # Architecture overview
│ ├── training/ # Training guides
│ └── contributing/ # Contributor guides + coding_agents.md
├── tests/ # Top-level tests (local_tests/)
├── AGENTS.md # Agent coding guidelines
└── .agents/ # Agent infrastructure (you are here)
```
## Key Training Entrypoints
### New framework (`fastvideo/train/`) — preferred
| Method | Config Example | Launch Pattern |
|--------|---------------|----------------|
| FineTune (Wan) | `examples/train/finetune_wan2.1_t2v_1.3B_vsa_*.yaml` | `torchrun -m fastvideo.train.entrypoint.train --config <yaml>` |
| DFSFT (Wan causal) | `examples/train/dfsft_wan_causal_t2v_1.3B.yaml` | `torchrun -m fastvideo.train.entrypoint.train --config <yaml>` |
| DMD2 distillation | `examples/train/distill_wan2.1_t2v_1.3B_dmd2.yaml` | `torchrun -m fastvideo.train.entrypoint.train --config <yaml>` |
| Self-Forcing | `examples/train/self_forcing_wan_causal_t2v_1.3B.yaml` | `torchrun -m fastvideo.train.entrypoint.train --config <yaml>` |
### Legacy pipelines (`fastvideo/training/`) — being phased out
| Pipeline | Entrypoint | Launch Pattern |
|----------|-----------|----------------|
| Wan T2V finetune | `fastvideo/training/wan_training_pipeline.py` | `torchrun --nproc_per_node N` |
| Wan I2V finetune | `fastvideo/training/wan_i2v_training_pipeline.py` | `torchrun --nproc_per_node N` |
| Wan distillation (DMD) | `fastvideo/training/wan_distillation_pipeline.py` | `torchrun --nproc_per_node N` |
| Self-forcing distill | `fastvideo/training/wan_self_forcing_distillation_pipeline.py` | `torchrun --nproc_per_node N` |
| LTX-2 finetune | `fastvideo/training/ltx2_training_pipeline.py` | `torchrun --nproc_per_node N` |
| Matrix-Game 2.0 | `fastvideo/training/matrixgame2_training_pipeline.py` | `torchrun --nproc_per_node N` |
## W&B Integration
- **Tracker classes**: `fastvideo/training/trackers.py`
- `WandbTracker` — logs metrics, videos, timing
- `SequentialTracker` — fan-out to multiple trackers
- `DummyTracker` — no-op for offline/test
- **Run summary location**: `<output_dir>/tracker/wandb/latest-run/files/wandb-summary.json`
- **Reference summaries**: `fastvideo/tests/training/*/` (e.g., `a40_reference_wandb_summary.json`)
- **Environment**: `WANDB_API_KEY`, `WANDB_BASE_URL`, `WANDB_MODE`
## Critical Environment Variables
| Variable | Purpose |
|----------|---------|
| `WANDB_API_KEY` | W&B authentication |
| `WANDB_MODE` | `online` / `offline` |
| `FASTVIDEO_ATTENTION_BACKEND` | `FLASH_ATTN` / `TORCH_SDPA` |
| `TOKENIZERS_PARALLELISM` | Set `false` to avoid fork warnings |
| `HF_HOME` | HuggingFace cache directory |
## Build & Test Commands
```bash
uv pip install -e ".[dev]" # Editable install
pre-commit run --all-files # Lint/format/spell
pytest tests/ # Top-level tests
pytest fastvideo/tests/ -v # Package tests
pytest fastvideo/tests/training/Vanilla -srP # Training loss regression
pytest fastvideo/tests/ssim/ -vs # SSIM visual regression
cd fastvideo-kernel && ./build.sh # Build kernels
```
@@ -1,201 +0,0 @@
# Dreamverse Integration — Memory Index
Living knowledge base for the FastVideo ↔ Dreamverse ↔ Dynamo integration.
Tracks the public API refactor (PRs 0-17), the LTX-2 streaming server
upstream, the Dreamverse switch from `FastVideo-internal` to public
`FastVideo`, and the NVFP4 quantization landing.
**Last reconciled:** 2026-05-06 (**D-26** EXECUTED — rebased
`will/dreamverse-monorepo` directly onto `origin/main` (`c17d33bf`)
via `git rebase origin/main`; 66 commits cleanly replayed; force-
pushed via `--force-with-lease`. New tip `83829c5e`. Local backup
branch `will/dreamverse-monorepo-pre-main-rebase-backup-20260506`
preserved at the pre-rebase tip `2ee839a3`. PR #1288 on
`will/ltx2_sr_port` is untouched. The branch is now ready to open
as a single PR against main.
**Earlier — D-21** EXECUTED — chunk-stutter root cause
analysis + NVENC build path + opt-in `--nvenc` flag + benchmark regression
test. 3 parallel explore agents confirmed apps/dreamverse matches
FastVideo-internal byte-for-byte on NVFP4 + torch.compile coverage; stutter
is NOT a regression. Software libx264 encoding consumes ~22% of segment
wall-time. Built ffmpeg with NVENC support; B200 silicon doesn't have NVENC
encoder hardware so the path is currently moot on this dev host but works
on RTX 50-series / T4 / A10 deploys. Benchmark captured libx264 ultrafast
at 611ms median (8.25x realtime in isolation). Open follow-ups D-22/D-23/D-24.
**Earlier — D-20** EXECUTED — segment-2 BrokenPipe
root cause was a TWO-direction silent drop of LTX-2 audio kwargs in
public `VideoGenerator` (inbound `SamplingParam.update()` rejected
`audio_num_frames`/`ltx2_audio_clean_latent`/etc. as unknown fields and
`logger.error`'d, outbound result dict didn't surface
`ltx2_audio_latents` from `output_batch.extra`). Ported the
FastVideo-internal extra-overrides routing block + made `update()`
strict + added regression test (7 tests, all pass) + landed 4 commits
on `will/dreamverse-monorepo` @ `5eaf0a13` (11 commits ahead of
`fbd823df`). End-to-end verified on GPU4: `Cached audio latents shape
=(1, 8, 126, 16) for segment 2`, `Segment 2: relayed av chunks=22,
bytes=3.8MB`, no BrokenPipeError. Public-API fix needs cherry-pick to
`will/ltx2_sr_port` for PR #1288 — see open-threads.md item D-20-CP.
**D-19** EXECUTED previously: Dreamverse migration landed on
`will/dreamverse-monorepo` @ `c1fe5d4c` (5 commits ahead of
`will/ltx2_sr_port` HEAD `fbd823df`). 164 files, 53,294 LOC, 31,725
files under `apps/dreamverse/`. e2e PASSES against migrated code (8/8
Playwright in 5.1s, `/proc/$PID/cwd` verified). Significant deviation
from [integration-plan.md](integration-plan.md): the plan's "DELETE
generic-merged from Dreamverse, import public substitutes" assumption
was invalid (public APIs aren't drop-ins) — generic-merged files now
carried PRODUCT-LOCAL inside `apps/dreamverse/server/`. Public
`fastvideo.entrypoints.streaming.*` reverts to `fbd823df` state. See
[decisions-log.md D-19](decisions-log.md#d-19) +
[D-20](decisions-log.md#d-20) for full context.).
FastVideo `will/ltx2_sr_port` @ HEAD (post-D-17 STACK.md removal +
integration-review.md addition + integration-plan.md addition + D-18
reconciliation). Dreamverse `will/integrate-public-fastvideo` @ `ec8ef92`.
PRs #1257 / #1258 / #1284 / #1286 MERGED to main. **PR #1287 CLOSED
(in favor of consolidation); PR #1288 OPEN as the single mega-PR
landing the entire `will/ltx2_sr_port` chain at once** (LTX-2 SR
runtime + NVFP4 + `generate_async`/Dynamo contract + agents memory dir).
Split branches kept as historical bookmarks; STACK.md model **abandoned** —
see [decisions-log.md D-17](decisions-log.md#d-17). Local backup
`will/ltx2_sr_port-pre-1286-rebase` @ `1baa60bb` preserves the
pre-rebase chain.
## Fresh-context onboarding (read in order)
If you're an agent picking up this work for the first time, do these
**5 things in this order**. Once done, you have full context to continue
any open thread, commit correctly, push, and propagate to the open PR.
1. **Confirm worktree state** — run the "First 60 seconds" block in
[runbook.md](runbook.md). Tells you the branch is right, services
are up, and PR #1286's head matches what this dir claims.
2. **Read [state.md](state.md)** — single-page snapshot of branch tips,
live services, test status, pre-existing failures, "do not pop"
stashes.
3. **Read [pr-roadmap.md](pr-roadmap.md)** — what PRs landed, what's in
flight, what's planned. Identifies the active open PR (currently
#1286) and where it sits in the dependency chain.
4. **Read [open-threads.md](open-threads.md)** — prioritized work items
with effort estimates and dependencies. The "Recommended pull order"
section is a ready-made TODO list if you need one.
5. **Skim [runbook.md](runbook.md) end-to-end** — operational how-to:
verify, commit (with co-author trailers), push, propagate to PR
#1286, maintain the memory dir, and the "Common pitfalls" section
that catches the recurring traps.
Skip the deep-context docs (design / streaming-server / cross-repo /
quantization / decisions-log) until you need them — they're indexed in
the "Deep-dive reading guide" below.
Final check: run the "Self-test" block at the bottom of
[runbook.md](runbook.md). If you can answer all 8 questions from this
dir alone, you're ready. If you can't, the gap is a memory-dir bug —
file it in [open-threads.md](open-threads.md) before continuing.
## Deep-dive reading guide
| Question / task | File |
|---|---|
| "What's running right now? What just landed?" | [state.md](state.md) |
| "How do I commit / push / propagate to PR #1286?" | [runbook.md](runbook.md) |
| "Why is the schema typed this way? What's the philosophy?" | [design.md](design.md) |
| "What PRs landed? In flight? Planned?" | [pr-roadmap.md](pr-roadmap.md) |
| "Streaming server, `generate_async`, `build_app` routes?" | [streaming-server.md](streaming-server.md) |
| "How does Dreamverse use FastVideo? What about Dynamo?" | [cross-repo-surfaces.md](cross-repo-surfaces.md) |
| "NVFP4? Layer profiles? `LinearBase` fallback? AbsMaxFP8?" | [quantization.md](quantization.md) |
| "Why was decision X made? What's resolved vs. open?" | [decisions-log.md](decisions-log.md) |
| "What should I work on next? Priority order?" | [open-threads.md](open-threads.md) |
| "Who should be co-authored on commits in this scope?" | [authors.md](authors.md) |
| "How do we execute the Dreamverse → FastVideo monorepo merge?" | [integration-plan.md](integration-plan.md) ← **CURRENT** |
| "Historical drift audit + Option-D evaluation (deprecated by D-18)" | [integration-review.md](integration-review.md) (DEPRECATED) |
## Repo + worktree paths
| Repo | Path | Active branch |
|---|---|---|
| FastVideo (public) | `/home/william5lin/FastVideo` | `will/ltx2_sr_port` |
| Dreamverse | `/home/william5lin/Dreamverse` | `will/integrate-public-fastvideo` |
| FastVideo-internal (read-only ref) | `/home/william5lin/FastVideo-internal` | their `main` |
| Dynamo (read-only ref) | `/home/william5lin/dynamo` | upstream |
## Glossary
- **NVFP4**: NVIDIA's specific block-scaled FP4 (e2m1 mantissa, fp32 alpha,
`layout_128x4` scale layout, group size 16). Distinct from MX-FP4 / OCP-FP4.
- **`GeneratorConfig`**: typed init-time public config (model_path, engine,
pipeline). Replaces flat `from_pretrained(**kwargs)`.
- **`GenerationRequest`**: typed per-call request (prompt, inputs, sampling,
runtime, output, stage_overrides, state, plan, extensions). Replaces flat
`generate_video(**kwargs)`.
- **`ServeConfig`** / **`RunConfig`**: top-level YAML envelopes. ServeConfig
for `fastvideo serve`; RunConfig for offline `fastvideo generate`.
- **`InferencePreset`**: model-owned named preset (e.g. `ltx2_two_stage`)
defining stage topology + per-stage defaults + valid override types.
- **`ContinuationState`**: opaque round-trip state envelope `{kind, payload}`.
Hybrid: server-held for streaming WS, client-round-trip for stateless HTTP.
- **`generate_async`**: future canonical async exec API (PR 7.10) yielding
`VideoProgressEvent` / `VideoPartialEvent` / `VideoFinalEvent`. Substrate
for streaming server, OpenAI server, AND Dynamo backend.
- **`build_app`**: FastAPI app factory in
`fastvideo.entrypoints.streaming.server`. Currently exposes only
`/health` + `/v1/stream`. FE-required `/healthz`+`/readyz`+`/status`
migration is open follow-up #1.
- **`LLMProvider`**: protocol abstraction for prompt enhancer providers
(cerebras, cerebras_ifm, groq). Public schema currently restricts to
`Literal["cerebras", "groq"]`; `cerebras_ifm` is internal-only.
- **`compat.py`**: legacy kwargs translation layer (~370 lines). Scheduled
for death across PRs 14-17.
- **`prepare_for_compile`**: duck-type protocol method called via
`getattr(module, "prepare_for_compile", None)` before `torch.compile`.
Currently only Gemma3 implements it.
- **`SubprocessGpuPool`**: PR 7.6 public replacement for the internal
`realtime/local_runtime.GPUPool`. Per-GPU subprocess workers, typed
`GeneratorConfig` boundary.
- **PR 5.5**: streaming server subpackage skeleton — adds
`fastvideo/entrypoints/streaming/` parallel to `openai/`.
- **PR 7.10**: the unlock PR. Closes Q-5 (audio re-encode), Q-9 (Dynamo
progress), and PR 7.5's mid-segment cancellation TODO simultaneously.
## Live process map (as of 2026-05-03)
| Port | Service | Source |
|---|---|---|
| 8009 | `dreamverse-server` | running, `/readyz` 200, 1 warmed GPU worker |
| 5274 | `next-server` (dev) | running |
| 8000 | unknown FastAPI | not in handoff — verify before launching new BE |
## How this directory is maintained
- Source of truth for the integration story. Update when state changes.
- Each file has a "Last updated" header; bump when you edit.
- Cross-reference siblings via relative links; do NOT duplicate content.
- New entries: register in `../index.jsonl`.
- These files supersede the untracked source docs in the repo root and
`.agents/exploration/` — see [state.md](state.md) "Untracked but
present" section for disposition.
## Source documents (archived 2026-05-03)
The 7 source docs that this directory consolidates have been moved into
[`source-archive/`](source-archive/). They remain available for agents
who want the full unsynthesized rationale, but the synthesized memory
files in this dir are the canonical source of truth.
| Source doc | Lines | Synthesized into |
|---|---|---|
| [`source-archive/apirefactor.md`](source-archive/apirefactor.md) | 838 | [design.md](design.md) |
| [`source-archive/PR-plan.md`](source-archive/PR-plan.md) | 1145 | [pr-roadmap.md](pr-roadmap.md) |
| [`source-archive/dreamverse_review.md`](source-archive/dreamverse_review.md) | 390 | [state.md](state.md) + [decisions-log.md](decisions-log.md) |
| [`source-archive/handoff-nvfp4-launch-demo.md`](source-archive/handoff-nvfp4-launch-demo.md) | 518 | [state.md](state.md) + [quantization.md](quantization.md) + [open-threads.md](open-threads.md) |
| [`source-archive/streaming-server-upstream-plan.md`](source-archive/streaming-server-upstream-plan.md) | 539 | [streaming-server.md](streaming-server.md) + [decisions-log.md](decisions-log.md) |
| [`source-archive/dreamverse_integration.md`](source-archive/dreamverse_integration.md) | 285 | [cross-repo-surfaces.md](cross-repo-surfaces.md) |
| [`source-archive/video-generator-config-api-design.md`](source-archive/video-generator-config-api-design.md) | 93 | [design.md](design.md) (early-draft material) |
| `.agents/exploration/pr-link-review.md` | 29 | already promoted to `.agents/skills/review-pr-link/` (kept in exploration dir) |
See [`source-archive/README.md`](source-archive/README.md) for the
archive policy.
@@ -1,149 +0,0 @@
# Authors — Dreamverse Integration
**Status:** PERMANENT — keep around as the source of truth for who collaborated
on the dreamverse-integration work, even after every PR in the integration
scope has merged.
**Last updated:** 2026-05-05 (strategy reversal — single mega-PR #1288 on `will/ltx2_sr_port` replaces planned 6-PR split; #1287 closed; per [decisions-log.md D-17](decisions-log.md#d-17))
This file documents the human co-authors credited on every commit in the
dreamverse-integration scope (FastVideo public-API refactor, streaming server
upstream, GPU pool, prompt enhancer, NVFP4 wire-up, LTX-2 SR port). The
4 collaborators below worked on the FastVideo-internal precursor of this code
and are credited as co-authors on every public-side upstream commit via Git's
standard
[`Co-authored-by`](https://docs.github.com/en/pull-requests/committing-changes-to-your-project/creating-and-editing-commits/creating-a-commit-with-multiple-authors)
trailer convention.
Scope-wise this is the dreamverse-integration-flavored mirror of
[co-authors.md](co-authors.md), which is scoped to the broader
`will/ltx2_sr_port` 10-PR stack. The roster is identical; both files now live
in this memory dir so the provenance docs stay self-contained and discoverable.
## Co-author roster
| GitHub user | Real name | GitHub ID | Trailer email |
|---|---|---|---|
| [`@Davids048`](https://github.com/Davids048) | Junda (David) Su | 90978028 | `90978028+Davids048@users.noreply.github.com` |
| [`@RandNMR73`](https://github.com/RandNMR73) | Matthew Noto | 99706358 | `99706358+RandNMR73@users.noreply.github.com` |
| [`@XOR-op`](https://github.com/XOR-op) | (unset) | 17672363 | `17672363+XOR-op@users.noreply.github.com` |
| [`@jzhang38`](https://github.com/jzhang38) | Zhang Peiyuan | 42993249 | `42993249+jzhang38@users.noreply.github.com` |
## Verification — where these trailers appear
Verified via `gh pr view <PR> --json commits --jq '.commits[].messageBody'`
across every PR in the integration scope:
| PR | Branch | Status | Trailers present on every commit |
|---|---|---|---|
| #1257 | `will/api_7.6` (GPU pool upstream) | ✅ merged 2026-05-04 | yes (4/4) |
| #1258 | `will/api_7.7` (prompt enhancer + LLMProvider) | ✅ merged 2026-05-04 | yes (3/3) |
| #1284 | `will/api_7.8` (streaming auxiliaries) | ✅ merged 2026-05-04 | yes (2/2) |
| #1286 | `will/api_7.9` (streaming router) | ✅ merged 2026-05-05 at `2aaeee2a` (squash) | yes on commits 1-3; commit `a152cb77` (`[fix] streaming: router polish`) was missing trailers but got squashed into the merge commit, so the merge commit on main inherits the trailers from the other 3. The trailerless cherry-pick partner (`40e265b8` on `will/ltx2_sr_port`) was dropped by the post-#1286 rebase — gap permanently resolved. |
| #1287 | `will/api_7.10` (`generate_async` + `VideoEvent`) | ❌ CLOSED 2026-05-05 — superseded by #1288 per [D-17](decisions-log.md#d-17) | yes on all 3 commits (now part of #1288's chain) |
| **#1288** | **`will/ltx2_sr_port`** (mega-PR — full stack: SR runtime + NVFP4 + generate_async + Dynamo contract + agents memory + integration-review) | 🟢 OPEN, MERGEABLE at `b36bdbc9`, 36 commits / 70 files / ~+13.0k LOC (post STACK.md removal) | yes on all 36 commits |
Aggregate count across `will/ltx2_sr_port` (top of stack) at the time of
writing: 32-33 commits per co-author, matching the 32 commits in the stack
on top of base `cfccd292`. Numbers stay consistent because the rebase
command (see "How the trailers were applied" below) walks every commit.
## Trailer block (copy-paste ready)
The trailers added to every commit on `will/ltx2_sr_port` and every
dreamverse-integration PR:
```
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
```
For one-off `git commit -m` invocations, use `--trailer` flags:
```bash
git commit -m "..." \
--trailer "Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>" \
--trailer "Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>" \
--trailer "Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>" \
--trailer "Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>"
```
`--trailer` is idempotent (dedupes by full `key: value`) so re-running is safe.
## Why no-reply emails
GitHub's `<id>+<username>@users.noreply.github.com` form is the most reliable
way to link a `Co-authored-by` trailer to a GitHub account. It:
- Always works regardless of whether the user has a public verified email
- Survives the user changing their primary email
- Doesn't expose anyone's personal email to git history
- Is the format GitHub itself produces when you click "Add co-author" in the
web UI
(All 4 collaborators have this email already used in `FastVideo-internal`
git history, verified via `git log --all` on that repo.)
## How the trailers were applied (bulk rebase)
```bash
git rebase --exec '
git commit --amend --no-edit \
--trailer "Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>" \
--trailer "Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>" \
--trailer "Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>" \
--trailer "Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>"
' origin/main will/ltx2_sr_port
```
After running, re-slice all 10 split branches per [`STACK.md`](../../../STACK.md)
and force-push the published branches (`will/api_7.9`, `will/ltx2_sr_port`).
## How to add a new co-author later
1. Add the user to the roster table above (and [co-authors.md](co-authors.md)
— keep them in sync).
2. Append their `Co-authored-by` line to the trailer block above.
3. Re-run the bulk rebase command on `will/ltx2_sr_port` — git's trailer
dedupe handles the existing 4; the new one gets appended.
4. Re-slice all split branches per [`STACK.md`](../../../STACK.md).
5. Force-push the published branches.
## What we do NOT add
Per the repo's top-level [`AGENTS.md`](../../../AGENTS.md):
> Never add any coding agent or models such as Claude (or Claude Code), GPT,
> Codex or others as a co-author in commits or PRs. Do not include
> `Co-Authored-By: Claude ...` trailers or "Generated with Claude Code" and
> other such lines.
So no `Co-authored-by: Claude <noreply@anthropic.com>`, no
`Generated with Claude Code` footer, no `Cursor <cursoragent@cursor.com>`
trailer (one such commit exists on `will/ltx2_sr_port` from a pre-policy
external contribution and stays grandfathered; new commits MUST NOT introduce
the pattern). Only human collaborators.
## Known gaps
**Resolved 2026-05-05 by the post-#1286 rebase.** The two trailerless
commits (`a152cb77` on `will/api_7.9` and `40e265b8` on
`will/ltx2_sr_port`) are no longer reachable from any active branch:
- `a152cb77` was absorbed into squash merge `2aaeee2a` on main, which
inherits the trailers from the other 3 commits in the squash.
- `40e265b8` was dropped by the post-#1286 rebase of
`will/ltx2_sr_port`.
Both still exist on the local backup `will/ltx2_sr_port-pre-1286-rebase`
for archeological reference. No further action needed.
## See also
- [co-authors.md](co-authors.md) — stack-scoped co-authors file (same roster,
broader scope)
- [`../../../STACK.md`](../../../STACK.md) — 10-PR split layout for
`will/ltx2_sr_port` (re-slice commands live here)
- [`pr-roadmap.md`](pr-roadmap.md) — per-PR status within the
dreamverse-integration scope
@@ -1,182 +0,0 @@
# `.agents/` Cleanup Log — Phase 1 (Deletes Only)
**Status:** TEMPORARY — delete this file after the cleanup is reviewed/committed.
**Date:** 2026-05-04
**Branch:** `will/ltx2_sr_port`
**Scope:** Phase 1 of the `.agents/` cleanup plan (deletes only; no rewrites or additions).
For the full multi-phase plan, see the prior session analysis. This file tracks
exactly what got deleted, why, and what cross-references still point at deleted
content (to fix in a future phase).
---
## Deletions executed
### Files deleted
| Path | Size | Reason |
|---|---|---|
| `.agents/STATUS.md` | 3.85 KB | Stale dashboard, last synced 2026-03-02. Counts wrong (claimed 8 skills/4 workflows/4 memory; actual 9/5/5). References old snake_case filenames (`codebase_map.md`/`experiment_journal.md`) that don't exist. Hand-maintained derivative of `.agents/{memory,skills}/index.jsonl` — strictly redundant. |
| `.agents/exploration/pr-link-review.md` | 1.11 KB | Status: "promoted" to `.agents/skills/review-pr-link/`. Per `.agents/exploration/README.md` lifecycle, promoted exploration logs should not linger after the skill exists. |
| `.agents/workflows/sync-dashboard.md` | 1.87 KB | SOP for maintaining `STATUS.md` (which is also deleted). Contained obsolete file paths (`.agents/skills/launch-experiment.md` flat layout vs. actual `<skill>/SKILL.md` per-dir layout). Has never been run successfully (judging by stale dates everywhere). |
### Skill directories deleted
| Path | Size | Reason |
|---|---|---|
| `.agents/skills/index-related-work/` | 2.18 KB | Vapor-skill operating on the empty `.agents/memory/related-work/` registry. Never used (the registry has zero entries despite ~6 weeks since skill creation). Re-add when the related-work catalog gains entries. |
| `.agents/skills/search-related-work/` | 1.91 KB | Same: vapor-skill against empty registry. The skill description literally requires "The related work index has entries" as a prerequisite, and there are none. |
**Total deleted: 5 items, ~10.9 KB.**
### Registry updates
| File | Change |
|---|---|
| `.agents/skills/index.jsonl` | Removed entries for `index-related-work` and `search-related-work`. Was 9 entries; now 7. |
### Symlink hygiene
`.agents/scripts/sync-skills.sh` was run to prune now-stale symlinks under
`.claude/skills/` that pointed at the deleted skill directories. Output captured
in the run log.
---
## What was KEPT (despite being candidates)
| Path | Why kept |
|---|---|
| `.agents/scripts/sync-skills.sh` | User explicitly requested keep. **Verified**: this script is INDEPENDENT of STATUS.md / sync-dashboard.md. It mirrors `.agents/skills/` → `.claude/skills/` via symlinks for Claude Code skill discovery. Self-contained, useful, prunes its own stale symlinks. |
| `.agents/memory/related-work/README.md` | Empty placeholder, but the schema/template is reusable. Kept for when first related-work entry is added. |
| `.agents/memory/experiment-journal/README.md` | Same: empty placeholder with template; kept for when journaling begins. |
| `.agents/lessons/README.md` | Same: empty placeholder, reusable schema. |
| `.agents/exploration/README.md` | Active template for new exploration logs. Kept. |
---
## Remaining broken cross-references (FOLLOW-UP NEEDED)
These files still reference deleted content. **NOT fixed in Phase 1** — track for
the next pass (Phase 2: rewrites/dedupe).
### References to deleted `STATUS.md`
| Referencing file | Action needed |
|---|---|
| `.agents/onboarding/README.md` | Quick-reference tree (line ~65) lists `STATUS.md ← dashboard: completeness & trust of all components`. Remove that line + the `ONBOARDING.md` typo (file is `README.md`). |
### References to deleted `pr-link-review.md`
| Referencing file | Action needed |
|---|---|
| `.agents/memory/dreamverse-integration/state.md` | "Untracked but present" / "Source docs (archived)" sections still mention `pr-link-review.md` as kept. Update to reflect deletion. |
| `.agents/memory/dreamverse-integration/README.md` | Same — table row for `pr-link-review.md` says "kept in exploration dir". Update or remove the row. |
### References to deleted skills (`index-related-work`, `search-related-work`)
| Referencing file | Action needed |
|---|---|
| `.agents/memory/related-work/README.md` | Says "Use the `index-related-work` skill". Either remove that hint or note "skill removed; re-add when registry has entries". |
| `.agents/workflows/evaluation-development.md` | Step 1 says "Search `.agents/memory/related-work/` for existing evaluation approaches" — that's still valid (manual search). No change needed. |
### References to deleted `sync-dashboard.md`
| Referencing file | Action needed |
|---|---|
| `.agents/memory/evaluation-registry/README.md` | Doesn't reference sync-dashboard directly. No change. |
| `.agents/STATUS.md` | Already being deleted. |
---
## Other registry inconsistencies discovered (NOT FIXED in Phase 1)
While editing `.agents/skills/index.jsonl`, two skill directories were found
that exist on disk but **are not registered** in `index.jsonl`:
| Skill dir | Status | Why missing from index |
|---|---|---|
| `.agents/skills/diagnose-ssim-failure/` | Untracked locally; NOT on `origin/main`. 12.3 KB SKILL.md + `scripts/compare_latent_pt.py`. Recent mtime (2026-05-01). | Created in a prior session but the registration step was skipped. |
| `.agents/skills/review-pr-link/` | Untracked locally; NOT on `origin/main`. 2.9 KB SKILL.md + `scripts/prepare_pr_review.py` + `agents/openai.yaml`. The promotion target of the deleted `pr-link-review.md` exploration log. | Skipped registration when promoted from exploration log. |
Both skills are functional and exposed via `sync-skills.sh` symlinks (just verified in
`.claude/skills/`), but agents reading `index.jsonl` to discover skills will miss them.
**Action for Phase 2**: Add entries to `.agents/skills/index.jsonl` for both,
likely with `trust: medium` since they have working scripts and recent use.
---
## Skill registry parity check
After Phase 1, `.agents/skills/` contains 9 directories but `index.jsonl` lists 7:
| In `index.jsonl` | On disk |
|---|---|
| ✓ launch-experiment | ✓ launch-experiment/ |
| ✓ monitor-experiment | ✓ monitor-experiment/ |
| ✓ summarize-run | ✓ summarize-run/ |
| ✓ log-experiment | ✓ log-experiment/ |
| ✓ evaluate-video-quality | ✓ evaluate-video-quality/ |
| ✓ seed-ssim-references | ✓ seed-ssim-references/ |
| ✓ reseed-ssim-references | ✓ reseed-ssim-references/ |
| ❌ (missing) | ⚠ diagnose-ssim-failure/ |
| ❌ (missing) | ⚠ review-pr-link/ |
`.claude/skills/` symlinks (the runtime-discoverable surface) include all 9 ✓.
---
## Phase 2+ items (NOT executed in this session)
For future cleanup sessions, the prior plan identified:
**Phase 2 (rewrites)**:
- Rewrite `.agents/onboarding/worldmodel-training/README.md` to drop ~50% structural duplication with `codebase-map/README.md`
- Refresh `.agents/memory/codebase-map/README.md` (last updated 2026-03-08; missing `fastvideo/api/`, `fastvideo/entrypoints/streaming/`, etc.)
- Refresh `.agents/memory/evaluation-registry/README.md` (last updated 2026-03-02; references old `evaluation_registry.md` filename)
- Merge `.agents/workflows/experiment-journaling.md` into `experiment-lifecycle.md` (one SOP per workflow)
- Fix the broken cross-references listed above
**Phase 3 (additions)**:
- `fastvideo/api/AGENTS.md`
- `fastvideo/entrypoints/AGENTS.md`
- `tests/AGENTS.md` (top-level, distinct from `fastvideo/tests/AGENTS.md`)
- `fastvideo/distributed/AGENTS.md`
- `examples/AGENTS.md`
- `docs/AGENTS.md`
- `benchmarks/AGENTS.md`
**Phase 4 (registry)**:
- Add `.agents/workflows/index.jsonl`
- Standardize all three index.jsonl schemas
**Phase 5 (skills quality)**:
- Promote tested skills (`seed-ssim-references`, `reseed-ssim-references`, `diagnose-ssim-failure`, `review-pr-link`) from `trust: low` to `trust: medium`
- Mark untested skills (`launch-experiment`, `monitor-experiment`, `summarize-run`, `log-experiment`, `evaluate-video-quality`) explicitly with their gating prerequisite (e.g. "operates on empty registry")
---
## Recovery
All deletions are local (`will/ltx2_sr_port`, not committed). To restore any
deleted file:
```bash
git restore --source=HEAD .agents/STATUS.md
git restore --source=HEAD .agents/exploration/pr-link-review.md
git restore --source=HEAD .agents/workflows/sync-dashboard.md
git restore --source=HEAD .agents/skills/index-related-work/SKILL.md
git restore --source=HEAD .agents/skills/search-related-work/SKILL.md
```
---
## When to delete THIS file
Once:
1. The Phase 1 deletions are committed (or merged), AND
2. Phase 2 (broken cross-reference cleanup) is also committed,
remove this file. Its purpose is transient bookkeeping for a multi-phase cleanup.
@@ -1,88 +0,0 @@
# Co-Authors — `will/ltx2_sr_port` Stack
**Status:** PERMANENT — keep around as the source of truth for who collaborated on this work, even after the stack merges.
**Last updated:** 2026-05-04
This file documents the human co-authors credited on every commit in the
`will/ltx2_sr_port` stack and its 10 split PRs. The 4 collaborators below
worked on the FastVideo-internal precursor of this code (LTX-2 streaming
server, NVFP4 wire-up, GPU pool, prompt enhancer, etc.) and are credited as
co-authors on the public-side upstream commits via Git's standard
[`Co-authored-by`](https://docs.github.com/en/pull-requests/committing-changes-to-your-project/creating-and-editing-commits/creating-a-commit-with-multiple-authors)
trailer convention.
The trailers are added to every commit on `will/ltx2_sr_port` (see
[`STACK.md`](STACK.md)), which means GitHub will:
- Show the 4 co-authors on every commit detail page
- Show them on the merge commit / squash commit summary
- Display their avatars in the PR's "Contributors" sidebar
- Surface them in [`/contributors`](https://github.com/hao-ai-lab/FastVideo/contributors) once the stack lands
## Co-author roster
| GitHub user | Real name | GitHub ID | Trailer email |
|---|---|---|---|
| [`@Davids048`](https://github.com/Davids048) | Junda (David) Su | 90978028 | `90978028+Davids048@users.noreply.github.com` |
| [`@RandNMR73`](https://github.com/RandNMR73) | Matthew Noto | 99706358 | `99706358+RandNMR73@users.noreply.github.com` |
| [`@XOR-op`](https://github.com/XOR-op) | (unset) | 17672363 | `17672363+XOR-op@users.noreply.github.com` |
| [`@jzhang38`](https://github.com/jzhang38) | Zhang Peiyuan | 42993249 | `42993249+jzhang38@users.noreply.github.com` |
## Why no-reply emails
GitHub's `<id>+<username>@users.noreply.github.com` form is the most reliable
way to link a `Co-authored-by` trailer to a GitHub account. It:
- Always works regardless of whether the user has a public verified email
- Survives the user changing their primary email
- Doesn't expose anyone's personal email to git history
- Is the format GitHub itself produces when you click "Add co-author" in the
web UI
(All 4 collaborators have this email already used in `FastVideo-internal`
git history, verified via `git log --all` on that repo.)
## Trailer block (copy-paste ready)
The trailers added to every commit on `will/ltx2_sr_port`:
```
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
```
## How the trailers were applied
```bash
git rebase --exec '
git commit --amend --no-edit \
--trailer "Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>" \
--trailer "Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>" \
--trailer "Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>" \
--trailer "Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>"
' origin/main will/ltx2_sr_port
```
Git's `--trailer` flag is idempotent (it dedupes by the full `key: value`
string), so re-running the rebase is safe and won't add duplicates.
## How to add a new co-author later
1. Add the user to the roster table above.
2. Append their `Co-authored-by` line to the trailer block.
3. Re-run the rebase command above on `will/ltx2_sr_port` — git's
trailer dedupe handles the existing 4; the new one gets appended.
4. Re-slice all 10 split branches per [`STACK.md`](STACK.md).
5. Force-push `will/api_7.6`, `will/api_7.7`, and `will/ltx2_sr_port`.
## What we do NOT add
Per [`AGENTS.md`](AGENTS.md):
> Never add any coding agent or models such as Claude (or Claude Code), GPT,
> Codex or others as a co-author in commits or PRs.
So no `Co-authored-by: Claude <noreply@anthropic.com>` or similar. Only
human collaborators.
@@ -1,277 +0,0 @@
# Cross-Repo Surfaces — Dreamverse + Dynamo
How Dreamverse consumes FastVideo today, what's already shared, what's
ad hoc, and what migrations land alongside each PR. Plus the Dynamo
backend contract.
For the streaming-server side see [streaming-server.md](streaming-server.md).
For the API design see [design.md](design.md). For PR sequence see
[pr-roadmap.md](pr-roadmap.md).
**Last updated:** 2026-05-03.
## The three surfaces
Dreamverse depends on FastVideo across three surfaces (in order of
stability):
1. **Pipeline construction** (stable)
2. **Realtime runtime** (in flight: PRs 7.5/7.6)
3. **Continuation state** (PR 7 typed; PR 7.6 wires server-held)
## Surface 1: Pipeline construction (stable)
`Dreamverse/server/video_generation.py:VideoGenerationWorker` calls
`VideoGenerator.from_pretrained(...)`.
After PR 6 the typed `GeneratorConfig` path exists; **as of `d80c2a8`
(May 2)** Dreamverse migrated to the typed path:
| Dreamverse usage | FastVideo public surface (post-PR 6) |
|---|---|
| `VideoGenerator.from_pretrained(model_path, ltx2_refine_enabled=…, …)` | `VideoGenerator.from_pretrained(config=GeneratorConfig(...))` |
| Flat `torch_compile_kwargs={…}` dict | `engine.compile.{backend,fullgraph,mode,dynamic,extras}` |
| `ltx2_vae_tiling=True` | `pipeline.vae_tiling=True` |
| `ltx2_refine_*` family | `pipeline.preset_overrides.refine.*` + `pipeline.components.upsampler_weights` |
| `enable_torch_compile_text_encoder` | `engine.compile.text_encoder_enabled` |
Refine knobs moved from `ltx2_refine_*` flat kwargs into
`preset_overrides["refine"]`. **The in-memory `pipeline_config` pin**
(`dit_config.quant_config = NVFP4Config()`) keeps using the legacy
`experimental["pipeline_config"]` carrier because typed
`transformer_quant: "NVFP4"` doesn't yet support setting
`layer_profile` (see [open-threads.md](open-threads.md) follow-up #4 +
[quantization.md](quantization.md)).
Legacy flat-kwarg path stays supported via `compat.py`; migration is
opt-in. PR 13's deprecation warnings are the eventual nudge.
## Surface 2: Realtime runtime (in flight: PRs 7.5–7.6)
`Dreamverse/server/runtime/factory.py` selects a runtime backend at
process start:
```python
def create_runtime_pool() -> RuntimePool:
if os.getenv("FASTVIDEO_REALTIME_BASE_URL"):
return FastVideoRealtimePool(base_url=..., ws_url=..., default_model_id=...)
return GPUPool(get_available_gpus()) # in-process, wraps
# fastvideo.entrypoints.realtime.local_runtime
```
Both backends speak the same `RuntimePool` / `RuntimeSlot` Protocol
(`server/runtime/interfaces.py`):
- `acquire(client_id, websocket=None) -> (gpu_id, RuntimeSlot)`
- `release(client_id)`
- `RuntimeSlot.{join_user, user_step, leave_user, register_stream_queue, ...}`
Today both impls reach into FastVideo-internal's
`fastvideo.entrypoints.realtime.local_runtime` (which exposes
`RealtimeRuntimeConfig`, `GPUPool`, `GPUSlot`). The remote backend talks
HTTP+WS to a separately-deployed runtime of the same shape.
**Contract that PR 7.5/7.6 must preserve:**
- `RealtimeRuntimeConfig` accepts `model_registry`, `default_model_id`,
`default_height/width/num_frames/fps/num_inference_steps/guidance_scale/seed/negative_prompt`,
`default_ltx2_image_crf`, `startup_warmup_{enabled,prompt,timeout_seconds}`.
- `GPUPool(gpu_ids: list[int], config: RealtimeRuntimeConfig)` constructor.
- `pool.initialize() / shutdown() / acquire() / release() / get_status()`.
- HTTP endpoints on the remote variant: `GET /healthz`, `GET /readyz`,
`GET /status`, `WS /ws`. (Already match what
`Dreamverse/server/routes/health.py` consumes.)
**These three health routes still need to migrate into FastVideo's
`build_app` to make `BE_FLAVOR=fastvideo` FE-compatible** — see
[streaming-server.md](streaming-server.md) "build_app route contract" +
[open-threads.md](open-threads.md) follow-up #1.
When PR 7.6 lands the upstream of `fastvideo/entrypoints/realtime/`,
Dreamverse should not need any code change unless the import path
renames. Decided: keep `streaming/` (post-PR-5.5 public name); ship
`realtime/__init__.py` as a re-export with `DeprecationWarning` for one
release cycle.
### Note on `default_ltx2_image_crf`
Dreamverse's `RealtimeRuntimeConfig` includes `default_ltx2_image_crf`.
The April 26 Dreamverse review (D-8) showed this getting passed to
`SamplingParam(...)` and **silently dropped** by the public schema. Post
`d80c2a8` (May 2 typed-config refactor), the migration target is
`request.stage_overrides.refine.image_crf` (per
[design.md](design.md) compatibility mapping table).
**Whether `d80c2a8` actually wired this through, or it's still latent,
is unverified.** See [open-threads.md](open-threads.md) item D-8.
## Surface 3: Continuation state (PR 7)
`Dreamverse/server/video_generation.py:89 ContinuationState` is
Dreamverse's hand-rolled per-session state holder. PR 7 introduced the
typed equivalent at
[`fastvideo/pipelines/basic/ltx2/continuation.py`](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/continuation.py).
### Field mapping
| Dreamverse | PR 7 `LTX2ContinuationState` | Notes |
|---|---|---|
| `video_images: list[PIL.Image]` | `video_frames: list[np.ndarray]` (uint8 H×W×3) | numpy is leaner; Dreamverse already round-trips PIL→numpy→PIL just to add noise |
| `audio_latents: torch.Tensor` `[B, C, T, mel]` | `audio_latents: torch.Tensor` (safetensors-serialized; bf16-safe) | unchanged shape; safetensors preserves bf16 |
| `LTX2_VIDEO_CONDITIONING_FRAME_IDX` (env) | `video_conditioning_frame_idx: int` | env constant → per-state field |
| `LTX2_VIDEO_CONDITIONING_STRENGTH` (env) | `video_conditioning_strength: float` | env constant → per-state field |
| `AUDIO_CONDITIONING_NUM_FRAMES` (env) | `audio_conditioning_num_frames: int` | env constant → per-state field |
| `AUDIO_CONDITIONING_STRENGTH` (env) | `audio_conditioning_strength: float` | env constant → per-state field |
| `audio_lps` (passed into `apply_audio`) | `audio_sample_rate: int \| None` | analogous; rename worth confirming with audio team |
| Computed `prefix_sec` per segment | `video_position_offset_sec: float` | **see open question below** |
| `segment_idx` (param to `apply_*`) | `segment_index: int` | per-state field |
| `VIDEO_CONTEXT_NOISE`, `AUDIO_CONTEXT_NOISE`, `ENABLE_AUDIO_COND` | not on state | runtime policy / regularization knobs, not portable session data |
| `apply_video / apply_audio / save_video / save_audio_latents / clear` | not on PR-7 state class | state is a pure data carrier; runtime owns lifecycle policy |
PR 7 is a strict superset of Dreamverse's data model **plus** lifts
several env globals into per-session typed fields.
### Lifecycle mapping
| Dreamverse pattern | `SessionStore` API |
|---|---|
| `self.continuation = ContinuationState()` per session | `state = session_store.snapshot(sid) or LTX2ContinuationState()` |
| `apply_video(req_kwargs, segment_idx)` + `apply_audio(req_kwargs, segment_idx, audio_lps)` | `state = session_store.snapshot(sid)`; runtime builds request from `state.video_frames` / `state.audio_latents` |
| `save_video(frames)` + `save_audio_latents(latents)` | runtime constructs new `LTX2ContinuationState`, `session_store.store(sid, ...)` |
| `clear()` at end of session | `session_store.drop(sid)` |
`SessionStore` and `BlobStore` ABCs ship with thread-safe in-memory
defaults (`InMemorySessionStore`, `InMemoryBlobStore`). Dreamverse can
adopt them as-is for the local runtime; remote runtimes can plug in
redis-backed implementations later.
### Wire format (HTTP/WS round-trip)
Dreamverse's `FastVideoRealtimePool` already speaks the realtime
runtime's HTTP+WS protocol. When PR 7.5/7.6 land state emission on the
server side, the on-the-wire payload is the public envelope:
```json
{
"kind": "ltx2.v1",
"payload": {
"schema_version": 1,
"segment_index": 3,
"video_conditioning_frame_idx": 9,
"video_conditioning_strength": 0.75,
"audio_sample_rate": 24000,
"audio_conditioning_num_frames": 5,
"audio_conditioning_strength": 0.5,
"video_position_offset_sec": 0.2,
"video": {"frames_b64": ["..."]},
"audio": {"safetensors_b64": "..."},
"metadata": {}
}
}
```
JSON-serializable end-to-end; safetensors blob preserves audio dtype
(incl. bf16). For payloads above the inline threshold a `BlobStore`
indirection replaces the b64-encoded body with `{"blob_id": "..."}`;
the blob itself stays inside the runtime that produced it.
## Migration plan per PR
| PR | Dreamverse action |
|---|---|
| PR 6 (landed) | Typed `GeneratorConfig` available; flat-kwarg path still works via compat. Optional migration. |
| PR 7 (landed) | Typed `LTX2ContinuationState` available. ~50-line Dreamverse PR: replace `server/video_generation.py:89` import; move `apply_*`/`save_*`/`clear` off the state class onto `VideoGenerationWorker`; read knobs from typed state instead of env globals; swap `list[PIL.Image]` → `list[np.ndarray]`. |
| PR 7.5 (open) | Streaming server skeleton — Dreamverse's `runtime/factory.py` either keeps building `GPUPool` from `RealtimeRuntimeConfig` (current path), or migrates to `ServeConfig.streaming` shape and invokes `fastvideo serve --config realtime.yaml`. Dreamverse's `RuntimePool`/`RuntimeSlot` Protocol can stay in place. |
| PR 7.6 (branch ready) | GPU pool upstream — `local_runtime.py` import becomes a public import with same symbols (`RealtimeRuntimeConfig`, `GPUPool`, `get_available_gpus`). Per-GPU continuation state inside the worker becomes a `SessionStore` reference (Dreamverse doesn't see this). `request.state` / `result.state` round-trip starts working end-to-end on the local runtime. |
| PR 7.10 (planned) | `generate_async` is canonical. Dreamverse's per-segment `user_step` flow can migrate from sync `generate_video(..., **kwargs)` to consuming the typed event stream. Optional; sync wrapper stays. |
## Dynamo backend contract
**FastVideo does not host any Dynamo code.** The backend package
(`args.py`, `main.py`, `backend.py`, `register.py`, `health_check.py`,
adapter, Dockerfile) lives entirely in the Dynamo repo at
`components/src/dynamo/fastvideo/`, modeled on
`components/src/dynamo/sglang/`.
FastVideo's only obligation is to expose a stable, typed Python API
that Dynamo's backend package imports.
### Contract surface
| Surface | Exposed as |
|---|---|
| Construction | `VideoGenerator.from_pretrained(model_path, **typed_kwargs)` (typed_kwargs = a stable subset from `GeneratorConfig`; no flat LTX2 legacy) |
| Sync execution | `generator.generate_video(request: GenerationRequest) -> VideoResult` |
| Async execution | `generator.generate_async(request: GenerationRequest) -> AsyncGenerator[VideoEvent, None]` (PR 7.10) |
| Typed request | `fastvideo.api.GenerationRequest`, `SamplingConfig`, `InputConfig` |
| Typed result | `VideoResult` with `video_bytes` or tensor frames + optional `ContinuationState` |
| Continuation | `ContinuationState(kind, payload)` — schema-versioned payloads |
| Health-check input | `VideoGenerator.default_health_check_request() -> GenerationRequest` (256x256 / 8 frames / 1 step) |
| Config dump | `GeneratorConfig.to_dict()` / `ServeConfig.to_dict()` |
### Request/response mapping (Dynamo ↔ FastVideo)
```
NvCreateVideoRequest -> fastvideo.api.GenerationRequest
prompt -> sampling.prompt
size="WxH" -> sampling.width, sampling.height
seconds -> (seconds * nvext.fps) -> sampling.num_frames
input_reference -> input.image_path / input.video_path
nvext.fps -> sampling.fps
nvext.num_frames -> sampling.num_frames (overrides seconds*fps)
nvext.num_inference_steps -> sampling.num_inference_steps
nvext.guidance_scale -> sampling.guidance_scale
nvext.seed -> sampling.seed
nvext.negative_prompt -> sampling.negative_prompt
response_format -> (handled by adapter at output)
VideoFinalEvent -> NvVideosResponse
video_bytes -> data[0].b64_json (if response_format=b64_json)
video_url (after upload) -> data[0].url (if response_format=url)
metadata.inference_time_s -> inference_time_s
```
All fields exist on FastVideo's typed schema after PR 6 expansion (typed
LTX2 kwargs) + PR 7.10 (`generate_async` + health check).
### Reference: PR ai-dynamo/dynamo#7544
Closed draft establishing the Dynamo backend shape. Two frictions
identified:
1. Flat legacy LTX2 kwargs — solved by PR 6.
2. Sync-only generation — solved by PR 7.10's `generate_async`.
Next iteration of this PR (or its successor) will reopen against PR 8's
docs reference and land cleanly.
## Open questions across surfaces
| # | Question | Source | Status |
|---|---|---|---|
| Q-1 | Multi-model GPU pool | dreamverse_review D-1 | Deferred (production single-model) |
| Q-2 | LTX-2 prompt orchestration promotion to public | dreamverse_review D-2 | Open; consumer-side until 2nd consumer |
| Q-3 | Race-based provider fallback | dreamverse_review D-3 | Open; sequential is current public |
| Q-4 | Router upstream skip on Dreamverse | dreamverse_review D-4 | Resolved (PR 7.9 lands publicly, Dreamverse doesn't consume) |
| Q-5 / D-5 | `generate_async` cutover (audio re-encode) | dreamverse_review | **Blocked on PR 7.10** |
| D-6 | Don't upstream `realtime/local_runtime.py` | dreamverse_review | Resolved (Dreamverse switches to `streaming.gpu_pool.SubprocessGpuPool`) |
| D-7 / Q-6 | FP4Config public colocation | dreamverse_review | **Resolved May 2** — public NVFP4 landed with lazy flashinfer |
| D-8 | `ltx2_image_crf` silently dropped | dreamverse_review | **Unverified post-`d80c2a8`** — see [open-threads.md](open-threads.md) |
| D-9 | `aarch64-conda-linux-gnu-cc` triton compile failure | dreamverse_review | Operational; `ENABLE_TORCH_COMPILE=0` workaround |
| D-10 | Warmup OOM on shared GPU | dreamverse_review | Operational; idle-GPU pre-warm probe |
| D-11 | ffmpeg `Broken pipe` on disconnect | dreamverse_review | Cosmetic logging cleanup |
| — | `video_position_offset_sec` semantics (persistent vs per-segment) | dreamverse_integration | **Open — needs decision before PR 7.6 emits state** |
| — | `SessionStore` / `BlobStore` lifecycle (TTL/eviction/blob-drop) | dreamverse_integration | Open — defer to PR 7.5 design pass |
See [decisions-log.md](decisions-log.md) for full rationale per
decision.
## Don't / Cautions
- **Don't pop the Dreamverse stash on this branch.** It's 3867 lines of
orphan modular refactor with broken absolute imports.
- **Don't change `RealtimeRuntimeConfig` shape without coordinating
with Dreamverse `runtime/factory.py`.**
- **Don't promise public compatibility for private Dreamverse-only
field aliases.** Those belong in the private adapter layer per design
spec.
@@ -1,980 +0,0 @@
# Decisions Log — D + Q Resolutions
Cross-doc consolidated decision log. Each entry: ID, source doc,
question/decision, rationale, current status.
For implementation status see [pr-roadmap.md](pr-roadmap.md). For
follow-up actions see [open-threads.md](open-threads.md).
**Last updated:** 2026-05-06 (added D-21 — chunk-stutter root cause is software libx264 encoding consuming ~22% of segment wall-time, NOT a migration regression; verified by 3 parallel explore agents that NVFP4 + torch.compile coverage matches FastVideo-internal exactly; landed opt-in NVENC build path in install_native_ffmpeg.sh + `--nvenc`/`--no-nvenc` flag in dreamverse-deploy.sh + `apps/dreamverse/server/benchmarks/benchmark_av_streaming.py` regression test + memory dir update; default codec stays `libx264` for backward compat, opt-in via `--nvenc`. Earlier: added D-12 — GpuPool layer separation, Oracle review post-#1257-merge; added D-13 — prompt enhancer / LLMProvider abstraction shape, Oracle review pre-#1258-merge; added D-14 — streaming auxiliaries cohesion, Oracle review during #1284 review cycle; added D-15 — streaming router placement + sticky/active-active deferral, Oracle review during #1286 review cycle; added D-16 — streaming router polish round 2, second-pass review on top of D-15 covering bridge cancellation hygiene, registry state machine, httpx hard-fail, replica YAML parsing, and `websockets` dep; added D-17 — strategy reversal: abandon 6-PR split in favor of single mega-PR #1288 on `will/ltx2_sr_port`; added D-18 — Option B+ chosen: Dreamverse FE+product-server move into FastVideo as `apps/dreamverse/` subfolder while generic backend stays at `fastvideo.entrypoints.streaming.*`; integration-review.md deprecated, integration-plan.md is the executable migration plan; added D-19 — D-18 executed: 5 commits land on `will/dreamverse-monorepo`, fix-up commits corrected the integration-plan's invalid "delete generic-merged, import public substitutes" assumption — generic-merged files carried product-local instead, e2e passes against migrated code with /proc-verified evidence; added D-20 — segment-2 BrokenPipe root cause was a TWO-direction silent drop of LTX-2 audio kwargs in public `VideoGenerator`).
**Update 2026-05:** Dreamverse frontend tooling migrated from standalone pnpm to standalone npm. `apps/dreamverse/web/package-lock.json` is authoritative; see PR #1385.
## Status legend
- ✅ **Resolved** — decision made and implementation complete (or no implementation needed)
- 🟡 **Deferred** — decision made, implementation deferred to a known PR
- 🔴 **Open** — needs decision
## Post-merge architecture decisions
### D-26: Rebase `will/dreamverse-monorepo` directly onto `origin/main` (PR base flip)
**Status:** ✅ Resolved 2026-05-06 (UTC; 2026-05-07 local). 66 commits cleanly replayed via `git rebase origin/main`; force-pushed via `--force-with-lease`. Local backup branch `will/dreamverse-monorepo-pre-main-rebase-backup-20260506` preserved at the pre-rebase tip `2ee839a3`. PR #1288 on `will/ltx2_sr_port` is untouched.
**Source:** User directive: "update this branch will/dreamverse-monorepo to be directly against origin/main for PR purposes instead of ltx_sr_port (don't open new PRs)".
**Question:** `will/dreamverse-monorepo` was forked from `will/ltx2_sr_port` (PR #1288's head). Should it stay stacked on top of `will/ltx2_sr_port`, or be rebased so its PR base becomes `origin/main` directly?
**Decision:** Rebase `will/dreamverse-monorepo` onto `origin/main` directly. This makes the branch openable as a single PR against main containing the entire 66-commit chain (40 from the legacy `will/ltx2_sr_port` work + 26 dreamverse-monorepo-specific commits including D-19/D-20/D-21/D-22).
**Pre-rebase state:**
```
origin/main (c17d33bf) ← 1 new commit since 2aaeee2a (PR #1253 SSIM cosine)
│
└─ 2aaeee2a (PR #1286 squash merge; old shared base)
├─ 40 commits → will/ltx2_sr_port @ fbd823df (PR #1288)
└─ 40 + 26 = 66 commits → will/dreamverse-monorepo @ 2ee839a3
```
**Post-rebase state:**
```
origin/main (c17d33bf)
│
└─ 66 commits → will/dreamverse-monorepo @ 83829c5e (NEW SHAs, same content)
origin/main (c17d33bf) -- still in main lineage
│
└─ 2aaeee2a
└─ 40 commits → will/ltx2_sr_port @ fbd823df (PR #1288, UNCHANGED)
```
**Operation:**
1. **Backup**: `git branch will/dreamverse-monorepo-pre-main-rebase-backup-20260506 will/dreamverse-monorepo` (local-only safety net pinning `2ee839a3`).
2. **Rebase**: `git rebase origin/main` while on `will/dreamverse-monorepo`. All 66 commits replayed conflict-free in ~30s. The single new origin/main commit `c17d33bf` (#1253 SSIM cosine regression) touches only `fastvideo/tests/ssim/*`, `fastvideo/tests/modal/ssim_test.py`, `.agents/skills/seed-ssim-references/SKILL.md`, `fastvideo/pipelines/basic/stable_audio/stages/decoding.py`, and a 56-line block in `fastvideo/entrypoints/video_generator.py`. None of our 66 commits touch the SSIM files; the `video_generator.py` overlap auto-merged because our D-20 changes (`_BATCH_EXTRA_PASSTHROUGH_KEYS`, result-dict surface) and #1253's changes are in different parts of the file.
3. **Verification (file-level, before push)**:
- Commit count: 66 (PASS)
- Co-author trailer count: 231 across 66 commits — within historical norm (some legacy `will/ltx2_sr_port` commits had partial trailers per [`authors.md`](authors.md) "Known gaps").
- Tree-diff vs backup: only the 845-line forward delta from `c17d33bf` (expected). All D-20/D-21/D-22-introduced content (`_BATCH_EXTRA_PASSTHROUGH_KEYS` x2, "unknown field" x1, `av_chunk_interval_ms` x4, `av_chunk_publish_ms` x2, `enable-nvenc` x2, `NVENC_OVERRIDE` x4) survives intact.
- `pre-commit run --files` on the 7 most-edited files: PASS (yapf/ruff/codespell/mypy/spaces).
- Runtime pytest hung mid-import on this shared dev box (pre-existing CUDA/torch init issue affecting other commands too) — NOT a rebase regression. The same `fastvideo/tests/api/` suite passed 185/185 pre-rebase.
4. **Force-push**: `git push --force-with-lease=will/dreamverse-monorepo:2ee839a3... origin will/dreamverse-monorepo`. Updated `2ee839a3...83829c5e (forced update)`.
**Implications:**
- The branch can now be opened as a PR against `origin/main` containing all the LTX-2 SR port + NVFP4 + Dreamverse monorepo migration + audio kwarg fix + warmup + NVENC build + benchmarks + integration memory dir, in a single review unit.
- PR #1288 on `will/ltx2_sr_port` is untouched and continues to track its own subset of the work. If PR #1288 lands first, the duplicated commits on `will/dreamverse-monorepo` will be reconciled at the next rebase by `git rebase origin/main` dropping commits whose content is now in main (same mechanism as the post-#1286 rebase recorded in [state.md](state.md) "Post-#1286 rebase summary").
- Local backup `will/dreamverse-monorepo-pre-main-rebase-backup-20260506 @ 2ee839a3` keeps the old chain available; recommend deleting after the new branch state is verified by running on a non-stuck dev box (or after the PR merges).
**Watch-outs:**
- Anyone with a checkout of the OLD `will/dreamverse-monorepo` (pre-rebase) needs to `git fetch origin will/dreamverse-monorepo --force` and discard local commits, OR rebase their local commits onto the new `83829c5e`. None observed in the runbook's worktree-sharing model.
- The trailer-count 231 (vs expected 264 for full coverage) is NOT introduced by this rebase — it's the historical gap from [`authors.md`](authors.md). Verifiable by counting trailers on the backup branch (same 231).
### D-21: AV chunk stutter root cause — software libx264 encoding consumes ~22% of segment wall time; opt-in NVENC fix
**Status:** 🟡 Deferred — fix landed (NVENC build + `--nvenc` flag), default unchanged so no behavior regression for operators who haven't rebuilt ffmpeg yet. Switch default to `h264_nvenc` once benchmark numbers are captured + a release notes entry is published.
**Source:** Live debug session 2026-05-06 prompted by user observation: "stuttering still between chunks" after D-20 + warmup r3 fix. 3 explore agents fanned out (NVFP4 config, torch.compile config, AV chunk pacing).
**Question:** With D-20 (audio kwarg routing fix) + r3 warmup pass landed and warmup_success=true, why is the live deploy still stuttering between chunks? Is the migration missing some quantization or compile coverage that the FastVideo-internal reference has?
**Investigation (3 parallel explore agents):**
1. **NVFP4 config diff** (`bg_0629273d`): No divergence. apps/dreamverse, original Dreamverse, and FastVideo-internal ALL use `layer_profile="refine"` with the same 48-block `fp4_layers` superset, e2m1 mantissa, fp32 alpha, layout_128x4 scale, group size 16. Stage gating (`base` vs `refine`) and `transformer_refine_quant` carrier match. The only naming difference is the public surface rename `FP4Config` → `NVFP4Config`.
2. **torch.compile config diff** (`bg_97081cc5`): No divergence. ALL THREE repos use `inductor / fullgraph=True / max-autotune-no-cudagraphs / dynamic=False` and compile only the transformer + text_encoder. **VAE / audio_vae / vocoder are eager bf16 in all three** — so the migrated repo is not missing a compile pass that the original had. Only behavioral difference: FastVideo-internal calls `target.eval()` on its audio encoder; the migrated worker does not (separate D-22 follow-up).
3. **AV chunk pacing** (`bg_574e831e`): `stream_fmp4` has only one timer (`av_encode_stream_ms`) — no per-chunk instrumentation. Between `worker.generate_step()` returning and `stream_fmp4()` returning there is no significant work other than ffmpeg's own encode + chunk emission. The `main_user_step − worker_e2e ≈ 1300ms` gap is therefore mostly ffmpeg + controller relay loop, NOT IPC. Per-chunk emission cadence is driven by stdout read size (1 MiB), not fragment boundary, so non-uniform pacing is plausible.
**Root cause synthesis:**
| Phase | Wall-time | Compute | Compiled? | NVFP4? |
|---|---|---|---|---|
| transformer denoise (gen) | ~4500 ms | GPU | YES | YES |
| save_conditioning | ~100 ms | GPU/CPU | n/a | n/a |
| **stream_fmp4 (libx264 software encode)** | **~1300 ms** | **CPU** | **NO** | **NO** |
| TOTAL wall | ~6000 ms | producing 4.67 s playable | | |
Realtime ratio = 4.67 s playable / 6.0 s wall = **0.78x**. FE buffer drains 1.3 s per segment until empty → inter-segment stutter. **The cause is not a migration regression** (FastVideo-internal exhibits the same architectural limit); it is the choice to use `libx264 ultrafast` software encoding on the segment-streaming hot path on machines that have idle hardware encoders.
**Build-time finding:** `apps/dreamverse/scripts/install_native_ffmpeg.sh` was building ffmpeg with only `--enable-libx264 --enable-lto`. No `--enable-nvenc`, no nv-codec-headers prereq. So even setting `FASTVIDEO_VIDEO_CODEC=h264_nvenc` would have failed at runtime with "encoder not found".
**Resolution (this round, default-preserving):**
1. **`apps/dreamverse/scripts/install_native_ffmpeg.sh`** now:
- Clones `nv-codec-headers` (NVENC/NVDEC API headers; no CUDA libs) into `$INSTALL_PREFIX` so `ffnvcodec.pc` is on `pkg-config`'s search path.
- Configures ffmpeg with `--enable-cuda --enable-nvenc --enable-cuvid --enable-nvdec` plus `--extra-cflags=-I$CUDA_PREFIX/include` and `--extra-ldflags=-L$CUDA_PREFIX/lib64`.
- Drops `--enable-cuda-nvcc` and `--enable-libnpp` so the build does NOT require `--enable-nonfree` (we only need hardware encode/decode, not GPU-side filters).
- New env knobs: `ENABLE_NVENC=0|1` (default 1), `CUDA_PREFIX` (default `/usr/local/cuda`), `NV_CODEC_REF`.
- Sanity check verifies the resulting binary exposes `h264_nvenc` / `hevc_nvenc` encoders before exiting 0.
- **Emits `FASTVIDEO_VIDEO_CODEC=libx264` in the env file by default** so existing deploys are runtime-unchanged.
2. **`.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh`** now:
- Accepts `--nvenc` / `--no-nvenc` flag (defaults to off, matches existing behavior). When `--nvenc`, exports `FASTVIDEO_VIDEO_CODEC=h264_nvenc` into the backend setsid block.
- Validates the binary actually has `h264_nvenc` encoder via `ffmpeg -encoders | grep h264_nvenc`. Fails fast if `--nvenc` is requested against a libx264-only ffmpeg with a clear hint to rerun the install script.
- `DREAMVERSE_NVENC` env var as the env-var counterpart (flag overrides).
- Banner now prints `nvenc=true|false` alongside `warmup` and `torch_compile`.
3. **`apps/dreamverse/server/benchmarks/benchmark_av_streaming.py`** is the regression test:
- Drives `stream_fmp4` directly with synthetic 121-frame 1920x1088 video + 5-second 24kHz stereo audio (production shape).
- Sweeps `libx264` and `h264_nvenc` (auto-skips encoders not present in the active ffmpeg).
- Reports per-codec: wall_ms_min/median/p95/max, bytes_median, chunks_median, and **realtime_ratio_median** (playable / wall, must be ≥ 1.0 to avoid buffer drain).
- Exit code 1 if any codec produces realtime_ratio < 1.0; useful as a CI gate.
**Empirical findings (this dev host, B200):**
- `libx264 ultrafast` benchmark on production-shape input (121 frames, 1920x1088, 5s 24kHz stereo audio): wall_median=611ms, wall_p95=644ms, 8.25x realtime ratio in isolation. Captured via `apps/dreamverse/server/benchmarks/benchmark_av_streaming.py --runs 5 --codecs libx264`.
- **B200 has NO NVENC silicon.** Direct ffmpeg probe (`-c:v h264_nvenc` against a 64x64 0.2s color frame) fails with `OpenEncodeSessionEx failed: unsupported device (2): No capable devices found`. This is a hardware omission — datacenter Blackwell (B200) and some H100 SKUs prioritize compute density and ship without NVENC. Only consumer Blackwell (RTX 50-series) and select datacenter SKUs (T4, A10, A100 PCIe) have NVENC.
- The deploy script's `--nvenc` path now does TWO checks: (a) `h264_nvenc` is in the encoder list (build-time); (b) a 64x64 ffmpeg probe actually succeeds (runtime-time). On a B200, the runtime probe fails fast with a clear "no NVENC silicon on this host" error pointing at SKU-level alternatives.
- **Stutter on B200 is NOT primarily ffmpeg encoding.** The benchmark shows libx264 ultrafast at 611ms, but production logs show a 1300ms gap between `worker_e2e` and `main_user_step`. The other ~700ms is IPC + controller WebSocket relay (not measured per-chunk in current `stream_fmp4`).
**Open / deferred:**
- ~~D-22~~ ✅ **Resolved 2026-05-06** in `bade2c0a` — `stream_fmp4` is now fully instrumented with per-phase + per-chunk timings; gpu_pool prints a per-segment summary line. The controller-side WS-send slice is tracked separately as D-22-CTL.
- D-22-CTL: per-chunk timing in the controller's AV relay loop in `session/controller.py` (between media event arrival and `ws_send_bytes`). Worker side is captured by D-22; the controller side is the still-unmeasured remainder of the 700ms IPC+relay gap.
- D-23: B200 + RTX 5090 split — for B200-class deploys without NVENC, the stutter fix needs a different approach (pipeline gen N+1 with encode N, or larger initial FE buffer pre-fill). For RTX 5090 / T4 / A10 deploys, `--nvenc` is the answer once tested. Make the deploy default conditional: probe NVENC at boot, default `--nvenc` if available.
- D-24: re-evaluate flipping the deploy default from `libx264` to `h264_nvenc` after a real-NVENC host (RTX 5090 or H100 PCIe) benchmark is run. Tradeoff: NVENC has ~5-10% lower compression at same quality but is hardware-accelerated. For real-time streaming the latency win dominates IF the host has NVENC.
- D-25: pipeline benchmark via Python SDK (`apps/dreamverse/server/benchmarks/benchmark_pipeline.py`, landed in `f98811e0`) — captures per-stage timings via `FASTVIDEO_STAGE_LOGGING=1`. Cold-run baseline on B200 NVFP4 (no compile, no warmup): 6.99s for 5.04s playable = 0.72x realtime. Per-stage profile dominated by LTX2RefineLoRAStage (40% on cold, includes one-time LoRA conversion + adapter load), LTX2DenoisingStage (20% — base DiT 5-step denoise), LTX2TextEncodingStage (11%), DecodingStage (4%), audio_decoding/upsample/init each <1%. Future work: rerun under `compile_warm` scenario for steady-state numbers.
- The `target.eval()` call on the audio encoder that FastVideo-internal does and the migrated Dreamverse does not — surface as a separate decision once observed empirically (probably a no-op in inference path but worth the symmetry).
**Cross-references:**
- [D-20](decisions-log.md#d-20) (audio kwarg routing fix) made segment 2 generate cleanly. D-21 makes the segment-to-segment chain stutter-free in real-time.
- [`apps/dreamverse/server/av_streaming.py::stream_fmp4`](file:///home/william5lin/FastVideo/apps/dreamverse/server/av_streaming.py) is the hot path; the cmd builder there respects `FASTVIDEO_VIDEO_CODEC` and branches on `*_nvenc` codecs to use NVENC presets (`p1`/`p2`/...) and `-rc constqp -qp 28` instead of libx264 presets (`ultrafast`/...).
- The benchmark cross-references D-21 in its module docstring so a future contributor reading the script alone has the context.
### D-20: Segment-2 BrokenPipe root cause — two-direction silent drop of LTX-2 audio kwargs in public `VideoGenerator`
**Status:** ✅ Resolved 2026-05-05. 4 commits on `will/dreamverse-monorepo` (`1b686f4e`..`5eaf0a13`); pushed to origin. End-to-end verified on GPU4 (`/proc/$BE_PID/environ` + `Cached audio latents shape=(1, 8, 126, 16) for segment 2` log line + `Segment 2: relayed av chunks=22, bytes=3.8MB`).
**Source:** Live debugging session triggered by recurring `RuntimeError: ffmpeg frame writer failed: [Errno 32] Broken pipe` on segment 2 in `/tmp/opencode/dreamverse-deploy/backend-gpu4.log`.
**Question:** Why does segment 2 of every Dreamverse session crash with a writer-side EPIPE, while segment 1 streams fine?
**Symptom chain decoded:**
1. ffmpeg writer thread in `apps/dreamverse/server/av_streaming.py:307-317` raises `BrokenPipeError` mid-frame.
2. ffmpeg's exit code is `0` — the `if rc != 0` branch of `stream_fmp4` never fires; only `if writer_error[0] is not None` does.
3. `rc=0 + BrokenPipeError` means ffmpeg exited cleanly *before* the writer finished pushing all frames → ffmpeg closed stdin early due to `-shortest` + audio shorter than video.
4. Manual repro confirmed: with audio 71240 samples (~2.97s) and video 112 frames (~4.67s), ffmpeg `-shortest` + 1MB pipe + 1920x1088x3 frames produces exactly this signature (`rc=0 out_bytes=751028 writer_error=BrokenPipeError(32)`).
5. The 1.7s audio undershoot is suspicious because Dreamverse's `apply_audio()` in `apps/dreamverse/server/video_generation.py:132-186` explicitly extends `audio_num_frames = NUM_FRAMES + audio_extra` (=`121 + 40 = 161`) for continuation segments. Audio should be 6.71s, not 5.01s. The kwarg was being silently dropped.
**Root cause (TWO directions):**
| Direction | Where | What was missing | What it broke |
|---|---|---|---|
| **Inbound** (kwargs → `batch.extra`) | `fastvideo/entrypoints/video_generator.py::_generate_video_impl` | The 5-key extraction block (`ltx2_audio_latents`, `ltx2_audio_clean_latent`, `ltx2_audio_denoise_mask`, `audio_num_frames`, `video_position_offset_sec`) that FastVideo-internal has at lines 168-183. Without it, `apply_audio`'s kwargs landed in `sampling_param.update(kwargs)` which `logger.error`'d "%s has no field %s" and dropped them. | `audio_num_frames` never reached `batch.extra` → `ltx2_denoising.py:325` fell back to `batch.num_frames=121` → audio generated for 5.01s instead of 6.71s. After `head_trim_audio_frames=49` removed the leading 2.04s, only 2.97s of audio remained for 4.67s of video. |
| **Outbound** (`batch.extra` → result dict) | `fastvideo/entrypoints/video_generator.py::_generate_single_video` | `"ltx2_audio_latents": output_batch.extra.get("ltx2_audio_latents")` in the result dict. Internal exposes it at line 556; public didn't expose it at all. | `Dreamverse._derive_next_audio_latents()` always saw `None` → `self.continuation.audio_latents` was never set → segment 2's `apply_audio` short-circuited (`if not (… and self.audio_latents is not None): return`) → no audio continuation, no `audio_num_frames` extension, no `ltx2_audio_clean_latent` carry-over. |
The two directions hid each other: even after the inbound block was ported, segment 2 still broke until the outbound surface was added. Both must be present for audio continuation to round-trip.
**Resolution — 4 commits on `will/dreamverse-monorepo`:**
| SHA | Subject |
|---|---|
| `1b686f4e` | `[fix] dreamverse: pin ffmpeg native build toolchain by uname -m` |
| `265ce1a6` | `[fix] api: route LTX-2 audio kwargs through batch.extra; strict update` |
| `dab9499c` | `[feat] dreamverse-deploy: native ffmpeg + compile-off defaults` |
| `5eaf0a13` | `[feat] dreamverse-deploy: --warmup / --torch-compile CLI flags` |
`265ce1a6` is the substantive public-API fix: ports `_BATCH_EXTRA_PASSTHROUGH_KEYS` extraction in `_generate_video_impl`, ports `_extra_overrides` consumption in `_generate_single_video` (writes into `batch.extra`), surfaces `ltx2_audio_latents` in the result dict, and converts `SamplingParam.update()` from `logger.error`-and-drop to `raise ValueError` on unknown keys so the next contributor who adds an unrecognized kwarg hits a loud failure pointing at `_BATCH_EXTRA_PASSTHROUGH_KEYS` instead of debugging a broken-pipe-shaped symptom hours later. Adds `fastvideo/tests/api/test_extra_overrides_routing.py` (7 tests pinning the contract).
**Why the strict-update is non-negotiable:** the `logger.error`-and-continue pattern is the exact mechanism that hid this bug for the entire Dreamverse-monorepo migration window. It silently traded loud-failure-now for silent-corruption-later. Strict raise + a clear "route via `_BATCH_EXTRA_PASSTHROUGH_KEYS`" hint converts the next regression of this shape from "broken pipe in production" to "ValueError at first call".
**Cross-references:**
- [D-11](decisions-log.md#d-11) (Apr 26) noted "ffmpeg fragment write Broken pipe" but framed it as cosmetic (client-disconnect race). Today's was a different code path: server-side EPIPE caused by a server-internal A/V duration mismatch. Both are now resolved; D-11's fix domain (swallow on intentional disconnect) remains unchanged but lower priority.
- The `[fix] api: route LTX-2 audio kwargs ...` commit (`265ce1a6`) is on `will/dreamverse-monorepo` only. Cherry-picking onto `will/ltx2_sr_port` so it lands as part of PR #1288's mega-PR is a deferred follow-up — see open-threads.md "Cherry-pick API audio routing fix to PR #1288".
**Side effects of the fix:**
- `[feat] dreamverse-deploy: native ffmpeg + compile-off defaults` (`dab9499c`) wires the native LTO+libx264+native-arch ffmpeg the team playbook mandates into every deploy via `FASTVIDEO_FFMPEG_BIN=$HOME/opt/ffmpeg-native/bin/ffmpeg` and disables `torch.compile` by default (`ENABLE_TORCH_COMPILE=0`) so segment-1 cold start drops from ~3-4min (max-autotune) to ~45s (pure inference) — required for any iterative debugging cycle to fit inside the 300s session timeout.
- `[fix] dreamverse: pin ffmpeg native build toolchain by uname -m` (`1b686f4e`) makes `apps/dreamverse/scripts/install_native_ffmpeg.sh` immune to conda envs that activate both `gcc_linux-64` and `gcc_linux-aarch64` (the aarch64 activation script sorts later and wins, so the inherited `CC=aarch64-conda-linux-gnu-cc` defeated the script's `: "${CC:=...}"` deferred default and tripped x264's compiler probe with "unknown value 'native' for '-march'").
- `[feat] dreamverse-deploy: --warmup / --torch-compile CLI flags` (`5eaf0a13`) adds `--warmup`/`--no-warmup` and `--torch-compile`/`--no-torch-compile` flags that override the env-var defaults and accept any position relative to the positional GPU/port args (verified across 13 parser permutations).
**Watch-outs for downstream contributors:**
- Adding a new pipeline-specific kwarg now requires either making it a `SamplingParam` field OR adding it to `_BATCH_EXTRA_PASSTHROUGH_KEYS` in `fastvideo/entrypoints/video_generator.py`. Strict `update()` will raise `ValueError` on unknown kwargs that aren't routed through one of those paths.
- The `[fix] api:` commit (`265ce1a6`) is BACKWARD-INCOMPATIBLE for any caller that was relying on `SamplingParam.update()` to silently swallow unknown keys. If pre-existing CI breaks elsewhere on this surface, the fix is to add the legitimate kwarg to `SamplingParam` or `_BATCH_EXTRA_PASSTHROUGH_KEYS` (not to revert the strict mode).
### D-19: D-18 executed — Dreamverse migration on `will/dreamverse-monorepo`, generic-merged files carried product-local
**Status:** ✅ Resolved 2026-05-05. 5 commits on top of `will/ltx2_sr_port` HEAD `fbd823df`. Branch pushed to origin at `c1fe5d4c`. e2e definitively passes.
**Source:** Execution of [integration-plan.md](integration-plan.md) (D-18 plan), with one significant deviation surfaced by Oracle review.
**Commits:**
| SHA | Message |
|---|---|
| `08828d96` | `[feat] dreamverse-monorepo: Phase 1 — skeleton + tooling` |
| `f3a863ba` | `[feat] dreamverse-monorepo: Phase 2 — backend move + import rewires` |
| `876f7eb3` | `[feat] dreamverse-monorepo: Phase 3 — frontend + public assets` |
| `1d47ede6` | `[fix] dreamverse-monorepo: carry generic-merged files product-local` |
| `c1fe5d4c` | `[fix] dreamverse-monorepo: entrypoint + audio re-encode + verification` |
Final stats: 164 files changed, +53294/-2 LOC, 31725 files under `apps/dreamverse/`.
**Question:** [integration-plan.md](integration-plan.md) Phase 2 instructed "DELETE generic-merged files (`gpu_pool.py`, `av_streaming.py`, `worker_ipc.py`, `mock_server.py`, `session_init_image.py`, `session_logger.py`) from the Dreamverse copy and rewire imports to public `fastvideo.entrypoints.streaming.*` substitutes". The first execution attempt followed this instruction. Was that the right call?
**Decision (forced by Oracle finding):** No. The "delete + import substitute" strategy assumed the public modules were API-compatible drop-ins for the Dreamverse-product modules. They are NOT. Public `fastvideo.entrypoints.streaming.GpuPool` is an abstract base class with `acquire/run/release/shutdown/health` semantics designed for a future Phase 4 streaming-runtime API. Dreamverse's product `GPUPool(gpu_ids).initialize()` has totally different shape (per-GPU subprocess workers with `slot.user_step / register_stream_queue / acquire(client_id, websocket) -> (gpu_id, slot)` semantics). Substituting one for the other made `apps/dreamverse/server/main.py:63` fail at boot with `TypeError: GpuPool() takes no arguments`.
**Fix:** Carry all 7 generic-merged files (gpu_pool, av_streaming, worker_ipc, mock_server, session_init_image, session_logger, server_entry) into `apps/dreamverse/server/` as PRODUCT-LOCAL — same status as the rest of the Dreamverse server tree. Imports stay flat (`from gpu_pool import GPUPool`, etc.), matching Dreamverse's existing sys-path-injection convention. Public `fastvideo/entrypoints/streaming/*` reverts cleanly to its `fbd823df` state — `git diff fbd823df..c1fe5d4c -- fastvideo/entrypoints/streaming/` is empty.
The "import public substitutes" promise becomes a future Phase 4 task: actually harmonize the APIs so Dreamverse can drop its product-local copies. Not in scope for this migration.
**Why the e2e initially gave a false positive:**
The first run of all 8 Playwright tests passed (5.1s) — Oracle caught that this was misleading. The `dreamverse-server` console script in `/home/william5lin/miniconda3/envs/fv-main/bin/dreamverse-server` was installed by a prior `pip install -e /home/william5lin/Dreamverse`, so `from server_entry import cli` resolved to `/home/william5lin/Dreamverse/server/...` (the canonical install), not `apps/dreamverse/server/...` (the migrated tree). The migrated code was never actually exercised.
**Fix to entrypoint resolution** (in `c1fe5d4c`): wrapper scripts at `apps/dreamverse/scripts/dreamverse-server` (and `dreamverse-mock-server`) explicitly do:
```bash
#!/usr/bin/env bash
REPO_ROOT="$(cd "$(dirname "$0")/../../.." && pwd)"
cd "${REPO_ROOT}/apps/dreamverse/server"
exec "${REPO_ROOT}/.venv/bin/python" main.py "$@"
```
This guarantees the migrated tree is what runs. Verified end-to-end via `/proc/$PID/cwd = /home/william5lin/FastVideo/apps/dreamverse/server` during the second e2e run. All 8 Playwright tests pass against this confirmed-migrated backend.
**Phase 0 environment prereqs (validated 2026-05-05):**
- `flashinfer-python` in FastVideo `.venv` — required for NVFP4 path. Without it, model load fails with `ImportError: NVFP4 quantization requires flashinfer`.
- `cerebras-cloud-sdk` and `openai` in `.venv` — required by the migrated prompt enhancer.
- For B200 / sm_100a + gcc-15 conda toolchain: nvcc rejects host compiler. Workaround:
```bash
CUDAHOSTCXX=/usr/bin/g++-13
NVCC_PREPEND_FLAGS="-ccbin /usr/bin/gcc-13 -allow-unsupported-compiler"
```
Without these, flashinfer JIT compilation fails with `error: #error -- unsupported GNU version! gcc versions later than 14 are not supported!`.
These are **operator-side prerequisites**, not migration code defects. Documented in `apps/dreamverse/README.md` and `docs/contributing/dreamverse-development.md`.
**E2E evidence (live, post-fix-up):**
```
PID 179112 cwd: /home/william5lin/FastVideo/apps/dreamverse/server
/healthz → 200 {"status":"ok","service":"ltx2-streaming-backend",...}
/readyz → 200 {"status":"ready","ready_gpu_workers":1,"total_gpus":1,...}
GPU4 mem: 50.9 GiB (NVFP4 model loaded)
Playwright (8/8 PASS in 5.1s):
✓ backend-health/healthz returns ok via the next.js rewrite (32ms)
✓ backend-health/readyz reports gpu pool state (11ms)
✓ backend-health/status endpoint exposes gpu pool snapshot (12ms)
✓ backend-health/prompt-system-config exposes operator-tunable prompts (15ms)
✓ backend-health/curated presets endpoint serves a non-empty list (devtools only) (19ms)
✓ frontend-shell/main page loads and exposes the FastVideo brand chip (1.4s)
✓ frontend-shell/composer hydrates with curated preset cards (1.3s)
✓ preset-prompt-generation/generates the first segment from a curated preset prompt (1.8s)
```
**Implications:**
- The integration-plan.md is **partially superseded** by D-19 outcome:
- Phase 2's "DELETE generic-merged" prescription is invalid; replace with "carry product-local".
- Phase 0 prereqs need the gcc-13 / `NVCC_PREPEND_FLAGS` workaround documented for B200 hosts.
- Phase 4 (in a future PR) is now responsible for actually harmonizing public-vs-Dreamverse pool APIs so the carried product-local modules can be deleted.
- `dreamverse-mock-server` script under `apps/dreamverse/scripts/` is the only canonical launcher. The conda env's legacy `/home/william5lin/miniconda3/envs/fv-main/bin/dreamverse-server` should NOT be used (it points at canonical Dreamverse repo).
- The audio re-encode handling in `apps/dreamverse/server/video_generation.py` was carried per `c1fe5d4c`; verify the exact shape (restored from source vs deferred) when reading the commit.
- Once this branch lands as a PR + merges, the Dreamverse repo can be archived per [integration-plan.md](integration-plan.md) Phase 7.
**Open follow-ups:**
- Open a PR for `will/dreamverse-monorepo` (target main, base on `will/ltx2_sr_port` until #1288 merges).
- Re-Oracle the post-fix-up state to confirm the prior FAIL is now PASS.
- Eventually move audio re-encode into a public module (Phase 4) so the carried product file can be slimmed.
- Eventually do real Phase 4 API harmonization between Dreamverse pool and public `fastvideo.entrypoints.streaming.GpuPool` so the 7 carried product files can be deleted.
### D-18: Option B+ — Dreamverse becomes `apps/dreamverse/` subfolder under FastVideo
**Status:** ✅ Resolved 2026-05-05. [integration-plan.md](integration-plan.md) is the executable migration plan; [integration-review.md](integration-review.md) is deprecated but kept for drift audit + OSS precedents.
**Source:** User decision after reviewing [integration-review.md](integration-review.md)'s Option D recommendation.
**Question:** [integration-review.md](integration-review.md) recommended **Option D** — Dreamverse stays a separate repo, generic backend (streaming runtime, GPU pool, prompt enhancer, router) merges into `fastvideo.entrypoints.streaming.*`. The user reviewed this and chose a different shape: keep the generic-backend principle from Option D but ALSO move the Dreamverse FE + product server into FastVideo as a subfolder (`apps/dreamverse/`). Combination is "Option B+" (Option B layout with Option D's backend principle).
**Decision:** Option B+. Concrete shape:
- **One repo**: `hao-ai-lab/FastVideo`. Dreamverse repo gets archived after migration completes.
- **Python ML library** stays at root: `fastvideo/`, `fastvideo-kernel/`.
- **Generic backend** stays at `fastvideo.entrypoints.streaming.*` (already there per #1257/#1258/#1284/#1286/#1288).
- **Dreamverse product** moves into `apps/dreamverse/{server,web,prompts,serve_configs,scripts}/`.
- **Tooling**: uv workspace for Python (`[tool.uv.workspace] members = ["apps/dreamverse/server"]`), standalone npm for the FE (no root `package.json`), split CI workflows with path-filter triggers.
**Rationale:**
- Drops the cross-repo coordination overhead identified in the post-#1286 rebase cycle (D-17 handled by consolidating into mega-PR; D-18 prevents the next round of cross-repo coordination from happening).
- Keeps the architectural separation Option D recommended (FastVideo owns reusable runtime; product owns product). The boundary is now `apps/dreamverse/` directory rather than two repos.
- Single repo means atomic cross-cutting refactors (e.g. GpuPool API change + Dreamverse adoption) ship as one PR.
- OSS precedents support the shape (chainlit uv-workspace + frontend package manager; open-webui Python + Svelte with paths-ignore CI). The librarian explicitly noted no precedent for "Python ML library + Next.js product merged into library namespace" — but this isn't that pattern. Dreamverse goes into a sibling directory, NOT into `fastvideo.entrypoints.dreamverse.*`. Library namespace stays clean.
**Why not Option D (separate repos):**
- Each upstream merge into FastVideo invalidates Dreamverse's lockfile/imports; the post-#1286 rebase showed this requires coordination overhead that scales with feature velocity.
- Cross-repo contract tests catch shape drift but not behavior drift.
- Two repos means two `AGENTS.md`, two CI configs, two release stories, two Dependabot dashboards.
**Why not Option C (full merge into `fastvideo.entrypoints.dreamverse.*`):**
- Forces FastVideo to ship Tailwind config + curated preset JSON + Next.js build artifacts.
- Locks Dreamverse product cadence to FastVideo PyPI releases.
- Librarian: "no 1:1 precedent for Python ML library + Next.js product merged into library namespace" — argues against this.
**Why not Option B (subfolder, but generic backend folded into `apps/dreamverse/server/`):**
- Other consumers (Dynamo, future streaming clients) need the backend without the Dreamverse product. Folding the backend under `apps/dreamverse/server/` would force Dynamo to either depend on `apps/` paths (ugly) or carry a fork.
**Implications:**
- [integration-review.md](integration-review.md) is **deprecated** (banner header + reading-guide demotion). Kept in tree for drift audit + OSS precedent reference.
- [integration-plan.md](integration-plan.md) is the **canonical executable plan** with 7 phases (Phase 0: land #1288; Phase 1: skeleton + tooling; Phase 2: backend move; Phase 3: FE move; Phase 4: promote generic-pending; Phase 5: prompt enhancer fork retirement; Phase 6: CI/release cutover; Phase 7: archive Dreamverse repo).
- Dreamverse repo will be **archived** at end of Phase 7 — not before.
- Dreamverse history does NOT migrate cross-repo via `git mv` (technical limitation); original history stays in archived Dreamverse repo, and Phase 2 PR body records the source SHA(s).
- New top-level `apps/` directory created — must be excluded from FastVideo PyPI wheel via `[tool.setuptools.packages.find] exclude = ["apps*", ...]`.
- Drift items from [integration-review.md](integration-review.md) get folded into specific phases of [integration-plan.md](integration-plan.md) (e.g. health routes → Phase 4, DR-1 → Phase 5).
**Open questions deferred to phase planning:**
- DR-2 (`cerebras_ifm`): public Literal vs Dreamverse-side custom provider — decide before Phase 5.
- VPO (`video_position_offset_sec` semantics): persistent vs per-segment — decide in Phase 4.
- Cross-repo history: fresh import vs `git subtree` import — decide before Phase 2.
- CORS / write-endpoint security policy: dev-only vs auth vs firewall — decide before Phase 6.
### D-17: Abandon 6-PR split — land everything as single mega-PR #1288
**Status:** ✅ Resolved 2026-05-05. PR #1287 closed; PR #1288 opened on `will/ltx2_sr_port` covering the full chain.
**Source:** User decision after observing the post-#1286 rebase + re-slice cycle.
**Question:** The original plan ([STACK.md](../../../STACK.md), [pr-roadmap.md](pr-roadmap.md)) called for the remaining `will/ltx2_sr_port` content (after PRs 7.5/7.6/7.7/7.8/7.9 landed) to ship as 6 stacked PRs: 7.10 (#1287, generate_async), 8 (server contract docs), LTX-2 SR runtime, NVFP4, post-fixes, agents-cleanup. PR #1287 was opened on 2026-05-05 as the first slice. Should the remaining 5 slices be opened sequentially as planned, or should everything be consolidated into one PR?
**Decision:** Consolidate. Close #1287; open one mega-PR (#1288) on `will/ltx2_sr_port` covering all 34 commits / 71 files / +13,074 LOC at once.
**Rationale:**
- The post-#1286 rebase + re-slice cycle exposed real overhead: backup branch, interactive rebase with manual `drop` directives, force-push, re-slice 6 bookmarks, push next slice as new remote, open new PR, update memory dir. Repeating that 6 more times for the remaining slices accumulates substantial review-coordination overhead with diminishing structural benefit.
- The 6 layers are not independent in the way that landed PRs 7.5-7.9 were. PR 7.10 (`generate_async`) is the only API-shape change; PR 8 is docs+tests on top; LTX-2 SR / NVFP4 / post-fixes / agents-cleanup are feature/fix/docs work that doesn't shape the public API. Reviewing them as one ordered diff is at least as easy as reviewing 6 stacked PRs whose dependencies must be tracked manually.
- Single PR keeps CI / merge queue simpler and avoids the 6-PR cascade where every upstream merge invalidates the chain below it.
**Implications:**
- [STACK.md](../../../STACK.md) (top-level, 10-PR split tracker) is **deprecated**. Kept in tree as a historical artifact with the merged half (PRs 1-4 of the 10) accurate. Safe to delete in a follow-up.
- [authors.md](authors.md), [co-authors.md](co-authors.md) — co-author roster is unchanged; trailers still apply per-commit on every commit in the consolidated PR.
- [runbook.md](runbook.md) — "After a PR merges (re-slice protocol)" section replaced by a simpler "After PR #1288 merges" section.
- Local split bookmarks (`will/api_7.10`, `will/api_8`, `will/ltx2_sr_runtime`, `will/ltx2_nvfp4`, `will/ltx2_post_fixes`, `will/agents_cleanup`) are no longer maintained; safe to delete locally.
- `origin/will/api_7.10` — pushed during the #1287 cycle; can be deleted on origin once #1287 close-cleanup completes.
**Watch outs:**
- The PR is large (71 files, +13,074 LOC). Reviewers will need commit-by-commit review; the PR body structures the layers in commit order to make this tractable.
- If #1288 becomes too large to merge cleanly later (e.g. main moves significantly underneath it), the fallback is to re-split — but the current expectation is to land it as-is.
### D-12: `GpuPool` layer separation — keep distinct from `VideoGenerator`
**Status:** ✅ Resolved (interim) + 🟡 Deferred long-term shape to PR 7.10.
**Source:** Oracle review on 2026-05-04, post-PR-#1257 merge.
**Question:** Should `fastvideo.entrypoints.streaming.GpuPool` (PR #1257) be
folded into `fastvideo.entrypoints.video_generator.VideoGenerator`, or kept
separate? Three alternatives were evaluated:
| Alt | Approach | Verdict |
|---|---|---|
| A | Status quo — `VideoGenerator` (single inference call) and `GpuPool` (multi-session orchestration) stay separate | ✅ Correct as **interim** |
| B | `VideoGenerator` absorbs the pool's role (`from_pretrained_pool`, `acquire/release/run`) | ❌ **Wrong layer.** Conflates execution with serving scheduler. |
| C | `GpuPool` becomes a thin **session-aware async executor** over PR 7.10's `generate_async` | ✅ Correct **long-term destination** |
**Decision:** Alt A as interim; evolve toward Alt C once PR 7.10 lands
`generate_async`. Do NOT pursue Alt B.
**Rationale:**
- `VideoGenerator` is a library handle — "execute one request, possibly
across ranks via `MultiprocExecutor`/`RayDistributedExecutor`."
- `GpuPool` is serving infrastructure — "schedule N concurrent sessions
across N independent replicas, with sticky session-to-GPU affinity for
cache locality."
- These are different layers driven by different consumers (a Python
script doing `gen.generate(req)` vs. a WebSocket server with sticky
sessions). Folding them muddies both surfaces.
**Key finding — `MultiprocExecutor` and `SubprocessGpuPool` are orthogonal,
not redundant:**
| Layer | Job | Granularity |
|---|---|---|
| `MultiprocExecutor` (`fastvideo/worker/`) | TP/SP shard ONE inference call across N GPU ranks | per-call |
| `streaming_generator.py` (existing real-time path) | Per-frame streaming via `MultiprocExecutor.submit_step`/`get_result` | per-step within one generator |
| `SubprocessGpuPool` (`entrypoints/streaming/`, PR #1257) | Serve N concurrent sessions on N replicas, sticky-bound | per-session |
Both spawn subprocesses because **CUDA contexts demand process boundaries**,
not because they solve the same problem. Sharing low-level lifecycle
utilities (process spawn, queue plumbing, shutdown) is a future refactor;
unifying the abstractions is wrong.
**Sticky binding stays in the pool, NOT in `VideoGenerator`:** sticky
session-to-GPU affinity is a serving policy driven by LTX-2's per-GPU
continuation cache (last-9-decoded-frames + audio-latents). Different
consumers want different policies — stateless OpenAI HTTP wants
per-request leasing; LTX-2 streaming wants sticky affinity; per-frame
real-time streaming wants a continuous queue. Keeping policy in the pool
keeps `VideoGenerator` policy-free.
**Specific risks flagged in PR #1257 (already merged):**
| Risk | Mitigation (when relevant) |
|---|---|
| `GpuPool.run() -> Any` is sync — fine for whole-segment dispatch, blocks on cancellation | Replace with `run_async() -> AsyncIterator[VideoEvent]` in PR 7.10 cycle (`generate_async` makes this trivial) |
| `PoolAssignment.gpu_id: int` assumes one-GPU-per-worker | Don't lock as public API. Future may need `device_ids: list[int]` for topology-aware pooling (one worker = group of GPUs running internal `MultiprocExecutor`) |
| `GpuPool` could be documented as the canonical FastVideo serving API | Mark as **experimental / server-internal** in docstring until PR 7.10 lands. Don't include in user-facing API docs yet |
| Memory: N processes = N model replicas (~10-50 GB each) | Expected for concurrent serving with crash isolation. CUDA IPC weight sharing loses isolation; CPU-shared-memory loading helps host RAM not device. Real scalable path is topology-aware pooling later. |
**Action items (carried into post-7.10 cycle):**
- [ ] Update `GpuPool` ABC docstring to note "API may change post-PR-7.10"
- [ ] Plan to replace `run()` with `run_async() -> AsyncIterator[VideoEvent]` in PR 7.10 cycle
- [ ] Don't promote `gpu_id: int` to public API; revisit shape post-7.10
- [ ] Consider clarifying field naming (e.g. `worker_id` is the stable identifier; `gpu_id` is current-impl detail)
- [ ] When opening 7.10's PR, have it consume `generate_async` from `GpuPool.run_async` end-to-end
**Open thread it touches:** PR 7.10 (`open-threads.md` item D — generate_async)
unblocks Alt C and is the natural place to land the API shape change.
### D-15: Streaming router (PR #1286) — keep in-repo, defer sticky / active-active
**Status:** ✅ Resolved (interim). Pre-merge polishes applied. Three follow-up
items tracked.
**Source:** Oracle review on 2026-05-05, during PR #1286 review cycle.
**Question:** Where should the multi-replica WebSocket router live? Should it
ship at all (vs. delegating to nginx/envoy)? Should sticky session routing
or weighted/round-robin balancing be in the initial PR?
| Alt | Approach | Verdict |
|---|---|---|
| A | Status quo — `fastvideo/entrypoints/streaming/router/`, FastAPI-based, single-primary failover, lazy `httpx`/`websockets` imports | ✅ **Keep** |
| B | Move to separate package `fastvideo-router/` | ❌ **Premature** — adds packaging/release/compat overhead before evidence of independent adoption |
| C | Fold router into the streaming server itself (one app, mode flag) | ❌ Conflates router/generator lifecycles, mode-dependent config, drags inference deps into routing deployments |
| D | Replace with reverse proxy (nginx/envoy/HAProxy) recipes | ❌ Not as the SOLE answer — mature proxies don't naturally emit FastVideo typed `gpu_unavailable` frames or evolve with FastVideo session semantics. Recommend external proxies as a complement at high scale. |
| E | Add sticky session routing now | ❌ **Defer** — implementing correctly depends on where `session_id` is available (URL/header is easy, first JSON frame is invasive). Reconnects are rare today. |
| F | Add weighted / round-robin now | ❌ **Defer** — active-active without sticky routing is worse for LTX-2 continuation locality than active-passive failover |
**Decision:** Alt A — keep current shape. Apply pre-merge polishes; preserve
forward-compat for sticky routing.
**Rationale:**
- Python router is justified as a FastVideo-aware control-plane component,
not a replacement for Envoy/HAProxy. It can emit typed
`gpu_unavailable` frames, evolve with FastVideo session semantics,
and ship local/dev deployment without ceremony.
- The current abstraction is small + testable: `RouterConfig`,
`ReplicaRegistry`, `ReplicaStatus`, `HttpProbe` (Protocol/structural alias).
Adding strategy registries / telemetry interfaces / active-active policies
now would be over-engineering.
- Active-passive (single primary) is the right MVP for LTX-2 streaming —
preserves continuation cache locality (D-12 sticky binding rationale)
better than naive active-active.
- The biggest architectural risk isn't placement; it's accidentally baking
in unstated semantics. Define single-primary behavior + config validation
now so future active-active or sticky routing becomes additive.
**Pre-merge polishes applied (per gemini + Oracle review):**
| # | What | Why |
|---|---|---|
| 1 | `ReplicaRegistry.select()` docstring rewrite | gemini flagged "round-robin via insertion order" claim was misleading — implementation always returns `[0]`. Replaced with explicit "first healthy primary, else first healthy non-primary; this MVP picks first match within tier; round-robin/weighted deferred". |
| 2 | Refactored `run_health_check_loop` to share single `httpx.AsyncClient` across the loop's lifetime via `_build_default_probe()` async context manager | gemini flagged per-probe client instantiation as inefficient. With ~1 probe/second default polling, TCP/TLS handshake overhead is non-trivial; now reuses connection. Tests inject probes directly so the path stays bypassable. |
| 3 | Probe all replicas concurrently per cycle via `asyncio.gather(..., return_exceptions=True)` | gemini flagged sequential probes risk falling behind `health_check_interval_seconds` if replicas time out. Now per-cycle wall time = max(probe latencies), not sum. |
| 4 | `RouterConfig.__post_init__` validation | Oracle recommended: empty replicas, non-positive intervals/timeouts, thresholds < 1, non-`http(s)://` URLs, and >1 primary all `raise ValueError`. Surfaces misconfiguration at config-load instead of confusing runtime failures. |
| 5 | Migrated `@app.on_event("startup"/"shutdown")` to `@contextlib.asynccontextmanager`-based `_lifespan()` | Pre-merge — FastAPI deprecated the old API. Was tracked as the 7.9 caveat in pr-roadmap.md. |
**One review comment intentionally not implemented:**
| Comment | Decision |
|---|---|
| gemini medium: `_load_router_config` duplicates `fastvideo.api.parser.parse_config` logic | Kept manual flat-from-nested mapping. The YAML schema has nested `health_check:` block but `RouterConfig` is flat; using `parse_config` directly would require either restructuring `RouterConfig` to have a nested `HealthCheckConfig` (schema change beyond this PR's scope) or accepting incomplete parsing. Manual mapping is intentional and well-typed. |
All 4 review threads marked resolved on the GitHub PR.
**Action items (deferred):**
- [ ] Track sticky session routing extensibility — when needed, add
`ReplicaRegistry.select(routing_key: str | None = None)` so registry
evolution is additive; document upfront where `session_id` should
appear (URL/header preferred over first JSON frame to avoid
buffering/peeking)
- [ ] Track `_bridge_session()` backpressure note — fine for MVP because
`websockets` library provides basic transport backpressure, but at
high scale add max_size/timeouts or recommend Envoy/HAProxy in front
- [ ] If active-active multi-primary becomes a requirement, define
behavior (round-robin within healthy primaries, weighted, sticky-by-key)
rather than letting `select()` silently pick `[0]`
**Watch outs:**
- `session_id` in WebSocket URL/headers is the cleanest sticky-routing
hook. If it ends up only in the first JSON message, sticky routing
later will require buffering/peeking before backend selection.
- Multi-primary configs are now explicitly rejected by validation;
documented + enforced.
- `_bridge_session()` is fine for MVP (the libraries provide basic
backpressure), but not production-grade for edge load. Document the
limit.
**Open thread it touches:** open-threads.md items #13 (sticky routing),
#14 (bridge backpressure), #15 (multi-primary semantics).
### D-16: Streaming router polish round 2 — second-pass fixes on top of D-15
**Status:** ✅ Resolved. Applied as `[fix] streaming: router polish — bridge
cancel + state machine + deps` (`a152cb77` on `will/api_7.9`, `40e265b8` on
`will/ltx2_sr_port`).
**Source:** Second-pass review on PR #1286, 2026-05-05, after D-15's pre-merge
polishes landed.
**Question:** D-15 closed the structural review (placement, sticky/active-active
deferral, basic `__post_init__` validation). On a second pass through the same
files, five latent issues surfaced that weren't covered by gemini's first pass
or Oracle's structural review. Apply them on top of the merged D-15 polishes,
or queue for a follow-up PR?
**Decision:** Apply on top of `will/api_7.9` directly. All five are bug-class
or DX-class — none are scope-expanding architecture changes — so folding them
into PR #1286 keeps the router landing in one reviewable unit instead of
shipping a router PR plus an immediate follow-up fix PR.
**Fixes applied:**
| # | File | What | Why |
|---|---|---|---|
| 1 | `router/main.py::_bridge_session` | Replaced `asyncio.gather()` with `wait(FIRST_COMPLETED)` + explicit `cancel()`/drain + `_is_normal_disconnect()` classifier | `gather` waited for both directions; on client disconnect, the backend-reader task leaked and stayed pending. Backend `ConnectionClosed` also surfaced as an unhandled exception in server logs. New shape: first task to finish triggers explicit cancel of the other, both are drained, and only non-routine exceptions re-raise. |
| 2 | `router/registry.py::record_success` | Split state transitions: `UNKNOWN -> HEALTHY` is now immediate on first successful probe; only `UNHEALTHY -> HEALTHY` remains gated by `recovery_threshold` | Previously a fresh registry needed `recovery_threshold` consecutive successes before any replica was selectable. With default `recovery_threshold=2` and `health_check_interval=1s`, that meant 2-3s of `gpu_unavailable` rejections at startup. Now the first probe promotes immediately; recovery gating still protects against flapping replicas. |
| 3 | `router/registry.py::_build_default_probe` | Missing `httpx` now raises `RuntimeError` with install hint instead of yielding a "disabled" probe stub | Previous behavior: silently returned `(0.0, "httpx not installed; ...")` for every probe, which `record_failure` then folded into `UNHEALTHY` after `failure_threshold` cycles. Operators saw replicas drop UNHEALTHY with a confusing reason and no clear remediation. Hard-fail at startup is the right surface. |
| 4 | `router/config.py::__post_init__` | Extended D-15 polish #4 with: rejects `urlparse(url).path not in ("", "/")`, rejects `query`/`fragment`, rejects duplicate URLs across replicas | D-15's validation rejected non-`http(s)://` URLs and >1 primary; it didn't catch `http://host/api` (the router appends `/health` and `/v1/stream` itself, so a base-URL with path yields malformed routes) or `[{url: x}, {url: x}]` (replica registry keys by URL — duplicates would silently collapse to one entry, masking the misconfiguration). |
| 5 | `cli/router_serve.py::_load_router_config` | Replaced silent list-comprehension filter (`for r in replicas_raw if isinstance(r, dict) and r.get("url")`) with per-index `raise ValueError` | Original parser silently dropped malformed YAML entries. A single typo in `replicas[2].url` would yield 2 replicas instead of 3 with no log line. New shape: explicit per-index error message ("missing required key 'url'", "must be a mapping"). |
| 6 | `pyproject.toml::[streaming]` extra | Added `websockets` as explicit dep | `router/main.py::_bridge_session` does `import websockets` lazily and raises `RuntimeError` if missing. The `[streaming]` extra was an implicit transitive — anyone installing only `[streaming]` (and not the broader requirements) hit the runtime error. Now explicit. |
**Tests added (7 cases in `fastvideo/tests/entrypoints/streaming/test_router.py`):**
- `TestUnknownToHealthyImmediate.test_first_success_promotes_unknown` — first probe success transitions `UNKNOWN -> HEALTHY` regardless of `recovery_threshold`
- `TestUnknownToHealthyImmediate.test_unhealthy_recovery_still_gated_by_threshold` — `UNHEALTHY -> HEALTHY` still requires `recovery_threshold` successes
- `TestConfigValidation.test_rejects_path_in_url` / `test_rejects_query_in_url` / `test_rejects_fragment_in_url` / `test_rejects_duplicate_urls` / `test_accepts_trailing_slash` — `__post_init__` URL validation matrix
**Verification:** 17/17 router tests pass on both branches. `pre-commit run`
clean (yapf / ruff / codespell / mypy). `lsp_diagnostics` clean on changed
regions; the one pre-existing `Task` generic-type warning at `main.py:37` is
unrelated and predates this commit.
**In-flight pre-commit corrections (not part of the 6 fixes themselves):**
- yapf auto-reformatted 4 files (kept verbatim).
- ruff `UP038`: rewrote `isinstance(exc, (CancelledError, WebSocketDisconnect))`
to `isinstance(exc, CancelledError | WebSocketDisconnect)`.
- mypy `[misc]`: renamed loop var `exc` (inside `for task in done`) to
`task_exc` to avoid name collision with the outer
`except ImportError as exc` binding.
**Open thread it touches:** None new. Item #14 (bridge backpressure) and
item #13 (sticky routing) from D-15 remain deferred — this round addressed
**cancellation/disconnect** semantics on the bridge, which is distinct from
**throughput backpressure**. Item #14 still applies: at higher load, add
`_bridge_session()` max-size + timeout limits or recommend Envoy/HAProxy
in front.
### D-14: Streaming auxiliaries (PR #1284) — cohesion + concrete-vs-Protocol scoping
**Status:** ✅ Resolved (interim). Two polish items applied during review; one
operational caveat tracked.
**Source:** Oracle review on 2026-05-04, during PR #1284 review cycle.
**Question:** Is PR #1284's bundle of 4 streaming-server auxiliary modules
(`prompt/safety.py`, `prompt/rewrite.py`, `session_logger.py`,
`mock_server.py`) correctly scoped? Should `mock_server` live in production
module path? Should `PromptSafetyFilter` be a Protocol? Should the bundle
have been split into 4 PRs?
| Alt | Approach | Verdict |
|---|---|---|
| A | Status quo — single PR, 4 modules under `streaming/`, mock_server in production path, concrete safety filter | ✅ **Keep** |
| B | Split into 4 separate PRs | ❌ Process overhead, not architectural improvement |
| C | Move `mock_server.py` into `tests/` | ❌ Would reduce discoverability + install-time usability of `python -m fastvideo.entrypoints.streaming.mock_server` |
| D | Move `session_logger.py` to `streaming/observability/` (or top-level `fastvideo/observability/`) | ❌ Premature — currently session-shaped + streaming-specific; promote when a non-streaming consumer appears |
| E | Convert `PromptSafetyFilter` to Protocol (like `LLMProvider`) | ❌ Premature abstraction — only one classifier exists; small duck-typed surface preserves future Protocol introduction without breaking the concrete |
| F | Convert `MockGenerator` to Protocol | ❌ Same — small duck-typed surface; no second mock generator exists |
**Decision:** Alt A — keep current shape. Apply two polish items from
Oracle's review before merge.
**Rationale:**
- "Streaming-server auxiliaries" is cohesive enough at 730 LOC with
isolated modules + tests. Each module has independent code path but
shared deployment context (the streaming server boots them all).
- `mock_server.py` in production path is a strength: reuses
`build_app()` for protocol parity. Hiding it under `tests/` would lose
`python -m fastvideo.entrypoints.streaming.mock_server` CLI access for
FE devs.
- Concrete `PromptSafetyFilter` matches "ship what we have, abstract
later" pattern. Internal had multi-classifier composition; public
ships single + leaves chaining as a Dreamverse-side concern (per D-2).
- Same pattern for `MockGenerator`: small duck-typed `_GeneratorLike`
surface lets a second mock implementation drop in without inheritance.
- `threading.Lock` (not `asyncio.Lock`) in `session_logger.py` is
correct — writes come from real encoder/control threads via
`run_in_executor`, not from coroutines directly. `asyncio.Lock` would
be the wrong primitive for cross-thread concurrency.
**Pre-merge polishes applied (per Oracle):**
| Polish | What | Why |
|---|---|---|
| 1 | Removed `RewriteOptions.user_system_prompt_override` | Inert public field — was declared but never threaded through to `enhancer.rewrite()`. Shipping unused public options is more likely to bite than any structural choice. Re-add when actually wired through. |
| 2 | Sanitized `session_id` filename in `session_logger.SessionLogger._get_file()` | Defense-in-depth: today session_id is server-generated UUID, but a future code path that accepts client-supplied ids would otherwise allow path traversal via `../`. Added `_FILENAME_SANITIZE_RE = re.compile(r"[^A-Za-z0-9._-]")` + sub before `os.path.join`. |
**Operational caveat tracked (not a code change):**
- `SafetyDecision.UNAVAILABLE` is treated as `ALLOW` by callers — a
policy choice that's correct for an opt-in safety filter, but
callers should log loudly so operators know the filter is degraded.
Tracked as open-threads.md item #12.
**Pre-merge review feedback (4 of 4 resolved on the GitHub PR):**
| # | File:Line | Severity | Issue | Fix applied |
|---|---|---|---|---|
| 1 | `session_logger.py:57` | High | `log()` race vs `close()` — `KeyError` on `_locks[session_id]` | Atomic capture in `_get_file()`; master `_registry_lock`; `with lock, contextlib.suppress(ValueError):` |
| 2 | `rewrite.py:71` | Medium | `re.compile()` in hot path | Module-level `_LEADING_MARKER_RE`, top-level `import re` |
| 3 | `safety.py:105` | Medium | `_ensure_loaded()` race on concurrent fastText load | `_load_lock = threading.Lock()` + double-check pattern |
| 4 | `pyproject.toml:145` | Medium | `streaming` extra missing `prompt-safety` | Added to aggregator |
All 4 review threads marked resolved via GraphQL `resolveReviewThread`.
**Action items (deferred):**
- [ ] Track `SafetyDecision.UNAVAILABLE` log loudness in
open-threads.md item #12 — when streaming server starts using the
safety filter, ensure operator-visible logging on `UNAVAILABLE`
results
- [ ] If a second safety classifier appears (Perspective API, Detoxify,
custom rules), promote `PromptSafetyFilter` to a Protocol — same
pattern as `LLMProvider` per D-13
- [ ] If a second mock generator appears (different frame patterns,
different latency models), promote `MockGenerator` to a Protocol
**Open thread it touches:** PR #1284 itself; future safety-classifier
Protocol promotion; future observability module extraction.
### D-13: Prompt enhancer / `LLMProvider` abstraction shape — keep streaming-scoped
**Status:** ✅ Resolved (interim) + 🟡 Three deferred polishes after metrics or 2nd consumer.
**Source:** Oracle review on 2026-05-04, pre-PR-#1258-merge.
**Question:** Is PR #1258's `fastvideo.entrypoints.streaming.prompt.*` module
correctly designed? Should it be (a) Protocol-based vs ABC, (b) under
`streaming/` vs top-level `fastvideo.prompt.*`, (c) closed 3-op enum vs
open `complete()` API?
| Alt | Approach | Verdict |
|---|---|---|
| A | Status quo — `streaming/prompt/*`, Protocol provider, fixed 3 ops, lazy `httpx`, per-call `AsyncClient` | ✅ **Keep** |
| B | Move to top-level `fastvideo.prompt.*` (decouple from streaming) | ❌ **Premature.** No second consumer exists yet. |
| C | Convert `LLMProvider` Protocol → ABC with default impls + retry classification | ❌ **Wrong direction.** Biases extension toward OpenAI shape; `_openai_compat.py` already factors that as helper not inheritance. |
**Decision:** Alt A as interim. Promote to Alt B only when a second
non-streaming consumer (OpenAI server, batch generation, tooling) actually
needs the prompt enhancer. Don't pursue Alt C.
**Rationale:**
- Public contract is tiny — `name: str` + `async complete(LLMRequest) -> LLMResponse`. ABC adds zero value.
- `_openai_compat.py` is the right place for shared logic — helper, not base class. Anthropic / local / custom providers stay first-class.
- The 3 ops (enhance / auto_extend / rewrite) are LTX-2 streaming concepts. `auto_extend` (continue prompt sequence) and `rewrite` (multi-line alternatives) come directly from session UX. Calling this "the FastVideo prompt API" misrepresents that.
**Specific risks flagged in PR #1258 (already merged-pending review):**
| Risk | Mitigation (when relevant) |
|---|---|
| API publicity — calling this "the FastVideo prompt API" before a second consumer exists | Document module as "streaming-server prompt enhancement" in user-facing docs; keep it nested under `entrypoints/streaming/` |
| `httpx.AsyncClient` per-call (no connection pooling) | Acceptable for ~6-10 calls per LTX-2 session; LLM latency dominates. Add optional `client_factory` parameter LATER if metrics show connect overhead is meaningful. |
| 3 fixed operations could constrain future generic use | Closed enum is right for application-level orchestration. Future generic consumers should either call `provider.complete()` directly, or get a thin separate enhancer that shares the provider/fallback machinery. |
| `register_provider(priority=-1)` semantics rely on Python's negative-index `list.insert` | Cosmetic concern; docstring is clear. Could be tightened to explicit branch later. |
| `runtime_checkable` Protocol with `name: str` instance attribute — static type checkers may miss missing `name` | Acceptable; runtime check via `isinstance(p, LLMProvider)` works for plugin discovery. |
**Action items (deferred):**
- [ ] Document `fastvideo.entrypoints.streaming.prompt.*` as streaming-scoped in user-facing docs (PR 12 docs migration); avoid promoting as framework-level
- [ ] Add optional `client_factory` parameter to providers when metrics justify pooling
- [ ] Plan future move to `fastvideo.prompt.*` (with import shim) when second non-streaming consumer materializes
- [ ] Track Q-2 reactivation: promote LTX-2 prompt orchestration (locked segments, segment-prompts JSON parsing) to public `fastvideo.entrypoints.streaming.prompt.ltx2_orchestration` when a second LTX-2-style consumer appears
**Open thread it touches:** Dreamverse migration (open-threads.md DR-1)
will be the first real test of the public surface. Lessons learned there
inform whether Alt B becomes feasible.
## D-decisions (from `dreamverse_review.md`, Apr 26)
### D-1: Realtime runtime → streaming GpuPool migration shape
**Status:** ✅ Resolved.
Internal `RealtimeRuntimeConfig` had a multi-model registry +
flattened sampling defaults. Public `SubprocessGpuPool` is single-model
+ uses per-request `SamplingConfig`.
**Decision:** Drop multi-model registry on integration branch (not used
in production). Construct `GeneratorConfig` for chosen model and pass to
`SubprocessGpuPool`. Move sampling defaults to a server-side
`default_request: GenerationRequest` template.
**Risk:** Migration branch surfaces missing-model errors if a flow
silently relied on registry to swap models per-session. Integration
tests exercise at least one segment per supported model id before
merging.
### D-2: PR 7.7 prompt enhancer API surface narrower than internal
**Status:** ✅ Resolved.
Public `PromptEnhancer.enhance/auto_extend/rewrite` returns
`LLMResponse(content, provider, model, latency_ms, fallback_used)`.
Internal returns `EnhanceResult(prompt, fallback_used, error, ...)` /
`RewriteResult(prompts, ..., rollout_id, rollout_label, ...)`.
**Decision:** Adapt at the call site via
`Dreamverse/server/prompting/_internal_compat.py` shim. Locked-segment /
next-segment-index plumbing stays Dreamverse-side. Public stays minimal
and provider-agnostic.
**Open question (Q-2):** Promote LTX-2-specific orchestration into
`fastvideo.entrypoints.streaming.prompt.ltx2_orchestration` once a
second consumer appears. Logged for future review.
### D-3: Multi-stage provider race vs. sequential fallback
**Status:** ✅ Resolved (public stays sequential).
Internal enhancer runs all providers in a stage in parallel
(`_run_provider_race`). Public enhancer runs sequentially with
retryable-error fallback.
**Decision:** Public stays sequential for PR 7.7. Race is a
Dreamverse-specific tail-latency optimization that depends on parallel
API budgets.
**Risk / Q-3:** First-segment latency on Dreamverse may regress
slightly when Cerebras has a bad minute (sequential waits 20s before
trying Groq). If real production concern, add public
`concurrency: int = 1` knob behind a race path — but only after measuring.
### D-4: Skip PR 7.9 router for the integration branch
**Status:** ✅ Resolved.
Internal stack ships `router/main.py` for multi-replica load balancing.
Dreamverse deployment uses single replica per region.
**Decision:** Land PR 7.9 publicly (upstream the surface). Skip wiring
into Dreamverse integration branch. Dreamverse's `server/main.py` does
not import from `router/`.
### D-5: Audio re-encode (PR 7.10) needed for streaming, deferred
**Status:** 🟡 Deferred to PR 7.10.
Internal streaming server's per-step path runs `_re_encode_audio` inside
`_stream_av_fmp4_events` so each fMP4 segment ships with
continuation-conditioning audio. Whole-segment `pool.run()` path doesn't
need this.
**Decision:** Land PR 7.10's `generate_async` publicly. Dreamverse
integration branch initially keeps using `pool.run()` (whole segment, no
re-encode). Follow-up branch swaps to `generate_async` + audio re-encode.
**Open question (Q-5):** Acceptable for first switch, or does
Dreamverse audio quality regress vs. internal until 7.10 wires in?
### D-6: `realtime/local_runtime.py` is NOT upstreamed
**Status:** ✅ Resolved.
It was the FastVideo-internal precursor to `streaming.gpu_pool`.
Upstreaming both would create two GPU pool implementations in public.
**Decision:** Don't upstream `realtime/local_runtime.py`. Dreamverse
switches to `streaming.gpu_pool.SubprocessGpuPool` on integration
branch. Internal module can be deleted at follow-up.
### D-7 / Q-6: `FP4Config` is private-only
**Status:** ✅ **Resolved May 2.**
April 26: `Dreamverse/server/video_generation.py:271` imported
`fastvideo.layers.quantization.fp4_config.FP4Config` from
FastVideo-internal only. The 411-line module hard-imported `flashinfer`.
**Two options at the time:**
1. Colocate publicly with `flashinfer` as optional extra
`pip install fastvideo[fp4]`; refactor `FP4QuantizeMethod` to take
layer-prefix list from a pipeline-config field instead of hardcoding
ltx2 paths.
2. Keep private — Dreamverse imports from internal via thin shim.
**Recommendation at the time:** option 1 once API refactor settles.
**Resolution:** May 2 work chose option 1.
- `365a66c7` upstreamed FP4Config with lazy `flashinfer` import in
loader helper (no public hard-dep)
- `94c983a2` renamed FP4 → NVFP4 to disambiguate from MX-FP4 / OCP-FP4
- `42b30bf9` wired through `fastvideo.layers.quantization`
See [quantization.md](quantization.md) for full details.
### D-8: `ltx2_image_crf` silently dropped by public schema
**Status:** 🔴 **Unverified post-`d80c2a8`.**
April 26: Dreamverse's `server/video_generation.py:406` passed
`ltx2_image_crf=0.0` to `SamplingParam(...)`. Public
`fastvideo.api.sampling_param.SamplingParam` did NOT have this field;
the BE logged ERROR and silently dropped the kwarg.
**Migration target** (per [design.md](design.md) compatibility map):
`request.stage_overrides.refine.image_crf`.
**Resolution status:** `d80c2a8` (May 2) refactored
`server/video_generation.py` to use typed `GeneratorConfig` +
`preset_overrides["refine"]`. Whether this PR routed `image_crf`
through the typed `stage_overrides` path or left it silently dropped is
unverified. See [open-threads.md](open-threads.md).
### D-9: `aarch64-conda-linux-gnu-cc` triton compile failure
**Status:** ✅ Resolved (operational).
Conda env injected an ARM cross-compiler ahead of `gcc` on `$PATH`, so
`torch._inductor`'s triton launcher failed compilation. Setting
`ENABLE_TORCH_COMPILE=0` bypasses it.
**Long-term fix:** clean conda env's compiler shadowing or add
`CC=gcc` override in Dreamverse's worker bootstrap.
### D-10: Warmup OOM on shared GPU
**Status:** ✅ Resolved (operational).
When `CUDA_VISIBLE_DEVICES` lands on a GPU another tenant uses, LTX-2
warmup fails with OOM. Picking an idle GPU (4-7 in test setup) is a
manual step.
**Improvement:** pre-warm probe that checks free memory before booting
the pool would prevent this.
### D-11: ffmpeg fragment write `Broken pipe`
**Status:** ✅ Resolved (cosmetic).
When WS client closes before backend finishes streaming first segment,
ffmpeg hits `[Errno 32] Broken pipe`. Currently propagates to
"User step failed". Cosmetic — swallowing pipe-broken on intentional
disconnect would clean up logs.
## Q-questions (from `streaming-server-upstream-plan.md`, Apr 17)
### Q-1: Router placement (in-repo or separate package)
**Status:** ✅ Resolved (in-tree).
**Recommendation at the time:** separate package `fastvideo-router/` or
`fastvideo/contrib/router/`; defer final call to PR 7.9.
**Resolution:** PR 7.9 implementation places router in-tree at
`fastvideo/entrypoints/streaming/router/`.
### Q-2: Session ID authority
**Status:** ✅ Resolved (server-generated).
**Recommendation:** server-generated UUID; accept externally provided
session ID only for resume flows.
### Q-3: Torch compile kwargs typing (opaque vs full vs hybrid)
**Status:** ✅ Resolved (hybrid).
**Recommendation:** hybrid — type the common four (`backend`,
`fullgraph`, `mode`, `dynamic`) + allow `extras: dict[str, Any]`.
**Resolution:** PR 6 + NVFP4 `221cb20a` shipped exactly this hybrid.
### Q-4: Prompt safety / fasttext dependency
**Status:** ✅ Resolved (optional extra).
**Recommendation:** ship as optional extra `pip install fastvideo[prompt-safety]`.
**Resolution:** PR 7.8 implements as optional extra.
### Q-5: Audio-specific tensor payloads in continuation
**Status:** ✅ Resolved (typed `LTX2ContinuationState`).
`ltx2_audio_clean_latent`, `ltx2_audio_denoise_mask`,
`ltx2_audio_latents` not in pre-refactor public schema.
**Recommendation:** classify as opaque fields inside
`LTX2ContinuationState.payload`, not top-level sampling fields.
**Resolution:** PR 7's typed `LTX2ContinuationState` lifts these into
typed fields (see [cross-repo-surfaces.md](cross-repo-surfaces.md)
field mapping table).
### Q-6: Dynamo subpackage home
**Status:** ✅ Resolved (lives in Dynamo repo).
**Resolution:** No Dynamo code in FastVideo. Full backend package
(handler, adapter, registration, health check) owned by Dynamo repo at
`components/src/dynamo/fastvideo/`, same pattern as vllm/sglang.
FastVideo only guarantees the public API contract.
### Q-7 (was Q-6 in dreamverse_review): How to land FP4Config publicly
**Status:** ✅ Resolved May 2 — option 1 (colocate publicly).
See D-7 above.
### Q-8: Disaggregation readiness contract test
**Status:** 🟡 Recommended; not yet shipped.
PR ai-dynamo/dynamo#7544 is aggregated-only. `ContinuationState` hybrid
already supports future prefill/decode split.
**Recommendation:** PR 7.10 explicitly validate `ContinuationState`
survives round-trip through Dynamo-style RPC (pickle or JSON), even
though Dynamo isn't using it today. Cheap regression guard.
### Q-9: Dynamo progress/status passthrough
**Status:** 🟡 Deferred until Dynamo clarifies.
`NvVideosResponse` has `status` and `progress` fields.
**Recommendation:** PR 7.10 stays aggregated-final-only to match PR
#7544 shape; revisit after Dynamo clarifies their streaming/progress
semantics.
## Cross-doc questions still 🔴 OPEN
These need decisions; tracked also in [open-threads.md](open-threads.md):
| ID | Question | Source | Why it matters |
|---|---|---|---|
| **D-8** | Did `d80c2a8` route `ltx2_image_crf` correctly, or is it still silently dropped? | dreamverse_review | Latent silent-drop bug; FP4-disabled paths may degrade |
| **VPO** | `video_position_offset_sec` — persistent accumulation (a) vs per-segment hint (b) | dreamverse_integration | Needs decision before PR 7.6 emits state |
| **SBS** | `SessionStore` / `BlobStore` lifecycle (TTL/eviction/blob-drop on state replacement) | dreamverse_integration | Needs decision in PR 7.5 design pass |
| **#1** | Migrate `/healthz`+`/readyz`+`/status` into FastVideo `build_app` | streaming-upstream-plan + handoff | Closes BE_FLAVOR=fastvideo FE-compatibility |
| **#3** | Add `cerebras_ifm` to public `PromptEnhancerConfig.provider` Literal | handoff | Internal supports it; public schema doesn't |
| **#4** | Expose `layer_profile` on typed `engine.quantization` | handoff | Removes Dreamverse's `experimental["pipeline_config"]` dodge |
| **#5** | Typed `dit_config.quant_config` carrier (design TBD) | handoff | Eliminates the `experimental["pipeline_config"]` escape hatch entirely |
@@ -1,332 +0,0 @@
# Design — Typed Public Inference API
Synthesis of the FastVideo public inference API refactor design philosophy.
For PR-by-PR execution see [pr-roadmap.md](pr-roadmap.md). For the streaming
extension see [streaming-server.md](streaming-server.md).
**Last updated:** 2026-05-03.
## Why the refactor
The pre-refactor public boundary mixed three concerns through `**kwargs`:
- `VideoGenerator.from_pretrained(..., **kwargs)` mixed engine/runtime,
pipeline init, and component overrides.
- `VideoGenerator.generate_video(..., **kwargs)` mixed prompt+inputs,
sampling, output, and model-specific workflow knobs.
- Unknown keys silently filtered or merely logged → API drift hard to detect.
- Multi-stage models (LTX-2 two-stage, Hunyuan15 SR, LongCat distill+refine)
exposed via ad hoc top-level flags.
This was already painful for LTX2/Dreamverse and would worsen as more
multi-stage pipelines came in.
## Core decision
FastVideo has:
1. **Typed nested public schema** — `RunConfig`, `ServeConfig`,
`GeneratorConfig`, `GenerationRequest`, `ContinuationState`.
2. **Model-owned named pipeline presets** — `ltx2_two_stage`,
`longcat_distill_refine`, `hunyuan15_sr_1080p`, etc. All 13 model families
landed presets in PR 4.
3. **Semantic stage overrides by stage name** —
`request.stage_overrides["refine"] = LTX2RefineStageOverride(...)`.
4. **Optional advanced explicit plans** for power users — `GenerationPlan`
(escape hatch only; not the canonical surface).
5. **YAML-first CLI** with dotted overrides —
`fastvideo generate --config run.yaml --request.sampling.seed 42`.
The canonical user experience: choose a model → choose a preset → override
a few typed fields → generate. Dicts/YAML/JSON are supported as
serialization, but parse immediately into typed objects with strict
unknown-key validation.
## Schema surface
Implemented in [`fastvideo/api/`](file:///home/william5lin/FastVideo/fastvideo/api/):
| Type | Role |
|---|---|
| `RunConfig` | Offline envelope: `generator` + `request` |
| `ServeConfig` | Serving envelope: `generator` + `server` + `default_request` + optional `streaming` |
| `GeneratorConfig` | `model_path`, `revision`, `trust_remote_code`, `engine`, `pipeline` |
| `EngineConfig` | parallelism / offload / compile / quantization / flags |
| `PipelineSelection` | `workload_type`, `preset`, `preset_version`, `components`, `preset_overrides`, `experimental` |
| `GenerationRequest` | `prompt`, `negative_prompt`, `inputs`, `sampling`, `runtime`, `output`, `stage_overrides`, `state`, `plan`, `extensions` |
| `ContinuationState` | Opaque envelope `{kind: str, payload: dict[str, Any]}` |
| `GenerationPlan` | Advanced/escape-hatch only; `{stages: list[PlannedStage], final_stage: str|None}` |
Files:
| File | Role |
|---|---|
| [`schema.py`](file:///home/william5lin/FastVideo/fastvideo/api/schema.py) | All public dataclasses |
| [`parser.py`](file:///home/william5lin/FastVideo/fastvideo/api/parser.py) | `from_dict`, `to_dict`, `load_yaml`, `load_json`, validation |
| [`overrides.py`](file:///home/william5lin/FastVideo/fastvideo/api/overrides.py) | Dotted override application |
| [`compat.py`](file:///home/william5lin/FastVideo/fastvideo/api/compat.py) | Legacy kwargs translation (~370 lines, scheduled for death PRs 14-17) |
| [`presets.py`](file:///home/william5lin/FastVideo/fastvideo/api/presets.py) | Preset registry |
| [`sampling_param.py`](file:///home/william5lin/FastVideo/fastvideo/api/sampling_param.py) | Internal `SamplingParam` adapter (canonical home since PR 4) |
| [`results.py`](file:///home/william5lin/FastVideo/fastvideo/api/results.py) | `GenerationResult` / `VideoResult` |
| [`errors.py`](file:///home/william5lin/FastVideo/fastvideo/api/errors.py) | Path-aware validation errors |
## Boundary normalization rule
Every public inference entrypoint normalizes into typed config objects
before touching legacy internals (`FastVideoArgs`, `SamplingParam`).
Includes Python constructors, `generate*` calls, CLI `generate`, CLI
`serve`, OpenAI server request translation, streaming server request
translation.
Legacy internals (`FastVideoArgs`, `SamplingParam`) may remain temporarily,
but only behind a typed normalization boundary.
## Strict-by-default validation
All structured inputs are strict:
- Unknown keys → error
- Wrong types → error
- Invalid stage names → error
- Incompatible state/preset combinations → error
The only intentional escape hatches:
- `generator.pipeline.experimental` — for in-flight features without typed home
- `request.extensions` — same, request-side
These bypass validation by design. Intent: shrink as presets absorb
model-specific fields. New fields should not land in `experimental` /
`extensions` without a plan to either promote them to typed fields or
remove them within two PR cycles.
Error format includes nested path:
```
Invalid field: request.stage_overrides.refine.num_inference_steps
Expected int, got "two"
Preset: ltx2_two_stage
Stage: refine
```
## Request mutation tracking
When a `GenerationRequest` is parsed from raw dict (YAML/JSON/Python),
FastVideo tracks which fields the user explicitly provided vs. which got
schema defaults. Matters for `request_to_sampling_param()` — explicit
values override model defaults; schema defaults do NOT.
Mechanics:
- At parse time, original raw dict + baseline snapshot stored on the request.
- Dataclass field mutations (e.g. `request.sampling.seed = 7`) captured via
lightweight `__setattr__` dirty-path recording.
- Dict-typed field mutations (e.g. `del request.stage_overrides["refine"]`)
detected at access time by diffing current dict vs. baseline.
- Setting a field to its schema default value IS captured as explicit, so
it overrides model defaults.
- Raw dict reconciled lazily when `normalize_generation_request()` is called.
## Schema purity (model-specific fields still in shared schema)
Remain for back-compat during initial migration; targeted for migration
into preset-owned typed override classes:
| Field | Owner | Migration target |
|---|---|---|
| `SamplingConfig.height_sr` / `width_sr` / `num_inference_steps_sr` | Hunyuan15 SR | `HunyuanSRStageOverride` (PR 10) |
| `SamplingConfig.guidance_scale_2`, `boundary_ratio` | Wan2.2, LingBotWorld | preset-owned (per-family PR) |
| `InputConfig.mouse_cond`, `keyboard_cond`, `grid_sizes` | MatrixGame | `request.extensions` or typed input config |
| `InputConfig.c2ws_plucker_emb` | LingBotWorld | `request.extensions` or typed input config |
| `InputConfig.refine_from`, `stage1_video` | LongCat | `LongCatRefineStageOverride` inputs (PR 9) |
LTX-2 multi-modal CFG knobs (`ltx2_modality_scale_video/_audio`,
`ltx2_rescale_scale`, `ltx2_stg_scale_video/_audio`,
`ltx2_stg_blocks_video/_audio`) still leak into shared `SamplingParam` but
only LTX-2 reads them today. Migration to typed `LTX2SamplingOverride` is
deferred to per-model migration sweep.
**LTX-2 CFG-force fix landed in PR 6**: defaults moved from `3.0/7.0` to
`1.0/1.0` to stop force-enabling CFG for non-LTX-2 families.
`ltx2_base` preset still sets `3.0/7.0` explicitly. Regression guard:
`test_presets.py::TestPresetDefaultTypes::test_ltx2_cfg_defaults_are_off`.
## Continuation state
Public surface:
```python
@dataclass
class ContinuationState:
kind: str # e.g. "ltx2.v1"
payload: dict[str, Any]
```
Internally, model-specific typed subclasses (e.g. `LTX2ContinuationState`
at [`fastvideo/pipelines/basic/ltx2/continuation.py`](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/continuation.py)).
Payload must be JSON-serializable or use opaque blob-ID indirection for
large tensors — supports both stateless OpenAI client round-trip AND
future Dynamo prefill/decode disaggregation.
Hybrid model: server-held for streaming WS, client-round-trip for
stateless HTTP. See [streaming-server.md](streaming-server.md) D-1.
## Pipeline package structure (target)
Per-family colocation under `pipelines/basic/<family>/`:
```
fastvideo/pipelines/basic/<family>/
├── <family>_pipeline.py # pipeline implementation(s)
├── presets.py # user-facing presets (DONE in PR 4)
├── pipeline_configs.py # engine/arch config (from configs/pipelines/)
└── stages/ # model-specific stages (optional, if >2 files)
```
What stays shared:
- `configs/pipelines/base.py` — `PipelineConfig` base class
- `configs/models/` — architecture defs (dits/, vaes/, encoders/)
- `pipelines/stages/` — shared stages only (denoising, encoding, decoding,
text_encoding, timestep_preparation, ...)
What's gone (PR 4):
- `fastvideo/configs/sample/` — directory removed entirely; defaults
absorbed into per-family `presets.py`.
- All 12 `*_SamplingParam` subclass files — `SamplingParam` lives at
`fastvideo/api/sampling_param.py`; defaults flow through
`SamplingParam.from_pretrained()` → `_from_preset()`.
What's pending: `configs/pipelines/<family>.py` colocation, optional
`pipelines/stages/<family>_*.py` colocation. Per-model migration PRs
(6/9/10) include the colocation step for that family.
## YAML examples
### Run config
```yaml
generator:
model_path: /models/ltx2
engine:
num_gpus: 1
parallelism: {tp_size: -1, sp_size: -1}
offload: {dit: false, text_encoder: false, vae: false, pin_cpu_memory: true}
pipeline:
workload_type: t2v
preset: ltx2_two_stage
components:
config_root: /models/ltx2-config
upsampler_weights: /models/ltx2-refine
lora_path: /models/ltx2-refine-lora
preset_overrides:
refine: {enabled: true, add_noise: true}
request:
prompt: "a fox running through snow"
sampling: {num_frames: 121, height: 1024, width: 1536, num_inference_steps: 8, seed: 42}
output: {save_video: true, return_state: true}
stage_overrides:
refine: {num_inference_steps: 2, guidance_scale: 1.0}
```
### Serve config
See [`Dreamverse/serve_configs/streaming_demo.yaml`](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml)
for a canonical example matching internal/ui defaults (LTX-2 distilled,
NVFP4, 121 frames @ 1088×1920 24fps, 5 inference steps, 2-step refine).
## Compatibility mapping (legacy → typed)
| Legacy field | New path |
|---|---|
| `model_path` | `generator.model_path` |
| `num_gpus` | `generator.engine.num_gpus` |
| `tp_size` / `sp_size` | `generator.engine.parallelism.{tp_size,sp_size}` |
| `dit_cpu_offload` | `generator.engine.offload.dit` |
| `enable_torch_compile` | `generator.engine.compile.enabled` |
| `torch_compile_kwargs` | split: `generator.engine.compile.{backend,fullgraph,mode,dynamic}` + `.extras` |
| `enable_torch_compile_text_encoder` | `generator.engine.compile.text_encoder_enabled` |
| `prompt_txt` | `request.inputs.prompt_path` |
| `image_path` / `video_path` | `request.inputs.{image_path,video_path}` |
| `output_path` / `save_video` / `return_frames` | `request.output.*` |
| `seed` / `num_frames` / `height` / `width` / `fps` / `num_inference_steps` / `guidance_scale` | `request.sampling.*` |
| `enable_teacache` / `return_trajectory_*` | `request.runtime.*` |
LTX-2 specific (private adapter, NOT public compat promise):
| Legacy LTX-2 field | New path |
|---|---|
| `config_model_path` | `generator.pipeline.components.config_root` |
| `ltx2_refine_enabled` | `generator.pipeline.preset_overrides.refine.enabled` |
| `ltx2_refine_upsampler_path` | `generator.pipeline.components.upsampler_weights` |
| `ltx2_refine_lora_path` | `generator.pipeline.components.lora_path` |
| `ltx2_refine_num_inference_steps` | `request.stage_overrides.refine.num_inference_steps` |
| `ltx2_refine_guidance_scale` | `request.stage_overrides.refine.guidance_scale` |
| `ltx2_refine_add_noise` | `generator.pipeline.preset_overrides.refine.add_noise` |
| `ltx2_image_crf` | `request.stage_overrides.refine.image_crf` |
| `return_continuation_state` | `request.output.return_state` |
LongCat:
| Legacy | New |
|---|---|
| `refine_from` / `stage1_video` | `request.inputs.{refine_from,stage1_video}` |
| `t_thresh` / `spatial_refine_only` / `num_cond_frames` | `request.stage_overrides.refine.*` |
## External inspirations (and limits)
| Source | Useful idea | Don't copy |
|---|---|---|
| Ray | YAML-first config interchange | Ray's package layout |
| SGL `multimodal_gen` | Split instance/request config; dict input parsed into typed objects; merge user overrides on model defaults | `SamplingParams._adjust(ServerArgs)` (request depending on engine config); broad weakly-typed request bags |
| vLLM-Omni | Model-owned pipeline presets; explicit stage topology; per-stage default sampling | Positional `sampling_params_list`; serving-engine stage-index semantics in primary Python API |
## Naming guidance
- Public schema names namespaced under `fastvideo.api`
- Don't export from top-level `fastvideo/__init__.py` until migration further along
- `RunConfig` / `ServeConfig` get sufficient disambiguation from training
config via the namespace
- Future rename to `EngineQuantizationConfig` reserved if a collision
arises (deferred)
## Public Python API (canonical form)
```python
from fastvideo import VideoGenerator
from fastvideo.api import (
GeneratorConfig, GenerationRequest,
EngineConfig, OutputConfig,
PipelineSelection, SamplingConfig,
)
generator = VideoGenerator.from_pretrained(
config=GeneratorConfig(
model_path="/models/ltx2",
engine=EngineConfig(num_gpus=1),
pipeline=PipelineSelection(workload_type="t2v", preset="ltx2_two_stage"),
)
)
result = generator.generate(
GenerationRequest(
prompt="a fox running through snow",
sampling=SamplingConfig(num_frames=121, height=1024, width=1536,
num_inference_steps=8, seed=42),
output=OutputConfig(save_video=True, return_state=True),
)
)
```
Accepted constructor forms:
```python
VideoGenerator.from_pretrained(config=GeneratorConfig(...))
VideoGenerator.from_config(GeneratorConfig(...))
VideoGenerator.from_file("run.yaml")
VideoGenerator.from_pretrained("model-id", num_gpus=2, ...) # stable convenience
VideoGenerator.from_pretrained(model_path, **legacy_kwargs) # compat (deprecated PR 13)
```
File diff suppressed because it is too large Load Diff
@@ -1,888 +0,0 @@
# Integration Review — Drift Audit + Path Forward
> # ⚠️ DEPRECATED — superseded by [integration-plan.md](integration-plan.md)
>
> This document recommended **Option D** (Dreamverse stays a separate repo;
> generic backend merges into FastVideo). On 2026-05-05 the team chose
> **Option B+** instead (Dreamverse FE + product server move into FastVideo
> as `apps/dreamverse/`; generic backend stays at
> `fastvideo.entrypoints.streaming.*` per Option D's principle).
> See [decisions-log.md D-18](decisions-log.md#d-18) for the strategy
> reversal rationale and [integration-plan.md](integration-plan.md) for the
> executable migration plan.
>
> **What's still authoritative in this file:**
> - **Part 1 — Drift audit** (the 17-row drift summary table). The drift
> findings remain valid; the migration plan in `integration-plan.md`
> folds them into specific phases.
> - **OSS precedent citations** (vLLM, BentoML, Ray Serve, TGI+ChatUI,
> Transformers.js, ComfyUI, AUTOMATIC1111). Reused in `integration-plan.md`.
>
> **What's superseded:**
> - **Part 2 — Recommendation (Option D)**. Replaced by Option B+ in the
> new plan. Read `integration-plan.md` for the current decision.
> - **Part 3 — Action items**. Replaced by the phased migration plan.
>
> Kept in tree for historical reference and audit trail. Do not delete.
**Last updated:** 2026-05-05 (deprecated header added).
**Scope:** FastVideo public `will/ltx2_sr_port` at the requested audit
anchor `b36bdbc9`; Dreamverse `will/integrate-public-fastvideo` at
`ec8ef92`; FastVideo-internal `will/rebase-nbv` as read-only comparison.
**Memory-dir context:** the current integration memory snapshot tracks the
same public mega-PR lineage as `will/ltx2_sr_port`, with PRs #1257,
#1258, #1284, and #1286 already merged, #1287 closed, and #1288 open as
the consolidated landing vehicle for LTX-2 SR runtime, NVFP4,
`generate_async`, Dynamo contract, and memory-dir cleanup. Source:
[memory index](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/README.md#L8-L19)
and [D-17](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L19-L45).
**Bottom line:** Zero core typed API drift — typed construction, typed
continuation state, NVFP4 wiring, and Dynamo-facing async events are either
already public or in #1288. **Real drift remains on the realtime-runtime
contract surface (`/healthz` / `/readyz` / `/status` routes per
[cross-repo-surfaces.md](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L74-L88))
and on operational/product edges**: stale Dreamverse docs/scripts, a
1933-LOC Dreamverse prompt-enhancer fork, two unresolved per-session
fields (`ltx2_image_crf` D-8, `video_position_offset_sec` VPO), one
missing example config, and two internal-only utilities whose product
relevance is not yet proven.
---
## Part 1 — Drift audit
### Methodology
1. **Compared three repositories and branches.**
- FastVideo public: `/home/william5lin/FastVideo`, branch
`will/ltx2_sr_port`.
- Dreamverse: `/home/william5lin/Dreamverse`, branch
`will/integrate-public-fastvideo`.
- FastVideo-internal: `/home/william5lin/FastVideo-internal`, branch
`will/rebase-nbv`.
- Canonical repo paths are listed in the integration memory index:
[repo paths](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/README.md#L72-L79).
2. **Scoped the audit to the ultimate integration goal.**
- Dreamverse should depend on public `fastvideo`, not
`FastVideo-internal`.
- FastVideo should own the reusable backend subset that Dreamverse
currently needs from internal: streaming runtime, GPU pool, router,
prompt enhancer, NVFP4, continuation state, and typed generation.
- Dynamo should consume FastVideo through typed public Python APIs, not
through private modules.
- The three Dreamverse surfaces are documented as pipeline construction,
realtime runtime, and continuation state:
[cross-repo surfaces](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L13-L20).
3. **Separated intentional refactor from drift.**
- A path rename is not drift if the public branch contains the same
responsibility under the typed design.
- A deleted file is not drift if the public design intentionally
consolidated it.
- A private alias is not drift if the public schema exposes a typed
replacement with contract tests.
- This matches the typed-public-boundary rule in
[design.md](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L24-L43).
4. **Used memory docs for rationale and worktree files for concrete proof.**
- API schema and public exports:
[schema](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L68-L85),
[api exports](file:///home/william5lin/FastVideo/fastvideo/api/__init__.py#L49-L109).
- Streaming server current routes:
[build_app](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/server.py#L88-L160).
- Dreamverse dependency state:
[pyproject server extra](file:///home/william5lin/Dreamverse/pyproject.toml#L17-L22),
[uv lock editable source](file:///home/william5lin/Dreamverse/uv.lock#L716-L722).
- Contract tests:
[Dreamverse shape](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dreamverse_shape.py#L1-L26),
[Dynamo shape](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dynamo_shape.py#L1-L19),
[generate_async](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_generate_async.py#L1-L7).
5. **Did not treat product-only Dreamverse behavior as FastVideo drift.**
- Dreamverse keeps a local product server and Next.js UI today:
[README baseline](file:///home/william5lin/Dreamverse/README.md#L5-L16).
- Product-only routes, curated presets, devtools, and frontend-specific
behavior belong in Dreamverse unless a second non-Dreamverse consumer
needs them.
6. **Risk scale used below.**
- **P0:** blocks Dreamverse from running without FastVideo-internal.
- **P1:** blocks clean `BE_FLAVOR=fastvideo` or Dynamo/public API use.
- **P2:** reproducibility or maintenance drag.
- **P3:** optional parity or future memory/perf improvement.
### Findings: zero core typed API drift
The public branch is aligned with the goal on the **core typed API
surface** (construction, request, continuation state, async events).
The table below lists items that look like drift only if compared by
path name or legacy field name. They are intentional public refactors
or already guarded by tests. **Note:** the realtime-runtime _contract_
surface (FE-required health routes) is a separate matter — see "real
drift items" §4 below.
| Investigated item | Drift? | Evidence | Conclusion |
|---|---:|---|---|
| Dreamverse surface 1: pipeline construction | No | Dreamverse migrated from flat kwargs to typed `GeneratorConfig` at `d80c2a8`; mapping documented in [cross-repo surfaces](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L22-L47). | Stable public surface exists. |
| Dreamverse surface 2: realtime runtime | No on architecture; some route work remains | Runtime migration target is public `streaming/`, not internal `realtime/`: [streaming upstream](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/streaming-server.md#L14-L31). | Rename/refactor is intentional. |
| Dreamverse surface 3: continuation state | No | Public typed `ContinuationState` plus LTX-2 state mapping are documented in [cross-repo surfaces](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L108-L146). | Public state is a superset of Dreamverse's data carrier. |
| Internal `fastvideo/entrypoints/realtime/` | No | Public design chooses parallel `fastvideo/entrypoints/streaming/`: [layout decision](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/streaming-server.md#L55-L60), [current build_app](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/server.py#L88-L160). | Intentional rename plus typed-config rewrite. |
| Internal `configs/sample/` presets | No | Public PR 4 intentionally deleted `configs/sample/` and moved defaults to per-family presets plus `fastvideo/api/sampling_param.py`: [design](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L194-L205), [PR roadmap](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/pr-roadmap.md#L21-L29). | Intentional consolidation. |
| LTX-2 pipeline presets | No | Public target is model-owned named presets and per-family colocation: [design](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L175-L205). | Public layout matches design. |
| Internal `use_fp4_linear` flag | No | Public typed quant carrier is `engine.quantization.transformer_quant`; schema field exists in [schema](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L68-L85), compat resolves it in [compat.py](file:///home/william5lin/FastVideo/fastvideo/api/compat.py#L267-L279). | Replaced by typed NVFP4 surface. |
| Public-only `transformer_quant` field | No | Public `FastVideoArgs` pins typed quant to `dit_config.quant_config`: [fastvideo_args](file:///home/william5lin/FastVideo/fastvideo/fastvideo_args.py#L220-L228), [apply logic](file:///home/william5lin/FastVideo/fastvideo/fastvideo_args.py#L260-L279). | Public superset, not drift. |
| Internal `config_model_path` | No | Public typed home is `generator.pipeline.components.config_root`: [design mapping](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L261-L269), [compat mapping](file:///home/william5lin/FastVideo/fastvideo/api/compat.py#L295-L299). | Alias is covered. |
| Internal flat video request fields | No | Public `GenerationRequest` nests `inputs`, `sampling`, `runtime`, `output`, `state`, `extensions`: [schema](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L193-L204). Internal legacy fields live in internal protocol at [protocol.py](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/openai/protocol.py#L64-L82). | Intentional request refactor. |
| Dreamverse typed init kwargs | No | Contract test asserts current Dreamverse load kwargs all land on typed fields, not `experimental`: [test_dreamverse_shape](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dreamverse_shape.py#L44-L135). | Guard in place. |
| Dreamverse request path | No | Contract test asserts request fields round-trip through typed `GenerationRequest`: [test_dreamverse_shape](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dreamverse_shape.py#L153-L197). | Guard in place. |
| Dynamo native backend shape | No | FastVideo's only obligation is stable typed Python API: [cross-repo contract](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L188-L210). | Dynamo should stay out of FastVideo. |
| `generate_async` event API | No | API exists in [video_generator](file:///home/william5lin/FastVideo/fastvideo/entrypoints/video_generator.py#L264-L332), event types exist in [results.py](file:///home/william5lin/FastVideo/fastvideo/api/results.py#L109-L164). | #1288 covers the async contract. |
| Dynamo request mapping | No | Authoritative source is the contract test [test_dynamo_shape](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dynamo_shape.py#L90-L175) which asserts `req.prompt`, `req.sampling.{height,width,num_frames,fps,num_inference_steps,guidance_scale,seed,negative_prompt}`, and `req.inputs.{image_path,video_path}` against the actual nested [`GenerationRequest` schema](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L193-L204). The `streaming-server.md` Dynamo mapping table mis-cites a `prompt -> sampling.prompt` path that no longer exists; the test is correct, the doc is stale and tracked for refresh. | Guard in place; companion doc needs minor refresh. |
| Public API exports | No | `VideoEvent`, `VideoResult`, and typed schema classes are exported from [fastvideo.api](file:///home/william5lin/FastVideo/fastvideo/api/__init__.py#L49-L109). | Integration imports resolve. |
| FastVideo-internal FP4/NVFP4 paths | No | Public NVFP4 files and roles are documented in [quantization](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/quantization.md#L24-L35); actual `NVFP4Config` documents lazy FlashInfer and public naming in [nvfp4_config.py](file:///home/william5lin/FastVideo/fastvideo/layers/quantization/nvfp4_config.py#L1-L19). | Public is typed superset. |
| AbsMaxFP8 refactor | No for Dreamverse | Public quant registry includes `AbsMaxFP8` and `NVFP4`: [quantization init](file:///home/william5lin/FastVideo/fastvideo/layers/quantization/__init__.py#L1-L8). AbsMaxFP8 failure is tracked as separate tech debt: [open threads](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L100-L117). | Not Dreamverse blocker. |
| Internal realtime API regression test | No | Public contract tests replace it: [Dreamverse contract](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dreamverse_shape.py#L1-L26), [Dynamo contract](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dynamo_shape.py#L1-L19), [generate_async tests](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_generate_async.py#L91-L230). | Better scoped guards exist. |
| Dreamverse dependency declaration | No | `server` extra declares `fastvideo>=0.1.7`: [pyproject](file:///home/william5lin/Dreamverse/pyproject.toml#L17-L22). Dev lock resolves editable public `../FastVideo`: [uv.lock](file:///home/william5lin/Dreamverse/uv.lock#L716-L722), [package source](file:///home/william5lin/Dreamverse/uv.lock#L777-L780). | Dependency is already switched in metadata/lock. |
#### Core conclusion for the zero-typed-drift section
The public typed API no longer needs to mirror `FastVideo-internal` file
paths. The correct test is whether Dreamverse and Dynamo can express their
needs through public typed objects and public entrypoints. On that test,
the **typed core** is covered (construction, request, continuation,
async events). The **runtime contract** still has health-route gaps —
see real drift §4. On the **typed core**:
- `GeneratorConfig` and `GenerationRequest` cover construction and calls:
[schema surface](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L45-L72).
- `ServeConfig.streaming` covers the server envelope:
[schema](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L244-L279).
- `generate_async` covers streaming, OpenAI, and Dynamo on one substrate:
[video_generator](file:///home/william5lin/FastVideo/fastvideo/entrypoints/video_generator.py#L264-L332).
- Contract tests now encode the cross-repo shapes:
[Dreamverse](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dreamverse_shape.py#L70-L214),
[Dynamo](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dynamo_shape.py#L170-L331),
[async events](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_generate_async.py#L91-L273).
### Findings: real drift items requiring action
#### 1. Dreamverse README and bootstrap script still point at FastVideo-internal
- **Priority:** P0 for a clean public-dependency story.
- **Effort:** Small.
- **Owner:** Dreamverse repo.
- **Evidence:** Dreamverse metadata already points at public FastVideo:
[pyproject](file:///home/william5lin/Dreamverse/pyproject.toml#L17-L22),
[uv source](file:///home/william5lin/Dreamverse/pyproject.toml#L54-L61),
[uv.lock](file:///home/william5lin/Dreamverse/uv.lock#L716-L722).
- **Drift:** README still tells users that `uv` resolves from
`../FastVideo-internal` and that bootstrap expects `../FastVideo-internal`:
[README](file:///home/william5lin/Dreamverse/README.md#L76-L109).
- **Drift:** bootstrap script still defaults to cloning the private repo and
verifying imports from that clone:
[script defaults](file:///home/william5lin/Dreamverse/.agents/skills/bootstrap-fastvideo-private-fork/scripts/bootstrap_fastvideo_private.sh#L7-L11),
[script clone flow](file:///home/william5lin/Dreamverse/.agents/skills/bootstrap-fastvideo-private-fork/scripts/bootstrap_fastvideo_private.sh#L33-L63),
[script import assertion](file:///home/william5lin/Dreamverse/.agents/skills/bootstrap-fastvideo-private-fork/scripts/bootstrap_fastvideo_private.sh#L66-L88).
- **Action:** Replace private-fork bootstrap with public FastVideo bootstrap
or delete the bootstrap once PyPI publication is the default path.
- **Do not overreach:** no FastVideo code change required.
#### 2. Dreamverse carries a 1933-line prompt-enhancer fork
- **Priority:** P1.
- **Effort:** Medium.
- **Owner:** Dreamverse repo, after public prompt enhancer is available.
- **Evidence:** Dreamverse local fork starts at
[server/prompt_enhancer.py](file:///home/william5lin/Dreamverse/server/prompt_enhancer.py#L1-L80).
- **Public replacement:** FastVideo now has provider-agnostic
`PromptEnhancer` with `enhance`, `auto_extend`, `rewrite`, and
`register_provider`:
[public enhancer](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/prompt/enhancer.py#L66-L142).
- **Provider extension point:** custom providers implement `LLMProvider`:
[provider protocol](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/prompt/providers/base.py#L63-L75).
- **Tracking:** DR-1 in open threads already defines the compat-shim shape:
[DR-1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L174-L206).
- **Action:** Replace the fork with a small Dreamverse shim that adapts
public `LLMResponse` to Dreamverse's product response objects and keeps
only product-only extras.
- **Do not overreach:** do not merge Dreamverse's full prompt product layer
into FastVideo unless a second consumer needs the same semantics.
#### 3. `cerebras_ifm` provider is unresolved
- **Priority:** P1 if Dreamverse needs IFM in production; P2 otherwise.
- **Effort:** Small decision plus small/medium implementation.
- **Owner:** Team decision; implementation either Dreamverse-side or public.
- **Public state:** `PromptEnhancerConfig.provider` is currently
`Literal["cerebras", "groq"]`:
[schema](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L229-L235).
- **Design note:** public Literal excludes `cerebras_ifm` today:
[streaming-server D-3](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/streaming-server.md#L61-L100).
- **Tracking:** DR-2 already frames the public-vs-Dreamverse decision:
[DR-2](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L211-L229).
- **Recommended default:** implement IFM as a Dreamverse-side custom provider
registered through `enhancer.register_provider(...)` unless there is a
non-Dreamverse public user.
#### 4. `/healthz`, `/readyz`, and `/status` are not in public `build_app`
- **Priority:** P1 for `BE_FLAVOR=fastvideo` frontend compatibility.
- **Effort:** Medium/Large because route shapes need tests.
- **Owner:** FastVideo public.
- **Public current state:** `build_app` exposes `GET /health` and
`WS /v1/stream`:
[server.py](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/server.py#L126-L160).
- **Dreamverse expected state:** Dreamverse exposes `GET /healthz`,
`GET /readyz`, and `GET /status`:
[routes/health.py](file:///home/william5lin/Dreamverse/server/routes/health.py#L34-L79).
- **Tracking:** open item #1 documents route ownership and files likely to
touch:
[open threads](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L69-L99).
- **Design note:** `/curated-presets`, `/prompt-system-config`, and devtools
stay Dreamverse-side, with feature detection:
[streaming route contract](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/streaming-server.md#L232-L259).
#### 5. `fastvideo/models/layerwise_offload.py` exists only internally
- **Priority:** P3 unless memory-tight Dreamverse deployments require it.
- **Effort:** Medium if adopted; low if documented as deferred.
- **Owner:** FastVideo public only if a concrete deployment needs it.
- **Internal evidence:** internal file defines async layerwise CPU offload
manager with pinned CPU memory and prefetch stream:
[layerwise_offload.py](file:///home/william5lin/FastVideo-internal/fastvideo/models/layerwise_offload.py#L1-L20),
[prefetch path](file:///home/william5lin/FastVideo-internal/fastvideo/models/layerwise_offload.py#L127-L180).
- **Public state:** no equivalent public file was identified in this audit.
- **Action:** defer unless Dreamverse or another public deployment hits a
memory ceiling that cannot be handled by existing offload knobs.
- **Decision rule:** if adopted, port as a generic offload utility with
tests; do not make it Dreamverse-specific.
#### 6. Standalone LTX-2 upsampler CLI exists only internally
- **Priority:** P2 for reproducibility; P3 for product runtime.
- **Effort:** Small/Medium after scope decision.
- **Owner:** FastVideo public if standalone upsampling is a supported user
workflow.
- **Internal utility:** `upscale_video_file(...)` reads an existing video,
prepares frame count/resolution, loads VAE + upsampler, and writes an mp4:
[upsample.py](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/upsample.py#L120-L180),
[write tail](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/upsample.py#L181-L202).
- **Internal CLI:** `fastvideo upsample` wrapper exists internally:
[cli/upsample.py](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/cli/upsample.py#L15-L35),
[CLI args](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/cli/upsample.py#L48-L130).
- **Public related functionality:** LTX-2 SR refine stage covers the
in-pipeline latent upsample/refine path:
[ltx2_refine.py](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L1-L22),
[upsample stage](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L116-L180).
- **Action:** decide whether standalone file-to-file upsampling is a public
CLI promise or whether the SR refine stage is sufficient.
#### 7. Reproducible streaming demo config lives only in Dreamverse
- **Priority:** P2.
- **Effort:** Small.
- **Owner:** FastVideo public.
- **Evidence:** canonical demo config currently lives at
[Dreamverse/serve_configs/streaming_demo.yaml](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml#L1-L12).
- **Config content:** it documents LTX-2 distilled model, one GPU,
no offload, compile settings, NVFP4, refine overrides, default request,
and streaming settings:
[generator block](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml#L31-L87),
[streaming block](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml#L108-L149).
- **Memory pointer:** design.md already treats this as the canonical
example:
[design YAML example](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L235-L239).
- **Action:** copy/adapt it into
`examples/serving/streaming_demo.yaml` with public-safe comments.
#### 8. LTX-2 stage equivalence is a verification gap, not proven drift
- **Priority:** P2.
- **Effort:** Medium if parity checks are added; small if only manual audit.
- **Owner:** FastVideo public.
- **Public state:** model-specific LTX-2 stages are colocated under
`fastvideo/pipelines/basic/ltx2/stages/`, consistent with the target
layout in [design.md](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/design.md#L175-L205).
- **Example public stage:** `ltx2_refine.py` explicitly says it is a
public-side port of the internal stage and describes the three-stage SR
flow:
[ltx2_refine.py](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L1-L22).
- **Action:** verify behavior for the six internal `ltx2_*` stage files
against public colocated stages. If a mismatch is found, file it as a
real drift item with a failing parity test.
### Findings: deferred / accepted residual
These items should not block the public-dependency transition.
1. **StepVideo residual.**
- Dreamverse's model registry is LTX-2/LTX-2.3 only:
[Dreamverse config](file:///home/william5lin/Dreamverse/server/config.py#L28-L45).
- Internal local tests even stub StepVideo modules to keep LTX registry
tests focused:
[test_ltx2_registry.py](file:///home/william5lin/FastVideo-internal/tests/local_tests/test_ltx2_registry.py#L38-L61).
- Conclusion: accepted low-priority deferral unless Dreamverse adds a
StepVideo model.
2. **Internal debug-only `FastVideoArgs` fields.**
- Internal debug fields exist around `FastVideoArgs` and stage/model sums:
[internal grep source](file:///home/william5lin/FastVideo-internal/fastvideo/fastvideo_args.py#L200-L203).
- They are debug-only and not a public user-facing integration surface.
- Conclusion: low-priority; do not add to public schema unless a debug
workflow requires them.
3. **Private request aliases.**
- Public request schema is nested and strict:
[GenerationRequest](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L193-L204).
- Legacy OpenAI flat fields are compatibility input, not the canonical
public API:
[internal protocol](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/openai/protocol.py#L64-L82).
- Conclusion: no action beyond current compat tests.
4. **`experimental["pipeline_config"]` escape hatch.**
- Dreamverse currently uses an explicit in-memory quant config because
typed `transformer_quant: "NVFP4"` does not expose `layer_profile`:
[quantization](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/quantization.md#L86-L97).
- Open thread #4 tracks `layer_profile`:
[open threads](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L244-L260).
- Conclusion: defer broader typed carrier design; add `layer_profile`
first if Dreamverse needs base/refine profile selection.
5. **Router sticky routing and active-active semantics.**
- Public router intentionally ships active-passive first and defers
sticky/weighted routing:
[D-15](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L118-L155).
- Follow-ups are tracked:
[D-15 action items](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L175-L201).
- Conclusion: not drift; defer until load-balancing needs are real.
6. **AbsMaxFP8 failure.**
- Pre-existing and not introduced by NVFP4:
[state](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/state.md#L154-L159),
[quantization](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/quantization.md#L202-L214).
- Conclusion: fix separately; not a Dreamverse public-dependency blocker.
### Drift summary table
| # | Item | Priority | Effort | Status | Tracked where | Next action |
|---:|---|---|---|---|---|---|
| 1 | Dreamverse README still names `../FastVideo-internal` | P0 | S | Real drift | [README lines](file:///home/william5lin/Dreamverse/README.md#L76-L109) | Update docs to public FastVideo / PyPI path. |
| 2 | Dreamverse private bootstrap clones internal repo | P0 | S | Real drift | [bootstrap script](file:///home/william5lin/Dreamverse/.agents/skills/bootstrap-fastvideo-private-fork/scripts/bootstrap_fastvideo_private.sh#L7-L11) | Replace or delete private bootstrap. |
| 3 | Dreamverse `prompt_enhancer.py` fork | P1 | M | Real drift | [DR-1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L174-L206) | Build compat shim over public enhancer. |
| 4 | `cerebras_ifm` provider path | P1/P2 | S-M | Real drift / decision | [DR-2](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L211-L229) | Choose public provider vs Dreamverse custom provider. |
| 5 | Health route mismatch | P1 | M-L | Real drift | [open item #1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L69-L99) | Add `/healthz`, `/readyz`, `/status` to public build_app. |
| 6 | Missing public streaming demo config | P2 | S | Real drift | [Dreamverse config](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml#L1-L12) | Add `examples/serving/streaming_demo.yaml`. |
| 7 | Standalone upsampler CLI | P2/P3 | S-M | Real drift if standalone CLI is desired | [internal CLI](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/cli/upsample.py#L15-L35) | Decide CLI promise; port or defer. |
| 8 | Layerwise offload utility | P3 | M | Optional internal-only residual (no Dreamverse deployment requires it today) | [internal manager](file:///home/william5lin/FastVideo-internal/fastvideo/models/layerwise_offload.py#L15-L20) | Defer until memory-tight deployment needs it. |
| 9 | LTX-2 stage equivalence | P2 | S-M | Verification gap | [public refine stage](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L1-L22) | Add targeted parity audit/test if needed. |
| 10 | StepVideo | P3 | M | Accepted residual | [Dreamverse model registry](file:///home/william5lin/Dreamverse/server/config.py#L28-L45) | No action unless Dreamverse adds StepVideo. |
| 11 | Debug-only fields | P3 | S | Accepted residual | [internal args](file:///home/william5lin/FastVideo-internal/fastvideo/fastvideo_args.py#L200-L203) | Do not publicize unless needed. |
| 12 | `layer_profile` typed quant knob | P2 | M | Tracked gap | [open item #4](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L244-L260) | Add typed layer profile if Dreamverse drops escape hatch. |
| 13 | `ltx2_image_crf` per-segment field flow (D-8) | P1 | S | Open verification gap — Dreamverse still passes `ltx2_image_crf=0.0` per [Dreamverse video_generation.py](file:///home/william5lin/Dreamverse/server/video_generation.py#L420-L435); needs trace-through to confirm it lands on `request.stage_overrides.refine.image_crf` rather than being silently dropped | [D-8 in open-threads](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L52-L68) | 10-min trace + add a Dreamverse-shape contract test pinning the field. |
| 14 | `video_position_offset_sec` semantics (VPO) | P1 | S | Open decision — persistent-vs-per-segment ambiguity unresolved | [VPO in open-threads](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L120-L144) | Confirm semantics with audio team; document + add test. Decision deadline was "before PR 7.6 emits state" — that PR (7.6 / #1257) is now MERGED, so the decision is overdue. |
| 15 | `GpuPool` ABC docstring missing experimental caveat (D-12-A) | P3 | trivial | Tracked gap — `GpuPool` ABC at [gpu_pool.py:74-83](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/gpu_pool.py#L74-L83) lacks the "API may change post-PR-7.10; experimental / server-internal" caveat | [D-12-A](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L301-L313) | Edit docstring; trivial. |
| 16 | `GpuPool.run_async()` migration (D-12-B) | P2 | M | Tracked gap — `GpuPool.run() -> Any` should become `run_async() -> AsyncIterator[VideoEvent]` per D-12 / D-12-B | [D-12-B](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L317-L327) | Land alongside #1288 merge or in immediate follow-up. |
| 17 | `SessionStore` / `BlobStore` lifecycle policy (SBS) | P2 | M | Tracked gap — in-memory defaults have no eviction/TTL/blob-cleanup policy | [SBS](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L278-L294) | Streaming server design pass needed before high-traffic deployment. |
| 13 | Router sticky / active-active | P3 | M | Deferred | [D-15](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L175-L201) | Defer until reconnect/load evidence. |
| 14 | AbsMaxFP8 test failure | P2 | S | Separate tech debt | [state](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/state.md#L154-L159) | Fix outside Dreamverse migration. |
| 15 | Dynamo backend package | P1 | External | Not FastVideo drift | [Dynamo contract](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L188-L210) | Reopen Dynamo-side PR after public API lands. |
---
## Part 2 — Integration path tradeoffs
### The four options
#### Option A — Status quo: Dreamverse stays separate and depends on `fastvideo`
**Shape**
- FastVideo remains the Python library and reusable backend runtime.
- Dreamverse remains the product repo with FastAPI product glue and Next.js
frontend.
- Dreamverse `server` extra depends on `fastvideo>=0.1.7`:
[pyproject](file:///home/william5lin/Dreamverse/pyproject.toml#L17-L22).
- Local development can keep using editable `../FastVideo` until PyPI
publication catches up:
[uv.lock](file:///home/william5lin/Dreamverse/uv.lock#L716-L722).
**What it solves**
- Directly satisfies "Dreamverse depends on public FastVideo".
- Keeps frontend release cadence independent.
- Keeps product-specific prompts, routes, and UI in the product repo.
- Minimizes FastVideo packaging and CI growth.
**What it does not solve by itself**
- Does not remove Dreamverse prompt-enhancer fork unless DR-1 is executed.
- Does not give Dreamverse FE compatibility with public `build_app` until
health routes migrate.
- Does not make Dreamverse server itself reusable as a public entrypoint.
**Best fit**
- Default for the next release if the goal is to stop using
FastVideo-internal quickly and safely.
#### Option B — Dreamverse as a subfolder under FastVideo
**Shape**
- One repository: FastVideo contains `dreamverse/server/` and
`dreamverse/apps/web/`.
- Dreamverse can remain a separate package in the same repo, or FastVideo's
build can ignore Dreamverse by default.
- CI must understand Python library tests plus Next.js install/build/test.
**What it solves**
- Eliminates sibling-checkout drift.
- Makes cross-repo integration changes atomic.
- Easier for a single reviewer to see library and product changes together.
**Costs**
- Adds frontend dependency management to a Python ML library repo.
- Couples clone size, CI setup, issue tracking, and review load.
- Forces maintainers to decide whether product assets are included in source
distributions, wheels, docs, and release notes.
**Best fit**
- Only if Dreamverse becomes the primary FastVideo product surface and the
team accepts a product monorepo.
#### Option C — Full merge into `fastvideo.entrypoints.dreamverse.*`
**Shape**
- Dreamverse backend becomes FastVideo code.
- Public import becomes something like
`from fastvideo.entrypoints.dreamverse import build_app`.
- CLI becomes `fastvideo dreamverse-serve --config dreamverse.yaml`.
- Frontend either ships as static assets in the package or as a frontend
extra.
**What it solves**
- One namespace and one release train for library plus product backend.
- No dependency boundary between Dreamverse server and FastVideo internals.
- Product route contract can be tested entirely inside FastVideo CI.
**Costs**
- Maximally expands FastVideo's public/security surface.
- Locks product experiments to FastVideo release cadence.
- Makes private prompt/provider/product assumptions look like framework API.
- Has weak precedent for a Python ML library plus Next.js product being merged
into the library namespace.
**Best fit**
- Only if Dreamverse is no longer a separate product and becomes the
canonical FastVideo UI/serving mode.
#### Option D — Hybrid: backend merges, frontend stays separate
**Shape**
- Reusable backend components merge into public FastVideo.
- Frontend stays in a separate Dreamverse UI repo or Dreamverse product repo.
- The backend should be generic where possible: `fastvideo.entrypoints.streaming`,
not product-only names, unless product-only routes are intentionally
accepted as public API.
- This matches the current trajectory: streaming server, GPU pool, prompt
enhancer, safety/rewrite/session logging, router, NVFP4, and
`generate_async` are public-side work already tracked in the PR roadmap:
[pr-roadmap](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/pr-roadmap.md#L19-L42).
**What it solves**
- Removes FastVideo-internal dependency for reusable backend pieces.
- Keeps product frontend cadence independent.
- Gives non-Dreamverse users a streaming backend and typed API without
carrying the Dreamverse app.
- Gives Dynamo a stable library API while leaving Dynamo package code in
Dynamo:
[Dynamo contract](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L188-L210).
**Costs**
- Requires careful boundary discipline: generic streaming/server code in
FastVideo; product routes/prompts/presets in Dreamverse.
- Requires contract tests to prevent drift.
- Some Dreamverse compatibility routes may become public and need support.
**Best fit**
- Best long-term target if the team wants FastVideo to own serving/runtime
infrastructure while keeping Dreamverse as a separately evolving product.
### Comparison matrix
| Criterion | A. Separate dep | B. Subfolder monorepo | C. Full namespace merge | D. Hybrid backend merge |
|---|---|---|---|---|
| Alignment with stated goal | High: Dreamverse depends on public package | Medium: no external dep, but product becomes repo-local | Medium: dependency disappears by absorption | High: reusable backend in public, product separate |
| Time to remove `FastVideo-internal` | Fastest | Medium | Slowest | Medium-fast |
| Build complexity | Low | High: Python + Next.js in one repo | High: Python package plus static/frontend extras | Medium: Python backend only in FastVideo |
| Release cadence | Independent | Coupled clone; releases can still be separate but more friction | Fully coupled | Backend coupled to FastVideo, frontend independent |
| Security surface in FastVideo | Low | Medium/High | Highest | Medium |
| Contributor friction | Low for both repos | Higher for library contributors | Highest; product assumptions in library | Medium; clear backend boundary needed |
| Dependency management | Normal package pin | Workspace/monorepo tooling needed | FastVideo extras/static asset decisions needed | FastVideo extras for backend; FE out-of-tree |
| CI cost | Low/medium | High | High | Medium |
| Contract-test value | High; cross-repo contract tests are essential | Medium; same repo but still useful | Medium; less boundary pressure | High; generic backend vs product boundary |
| Precedent strength | Strong: library/server plus external UI patterns exist | Mixed | Weak for Python ML library + Next.js inside namespace | Strongest match: in-tree server/backend, external UI |
| Packaging risk | Low | Medium/high | High | Medium |
| Future Dynamo fit | Strong | Strong if API remains clean | Risky if product API bleeds in | Strong |
| Frontend iteration speed | Highest | Lower | Lowest | Highest |
| Risk of product-specific API leakage | Low | Medium | High | Medium; controllable with naming discipline |
| Reversibility | High | Medium | Low | Medium/high |
### OSS precedents (with citations)
| Pattern | Project | What it supports | Citation |
|---|---|---|---|
| Library plus in-tree server | vLLM | A Python ML library can ship an in-tree OpenAI-compatible server while clients remain external. | https://github.com/vllm-project/vllm/blob/bcf5cac9fb956788f649d1f5297b74c886a9d6d3/README.md#L64-L74 |
| Service packaging | BentoML | Packaging model + service + dependencies is supported, but CWD packaging creates discipline needs. | https://github.com/bentoml/BentoML/blob/32230a5276a8da8b23c4a06a9ec6272c1993451a/docs/source/build-with-bentoml/asgi.rst#L5-L18 |
| YAML-driven production serving | Ray Serve | Production updates should avoid in-place mutation; use new deployment/traffic switch. | https://docs.ray.io/en/latest/serve/advanced-guides/inplace-updates.html |
| Library/server plus external UI | TGI + ChatUI | Server can live with backend project while UI is separate. | https://github.com/huggingface/text-generation-inference/blob/b4adbf2f6e2e721280bd0ea5f91d70f7d033f5ed/docs/source/basic_tutorials/consuming_tgi.md#L182-L186 |
| Lean library plus examples elsewhere | Transformers.js | Library stays lean; demos/examples can live outside core. | https://github.com/huggingface/transformers.js/blob/f7487c737aa8cafbc106c9adf69dc9578c8f3fe0/README.md#L26-L34 |
| Product monorepo that later split frontend | ComfyUI | Product UI/server monorepo can hit release-cadence mismatch and split FE later. | https://github.com/comfyanonymous/ComfyUI/blob/fed8d5efa6b70d5b24c4c33cb643bfccc39d45b5/README.md#L131-L149 and https://github.com/Comfy-Org/ComfyUI_frontend/blob/60f789d58070a9d1d789b260f83c36d7293a39f0/README.md#L31-L60 |
| Tightly coupled UI/server product | AUTOMATIC1111 SD WebUI | Product repos can couple UI/server tightly, but security surface becomes product-sized. | https://github.com/AUTOMATIC1111/stable-diffusion-webui/blob/82a973c04367123ae98bd9abdf80d9eda9b910e2/webui.py#L48-L104 |
#### Precedent synthesis
- Strong precedents exist for a Python ML library shipping a server entrypoint.
- Strong precedents exist for keeping frontend/product UI out of the backend
library repo.
- The cited set does not contain a clean precedent for merging a Next.js
product into a Python ML library namespace.
- The most applicable pattern is **backend/server in the ML project,
product UI outside**.
### Recommendation
#### Recommend Option D, constrained: backend merges as generic FastVideo streaming; frontend stays separate
Recommendation: follow **Option D** as the long-term architecture, but keep
the backend merge generic. In practice, this means continuing the current
public FastVideo path:
- `fastvideo.entrypoints.streaming.*` owns reusable streaming runtime.
- `fastvideo.entrypoints.streaming.gpu_pool` owns generic GPU worker pools.
- `fastvideo.entrypoints.streaming.prompt.*` owns provider-agnostic prompt
operations.
- `fastvideo.entrypoints.streaming.router.*` owns FastVideo-aware routing.
- `fastvideo.api` owns typed construction, requests, results, events, and
continuation state.
- Dreamverse keeps product-only FE, curated presets, prompt UX, product
routes, and launch scripts.
This is effectively the path already underway in PRs #1257, #1258, #1284,
#1286, and #1288:
[PR roadmap](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/pr-roadmap.md#L21-L42).
#### Why not Option A as the final answer?
Option A is the fastest near-term release posture and should be used as the
immediate migration posture. However, plain status quo is not enough for
the ultimate goal because reusable backend pieces still need to live in
public FastVideo so Dreamverse can stop reaching into internal code. That
work is already partly complete:
- GPU pool: [D-12](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L47-L116).
- Prompt enhancer: [PR roadmap 7.7](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/pr-roadmap.md#L32-L35).
- Streaming auxiliaries: [PR roadmap 7.8](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/pr-roadmap.md#L35-L36).
- Router: [D-15](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L118-L155).
- `generate_async`: [streaming-server unlock PR](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/streaming-server.md#L314-L345).
So the practical answer is:
- **Near term:** Option A operationally, after docs/scripts are fixed.
- **Architecture target:** Option D, with generic backend ownership in
FastVideo and product ownership in Dreamverse.
#### Why not Option B?
Option B makes cross-repo coordination easier but imports frontend build,
package, and CI complexity into FastVideo. That is unnecessary while a
normal package dependency plus contract tests can guard the integration.
FastVideo's current repo structure is a Python package with examples and
docs, not a product monorepo:
[codebase map](file:///home/william5lin/FastVideo/.agents/memory/codebase-map/README.md#L5-L75).
#### Why not Option C?
Option C makes the product backend a public FastVideo namespace. That is
only appropriate if the team wants to support Dreamverse as a first-class
FastVideo product surface. Today the known public obligations are generic:
typed requests, streaming server, GPU pool, prompt provider protocol,
router, NVFP4, and Dynamo event APIs. Product-only Dreamverse behavior does
not need to become framework API.
#### Conditions that would change the recommendation
Move from constrained D toward **C** only if all of these become true:
1. Dreamverse is declared the canonical FastVideo serving product.
2. Product routes such as curated presets and prompt-system config are
accepted as public FastVideo API.
3. FastVideo maintainers accept the security and support surface.
4. Release cadence for product UX and FastVideo core is intentionally
coupled.
5. Frontend packaging/static asset strategy is explicitly owned by
FastVideo.
Move from constrained D back toward **A** if any of these become true:
1. Prompt enhancement, router, or GPU pool turn out to be Dreamverse-only.
2. No second user appears for the streaming backend outside Dreamverse.
3. FastVideo maintainers want to minimize serving surface and publish only
Python library APIs.
4. Dreamverse needs product changes faster than FastVideo can release.
5. Security review rejects in-tree serving/router responsibilities.
### Migration sketch for the recommended path
#### Phase 0 — Land the public backend stack
- **Effort:** Large, already in flight.
- **Owner:** FastVideo public.
- **Files:** #1288 scope, especially `fastvideo/api/`,
`fastvideo/entrypoints/video_generator.py`,
`fastvideo/entrypoints/streaming/`, LTX-2 pipeline stages, NVFP4 files,
and contract tests.
- **Exit criteria:** #1288 merges; public `fastvideo.api.VideoEvent` and
`VideoGenerator.generate_async` are available:
[results.py](file:///home/william5lin/FastVideo/fastvideo/api/results.py#L109-L164),
[video_generator.py](file:///home/william5lin/FastVideo/fastvideo/entrypoints/video_generator.py#L264-L332).
#### Phase 1 — Fix Dreamverse dependency docs and bootstrap
- **Effort:** Small.
- **Owner:** Dreamverse.
- **Files:**
- [README.md](file:///home/william5lin/Dreamverse/README.md#L76-L109)
- [bootstrap script](file:///home/william5lin/Dreamverse/.agents/skills/bootstrap-fastvideo-private-fork/scripts/bootstrap_fastvideo_private.sh#L7-L11)
- [pyproject.toml](file:///home/william5lin/Dreamverse/pyproject.toml#L17-L22)
- [uv.lock](file:///home/william5lin/Dreamverse/uv.lock#L716-L722)
- **Exit criteria:** no user-facing docs or scripts mention
`FastVideo-internal` as the expected dependency path.
#### Phase 2 — Add public health/readiness/status route compatibility
- **Effort:** Medium/Large.
- **Owner:** FastVideo public.
- **Files likely to touch:**
- `fastvideo/entrypoints/streaming/server.py::build_app`
- new `fastvideo/entrypoints/streaming/health.py`
- tests under `fastvideo/tests/entrypoints/streaming/`
- **Source route shapes:**
[Dreamverse health routes](file:///home/william5lin/Dreamverse/server/routes/health.py#L34-L79).
- **Tracking:** [open item #1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L69-L99).
- **Exit criteria:** Dreamverse FE can target public `build_app` for
`/healthz`, `/readyz`, `/status`, and `/v1/stream`; product-only routes
remain feature-detected.
#### Phase 3 — Replace Dreamverse prompt enhancer fork
- **Effort:** Medium.
- **Owner:** Dreamverse.
- **Files likely to touch:**
- new `Dreamverse/server/prompting/_internal_compat.py`
- `Dreamverse/server/runtime.py`
- `Dreamverse/server/main.py`
- `Dreamverse/server/prompt_enhancer.py`
- **Public API:**
[PromptEnhancer](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/prompt/enhancer.py#L66-L142),
[LLMProvider](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/prompt/providers/base.py#L63-L75).
- **Tracking:** [DR-1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L174-L206).
- **Exit criteria:** Dreamverse no longer carries a full local fork for the
generic prompt operations public FastVideo already owns.
#### Phase 4 — Decide and implement `cerebras_ifm`
- **Effort:** Small decision plus small/medium implementation.
- **Owner:** Team decision, then Dreamverse or FastVideo.
- **Default recommendation:** Dreamverse-side custom provider.
- **Public schema source:**
[PromptEnhancerConfig](file:///home/william5lin/FastVideo/fastvideo/api/schema.py#L229-L235).
- **Tracking:** [DR-2](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L211-L229).
- **Exit criteria:** Dreamverse IFM provider works after prompt fork removal.
#### Phase 5 — Move streaming demo config into FastVideo examples
- **Effort:** Small.
- **Owner:** FastVideo public.
- **Source:**
[Dreamverse streaming_demo.yaml](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml#L1-L149).
- **Target:** `examples/serving/streaming_demo.yaml`.
- **Exit criteria:** users can reproduce the typed streaming path from the
FastVideo repo without checking out Dreamverse.
#### Phase 6 — Remove `experimental["pipeline_config"]` where practical
- **Effort:** Medium for `layer_profile`; Large for a full typed
`dit_config.quant_config` carrier.
- **Owner:** FastVideo public, then Dreamverse cleanup.
- **Tracking:**
[open item #4](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L244-L260),
[quantization follow-up](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/quantization.md#L175-L200).
- **Exit criteria:** Dreamverse can express its quant layer profile through
typed config instead of in-memory mutation.
#### Phase 7 — Decide standalone upsampler CLI
- **Effort:** Small/Medium.
- **Owner:** FastVideo public.
- **Input:** internal standalone utility
[upsample.py](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/upsample.py#L120-L180)
and internal CLI
[cli/upsample.py](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/cli/upsample.py#L48-L130).
- **Public alternative:** SR refine stage already covers in-pipeline latent
upsampling:
[ltx2_refine.py](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L116-L180).
- **Exit criteria:** explicit decision: port CLI, document refine-stage-only
support, or defer.
#### Phase 8 — Validate LTX-2 stage parity and offload residuals
- **Effort:** Small/Medium for stage parity; Medium for layerwise offload.
- **Owner:** FastVideo public.
- **Stage source:**
[public LTX-2 stages](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/).
- **Offload source:**
[internal layerwise offload](file:///home/william5lin/FastVideo-internal/fastvideo/models/layerwise_offload.py#L15-L20).
- **Exit criteria:** no known behavior gap between internal and public LTX-2
stages; offload is either deliberately deferred or ported with tests.
### Open questions
1. **Which provider path for `cerebras_ifm`?**
- Public provider or Dreamverse-side custom provider?
- Default recommendation: Dreamverse-side unless there is another user.
- Source: [DR-2](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L211-L229).
2. **Should public FastVideo support standalone LTX-2 file upsampling?**
- If yes, port internal CLI.
- If no, document that SR support is pipeline-refine only.
- Sources: [internal CLI](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/cli/upsample.py#L15-L35),
[public refine stage](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L1-L22).
3. **Does Dreamverse need layerwise CPU offload?**
- If memory-tight deployments require it, port as generic FastVideo.
- Otherwise defer.
- Source: [internal offload manager](file:///home/william5lin/FastVideo-internal/fastvideo/models/layerwise_offload.py#L15-L20).
4. **How much Dreamverse route surface should FastVideo own?**
- Health/readiness/status should migrate because they are part of
streaming-server compatibility.
- Curated presets and prompt-system config should stay Dreamverse-side
unless product policy changes.
- Source: [route contract](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/streaming-server.md#L232-L259).
5. **Should `layer_profile` be the only near-term quant typed addition?**
- Adding `layer_profile` is bounded.
- A typed carrier for arbitrary mutated `PipelineConfig` is larger design
work.
- Source: [quantization follow-ups](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/quantization.md#L175-L200).
6. **When does Option D become Option C?**
- Only if Dreamverse backend routes become public FastVideo product API.
- Until then, keep generic streaming code in FastVideo and product code in
Dreamverse.
---
## Part 3 — Action items
1. **P0 / S — Update Dreamverse README dependency notes.**
- Replace `../FastVideo-internal` with public FastVideo instructions.
- Preserve local editable `../FastVideo` dev flow where useful.
- Source: [README stale lines](file:///home/william5lin/Dreamverse/README.md#L76-L109).
2. **P0 / S — Replace or remove private FastVideo bootstrap script.**
- Current script clones `FastVideo-internal` and verifies imports from it.
- Source: [script](file:///home/william5lin/Dreamverse/.agents/skills/bootstrap-fastvideo-private-fork/scripts/bootstrap_fastvideo_private.sh#L7-L11).
3. **P1 / M-L — Add `/healthz`, `/readyz`, and `/status` to public `build_app`.**
- Keep `/curated-presets` and prompt-system config in Dreamverse.
- Source: [open item #1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L69-L99).
4. **P1 / M — Replace Dreamverse prompt-enhancer fork with compat shim.**
- Wrap public `PromptEnhancer`.
- Keep only Dreamverse-specific metadata and product fallback behavior.
- Source: [DR-1](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L174-L206).
5. **P1 / S-M — Decide `cerebras_ifm` provider path.**
- Default: Dreamverse custom provider via `register_provider`.
- Source: [DR-2](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L211-L229).
6. **P2 / S — Add `examples/serving/streaming_demo.yaml` to FastVideo.**
- Start from Dreamverse config and remove Dreamverse-private comments.
- Source: [streaming_demo.yaml](file:///home/william5lin/Dreamverse/serve_configs/streaming_demo.yaml#L1-L149).
7. **P2 / M — Verify each public LTX-2 colocated stage against internal behavior.**
- Start with refine, denoising, latent prep, image conditioning, text
encoding, and audio decoding.
- Source: [public refine stage](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/stages/ltx2_refine.py#L1-L22).
8. **P2 / M — Add typed `transformer_quant_layer_profile` if Dreamverse needs it.**
- Thread schema → compat → `FastVideoArgs._apply_transformer_quant`.
- Source: [open item #4](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L244-L260).
9. **P2 / S-M — Decide standalone LTX-2 upsampler CLI support.**
- Port internal CLI only if file-to-file upsampling is a public workflow.
- Source: [internal upsample CLI](file:///home/william5lin/FastVideo-internal/fastvideo/entrypoints/cli/upsample.py#L15-L35).
10. **P2 / S — Fix pre-existing AbsMaxFP8 test failure separately.**
- Do not block Dreamverse migration on it.
- Source: [open item #2](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/open-threads.md#L100-L117).
11. **P2 / S-M — Document public streaming install extras and dependencies.**
- Include router `websockets`, prompt enhancer provider SDKs, and optional
safety classifier extras.
- Source: [D-16 dependency note](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L224-L242).
12. **P2 / S — Keep contract tests in the FastVideo CI path.**
- Guard Dreamverse shape, Dynamo shape, and async events.
- Sources: [Dreamverse test](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dreamverse_shape.py#L1-L26),
[Dynamo test](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_dynamo_shape.py#L1-L19),
[async test](file:///home/william5lin/FastVideo/fastvideo/tests/contract/test_generate_async.py#L1-L7).
13. **P3 / M — Defer layerwise offload until a deployment needs it.**
- Port only as generic FastVideo utility with tests.
- Source: [internal offload](file:///home/william5lin/FastVideo-internal/fastvideo/models/layerwise_offload.py#L15-L20).
14. **P3 / M — Defer StepVideo public parity for this integration.**
- Dreamverse model registry is LTX-2/LTX-2.3 only.
- Source: [Dreamverse config](file:///home/william5lin/Dreamverse/server/config.py#L28-L45).
15. **P3 / S — Do not add debug-only fields to public schema by default.**
- Keep them private unless there is a user-facing debugging workflow.
- Source: [internal debug args](file:///home/william5lin/FastVideo-internal/fastvideo/fastvideo_args.py#L200-L203).
16. **P3 / M — Keep router active-active and sticky routing deferred.**
- Add only when session-routing evidence justifies it.
- Source: [D-15 action items](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L175-L201).
17. **P3 / S — Preserve Dynamo as an external backend package.**
- FastVideo should expose typed API; Dynamo code lives in Dynamo.
- Source: [Dynamo contract](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/cross-repo-surfaces.md#L188-L210).
18. **P3 / S — After #1288 merges, update memory-dir state.**
- Mark item D resolved and update branch tips.
- Source: [runbook post-merge steps](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/runbook.md#L51-L70).
19. **P3 / S — Remove stale split-PR mental model from follow-up docs.**
- #1288 is the current vehicle; split bookmarks are historical.
- Source: [D-17 implications](file:///home/william5lin/FastVideo/.agents/memory/dreamverse-integration/decisions-log.md#L34-L45).
20. **P3 / S — Keep product-only Dreamverse frontend out of FastVideo unless explicitly re-scoped.**
- This preserves release cadence and avoids packaging bloat.
- Source: [Dreamverse baseline](file:///home/william5lin/Dreamverse/README.md#L5-L16).
@@ -1,618 +0,0 @@
# Open Threads — Active Follow-Ups
Live work items with priority, effort estimate, dependencies, and
recommended next action.
For why each item is open see [decisions-log.md](decisions-log.md). For
PR-level context see [pr-roadmap.md](pr-roadmap.md).
**Last updated:** 2026-05-14 (added DR-7 follow-up to remove LTX2 debug
logging env vars and replace them with per-pipeline/model state. Earlier: 2026-05-14 added DR-6 follow-up for PR #1335's Gemma
lazy-load/device-placement behavior outside compiled `forward`. Earlier: 2026-05-13 added DR-5 follow-up for PR #1333's LTX2
distilled SSIM reference refresh. Earlier: 2026-05-12 added DR-4 follow-up for PR #1330's skipped
`App websocket integration` suite. Earlier: 2026-05-05 D-20 broken-pipe root cause + fix landed on
`will/dreamverse-monorepo` @ `5eaf0a13`; added new thread D-20-CP for
cherry-picking the public-API audio routing fix to `will/ltx2_sr_port`
so it lands in PR #1288. Earlier: strategy reversal — PR #1287 CLOSED,
replaced by mega-PR #1288 on `will/ltx2_sr_port` @ `b36bdbc9` covering
the full 6-layer stack at once. See [decisions-log.md D-17](decisions-log.md#d-17).
Item D resolution gate is now #1288 merge instead of #1287; same content,
different vehicle.).
## Priority overview
| # | Pri | Item | Effort | Unblocks |
|---|---|---|---|---|
| **D-20-CP** | High | Cherry-pick `[fix] api: route LTX-2 audio kwargs through batch.extra; strict update` (`265ce1a6`) onto `will/ltx2_sr_port` so it lands in PR #1288 | 15 min | Surfaces the public-API audio-conditioning fix in the mega-PR rather than waiting for `will/dreamverse-monorepo` to fold in. Tests already pass (185/185 api). |
| **~~D-22~~** | ~~Med~~ | ~~Per-chunk timing instrumentation in `stream_fmp4` and the controller's AV relay loop~~ | ~~M~~ | ✅ **Resolved 2026-05-06** in `bade2c0a`. `av_streaming.stream_fmp4` now records `av_wav_write_ms`, `av_ffmpeg_spawn_ms`, `av_first_chunk_ms`, `av_chunk_interval_ms_{min,median,p95,max}`, `av_chunk_publish_ms_{median,p95}`, `av_chunk_read_ms_{median,p95}` into the timings dict; `gpu_pool.handle_command` prints a one-line summary per segment so they show up directly in the deploy log. **Controller-side WS-send instrumentation is the remaining unaddressed slice** — see new D-22-CTL. |
| **D-22-CTL** | Low | Controller-side AV relay timing in `apps/dreamverse/server/session/controller.py` (between media event arrival and ws_send_bytes). | S | The worker side is now fully attributed by D-22. The controller-side per-chunk WS-send overhead is the still-unmeasured slice of the 700ms gap between `worker_e2e` and `main_user_step`. Add `t_ws_send_ms` per chunk in the AV relay loop, surface as `controller_ws_send_ms_{median,p95}` in the segment summary log line. |
| **D-23** | Med | On NVENC-capable hosts (RTX 50-series / T4 / A10 / H100 PCIe), benchmark `h264_nvenc` vs `libx264` and decide whether to default `--nvenc=true` for that SKU | S benchmark + S decision | Stutter mitigation depends on hardware; B200 needs different fix path (D-24). |
| **D-24** | Low | For B200-class deploys without NVENC, prototype gen-N+1 // encode-N pipelining or FE buffer pre-fill | L | Architectural change required; benchmark suggests ~700ms can be hidden if we pipeline. Ship-blocker only when realtime stutter becomes user-visible on B200 deploys. |
| **D-8** | High | Verify `ltx2_image_crf` post-`d80c2a8` | 10 min | Confirms typed stage-override path actually flows; closes a latent silent-drop bug |
| **1** | High | Migrate `/healthz`+`/readyz`+`/status` into FastVideo `build_app` | M-L | Closes BE_FLAVOR=fastvideo FE-compatibility; closes streaming-upstream contract debt |
| **2** | High | Fix pre-existing AbsMaxFP8 test failure | S | Self-contained quantization tech debt |
| **VPO** | High | Decide `video_position_offset_sec` semantics (a vs b) | 30 min | Unblocks PR 7.6 state emission |
| **D** | 🟢 in flight | Implement `generate_async` — content shipped in mega-PR **#1288** on `will/ltx2_sr_port` @ `b36bdbc9` (was #1287, CLOSED + re-routed per [D-17](decisions-log.md#d-17)) | L | Closes Q-5/Q-9/PR-7.5 TODOs simultaneously; enables Dynamo backend; unblocks audio re-encode; enables `GpuPool.run_async()` migration (D-12-B). Resolution gate: #1288 merge. |
| **DR-1** | High | Dreamverse: create `prompting/_internal_compat.py` shim + replace local `prompt_enhancer.py` (1933 LOC) — **PR #1258 has merged (`f673423b`); now actionable** | M (~150-200 LOC shim, replace upstream wiring) | Lets Dreamverse stop carrying a 1933-LOC fork |
| **DR-2** | Med | Decide `cerebras_ifm` provider path: (a) public Literal + `CerebrasIFMProvider` shipped, OR (b) Dreamverse-side custom provider via `enhancer.register_provider(...)` | S (decision) + S-M (impl) | Resolves the cerebras_ifm gap left by PR #1258. Same item as legacy #3 below; DR-2 is the Dreamverse-side framing. |
| **DR-3** | Low | Replace Dreamverse `PromptEnhancer._run_blocking_request` manual thread/queue polling with `asyncio.to_thread` after the PR #1327 prompt-enhancer compatibility surface is retired or isolated | S | Review comment #1327 (`prompt_enhancer.py`) is valid, but deferred to avoid patching the local fork in this PR. |
| **DR-4** | Low | Investigate unskipping PR #1330's skipped public `App websocket integration` suite | S-M | Public PR #1330 has `describe.skip(...)` around 27 websocket tests while the internal equivalent suite is active with 25 tests. The 2 public-only tests cover backend unreachable / GPU workers not ready. Not blocking while skipped, but stale assertions may need safe refresh before unskip. |
| **DR-5** | Low | Regenerate LTX2-Distilled latent SSIM references under the intended neutral/distilled defaults, then remove the PR #1333 historical full-guidance pins | S-M | PR #1333 changed public LTX2 distilled defaults to neutral/distilled values, but existing LTX2 latent SSIM references appear to have been generated with historical full-guidance defaults. The current PR pins the SSIM test to old values to keep CI compatible until references are refreshed. |
| **DR-6** | Low | Decide whether `LTX2GemmaTextEncoderModel` needs a non-forward device-placement hook after lazy Gemma load | S | PR #1335 should remove the `model.device` / `model.to(...)` guard from `forward` for Dynamo/fullgraph compatibility. Non-compiled runs probably do not need it because `gemma_model` moves Gemma at first load, but a later wrapper `.to(...)` after lazy load could leave Gemma on the old device unless lifecycle placement handles it. |
| **DR-7** | Low | Remove LTX2 debug logging env-var plumbing and replace it with per-pipeline/model debug state | S-M | PR #1335 review flagged that `initialize_pipeline()` mutates process-global LTX2 debug env vars, which can leak/race across pipeline instances. Deleting only the mutation is low-risk but loses config-driven debug logging; deleting all reads without replacement would remove useful SSIM/latent drift diagnostics. |
| **3** | Med | Add `cerebras_ifm` to `PromptEnhancerConfig.provider` Literal + provider | S-M | Public-side resolution if DR-2 picks (a) |
| **4** | Med | Expose `layer_profile` on typed `engine.quantization` | M | Removes Dreamverse's `experimental["pipeline_config"]` dodge for stage profiles |
| **5** | Med | Design typed `dit_config.quant_config` carrier | L design + L impl | Removes broader `experimental["pipeline_config"]` escape hatch |
| **SBS** | Med | `SessionStore` / `BlobStore` lifecycle policy | M design | Needed in PR 7.5 design pass |
| **D-12-A** | Med | Update `GpuPool` ABC docstring: mark "API may change post-PR-7.10; experimental / server-internal" | trivial | Prevents accidental promotion of streaming-internal API to framework-level |
| **D-12-B** | Med | Replace `GpuPool.run() -> Any` with `run_async() -> AsyncIterator[VideoEvent]` in PR 7.10 cycle | M | Closes the streaming-server cancellation TODO; converges with `generate_async` |
| **D-13-A** | Med | Document `fastvideo.entrypoints.streaming.prompt.*` in user-facing docs as "streaming-server scoped"; avoid framework-level framing | trivial (docs only) | Keeps future move to `fastvideo.prompt.*` cheap |
| **D-13-B** | Low | Add optional `client_factory` parameter to `LLMProvider` for `httpx.AsyncClient` pooling | S | Only if metrics show connect/TLS overhead is meaningful |
| **D-12-C** | Low | Avoid locking `PoolAssignment.gpu_id: int` as public; rename to `worker_id` (already exists) or add `device_ids: list[int]` for topology-aware pooling | S | Future multi-GPU-per-worker refactor stays cheap |
| **6** | Low | Audio attention quantization profile + test update | S | Future audio quant exploration |
| **7** | Low | Schema parity inventory cleanup (env-driven prompt fields) | S-M | Long-term consistency |
| **8** | Low | Stale `apps/web/test-results/` dir cleanup | trivial | Cosmetic |
| **11** | Low | Promote LTX-2 prompt orchestration (locked segments, segment_prompts JSON shape, rollout id/label) to `fastvideo.entrypoints.streaming.prompt.ltx2_orchestration` | M | Resolves Q-2 from decisions-log when a second LTX-2-style consumer appears |
| **12** | Low | When streaming server starts using `PromptSafetyFilter`, ensure operator-visible logging on `SafetyDecision.UNAVAILABLE` results | trivial | Surfaces degraded-safety state to operators (per D-14 Watch-Out item) |
| **13** | Low | When sticky session routing is needed, add `ReplicaRegistry.select(routing_key: str | None = None)` and document where `session_id` lives (WS URL/header preferred over first JSON frame) | M | Forward-compat from D-15 — keeps the door open without buffering/peeking |
| **14** | Low | At higher load, add `_bridge_session()` max-size + timeout limits OR recommend Envoy/HAProxy in front | S-M | The libraries' basic backpressure suffices for MVP; document the limit per D-15 |
| **15** | Low | If active-active multi-primary becomes a requirement, define behavior (round-robin within healthy primaries, weighted, sticky-by-key) | M | Currently `RouterConfig.__post_init__` rejects multi-primary; D-15 deferred until evidence |
| **~~Source-doc disposition~~** | ~~Med~~ | ~~Disposition of 7 untracked source docs~~ | ~~trivial~~ | ✅ **Resolved 2026-05-03** — moved into [source-archive/](source-archive/) |
| **~~9~~** | ~~Low~~ | ~~Commit-message cleanup: PR 8's 3 commits still have `[8/n] Improve API:` prefix~~ | ~~S~~ | ✅ **Resolved 2026-05-04** — bundled into the will/api_7.8 prep rebase. PR 8's 3 commits now read `[type] streaming: ...` |
| **~~10~~** | ~~Low~~ | ~~Commit-message cleanup: PR 7.8/7.9 commits have `streaming: streaming X` duplication~~ | ~~S~~ | ✅ **Resolved 2026-05-04** — bundled into the will/api_7.8 prep rebase. 3 commits dedup'd. |
---
## High priority
### D-20-CP: Cherry-pick public-API audio routing fix to `will/ltx2_sr_port`
**Why:** Commit `265ce1a6` (`[fix] api: route LTX-2 audio kwargs through batch.extra; strict update`) lives only on `will/dreamverse-monorepo` today. It touches public FastVideo surface (`fastvideo/entrypoints/video_generator.py`, `fastvideo/api/sampling_param.py`, plus a new regression test). For PR #1288 to ship a coherent public API — including the strict `SamplingParam.update()` — this commit needs to also land on `will/ltx2_sr_port`.
**Action:**
1. `git checkout will/ltx2_sr_port`
2. `git cherry-pick 265ce1a6` (clean; only touches `fastvideo/` paths that exist on both branches)
3. `pre-commit run --files fastvideo/entrypoints/video_generator.py fastvideo/api/sampling_param.py fastvideo/tests/api/test_extra_overrides_routing.py`
4. `pytest fastvideo/tests/api/ -q` (expect 185 passed)
5. `git push origin will/ltx2_sr_port` (fast-forward, no force)
6. `git checkout will/dreamverse-monorepo` (return to default working branch per runbook)
**Outcome:** PR #1288 picks up the fix automatically (its head IS `will/ltx2_sr_port`). The cherry-pick lives on both branches as separate SHAs; they'll dedupe naturally on any future rebase.
**Effort:** 15 minutes (cherry-pick + lint + test + push).
**Dependencies:** None. Tests already pass; no rebase conflicts expected.
**Files touched (same on both branches):**
- `fastvideo/entrypoints/video_generator.py`
- `fastvideo/api/sampling_param.py`
- `fastvideo/tests/api/test_extra_overrides_routing.py` (new file)
### D-8: Verify `ltx2_image_crf` typed flow post-`d80c2a8`
**Why:** Apr 26 dreamverse_review documented this field getting silently
dropped by the public `SamplingParam`. May 2 `d80c2a8` (Dreamverse)
refactored to typed `GeneratorConfig` + `preset_overrides`. Whether
`image_crf` now flows through `request.stage_overrides.refine.image_crf`
(per [design.md](design.md) mapping) or is still dropped is unverified.
**Action:**
1. Read [`Dreamverse/server/video_generation.py`](file:///home/william5lin/Dreamverse/server/video_generation.py)
post-`d80c2a8` for `image_crf` handling
2. Trace through to FastVideo's `request.stage_overrides.refine.image_crf`
3. Confirm runtime consumption in [`fastvideo/pipelines/basic/ltx2/`](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/)
**Effort:** 10 min, no code changes.
**Outcome:** Either confirms working OR identifies bug → opens fix item.
### Item #1: Migrate `/healthz`+`/readyz`+`/status` into `build_app`
**Why:** Today
[`fastvideo.entrypoints.streaming.server.build_app`](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/server.py)
exposes only `/health` + `/v1/stream`. Dreamverse FE expects all of
`/healthz`, `/readyz`, `/status`, `/curated-presets`,
`/prompt-system-config`, devtools.
The streaming-server-upstream plan (line 84) explicitly lists
`/healthz`+`/readyz`+`/status` as part of the contract that the upstream
of `realtime/` → `streaming/` must preserve. They were deferred from
PR 7.5's MVP. `/curated-presets` and `/prompt-system-config` are
operator-side and stay in Dreamverse (FE feature-detects).
**Action:**
1. Read PR 7.5 (#1251) `build_app` to scope what's there
2. Read [`Dreamverse/server/routes/health.py`](file:///home/william5lin/Dreamverse/server/routes/health.py)
for the route shapes Dreamverse already consumes
3. Propose route migration as commit on top of `will/api_7.5` or as
part of PR 7.10 cycle
4. Land
**Effort:** Medium-Large (route shapes need preservation; tests).
**Dependencies:** None blocking; can land anytime.
**Files likely to touch:**
- `fastvideo/entrypoints/streaming/server.py::build_app`
- New `fastvideo/entrypoints/streaming/health.py`
- Tests in `fastvideo/tests/entrypoints/streaming/`
### Item #2: AbsMaxFP8 pre-existing test failure
**Why:** [`fastvideo/tests/ops/quantization/test_absmax_fp8.py::test_create_weights_rejects_invalid_dtype`](file:///home/william5lin/FastVideo/fastvideo/tests/ops/quantization/test_absmax_fp8.py)
fails with `AssertionError not raised`. Pre-existing on `main`; verified
NOT introduced by NVFP4 work via `git stash`.
**Action:**
1. `git log --oneline fastvideo/tests/ops/quantization/test_absmax_fp8.py`
to find when it last passed
2. Either:
- Restore the assert in `AbsMaxFP8LinearMethod.create_weights` if
intentional behavior was lost
- Drop the test if assert is no longer correct
3. Verify
**Effort:** Small.
**Dependencies:** None.
### Item VPO: `video_position_offset_sec` semantics
**Why:** Per
[`fastvideo/pipelines/basic/ltx2/continuation.py`](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/continuation.py),
`LTX2ContinuationState.video_position_offset_sec` exists as a state
field. Two valid interpretations:
- **(a) Persistent across segments** — accumulating time offset for long
sessions; useful for time-coherent audio chaining.
- **(b) Per-segment hint that rides on the carrier** — runtime
overwrites every time; field is harmless redundancy.
Dreamverse computes `prefix_sec = float(audio_extra) / 24.0` per segment
in `apply_audio` and currently does NOT persist it on
`ContinuationState`. Field's docstring leans toward (b).
**Decision deadline:** before PR 7.6 starts emitting/consuming the
field (PR 7.6 branch is ready, not yet PR'd).
**Action:**
1. Confirm field's intended semantics with audio team
2. If (a): document the accumulation rule explicitly + add tests
3. If (b): leave docstring as-is + add test confirming overwrite
**Effort:** 30 min discussion + small implementation.
### Item D: Implement `generate_async` (PR 7.10)
**Why:** Highest leverage. Closes:
- D-5 / Q-5: audio re-encode for cross-segment continuity
- Q-9: Dynamo progress passthrough (deferred)
- PR 7.5's mid-segment cancellation TODO
- Unblocks Dynamo native backend integration
- **D-12-B**: enables `GpuPool.run() -> run_async() -> AsyncIterator[VideoEvent]` migration
**Action:** See [streaming-server.md](streaming-server.md) "PR 7.10 — the
unlock PR" section for scoping.
**Effort:** Large.
**Dependencies:** Best after PR 7.6 lands (gpu_pool upstream).
**Files:**
- `fastvideo/entrypoints/video_generator.py` — add `generate_async`,
refactor `generate_video` as wrapper
- `fastvideo/api/results.py` — add `VideoEvent`/`VideoProgressEvent`/
`VideoPartialEvent`/`VideoFinalEvent`
- `fastvideo/entrypoints/streaming/server.py` — consume `generate_async`,
remove TODO markers
- `fastvideo/entrypoints/streaming/gpu_pool.py` — add `run_async()`
forwarding events from worker to caller
- New `fastvideo/tests/entrypoints/test_generate_async.py`
- New `fastvideo/tests/contract/test_dynamo_shape.py` (already in PR 8)
### Item DR-1: Dreamverse — replace local `prompt_enhancer.py` with public + compat shim
**Why:** Today Dreamverse carries `Dreamverse/server/prompt_enhancer.py`
(1933 LOC) — a local copy/derivative of the FastVideo-internal version.
After PR #1258 merges, Dreamverse should switch to the public
`fastvideo.entrypoints.streaming.prompt.PromptEnhancer` and delete most
of the local module.
**Migration shape:**
1. **Create** `Dreamverse/server/prompting/_internal_compat.py` (~150-200 LOC):
- Wraps public `PromptEnhancer.enhance()` → returns `EnhanceResult` shape Dreamverse expects
- Wraps public `PromptEnhancer.auto_extend()` — JSON-parses `LLMResponse.content` into `{"next_prompt": "..."}`
- Wraps public `PromptEnhancer.rewrite()` — JSON-parses into `{"segment_prompts": [...]}` with lenient fallback for malformed JSON
- Layers locked-segment + rollout_id + rollout_label metadata back on top
2. **Update** `Dreamverse/server/runtime.py + main.py` — replace `from prompt_enhancer import PromptEnhancer` with `from prompting._internal_compat import PromptEnhancer`
3. **Delete most of** `Dreamverse/server/prompt_enhancer.py` (1933 LOC). Keep only the bits that don't have a public equivalent:
- Race-based parallel fallback (`_run_provider_race`) — Dreamverse-specific tail-latency optimization
- `cerebras_ifm` provider — pending DR-2 decision
- Multi-classifier prompt safety (NSFW + hate-speech chained) — public ships single classifier
4. **Tests** — verify Dreamverse session controllers still see the expected response shapes through the shim
**Effort:** Medium (~150-200 LOC shim + replace upstream wiring + delete 1700+ LOC local module + test fixture updates).
**Dependencies:**
- PR #1258 must merge first (publishes `fastvideo.entrypoints.streaming.prompt.*`)
- DR-2 informs the cerebras_ifm path
**Files:**
- New: `Dreamverse/server/prompting/_internal_compat.py`
- Modified: `Dreamverse/server/runtime.py`, `Dreamverse/server/main.py`
- Mostly deleted: `Dreamverse/server/prompt_enhancer.py`
---
## Medium priority
### Item DR-2: Decide `cerebras_ifm` provider path
**Why:** Public PR #1258's `PromptEnhancerConfig.provider` is
`Literal["cerebras", "groq"]`. Internal supports `"cerebras_ifm"` (the
Cerebras IFM API endpoint with different auth). Dreamverse needs
`cerebras_ifm` working post-migration.
Two options:
| Option | Approach | Pros | Cons |
|---|---|---|---|
| **(a) Public** | Add `"cerebras_ifm"` to public Literal + ship `CerebrasIFMProvider` in `fastvideo/entrypoints/streaming/prompt/providers/cerebras_ifm.py` | Discoverable; users with IFM access can use typed config | Adds ~50 LOC + Literal extension to public surface |
| **(b) Dreamverse-side** | Implement `CerebrasIFMProvider` Dreamverse-side as a custom `LLMProvider`, register via `enhancer.register_provider(CerebrasIFMProvider())` | Zero public surface change; private endpoint stays private | Slightly more boilerplate Dreamverse-side; not surfaced to non-Dreamverse users |
**Recommendation:** Option (b) is more contained. Option (a) is more
discoverable. Default to (b) unless there's a third-party user who needs
IFM access. The Dreamverse-side PR carrying DR-1 is the natural place to
make this decision.
**Effort:** Small (decision) + Small-Medium (implementation).
**Dependencies:** DR-1 (compat shim creation).
### Item #3: `cerebras_ifm` provider in public Literal
**Why:** Same item as DR-2 from the public-side framing. If DR-2 picks
option (a), this is the implementation. If DR-2 picks option (b), this
item is closed without implementation.
**Action:** See DR-2.
**Effort:** S-M.
### Item #4: Expose `layer_profile` on typed `engine.quantization`
**Why:** Today `transformer_quant: "NVFP4"` always constructs
`NVFP4Config()` with default `layer_profile="refine"`. Dreamverse
dodges via `experimental["pipeline_config"]`.
**Action:**
1. Add `transformer_quant_layer_profile: str | None = None` to
`QuantizationConfig` in [`schema.py`](file:///home/william5lin/FastVideo/fastvideo/api/schema.py)
2. Thread through [`compat.py`](file:///home/william5lin/FastVideo/fastvideo/api/compat.py)
3. Update `_apply_transformer_quant` in
[`fastvideo_args.py`](file:///home/william5lin/FastVideo/fastvideo/fastvideo_args.py)
to pass profile
4. Update Dreamverse to drop the `experimental["pipeline_config"]`
dodge in favor of typed knob
5. Tests in [`test_typed_quant_flow.py`](file:///home/william5lin/FastVideo/fastvideo/tests/api/test_typed_quant_flow.py)
**Effort:** Medium.
**Files:** schema.py, compat.py, fastvideo_args.py, test_typed_quant_flow.py,
+ Dreamverse/server/video_generation.py.
### Item #5: Typed `dit_config.quant_config` carrier
**Why:** The `experimental["pipeline_config"]` escape hatch in
Dreamverse should eventually become a typed field. Design TBD.
**Action:** Heaviest design work. Should consult Oracle.
**Effort:** Large design + Large implementation.
**Dependencies:** #4 should land first; this is the "final form" of #4.
### Item SBS: `SessionStore` / `BlobStore` lifecycle policy
**Why:** PR 7's in-memory implementations have no eviction, no TTL, no
automatic blob cleanup on state replacement. Documented as per-deployment
policy decision.
When PR 7.5/7.6 land the live consumer, who owns:
- bounded session capacity (LRU? TTL? hard max?)
- blob `drop()` chained when state is replaced
- session expiry on websocket disconnect
**Recommendation:** streaming server's session manager. Worth stating
explicitly in PR 7.5's design.
**Effort:** Medium design + small implementation.
### Item D-12-A: Update `GpuPool` ABC docstring — mark experimental
**Why:** Per D-12 in [decisions-log.md](decisions-log.md), `GpuPool`
should be documented as "API may change post-PR-7.10; experimental /
server-internal" to prevent accidental promotion of streaming-internal
API to framework-level. PR #1257 merged without this caveat.
**Action:** Edit
[`fastvideo/entrypoints/streaming/gpu_pool.py`](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/gpu_pool.py)
class docstring on `GpuPool` ABC. Add a note: "API may change post-PR-7.10
when run_async() lands; treat as server-internal for now."
**Effort:** Trivial.
**Dependencies:** None.
### Item D-12-B: Replace `GpuPool.run() -> Any` with `run_async() -> AsyncIterator[VideoEvent]`
**Why:** Per D-12, this is the canonical evolution post-PR-7.10. Closes
the streaming server's cancellation TODO and converges the streaming +
OpenAI + Dynamo consumers on a single async API.
**Action:** As part of PR 7.10 cycle:
1. Add `GpuPool.run_async(session_id, request) -> AsyncIterator[VideoEvent]`
2. Worker forwards events through `result_queue` with type discriminator
3. Streaming server replaces `await pool.run(...)` with `async for event in pool.run_async(...)`
4. Sync `run()` becomes a thin compat wrapper that collects events and returns the final
5. Cancellation propagates: client disconnect → `asyncio.CancelledError` → worker stops mid-step
**Effort:** Medium. Adds ~50-100 LOC + tests.
**Dependencies:** Item D (PR 7.10 — `generate_async` on `VideoGenerator`).
### Item D-13-A: Document `streaming/prompt/*` as streaming-scoped
**Why:** Per D-13 in [decisions-log.md](decisions-log.md), the prompt
enhancer is currently scoped to streaming-server use even though the
abstraction is general. Phrase user-facing docs as "streaming-server
prompt enhancement" to keep future move to `fastvideo.prompt.*` cheap.
**Action:** When PR 12 (docs migration) is written, the prompt enhancer
section should:
- Be titled "Streaming Server Prompt Enhancement", not "Prompt API"
- Note the 3 fixed operations (`enhance` / `auto_extend` / `rewrite`) are
shaped by LTX-2 streaming session needs
- Note that consumers wanting custom prompt operations can use
`provider.complete()` directly with their own LLMRequest
- Avoid `from fastvideo import LLMProvider` exports until a second
consumer exists
**Effort:** Trivial (docs only).
**Dependencies:** PR 12 (docs migration).
---
## Low priority
### Item D-13-B: Optional `client_factory` parameter for `httpx.AsyncClient` pooling
**Why:** Today `_openai_compat.py` instantiates `httpx.AsyncClient` per
call (no connection pooling). Reviewer flagged inefficient. Team chose
simplicity for the expected scale (~6-10 enhancer calls per LTX-2
session). If real-world metrics show connect/TLS overhead is meaningful,
add an optional `client_factory: Callable[[], httpx.AsyncClient] | None`
parameter to providers so they can share a pool.
**Action:** Only when metrics justify. Add `client_factory=None` parameter
to `CerebrasProvider` / `GroqProvider` constructors and pass through to
`complete_openai_compatible()`. Default to current per-call behavior.
**Effort:** Small.
**Dependencies:** None blocking; only act on real perf data.
### Item DR-4: Unskip PR #1330 `App websocket integration` suite safely
**Why:** Public PR #1330 currently wraps `App websocket integration` in
`describe.skip(...)`, so the 27-test websocket integration suite does not
run. The internal repo has the equivalent suite active via `describe(...)`
with 25 tests. The two public-only cases cover backend unreachable and GPU
workers not ready readiness/reachability behavior.
**Action:** Later, investigate whether the public suite can be unskipped and
refresh any stale assertions without changing the intended websocket contract.
Because the suite is skipped today, this is not blocking the current stack or
current PR #1330 review.
**Effort:** Small-Medium.
**Dependencies:** Best handled when someone can run the frontend integration
stack end-to-end and compare the public assertions against the internal active
suite.
### Item DR-6: Gemma lazy-load device placement outside `forward`
**Why:** PR #1335 review flagged this pattern in
`fastvideo/models/encoders/gemma.py::LTX2GemmaTextEncoderModel.forward`:
```py
if model.device != target_device:
model.to(device=target_device)
```
The immediate concern is the compiled/Dynamo path: `model.device` and
`model.to(...)` inside `forward` can introduce non-tensor/device parsing work
that fullgraph tracing should not see. Removing the guard from `forward` is the
right PR #1335 review fix.
For non-compiled execution, the guard is mostly defensive rather than required:
`gemma_model` already moves the lazily loaded HF Gemma model to the wrapper's
current parameter device on first load. The remaining edge case is a lifecycle
sequence where Gemma is loaded, then the parent wrapper is later moved to a
different device; in that case Gemma could stay behind unless placement is
handled outside `forward`.
**Action:** After PR #1335 review is unblocked, decide whether FastVideo needs a
small lifecycle hook/helper for this class so lazy Gemma is moved whenever the
wrapper/device placement changes. If yes, implement it outside `forward`; if no,
document that Gemma must be loaded after final device placement.
**Effort:** Small.
**Dependencies:** Not blocking PR #1335 if the forward-path guard is removed and
the normal load path keeps placing Gemma on the wrapper's current device.
### Item DR-7: Remove LTX2 debug logging env-var plumbing
**Why:** PR #1335 review flagged that
`fastvideo/pipelines/basic/ltx2/ltx2_pipeline.py::initialize_pipeline()` sets
and pops LTX2 debug env vars such as `LTX2_PIPELINE_DEBUG_LOG`,
`LTX2_PIPELINE_DEBUG_PATH`, `LTX2_DEBUG_DETAIL`, and
`LTX2_PIPELINE_DEBUG_DETAIL_PATH`. Those env vars are process-global, so one
pipeline instance can enable, overwrite, or clear debug behavior for another
pipeline instance running in the same process.
Deleting only the `initialize_pipeline()` env mutation is low risk for normal
generation, but it would stop config-driven debug logging unless replaced.
Deleting all env-var reads without replacement is riskier because these logs are
useful for SSIM/latent drift diagnosis and may be used by local debug scripts.
**Action:** Replace LTX2 debug env-var plumbing with per-pipeline/model debug
state. Use pipeline/model config for construction-time hooks and a
`ForwardContext`/`ForwardBatch`-style carrier for forward-time logging. Keep
external env-var compatibility only if there is a documented operator workflow
that still needs it.
**Effort:** Small-Medium.
**Dependencies:** Not blocking PR #1335 if the immediate fix is limited to
removing process-global mutation from pipeline initialization while preserving
existing externally supplied env-var reads.
### Item D-12-C: Avoid locking `PoolAssignment.gpu_id: int` as public
**Why:** Today `PoolAssignment` exposes `gpu_id: int`, assuming
one-GPU-per-worker. Future topology-aware pooling may need
`device_ids: list[int]` (one worker = group of GPUs running internal
`MultiprocExecutor`). Don't freeze the int field as public API.
**Action:**
- Treat `gpu_id` as a current-impl detail; prefer `worker_id` (already
exists, is stable identifier)
- When a worker actually spans multiple GPUs, add
`PoolAssignment.device_ids: list[int]` and let `gpu_id` be `device_ids[0]`
for backward compat
- Or rename to `gpu_id` → `device_id` with deprecation alias
**Effort:** Small (1 field rename + alias).
**Dependencies:** Driven by an actual future "one worker = many GPUs" use case. Don't act preemptively.
### Item #6: Audio attention quantization profile
**Why:** Today audio attn and FFN are bf16. If an audio-quant profile
is added to `NVFP4Config.fp4_layers`, update
[`test_basic_av_block_propagates_quant_config_to_all_children`](file:///home/william5lin/FastVideo/fastvideo/tests/ops/quantization/test_nvfp4_ltx2_wiring.py).
**Effort:** Small (one test + one config field).
### Item #7: Schema parity inventory cleanup
**Why:** A few internal-only fields are not exposed publicly:
- `PROMPT_HTTP_TIMEOUT_MS`
- `PROMPT_INITIAL_STAGE_TIMEOUT_MS`
- `PROMPT_TEMPERATURE`
- `PROMPT_MAX_COMPLETION_TOKENS`
- `PROMPT_AUTO_SLEEP_MS`
- `PROMPT_AUTO_TIMEOUT_MS`
- curated-presets file paths
These flow via env vars on `dreamverse-server` today. If
`fastvideo serve --config` becomes the canonical entrypoint, they need
typed homes.
**Effort:** Small-Medium.
### Item #8: Stale `apps/web/test-results/` directory
**Why:** Cosmetic. `.gitignore` entry hides it from `git status`, but
the dir has stale `.last-run.json` (45 bytes) from a prior Playwright
run.
**Action:** `rm -rf apps/web/test-results` whenever convenient.
**Effort:** Trivial.
### ~~Item #9~~ + ~~#10~~: Commit-message cleanups — ✅ Resolved 2026-05-04
Both items resolved during the `will/api_7.8` prep rebase. A targeted
conditional script (`/tmp/opencode/cleanup_subjects_v2.sh` — only amends
when text actually changes) ran across 33 commits, modified 6:
- **#9 fix**: extended the regex from `\[\d+\.\d+/n\]` to
`\[\d+(\.\d+)?/n\]` so single-digit prefixes match. PR 8's 3 commits
now read `[type] streaming: ...` instead of `[type] [8/n] Improve API: ...`.
- **#10 fix**: added second substitution `streaming: streaming X` →
`streaming: X`. PR 7.8 / 7.9 commits no longer have the duplication.
The conditional check skipped pre-commit-hook flakiness on no-op amends
(unlike the earlier first attempt). All affected commits verified clean
post-rebase.
### Item #11: Promote LTX-2 prompt orchestration to public (when 2nd consumer exists)
**Why:** Per Q-2 in [decisions-log.md](decisions-log.md) and D-13's
"missing alternative", the LTX-2-specific orchestration (locked
segments, segment_prompts JSON shape, rollout id/label, lenient JSON
parsing) currently stays Dreamverse-side per DR-1. If a second
LTX-2-style consumer appears (e.g. another video model with multi-segment
continuation needing the same prompt orchestration), promote this layer
to `fastvideo.entrypoints.streaming.prompt.ltx2_orchestration`.
**Action:** Wait for a second consumer to materialize. Until then, the
orchestration stays in Dreamverse's `_internal_compat.py` shim (DR-1).
**Effort:** Medium when triggered.
**Dependencies:** A second consumer.
---
## Recommended pull order
If you have unbounded time and want to maximize forward progress:
1. **D-8 verify** (10 min) — eliminates uncertainty
2. **D-12-A docstring** (trivial) — caveat the GpuPool API publicly
3. **Item #2 AbsMaxFP8** (S) — clears tech debt
4. **Item VPO video_position_offset_sec** (30 min) — unblocks PR 7.10 (since 7.6 has merged, this is now scoped to whatever consumer first reads the field)
5. **DR-1 + DR-2 Dreamverse migration** (M) — **now unblocked since PR #1258 merged**; replaces 1700+ LOC of local fork
6. **Item #4 layer_profile** (M) — closes Dreamverse quant escape hatch
7. **Item #1 build_app routes** (M-L) — closes FE-compat
8. **Item D generate_async** (L) — unlock PR; brings along D-12-B (run_async) + closes Q-5/Q-9/PR-7.5 TODOs
9. **Item #5 typed quant_config carrier** (L+L) — final form
10. **Items #6/#7/#8 + D-12-C/D-13-A/D-13-B + #11** — cleanup polish (#9, #10 resolved 2026-05-04)
If you have a specific user goal (e.g. "ship `BE_FLAVOR=fastvideo`
flavor end-to-end"), that goal dictates the order — read this list as a
menu, not a prescription.
---
## Verification gates per item
When implementing any item above, evidence required:
| Phase | Check |
|---|---|
| Build | `lsp_diagnostics` clean on changed files |
| Test | new + relevant existing tests pass; output captured |
| Manual QA | actually run the affected feature end-to-end (per AGENTS.md MANUAL_QA_MANDATE) |
| Regression | full `fastvideo/tests/api/` + `contract/` + relevant SSIM (if NVFP4 touch) |
For NVFP4 touches: re-run `test_nvfp4_ltx2_wiring.py` +
`test_typed_quant_flow.py` (CPU) + ideally a flashinfer-enabled path
test (manual, not in CI).
For Dreamverse-side items (DR-1, DR-2): re-run
`Dreamverse/apps/web/npx playwright test e2e/preset-prompt-generation.spec.ts`
end-to-end against the live BE+FE — this is the contract test that
exercises the prompt enhancer through a real session.
@@ -1,141 +0,0 @@
# PR Roadmap
Status of all 17 PRs in the FastVideo public API refactor + streaming
server upstream + Dynamo backend contract + post-deprecation cleanup.
For design rationale see [design.md](design.md). For streaming-specific
PRs (7.5-7.10) see [streaming-server.md](streaming-server.md). For NVFP4
work that runs parallel to this sequence see [quantization.md](quantization.md).
**Last updated:** 2026-05-05 (strategy reversal — single mega-PR #1288 replaces planned splits 7.10/8/LTX-2/NVFP4/post-fixes/agents_cleanup; see [decisions-log.md D-17](decisions-log.md#d-17)).
## Status legend
- ✅ **Landed on `origin/main`**
- 🟢 **Open / in flight** — branch exists, may have open PR
- 🟡 **Planned** — designed, not started
- 🔵 **Future** — deferred to post-PR-13 cleanup
## Landed PRs (0 → 7.7)
| # | PR | Status | Merge commit | Scope |
|---|---|---|---|---|
| 0 | #1218 [1/n] | ✅ | merged | Parity inventory + typed inference schema |
| 1 | #1218 [1/n] | ✅ | merged | Strict parser/validation/overrides + API tests |
| 2 | #1220 [2/n] | ✅ | merged | Typed `VideoGenerator` constructors + request path + compat |
| 3 | #1226 [3/n] | ✅ | merged | CLI/YAML-first typed config loading for `generate` and `serve` |
| 4 | #1234 [4/n] | ✅ | merged | Preset registry + presets for all 13 model families; `SamplingParam` moved to `fastvideo/api/`; `configs/sample/` deleted entirely |
| 5 | #1237 [5/n] | ✅ | merged | `ServeConfig.default_request` wired into stateless OpenAI server |
| 5.5 | (`5d1d71fc`) | ✅ | merged | Streaming server package skeleton, typed `StreamingConfig`/`GpuPoolConfig`/`PromptEnhancerConfig`/`PromptSafetyConfig`/`WarmupConfig`, `streaming-serve` CLI stub |
| 6 | #1239 [6/n] | ✅ | merged | LTX2 public preset + asset wiring + `gpu_pool.py` typed-kwarg translation |
| 7 | #1250 [7/n] | ✅ | merged | Typed LTX2 continuation state + streaming session store + blob store |
| **7.5** | **#1251** | ✅ | `95fd29e0` (merged 2026-04-26) | Streaming server skeleton (WebSocket + fMP4 + single generator). 8 commits. Deferred TODOs (per-step progress, mid-segment cancellation) carried forward to PR 7.10. |
| **7.6** | **#1257** | ✅ | `eb0a4152` (merged 2026-05-04) | GPU pool upstream + worker subprocess + two-segment warmup. 7 commits squashed. APPROVED by Eigensystem. See [decisions-log.md D-12](decisions-log.md#d-12) for the architectural review. |
| **7.7** | **#1258** | ✅ | `f673423b` (merged 2026-05-04) | Prompt enhancer with `LLMProvider` abstraction. Built-in providers: cerebras, groq. 3 commits squashed. **Public Literal does NOT include `cerebras_ifm`** — open-threads.md item DR-2 covers the gap. See [decisions-log.md D-13](decisions-log.md#d-13) for the architectural review. |
| **7.8** | **#1284** | ✅ | `eb3a3942` (merged 2026-05-04) | Streaming auxiliaries — `prompt/safety.py` (optional fasttext, lazy import), `prompt/rewrite.py`, `session_logger.py` (thread-safe JSONL), `mock_server.py` (build_mock_app + MockGenerator for FE dev). 730 LOC, 2 commits. See [decisions-log.md D-14](decisions-log.md#d-14). |
| **7.9** | **#1286** | ✅ | `2aaeee2a` (merged 2026-05-05) | Streaming router (multi-replica load balancer + WS proxy + `fastvideo router-serve` CLI). Squashed `cd76cf51 + 1ac1e732 + b0b7f59c + a152cb77` (router-polish second-pass; cherry-pick of `40e265b8` from `will/ltx2_sr_port`). See [decisions-log.md D-15](decisions-log.md#d-15) (structural review) + [D-16](decisions-log.md#d-16) (second-pass polish). |
## In flight (mega-PR #1288)
| # | PR | Status | Branch | Scope |
|---|---|---|---|---|
| **mega** | **#1288** | 🟢 OPEN, MERGEABLE | `will/ltx2_sr_port` (head `b36bdbc9`) | **Single consolidated landing of the full `will/ltx2_sr_port` chain.** Was originally planned as 6 stacked PRs (slices 1-3 / 4-6 / 7-15 / 16-21 / 22-23 / 24-34). Now landing as one PR — see [decisions-log.md D-17](decisions-log.md#d-17) for the strategy decision. **Contents** (commit-ordered): (1) streaming `generate_async` + `VideoEvent` + Dynamo backend contract (3 commits, was PR 7.10/#1287 closed); (2) server contract docs + Dreamverse/Dynamo shape tests (3 commits, was PR 8); (3) LTX-2 SR runtime port + i2v conditioning + alignment harness (9 commits); (4) NVFP4 wire-up + per-component compile + typed `transformer_quant` flow (6 commits); (5) LTX-2 post-handoff parity fixes — Gemma `to()`, list-of-generators (2 commits); (6) `.agents/memory/dreamverse-integration/` knowledge base + agents Phase 1 cleanup (11 commits). 34 commits total, 71 files, +13,074/-583 LOC. |
## Closed PRs in this scope
| # | PR | Status | Why closed |
|---|---|---|---|
| **7.10** | **#1287** | ❌ CLOSED 2026-05-05 | Superseded by mega-PR #1288 — strategy reversal to land everything in one go. Same 3 commits now form the head of #1288. |
## Deprecated split bookmarks (D-17)
`will/api_7.10` / `will/api_8` / `will/ltx2_sr_runtime` / `will/ltx2_nvfp4` / `will/ltx2_post_fixes` / `will/agents_cleanup` were the split-PR bookmarks under the abandoned 6-PR plan. They remain locally as historical references but are no longer maintained. STACK.md (top-level) is similarly deprecated.
## Planned (post-#1288 merge)
| # | Status | Branch | Scope |
|---|---|---|---|
| 9 | 🟡 | — | LongCat preset migration + colocation (9 model-specific stage files) |
| 10 | 🟡 | — | Hunyuan15 SR preset migration + colocation + SR field migration POC |
| 11 | 🟡 | — | SSIM/performance test migration off legacy `generate_video(..., **kwargs)` |
| 12 | 🟡 | — | Docs + examples migration (includes streaming server + Dynamo) |
| 13 | 🟡 | — | Deprecation cleanup (includes flat LTX2 kwargs the internal `gpu_pool.py` used to consume) |
## Future (compat.py death sequence)
After PR 13 lands deprecation warnings, `fastvideo/api/compat.py` (~370
lines) is the last translation shim between typed public API and legacy
internals (`FastVideoArgs`, `SamplingParam`).
| # | Status | Scope | Lines removed |
|---|---|---|---|
| 14 | 🔵 reachable | Strip forward translation: `legacy_from_pretrained_to_config`, `legacy_generate_call_to_request`, `_sampling_param_to_request_raw`, `_LEGACY_REQUEST_ALIASES`, `_LTX2_REFINE_FLAT_KEYS`. Depends on PRs 11/12/7.6 callers being migrated. | ~100 |
| 15 | 🔵 | `FastVideoArgs` becomes a `@dataclass` view over `GeneratorConfig` with `@property` accessors backing legacy field names. ~600-line god-object refactor. Depends on PR 14. | reverse-translation half (~150) trivial |
| 16 | 🔵 | `ForwardBatch` reads `GenerationRequest` by reference; kills `request_to_sampling_param` and the `ForwardBatch(**shallow_asdict(sampling_param), …)` spread. `SamplingParam` demoted or deleted. Depends on PR 15. | rest |
| 17 | 🔵 | Move `normalize_generator_config`, `normalize_generation_request`, `load_generator_config_from_file` to `parser.py`. Delete `compat.py`. | file gone |
PRs 15-17 touch training, distributed, and worker code in addition to
inference path; realistically 1-2 quarters beyond the current plan.
## Dependency chain
```
PR 13 (deprecation)
↓
PRs 11, 12, 7.6 (migrate callers)
↓
PR 14 (forward translation gone) ─── ~100 lines out of compat.py
↓
PR 15 (FastVideoArgs as view) ─── reverse-translation trivial
↓
PR 16 (ForwardBatch reads request) ─── SamplingParam demoted
↓
PR 17 (move normalizers, delete file)
```
## NVFP4 work (out-of-band, parallel to PR 7.5+)
NOT in the canonical PR sequence. Lives on `will/ltx2_sr_port`
(currently @ `156103b9`) — a separate stack alongside the public-API
upstreaming. See [quantization.md](quantization.md) for what each commit
locks in.
| Commit range | Topic |
|---|---|
| `cfccd292..b6ac7630` | LTX-2 i2v + SR runtime port + alignment harness |
| `a4760bae..c6c14c55` | NVFP4 LTX-2 wire-up + per-component compile + parity fixes (May 2 handoff) |
| `a5fcd19c..156103b9` | Post-handoff parity/perf fixes |
## Key landed artifacts (reference points)
- Parity inventory: [`docs/design/inference_schema_parity_inventory.yaml`](file:///home/william5lin/FastVideo/docs/design/inference_schema_parity_inventory.yaml) + guard [`fastvideo/tests/api/test_schema_parity_inventory.py`](file:///home/william5lin/FastVideo/fastvideo/tests/api/test_schema_parity_inventory.py)
- Typed schema: [`fastvideo/api/schema.py`](file:///home/william5lin/FastVideo/fastvideo/api/schema.py)
- Compat layer: [`fastvideo/api/compat.py`](file:///home/william5lin/FastVideo/fastvideo/api/compat.py)
- Preset system: [`fastvideo/api/presets.py`](file:///home/william5lin/FastVideo/fastvideo/api/presets.py) + per-family `pipelines/basic/<family>/presets.py`
- Streaming package skeleton (PR 5.5): [`fastvideo/entrypoints/streaming/`](file:///home/william5lin/FastVideo/fastvideo/entrypoints/streaming/)
- LTX2 typed continuation state (PR 7): [`fastvideo/pipelines/basic/ltx2/continuation.py`](file:///home/william5lin/FastVideo/fastvideo/pipelines/basic/ltx2/continuation.py)
## Known notable decisions carried forward
- **Public inference boundary stays plain dataclasses + plain dict/YAML/JSON**
— not OmegaConf, not runtime config wrappers.
- **Every public entrypoint normalizes into typed config objects** before
touching legacy `FastVideoArgs` or `SamplingParam`.
- **Legacy `generate_video(..., **kwargs)` stays on direct legacy execution
path until PR 11**'s SSIM/performance migration. Prevents golden
baselines from drifting during compat period.
- **Typed requests use schema defaults**; legacy `generate_video(...)`
continues to inherit model-specific `SamplingParam` defaults during
compat period.
- **Preset registry uses explicit `_register_presets()` pattern** matching
`_register_configs()`; lookup keyed by `model_family`.
- **Stateless OpenAI server clones `ServeConfig.default_request`** and
merges user overrides; preset validation runs before legacy generation.
- **Streaming server added as sibling `fastvideo/entrypoints/streaming/`**
rather than extending `fastvideo/entrypoints/openai/` (PR 5.5).
## Per-PR commit-level detail
For per-PR commit lists, test plans, and merge criteria, the archived
source [`source-archive/PR-plan.md`](source-archive/PR-plan.md) (1145 lines)
remains the deepest reference. This file is the navigable summary.
@@ -1,229 +0,0 @@
# Quantization — NVFP4, LinearBase Fallback, Layer Profiles
What landed in the May 2 NVFP4 stack, why it's load-bearing, and what's
still owed (`layer_profile`, typed quant carrier, AbsMaxFP8 cleanup).
For overall API design see [design.md](design.md). For the open
follow-ups see [open-threads.md](open-threads.md).
**Last updated:** 2026-05-03.
## NVFP4 — what it is
NVIDIA's specific block-scaled FP4 format:
- e2m1 mantissa
- fp32 alpha
- `layout_128x4` scale layout
- group size 16
Distinct from MX-FP4 / OCP-FP4 / generic e3m0. The May 2 rename
(`94c983a2`) disambiguated the naming throughout FastVideo's public
surface.
## Files (current)
| File | Role |
|---|---|
| [`fastvideo/layers/quantization/nvfp4_config.py`](file:///home/william5lin/FastVideo/fastvideo/layers/quantization/nvfp4_config.py) | `NVFP4Config`, `NVFP4QuantizeMethod`, `convert_model_to_nvfp4` |
| [`fastvideo/layers/quantization/__init__.py`](file:///home/william5lin/FastVideo/fastvideo/layers/quantization/__init__.py) | `QuantizationMethods` literal includes `"NVFP4"`; `get_quantization_config` resolves it |
| [`fastvideo/layers/linear.py`](file:///home/william5lin/FastVideo/fastvideo/layers/linear.py) | `LinearBase.__init__` falls back to `UnquantizedLinearMethod` when `quant_config.get_quant_method` returns None — **load-bearing** |
| [`fastvideo/models/loader/fsdp_load.py`](file:///home/william5lin/FastVideo/fastvideo/models/loader/fsdp_load.py) | `_maybe_convert_model_to_nvfp4` helper detects via `isinstance(quant_method, NVFP4QuantizeMethod)`; calls `convert_model_to_nvfp4` to materialize buffers |
| [`fastvideo/models/dits/ltx2.py`](file:///home/william5lin/FastVideo/fastvideo/models/dits/ltx2.py) | `nn.Linear` → `ReplicatedLinear` for FP4-eligible subset; `_supports_prequantized_input` + `_linear_project_with_optional_prequant` helpers; quant_config + prefix= plumbing |
| [`fastvideo/api/compat.py`](file:///home/william5lin/FastVideo/fastvideo/api/compat.py) | Typed `engine.quantization.transformer_quant: "NVFP4"` resolves to `NVFP4Config()` instance |
| [`fastvideo/fastvideo_args.py`](file:///home/william5lin/FastVideo/fastvideo/fastvideo_args.py) | `__post_init__._apply_transformer_quant` pins `pipeline_config.dit_config.quant_config = NVFP4Config()` |
## Buffer naming (post-rename)
| Old | New |
|---|---|
| `_fp4_weight` / `_fp4_alpha` | `_nvfp4_weight` / `_nvfp4_alpha` |
| `_weight_global_sf` | unchanged |
| `convert_model_to_fp4` | `convert_model_to_nvfp4` |
| `FP4QuantizeMethod` | `NVFP4QuantizeMethod` |
| `QuantizationMethods` literal `"FP4"` | `"NVFP4"` |
Internal-scope torch op namespace `fastvideo_fp4::*` and
`_get_ltx2_fp4_stage_profile` deliberately left as-is — purely internal
naming that mirrors FastVideo-internal.
## Layer set asymmetry — by design
`NVFP4Config.fp4_layers` (default `layer_profile="refine"`) covers:
- `attn1.{to_q,to_k,to_v,to_out}` — full self-attention
- `attn2.{to_q,to_out}` — cross-attn Q + out only (text context not quantized)
- `audio_to_video_attn.{to_q,to_out}` — AV cross Q + out
- `video_to_audio_attn.{to_k,to_v}` — VA cross K + V
- `ffn.{fc_in,fc_out}` — video FFN
- `adaln_single.linear` — but this is `nn.Linear` (not `LinearBase`),
so it never actually gets FP4'd. List entry has no effect; matches
internal.
**NOT in the set:**
- audio self-attention (`audio_attn1.*`)
- audio cross-attention (`audio_attn2.*`)
- audio FFN (`audio.ffn.*`)
Audio path is cheap enough that quant overhead isn't worth it. Test
[`test_basic_av_block_propagates_quant_config_to_all_children`](file:///home/william5lin/FastVideo/fastvideo/tests/ops/quantization/test_nvfp4_ltx2_wiring.py)
locks this in — if you add audio quantization later, update the test.
## `LinearBase` fallback — DO NOT REMOVE
[`fastvideo/layers/linear.py:191-202`](file:///home/william5lin/FastVideo/fastvideo/layers/linear.py#L191-L202): when `quant_config.get_quant_method` returns
`None` (layer not in the quant config's set), we fall back to
`UnquantizedLinearMethod`.
**Removing this fallback would break every non-tagged
`ReplicatedLinear` constructed with an `NVFP4Config`** — the previous
`assert quant_method is not None` would crash on unmatched layers (e.g.
text-encoder K/V projections, audio attention, etc.).
This is one of the load-bearing changes from `42b30bf9`.
## `transformer_quant` precedence rules
`FastVideoArgs._apply_transformer_quant` only writes
`dit_config.quant_config` when it's currently `None`. **If a caller has
explicitly set** `pipeline_config.dit_config.quant_config = NVFP4Config(...)`,
the explicit setter wins.
Dreamverse's `video_generation.py` relies on this precedence — it sets
`NVFP4Config()` directly via `experimental["pipeline_config"]` because
typed `transformer_quant: "NVFP4"` doesn't yet expose `layer_profile`.
See "Open follow-ups" below.
## Attention forward optimization
[`models/dits/ltx2.py`](file:///home/william5lin/FastVideo/fastvideo/models/dits/ltx2.py)
ports `_supports_prequantized_input` and
`_linear_project_with_optional_prequant`. Attention forward
pre-quantizes input once (`quantize_input`), reuses the
`(x_fp4, x_scale, x_global_sf)` tuple for k/v projections when
`context is x` — bit-matches internal's fused path.
## `prepare_for_compile` protocol
[`composed_pipeline_base._maybe_compile_pipeline_module`](file:///home/william5lin/FastVideo/fastvideo/pipelines/composed_pipeline_base.py)
calls `getattr(module, "prepare_for_compile", None)` before invoking
`torch.compile`. Defined as a duck-type protocol — no base class method.
Currently only **Gemma3** implements it (to materialize HF weights
outside Dynamo's tracer). Add to other models that have lazy external
state if you observe compile-time graph breaks.
## Per-component compile flags
`CompileConfig` (in
[`fastvideo/api/schema.py`](file:///home/william5lin/FastVideo/fastvideo/api/schema.py))
gained per-component knobs in `221cb20a`:
```python
@dataclass
class CompileConfig:
enabled: bool = False # master DiT switch
backend: str = "inductor"
fullgraph: bool = False
mode: str | None = None
dynamic: bool | None = None
extras: dict = field(default_factory=dict)
# Per-component overlays, None = inherit master `enabled`
text_encoder_enabled: bool | None = None
vae_enabled: bool | None = None
audio_vae_enabled: bool | None = None
# Per-component kwargs override master when non-empty
dit_kwargs: dict = field(default_factory=dict)
text_encoder_kwargs: dict = field(default_factory=dict)
vae_kwargs: dict = field(default_factory=dict)
audio_vae_kwargs: dict = field(default_factory=dict)
```
**`transformer_refine` is auto-compiled with the master DiT flag.** No
separate `enable_torch_compile_refine` flag — by design, refine inherits
DiT compile state to keep typed surface small. Decoupling would add a
new flag, not repurpose existing ones.
## Quantization commit chain (`will/ltx2_sr_port`)
| Commit | Locks in |
|---|---|
| `365a66c7 feat(quantization): upstream LTX-2 FP4Config with lazy flashinfer` | Public colocation of FP4Config (resolves dreamverse_review Q-6 option 1); flashinfer lazy-imported in loader helper, no public hard-dep |
| `a4760bae fix(api): propagate generic refine_*` | `_resolve_refine_args()` copies generic `refine_*` knobs onto `ltx2_refine_*` runtime carriers; `_randn_ltx2_video_latents` reverts to `torch.randn` to bit-match internal under single-generator inference |
| `221cb20a feat(api): typed per-component CompileConfig` | `CompileConfig` per-component knobs; matching `FastVideoArgs` carriers; compat layer round-trip |
| `6da342ba feat(compile): per-component compile + transformer_refine + prepare hook` | `composed_pipeline_base.post_init` compiles `transformer_refine` alongside `transformer`/`transformer_2`; per-component compile loops; `prepare_for_compile` hook on Gemma3 |
| `42b30bf9 feat(ltx2): wire FP4 inference` (largest) | `nn.Linear` → `ReplicatedLinear` for FP4-eligible LTX2 subset; `quant_config` + `prefix=` plumbing; `_maybe_convert_model_to_nvfp4` helper; `LinearBase` fallback to `UnquantizedLinearMethod`; typed `transformer_quant` resolution |
| `94c983a2 refactor(quant): rename FP4 → NVFP4` | Mechanical rename across config, methods, buffers, tests |
| `c6c14c55 test(nvfp4): lock LTX-2 wiring + typed transformer_quant flow` | 6+4 tests in `test_nvfp4_ltx2_wiring.py` + `test_typed_quant_flow.py` |
| `a5fcd19c [fix]: lazy-import flash_attn 2 fallback in attention backend` | post-handoff: lazy import to avoid hard flash_attn 2 dep |
| `d4ee5be2 [fix]: avoid model.to() round-trip in Gemma encoder forward` | post-handoff: parity / perf fix |
| `156103b9 [fix]: unwrap list-of-generator before torch.randn in LTX-2 latent prep` | post-handoff: parity fix for list-of-generators (was bit-matching only single-generator path) |
## Tests
| Test | Asserts |
|---|---|
| [`fastvideo/tests/ops/quantization/test_nvfp4_ltx2_wiring.py`](file:///home/william5lin/FastVideo/fastvideo/tests/ops/quantization/test_nvfp4_ltx2_wiring.py) (6 tests) | `LTXSelfAttention.to_q/to_k/to_v/to_out` are `ReplicatedLinear`; `NVFP4Config()` attaches `NVFP4QuantizeMethod` on the quantized subset with correct `layer_prefix`; non-tagged projections (cross-attn K/V, audio attn, audio FFN) fall back to `UnquantizedLinearMethod`; `BasicAVTransformerBlock` propagates `quant_config`+`prefix` correctly to all 4 attention modules + FFN |
| [`fastvideo/tests/api/test_typed_quant_flow.py`](file:///home/william5lin/FastVideo/fastvideo/tests/api/test_typed_quant_flow.py) (4 tests) | typed `engine.quantization.transformer_quant: "NVFP4"` → `NVFP4Config()` instance flow; default leaves `transformer_quant` None; explicit `dit_config.quant_config = ...` wins over typed carrier |
CPU-only by design; do NOT exercise actual FP4 kernels (no flashinfer in
CI). Real kernel coverage requires a CI run with flashinfer installed.
## Open follow-ups (quantization-specific)
### #4: Expose `layer_profile` on typed `engine.quantization`
Today `transformer_quant: "NVFP4"` always constructs `NVFP4Config()`
with default `layer_profile="refine"`. To support stage-1 profiles (no
`attn2.to_out`, no cross-modal AV) via typed config, add
`transformer_quant_layer_profile: str | None = None` and thread it
through:
- `fastvideo/api/schema.py` — `QuantizationConfig` field
- `fastvideo/api/compat.py` — typed → flat translation
- `fastvideo/fastvideo_args.py` — `_apply_transformer_quant` consumes it
Dreamverse currently dodges this by setting `NVFP4Config()` directly via
`experimental["pipeline_config"]`. Exposing `layer_profile` removes the
dodge. See [open-threads.md](open-threads.md) #4.
### #5: Typed `dit_config.quant_config` carrier (replace `experimental["pipeline_config"]`)
Long-term: design a typed home for an in-memory `PipelineConfig`
instance with mutated `dit_config`. Today `compat.py` recognizes the
`pipeline_config` key in `experimental` and threads it through to
`FastVideoArgs.from_kwargs`. This is fine for short-term but not pretty.
Heaviest design work in the open queue. May need Oracle consult.
### #2: AbsMaxFP8 pre-existing test failure
`fastvideo/tests/ops/quantization/test_absmax_fp8.py::test_create_weights_rejects_invalid_dtype`
fails on `main` and on `will/ltx2_sr_port` with the same error
(`AssertionError not raised`). Verified via `git stash` that the
failure pre-dates NVFP4 work.
Either:
- Fix the test (`AbsMaxFP8LinearMethod.create_weights` no longer
asserts on invalid dtype — restore the assert if intentional, or drop
the test).
Self-contained tech debt; small fix.
## Don't / Cautions
- **Don't change `NVFP4Config` buffer names back to `_fp4_*`.** Rename
is intentional to disambiguate from MX-FP4 / OCP-FP4.
- **Don't remove the `LinearBase` `UnquantizedLinearMethod` fallback.**
Load-bearing for non-tagged layers when a `quant_config` is set.
- **Don't repurpose `enable_torch_compile` to mean DiT-only.** It also
drives `transformer_refine` and `transformer_2` compile.
- **Don't bypass the typed surface for new options.** New compile /
quant / refine knobs should land on the dataclass + compat.py +
parity inventory together. The existing test suite locks this in.
- **Don't merge to main without a CI run that covers FP4.** Current CI
doesn't run flashinfer-dependent paths; the wiring tests are CPU-only
by design.
@@ -1,385 +0,0 @@
# Runbook — How to Do Work in This Scope
Operational how-to for the dreamverse-integration scope. Read after
[state.md](state.md) and [open-threads.md](open-threads.md).
For design rationale see [design.md](design.md). For who to credit see
[authors.md](authors.md). For PR status see [pr-roadmap.md](pr-roadmap.md).
**Last updated:** 2026-05-05 (strategy reversed to single mega-PR #1288 on `will/ltx2_sr_port`; #1287 closed; STACK.md split model deprecated per [decisions-log.md D-17](decisions-log.md#d-17)).
## Worktree contract
```
Repo: /home/william5lin/FastVideo
Branch: will/ltx2_sr_port
```
Other agents and the user share this worktree concurrently. If `git status`
shows changes you don't recognize, they belong to **someone else's work** —
don't revert, don't `git stash drop`, don't `git checkout -- <file>`.
Switch to `will/ltx2_sr_port` cleanly with `git checkout will/ltx2_sr_port`
(safe if your own working tree is clean) and proceed.
If your task requires a different branch (e.g. cherry-pick to
`will/api_7.9` for PR #1286 propagation), return to `will/ltx2_sr_port`
when done — that is the assumed default.
## Branch topology (single mega-PR model)
The dreamverse-integration work now ships as one PR (#1288) off
`will/ltx2_sr_port`. The split-PR model documented in earlier revisions
of this runbook (and in top-level `STACK.md`) is **abandoned** —
see [decisions-log.md D-17](decisions-log.md#d-17).
```
origin/main
↓ [public-API refactor: PRs 0..7.9 merged on main, latest #1286 = 2aaeee2a]
will/ltx2_sr_port (**PR #1288 head** — single mega-PR, 34 commits, 71 files, +13,074/-583)
```
| Branch | Role | Status |
|---|---|---|
| `will/ltx2_sr_port` | **PR #1288 head**, default working branch | OPEN, MERGEABLE |
| `will/api_7.10` / `will/api_8` / `will/ltx2_sr_runtime` / `will/ltx2_nvfp4` / `will/ltx2_post_fixes` / `will/agents_cleanup` | deprecated split-PR bookmarks | local-only historical references; safe to delete |
| `will/ltx2_sr_port-pre-1286-rebase` | safety backup | local-only; preserves the 4 commits dropped during the post-#1286 rebase |
**Sanity check:** `git merge-base --is-ancestor origin/main will/ltx2_sr_port`
should exit 0. If it doesn't, the branch is in an unexpected state — read
[state.md](state.md) before continuing.
## After PR #1288 merges
When the mega-PR squash-merges into `main`:
1. `git fetch origin main` to pull the merge commit.
2. The entire `will/ltx2_sr_port` content is now on main; the branch can
be deleted (locally + on origin) once all consumers are notified.
3. Delete deprecated split bookmarks: `git branch -D will/api_7.10
will/api_8 will/ltx2_sr_runtime will/ltx2_nvfp4 will/ltx2_post_fixes
will/agents_cleanup` (local-only, no remote).
4. Optionally remove top-level `STACK.md` (now a historical artifact).
Keep [co-authors.md](co-authors.md) — still the canonical roster reference.
5. Decide whether to keep `will/ltx2_sr_port-pre-1286-rebase` (safety
backup of the pre-rebase chain) — recommend deleting once #1288 is
merged and verified on main.
6. Update memory dir to reflect the post-merge state — bump
`Last reconciled` headers, mark Item D resolved in
[open-threads.md](open-threads.md), record the merge commit in
[decisions-log.md](decisions-log.md).
## Historical: split-PR re-slice protocol (deprecated)
Prior revisions of this runbook documented a 10-step re-slice protocol
for the abandoned 6-PR split model. That protocol is now obsolete.
The post-#1286 rebase (2026-05-05) was the last execution of it; details
are preserved in [state.md](state.md) "Post-#1286 rebase summary" and
git history at commit `b34d9704`.
## Verification
### Lint (pre-commit)
```bash
pre-commit run --files <changed-paths...>
```
- Binary: `/home/william5lin/miniconda3/envs/fv-main/bin/pre-commit`.
NOT `.venv/bin/pre-commit` — that doesn't exist in this worktree.
- Auto-applies yapf reformatting; re-stage modified files after.
- Hook chain: yapf → ruff → codespell → mypy → spaces-check.
- Memory dir (`.agents/memory/`) is yapf/ruff/mypy excluded — only
"spaces" runs. Memory edits don't need lint, but DO use UTF-8 and
consistent line endings.
### Tests
Router tests (PR #1286 scope):
```bash
.venv/bin/python -m pytest fastvideo/tests/entrypoints/streaming/test_router.py -v --no-header
```
Stack baseline (May 2 handoff suite — re-run when you change anything in
api/, contract/, or LTX-2 paths):
```bash
.venv/bin/python -m pytest \
fastvideo/tests/api/ \
fastvideo/tests/contract/ \
fastvideo/tests/ops/quantization/test_nvfp4_*.py \
tests/local_tests/pipelines/test_ltx2_pipeline_smoke.py \
-q --no-header
```
Expected baselines:
- May 2 handoff (`156103b9`): 222 passed, 1 skipped.
- Post-D-16 (`a152cb77` / `09647a30`): +7 router tests pass on top.
### LSP
Use `lsp_diagnostics` on changed files BEFORE running build. Pre-existing
warnings to ignore (predate this work):
- `fastvideo/entrypoints/streaming/router/main.py:37` — `Task` generic.
- `fastvideo/entrypoints/cli/router_serve.py:55` — `_SubParsersAction` generic.
### gh CLI for PR status
```bash
# PR #1286 quick status
gh pr view 1286 --json headRefOid,mergeable,statusCheckRollup \
--jq '{headRefOid, mergeable, checks: [.statusCheckRollup[] | {name, status, conclusion}]}'
# All commits in a PR + co-author check
gh pr view 1286 --json commits \
--jq '.commits[] | {oid: .oid[0:8], msg: .messageHeadline, author: .authors[0].login}'
```
## Commit workflow
### Subject convention
`[type] <scope>: <imperative summary>` — keep ≤ 72 chars.
Types observed in this scope: `feat`, `fix`, `test`, `docs`, `chore`,
`refactor`. Scopes observed: `streaming`, `dreamverse-integration`,
`api`, `quant`, `ltx2`, `nvfp4`, etc.
Examples:
- `[fix] streaming: router polish — bridge cancel + state machine + deps`
- `[docs] dreamverse-integration: add authors.md + track D-16 router polish`
### Body convention
Bullet list, one bullet per file or concern. Why-before-what. Wrap at
~80 chars (yapf doesn't reformat commit messages; readability is on you).
### Co-author trailers (REQUIRED on every commit)
The 4 trailers in [authors.md](authors.md) MUST appear on every commit
in this scope. Use `--trailer` flags or write the body to a file with
`-F` — DO NOT use multiple `-m` blocks for the trailers (each `-m` is
its own paragraph and git's trailer parser only reads the LAST paragraph,
yielding 1 trailer parsed instead of 4).
**Inline `--trailer` form (preferred for short commits):**
```bash
git commit -m "subject" -m "body..." \
--trailer "Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>" \
--trailer "Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>" \
--trailer "Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>" \
--trailer "Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>"
```
**File form (preferred for multi-paragraph bodies):**
```bash
cat > /tmp/opencode/msg.txt <<'EOF'
[type] scope: subject
* Bullet one with rationale.
* Bullet two with rationale.
Co-authored-by: Junda (David) Su <90978028+Davids048@users.noreply.github.com>
Co-authored-by: Matthew Noto <99706358+RandNMR73@users.noreply.github.com>
Co-authored-by: XOR-op <17672363+XOR-op@users.noreply.github.com>
Co-authored-by: Zhang Peiyuan <42993249+jzhang38@users.noreply.github.com>
EOF
git commit -F /tmp/opencode/msg.txt
```
The trailers MUST be a single block at the end of the message with no
blank lines between them.
**Verify trailers parsed:**
```bash
git log -1 --format='%(trailers:key=Co-authored-by,valueonly)'
```
Should print 4 lines (one per author). If only 1 line, you have the
multi-`-m` bug — amend with `-F` to fix (allowed if commit is unpushed
and you authored it in this session per AGENTS.md amend rules).
### NEVER add to commits
Per [`AGENTS.md`](../../../AGENTS.md):
- AI co-authors (Claude, GPT, Codex, Cursor, etc.) — explicitly forbidden
- "Generated with Claude Code" footer — explicitly forbidden
- `--no-verify` to skip pre-commit — explicitly forbidden
## Push + PR propagation
### Pushing `will/ltx2_sr_port` (top of stack)
```bash
git push origin will/ltx2_sr_port # fast-forward, no force needed
```
If git wants to force-push, you've rewritten history. STOP and verify:
```bash
git log origin/will/ltx2_sr_port..will/ltx2_sr_port # local-only commits
git log will/ltx2_sr_port..origin/will/ltx2_sr_port # remote-only commits
```
Force-push requires explicit user confirmation per `AGENTS.md`.
### Propagating fixes to PR #1286 (`will/api_7.9`)
When a fix is in router code (`fastvideo/entrypoints/streaming/router/`,
`cli/router_serve.py`, `tests/entrypoints/streaming/test_router.py`,
or `pyproject.toml` router-related), it must land on BOTH branches.
Cherry-pick avoids any force-push:
```bash
# 1. Commit on will/ltx2_sr_port first (working branch)
git add <files...>
git commit -F /tmp/opencode/msg.txt # with trailers per above
# 2. Cherry-pick onto will/api_7.9 (creates a separate SHA, identical diff)
git checkout will/api_7.9
git cherry-pick <ltx2_sr_port-sha>
git push origin will/api_7.9 # fast-forward, no force
# 3. Return to working branch
git checkout will/ltx2_sr_port
# 4. Verify PR #1286 picked it up
gh pr view 1286 --json headRefOid --jq '.headRefOid'
```
Two SHAs for the same diff — they'll dedupe naturally on the next
bulk-rebase via the trailer-injection rebase command in
[authors.md](authors.md).
### When a fix is memory-dir-only
`.agents/memory/dreamverse-integration/` lives in the `agents_cleanup`
layer of the stack — it does NOT belong on `will/api_7.9`. Memory updates
stay on `will/ltx2_sr_port` only.
### When a fix is non-router code in the integration scope
Land on `will/ltx2_sr_port`. If that fix needs to ship as a separate PR
(e.g. extending PR 7.10 or starting PR 9), open a new branch off the
right base per [pr-roadmap.md](pr-roadmap.md).
## Memory dir maintenance
When state changes, update the memory dir BEFORE moving on. Every file
has a "Last updated" header — bump when you edit.
| Change | File to update |
|---|---|
| Branch tip moves | [state.md](state.md) "Branch tips" + "Last reconciled" |
| PR opens / merges | [pr-roadmap.md](pr-roadmap.md) status table |
| New decision made | [decisions-log.md](decisions-log.md) — add D-N entry, bump header |
| Open thread resolved | [open-threads.md](open-threads.md) — strikethrough + "Resolved" note |
| New open thread | [open-threads.md](open-threads.md) — priority overview + section |
| New collaborator credited | [authors.md](authors.md) roster + trailer block + bulk-rebase |
| Source doc archived | [source-archive/README.md](source-archive/README.md) + [README.md](README.md) sources table |
| Process / runbook detail changes | [runbook.md](runbook.md) (this file) |
Cross-link siblings via relative paths. Never duplicate content — link.
## Common pitfalls
### `pre-commit` not in `.venv/bin`
`pre-commit` lives at `/home/william5lin/miniconda3/envs/fv-main/bin/pre-commit`.
The `.venv` here is for the FastVideo package itself, not pre-commit.
### Trailers split across paragraphs
`git commit -m A -m B -m C` makes A, B, C separate paragraphs. Git's
trailer parser only reads the LAST paragraph — multiple `-m
"Co-authored-by: ..."` produces 1 trailer parsed, not 4. Use `--trailer`
flags or `-F` with the trailers in a single block at the end.
### Stash 0 on FastVideo IS NOT yours
`stash@{0}: WIP on main: 71bfc13d HunyuanVideo plugin` predates this work.
**DO NOT POP.** See [state.md](state.md) "Stashes — DO NOT POP".
### `AbsMaxFP8` test "failure" is pre-existing
`fastvideo/tests/ops/quantization/test_absmax_fp8.py::test_create_weights_rejects_invalid_dtype`
fails on `main` and on every branch in this scope. NOT introduced by
integration work. See [open-threads.md](open-threads.md) item #2.
### Untracked nested clones at repo root
`dynamo/`, `ray/`, `vllm-omni/` are untracked nested git clones at the
FastVideo repo root. Reference repos for cross-repo work. **Do not
`rm -rf`** — they're someone else's working state.
### Live services on 8009 / 5274
`dreamverse-server` runs on 8009 (warmed GPU worker), Next.js dev server
on 5274. Don't start new instances on those ports without checking
[state.md](state.md) "Live services" first.
### Branch may have been switched by another agent
Other agents share this worktree. If `git branch --show-current` returns
something other than `will/ltx2_sr_port`, switch back cleanly with
`git checkout will/ltx2_sr_port` — don't disturb their work, don't
discard their uncommitted changes.
### Force-push policy
Per `AGENTS.md`: never force-push without explicit user confirmation.
For trailer fixes on already-pushed commits, prefer the bulk-rebase
command in [authors.md](authors.md) — safe to re-run.
### Two trailerless commits in PR #1286
`a152cb77` (on `will/api_7.9`) and `40e265b8` (now-superseded ancestor
on `will/ltx2_sr_port`) lack the 4 co-author trailers. **Accepted gap**
per user decision — see [authors.md](authors.md) "Known gaps".
## Self-test (verify your context is loaded)
After reading the memory dir, you should be able to answer:
1. What branch should I be on? → `will/ltx2_sr_port`
2. What's the active open PR in this scope? → #1286 on `will/api_7.9`
3. Where does PR #1286 land in the stack? → Bottom; ancestor of `will/ltx2_sr_port`
4. Who do I credit on every commit? → 4 authors per [authors.md](authors.md)
5. Where do memory updates land? → `will/ltx2_sr_port` only (NOT api_7.9)
6. What's the next-priority open thread? → See [open-threads.md](open-threads.md) "Recommended pull order" — D-8 verify is current top
7. What pre-existing failure can I ignore? → AbsMaxFP8 test (item #2)
8. What's the bulk-rebase command for adding trailers across the stack? → See [authors.md](authors.md) "How the trailers were applied"
If you can't answer one of these from the memory dir alone, the dir has
a gap — file it as a new entry in [open-threads.md](open-threads.md)
before continuing.
## First 60 seconds — copy-paste orientation
```bash
# 1. Confirm branch
cd /home/william5lin/FastVideo
git branch --show-current # should print: will/ltx2_sr_port
# If not, recover: git checkout will/ltx2_sr_port
# 2. Confirm worktree clean (untracked nested clones expected)
git status --short
# 3. Confirm PR #1286 head matches expected api_7.9 tip
gh pr view 1286 --json headRefOid --jq '.headRefOid'
git rev-parse will/api_7.9 # should match PR head
# 4. Confirm your context vs the memory dir
git log -1 --oneline
cat .agents/memory/dreamverse-integration/state.md | head -30
# 5. Confirm live services still running
curl -s http://localhost:8009/readyz | head -c 200
curl -s http://localhost:5274/ -o /dev/null -w "%{http_code}\n"
```
If any of those produce unexpected output, read [state.md](state.md)
before changing anything.
File diff suppressed because it is too large Load Diff
@@ -1,44 +0,0 @@
# Source Archive
These are the original unsynthesized design and integration docs that
predate the consolidation in
[`../`](../). They are **NOT** the source of truth — the synthesized
sibling files in the parent directory are.
Archived 2026-05-03. All previously untracked.
## Contents
| File | Original location | Date | Synthesized into |
|---|---|---|---|
| `apirefactor.md` | `FastVideo/` (repo root) | 2026-04-21 | [`../design.md`](../design.md) |
| `PR-plan.md` (was `PR plan.md` at repo root) | `FastVideo/` (repo root) | 2026-04-25 | [`../pr-roadmap.md`](../pr-roadmap.md) |
| `dreamverse_review.md` | `FastVideo/` (repo root) | 2026-04-26 | [`../decisions-log.md`](../decisions-log.md) + [`../state.md`](../state.md) |
| `handoff-nvfp4-launch-demo.md` | `.agents/exploration/` | 2026-05-02 | [`../state.md`](../state.md) + [`../quantization.md`](../quantization.md) + [`../open-threads.md`](../open-threads.md) |
| `streaming-server-upstream-plan.md` | `.agents/exploration/` | 2026-04-17 | [`../streaming-server.md`](../streaming-server.md) + [`../decisions-log.md`](../decisions-log.md) |
| `dreamverse_integration.md` | `.agents/exploration/` | 2026-04-23 | [`../cross-repo-surfaces.md`](../cross-repo-surfaces.md) |
| `video-generator-config-api-design.md` | `.agents/exploration/` | 2026-04-02 | [`../design.md`](../design.md) (early-draft material) |
## Why archived (not deleted)
- Future agents may want the **full unsynthesized rationale** for a
decision the synthesis abbreviated.
- The originals remain useful as a **time machine** for understanding
how the design evolved.
- These docs were never committed to git, so leaving them on disk costs
nothing.
## When to read the archive vs. the synthesis
- **Read the synthesis (`../*.md`)** for: current state, decision
status, action items, design rationale at the conceptual level.
- **Read the archive (here)** for: deep historical context, exact wording
of design decisions, full PR plan with all sub-PR commit details,
the original Q-1..Q-9 / D-1..D-11 prose.
## Maintenance rule
Do NOT edit files in this archive. They are point-in-time snapshots.
If new design material appears that supersedes an entry here, update the
synthesis (the parent dir) and append a note to that synthesis file —
do not mutate this archive.
@@ -1,838 +0,0 @@
# FastVideo API Refactor Design
## Related Documents
- [PR plan.md](PR%20plan.md) — PR-by-PR implementation plan for this design
- [.agents/exploration/streaming-server-upstream-plan.md](.agents/exploration/streaming-server-upstream-plan.md) — streaming-server upstream + Dynamo backend contract (shapes PRs 5.5-7.10)
- `../FastVideo-internal/.agents/exploration/rebase-upstream-fastvideo.md` — rebasing FastVideo-internal onto upstream (enables PRs 6-8)
- `../FastVideo-internal/ui/ltx2-streaming/` — source for the streaming server being upstreamed (PRs 7.5-7.9)
- `../dynamo/` — local clone of ai-dynamo/dynamo; `components/src/dynamo/sglang/` is the template for FastVideo's native backend landed in PR 7.10
- https://github.com/ai-dynamo/dynamo/pull/7544 — closed draft PR that establishes the Dynamo backend shape this design must satisfy
## Status
Design spec for the public inference API refactor. PRs 0-5.5 are landed; see [PR plan.md](PR%20plan.md) for rollout status and the PR 6+ roadmap. The typed schema, strict parser, preset system, typed VideoGenerator, typed CLI, and stateless OpenAI server default-request merge are all implemented. Streaming package skeleton + typed streaming config types are in place; live streaming server + Dynamo contract are the next milestones.
## Executive Summary
FastVideo should move to a single typed nested inference schema that is shared across:
- Python API
- CLI
- YAML/JSON config files
- OpenAI/server request translation
The core split is:
- `GeneratorConfig`: generator-instance lifetime settings
- `GenerationRequest`: per-call inputs, sampling, outputs, and continuation
- `InferencePreset`: model-owned named multi-stage defaults
The canonical user experience should be:
1. Choose a model.
2. Choose a pipeline preset.
3. Override a few typed fields.
4. Generate.
FastVideo should not make a raw free-form string dict the primary API. Dicts and YAML/JSON should be supported as serialization/interchange layers, but they must be parsed immediately into typed config objects with strict unknown-key validation.
The repo should also shift model-specific preset/default definitions closer to their pipeline implementations, while keeping the shared public schema and parsers centralized.
## Why This Refactor Is Needed
Today the public inference boundary is too flat and too forgiving.
- `VideoGenerator.from_pretrained(..., **kwargs)` mixes:
- engine/runtime settings
- pipeline init settings
- component overrides
- `VideoGenerator.generate_video(..., **kwargs)` mixes:
- prompt and inputs
- sampling parameters
- output settings
- model-specific workflow knobs
- unknown or drifting keys can be silently filtered or merely logged instead of failing fast
- model-specific multi-stage behavior is exposed through ad hoc top-level flags instead of a stable preset/stage abstraction
This is already painful in LTX2/Dreamverse, and it will get worse as more multi-stage pipelines are upstreamed.
## Design Goals
- Keep the Python API typed and editor-friendly.
- Make YAML/JSON a first-class serialization of the same schema.
- Support CLI overrides cleanly without flattening the schema into hundreds of canonical flags.
- Separate init-time config from request-time config.
- Provide a stable public abstraction for multi-stage pipelines.
- Support LTX2 two-stage and continuation behavior cleanly.
- Keep the simple case simple.
- Co-locate model-owned defaults and stage topology with the relevant pipeline.
- Protect current public/server behavior with an explicit schema parity audit before freezing the new surface.
- Preserve backward compatibility long enough to migrate examples, internal users, and servers safely.
## Non-Goals
- Do not make Ray a structural dependency or copy its package layout.
- Do not make a raw free-form dict the primary Python API.
- Do not force all models into one universal `RefineConfig`.
- Do not expose stage indices as the primary user interface.
- Do not move every shared config class into per-model directories.
## External Inspiration
### Ray
Borrow only the ergonomic idea that user-facing config can be expressed as a string-keyed dict or YAML/JSON config. Do not copy Ray's structure into FastVideo.
### SGL Multimodal Gen
Useful ideas: split instance config from request config; allow dict input at the boundary; parse dicts immediately into typed request objects; merge request overrides onto model defaults; validate request params against pipeline/task type. Do not copy: request objects depending on server/engine config; broad weakly typed request bags as the canonical API.
### vLLM-Omni
Useful ideas: model-owned pipeline presets; explicit stage topology; per-stage default sampling params; clean separation between stage topology, engine defaults, and runtime overrides. Do not copy: positional `sampling_params_list` as the primary public API; serving-engine-oriented stage index semantics in the main Python interface.
## Core Decision
FastVideo should have:
1. A shared typed public schema.
2. Model-owned named pipeline presets.
3. Semantic stage overrides by stage name.
4. Optional advanced explicit plans for power users.
5. YAML-first config loading with dotted CLI overrides.
The public API should be stable at the schema level, while model-specific behavior should be contained in preset definitions and model-specific typed override classes.
## Schema Parity Requirement
Before the new schema is declared canonical, FastVideo should build a parity inventory across all current public inference surfaces (Python `VideoGenerator` kwargs, CLI flags, YAML/JSON config inputs, OpenAI/server request models, model-specific sampling/runtime fields). Each field must be marked: kept as-is, renamed, moved to a nested path, preset-owned, private-only adapter field, or intentionally dropped. No field should disappear implicitly.
For any public field that remains supported, there should be either a normalized-config equivalence test, or an explicit parser/translation test. Fields that exist only in private Dreamverse integration code should be handled by a private adapter layer, not quietly converted into public FastVideo compatibility guarantees.
Landed artifact: [inference_schema_parity_inventory.yaml](docs/design/inference_schema_parity_inventory.yaml) + guard [test_schema_parity_inventory.py](fastvideo/tests/api/test_schema_parity_inventory.py).
## Canonical Public Schema
The typed schema is implemented in [fastvideo/api/schema.py](fastvideo/api/schema.py). Envelope types:
- `RunConfig` — offline: `generator` (GeneratorConfig) + `request` (GenerationRequest)
- `ServeConfig` — serving: `generator` + `server` (ServerConfig) + `default_request` (GenerationRequest) + optional `streaming` (StreamingConfig)
Key nested types (summary; full fields in `schema.py`):
- `GeneratorConfig` → `model_path`, `revision`, `trust_remote_code`, `engine` (EngineConfig: parallelism/offload/compile/quantization/flags), `pipeline` (PipelineSelection: workload_type, preset, preset_version, components, preset_overrides, experimental)
- `GenerationRequest` → `prompt`, `negative_prompt`, `inputs` (InputConfig), `sampling` (SamplingConfig), `runtime` (RequestRuntimeConfig), `output` (OutputConfig), `stage_overrides`, `state` (ContinuationState), `plan` (GenerationPlan), `extensions`
- `ContinuationState` → opaque `{kind: str, payload: dict[str, Any]}`
- `GenerationPlan` → `{stages: list[PlannedStage], final_stage: str | None}`; advanced/escape-hatch only
### Important Semantics
- Dataclasses are canonical for Python users.
- Dict and YAML/JSON are parsed into these dataclasses immediately.
- Unknown keys must raise validation errors.
- Typed `GenerationRequest` defaults come from the public schema, not from model-specific `SamplingParam.from_pretrained(...)` defaults.
- Legacy `generate_video(...)` continues to inherit model-specific sampling defaults until the SSIM/performance migration lands (PR 11).
- The only open-ended escape hatches are:
- `generator.pipeline.experimental`
- `request.extensions`
That keeps the public contract strict without blocking experimental work.
### Request Mutation Tracking
When a `GenerationRequest` is parsed from a raw dict (YAML, JSON, or Python mapping), FastVideo records which fields the user explicitly provided versus which received schema defaults. This matters because `request_to_sampling_param()` must distinguish user-provided values (which should override model defaults) from schema defaults (which should NOT override model defaults).
The tracking contract:
- At parse time, the original raw dict and a baseline snapshot of the parsed object are stored on the request.
- Dataclass field mutations after parsing (e.g., `request.sampling.seed = 7`) are captured via lightweight `__setattr__` dirty-path recording.
- Dict-typed field mutations (e.g., `del request.stage_overrides["refine"]`) are detected at access time by diffing the current dict against the baseline snapshot.
- Setting a field to the schema default value IS captured as explicit, so it will override model defaults.
- The raw dict is reconciled lazily when `normalize_generation_request()` is called, not on every individual mutation.
### Schema Purity and Model-Specific Fields
The shared schema currently contains fields that are specific to one or two model families. These remain for backward compatibility during the initial migration (PRs 0-3) but should migrate to preset-owned typed override classes as the preset system lands (PRs 4-10).
**SamplingConfig fields to migrate:**
- `height_sr`, `width_sr`, `num_inference_steps_sr`: Hunyuan15 SR only. Target: `HunyuanSRStageOverride` in PR 10.
- `guidance_scale_2`, `boundary_ratio`: Wan2.2 and LingBotWorld only. Target: preset-owned overrides in the relevant model migration PR.
**InputConfig fields to migrate:**
- `mouse_cond`, `keyboard_cond`, `grid_sizes`: MatrixGame action control only. Target: `request.extensions` or a typed MatrixGame input config.
- `c2ws_plucker_emb`: LingBotWorld camera control only. Target: `request.extensions` or a typed LingBotWorld input config.
- `refine_from`, `stage1_video`: LongCat refinement only. Target: `LongCatRefineStageOverride` inputs or keep in `InputConfig` if they remain a public contract.
**Universal fields that stay in the shared schema:**
- `guidance_rescale`: used by multiple denoising stages across models, default 0.0. Universally applicable.
- `true_cfg_scale`: OpenAI adapter surface. Keep for protocol compatibility.
### Escape Hatch Sunset
`generator.pipeline.experimental` and `request.extensions` are intentional escape hatches for experimental and private work. They bypass strict validation by design.
Rules for escape hatch usage:
- New fields should not be added to `experimental` or `extensions` without a plan to either promote them to typed fields or remove them within two PR cycles.
- Each model migration PR (PRs 6-10) should review and shrink escape hatch usage for that model family.
- The compatibility layer currently routes unrecognized legacy kwargs into `experimental`. This pass-through should shrink as presets absorb model-specific fields.
## Public Python API
### New Canonical API
```python
from fastvideo import VideoGenerator
from fastvideo.api import (
GeneratorConfig, GenerationRequest,
EngineConfig, OutputConfig,
PipelineSelection, SamplingConfig,
)
generator = VideoGenerator.from_pretrained(
config=GeneratorConfig(
model_path="/models/ltx2",
engine=EngineConfig(num_gpus=1),
pipeline=PipelineSelection(
workload_type="t2v",
preset="ltx2_two_stage",
),
)
)
result = generator.generate(
GenerationRequest(
prompt="a fox running through snow",
sampling=SamplingConfig(
num_frames=121, height=1024, width=1536,
num_inference_steps=8, seed=42,
),
output=OutputConfig(save_video=True, return_state=True),
)
)
```
### Accepted Construction Forms
Canonical:
```python
VideoGenerator.from_pretrained(config=GeneratorConfig(...))
VideoGenerator.from_config(GeneratorConfig(...))
VideoGenerator.from_file("run.yaml")
```
Stable convenience constructor:
```python
VideoGenerator.from_pretrained("model-id")
VideoGenerator.from_pretrained("model-id", num_gpus=2, use_fsdp_inference=False, ...)
```
Legacy compatibility:
```python
VideoGenerator.from_pretrained(model_path, **legacy_kwargs)
```
All constructor forms normalize through the same typed path. Stable convenience kwargs remain supported with no deprecation warning. Advanced model/pipeline-specific kwargs are accepted during migration but only as compatibility inputs that normalize into `GeneratorConfig`. The thing being deprecated over time is the unbounded legacy kwarg surface, not the `from_pretrained(...)` entrypoint itself.
### Generation Entry Point
Canonical: `generator.generate(request: GenerationRequest) -> GenerationResult`.
Compatibility alias: `generator.generate_video(prompt=..., **legacy_kwargs)` — converts legacy calls into a `GenerationRequest` and emits a deprecation warning.
During the compat period, `generate(request=...)` uses schema defaults while `generate_video(...)` preserves legacy model-default behavior. These paths intentionally differ until preset-owned defaults replace the remaining `SamplingParam` default logic (migrated in PR 11).
### Boundary Normalization Rule
Every public inference entrypoint normalizes into typed config objects before touching legacy internals. That includes Python constructors, generation calls, CLI `generate`, CLI `serve`, and OpenAI/server request translation. Legacy internals (`FastVideoArgs`, `SamplingParam`) may remain temporarily, but only behind a typed normalization boundary.
## Pipeline Presets
### Definition
An `InferencePreset` is a named model-owned preset that defines:
- workload selection
- stage topology
- per-stage defaults
- stage names
- allowed stage override types
- init-time feature requirements
The preset is not user-authored by default. It is supplied by the model integration.
### Why Presets Are The Right Abstraction
Users usually do not want to assemble a stage graph by hand. They want to say:
- use LongCat distill + refine
- use Hunyuan 1080p SR
- use LTX2 two-stage continuation mode
Presets provide a stable public noun for that behavior.
### Preset Naming Rules
- Use semantic names, not stage indices.
- Keep names stable across releases.
- If semantics change incompatibly, change `preset_version` or create a new preset name.
Examples: `ltx2_base`, `ltx2_two_stage`, `longcat_distill_refine`, `hunyuan15_sr_720p`, `hunyuan15_sr_1080p`.
### Preset-Owned Stage Names
Stage names are public and stable within a preset.
- LTX2: `base`, `refine`
- LongCat: `distill`, `refine`
- Hunyuan15: `base`, `sr_720p`, `sr_1080p`
Public overrides should reference these stage names, never stage indices.
## Stage Overrides
The main user override surface for multi-stage pipelines is:
```python
request.stage_overrides["refine"] = ...
```
Each model family should expose typed override classes for its stage names. Examples for the model families that land in PRs 6/9/10:
```python
@dataclass
class LTX2RefineStageOverride:
enabled: bool | None = None
num_inference_steps: int | None = None
guidance_scale: float | None = None
add_noise: bool | None = None
image_crf: int | None = None
video_position_offset_sec: float | None = None
@dataclass
class LongCatRefineStageOverride:
t_thresh: float | None = None
spatial_refine_only: bool | None = None
num_cond_frames: int | None = None
@dataclass
class HunyuanSRStageOverride:
num_inference_steps: int | None = None
guidance_scale: float | None = None
```
### Strictness Rules
- Stage names must exist in the selected preset.
- Override fields must be valid for that stage type.
- Unknown stage names and unknown fields must error.
## Advanced Explicit Plans
Presets should be the default API. `GenerationPlan` exists only for advanced composition or experimentation:
- building a custom workflow that is not yet standardized as a preset
- debugging or benchmarking stage combinations
- prototyping a future preset
Do not require `GenerationPlan` for normal users.
## Continuation State
Continuation must be a first-class part of the API.
### Public Contract
- `GenerationResult.state` may return a `ContinuationState`.
- `GenerationRequest.state` may accept a previously returned state.
- Most users should treat `state` as opaque and round-trip it back into the next request.
### Why This Matters
Dreamverse/LTX2 currently leaks continuation internals into app-level request fields like video conditions, audio clean latent, audio denoise mask, and segment offsets. Those should not remain top-level app-owned public API.
### State Design
Public surface:
```python
@dataclass
class ContinuationState:
kind: str
payload: dict[str, Any]
```
Internally, FastVideo should also define typed model-specific state subclasses, e.g. `LTX2ContinuationState` (PR 7) and `LongCatIntermediateState` if ever needed. Minimal stable surface: return state, pass state back in, validate that the state is compatible with the active preset.
Payload serialization: fields must be JSON-serializable or use an opaque blob-ID indirection for large tensors. This supports both the stateless OpenAI client round-trip AND future Dynamo prefill/decode disaggregation where prefill yields a state that decode hydrates across workers.
## YAML / JSON Design
YAML and JSON should be exact serializations of the typed schema, not a second unrelated config system. YAML is the primary documented format. JSON is accepted with the same schema.
### Run Config Example
```yaml
generator:
model_path: /models/ltx2
engine:
num_gpus: 1
parallelism: {tp_size: -1, sp_size: -1}
offload: {dit: false, text_encoder: false, vae: false, pin_cpu_memory: true}
pipeline:
workload_type: t2v
preset: ltx2_two_stage
components:
config_root: /models/ltx2-config
upsampler_weights: /models/ltx2-refine
lora_path: /models/ltx2-refine-lora
preset_overrides:
refine: {enabled: true, add_noise: true}
request:
prompt: "a fox running through snow"
sampling:
num_frames: 121
height: 1024
width: 1536
num_inference_steps: 8
seed: 42
output: {save_video: true, return_state: true}
stage_overrides:
refine: {num_inference_steps: 2, guidance_scale: 1.0}
```
### Serve Config Example
```yaml
generator:
model_path: /models/ltx2
engine: {num_gpus: 1}
pipeline: {workload_type: t2v, preset: ltx2_two_stage}
server: {host: 0.0.0.0, port: 8000, output_dir: outputs/}
default_request:
sampling: {num_frames: 121, height: 1024, width: 1536, num_inference_steps: 8}
output: {save_video: false, return_frames: false}
```
### Validation Rules
- top-level schema must match `RunConfig` or `ServeConfig`
- unknown keys must fail
- dotted CLI overrides are applied to the nested config before typed parsing
- parse errors must include the exact nested path that failed
## CLI Design
Inference CLI reuses the best parts of the current training authoring flow (YAML-first authoring, dotted nested overrides, typed parsing after merge) but stays stricter than training at the public boundary because it is a user-facing API surface for Python, CLI, YAML/JSON, and serving.
### Canonical CLI Forms
```bash
fastvideo generate --config run.yaml
fastvideo generate --config run.yaml --request.sampling.seed 42
fastvideo generate --config run.yaml --generator.engine.num_gpus 2
fastvideo serve --config serve.yaml
fastvideo serve --config serve.yaml --server.port 8090
```
The CLI is config-only. Beyond `--config`, CLI input uses dotted override paths into the nested schema rather than maintaining a second flat flag surface.
Implementation: YAML/JSON is loaded into a nested dict, dotted CLI overrides are applied to the nested dict, then the result is parsed into typed config objects. Flat CLI flags are rejected so the nested schema stays canonical.
## OpenAI / Server Mapping
`fastvideo serve` loads `ServeConfig`. Incoming HTTP requests are translated into `GenerationRequest` by:
1. cloning `default_request`
2. applying API request fields onto that request
3. validating against the selected preset
This is similar in spirit to the SGL pattern of merging user overrides onto model defaults.
Rules:
- HTTP request translation must not bypass typed validation.
- multi-stage defaults should come from the preset and `default_request`, not from ad hoc server-local logic.
- stateful continuation requests should accept and return typed `ContinuationState` payloads.
Landed in PR 5 for the stateless OpenAI server at `fastvideo/entrypoints/openai/`. The streaming/session server (PRs 7.5-7.9) uses the same preset/default_request merge through `ServeConfig.streaming`.
## Streaming Server + Dynamo Backend
The typed public API is consumed by three server-class integrations. They must share one execution substrate so we don't grow three near-duplicate progress loops.
### The three consumers
| Consumer | Transport | Request shape | State |
|---|---|---|---|
| Stateless OpenAI (`fastvideo/entrypoints/openai/`) | HTTP POST | `GenerationRequest` merged onto `ServeConfig.default_request` | Stateless; continuation via opaque payload if needed |
| Streaming WebSocket (`fastvideo/entrypoints/streaming/`) | WebSocket JSON + binary fMP4 | `GenerationRequest` per segment, session-scoped | Server-held session (per-GPU continuation cache); snapshot on demand |
| Dynamo native backend (`ai-dynamo/dynamo/components/src/dynamo/fastvideo/`) | Dynamo RPC endpoint | `NvCreateVideoRequest` ↔ adapter ↔ `GenerationRequest` | Aggregated today; disaggregated prefill/decode later via `ContinuationState` |
### Shared execution substrate: `VideoGenerator.generate_async`
The OpenAI server, streaming server, and Dynamo backend all want the same thing: a typed async API that yields progress events and a typed final result. FastVideo exposes exactly one canonical entry point:
```python
async def generate_async(
self,
request: GenerationRequest,
) -> AsyncGenerator[VideoEvent, None]: ...
```
Events:
```python
@dataclass
class VideoProgressEvent:
step: int
total_steps: int
stage: str # "denoise" | "refine" | "decode" | ...
@dataclass
class VideoPartialEvent:
frames: np.ndarray # shape: (num_frames, H, W, 3)
index: int # monotonic chunk index
@dataclass
class VideoFinalEvent:
video_bytes: bytes | None # mp4-encoded if requested
tensor: torch.Tensor | None # raw if requested
metadata: dict[str, Any]
continuation_state: ContinuationState | None
VideoEvent = VideoProgressEvent | VideoPartialEvent | VideoFinalEvent
```
The sync `generate_video(request=...) -> VideoResult` becomes a thin `asyncio.run` wrapper over `generate_async` that collects events and returns the final.
### Streaming server mapping
`fastvideo/entrypoints/streaming/` owns per-session state:
- `SessionStore.hydrate(state: ContinuationState) -> session_id`
- `SessionStore.snapshot(session_id) -> ContinuationState`
- Per-GPU implicit continuation cache (today's internal behavior) is wrapped as a `SessionStore` implementation.
Per-segment, the session writes a `GenerationRequest`, pipes the event stream to the WebSocket (progress → JSON messages, partial → fMP4 frames), and persists the final's `ContinuationState` into the session.
### Dynamo backend mapping
Dynamo's backend pattern (from `components/src/dynamo/sglang/`) is a pure Python import. FastVideo does not host a `fastvideo/entrypoints/dynamo/` subpackage; the integration lives in the Dynamo repo. FastVideo exposes a stable contract:
| Surface | Exposed as |
|---|---|
| Construction | `VideoGenerator.from_pretrained(model_path, **typed_kwargs)` |
| Execution (async) | `VideoGenerator.generate_async(request) -> AsyncGenerator[VideoEvent, None]` |
| Execution (sync) | `VideoGenerator.generate_video(request=...) -> VideoResult` |
| Typed request | `fastvideo.api.GenerationRequest`, `SamplingConfig`, `InputConfig` |
| Typed result | `fastvideo.api.VideoResult`, `VideoEvent`, `ContinuationState` |
| Health-check input | `VideoGenerator.default_health_check_request() -> GenerationRequest` |
| Config dump | `config_to_dict(cfg)` (already exists) |
Request/response mapping the Dynamo adapter must perform:
```
NvCreateVideoRequest -> fastvideo.api.GenerationRequest
prompt -> sampling.prompt
size="WxH" -> sampling.width, sampling.height
seconds -> seconds * nvext.fps -> sampling.num_frames
input_reference -> input.image_path | input.video_path
nvext.fps -> sampling.fps
nvext.num_frames -> sampling.num_frames (overrides seconds*fps)
nvext.num_inference_steps -> sampling.num_inference_steps
nvext.guidance_scale -> sampling.guidance_scale
nvext.seed -> sampling.seed
nvext.negative_prompt -> sampling.negative_prompt
response_format -> (handled at the adapter's output stage)
VideoFinalEvent -> NvVideosResponse
video_bytes -> data[0].b64_json (if response_format=b64_json)
uploaded URL -> data[0].url (if response_format=url)
metadata.inference_time_s -> inference_time_s
continuation_state -> (reserved for future disaggregation)
```
All fields already exist (or will exist after PR 6's typed-kwarg expansion) on FastVideo's typed schema. **The adapter lives entirely in the Dynamo repo** at `components/src/dynamo/fastvideo/` — FastVideo does not host any Dynamo subpackage, dep, or CLI. The only FastVideo obligation is the stable public Python API listed above.
### Constraints this places on other sections
- **Continuation State** (see earlier section): `ContinuationState.payload` must be JSON-serializable or use an opaque blob-ID indirection for large tensors. This supports both the stateless OpenAI client round-trip *and* future Dynamo prefill/decode disaggregation, where prefill yields a state that decode hydrates across workers.
- **Typed GeneratorConfig** (see Public Python API): every flat legacy LTX2 kwarg currently used by the internal `gpu_pool.py` must have a typed home reachable from `GeneratorConfig`. Dynamo's `FastVideoArgGroup` builds the config from its CLI and must not have to know any legacy LTX2 name.
- **Public exports**: `from fastvideo import VideoGenerator`; `from fastvideo.api import GenerationRequest, SamplingConfig, ContinuationState, VideoResult, VideoEvent, VideoProgressEvent, VideoPartialEvent, VideoFinalEvent`.
## Repo Layout
### Shared Public API
`fastvideo/api/` contains the shared public API package. Current files:
- `schema.py` — `RunConfig`, `ServeConfig`, `ServerConfig`, `GeneratorConfig`, and all nested typed config dataclasses
- `sampling_param.py` — `SamplingParam` + `CacheParams` (canonical home since PR 4; former `configs/sample/base.py` location removed)
- `presets.py` — `InferencePreset`, `PresetStageSpec`, registry APIs
- `results.py` — `GenerationResult` / `VideoResult`
- `parser.py` — `from_dict`, `to_dict`, `load_yaml`, `load_json`, validation
- `overrides.py` — dotted override application
- `compat.py` — legacy Python kwargs translation
- `errors.py` — path-aware validation errors
May split further by concern in a future cleanup.
### Pipeline-Local Model-Owned Config
Model-owned presets and override types live next to the model pipeline:
```text
fastvideo/pipelines/basic/ltx2/
ltx2_pipeline.py, presets.py, stage_overrides.py, continuation.py
fastvideo/pipelines/basic/longcat/
longcat_pipeline.py, presets.py, stage_overrides.py
fastvideo/pipelines/basic/hunyuan15/
hunyuan15_pipeline.py, hunyuan15_sr_pipeline.py, hunyuan15_2sr_pipeline.py,
presets.py, stage_overrides.py
```
PR 4 landed `presets.py` for all 13 model families. Remaining colocation targets are `pipeline_configs.py` (moving `configs/pipelines/<family>.py`) and model-specific stages (moving `pipelines/stages/<family>_*.py`); see [PR plan.md](PR%20plan.md) "Pipeline Package Structure".
### Registry
Central registry (`fastvideo/registry.py`) registers preset providers rather than owning all model-specific defaults directly. It answers:
- which pipeline class corresponds to a model path
- which presets are available for that model family
- which override/state classes are valid for a selected preset
## Relationship To Current Internal Classes
This refactor does not require deleting current internals immediately.
- `FastVideoArgs` is an internal compatibility/input adapter, no longer the primary public inference type.
- `SamplingParam` now lives in `fastvideo/api/sampling_param.py` and gets model-specific defaults from presets via `_from_preset()`. All 12 `SamplingParam` subclasses have been removed and the former `fastvideo/configs/sample/` directory has been deleted entirely (PR 4). It remains an internal adapter between the preset system and the runtime.
- current `PipelineConfig` classes can remain temporarily as internal component config carriers
- the new public schema is the stable boundary above them
`VideoGenerator` accepts the new schema and translates down into current execution internals. Legacy `generate_video(..., **kwargs)` stays on the direct execution path during the compat period until SSIM/performance tests migrate in PR 11.
## Model-Specific Design
### LTX2 / Dreamverse
LTX2 needs both:
- init-time two-stage feature wiring
- request-time continuation/refine behavior
Expressed as:
- preset: `ltx2_two_stage`
- init-time fields: refine assets, optional config root, stage enablement
- request-time fields: stage override for refine behavior, optional returned continuation state
#### LTX2 Preset Example
```yaml
generator:
pipeline:
preset: ltx2_two_stage
components:
config_root: /models/ltx2-config
upsampler_weights: /models/ltx2-refine
lora_path: /models/ltx2-refine-lora
preset_overrides:
refine: {enabled: true, add_noise: true}
```
#### LTX2 Request Example
```yaml
request:
prompt: "continue the previous sequence"
state: ${previous_result.state}
stage_overrides:
refine:
num_inference_steps: 2
guidance_scale: 1.0
image_crf: 18
output:
return_state: true
```
#### LTX2 Explicit Decisions
- `config_model_path` becomes `generator.pipeline.components.config_root`
- `ltx2_refine_*` stops being a pile of top-level kwargs
- continuation internals move into `ContinuationState`
- app-level code should pass `state`, not raw latent/audio condition payloads
### LongCat
LongCat should expose a named preset like `longcat_distill_refine` with stage topology `distill` and `refine`.
User-facing override knobs remain model-specific (`t_thresh`, `spatial_refine_only`, `num_cond_frames`) but live under:
```yaml
request:
stage_overrides:
refine:
t_thresh: 0.5
spatial_refine_only: false
num_cond_frames: 8
```
### Hunyuan 1.5 SR
Hunyuan already behaves like an integrated multi-stage pipeline. Expose it via presets: `hunyuan15_sr_720p`, `hunyuan15_sr_1080p`. Users should not need to know the exact internal pipeline class split between base and SR stages. Per-stage override surface should stay small and mostly sampling-focused.
Hunyuan15 presets (`hunyuan15_t2v_480p`, `hunyuan15_i2v_480p_distilled`, `hunyuan15_t2v_720p`, `hunyuan15_i2v_720p_distilled`, `hunyuan15_sr_1080p`) are implemented (PR 4). The `Hunyuan15_*_SamplingParam` subclasses have been removed; defaults (including precomputed sigmas) come from preset `defaults` dicts. Remaining work: adding typed `HunyuanSRStageOverride` classes and colocating PipelineConfig (PR 10).
## Exact Compatibility Mapping
Intended translation layer for common current fields.
| Legacy Field | New Path |
| --- | --- |
| `model_path` | `generator.model_path` |
| `revision` | `generator.revision` |
| `trust_remote_code` | `generator.trust_remote_code` |
| `workload_type` | `generator.pipeline.workload_type` |
| `num_gpus` | `generator.engine.num_gpus` |
| `tp_size` | `generator.engine.parallelism.tp_size` |
| `sp_size` | `generator.engine.parallelism.sp_size` |
| `dit_cpu_offload` | `generator.engine.offload.dit` |
| `dit_layerwise_offload` | `generator.engine.offload.dit_layerwise` |
| `text_encoder_cpu_offload` | `generator.engine.offload.text_encoder` |
| `image_encoder_cpu_offload` | `generator.engine.offload.image_encoder` |
| `vae_cpu_offload` | `generator.engine.offload.vae` |
| `pin_cpu_memory` | `generator.engine.offload.pin_cpu_memory` |
| `enable_torch_compile` | `generator.engine.compile.enabled` |
| `torch_compile_kwargs` | split across `generator.engine.compile.backend`, `.fullgraph`, `.mode`, `.dynamic`; uncommon keys land in `.extras` |
| `enable_torch_compile_text_encoder` | `generator.engine.compile.text_encoder_enabled` |
| `enable_stage_verification` | `generator.engine.enable_stage_verification` |
| `prompt_txt` | `request.inputs.prompt_path` |
| `prompt` | `request.prompt` |
| `negative_prompt` | `request.negative_prompt` |
| `image_path` | `request.inputs.image_path` |
| `video_path` | `request.inputs.video_path` |
| `output_path` | `request.output.output_path` |
| `output_video_name` | `request.output.output_video_name` |
| `save_video` | `request.output.save_video` |
| `return_frames` | `request.output.return_frames` |
| `num_videos_per_prompt` | `request.sampling.num_videos_per_prompt` |
| `seed` | `request.sampling.seed` |
| `num_frames` | `request.sampling.num_frames` |
| `height` | `request.sampling.height` |
| `width` | `request.sampling.width` |
| `fps` | `request.sampling.fps` |
| `num_inference_steps` | `request.sampling.num_inference_steps` |
| `guidance_scale` | `request.sampling.guidance_scale` |
| `guidance_scale_2` | `request.sampling.guidance_scale_2` |
| `guidance_rescale` | `request.sampling.guidance_rescale` |
| `true_cfg_scale` | `request.sampling.true_cfg_scale` |
| `boundary_ratio` | `request.sampling.boundary_ratio` |
| `sigmas` | `request.sampling.sigmas` |
| `enable_teacache` | `request.runtime.enable_teacache` |
| `return_trajectory_latents` | `request.runtime.return_trajectory_latents` |
| `return_trajectory_decoded` | `request.runtime.return_trajectory_decoded` |
### Private Dreamverse Adapter Mapping
The mappings below are useful for private Dreamverse migration, but they should not be treated as a public FastVideo backward-compatibility promise unless and until those fields actually exist in the public repo surfaces.
| Private Adapter Field | New Path |
| --- | --- |
| `config_model_path` | `generator.pipeline.components.config_root` |
| `ltx2_refine_enabled` | `generator.pipeline.preset_overrides.refine.enabled` |
| `ltx2_refine_upsampler_path` | `generator.pipeline.components.upsampler_weights` |
| `ltx2_refine_lora_path` | `generator.pipeline.components.lora_path` |
| `ltx2_refine_num_inference_steps` | `request.stage_overrides.refine.num_inference_steps` |
| `ltx2_refine_guidance_scale` | `request.stage_overrides.refine.guidance_scale` |
| `ltx2_refine_add_noise` | `generator.pipeline.preset_overrides.refine.add_noise` |
| `ltx2_image_crf` | `request.stage_overrides.refine.image_crf` |
| `return_continuation_state` | `request.output.return_state` |
### LongCat Legacy Mapping
| Legacy Field | New Path |
| --- | --- |
| `refine_from` | `request.inputs.refine_from` |
| `stage1_video` | `request.inputs.stage1_video` |
| `t_thresh` | `request.stage_overrides.refine.t_thresh` |
| `spatial_refine_only` | `request.stage_overrides.refine.spatial_refine_only` |
| `num_cond_frames` | `request.stage_overrides.refine.num_cond_frames` |
## Validation and Error Handling
### Strict by Default
All structured inputs should be strict by default: unknown keys error, wrong types error, invalid stage names error, incompatible state/preset combinations error.
### Exceptions
The only intentionally open-ended fields are `generator.pipeline.experimental` and `request.extensions`. These must be clearly documented as unstable and unsupported for long-term API compatibility.
### Error Quality
Validation errors should include the full nested path, expected type or valid choices, and preset/stage context when relevant:
```text
Invalid field: request.stage_overrides.refine.num_inference_steps
Expected int, got "two"
Preset: ltx2_two_stage
Stage: refine
```
## Implementation Plan
### Phases 0-5: Landed
- Phase 0 — Schema Parity Inventory: inventory complete; field classifications live in `docs/design/inference_schema_parity_inventory.yaml`; parity test guard in `fastvideo/tests/api/test_schema_parity_inventory.py`.
- Phase 1 — Shared Schema: `fastvideo/api/` with typed dataclasses, parser, validation, dotted overrides, `RunConfig`/`ServeConfig`.
- Phase 2 — VideoGenerator Compat: `from_config`, `from_file`, `generate(request=...)`, legacy `from_pretrained(..., **kwargs)` and `generate_video(..., **kwargs)` as compat shims routed through typed normalization.
- Phase 3 — CLI Refactor: `fastvideo generate` and `fastvideo serve` parse nested YAML/JSON with training-style dotted overrides; flat flag expansion removed as the canonical path.
- Phase 4 — Preset System: shared registry + pipeline-local `presets.py` for all 13 families; all 12 `SamplingParam` subclasses removed; `SamplingParam` moved to `fastvideo/api/sampling_param.py`.
- Phase 5 — Server Request Translation: `fastvideo serve` loads `ServeConfig`; stateless OpenAI endpoint clones `default_request` and merges validated user overrides.
### Remaining Phases
- **Phase 6 — LTX2 Public Upstream Path** (PR 6): upstream `ltx2_two_stage` preset; upstream continuation-state contract; upstream only repo-visible/public LTX2 surfaces into FastVideo.
- **Phase 7 — Dreamverse Adapter Migration** (PR 7 + private repo work): translate private Dreamverse-only request/config fields in a private adapter; replace raw app-owned continuation kwargs with `state` in the private server; do not expand the public FastVideo compatibility promise just to match private adapter fields.
- **Phase 7.5-7.10 — Streaming Server and Dynamo Contract** (PRs 7.5-7.10): upstream the streaming server (skeleton, GPU pool, prompt enhancer, auxiliaries, router) consuming `generate_async`; land the Dynamo backend contract (`VideoGenerator.generate_async`, health-check helper) with the Dynamo backend package itself living in the Dynamo repo.
- **Phase 8 — Model Migration and Docs** (PRs 9-10, 12): colocate `configs/pipelines/<family>.py` with pipeline implementations; add typed stage override classes for multi-stage models; update basic examples to the new API; document YAML-first inference config and migration guidance.
- **Phase 8.5 — Golden-Test Migration** (PR 11): keep SSIM/performance regression tests on legacy Python generation while preset defaults are still settling; one dedicated migration pass after the preset system and model-default behavior are stable; complete this migration before removing legacy Python inference entrypoints or kwargs.
- **Phase 9 — Deprecation and Cleanup** (PR 13): deprecate direct public use of `FastVideoArgs`; deprecate direct public use of `SamplingParam`; gradually reduce public documentation for flat flags; eventually remove legacy kwargs after downstream migration is complete.
## Final Recommendation
The public FastVideo inference API is being rebuilt around:
- typed nested configs
- model-owned named presets
- semantic stage overrides
- first-class continuation state
- YAML-first CLI with dotted overrides
The primary abstraction is `InferencePreset`, not raw kwargs and not a fully manual stage graph.
The repo is moving model-specific defaults closer to each pipeline, while keeping the public schema and parsing logic centralized.
Regression and quality tests follow the rollout. Unit/entrypoint tests migrated to the typed API early, but SSIM/performance suites only move once the typed path can express all current knobs without compatibility exceptions and produces stable defaults through presets (PR 11).
End state:
- stable Python typing
- clean YAML/JSON support
- a much better CLI story
- a sane path for Dreamverse/LTX2
- a unified abstraction for LongCat, Hunyuan, and future multi-stage models
@@ -1,285 +0,0 @@
# Dreamverse ↔ FastVideo Integration
## Status
Working integration record. Captures how Dreamverse consumes the
FastVideo public API today, what's already shared, what's still ad
hoc, and what migrations land alongside each PR in the API refactor
sequence.
Pinned versions (last reconciled this session):
| Repo | Branch | Commit | Note |
|---|---|---|---|
| FastVideo (public) | `origin/main` | `70ee5d23` | PR 6 merged |
| FastVideo (public) | `will/api_7` | `3de5f833` | PR 7 in flight (typed continuation state) |
| FastVideo-internal | `will/rebase-nbv` | `1adc513e` | pre-PR-1 on the API refactor; has live realtime runtime |
| Dreamverse | `master` | `dc500330` | uses local + remote FastVideo runtimes via `server/runtime/` |
## Related Documents
- [PR plan.md](../../PR%20plan.md) — PR-by-PR sequence for the API refactor
- [apirefactor.md](../../apirefactor.md) — design spec
- [streaming-server-upstream-plan.md](streaming-server-upstream-plan.md) — upstream plan for `ui/ltx2-streaming/server/`
- `../../../Dreamverse/server/video_generation.py` — Dreamverse's worker + local `ContinuationState`
- `../../../Dreamverse/server/runtime/{factory,backend,gpu_pool,interfaces}.py` — runtime abstraction
- `../../../FastVideo-internal/fastvideo/entrypoints/realtime/{api_server,local_runtime}.py` — internal's realtime runtime (PR 7.5/7.6 upstream source)
## Surface Area
Dreamverse depends on FastVideo across three surfaces. Listed in order
of how stable each is.
### 1. Pipeline construction (stable)
`Dreamverse/server/video_generation.py:VideoGenerationWorker` calls
`VideoGenerator.from_pretrained(...)` with flat LTX-2 kwargs today.
After PR 6 the typed `GeneratorConfig` path exists; Dreamverse can
migrate at its own pace.
| Dreamverse usage | FastVideo public surface (post-PR 6) |
|---|---|
| `VideoGenerator.from_pretrained(model_path, ltx2_refine_enabled=…, …)` | `VideoGenerator.from_pretrained(config=GeneratorConfig(...))` |
| Flat `torch_compile_kwargs={…}` dict | `engine.compile.{backend,fullgraph,mode,dynamic,extras}` |
| `ltx2_vae_tiling=True` | `pipeline.vae_tiling=True` |
| `ltx2_refine_*` family | `pipeline.preset_overrides.refine.*` + `pipeline.components.upsampler_weights` |
| `enable_torch_compile_text_encoder` | `engine.compile.text_encoder_enabled` |
The legacy flat-kwarg path stays supported via `compat.py`; migration
is opt-in. PR 13's deprecation warnings are the eventual nudge.
### 2. Realtime runtime (in flight: PRs 7.5–7.6)
`Dreamverse/server/runtime/factory.py` selects a runtime backend at
process start:
```python
def create_runtime_pool() -> RuntimePool:
if os.getenv("FASTVIDEO_REALTIME_BASE_URL"):
return FastVideoRealtimePool(base_url=..., ws_url=..., default_model_id=...)
return GPUPool(get_available_gpus()) # in-process, wraps fastvideo.entrypoints.realtime.local_runtime
```
Both backends speak the same `RuntimePool` / `RuntimeSlot` Protocol
(`server/runtime/interfaces.py`):
- `acquire(client_id, websocket=None) -> (gpu_id, RuntimeSlot)`
- `release(client_id)`
- `RuntimeSlot.{join_user, user_step, leave_user, register_stream_queue, …}`
Today both impls reach into FastVideo-internal's
`fastvideo.entrypoints.realtime.local_runtime` (which exposes
`RealtimeRuntimeConfig`, `GPUPool`, `GPUSlot`). The remote backend
talks HTTP+WS to a separately-deployed runtime of the same shape.
**Contract that PR 7.5/7.6 must preserve:**
- `RealtimeRuntimeConfig` accepts `model_registry`, `default_model_id`,
`default_height/width/num_frames/fps/num_inference_steps/guidance_scale/seed/negative_prompt`,
`default_ltx2_image_crf`, `startup_warmup_{enabled,prompt,timeout_seconds}`.
- `GPUPool(gpu_ids: list[int], config: RealtimeRuntimeConfig)` constructor.
- `pool.initialize() / shutdown() / acquire() / release() / get_status()`.
- HTTP endpoints on the remote variant: `GET /healthz`, `GET /readyz`,
`GET /status`, `WS /ws`. (These already match what
`Dreamverse/server/routes/health.py` consumes.)
When PR 7.6 lands the upstream of `fastvideo/entrypoints/realtime/`,
Dreamverse should not need any code change unless we rename the import
path. **Open: do we rename `realtime/` → `streaming/` to match the
public package introduced in PR 5.5?** A deprecation alias module
keeps both working during transition.
### 3. Continuation state (PR 7)
`Dreamverse/server/video_generation.py:89 ContinuationState` is
Dreamverse's hand-rolled per-session state holder. PR 7 introduces
the typed equivalent at `fastvideo/pipelines/basic/ltx2/continuation.py`.
#### Field mapping
| Dreamverse | PR 7 `LTX2ContinuationState` | Notes |
|---|---|---|
| `video_images: list[PIL.Image]` | `video_frames: list[np.ndarray]` (uint8 H×W×3) | numpy is leaner; Dreamverse already round-trips PIL→numpy→PIL just to add noise |
| `audio_latents: torch.Tensor` `[B, C, T, mel]` | `audio_latents: torch.Tensor` (safetensors-serialized; bf16-safe) | unchanged shape; safetensors preserves dtype incl. `bfloat16` |
| `LTX2_VIDEO_CONDITIONING_FRAME_IDX` (env) | `video_conditioning_frame_idx: int` | env constant → per-state field |
| `LTX2_VIDEO_CONDITIONING_STRENGTH` (env) | `video_conditioning_strength: float` | env constant → per-state field |
| `AUDIO_CONDITIONING_NUM_FRAMES` (env) | `audio_conditioning_num_frames: int` | env constant → per-state field |
| `AUDIO_CONDITIONING_STRENGTH` (env) | `audio_conditioning_strength: float` | env constant → per-state field |
| `audio_lps` (passed into `apply_audio`) | `audio_sample_rate: int \| None` | analogous; rename worth confirming with audio team |
| Computed `prefix_sec` per segment | `video_position_offset_sec: float` | **see open question below** |
| `segment_idx` (param to apply_*) | `segment_index: int` | per-state field |
| `VIDEO_CONTEXT_NOISE`, `AUDIO_CONTEXT_NOISE`, `ENABLE_AUDIO_COND` | not on state | runtime policy / regularization knobs, not portable session data |
| `apply_video / apply_audio / save_video / save_audio_latents / clear` | not on PR-7 state class | state is a pure data carrier; runtime owns lifecycle policy |
PR-7 is a strict superset of Dreamverse's data model **plus** lifts
several env globals into per-session typed fields.
#### Lifecycle mapping
| Dreamverse pattern | `SessionStore` API |
|---|---|
| `self.continuation = ContinuationState()` per session | `state = session_store.snapshot(sid) or LTX2ContinuationState()` |
| `apply_video(req_kwargs, segment_idx)` + `apply_audio(req_kwargs, segment_idx, audio_lps)` | `state = session_store.snapshot(sid)`; runtime builds request from `state.video_frames` / `state.audio_latents` etc. |
| `save_video(frames)` + `save_audio_latents(latents)` | runtime constructs new `LTX2ContinuationState`, then `session_store.store(sid, new_state.to_continuation_state())` |
| `clear()` at end of session | `session_store.drop(sid)` |
`SessionStore` and `BlobStore` ABCs ship with thread-safe in-memory
defaults (`InMemorySessionStore`, `InMemoryBlobStore`). Dreamverse can
adopt them as-is for the local runtime; remote runtimes can plug in
redis-backed implementations later.
#### Wire format (HTTP/WS round-trip)
Dreamverse's `FastVideoRealtimePool` already speaks the realtime
runtime's HTTP+WS protocol. When PR 7.5/7.6 land state emission on
the server side, the on-the-wire payload is the public envelope:
```json
{
"kind": "ltx2.v1",
"payload": {
"schema_version": 1,
"segment_index": 3,
"video_conditioning_frame_idx": 9,
"video_conditioning_strength": 0.75,
"audio_sample_rate": 24000,
"audio_conditioning_num_frames": 5,
"audio_conditioning_strength": 0.5,
"video_position_offset_sec": 0.2,
"video": {"frames_b64": ["..."]},
"audio": {"safetensors_b64": "..."},
"metadata": {}
}
}
```
JSON-serializable end-to-end; safetensors blob preserves audio dtype
(incl. bf16). For payloads above the inline threshold a `BlobStore`
indirection replaces the b64-encoded body with `{"blob_id": "..."}`;
the blob itself stays inside the runtime that produced it.
## Migration Plan
Per PR landed, Dreamverse adoption is opt-in.
### After PR 7 merges
Single-file change in Dreamverse, ~50-line PR:
1. Replace `server/video_generation.py:89 ContinuationState` import
with `from fastvideo.pipelines.basic.ltx2.continuation import LTX2ContinuationState`.
2. Move `apply_video`, `apply_audio`, `save_video`, `save_audio_latents`,
`clear` off the state class onto `VideoGenerationWorker` (these are
runtime policy that uses the state, not part of the state itself).
3. Update `apply_audio` to read knobs from `state.audio_conditioning_num_frames`
and `state.audio_conditioning_strength` instead of the env globals
`AUDIO_CONDITIONING_NUM_FRAMES` / `AUDIO_CONDITIONING_STRENGTH`. The
env globals can stay as defaults that populate the state when a new
session starts.
4. Same treatment for video knobs: `state.video_conditioning_frame_idx`,
`state.video_conditioning_strength`.
5. Frame storage swaps `list[PIL.Image]` for `list[np.ndarray]` —
simpler `save_video` (no PIL conversion) and simpler `clear` (no
`.close()` loop).
### After PR 7.5 lands streaming server skeleton
Dreamverse's runtime/factory.py either:
- Continues to construct `GPUPool` from `RealtimeRuntimeConfig` (the
current path), now backed by the upstreamed `fastvideo/entrypoints/realtime/`.
- Or migrates to the upstream's `ServeConfig.streaming` shape and
invokes `fastvideo serve --config realtime.yaml` as the launch path.
Either way, `Dreamverse/server/runtime/interfaces.py` `RuntimePool` /
`RuntimeSlot` Protocol can stay in place — it was modeled after the
realtime runtime's surface. No interface change needed.
### After PR 7.6 lands the GPU pool upstream
- The `local_runtime.py` import in
`Dreamverse/server/runtime/gpu_pool.py:24` becomes a public import
with the same symbols (`RealtimeRuntimeConfig`, `GPUPool`,
`get_available_gpus`).
- Per-GPU continuation state inside the worker (`ltx2_continuation_images`,
`ltx2_continuation_audio_latents`) gets replaced by a `SessionStore`
reference. Dreamverse doesn't see this change — it's runtime-internal.
- `request.state` / `result.state` round-trip starts working end-to-end
on the local runtime. Dreamverse's worker can begin reading
`result.state` and feeding `request.state` between segments.
### After PR 7.10 lands the Dynamo backend contract
- `VideoGenerator.generate_async(...) -> AsyncGenerator[VideoEvent, None]`
is the canonical API.
- Dreamverse's per-segment `user_step` flow can migrate from the legacy
sync `generate_video(..., **kwargs)` path to consuming the typed
event stream. Optional; the sync wrapper stays.
## Open Questions
### `video_position_offset_sec` semantics
Dreamverse computes `prefix_sec = float(audio_extra) / 24.0` per
segment in `apply_audio`. Not persisted on `ContinuationState`.
PR-7 has `video_position_offset_sec` as a **state field**. Two valid
interpretations:
(a) **Persistent across segments** — accumulating time offset for
long sessions; useful for time-coherent audio chaining.
(b) **Per-segment hint that rides on the carrier** — runtime
overwrites every time; field is harmless redundancy.
Field's docstring leans toward (b). Decide before PR 7.6 starts
emitting/consuming it. If we land on (a), document the accumulation
rule explicitly.
### `BlobStore` / `SessionStore` lifecycle ownership
PR 7's in-memory implementations have no eviction, no TTL, no
automatic blob cleanup on state replacement. Documented as a
per-deployment policy decision.
When PR 7.5/7.6 land the live consumer, who owns:
- bounded session capacity (LRU? TTL? hard max?)
- blob `drop()` chained when a state is replaced
- session expiry on websocket disconnect
Probably the streaming server's session manager, but worth stating
explicitly in PR 7.5's design.
### `realtime/` vs `streaming/` package naming
Currently:
- Public PR 5.5 introduced `fastvideo/entrypoints/streaming/` (skeleton + typed config).
- Internal has `fastvideo/entrypoints/realtime/` (live runtime).
- Dreamverse imports from `fastvideo.entrypoints.realtime` (per the internal name).
PR 7.5 either picks one or ships a deprecation alias module.
Recommendation in `streaming-server-upstream-plan.md`: keep
`streaming/` (it's the post-PR-5.5 public name), provide
`realtime/__init__.py` as a re-export with a `DeprecationWarning` for
one release cycle so internal/Dreamverse can land import updates.
## Test Coverage on the FastVideo Side
PR 7 ships:
- `fastvideo/tests/api/test_ltx2_continuation.py` — typed
state round-trip (inline + blob), bf16 preservation, JSON
serializability, kind/version validation, schema_version guard.
- `fastvideo/tests/entrypoints/streaming/test_session_store.py` —
store/snapshot/hydrate/drop behavior on `InMemorySessionStore`;
put/get/drop on `InMemoryBlobStore`; thread-safety of both.
PR 7.5+ should add a contract test that exercises the round-trip via
the same wire format Dreamverse's `FastVideoRealtimePool` consumes.
## Changelog
| Date | Change |
|------|--------|
| 2026-04-23 | Initial draft. Captures PR 6 / PR 7 mapping; open questions on `video_position_offset_sec`, lifecycle ownership, and `realtime/` vs `streaming/` naming. |
@@ -1,390 +0,0 @@
# Dreamverse Integration Review Log
This document tracks design decisions, open questions, and integration-time
choices made while landing the public-side stacked PRs (7.7 → 8) and switching
Dreamverse from `FastVideo-internal` to public `FastVideo`. The user will
review this carefully — entries are deliberately verbose about *why*.
## Goal
Replace Dreamverse's dependency on `FastVideo-internal` with the public
`FastVideo` package, using the upstreamed streaming server stack
(`fastvideo.entrypoints.streaming.*`) where Dreamverse currently has local
copies or imports private modules.
## Surfaces Dreamverse currently uses from FastVideo-internal
(from `/home/william5lin/Dreamverse/server/`, scanned 2026-04-26):
| Dreamverse import | Internal path | Public replacement |
|---|---|---|
| `fastvideo.entrypoints.realtime.local_runtime.RealtimeRuntimeConfig` | `FastVideo-internal/fastvideo/entrypoints/realtime/local_runtime.py` | (none) — Dreamverse rewires through `streaming.gpu_pool.SubprocessGpuPool` |
| `fastvideo.entrypoints.realtime.local_runtime.GPUPool` | same as above | `fastvideo.entrypoints.streaming.gpu_pool.SubprocessGpuPool` (PR 7.6) |
| `fastvideo.configs.pipelines.base.PipelineConfig` | already in public | unchanged |
| `fastvideo.entrypoints.video_generator.VideoGenerator` | already in public | unchanged |
| `fastvideo.layers.quantization.fp4_config.FP4Config` | already in public | unchanged |
| `fastvideo.utils.maybe_download_model` | already in public | unchanged |
| `fastvideo.models.audio.ltx2_audio_processing.AudioProcessor` | already in public | unchanged |
| `fastvideo.models.loader.component_loader.ComponentLoader` | already in public | unchanged |
| `fastvideo.models.dits.ltx2.*` | already in public | unchanged |
| local copy: `Dreamverse/server/prompt_enhancer.py` (1933 lines) | mirrors `FastVideo-internal/.../prompt_enhancer.py` | `fastvideo.entrypoints.streaming.prompt.*` (PR 7.7) |
| local copy: `Dreamverse/server/prompt_safety.py` | mirrors `FastVideo-internal/.../prompt_safety.py` | `fastvideo.entrypoints.streaming.prompt.safety` (PR 7.8) |
| local copy: `Dreamverse/server/session_logger.py` | mirrors `FastVideo-internal/.../session_logger.py` | `fastvideo.entrypoints.streaming.session_logger` (PR 7.8) |
| local copy: `Dreamverse/server/rewrite_prompt_payload.py` | mirrors `FastVideo-internal/.../rewrite_prompt_payload.py` | `fastvideo.entrypoints.streaming.prompt.rewrite` (PR 7.8) |
| local copy: `Dreamverse/server/mock_server.py` (1200 lines) | mirrors `FastVideo-internal/.../mock_server.py` | `fastvideo.entrypoints.streaming.mock_server` (PR 7.8) |
| local copy: `Dreamverse/server/session_init_image.py` | mirrors `FastVideo-internal/.../session_init_image.py` | `fastvideo.entrypoints.streaming.session_init_image` (PR 7.5 — already public) |
## Design decisions made (auto-resolved)
### D-1: Realtime runtime → streaming GpuPool migration shape
**Context.** Dreamverse's `server/runtime/gpu_pool.py` thin-wraps
`fastvideo.entrypoints.realtime.local_runtime.GPUPool`, which takes a
`RealtimeRuntimeConfig(model_registry=…, default_model_id=…, default_height=…,
default_width=…, default_num_frames=…, default_num_inference_steps=…,
startup_warmup_*…)`. The public `streaming.gpu_pool.SubprocessGpuPool` takes a
typed `GeneratorConfig` + `GpuPoolConfig` + `WarmupConfig`.
The shapes differ in two important ways:
1. The internal version had a multi-model registry (`model_id → model_config`
dict). The public version is single-model (one `GeneratorConfig`).
2. The internal version flattened a few sampling defaults (height/width/frames/
steps) into the runtime config. The public version expects them as part of
the per-request `SamplingConfig`.
**Decision.** Dreamverse will:
1. Drop the multi-model registry on the integration branch (it is not used in
production today — Dreamverse boots one model per replica).
2. Construct a `GeneratorConfig` for the chosen model from `MODEL_REGISTRY[id]`
and pass it to `SubprocessGpuPool`.
3. Move the `default_height` / `default_width` / `default_num_frames` /
`default_num_inference_steps` defaults into a server-side
`default_request: GenerationRequest` template the session controller fills
from per-request input.
**Why.** Multi-model is feasible to add back later (one pool per model id,
acquire by `(session_id, model_id)`), but not on the migration branch — that
would couple the upstream switch to a feature redesign. Punting keeps the
upstream switch a pure mechanical refactor.
**Risk.** If a Dreamverse code path silently relied on the registry to swap
models per-session, the migration branch will surface that as a missing-model
error. The integration tests must exercise at least one segment per supported
model id before merging the Dreamverse branch.
### D-2: PR 7.7 prompt enhancer API surface narrower than the internal one
**Context.** The upstreamed `PromptEnhancer.enhance/auto_extend/rewrite` returns
`LLMResponse(content, provider, model, latency_ms, fallback_used)`. The internal
`enhance_prompt` / `generate_auto_prompt` / `rewrite_prompt_sequence` returns
`EnhanceResult(prompt, fallback_used, error, provider, model, latency_ms)` /
`RewriteResult(prompts, …, rollout_id, rollout_label, raw_response_text)`.
**Decision.** The Dreamverse integration branch will adapt at the call site:
- `enhancer.enhance_prompt(...)` → `enhancer.enhance(prompt)` + a thin shim
that maps the structured response into the existing `EnhanceResult` shape
for the session-controller code path. Move the shim to
`Dreamverse/server/prompting/_internal_compat.py`.
- The locked-segment / next-segment-index plumbing the internal version
built into the user payload becomes Dreamverse-side template logic in
the shim.
- The JSON-shaped responses the internal prompts assume (`{"next_prompt":
"..."}` / `{"segment_prompts": [...]}`) become Dreamverse-side
parsing in the shim, since the public `LLMResponse` is intentionally raw.
**Why.** The public surface stays minimal and provider-agnostic; the
LTX-2-specific orchestration (locked segments, rollout id/label, JSON
schemas) is an internal-UI concern, not something every public consumer
should wear. Dreamverse keeps its existing call shape; the public stays
clean.
**Open question for review:** Should we promote some of this into
`fastvideo.entrypoints.streaming.prompt.ltx2_orchestration` (or similar)
once a second consumer appears? Logging here so we have the option.
### D-3: Multi-stage provider race (Dreamverse) vs sequential fallback (public)
**Context.** The internal enhancer runs all providers in a stage in parallel
and returns the first to succeed (`_run_provider_race`). The public
enhancer runs providers strictly sequentially with retryable-error fallback.
**Decision.** Public stays sequential for PR 7.7. The race-based fallback is
a Dreamverse-specific tail-latency optimization that depends on parallel API
budgets; promoting it would force every public consumer to have multiple
provider keys configured. Dreamverse can keep `_run_provider_race` as an
internal optimization on its side.
**Risk.** First-segment latency on Dreamverse may regress slightly when
Cerebras is having a bad minute (sequential fallback waits the full
20s timeout before trying Groq). If this is a real production concern,
add a public knob like `concurrency: int = 1` on `PromptEnhancer` that
gates a race path — but only after measuring.
### D-4: Skipping PR 7.9 router for the integration branch
**Context.** The internal stack ships a `router/main.py` that load-balances
across replicas with health checks. Dreamverse's deployment uses a single
replica per region (per `gpu_pool.py:_parse_requested_gpu_limit`).
**Decision.** Land PR 7.9 on the public side (so the surface is upstreamed)
but skip wiring it into the Dreamverse integration branch. Dreamverse's
`server/main.py` does not import from `router/`.
### D-5: Audio re-encode (PR 7.10) needed for streaming, deferred
**Context.** The internal streaming server's per-step path runs an audio
re-encode (`_re_encode_audio` inside `_stream_av_fmp4_events` /
`do_step_ltx2`) so each fMP4 segment ships with continuation-conditioning
audio. The whole-segment `pool.run()` path the public streaming server
currently uses doesn't need this. The PR plan defers re-encode integration
to PR 7.10 (`generate_async` / per-step streaming).
**Decision.** Land PR 7.10's `generate_async` on the public side. The
Dreamverse integration branch initially keeps using `pool.run()` (whole
segment, no re-encode); a follow-up branch swaps it to
`generate_async` + audio re-encode once that path is exercised end-to-end.
### D-6: `realtime/local_runtime.py` is *not* upstreamed
**Context.** It is the FastVideo-internal precursor to `streaming.gpu_pool`.
Upstreaming both would create two GPU pool implementations in the public
repo.
**Decision.** Don't upstream `realtime/local_runtime.py`. Dreamverse switches
to `streaming.gpu_pool.SubprocessGpuPool` on the integration branch. The
internal module can be deleted from FastVideo-internal at a follow-up.
## Open questions for user review
Each section below is a place the auto-decision could plausibly be wrong.
Please flip / annotate these in review.
### Q-1 Multi-model GPU pool (D-1)
Does any current Dreamverse production flow load multiple model ids
concurrently? If yes, we need to either (a) keep `realtime/local_runtime`
alive on the internal side until the public side gains a multi-model pool,
or (b) build the multi-model abstraction upstream as part of PR 7.6 follow-up
work.
### Q-2 Promoting LTX-2 prompt orchestration (D-2)
The locked-segments / next-segment-index / JSON-response orchestration is
LTX-2-specific. If Cosmos / Wan / Hunyuan ever grow a similar continuation
flow, we'll regret keeping the orchestration on the consumer side. Worth
promoting now?
### Q-3 Race-based provider fallback (D-3)
The sequential fallback in the public enhancer adds up to `timeout_ms` of
extra latency per failing provider before the next is tried. For Dreamverse
that's 20s. Should we land the race path now behind a `concurrency: int = 1`
knob, or wait until we have data?
### Q-4 Router upstream skip on Dreamverse branch (D-4)
We're upstreaming PR 7.9 (router) but not consuming it in the Dreamverse
integration branch. Is that right? Dreamverse currently has no router
component, so the answer is probably yes — but flagging.
### Q-5 generate_async cutover for the streaming path (D-5)
The plan leaves Dreamverse using `pool.run` (whole segment) initially.
Audio re-encode for cross-segment continuity is deferred to a follow-up.
Is that acceptable for the first switch, or does Dreamverse audio quality
regress relative to the internal path until 7.10 is wired in?
## PR-by-PR execution log
### PR 7.6 — already opened (#1257)
`will/api_7.6` rebased onto `origin/main`, with subprocess-pool robustness
review fixes pushed (boot_ok event, dead-worker detection, parallel shutdown,
reader-exit pending-job cleanup). 17/17 gpu_pool tests + 89/89 streaming
tests green at head.
### PR 7.7 — already opened (#1258)
`will/api_7.7` rebased onto the new 7.6 + LLM provider review fixes applied
locally (per-instance `retryable`, 4xx-non-retryable, json-decode wrap,
shared `_openai_compat.complete_openai_compatible`, `dataclasses.replace`
for the fallback marker). 29/29 prompt tests + 120/120 streaming tests green.
**Pending push** — the user opted to push this branch themselves.
### PR 7.8 — rebased onto new 7.7
`will/api_7.8` two commits replayed cleanly on the new 7.7. Adds
`fastvideo/entrypoints/streaming/{prompt/safety,prompt/rewrite,session_logger,
mock_server}.py` plus `test_auxiliaries.py`. 141/141 streaming tests green.
Notable gap vs internal version: the public `PromptSafetyFilter` ships one
classifier slot (`unsafe` label, single threshold) whereas the internal
version chained an NSFW filter and a hate-speech filter with marker-based
label matching. Multi-classifier composition is left to Dreamverse —
operators chain two filters explicitly. See **D-7** below.
### PR 7.9 — rebased onto new 7.8
`will/api_7.9` three commits replayed cleanly. Adds streaming router
(`router/{config,registry,main}.py`), `fastvideo router-serve` CLI
subcommand, and `test_router.py`. 151/151 streaming tests green.
Caveat: router/main.py uses the deprecated FastAPI `app.on_event("shutdown")`
hook — emits a DeprecationWarning. Migration to lifespan handlers is a
pre-merge cleanup item but not a blocker.
### PR 7.10 — rebased onto new 7.9
`will/api_7.10` three commits replayed with two trivial conflicts (line
wrap in `server.py`, redundant test in `test_cli_translation.py`). Adds
`VideoEvent` hierarchy, `VideoGenerator.generate_async`,
`default_health_check_request`, plus `test_generate_async.py` (273-line
contract test). 184/184 streaming + contract tests green.
### PR 8 — rebased onto new 7.10
`will/api_8` four commits → three (the 4th was a duplicate
`streaming.md` doc that 7.5 already shipped, dropped during rebase).
Adds `docs/design/server_contracts/{dynamo,index,openai}.md`,
`mkdocs.yml` entries, and `fastvideo/tests/contract/test_{dreamverse,
dynamo}_shape.py`. 206/206 streaming + contract tests green.
### Dreamverse `will/integrate-public-fastvideo`
Branch created from Dreamverse `master`. Single change: `pyproject.toml`
swaps `fastvideo = { path = "../FastVideo-internal", editable = true }`
to point at `../FastVideo`. Comment added linking back to this review
doc.
**Verified:** every TRACKED `from fastvideo.*` import in Dreamverse
(`server/video_generation.py` only) resolves against the public
package — except `fastvideo.layers.quantization.fp4_config.FP4Config`
(see **D-7** / Q-6 below).
**Untracked WIP** in `Dreamverse/server/{config,prompting,runtime,session}/`
imports `fastvideo.entrypoints.realtime.local_runtime` (D-6); this
branch does not migrate that WIP. The user's existing untracked work
stays untouched and will need a separate follow-up to consume
`streaming.gpu_pool.SubprocessGpuPool`.
## Test ladder (built-up to e2e per user request)
Each rung verifies the integration switch at one layer. Run from the
narrowest to the broadest before running the full e2e against real
GPU + model weights.
| # | Layer | Command | Status against the switched stack |
|---|---|---|---|
| 1 | Public FastVideo unit + contract tests | `pytest fastvideo/tests/api/ fastvideo/tests/entrypoints/streaming/ fastvideo/tests/contract/` | 358/358 passing on `will/api_8` |
| 2 | Public FastVideo FP4 lazy-import | `pytest fastvideo/tests/ops/quantization/test_fp4_config.py` | 3/3 passing |
| 3 | Dreamverse Python tests | `cd Dreamverse && uv run pytest server/tests/ -k "not stress and not benchmark and not health_endpoint"` | 73/73 passing against public FastVideo |
| 4 | Dreamverse FE unit/integration (vitest) | `cd Dreamverse/apps/web && npm test` | 54/86 passing — 32 failures are pre-existing copy-mismatches in `reducer.test.ts` etc., not caused by the switch |
| 5 | Backend HTTP smoke (Playwright) | `cd Dreamverse/apps/web && PLAYWRIGHT_SKIP_WEBSERVER=1 PLAYWRIGHT_BASE_URL=http://127.0.0.1:8009 npx playwright test e2e/backend-health.spec.ts` | 4/4 passing (5th correctly skipped because devtools-only route is off) |
| 6 | Frontend shell smoke (Playwright) | `npx playwright test e2e/frontend-shell.spec.ts` | Pending — requires Next.js dev server to be reachable; was stuck during this run, needs a clean restart |
| 7 | Full e2e preset generation | `npx playwright test e2e/preset-prompt-generation.spec.ts` | **8/8 passing** end-to-end after restart with `CUDA_VISIBLE_DEVICES=4 ENABLE_TORCH_COMPILE=0 FASTVIDEO_GPU_COUNT=1 FASTVIDEO_ENABLE_DEVTOOLS=1`. BE warmup + GPU 4 idle slot let `/readyz` flip green; the spec verifies preset → WS → backend handshake → "Generating video…" state. |
### How to reproduce e2e tier 7 from cold
```
# 1. BE — picks an idle GPU and skips torch.compile (avoids the
# aarch64 cross-compiler bug in the conda env's triton stack).
cd ~/Dreamverse
set -a; source ~/.env; set +a
CUDA_VISIBLE_DEVICES=4 ENABLE_TORCH_COMPILE=0 \
FASTVIDEO_ENABLE_DEVTOOLS=1 FASTVIDEO_GPU_COUNT=1 \
uv run dreamverse-server &
# 2. Wait for /readyz (~2 min for warmup x2 segments)
until curl -fsS http://127.0.0.1:8009/readyz >/dev/null; do sleep 5; done
# 3. FE
cd ~/Dreamverse/apps/web
BACKEND_URL=http://127.0.0.1:8009 NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
npm run dev:devtools &
# 4. Playwright
cd ~/Dreamverse/apps/web
PLAYWRIGHT_SKIP_WEBSERVER=1 \
PLAYWRIGHT_BASE_URL=http://127.0.0.1:5274 \
BACKEND_URL=http://127.0.0.1:8009 \
npx playwright test --project=chromium --reporter=list
```
### Surfaced during the e2e debug pass (logged here for follow-up)
* **`SamplingParam has no field ltx2_image_crf`** — Dreamverse's
`server/video_generation.py:406` passes `ltx2_image_crf=0.0` to a
`SamplingParam(...)` constructor. The internal SamplingParam (in
`fastvideo/configs/sample/base.py`) declared this field; the public
`fastvideo.api.sampling_param.SamplingParam` does not. Currently
the BE logs an `ERROR` and silently drops the kwarg; warmup still
succeeds because the field is non-load-bearing for FP4-disabled
inference. Either re-add the field to the public schema or update
Dreamverse to stop passing it. **D-8.**
* **`aarch64-conda-linux-gnu-cc` triton compile failure** — the conda
env we boot from injects an ARM cross-compiler ahead of `gcc` on
`$PATH`, so `torch._inductor`'s triton launcher fails compilation.
Setting `ENABLE_TORCH_COMPILE=0` bypasses it. Long-term fix: clean
the conda env's compiler shadowing or add a `CC=gcc` override in
Dreamverse's worker bootstrap. **D-9.**
* **GPU pool starts but warmup OOMs on a shared GPU** — when
`CUDA_VISIBLE_DEVICES` lands on a GPU another tenant is using
(107 GiB-pegged training run on GPU 0 in this case), LTX-2 warmup
fails with OOM. Picking an idle GPU (4-7 here) is a manual step.
A pre-warm probe that checks free memory before booting the pool
would prevent this. **D-10.**
* **ffmpeg fragment write `Broken pipe`** — when the WS client closes
before the backend finishes streaming the first segment, ffmpeg
hits `[Errno 32] Broken pipe`. Currently Dreamverse's
`gpu_pool.handle_command` re-raises this as a session error,
which then propagates to "User step failed". Cosmetic for now —
swallowing pipe-broken on intentional disconnect would clean up
the logs. **D-11.**
## Additional integration gaps surfaced during the switch
### D-7: `FP4Config` is private-only
**Context.** `Dreamverse/server/video_generation.py:271` imports
`fastvideo.layers.quantization.fp4_config.FP4Config` and assigns it to
`pipeline_config.dit_config.quant_config`. The 411-line module lives only
in `FastVideo-internal/fastvideo/layers/quantization/fp4_config.py` and
hard-imports `flashinfer` at module top — it never made the public
upstream pass. Public has `base_config.py` and `absmax_fp8.py` only.
**Decision (provisional).** Don't upstream `fp4_config.py` in this
session. Reasons:
1. It introduces a new external dependency (`flashinfer`) the public
package has avoided so far.
2. The class hard-codes LTX-2 layer paths
(`ltx2.blocks.{i}.attn1.to_q` etc.) — this is "LTX-2-specific FP4",
not generic FP4. Belongs colocated with `pipelines/basic/ltx2/` if
it goes anywhere.
3. The FP4 pre-quantize/forward op surface is the kind of thing where
a careful review pass matters more than a bulk copy.
**What this means for the integration branch.** Dreamverse will boot
fine; only the FP4-quantized path inside `video_generation.py:283`
will fail (lazy import). For workflows that don't enable FP4
quantization, the integration is complete.
### Q-6 (review): how to land FP4Config publicly?
Two reasonable next steps:
1. **Colocate.** Move FP4 code to `fastvideo/pipelines/basic/ltx2/quantization.py`
with `flashinfer` as an optional extra: `pip install fastvideo[fp4]`.
Refactor `FP4QuantizeMethod` to take its layer-prefix list from a
pipeline-config field instead of hardcoding ltx2 paths so the
approach generalizes.
2. **Keep private.** Treat FP4 as a Dreamverse-side concern — Dreamverse
imports `fp4_config` from the internal repo via a thin shim. Public
FastVideo stays focused on generic surfaces. This means the
"FastVideo-internal removable" goal is partially undone.
Recommendation: option 1 once the API refactor settles — wait until
the LTX-2 colocation step (PR 9 / 10 territory) and land FP4 there.
@@ -1,518 +0,0 @@
# Handoff: LTX-2 NVFP4 wire-up + Dreamverse launch-demo skill
This document hands off in-flight work to the next coding agent. It covers
two related streams that landed across two repos:
1. **FastVideo** (`will/ltx2_sr_port`): wire NVFP4 (NVIDIA's block-scaled
FP4) inference + per-component torch.compile + supporting parity fixes
so the public package matches `FastVideo-internal` for the LTX-2
distilled streaming path used by Dreamverse.
2. **Dreamverse** (`will/integrate-public-fastvideo`): switch the GPU
worker to the typed `GeneratorConfig` API, rename `FP4Config` →
`NVFP4Config`, add a `launch-demo` skill + canonical
`serve_configs/streaming_demo.yaml` for `fastvideo serve --config`.
Stack remains green: 222/222 FastVideo unit/contract/api tests pass; 8/8
Playwright e2e tests pass against the live `dreamverse-server` + Next.js
stack.
---
## Repo + branch state
| Repo | Path | Branch | Tip |
| --- | --- | --- | --- |
| FastVideo | `/home/william5lin/FastVideo` | `will/ltx2_sr_port` | `c6c14c55` |
| Dreamverse | `/home/william5lin/Dreamverse` | `will/integrate-public-fastvideo` | `3d7fd89` |
| Reference (read-only) | `/home/william5lin/FastVideo-internal` | (their) `main` | source of truth for parity |
> **Working branch on FastVideo is `will/ltx2_sr_port`, not the default checkout.**
> The shell may report `will/uv-pip-install-everywhere` because that was
> the earlier checkout. Run `git checkout will/ltx2_sr_port` before
> picking up FastVideo work.
### Live processes (do not duplicate)
```
:8009 dreamverse-server pid 2453227 (warmed, /readyz returns 200)
:5274 next-server (dev) pid 2399103 (devtools build)
```
### Stashes
* FastVideo: `stash@{0}: WIP on main: …HunyuanVideo plugin…` — pre-existing,
unrelated to this work, do not pop.
* Dreamverse: `stash@{0}: wip: server modular refactor (split
config/prompting/runtime/session)` — 3867 lines of orphan modular split
off this branch. Do not pop on this branch; recover on a separate
feature branch if anyone wants to resurrect it.
---
## What landed (FastVideo: `cfccd292..c6c14c55`)
Six commits on top of the i2v / continuation latent port:
```
c6c14c55 test(nvfp4): lock LTX-2 wiring + typed transformer_quant flow
94c983a2 refactor(quant): rename FP4 → NVFP4 to disambiguate from other FP4 variants
42b30bf9 feat(ltx2): wire FP4 inference through fastvideo.layers.quantization
6da342ba feat(compile): per-component compile + transformer_refine + prepare hook
221cb20a feat(api): typed per-component CompileConfig + FastVideoArgs carriers
a4760bae fix(api): propagate generic refine_* args + match internal randn
```
Each commit message has the rationale. Highlights below.
### `a4760bae` — three small parity fixes
* `FastVideoArgs.__post_init__` now calls `_resolve_refine_args()` which
copies the public-facing generic `refine_*` knobs onto their
`ltx2_refine_*` runtime carriers. Was missing → callers that set
`refine_lora_path=...` saw "applied to 0 layers" warnings as the value
was silently dropped.
* `_randn_ltx2_video_latents` patch path reverted from `randn_tensor` →
`torch.randn` to bit-match internal under single-generator inference.
Identical for a single `torch.Generator` but diverges for
`list[Generator]` (per-sample seeds).
* Classified 19 `refine_*` / `ltx2_refine_*` / i2v / `ltx2_audio_*` /
`ltx2_conditioning_latent_*` / `ltx2_video_conditions` fields in the
schema-parity inventory yaml.
### `221cb20a` — typed CompileConfig + FastVideoArgs carriers
`CompileConfig` (in `fastvideo/api/schema.py`) gained per-component knobs:
```python
@dataclass
class CompileConfig:
enabled: bool = False # master DiT switch
backend / fullgraph / mode / dynamic / extras # master kwargs
# Per-component overlays, None = inherit master `enabled`
text_encoder_enabled: bool | None = None
vae_enabled: bool | None = None
audio_vae_enabled: bool | None = None
# Per-component kwargs override master when non-empty
dit_kwargs: dict = ...
text_encoder_kwargs: dict = ...
vae_kwargs: dict = ...
audio_vae_kwargs: dict = ...
```
Matching carrier fields on `FastVideoArgs`:
`enable_torch_compile_text_encoder/vae/audio_vae` and
`torch_compile_kwargs_dit/text_encoder/vae/audio_vae`. Compat layer
round-trips them through `legacy_from_pretrained_to_config` and
`generator_config_to_fastvideo_args`. **No behavior change yet** — these
are surface ports only; consumed in the next commit.
### `6da342ba` — refine + per-component compile + prepare_for_compile
`composed_pipeline_base.post_init` now:
* compiles `transformer_refine` alongside `transformer` and
`transformer_2` whenever the DiT compile flag is on (closes the LTX-2
stage-2 silent-eager bug);
* dispatches per-component compile loops (text encoder, VAE, audio VAE)
with per-component kwargs falling back to master when empty;
* calls `module.prepare_for_compile()` on each compiled submodule
before invoking `torch.compile` (hook protocol — model-specific).
Implemented on `Gemma3` to materialize HF weights outside Dynamo's
tracer.
### `42b30bf9` — NVFP4 LTX-2 inference wire-up *(largest)*
End-to-end:
1. `models/dits/ltx2.py` — swap `nn.Linear` → `ReplicatedLinear` for the
FP4-eligible subset (`LTXSelfAttention`, `LTXDistributedSelfAttention`,
`FeedForward`/`GELUApprox`); plumb `quant_config` and `prefix=` from
`BasicAVTransformerBlock` → `_init_transformer_blocks` → `LTXModel`
→ `LTX2Transformer3DModel`. Other linears
(`TimestepEmbedding`, `PixArtAlphaTextProjection`, `patchify_proj`,
`proj_out`, `AdaLayerNormSingle.linear`) stay `nn.Linear` —
matches internal exactly.
2. Port `_supports_prequantized_input` and
`_linear_project_with_optional_prequant` helpers. Attention forward
pre-quantizes input once (`quantize_input`), reuses the
`(x_fp4, x_scale, x_global_sf)` tuple for k/v projections when
`context is x` — bit-matches internal's fused path.
3. `models/loader/fsdp_load.py` — new `_maybe_convert_model_to_nvfp4`
helper detects via `isinstance(quant_method, NVFP4QuantizeMethod)`
(no flag); calls `convert_model_to_nvfp4` to materialize
`_nvfp4_weight*` / `_nvfp4_alpha` / `_weight_global_sf` buffers.
`flashinfer` import is lazy (inside the helper), so the loader is a
no-op on hosts without flashinfer.
4. `layers/quantization/__init__.py` — registered `"NVFP4"` in
`QuantizationMethods` literal + `get_quantization_config`.
5. `api/compat.py` + `fastvideo_args.py` — typed
`engine.quantization.transformer_quant: "NVFP4"` resolves to a
concrete `NVFP4Config()` instance, carried on `FastVideoArgs.transformer_quant`,
pinned onto `pipeline_config.dit_config.quant_config` in
`__post_init__._apply_transformer_quant`. **The explicit setter
(legacy mutation pattern) wins** if `dit_config.quant_config` is
already non-None.
6. `layers/linear.py` — `LinearBase.__init__` now falls back to
`UnquantizedLinearMethod` when `quant_config.get_quant_method` returns
`None`. `NVFP4Config` only tags a curated subset of LTX-2 layers, and
the previous `assert quant_method is not None` would crash any
non-tagged layer that received a quant_config.
### `94c983a2` — FP4 → NVFP4 rename
NVIDIA's specific block-scaled fp4 format (e2m1 mantissa, fp32 alpha,
`layout_128x4` scale layout, group size 16) — distinct from MX-FP4 /
OCP-FP4 / generic e3m0. Mechanical rename, no behavior change:
* `fp4_config.py` → `nvfp4_config.py`
* `FP4Config` → `NVFP4Config`; `get_name()` returns `"nvfp4"`
* `FP4QuantizeMethod` → `NVFP4QuantizeMethod`
* `convert_model_to_fp4` → `convert_model_to_nvfp4`
* `QuantizationMethods` literal: `"FP4"` → `"NVFP4"`
* registered buffer names: `_fp4_weight`/`_fp4_alpha` →
`_nvfp4_weight`/`_nvfp4_alpha`
* loader helper renamed
* test file rename + symbol updates
Internal-scope torch op namespace `fastvideo_fp4::*` and
`_get_ltx2_fp4_stage_profile` deliberately left as-is — purely
internal naming that mirrors FastVideo-internal.
### `c6c14c55` — contract + numerical lock-in tests
* `fastvideo/tests/ops/quantization/test_nvfp4_ltx2_wiring.py` (6 tests):
asserts that `LTXSelfAttention.to_q/to_k/to_v/to_out` are
`ReplicatedLinear`; `NVFP4Config()` attaches `NVFP4QuantizeMethod`
on the quantized subset with the correct `layer_prefix`; non-tagged
projections (cross-attn K/V, audio attn, audio FFN) fall back to
`UnquantizedLinearMethod`; `BasicAVTransformerBlock` propagates
`quant_config` and `prefix` correctly to all 4 attention modules +
FFN at once.
* `fastvideo/tests/api/test_typed_quant_flow.py` (4 tests): asserts
typed `engine.quantization.transformer_quant: "NVFP4"` →
`NVFP4Config()` instance flow; default leaves `transformer_quant`
None; explicit `dit_config.quant_config = …` wins over typed carrier.
---
## What landed (Dreamverse: `248060b..3d7fd89`)
Three commits on top of the e2e tier:
```
3d7fd89 feat(skill): launch-demo orchestrator + fastvideo serve YAML
d80c2a8 refactor(server): drive FP4 + per-component compile via typed GeneratorConfig
4cc6b30 chore: gitignore Playwright + Next.js build artifacts under apps/web
```
### `d80c2a8` — server/video_generation.py refactor
Three coordinated changes in the GPU worker:
* Replace legacy `load_kwargs` dict + `VideoGenerator.from_pretrained(model_root, **kwargs)`
call with the typed `GeneratorConfig` (`EngineConfig` /
`OffloadConfig` / `CompileConfig` / `PipelineSelection` /
`ComponentConfig`). Refine knobs move from `ltx2_refine_*` flat
kwargs into `preset_overrides["refine"]`. **The in-memory
`pipeline_config` pin** (`dit_config.quant_config = NVFP4Config()`)
keeps using the legacy `experimental["pipeline_config"]` carrier
because typed `transformer_quant: "NVFP4"` doesn't yet support
setting `layer_profile`.
* Rename FP4 → NVFP4.
* Re-enable `"mode": "max-autotune-no-cudagraphs"` (was commented out).
Closes the last known divergence vs FastVideo-internal in the
worker-level path trace.
### `4cc6b30` — gitignore Playwright/Next.js artifacts
Added `apps/web/{node_modules,.next,test-results,playwright-report}` to
`.gitignore`. Mirror of the existing `prod-ui/` ignore set.
### `3d7fd89` — launch-demo skill
```
.agents/skills/launch-demo/
├── SKILL.md
└── scripts/
├── launch_demo.sh # orchestrator: BE + FE + health probes + Ctrl-C trap
├── launch_backend_dreamverse.sh # uv run dreamverse-server (default)
├── launch_backend_fastvideo.sh # uv run fastvideo serve --config (typed path)
└── launch_frontend.sh # next dev (devtools/dev/single5s)
serve_configs/
└── streaming_demo.yaml # canonical ServeConfig matching internal/ui
```
YAML has every field annotated with the internal source line it mirrors:
LTX-2 distilled, NVFP4, 121 frames @ 1088×1920 24fps, 5 inference steps,
2-step refine gs=1.0 add_noise=true, max-autotune-no-cudagraphs compile,
121-frame default request, 300s session timeout, 6 segment cap, av_fmp4
streaming, cinematic-drone warmup prompt, 2400s warmup timeout, 9
conditioning frames + 0 end-offset, prompt enhancer on with cerebras /
gpt-oss-120b / 20s timeout.
**Two BE flavors documented in SKILL.md:**
| `BE_FLAVOR=` | Boots | Routes served | FE compatible |
| --- | --- | --- | --- |
| `dreamverse` (default) | `dreamverse-server` | `/healthz`, `/readyz`, `/curated-presets`, `/v1/stream`, devtools, session monitor | ✓ full |
| `fastvideo` | `fastvideo serve --config <yaml>` | `/health`, `/v1/stream` | ⚠ FE will surface fetch errors for `/curated-presets`, `/readyz` until those routes migrate into FastVideo's `build_app` |
The fastvideo flavor exists today as the verifiable typed-config path
(YAML parses, streaming worker boots, dotted overrides work). It is not
yet a drop-in for the FE — see "Open follow-ups" below.
---
## Verified
* `222 passed, 1 skipped` across `fastvideo/tests/api/`,
`fastvideo/tests/contract/`,
`fastvideo/tests/ops/quantization/test_nvfp4_*`,
`tests/local_tests/pipelines/test_ltx2_pipeline_smoke.py`.
* `8 passed` Playwright e2e (backend-health 5, frontend-shell 2,
preset-prompt-generation 1) against the live `dreamverse-server`
+ Next.js stack.
* `streaming_demo.yaml` parses cleanly against `ServeConfig`; the
validation path of `fastvideo serve --config <yaml>` runs without
error and accepts dotted overrides like `--server.port 8010`.
* FastVideo `bash -n` clean across all four launch scripts.
---
## Critical context (gotchas a successor should know)
### NVFP4 layer set is asymmetric — by design
`NVFP4Config.fp4_layers` covers:
* `attn1.{to_q,to_k,to_v,to_out}` — full self-attention
* `attn2.{to_q,to_out}` — cross-attn Q + out only (text context not quantized)
* `audio_to_video_attn.{to_q,to_out}` — AV cross Q + out
* `video_to_audio_attn.{to_k,to_v}` — VA cross K + V
* `ffn.{fc_in,fc_out}` — video FFN
* `adaln_single.linear` — but this is `nn.Linear` (not `LinearBase`),
so it never actually gets FP4'd. List entry has no effect; matches
internal.
**NOT in the set:** audio self-attention (`audio_attn1.*`), audio
cross-attention (`audio_attn2.*`), audio FFN (`audio.ffn.*`). Audio
path is cheap enough that quant overhead isn't worth it. Test
`test_basic_av_block_propagates_quant_config_to_all_children` locks
this in — if you add audio quantization later, update the test.
### `LinearBase` fallback is load-bearing
`fastvideo/layers/linear.py:191-202`: when `quant_config.get_quant_method`
returns `None` (layer not in the quant config's set), we fall back to
`UnquantizedLinearMethod`. **Do not remove this fallback** — it would
break every non-tagged `ReplicatedLinear` constructed with a
`NVFP4Config`, and `assert quant_method is not None` in
`ReplicatedLinear.__init__` would fire on unmatched layers.
### Typed `transformer_quant` precedence
`FastVideoArgs._apply_transformer_quant` only writes
`dit_config.quant_config` when it's currently `None`. If a caller has
explicitly set `pipeline_config.dit_config.quant_config = NVFP4Config(...)`,
the explicit setter wins. Dreamverse's `video_generation.py` relies on
this — it sets `NVFP4Config()` directly because the typed
`transformer_quant: "NVFP4"` doesn't expose `layer_profile`.
### Pre-existing AbsMaxFP8 test failure is NOT mine
`fastvideo/tests/ops/quantization/test_absmax_fp8.py::test_create_weights_rejects_invalid_dtype`
fails on `main` and on this branch with the same error
("AssertionError not raised"). I confirmed via `git stash` that the
failure pre-dates my changes. Not blocking; tracked as separate tech
debt.
### `transformer_refine` is auto-compiled with the master DiT flag
Set `enable_torch_compile=True` and `transformer_refine` compiles
along with `transformer` and `transformer_2`. There is **no separate
`enable_torch_compile_refine` flag** — by design, refine inherits the
DiT compile state to keep the typed surface small. If you need them
decoupled, add a new field; don't repurpose existing ones.
### `prepare_for_compile` is a duck-type protocol, not a base class method
Defined nowhere; called via `getattr(module, "prepare_for_compile", None)`
in `composed_pipeline_base._maybe_compile_pipeline_module`. Currently
only Gemma implements it (to materialize HF weights outside Dynamo).
Add to other models that have lazy external state if you observe
compile-time graph breaks.
### Public typed `PromptEnhancerConfig.provider` is `Literal["cerebras", "groq"]`
Internal supports `"cerebras_ifm"` (config.py:143). The public typed
schema does not. The `streaming_demo.yaml` defaults to `"cerebras"`.
For agents that need `cerebras_ifm`, the `dreamverse-server` flavor
respects the `FASTVIDEO_PROMPT_PROVIDER` env var (legacy path);
`fastvideo serve --config` does not currently expose it.
### Dreamverse `pipeline_config` is still a Python object passed via `experimental`
The typed `GeneratorConfig` doesn't have a clean home for an
in-memory `PipelineConfig` instance with mutated `dit_config`. We
pass it via `pipeline.experimental["pipeline_config"]` — the
`compat.py` legacy adapter recognizes that key and threads it through
to `FastVideoArgs.from_kwargs`. This is fine but not pretty; if
someone designs a typed `dit_config` carrier later, this becomes
obsolete.
### `fastvideo serve --config` is not yet a drop-in for the FE
`fastvideo.entrypoints.streaming.server.build_app` exposes only
`/health` and `/v1/stream`. The Dreamverse Next.js shell expects
`/healthz`, `/readyz`, `/status`, `/curated-presets`,
`/curated-presets/append`, `/prompt-system-config`, and the devtools
routes. These all live in `Dreamverse/server/main.py` +
`Dreamverse/server/routes/`. Until they migrate into FastVideo's
`build_app` (or are exposed via a Dreamverse-side proxy), the
`BE_FLAVOR=fastvideo` flavor is for verifying the typed serve config
path only — not for full FE compatibility.
---
## Open follow-ups (prioritized)
### High
1. **Migrate FE-required routes into FastVideo's `build_app`.**
`/healthz`, `/readyz`, `/status` look obviously fastvideo-side
(they're streaming-server health). `/curated-presets` and
`/prompt-system-config` are operator-side surfaces and should
probably stay in Dreamverse (or migrate as opt-in routes that the
FE feature-detects). Without this, `BE_FLAVOR=fastvideo` is
permanently a "diagnostic" flavor. Closes the
`launch-demo` skill TODO.
2. **AbsMaxFP8 test failure cleanup.** Pre-existing. Either fix the
test (`AbsMaxFP8LinearMethod.create_weights` no longer asserts on
invalid dtype — restore the assert if intentional, otherwise drop
the test).
### Medium
3. **Add `cerebras_ifm` to public `PromptEnhancerConfig.provider`
Literal.** Trivial schema change; needs paired enhancer-side
provider implementation in
`fastvideo/entrypoints/streaming/prompt/providers/`.
4. **Expose `layer_profile` on typed `engine.quantization`.** Today
`transformer_quant: "NVFP4"` always constructs `NVFP4Config()`
with the default `layer_profile="refine"`. To support stage-1
profiles (no `attn2.to_out`, no cross-modal AV) via typed config,
add `transformer_quant_layer_profile: str | None = None` and
thread it through `compat.py`. Dreamverse currently dodges this
by setting `NVFP4Config()` directly via `experimental`.
5. **Typed `dit_config.quant_config` carrier.** The
`experimental["pipeline_config"]` escape hatch in Dreamverse
should eventually become a typed field. Design TBD.
### Low
6. **Audio attention quantization profile.** If an audio-quant
profile is added to `NVFP4Config.fp4_layers` (currently audio attn
and FFN are bf16), update
`test_basic_av_block_propagates_quant_config_to_all_children`.
7. **Schema parity inventory.** A few internal-only fields are not
exposed publicly (`PROMPT_HTTP_TIMEOUT_MS`,
`PROMPT_INITIAL_STAGE_TIMEOUT_MS`, `PROMPT_TEMPERATURE`,
`PROMPT_MAX_COMPLETION_TOKENS`, `PROMPT_AUTO_SLEEP_MS`,
`PROMPT_AUTO_TIMEOUT_MS`, the curated-presets file paths).
These all flow via env vars on `dreamverse-server` today; if
`fastvideo serve --config` becomes the canonical entrypoint,
they'll need typed homes.
8. **Empty `apps/web/test-results/` directory locally.** The
`.gitignore` entry I added makes it invisible to `git status`,
but the dir itself still has a stale `.last-run.json` (45 bytes)
from a prior Playwright run. Harness blocked auto-cleanup
("pre-existing files"); the user can `rm -rf
apps/web/test-results` whenever convenient.
---
## How to pick up work
### Quick orientation (run these first)
```bash
# FastVideo state
cd /home/william5lin/FastVideo
git checkout will/ltx2_sr_port
git log --oneline cfccd292..HEAD # six commits added this round
.venv/bin/python -m pytest fastvideo/tests/api/ \
fastvideo/tests/contract/ \
fastvideo/tests/ops/quantization/test_nvfp4_*.py \
tests/local_tests/pipelines/test_ltx2_pipeline_smoke.py \
-q --no-header # expect 222 passed, 1 skipped
# Dreamverse state
cd /home/william5lin/Dreamverse
git log --oneline 248060b..HEAD # three commits added this round
cat serve_configs/streaming_demo.yaml | head -40
ls .agents/skills/launch-demo/
# Live stack health (already running on this host)
curl -s http://localhost:8009/readyz | head -c 200
curl -s http://localhost:5274/ | head -c 100
( cd apps/web && npx playwright test --reporter=line ) # expect 8 passed
```
### Reference docs
* **FastVideo internal/ui parity source:** `../FastVideo-internal/ui/ltx2-streaming/server/config.py`
* **NVFP4 source on internal:** `../FastVideo-internal/fastvideo/layers/quantization/fp4_config.py`
* **Worker-trace audit:** `../FastVideo/dreamverse_review.md` (D-1
multi-model, D-5 audio re-encode, prior gap inventory)
* **Schema parity inventory:** `docs/design/inference_schema_parity_inventory.yaml`
* **PR-plan for the broader migration:** `../FastVideo/PR plan.md`
### Files most likely to need touches in follow-ups
* `fastvideo/api/schema.py` — `CompileConfig`, `QuantizationConfig`,
`PromptEnhancerConfig` Literal extension.
* `fastvideo/api/compat.py` — typed → flat translation.
* `fastvideo/fastvideo_args.py` — carrier fields and
`_apply_transformer_quant`.
* `fastvideo/entrypoints/streaming/server.py::build_app` — add
`/healthz`, `/readyz`, `/status` routes for FE compatibility (high
priority follow-up #1).
* `Dreamverse/server/video_generation.py` — typed `GeneratorConfig`
builder (current).
* `Dreamverse/serve_configs/streaming_demo.yaml` — every parity
knob; edit here, not in shell scripts.
---
## Don't / Cautions
* **Don't pop the Dreamverse stash on this branch.** It's 3867 lines
of orphan modular refactor (server/{config,prompting,runtime,session}/)
with broken absolute imports. If anyone wants to resurrect it, do so
on a separate feature branch.
* **Don't remove the `LinearBase` `UnquantizedLinearMethod` fallback.**
See "Critical context" above.
* **Don't repurpose `enable_torch_compile` to mean DiT-only.** It also
drives `transformer_refine` and `transformer_2` compile. Add a new
flag if decoupling is needed.
* **Don't change `NVFP4Config` buffer names back to `_fp4_*`.** The
rename is intentional to disambiguate from MX-FP4 / OCP-FP4.
* **Don't bypass the typed surface for new options.** New compile /
quant / refine knobs should land on the dataclass + compat.py +
parity inventory together. The existing test suite locks this in.
* **Don't merge to main without a CI run that covers FP4.** Current
CI doesn't run flashinfer-dependent paths; the wiring tests in
`test_nvfp4_ltx2_wiring.py` are CPU-only by design and don't
exercise the actual FP4 kernels.
---
*Last updated: end of session that landed `c6c14c55` on FastVideo and
`3d7fd89` on Dreamverse. Stack remains green; no dirty state.*
@@ -1,539 +0,0 @@
# FastVideo Streaming Server Upstream — Design & Plan
## Status
Exploration / design draft. Captures the re-evaluation triggered by the
decision to upstream `FastVideo-internal/ui/ltx2-streaming/server/` into
the public repo. Not yet approved for execution.
## Related Documents
- [PR plan.md](../../PR%20plan.md) — PR-by-PR implementation plan for the API refactor
- [apirefactor.md](../../apirefactor.md) — design spec this plan implements
- `../../../FastVideo-internal/ui/ltx2-streaming/` — upstream source (server side)
- `../../../dynamo/` — local clone of ai-dynamo/dynamo; backend patterns at
`components/src/dynamo/{vllm,sglang,trtllm}/` and `CLAUDE.md` files
- https://github.com/ai-dynamo/dynamo/pull/7544 — draft PR that promotes
FastVideo to a native Dynamo backend (CLOSED, superseded — but establishes
the integration shape)
## Context
The internal `FastVideo-internal/ui/ltx2-streaming/` directory contains a
complete LTX2 streaming service. The user has decided:
- **Frontend clients** (`client/`, `prod-ui/`) stay in the internal repo
- **Everything server-side** — FastAPI/WebSocket server, GPU pool, prompt
enhancer, router, auxiliaries — will be upstreamed to FastVideo
In parallel, FastVideo is becoming a **first-class Dynamo backend** (same
tier as vllm, sglang, trtllm). The refactor must produce an API that
Dynamo's `components/src/dynamo/fastvideo/` package can consume as a
pure Python import, without re-introducing the legacy flat-kwarg
surface. Draft PR ai-dynamo/dynamo#7544 defines the concrete integration
shape we need to support.
This materially changes the tail of the API refactor plan. The current
PR 5 ("wire `ServeConfig.default_request` into the OpenAI-compatible
HTTP server") addresses only the stateless endpoint; the real upstream
target is a much larger, session-based stack **plus** a clean Dynamo
backend contract.
This document captures:
- what's being upstreamed and where it lands
- four design decisions that shape the upstream (continuation model,
streaming server layout, LLM provider abstraction, Dynamo backend
integration)
- a revised PR sequence for the tail of the refactor
## What's being upstreamed
| Internal path | Size | Role | Upstream target |
|---|---|---|---|
| `server/main.py` | 94KB | FastAPI + WebSocket, session lifecycle, segment orchestration | `fastvideo/entrypoints/streaming/server.py` + handlers |
| `server/gpu_pool.py` | 66KB | GPU orchestration, subprocess workers | `fastvideo/entrypoints/streaming/gpu_pool.py` |
| `server/prompt_enhancer.py` | 69KB | LLM orchestration (cerebras_ifm, cerebras, groq) | `fastvideo/entrypoints/streaming/prompt/` package |
| `server/mock_server.py` | 45KB | Mock backend for dev/tests | `fastvideo/entrypoints/streaming/mock_server.py` |
| `server/prompt_safety.py` | 7KB | Optional fasttext-gated prompt safety | `fastvideo/entrypoints/streaming/prompt/safety.py` |
| `server/session_init_image.py` | 3KB | i2v init image handling | `fastvideo/entrypoints/streaming/session_init_image.py` |
| `server/rewrite_prompt_payload.py` | 3KB | Rewrite flow payload builder | `fastvideo/entrypoints/streaming/prompt/rewrite.py` |
| `server/session_logger.py` | 1KB | Session JSONL logs | `fastvideo/entrypoints/streaming/session_logger.py` |
| `server/config.py` | 9KB | Env-driven server config | Typed `ServeConfig` extensions |
| `router/main.py` | 27KB | Multi-replica load balancer + WS proxy | `fastvideo/entrypoints/streaming/router/` (or separate package) |
| `slurm/` | — | Deployment scripts | Likely stays internal |
## FastVideo contact surface today
Direct calls from the internal stack into FastVideo, all in `gpu_pool.py`:
| Location | Call | Notes |
|---|---|---|
| `gpu_pool.py:164` | `from fastvideo.entrypoints.video_generator import VideoGenerator` | Subprocess-level import, post-`CUDA_VISIBLE_DEVICES` setup |
| `gpu_pool.py:230` | `PipelineConfig.from_pretrained(config_model_path)` | Direct access to legacy `PipelineConfig` |
| `gpu_pool.py:231` | `pipeline_config.dit_config.quant_config = FP4Config()` | Direct internals mutation |
| `gpu_pool.py:264-267` | `VideoGenerator.from_pretrained(model_root, **load_kwargs)` | Flat legacy kwargs |
| `gpu_pool.py:837` | `generator.generate_video(**request_kwargs)` | Per-segment flat kwargs |
| `gpu_pool.py:282-288` | `LTX2AudioEncoder`, `AudioProcessor`, `get_diffusers_config` | Audio re-encode path |
`load_kwargs` at `gpu_pool.py:233-260` contains:
`ltx2_refine_enabled`, `ltx2_refine_upsampler_path`, `ltx2_refine_lora_path`,
`ltx2_refine_num_inference_steps`, `ltx2_refine_guidance_scale`,
`ltx2_refine_add_noise`, `pipeline_config`, `torch_compile_kwargs`,
`dit_cpu_offload`, `dit_layerwise_offload`, `vae_cpu_offload`,
`text_encoder_cpu_offload`, `pin_cpu_memory`, `ltx2_vae_tiling`,
`use_fsdp_inference`, `enable_torch_compile`.
`request_kwargs` at `gpu_pool.py:837` includes:
`ltx2_audio_clean_latent`, `ltx2_audio_denoise_mask`,
`ltx2_video_conditions`, `video_position_offset_sec`, standard sampling
fields.
**Implication**: upstreaming `gpu_pool.py` as-is perpetuates the flat
kwarg surface inside the public server. We need a typed translation
(PR 6 expansion) at the worker boundary before, or as part of, the
gpu_pool upstream.
## Session / continuation semantics today
Per-session state (in `server/main.py`):
- `locked_segment_prompts`, `curated_prompts`, `segment_idx`,
`generated_segment_count`, `loop_iteration`
Per-**GPU** (not per-session) continuation cache (in `gpu_pool.py`):
- `ltx2_continuation_images` — last 9 decoded frames for clip conditioning
- `ltx2_continuation_audio_latents` — denoised audio latents for audio conditioning
Segment N+1 automatically conditions on segment N's trailing frames and
audio. On session reset or handoff (`USER_JOIN`), the per-GPU cache is
cleared. There is currently **no way for a client to serialize and
resume continuation state elsewhere** — it lives on the GPU only.
## Design Decision 1: Continuation model
### Options
**A. Opaque client-round-trip payload** (current plan PR 7 design)
- Server returns `ContinuationState(kind, payload)`; client sends it back.
- Pro: stateless server, trivially load-balanceable, survives disconnects.
- Con: large payloads (frames + audio latents) over every request hop;
bandwidth heavy on multi-segment WebSocket sessions.
**B. Server-held session state** (internal reality)
- Continuation lives per-GPU; implicit between adjacent segments.
- Pro: zero client bandwidth; fast; matches today.
- Con: needs GPU affinity, no resume after disconnect, harder to scale horizontally.
**C. Hybrid** (recommended)
- Server-held is the default for streaming WebSocket sessions.
- Server exposes a `snapshot_state` message that returns the opaque
payload form for migration/retry.
- Stateless HTTP endpoints always use round-trip opaque payloads.
- One serialization format underlies both surfaces.
### Decision: **C (Hybrid)**
Rationale: matches both internal streaming use (server-held, fast) and
stateless API use (client-round-trip, resumable). Cost is one serialization
layer that serves both.
### Implications
- `ContinuationState.kind` identifies the payload schema
(e.g. `"ltx2.v1"`).
- `ContinuationState.payload` must cover:
- trailing conditioning frames (or a tensor reference)
- audio latents (or a tensor reference)
- segment index / rollout position
- any model-specific conditioning metadata (e.g. audio sample rate,
`video_position_offset_sec`)
- For large tensors, payload may reference a server-side blob by ID
rather than inline everything.
- Streaming server has a `SessionStore` keyed by session ID that holds
a typed `LTX2ContinuationState` object.
- `SessionStore.snapshot(session_id) -> ContinuationState` serializes
the current state for export.
- `SessionStore.hydrate(state: ContinuationState) -> session_id` loads
a state into a new session.
- Plan PR 7 expands to cover both surfaces and define the payload schema.
## Design Decision 2: Streaming server layout
### Options
- **A. `fastvideo/entrypoints/streaming/`** — parallel to
`fastvideo/entrypoints/openai/`
- **B. `fastvideo/entrypoints/server/{stateless,streaming}/`** — reorg both
- **C. `fastvideo/streaming/`** — top-level package, not under entrypoints
### Decision: **A (parallel subpackage)**
Rationale: lowest-friction, no existing code moves, both servers share
the same `fastvideo/entrypoints/*` namespace and import style. Shared
utilities can be factored into `fastvideo/entrypoints/server_common/`
later if needed. Option B creates churn across every openai/ import for
marginal organizational win.
### Target layout
```text
fastvideo/entrypoints/
├── openai/ # existing: stateless HTTP POST
│ ├── api_server.py
│ ├── video_api.py
│ ├── image_api.py
│ ├── common_api.py
│ ├── protocol.py
│ ├── state.py
│ ├── stores.py
│ └── utils.py
├── streaming/ # NEW: session WebSocket
│ ├── server.py # FastAPI + WebSocket entry
│ ├── session.py # session lifecycle, state machine
│ ├── session_store.py # typed session state + snapshot/hydrate
│ ├── protocol.py # JSON WebSocket message schemas
│ ├── stream.py # fMP4 encoding (av_fmp4 mode)
│ ├── gpu_pool.py # subprocess workers
│ ├── worker.py # per-GPU worker loop
│ ├── continuation.py # typed LTX2 state payload
│ ├── session_init_image.py
│ ├── session_logger.py
│ ├── mock_server.py
│ ├── prompt/
│ │ ├── enhancer.py # provider-agnostic prompt ops
│ │ ├── rewrite.py
│ │ ├── safety.py # optional fasttext
│ │ ├── payload.py # rewrite payload builder
│ │ └── providers/
│ │ ├── base.py # LLMProvider protocol
│ │ ├── cerebras.py
│ │ ├── cerebras_ifm.py
│ │ └── groq.py
│ └── router/ # or separate top-level package
│ ├── main.py
│ └── registry.py
├── cli/ # existing
└── video_generator.py # existing
```
### Config integration
`ServeConfig` gets an optional `streaming: StreamingConfig | None` field:
```python
@dataclass
class StreamingConfig:
session_timeout_seconds: int = 300
generation_segment_cap: int = 6
stream_mode: Literal["av_fmp4", "legacy_jpeg"] = "av_fmp4"
warmup: WarmupConfig = field(default_factory=WarmupConfig)
pool: GpuPoolConfig = field(default_factory=GpuPoolConfig)
prompt: PromptEnhancerConfig | None = None
safety: PromptSafetyConfig | None = None
@dataclass
class GpuPoolConfig:
num_workers: int | None = None # default: CUDA_VISIBLE_DEVICES count
enable_audio_reencode: bool = True
conditioning_num_frames: int = 9
conditioning_end_offset: int = 0
@dataclass
class PromptEnhancerConfig:
provider: Literal["cerebras_ifm", "cerebras", "groq"] = "cerebras_ifm"
model: str = "gpt-oss-120b"
timeout_ms: int = 20000
system_prompt_dir: str | None = None # hot-reloadable system prompts
@dataclass
class PromptSafetyConfig:
enabled: bool = False
classifier_path: str | None = None
```
## Design Decision 3: LLM provider abstraction
### Problem
`prompt_enhancer.py` (69KB) hard-codes three providers (cerebras_ifm,
cerebras, groq) with provider-specific request/response handling
scattered throughout. Upstreaming as-is locks FastVideo to those three
providers and couples the prompt operations to their response shapes.
### Shape
Introduce an `LLMProvider` protocol:
```python
from typing import Protocol, AsyncIterator, Literal
from dataclasses import dataclass
@dataclass
class LLMMessage:
role: Literal["system", "user", "assistant"]
content: str
@dataclass
class LLMRequest:
messages: list[LLMMessage]
model: str
max_tokens: int | None = None
temperature: float | None = None
timeout_ms: int | None = None
@dataclass
class LLMResponse:
content: str
provider: str
model: str
latency_ms: float
fallback_used: bool = False
class LLMProvider(Protocol):
name: str
async def complete(self, request: LLMRequest) -> LLMResponse: ...
```
### Decision: **Protocol + built-in implementations for cerebras, cerebras_ifm, groq**
Rationale: keeps the prompt enhancer free of provider-specific branching;
users (and future OpenAI/Anthropic/local additions) can register their
own provider without modifying FastVideo. Each built-in provider is
100-200 LOC; the enhancer becomes provider-agnostic prompt orchestration.
### Implications
- `prompt_enhancer.py` splits into `enhancer.py` (prompt operations) +
`providers/` (IO).
- Config moves from scattered env vars to typed `PromptEnhancerConfig`
under `ServeConfig.streaming.prompt`.
- Hot-reloadable system prompts stay — exposed as a management endpoint
on the streaming server.
- Fallback behavior (retry across providers in priority order) moves
into the enhancer layer, orthogonal to provider implementations.
## Design Decision 4 preamble: what Dynamo expects from FastVideo
Dynamo's backend pattern (observed in
`dynamo/components/src/dynamo/sglang/` and confirmed by PR #7544) is a
**pure Python import** pattern. Dynamo owns the backend subpackage in its
own repo; FastVideo only needs to expose a stable, typed, aggregated
and (later) streaming generation surface.
### Contract surface Dynamo consumes
| Surface | Shape | Notes |
|---|---|---|
| Constructor | `VideoGenerator.from_pretrained(model_path, **typed_kwargs)` | Already exists; `typed_kwargs` must be a stable subset from `GeneratorConfig` — no flat LTX2 legacy kwargs. |
| Sync execution | `generator.generate_video(request: GenerationRequest) -> VideoResult` | Aggregated mode; Dynamo wraps in `asyncio.to_thread` under an `asyncio.Lock`. |
| Async execution | `generator.generate_async(request: GenerationRequest) -> AsyncGenerator[VideoEvent, None]` | Needed for: (a) streaming server fMP4 chunks; (b) future Dynamo disaggregation. Events: `Progress`, `Partial?`, `Final`. |
| Typed request | `fastvideo.api.GenerationRequest`, `SamplingConfig`, `InputConfig` | Stable import path; Dynamo's adapter builds this from `NvCreateVideoRequest` + `VideoNvExt`. |
| Typed result | `VideoResult` with `video_bytes` or tensor frames, plus `ContinuationState?` | Must be picklable / JSON-serializable enough for Dynamo RPC. |
| Continuation | `ContinuationState(kind, payload)` with schema-versioned payloads | Used by FastVideo's session store today; tomorrow by Dynamo disaggregated workers. |
| Health check input | `VideoGenerator.default_health_check_request() -> GenerationRequest` | Minimal 256x256 / 8 frames / 1 step; lets Dynamo's `FastVideoHealthCheckPayload.to_dict()` produce the Dynamo `health_check_payload` kwarg without knowledge of FastVideo internals. |
| Config dump | `GeneratorConfig.to_dict()` / `ServeConfig.to_dict()` | Dynamo calls `dynamo.common.config_dump.dump_config(path, config)` at worker start; we already have `config_to_dict()`. |
### Request/response mapping (Dynamo ↔ FastVideo)
Dynamo's video protocol (`NvCreateVideoRequest` / `NvVideosResponse`):
```
NvCreateVideoRequest -> fastvideo.api.GenerationRequest
prompt -> sampling.prompt
size="WxH" -> sampling.width, sampling.height
seconds -> (seconds * nvext.fps) -> sampling.num_frames
input_reference -> input.image_path / input.video_path
nvext.fps -> sampling.fps
nvext.num_frames -> sampling.num_frames (overrides seconds*fps)
nvext.num_inference_steps -> sampling.num_inference_steps
nvext.guidance_scale -> sampling.guidance_scale
nvext.seed -> sampling.seed
nvext.negative_prompt -> sampling.negative_prompt
response_format -> (handled by adapter at output)
VideoFinalEvent -> NvVideosResponse
video_bytes -> data[0].b64_json (if response_format=b64_json)
video_url (after upload) -> data[0].url (if response_format=url)
metadata.inference_time_s -> inference_time_s
```
All fields already exist (or will exist after PR 6 expansion) on
FastVideo's typed schema. No FastVideo changes required beyond what the
rest of this plan already covers **except**:
1. `generate_async` must exist (new in PR 7.10).
2. `default_health_check_request()` helper (new in PR 7.10).
3. The sync `generate_video(request=...)` path must be reachable without
extra wrapping (exists since PR 2; confirm stability).
### Where the Dynamo subpackage lives
The Dynamo-side integration (`FastVideoHandler`, `register_fastvideo_model`,
`FastVideoHealthCheckPayload`, args parsing, main.py, Dockerfile,
request/response mapping) lives **entirely in the Dynamo repo** at
`components/src/dynamo/fastvideo/`, matching the pattern used by vllm
and sglang. FastVideo does **not** host any Dynamo-related subpackage,
Dynamo dependency, or Dynamo-specific CLI. FastVideo's only obligation
is to expose a clean, stable, typed Python API that Dynamo's backend
package can import.
## Design Decision 4: Dynamo as first-class backend target
### Problem
PR #7544 (closed) shows two frictions with the pre-refactor API:
1. **Flat legacy kwargs** — the Dynamo handler had to know about
LTX2-specific flat names.
2. **Sync-only generation** — Dynamo's async handler wrapped
`generator.generate(...)` in `asyncio.to_thread` under a lock; no
progress streaming, no disaggregation path.
The refactor's stateless OpenAI server, WebSocket streaming server, and
Dynamo backend all want the same thing: **a typed async API that yields
progress events and a typed final result**. If we build it once in
`VideoGenerator`, all three adapters become thin.
### Options
**A. Keep sync-only, each adapter wraps**
- Simple; matches PR #7544.
- Con: streaming server needs its own async runner; Dynamo loses progress
streaming; no path to disaggregation.
**B. Add async event stream to `VideoGenerator`**
- `generate_async(request) -> AsyncGenerator[VideoEvent, None]`.
- Sync `generate_video` becomes a thin `asyncio.run` wrapper internally.
- Pro: one canonical execution API; streaming server, OpenAI server,
and Dynamo all consume events directly.
- Con: larger delta in `VideoGenerator` — must thread async through the
pipeline step loop.
**C. Queue-based `generate(request, event_cb)` callback**
- Middle ground; callback receives events.
- Pro: no async rewrite needed.
- Con: callers have to invert control; awkward for Dynamo's async
handler.
### Decision: **B (async event stream)**
Rationale: one substrate serves all three consumers. The cost is a
`generate_async` implementation that runs the pipeline step loop in a
thread and bridges events back via an asyncio queue — standard pattern,
limited surface area.
### Implications
- New PR 7.10 adds `generate_async` on `VideoGenerator` with three event
types: `VideoProgressEvent(step, total_steps, stage)`,
`VideoPartialEvent(frames_ndarray, index)` (optional; emitted only in
the streaming path), `VideoFinalEvent(video_bytes_or_tensor, metadata,
continuation_state?)`.
- Sync `generate_video(request=...)` becomes `asyncio.run(...)` over
`generate_async`, collecting events and returning the final.
- Streaming server's fMP4 encoder consumes `VideoPartialEvent` frames
directly, never re-decoding through disk.
- Dynamo adapter consumes `generate_async` and yields one
`NvVideosResponse` per `VideoFinalEvent` (aggregated mode; ignores
intermediate events today; can surface progress via Dynamo's
status/progress fields in the future).
- `ContinuationState` can be attached to `VideoFinalEvent.metadata`,
giving Dynamo a first-class way to surface state for disaggregation
later.
- Stable public exports: `from fastvideo import VideoGenerator`;
`from fastvideo.api import GenerationRequest, SamplingConfig,
ContinuationState, VideoResult, VideoEvent`.
- No Dynamo subpackage, dep, or CLI lives in FastVideo. The adapter
(`NvCreateVideoRequest ↔ GenerationRequest` mapping, handler,
registration) lives entirely in the Dynamo repo at
`components/src/dynamo/fastvideo/`.
### Constraints this adds to earlier PRs
- **PR 6** (typed LTX2 kwargs): every flat kwarg must have a typed home
**reachable from `GeneratorConfig`**, so Dynamo can construct the
generator without importing internal compat paths.
- **PR 7** (continuation state): `ContinuationState.payload` must be
JSON/YAML serializable (no raw torch tensors inline; use blob
indirection) so it survives Dynamo RPC transport.
- **PR 7.5** (streaming skeleton): consume `generate_async` rather than
re-implementing a progress loop around `generate_video`.
- **PR 2/3/4 already landed**: the typed request shape is fixed and
matches Dynamo's mapping needs — no backtracking required.
## Revised PR sequence (PR 5 onwards)
PRs 0-4 are unchanged and already landed. PR 5 is narrowed; PRs 5.5-7.9
are new inserts; PRs 8-13 are reshaped or kept.
| # | Title | Change | Key deliverables |
|---|---|---|---|
| **5** | Stateless `ServeConfig.default_request` merge | **Narrowed.** Wire typed default-request into `fastvideo/entrypoints/openai/`. | `_merge_default_request` helper, validated-against-preset, tests for default+user-override precedence |
| **5.5** | Server architecture split | **NEW.** Introduce `fastvideo/entrypoints/streaming/` subpackage skeleton. No behavior change. | Empty subpackage + stub server.py; CLI subcommand `fastvideo streaming-serve` (raises NotImplementedError); doc on layout |
| **6** | LTX2 public preset + stage overrides + config colocation | **Expanded.** Also add typed replacements for every flat kwarg used by internal `gpu_pool.py`. | `ltx2_two_stage` preset, `LTX2RefineStageOverride`, `CompileConfig` field types, typed `FP4Config` integration, colocation |
| **7** | Continuation state (public + session) | **Expanded.** Define both opaque payload AND server-held session store. | `ContinuationState.payload` schema, `LTX2ContinuationState` typed subclass, `SessionStore` interface, snapshot/hydrate APIs |
| **7.5** | Streaming server skeleton | **NEW.** Minimum viable WebSocket server: session lifecycle, JSON messages, fMP4 output, single-generator. | `server.py`, `session.py`, `protocol.py`, `stream.py` (fMP4), typed `StreamingConfig` |
| **7.6** | GPU pool upstream | **NEW.** Upstream `gpu_pool.py` with typed config boundary. | `gpu_pool.py`, `worker.py`, job queue, session-to-GPU binding, session timeout handling |
| **7.7** | Prompt enhancer upstream | **NEW.** Upstream `prompt_enhancer.py` with `LLMProvider` abstraction. | `prompt/enhancer.py`, `prompt/providers/{base,cerebras,cerebras_ifm,groq}.py`, hot-reloadable system prompts |
| **7.8** | Streaming auxiliaries | **NEW.** Small, isolated. | `prompt/safety.py`, `session_init_image.py`, `prompt/rewrite.py`, `session_logger.py`, `mock_server.py` |
| **7.9** | Router upstream | **NEW.** Multi-replica load balancer + WS proxy. | `streaming/router/` (or separate top-level package), health checks, WS proxy |
| **7.10** | Dynamo backend contract | **NEW.** Add `VideoGenerator.generate_async` event stream + `default_health_check_request()` helper. FastVideo exposes the async API only; the Dynamo backend package (handler, adapter, registration) lives entirely in the Dynamo repo at `components/src/dynamo/fastvideo/`. Streaming server (PR 7.5) and Dynamo backend both consume the same async API. | `generate_async` with `VideoProgressEvent`/`VideoPartialEvent`/`VideoFinalEvent`; sync `generate_video` becomes a thin wrapper; contract tests against a mock Dynamo-style handler that imports only public FastVideo APIs |
| **8** | Internal-UI ↔ public-server contract docs & tests | **Reframed.** Was "Dreamverse Server Adaptation Layer." Also covers Dynamo integration reference. | WebSocket protocol reference, contract tests, migration examples, Dynamo adapter example that upstream PR can copy verbatim |
| **9** | LongCat preset migration + colocation | **Keep.** | Stage overrides, colocation |
| **10** | Hunyuan15 SR preset migration + colocation | **Keep.** | Stage overrides, SR field migration POC, colocation |
| **11** | SSIM / perf test migration | **Keep.** Now blocked on PR 6 expansion. | Typed API migration of golden tests |
| **12** | Docs + examples | **Keep, expand.** | Streaming server docs now part of scope |
| **13** | Deprecation + cleanup | **Keep, expand.** | Also deprecate flat kwargs that internal gpu_pool uses today |
Total PR count: 13 → ~20 (13 original + 5 streaming-upstream inserts +
1 architecture split + 1 Dynamo contract). Each new PR is small and
self-contained because the streaming components are already cleanly
separated in the internal repo, and the Dynamo contract rides on top of
the async API that the streaming server already needs.
## Open questions
1. **Router: in-repo or separate package?** — It's orthogonal to inference;
in-repo couples deploy cycles, separate leaves FastVideo cleaner.
Recommendation: separate package `fastvideo-router/` or
`fastvideo/contrib/router/`; defer final call to PR 7.9.
2. **Session ID authority** — internal uses ad-hoc client IDs.
Recommendation: server-generated UUID, accept externally provided
session ID only for resume flows.
3. **Torch compile kwargs typing** — `CompileConfig.kwargs: dict[str, Any]`
today accepts `mode`, `backend`, `fullgraph`, `dynamic`. Options: keep
as opaque dict; fully type; hybrid (type the common four + allow
extras). Recommendation: hybrid, type common fields.
4. **Prompt safety / fasttext dependency** — heavy for users who don't
need it. Recommendation: ship as optional extra
`pip install fastvideo[prompt-safety]`.
5. **Audio-specific tensor payloads** — `ltx2_audio_clean_latent`,
`ltx2_audio_denoise_mask`, `ltx2_audio_latents` are not in the current
public schema. PR 7 should classify them (probably as opaque fields
inside `LTX2ContinuationState.payload`, not top-level sampling fields).
6. **Batching behavior** — internal `test_batching.py` suggests batching
is exercised. Scope this into PR 7.5 or defer to a post-cleanup perf PR?
7. ~~**Dynamo subpackage home**~~ — **Resolved.** No Dynamo code lives
in FastVideo. The full backend package (handler, adapter,
registration, health check) is owned by the Dynamo repo at
`components/src/dynamo/fastvideo/`, same pattern as vllm/sglang.
FastVideo only guarantees the public API contract listed above.
8. **Disaggregation readiness** — PR #7544 is aggregated-only. Our
`ContinuationState` hybrid already supports a future prefill/decode
split (prefill yields state; decode hydrates it). Should PR 7.10
explicitly validate that `ContinuationState` survives round-trip
through a Dynamo-style RPC (pickle or JSON), even though Dynamo
isn't using it today? Recommendation: yes; cheap contract test that
prevents drift.
9. **Dynamo progress/status passthrough** — `NvVideosResponse` has
`status` and `progress` fields. Should PR 7.10's handler contract
emit intermediate `NvVideosResponse` chunks keyed off
`VideoProgressEvent`, or stay aggregated-final-only to match PR
#7544? Recommendation: stay aggregated-final for PR 7.10; revisit
after Dynamo clarifies their streaming/progress semantics.
## Immediate path forward
1. Land `will/api_5` cleanup commits — **done** (`e03ca7d9`, `41f93179`
force-pushed without Claude co-author).
2. Review this plan with a human — commit the doc to capture the state.
3. Execute PR 5 (narrow stateless merge) and PR 5.5 (subpackage split)
in parallel. Both small; both unblock the streaming upstream that
follows.
4. Start PR 6 expansion (typed replacements for flat LTX2 kwargs) as the
critical path for PR 7.6 (gpu_pool upstream).
@@ -1,93 +0,0 @@
# Exploration Log: Video Generator Config API Design
## Status: draft
## Context
FastVideo's Python inference API currently mixes generator-instance settings,
pipeline initialization settings, and per-request sampling/runtime settings
through broad `**kwargs` surfaces on `VideoGenerator.from_pretrained(...)` and
`VideoGenerator.generate_video(...)`.
This exploration compares the current FastVideo design with
`sglang/multimodal_gen` and examines how to upstream multi-stage LTX2 /
Dreamverse behavior without growing more ad hoc top-level flags.
## Progress
- [x] Read FastVideo onboarding, codebase map, and relevant design docs.
- [x] Inspect current FastVideo generator, args, sampling, registry, and
workflow abstractions.
- [x] Inspect internal LTX2 streaming server usage and current two-stage /
continuation requirements.
- [x] Inspect SGL diffusion generator, server args, sampling params, and
request preparation boundary.
- [x] Inspect vLLM-Omni stage config, stage metadata, request, and orchestration
surfaces for multi-stage pipeline ideas.
- [x] Inspect current FastVideo CLI/config-file loading and compare with the
training YAML-only entrypoint.
- [ ] Convert findings into a concrete implementation plan for FastVideo.
## Findings
- FastVideo already has the right internal separation points:
`FastVideoArgs`, `PipelineConfig`, `SamplingParam`, and `ForwardBatch`.
- The public boundary is the unstable part:
init-time and request-time knobs are mixed through `**kwargs`.
- Unknown init keys can be silently filtered, while unknown request keys can be
only logged rather than rejected. This makes API drift hard to detect.
- SGL's split is cleaner:
`ServerArgs` for engine/runtime, `PipelineConfig` for model-family wiring,
and `SamplingParams` for per-request settings.
- SGL also has better merge semantics for user request overrides:
it preserves model defaults, tracks explicitly provided fields, and validates
request params against pipeline task type.
- SGL still has a design smell worth avoiding in FastVideo:
`SamplingParams._adjust(...)` depends on `ServerArgs`, which leaks
engine/pipeline concerns back into the request object.
- vLLM-Omni contributes a useful extra abstraction beyond SGL:
model-owned multi-stage topology via `ModelPipeline` and `StageConfig`,
with per-stage defaults (`default_sampling_params`) and runtime override
layering.
- vLLM-Omni's best reusable idea for FastVideo is not the serving stack, but
the separation between:
1. model-defined stage topology and per-stage defaults,
2. runtime engine overrides,
3. request-time sampling/state handoff.
- vLLM-Omni also shows the downside of exposing stage-indexed request lists too
directly: `sampling_params_list` works for a serving engine, but is too
positional and low-level for FastVideo's higher-level Python API.
- FastVideo already supports YAML/JSON config files for inference CLI, but the
current mechanism flattens nested documents back into argparse flags. This
preserves backward compatibility but keeps the CLI surface as the canonical
schema instead of a typed document model.
- The training stack has a cleaner precedent: a YAML-first config loaded into a
typed schema, with dotted CLI overrides applied onto the nested document
before parsing. Inference can likely adopt a lighter variant of that pattern.
- Multi-stage generation should be unified at the orchestration layer, not by
forcing LongCat refine, Hunyuan SR, and LTX2 continuation into one leaf config.
## Mistakes / Dead Ends
- A fully free-form string-dict API would lose too much type safety and would
likely recreate the current drift problem under a different shape.
- A single universal `RefineConfig` for all models would become a sparse bag of
nullable fields and would not map cleanly to existing model families.
## Proposed Standardization
- Introduce a typed public split:
`GeneratorConfig` for instance-lifetime engine/init settings and
`GenerationRequest` for per-call inputs/sampling/output.
- Allow dict input only as an interchange layer that is parsed immediately into
typed configs with strict unknown-key validation.
- Add a typed `GenerationPlan` / multi-stage orchestration layer with
discriminated stage configs:
`SampleStageConfig`, `LongCatRefineStageConfig`,
`HunyuanSRStageConfig`, `LTX2ContinuationStageConfig`.
- Let model families own stage defaults and stage topology through named
profiles or model-defined stage plans, similar in spirit to vLLM-Omni's
pipeline YAMLs, but expose them through typed Python config objects rather
than raw stage-indexed lists in the primary API.
- Make YAML/JSON a first-class serialization of the same typed inference
schema, not just a file format that expands into CLI flags.
- Prefer a YAML-first CLI pattern for nested configs:
`fastvideo generate --config run.yaml --request.sampling.seed 42`,
while keeping a compatibility layer for existing flat flags during migration.
- Upstream LTX2 two-stage / continuation behavior as a first-class stage or
pipeline profile rather than more `ltx2_*` top-level kwargs.
@@ -1,196 +0,0 @@
# Current State — 2026-05-06 (D-26 rebase onto origin/main; PR base flipped)
Point-in-time snapshot of branches, commits, and live infrastructure.
Update whenever commits land or services restart.
For HOW to commit / push / verify see [runbook.md](runbook.md). For
roster of co-authors to credit on every commit see
[authors.md](authors.md).
## Branch tips
| Repo | Branch | Tip | Distance |
|---|---|---|---|
| FastVideo | `will/ltx2_sr_port` (**PR #1288 head**) | `fbd823df` | merged-into-main pending; latest tip post-D-16 + integration-review + integration-plan + GPU4 smoke validation |
| FastVideo | **`will/dreamverse-monorepo`** (**REBASED ONTO `origin/main`**, was forked from `will/ltx2_sr_port`) | `83829c5e` | 66 commits ahead of `origin/main` (`c17d33bf`). Contains the full LTX-2 SR port + NVFP4 + Dreamverse monorepo migration + audio kwarg fix + warmup + NVENC build + benchmarks + integration memory dir, all rebased to be openable as a single PR against main. PR #1288 on `will/ltx2_sr_port` is untouched. End-to-end verified on GPU4 pre-rebase with audio continuation across segments 1→2 (`Cached audio latents shape=(1, 8, 126, 16) for segment 2`, `Segment 2: relayed av chunks=22, bytes=3.8MB`, no BrokenPipeError). NVENC build supported but not usable on this dev host (B200 has no NVENC silicon — verified by direct ffmpeg probe). See [decisions-log.md D-19](decisions-log.md#d-19) + [D-20](decisions-log.md#d-20) + [D-21](decisions-log.md#d-21) + [D-26](decisions-log.md#d-26). |
| FastVideo | `will/dreamverse-monorepo-pre-main-rebase-backup-20260506` | `2ee839a3` | **local-only safety backup** of the pre-rebase chain (the same 66 commits stacked on `2aaeee2a`). Keep until the new chain is fully verified by the next round of e2e on a non-stuck dev box or until the PR merges. |
| FastVideo | `will/api_7.10` | `6ae7a99f` | **deprecated** — PR #1287 closed in favor of #1288. Branch can be deleted on origin and locally; kept for now as historical reference. |
| FastVideo | `will/api_8`, `will/ltx2_sr_runtime`, `will/ltx2_nvfp4`, `will/ltx2_post_fixes`, `will/agents_cleanup` | (various) | **deprecated** split bookmarks. Strategy reversed to single mega-PR (D-17). Safe to delete locally; not pushed to origin. |
| FastVideo | `will/ltx2_sr_port-pre-1286-rebase` | `1baa60bb` | **local-only safety backup** of pre-rebase chain (37 commits); keep until next slice merges |
| Dreamverse | `will/integrate-public-fastvideo` | `ec8ef92` | 10 commits ahead of `737f3c1` (the dep switch) |
| FastVideo-internal | their `main` | (read-only ref) | — |
FastVideo worktree default branch is `will/ltx2_sr_port`. Other agents
share this worktree — if `git branch --show-current` shows something
else, switch back cleanly with `git checkout will/ltx2_sr_port` (don't
disturb their uncommitted work). I observed this happen repeatedly in
the 2026-05-05 session — confirmed harmless; switching back was always
safe with a clean working tree.
## Post-#1286 rebase summary
PR #1286 merged at `2aaeee2a` (squash). `will/ltx2_sr_port` was rebased
onto new `origin/main`, dropping 4 commits whose content is now in main:
- `cd76cf51` `[feat] streaming: router (multi-replica load balancer)`
- `1ac1e732` `[feat] streaming: fastvideo router-serve CLI`
- `b0b7f59c` `[test] streaming: router registry + health loop ...`
- `40e265b8` `[fix] streaming: router polish — bridge cancel + state
machine + deps` (squashed into `2aaeee2a` via cherry-pick `a152cb77`)
Rebase was clean — no conflicts. All 33 surviving commits got new SHAs
(rebase rewrites). The pre-rebase tip `1baa60bb` is preserved on the
local backup branch `will/ltx2_sr_port-pre-1286-rebase`.
## New linearized chain (33 commits, slice indices for STACK.md)
| Slice | PR | Commits | Tip SHA | Subject |
|---|---|---|---|---|
| 1-3 | 7.10 (PR #1287) | 3 | `6ae7a99f` | `[test] streaming: generate_async coverage + refreshed streaming test` |
| 4-6 | 8 | 3 | `f32e31ec` | `[test] streaming: contract tests for Dreamverse + Dynamo shapes` |
| 7-15 | LTX-2 SR | 9 | `e7297519` | `feat(ltx2): full i2v conditioning + continuation latent port` |
| 16-21 | NVFP4 | 6 | `6793166b` | `test(nvfp4): lock LTX-2 wiring + typed transformer_quant flow` |
| 22-23 | LTX-2 post-fixes | 2 | `25897b67` | `[fix]: unwrap list-of-generator before torch.randn in LTX-2 latent prep` |
| 24-33 | agents_cleanup | 10 | `b34d9704` | `[docs] dreamverse-integration: add runbook + fresh-context onboarding` |
5 PRs landed (7.5, 7.6, 7.7, 7.8, 7.9), 1 in flight (7.10), 5 remaining
(8 / LTX-2 SR / NVFP4 / post-fixes / agents_cleanup).
## Historical commit chain analysis (pre-#1286 rebase)
The layered chain analysis below documented the pre-rebase SHAs (LTX-2
SR layer, NVFP4 layer, post-handoff fixes layer). Those SHAs no longer
exist on `will/ltx2_sr_port` — they live only on
`will/ltx2_sr_port-pre-1286-rebase`. Content semantics are unchanged;
SHAs were rewritten by the rebase. Kept here for narrative continuity.
## FastVideo: commit chain `cfccd292..156103b9`
Three layers since LTX-2 i2v port:
### Layer 1 — LTX-2 SR port + alignment harness (5 commits)
```
365a66c7 feat(quantization): upstream LTX-2 FP4Config with lazy flashinfer
433d26b2 feat(ltx2): port LTX-2 SR runtime — upsampler, refine stages, refine args
751d05de feat(ltx2): wire SR pipeline graph + port denoising/latent-prep stages
af6bbfea test(ltx2-sr): add numerical alignment harness — public vs internal
974cd430 fix(ltx2-sr): close port gaps surfaced by alignment harness retries
b6ac7630 test(ltx2-sr): pin ltx2 sampling knobs in harness for parity diff
b043d550 fix(api): align public SamplingParam ltx2 defaults with distilled
663dda80 fix(registry): order LTX-2 detectors so distilled wins for distilled paths
cfccd292 feat(ltx2): full i2v conditioning + continuation latent port (BASE)
```
(Predates the May 2 handoff.)
### Layer 2 — NVFP4 wire-up + per-component compile (6 commits, May 2 handoff)
```
a4760bae fix(api): propagate generic refine_* args + match internal randn
221cb20a feat(api): typed per-component CompileConfig + FastVideoArgs carriers
6da342ba feat(compile): per-component compile + transformer_refine + prepare hook
42b30bf9 feat(ltx2): wire FP4 inference through fastvideo.layers.quantization
94c983a2 refactor(quant): rename FP4 → NVFP4 to disambiguate from other FP4 variants
c6c14c55 test(nvfp4): lock LTX-2 wiring + typed transformer_quant flow
```
See [quantization.md](quantization.md) for what each commit locks in.
### Layer 3 — Post-handoff parity/perf fixes (3 commits, since May 2)
```
a5fcd19c [fix]: lazy-import flash_attn 2 fallback in attention backend
d4ee5be2 [fix]: avoid model.to() round-trip in Gemma encoder forward
156103b9 [fix]: unwrap list-of-generator before torch.randn in LTX-2 latent prep (HEAD)
```
Three small fixes — no new features. Continued parity tightening with internal.
## Dreamverse: commit chain `737f3c1..ec8ef92`
```
737f3c1 chore: switch fastvideo dep from FastVideo-internal to public FastVideo
4cc6b30 chore: gitignore Playwright + Next.js build artifacts under apps/web
33caa92 test(e2e): align Playwright specs with the actual production composer
6fd137c test(e2e): tighten frontend-shell + preset specs to match actual UI
248060b test(e2e): add Playwright tier with backend-health smoke + preset run
d80c2a8 refactor(server): drive FP4 + per-component compile via typed GeneratorConfig
3d7fd89 feat(skill): launch-demo orchestrator + fastvideo serve YAML
72f69b9 Update ffmpeg installation instructions.
1ba5635 fix(server): block startup on GPU warmup readiness, propagate failures
ec8ef92 fix(server): detect worker death in _send_command via proc.sentinel (HEAD)
```
The post-handoff trio (`72f69b9`, `1ba5635`, `ec8ef92`) hardens server
startup robustness — ffmpeg install docs, GPU warmup readiness gate, and
worker-death detection.
## Live services (do not duplicate)
| Port | Service | PID | Status |
|---|---|---|---|
| 8009 | `dreamverse-server` | 705513 | `/readyz` returns 200, 1 GPU worker on GPU 4, NVFP4 (50.9 GiB), `ENABLE_TORCH_COMPILE=0`, `FASTVIDEO_FFMPEG_BIN=$HOME/opt/ffmpeg-native/bin/ffmpeg`, `FASTVIDEO_VIDEO_CODEC=libx264` (post-D-20 deploy at 2026-05-05 14:??) |
| 5274 | `next-server` (dev) | 707746 | 200 |
| 8000 | unknown FastAPI | — | **Not in handoff.** Probably stray `fastvideo serve`. Verify with `lsof -i :8000` before launching a new BE on the default port. |
## Stashes — DO NOT POP
| Repo | Stash | Reason |
|---|---|---|
| FastVideo | `stash@{0}: WIP on main: 71bfc13d HunyuanVideo plugin` | Pre-existing, unrelated to integration work |
| Dreamverse | `stash@{0}: wip: server modular refactor (split config/prompting/runtime/session)` | 3867-line orphan modular split, **not part of `will/integrate-public-fastvideo`**. Recover on a separate branch if needed. |
## Test status
| Suite | Status |
|---|---|
| FastVideo `fastvideo/tests/api/` (post-D-20) | 185 passed (was 222 before some tests moved; new `test_extra_overrides_routing.py` adds 7) |
| FastVideo `contract/` + `nvfp4_*` + `ltx2_pipeline_smoke` (May 2 handoff) | 222 passed, 1 skipped |
| Playwright e2e against live BE+FE (D-19) | 8 passed (5 backend-health + 2 frontend-shell + 1 preset-prompt-generation) |
| Live segment-1→segment-2 audio continuation (D-20) | passes — `Cached audio latents shape=(1, 8, 126, 16) for segment 2` + `Segment 2: relayed av chunks=22, bytes=3.8MB`, no BrokenPipeError |
| `dreamverse-deploy.sh` flag parser standalone test (D-20) | 13/13 permutations pass + bad-flag rejection (defaults / single flags / both flags / `--no-*` overrides / env-only / flag-overrides-env / both-env+both-no-flags / flags interleaved with positional args) |
| `fastvideo serve --config streaming_demo.yaml` validation (May 2 handoff) | parses cleanly; dotted overrides work |
| `bash -n` on `apps/dreamverse/scripts/install_native_ffmpeg.sh` + `dreamverse-deploy.sh` (D-20) | clean |
## Pre-existing failures (NOT caused by this work)
| Test | Failure | Notes |
|---|---|---|
| `fastvideo/tests/ops/quantization/test_absmax_fp8.py::test_create_weights_rejects_invalid_dtype` | `AssertionError not raised` | Pre-existing on `main`. Verified via `git stash` that NVFP4 work doesn't introduce it. See [open-threads.md](open-threads.md) item #2. |
## Source docs (archived 2026-05-03)
The 7 source docs that this memory dir consolidates have been moved into
[`source-archive/`](source-archive/) — see the
[archive README](source-archive/README.md) for the archive policy and
synthesis mapping.
Other untracked items at the FastVideo repo root:
- Nested clones: `dynamo/`, `ray/`, `vllm-omni/`
- Lock files: `uv.lock`, `fastvideo/tests/ssim/.reference_videos_download.lock`
- Skill dirs: `.agents/skills/diagnose-ssim-failure/`, `.agents/skills/review-pr-link/`
- `.agents/exploration/pr-link-review.md` (kept; already promoted to a skill)
## Quick orientation commands
```bash
# FastVideo state
cd /home/william5lin/FastVideo
git log --oneline cfccd292..HEAD # 14 commits this round
# Dreamverse state
cd /home/william5lin/Dreamverse
git log --oneline 737f3c1..HEAD # 10 commits this round
# Live stack health (already running)
curl -s http://localhost:8009/readyz | head -c 300
curl -s http://localhost:5274/ -o /dev/null -w "%{http_code}\n"
# Re-verify test suite
.venv/bin/python -m pytest fastvideo/tests/api/ \
fastvideo/tests/contract/ \
fastvideo/tests/ops/quantization/test_nvfp4_*.py \
tests/local_tests/pipelines/test_ltx2_pipeline_smoke.py \
-q --no-header
```
@@ -1,389 +0,0 @@
# Streaming Server Upstream — PRs 5.5 → 7.10
The `FastVideo-internal/ui/ltx2-streaming/server/` stack is being
upstreamed into public FastVideo at `fastvideo/entrypoints/streaming/`.
In parallel, FastVideo is becoming a first-class Dynamo backend (same
tier as vllm, sglang, trtllm). This file covers both threads since they
share `generate_async` as the substrate.
For PR sequence/status see [pr-roadmap.md](pr-roadmap.md). For the
Dreamverse-side adoption see [cross-repo-surfaces.md](cross-repo-surfaces.md).
**Last updated:** 2026-05-03.
## What's being upstreamed
| Internal path | Size | Role | Public target |
|---|---|---|---|
| `server/main.py` | 94 KB | FastAPI + WebSocket, session lifecycle, segment orchestration | `fastvideo/entrypoints/streaming/server.py` + handlers |
| `server/gpu_pool.py` | 66 KB | GPU orchestration, subprocess workers | `fastvideo/entrypoints/streaming/gpu_pool.py` |
| `server/prompt_enhancer.py` | 69 KB | LLM orchestration (cerebras_ifm, cerebras, groq) | `fastvideo/entrypoints/streaming/prompt/` package |
| `server/mock_server.py` | 45 KB | Mock backend for dev/tests | `fastvideo/entrypoints/streaming/mock_server.py` |
| `server/prompt_safety.py` | 7 KB | Optional fasttext-gated prompt safety | `prompt/safety.py` |
| `server/session_init_image.py` | 3 KB | i2v init image handling | `streaming/session_init_image.py` (PR 7.5, already public) |
| `server/rewrite_prompt_payload.py` | 3 KB | Rewrite flow payload builder | `prompt/rewrite.py` |
| `server/session_logger.py` | 1 KB | Session JSONL logs | `streaming/session_logger.py` |
| `server/config.py` | 9 KB | Env-driven server config | typed `ServeConfig.streaming` extensions |
| `router/main.py` | 27 KB | Multi-replica load balancer + WS proxy | `fastvideo/entrypoints/streaming/router/` |
| `slurm/` | — | Deployment scripts | Stays internal |
Frontend clients (`client/`, `prod-ui/`) stay in the internal repo.
## Four design decisions that shape the upstream
### D-1: Continuation model — Hybrid (server-held + opaque client-round-trip)
Streaming WebSocket sessions hold continuation per-GPU (matches today's
internal behavior, fast, zero client bandwidth). Stateless HTTP endpoints
use opaque round-trip payloads. Server exposes a `snapshot_state` message
that returns the opaque form for migration/retry.
One serialization layer underlies both surfaces.
Implementation: `SessionStore` (in-memory default, pluggable for
redis/etc.) keyed by session ID, holds typed `LTX2ContinuationState`.
- `snapshot(session_id) -> ContinuationState` exports for migration
- `hydrate(state: ContinuationState) -> session_id` loads state into new session
Payload schema covers: trailing conditioning frames (or tensor-blob ID),
audio latents (or blob ID), segment index, audio sample rate,
`video_position_offset_sec`, model-specific metadata.
Landed in PR 7. See [cross-repo-surfaces.md](cross-repo-surfaces.md) for
the full wire format.
### D-2: Streaming server layout — Parallel subpackage `fastvideo/entrypoints/streaming/`
Sits next to `fastvideo/entrypoints/openai/`. No existing code moves.
Both servers share the `entrypoints/*` namespace. Shared utilities can be
factored into `fastvideo/entrypoints/server_common/` later if needed.
### D-3: LLM provider abstraction — `LLMProvider` protocol + built-in providers
`prompt_enhancer.py` (69 KB) hard-coded three providers (cerebras_ifm,
cerebras, groq) with provider-specific request/response handling
scattered throughout. Upstreaming as-is would lock FastVideo to those
providers.
Protocol shape:
```python
@dataclass
class LLMRequest:
messages: list[LLMMessage]
model: str
max_tokens: int | None = None
temperature: float | None = None
timeout_ms: int | None = None
@dataclass
class LLMResponse:
content: str
provider: str
model: str
latency_ms: float
fallback_used: bool = False
class LLMProvider(Protocol):
name: str
async def complete(self, request: LLMRequest) -> LLMResponse: ...
```
PR 7.7 ships built-in providers for cerebras, groq. **Public Literal
currently restricts to `Literal["cerebras", "groq"]`** — `cerebras_ifm`
is internal-only and remains environment-driven on `dreamverse-server`.
See [open-threads.md](open-threads.md) follow-up #3.
Hot-reloadable system prompts via management endpoint. Sequential
fallback across providers in priority order — race-based fallback (the
internal optimization) deferred per [decisions-log.md](decisions-log.md)
D-3.
### D-4: Dynamo as first-class backend target — async event stream
PR ai-dynamo/dynamo#7544 (closed draft) showed two frictions:
1. Flat legacy kwargs — Dynamo handler had to know LTX-2-specific names.
2. Sync-only generation — Dynamo wrapped `generator.generate(...)` in
`asyncio.to_thread` under a lock; no progress streaming, no
disaggregation path.
Decision: **add `generate_async`** as the canonical execution API.
```python
async def generate_async(
self,
request: GenerationRequest,
) -> AsyncGenerator[VideoEvent, None]: ...
```
Events:
```python
@dataclass
class VideoProgressEvent:
step: int
total_steps: int
stage: str # "denoise" | "refine" | "decode" | ...
@dataclass
class VideoPartialEvent:
frames: np.ndarray # (num_frames, H, W, 3)
index: int # monotonic chunk index
@dataclass
class VideoFinalEvent:
video_bytes: bytes | None
tensor: torch.Tensor | None
metadata: dict[str, Any]
continuation_state: ContinuationState | None
```
The sync `generate_video(request=...) -> VideoResult` becomes a thin
`asyncio.run` wrapper over `generate_async` that collects events and
returns the final.
**Three consumers, one substrate:**
| Consumer | Transport | Request shape | State |
|---|---|---|---|
| Stateless OpenAI (`fastvideo/entrypoints/openai/`) | HTTP POST | `GenerationRequest` merged onto `ServeConfig.default_request` | Stateless; opaque payload |
| Streaming WebSocket (`fastvideo/entrypoints/streaming/`) | WebSocket JSON + binary fMP4 | `GenerationRequest` per segment, session-scoped | Server-held; per-GPU continuation cache |
| Dynamo native backend (`ai-dynamo/dynamo/components/src/dynamo/fastvideo/`) | Dynamo RPC endpoint | `NvCreateVideoRequest` ↔ adapter ↔ `GenerationRequest` | Aggregated today; future disaggregated via `ContinuationState` |
**FastVideo does NOT host any Dynamo code.** The full backend package
(`args.py`, `main.py`, `backend.py`, `register.py`, `health_check.py`)
lives entirely in the Dynamo repo at `components/src/dynamo/fastvideo/`,
matching the vllm/sglang pattern. FastVideo's only obligation is the
stable public Python API.
PR 7.10 lands the FastVideo-side contract. Dynamo backend code lives in
ai-dynamo/dynamo (next iteration of #7544 reopens against PR 8 reference
docs).
## Target package layout
```
fastvideo/entrypoints/
├── openai/ # existing: stateless HTTP POST
├── streaming/ # NEW: session WebSocket
│ ├── server.py # FastAPI + WebSocket entry
│ ├── session.py # session lifecycle, state machine
│ ├── session_store.py # typed session state + snapshot/hydrate
│ ├── protocol.py # JSON WebSocket message schemas
│ ├── stream.py # fMP4 encoding (av_fmp4 mode)
│ ├── gpu_pool.py # subprocess workers (PR 7.6)
│ ├── worker.py # per-GPU worker loop
│ ├── continuation.py # typed LTX2 state payload
│ ├── session_init_image.py
│ ├── session_logger.py
│ ├── mock_server.py
│ ├── prompt/
│ │ ├── enhancer.py # provider-agnostic prompt ops
│ │ ├── rewrite.py
│ │ ├── safety.py # optional fasttext
│ │ └── providers/
│ │ ├── base.py # LLMProvider protocol
│ │ ├── cerebras.py
│ │ ├── cerebras_ifm.py
│ │ └── groq.py
│ └── router/
│ ├── main.py
│ └── registry.py
├── cli/
└── video_generator.py
```
## Typed config integration
`ServeConfig` gets an optional `streaming: StreamingConfig | None`:
```python
@dataclass
class StreamingConfig:
session_timeout_seconds: int = 300
generation_segment_cap: int = 6
stream_mode: Literal["av_fmp4", "legacy_jpeg"] = "av_fmp4"
warmup: WarmupConfig = field(default_factory=WarmupConfig)
pool: GpuPoolConfig = field(default_factory=GpuPoolConfig)
prompt: PromptEnhancerConfig | None = None
safety: PromptSafetyConfig | None = None
@dataclass
class GpuPoolConfig:
num_workers: int | None = None # default: CUDA_VISIBLE_DEVICES count
enable_audio_reencode: bool = True
conditioning_num_frames: int = 9
conditioning_end_offset: int = 0
@dataclass
class PromptEnhancerConfig:
provider: Literal["cerebras", "groq"] = "cerebras" # cerebras_ifm pending
model: str = "gpt-oss-120b"
timeout_ms: int = 20000
system_prompt_dir: str | None = None # hot-reloadable
@dataclass
class PromptSafetyConfig:
enabled: bool = False
classifier_path: str | None = None
```
## `build_app` route contract — open follow-up
Today `fastvideo.entrypoints.streaming.server.build_app` exposes only:
- `GET /health`
- `WS /v1/stream`
The Dreamverse Next.js shell expects these additional routes that the
upstream plan (and Dreamverse FE today) require:
| Route | Owner per upstream plan | Status |
|---|---|---|
| `GET /healthz` | Streaming-server-side health (FastVideo) | 🔴 NOT YET MIGRATED |
| `GET /readyz` | Streaming-server-side health (FastVideo) | 🔴 NOT YET MIGRATED |
| `GET /status` | Streaming-server-side health (FastVideo) | 🔴 NOT YET MIGRATED |
| `GET /curated-presets` | Operator-side surface (Dreamverse) | 🟡 stays in Dreamverse, FE feature-detects |
| `POST /curated-presets/append` | Operator-side surface (Dreamverse) | 🟡 stays in Dreamverse |
| `GET /prompt-system-config` | Operator-side surface (Dreamverse) | 🟡 stays in Dreamverse |
| Devtools routes | Dreamverse-only | 🟡 stays in Dreamverse |
Until the three health routes migrate into FastVideo's `build_app`, the
`BE_FLAVOR=fastvideo` flavor of `launch_demo.sh` is a "diagnostic" flavor
only (verifies typed serve-config path) — not FE-compatible. See
[open-threads.md](open-threads.md) follow-up #1.
The streaming-upstream plan listed `/healthz`, `/readyz`, `/status`,
`/ws` as the contract that the upstream of `realtime/` → `streaming/`
must preserve. They were deferred from PR 7.5's MVP.
## PR 7.5 status — open as #1251
Single-generator WebSocket end-to-end shipped (8 commits):
1. `feat(streaming): protocol schemas + session state machine`
2. `feat(streaming): fMP4 encoder + session init-image persistence`
3. `feat(streaming): single-generator WebSocket server entry`
4. `test(streaming): server lifecycle + protocol + fMP4 coverage`
5. `docs(streaming): server contract spec`
6. `fix(streaming): restore missing-streaming-block guard + retire stub-era test`
7. `simplify(streaming): review follow-ups (idle timeout via asyncio.wait_for, _send_error helper, _cleanup_session, Protocol-typed generator, cleanup-on-disconnect)`
8. `fix(streaming): enforce idle timeout on receive_json + flag generator-cancellation gap (TODO → PR 7.10)`
Deferred TODOs (intentionally) blocking on PR 7.10:
- **Per-step progress events** — only terminal `step_complete` today;
needs `generate_async` for per-step `VideoProgressEvent` emission.
- **Mid-segment cancellation on client disconnect** — TODO marker in
`server.py` near `pool.run`. Needs `generate_async`'s cancellation
propagation.
## PR 7.6 status — branch ready, not yet PR'd
`will/api_7.6` (5 commits, rebased on 7.5):
1. `feat [7.6/n]: GPU pool manager with typed worker boundary`
2. `refactor [7.6/n]: route streaming server through GpuPool`
3. `test [7.6/n]: GPU pool coverage (in-process + subprocess)`
4. `fix(streaming): restore missing asyncio import in server` (rebase fixup)
5. `feat(streaming): extract worker.py and add two-segment warmup`
Tests: 17/17 gpu_pool tests + 89/89 streaming tests green.
Ships:
- `GpuPool` ABC + `InProcessGpuPool` + `SubprocessGpuPool` +
`PoolAssignment` / `PoolHealth` / `PoolAcquireTimeout`
- `worker.py` — per-GPU `worker_main` and two-segment warmup helper
- Subprocess startup uses typed `GeneratorConfig`, NOT flat kwargs
- Session-to-GPU binding with timeout + queue for contention
- Two-segment startup warmup per worker (segment 1 fresh + segment 2
with returned `ContinuationState` so both compile branches are primed)
- `SessionStore` (from PR 7) wired for per-GPU continuation cache
Deferred to PR 7.10:
- **Audio re-encode (`LTX2AudioEncoder`, `AudioProcessor`)**: internal
`_re_encode_audio` runs *inside* the per-step streaming loop
(`_stream_av_fmp4_events` / `do_step_ltx2`). The whole-segment
`pool.run()` path PR 7.6 ships doesn't need it. Re-encode is a
per-step streaming concern that belongs with `generate_async`.
- **Deprecate `VideoGenerator.from_pretrained(**flat_kwargs)`**: belongs
with PR 13 cleanup.
## PR 7.10 — the unlock PR
PR 7.10 adds `generate_async` and closes three open threads
simultaneously:
- Q-5 / D-5: audio re-encode for cross-segment continuity
- Q-9: Dynamo progress passthrough
- PR 7.5's mid-segment cancellation TODO (client disconnect →
`asyncio.CancelledError` → GPU work stops)
Plus health-check helper:
```python
def default_health_check_request(self) -> GenerationRequest: ...
# Returns 256x256, 8 frames, 1 step. Lets Dynamo's
# FastVideoHealthCheckPayload.to_dict() produce a Dynamo
# health_check_payload kwarg without knowledge of FastVideo internals.
```
Stable public exports:
```python
from fastvideo import VideoGenerator
from fastvideo.api import (
GenerationRequest, SamplingConfig, ContinuationState,
VideoResult, VideoEvent,
VideoProgressEvent, VideoPartialEvent, VideoFinalEvent,
)
```
Streaming server (PR 7.5) gets rewired to consume `generate_async`
directly — no wrapper duplication.
## Dynamo request/response mapping
```
NvCreateVideoRequest -> fastvideo.api.GenerationRequest
prompt -> sampling.prompt
size="WxH" -> sampling.width, sampling.height
seconds -> seconds * nvext.fps -> sampling.num_frames
input_reference -> input.image_path | input.video_path
nvext.fps -> sampling.fps
nvext.num_frames -> sampling.num_frames (overrides seconds*fps)
nvext.num_inference_steps -> sampling.num_inference_steps
nvext.guidance_scale -> sampling.guidance_scale
nvext.seed -> sampling.seed
nvext.negative_prompt -> sampling.negative_prompt
response_format -> (handled at adapter's output stage)
VideoFinalEvent -> NvVideosResponse
video_bytes -> data[0].b64_json (response_format=b64_json)
uploaded URL -> data[0].url (response_format=url)
metadata.inference_time_s -> inference_time_s
continuation_state -> (reserved for future disaggregation)
```
## Open questions
1. **Router placement** — in-tree at `fastvideo/entrypoints/streaming/router/`
(current implementation per PR 7.9) or separate package
`fastvideo-router/` / `fastvideo/contrib/router/`. Effectively
resolved in-tree by the PR 7.9 implementation.
2. **Session ID authority** — server-generated UUID; accept externally
provided session ID only for resume flows.
3. **Disaggregation readiness contract test** — should PR 7.10 validate
`ContinuationState` survives round-trip through Dynamo-style RPC
(pickle or JSON), even though Dynamo isn't using it today?
Recommended: yes; cheap regression guard.
4. **Dynamo progress/status passthrough** — should PR 7.10's handler
contract emit intermediate `NvVideosResponse` chunks keyed off
`VideoProgressEvent`, or stay aggregated-final-only? Recommended:
stay aggregated-final for PR 7.10; revisit after Dynamo clarifies.
5. **`video_position_offset_sec` semantics** — see [decisions-log.md](decisions-log.md)
open question; needs decision before PR 7.6 emits state.
6. **`SessionStore` / `BlobStore` lifecycle** — eviction, TTL, blob-drop
on state replacement; defer to PR 7.5 design pass.
@@ -1,336 +0,0 @@
# Evaluation Metrics Registry
Living catalog of all evaluation metrics for FastVideo-WorldModel video quality
assessment. Each metric includes a detailed explanation, implementation status,
usage instructions, and interpretation guide.
_Last updated: 2026-03-02_
---
## Metric Summary
| Metric | Category | Status | Location | Trust |
|--------|----------|--------|----------|-------|
| **FVD** | Distribution | ✅ Implemented | `fastvideo/eval/metrics/common/fvd/` | High |
| **SSIM** | Reference | ✅ Implemented | `fastvideo/tests/ssim/` | High |
| **LPIPS** | Perceptual | ✅ Implemented | `scripts/lora_extraction/` | Medium |
| **Loss trajectory** | Training signal | ✅ Implemented | W&B `train_loss` | Medium |
| **Grad norm stability** | Training signal | ✅ Implemented | W&B `grad_norm` | Medium |
| **GameWorld Score** | Multi-dim benchmark | 🟡 External | Matrix-Game repo | Low |
| **Human preference** | Gold standard | 🔴 Manual | N/A | Highest |
---
## Implemented Metrics
### FVD — Fréchet Video Distance
**Category**: Distribution-level quality metric
**Status**: ✅ Registered as the `common.fvd` eval metric in `fastvideo/eval/metrics/common/fvd/`
**Trust**: High — standard protocol, I3D feature extractor (CLIP / VideoMAE backbones also available, research-grade)
#### What It Measures
FVD measures the distance between the **distribution** of generated videos and
a distribution of real/reference videos. It works by:
1. Extracting spatiotemporal features from both real and generated video sets
using a pretrained **I3D** (Inflated 3D ConvNet) model.
2. Modeling each set of features as a multivariate Gaussian (mean + covariance).
3. Computing the **Fréchet distance** between the two Gaussians.
Lower FVD = generated videos are more statistically similar to real videos.
#### Why It Matters
- FVD is the **de facto standard** for benchmarking video generation models.
- It captures both **visual quality** (are individual frames realistic?) and
**temporal coherence** (do frames flow naturally?).
- Matrix-Game 2.0, Open-Sora, and most video generation papers report FVD.
#### Limitations
- Requires a **large sample set** (standard protocol uses 2048 videos) to
produce stable statistics. Small sample sizes yield noisy results.
- Measures **distributional similarity**, not per-video quality. A model could
have low FVD by generating a diverse set of "roughly okay" videos.
- The I3D model was trained on Kinetics-400 (human actions). It may be less
sensitive to domain-specific artifacts in non-human-action videos (e.g.,
driving, game environments).
- Does not directly measure text-video alignment or action controllability.
#### How to Use
```python
# Programmatic — drive the metric directly for custom kwargs
from fastvideo.eval import get_metric
metric = get_metric("common.fvd", extractor="i3d") # or "clip" / "videomae"
metric.to("cuda")
metric.setup()
metric.reset()
# First sample carries the reference set; later samples reuse the cache.
metric.accumulate({"video": gen_tensors[0], "reference": real_tensors})
for gen in gen_tensors[1:]:
metric.accumulate({"video": gen})
result = metric.finalize()
print(f"FVD: {result.score:.2f}")
```
```bash
# CLI — folder of generated mp4s vs a reference folder
python examples/inference/eval/eval_fvd.py \
--gen-dir outputs/gen/ \
--reference-dir data/real/ \
--extractor i3d \
--output fvd_scores.json
```
**Feature extractors**: `i3d` (default, standard FVD spec used in papers),
`clip`, `videomae` (research-grade; not directly comparable to published
FVD numbers).
**Protocol**: standard FVD uses 2048 generated + 2048 reference videos at
16 frames each. A warning fires below 256 — the score becomes
statistically unreliable.
#### Interpretation
| FVD Range | Interpretation |
|-----------|---------------|
| < 100 | Excellent — near-real quality |
| 100–300 | Good — competitive with SOTA |
| 300–600 | Fair — noticeable gap from real |
| > 600 | Poor — significant quality issues |
> FVD values are dataset-dependent. Always compare against baselines evaluated
> on the same real video distribution.
---
### SSIM — Structural Similarity Index
**Category**: Per-frame reference comparison
**Status**: ✅ Implemented in `fastvideo/tests/ssim/`
**Trust**: High — used in CI regression tests
#### What It Measures
SSIM compares two images (or video frames) based on three components:
1. **Luminance**: brightness similarity
2. **Contrast**: dynamic range similarity
3. **Structure**: spatial pattern similarity
The final score is a value in [0, 1] where 1.0 = identical.
#### Why It Matters
- Used as a **regression guard** in CI: ensures model updates don't degrade
visual output below a threshold.
- More perceptually meaningful than raw pixel MSE.
- Fast to compute — suitable for automated testing.
#### Limitations
- Requires a **pixel-aligned reference** video. Cannot compare videos with
different seeds, prompts, or angles.
- Operates **per-frame** — does not capture temporal coherence.
- Insensitive to some perceptual artifacts (color shifts, high-frequency noise).
#### How to Use
```bash
pytest fastvideo/tests/ssim/ -vs
```
#### Interpretation
| SSIM Range | Quality |
|------------|---------|
| > 0.90 | Excellent — very close to reference |
| 0.80–0.90 | Good — acceptable for most uses |
| 0.70–0.80 | Fair — noticeable differences |
| < 0.70 | Poor — significant divergence |
---
### LPIPS — Learned Perceptual Image Patch Similarity
**Category**: Per-frame perceptual distance
**Status**: ✅ Implemented in `scripts/lora_extraction/lora_inference_comparison.py`
**Trust**: Medium — available but only used for LoRA comparison currently
#### What It Measures
LPIPS uses a pretrained neural network (AlexNet by default) to extract
deep features from two images and computes the distance between them in
feature space. Unlike SSIM, LPIPS correlates much more strongly with
**human perceptual judgments**.
Lower LPIPS = more perceptually similar.
#### Why It Matters
- Best available automated proxy for **human visual judgments** at the frame
level.
- Captures semantic and structural differences that SSIM misses (e.g., texture
changes, minor recoloring).
- Used for validating LoRA merge quality.
#### Limitations
- Per-frame metric — no temporal awareness.
- Requires reference video (paired comparison only).
- Slightly slower than SSIM due to neural network forward pass.
#### How to Use
```bash
python scripts/lora_extraction/lora_inference_comparison.py \
--base merged_model \
--ft path/to/finetuned \
--adapter NONE \
--output-dir results \
--prompt "A cat" \
--compute-lpips
```
#### Interpretation
| LPIPS Range | Quality |
|-------------|---------|
| < 0.10 | Excellent — nearly indistinguishable |
| 0.10–0.20 | Good — minor perceptual differences |
| 0.20–0.40 | Fair — noticeable differences |
| > 0.40 | Poor — clearly different |
---
### Loss Trajectory
**Category**: Training signal proxy
**Status**: ✅ Active (from W&B `train_loss`)
**Trust**: Medium — proxy, not direct quality measure
#### What It Measures
Tracks the training loss over time. A healthy training run shows:
- **Decreasing loss** over the first hundreds of steps.
- **Stable gradient norms** (no wild spikes).
- **Consistent step times** (no infrastructure issues).
#### Why It Matters
- Cheapest evaluation signal — available in real-time from W&B.
- Critical for the **30-minute quality check** workflow.
- At later training stages (when loss becomes meaningful), trajectory shape
can predict final model quality.
#### Context: How This Evolves
The team's experience shows evaluation signals change during a project:
- **Early stage**: Loss may be flat or meaningless → focus on SSIM & visual
inspection instead.
- **Mid stage**: Loss starts decreasing → trajectory shape becomes useful.
- **Late stage**: Loss is meaningful → can compare trajectories across runs.
This dynamic is a key insight from the team's workflow: don't over-rely on
loss early; don't ignore it late.
---
### Grad Norm Stability
**Category**: Training health diagnostic
**Status**: ✅ Active (from W&B `grad_norm`)
**Trust**: Medium — diagnostic, not quality metric
#### What It Measures
The magnitude of gradients during training. Stable grad norms indicate
healthy optimization. Spikes or NaN values indicate training instability.
#### Alert Thresholds
| Condition | Meaning |
|-----------|---------|
| Stable ~0.3–0.5 | Normal training |
| Single spike > 3× average | Possible bad batch, monitor |
| NaN or Inf | 🔴 Training has diverged — stop run |
| Increasing trend | Learning rate may be too high |
---
## External Benchmarks
### GameWorld Score Benchmark (Matrix-Game)
**Category**: Multi-dimensional evaluation framework for interactive world models
**Status**: 🟡 External — not implemented in-repo
**Source**: [Matrix-Game 1.0 benchmark](https://github.com/SkyworkAI/Matrix-Game), used in [Matrix-Game 2.0 paper](https://arxiv.org/abs/2508.13009)
#### What It Measures
A comprehensive benchmark examining **four critical capabilities**:
| Dimension | What It Evaluates | Example Signals |
|-----------|-------------------|-----------------|
| **Visual quality** | Frame-level realism, absence of artifacts | Color fidelity, sharpness, coherence |
| **Temporal quality** | Smoothness across frames, motion consistency | Jitter, flickering, temporal aliasing |
| **Action controllability** | Response to input actions (keyboard/mouse) | Action delay, correctness, smoothness |
| **Physical rule understanding** | Adherence to physics (gravity, collision) | Object persistence, plausible motion |
#### Context from Matrix-Game 2.0
- Evaluation uses **597-frame composite action sequences** over 32 Minecraft
scenes and 16 wild scenes.
- Action controllability assessment is **Minecraft-specific** — cannot be
directly applied to wild/general scenes.
- The paper notes that models that "collapse" to static frames can
paradoxically score higher on consistency metrics — beware of this confound.
#### Relevance to FastVideo
- Matrix-Game 2.0 is built on SkyReels-V2/Wan2.1 architecture — **same model
family as FastVideo**.
- Their distillation uses DMD-based Self-Forcing — **same technique** as our
`self_forcing_distillation_pipeline.py`.
- GameWorld Score dimensions are a useful framework for thinking about world
model quality even outside gaming contexts.
---
## Human Preference Evaluation
**Category**: Gold-standard quality assessment
**Status**: 🔴 Manual process — no automated implementation
**Priority**: **Highest** — this is the most important evaluation signal
**Trust**: Highest — but expensive
### What It Measures
Human evaluators compare generated videos and rate them on dimensions like:
- Overall quality and realism
- Temporal coherence and smoothness
- Prompt adherence / action correctness
- Absence of artifacts
#### Why It's the Most Important Metric
All automated metrics are **proxies** for human judgment. They can be gamed
or may miss artifacts that humans easily notice. Human preference is the
ultimate ground truth for video generation quality.
#### Cost & Practicality
| Approach | Cost | Scale | When to Use |
|----------|------|-------|-------------|
| Internal team review | Low | ~10–50 videos | Every major checkpoint |
| Crowdsource (MTurk, Scale) | Medium | 100+ videos | Pre-release validation |
| A/B preference test | Medium | Pairs | Comparing two model versions |
#### Recommended Protocol
1. Sample 10–20 videos from the model at a checkpoint.
2. Include diverse prompts (easy + hard, short + long).
3. Have 2–3 evaluators score each video 1–5 on: quality, coherence, fidelity.
4. Record scores in the experiment journal.
---
## Metrics NOT Used
| Metric | Reason |
|--------|--------|
| ~~CLIP-Score~~ | Not used by the team. Measures text-image alignment using CLIP embeddings, but not well-suited for video temporal quality. |
| Inception Score (IS) | Less informative than FVD for video; primarily an image metric. |
| PSNR | Pixel-level metric; less perceptually meaningful than SSIM/LPIPS. |
---
## Adding a New Metric
Follow the SOP: `.agents/workflows/evaluation-development.md`
1. Prototype in `.agents/exploration/`
2. Validate on known-good and known-bad samples
3. Add to this registry
4. Update the `evaluate-video-quality` skill
@@ -1,21 +0,0 @@
# Experiment Journal
Living log of all experiments. Each entry captures what was tried, the result,
and any insights. Newest entries go at the top.
_No experiments logged yet. Use the `log-experiment` skill to add entries._
<!-- TEMPLATE — copy and fill for each new experiment:
## [YYYY-MM-DD] Experiment: <name>
- **Hypothesis**: <what you expected to learn>
- **Config**: model=..., lr=..., sp_size=..., gpus=..., script=...
- **W&B run**: <run_id or URL>
- **Duration**: <total wall time>
- **Key metrics**: loss=..., step_time=..., grad_norm=...
- **Checkpoint**: <path>
- **Insight**: <what was learned>
- **Status**: running | completed | failed | abandoned
- **Related lessons**: `.agents/lessons/<filename>.md`
-->
-5
View File
@@ -1,5 +0,0 @@
{"name": "codebase-map", "description": "High-level structural index of the FastVideo-WorldModel repository", "path": "codebase-map/README.md", "status": "ready", "trust": "high"}
{"name": "evaluation-registry", "description": "Catalog of all evaluation metrics with detailed explanations, implementation status, and usage guides", "path": "evaluation-registry/README.md", "status": "draft", "trust": "medium"}
{"name": "experiment-journal", "description": "Living log of all experiments with hypotheses, configs, metrics, and insights", "path": "experiment-journal/README.md", "status": "draft", "trust": "medium"}
{"name": "related-work", "description": "Index of related papers, repos, and blog posts with structured comparisons to FastVideo", "path": "related-work/README.md", "status": "draft", "trust": "low"}
{"name": "dreamverse-integration", "description": "Consolidated knowledge base for the FastVideo public API refactor (PRs 0-17), LTX-2 streaming server upstream, Dreamverse migration from FastVideo-internal, and NVFP4 quantization landing", "path": "dreamverse-integration/README.md", "status": "ready", "trust": "high"}
-34
View File
@@ -1,34 +0,0 @@
# Related Work Index
Each file in this directory is a structured summary of a related paper, repo,
or blog post relevant to FastVideo-WorldModel training.
## File Format
Each file is named `<slug>.md` and follows this structure:
```markdown
---
title: <paper/repo title>
source: <URL or citation>
type: paper | repo | blog
date_indexed: <ISO-8601>
tags: [world-model, distillation, evaluation, reward-shaping, ...]
---
## Summary
<1-2 paragraph summary of the work.>
## Key Differences from FastVideo
- <Bullet points comparing their approach to ours.>
## Actionable Insights
- <What we could adopt or adapt.>
```
## How to Add New Entries
Use the `index-related-work` skill, or manually create a file following the
template above.
_No related work indexed yet._
-76
View File
@@ -1,76 +0,0 @@
# Agent Onboarding — FastVideo-WorldModel
Welcome, agent. This is the **master onboarding** guide. Follow the steps below,
then check if a **domain-specific onboarding** exists for your task.
## Domain-Specific Onboarding
If your task falls into one of these areas, read the specialized guide **after**
completing the general steps below:
| Domain | Guide | When to Use |
|--------|-------|-------------|
| **WorldModel Training** | `worldmodel-training/README.md` | Training, finetuning, distillation, experiment management |
---
## Step 1: Understand the Codebase
Read these files to build your context:
| Priority | File | What you learn |
|----------|------|----------------|
| 1 | `AGENTS.md` | Coding guidelines, build/test commands, PR conventions |
| 2 | `docs/design/overview.md` | Architecture: models, pipelines, configs, registry |
| 3 | `fastvideo/train/` | Refactored training framework (YAML-driven, modular methods/models/callbacks) |
| 4 | `docs/training/overview.md` | Training data flow and preprocessing |
| 5 | `docs/training/finetune.md` | Training arguments, parallelism, LoRA, validation |
| 6 | `docs/contributing/coding_agents.md` | How to add model pipelines with agent assistance |
## Step 2: Discover Available Resources
Read these two index files to see what skills and memory modules exist:
- **`.agents/skills/index.jsonl`** — catalog of all agent skills (name + description)
- **`.agents/memory/index.jsonl`** — catalog of all memory modules (name + description)
Each entry has a `path` field pointing to the full content. Only load the
full README.md for modules relevant to your current task.
## Step 3: Check for Existing Skills & SOPs
Before writing new code or procedures:
1. **Skills**: Read `.agents/skills/index.jsonl` — find a matching skill by description.
2. **Workflows/SOPs**: Browse `.agents/workflows/` — step-by-step procedures for common tasks.
3. **Lessons**: Browse `.agents/lessons/` — known pitfalls and their fixes.
If a skill or SOP exists for your task, **use it**. If not, you are in **exploration mode** — see Step 4.
## Step 4: Exploration Mode
If no existing skill/SOP covers your task:
1. Document your progress in `.agents/exploration/<topic>.md` using the template in `.agents/exploration/README.md`.
2. At the end of your session, reflect:
- **What worked** → propose a new skill or SOP in the exploration log.
- **What failed** → create a lesson in `.agents/lessons/`.
3. Flag the exploration log for human review.
## Quick Reference
```
.agents/
├── ONBOARDING.md ← you are here
├── STATUS.md ← dashboard: completeness & trust of all components
├── skills/ ← reusable agent skills
├── workflows/ ← SOPs and procedures
├── memory/ ← persistent context (folder per topic + index.jsonl)
│ ├── index.jsonl
│ ├── codebase-map/
│ ├── experiment-journal/
│ ├── evaluation-registry/
│ └── related-work/
├── lessons/ ← mistakes and fixes
└── exploration/ ← draft procedures
```
@@ -1,305 +0,0 @@
# WorldModel Training — Agent Onboarding
Specialized onboarding for agents working on FastVideo-WorldModel training,
distillation, and evaluation. Read the master onboarding (`.agents/onboarding/README.md`)
first, then come here.
---
## Domain Context
FastVideo-WorldModel trains **interactive world models** — video generation systems
that respond to user actions (keyboard/mouse) in real-time. The architecture is
based on **Wan2.1** (SkyReels-V2) DiT models with causal attention for
auto-regressive streaming generation.
**Key techniques you will work with:**
- Full finetuning and LoRA on Wan / LTX-2 / Matrix-Game 2.0 models
- DMD-based distillation (few-step generation)
- Self-Forcing distillation (causal streaming)
- Diffusion-Forcing SFT (DFSFT) for causal models
- VSA (Variable Sparsity Acceleration) for efficient training
---
## Training Code: Two Generations
### New modular framework: `fastvideo/train/` (preferred)
The refactored training code uses a **YAML-only config-driven** architecture
with composable methods, per-role models, and a callback system. All new
training work should use this framework.
### Legacy pipelines: `fastvideo/training/` (deprecated)
The old monolithic pipeline classes (`WanTrainingPipeline`,
`DistillationPipeline`, etc.) still exist but are being phased out. The new
framework imports select utilities from `fastvideo/training/` for backward
compatibility (EMA, gradient clipping, checkpoint wrappers).
---
## Essential Reading (Training-Specific)
Read these **in order** before touching any training code:
| # | File | What You Learn |
|---|------|----------------|
| 1 | `docs/training/overview.md` | Training data flow: raw video → text embeddings + video latents → training |
| 2 | `docs/training/finetune.md` | Training arguments, parallelism (SP/TP), LoRA, validation settings |
| 3 | `docs/training/data_preprocess.md` | How to preprocess datasets into the expected format |
| 4 | `docs/design/overview.md` | Architecture: models, pipelines, configs, registry |
---
## New Training Framework (`fastvideo/train/`)
### Architecture Overview
```
fastvideo/train/
├── __init__.py → exports Trainer
├── trainer.py → main training loop coordinator
├── entrypoint/
│ ├── train.py → YAML-only training entrypoint
│ └── dcp_to_diffusers.py → checkpoint conversion utility
├── methods/ → training algorithms (TrainingMethod ABC)
│ ├── base.py → TrainingMethod base class
│ ├── fine_tuning/
│ │ ├── finetune.py → FineTuneMethod (supervised finetuning)
│ │ └── dfsft.py → DiffusionForcingSFTMethod (causal)
│ ├── distribution_matching/
│ │ ├── dmd2.py → DMD2Method (distribution matching distill)
│ │ └── self_forcing.py → SelfForcingMethod (causal streaming)
│ ├── knowledge_distillation/ → (stub, not yet implemented)
│ └── consistency_model/ → (stub, not yet implemented)
├── models/ → per-role model instances
│ ├── base.py → ModelBase & CausalModelBase (ABC)
│ └── wan/
│ ├── wan.py → WanModel (non-causal)
│ └── wan_causal.py → WanCausalModel (causal streaming)
├── callbacks/ → training hooks & monitoring
│ ├── callback.py → Callback base class + CallbackDict
│ ├── grad_clip.py → GradNormClipCallback
│ ├── ema.py → EMACallback (shadow weights)
│ └── validation.py → ValidationCallback (sampling + eval)
└── utils/ → configuration, building, checkpointing
├── builder.py → build_from_config() (config → runtime)
├── checkpoint.py → CheckpointManager (DCP-based)
├── config.py → load_run_config() (YAML → RunConfig)
├── training_config.py → TypedConfig dataclasses
├── optimizer.py → build_optimizer_and_scheduler()
├── instantiate.py → resolve_target() + instantiate()
├── tracking.py → build_tracker() (W&B, etc.)
├── dataloader.py → dataloader utilities
├── module_state.py → apply_trainable()
└── moduleloader.py → load_module_from_path()
```
### Key Concepts
**TrainingMethod** (`methods/base.py`): Abstract base class for all training
algorithms. Owns role models (student, teacher, critic), manages checkpoint
state, and defines the training step interface.
**ModelBase** (`models/base.py`): Per-role model wrapper. Each role (student,
teacher, critic) gets its own `ModelBase` instance owning a `transformer` and
`noise_scheduler`. `CausalModelBase` extends this for streaming models.
**Callback system** (`callbacks/`): Composable hooks for gradient clipping,
EMA, validation, etc. Configured via YAML, dispatched by `CallbackDict`.
**Config system** (`utils/config.py`, `utils/training_config.py`): YAML files
are parsed into typed `RunConfig` dataclass trees. Models and methods use
`_target_` fields for instantiation (similar to Hydra).
### Training Flow
```
run_training_from_config(config_path)
→ load_run_config() # YAML → RunConfig
→ init_distributed() # TP/SP setup
→ build_from_config() # instantiate models, method, dataloader
→ Trainer.run() # main loop:
├─ callbacks.on_train_start()
├─ checkpoint_manager.maybe_resume()
├─ for step in range(max_steps):
│ ├─ method.single_train_step(batch)
│ ├─ method.backward()
│ ├─ callbacks.on_before_optimizer_step()
│ ├─ method.optimizers_schedulers_step()
│ ├─ tracker.log(metrics, step)
│ ├─ callbacks.on_training_step_end()
│ └─ checkpoint_manager.maybe_save(step)
├─ callbacks.on_train_end()
└─ checkpoint_manager.save_final()
```
### Training Methods
| Method | Class | Use Case |
|--------|-------|----------|
| **FineTune** | `FineTuneMethod` | Single-role supervised finetuning |
| **DFSFT** | `DiffusionForcingSFTMethod` | Diffusion-forcing SFT with inhomogeneous timesteps |
| **DMD2** | `DMD2Method` | Multi-role distribution matching distillation (student + teacher + critic) |
| **Self-Forcing** | `SelfForcingMethod` | Extends DMD2 for causal student rollouts |
### Launching Training (New Framework)
Training is launched via `torchrun` with a single YAML config:
```bash
torchrun --nproc_per_node <N_GPUS> \
-m fastvideo.train.entrypoint.train \
--config examples/train/<config>.yaml
```
### Example YAML Configs
| Config | Method | Description |
|--------|--------|-------------|
| `examples/train/finetune_wan2.1_t2v_1.3B_vsa_phase3.4_0.9sparsity.yaml` | FineTune | Wan 1.3B finetuning with VSA sparsity |
| `examples/train/distill_wan2.1_t2v_1.3B_dmd2.yaml` | DMD2 | Wan 1.3B distillation (student + teacher + critic) |
| `examples/train/dfsft_wan_causal_t2v_1.3B.yaml` | DFSFT | Causal Wan 1.3B diffusion-forcing SFT |
| `examples/train/self_forcing_wan_causal_t2v_1.3B.yaml` | Self-Forcing | Causal streaming distillation |
### Checkpointing (New Framework)
**CheckpointManager** (`utils/checkpoint.py`) saves via `torch.distributed.checkpoint`:
```
output_dir/
└─ checkpoint-{step}/
├─ dcp/ # DCP state dict
├─ config.json # resolved training config
└─ .fastvideo_metadata.json
```
Checkpoint state includes: role model weights, per-role optimizers/schedulers,
CUDA RNG state, and callback state (e.g., EMA shadow weights).
### Config Structure
A YAML config defines the full training pipeline:
```yaml
models:
student:
_target_: fastvideo.train.models.wan.WanModel
model_path: ...
trainable: true
teacher: # optional, for distillation
_target_: fastvideo.train.models.wan.WanModel
model_path: ...
trainable: false
method:
_target_: fastvideo.train.methods.fine_tuning.FineTuneMethod
# method-specific params...
training:
distributed: { num_gpus: 8, tp_size: 1, sp_size: 8 }
data: { data_path: ..., batch_size: 1 }
optimizer: { lr: 1e-5, lr_scheduler: constant_with_warmup }
loop: { max_train_steps: 1000 }
checkpoint: { output_dir: ./outputs }
tracker: { trackers: [wandb], project_name: ... }
callbacks:
grad_clip:
_target_: fastvideo.train.callbacks.GradNormClipCallback
max_grad_norm: 1.0
validation:
_target_: fastvideo.train.callbacks.ValidationCallback
validation_steps: 100
```
---
## Legacy Training Pipelines (`fastvideo/training/`)
> **Note:** Use the new `fastvideo/train/` framework for new work. This section
> is retained for reference on existing pipelines not yet migrated.
| Pipeline | Entrypoint | Use Case |
|----------|-----------|----------|
| Wan T2V finetune | `fastvideo/training/wan_training_pipeline.py` | Standard text-to-video finetune / LoRA |
| Wan I2V finetune | `fastvideo/training/wan_i2v_training_pipeline.py` | Image-to-video (first frame conditioned) |
| Matrix-Game 2.0 finetune | `fastvideo/training/matrixgame2_training_pipeline.py` | Action-conditioned world model |
| Matrix-Game 2.0 AR diffusion | `fastvideo/training/matrixgame2_ar_diffusion_pipeline.py` | AR diffusion-forcing training |
| Matrix-Game 2.0 ODE-init | `fastvideo/training/matrixgame2_ode_causal_pipeline.py` | ODE-trajectory init |
| Matrix-Game 2.0 self-forcing distill | `fastvideo/training/matrixgame2_self_forcing_distillation_pipeline.py` | Self-forcing distillation |
| LTX-2 finetune | `fastvideo/training/ltx2_training_pipeline.py` | LTX-2 architecture finetuning |
| Wan DMD distillation | `fastvideo/training/wan_distillation_pipeline.py` | Few-step distillation via DMD |
| Self-Forcing distill | `fastvideo/training/wan_self_forcing_distillation_pipeline.py` | Causal streaming distillation |
---
## Key Infrastructure
### W&B Integration
- **Tracker**: `fastvideo/training/trackers.py` — `WandbTracker` class
- **New framework tracker**: `fastvideo/train/utils/tracking.py` — `build_tracker()`
- **Env vars**: `WANDB_API_KEY`, `WANDB_BASE_URL`, `WANDB_MODE`
### Parallelism
- **SP** (Sequence Parallel): splits video frames across GPUs — `sp_size: N`
- **TP** (Tensor Parallel): splits model layers across GPUs — `tp_size: N`
- Typical configs: SP=2–8, TP=1–2
---
## Evaluation (for training runs)
Read `.agents/memory/evaluation-registry/README.md` for the full metric catalog.
**Quick summary for training agents:**
| Metric | When to Use | Trust |
|--------|-------------|-------|
| **Loss trajectory** | Every run, real-time from W&B | Medium |
| **SSIM** | When comparing against reference outputs | High |
| **FVD** | For benchmarking model quality (`common.fvd` eval metric; example: `examples/inference/eval/eval_fvd.py`) | High |
| **LPIPS** | LoRA merge validation | Medium |
| **Human preference** | Major checkpoints | Highest |
---
## Common Workflows
| Task | Skill / SOP |
|------|-------------|
| Launch a training run | `.agents/skills/launch-experiment/SKILL.md` |
| Monitor a running experiment | `.agents/skills/monitor-experiment/SKILL.md` |
| Summarize final results | `.agents/skills/summarize-run/SKILL.md` |
| Full experiment lifecycle | `.agents/workflows/experiment-lifecycle.md` |
| Capture lessons from failures | `.agents/workflows/lesson-capture.md` |
---
## World Model–Specific Concepts
### Action Injection (Matrix-Game 2.0)
The Matrix-Game 2.0 pipeline adds **action modules** to each DiT block, enabling
frame-level mouse/keyboard input conditioning. The action sequence is injected
per-frame alongside the latent video tokens.
### Causal Architecture
For streaming generation, the model uses **causal attention** (each frame only
attends to previous frames). This enables auto-regressive chunk-by-chunk
generation — critical for real-time interactive world models.
### Self-Forcing Distillation
A **data-free** distillation method where the student model is trained to
generate coherent video sequences by being forced to use its own previous
outputs (rather than ground-truth) as context. This produces models robust to
their own error accumulation during long auto-regressive generation.
### DMD Distillation (Distribution Matching Distillation)
Reduces inference steps from ~50 to 3–4 by training a student model to match
the output distribution of the teacher model. Uses a critic network to estimate
distribution divergence.
### Diffusion-Forcing SFT (DFSFT)
Supervised finetuning with **inhomogeneous timesteps** across chunks — each
chunk in a causal sequence can have a different noise level, training the model
to handle mixed-fidelity contexts.
+3 -5
View File
@@ -50,8 +50,6 @@ Each skill lives in its own directory under `.agents/skills/`:
└── assets/ # Optional: templates, resources
```
After creating a new skill, add an entry to `.agents/skills/index.jsonl`:
```json
{"name": "<skill-name>", "description": "<description>", "path": "<skill-name>/SKILL.md", "status": "draft", "trust": "low"}
```
Skill discovery is directory-based; no hand-maintained registry entry is
required. Run `.agents/scripts/sync-skills.sh` if a local Claude Code checkout
needs refreshed `.claude/skills/` symlinks.
-26
View File
@@ -1,26 +0,0 @@
# add-model skill review backlog (historical)
Updated 2026-04-30 after the phase-based `/add-model` rewrite.
All skill-text items from the prior review were incorporated into the current
split skill stack under `.agents/skills/add-model*`. This file is kept
only for codebase-owner follow-ups that are not blockers for the skill workflow.
### 1. Audit `wan_to_diffusers.py` usage
`SKILL.md` now treats `scripts/checkpoint_conversion/wan_to_diffusers.py` as a
legacy regex-reference file, not a conversion-script template.
Open codebase question: is this module still imported by live code? If yes,
document the caller near the script or in developer docs. If no, delete it in a
separate cleanup PR.
### 2. Decide Whether To Add Audio Workload Enums
The current pipeline skill documents the repository's compatibility workaround:
until `WorkloadType` grows audio values, audio-only pipelines may register as
`T2V` with explicit rationale and minimal video-shaped placeholders when shared
`VideoGenerator` paths require them.
Open codebase question: should `WorkloadType` be extended now with audio and
joint AV variants, or should the first audio pipeline PR own that enum change?
+1 -3
View File
@@ -422,8 +422,6 @@ matching `*secret*`.
script shape.
- `scripts/checkpoint_conversion/wan_to_diffusers.py` for legacy regex mapping
reference only.
- `REVIEW.md` is historical; its decisions are incorporated here as of
2026-04-30.
## Changelog
@@ -431,7 +429,7 @@ matching `*secret*`.
|---|---|
| 2026-04-24 | Initial FastVideo add-model workflow. |
| 2026-04-30 | Split external setup into `add-model-01-prep`. |
| 2026-04-30 | Rewrote as manual `/add-model` phase workflow and incorporated `REVIEW.md` decisions. |
| 2026-04-30 | Rewrote as manual `/add-model` phase workflow and incorporated prior review decisions. |
| 2026-04-30 | Extracted early parity scaffolding into `add-model-02-parity` and moved it before conversion/component implementation. |
| 2026-04-30 | Added component reuse proof gate, bucket-specific porting skills, and parity PASS requirement for reused components. |
| 2026-04-30 | Split prototype, conversion, and parity-debug phases; added conversion skill for monolithic and separate checkpoint layouts. |
-32
View File
@@ -1,32 +0,0 @@
---
name: add-reward-model
description: Use when adding reusable reward models under fastvideo/train/methods/rl/rewards for RLHF or online RL training.
---
# Add Reward Model
Use for reward models consumed by RL methods.
## Placement
- Put reusable reward code under `fastvideo/train/methods/rl/rewards/`.
- Expose public builders from `fastvideo/train/methods/rl/rewards/__init__.py`.
- Keep method-specific aggregation or advantage logic out of reward classes.
## Media Inputs
- Reward callables receive decoded media tensors.
- Accept single-frame tensors as `[B, C, H, W]` and multi-frame tensors as `[B, C, T, H, W]` when practical.
- Frame selection is reward-specific. Frame scorers such as PickScore and CLIPScore should explicitly select frame `0`; temporal rewards should inspect whichever frames they need.
- Return one scalar reward per prompt/sample.
## Attribution
- If code is ported or closely adapted from another repo, add a short comment or docstring naming the source file/function.
- Preserve SPDX headers used by FastVideo files.
## Tests
- Unit-test tensor layout handling without loading large reward checkpoints.
- Allow fake scorer injection for multi-reward tests.
- Test weighted reward aggregation and metric keys.
-38
View File
@@ -1,38 +0,0 @@
---
name: add-rl-method
description: Use when adding or modifying an RL/RLHF method under fastvideo/train/methods/rl, including DiffusionNFT-like methods.
---
# Add RL Method
Use for new RL methods in the modular `fastvideo/train` stack.
## Required Shape
- Add the method under `fastvideo/train/methods/rl/`.
- Subclass `TrainingMethod`.
- Keep model-family logic in `ModelBase` wrappers.
- Decode generated latents through `ModelBase.decode_latents`; add that hook to the new model wrapper instead of decoding inside the RL method.
- Use `fastvideo/train/methods/rl/common/sampling.py` for generation unless the method has a documented reason to avoid sampling.
- Use `fastvideo/train/methods/rl/common/prompt_sampling.py` for reusable grouped prompt sampling patterns such as DiffusionNFT K-repeat.
- Use `fastvideo/train/methods/rl/rewards/` for reward models.
## Optimization
- Return `manages_optimization() == True` only when the method must own a nonstandard outer/inner loop.
- If using managed optimization, implement `managed_train_step(data_stream, iteration)`.
- Existing trainer callbacks, checkpointing, tracking, and validation should still work.
## Config
- Put method knobs under `method`.
- Put sampler knobs under `method.sampling`.
- Do not put scheduler or trajectory policy into model configs.
- Do not split a diffusers-style scheduler from its built-in `step()` solver in YAML; use `trajectory` only for higher-level ODE vs re-noise behavior.
- Avoid fixed timestep lists in examples unless reproducing a known baseline; prefer scheduler-generated defaults.
## Tests
- Add fake-model tests for sampler/method behavior.
- Add config parse tests for the public YAML.
- Confirm existing train methods stay on the default Trainer path.
+16 -11
View File
@@ -44,7 +44,7 @@ pipeline PR.
|-----------|----------|-------------|
| PR number or URL | Yes | E.g. `1280` or `https://github.com/hao-ai-lab/FastVideo/pull/1280` |
| Max desired PR size | No | Defaults to ~2,500 LOC of code per stack PR (excluding generated/journal files) |
| Output dir | No | Defaults to `.agents/exploration/decompose-<pr-number>.md` |
| Output directory | No | Defaults to `.agents/tmp/decompose-<pr-number>/` (gitignored) |
## Steps
@@ -54,8 +54,10 @@ pipeline PR.
Always cross-check against the authoritative `git diff`:
```bash
mkdir -p .agents/tmp/decompose-<N>
git fetch origin pull/<N>/head:<feature-branch>
git diff origin/main..origin/<feature-branch> --name-status > /tmp/pr-<N>-files.txt
git diff origin/main..origin/<feature-branch> --name-status \
> .agents/tmp/decompose-<N>/files.txt
git diff origin/main..origin/<feature-branch> --stat
```
@@ -68,7 +70,7 @@ Classify every changed file into one of four tiers:
| Tier | Description | Examples |
|---|---|---|
| **Tier 0 — Invisible** | Lint/style/CI configs that don't affect runtime | `.gitignore`, `pyproject.toml` (codespell only), `.agents/skills/index.jsonl` stubs |
| **Tier 0 — Invisible** | Lint/style/CI configs that don't affect runtime | `.gitignore`, `pyproject.toml` (codespell only), agent documentation |
| **Tier 1 — Dead code** | New files in their own dirs; aggregator one-liners | `fastvideo/models/dits/<new>/`, `fastvideo/pipelines/basic/<new>/`, `examples/inference/basic/basic_<new>*.py`, `tests/local_tests/<new>/`, `__init__.py` exports |
| **Tier 2 — Cross-cutting infra** | Modifications to files used by every pipeline | See protected-paths list below |
| **Tier 3 — Activation switch** | `register_configs(...)` calls + the example scripts that demo them | `fastvideo/registry.py` |
@@ -235,9 +237,11 @@ git -C "$REPO" fetch origin main:main
git -C "$REPO" fetch "origin/$SOURCE_BRANCH"
git -C "$REPO" worktree add "$WORKTREE" origin/main
# Capture baseline for provenance
mkdir -p "$REPO/.agents/exploration"
cat > "$REPO/.agents/exploration/<feature>-baseline-${SOURCE_SHA:0:8}.txt" <<EOF
# Capture baseline for provenance. Everything under .agents/tmp is transient
# and ignored by git.
OUTPUT_DIR="$REPO/.agents/tmp/decompose-$SOURCE_PR"
mkdir -p "$OUTPUT_DIR"
cat > "$OUTPUT_DIR/<feature>-baseline-${SOURCE_SHA:0:8}.txt" <<EOF
Source PR: <repo>#$SOURCE_PR
Source SHA: $SOURCE_SHA
Authoritative file count: $(git -C "$REPO" diff origin/main..origin/$SOURCE_BRANCH --name-only | wc -l)
@@ -272,12 +276,13 @@ Notes:
## Outputs
The skill produces:
The skill produces all transient planning artifacts under
`.agents/tmp/decompose-<pr>/`:
1. A markdown decomposition plan (`.agents/exploration/decompose-<pr>.md`)
1. A markdown decomposition plan (`plan.md`)
2. A proposed branch graph
3. A worktree-bootstrap script
4. Per-PR file allocation lists (under `/tmp/<feature>-stack/`)
3. A worktree-bootstrap script (`bootstrap.sh`)
4. Per-PR file allocation lists (under `stack/`)
5. AGENTS.md scaffolds for any new pipeline packages
6. Draft lesson files (placed alongside the PR that owns the code they concern)
7. A finalized provenance table for the package AGENTS.md
@@ -312,7 +317,7 @@ The skill should warn against:
```
User: split PR 1280
Agent: [invokes decompose-pipeline-pr]
→ produces .agents/exploration/decompose-1280.md with:
→ produces .agents/tmp/decompose-1280/plan.md with:
- tiered file table (56 files: 3 tier-0, 35 tier-1, 9 tier-2,
9 tier-3)
- branch graph (PR-A + PR-B + 8-PR stack)
+50 -44
View File
@@ -8,29 +8,28 @@ description: Use when redeploying the migrated Dreamverse app backend and fronte
**Scope:** project (lives in this repo at `.agents/skills/dreamverse-deploy/`)
**When to use:** you want to (re)launch the migrated `apps/dreamverse/` backend
+ frontend on this dev node, pinned to a specific physical GPU. Tears down
and frontend on this dev node, pinned to a specific physical GPU. Tears down
any existing deploy on the same ports first, then boots fresh and waits for
both `/readyz` and the FE root to return 200.
**Pairs with:** [`integration-plan.md`](../../memory/dreamverse-integration/integration-plan.md)
"Local GPU4 verification hook" + [`decisions-log.md D-19`](../../memory/dreamverse-integration/decisions-log.md#d-19).
## Prerequisites
- Working tree on a branch that has `apps/dreamverse/` (e.g. `will/dreamverse-monorepo`)
- Working tree containing `apps/dreamverse/`
- `dreamverse-server` installed from this checkout; if missing, run
`uv pip install -e ".[dreamverse]"`
- Local conda env at `~/miniconda3/envs/fv-main/` with `flashinfer-python`,
`cerebras-cloud-sdk`, `openai` installed (override the default path with
`DREAMVERSE_PYTHON=/path/to/python`)
- `~/.env` exporting `CEREBRAS_API_KEY`, `GROQ_API_KEY`, etc.
- npm available in `$PATH` (or set `NPM=/path/to/npm`)
- `gcc-13` + `g++-13` at `/usr/bin/` (workaround for nvcc gcc-15 rejection)
- **Recommended:** native ffmpeg env file at `apps/dreamverse/scripts/ffmpeg-env.sh`
(built once via `bash apps/dreamverse/scripts/install_native_ffmpeg.sh`).
When present, the deploy sources it inside the backend setsid block so the
worker spawns ffmpeg from `$HOME/opt/ffmpeg-native/bin/ffmpeg` (LTO + libx264
+ native arch) instead of the system `/usr/bin/ffmpeg`. When missing, the
deploy falls back to system ffmpeg with a warning. Set
`DREAMVERSE_REQUIRE_NATIVE_FFMPEG=true` to make the missing env file a hard
- **Recommended:** native ffmpeg at `$HOME/opt/ffmpeg-native/bin/ffmpeg`, built
via `bash apps/dreamverse/scripts/install_native_ffmpeg.sh`. The deploy
detects that binary directly and exports it for the backend. The installer's
generated `apps/dreamverse/scripts/ffmpeg-env.sh` is for manual launches.
When the binary is missing, the deploy falls back to system ffmpeg with a
warning. Set
`DREAMVERSE_REQUIRE_NATIVE_FFMPEG=true` to make the missing binary a hard
failure.
If any required prereq is missing, the script fails fast with a clear message.
@@ -38,23 +37,22 @@ If any required prereq is missing, the script fails fast with a clear message.
## Usage
```bash
# Deploy on GPU 4 with default ports (backend 8009, FE 5274) — torch.compile
# and warmup are both OFF by default so first-segment cold start is ~45s
# instead of ~3-4min.
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh 4
# Deploy on GPU 4 with the current web port. The legacy helper default remains
# 5274, so pass 5299 explicitly. Torch compile and warmup are both off.
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh 4 8009 5299
# Deploy on GPU 6 with custom ports
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh 6 8089 5275
# Deploy on GPU 0 with warmup enabled
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh --warmup 0
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh --warmup 0 8009 5299
# Deploy with torch.compile enabled (max-autotune; first segment ~3-4min,
# subsequent segments save ~3s — only worth it for benchmarking)
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh --torch-compile 4
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh --torch-compile 4 8009 5299
# Deploy with both warmup AND torch.compile enabled
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh --warmup --torch-compile 4
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh --warmup --torch-compile 4 8009 5299
# Flags can appear before, between, or after positional args
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh 4 8089 5275 --warmup
@@ -86,9 +84,9 @@ Flags can appear in any position relative to the positional args. Explicit flag
| `DREAMVERSE_WARMUP` | `false` | Same as `--warmup`/`--no-warmup`. Flag takes precedence |
| `DREAMVERSE_TORCH_COMPILE` | `false` | Same as `--torch-compile`/`--no-torch-compile`. Flag takes precedence |
| `DREAMVERSE_NVENC` | `false` | Same as `--nvenc`/`--no-nvenc`. Flag takes precedence |
| `DREAMVERSE_PYTHON` | `~/miniconda3/envs/fv-main/bin/python` | Conda env python used for prereq probes (flashinfer import). The wrapper at `apps/dreamverse/scripts/dreamverse-server` still resolves python via the `.venv` symlink, which points at the same interpreter on this dev node |
| `DREAMVERSE_PYTHON` | `~/miniconda3/envs/fv-main/bin/python` | Conda environment used for the flashinfer prerequisite probe; `dreamverse-server` itself is resolved from `PATH` |
| `DREAMVERSE_REPO_ROOT` | git rev-parse | Repo root override |
| `DREAMVERSE_LOG_DIR` | `/tmp/opencode/dreamverse-deploy` | Where to write `backend.log` / `frontend.log` |
| `DREAMVERSE_LOG_DIR` | `/tmp/opencode/dreamverse-deploy` | Directory for the per-GPU backend and per-port frontend logs |
| `DREAMVERSE_REQUIRE_NATIVE_FFMPEG` | `false` | If `true`, fail when `$HOME/opt/ffmpeg-native/bin/ffmpeg` is absent |
## What it does
@@ -106,12 +104,14 @@ Flags can appear in any position relative to the positional args. Explicit flag
- `CC=/usr/bin/gcc-13 CXX=/usr/bin/g++-13 CUDAHOSTCXX=/usr/bin/g++-13`
- `NVCC_PREPEND_FLAGS="-ccbin /usr/bin/gcc-13 -allow-unsupported-compiler"`
- `FASTVIDEO_FFMPEG_BIN=$HOME/opt/ffmpeg-native/bin/ffmpeg` +
`FASTVIDEO_VIDEO_CODEC=libx264` (when the native binary exists)
5. Launches the backend via `apps/dreamverse/scripts/dreamverse-server` in a
detached `setsid` session, captures PID.
6. Polls `/readyz` until 200 (max 5 min).
7. Launches the frontend via `npm run dev:devtools` in a detached session,
captures PID.
`FASTVIDEO_VIDEO_CODEC=<libx264|h264_nvenc>` (when the native binary exists)
5. Launches the installed `dreamverse-server` console command in a detached
`setsid` session and captures its PID.
6. Polls `/readyz` until 200. The budget is 5 minutes by default, 8 minutes
with one startup optimization enabled, and 15 minutes with both warmup and
`torch.compile` enabled.
7. Launches the devtools frontend through npm in a detached session and
captures its PID.
8. Polls FE `/` until 200 (max 60s).
9. Prints URLs, PIDs, and log paths.
@@ -122,13 +122,13 @@ Flags can appear in any position relative to the positional args. Explicit flag
- Does not run Playwright. Use the e2e wrapper separately:
```bash
cd apps/dreamverse/web
PLAYWRIGHT_SKIP_WEBSERVER=1 BACKEND_URL=http://127.0.0.1:8009 \
PLAYWRIGHT_BASE_URL=http://127.0.0.1:5274 \
PLAYWRIGHT_SKIP_WEBSERVER=1 BACKEND_HOST=127.0.0.1 BACKEND_PORT=8009 \
PLAYWRIGHT_BASE_URL=http://127.0.0.1:5299 \
NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
npm exec -- playwright test
```
The fast suite (8 specs, ~5s) runs by default; the long-running
two-segment audio-continuation spec is gated behind
The standard suite runs by default; the long-running two-segment
audio-continuation spec is gated behind
`PLAYWRIGHT_LONG_RUNNING=1` (see below).
## Long-running e2e (paired with `--warmup --torch-compile`)
@@ -136,19 +136,20 @@ Flags can appear in any position relative to the positional args. Explicit flag
[`apps/dreamverse/web/e2e/long-running-segments.spec.ts`](../../../apps/dreamverse/web/e2e/long-running-segments.spec.ts)
drives a real two-segment session through the FE, captures every WS
frame, and asserts segments 1 AND 2 both reach `media_segment_complete`
with at least one binary fMP4 chunk per segment — the canonical
regression guard against the D-20 BrokenPipe pattern documented in
[`decisions-log.md D-20`](../../memory/dreamverse-integration/decisions-log.md#d-20).
with at least one binary fMP4 chunk per segment. It guards against the
BrokenPipe regression previously caused by dropped LTX-2 audio continuation
kwargs.
Skipped by default. Enable with:
```bash
./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh \
--warmup --torch-compile 4
--warmup --torch-compile 4 8009 5299
cd apps/dreamverse/web
PLAYWRIGHT_SKIP_WEBSERVER=1 \
BACKEND_URL=http://127.0.0.1:8009 \
PLAYWRIGHT_BASE_URL=http://127.0.0.1:5274 \
BACKEND_HOST=127.0.0.1 \
BACKEND_PORT=8009 \
PLAYWRIGHT_BASE_URL=http://127.0.0.1:5299 \
NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
PLAYWRIGHT_LONG_RUNNING=1 \
npm exec -- playwright test e2e/long-running-segments.spec.ts
@@ -180,11 +181,16 @@ children survive, GPU stays full, next deploy OOMs.
## Notes
- The wrapper at `apps/dreamverse/scripts/dreamverse-server` is what makes
the migrated `apps/dreamverse/server/main.py` run instead of the legacy
conda-installed Dreamverse — see [decisions-log.md D-19](../../memory/dreamverse-integration/decisions-log.md#d-19) for why
this matters.
- The installed `dreamverse-server` console command enters
`apps/dreamverse/dreamverse/server_entry.py`, which loads the current
Dreamverse runtime from `apps/dreamverse/dreamverse/`.
- The B200 / sm_100a NVCC flags are mandatory on this dev node because the
conda toolchain ships gcc-15, which nvcc rejects. If you're on a machine
with a supported native gcc, those exports are still safe (no-op when the
paths don't exist; the script verifies them upfront).
conda toolchain ships gcc-15, which nvcc rejects. The script requires the
configured gcc-13 and g++-13 binaries during preflight.
## Deployment boundary
This skill is for a local checkout on a directly attached GPU. For a container
image, use `apps/dreamverse/docker/README.md`. For Modal, follow
`apps/dreamverse/scripts/modal/README.md`; do not adapt this process-killing
workflow to a remote deployment.
@@ -60,7 +60,7 @@ list_port_pids() {
}
if [[ "${1:-}" == "--stop" ]]; then
for pat in 'apps/dreamverse/server/main.py' 'main.py --host 0.0.0.0 --port' 'next dev --port' 'next-server (v'; do
for pat in 'apps/dreamverse/dreamverse/main.py' 'dreamverse-server --host 0.0.0.0 --port' 'next dev --port' 'next-server (v'; do
terminate_pattern "${pat}"
done
if [[ -n "${2:-}" ]] && [[ "${2}" =~ ^[0-9]+$ ]]; then
@@ -176,8 +176,9 @@ bail() { echo "error: $*" >&2; exit 3; }
[[ -d "${REPO_ROOT}/apps/dreamverse" ]] \
|| bail "REPO_ROOT '${REPO_ROOT}' does not contain apps/dreamverse/. Are you on a migration branch?"
[[ -x "${REPO_ROOT}/apps/dreamverse/scripts/dreamverse-server" ]] \
|| bail "wrapper script missing or not executable: apps/dreamverse/scripts/dreamverse-server"
DREAMVERSE_SERVER="$(command -v dreamverse-server 2>/dev/null || true)"
[[ -n "${DREAMVERSE_SERVER}" ]] && [[ -x "${DREAMVERSE_SERVER}" ]] \
|| bail "dreamverse-server not executable or not in PATH (run: uv pip install -e \".[dreamverse]\")"
CONDA_ENV_PYTHON="${DREAMVERSE_PYTHON:-${HOME}/miniconda3/envs/fv-main/bin/python}"
[[ -x "${CONDA_ENV_PYTHON}" ]] \
@@ -248,7 +249,7 @@ kill_port_pid() {
done
}
for pat in "main.py --host 0.0.0.0 --port ${BACKEND_PORT}" "next dev --port ${FRONTEND_PORT}" "NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 next dev --port ${FRONTEND_PORT}"; do
for pat in "dreamverse-server --host 0.0.0.0 --port ${BACKEND_PORT}" "next dev --port ${FRONTEND_PORT}" "NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 next dev --port ${FRONTEND_PORT}"; do
terminate_pattern "${pat}"
done
kill_port_pid "${BACKEND_PORT}"
@@ -310,14 +311,14 @@ setsid bash -c "
export CUDAHOSTCXX=${GPP13}
export NVCC_PREPEND_FLAGS=\"-ccbin ${GCC13} -allow-unsupported-compiler\"
cd \"${REPO_ROOT}\"
exec ./apps/dreamverse/scripts/dreamverse-server --host 0.0.0.0 --port ${BACKEND_PORT}
exec \"${DREAMVERSE_SERVER}\" --host 0.0.0.0 --port ${BACKEND_PORT}
" > "${backend_log}" 2>&1 < /dev/null &
disown
# Wait briefly, then resolve actual python PID (the inner process, not the
# wrapper bash).
sleep 4
backend_pid="$(pgrep -f "main.py --host 0.0.0.0 --port ${BACKEND_PORT}" | head -1 || true)"
backend_pid="$(pgrep -f "dreamverse-server --host 0.0.0.0 --port ${BACKEND_PORT}" | head -1 || true)"
if [[ -z "${backend_pid}" ]]; then
echo "error: backend failed to spawn. Last 30 lines of log:" >&2
@@ -1,128 +0,0 @@
---
name: evaluate-video-quality
description: Evaluate generated video quality using available metrics (SSIM, loss trajectory, caption consistency)
---
# Evaluate Video Quality
## Purpose
Assess the quality of videos generated by a training run. Combines multiple
signals to give a holistic quality assessment. This skill is **evolving** —
new metrics will be added as they are developed.
## Prerequisites
- Generated videos available locally or via W&B artifacts.
- For SSIM: reference videos from official implementations.
- For caption consistency: LLM access (optional, stub for now).
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| `video_paths` | Yes | List of paths to generated videos |
| `reference_paths` | No | Paths to reference videos (for SSIM) |
| `prompts` | No | Prompts used to generate videos (for caption check) |
| `loss_summary` | No | Path to W&B summary JSON (for loss trajectory) |
| `metrics` | No | Which metrics to run (default: all available) |
## Available Metrics
Check `.agents/memory/evaluation-registry/README.md` for the current catalog.
### SSIM (Active)
Leverages the existing infrastructure in `fastvideo/tests/ssim/`.
```bash
pytest fastvideo/tests/ssim/ -vs --video-path <generated> --reference-path <reference>
```
Or use the SSIM utility directly:
```python
from fastvideo.tests.ssim.ssim_utils import compute_ssim
score = compute_ssim(generated_video, reference_video)
# score > 0.85 is typically "acceptable"
```
**Interpretation**:
| SSIM Range | Quality |
|------------|---------|
| > 0.90 | Excellent — very close to reference |
| 0.80–0.90 | Good — acceptable for most uses |
| 0.70–0.80 | Fair — noticeable differences |
| < 0.70 | Poor — significant quality issues |
### Loss Trajectory (Active)
Analyze the loss curve shape from W&B summary:
```python
import json
with open(loss_summary_path) as f:
summary = json.load(f)
final_loss = summary["train_loss"]
runtime = summary["_runtime"]
steps = summary["_step"]
```
**Early-stage heuristics** (first 500 steps):
- Loss should be decreasing (even slightly).
- Grad norm should be stable (no wild oscillations).
- If loss is flat or increasing, flag for review.
### Caption Consistency (Draft — Not Yet Calibrated)
Use an LLM to evaluate whether the video content matches the input prompt.
```
Prompt: "A golden retriever playing in the snow"
Video: <path>
Score the video on:
1. Object presence (is there a golden retriever?)
2. Action accuracy (is it playing?)
3. Environment match (is there snow?)
4. Overall coherence (does it look natural?)
Each 1-5, total /20.
```
> ⚠️ This metric is in **draft** status. Results should not be treated as
> ground truth until calibrated against human judgments.
## Steps
1. **Identify available metrics** — Check `.agents/memory/evaluation-registry/README.md`.
2. **Run each metric** — Collect scores.
3. **Aggregate** — Produce a combined quality report.
4. **Log** — Update the experiment journal with quality results.
## Outputs
```markdown
## Video Quality Report: <experiment_name>
| Metric | Score | Threshold | Status |
|--------|-------|-----------|--------|
| SSIM (avg) | 0.87 | > 0.80 | ✅ Pass |
| Loss trajectory | decreasing | decreasing | ✅ Pass |
| Caption consistency | 16/20 | > 14/20 | ✅ Pass |
### Per-Video Scores
| Video | SSIM | Caption |
|-------|------|---------|
| video_001.mp4 | 0.89 | 17/20 |
| video_002.mp4 | 0.85 | 15/20 |
```
## References
- `fastvideo/tests/ssim/` — SSIM test infrastructure
- `fastvideo/tests/training/Vanilla/test_training_loss.py` — loss comparison
- `.agents/memory/evaluation-registry/README.md` — metric catalog
## Changelog
| Date | Change |
|------|--------|
| 2026-03-02 | Initial version with SSIM, loss trajectory, caption consistency stub |
-13
View File
@@ -1,13 +0,0 @@
{"name": "launch-experiment", "description": "Generate and execute a training launch command for FastVideo models", "path": "launch-experiment/SKILL.md", "status": "draft", "trust": "low"}
{"name": "monitor-experiment", "description": "Poll a running W&B training run for progress and emit structured alerts", "path": "monitor-experiment/SKILL.md", "status": "draft", "trust": "low"}
{"name": "summarize-run", "description": "Extract a W&B run summary into a structured experiment report", "path": "summarize-run/SKILL.md", "status": "draft", "trust": "low"}
{"name": "log-experiment", "description": "Append or update an experiment entry in the experiment journal", "path": "log-experiment/SKILL.md", "status": "draft", "trust": "low"}
{"name": "evaluate-video-quality", "description": "Evaluate generated video quality using available metrics (SSIM, loss trajectory, caption consistency)", "path": "evaluate-video-quality/SKILL.md", "status": "draft", "trust": "low"}
{"name": "seed-ssim-references", "description": "Run a new or updated fastvideo/tests/ssim/ test on Modal, pull generated videos, and upload them to FastVideo/ssim-reference-videos so the test has a regression baseline", "path": "seed-ssim-references/SKILL.md", "status": "draft", "trust": "low"}
{"name": "reseed-ssim-references", "description": "Re-seed (overwrite) HF reference videos for an existing fastvideo/tests/ssim/ test and a single model id on Modal L40S. Always backs up current refs first, regenerates on Modal, pauses for the user to eyeball before-vs-after, then uploads with --force scoped to --model-id. Sister skill to seed-ssim-references; use when intentional code change has invalidated existing refs", "path": "reseed-ssim-references/SKILL.md", "status": "draft", "trust": "low"}
{"name": "decompose-pipeline-pr", "description": "Decompose an oversized FastVideo pipeline PR into a stack of independently-reviewable PRs. Tiers the diff by blast radius (invisible / dead code / cross-cutting infra / activation), produces a branch graph and worktree bootstrap, drafts the AGENTS.md manifest, flags missing tests on cross-cutting infra changes, and extracts lessons from the PR body. Worked example: PR #1280 daVinci-MagiHuman (9.8k LOC) decomposed into 10 stacked PRs.", "path": "decompose-pipeline-pr/SKILL.md", "status": "tested", "trust": "medium"}
{"name": "reseed-performance-baseline", "description": "Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, or environment-caused benchmark shift. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline must be advanced by replicating one reviewed shifted source result into three success=true records, or five records when explicitly requested", "path": "reseed-performance-baseline/SKILL.md", "status": "draft", "trust": "low"}
{"name": "add-model", "description": "Add a new model (or variant) to FastVideo: DiT + configs + pipeline + presets + registry + tests. Walks through FastVideo's single stage-based pipeline architecture with exact file paths and registration hooks.", "path": "add-model/SKILL.md", "status": "draft", "trust": "low"}
{"name": "rlhf-training-abstractions", "description": "Use when changing FastVideo RLHF/RL training infrastructure, especially sampler, reward, scheduler trajectory, or method boundaries under fastvideo/train.", "path": "rlhf-training-abstractions/SKILL.md", "status": "draft", "trust": "low"}
{"name": "add-rl-method", "description": "Use when adding or modifying an RL/RLHF method under fastvideo/train/methods/rl, including DiffusionNFT-like methods.", "path": "add-rl-method/SKILL.md", "status": "draft", "trust": "low"}
{"name": "add-reward-model", "description": "Use when adding reusable reward models under fastvideo/train/methods/rl/rewards for RLHF or online RL training.", "path": "add-reward-model/SKILL.md", "status": "draft", "trust": "low"}
-127
View File
@@ -1,127 +0,0 @@
---
name: launch-experiment
description: Generate and execute a training launch command for FastVideo models
---
# Launch Experiment
## Purpose
Construct a fully-specified `torchrun` training command for a FastVideo model
given a target pipeline, dataset, and hyperparameter overrides. This skill
automates the boilerplate of setting environment variables, picking the right
entrypoint, and applying defaults from the closest example script.
## Prerequisites
- The repo is cloned and `fastvideo` is installed (`uv pip install -e ".[dev]"`).
- Dataset is preprocessed (see `docs/training/data_preprocess.md`).
- `WANDB_API_KEY` is set in the environment (or `WANDB_MODE=offline` for local).
- GPU resources are available (multi-GPU requires NCCL).
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| `pipeline` | Yes | Training pipeline type: `finetune`, `distill-dmd`, `self-forcing`, `lora`, `consistency` |
| `model` | Yes | Model family: `wan-t2v-1.3B`, `wan-i2v-14B`, `ltx2`, `matrixgame` |
| `data_path` | Yes | Path to preprocessed dataset (parquet) |
| `num_gpus` | Yes | Number of GPUs |
| `overrides` | No | Dict of hyperparameter overrides (any CLI arg) |
| `output_dir` | No | Output directory (default: `outputs/<model>_<pipeline>`) |
| `run_name` | No | W&B run name (default: auto-generated) |
## Steps
### 1. Identify the training entrypoint
| Pipeline | Entrypoint |
|----------|-----------|
| `finetune` (Wan T2V) | `fastvideo/training/wan_training_pipeline.py` |
| `finetune` (Wan I2V) | `fastvideo/training/wan_i2v_training_pipeline.py` |
| `finetune` (LTX-2) | `fastvideo/training/ltx2_training_pipeline.py` |
| `finetune` (Matrix-Game 2.0) | `fastvideo/training/matrixgame2_training_pipeline.py` |
| `distill-dmd` | `fastvideo/training/wan_distillation_pipeline.py` |
| `self-forcing` | `fastvideo/training/wan_self_forcing_distillation_pipeline.py` |
### 2. Resolve default hyperparameters
Find the closest example script in `examples/training/` for the model:
| Model | Example Script Directory |
|-------|-------------------------|
| `wan-t2v-1.3B` | `examples/training/finetune/wan_t2v_1.3B/crush_smol/` |
| `wan-i2v-14B` | `examples/training/finetune/wan_i2v_14B_480p/crush_smol/` |
| `ltx2` | `examples/training/finetune/ltx2/` |
| `matrixgame` | `examples/training/finetune/MatrixGame2.0/` |
| `distill-dmd` | `scripts/distill/v1_distill_dmd_wan.sh` |
Read the script to extract default values for:
- `--learning_rate`, `--train_batch_size`, `--sp_size`, `--tp_size`
- `--num_latent_t`, `--num_height`, `--num_width`, `--num_frames`
- `--gradient_accumulation_steps`, `--max_train_steps`
- `--mixed_precision`, `--weight_decay`, `--max_grad_norm`
- `--validation_steps`, `--validation_sampling_steps`
### 3. Set environment variables
```bash
export WANDB_API_KEY="${WANDB_API_KEY}"
export WANDB_BASE_URL="https://api.wandb.ai"
export FASTVIDEO_ATTENTION_BACKEND=FLASH_ATTN
export TOKENIZERS_PARALLELISM=false
export TRITON_CACHE_DIR=/tmp/triton_cache
```
### 4. Construct the torchrun command
```bash
torchrun --nnodes 1 --nproc_per_node <num_gpus> \
<entrypoint> \
--pretrained_model_name_or_path <model_hf_id> \
--data_path "<data_path>" \
--output_dir "<output_dir>" \
--wandb_run_name "<run_name>" \
--tracker_project_name "<project_name>" \
--log_validation \
<...all hyperparameters...>
```
### 5. Log to experiment journal
After launching, append an entry to `.agents/memory/experiment-journal/README.md`:
```markdown
## [YYYY-MM-DD] Experiment: <run_name>
- **Hypothesis**: <user-provided or auto-generated>
- **Config**: model=<model>, lr=<lr>, sp_size=<sp>, gpus=<n>, script=<entrypoint>
- **W&B run**: <pending — will be updated by monitor skill>
- **Status**: running
```
## Outputs
- A ready-to-execute shell command.
- An experiment journal entry.
## Example Usage
```
Launch a Wan T2V 1.3B finetune on 4 GPUs with lr=5e-5 and max_train_steps=1000:
pipeline: finetune
model: wan-t2v-1.3B
data_path: data/crush_smol_preprocessed/
num_gpus: 4
overrides:
learning_rate: 5e-5
max_train_steps: 1000
```
## References
- `examples/training/finetune/wan_t2v_1.3B/crush_smol/finetune_t2v.sh`
- `scripts/distill/v1_distill_dmd_wan.sh`
- `docs/training/finetune.md` (training arguments table)
- `fastvideo/training/trackers.py` (tracker initialization)
## Changelog
| Date | Change |
|------|--------|
| 2026-03-02 | Initial version |
-87
View File
@@ -1,87 +0,0 @@
---
name: log-experiment
description: Append or update an experiment entry in the experiment journal
---
# Log Experiment
## Purpose
Create or update an entry in `.agents/memory/experiment-journal/README.md` to maintain
a living record of all experiments and their outcomes.
## Prerequisites
- `.agents/memory/experiment-journal/README.md` exists.
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| `name` | Yes | Experiment name / identifier |
| `hypothesis` | No | What you expected to learn |
| `config` | Yes | Key config: model, lr, sp_size, gpus, script |
| `wandb_run` | No | W&B run ID or URL |
| `duration` | No | Total wall time |
| `metrics` | No | Key metrics dict (loss, step_time, grad_norm) |
| `checkpoint` | No | Path to checkpoint |
| `insight` | No | What was learned |
| `status` | Yes | `running`, `completed`, `failed`, `abandoned` |
| `lessons` | No | Paths to related lesson files |
## Steps
### 1. Check for existing entry
Search `.agents/memory/experiment-journal/README.md` for an entry with the same name.
If found, update it instead of creating a duplicate.
### 2. Format the entry
```markdown
## [YYYY-MM-DD] Experiment: <name>
- **Hypothesis**: <hypothesis or "N/A">
- **Config**: model=<model>, lr=<lr>, sp_size=<sp>, gpus=<n>, script=<script>
- **W&B run**: <wandb_run or "pending">
- **Duration**: <duration or "in progress">
- **Key metrics**: loss=<loss>, step_time=<step_time>, grad_norm=<grad_norm>
- **Checkpoint**: <checkpoint or "N/A">
- **Insight**: <insight or "pending">
- **Status**: <status>
- **Related lessons**: <lessons or "none">
```
### 3. Insert at the top of the journal
New entries go at the top of the file (after the header), so the most recent
experiments are always visible first.
### 4. Warn on duplicates
If a similar experiment name exists with `status: completed`, warn that this
may be a repeat. If it's `status: running`, assume this is an update.
## Outputs
- Updated `.agents/memory/experiment-journal/README.md`.
## Example Usage
```
Log a completed experiment:
name: wan-t2v-finetune-lr5e5-sp4
config: model=wan-t2v-1.3B, lr=5e-5, sp_size=4, gpus=4
wandb_run: fastvideo/training/run_abc123
duration: 2h 15m
metrics: {loss: 0.065, step_time: 2.3, grad_norm: 0.35}
checkpoint: outputs/wan_finetune/checkpoint-1000
insight: LR 5e-5 converges 30% faster than 1e-5 with no quality loss
status: completed
```
## References
- `.agents/memory/experiment-journal/README.md` — journal file
- `.agents/workflows/experiment-lifecycle.md` — when to log
## Changelog
| Date | Change |
|------|--------|
| 2026-03-02 | Initial version |
-134
View File
@@ -1,134 +0,0 @@
---
name: monitor-experiment
description: Poll a running W&B training run for progress and emit structured alerts
---
# Monitor Experiment
## Purpose
Continuously (or on-demand) check a running experiment's W&B metrics and emit
alerts for anomalies. Supports the "30-minute quality check" paradigm: after
the first 30 minutes of a long training run, produce a checkpoint quality
report before committing more resources.
## Prerequisites
- `WANDB_API_KEY` is set in the environment.
- The experiment is actively logging to W&B (not in `WANDB_MODE=offline`).
- For offline mode: read from local `wandb-summary.json` instead.
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| `run_id` | Yes* | W&B run ID (e.g., `entity/project/run_id`) |
| `output_dir` | Yes* | Local output directory (for offline mode fallback) |
| `poll_interval` | No | Seconds between polls (default: 60) |
| `alert_on` | No | List of alert conditions to enable (default: all) |
\* One of `run_id` or `output_dir` is required.
## Steps
### 1. Connect to the run
**Online mode** (preferred):
```python
import wandb
api = wandb.Api()
run = api.run("<run_id>")
```
**Offline fallback**:
```python
import json
summary_path = f"{output_dir}/tracker/wandb/latest-run/files/wandb-summary.json"
with open(summary_path) as f:
summary = json.load(f)
```
### 2. Track key metrics
| Metric | W&B Key | Description |
|--------|---------|-------------|
| Training loss | `train_loss` | Primary training loss |
| Gradient norm | `grad_norm` | Gradient magnitude |
| Step time | `step_time` | Wall-clock seconds per step |
| Learning rate | `learning_rate` | Current LR |
| Avg step time | `avg_step_time` | Running average step time |
| Validation videos | `validation_videos_*` | Generated validation samples |
### 3. Evaluate alert conditions
| Alert | Condition | Severity |
|-------|-----------|----------|
| **Loss spike** | `current_loss > 3 × rolling_avg_loss` | 🔴 Critical |
| **NaN/Inf gradient** | `grad_norm` is NaN or Inf | 🔴 Critical |
| **Step time regression** | `step_time > 2 × baseline_step_time` | 🟡 Warning |
| **No progress** | No new W&B logs for > 10 minutes | 🟡 Warning |
| **Loss plateau** | Loss change < 1% over last 100 steps | 🟢 Info |
### 4. Emit structured status
Output format (agent-consumable):
```json
{
"run_id": "...",
"step": 500,
"metrics": {
"train_loss": 0.078,
"grad_norm": 0.41,
"step_time": 2.5,
"learning_rate": 1e-6
},
"alerts": [
{"type": "loss_spike", "severity": "critical", "message": "Loss jumped to 0.45 (avg: 0.08)"}
],
"status": "running"
}
```
### 5. 30-Minute Quality Check
After the first 30 minutes of wall-clock time:
1. Summarize the loss curve shape (decreasing? at what rate?).
2. Check if validation videos have been generated.
3. Report step count, loss at start vs. current, and estimated time to completion.
4. Produce a go/no-go recommendation.
```markdown
## 30-Minute Check: <run_name>
- **Steps completed**: 150
- **Loss**: 0.12 → 0.08 (↓ 33%)
- **Grad norm**: stable at ~0.4
- **Step time**: 2.5s/step (consistent)
- **Validation videos**: 5 generated at step 100
- **Recommendation**: ✅ Continue — loss is decreasing normally
```
## Outputs
- Structured JSON status updates.
- Alert messages for anomalous conditions.
- 30-minute checkpoint quality report.
## Example Usage
```
Monitor W&B run "fastvideo/Wan_distillation/abc123":
run_id: fastvideo/Wan_distillation/abc123
poll_interval: 120
alert_on: [loss_spike, nan_gradient, step_time_regression]
```
## References
- `fastvideo/training/trackers.py` — `WandbTracker` implementation
- `fastvideo/tests/training/Vanilla/test_training_loss.py` — how summaries are compared
- `fastvideo/tests/training/Vanilla/a40_reference_wandb_summary.json` — reference summary format
## Changelog
| Date | Change |
|------|--------|
| 2026-03-02 | Initial version |
@@ -1,41 +0,0 @@
---
name: rlhf-training-abstractions
description: Use when changing FastVideo RLHF/RL training infrastructure, especially sampler, reward, scheduler trajectory, or method boundaries under fastvideo/train.
---
# RLHF Training Abstractions
Use this skill before editing RLHF-style training code in `fastvideo/train`.
## Boundaries
- RL methods live under `fastvideo/train/methods/rl/` and own algorithm logic: reward collection, advantage computation, policy loss, KL/reference terms, and optimizer cadence.
- Rewards live under `fastvideo/train/methods/rl/rewards/` and must be reusable across RL methods.
- RL methods pass decoded media to rewards; each reward decides whether to use the first frame, sampled frames, or the full video.
- Sampling lives under `fastvideo/train/methods/rl/common/` and must use `ModelBase` primitives plus scheduler math, not model-family inference pipelines.
- Model wrappers under `fastvideo/train/models/` own model-specific forward details.
- Model wrappers also own model-specific latent decoding via `ModelBase.decode_latents`; RL methods should not reach into VAE normalization internals.
- Shared RL helpers such as K-repeat prompt sampling belong under `fastvideo/train/methods/rl/common/` when they are reusable across RL methods.
## Anti-Patterns
- Do not bind RL methods to inference pipeline classes such as `WanDMDPipeline`.
- Do not hardcode timestep lists in a method when the scheduler can generate them.
- Do not put reward-model code inside one RL method.
- Do not make existing non-RL methods use method-managed optimization unless explicitly requested.
## Sampling Policy
- Prefer YAML-configured `method.sampling` with `scheduler`, `trajectory`, `num_steps`, `timesteps`, and `sigmas`.
- Treat diffusers-style scheduler classes as owning both the timestep schedule and their `step()` update rule; avoid a separate `solver` field unless a new sampler truly implements solver math outside the scheduler object.
- Missing `timesteps` means “ask the scheduler”; explicit `timesteps` or `sigmas` are overrides.
- ODE-style trajectories should not re-noise between denoising steps.
- SDE/re-noise behavior must be explicit in config.
## Validation
- Run focused local tests for sampler config and Trainer opt-in behavior.
- Verify existing train methods still report `manages_optimization() == False`.
- Keep fixed-prompt validation helpers in `fastvideo/train/methods/rl/common/validation.py` so new RL methods can reuse sharding and captions.
- Test distributed prompt grouping helpers separately from heavyweight model loading.
- Run `pre-commit run --files <changed paths>`; respect configured excludes.
+6 -4
View File
@@ -150,10 +150,12 @@ modal run fastvideo/tests/modal/ssim_test.py \
Env prefix rationale (parity with CI; see `.buildkite/pipeline.yml:1-3` and
`.buildkite/scripts/pr_test.sh:62-83`):
- `IMAGE_VERSION=py3.12-latest`: pins the Modal image tag to the same one CI
uses. Without this, `ssim_test.py:17` falls back to `latest`, which on
GHCR is built from `Dockerfile.python3.10` — different Python, torch, and
flash-attn wheel than CI's `py3.12-latest` (`infra-build-image.yml:51-67`,
`_template-build-image.yml:65-101`).
uses. The published `py3.12-latest` and `latest` tags point at Python 3.12 /
CUDA 12.6.3 / cu126; `py3.12-cuda12.6.3-latest` is the explicit alias for the
same image. CUDA 13 / cu130 is available under the explicit
`py3.12-cuda13.0.0-latest` tag. This tag policy comes from
`infra-build-image.yml`; the unparameterized `docker/Dockerfile` build itself
still defaults to CUDA 13 / cu130.
- `BUILDKITE_REPO`/`BUILDKITE_COMMIT`/`BUILDKITE_PULL_REQUEST`: mirror what
Buildkite exports. `ssim_test.py:38-46` bakes these into the image's
`.env(...)` block; mismatched values can perturb in-container code paths
-137
View File
@@ -1,137 +0,0 @@
---
name: summarize-run
description: Extract a W&B run summary into a structured experiment report
---
# Summarize Run
## Purpose
After a training run completes (or at any checkpoint), extract key metrics from
the W&B run summary and produce a structured markdown report. Supports both
online (W&B API) and offline (local `wandb-summary.json`) modes.
## Prerequisites
- Run has completed or reached a checkpoint with a saved summary.
- For online: `WANDB_API_KEY` set in environment.
- For offline: access to `<output_dir>/tracker/wandb/latest-run/files/wandb-summary.json`.
## Inputs
| Parameter | Required | Description |
|-----------|----------|-------------|
| `run_id` | Yes* | W&B run ID for online access |
| `output_dir` | Yes* | Local output dir for offline access |
| `reference_run` | No | Path to reference `wandb-summary.json` for comparison |
| `experiment_name` | No | Name for the journal entry (default: from W&B) |
\* One of `run_id` or `output_dir` is required.
## Steps
### 1. Load run summary
**Online**:
```python
import wandb
api = wandb.Api()
run = api.run("<run_id>")
summary = dict(run.summary)
config = dict(run.config)
```
**Offline** (existing codebase pattern from `fastvideo/tests/training/`):
```python
import json
summary_path = f"{output_dir}/tracker/wandb/latest-run/files/wandb-summary.json"
with open(summary_path) as f:
summary = json.load(f)
```
### 2. Extract key fields
| Field | Source | Description |
|-------|--------|-------------|
| `train_loss` | `summary["train_loss"]` | Final training loss |
| `avg_step_time` | `summary["avg_step_time"]` | Average seconds per step |
| `step_time` | `summary["step_time"]` | Last step time |
| `grad_norm` | `summary["grad_norm"]` | Final gradient norm |
| `learning_rate` | `summary["learning_rate"]` | Final LR |
| `_step` | `summary["_step"]` | Total steps completed |
| `_runtime` | `summary["_runtime"]` | Total wall-clock seconds |
| `validation_videos_*` | `summary[key]` | Validation video artifacts |
### 3. Compare against reference (optional)
Follow the pattern in `fastvideo/tests/training/Vanilla/test_training_loss.py`:
```python
# Fields to compare
compare_fields = ["train_loss", "grad_norm", "avg_step_time"]
tolerance = 0.05 # 5% relative tolerance
for field in compare_fields:
ref_val = reference_summary[field]
cur_val = summary[field]
diff_pct = abs(cur_val - ref_val) / abs(ref_val) * 100
status = "✅" if diff_pct < tolerance * 100 else "⚠️"
print(f"{status} {field}: {cur_val:.4f} (ref: {ref_val:.4f}, diff: {diff_pct:.1f}%)")
```
### 4. Generate report
```markdown
# Run Summary: <experiment_name>
| Metric | Value | Reference | Diff |
|--------|-------|-----------|------|
| Train Loss | 0.0788 | 0.0800 | -1.5% ✅ |
| Avg Step Time | 2.81s | 2.80s | +0.4% ✅ |
| Grad Norm | 0.408 | 0.410 | -0.5% ✅ |
| Total Steps | 500 | — | — |
| Wall Time | 23m 30s | — | — |
## Configuration
- Model: Wan-AI/Wan2.1-T2V-1.3B-Diffusers
- Learning Rate: 1e-6
- Batch Size: 1
- GPUs: 8 × (SP=1, TP=1)
- Mixed Precision: bf16
## Validation Videos
<list of validation video paths if available>
## Notes
<any observations or anomalies>
```
### 5. Update experiment journal
Append or update the experiment's entry in `.agents/memory/experiment-journal/README.md`
with the final metrics and status.
## Outputs
- Structured markdown report.
- Updated experiment journal entry.
## Example Usage
```
Summarize the run in output directory "outputs/wan_finetune":
output_dir: outputs/wan_finetune
reference_run: fastvideo/tests/training/Vanilla/a40_reference_wandb_summary.json
experiment_name: wan-t2v-finetune-lr1e6
```
## References
- `fastvideo/tests/training/Vanilla/test_training_loss.py` — reference comparison pattern
- `fastvideo/tests/training/Vanilla/a40_reference_wandb_summary.json` — example summary
- `fastvideo/tests/training/lora/test_lora_training.py` — LoRA summary comparison
- `fastvideo/training/trackers.py` — tracker summary generation
## Changelog
| Date | Change |
|------|--------|
| 2026-03-02 | Initial version |
@@ -1,92 +0,0 @@
---
description: How to develop, validate, and register a new evaluation metric
---
# Evaluation Development SOP
Standard procedure for adding new video quality evaluation metrics to
the FastVideo agent toolkit.
## When to use
- You need a metric that does not exist in
`.agents/memory/evaluation-registry/README.md`.
- An existing metric needs significant changes to its methodology.
- You are exploring a new evaluation approach.
## Steps
### 1. Research
- Search `.agents/memory/related-work/` for existing evaluation
approaches.
- Check `.agents/memory/evaluation-registry/README.md` for current
metrics and their limitations.
- Review literature: FVD, CLIP-Score, human preference, etc.
### 2. Prototype
- Write a standalone script in `.agents/exploration/<metric-name>.md`.
- Keep it simple: one script, minimal dependencies.
- Test on a few known-good and known-bad video samples.
### 3. Validate
- **Known-good test**: metric should score high on reference-quality
videos.
- **Known-bad test**: metric should score low on degraded or unrelated
videos.
- **Sensitivity test**: small quality differences should produce
meaningful score differences.
- Document thresholds and their justification.
### 4. Register
Update `.agents/memory/evaluation-registry/README.md`:
- Add the metric with status `Active`.
- Document location, thresholds, and trust level.
### 5. Integrate
Update `.agents/skills/evaluate-video-quality/SKILL.md`:
- Add the new metric as a section.
- Include code examples and interpretation guide.
### 6. Document
- Move the exploration log content into the skill.
- Clean up the exploration file or mark it as `promoted`.
- If anything went wrong during development, create a lesson.
## Where the metrics live
The eval suite is `fastvideo/eval/`. New metrics register themselves
via `@register("<group>.<name>")` and are auto-discovered when
`fastvideo.eval.metrics` is imported.
- **Native metrics** (SSIM, PSNR, LPIPS, optical flow, VLM): add a
file under the appropriate group dir
(`fastvideo/eval/metrics/common/`, `optical_flow/`, `videoscore2/`,
`physics_iq/`).
- **Metrics that wrap upstream research code**: follow the vbench
pattern in `fastvideo/eval/metrics/vbench/`. The contract is:
- Upstream lives as a git submodule under
`fastvideo/third_party/eval/<bench>/`, pinned to a SHA in repo-root
`.gitmodules`.
- The metric package's `__init__.py` inserts the submodule path on
`sys.path` and installs runtime compat shims (attribute-level
monkey-patches) for any modern-dep drift. Do not modify upstream
files on disk, and do not ship a `setup.sh`.
- See `fastvideo/eval/README.md` for the worked vbench example.
- Full porting guide:
[`docs/contributing/eval-metrics.md`](../../docs/contributing/eval-metrics.md).
## Out of scope of the initial eval port
The following land in follow-up PRs:
- **MIND** metrics (depends on a separate `vipe` submodule).
- **VBench-2.0** sibling package.
- The training-time `EvalCallback`.
@@ -1,47 +0,0 @@
---
description: When and how to log experiments in the experiment journal
---
# Experiment Journaling SOP
Ensures every experiment is properly recorded with context and outcomes.
## When to Log
**Always.** Every experiment — even quick tests — should be journaled.
## Steps
### 1. Before Launch — Create Draft Entry
Use the `log-experiment` skill with `status: running`:
- Include hypothesis and config.
- Leave metrics, duration, and insight blank.
### 2. After 30-Minute Check — Update with Initial Metrics
Update the entry with:
- Current loss and its trajectory direction.
- Step time.
- Number of validation videos generated.
- Preliminary go/no-go assessment.
### 3. On Completion — Fill Final Entry
Update the entry with `status: completed`:
- Final loss, grad norm, avg step time.
- Total duration and steps.
- Checkpoint path.
- Key insight.
### 4. On Failure — Document Failure Mode
Update the entry with `status: failed`:
- What went wrong (OOM, NaN, crash, etc.).
- At what step the failure occurred.
- Create a lesson in `.agents/lessons/` for non-trivial failures.
### 5. Cross-Reference
- Link related lessons: `**Related lessons**: .agents/lessons/<filename>.md`
- Link related experiments: if this is a follow-up, reference the prior entry.
-87
View File
@@ -1,87 +0,0 @@
---
description: End-to-end experiment lifecycle from hypothesis to lessons learned
---
# Experiment Lifecycle SOP
Standard operating procedure for running ML training experiments on
FastVideo-WorldModel. Every experiment should follow this flow.
## Overview
```
Plan → Launch → Monitor → Summarize → Journal → Reflect
```
## Steps
### 1. Plan the Experiment
Before launching:
- [ ] Define a clear **hypothesis** (what you expect to learn).
- [ ] Select the **model** and **pipeline** type (finetune, distill, lora, etc.).
- [ ] Prepare the **dataset** (preprocessed into parquet format).
- [ ] Review existing experiments in `.agents/memory/experiment-journal/README.md` for related work.
- [ ] Check `.agents/lessons/` for known pitfalls with this configuration.
- [ ] Document the plan in the experiment journal as a draft entry.
### 2. Launch the Experiment
Use the `launch-experiment` skill:
- Provide: pipeline, model, data_path, num_gpus, and any hyperparameter overrides.
- The skill generates the `torchrun` command and creates a journal entry.
- Verify the command looks correct before executing.
Reference: `.agents/skills/launch-experiment.md`
### 3. Monitor the Experiment
Use the `monitor-experiment` skill:
- Provide the W&B run ID (or output_dir for offline).
- Monitor alerts: loss spikes, NaN gradients, step time regressions.
- At the **30-minute mark**: perform the quality check.
- Is loss decreasing?
- Are validation videos reasonable?
- Is step time consistent?
- **Decision point**: Continue or abort based on the 30-min check.
Reference: `.agents/skills/monitor-experiment.md`
### 4. Summarize the Run
After completion (or at any checkpoint), use the `summarize-run` skill:
- Extract final metrics from W&B summary.
- Compare against reference runs if available.
- Generate a structured report.
Reference: `.agents/skills/summarize-run.md`
### 5. Update the Experiment Journal
Use the `log-experiment` skill to update the journal entry:
- Fill in final metrics, duration, checkpoint paths.
- Record the key insight learned.
- Set status to `completed`, `failed`, or `abandoned`.
Reference: `.agents/skills/log-experiment.md`
### 6. Reflect and Capture Lessons
After every experiment:
- **What went right?** → Note in the journal insight field.
- **What went wrong?** → Create a lesson in `.agents/lessons/`:
- Use the template in `.agents/lessons/README.md`.
- Cross-reference the experiment journal entry.
- **What was surprising?** → Consider creating an exploration log if this
warrants further investigation.
Reference: `.agents/workflows/lesson-capture.md`
## Validation Criteria
This SOP is validated when an agent can:
1. Follow steps 1–6 end-to-end for a minimal training run
(e.g., `examples/training/finetune/wan_t2v_1.3B/crush_smol/finetune_t2v.sh`
with `--max_train_steps 5`).
2. Produce a complete experiment journal entry.
3. Generate a run summary report.
-71
View File
@@ -1,71 +0,0 @@
---
description: Post-experiment reflection to capture lessons learned
---
# Lesson Capture SOP
Systematic procedure for turning experiment outcomes into persistent knowledge.
## When to Use
After **every** completed or failed experiment. Even successful experiments
can yield lessons (e.g., "LR 5e-5 works better than 1e-5 for LoRA").
## Steps
### 1. Review the Experiment
Read the experiment journal entry. Ask:
- Did anything go wrong?
- Was anything surprising?
- Did anything take longer than expected?
- Was a workaround needed?
### 2. Decide: Lesson or Not?
| Situation | Action |
|-----------|--------|
| Something broke | Create a lesson (category: `infrastructure` or `data`) |
| Hyperparameter choice mattered | Create a lesson (category: `hyperparameter`) |
| Porting issue found | Create a lesson (category: `porting`) |
| Evaluation metric was misleading | Create a lesson (category: `evaluation`) |
| Everything went smoothly | No lesson needed, but note in the journal insight |
### 3. Create the Lesson File
In `.agents/lessons/`, create `<YYYY-MM-DD>_<short-slug>.md`:
```markdown
---
date: <ISO-8601>
experiment: <journal entry reference>
category: hyperparameter | data | infrastructure | evaluation | porting
severity: critical | important | minor
---
# <Short Descriptive Title>
## What Happened
<description>
## Root Cause
<analysis>
## Fix / Workaround
<resolution>
## Prevention
<how to avoid in future>
```
### 4. Cross-Reference
- Update the experiment journal entry with a link to the lesson file.
- If a similar lesson already exists, add a reference or update it.
### 5. Periodic Pattern Review
Every ~10 lessons, scan for patterns:
- Multiple lessons in the same category → consider a new skill or SOP.
- Repeated mistakes → strengthen the relevant SOP with a checklist item.
- Infrastructure issues → propose a codebase fix.
+40 -17
View File
@@ -217,6 +217,17 @@ steps:
limit: 2
agents:
queue: "default"
- label: ":test_tube: Eval Metrics Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "eval"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
# ============================================================
# Fastcheck: Runs on every PR (~10-15 min parallel)
@@ -240,7 +251,7 @@ steps:
- "fastvideo/models/loader/**"
- "fastvideo/tests/encoders/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: ":microscope: Encoder Tests"
@@ -253,7 +264,7 @@ steps:
- "fastvideo/models/loader/**"
- "fastvideo/tests/vaes/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: ":microscope: VAE Tests"
@@ -268,7 +279,7 @@ steps:
- "fastvideo/layers/**"
- "fastvideo/attention/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Transformer Tests"
@@ -279,7 +290,7 @@ steps:
- path:
- "fastvideo-kernel/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Kernel Tests"
@@ -292,7 +303,7 @@ steps:
- ".buildkite/**"
- ".github/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Unit Tests"
@@ -333,7 +344,7 @@ steps:
- path:
- "fastvideo/**/*.py"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":bar_chart: SSIM Tests"
@@ -352,7 +363,7 @@ steps:
- "fastvideo/pipelines/**"
- "fastvideo/layers/lora/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Inference Tests"
@@ -363,7 +374,7 @@ steps:
- path:
- "fastvideo/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Training Tests"
@@ -374,7 +385,7 @@ steps:
- path:
- "fastvideo/training/*distillation_pipeline.py"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Distillation DMD Tests"
@@ -386,9 +397,9 @@ steps:
- "fastvideo/training/*self_forcing_distillation_pipeline.py"
- "fastvideo/tests/training/self-forcing/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Self-Forcing Tests"
env:
- TEST_TYPE=self_forcing
@@ -397,7 +408,7 @@ steps:
- path:
- "fastvideo/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Training Tests"
@@ -413,7 +424,7 @@ steps:
- "fastvideo/**"
- "fastvideo-kernel/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Training Tests VSA"
@@ -429,7 +440,7 @@ steps:
- "fastvideo-kernel/**"
- "fastvideo/attention/backends/vmoba.py"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Inference Tests VMoBA"
@@ -447,7 +458,7 @@ steps:
- "fastvideo/tests/performance/**"
- ".buildkite/performance-benchmarks/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Performance Tests"
@@ -460,7 +471,7 @@ steps:
- "fastvideo/entrypoints/cli/serve.py"
- "fastvideo/tests/entrypoints/test_openai_api_integration.py"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: API Server Tests"
@@ -475,7 +486,7 @@ steps:
- "fastvideo/models/dits/**"
- "fastvideo/models/loader/**"
- "pyproject.toml"
- "docker/Dockerfile.python3.12"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Train Framework Tests"
@@ -483,3 +494,15 @@ steps:
- TEST_TYPE=train_framework
agents:
queue: "default"
- path:
- "fastvideo/eval/**"
- "fastvideo/tests/eval/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Eval Metrics Tests"
env:
- TEST_TYPE=eval
agents:
queue: "default"
+5 -1
View File
@@ -76,7 +76,7 @@ EFFECTIVE_PR=${BUILDKITE_PULL_REQUEST:-false}
if [ "$EFFECTIVE_PR" = "false" ] && [ -n "${PR_NUMBER:-}" ]; then
EFFECTIVE_PR=$PR_NUMBER
fi
MODAL_ENV="BUILDKITE_REPO=$BUILDKITE_REPO BUILDKITE_COMMIT=$BUILDKITE_COMMIT BUILDKITE_PULL_REQUEST=$EFFECTIVE_PR BUILDKITE_BRANCH=${BUILDKITE_BRANCH:-} TEST_SCOPE=${TEST_SCOPE:-} IMAGE_VERSION=$IMAGE_VERSION"
MODAL_ENV="BUILDKITE_REPO=$BUILDKITE_REPO BUILDKITE_COMMIT=$BUILDKITE_COMMIT BUILDKITE_PULL_REQUEST=$EFFECTIVE_PR BUILDKITE_BRANCH=${BUILDKITE_BRANCH:-} TEST_SCOPE=${TEST_SCOPE:-} BUILDKITE_BUILD_URL=${BUILDKITE_BUILD_URL:-} BUILDKITE_BUILD_ID=${BUILDKITE_BUILD_ID:-} BUILDKITE_JOB_ID=${BUILDKITE_JOB_ID:-} IMAGE_VERSION=$IMAGE_VERSION"
POST_RUN_HOOK=""
@@ -219,6 +219,10 @@ case "$TEST_TYPE" in
log "Running fastvideo.train framework tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_train_framework_tests"
;;
"eval")
log "Running eval metric tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_eval_tests"
;;
"lora_extraction")
log "Running LoRA extraction tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_lora_extraction_tests"
+31
View File
@@ -0,0 +1,31 @@
# Build-context excludes: keep the context small and the `COPY . .` layer cache
# stable. Docker uploads everything here to the daemon and bakes it into a layer;
# without this, the 6.5 GB host .venv alone is shipped + cached on every build.
#
# IMPORTANT: do NOT ignore .git — fastvideo-kernel's build runs
# `git submodule update --init --recursive`, which needs the repo metadata.
# Virtualenvs — the image builds its own /opt/venv
.venv/
venv/
env/
# Python caches & build/test/lint artifacts
**/__pycache__/
*.py[cod]
*.egg-info/
.eggs/
.pytest_cache/
.mypy_cache/
.ruff_cache/
.cache/
# Local run outputs / logs (not needed in the image)
outputs/
wandb/
*.log
# Editor / OS cruft
.DS_Store
.idea/
.vscode/
+4 -14
View File
@@ -186,13 +186,13 @@ pull_request_rules:
label:
add: ["scope: docs"]
- name: "label scope: ui"
- name: "label scope: studio"
conditions:
- files~=^ui/
- files~=^apps/fastvideo_studio/
- -closed
actions:
label:
add: ["scope: ui"]
add: ["scope: studio"]
- name: "label scope: model"
conditions:
@@ -272,7 +272,7 @@ pull_request_rules:
remove: [needs-rebase]
# ============================================================
# Auto-merge and auto-rebase
# Auto-merge
# ============================================================
- name: auto-merge when ready and all checks pass
@@ -290,16 +290,6 @@ pull_request_rules:
merge:
method: squash
- name: auto-update when ready
conditions:
- label=ready
- "#approved-reviews-by>=1"
- -conflict
- -closed
- -draft
actions:
update: {}
# ============================================================
# PR title format help
# ============================================================
+85 -7
View File
@@ -24,10 +24,30 @@ on:
required: false
type: boolean
default: true
mark_as_latest:
required: false
type: boolean
default: false
runner:
required: false
type: string
default: ubuntu-latest
architecture:
required: false
type: string
default: amd64
push_by_digest:
required: false
type: boolean
default: false
digest_artifact_name:
required: false
type: string
default: ''
jobs:
build-and-push:
runs-on: ubuntu-latest
runs-on: ${{ inputs.runner }}
permissions:
contents: read
packages: write
@@ -35,6 +55,10 @@ jobs:
steps:
- name: Checkout code
uses: actions/checkout@v4
# The Docker context intentionally includes .git so the kernel build can
# initialize its pinned submodules. Do not copy the checkout token with it.
with:
persist-credentials: false
- name: Free up disk space
run: |
@@ -73,6 +97,12 @@ jobs:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Normalize image reference
id: image
env:
IMAGE: ghcr.io/${{ github.repository }}/${{ inputs.image_name }}
run: echo "name=${IMAGE,,}" >> "${GITHUB_OUTPUT}"
- name: Prepare tags
id: prepare-tags
@@ -85,8 +115,8 @@ jobs:
TAGS="type=raw,value=${{ inputs.tag_suffix }}-latest\n${TAGS}"
fi
# Set Python 3.10 as the default image
if [[ "${{ inputs.include_latest_tags }}" == "true" && "${{ inputs.python_version }}" == "3.10" ]]; then
# Tag the designated default variant as the global `latest` image
if [[ "${{ inputs.include_latest_tags }}" == "true" && "${{ inputs.mark_as_latest }}" == "true" ]]; then
TAGS="${TAGS}\ntype=raw,value=latest"
fi
@@ -100,23 +130,71 @@ jobs:
id: meta
uses: docker/metadata-action@v5
with:
images: ghcr.io/${{ github.repository }}/${{ inputs.image_name }}
images: ${{ steps.image.outputs.name }}
tags: ${{ steps.prepare-tags.outputs.tags }}
- name: Build and push Docker image
if: ${{ !inputs.push_by_digest }}
id: build-push
uses: docker/build-push-action@v6
with:
context: .
file: ${{ inputs.dockerfile_path }}
platforms: linux/${{ inputs.architecture }}
push: true
tags: ${{ steps.meta.outputs.tags }}
labels: ${{ steps.meta.outputs.labels }}
build-args: ${{ inputs.build_args }}
cache-from: type=gha
cache-to: type=gha,mode=max
cache-from: type=gha,scope=${{ inputs.image_name }}-${{ inputs.tag_suffix }}-${{ inputs.architecture }}
cache-to: type=gha,mode=max,scope=${{ inputs.image_name }}-${{ inputs.tag_suffix }}-${{ inputs.architecture }}
# Multi-architecture callers publish immutable manifests by digest here,
# then create the shared user-facing tags in a single downstream job.
# This prevents architecture jobs from racing to replace the same tag.
- name: Build and push Docker image by digest
if: ${{ inputs.push_by_digest }}
id: build-push-digest
uses: docker/build-push-action@v6
with:
context: .
file: ${{ inputs.dockerfile_path }}
platforms: linux/${{ inputs.architecture }}
labels: ${{ steps.meta.outputs.labels }}
build-args: ${{ inputs.build_args }}
outputs: type=image,name=${{ steps.image.outputs.name }},push-by-digest=true,name-canonical=true,push=true
cache-from: type=gha,scope=${{ inputs.image_name }}-${{ inputs.tag_suffix }}-${{ inputs.architecture }}
cache-to: type=gha,mode=max,scope=${{ inputs.image_name }}-${{ inputs.tag_suffix }}-${{ inputs.architecture }}
- name: Export digest
if: ${{ inputs.push_by_digest }}
env:
DIGEST: ${{ steps.build-push-digest.outputs.digest }}
run: |
if [[ -z "${{ inputs.digest_artifact_name }}" ]]; then
echo "digest_artifact_name is required when push_by_digest is true" >&2
exit 1
fi
mkdir -p /tmp/digests
touch "/tmp/digests/${DIGEST#sha256:}"
- name: Upload digest
if: ${{ inputs.push_by_digest }}
uses: actions/upload-artifact@v4
with:
name: ${{ inputs.digest_artifact_name }}
path: /tmp/digests/*
if-no-files-found: error
retention-days: 1
- name: Success message
if: ${{ !inputs.push_by_digest }}
run: |
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ghcr.io/${{ github.repository }}/${{ inputs.image_name }}:${{ inputs.tag_suffix }}-sha-${GITHUB_SHA::7}"
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ${{ steps.image.outputs.name }}:${{ inputs.tag_suffix }}-sha-${GITHUB_SHA::7}"
echo "To run tests with this image, manually trigger the 'Run Tests' workflow."
- name: Digest success message
if: ${{ inputs.push_by_digest }}
env:
DIGEST: ${{ steps.build-push-digest.outputs.digest }}
run: |
echo "✅ Python ${{ inputs.python_version }} linux/${{ inputs.architecture }} image pushed as ${DIGEST}"
+2 -2
View File
@@ -125,7 +125,7 @@ jobs:
set -euo pipefail
TEST_NAME=$(echo "$COMMENT" | grep -oP '(?<=/test\s)\S+' | head -1 || true)
VALID="encoder vae transformer kernel unit dreamverse ssim training lora-inference lora-training distillation self-forcing vsa vmoba performance api train-framework full fastcheck pre-commit"
VALID="encoder vae transformer kernel unit dreamverse ssim training lora-inference lora-training distillation self-forcing vsa vmoba performance api train-framework eval full fastcheck pre-commit"
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
echo "Unknown test: '$TEST_NAME'. Valid: $VALID"
exit 1
@@ -139,7 +139,7 @@ jobs:
[distillation]=distillation_dmd [self-forcing]=self_forcing
[vsa]=training_vsa [vmoba]=inference_vmoba
[performance]=performance [api]=api_server
[train-framework]=train_framework
[train-framework]=train_framework [eval]=eval
)
if [ "$TEST_NAME" = "full" ]; then
+168 -66
View File
@@ -3,28 +3,13 @@ name: Build and Push Docker Images
on:
workflow_dispatch:
inputs:
python_3_10:
description: 'Build Python 3.10 image'
build_cuda_matrix:
description: 'Build multi-arch CUDA images (12.6.3 default; 13.0.0 explicitly tagged)'
required: false
default: false
default: true
type: boolean
python_3_11:
description: 'Build Python 3.11 image'
required: false
default: false
type: boolean
python_3_12:
description: 'Build Python 3.12 image'
required: false
default: false
type: boolean
python_3_12_cuda_12_9:
description: 'Build Python 3.12 image Cuda 12.9'
required: false
default: false
type: boolean
dreamverse_cuda_12_9:
description: 'Build Dreamverse CUDA 12.9 backend-only and UI images'
build_dreamverse_matrix:
description: 'Build the amd64 Dreamverse matrix (backend + UI x CUDA 12.6.3/13.0.0)'
required: false
default: false
type: boolean
@@ -35,62 +20,179 @@ permissions:
packages: write
jobs:
build-python-3-10:
if: ${{ github.event.inputs.python_3_10 == 'true' }}
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.10'
dockerfile_path: docker/Dockerfile.python3.10
tag_suffix: py3.10
secrets: inherit
build-python-3-11:
if: ${{ github.event.inputs.python_3_11 == 'true' }}
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.11'
dockerfile_path: docker/Dockerfile.python3.11
tag_suffix: py3.11
secrets: inherit
build-python-3-12:
if: ${{ github.event.inputs.python_3_12 == 'true' }}
# CUDA matrix: Python 3.12 x {12.6.3, 13.0.0} x {amd64, arm64}. Each architecture
# builds natively and pushes only by digest; publish-cuda-manifests is the sole
# owner of the shared tags. CUDA 12.6 arm64 targets Hopper (sm_90a / GH200 class),
# because CUDA 12.6 cannot compile sm_121. CUDA 13 arm64 targets GB10 (sm_121).
# 12.6.3/cu126 owns the default `py3.12`/`latest` tags and keeps its versioned
# aliases; 13.0.0/cu130 is published under explicit versioned tags. Flash-attn
# 2.8.3 comes from the architecture-specific prebuilt releases.
build-cuda-images:
if: ${{ github.event.inputs.build_cuda_matrix == 'true' }}
strategy:
fail-fast: false
matrix:
cuda:
- version: '12.6.3'
suffix: '-cuda12.6.3'
torch_backend: 'cu126'
fa_tag: 'cu126torch2.12'
cmake_build_parallel_level: '4'
torch_cuda_arch_list:
amd64: '9.0a'
arm64: '9.0a'
- version: '13.0.0'
suffix: '-cuda13.0.0'
torch_backend: 'cu130'
fa_tag: 'cu130torch2.12'
cmake_build_parallel_level: '1'
torch_cuda_arch_list:
amd64: '9.0a'
arm64: '12.1'
architecture:
- name: amd64
runner: ubuntu-latest
- name: arm64
runner: ubuntu-24.04-arm
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.12'
dockerfile_path: docker/Dockerfile.python3.12
tag_suffix: py3.12
dockerfile_path: docker/Dockerfile
tag_suffix: py3.12${{ matrix.cuda.suffix }}
runner: ${{ matrix.architecture.runner }}
architecture: ${{ matrix.architecture.name }}
push_by_digest: true
digest_artifact_name: fastvideo-dev-cuda${{ matrix.cuda.version }}-${{ matrix.architecture.name }}
build_args: |
PYTHON_VERSION=3.12
CUDA_VERSION=${{ matrix.cuda.version }}
UV_TORCH_BACKEND=${{ matrix.cuda.torch_backend }}
TORCH_CUDA_ARCH_LIST=${{ matrix.cuda.torch_cuda_arch_list[matrix.architecture.name] }}
CMAKE_BUILD_PARALLEL_LEVEL=${{ matrix.cuda.cmake_build_parallel_level }}
FLASH_ATTN_WHEEL_TAG=${{ matrix.cuda.fa_tag }}
FLASH_ATTN_WHEEL_RELEASE=https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.9.17
FLASH_ATTN_WHEEL_RELEASE_ARM64=https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.9.22
secrets: inherit
build-python-3-12-cuda-12-9:
if: ${{ github.event.inputs.python_3_12_cuda_12_9 == 'true' }}
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.12'
dockerfile_path: docker/Dockerfile.python3.12.cuda12.9.1
tag_suffix: py3.12-cuda12.9.1
secrets: inherit
publish-cuda-manifests:
# !cancelled(): a failed sibling build leg must not skip the manifests for a
# CUDA lane whose own digests all exist; the digest-count check below fails
# the incomplete lane loudly instead.
if: ${{ !cancelled() && github.event.inputs.build_cuda_matrix == 'true' }}
needs: build-cuda-images
runs-on: ubuntu-latest
permissions:
contents: read
packages: write
strategy:
fail-fast: false
matrix:
cuda:
- version: '12.6.3'
suffix: '-cuda12.6.3'
is_default: true
- version: '13.0.0'
suffix: '-cuda13.0.0'
is_default: false
steps:
- name: Download architecture digests
uses: actions/download-artifact@v4
with:
path: /tmp/digests
pattern: fastvideo-dev-cuda${{ matrix.cuda.version }}-*
merge-multiple: true
build-dreamverse-backend-cuda-12-9:
if: ${{ github.event.inputs.dreamverse_cuda_12_9 == 'true' }}
- name: Set up Docker Buildx
uses: docker/setup-buildx-action@v3
- name: Login to GitHub Container Registry
uses: docker/login-action@v3
with:
registry: ghcr.io
username: ${{ github.repository_owner }}
password: ${{ secrets.GITHUB_TOKEN }}
- name: Normalize image reference
id: image
env:
IMAGE: ghcr.io/${{ github.repository }}/fastvideo-dev
run: echo "name=${IMAGE,,}" >> "${GITHUB_OUTPUT}"
- name: Prepare tags
id: prepare-tags
env:
VERSIONED_TAG_PREFIX: py3.12${{ matrix.cuda.suffix }}
PUBLISH_DEFAULT_TAGS: ${{ matrix.cuda.is_default }}
run: |
SHORT_SHA="${GITHUB_SHA::7}"
TAGS="type=raw,value=${VERSIONED_TAG_PREFIX}-latest\ntype=raw,value=${VERSIONED_TAG_PREFIX}-sha-${SHORT_SHA}"
if [[ "${PUBLISH_DEFAULT_TAGS}" == "true" ]]; then
TAGS="type=raw,value=latest\ntype=raw,value=py3.12-latest\ntype=raw,value=py3.12-sha-${SHORT_SHA}\n${TAGS}"
fi
{
echo "tags<<EOF"
echo -e "${TAGS}"
echo "EOF"
} >> "${GITHUB_OUTPUT}"
- name: Extract metadata for Docker
id: meta
uses: docker/metadata-action@v5
with:
images: ${{ steps.image.outputs.name }}
tags: ${{ steps.prepare-tags.outputs.tags }}
- name: Create multi-architecture manifests
env:
IMAGE: ${{ steps.image.outputs.name }}
run: |
mapfile -t DIGESTS < <(find /tmp/digests -maxdepth 1 -type f -printf '%f\n' | sort)
if [[ "${#DIGESTS[@]}" -ne 2 ]]; then
echo "Expected exactly two architecture digests, found ${#DIGESTS[@]}" >&2
exit 1
fi
mapfile -t TAGS < <(jq -r '.tags[]' <<< "${DOCKER_METADATA_OUTPUT_JSON}")
TAG_ARGS=()
for TAG in "${TAGS[@]}"; do
TAG_ARGS+=(--tag "${TAG}")
done
IMAGE_REFS=()
for DIGEST in "${DIGESTS[@]}"; do
IMAGE_REFS+=("${IMAGE}@sha256:${DIGEST}")
done
docker buildx imagetools create "${TAG_ARGS[@]}" "${IMAGE_REFS[@]}"
docker buildx imagetools inspect "${TAGS[0]}"
# Dreamverse matrix: {backend, UI} x {12.6.3, 13.0.0}, Python 3.12. Torch backend
# matches the base CUDA (cu126 / cu130). Keep these images amd64-only until the
# required FA4 dependency stack is available and validated on arm64.
build-dreamverse:
if: ${{ github.event.inputs.build_dreamverse_matrix == 'true' }}
strategy:
fail-fast: false
matrix:
cuda:
- version: '12.6.3'
torch_backend: 'cu126'
- version: '13.0.0'
torch_backend: 'cu130'
variant:
- name: backend
ui: '0'
- name: ui
ui: '1'
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.12'
dockerfile_path: apps/dreamverse/docker/Dockerfile
tag_suffix: dreamverse-backend-cuda12.9.1
tag_suffix: dreamverse-${{ matrix.variant.name }}-cuda${{ matrix.cuda.version }}
image_name: dreamverse
build_args: BUILD_DREAMVERSE_UI=0
include_latest_tags: false
secrets: inherit
build-dreamverse-ui-cuda-12-9:
if: ${{ github.event.inputs.dreamverse_cuda_12_9 == 'true' }}
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.12'
dockerfile_path: apps/dreamverse/docker/Dockerfile
tag_suffix: dreamverse-ui-cuda12.9.1
image_name: dreamverse
build_args: BUILD_DREAMVERSE_UI=1
build_args: |
CUDA_VERSION=${{ matrix.cuda.version }}
UV_TORCH_BACKEND=${{ matrix.cuda.torch_backend }}
BUILD_DREAMVERSE_UI=${{ matrix.variant.ui }}
include_latest_tags: false
secrets: inherit
+10 -5
View File
@@ -5,15 +5,21 @@ on:
branches: [ main ]
paths:
- 'docs/**'
- 'examples/**'
- 'mkdocs.yml'
- 'requirements-mkdocs.in'
- 'requirements-mkdocs.txt'
- 'scripts/check_docs_links.py'
- '.github/workflows/infra-docs.yml'
pull_request:
branches: [ main ]
paths:
- 'docs/**'
- 'examples/**'
- 'mkdocs.yml'
- 'requirements-mkdocs.in'
- 'requirements-mkdocs.txt'
- 'scripts/check_docs_links.py'
- '.github/workflows/infra-docs.yml'
permissions:
@@ -31,6 +37,8 @@ jobs:
steps:
- name: Checkout
uses: actions/checkout@v4
with:
fetch-depth: 0
- name: Setup Python
uses: actions/setup-python@v5
@@ -46,15 +54,12 @@ jobs:
- name: Setup Pages
uses: actions/configure-pages@v4
- name: Generate docs examples
run: python docs/generate_examples.py
- name: Build documentation
run: mkdocs build
- name: Check docs links
run: python scripts/check_docs_links.py
- name: Build documentation
run: mkdocs build
- name: Upload artifact
uses: actions/upload-pages-artifact@v3
with:
+69 -29
View File
@@ -47,29 +47,39 @@ jobs:
name: Build Wheel
needs: check-version-change
if: ${{ needs.check-version-change.outputs.version-changed == 'true' || github.event_name == 'workflow_dispatch' }}
runs-on: ${{ matrix.os }}
runs-on: ${{ matrix.platform.os }}
strategy:
fail-fast: false
matrix:
os: [ubuntu-22.04]
python-version: ['3.10', '3.11', '3.12']
python-version: ['3.12']
torch-cuda:
# - torch-version: '2.5.1'
# cuda-version: '12.4.1'
# torch-cuda-short: 'cu124'
# - torch-version: '2.6.0'
# cuda-version: '12.6.3'
# torch-cuda-short: 'cu126'
# - torch-version: '2.7.1'
# cuda-version: '12.8.0'
# torch-cuda-short: 'cu128'
# - torch-version: '2.9.1'
# cuda-version: '12.8.0'
# torch-cuda-short: 'cu128'
- torch-version: '2.10.0'
cuda-version: '12.8.0'
torch-cuda-short: 'cu128'
# torch 2.12 dropped cu128; cu130 (CUDA 13) is the default wheel, cu126 covers older drivers.
- torch-version: '2.12.0'
cuda-version: '12.6.3'
torch-cuda-short: 'cu126'
- torch-version: '2.12.0'
cuda-version: '13.0.0'
torch-cuda-short: 'cu130'
platform:
# x86_64 builds the full cu126 + cu130 set (cu130 ships the consumer
# Blackwell sm_120a FP4 kernels).
- os: ubuntu-22.04
arch: x86_64
wheel-plat: manylinux_2_35_x86_64
# aarch64 is Blackwell (GB200 sm_100a + DGX Spark / consumer sm_120a), not
# Hopper, and Blackwell needs CUDA >= 12.8 — so only the cu130 leg applies.
# Added via include so x86 keeps cu126 + cu130 while aarch64 stays cu130-only.
include:
- python-version: '3.12'
torch-cuda:
torch-version: '2.12.0'
cuda-version: '13.0.0'
torch-cuda-short: 'cu130'
platform:
os: ubuntu-22.04-arm
arch: aarch64
wheel-plat: manylinux_2_35_aarch64
steps:
- name: Free up disk space
@@ -104,7 +114,7 @@ jobs:
python-version: ${{ matrix.python-version }}
- name: Install CUDA ${{ matrix.torch-cuda.cuda-version }}
uses: Jimver/cuda-toolkit@v0.2.21
uses: Jimver/cuda-toolkit@v0.2.35
id: cuda-toolkit
with:
cuda: ${{ matrix.torch-cuda.cuda-version }}
@@ -152,9 +162,33 @@ jobs:
cd fastvideo-kernel
git submodule update --init --recursive # Ensure ThunderKittens submodule is initialized
# Release builds are produced on GPU-less runners, so force-enable TK and target Hopper.
export TORCH_CUDA_ARCH_LIST="9.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=ON -DCMAKE_CUDA_ARCHITECTURES=90a"
# Release builds run on GPU-less runners, so set kernels + arch explicitly:
# * aarch64 = Blackwell (GB200 sm_100a + DGX Spark/consumer sm_120a), NOT
# Hopper, so TK (sm_90a wgmma) is OFF. The C++ FP4 (attn_qat_infer, SM120)
# covers sm_120a; turbodiffusion covers sm_100a+sm_120a. The sm_100 FP4
# forward is the FA4 CuTe DSL path in the fastvideo package (PR #1221),
# JIT-compiled at runtime — not built into this wheel.
# * x86_64 cu130 = Hopper TK + consumer Blackwell sm_120a FP4.
# * x86_64 cu126 = Hopper TK only (older drivers; CUDA < 12.8 has no FP4).
# The per-arch split in CMakeLists pins the FP4 targets to sm_120a and builds
# the main extension for the full arch list. CMAKE_BUILD_PARALLEL_LEVEL caps
# Ninja so heavy CUTLASS/TK template TUs don't OOM the 16 GB runner (exit 143).
if [ "${{ matrix.platform.arch }}" = "aarch64" ]; then
export TORCH_CUDA_ARCH_LIST="10.0a;12.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=OFF -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON"
export CMAKE_BUILD_PARALLEL_LEVEL=1
elif [ "${{ matrix.torch-cuda.torch-cuda-short }}" = "cu130" ]; then
export TORCH_CUDA_ARCH_LIST="9.0a;12.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=ON -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON -DCMAKE_CUDA_ARCHITECTURES=90a"
# A single FP4 TU (attn_qat_infer) can use ~8-12 GB on its own, so serialize.
export CMAKE_BUILD_PARALLEL_LEVEL=1
else
export TORCH_CUDA_ARCH_LIST="9.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=ON -DCMAKE_CUDA_ARCHITECTURES=90a"
# Hopper-only: the two TK TUs are the heavy ones (single-arch) and fit side
# by side; -j4 overlaps the light TUs without OOMing (~12-14 GB peak).
export CMAKE_BUILD_PARALLEL_LEVEL=4
fi
# Build standard wheel (no local version suffix) for PyPI
python -m build --wheel --outdir dist
@@ -170,8 +204,8 @@ jobs:
PY
)
export LD_LIBRARY_PATH="${TORCH_LIB_DIR}:${LD_LIBRARY_PATH}"
# Target manylinux_2_35 (Ubuntu 22.04 native)
auditwheel repair dist/*.whl --plat manylinux_2_35_x86_64 -w fixed_dist \
# Target manylinux_2_35 (Ubuntu 22.04 native), per-arch plat tag.
auditwheel repair dist/*.whl --plat ${{ matrix.platform.wheel-plat }} -w fixed_dist \
--exclude libtorch_cuda.so \
--exclude libtorch_cpu.so \
--exclude libtorch.so \
@@ -183,11 +217,11 @@ jobs:
mv fixed_dist/*.whl dist/
- name: Upload wheel artifact
# Only upload if it's the "main" CUDA version we want on PyPI
# We upload all to artifacts for inspection/GH releases, but give them distinct artifact names
# Upload every matrix leg as a distinct artifact (for inspection / GitHub releases).
# The publish job below selects which CUDA build is pushed to PyPI.
uses: actions/upload-artifact@v4
with:
name: fastvideo_kernel-py${{ matrix.python-version }}-${{ matrix.torch-cuda.torch-cuda-short }}-torch${{ matrix.torch-cuda.torch-version }}
name: fastvideo_kernel-py${{ matrix.python-version }}-${{ matrix.torch-cuda.torch-cuda-short }}-${{ matrix.platform.arch }}-torch${{ matrix.torch-cuda.torch-version }}
path: fastvideo-kernel/dist/*.whl
retention-days: 90
@@ -204,13 +238,19 @@ jobs:
- uses: actions/setup-python@v5
with:
python-version: '3.10'
python-version: '3.12'
- name: Download PyPI wheels
# Publish the cu130 (CUDA 13) wheels to PyPI for both architectures:
# x86_64 — Hopper sm_90a TK + consumer Blackwell sm_120a FP4
# aarch64 — Blackwell: turbodiffusion (sm_100a/sm_120a) + C++ FP4 (sm_120a);
# no TK (Hopper). sm_100 FP4 forward is the FA4 CuTe DSL path in the
# fastvideo package (#1221), shipped/JIT separately.
# The x86_64 cu126 wheel stays available as a build artifact / GitHub-release asset.
uses: actions/download-artifact@v4
with:
path: fastvideo-kernel/dist/
pattern: 'fastvideo_kernel-py*'
pattern: 'fastvideo_kernel-py*-cu130-*'
merge-multiple: true
- name: Install uv
+8 -6
View File
@@ -53,6 +53,7 @@ eggs/
# MkDocs documentation
site/
docs/getting_started/examples/
docs/examples/
docs/inference/examples/
docs/training/examples/
docs/distillation/examples/
@@ -87,7 +88,7 @@ docs/distillation/examples/
dmd_t2v_output/
preprocess_output_text/
# Next.js / Node artifacts under ui/: see ui/.gitignore
# SvelteKit / Node artifacts under apps/fastvideo_studio/: see apps/fastvideo_studio/.gitignore
# Next.js / Node artifacts under apps/dreamverse/web/
apps/dreamverse/web/node_modules/
@@ -117,16 +118,17 @@ apps/dreamverse/web/.env.production.local
!apps/dreamverse/web/prompts/**/*.jpeg
!apps/dreamverse/web/prompts/**/*.mp4
!apps/dreamverse/web/prompts/**/*.gif
!apps/dreamverse/server/prompts/**/*.png
!apps/dreamverse/server/prompts/**/*.jpg
!apps/dreamverse/server/prompts/**/*.jpeg
!apps/dreamverse/server/prompts/**/*.mp4
!apps/dreamverse/server/prompts/**/*.gif
!apps/dreamverse/gpu-pool.svg
!apps/dreamverse/gpu-pool.drawio
.claude/
.codex/
.agents/tmp/
.sisyphus/
openspec/
fastvideo/tests/ssim/reference_videos/**
# Editor logs and local Python version pins (accidentally committed)
*.nvimlog
.nvimlog
.python-version
-1
View File
@@ -1 +0,0 @@
WRN 2026-03-26T13:46:33.469 ?.19646 server_start:193: Failed to start server: operation not permitted: /var/folders/z_/h_6myyk14d1b7z87z3vy4mjh0000gn/T/nvim.dsynkd/iSe0el/nvim.19646.0
+2
View File
@@ -10,6 +10,8 @@ exclude: |
scripts/.*|
fastvideo/dataset/.*|
fastvideo/models/.*|
v2/(layers|attention|platforms|configs|distributed|models|logging_utils|third_party|hooks|api)/.*|
v2/(envs|logger|utils|version|forward_context|fastvideo_args)\.py|
^apps/dreamverse/web/.*|
examples/.*|
\.agents/.*|
-1
View File
@@ -1 +0,0 @@
3.12
+11 -10
View File
@@ -11,7 +11,7 @@
- Static assets: `assets/` (including `assets/images/`, `assets/videos/`, and `assets/prompts/`) and `comfyui/assets/`.
## Build, Test, and Development Commands
- `uv pip install -e ".[dev]"`: editable install with lint/test extras.
- `UV_TORCH_BACKEND=cu126 uv pip install -e ".[dev]"`: editable CUDA 12 install with lint/test extras (`cu130` on CUDA 13).
- `pre-commit install --hook-type pre-commit --hook-type commit-msg`: enable local hooks.
- `pre-commit run --all-files`: run formatter/lint/type/spelling checks.
- `pytest tests/`: run top-level test suite.
@@ -44,17 +44,18 @@
## Agent Infrastructure
This repository is agent-friendly. Before doing any work, read:
This repository is agent-friendly. Before doing any work:
1. `.agents/onboarding/README.md` — full onboarding guide with step-by-step instructions.
2. `.agents/memory/codebase-map/README.md` — structural index of the entire repository.
3. `.agents/skills/` — available agent skills (check if one exists before writing code).
4. `.agents/workflows/` — SOPs for common procedures (experiment lifecycle, evaluation, etc.).
5. `.agents/lessons/` — known pitfalls and their documented fixes.
1. Read the nearest in-scope `AGENTS.md` for every directory you may edit.
2. Read the relevant user-facing design or contributor guide.
3. Check `.agents/skills/*/SKILL.md` for a task-specific workflow.
4. Search `.agents/lessons/` for relevant pitfalls.
If you are exploring a new procedure that has no existing SOP, document your
progress in `.agents/exploration/` and flag it for review at the end of your
session.
Architecture, commands, and operational guidance belong beside the code or
under `docs/`. Do not maintain static codebase maps, experiment journals,
branch-state snapshots, or duplicate user-facing documentation under
`.agents/`. Put a reusable procedure directly in a skill or contributor guide
and capture durable failures in `.agents/lessons/`.
## Per-Directory AGENTS.md
+30 -4
View File
@@ -9,8 +9,9 @@
**FastVideo is a unified post-training and real-time inference framework for accelerated video generation.**
## NEWS
- `2026/03/17`: Release Live demo: [Into the Dreamverse: Vibe Directing in FastVideo](https://dreamverse.fastvideo.org/), check out the [Blog](https://haoailab.com/blogs/dreamverse/).
- `2026/03/13`: Release Live demo: [Create a 5s 1080p Video in 4.5s with FastVideo on a Single GPU](https://1080p.fastvideo.org/), check out the [Blog](https://haoailab.com/blogs/fastvideo_realtime_1080p/).
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), check out the [Blog](https://haoailab.com/blogs/fastwan-qad/).
- `2026/03/17`: Release demo: Into the Dreamverse: Vibe Directing in FastVideo, check out the [Blog](https://haoailab.com/blogs/dreamverse/).
- `2026/03/13`: Release demo: Create a 5s 1080p Video in 4.5s with FastVideo on a Single GPU, check out the [Blog](https://haoailab.com/blogs/fastvideo_realtime_1080p/).
- `2025/11/19`: Release [CausalWan2.2 I2V A14B Preview](https://huggingface.co/FastVideo/CausalWan2.2-I2V-A14B-Preview-Diffusers) models, [Blog](https://hao-ai-lab.github.io/blogs/fastvideo_causalwan_preview/) and [Inference Code!](https://github.com/hao-ai-lab/FastVideo/blob/main/examples/inference/basic/basic_self_forcing_causal_wan2_2_i2v.py).
- `2025/08/04`: Release [FastWan](https://hao-ai-lab.github.io/FastVideo/distillation/dmd) models and [Sparse-Distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/).
@@ -54,12 +55,37 @@ We recommend using [uv](https://docs.astral.sh/uv/) to create a clean environmen
uv venv --python 3.12 --seed
source .venv/bin/activate
# Install FastVideo
uv pip install fastvideo
# Install FastVideo on NVIDIA CUDA 12
UV_TORCH_BACKEND=cu126 uv pip install fastvideo
```
Use `UV_TORCH_BACKEND=cu130` on CUDA 13. Apple silicon users should follow the
[MPS installation guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
Please see our [docs](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/) for more detailed installation instructions.
> **On an NVIDIA DGX Spark (GB10 / ARM64 + CUDA 13)?** There's no prebuilt ARM wheel for the FastVideo CUDA kernel, so it's an editable from-source install (`UV_TORCH_BACKEND=cu130 uv pip install -e .`, which compiles that kernel for you) rather than `UV_TORCH_BACKEND=cu130 uv pip install fastvideo`. A compatible prebuilt ARM64 FlashAttention wheel is available separately. Follow the [DGX Spark install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/spark/).
### Install with an AI coding agent
FastVideo is a monorepo with rich agent guidance (see [`AGENTS.md`](AGENTS.md)). If you use Claude Code, Cursor, or another coding agent, paste the prompt below — it detects your platform and follows the matching guide:
```text
Install FastVideo (https://github.com/hao-ai-lab/FastVideo) into a fresh uv virtual environment.
1. Detect the platform: run `uname -m`, `nvidia-smi`, and `nvcc --version`.
2. Read and follow the matching install guide exactly (in this repo, or at
https://hao-ai-lab.github.io/FastVideo/getting_started/installation/):
- NVIDIA GPU, x86_64 -> docs/getting_started/installation/gpu.md
- NVIDIA DGX Spark / GB10, aarch64, CUDA 13 -> docs/getting_started/installation/spark.md
- Apple Silicon, macOS -> docs/getting_started/installation/mps.md
3. Use uv for every step. If a command fails, debug it and tell me what you changed.
4. Verify the result:
python -c "import fastvideo, torch; print('cuda', torch.cuda.is_available())"
fastvideo --help
5. Report which platform you detected and any deviations you had to make.
```
## Sparse Distillation
For our sparse distillation techniques, please see our [distillation docs](https://hao-ai-lab.github.io/FastVideo/distillation/dmd/) and check out our [blog](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/).
+6 -7
View File
@@ -1,13 +1,12 @@
# Dreamverse Agent Notes
## Repo-local skills
## Repo-local workflows
- `.agents/skills/bootstrap-fastvideo-private-fork/`: temporary setup skill for
cloning `git@github.com:hao-ai-lab/FastVideo-internal.git` at
`will/rebase-nbv` into `../FastVideo-internal`, then running
`uv sync --extra server`.
- Prefer the bundled script in that skill instead of inventing a new private
FastVideo bootstrap flow.
- Use `.agents/skills/dreamverse-deploy/` to redeploy or stop the local
backend/frontend stack on a chosen physical GPU.
- Use `apps/dreamverse/scripts/modal/README.md` for Modal deployments and
`apps/dreamverse/docker/README.md` for image builds. The local deploy skill
intentionally does not manage remote deployments.
## Repo layout
+7 -4
View File
@@ -15,7 +15,7 @@ pip install --upgrade pip
pip install uv
uv venv .venv --python 3.12
source .venv/bin/activate
uv pip install "fastvideo[dreamverse]"
UV_TORCH_BACKEND=cu126 uv pip install "fastvideo[dreamverse]"
```
### Method 2: From source
@@ -28,9 +28,11 @@ pip install --upgrade pip
pip install uv
uv venv .venv --python 3.12
source .venv/bin/activate
uv pip install -e ".[dreamverse]"
UV_TORCH_BACKEND=cu126 uv pip install -e ".[dreamverse]"
```
Use `UV_TORCH_BACKEND=cu130` instead on CUDA 13.
### Method 3: Using Docker
```bash
@@ -244,8 +246,9 @@ npm run e2e
`dreamverse-server` exits with an install hint
- install the Dreamverse extra with `uv pip install -e ".[dreamverse]"` from a
source checkout, or `uv pip install "fastvideo[dreamverse]"` from PyPI.
- install the Dreamverse extra with `UV_TORCH_BACKEND=cu126 uv pip install -e ".[dreamverse]"` from a
source checkout, or `UV_TORCH_BACKEND=cu126 uv pip install "fastvideo[dreamverse]"` from PyPI
(`cu130` on CUDA 13).
Prompt-provider environment variable errors
+21 -10
View File
@@ -1,13 +1,23 @@
# syntax=docker/dockerfile:1.7
ARG CUDA_TAG=12.9.1-cudnn-devel-ubuntu22.04
FROM nvidia/cuda:${CUDA_TAG}
# CUDA base. CUDA_VERSION/UBUNTU_VERSION feed the default tag (matches the
# unified docker/Dockerfile); override BUILD_BASE_IMAGE wholesale for a mirror.
ARG CUDA_VERSION=13.0.0
ARG UBUNTU_VERSION=22.04
ARG BUILD_BASE_IMAGE=nvidia/cuda:${CUDA_VERSION}-cudnn-devel-ubuntu${UBUNTU_VERSION}
FROM ${BUILD_BASE_IMAGE}
ARG BUILD_FASTVIDEO_KERNEL_FROM_SOURCE=0
ARG BUILD_DREAMVERSE_UI=0
ENV DEBIAN_FRONTEND=noninteractive \
PYTHONUNBUFFERED=1 \
UV_LINK_MODE=copy
UV_LINK_MODE=copy \
UV_CACHE_DIR=/opt/uv/cache
# pyproject no longer pins a PyTorch index; GPU-less build host -> pin explicitly
# (auto would fall back to CPU). Matches the base CUDA: cu130 (13.0) / cu126 (12.6).
ARG UV_TORCH_BACKEND=cu130
ENV UV_TORCH_BACKEND=${UV_TORCH_BACKEND}
SHELL ["/bin/bash", "-c"]
@@ -20,7 +30,8 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
&& update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 \
--slave /usr/bin/g++ g++ /usr/bin/g++-11
ENV CUDA_HOME=/usr/local/cuda-12.9
# Version-agnostic /usr/local/cuda symlink so this works for any CUDA_VERSION.
ENV CUDA_HOME=/usr/local/cuda
ENV PATH=/root/.local/bin:/opt/venv/bin:${CUDA_HOME}/bin:${PATH}
ENV LD_LIBRARY_PATH=${CUDA_HOME}/lib64:${LD_LIBRARY_PATH}
ENV VIRTUAL_ENV=/opt/venv
@@ -34,17 +45,17 @@ WORKDIR /opt/FastVideo
COPY . /opt/FastVideo
RUN source /opt/venv/bin/activate \
&& uv pip install --no-cache-dir "/opt/FastVideo[dreamverse]"
RUN --mount=type=cache,target=/opt/uv/cache \
source /opt/venv/bin/activate \
&& uv pip install "/opt/FastVideo[dreamverse]"
# Standard docker build does not expose GPUs, while fastvideo-kernel/build.sh
# detects the CUDA architecture with torch at build time. The FastVideo package
# install above brings in the pinned fastvideo-kernel package; rebuild from the
# copied source only on hosts configured for build-time GPU access.
RUN if [[ "${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE}" == "1" ]]; then \
sed -i 's/^git submodule update --init --recursive$/if [[ -d ..\/.git ]]; then git submodule update --init --recursive; fi/' \
/opt/FastVideo/fastvideo-kernel/build.sh \
&& source /opt/venv/bin/activate \
RUN --mount=type=cache,target=/opt/uv/cache \
if [[ "${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE}" == "1" ]]; then \
source /opt/venv/bin/activate \
&& cd /opt/FastVideo/fastvideo-kernel \
&& ./build.sh; \
else \
+20 -5
View File
@@ -27,12 +27,27 @@ BUILD_DREAMVERSE_UI=1 DREAMVERSE_IMAGE=<image-tag> apps/dreamverse/docker/docker
Prefer SHA-specific tags for deployable images; avoid `latest`.
The Dockerfile builds a CUDA 12.9.1 image, installs FastVideo from this
checkout with the `dreamverse` extra, installs the FA4
flash-attention fork, builds native FFmpeg, and installs FlashInfer for NVFP4
quantization.
The Dockerfile defaults to CUDA 13.0.0 with the cu130 PyTorch backend. CI also
builds a CUDA 12.6.3 / cu126 image. Select that local build explicitly with:
FastVideo's pinned `fastvideo-kernel==0.2.6` package is installed by default.
```bash
CUDA_VERSION=12.6.3 apps/dreamverse/docker/docker_build.sh
```
The helper derives `UV_TORCH_BACKEND=cu126` for CUDA 12.x and `cu130` for CUDA
13.x. Set `UV_TORCH_BACKEND` explicitly for a custom CUDA version. The legacy
complete-image-tag override remains supported and selects the same matching
backend:
```bash
CUDA_TAG=12.6.3-cudnn-devel-ubuntu22.04 apps/dreamverse/docker/docker_build.sh
```
Do not set `CUDA_TAG` and `CUDA_VERSION` together. The image installs FastVideo
from this checkout with the `dreamverse` extra, including FA4 flash-attention
and FlashInfer for NVFP4 quantization, and builds native FFmpeg.
FastVideo's pinned `fastvideo-kernel==0.3.2` package is installed by default.
To rebuild `fastvideo-kernel` from this checkout during the image build, set:
```bash
+41 -1
View File
@@ -21,7 +21,47 @@ if [[ -z "${DOCKER_BUILDKIT:-}" ]] && docker buildx version >/dev/null 2>&1; the
fi
build_args=()
[[ -n "${CUDA_TAG:-}" ]] && build_args+=(--build-arg "CUDA_TAG=${CUDA_TAG}")
cuda_version="${CUDA_VERSION:-}"
torch_backend="${UV_TORCH_BACKEND:-}"
# CUDA_TAG was the Dockerfile's original override and contains the complete
# nvidia/cuda tag (for example, 12.6.3-cudnn-devel-ubuntu22.04). Keep accepting
# it while translating it to the parameterized Dockerfile inputs.
if [[ -n "${CUDA_TAG:-}" ]]; then
if [[ -n "${CUDA_VERSION:-}" ]]; then
printf 'CUDA_TAG and CUDA_VERSION cannot both be set. Use CUDA_VERSION for new builds.\n' >&2
exit 2
fi
cuda_version="${CUDA_TAG%%-*}"
if [[ ! "${cuda_version}" =~ ^[0-9]+\.[0-9]+(\.[0-9]+)?$ ]]; then
printf 'Cannot infer CUDA_VERSION from CUDA_TAG=%s. Use CUDA_VERSION and UV_TORCH_BACKEND instead.\n' \
"${CUDA_TAG}" >&2
exit 2
fi
build_args+=(--build-arg "BUILD_BASE_IMAGE=nvidia/cuda:${CUDA_TAG}")
fi
if [[ -n "${cuda_version}" ]]; then
build_args+=(--build-arg "CUDA_VERSION=${cuda_version}")
fi
if [[ -z "${torch_backend}" && -n "${cuda_version}" ]]; then
case "${cuda_version}" in
12.*) torch_backend=cu126 ;;
13.*) torch_backend=cu130 ;;
*)
printf 'No default UV_TORCH_BACKEND for CUDA_VERSION=%s. Set UV_TORCH_BACKEND explicitly.\n' \
"${cuda_version}" >&2
exit 2
;;
esac
fi
if [[ -n "${torch_backend}" ]]; then
build_args+=(--build-arg "UV_TORCH_BACKEND=${torch_backend}")
fi
[[ -n "${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE:-}" ]] && \
build_args+=(--build-arg "BUILD_FASTVIDEO_KERNEL_FROM_SOURCE=${BUILD_FASTVIDEO_KERNEL_FROM_SOURCE}")
build_args+=(--build-arg "BUILD_DREAMVERSE_UI=${BUILD_DREAMVERSE_UI:-0}")
@@ -9,17 +9,17 @@ ratio (5.04s of generated video produced in <=5.04s wall-time).
Usage::
python -m apps.dreamverse.server.benchmarks.benchmark_av_streaming
python -m apps.dreamverse.server.benchmarks.benchmark_av_streaming \\
python -m dreamverse.benchmarks.benchmark_av_streaming
python -m dreamverse.benchmarks.benchmark_av_streaming \\
--frames 121 --width 1920 --height 1088 --runs 3 \\
--codecs libx264 h264_nvenc --x264-preset ultrafast \\
--nvenc-preset p1
FASTVIDEO_FFMPEG_BIN=$HOME/opt/ffmpeg-native/bin/ffmpeg \\
python -m apps.dreamverse.server.benchmarks.benchmark_av_streaming
python -m dreamverse.benchmarks.benchmark_av_streaming
Skips ``h264_nvenc`` automatically if the binary lacks the encoder.
This is the regression guard documented in
`.agents/memory/dreamverse-integration/decisions-log.md` D-21.
Skips ``h264_nvenc`` automatically if the binary lacks the encoder. This is
the regression guard for software-encoding overhead that can drain the
inter-segment playback buffer.
"""
from __future__ import annotations
@@ -1,6 +1,6 @@
"""Benchmark the LTX-2 generation pipeline driven by the dreamverse Python SDK path.
Mirrors how ``apps/dreamverse/server/video_generation.py`` constructs
Mirrors how ``apps/dreamverse/dreamverse/video_generation.py`` constructs
``GeneratorConfig`` and calls ``VideoGenerator.generate()``, then
captures per-stage timings via the ``FASTVIDEO_STAGE_LOGGING=1`` log
hooks (same mechanism as ``FastVideo-internal/examples/inference/basic/
@@ -25,12 +25,12 @@ AV streaming path (use ``benchmark_av_streaming.py`` for that).
Usage::
python -m apps.dreamverse.server.benchmarks.benchmark_pipeline
python -m apps.dreamverse.server.benchmarks.benchmark_pipeline \\
python -m dreamverse.benchmarks.benchmark_pipeline
python -m dreamverse.benchmarks.benchmark_pipeline \\
--runs 3 --scenarios compile_warm cold --gpu 4
Cross-references D-21 / D-22 in
``.agents/memory/dreamverse-integration/decisions-log.md``.
This benchmark is the source of truth for the pipeline timing breakdown used
when investigating inter-segment buffer drain.
"""
from __future__ import annotations
+1 -1
View File
@@ -14,7 +14,7 @@ dependencies = [
[project.optional-dependencies]
server = [
"cerebras-cloud-sdk",
"flash-attn-cute @ git+https://github.com/XOR-op/flash-attention.git@fa4-compile#subdirectory=flash_attn/cute",
"flash-attn-4 @ git+https://github.com/Dao-AILab/flash-attention.git@940cd9680f3315f2f06b43ab5bea2c2cf2d96806#subdirectory=flash_attn/cute",
"flashinfer-python",
"openai>=1.40",
]
@@ -302,4 +302,4 @@ echo "[install_native_ffmpeg] ✓ done."
echo "[install_native_ffmpeg] binary: $ffmpeg_bin"
echo "[install_native_ffmpeg] env: $env_file"
echo "[install_native_ffmpeg] source it before running the demo:"
echo "[install_native_ffmpeg] source scripts/ffmpeg-env.sh"
echo "[install_native_ffmpeg] source apps/dreamverse/scripts/ffmpeg-env.sh"
+4 -2
View File
@@ -13,12 +13,14 @@ apps/dreamverse/scripts/launch/launch_demo.sh
The launcher starts the Dreamverse backend and frontend, polls readiness, and
prints the active URLs. It defaults to `dreamverse-server` on backend port
`8009` and frontend port `5274`.
`8009` and its legacy frontend setting is `5274`. The current web package
scripts bind `5299`, so pass that port explicitly until the launcher default is
updated in a separate behavioral change.
Useful overrides:
```bash
BE_PORT=8010 FE_PORT=5274 apps/dreamverse/scripts/launch/launch_demo.sh
BE_PORT=8010 FE_PORT=5299 apps/dreamverse/scripts/launch/launch_demo.sh
NO_FRONTEND=1 apps/dreamverse/scripts/launch/launch_demo.sh
NO_BROWSER=1 apps/dreamverse/scripts/launch/launch_demo.sh
```
@@ -7,12 +7,13 @@
# Defaults match internal/ui:
# * BE = dreamverse-server (8009) — full FE compatibility
# * FE = Next.js dev:devtools (5274)
# The current web package scripts bind 5299; pass FE_PORT=5299 explicitly.
#
# Switch BE to fastvideo serve --config (typed-only path):
# BE_FLAVOR=fastvideo bash launch_demo.sh
#
# Other env knobs:
# BE_PORT=8010 FE_PORT=5274 bash launch_demo.sh
# BE_PORT=8010 FE_PORT=5299 bash launch_demo.sh
# NO_FRONTEND=1 bash launch_demo.sh # backend only
# NO_BROWSER=1 bash launch_demo.sh # skip xdg-open
@@ -1,5 +1,5 @@
#!/usr/bin/env bash
# Launch the Dreamverse Next.js frontend in dev mode on port 5274
# Launch the Dreamverse Next.js frontend in dev mode on port 5299
# (the devtools-enabled build the e2e tests target).
#
# Usage:
@@ -40,7 +40,7 @@ fi
case "${FRONTEND_MODE}" in
devtools)
echo "[launch-demo] starting Next.js dev:devtools (port 5274)"
echo "[launch-demo] starting Next.js dev:devtools (port 5299)"
exec npm run dev:devtools -- "$@"
;;
dev)
@@ -48,7 +48,7 @@ case "${FRONTEND_MODE}" in
exec npm run dev -- "$@"
;;
single5s)
echo "[launch-demo] starting Next.js dev:single5s (port 5274)"
echo "[launch-demo] starting Next.js dev:single5s (port 5299)"
exec npm run dev:single5s -- "$@"
;;
esac
+11 -3
View File
@@ -36,9 +36,17 @@ DREAMVERSE_IMAGE=ghcr.io/<org>/<repo>/dreamverse:<tag> \
modal deploy apps/dreamverse/scripts/modal/modal_app.py
```
Use a SHA-specific tag, not `latest`. Use a `dreamverse-backend-cuda12.9.1-sha-*`
tag for backend-only deploys, or a `dreamverse-ui-cuda12.9.1-sha-*` tag for an
image that includes the static UI served by the backend.
Use a SHA-specific tag, not `latest`. The image workflow publishes backend and
UI variants for CUDA 12.6.3 and CUDA 13.0.0:
- `dreamverse-backend-cuda13.0.0-sha-*` or
`dreamverse-ui-cuda13.0.0-sha-*` for the CUDA 13 / cu130 images
- `dreamverse-backend-cuda12.6.3-sha-*` or
`dreamverse-ui-cuda12.6.3-sha-*` for the CUDA 12 / cu126 images
The Dockerfile defaults to the CUDA 13 / cu130 lane, which is the recommended
image for this B200 deployment. Choose the `ui` variant only when the backend
should serve the static UI.
For local image build details, see
`apps/dreamverse/docker/README.md` and `apps/dreamverse/docker/docker_build.sh`.
+3 -2
View File
@@ -10,8 +10,9 @@ IMAGE = os.environ.get("DREAMVERSE_IMAGE")
if not IMAGE:
raise RuntimeError(
"DREAMVERSE_IMAGE is required. Set it to a published SHA-specific Dreamverse image, "
"for example a dreamverse-backend-cuda12.9.1-sha-* tag or a "
"dreamverse-ui-cuda12.9.1-sha-* tag if serving the static UI."
"for example a dreamverse-backend-cuda13.0.0-sha-* tag or a "
"dreamverse-ui-cuda13.0.0-sha-* tag if serving the static UI. "
"CUDA 12 / cu126 images use the corresponding cuda12.6.3 tag."
)
# ``@modal.web_server`` invokes ``serve()`` directly and bypasses the image
@@ -4,18 +4,21 @@ import { test, expect, type WebSocket as PWWebSocket } from '@playwright/test';
* Long-running e2e: drive a real two-segment session end-to-end with
* torch.compile + warmup ENABLED on the backend, and assert audio
* conditioning carries from segment 1 → segment 2 without the
* BrokenPipe regression documented in
* `.agents/memory/dreamverse-integration/decisions-log.md` D-20.
* BrokenPipe regression previously caused by dropped LTX-2 audio continuation
* kwargs.
*
* Skipped by default. Opt in with PLAYWRIGHT_LONG_RUNNING=1, and
* boot the backend with both knobs on:
*
* # Terminal 1: backend + frontend
* ./.agents/skills/dreamverse-deploy/scripts/dreamverse-deploy.sh \
* --warmup --torch-compile 4
* --warmup --torch-compile 4 8009 5299
* # Terminal 2: test
* cd apps/dreamverse/web
* PLAYWRIGHT_SKIP_WEBSERVER=1 \
* BACKEND_HOST=127.0.0.1 \
* BACKEND_PORT=8009 \
* PLAYWRIGHT_BASE_URL=http://127.0.0.1:5274 \
* PLAYWRIGHT_BASE_URL=http://127.0.0.1:5299 \
* NEXT_PUBLIC_INCLUDE_DEVTOOLS=1 \
* PLAYWRIGHT_LONG_RUNNING=1 \
* npm exec -- playwright test e2e/long-running-segments.spec.ts
@@ -1,4 +1,4 @@
# FastVideo UI
# FastVideo Studio
A lightweight web-based UI for interacting with FastVideo.
@@ -16,7 +16,7 @@ The UI currently supports:
The easiest way to run the application is:
```bash
cd ui
cd apps/fastvideo_studio
npm install
npm run build
npm run start
@@ -33,10 +33,10 @@ The UI is composed of two separate components:
To run each component separately, you can use the commands `npm run start:web` and `npm run start:api`.
You can also run the API server with this command from the root directory:
You can also run the API server with this command from the `apps/` directory:
```bash
python -m ui.server --output-dir /path/to/videos --log-dir /path/to/logs
python -m fastvideo_studio.server --output-dir /path/to/videos --log-dir /path/to/logs
```
Running it this way allows you to pass command line parameters.
+7
View File
@@ -0,0 +1,7 @@
# SPDX-License-Identifier: Apache-2.0
"""Allow ``python -m fastvideo_studio`` to start the server."""
from fastvideo_studio.server import main
if __name__ == "__main__":
main()
@@ -1,6 +1,6 @@
# SPDX-License-Identifier: Apache-2.0
"""
SQLite persistence for FastVideo UI.
SQLite persistence for FastVideo Studio.
Stores jobs and default settings. Uses Python's built-in sqlite3.
"""
@@ -14,7 +14,7 @@ import threading
from pathlib import Path
from typing import Any
logger = logging.getLogger("fastvideo.ui.database")
logger = logging.getLogger("fastvideo.studio.database")
# Default options schema - used for settings table defaults
DEFAULT_SETTINGS: dict[str, Any] = {
@@ -24,14 +24,14 @@ from dataclasses import dataclass, field
from pathlib import Path
from typing import Any
from fastvideo.utils import get_mp_context
from ui.database import Database
from ui.training_config import (
from fastvideo_studio.database import Database
from fastvideo_studio.training_config import (
build_training_args,
get_training_env,
get_training_module_info,
)
logger = logging.getLogger("fastvideo.ui.job_runner")
logger = logging.getLogger("fastvideo.studio.job_runner")
# Regex patterns for parsing tqdm-style progress output.
# Matches e.g. " 40%|████ | 20/50 " or " 20/50 "
+14
View File
@@ -0,0 +1,14 @@
# SPDX-License-Identifier: Apache-2.0
"""Pydantic request/response models for the API."""
from fastvideo_studio.models.create_job_request import CreateJobRequest
from fastvideo_studio.models.settings_update import SettingsUpdate
from fastvideo_studio.models.create_dataset_request import CreateDatasetRequest
from fastvideo_studio.models.update_caption_request import UpdateCaptionRequest
__all__ = [
"CreateJobRequest",
"SettingsUpdate",
"CreateDatasetRequest",
"UpdateCaptionRequest",
]
@@ -1,11 +1,11 @@
{
"name": "fastvideo-ui",
"name": "fastvideo-studio",
"version": "0.1.0",
"lockfileVersion": 3,
"requires": true,
"packages": {
"": {
"name": "fastvideo-ui",
"name": "fastvideo-studio",
"version": "0.1.0",
"dependencies": {
"@sveltejs/kit": "2.57.1",
@@ -1,5 +1,5 @@
{
"name": "fastvideo-ui",
"name": "fastvideo-studio",
"version": "0.1.0",
"private": true,
"type": "module",
@@ -8,7 +8,7 @@
"build": "vite build",
"preview": "vite preview",
"start": "concurrently --kill-others-on-fail \"npm:start:api\" \"npm:start:web\"",
"start:api": "cd .. && python -m ui.server",
"start:api": "cd .. && python -m fastvideo_studio.server",
"start:web": "node build",
"lint": "eslint ."
},
@@ -13,7 +13,7 @@ import threading
from pathlib import Path
from collections.abc import Callable
logger = logging.getLogger("fastvideo.ui.preprocess_runner")
logger = logging.getLogger("fastvideo.studio.preprocess_runner")
def build_preprocess_args(

Before

Width:  |  Height:  |  Size: 119 KiB

After

Width:  |  Height:  |  Size: 119 KiB

Some files were not shown because too many files have changed in this diff Show More