Compare commits

...
Author SHA1 Message Date
SolitaryThinkerandClaude Opus 5.5 8a3f8c4838 [misc]: re-run pre-commit after the 2026-10-07 merges
Today's merges (33d81730..c520878f) left main failing the CI pre-commit
job, while start-of-day main 33d81730 passed every hook. This runs
`pre-commit run --all-files --hook-stage manual` again (Python 3.12, as
in CI) and fixes what it reports. Behaviour is unchanged.

yapf (the hook's in-place reformatting):
- apps/dreamverse/dreamverse/config.py
- apps/dreamverse/dreamverse/mock_server.py
- fastvideo/distributed/device_communicators/ulysses_a2a.py
- fastvideo/entrypoints/video_generator.py
- fastvideo/envs.py
- fastvideo/pipelines/stages/denoising.py
- fastvideo/registry.py
- fastvideo/train/models/hunyuan15/hunyuan15.py
- fastvideo/train/models/ltx2/ltx2.py
- fastvideo/train/utils/dataloader.py
- fastvideo/training/cosmos2_5_training_pipeline.py
- fastvideo/training/ltx2_training_pipeline.py
- fastvideo/training/training_utils.py
- fastvideo/worker/ray_distributed_executor.py
In fastvideo/registry.py the Cosmos 2.5 DFD detector lambda becomes a
small nested _cosmos25_dfd_detector(path) with the same logic, which
avoids yapf's deep hanging layout.

ruff (SIM108; not auto-fixed because of the comment):
- fastvideo/train/utils/performance.py: the if/else that sets
  max_frames is now a conditional expression, with the comment kept
  above it.

codespell (flags "THW" as a misspelling of THE/THAW):
- fastvideo/pipelines/stages/denoising.py: the two Cosmos25 RoPE
  comments now say "(T*H*W, D)".

pymarkdown, actionlint, mypy and check-filenames already passed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 12:40:35 +00:00
Zhang PeiyuanandClaude Opus 5.5 c520878fd2 [perf]: QAT/QAD add safe regional compile for modular training (#1718)
Adds opt-in regional torch.compile to the YAML-driven modular Wan trainer (training.model.enable_torch_compile / torch_compile_kwargs) with FA4 and Wan modulation forward/backward as opaque custom ops, compiled repeated DiT blocks after FSDP setup, and activation checkpointing applied before FSDP for trainable Wan transformers.

Merged with main's #1716 (WanModel resolves the checkpointing type in __init__; the PR's pre-FSDP wrap uses it and the post-load wrap is gone) and #1757 (pre_fsdp_model_transform runs before this PR's pre_fsdp_transform). Maintainer review (Swarm order #776): AdaLN skip names are canonicalized through checkpoint-wrapper prefixes, test_inference_regional_compile.py joins the unit lane, the Wan modulation custom-op test moved to the transformer lane, regional_compile is classified in the schema inventory, and a 1/2-GPU training smoke covers fully_shard(CheckpointWrapper) in eager and regional-compile modes. Validated on GB200: the FA4 CuTe forward/backward parity and fullgraph-backward tests and all four smoke cases pass.

Also exempts MMAudio recipes, whose student builds its pipeline config at init, from the fine-tuning recipe construction check that #1648 left failing on main.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 05:24:24 -07:00
99f8f3e7a7 [feat] Add streaming and GPU-accelerated LoRA extraction (#1784)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 04:48:35 -07:00
Raghav KandClaude Opus 5.5 14e4b7189b [feat] Add Cosmos Predict2.5 DFD continuation to DreamVerse (#1768)
Adds a DreamVerse generation backend for Cosmos Predict 2.5 DFD video-to-world continuation (cosmos25_dfd_generation.py), its config wiring and worker hook, README notes, and tests.

This PR stacked on #1767, which landed first with the review fixes for the Cosmos 2.5 pipelines; what lands here is the DreamVerse integration only. The fixes requested in this PR's maintainer review (Swarm order #748: shared fps, FloatingPointError on non-finite rollouts, DFD detector anchoring, 24 fps defaults, FP64 distilled rollout) all reached main through #1767.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 04:46:54 -07:00
8530fdc550 [feat] Add external-launcher executor for multi-node inference (#1746)
Co-authored-by: Will Lin <willlin@nvidia.com>
Co-authored-by: William Lin <8941107+SolitaryThinker@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 04:46:26 -07:00
Raghav KandClaude Opus 5.5 315b25d26d [new-model] Add Cosmos Predict2.5 distilled T2W and DFD V2W inference (#1767)
Adds the Cosmos Predict 2.5 distilled text-to-world sampler (sCM, per-frame timesteps) and the DFD video-to-world pipeline, with their schedulers, registry entries, presets and examples.

Maintainer review fixes (Swarm orders #718 and #748): the DiT receives one shared fps value (its RoPE table is batch-independent, so per-sample vectors crashed batch > 1; differing values now raise), non-finite rollouts raise FloatingPointError instead of being zeroed by nan_to_num, the DFD detector matches "dfd" only in the last path component, the distilled example defaults to 24 fps, and the distilled rollout keeps its state in FP64 through the scheduler step as the official driver does.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 04:42:44 -07:00
b8920ba50a [misc]: add env var to control process-aware logging (#1680)
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-10-07 04:31:52 -07:00
3467f8befb [feat] Add end-to-end MMAudio V2A training, inference, and evaluation (#1648)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 04:25:00 -07:00
Kun LinandClaude Opus 5.5 7320d9d350 [feat] Add HunyuanVideo 1.5 embedding preprocessing pipeline (#1663)
Adds preprocess_hunyuan15_overfit.py, which writes VAE latents plus Qwen and ByT5 text embeddings in the dual-text parquet schema (pyarrow_schema_t2v_dual_text) that main's HunyuanVideo 1.5 training (#1662) reads; main's guide and overfit config pointed at this script, but nothing in-tree produced text_embedding_2 before.

Reconciled with #1662 during maintainer review (Swarm orders #723 and #777): main's schema and collate are the base. On top of them, both text streams now share one CFG-dropout decision per row also when _sample_index is absent, and a row missing the primary text_embedding raises instead of training unconditioned. The PR's widening of the shared t2v schema and its writer changes were dropped. docs/training/hunyuan15_overfit.md now describes the in-tree script.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 04:20:58 -07:00
af78c759f1 [new-model] Add Wan2.2-Animate-14B character animation/replacement inference (#1765)
Co-authored-by: SuhaanCommits <suhaan@cedar.build>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-10-07 04:19:37 -07:00
f82c190cbb [kernel] Build + allow attn_qat_infer FP4 attention on sm_121a (DGX Spark) (#1598)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 04:09:39 -07:00
e722023bd6 [perf] Wan kernel fusion: Triton-fused residual+LayerNorm+modulate inference path (#1708)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 04:08:07 -07:00
bc0191d775 [perf] Harden Ulysses agreement and tune bounded H3 transfers (#1807)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 04:01:53 -07:00
2d6e7d9b6b [feat]: MiniMax-H3 encoder split: dedicated text-encoder node group for 720p on DGX Spark (#1877)
Co-authored-by: magicbear <magicbear@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:46:33 -07:00
Kevin LinandClaude Opus 5.5 5e61accc37 [misc] Clean up QAD 5090 example inference scripts (#1496)
Split the FastWan-QAD Wan2.1-1.3B examples into per-recipe scripts (fp8, nvfp4_qat, nvfp4_sa2) with a shared _qad_common.py and a standalone multi-run benchmark harness, replacing FastWan_QAD_TAEHV.py.

The PR's fastvideo-kernel build changes (sm_121a arch probe and build.sh mapping) were dropped during maintainer review (Swarm order #664) in favour of #1598, which owns the kernel arch policy; the RTX 5090 (sm_120a) recipes need nothing beyond main's kernels.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:44:49 -07:00
f60514183c [ci] Separate performance behavior tests (#1620)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:43:12 -07:00
William LinandClaude Opus 5.5 156611bab3 [ci]: run the full attention test directory in the unit lane (#1658)
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:37:46 -07:00
William LinandClaude Opus 5.5 6d6195a9f1 [feat]: fp8 PV mode for the FA4-FP4 attention path (#1654)
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:36:54 -07:00
8348e83e80 [feat] Import MiniMax H3 Comfy NVFP4-AWQ text encoder (#1865)
Co-authored-by: 武垚乐 <wuyaole@mininglamp.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:34:40 -07:00
2452022f15 [new-model] Add LTX-2.5 inference support (#1704)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: CodeRabbit <noreply@coderabbit.ai>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:34:18 -07:00
4b3f99b224 [new-model] Added LingBot-World-Fast image-to-video support (#1665)
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-10-07 03:32:38 -07:00
William LinandClaude Opus 5.5 d967b99928 [bugfix]: registry ambiguity warning by specificity + QAD weights file path (#1649)
Registry: when several model detectors match a path, the first registered match still wins (main's pinned precedence, e.g. test_local_manifest_detectors_keep_first_match). Path specificity against each entry's registered HF paths now only decides whether to warn: the resolution is silent when the first match is strictly more specific than every other match (e.g. LTX-2.3 distilled directories that the LTX-2 base detector also claims through the shared pipeline class name), and warns otherwise, including when a later match is more specific than the one chosen. The PR originally re-ranked matches by specificity; that changed main's tested first-match resolution for local manifest directories, so it was narrowed during maintainer review (Swarm order #720).

Also fixes the QAD weights file path.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:30:55 -07:00
8f1eb2992c [bugfix]: retain NVFP4 weights for LTX2 refinement LoRA (#1705)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:29:46 -07:00
pkisfaludi-nvandClaude Opus 5.5 683689de4f [bugfix] Fix Ulysses sequence-parallel crash in self-forcing causal Wan DiT (#1627)
Retargeted onto fastvideo/models/wan/causal_transformer.py, where main moved the causal Wan implementation (fastvideo/models/dits/causal_wanvideo.py is now a compatibility shim), with the guarded flex-attention trim prepared during the maintainer review (Swarm order #657).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:29:22 -07:00
d5287f8352 [feat]: add MiniMax H3 Ref2VA and LoRA training support (#1757)
Co-authored-by: William Lin <8941107+SolitaryThinker@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:21:45 -07:00
cdfbd64b04 [perf] FP8: fuse the dynamic activation quantization into Triton kernels (#1861)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:20:04 -07:00
Haochen JiangandClaude Opus 5.5 299d0f6838 [bugfix] wrap the generic DenoisingStage in the inference_denoising profiler region (#1707)
profiler_region_inference_denoising was wrapped in exactly one place (minimax_h3_denoising.py), so every model using the generic DenoisingStage (Wan among them) exported a trace with no denoising region. Wrap the denoising loops in DenoisingStage and DmdDenoisingStage with profiler_region("inference_denoising"); the loop bodies are untouched.

These are the two classes STAGE_METRIC_MAP names directly. The perf harness also reports dit_time_s for every stage that inherits performance_component_metric, and those forward overrides (Cosmos*, LongCat*, Gen3C, GameCraft, HYWorld, MatrixGame2/3, Causal*, DreamXWorldAR, LingBot, GlmImage) remain outside the region for now. With no profiler configured the wrap is a no-op; SSIM is exactly 1.0 on FLASH_ATTN and TORCH_SDPA.

Caveat: the region span now appears in exported traces, but a pre-existing profiler behaviour (region exit toggles CUDA collection off, fastvideo/profiler.py) may leave GPU kernel rows out of the trace.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:19:42 -07:00
c18436125a [misc]: Refactor separate training pipeline runner (#1696)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:19:21 -07:00
5d515ee617 [perf]: remove per-step host syncs and repeated work from causal inference (#1687)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:18:56 -07:00
37b9dc1836 [new-model] Add native Helios-Distilled T2V pipeline (#1670)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 03:18:24 -07:00
a0291a57c1 [ci] Add Wan2.2-TI2V-5B SSIM regression test (#1772)
Co-authored-by: SuhaanCommits <suhaan@cedar.build>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-10-07 02:44:24 -07:00
6102fac00d [bugfix]: fix modular ops activation checkpoint recomputation (#1716)
Co-authored-by: William Lin <8941107+SolitaryThinker@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 02:41:12 -07:00
1ad33c1832 [feat] Support training_cfg_rate > 0 for LTX-2 (#1752)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 02:38:42 -07:00
da87973412 [ci] Clarify dashboard cohort grouping for legacy v1 records (#1613)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 02:33:18 -07:00
William LinandClaude Opus 5.5 a3fd1d0bab [ci]: prevent direct tests from minting suite gates (#1615)
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 02:32:48 -07:00
3e87d6f4ae [feat] ROCm: route the video sparse attention backends to their Triton kernels (#1851)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 02:32:27 -07:00
4bb5627fa8 [feat] MiniMax H3: block-FP8 text encoder on Hopper, its converter, and a 4xH100 FastH3 profile (#1829)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 02:31:52 -07:00
08ab8b5c00 [bugfix] ROCm: detect the platform from a HIP torch build when amdsmi is unavailable (#1849)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 02:20:47 -07:00
fb5efebdab [feat] Add HunyuanVideo 1.5 T2V training support (#1662)
Co-authored-by: William Lin <willlin@nvidia.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 02:19:22 -07:00
6e4813f8b4 [bugfix] use in-process executor for num_gpus=1 (#1755) (#1879)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 02:18:23 -07:00
13ef81e6da [ci] Surface worker process logs on perf hard regressions (#1701)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-10-07 02:14:02 -07:00
13b3b36b8b [feat] Add Infinite Livestream app (#1878)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 02:07:37 -07:00
72795d0af8 [perf] avoid full-sequence materialization in VSA coarse/sparse combine (#1813)
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-10-07 01:54:35 -07:00
af63a92f6a [bugfix] Shut down Inductor workers gracefully (#1786)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 01:54:03 -07:00
a4e7322ba6 [bugfix]: synchronize post-FSDP LoRA replicas (#1666)
Co-authored-by: Suckl <Suckl@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 01:53:30 -07:00
d6e13f9cc8 [feat] Add VQeval long-video quality metric (#1727)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 01:39:11 -07:00
6c4c18690e [bugfix] Enable SP for LongCat BSA refinement (#1862)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 01:38:43 -07:00
a8688dddd9 [bugfix]: requantize MXFP8 weights after LoRA unmerge (#1876)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 01:38:02 -07:00
cacc8cfcb3 [bugfix]: condition HunyuanVideo 1.5 i2v on the reference image (#1693)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 01:36:35 -07:00
ef62ac48d1 [new-model] Add Wan2.2-S2V-14B speech-to-video (model port + checkpoint converter) (#1683)
Co-authored-by: SuhaanCommits <suhaan@cedar.build>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-10-07 01:36:01 -07:00
00cb7d0ba2 [ci] Add DGX Spark (GB10) single-GPU perf config + gpu_types gating (#1679)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 01:30:42 -07:00
69a6215e7f [feat] Add Dreamverse multimodal inputs and H3 routing (#1835)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 01:30:26 -07:00
abe98e87f5 [feat] Add phase-aware modular training performance metrics (#1826)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 01:30:11 -07:00
274b922d39 [misc]: surface GenerationResult.peak_memory_mb in the MiniMax H3 examples (#1785)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 01:29:56 -07:00
38966056d0 [ci]: make merge-gate startup reliable (#1762)
Co-authored-by: William Lin <8941107+SolitaryThinker@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 01:29:42 -07:00
d8ca702dfd [bugfix] registry: register magi_human presets (#1751)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
2026-10-07 01:28:53 -07:00
William Lin 33d81730e9 [feat]: Kandinsky 6 Lite: register the 3.2B checkpoints and add docs + cookbook recipes (#1927) 2026-10-06 17:09:53 -07:00
02a027c49f [feat]: FastH3 on DGX Spark and Apple Silicon (#1920)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: Aryan Kumar <aryank@Aryans-Mac-Studio.local>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 12:17:11 -07:00
Aryan Kumar e235b6a332 [docs] Announce FastH3 on consumer hardware and FastH3 Trim (#1923) 2026-10-06 11:39:24 -07:00
8b51380466 [feat]: FastH3 on single RTX GPUs (#1919)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: Aryan Kumar <aryank@Aryans-Mac-Studio.local>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 11:37:01 -07:00
William Lin 527a3165b9 [docs]: add Kandinsky 6 to the README news (#1925) 2026-10-06 11:05:40 -07:00
f06d824054 [feat]: add Kandinsky-6.0 TI2VA and VSR pipelines (#1924)
Co-authored-by: leffff <levnovitskiy@gmail.com>
Co-authored-by: Kirill <kykozlov@gmail.com>
Co-authored-by: shaoxiongduan <shaoxiongduan@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-06 10:50:17 -07:00
Kyle Hu 6ded84ee0f [ci]: improve CI efficiency by reordering lanes (#1914) 2026-10-04 12:35:12 -07:00
a4d6416c52 [docs] Add OpenAI serving recipes for Wan cookbook pages (#1906)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
Co-authored-by: Claude Sonnet 5.5 <noreply@anthropic.com>
2026-10-04 12:20:06 -07:00
lpc0220andClaude Opus 5.5 0cc41a22dc [bugfix] sm_100a VSA: fence the CLC response read before the slot release; forward PV descriptor lifetime (#1904)
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 12:37:34 -07:00
8444c0897a [feat] FastH3 Ref2VA PDD inference with reference-video VSA sparsity (#1907)
Co-authored-by: Davids048 <jundasu@ucsd.edu>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
2026-10-02 11:28:07 -07:00
Kyle Hu 9491c8638a [bugfix]: keep the full audio track when saving through the ffmpeg pipe (#1910) 2026-10-01 22:22:17 -07:00
Ishan 6809a751fb [bugfix] Fixed the MiniMax-H3 pin-fallback decode test setup (#1909) 2026-10-01 22:17:46 -07:00
Kyle Huandaryan5v 8322b01815 [bugfix]: restore the NVFP4 module after test_nvfp4_config re-imports it (#1908)
Co-authored-by: aryan5v <email.aryan1005@gmail.com>
2026-10-01 22:17:32 -07:00
Junda Su 9edc8adf5f [misc] Move test-suite environment access into the registry and allowlists (#1898) 2026-09-30 00:15:39 -07:00
e1f3904799 [kernel] sm_100a CUDA backward for VSA block-sparse attention, 128-token blocks (#1866)
Co-authored-by: Pengcheng Li <pengchengl@ptyche0203.ptyche.clusters.nvidia.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-29 16:57:22 -07:00
Junda Su 02f1ce11ae [misc] Move environment var reads into env registry and rename unprefixed variables (#1897) 2026-09-29 14:36:16 -07:00
Aryan KumarandAryan Kumar 7f03e03dc6 [docs]: FastH3 V2 cookbook serving and MLX install guide (#1884)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-29 13:03:18 -07:00
Junda Su e3b88bb12a [feat] Add a typed environment-variable registry, policy doc, and contract test (#1896) 2026-09-29 12:50:45 -07:00
Junda Su cb66acd400 [bugfix] Fix environment-variable bugs in LTX-2 debug logging, HF token lookup, and attention backend reads (#1895) 2026-09-29 11:48:31 -07:00
li-lizhe 442e2d2e18 [bugfix] fix(metrics): set SceneMetric device_map for any non-CPU device (#1817) 2026-09-28 10:24:00 -07:00
Max LI dd35763ad6 [bugfix] Fix MLX prompt enhancement sampler compatibility (#1891) 2026-09-27 14:53:25 -07:00
Keith e90be598e5 [perf] MiniMax H3: return uint8 frames from the decode worker (#1828) 2026-09-24 18:08:56 -07:00
Leleand武垚乐 ba5e81083c [bugfix] Preserve video frames when decoded audio is shorter (#1857)
Signed-off-by: 武垚乐 <wuyaole@mininglamp.com>
Co-authored-by: 武垚乐 <wuyaole@mininglamp.com>
2026-09-24 17:20:37 -07:00
YZJF 76ce9c7fd6 [bugfix] Reject mismatched model names on image generation routes (#1881) 2026-09-24 17:19:29 -07:00
Junda Su 08d99c089e [perf] Reuse cached VAE offload for H3 conditioning (#1882) 2026-09-24 17:18:39 -07:00
Jerry-XY 20751a21aa [bugfix] Keep the VSA-H3 tile buffer out of the autograd graph (#1858) 2026-09-24 17:16:49 -07:00
alanhuangyooandSolitaryThinker 9dd2a837f4 [bugfix] encoders: keep Qwen2.5-VL importable on transformers 5 (#1791)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-24 17:13:23 -07:00
alanhuangyoo 93aab45ac2 [bugfix] tests: include ltx2_3_base in local LTX-2 preset set (#1750) 2026-09-24 17:11:31 -07:00
Yogya Mehrotra 017ce6602d [bugfix] Fix dead NVFP4 QAT plumbing: bench imports, int8 error string, smooth_q crash (#1723) 2026-09-24 17:09:33 -07:00
Yogya Mehrotra a575055eec [perf] benchmark_attn_qat_train: device peak-TFLOPS table instead of silent RTX 5090 default (#1724) 2026-09-24 17:08:03 -07:00
Kyle Hu 81f3fec7fd [bugfix]: workers killed by a signal now log the reason (#1725) 2026-09-24 17:06:55 -07:00
Raghav K d265a454bf [bugfix]: Fall back when large pinned output allocation fails (#1759) 2026-09-24 17:05:59 -07:00
c100c66578 [bugfix]: isolate cancelled streaming requests from GPU result readers (#1848)
Co-authored-by: Gxj230958 <222823329+Gxj230958@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-24 17:04:59 -07:00
Aryan Kumar 8760eb7a06 [feat] Add FastH3 8-Step V2 MLX inference (#1863)
Keep the four-step uniform AdaLN cache, and use explicit DMD rungs only when the snapshot declares its own scheduler shifts.
2026-09-23 12:42:03 -07:00
Aryan Kumar 361f919c88 [feat]: add a narrow FastH3 8-step MLX ladder
Keep the four-step uniform AdaLN cache, and use explicit DMD rungs only when the snapshot declares its own scheduler shifts.
2026-09-23 11:06:55 -07:00
Aryan Kumar 8b5377aab2 Merge remote-tracking branch 'origin/main' into aryan/pr-1863-fasth3-8step 2026-09-23 10:52:20 -07:00
Yixing Wangandleo d995516da0 [bugfix] Restore page interaction after dismissing Create Job or successfully creating a job (#1869)
Co-authored-by: leo <yixingwang@YIXINGs-MacBook-Pro.local>
2026-09-21 12:03:54 -07:00
Keith f47ad3f5b7 [bugfix] fastvideo-kernel: install the Python + Triton package when the HIP toolchain cannot be configured (#1871) 2026-09-21 11:59:13 -07:00
Hexu ZhaoandClaude Opus 5 10bdf5e076 [bugfix] frozen offload: no host copy for an already resident module
`load` took the host copies before checking which tensors were missing from
the device, so with `vae_cpu_offload=False` — where the loader builds the VAEs
on CUDA and `unload` is never called — it pulled about 11 GB per rank back over
PCIe once and kept a pinned mirror of it for the life of the process, for a
module that never leaves the device.

It now asks what is absent first and returns early when nothing is. The
offloaded path is unchanged: weights still round-trip and come back bit
identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-20 22:46:13 -07:00
Hexu ZhaoandClaude Opus 5 3e26db40b0 [perf] H3 decode: pin the frozen VAEs' host copies
With the device-to-host copy gone, the remaining host-to-device copy is the
stage's cost, and from pageable memory it is staged through a bounce buffer at
roughly 2.6 GB/s instead of full PCIe speed. The host copies are made once, so
pinning them costs nothing per request.

It does cost the module's size in non-pageable host memory (10.4 GB for the
H3 video VAE, 0.6 GB for the audio VAE, per rank) for as long as the pipeline
is alive. It follows --pin-cpu-memory, which is on by default; --no-pin-cpu-
memory keeps the copies pageable and only loses the bandwidth. Note that this
flag did not already reach the VAE: its other use is FSDP's own CPUOffloadPolicy
and the VAE does not go through FSDP.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ATw7g6cxq9N2LtnMratsru
2026-09-20 22:46:13 -07:00
Hexu ZhaoandClaude Opus 5 c73dd0ab55 [perf] H3 decode: offload the frozen VAEs without the device-to-host copy
The decode stages move the 10.4 GB fp32 video VAE (and the 0.6 GB audio VAE)
with module.to(device) / module.to("cpu") on every request. The weights are
frozen -- the whole stage runs under no_grad -- so the copy back to the host
returns bytes the host already has.

Keep one host copy per tensor, made the first time the module is loaded, and
on unload point .data back at it instead of copying. One H2D per request, no
D2H at all; the module still lives on the host between requests. Where the old
path allocated a fresh host tensor per cycle and freed the previous one, this
reuses one buffer, so peak host memory can only go down.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01ATw7g6cxq9N2LtnMratsru
2026-09-20 22:46:13 -07:00
430e52154e [bugfix]: Restore exact Qwen3-VL vision interpolation (#1737)
Co-authored-by: William Lin <8941107+SolitaryThinker@users.noreply.github.com>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-18 05:00:05 -07:00
vanch 384eee8aef feat(mlx): add FastH3-8-Step-V2 DMD schedule and thermal-safe MLX inference
- Support official 8-step DMD schedule from fastvideo_inference.json in MLX pipeline
- Implement MiniMaxH3SchedulerState.from_dmd_steps for exact sigma ladder calculation
- Add auto-detection of fastvideo_inference.json in checkpoint converter AdaLN precomputation
- Add inter_step_cooldown_s and per-step MLX graph/cache cleanup to prevent thermal throttling
- Support configurable VAE tile sizes to guarantee 256px tiled decode without grid artifacts
- Skip non-existent shard files gracefully in Qwen3-VL conditioner
2026-09-17 16:12:02 +08:00
KeithandSolitaryThinker c4824c7764 [bugfix] FP8: gate the ROCm _scaled_mm path on CDNA4 (gfx950) instead of the capability tuple (#1859)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-16 19:58:55 -07:00
Raghav K 9b0e57fe4b [ci] Make Dreamverse provider race test deterministic (#1729) 2026-09-15 14:13:56 -07:00
William Lin 0100218594 [feat] Support the FastH3 8-Step V2 checkpoint: checkpoint-defined shifts, explicit DMD schedule, new example (#1852) 2026-09-15 14:13:45 -07:00
li-lizhe 39718cd54d [bugfix] fix(cosmos): make AdaLayerNorm autocast device-agnostic (#1818) 2026-09-15 07:44:51 -07:00
IshanandSolitaryThinker 37d06a832f [feat] Add fastvideo serve configs for Wan CUDA models (#1801)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-14 18:23:35 -07:00
sudhirpol522 8839ba8d4d [bugfix] Add OpenAI-compatible image generation endpoint (#1840) 2026-09-14 17:55:56 -07:00
sudhirpol522andSolitaryThinker 61b91220c0 [bugfix] Validate image response format before generation (#1841)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-09-14 17:01:40 -07:00
Lele 316f3876c2 [bugfix]: allow MiniMax H3 frame padding at the 15-second limit
Accept the causal-VAE-aligned 362-frame bucket (15.083 s) for 15-second H3 requests.

- Hoist MINIMAX_H3_MIN/MAX_ALIGNED_FRAMES next to align_num_frames and reuse them in the CUDA stage, the MLX runtime, and the LoRA example so the bound is consistent across entry points.
- State the accepted frame range in the error messages.
- Cover seconds="15" -> 362 at OpenAI admission and the MLX resolve_geometry bound.
- Update the Spark/openai cookbook frame caps from 345 to 362.

Known gaps: no GPU-lane coverage for the 362 bucket (SSIM runs 124 frames; the golden gate is a fixed-geometry fingerprint), and explicit num_frames=360 still hits the pre-existing grid gate.
2026-09-14 16:54:41 -07:00
Yaegaki1Erika 614b59543c [bugfix]: respect serialized tensor dtype in parquet dataloader (#1843) 2026-09-14 16:34:25 -07:00
William Lin 1c14afd559 [docs]: refresh AGENTS.md maps and add Wan SP/I2V and CI test notes (#1845) 2026-09-14 16:06:53 -07:00
Yaegaki1Erika 9a3c45779c [bugfix]: fix Wan I2V DMD conditioning under sequence parallelism (#1844)
The pipeline pre-sharded only the image conditioning along the temporal axis while the noise input stayed full-length, so concatenation could not match for sp_world_size > 1. WanTransformer3DModel shards the flattened token sequence after patch embedding, so pass full-length mask+latent conditioning and let the transformer shard.

Also reads temporal_compression_ratio from the VAE config, simplifies the conditioning mask, concatenates in the transformer's (bs, c, t, h, w) layout, and adds unit coverage.

Co-authored-by: Yaegaki1Erika <70182590+Yaegaki1Erika@users.noreply.github.com>
2026-09-14 15:58:13 -07:00
Kyle Hu bfc9c01797 [feat]: convert the MiniMax H3 text encoder to NVFP4 (#1838) 2026-09-12 18:48:07 -07:00
Kyle Hu 3a3ad3d209 [feat]: NVFP4 text encoder for MiniMax H3 (#1837) 2026-09-12 18:31:59 -07:00
lpc0220andClaude Fable 5.1 aef4e9b3b1 [kernel] VSA kernel: one sm_100a / sm_103a image per listed arch; un-gate backward on sm_103a (#1833)
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-12 16:49:36 -07:00
a943220c11 [bugfix]: drop dead h3_sequential_load from Spark FastH3 presets (#1831)
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-08 12:38:06 -07:00
William Lin 556ac7088e [refactor] Simplify Wan sampling and tests (#1825) 2026-09-07 15:57:19 -07:00
Junda Su e7456f1b75 Add H3 support into Dreamverse (#1800) 2026-09-07 13:09:44 -07:00
lpc0220 c993d7393e [kernel] sm_100a CUDA backward for VSA block-sparse attention (blk64) (#1819) 2026-09-06 21:00:18 -07:00
William Lin 7f83164233 [refactor] Move Wan VAE into the Wan package (#1824) 2026-09-05 18:43:55 -07:00
Junda Su 4e52f47d1e [feat] add MXFP8 support on H3 (#1796) 2026-09-05 17:08:58 -07:00
William Lin e19913f6e9 [refactor] Group Wan transformer and config (#1823) 2026-09-05 17:05:32 -07:00
William Lin 2413a57651 Disable old SSIM models (#1820) 2026-09-04 21:20:48 -07:00
SYLAR 7bb76b5ec9 [feat] Add native SM103a VSA support (#1812)
Signed-off-by: lishunyang12 <lishunyang12@163.com>
2026-09-03 15:00:28 -07:00
Aryan KumarandAryan Kumar 0bd19a976b [docs]: add one-Spark FastH3 cookbook runtime with a device-count row (#1811)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-02 10:05:59 -07:00
William Lin 40b93784d2 [docs]: add cookbook link to README (#1810) 2026-09-01 12:42:19 -07:00
Aryan Kumar 33d3478bad [docs] Announce local FastH3 support (#1809) 2026-09-01 12:28:04 -07:00
3d8ac9d14b [feat]: collapse cookbook recipe pages into an accordion layout, add … (#1805)
Co-authored-by: Vaish, Ishan <isvaish@UCSD.EDU>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-01 01:51:51 -07:00
aaef49bfc6 [feat]: run FastH3 across two DGX Sparks with Ray sequence parallel (#1803)
Co-authored-by: Kyle <shh075@ucsd.edu>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Satyam Srivastava <srivastavasatyam53@gmail.com>
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-01 01:37:23 -07:00
Aryan KumarandAryan Kumar cf6a00b9be [feat]: add opt-in CUDA TAEH3 preview decode for FastH3 (#1795)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 23:34:56 -07:00
Shahrad ZomorrodiandShahrad Zomorrodi 1ae39562dd [bugfix] Write generated documentation as UTF-8 (#1797)
Co-authored-by: Shahrad Zomorrodi <264690209+shahradzomorrodi@users.noreply.github.com>
2026-08-31 22:45:03 -07:00
Aryan KumarandAryan Kumar 26064193e2 [feat] Add an H3 server cookbook and prompt playground (#1798)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 22:14:33 -07:00
William Lin 8446fc003e [docs]: add FastH3 Preview v1 news links (#1804) 2026-08-31 22:12:02 -07:00
Aryan KumarandAryan Kumar a28f2bab4b [feat] Add an optional MLX TAEH3 preview decoder (#1794)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 05:08:06 -07:00
Aryan KumarandAryan Kumar f82d8be4bf [perf] Sequential MiniMax H3 start with GPU-direct DiT load (#1793)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 04:20:21 -07:00
Aryan KumarandAryan Kumar 8e1775183e [perf] Speed up exact MiniMax H3 MLX inference (#1792)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 04:14:55 -07:00
Aryan KumarandAryan Kumar 620bc36dc4 [docs]: Cookbook catalog improvements (#1790)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-30 17:50:57 -07:00
Suhaan Khurana 29ff16ec96 [feat] Add MiniMax H3 MLX spatial fast mode (#1789) 2026-08-30 17:47:02 -07:00
Aryan KumarandAryan Kumar 8f9d76a80d [perf]: dispatch wide-M affine H3 MLX linears through dequant plus dense GEMM (#1788)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-30 14:42:01 -07:00
a4d9a75e2c [perf] Add MiniMax H3 MLX VSA and SIMD attention (#1776)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-08-30 13:44:54 -07:00
KyleNeverGivesUp b2db0c0a13 [ci]: seed stable GB10 grad-norm references (#1756) 2026-08-30 03:09:18 -07:00
Kevin Lin 6aa7d8a278 [misc] FastVideo Studio UI Additions (H3 Ref2V support) (#1783) 2026-08-30 03:06:47 -07:00
Aryan KumarandAryan Kumar ccc9014430 [docs] Add model-family inference cookbook (#1787)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-30 02:26:10 -07:00
William Lin a159b63c67 [bugfix] Harden OpenAI serving after post-merge review (#1782) 2026-08-28 22:29:24 -07:00
ac48bb3cd1 [feat] Add MiniMax H3 MLX T2VA inference (#1770)
Co-authored-by: Aryan Kumar <aryank@Aryans-Mac-Studio.local>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: CodeRabbit <noreply@coderabbit.ai>
2026-08-28 12:54:33 -07:00
William Lin 3987b9ddcd [feat] Align multimodal OpenAI serving APIs (#1781) 2026-08-28 10:09:58 -07:00
William Lin c7da2f5d60 [chore]: release v0.2.1 (#1778) 2026-08-28 02:03:18 -07:00
William Lin 39ae1decc0 [misc] pin fastvideo-kernel to exact 0.3.5 (#1777) 2026-08-28 02:02:27 -07:00
William Lin 1aed667377 [chore] release fastvideo-kernel 0.3.5 (#1775) 2026-08-27 23:18:59 -07:00
William Lin c1612ff397 [bugfix]: pin fastvideo-kernel to Torch 2.12.0 (#1774) 2026-08-27 23:13:54 -07:00
William Linandshaoxiongduan a534ba20a0 [feat] Add MiniMax H3 LoRA inference and preview launchers (#1771)
Co-authored-by: shaoxiongduan <shaoxiongduan@gmail.com>
2026-08-27 14:43:39 -07:00
KyleNeverGivesUp e9bbaca07d [perf] Disable every offload path on unified memory, unblocking MiniMax H3 generation on one GB10 (#1715) 2026-08-26 15:32:56 -07:00
KyleNeverGivesUp 9bfa585448 [perf]: stop holding the whole checkpoint during DiT load, unblocking MiniMax H3 on one GB10 (#1714) 2026-08-26 15:23:51 -07:00
Raghav K b2062556a9 [perf] VSA Triton: widen the autotune num_stages range (the optimum was outside it) (#1706) 2026-08-26 12:30:56 -07:00
KyleNeverGivesUp c9c5585758 [perf]: MiniMax H3 on GB10 - skip text encoder CPU offload on unified memory (5m49s to 30ms) (#1710) 2026-08-26 12:03:14 -07:00
William Lin 9212f4f218 [ci] make GPU validation change-aware (#1747) 2026-08-25 21:26:25 -07:00
Aryan KumarandAryan Kumar 6388db815b [bugfix] FastMetal-QAD MLX support: refuse CUDA QAD trees, use packed mlx_dit config, stream loads (#1736) (#1758)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-25 14:51:40 -07:00
lpc0220 7a4285189f [kernel] Route block-sparse VSA to the sm_100a forward behind FASTVIDEO_VSA_SM100A (opt-in) (#1754) 2026-08-24 15:26:56 -07:00
William Lin a837fe841a [docs] Update FastH3 README (#1749) 2026-08-23 05:05:58 -07:00
William Lin f9e3680f11 [perf] Align FastH3 optimized inference profile (#1748) 2026-08-23 02:01:03 -07:00
William Lin 98f761ec45 [bugfix] validation: inherit the trained denoising ladder (#1738) 2026-08-22 23:09:26 -07:00
Shao Duan c041318f2c [perf] Add fused NVLink all-to-all for Ulysses (#1740) 2026-08-22 23:06:24 -07:00
William Lin 604e0205a4 [perf] Keep odd MiniMax-H3 VSA tiles on sm100a (#1745) 2026-08-22 18:29:30 -07:00
William Lin 13213395b4 [perf] Parallelize MiniMax-H3 VAE over sequence ranks (#1744) 2026-08-22 18:05:31 -07:00
William Lin 46afee5998 [bugfix] Classify MiniMax-H3 inference controls in schema inventory (#1743) 2026-08-22 17:17:02 -07:00
William Lin c488fa1211 [perf] Add opt-in packed-varlen FA4 for MiniMax-H3 (#1742) 2026-08-22 17:16:47 -07:00
William Lin d3cff517cd [perf] Add opt-in regional fullgraph compile for DiT inference (#1741) 2026-08-22 12:14:39 -07:00
Junda Su 2f3d407406 [perf] Optimize MiniMax H3 VAE decoding (#1734) 2026-08-21 14:57:32 -07:00
Kaiqin Kong bcffa4026e [perf] Optimize MiniMax-H3 text encoder memory (#1732) 2026-08-21 14:57:06 -07:00
William Lin 6d6a10be7a [feat] FastVideo-Minimax-FastH3-Preview few-step example + 64-token-tile VSA-H3 inference path (#1731) 2026-08-21 12:40:09 -05:00
Kaiqin Kong 73dd105f3d [perf] Add opt-in MiniMax-H3 Sol-Engine fusions (#1735) 2026-08-21 12:39:28 -05:00
William LinandClaude Fable 5 56d4a6074f [bugfix] fastvideo-kernel: fix Triton block-sparse backward logit scaling (bf16 K pre-scaling) (#1730)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-21 04:50:06 -05:00
Shao Duan c4ad4227c0 [misc] MiniMax-H3: move the AdaLN converter into scripts/checkpoint_conversion (#1712) 2026-08-21 02:15:00 -05:00
KyleNeverGivesUpandSolitaryThinker a63ccce73d [docs]: add a maintained inference cookbook (#1290)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-08-21 01:53:44 -05:00
KyleNeverGivesUp 0462e1b0e7 [perf]: MiniMax H3 - build the Qwen3-VL encoder only as far as it is read (-13.7 GB) (#1711) 2026-08-20 23:19:25 -05:00
lpc0220 907f2100ec [kernel] sm_100a CUDA block-sparse VSA forward (Blackwell), 64- and 128-token blocks (#1719) 2026-08-20 23:16:25 -05:00
Kaiqin Kong e0a3db5651 [perf] Reduce MiniMax-H3 VAE peak memory (#1703) 2026-08-20 23:13:23 -05:00
Raghav K fca45bc8e1 [perf] Quantize frames to uint8 on-device before the post-decode D->H copy (#1362) 2026-08-20 21:55:55 -05:00
Aryan Kumar 86d639c848 [docs] Announce FastMetal-QAD (#1721) 2026-08-19 13:45:49 -07:00
00338aa9ca [perf] Add FA4 CuTe backward support for VSA-256 (#1639)
Co-authored-by: Hyunsung Lee <hyunsungl@sizigistudios.com>
Co-authored-by: alexzms <3036648523@qq.com>
2026-08-19 11:46:49 -07:00
Aryan KumarandAryan Kumar 8537dcd6de [feat]: Apple Silicon MLX runtime — INT8 Wan2.1 and Wan2.2 inference (#1638)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-18 15:08:00 -07:00
William Lin 8208536cd1 [bugfix] profiler region system: record + export actually work, usability roll-up (#1691) 2026-08-09 20:35:53 -07:00
Kai 0653f8f3af [new-model] Add V2A: native MMAudio inference pipeline (#1622) 2026-08-09 17:29:27 -07:00
William Lin e0d702decb [feat] VSA for MiniMax H3: packed mixed-modality sparse attention (#1695) 2026-08-09 13:10:51 -07:00
Shao Duan 541ef014ee [perf] MiniMax-H3: rank-reduced AdaLN pruned model option (-39% params, -23 GiB VRAM) (#1699) 2026-08-09 12:31:57 -07:00
William Lin ffc1a7a58b [refactor] H3 pipeline cleanup: shared helpers, dead machinery, loop-invariant hoists (#1698) 2026-08-09 04:51:31 -07:00
William Lin 9028953625 [misc] yapf pass under CI's interpreter (3.12) + pin hook language_version (#1702) 2026-08-08 22:24:17 -07:00
KyleNeverGivesUpandClaude Opus 5 6eb95693a1 [misc]: re-run yapf on main so pre-commit passes again (#1700)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-07 16:17:52 -07:00
Junda Su c3567eb468 [feat] add Minimax H3 sft pipeline (#1688) 2026-08-07 16:12:04 -07:00
William Lin 15568f27db [perf]: H3 torch.compile + CUDA graphs (1.2-1.3x) with denoising step marking (#1689) 2026-08-06 16:23:33 -07:00
Kaiqin Kong 126a75ad63 [misc] Support partial Hugging Face model downloads (#1684) 2026-08-06 16:09:46 -07:00
Raghav KandSolitaryThinker a2bfc7cdb2 [docs] DGX Spark (GB10) performance & tuning guide + reproduction examples (#1631)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-08-06 14:36:38 -07:00
William Lin b963a24612 [bugfix]: hard-fail when ATTN_QAT_INFER is selected but the kernel is unusable (#1690) 2026-08-06 13:12:27 -07:00
William Lin fb7be2fe2c [ci]: add golden-gate lane — single-layer bitwise DiT fingerprints for all SSIM-covered families (#1682) 2026-08-05 12:47:03 -07:00
William Lin ab00392664 [bugfix]: add GB200 to the inline SSIM device tables #1676 missed (#1681) 2026-08-05 10:33:46 -07:00
Shao Duan 9f1e7c19d2 [bugfix] Wan I2V: CLIP image conditioning silently dropped when passed as a tensor during training (#1673) 2026-08-05 01:35:09 -07:00
Kaiqin Kong e8b0e4c61e [feat] Add MiniMax H3 (#1674) 2026-08-04 13:54:39 -07:00
Haochen Jiang 9145ffdc46 [bugfix]: keep _resolved_attention_backend out of the positional config signature (#1678) 2026-08-03 17:09:15 -07:00
William Lin c3d07c870b [bugfix]: give GB200 its own SSIM reference folder instead of B200's (#1676) 2026-08-03 15:14:25 -07:00
William Lin e8812bef0b [docs]: batched docs cleanup (landing page, links, requirements, nav) (#1644) 2026-08-02 18:03:26 -07:00
William Lin b9be2449dc [refactor]: delete the dead global attention-backend override (#1672) 2026-08-02 18:00:42 -07:00
Adhvay IyerandSolitaryThinker bc7a804618 [bugfix]: harden FastVideo Studio UI reliability and accessibility (#1659)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-08-02 15:26:18 -07:00
William Lin 7b094c945b [refactor]: resolve attention backend once per component at load time (#1657) 2026-08-02 15:07:42 -07:00
William Lin eeb3e8a597 [bugfix]: left-align Gemma connector tokens per batch row (#1664) 2026-08-02 15:01:34 -07:00
Suhaan Khurana 05406c5d1b [misc]: consolidate dataset download scripts under examples/datasets/ (#1667) 2026-08-02 14:21:24 -07:00
KyleNeverGivesUpandClaude Opus 5 99d04a7f98 [bugfix] keep loader-populated text encoder configs in the validation pipeline (#1669)
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-02 14:10:16 -07:00
William Lin 1b2b2a0161 [bugfix]: copy text encoder outputs out of CUDAGraph static buffers (#1650) 2026-07-27 21:52:53 -07:00
William Lin 98d65835b5 [misc]: refresh stale sm_120-only validation notes in QAT recipes (#1655) 2026-07-27 19:35:29 -07:00
William Lin 422585d08f [misc]: add LTX-2.3 fine-tuning example recipes (#1651) 2026-07-27 18:42:44 -07:00
ryanM154 e59a1ce16a [bugfix] Report actual package version in fastvideo --version (#1652) 2026-07-27 16:32:01 -07:00
William Lin d71acc0eb5 [bugfix]: fix stale imports in LTX-2.3 gradio local demo (#1640) 2026-07-27 16:31:07 -07:00
William Lin af2934dd6b [feat]: FA4-FP4 ATTN_QAT_INFER on sm_100/sm_103 + NVFP4 weight purge (#1647) 2026-07-27 15:46:52 -07:00
William Lin 1801512818 [docs]: cover all registered models in the support matrix (#1641) 2026-07-27 11:13:53 -07:00
William Lin 5ae05b032e [misc]: add LTX-2 fine-tuning example recipes (#1645) 2026-07-27 11:11:49 -07:00
Mac Lee 7a592ff09a [ci]: cache FastVideo kernel builds in Modal (#1562) 2026-07-26 04:21:09 -07:00
Lev NovitskiyandClaude Sonnet 5 8b23984c79 [feat] Add Kandinsky5 QAD training pipeline: data preprocessing, QAT finetune, QAT-aware DMD distillation (#1601)
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-26 02:59:21 -07:00
William Lin bf18371afe [feat] Enable LTX-2 NVFP4 linear and attention QAT fine-tuning (#1626) 2026-07-26 02:19:34 -07:00
William Lin 69349dd2aa [docs]: unify community links on the README Slack invite (#1643) 2026-07-25 17:26:58 -07:00
Yogya MehrotraandClaude Sonnet 5 8d89f30d3f [bugfix] Skip CUDA-only fastvideo-kernel/flashinfer-python deps on non-Linux (#1574)
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-23 14:23:45 -07:00
William Lin 10546353da [ci]: extend Full Suite training lane timeouts (#1616) 2026-07-23 13:00:00 -07:00
pkisfaludi-nvandClaude Opus 4.8 9fb74b9732 Make LTX-2 RMSNorm out-of-place so torch_tensorrt + Ulysses SP compiles (#1623)
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-22 04:32:04 -07:00
Adriel FungandSolitaryThinker 521dee0e82 [perf]: enable per-block torch.compile for LTX2 with persistent cache (#1602)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-20 15:59:13 -07:00
William Lin 229419208e [ci]: opt in to fork-PR head checkout after actions/checkout guard change (#1625) 2026-07-20 13:12:22 -07:00
William Lin 65f3b946b9 [feat] Add LTX-2 and LTX-2.3 fine-tuning to the modular trainer (#1624) 2026-07-20 11:58:39 -07:00
Zhang Peiyuan 191fcbf46c [feat] add qat docs (#1621) 2026-07-19 18:37:47 -07:00
Junda Su 755a4e4470 [new-model] Add LingBot-Video Dense and MoE/refiner T2V inference (#1595) 2026-07-18 20:18:16 -07:00
9709b7513b [feat] Port NVFP4 QAT/QAD to modular train framework (#1619)
Co-authored-by: Peiyuan Zhang <email>

Co-authored-by: Peiyuan <a>
2026-07-18 15:01:09 -07:00
William Lin 32cd603515 Revert docs trusted-branch-only workflow (#1618) 2026-07-17 00:22:56 -07:00
Junda Su d4bdd3621a [new-model] Port LingBot-World-v2 (#1579) 2026-07-16 19:51:53 -07:00
William Lin e2f8322842 [ci]: skip unused Buildkite submodule checkout (#1614) 2026-07-16 18:37:57 -07:00
6966f9e0bc [fix] Z-Image (#1236) draft port: rebase + strict-load contract + bf16 encoder parity + PORT_STATUS (#1339)
Co-authored-by: Mrinaal Dogra <mdogra@ucsd.edu>
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-16 18:18:25 -07:00
Mac Lee 743f4ed5f9 [ci]: harden Modal repository checkout (#1590) 2026-07-16 18:05:50 -07:00
Mac Lee 1c04ace573 [ci]: pin VSA training regression to H100 (#1591) 2026-07-15 22:04:57 -07:00
Mac LeeandSatyam Srivastava 6cbff73687 [bugfix]: skip unused output materialization (#1567)
Co-authored-by: Satyam Srivastava <srivastavasatyam53@gmail.com>
2026-07-16 04:03:07 +00:00
William Lin dec8b10939 [docs]: document automatic Docker image builds (#1608) 2026-07-15 18:52:21 -07:00
William Lin 133a5278af [ci] Run docs only for trusted PR branches (#1610) 2026-07-15 18:51:59 -07:00
William Lin da856274cc [feat]: FastVideo Studio — SvelteKit → Next.js port + review fixes (#1612) 2026-07-15 18:51:33 -07:00
Mac LeeandSolitaryThinker a253856147 [ci] Add exact identity performance statuses (#1560)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-15 14:59:09 -07:00
Satyam Srivastavaandgemini-code-assist[bot] 6e25d94ebc [ci]: enable scheduled perf runs to update rolling baseline (#1599)
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2026-07-14 16:55:12 -07:00
Mac Lee cae8fa18dc [bugfix]: propagate Qwen2.5-VL visual dtype (#1580) 2026-07-13 18:37:11 -07:00
William Lin 821e5a0832 [bugfix]: fix FlashAttention resolver tests after tuple return (#1597) 2026-07-13 16:29:03 -07:00
William Lin c1abc42782 [bugfix]: allow unrestricted head sizes in SDPA (#1596) 2026-07-13 16:04:51 -07:00
William Lin ef15ea2391 [bugfix]: keep LTX2 rms_norm outputs bf16 under torch 2.12 autocast (#1587) 2026-07-13 16:04:21 -07:00
Mac Lee 1ea2517e22 [ci]: extend LoRA training CI timeout (#1589) 2026-07-13 02:47:28 -07:00
MookandSolitaryThinker 0c63528c59 [perf] Cache RoPE position-embedding tables across denoising steps (#1442)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-13 02:34:51 -07:00
b063f8ca41 [feat] Fix FLUX.1-dev port: native RoPE, parity tests, SSIM reference (#1321)
Co-authored-by: Ishan Vaish <ivaish@ucsd.edu>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-13 01:49:42 -07:00
William Lin e7fff0173a [bugfix]: benchmark_weight_loading_comparison.py — iterate safe_open via .keys() (#1378) 2026-07-12 22:50:45 -07:00
Shreejith SGandH1yori233 d82abc271e [feat] Add GLM-Image inference support (#1030)
Co-authored-by: H1yori233 <k1kong@ucsd.edu>
2026-07-12 22:42:42 -07:00
Guian FangandSolitaryThinker 970409962f [feat] Add AnyFlow any-step video distillation (pretrain + on-policy) (#1371)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-12 02:32:08 +00:00
Raghav K 055586703d [perf]: register a real backward for FA2 default + masked/varlen custom ops (training-under-compile) (#1388) 2026-07-12 00:45:27 +00:00
Mac Lee 5d89f86675 [ci] Stop forcing FA4 in model-load lanes (#1561) 2026-07-11 14:10:36 -07:00
Satyam Srivastava 19a51a1fe6 [ci] Trigger performance benchmarks for performance code changes (#1583) 2026-07-10 20:21:33 -07:00
William Lin d3232cea5a [ci]: gate the full-suite trigger on pre-commit and docs build (#1572) 2026-07-11 02:56:49 +00:00
Raghav KandSolitaryThinker 0c90c8c24d [bugfix] nvfp4: cast fp32 inputs to bf16 instead of asserting (#1488)
Co-authored-by: SolitaryThinker <wlsaidhi@gmail.com>
2026-07-10 21:39:17 +00:00
Mingjia HuoandClaude Fable 5 4c08ffce49 [feat] World model training using third person games (#1443)
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-07-10 05:09:07 +00:00
Atharv Ramesh af4a77553c [ci]: add SSIM reference bootstrap flow (#1522) (#1547) 2026-07-10 01:49:38 +00:00
2128 changed files with 264305 additions and 40330 deletions
@@ -10,9 +10,7 @@ from pathlib import Path
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Clone a reference repo for FastVideo parity tests."
)
parser = argparse.ArgumentParser(description="Clone a reference repo for FastVideo parity tests.")
parser.add_argument("repo_url", help="Official reference repository URL")
parser.add_argument("target_dir", help="Directory to clone into")
parser.add_argument("--branch", help="Branch or tag to clone")
@@ -62,9 +60,7 @@ def gitignore_entry_for(target: Path) -> str:
try:
relative = resolved.relative_to(root)
except ValueError as exc:
raise ValueError(
"--update-gitignore requires target_dir to be under the current directory"
) from exc
raise ValueError("--update-gitignore requires target_dir to be under the current directory") from exc
text = relative.as_posix().rstrip("/")
return "/" + text + "/"
@@ -8,14 +8,12 @@ import os
import sys
from pathlib import Path
HF_TOKEN_ENV_KEYS = ("HF_TOKEN", "HUGGINGFACE_HUB_TOKEN", "HF_API_KEY")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Download a HF model snapshot or selected files into a local directory."
)
description="Download a HF model snapshot or selected files into a local directory.")
parser.add_argument("repo_id", help="HF repo id, for example Org/Model")
parser.add_argument("local_dir", help="Destination directory")
parser.add_argument("--repo-type", default="model", help="HF repo type (default: model)")
@@ -10,7 +10,6 @@ import sys
from pathlib import Path
from typing import Any
HF_TOKEN_ENV_KEYS = ("HF_TOKEN", "HUGGINGFACE_HUB_TOKEN", "HF_API_KEY")
RAW_WEIGHT_SUFFIXES = (".safetensors", ".pt", ".pth", ".ckpt", ".bin")
KNOWN_COMPONENTS = {
@@ -34,8 +33,7 @@ KNOWN_COMPONENTS = {
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Classify a HF repo or local directory as Diffusers, raw, custom, or unknown."
)
description="Classify a HF repo or local directory as Diffusers, raw, custom, or unknown.")
parser.add_argument("source", help="HF repo id or local weights directory")
parser.add_argument("--repo-type", default="model", help="HF repo type (default: model)")
parser.add_argument("--revision", help="HF revision to inspect")
@@ -94,14 +92,12 @@ def load_remote_files(
) -> list[str]:
from huggingface_hub import list_repo_files
return sorted(
list_repo_files(
repo_id,
repo_type=repo_type,
revision=revision,
token=token,
)
)
return sorted(list_repo_files(
repo_id,
repo_type=repo_type,
revision=revision,
token=token,
))
def load_remote_model_index(
@@ -215,24 +211,24 @@ def build_result(args: argparse.Namespace) -> dict[str, Any]:
"components_seen": components,
"file_count": len(files),
"file_scan_truncated": truncated,
"files_sample": files[: args.sample_limit],
"files_sample": files[:args.sample_limit],
}
def print_human(result: dict[str, Any]) -> None:
for key in (
"source",
"source_kind",
"repo_type",
"revision",
"token_env",
"source_layout",
"needs_conversion",
"model_index_class",
"model_index_diffusers_version",
"model_index_error",
"file_count",
"file_scan_truncated",
"source",
"source_kind",
"repo_type",
"revision",
"token_env",
"source_layout",
"needs_conversion",
"model_index_class",
"model_index_diffusers_version",
"model_index_error",
"file_count",
"file_scan_truncated",
):
value = result.get(key)
if value is not None:
@@ -18,7 +18,6 @@ import pytest
import torch
from torch.testing import assert_close
os.environ.setdefault("MASTER_ADDR", "localhost")
os.environ.setdefault("MASTER_PORT", "29519")
os.environ.setdefault("DISABLE_SP", "1")
@@ -35,15 +34,10 @@ FASTVIDEO_CONFIG_CLASS = "<FastVideoConfig>" # TODO.
FASTVIDEO_MODEL_MODULE = "fastvideo.models.<bucket>.<module>" # TODO.
FASTVIDEO_MODEL_CLASS = "<FastVideoModel>" # TODO.
OFFICIAL_REF_DIR = Path(
os.getenv("<FAMILY_UPPER>_OFFICIAL_REF_DIR", REPO_ROOT / "<ReferenceDir>")
)
LOCAL_WEIGHTS_DIR = Path(
os.getenv("<FAMILY_UPPER>_LOCAL_WEIGHTS_DIR", REPO_ROOT / "official_weights" / FAMILY)
)
CONVERTED_WEIGHTS_DIR = Path(
os.getenv("<FAMILY_UPPER>_CONVERTED_WEIGHTS_DIR", REPO_ROOT / "converted_weights" / FAMILY)
)
OFFICIAL_REF_DIR = Path(os.getenv("<FAMILY_UPPER>_OFFICIAL_REF_DIR", REPO_ROOT / "<ReferenceDir>"))
LOCAL_WEIGHTS_DIR = Path(os.getenv("<FAMILY_UPPER>_LOCAL_WEIGHTS_DIR", REPO_ROOT / "official_weights" / FAMILY))
CONVERTED_WEIGHTS_DIR = Path(os.getenv("<FAMILY_UPPER>_CONVERTED_WEIGHTS_DIR",
REPO_ROOT / "converted_weights" / FAMILY))
def _resolve_hf_token() -> str | None:
@@ -99,18 +93,14 @@ def _load_official_model(device: torch.device, dtype: torch.dtype) -> torch.nn.M
model = OfficialClass() # TODO: pass official config kwargs.
state_dict = {} # TODO: load official state dict from LOCAL_WEIGHTS_DIR.
missing, unexpected = model.load_state_dict(state_dict, strict=True)
assert not missing and not unexpected, (
f"official load mismatch missing={missing[:5]} unexpected={unexpected[:5]}"
)
assert not missing and not unexpected, (f"official load mismatch missing={missing[:5]} unexpected={unexpected[:5]}")
return model.to(device=device, dtype=dtype).eval()
def _load_fastvideo_model(device: torch.device, dtype: torch.dtype) -> torch.nn.Module:
"""Load the FastVideo component with the same tensor content."""
if not CONVERTED_WEIGHTS_DIR.exists() and not LOCAL_WEIGHTS_DIR.exists():
pytest.skip(
f"No FastVideo loadable weights: {CONVERTED_WEIGHTS_DIR} or {LOCAL_WEIGHTS_DIR}"
)
pytest.skip(f"No FastVideo loadable weights: {CONVERTED_WEIGHTS_DIR} or {LOCAL_WEIGHTS_DIR}")
# TODO: replace with the bucket-specific FastVideo config/class/loader.
# DiT examples:
@@ -127,8 +117,7 @@ def _load_fastvideo_model(device: torch.device, dtype: torch.dtype) -> torch.nn.
state_dict = {} # TODO: load converted or directly mapped state dict.
missing, unexpected = model.load_state_dict(state_dict, strict=True)
assert not missing and not unexpected, (
f"FastVideo load mismatch missing={missing[:5]} unexpected={unexpected[:5]}"
)
f"FastVideo load mismatch missing={missing[:5]} unexpected={unexpected[:5]}")
return model.to(device=device, dtype=dtype).eval()
@@ -187,11 +176,9 @@ def test_component_parity():
assert official_out.shape == fastvideo_out.shape
diff = (official_out - fastvideo_out).abs()
print(
f"official abs_mean={official_out.abs().mean().item():.6f} "
f"fastvideo abs_mean={fastvideo_out.abs().mean().item():.6f} "
f"diff_max={diff.max().item():.6f} diff_mean={diff.mean().item():.6f}"
)
print(f"official abs_mean={official_out.abs().mean().item():.6f} "
f"fastvideo abs_mean={fastvideo_out.abs().mean().item():.6f} "
f"diff_max={diff.max().item():.6f} diff_mean={diff.mean().item():.6f}")
# TODO: pick tolerance by scope:
# - single block / same kernel: 1e-4
@@ -27,7 +27,6 @@ try:
except ImportError: # pragma: no cover - optional local conversion dependency
snapshot_download = None
# TODO: fill with authoritative component prefixes for monolithic checkpoints.
# Example: {"model.model.": "transformer", "pretransform.model.": "vae"}
COMPONENT_PREFIXES: dict[str, str] = {}
@@ -47,10 +46,7 @@ SKIP_PATTERNS: tuple[str, ...] = ()
def _hf_token() -> str | None:
return (
os.environ.get("HF_TOKEN") or os.environ.get("HUGGINGFACE_HUB_TOKEN")
or os.environ.get("HF_API_KEY")
)
return (os.environ.get("HF_TOKEN") or os.environ.get("HUGGINGFACE_HUB_TOKEN") or os.environ.get("HF_API_KEY"))
def resolve_src(src: str, revision: str | None) -> Path:
@@ -95,11 +91,10 @@ def apply_mapping(key: str) -> str | None:
return key
def split_monolithic(
state: dict[str, torch.Tensor],
) -> dict[str, OrderedDict[str, torch.Tensor]]:
def split_monolithic(state: dict[str, torch.Tensor], ) -> dict[str, OrderedDict[str, torch.Tensor]]:
components: dict[str, OrderedDict[str, torch.Tensor]] = {
name: OrderedDict() for name in set(COMPONENT_PREFIXES.values())
name: OrderedDict()
for name in set(COMPONENT_PREFIXES.values())
}
intentionally_skipped: list[str] = []
unowned: list[str] = []
@@ -117,10 +112,8 @@ def split_monolithic(
unowned.append(key)
if unowned:
sample = ", ".join(unowned[:10])
raise ValueError(
f"Unowned monolithic keys: {len(unowned)}. "
f"Add COMPONENT_PREFIXES or SKIP_PATTERNS entries. Sample: {sample}"
)
raise ValueError(f"Unowned monolithic keys: {len(unowned)}. "
f"Add COMPONENT_PREFIXES or SKIP_PATTERNS entries. Sample: {sample}")
if intentionally_skipped:
print(f"Intentionally skipped {len(intentionally_skipped)} keys")
return {name: weights for name, weights in components.items() if weights}
@@ -143,8 +136,12 @@ def build_component_configs(_src_dir: Path) -> dict[str, dict[str, Any]]:
# TODO: emit config content accepted by FastVideo loaders. Most components use
# config.json; schedulers use scheduler_config.json.
return {
"transformer": {"_class_name": "<FastVideoTransformerClass>"},
"vae": {"_class_name": "<FastVideoVAEClass>"},
"transformer": {
"_class_name": "<FastVideoTransformerClass>"
},
"vae": {
"_class_name": "<FastVideoVAEClass>"
},
}
@@ -177,19 +174,13 @@ def build_model_index(
}
if revision:
index["_fastvideo_converted_revision"] = revision
return {
key: value
for key, value in index.items()
if key.startswith("_") or key in available_components
}
return {key: value for key, value in index.items() if key.startswith("_") or key in available_components}
def validate_component_configs(configs: dict[str, dict[str, Any]]) -> None:
# TODO: instantiate each FastVideo config and call update_model_arch(...) or
# update_model_config(...) with this JSON so unknown emitted keys fail here.
placeholder_configs = [
name for name, config in configs.items() if "<" in json.dumps(config)
]
placeholder_configs = [name for name, config in configs.items() if "<" in json.dumps(config)]
if placeholder_configs:
raise ValueError(f"Replace config placeholders for: {placeholder_configs}")
@@ -201,9 +192,7 @@ def verify_conversion(
del dst_dir, components
# TODO: load each emitted stateful component through its production loader and
# assert strict load, or document exact allowed missing/unexpected keys.
raise NotImplementedError(
"Implement production config validation and strict-load checks"
)
raise NotImplementedError("Implement production config validation and strict-load checks")
def write_component(
@@ -216,9 +205,7 @@ def write_component(
if component_dir.exists() and any(component_dir.iterdir()):
shutil.rmtree(component_dir)
component_dir.mkdir(parents=True, exist_ok=True)
save_file(
dict(state), str(component_dir / "diffusion_pytorch_model.safetensors")
)
save_file(dict(state), str(component_dir / "diffusion_pytorch_model.safetensors"))
if config is not None:
config_path = component_dir / config_filename(name)
with config_path.open("w", encoding="utf-8") as f:
@@ -261,9 +248,7 @@ def convert(
if layout in {"monolithic", "raw_official"}:
# TODO: replace model.safetensors with the official monolithic file name.
components = split_monolithic(
load_checkpoint(default_monolithic_checkpoint(src_path))
)
components = split_monolithic(load_checkpoint(default_monolithic_checkpoint(src_path)))
elif layout in {"separate_components", "mixed"}:
if not src_path.is_dir():
raise ValueError(f"{layout} layout requires a source directory: {src_path}")
@@ -271,9 +256,7 @@ def convert(
else:
raise ValueError(f"Unsupported template layout: {layout}")
copied = (
copy_passthrough(src_path, dst_dir) if src_path.is_dir() else []
)
copied = (copy_passthrough(src_path, dst_dir) if src_path.is_dir() else [])
configs = build_component_configs(src_path if src_path.is_dir() else src_path.parent)
validate_component_configs(configs)
for name, state in components.items():
@@ -289,9 +272,7 @@ def convert(
def main() -> None:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument(
"--src", required=True, help="HF repo id, local dir, or checkpoint path"
)
parser.add_argument("--src", required=True, help="HF repo id, local dir, or checkpoint path")
parser.add_argument("--revision", help="HF branch, tag, or commit for repo sources")
parser.add_argument(
"--dst",
@@ -24,8 +24,8 @@ from typing import Any
import torch
FAMILY: str = "<family>" # e.g. "magi_human", "ltx2", "wan"
COMPONENT: str = "<component>" # e.g. "dit", "vae", "encoder"
FAMILY: str = "<family>" # e.g. "magi_human", "ltx2", "wan"
COMPONENT: str = "<component>" # e.g. "dit", "vae", "encoder"
DRILL_LAYER_ENV: str = "<FAMILY>_DEBUG_DRILL_LAYER"
HYPOTHESIS_ENV: str = "<FAMILY>_DEBUG_PATCH_<HYPOTHESIS>"
REL_THRESHOLD: float = 0.005 # 0.5% abs_mean drift flags a block as divergent
@@ -94,6 +94,7 @@ def _attach_block_hooks(
handles: list[Any] = []
def _hook(name: str):
def fn(_module, _inputs, outputs):
t = outputs[0] if isinstance(outputs, tuple) else outputs
if not torch.is_tensor(t):
@@ -101,6 +102,7 @@ def _attach_block_hooks(
log.append({"side": label, **_stat(name, t)})
if tensors is not None:
tensors[name] = t.detach().float().cpu()
return fn
def _pre_hook(name: str):
@@ -114,6 +116,7 @@ def _attach_block_hooks(
log.append({"side": label, **_stat(key, t)})
if tensors is not None:
tensors[key] = t.detach().float().cpu()
return fn
# TODO: adapt attribute paths to your model. Remove adapter block if absent.
@@ -131,43 +134,21 @@ def _attach_block_hooks(
# magi-human uses: attention, mlp.pre_norm, mlp.up_gate_proj,
# mlp.down_proj (pre+post), mlp, attn_post_norm, mlp_post_norm.
if hasattr(layer, "attention"):
handles.append(
layer.attention.register_forward_hook(_hook(f"{tag}.attention"))
)
handles.append(layer.attention.register_forward_hook(_hook(f"{tag}.attention")))
if hasattr(layer, "mlp"):
mlp = layer.mlp
if hasattr(mlp, "pre_norm"):
handles.append(
mlp.pre_norm.register_forward_hook(_hook(f"{tag}.mlp.pre_norm"))
)
handles.append(mlp.pre_norm.register_forward_hook(_hook(f"{tag}.mlp.pre_norm")))
if hasattr(mlp, "up_gate_proj"):
handles.append(
mlp.up_gate_proj.register_forward_hook(
_hook(f"{tag}.mlp.up_gate_proj")
)
)
handles.append(mlp.up_gate_proj.register_forward_hook(_hook(f"{tag}.mlp.up_gate_proj")))
if hasattr(mlp, "down_proj"):
handles.append(
mlp.down_proj.register_forward_pre_hook(
_pre_hook(f"{tag}.mlp.down_proj")
)
)
handles.append(
mlp.down_proj.register_forward_hook(_hook(f"{tag}.mlp.down_proj"))
)
handles.append(mlp.down_proj.register_forward_pre_hook(_pre_hook(f"{tag}.mlp.down_proj")))
handles.append(mlp.down_proj.register_forward_hook(_hook(f"{tag}.mlp.down_proj")))
handles.append(mlp.register_forward_hook(_hook(f"{tag}.mlp")))
if hasattr(layer, "attn_post_norm"):
handles.append(
layer.attn_post_norm.register_forward_hook(
_hook(f"{tag}.attn_post_norm")
)
)
handles.append(layer.attn_post_norm.register_forward_hook(_hook(f"{tag}.attn_post_norm")))
if hasattr(layer, "mlp_post_norm"):
handles.append(
layer.mlp_post_norm.register_forward_hook(
_hook(f"{tag}.mlp_post_norm")
)
)
handles.append(layer.mlp_post_norm.register_forward_hook(_hook(f"{tag}.mlp_post_norm")))
return handles
@@ -193,11 +174,9 @@ def _write_log(entries: list[dict], path: Path) -> None:
path.parent.mkdir(parents=True, exist_ok=True)
with open(path, "w") as f:
for e in entries:
f.write(
f"{e['name']} {e['shape']} "
f"{e['abs_mean']:.8f} {e['sum']:.4f} "
f"{e['min']:.6f} {e['max']:.6f}\n"
)
f.write(f"{e['name']} {e['shape']} "
f"{e['abs_mean']:.8f} {e['sum']:.4f} "
f"{e['min']:.6f} {e['max']:.6f}\n")
def _sort_key(name: str, drill_layer: int) -> tuple:
@@ -205,9 +184,14 @@ def _sort_key(name: str, drill_layer: int) -> tuple:
return (0, "")
if name.startswith(f"L{drill_layer:02d}."):
sub_order = {
"attention": 0, "attn_post_norm": 1, "mlp.pre_norm": 2,
"mlp.up_gate_proj": 3, "mlp.down_proj<in>": 4,
"mlp.down_proj": 5, "mlp": 6, "mlp_post_norm": 7,
"attention": 0,
"attn_post_norm": 1,
"mlp.pre_norm": 2,
"mlp.up_gate_proj": 3,
"mlp.down_proj<in>": 4,
"mlp.down_proj": 5,
"mlp": 6,
"mlp_post_norm": 7,
}.get(name.split(".", 1)[1], 9)
return (1, f"block[{drill_layer:02d}]", sub_order)
if name.startswith("block["):
@@ -216,10 +200,8 @@ def _sort_key(name: str, drill_layer: int) -> tuple:
def _print_table(by_name: dict[str, dict], drill_layer: int) -> int | None:
hdr = (
f"{'name':<18} {'up_shape':<22} {'up_absmean':>12} {'fv_absmean':>12} "
f"{'absmean_diff':>14} {'rel%':>8} {'up_sum':>14} {'fv_sum':>14} {'sum_diff':>12}"
)
hdr = (f"{'name':<18} {'up_shape':<22} {'up_absmean':>12} {'fv_absmean':>12} "
f"{'absmean_diff':>14} {'rel%':>8} {'up_sum':>14} {'fv_sum':>14} {'sum_diff':>12}")
print(f"\n{hdr}\n{'-' * len(hdr)}")
first_div: int | None = None
for name in sorted(by_name.keys(), key=lambda n: _sort_key(n, drill_layer)):
@@ -235,11 +217,9 @@ def _print_table(by_name: dict[str, dict], drill_layer: int) -> int | None:
flag = " <<< DIVERGE"
if first_div is None:
first_div = int(name[len("block["):-1])
print(
f"{name:<18} {str(up['shape']):<22} {up['abs_mean']:>12.6f} "
f"{fv['abs_mean']:>12.6f} {am_diff:>14.6f} {am_rel * 100:>7.3f}% "
f"{up['sum']:>14.4f} {fv['sum']:>14.4f} {sum_diff:>12.4f}{flag}"
)
print(f"{name:<18} {str(up['shape']):<22} {up['abs_mean']:>12.6f} "
f"{fv['abs_mean']:>12.6f} {am_diff:>14.6f} {am_rel * 100:>7.3f}% "
f"{up['sum']:>14.4f} {fv['sum']:>14.4f} {sum_diff:>12.4f}{flag}")
return first_div
@@ -255,10 +235,8 @@ def _print_elementwise(up_t: dict[str, torch.Tensor], fv_t: dict[str, torch.Tens
continue
diff = (a - b).abs()
rel = (diff.mean().item() / max(a.abs().mean().item(), 1e-9)) * 100
print(
f"{name:<30} {str(tuple(a.shape)):<22} "
f"{diff.max().item():>12.6f} {diff.mean().item():>12.6f} {rel:>9.4f}%"
)
print(f"{name:<30} {str(tuple(a.shape)):<22} "
f"{diff.max().item():>12.6f} {diff.mean().item():>12.6f} {rel:>9.4f}%")
def main() -> None:
@@ -43,12 +43,10 @@ def _add_official_to_path() -> Path:
def _log_tensor_stats(label: str, tensor: torch.Tensor) -> None:
value = tensor.detach().float()
print(
f"[{_MODEL_FAMILY} PIPELINE] {label}: shape={tuple(tensor.shape)} "
f"dtype={tensor.dtype} device={tensor.device} "
f"min={value.min().item():.6f} max={value.max().item():.6f} "
f"mean={value.mean().item():.6f} std={value.std().item():.6f}"
)
print(f"[{_MODEL_FAMILY} PIPELINE] {label}: shape={tuple(tensor.shape)} "
f"dtype={tensor.dtype} device={tensor.device} "
f"min={value.min().item():.6f} max={value.max().item():.6f} "
f"mean={value.mean().item():.6f} std={value.std().item():.6f}")
def _extract_tensor(output: Any, key: str) -> torch.Tensor:
@@ -73,10 +71,8 @@ def _run_official_pipeline(
device: torch.device,
) -> Any:
del official_path, params, device
pytest.skip(
"TODO: import the official pipeline/factory, load official weights, "
"run with params, and return the comparison target."
)
pytest.skip("TODO: import the official pipeline/factory, load official weights, "
"run with params, and return the comparison target.")
def _run_fastvideo_pipeline(model_path: Path, params: dict[str, Any]) -> Any:
@@ -146,8 +142,6 @@ def test_todo_model_family_pipeline_official_parity() -> None:
assert official_tensor.shape == fastvideo_tensor.shape
diff = (official_tensor - fastvideo_tensor).abs()
print(
f"diff max={diff.max().item():.6f} "
f"mean={diff.mean().item():.6f} median={diff.median().item():.6f}"
)
print(f"diff max={diff.max().item():.6f} "
f"mean={diff.mean().item():.6f} median={diff.median().item():.6f}")
assert_close(fastvideo_tensor, official_tensor, atol=1e-2, rtol=1e-2)
+99
View File
@@ -0,0 +1,99 @@
---
name: ci-runner
description: Work on FastVideo's Slurm-only, change-aware GPU CI lanes, static Buildkite graph, trusted ci-runner policy, lane scripts, and GB200 validation.
---
# Slinky Slurm CI lanes
FastVideo's `ci-runner` Buildkite queue is the control plane for all active
GPU CI. A private host-owned dispatcher leases GPUs from the Slinky Slurm tray
and runs the immutable PR SHA inside an isolated Enroot container. Buildkite
pipeline upload and Slurm submission occur on the login plane; every test
payload executes on Slurm compute.
The files under `fastvideo/tests/modal/` and `.buildkite/scripts/pr_test.sh`
are dormant rollback code. Never add an active Buildkite or slash-command
route to them. `pr_test.sh` must continue to reject Buildkite invocations.
The private operator bundle is deliberately outside this repository because
it contains site paths and credentials. See
`docs/contributing/ci_architecture.md`; this skill covers the repository half
and the coordination contract with that bundle.
## Invariants
- `.buildkite/pipeline.yml` contains exactly one static step for every active
GPU lane. Each step pins a unique key and label, a 90-minute timeout, the
trusted `/opt/fastvideo-ci-runner/run-ci` command (`run-unit` is the one
compatibility wrapper), step-level internal `TEST_TYPE`, and
`queue: "ci-runner"`.
- Active CI contains no `pr_test.sh` command, Modal invocation, default queue,
Buildkite plugin, `soft_fail`, or job-controlled artifact glob.
- The six Fastcheck lanes use `:microscope:` labels. Full-Suite-only lanes use
`:test_tube:` or `:bar_chart:` so direct reruns update the right aggregate.
- SSIM and vanilla training request all four GPUs. Keep both in the
`fastvideo/slinky/whole-tray` Buildkite concurrency group with a limit of one
so the second job does not consume an agent or command timeout while waiting
for the same tray.
- `/test full` schedules all twenty lanes. `/merge`, `ready`, and new pushes to
ready PRs use the trusted base-branch planner in
`.github/scripts/plan_merge_ci.py`: automatic Fastcheck remains the universal
six-lane baseline, and the merge build adds only path-relevant integration
lanes. Unknown source/build paths fail closed to all fourteen additive lanes.
The trusted uploader still normalizes and validates the complete static graph
before Buildkite evaluates its plan conditions.
- Focused merge builds may pass allowlisted golden-gate and SSIM test basenames.
The private host validates the lane plan and basenames before staging them,
and the in-container scripts validate them again. Direct `/test ssim`,
explicit `/test full`, and the weekly main-branch schedule run the complete
SSIM matrix.
- The trusted uploader serves exactly three entry pipelines:
`pr-fastcheck` for automatic PR builds, `ci` for slash-command/ready-label
API builds, and `fastvideo-performance-lane` for the weekly schedule. Keep
incoming GitHub webhook processing disabled on `ci` so it cannot duplicate
`pr-fastcheck` on every PR update.
- Test payloads live in `.buildkite/scripts/unit_test.sh` or executable
`.buildkite/scripts/lanes/<lane>.sh`. Backend policy (GPU count, extras,
secrets, kernel build, artifacts) stays in the agent-owned lane table.
- Tests must preserve an inherited `MASTER_PORT`. Packed containers share the
tray network namespace, so the private runner assigns a distinct port range
per GPU lease and the SSIM scheduler assigns task offsets within its range.
- The ARM64 runner image includes the pinned FA4 CuTe overlay validated on
GB200. Keep SSIM at `FASTVIDEO_FA4=1` because its references were seeded with
FA4; keep lanes with FA2 baselines at `FASTVIDEO_FA4=0`. A runner image change
must revalidate both the FA4 import and an actual GB200 forward kernel.
- `fastvideo/tests/ssim/ci_runner.py` is the active four-GPU SSIM scheduler.
New SSIM files are discovered through `REQUIRED_GPUS` and
`*_MODEL_TO_PARAMS`; do not wire them through the dormant Modal scheduler.
- The host policy fail-closes unknown tuples. A repository-side lane change is
inert until the operator updates the private lane table and uploader policy
in the same rollout.
## Adding or changing a lane
1. Read the closest `AGENTS.md` and the domain-specific testing guide.
2. Add or update the executable lane payload under `.buildkite/scripts/`.
Keep it deterministic and free of host-specific paths or credential fetches.
3. Add the static pipeline step and canonical `/test <name>` mapping. Keep the
`<name>-ci` alias only when compatibility requires it.
4. Add its source/test path ownership to `.github/scripts/plan_merge_ci.py`.
Prefer the narrowest correctness-preserving lane set; leave unknown paths
fail-closed. Extend `fastvideo/tests/contract/test_ci_test_collection.py`,
`test_merge_ci_plan.py`, and focused CPU-only scheduler/policy tests.
5. Coordinate the private lane row: GPU count (1-4), wall time, script, scope
pairs, step key, command, HF cache/token, tracking mode, extras, attention
backend policy, kernel policy, and artifact relay. Active training lanes
keep W&B offline and do not stage a W&B credential.
6. Update the trusted pipeline-uploader schema. A mismatch must reject the
pipeline rather than silently skip a lane.
7. Run `pre-commit run --files <changed paths>`, the planner's representative
diff matrix, contract tests, private driver tests, and a real GB200 canary.
Multi-GPU, hardware-reference, training, performance, and SSIM changes need
their own target-hardware evidence.
## Rollback
Rollback the Slurm routing/configuration change or pause the `ci-runner` queue.
Do not silently reactivate Modal. A manual Modal experiment requires the
explicit local opt-in documented in `ci_architecture.md`; returning it to
production CI needs a separate reviewed decision.
@@ -0,0 +1,79 @@
---
name: env-var-conventions
description: Add, read, rename, or remove an environment variable in FastVideo, or change the environment-variable policy. Use before touching fastvideo/envs.py, os.environ, os.getenv, or monkeypatch.setenv in fastvideo/, and when fastvideo/tests/contract/test_env_policy.py fails.
---
# Environment Variable Conventions
## Purpose
FastVideo registers its environment variables as typed fields in
`fastvideo/envs.py`. The policy that governs them is
`docs/contributing/env_vars.md`, and the contract test
`fastvideo/tests/contract/test_env_policy.py` enforces the policy in the unit
CI lane. This skill routes an environment-variable change through that policy.
The policy doc is the single source of the rules; read it instead of relying
on a summary here.
## Prerequisites
- Read `docs/contributing/env_vars.md` in full.
- Decide whether the setting belongs in an environment variable or an argument
(rule 5 in the policy doc). Settings that users change per deployment are
arguments; add them through `fastvideo/fastvideo_args.py` instead.
## Inputs
| Parameter | Required | Description |
| ---------- | -------- | -------------------------------------------------------------- |
| `change` | Yes | Add, read, rename, or remove a variable, or change the policy. |
| `variable` | Yes | The variable name, with the `FASTVIDEO_` prefix. |
## Steps
1. **Declare or edit the variable in `fastvideo/envs.py`.**
- Pick the field type and category that the policy doc lists.
- Write a description that states what the variable does and its units.
- To rename, keep the old name in `deprecated_names`. To remove, add the
name to `DEPRECATED_VARIABLES`. Update the uses in `examples/`,
`scripts/`, `docs/`, `apps/`, and the tests.
2. **Read the variable with `envs.NAME.get()` inside a function.**
- In tests, change the value with `envs.NAME.override(value)`, and a variable
outside the registry with `envs.override_external(name, value)`; the
`env_overrides` fixture keeps either until the end of the test.
- Name a variable that only tests read `FASTVIDEO_TEST_*`.
- Do not call `os.environ`, `os.getenv`, or `monkeypatch.setenv` for a
FastVideo variable.
- To set a variable that another tool reads, call `envs.set_external`,
`envs.setdefault_external`, or `envs.unset_external`.
3. **Regenerate the table in the policy doc.**
- Run `python fastvideo/tests/contract/test_env_policy.py`.
4. **Run the contract test.**
- Run `pytest fastvideo/tests/contract/test_env_policy.py`.
- When the test reports a fixed known violation, delete or lower its entry
in `KNOWN_VIOLATIONS`. Never add an entry to `KNOWN_VIOLATIONS`.
5. **When the policy itself changes, update the policy doc and the contract
test in the same pull request.**
- The rules in `docs/contributing/env_vars.md`, the checks and allowlist in
`fastvideo/tests/contract/test_env_policy.py`, and this skill must agree.
## Outputs
- A registry entry in `fastvideo/envs.py` and call sites that use
`envs.NAME.get()`.
- A regenerated table in `docs/contributing/env_vars.md`.
- A passing `fastvideo/tests/contract/test_env_policy.py`.
## Example Usage
```
Add a FASTVIDEO_DEBUG_MY_STAGE switch that logs MyStage inputs.
```
## References
- `docs/contributing/env_vars.md`: the policy, the field types, and the
violation kinds that the contract test reports.
- `fastvideo/envs.py`: the registry.
- `fastvideo/tests/contract/test_env_policy.py`: the contract test,
`EXTERNAL_ALLOWLIST`, and `KNOWN_VIOLATIONS`.
@@ -1,20 +1,23 @@
---
name: reseed-performance-baseline
description: Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, or environment-caused benchmark shift using one or more reviewed normalized performance JSONs. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline in FastVideo/performance-tracking must be advanced from a consistent batch of reviewed source results. The workflow backs up existing history under /tmp, validates all source JSONs for the same (model_id, gpu_type), rejects internally inconsistent source batches, uploads one success=true reseed record per accepted source JSON, and offers to clean local temp state after a successful upload.
description: Re-seed the HF performance-tracking baseline for an intentional runtime, dependency, environment-caused benchmark shift, or reviewed v2 calibration using one or more reviewed normalized performance JSONs. Use when performance CI fails because metrics such as latency, throughput, component time, or peak memory changed for an accepted reason and the rolling median baseline in FastVideo/performance-tracking must be advanced, or when a new v2 exact comparable identity needs its first approved baseline. The workflow backs up existing history under /tmp, validates all source JSONs for the same legacy (model_id, gpu_type) target or the same v2 exact identity, rejects internally inconsistent source batches, uploads one success=true baseline record per accepted source JSON, and offers to clean local temp state after a successful upload.
---
# Re-seed Performance Baseline
## Purpose
Replace or advance the rolling performance baseline for a single
`(model_id, gpu_type)` pair in the HF dataset
`FastVideo/performance-tracking`.
Replace or advance the rolling performance baseline in the HF dataset
`FastVideo/performance-tracking`. Legacy targets are scoped by
`(model_id, gpu_type)`. V2 targets are scoped by exact comparable identity:
`workload_id`, `variant_id`, `benchmark_version`, `hardware_profile_id`,
`software_profile_id`, and `recipe_fingerprint`.
Performance comparison uses the median of up to the last 5 successful records
for the same model and GPU. Failed records are useful audit history, but they
do not move the future baseline because `compare_baseline.py` loads records
with `successful_only=True`.
Performance comparison uses the median of up to the last 5 successful,
baseline-eligible records for the same target. Failed or calibration-only
records are useful audit history, but they do not move the future baseline
because `compare_baseline.py` loads records with `successful_only=True` and
`baseline_eligible_only=True`.
This skill now reseeds from a reviewed batch of one or more source performance
JSONs. It uploads one new `success=true` record per accepted source JSON; it
@@ -22,11 +25,13 @@ does not blindly replicate one measurement into 3 or 5 records. The effective
reseed size is therefore dynamic and equals the number of provided, validated,
internally consistent source JSONs.
If the operator provides fewer than 3 records, call out that the last-5 rolling
median may not move immediately. If the operator provides 3 consistent shifted
records, the rolling median usually moves immediately. If the operator provides
5 consistent shifted records, the last-5 window is effectively reset to the new
runtime profile.
For baseline shifts with existing history, if the operator provides fewer than
3 records, call out that the last-5 rolling median may not move immediately. If
the operator provides 3 consistent shifted records, the rolling median usually
moves immediately. If the operator provides 5 consistent shifted records, the
last-5 window is effectively reset to the new runtime profile. For the first
approved v2 baseline of a new exact identity, one reviewed calibration seed is
enough for the next comparable run to leave `CALIBRATION_NEEDED`.
These records are intentional operator-approved baseline resets, not ordinary
independent main-branch persistence. Mark them clearly with provenance fields
@@ -66,8 +71,8 @@ approval, then upload reviewed accepted baseline records.
| Parameter | Required | Description |
|-----------|----------|-------------|
| `model_id` | Yes | Benchmark id, e.g. `wan-t2v-1.3b-2gpu`. This maps to the HF subdirectory after `sanitize(model_id)`. |
| `gpu_type` | Yes | Exact GPU device string from the performance record, e.g. the L40S device name emitted by CI. Baselines are GPU-specific. |
| `model_id` | Legacy required; v2 inferred | Benchmark id, e.g. `wan-t2v-1.3b-2gpu`. This maps to the HF subdirectory after `sanitize(model_id)`. For v2 records, use the `model_id` from each source artifact only as the upload directory; comparison is by exact identity. |
| `gpu_type` | Legacy required; v2 inferred | Exact GPU device string from the performance record, e.g. the L40S device name emitted by CI. V2 hardware matching uses `hardware_profile_id`; preserve `gpu_type` as display metadata. |
| `source_results` | Yes | One or more local paths or Buildkite artifact URLs for accepted shifted performance JSONs. Prefer normalized `normalized_perf_*.json` artifacts emitted by `compare_baseline.py`. Accept `source_result` as an alias only for a single JSON. |
| `max_intra_batch_regression` | No | Maximum allowed regression of any source JSON against the source batch median. Default: `0.05` (5%). |
| `intent_rationale` | Yes | One-line explanation for why the baseline shift is legitimate. This is written into provenance and should be reused in the PR. |
@@ -78,10 +83,14 @@ Hardcoded defaults:
supported by the code, but use the default unless the user explicitly asks).
- Local sync root: `/tmp/perf-tracking` (`PERFORMANCE_TRACKING_ROOT` override
is supported).
- Prepared-record staging root: `/tmp/performance_reseed_prepared`
(`PERFORMANCE_RESEED_STAGING_ROOT` override is supported). Keep it separate
and non-nested from the sync root.
- Backup root: `/tmp/performance_reseed_backup`.
- Download scratch root for source artifact URLs: `/tmp/performance_reseed_source`.
- Baseline window: last 5 `success=true` records for the same
`(model_id, gpu_type)`.
- Baseline window: last 5 `success=true`, `baseline_eligible=true` records
for the same legacy `(model_id, gpu_type)` target or the same v2 exact
comparable identity.
- Reseed count: dynamic. Upload exactly one accepted seed record per validated
source JSON.
@@ -115,12 +124,24 @@ with open(source_result, encoding="utf-8") as f:
record = json.load(f)
```
Stop if any normalized record's `model_id` or `gpu_type` does not match the
requested `model_id` and `gpu_type`.
Classify the source batch before continuing:
- **Legacy source records** have no v2 exact identity fields. Stop if any
normalized record's `model_id` or `gpu_type` does not match the requested
`model_id` and `gpu_type`.
- **V2 source records** have exact identity fields. Stop unless every source
record has all six comparable identity fields and they are identical across
the batch: `workload_id`, `variant_id`, `benchmark_version`,
`hardware_profile_id`, `software_profile_id`, and `recipe_fingerprint`.
Do not fall back to legacy `(model_id, gpu_type)` matching for v2 records.
The source records may have `success: false` when they came from failed
rolling baseline comparisons. That is expected; only the reviewed reseed
records become new `success: true` baseline records after explicit approval.
For a first v2 baseline seed, the source records must instead be successful
scheduled-main full-suite `CALIBRATION_NEEDED` normalized artifacts. Reject PR,
local, direct-run, non-main-branch, or non-full-suite calibration artifacts as
seed sources.
Sort validated source records by their original `timestamp` ascending before
preparing the seed records. If a source timestamp is missing or unparsable,
@@ -194,7 +215,7 @@ export HF_REPO_ID="${HF_REPO_ID:-FastVideo/performance-tracking}"
python -c 'from fastvideo.performance.hf_store import sync_from_hf; import os; sync_from_hf(os.environ["PERFORMANCE_TRACKING_ROOT"], strict=True)'
```
Then back up only the sanitized model directory under `/tmp`:
For legacy records, back up the sanitized model directory under `/tmp`:
```bash
SHORT_COMMIT=$(git rev-parse --short=12 HEAD)
@@ -209,6 +230,16 @@ mkdir -p "$BACKUP_DIR"
cp -R "${PERFORMANCE_TRACKING_ROOT}/${MODEL_SAFE}" "$BACKUP_DIR/" 2>/dev/null || true
```
For v2 records, back up the full local tracking root after sync. Exact identity
lookup scans across model directories, so a benchmark rename may have relevant
history outside the current source artifact's `model_id` directory:
```bash
BACKUP_DIR="/tmp/performance_reseed_backup/${TIMESTAMP}_${SHORT_COMMIT}_v2_exact_identity"
mkdir -p "$BACKUP_DIR"
cp -R "${PERFORMANCE_TRACKING_ROOT}" "$BACKUP_DIR/tracking-root"
```
Write provenance next to the backup:
```bash
@@ -231,7 +262,9 @@ first baseline seed. Continue, but report that baseline history was empty.
### 3. Compute old baseline and candidate shift
Load the last 5 successful records for the target:
Load the last 5 successful baseline records for the target.
For legacy targets:
```python
from fastvideo.performance.hf_store import load_records_for_model
@@ -242,6 +275,28 @@ records = load_records_for_model(
"<gpu_type>",
last_n=5,
successful_only=True,
baseline_eligible_only=True,
)
```
For v2 exact-identity targets:
```python
from fastvideo.performance.hf_store import load_records_for_identity
records = load_records_for_identity(
"/tmp/perf-tracking",
{
"workload_id": "<workload_id>",
"variant_id": "<variant_id>",
"benchmark_version": "<benchmark_version>",
"hardware_profile_id": "<hardware_profile_id>",
"software_profile_id": "<software_profile_id>",
"recipe_fingerprint": "<recipe_fingerprint>",
},
last_n=5,
successful_only=True,
baseline_eligible_only=True,
)
```
@@ -257,7 +312,8 @@ medians after appending the proposed seed records, and source batch spread for:
Also print how many successful old records exist. Make clear:
- 1 seed record usually does not move a last-5 median by itself.
- 1 seed record usually does not move an existing last-5 median by itself, but
it is enough to establish the first v2 baseline for a new exact identity.
- 3 consistent seed records usually move the last-5 median immediately.
- 5 consistent seed records effectively reset the last-5 window.
- The records are intentional approved baseline resets and must be labeled
@@ -267,10 +323,10 @@ Also print how many successful old records exist. Make clear:
Require an explicit confirmation phrase before preparing the upload:
> About to RE-SEED performance baseline for `<model_id>` on `<gpu_type>`.
> About to RE-SEED performance baseline for `<target description>`.
> This will upload `<N>` new `success=true` records to
> `FastVideo/performance-tracking/<sanitize(model_id)>/`, one per accepted
> source JSON.
> `FastVideo/performance-tracking/<sanitize(model_id)>/` or the source
> artifact's v2 model directory, one per accepted source JSON.
>
> Reason: `<intent_rationale>`
> Source results: `<source_results>`
@@ -288,25 +344,63 @@ Do not continue unless the user types exactly `confirm performance reseed`.
### 5. Create the accepted seed records
Create one seed record from each normalized source result. Do not copy the
Create one seed record from each normalized source result.
For first v2 baseline seeds, use the scoped utility. It validates exact
identity, requires successful scheduled-main full-suite `CALIBRATION_NEEDED`
source artifacts, preserves the normalized v2 identity and metadata fields,
and writes seed records with `success=true`, `baseline_eligible=true`, and
`comparison_status=PASS`:
```bash
python fastvideo/tests/performance/seed_baseline.py \
--source-result <normalized_perf_1.json> \
--source-result <normalized_perf_2.json> \
--intent-rationale "<intent_rationale>" \
--max-intra-batch-regression 0.05 \
--tracking-root "${PERFORMANCE_TRACKING_ROOT}" \
--staging-root "${PERFORMANCE_RESEED_STAGING_ROOT:-/tmp/performance_reseed_prepared}"
```
The utility is prepare-only and intentionally has no upload option. Upload the
scoped records only after the separate confirmation in step 6.
The utility validates against an isolated fresh HF snapshot and leaves
`PERFORMANCE_TRACKING_ROOT` untouched; that argument only proves the staging
root is separate from the operator's tracking mirror. Before writing, it stops
if the exact identity already has a successful baseline-eligible record or if
the workload/variant/version already trusts another recipe. It atomically
reserves the exact identity and writes a digest-protected upload manifest bound
to the current HF endpoint, repository id, and repository type. Keep the
prepared records, manifest, source files, and reservation unchanged until the
operation is uploaded or explicitly cleaned up.
If the prepared seed records look correct, upload only those scoped records in
step 7. Do not rerun the utility with a different source list after approval.
For legacy reseeds or accepted v2 baseline shifts from regression artifacts,
create one seed record from each normalized source result. Do not copy the
source JSON wholesale.
Infer the baseline field allowlist from all existing HF records for the target
`(model_id, gpu_type)` after syncing, including both `success=true` and
`success=false` records. Use the union of non-provenance keys present in those
target records, preserving only fields that also exist in the normalized
source record or are explicitly set by the reseed workflow. Always include
`model_id`, `timestamp`, and `success` because the upload path and baseline
loader depend on them. Always set `timestamp` to a fresh reseed timestamp and
`success` to `true`. Do not include unrelated source-only fields that are
absent from existing HF records.
after syncing, including both `success=true` and `success=false` records. For
legacy targets the target is `(model_id, gpu_type)`. For v2 baseline-shift
reseeds the target is the exact comparable identity. Use the union of
non-provenance keys present in those target records, preserving only fields
that also exist in the normalized source record or are explicitly set by the
reseed workflow. Always include `model_id`, `timestamp`, `success`,
`baseline_eligible`, and `comparison_status` because the upload path and
baseline loader depend on them. Always set `timestamp` to a fresh reseed
timestamp, `success` to `true`, `baseline_eligible` to `true`, and
`comparison_status` to `PASS`. Do not include unrelated source-only fields
that are absent from existing HF records.
Exclude existing provenance or operator metadata from the inferred baseline
field allowlist. At minimum, exclude keys prefixed with `baseline_reseed` and
any fields known to be local-only audit metadata.
If there are no previous HF records for the target model/GPU, fall back to this
default baseline field list:
If there are no previous HF records for the target, fall back to this default
baseline field list:
- `model_id`
- `timestamp`
@@ -319,6 +413,22 @@ default baseline field list:
- `dit_time_s`
- `vae_decode_time_s`
- `success`
- `baseline_eligible`
- `comparison_status`
For v2 baseline-shift reseeds with no previous HF records for the exact
identity, also preserve:
- `workload_id`
- `variant_id`
- `benchmark_version`
- `recipe_fingerprint`
- `hardware_profile_id`
- `software_profile_id`
- `recipe`
- `hardware_profile`
- `software_profile`
- `software_comparison_profile`
Do not upload extra fields from the source artifact.
@@ -334,6 +444,22 @@ Optional provenance fields are allowed and useful:
- `baseline_reseed_operator`
- `baseline_reseed_max_intra_batch_regression`
The v2 calibration seed utility writes analogous first-seed provenance:
- `baseline_seed: true`
- `baseline_seed_reason`
- `baseline_seed_source_result`
- `baseline_seed_source_status`
- `baseline_seed_source_timestamp`
- `baseline_seed_source_success`
- `baseline_seed_source_run_source`
- `baseline_seed_source_branch`
- `baseline_seed_source_test_scope`
- `baseline_seed_source_pr_number`
- `baseline_seed_batch_size`
- `baseline_seed_batch_index`
- `baseline_seed_operator`
Use a fresh reseed timestamp for each seed record, not the original source
result timestamp. This is required because
`load_records_for_model(..., last_n=5)` keeps the last records after loading
@@ -356,7 +482,8 @@ Prefer uploading new accepted seed records so failed history remains visible.
Print:
- Backup directory path under `/tmp`.
- Prepared local record paths under `PERFORMANCE_TRACKING_ROOT`.
- Prepared local record paths under `PERFORMANCE_RESEED_STAGING_ROOT`.
- Prepared upload-manifest path under the identity reservation.
- HF paths that will receive the new records.
- Old rolling medians.
- Source batch medians, source batch spread, reseed count, and candidate
@@ -368,22 +495,36 @@ prepared records plus backup on disk.
### 7. Upload only the scoped records
Use the shared storage helper so the path and repo type match CI:
For a first v2 calibration seed, use the manifest uploader after the user
replies exactly `upload`:
```python
from fastvideo.performance.hf_store import upload_record
upload_record("<local_record_path>", record, strict=True)
```bash
python -c 'from fastvideo.tests.performance.seed_baseline import upload_prepared_seed_manifest; print(upload_prepared_seed_manifest("<prepared_manifest>"))'
```
Run it once per prepared record. Each upload goes to:
The uploader verifies the source and prepared-record digests, pins and scans
the current HF revision, rechecks exact-identity and recipe-cohort conflicts,
and writes the entire batch in one commit whose `parent_commit` must still be
current. A concurrent Hub update makes the commit fail. Do not retry
automatically: preserve staging, refresh/review remote state, and request a new
explicit `upload` after the conflict is understood. Each record goes to:
```text
FastVideo/performance-tracking/<sanitize(model_id)>/<record_filename>.json
```
Never bulk upload the whole tracking root. Never modify another model's
directory in the same operation.
Never call `upload_record()` once per first-seed record: that can partially
land the batch and has no compare-and-swap guard.
For a legacy reseed or an accepted v2 baseline shift, the first-seed manifest
validator does not apply because an eligible baseline already exists. Upload
only the individually reviewed records prepared in step 5 with the shared
`upload_record(local_path, record, strict=True)` helper. Stop on the first
failure and report exactly which records reached HF; do not silently rerun or
replicate the remainder.
Never bulk upload the tracking or staging root, and never modify another
model's directory in the same operation.
### 8. Report outcome and offer cleanup
@@ -405,9 +546,14 @@ distinguish an accepted baseline shift from a hidden regression.
After the upload is verified, ask whether the user wants to clear temporary
local state. Explain what each directory is for:
- `PERFORMANCE_TRACKING_ROOT`, usually `/tmp/perf-tracking`: local synced
mirror of `FastVideo/performance-tracking` plus the prepared local seed
records used for scoped upload.
- `PERFORMANCE_TRACKING_ROOT`, usually `/tmp/perf-tracking`: read-only local
synced mirror used for operator review and reporting. First-v2 preparation
independently proves remote state from a fresh temporary HF snapshot.
- `PERFORMANCE_RESEED_STAGING_ROOT`, usually
`/tmp/performance_reseed_prepared`: prepared local seed records used for the
scoped upload, plus the identity reservation and digest manifest. Keeping
this separate prevents aborted preparations from appearing in later
baseline reads.
- `/tmp/performance_reseed_backup/<...>`: local backup of the target model's
pre-reseed HF history plus `PROVENANCE.txt`, kept so a bad reseed can be
audited or corrected.
@@ -417,14 +563,18 @@ local state. Explain what each directory is for:
Ask:
> Reseed succeeded. Do you want me to delete the local temp tracking mirror,
> source downloads, and reseed backup under `/tmp`? These files are local
> safety/audit artifacts only; HF already has the uploaded records.
> this reseed's prepared staging records, source downloads, and reseed backup
> under `/tmp`? These files are local safety/audit artifacts only; HF already
> has the uploaded records.
>
> Reply `cleanup reseed temp` to delete them, anything else to keep them.
Do not delete anything unless the user replies exactly
`cleanup reseed temp`. If cleanup is requested, remove only the specific
directories created for this reseed. Never remove unrelated `/tmp` contents.
directories and prepared record paths created for this reseed. Do not remove
the shared staging root when it contains other records. Remove this operation's
identity reservation only with its prepared records and manifest, and never
remove unrelated `/tmp` contents.
## Failure modes and handling
@@ -436,19 +586,34 @@ directories created for this reseed. Never remove unrelated `/tmp` contents.
against the source batch median by more than `max_intra_batch_regression`.
Ask for cleaner sources or a reviewed explanation before continuing.
- **Too few source records to move the median.** Continue only after making
clear that one or two records may not immediately move the last-5 median.
clear that one or two records may not immediately move an existing last-5
median. This warning does not block a first v2 calibration seed for an exact
identity with no eligible baseline yet.
- **The source results are noisy or suspicious.** Stop. Reseeding amplifies
those measurements into the baseline, so they must be reviewed first.
- **HF sync fails.** Stop for destructive reseeds. A stale or empty sync can
make the old baseline look missing.
- **The exact v2 identity already has an eligible baseline.** Stop. The
`CALIBRATION_NEEDED` artifact is stale; use the reviewed baseline-shift path
instead of the first-seed utility.
- **The workload/variant/version trusts another recipe.** Stop. The source is
stale relative to the current recipe cohort and must not bypass
`RECIPE_MISMATCH` by creating a second trusted recipe.
- **The staging root already has a prepared seed for the exact identity.**
Stop and reuse, upload, or explicitly clean that preparation. Do not prepare
another copy of the same measurement.
- **The conditional Hub commit loses its parent race.** Stop without retrying.
Keep the preparation, refresh and review the new remote state, then request
a new explicit `upload` only if the seed is still valid.
- **Candidate still violates fixed thresholds.** Report that this skill only
handles the rolling HF baseline; update benchmark JSON thresholds in code
review if maintainers accept the new absolute limit.
- **The user aborts at either confirmation.** Leave the backup and prepared
records on disk. Nothing should be uploaded.
- **The user declines cleanup.** Keep `/tmp/perf-tracking`, the source
download directory if any, and `/tmp/performance_reseed_backup/<...>` in
place for audit/debugging.
- **The user declines cleanup.** Keep `/tmp/perf-tracking`, the prepared seed
records under `/tmp/performance_reseed_prepared`, the source download
directory if any, and `/tmp/performance_reseed_backup/<...>` in place for
audit/debugging.
- **A bad seed was uploaded.** Use the backup and HF history to identify the
uploaded file, then remove or supersede it with an explicitly reviewed
corrective record. Do not silently rewrite unrelated history.
@@ -459,8 +624,9 @@ directories created for this reseed. Never remove unrelated `/tmp` contents.
intentional baseline replacement.
- `fastvideo/tests/performance/compare_baseline.py` — normalization, rolling
median comparison, and persistence rules.
- `fastvideo/performance/hf_store.py` — HF sync, record loading,
`sanitize()`, and `upload_record()`.
- `fastvideo/performance/hf_store.py` — HF sync and record loading helpers.
- `fastvideo/tests/performance/seed_baseline.py` — first-seed preparation,
staging reservation, manifest validation, and conditional batch upload.
- `fastvideo/tests/performance/test_inference_performance.py` — source result
JSON schema.
- `.buildkite/performance-benchmarks/tests/*.json` — fixed absolute benchmark
@@ -473,3 +639,4 @@ directories created for this reseed. Never remove unrelated `/tmp` contents.
| 2026-05-03 | Initial version. Sister workflow to `reseed-ssim-references`, scoped to one performance `(model_id, gpu_type)` baseline seed with backup, confirmation, provenance, and `success=true` upload. |
| 2026-05-03 | Previous policy: replicate one approved shifted source result into 3 success records by default, or 5 only when explicitly requested. Add provenance marker for replicated-source reseeds. Superseded by the 2026-05-08 dynamic multi-source policy. |
| 2026-05-08 | Replace fixed 3/5 replication with dynamic multi-source reseeding: upload one seed record per reviewed source JSON, validate intra-batch consistency, move backup/source scratch under `/tmp`, and ask whether to clean temp state after successful upload. |
| 2026-07-13 | Keep first-v2-seed preparation outside the canonical mirror, reserve staging identities atomically, reject stale or replayed calibration seeds, and upload reviewed manifests with a single parent-guarded Hub commit. |
@@ -1,6 +1,6 @@
---
name: reseed-ssim-references
description: Re-seed HF reference videos for a single existing SSIM test on Modal L40S. Always backs up current refs locally first, regenerates on Modal, pauses for the user to eyeball before-vs-after quality, then overwrites the targeted `<model_id>` subtree on `FastVideo/ssim-reference-videos` with `--force`. Use when an intentional code change (model port fix, attention backend swap, kernel upgrade, hyperparameter change) has invalidated existing refs and they need to be regenerated. Pairs with `seed-ssim-references`, which is for first-time seeding only.
description: Re-seed HF reference videos for a single existing SSIM test on Modal L40S. Always backs up current refs locally first, regenerates on Modal, pauses for the user to eyeball before-vs-after quality, then overwrites the targeted model subtree on `FastVideo/ssim-reference-videos` with `--force`. Use when an intentional code change (model port fix, attention backend swap, kernel upgrade, hyperparameter change) has invalidated existing refs and they need to be regenerated. Pairs with `seed-ssim-references`, which is for first-time seeding only.
---
# Re-seed SSIM Reference Videos
@@ -13,7 +13,7 @@ on HF — the old refs are overwritten — so the skill always:
1. Confirms intent with a one-liner the user has to type.
2. Downloads the existing refs as a local, timestamped backup.
3. Regenerates on Modal L40S (same code path that CI uses).
3. Regenerates through the manual legacy Modal L40S maintenance path.
4. Pauses for a side-by-side eyeball of backup vs new mp4s.
5. Uploads with `--force`, scoped to the single `--model-id`.
6. Reminds the user to keep the backup until the PR lands.
@@ -51,12 +51,13 @@ harder to recover from than failing closed.
Hardcoded:
- Modal GPU: **L40S** (matches CI; re-seeding from another SKU produces refs
that L40S CI cannot match).
- Modal GPU: **L40S**. This is a manual reference-maintenance target, not the
active Slurm CI compute path; changing the SKU also changes the historical
`L40S_reference_videos` contract.
- Quality tier: **`default`**. `full_quality` is a separate, deliberate
operation.
- HF repo: `FastVideo/ssim-reference-videos` (override via
`FASTVIDEO_SSIM_REFERENCE_HF_REPO`).
`FASTVIDEO_TEST_SSIM_REFERENCE_HF_REPO`).
- Device folder: `L40S_reference_videos`.
## Prerequisites
+4 -2
View File
@@ -35,7 +35,8 @@ The skill is run **manually**, once per new test. Before invoking it, the user
has already sanity-tested the new test locally — it launches `VideoGenerator`
and writes an artefact without crashing (the missing-reference assertion at
the end is expected). The skill does not re-test locally; it goes straight
to Modal L40S (which is what CI uses).
to the manual legacy Modal L40S reference-maintenance target. Active CI runs
on the Slinky Slurm cluster and only consumes the resulting references.
## When to use
@@ -61,7 +62,8 @@ Prompt the user for it if they didn't supply it.
Everything else is fixed:
- Modal runner GPU: **L40S** (hardcoded in `fastvideo/tests/modal/ssim_test.py`).
- Modal maintenance GPU: **L40S** (hardcoded in
`fastvideo/tests/modal/ssim_test.py`; this is not the active CI compute path).
- Device folder: `L40S_reference_videos`.
- Quality tier: `default` (the tier CI runs). The `full_quality` tier is not
seeded by this skill.
@@ -0,0 +1,51 @@
{
"benchmark_id": "wan-t2v-1.3b-1gpu-gb10",
"config_schema_version": 2,
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp1",
"benchmark_version": 3,
"description": "Wan2.1 T2V 1.3B single-GPU inference performance on NVIDIA DGX Spark (GB10). Single-GPU variant of wan-t2v-1.3b (same workload_id for dashboard comparability). Gated to the GB10 via run_config.gpu_types so it does not run on the shared H100/L40S lanes.",
"model": {
"model_path": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
"model_short_name": "Wan2.1-T2V-1.3B"
},
"init_kwargs": {
"num_gpus": 1,
"flow_shift": 7.0,
"sp_size": 1,
"tp_size": 1,
"vae_sp": false,
"vae_tiling": true,
"text_encoder_precisions": ["fp32"]
},
"generation_kwargs": {
"height": 480,
"width": 832,
"num_frames": 45,
"num_inference_steps": 4,
"guidance_scale": 3,
"embedded_cfg_scale": 6,
"seed": 1024,
"fps": 24,
"neg_prompt": "Bright tones, overexposed, static, blurred details, subtitles, style, works, paintings, images, static, overall gray, worst quality, low quality, JPEG compression residue, ugly, incomplete, extra fingers, poorly drawn hands, poorly drawn faces, deformed, disfigured, misshapen limbs, fused fingers, still picture, messy background, three legs, many people in the background, walking backwards"
},
"test_prompts": [
"Will Smith casually eats noodles, his relaxed demeanor contrasting with the energetic background of a bustling street food market. The scene captures a mix of humor and authenticity. Mid-shot framing, vibrant lighting."
],
"run_config": {
"num_warmup_runs": 2,
"num_measurement_runs": 5,
"required_gpus": 1,
"gpu_types": ["GB10"]
},
"thresholds": {
"GB10": {
"max_generation_time_s": 55.0,
"max_peak_memory_mb": 12000.0
},
"default": {
"max_generation_time_s": 120.0,
"max_peak_memory_mb": 40000.0
}
}
}
@@ -3,7 +3,7 @@
"config_schema_version": 2,
"workload_id": "wan-t2v",
"variant_id": "1.3b-sp2",
"benchmark_version": 2,
"benchmark_version": 3,
"description": "Wan2.1 T2V 1.3B inference performance",
"model": {
"model_path": "Wan-AI/Wan2.1-T2V-1.3B-Diffusers",
+465 -515
View File
@@ -1,6 +1,9 @@
env:
IMAGE_VERSION: "py3.12-latest"
BUILDKITE_CLEAN_CHECKOUT: true
# Slurm workers clone the immutable commit and initialize submodules inside
# their isolated container. The Buildkite login-plane checkout is a no-op.
BUILDKITE_GIT_SUBMODULES: false
notify:
- github_commit_status:
@@ -8,527 +11,474 @@ notify:
if: build.env("TEST_SCOPE") == "fastcheck" || build.env("TEST_SCOPE") == null
- github_commit_status:
context: "full-suite-passed"
if: build.env("TEST_SCOPE") == "full"
if: build.env("TEST_SCOPE") == "full" || build.env("TEST_SCOPE") == "merge"
- github_commit_status:
context: "direct-test-completed"
if: build.env("TEST_SCOPE") == "direct"
- github_commit_status:
context: "scheduled-ssim-passed"
if: build.env("TEST_SCOPE") == "scheduled"
# This is the complete active GPU CI surface. Every command is a trusted host
# dispatcher, and every test payload executes inside the Slinky Slurm tray.
# fastvideo/tests/modal remains available only for an explicit manual rollback;
# no active pipeline or slash-command route invokes it.
# Buildkite hands jobs to free agents in the order they appear here. Golden-gate comes first
# because every later merge lane waits for it; the fastcheck lanes follow from longest to
# shortest measured runtime, so the longest lane never starts last and stretches the build.
steps:
# ============================================================
# Direct test: triggered by /test <name> slash command.
# Labels match fastcheck/full-suite counterparts so the GitHub
# check status overwrites the original failed check.
# Only ONE step executes per build (gated by TEST_TYPE).
# ============================================================
- label: ":test_tube: Golden-Gate Tests"
key: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,golden-gate,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "golden_gate" || build.env("TEST_TYPE") == "golden_gate_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "golden_gate_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
# --- Fastcheck-scope direct tests ---
- label: ":microscope: Encoder Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "encoder"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: VAE Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "vae"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Transformer Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "transformer"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Kernel Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "kernel_tests"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Unit Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "unit_test"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: DreamVerse App Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "dreamverse_app"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Unit Tests"
key: "unit"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,unit,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "unit_test" || build.env("TEST_TYPE") == "unit_test_ci"))
command: "/opt/fastvideo-ci-runner/run-unit"
timeout_in_minutes: 90
env:
TEST_TYPE: "unit_test_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
# --- Full-suite-scope direct tests ---
- label: ":bar_chart: SSIM Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "ssim"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "default"
- label: ":test_tube: LoRA Inference Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "inference_lora"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: LoRA Extraction Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "lora_extraction"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Training Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Distillation DMD Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "distillation_dmd"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Self-Forcing Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "self_forcing"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: LoRA Training Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training_lora"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Training Tests VSA"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training_vsa"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Inference Tests VMoBA"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "inference_vmoba"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Performance Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "performance"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: API Server Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "api_server"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Train Framework Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "train_framework"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Eval Metrics Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "eval"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Kernel Tests"
key: "kernel-tests"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,kernel-tests,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "kernel_tests" || build.env("TEST_TYPE") == "kernel_tests_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "kernel_tests_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
# ============================================================
# Fastcheck: Runs on every PR (~10-15 min parallel)
# Core component validation: encoders, VAEs, transformers,
# CUDA kernels, and unit tests.
# ============================================================
- label: "Trigger Fastcheck"
if: build.env("TEST_SCOPE") == "fastcheck" || build.env("TEST_SCOPE") == null
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
plugins:
- monorepo-diff#v1.4.0:
diff: 'git fetch origin "${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}" && git diff --name-only "origin/${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}...HEAD"'
watch:
- path:
- "fastvideo/models/encoders/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/encoders/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: ":microscope: Encoder Tests"
env:
- TEST_TYPE=encoder
agents:
queue: "default"
- path:
- "fastvideo/models/vaes/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/vaes/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: ":microscope: VAE Tests"
env:
- TEST_TYPE=vae
agents:
queue: "default"
- path:
- "fastvideo/models/dits/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/transformers/**"
- "fastvideo/layers/**"
- "fastvideo/attention/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Transformer Tests"
env:
- TEST_TYPE=transformer
agents:
queue: "default"
- path:
- "fastvideo-kernel/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Kernel Tests"
env:
- TEST_TYPE=kernel_tests
agents:
queue: "default"
- path:
- "fastvideo/**"
- ".buildkite/**"
- ".github/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Unit Tests"
env:
- TEST_TYPE=unit_test
agents:
queue: "default"
- path:
- "apps/dreamverse/**"
- "pyproject.toml"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":microscope: DreamVerse App Tests"
env:
- TEST_TYPE=dreamverse_app
agents:
queue: "default"
- label: ":microscope: DreamVerse App Tests"
key: "dreamverse"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,dreamverse,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "dreamverse_app" || build.env("TEST_TYPE") == "dreamverse_app_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "dreamverse_app_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
# ============================================================
# Full Suite: Runs when TEST_SCOPE=full
# Triggered by adding the 'ready' label (via ci-trigger-full-suite.yml)
# or on-demand via /test full slash command.
# Includes integration tests, SSIM regression, training pipelines,
# and performance benchmarks.
# ============================================================
- label: "Trigger Full Suite"
if: build.env("TEST_SCOPE") == "full"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
plugins:
- monorepo-diff#v1.4.0:
diff: 'git fetch origin "${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}" && git diff --name-only "origin/${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}...HEAD"'
watch:
- path:
- "fastvideo/**/*.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":bar_chart: SSIM Tests"
env:
- TEST_TYPE=ssim
retry:
automatic:
- exit_status: 1
limit: 2
agents:
queue: "default"
- path:
- "fastvideo/tests/lora/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/transformers/**"
- "fastvideo/pipelines/**"
- "fastvideo/layers/lora/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Inference Tests"
env:
- TEST_TYPE=inference_lora
agents:
queue: "default"
- path:
- "scripts/lora_extraction/**"
- "fastvideo/tests/lora_extraction/**"
- "fastvideo/models/loader/**"
- "fastvideo/training/training_utils.py"
- "fastvideo/layers/lora/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Extraction Tests"
env:
- TEST_TYPE=lora_extraction
agents:
queue: "default"
- path:
- "fastvideo/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Training Tests"
env:
- TEST_TYPE=training
agents:
queue: "default"
- path:
- "fastvideo/training/*distillation_pipeline.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Distillation DMD Tests"
env:
- TEST_TYPE=distillation_dmd
agents:
queue: "default"
- path:
- "fastvideo/training/*self_forcing_distillation_pipeline.py"
- "fastvideo/tests/training/self-forcing/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Self-Forcing Tests"
env:
- TEST_TYPE=self_forcing
agents:
queue: "default"
- path:
- "fastvideo/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Training Tests"
env:
- TEST_TYPE=training_lora
retry:
automatic:
- exit_status: 1
limit: 2
agents:
queue: "default"
- path:
- "fastvideo/**"
- "fastvideo-kernel/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Training Tests VSA"
env:
- TEST_TYPE=training_vsa
retry:
automatic:
- exit_status: 1
limit: 2
agents:
queue: "default"
- path:
- "fastvideo-kernel/**"
- "fastvideo/attention/backends/vmoba.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Inference Tests VMoBA"
env:
- TEST_TYPE=inference_vmoba
agents:
queue: "default"
- path:
- "fastvideo/models/dits/**"
- "fastvideo/pipelines/**"
- "fastvideo/attention/**"
- "fastvideo/layers/**"
- "fastvideo/worker/**"
- "fastvideo/entrypoints/**"
- "fastvideo/tests/performance/**"
- ".buildkite/performance-benchmarks/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Performance Tests"
env:
- TEST_TYPE=performance
agents:
queue: "default"
- path:
- "fastvideo/entrypoints/openai/**"
- "fastvideo/entrypoints/cli/serve.py"
- "fastvideo/tests/entrypoints/test_openai_api_integration.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: API Server Tests"
env:
- TEST_TYPE=api_server
agents:
queue: "default"
- path:
- "fastvideo/train/**"
- "fastvideo/tests/train/models/**"
- "fastvideo/tests/train/fixtures/**"
- "fastvideo/models/dits/**"
- "fastvideo/models/loader/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Train Framework Tests"
env:
- TEST_TYPE=train_framework
agents:
queue: "default"
- path:
- "fastvideo/eval/**"
- "fastvideo/tests/eval/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Eval Metrics Tests"
env:
- TEST_TYPE=eval
agents:
queue: "default"
- label: ":microscope: Encoder Tests"
key: "encoder"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,encoder,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "encoder" || build.env("TEST_TYPE") == "encoder_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "encoder_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":microscope: VAE Tests"
key: "vae"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,vae,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "vae" || build.env("TEST_TYPE") == "vae_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "vae_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":microscope: Transformer Tests"
key: "transformer"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,transformer,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "transformer" || build.env("TEST_TYPE") == "transformer_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "transformer_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":bar_chart: SSIM Tests"
key: "ssim"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
build.env("TEST_SCOPE") == "scheduled" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,ssim,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "ssim" || build.env("TEST_TYPE") == "ssim_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
concurrency: 1
concurrency_group: "fastvideo/slinky/whole-tray"
env:
TEST_TYPE: "ssim_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: LoRA Inference Tests"
key: "lora-inference"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,lora-inference,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "inference_lora" || build.env("TEST_TYPE") == "inference_lora_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "inference_lora_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: LoRA Extraction Tests"
key: "lora-extraction"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,lora-extraction,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "lora_extraction" || build.env("TEST_TYPE") == "lora_extraction_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "lora_extraction_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Training Tests"
key: "training"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,training,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "training" || build.env("TEST_TYPE") == "training_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
concurrency: 1
concurrency_group: "fastvideo/slinky/whole-tray"
env:
TEST_TYPE: "training_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Distillation DMD Tests"
key: "distillation"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,distillation,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "distillation_dmd" || build.env("TEST_TYPE") == "distillation_dmd_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "distillation_dmd_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Self-Forcing Tests"
key: "self-forcing"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,self-forcing,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "self_forcing" || build.env("TEST_TYPE") == "self_forcing_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "self_forcing_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: LoRA Training Tests"
key: "lora-training"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,lora-training,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "training_lora" || build.env("TEST_TYPE") == "training_lora_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "training_lora_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Training Tests VSA"
key: "training-vsa"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,training-vsa,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "training_vsa" || build.env("TEST_TYPE") == "training_vsa_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "training_vsa_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Inference Tests VMoBA"
key: "inference-vmoba"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,inference-vmoba,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "inference_vmoba" || build.env("TEST_TYPE") == "inference_vmoba_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "inference_vmoba_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Performance Tests"
key: "performance"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,performance,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "performance" || build.env("TEST_TYPE") == "performance_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "performance_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: API Server Tests"
key: "api-server"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,api-server,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "api_server" || build.env("TEST_TYPE") == "api_server_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "api_server_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Train Framework Tests"
key: "train-framework"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,train-framework,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "train_framework" || build.env("TEST_TYPE") == "train_framework_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "train_framework_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Eval Metrics Tests"
key: "eval"
depends_on: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,eval,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "eval" || build.env("TEST_TYPE") == "eval_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "eval_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the OpenAI-compatible API lane.
set -euo pipefail
exec pytest ./fastvideo/tests/entrypoints/test_openai_api_integration.py -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the distillation-DMD lane.
set -euo pipefail
exec pytest ./fastvideo/tests/training/distill/test_distill_dmd.py -vs
+87
View File
@@ -0,0 +1,87 @@
#!/usr/bin/env bash
# DreamVerse needs a GPU for import-time device resolution, but it does not
# build or exercise fastvideo-kernel. A checksummed Node archive is installed
# in the disposable Slurm container because the shared CI image is
# Python/CUDA focused.
set -euo pipefail
node_version=v22.23.2
case $(uname -m) in
aarch64 | arm64)
node_arch=arm64
node_archive_sha256=013b59cfd2819703a6f4a14ab891fc46fc2a4e3f5bcd92de3fb4929b43e35b30
;;
x86_64 | amd64)
node_arch=x64
node_archive_sha256=b294a556e639d64338823920e5866c21c02741742d2e1529ee1a225c1ec9252a
;;
*)
echo "Unsupported architecture for DreamVerse Node runtime: $(uname -m)" >&2
exit 2
;;
esac
node_archive="node-${node_version}-linux-${node_arch}.tar.gz"
node_runtime_root=$(mktemp -d -t fastvideo-node.XXXXXX)
node_archive_path="${node_runtime_root}/${node_archive}"
node_install_dir="${node_runtime_root}/${node_archive%.tar.gz}"
curl --proto '=https' --tlsv1.2 --retry 5 --retry-all-errors \
--location --fail --silent --show-error \
"https://nodejs.org/dist/${node_version}/${node_archive}" \
--output "$node_archive_path"
printf '%s %s\n' "$node_archive_sha256" "$node_archive_path" | sha256sum --check --status
tar -xzf "$node_archive_path" -C "$node_runtime_root"
export PATH="${node_install_dir}/bin:${PATH}"
node --version
npm --version
export PYTHONPATH="$(pwd)/apps/dreamverse${PYTHONPATH:+:$PYTHONPATH}"
pytest apps/dreamverse/dreamverse/tests -q
cd apps/dreamverse/web
npm ci
npm run typecheck
npm test
machine_arch=$(uname -m)
if [[ $machine_arch =~ ^(aarch64|arm64)$ ]]; then
npx playwright install --with-deps chromium firefox
else
npx playwright install --with-deps chromium webkit firefox
fi
master_port=${MASTER_PORT:-7959}
BACKEND_PORT=${BACKEND_PORT:-$((master_port + 50))}
python -m uvicorn dreamverse.mock_server:app --host 127.0.0.1 --port "$BACKEND_PORT" &
mock_server_pid=$!
cleanup() {
kill "$mock_server_pid" 2>/dev/null || true
wait "$mock_server_pid" 2>/dev/null || true
}
trap cleanup EXIT INT TERM
for _ in {1..30}; do
curl -fsS "http://127.0.0.1:$BACKEND_PORT/healthz" && break
sleep 1
done
curl -fsS "http://127.0.0.1:$BACKEND_PORT/healthz"
if [[ $machine_arch =~ ^(aarch64|arm64)$ ]]; then
# Playwright WebKit traps before opening a page on Linux ARM64, and its
# bundled Chromium lacks the H.264/AAC codecs used by the fMP4 assertions.
# Firefox covers every flow, including streaming. Chromium and its mobile
# profile still cover all codec-independent UI behavior on GB200.
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
npm run e2e -- --project=firefox
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
npm run e2e -- \
--project=chromium \
--project=mobile-chromium \
--grep-invert='streams, plays, and surfaces a downloadable clip|starts a new project and switches back to the prior session|saved projects persist across a page reload'
else
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
npm run e2e -- \
--project=chromium \
--project=webkit \
--project=firefox \
--project=mobile-safari \
--project=mobile-chromium
fi
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the encoder lane.
set -euo pipefail
exec pytest ./fastvideo/tests/encoders -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the evaluation lane.
set -euo pipefail
exec pytest ./fastvideo/tests/eval -vs
+35
View File
@@ -0,0 +1,35 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the golden-gate lane. Environment (HF_HOME
# and authentication) is the runner's responsibility.
set -euo pipefail
golden_root=./fastvideo/tests/golden_gate
selected=${FASTVIDEO_GOLDEN_TEST_FILES-}
if [ -z "$selected" ]; then
if [ "${TEST_SCOPE:-}" = merge ]; then
echo "Missing FASTVIDEO_GOLDEN_TEST_FILES for merge scope" >&2
exit 2
fi
selected=all
fi
if [ "$selected" = all ]; then
exec pytest "$golden_root" -xvs
fi
[[ $selected =~ ^test_[a-z0-9_]+\.py(,test_[a-z0-9_]+\.py)*$ ]] || {
echo "Invalid FASTVIDEO_GOLDEN_TEST_FILES selection" >&2
exit 2
}
IFS=, read -r -a golden_files <<< "$selected"
golden_paths=()
for golden_file in "${golden_files[@]}"; do
golden_path="$golden_root/$golden_file"
[ -f "$golden_path" ] || {
echo "Selected golden test does not exist: $golden_file" >&2
exit 2
}
golden_paths+=("$golden_path")
done
exec pytest "${golden_paths[@]}" -xvs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the LoRA-inference lane.
set -euo pipefail
exec pytest ./fastvideo/tests/inference/lora/test_lora_inference_similarity.py -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the VMoBA-inference lane.
set -euo pipefail
exec python fastvideo/tests/inference/vmoba/test_vmoba_inference.py
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the custom-kernel lane.
set -euo pipefail
exec pytest fastvideo-kernel/tests/ -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the LoRA-extraction lane.
set -euo pipefail
exec pytest ./fastvideo/tests/lora_extraction/ -vs
+64
View File
@@ -0,0 +1,64 @@
#!/usr/bin/env bash
# Canonical Slurm performance lane. Reports are written outside the checkout
# so the trusted host driver can upload them after untrusted code exits.
set -uo pipefail
export PERFORMANCE_TRACKING_ROOT=/tmp/perf-tracking
export PERF_REPORTS_DIR=/workspace/artifacts/performance
mkdir -p "$PERF_REPORTS_DIR"
if [[ ${BUILDKITE_PULL_REQUEST:-false} =~ ^[1-9][0-9]*$ ]]; then
export PERF_RUN_SOURCE=pr
export PERF_UPLOAD_POLICY=pass
elif [ "${BUILDKITE_BRANCH:-}" = main ] \
&& { [ "${BUILDKITE_SOURCE:-}" = schedule ] || [ "${TEST_SCOPE:-}" = full ]; }; then
export PERF_RUN_SOURCE=scheduled_main
export PERF_UPLOAD_POLICY=always
elif [ "${TEST_SCOPE:-}" = direct ]; then
export PERF_RUN_SOURCE=unknown
export PERF_UPLOAD_POLICY=pass
else
export PERF_RUN_SOURCE=unknown
export PERF_UPLOAD_POLICY=never
fi
nvidia-smi \
--query-gpu=index,timestamp,clocks.sm,clocks.max.sm,power.draw,power.limit,temperature.gpu \
--format=csv -l 10 > "$PERF_REPORTS_DIR/gpu_telemetry.csv" 2>/dev/null &
telemetry_pid=$!
cleanup() {
kill "$telemetry_pid" 2>/dev/null || true
wait "$telemetry_pid" 2>/dev/null || true
}
trap cleanup EXIT INT TERM
pytest ./fastvideo/tests/performance/test_inference_performance.py -vs
pytest_rc=$?
compare_rc=0
if [ "$pytest_rc" -eq 0 ] || [ "$PERF_UPLOAD_POLICY" = always ]; then
PERF_PYTEST_RC=$pytest_rc python ./fastvideo/tests/performance/compare_baseline.py
compare_rc=$?
fi
python ./fastvideo/tests/performance/dashboard.py || true
cp -f fastvideo/tests/performance/results/*.json "$PERF_REPORTS_DIR/" 2>/dev/null || true
# The trusted host relays only .md/.html/.json/.csv from PERF_REPORTS_DIR, so
# mirror each captured worker log with an allowlisted extension.
for worker_log in fastvideo/tests/performance/results/worker_logs/*.log; do
[ -f "$worker_log" ] || continue
base=$(basename "${worker_log%.log}")
# WorkerLogCapture keeps a .log.1 backup after rollover, and read_log_tail
# includes it; mirror that retained history too so the artifact is complete.
if [ -f "$worker_log.1" ]; then
cp -f "$worker_log.1" "$PERF_REPORTS_DIR/${base}.1.md" 2>/dev/null || true
fi
cp -f "$worker_log" "$PERF_REPORTS_DIR/${base}.md" 2>/dev/null || true
done
echo "--- GPU telemetry (clocks.sm vs clocks.max.sm reveals capped hosts) ---"
cat "$PERF_REPORTS_DIR/gpu_telemetry.csv" || true
final_rc=$pytest_rc
if [ "$final_rc" -eq 0 ]; then
final_rc=$compare_rc
fi
exit "$final_rc"
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the self-forcing lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/self-forcing/test_self_forcing.py -vs
+40
View File
@@ -0,0 +1,40 @@
#!/usr/bin/env bash
# Canonical four-GPU SSIM lane for the Slinky Slurm worker.
set -euo pipefail
args=()
if [ "${FASTVIDEO_SSIM_BOOTSTRAP_MODE:-0}" = 1 ]; then
args+=(--bootstrap-mode)
fi
selected=${FASTVIDEO_SSIM_TEST_FILES-}
if [ -z "$selected" ]; then
if [ "${TEST_SCOPE:-}" = merge ]; then
echo "Missing FASTVIDEO_SSIM_TEST_FILES for merge scope" >&2
exit 2
fi
selected=all
fi
if [ "$selected" != all ]; then
[[ $selected =~ ^test_[a-z0-9_]+\.py(,test_[a-z0-9_]+\.py)*$ ]] || {
echo "Invalid FASTVIDEO_SSIM_TEST_FILES selection" >&2
exit 2
}
IFS=, read -r -a ssim_files <<< "$selected"
for ssim_file in "${ssim_files[@]}"; do
args+=(--test-file "$ssim_file")
done
fi
# MoGe's utils3d dependency builds glcontext from source on ARM64. The current
# runner image predates the baked-in X11 headers below, so keep this guarded
# bootstrap until every deployed image digest contains libx11-dev.
if [ ! -f /usr/include/X11/Xlib.h ]; then
apt-get -o Acquire::Retries=5 update
apt-get -o Acquire::Retries=5 install -y --no-install-recommends libx11-dev
rm -rf /var/lib/apt/lists/*
fi
uv pip install git+https://github.com/microsoft/MoGe.git
uv pip install k_diffusion einops_exts alias_free_torch torchsde
exec python fastvideo/tests/ssim/ci_runner.py "${args[@]}"
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the modular training-framework lane.
set -euo pipefail
exec pytest ./fastvideo/tests/train/models ./fastvideo/tests/train/methods -vs
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the legacy vanilla-training lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/Vanilla -srP
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the legacy LoRA-training lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/lora/test_lora_training.py -srP
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the legacy VSA-training lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/VSA -srP
+9
View File
@@ -0,0 +1,9 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the transformer lane.
set -euo pipefail
# The existing block reference records an absent FASTVIDEO_FA4 (FA2). Keep
# that reference identity; the component lane also selects FA2 explicitly.
env -u FASTVIDEO_FA4 pytest ./fastvideo/tests/golden_gate/test_wan_t2v.py -xvs
pytest ./fastvideo/tests/golden_gate/test_wan_causal.py -xvs
exec pytest ./fastvideo/tests/transformers -vs
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the VAE lane.
set -euo pipefail
pytest ./fastvideo/tests/golden_gate/test_wan_vae.py -xvs
exec pytest ./fastvideo/tests/vaes -vs
+41 -2
View File
@@ -1,6 +1,19 @@
#!/bin/bash
set -uo pipefail
# DORMANT ROLLBACK ONLY. Active CI is Slurm-only and pipeline.yml never calls
# this launcher. Refuse every Buildkite invocation even if a stale step or
# operator typo reaches this file; local rollback experiments require an
# explicit opt-in.
if [ -n "${BUILDKITE:-}" ]; then
echo "Legacy Modal CI is disabled; use the Slinky Slurm runner." >&2
exit 2
fi
if [ "${FASTVIDEO_ENABLE_LEGACY_MODAL_CI:-0}" != 1 ]; then
echo "Legacy Modal CI is dormant. Set FASTVIDEO_ENABLE_LEGACY_MODAL_CI=1 only for a manual rollback test." >&2
exit 2
fi
log() {
echo "[$(date '+%Y-%m-%d %H:%M:%S')] $1"
}
@@ -76,10 +89,27 @@ EFFECTIVE_PR=${BUILDKITE_PULL_REQUEST:-false}
if [ "$EFFECTIVE_PR" = "false" ] && [ -n "${PR_NUMBER:-}" ]; then
EFFECTIVE_PR=$PR_NUMBER
fi
MODAL_ENV="BUILDKITE_REPO=$BUILDKITE_REPO BUILDKITE_COMMIT=$BUILDKITE_COMMIT BUILDKITE_PULL_REQUEST=$EFFECTIVE_PR BUILDKITE_BRANCH=${BUILDKITE_BRANCH:-} TEST_SCOPE=${TEST_SCOPE:-} BUILDKITE_BUILD_URL=${BUILDKITE_BUILD_URL:-} BUILDKITE_BUILD_ID=${BUILDKITE_BUILD_ID:-} BUILDKITE_JOB_ID=${BUILDKITE_JOB_ID:-} IMAGE_VERSION=$IMAGE_VERSION"
MODAL_ENV="BUILDKITE_REPO=$BUILDKITE_REPO BUILDKITE_COMMIT=$BUILDKITE_COMMIT BUILDKITE_PULL_REQUEST=$EFFECTIVE_PR BUILDKITE_BRANCH=${BUILDKITE_BRANCH:-} BUILDKITE_SOURCE=${BUILDKITE_SOURCE:-} TEST_SCOPE=${TEST_SCOPE:-} BUILDKITE_BUILD_URL=${BUILDKITE_BUILD_URL:-} BUILDKITE_BUILD_ID=${BUILDKITE_BUILD_ID:-} BUILDKITE_JOB_ID=${BUILDKITE_JOB_ID:-} IMAGE_VERSION=$IMAGE_VERSION"
POST_RUN_HOOK=""
is_truthy() {
case "${1:-}" in
1|true|TRUE|yes|YES|on|ON) return 0 ;;
*) return 1 ;;
esac
}
ssim_bootstrap_args() {
local title="${PR_TITLE:-}"
local message="${BUILDKITE_MESSAGE:-}"
if is_truthy "${FASTVIDEO_SSIM_BOOTSTRAP_MODE:-}" \
|| [[ "$title" == *"[new-model]"* ]] \
|| [[ "$message" == *"[new-model]"* ]]; then
printf ' --bootstrap-mode'
fi
}
upload_performance_artifacts() {
SHORT_SHA=${BUILDKITE_COMMIT:0:7}
LOCAL_DIR="downloaded_reports"
@@ -170,9 +200,18 @@ case "$TEST_TYPE" in
log "Running transformer tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_transformer_tests"
;;
"golden_gate")
log "Running golden-gate tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_TEST_FILE::run_golden_gate_tests"
;;
"ssim")
log "Running SSIM tests..."
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run $MODAL_SSIM_TEST_FILE::run_ssim_tests"
SSIM_BOOTSTRAP_ARGS=$(ssim_bootstrap_args)
if [ -n "$SSIM_BOOTSTRAP_ARGS" ]; then
log "SSIM bootstrap mode enabled for new-model reference draft generation"
fi
MODAL_COMMAND="$MODAL_ENV HF_API_KEY=$HF_API_KEY python3 -m modal run "
MODAL_COMMAND+="$MODAL_SSIM_TEST_FILE::run_ssim_tests$SSIM_BOOTSTRAP_ARGS"
;;
"training")
log "Running training tests..."
+41
View File
@@ -0,0 +1,41 @@
#!/usr/bin/env bash
set -euo pipefail
# Collect the whole attention directory so new files cannot land uncovered.
# Its FA2/FA3 regression files skip when FA4 is selected (the Modal image
# enables FA4 by default), so pin FA4 off for the directory to be real
# coverage on every runner rather than a nominal collection.
export FASTVIDEO_FA4=0
# The livestream app's tests are CPU-only; its single gpu-marked module is
# deselected, and DreamVerse's GPU tests have their own lane.
exec pytest \
./apps/infinite_livestream/infinite_livestream/tests \
./fastvideo/tests/api/ \
./fastvideo/tests/contract/ \
./fastvideo/tests/dataset/ \
./fastvideo/tests/workflow/ \
./fastvideo/tests/entrypoints/ \
./fastvideo/tests/loader/ \
./fastvideo/tests/pipelines/ \
./fastvideo/tests/platforms/ \
./fastvideo/tests/schedulers/ \
./fastvideo/tests/train/ \
./fastvideo/tests/stages/ \
./fastvideo/tests/ops/ \
./fastvideo/tests/worker/ \
./fastvideo/tests/training/test_runner.py \
./fastvideo/tests/training/test_trackers.py \
./fastvideo/tests/inference/test_basic_fasth3_omniref_pdd.py \
./fastvideo/tests/inference/test_inference_regional_compile.py \
./fastvideo/tests/attention/ \
./fastvideo/tests/layers/test_pdd_linear.py \
./fastvideo/tests/layers/test_triton_fused_norm.py \
./fastvideo/tests/modal/test_kernel_build_cache.py \
./fastvideo/tests/modal/test_pr_test.py \
./fastvideo/tests/modal/test_ssim_test.py \
--ignore=./fastvideo/tests/entrypoints/test_openai_api_integration.py \
--ignore=./fastvideo/tests/train/models \
--ignore=./fastvideo/tests/train/methods \
-m "not gpu" \
-vs
+2 -2
View File
@@ -8,10 +8,10 @@ PR TITLE: Must start with a type tag, e.g.:
MERGE WORKFLOW:
1. Ensure pre-commit passes and you have at least 1 approval
2. Comment /merge (or add the "ready" label) to enter the Merge Queue
3. Full Test Suite runs automatically on a staging branch → auto-merge on success
3. A path-aware merge gate runs only relevant integration tests → auto-merge on success
ON-DEMAND TESTING (write access required):
/test full — Full Test Suite /test ssim — SSIM regression
/test full — Explicit all-lane run /test ssim — Full SSIM regression
/test training — Training pipeline /test encoder — Encoder tests
/test transformer — Transformer tests /test vae — VAE tests
/test kernel — CUDA kernel tests /test unit — Unit tests
+133
View File
@@ -0,0 +1,133 @@
#!/usr/bin/env bash
# Gate the path-aware Buildkite merge plan on the cheap GitHub checks.
#
# Polls the workflow runs for the PR head commit and only exits 0 once the
# watched cheap workflows (pre-commit, docs build) have succeeded, so the
# 'ready' label cannot burn path-selected GPU lanes on a head that a cheap
# check has already doomed.
#
# Semantics:
# - watched run completed with a bad conclusion -> exit 1 (fail CLOSED:
# no merge gate; the next push re-arms via the 'synchronize' trigger)
# - watched run cancelled -> still pending: the docs
# workflow's repo-global 'pages' concurrency group cancels runs superseded
# by unrelated pushes, so 'cancelled' is not a verdict on this PR
# - watched runs pending -> poll until done
# - docs run absent -> not applicable after a
# short grace period ('Deploy Documentation' is path-filtered on PRs)
# - pre-commit run absent -> keep polling: pre-commit
# is never path-filtered, so its absence is always anomalous
# - 'ready' label removed while waiting -> exit 1 (fail CLOSED:
# un-labeling is a deliberate maintainer action)
# - GitHub API unreachable or timeout -> exit 0 (fail OPEN,
# loud warning: never brick CI on a GitHub outage)
#
# Required env: PR_SHA (PR head commit), PR_NUMBER, GITHUB_REPOSITORY, GH_TOKEN.
set -euo pipefail
: "${PR_SHA:?PR_SHA (PR head commit) is required}"
: "${PR_NUMBER:?PR_NUMBER (pull request number) is required}"
: "${GITHUB_REPOSITORY:?GITHUB_REPOSITORY is required}"
# Workflow-level `name:` values that must be green before the merge gate
# may start. "Deploy Documentation" is path-filtered on PRs, so its run may
# legitimately never exist; pre-commit always runs, so it must appear.
WATCHED_NAMES='["pre-commit", "Deploy Documentation"]'
WATCHED_REGEX='^(pre-commit|Deploy Documentation)$'
POLL_SECS="${POLL_SECS:-20}"
GRACE_SECS="${GRACE_SECS:-60}"
MAX_WAIT_SECS="${MAX_WAIT_SECS:-1500}"
# Bound each API call so a hung connection hits the 3-strike fail-open path
# instead of pinning the loop until the job timeout (which would fail closed
# on exactly the GitHub-outage case this script is meant to survive).
if command -v timeout >/dev/null 2>&1; then
gh_api() { timeout 30 gh api "$@"; }
else
gh_api() { gh api "$@"; } # macOS dev boxes; CI always has coreutils timeout
fi
# The workflow checked the label before starting the gate, but the wait can
# last ~25 min: re-check once before any exit 0 and fail closed if 'ready'
# was removed in the meantime. An API error here proceeds (the label was
# present when the gate started; never brick CI on an outage).
recheck_ready_label() {
local pr_json
if pr_json=$(gh_api "repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}" 2>/dev/null); then
if ! jq -e '[.labels[]?.name] | index("ready")' <<<"$pr_json" >/dev/null 2>&1; then
echo "::error::PR #${PR_NUMBER} no longer has the 'ready' label —" \
"NOT triggering the Buildkite merge gate. Re-add the label to re-arm."
exit 1
fi
else
echo "::warning::Could not re-check the 'ready' label on PR #${PR_NUMBER}; proceeding (it was present when the gate started)."
fi
}
start=$(date +%s)
api_fails=0
missing=""
while true; do
elapsed=$(( $(date +%s) - start ))
if runs_json=$(gh_api "repos/${GITHUB_REPOSITORY}/actions/runs?head_sha=${PR_SHA}&per_page=100" 2>/dev/null) \
&& state=$(jq --arg re "$WATCHED_REGEX" '
[.workflow_runs[]? | select(.name // "" | test($re))]
| group_by(.name) | map(max_by(.id))
| map({name, status, conclusion})' <<<"$runs_json" 2>/dev/null); then
api_fails=0
echo "t+${elapsed}s watched checks: $(jq -c . <<<"$state")"
failed=$(jq -r '[.[] | select(.status == "completed"
and (.conclusion | IN("success", "skipped", "neutral", "cancelled") | not))]
| map(.name) | join(", ")' <<<"$state")
if [ -n "$failed" ]; then
echo "::error::Cheap check(s) failed on ${PR_SHA}: ${failed}." \
"NOT triggering the Buildkite merge gate. Push a fix (the 'ready'" \
"label re-arms on every push), or re-run the failed check and then" \
"re-run this workflow."
exit 1
fi
# 'cancelled' counts as pending: wait for a re-run to reach a real verdict
# (bounded by MAX_WAIT, then the fail-open below).
pending=$(jq '[.[] | select(.status != "completed" or .conclusion == "cancelled")] | length' <<<"$state")
missing=$(jq -r --argjson watched "$WATCHED_NAMES" '($watched - map(.name)) | join(", ")' <<<"$state")
if [ "$pending" -eq 0 ]; then
if [ -z "$missing" ]; then
recheck_ready_label
echo "All watched cheap checks are green — merge gate may proceed."
exit 0
fi
case "$missing" in
*pre-commit*)
echo "pre-commit run not found for ${PR_SHA} yet; waiting (pre-commit is never path-filtered, so its absence is anomalous)."
;;
*)
if [ "$elapsed" -ge "$GRACE_SECS" ]; then
recheck_ready_label
echo "::warning::Watched run(s) never appeared for ${PR_SHA}: ${missing} (path-filtered, likely not applicable). Proceeding on the checks that did run."
exit 0
fi
echo "Waiting up to ${GRACE_SECS}s grace for path-filtered run(s) to appear: ${missing}."
;;
esac
fi
else
api_fails=$(( api_fails + 1 ))
echo "::warning::GitHub API error querying workflow runs for ${PR_SHA} (attempt ${api_fails}/3)."
if [ "$api_fails" -ge 3 ]; then
recheck_ready_label
echo "::warning::FAILING OPEN: cannot query GitHub check status — triggering the merge gate WITHOUT the cheap-check gate."
exit 0
fi
fi
if [ "$elapsed" -ge "$MAX_WAIT_SECS" ]; then
recheck_ready_label
echo "::warning::FAILING OPEN: watched checks still pending after $(( MAX_WAIT_SECS / 60 )) min${missing:+ (never appeared: ${missing})} — triggering the merge gate anyway."
exit 0
fi
sleep "$POLL_SECS"
done
+592
View File
@@ -0,0 +1,592 @@
#!/usr/bin/env python3
"""Select the additive GPU integration lanes needed by a PR diff.
Fastcheck is the universal six-lane baseline and is intentionally not repeated
here. This planner selects only the more expensive merge-gate lanes. Unknown
source/build paths fail closed to the complete integration set, while explicit
documentation and repository-metadata paths require no additional GPU work.
"""
from __future__ import annotations
import argparse
import fnmatch
import re
from dataclasses import dataclass, field
from pathlib import Path
from typing import TextIO
MERGE_LANES = (
"golden-gate",
"ssim",
"lora-inference",
"lora-extraction",
"training",
"distillation",
"self-forcing",
"lora-training",
"training-vsa",
"inference-vmoba",
"performance",
"api-server",
"train-framework",
"eval",
)
LANE_SCRIPT_TO_KEY = {
"api_server.sh": "api-server",
"distillation_dmd.sh": "distillation",
"eval.sh": "eval",
"golden_gate.sh": "golden-gate",
"inference_lora.sh": "lora-inference",
"inference_vmoba.sh": "inference-vmoba",
"lora_extraction.sh": "lora-extraction",
"performance.sh": "performance",
"self_forcing.sh": "self-forcing",
"ssim.sh": "ssim",
"train_framework.sh": "train-framework",
"training.sh": "training",
"training_lora.sh": "lora-training",
"training_vsa.sh": "training-vsa",
}
FASTCHECK_LANE_SCRIPTS = {
"dreamverse.sh",
"encoder.sh",
"kernel_tests.sh",
"transformer.sh",
"vae.sh",
}
LEGACY_TRAINING_LANES = (
"training",
"distillation",
"self-forcing",
"lora-training",
"training-vsa",
)
ALL_TRAINING_LANES = (*LEGACY_TRAINING_LANES, "train-framework")
SSIM_SMOKE_TESTS = (
"test_flux_t2i_similarity.py",
"test_wan_t2v_similarity.py",
)
SAFE_PATTERNS = (
"*.md",
"*.rst",
".agents/**",
".claude/**",
".codex/**",
".github/ISSUE_TEMPLATE/**",
".github/PULL_REQUEST_TEMPLATE.md",
".github/dependabot.yml",
".github/mergify.yml",
".github/scripts/**",
".github/workflows/**",
".buildkite/scripts/pre_commit.sh",
".git-blame-ignore-revs",
".gitattributes",
".gitignore",
".pre-commit-config.yaml",
"AGENTS.md",
"CITATION.cff",
"CODE_OF_CONDUCT.md",
"CONTRIBUTING.md",
"LICENSE",
"NOTICE",
"__init__.py",
"collect_env.py",
"SECURITY.md",
"assets/**",
"comfyui/**",
"docs/**",
"examples/**",
"mkdocs.yml",
"requirements-mkdocs.in",
"requirements-mkdocs.txt",
"scripts/**",
"tests/__init__.py",
"tests/local_tests/**",
)
ALL_IMPACT_PATTERNS = (
".buildkite/pipeline.yml",
"docker/**",
"pyproject.toml",
"requirements*.txt",
"setup.cfg",
"setup.py",
"uv.lock",
)
@dataclass(frozen=True)
class FamilyCoverage:
pattern: re.Pattern[str]
golden_tests: tuple[str, ...]
ssim_tests: tuple[str, ...]
FAMILY_COVERAGE = (
FamilyCoverage(
re.compile(r"(^|[/_.-])dreamx(_world)?([/_.-]|$)"),
("test_dreamx.py", ),
("test_dreamx_world_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])flux[_-]?2([/_.-]|$)"),
("test_flux2_klein.py", ),
("test_flux2_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])flux(?![_-]?2)([/_.-]|$)"),
("test_flux.py", ),
("test_flux_t2i_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])(hunyuan)?gamecraft([/_.-]|$)"),
("test_gamecraft.py", ),
("test_gamecraft_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])gen3c([/_.-]|$)"),
("test_gen3c.py", ),
("test_gen3c_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])glm[_-]?image([/_.-]|$)"),
("test_glm_image.py", ),
("test_glm_image_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])hunyuan(video)?15([a-z0-9_-]*)([/_.-]|$)"),
(),
("test_hunyuan15_i2v_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])kandinsky[_-]?5([/_.-]|$)"),
("test_kandinsky5.py", ),
("test_kandinsky5_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])lingbot([a-z0-9_-]*)([/_.-]|$)"),
("test_lingbot.py", ),
("test_lingbot_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])longcat([/_.-]|$)"),
("test_longcat.py", ),
("test_longcat_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])ltx[_-]?2([/_.-]|$)"),
("test_ltx2.py", ),
("test_ltx2_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])matrixgame[_-]?2([/_.-]|$)"),
("test_matrixgame.py", ),
("test_matrixgame2_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])matrixgame[_-]?3([/_.-]|$)"),
("test_matrixgame.py", ),
("test_matrixgame3_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])minimax[_-]?h3([/_.-]|$)"),
("test_minimax_h3_t2v.py", ),
("test_minimax_h3_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])sd[_-]?3([._-]?5)?([/_.-]|$)"),
("test_sd35.py", ),
("test_sd35_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])stable[_-]?audio([/_.-]|$)"),
("test_stable_audio.py", ),
("test_stable_audio_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])turbo(diffusion)?([/_.-]|$)"),
(),
("test_turbodiffusion_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])wan(video|vae)?([/_.-]|$)"),
("test_wan_t2v.py", "test_wan_vae.py", "test_wan_causal.py", "test_wan_denoising.py"),
(
"test_causal_similarity.py",
"test_wan_i2v_similarity.py",
"test_wan_t2v_similarity.py",
"test_wan_ti2v_similarity.py",
),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])z[_-]?image([/_.-]|$)"),
("test_zimage.py", ),
("test_zimage_similarity.py", ),
),
)
@dataclass
class MergePlan:
lanes: set[str] = field(default_factory=set)
golden_tests: set[str] = field(default_factory=set)
ssim_tests: set[str] = field(default_factory=set)
golden_all: bool = False
ssim_all: bool = False
reasons: list[str] = field(default_factory=list)
def add_lanes(self, *lanes: str, reason: str) -> None:
unknown = set(lanes) - set(MERGE_LANES)
if unknown:
raise ValueError(f"Unknown merge lanes: {sorted(unknown)}")
self.lanes.update(lanes)
self.reasons.append(reason)
def add_golden(self, tests: tuple[str, ...], reason: str) -> None:
self.add_lanes("golden-gate", reason=reason)
self.golden_tests.update(tests)
def add_ssim(self, tests: tuple[str, ...], reason: str) -> None:
self.add_lanes("ssim", reason=reason)
self.ssim_tests.update(tests)
def require_all(self, reason: str) -> None:
self.lanes.update(MERGE_LANES)
self.golden_all = True
self.ssim_all = True
self.reasons.append(reason)
def ordered_lanes(self) -> tuple[str, ...]:
return tuple(lane for lane in MERGE_LANES if lane in self.lanes)
def encoded_lanes(self) -> str:
lanes = self.ordered_lanes()
return "," + ",".join(lanes or ("none", )) + ","
def encoded_golden_tests(self) -> str:
if "golden-gate" not in self.lanes:
return "none"
if self.golden_all or not self.golden_tests:
return "all"
return ",".join(sorted(self.golden_tests))
def encoded_ssim_tests(self) -> str:
if "ssim" not in self.lanes:
return "none"
if self.ssim_all or not self.ssim_tests:
return "all"
return ",".join(sorted(self.ssim_tests))
def _matches_any(path: str, patterns: tuple[str, ...]) -> bool:
return any(fnmatch.fnmatchcase(path, pattern) for pattern in patterns)
def _family_coverage(path: str) -> tuple[set[str], set[str]]:
normalized = path.lower()
golden: set[str] = set()
ssim: set[str] = set()
for family in FAMILY_COVERAGE:
if family.pattern.search(normalized):
golden.update(family.golden_tests)
ssim.update(family.ssim_tests)
# Select the component actually touched, including compatibility paths.
# Family configs/pipeline wiring can affect all four Wan gates.
if re.search(r"(^|[/_.-])wan(video|vae)?([/_.-]|$)", normalized):
if (normalized.endswith(("/wan/vae.py", "/wan/vae_config.py", "/vaes/wanvae.py"))
or normalized.endswith("/wan/stages/conditioning.py")):
golden = {"test_wan_vae.py"}
elif normalized.endswith(("/wan/causal_transformer.py", "/dits/causal_wanvideo.py",
"/wan/stages/causal_denoising.py")):
golden = {"test_wan_causal.py"}
elif (normalized == "fastvideo/models/dits/wanvideo.py"
or normalized.endswith(("/wan/transformer.py", "/wan/stages/denoising.py", "/wan/stages/dmd.py"))):
golden = {"test_wan_t2v.py", "test_wan_denoising.py"}
return golden, ssim
def _select_output_coverage(plan: MergePlan, path: str) -> None:
golden, ssim = _family_coverage(path)
if golden:
plan.add_golden(tuple(sorted(golden)), reason=f"model-family golden coverage: {path}")
else:
plan.golden_all = True
plan.add_lanes("golden-gate", reason=f"shared output golden coverage: {path}")
if ssim:
plan.add_ssim(tuple(sorted(ssim)), reason=f"model-family SSIM coverage: {path}")
else:
plan.add_ssim(SSIM_SMOKE_TESTS, reason=f"shared output SSIM smoke coverage: {path}")
def classify_paths(paths: list[str]) -> MergePlan:
plan = MergePlan()
normalized_paths: list[str] = []
for raw_path in paths:
path = raw_path.strip()
while path.startswith("./"):
path = path[2:]
if path:
normalized_paths.append(path)
normalized_paths = sorted(set(normalized_paths))
if not normalized_paths:
plan.require_all("changed-file list was empty; failing closed")
return plan
for path in normalized_paths:
if path == "__FASTVIDEO_CI_PLAN_ALL__":
plan.require_all("changed-file API failed; failing closed")
continue
if path in {"requirements-mkdocs.in", "requirements-mkdocs.txt"}:
plan.reasons.append(f"documentation dependencies need no GPU integration: {path}")
continue
if _matches_any(path, ALL_IMPACT_PATTERNS):
plan.require_all(f"cross-cutting build/runtime surface: {path}")
continue
lane_script_prefix = ".buildkite/scripts/lanes/"
if path.startswith(lane_script_prefix):
script_name = Path(path).name
lane = LANE_SCRIPT_TO_KEY.get(script_name)
if lane is None:
if script_name in FASTCHECK_LANE_SCRIPTS:
plan.reasons.append(f"covered by automatic Fastcheck lane: {path}")
else:
plan.require_all(f"unknown lane script: {path}")
elif lane == "golden-gate":
plan.golden_all = True
plan.add_lanes(lane, reason=f"golden lane implementation: {path}")
elif lane == "ssim":
plan.ssim_all = True
plan.add_lanes(lane, reason=f"SSIM lane implementation: {path}")
else:
plan.add_lanes(lane, reason=f"lane implementation: {path}")
continue
if path.startswith("fastvideo/tests/golden_gate/"):
name = Path(path).name
if name.startswith("test_") and name.endswith(".py"):
plan.add_golden((name, ), reason=f"changed golden test: {path}")
elif name in {"AGENTS.md", "README.md"}:
plan.reasons.append(f"golden documentation only: {path}")
else:
plan.golden_all = True
plan.add_lanes("golden-gate", reason=f"shared golden harness/reference: {path}")
continue
if path.startswith("fastvideo/tests/ssim/"):
name = Path(path).name
if name.startswith("test_") and name.endswith(".py"):
plan.add_ssim((name, ), reason=f"changed SSIM test: {path}")
elif path.endswith((".py", ".json", ".pt", ".png", ".mp4")):
plan.ssim_all = True
plan.add_lanes("ssim", reason=f"shared SSIM harness/reference: {path}")
continue
if path.startswith("fastvideo/tests/performance/") or path.startswith(".buildkite/performance-benchmarks/"):
plan.add_lanes("performance", reason=f"performance coverage: {path}")
continue
if path.startswith(("fastvideo/performance/", "fastvideo/performance_dashboard/",
"apps/performance_dashboard/")):
plan.add_lanes("performance", reason=f"performance implementation: {path}")
continue
if path.startswith("fastvideo/benchmarks/"):
if "/mlx_" in path or Path(path).name.startswith("mlx_"):
plan.reasons.append(f"covered by the path-filtered macOS MLX workflow: {path}")
else:
plan.add_lanes("performance", reason=f"benchmark implementation: {path}")
continue
if path.startswith("fastvideo/tests/eval/") or path.startswith("fastvideo/eval/"):
plan.add_lanes("eval", reason=f"evaluation coverage: {path}")
continue
if path.startswith("fastvideo/third_party/eval/"):
plan.add_lanes("eval", reason=f"vendored evaluation implementation: {path}")
continue
if path.startswith("fastvideo/tests/lora_extraction/") or path.startswith("scripts/lora_extraction/"):
plan.add_lanes("lora-extraction", reason=f"LoRA extraction coverage: {path}")
continue
if path.startswith("fastvideo/tests/inference/lora/"):
plan.add_lanes("lora-inference", reason=f"LoRA inference coverage: {path}")
continue
if path.startswith("fastvideo/tests/inference/vmoba/"):
plan.add_lanes("inference-vmoba", reason=f"VMoBA inference coverage: {path}")
continue
if path.startswith(("fastvideo/dataset/", "fastvideo/workflow/", "fastvideo/pipelines/preprocess/",
"fastvideo/pipelines/training/")):
plan.add_lanes(*ALL_TRAINING_LANES, reason=f"shared data/training input surface: {path}")
continue
if path.startswith("fastvideo/tests/train/") or path.startswith("fastvideo/train/"):
plan.add_lanes("train-framework", reason=f"modular training coverage: {path}")
continue
if path.startswith("fastvideo/tests/training/"):
lowered = path.lower()
if "/vanilla/" in lowered:
plan.add_lanes("training", reason=f"vanilla training coverage: {path}")
elif "/distill/" in lowered:
plan.add_lanes("distillation", reason=f"distillation coverage: {path}")
elif "/self-forcing/" in lowered:
plan.add_lanes("self-forcing", reason=f"self-forcing coverage: {path}")
elif "/lora/" in lowered:
plan.add_lanes("lora-training", reason=f"LoRA training coverage: {path}")
elif "/vsa/" in lowered:
plan.add_lanes("training-vsa", reason=f"VSA training coverage: {path}")
else:
plan.add_lanes(*LEGACY_TRAINING_LANES, reason=f"shared legacy training coverage: {path}")
continue
if path.startswith("fastvideo/training/"):
lowered = path.lower()
if "self_forcing" in lowered:
plan.add_lanes("self-forcing", reason=f"self-forcing implementation: {path}")
elif "distill" in lowered:
plan.add_lanes("distillation", reason=f"distillation implementation: {path}")
elif "lora" in lowered:
plan.add_lanes("lora-training", reason=f"LoRA training implementation: {path}")
else:
plan.add_lanes(*LEGACY_TRAINING_LANES, reason=f"shared legacy training implementation: {path}")
continue
lowered = path.lower()
if "vmoba" in lowered and path.startswith(("fastvideo/", ".buildkite/")):
plan.add_lanes("inference-vmoba", reason=f"VMoBA implementation: {path}")
plan.add_golden(("test_wan_t2v.py", ), reason=f"VMoBA end-to-end coverage: {path}")
continue
if "lora" in lowered and path.startswith("fastvideo/"):
plan.add_lanes(
"lora-inference",
"lora-extraction",
"lora-training",
reason=f"shared LoRA implementation: {path}",
)
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/entrypoints/") or path.startswith("fastvideo/api/"):
plan.add_lanes("api-server", reason=f"API/entrypoint integration: {path}")
if "openai" not in lowered and "/cli/" not in lowered:
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/worker/"):
plan.add_lanes("api-server", reason=f"worker/API integration: {path}")
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/distributed/"):
plan.add_lanes(
"training",
"train-framework",
reason=f"distributed runtime integration: {path}",
)
_select_output_coverage(plan, path)
continue
if path.startswith(("fastvideo/hooks/", "fastvideo/platforms/", "fastvideo/third_party/")):
_select_output_coverage(plan, path)
continue
if path.startswith(("fastvideo/models/", "fastvideo/pipelines/", "fastvideo/configs/",
"fastvideo/layers/", "fastvideo/attention/")):
_select_output_coverage(plan, path)
continue
if path in {
"fastvideo/fastvideo_args.py",
"fastvideo/forward_context.py",
"fastvideo/image_processor.py",
"fastvideo/registry.py",
"fastvideo/utils.py",
}:
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/mlx_runtime/"):
plan.reasons.append(f"covered by the path-filtered macOS MLX workflow: {path}")
continue
if path.startswith("fastvideo/logging_utils/") or path in {
"fastvideo/__init__.py",
"fastvideo/envs.py",
"fastvideo/logger.py",
"fastvideo/profiler.py",
"fastvideo/version.py",
}:
plan.reasons.append(f"covered by automatic Fastcheck: {path}")
continue
if path.startswith(("fastvideo-kernel/", "csrc/")):
plan.add_golden(("test_wan_t2v.py", ), reason=f"kernel integration smoke: {path}")
plan.add_ssim(("test_wan_t2v_similarity.py", ), reason=f"kernel numerical smoke: {path}")
continue
if path.startswith("apps/dreamverse/"):
# DreamVerse is already one of the six automatic Fastcheck lanes.
plan.reasons.append(f"covered by automatic DreamVerse Fastcheck: {path}")
continue
if path.startswith("apps/infinite_livestream/"):
# The app's CPU-only tests run in the automatic unit Fastcheck lane.
plan.reasons.append(f"covered by automatic unit Fastcheck: {path}")
continue
if path.startswith("fastvideo/tests/"):
# The automatic unit/component Fastcheck lanes own the remaining
# package tests. Domain-specific expensive test roots were handled
# above.
plan.reasons.append(f"covered by automatic Fastcheck: {path}")
continue
if path in {".buildkite/scripts/unit_test.sh", ".buildkite/scripts/pr_test.sh"}:
plan.reasons.append(f"covered by automatic unit Fastcheck: {path}")
continue
if _matches_any(path, SAFE_PATTERNS):
plan.reasons.append(f"no additional GPU integration needed: {path}")
continue
plan.require_all(f"unclassified path; failing closed: {path}")
return plan
def _write_github_output(output: TextIO, plan: MergePlan) -> None:
output.write(f"merge_test_plan={plan.encoded_lanes()}\n")
output.write(f"merge_golden_tests={plan.encoded_golden_tests()}\n")
output.write(f"merge_ssim_tests={plan.encoded_ssim_tests()}\n")
output.write(f"merge_plan_label={','.join(plan.ordered_lanes()) or 'none'}\n")
def _write_summary(output: TextIO, plan: MergePlan) -> None:
output.write("## Change-aware merge test plan\n\n")
output.write("| Selection | Value |\n|---|---|\n")
output.write(f"| Additional Slurm lanes | `{','.join(plan.ordered_lanes()) or 'none'}` |\n")
output.write(f"| Golden tests | `{plan.encoded_golden_tests()}` |\n")
output.write(f"| SSIM tests | `{plan.encoded_ssim_tests()}` |\n\n")
output.write("Fastcheck remains the universal six-lane baseline.\n")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--paths-file", type=Path, required=True)
parser.add_argument("--github-output", type=Path)
parser.add_argument("--summary-file", type=Path)
return parser.parse_args()
def main() -> int:
args = parse_args()
paths = args.paths_file.read_text(encoding="utf-8").splitlines()
plan = classify_paths(paths)
print(f"MERGE_TEST_PLAN={plan.encoded_lanes()}")
print(f"MERGE_GOLDEN_TESTS={plan.encoded_golden_tests()}")
print(f"MERGE_SSIM_TESTS={plan.encoded_ssim_tests()}")
for reason in plan.reasons:
print(f"- {reason}")
if args.github_output:
with args.github_output.open("a", encoding="utf-8") as output:
_write_github_output(output, plan)
if args.summary_file:
with args.summary_file.open("a", encoding="utf-8") as output:
_write_summary(output, plan)
return 0
if __name__ == "__main__":
raise SystemExit(main())
+122
View File
@@ -0,0 +1,122 @@
#!/usr/bin/env bash
# Self-test for gate_full_suite.sh using a mocked `gh`. No network, runs on
# any dev box: bash .github/scripts/test_gate_full_suite.sh
set -u
here=$(cd "$(dirname "$0")" && pwd)
tmp=$(mktemp -d)
trap 'rm -rf "$tmp"' EXIT
# Mock gh. Asserts the exact endpoint (including head_sha) it is called
# with — an endpoint typo in the gate script fails the test rather than
# silently serving canned data. On the runs endpoint it serves
# $MOCK_DIR/response_<call#>.json, sticking on the highest existing file,
# and exits 1 if none exist (simulates a GitHub API outage). On the pulls
# endpoint it serves $MOCK_DIR/pr.json, defaulting to a 'ready'-labeled PR.
cat > "$tmp/gh" <<'EOF'
#!/usr/bin/env bash
if [ "${1:-}" != "api" ]; then
echo "unexpected gh invocation: $*" >> "$MOCK_DIR/endpoint_error"
exit 2
fi
case "${2:-}" in
"repos/o/r/actions/runs?head_sha=deadbeef&per_page=100")
n=$(( $(cat "$MOCK_DIR/count" 2>/dev/null || echo 0) + 1 ))
echo "$n" > "$MOCK_DIR/count"
while [ "$n" -gt 0 ]; do
if [ -f "$MOCK_DIR/response_$n.json" ]; then
cat "$MOCK_DIR/response_$n.json"
exit 0
fi
n=$(( n - 1 ))
done
echo "api outage" >&2
exit 1
;;
"repos/o/r/pulls/42")
if [ -f "$MOCK_DIR/pr.json" ]; then
cat "$MOCK_DIR/pr.json"
else
echo '{"labels": [{"name": "ready"}]}'
fi
;;
*)
echo "unexpected gh endpoint: $2" >> "$MOCK_DIR/endpoint_error"
exit 2
;;
esac
EOF
chmod +x "$tmp/gh"
PC_OK='{"name": "pre-commit", "id": 1, "status": "completed", "conclusion": "success"}'
PC_BAD='{"name": "pre-commit", "id": 1, "status": "completed", "conclusion": "failure"}'
PC_PENDING='{"name": "pre-commit", "id": 1, "status": "in_progress", "conclusion": null}'
DOCS_OK='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "success"}'
DOCS_BAD='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "failure"}'
DOCS_CANCELLED='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "cancelled"}'
OTHER='{"name": "Trigger Merge Gate", "id": 3, "status": "in_progress", "conclusion": null}'
NULL_NAME='{"name": null, "id": 4, "status": "completed", "conclusion": "failure"}'
PC_OK_RERUN='{"name": "pre-commit", "id": 5, "status": "completed", "conclusion": "success"}'
fails=0
want_log="" # optional: expect() also greps out.log for this regex, then resets
pr_json="" # optional: served for the pulls (label re-check) endpoint, then resets
raw_body="" # optional: serve responses verbatim instead of wrapping in workflow_runs
expect() { # <name> <expected-exit> <response json>...
local name=$1 want=$2 dir i=1
shift 2
dir=$(mktemp -d "$tmp/test_XXXXXX")
for body in "$@"; do
if [ -n "$raw_body" ]; then
printf '%s' "$body" > "$dir/response_$i.json"
else
printf '{"workflow_runs": [%s]}' "$body" > "$dir/response_$i.json"
fi
i=$(( i + 1 ))
done
[ -n "$pr_json" ] && printf '%s' "$pr_json" > "$dir/pr.json"
( export PATH="$tmp:$PATH" MOCK_DIR="$dir" PR_SHA=deadbeef PR_NUMBER=42 \
GITHUB_REPOSITORY=o/r POLL_SECS=0 GRACE_SECS=1 MAX_WAIT_SECS=3
bash "$here/gate_full_suite.sh" > "$dir/out.log" 2>&1 )
local rc=$?
if [ "$rc" -ne "$want" ]; then
echo "FAIL: $name (exit $rc, want $want)"
cat "$dir/out.log"
fails=1
elif [ -f "$dir/endpoint_error" ]; then
echo "FAIL: $name (mock gh got an unexpected call)"
cat "$dir/endpoint_error"
fails=1
elif [ -n "$want_log" ] && ! grep -Eq "$want_log" "$dir/out.log"; then
echo "FAIL: $name (log does not match: $want_log)"
cat "$dir/out.log"
fails=1
else
echo "ok: $name"
fi
want_log="" pr_json="" raw_body=""
}
expect "both green -> proceed" 0 "$PC_OK, $DOCS_OK, $OTHER, $NULL_NAME"
expect "docs build failed -> blocked" 1 "$PC_OK, $DOCS_BAD"
expect "pre-commit failed -> blocked" 1 "$PC_BAD"
expect "pending then green -> proceed" 0 "$PC_PENDING" "$PC_OK, $DOCS_OK"
want_log="never appeared.*Deploy Documentation"
expect "docs run absent (path-filtered) -> proceed after grace" 0 "$PC_OK"
expect "API outage -> fail open" 0
want_log="FAILING OPEN"
expect "pending past MAX_WAIT -> fail open" 0 "$PC_PENDING"
want_log="FAILING OPEN"
expect "unrelated runs only -> no grace, fail open at MAX_WAIT" 0 "$OTHER"
expect "cancelled docs then green -> proceed" 0 \
"$PC_OK, $DOCS_CANCELLED" "$PC_OK, $DOCS_OK"
want_log="FAILING OPEN"
expect "cancelled docs forever -> fail open at MAX_WAIT" 0 "$PC_OK, $DOCS_CANCELLED"
want_log="FAILING OPEN"
expect "pre-commit absent -> no grace, fail open at MAX_WAIT" 0 "$DOCS_OK"
expect "duplicate run names -> latest wins" 0 "$PC_BAD, $PC_OK_RERUN, $DOCS_OK"
raw_body=1
expect "garbage response body -> fail open" 0 "this is not json"
pr_json='{"labels": [{"name": "other"}]}'
expect "ready label removed mid-gate -> blocked" 1 "$PC_OK, $DOCS_OK"
exit "$fails"
@@ -190,6 +190,7 @@ jobs:
if: ${{ !inputs.push_by_digest }}
run: |
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ${{ steps.image.outputs.name }}:${{ inputs.tag_suffix }}-sha-${GITHUB_SHA::7}"
echo "Digest: ${{ steps.build-push.outputs.digest }}"
echo "To run tests with this image, manually trigger the 'Run Tests' workflow."
- name: Digest success message
+44 -24
View File
@@ -26,29 +26,54 @@ jobs:
per_page: 100,
});
const bkStatuses = data.statuses.filter(
s => s.context.startsWith('buildkite/ci/')
);
const FASTCHECK_PREFIX = 'buildkite/ci/microscope-';
// Buildkite derives the GitHub context prefix from the label emoji.
// Keep hard Full Suite lanes in test-tube/bar-chart namespaces and
// Fastcheck lanes in microscope so targeted reruns cannot clear the
// wrong aggregate status. Automatic PR jobs use pr-fastcheck while
// slash-command and Full Suite jobs use ci; normalize the suffix
// and keep the newest status for each logical lane.
const FASTCHECK_PREFIXES = [
'buildkite/pr-fastcheck/microscope-',
'buildkite/ci/microscope-',
];
const FULL_SUITE_PREFIXES = [
'buildkite/ci/test-tube-',
'buildkite/ci/bar-chart-',
];
const fastcheck = bkStatuses.filter(
s => s.context.startsWith(FASTCHECK_PREFIX)
);
const fullSuite = bkStatuses.filter(
s => FULL_SUITE_PREFIXES.some(p => s.context.startsWith(p))
function newestByLane(prefixes) {
const statuses = new Map();
for (const status of data.statuses) {
const prefix = prefixes.find(p => status.context.startsWith(p));
if (!prefix) continue;
const lane = status.context.slice(prefix.length);
const previous = statuses.get(lane);
if (!previous || Date.parse(status.updated_at) > Date.parse(previous.updated_at)) {
statuses.set(lane, status);
}
}
return statuses;
}
const fastcheck = newestByLane(FASTCHECK_PREFIXES);
const fullSuiteOnly = newestByLane(FULL_SUITE_PREFIXES);
const fastcheckPassed =
fastcheck.size === 6
&& [...fastcheck.values()].every(s => s.state === 'success');
const fullSuitePassed =
fastcheckPassed
&& fullSuiteOnly.size === 14
&& [...fullSuiteOnly.values()].every(s => s.state === 'success');
// Direct reruns may repair a failed suite, never create a gate for
// a suite that did not run.
const failedAggregate = context => data.statuses.some(
s => s.context === context && s.state === 'failure'
);
if (
fastcheck.length > 0
&& fastcheck.every(s => s.state === 'success')
) {
if (failedAggregate('fastcheck-passed') && fastcheckPassed) {
core.info(
`All ${fastcheck.length} fastcheck tests passed — updating fastcheck-passed`
`All ${fastcheck.size} fastcheck tests passed — updating fastcheck-passed`
);
await github.rest.repos.createCommitStatus({
owner: context.repo.owner,
@@ -56,17 +81,13 @@ jobs:
sha,
state: 'success',
context: 'fastcheck-passed',
description:
`All ${fastcheck.length} fastcheck tests passed`,
description: `All ${fastcheck.size} fastcheck tests passed`,
});
}
if (
fullSuite.length > 0
&& fullSuite.every(s => s.state === 'success')
) {
if (failedAggregate('full-suite-passed') && fullSuitePassed) {
core.info(
`All ${fullSuite.length} full suite tests passed — updating full-suite-passed`
'All 20 full suite tests passed — updating full-suite-passed'
);
await github.rest.repos.createCommitStatus({
owner: context.repo.owner,
@@ -74,7 +95,6 @@ jobs:
sha,
state: 'success',
context: 'full-suite-passed',
description:
`All ${fullSuite.length} full suite tests passed`,
description: 'All 20 full suite tests passed',
});
}
+170
View File
@@ -0,0 +1,170 @@
name: macOS MLX Smoke
on:
pull_request:
branches: [main]
paths:
- ".github/workflows/ci-macos-mlx.yml"
- "fastvideo/mlx_runtime/**"
- "fastvideo/tests/mlx/**"
- "fastvideo/tests/platforms/test_mps_vsa_error.py"
- "fastvideo/tests/platforms/test_cpu_sdpa.py"
- "fastvideo/platforms/cpu.py"
- "fastvideo/platforms/mps.py"
- "fastvideo/platforms/__init__.py"
- "fastvideo/__init__.py"
- "examples/inference/basic/mlx_*.py"
- "fastvideo/benchmarks/mlx_*.py"
- "pyproject.toml"
workflow_dispatch:
permissions:
contents: read
concurrency:
group: macos-mlx-${{ github.ref }}
cancel-in-progress: true
jobs:
mlx-smoke:
if: github.event_name == 'workflow_dispatch' || github.event.pull_request.draft != true
runs-on: macos-15
timeout-minutes: 25
env:
FASTVIDEO_ATTENTION_BACKEND: TORCH_SDPA
TOKENIZERS_PARALLELISM: "false"
MASTER_ADDR: "127.0.0.1"
MASTER_PORT: "29513"
GLOO_SOCKET_IFNAME: lo0
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- uses: astral-sh/setup-uv@v3
- name: Install lightweight MLX smoke dependencies
run: |
uv pip install --system \
--index-url https://download.pytorch.org/whl/cpu \
torch==2.12.0 torchvision torchaudio
uv pip install --system \
pytest pytest-timeout numpy scipy pillow imageio einops cloudpickle filelock \
PyYAML diffusers huggingface_hub remote-pdb safetensors loguru mlx \
"ftfy>=6.3.1" "opencv-python>=4.10.0.84" psutil "transformers>=5.0.0"
- name: Show Apple runtime
run: |
python - <<'PY'
import platform
import mlx.core as mx
import torch
print("machine:", platform.machine())
print("processor:", platform.processor())
print("mlx default device:", mx.default_device())
device_info = mx.metal.device_info() if mx.metal.is_available() else "metal unavailable"
print("mlx device_info:", device_info)
print("torch:", torch.__version__)
print("torch mps available:", torch.backends.mps.is_available())
PY
- name: Run MLX smoke tests
run: |
python -m pytest \
fastvideo/mlx_runtime/tests/ \
fastvideo/tests/mlx/test_dmd_sampling.py \
fastvideo/tests/mlx/test_memory_limits.py \
fastvideo/tests/mlx/test_quant_capability.py \
fastvideo/tests/mlx/test_mlx_dit_parity.py \
fastvideo/tests/mlx/test_mlx_compile_parity.py \
fastvideo/tests/mlx/test_mlx_checkpoint.py \
fastvideo/tests/mlx/test_mlx_checkpoint_compat.py \
fastvideo/tests/mlx/test_mlx_affine_dq_gemm.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_parity.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa_regressions.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_mode.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_fastwan_benchmark.py \
fastvideo/tests/mlx/test_taehv_decode.py \
fastvideo/tests/mlx/test_frame_upsample.py \
fastvideo/tests/mlx/test_mlx_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_refine.py \
fastvideo/tests/mlx/test_mlx_prompt_enhance.py \
fastvideo/tests/mlx/test_mlx_prompt_to_video_decode.py \
fastvideo/tests/mlx/test_mlx_wan22_prompt_cache_fingerprint.py \
fastvideo/tests/mlx/test_wan22_sample.py \
fastvideo/tests/mlx/test_windowed_attention.py \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_download_unavailable_has_specific_error \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
fastvideo/tests/platforms/test_mps_vsa_error.py \
fastvideo/tests/platforms/test_cpu_sdpa.py \
-v -s --timeout=120 -o faulthandler_timeout=120
# Same tests on MLX's CPU backend. Hosted macOS runners are scarce and
# slower to schedule; this Linux job gives fast PR signal on the identical
# graph (the parity tests were designed to be backend-agnostic), while the
# macOS job above stays the source of truth for Metal behavior.
mlx-smoke-linux-cpu:
if: github.event_name == 'workflow_dispatch' || github.event.pull_request.draft != true
runs-on: ubuntu-latest
timeout-minutes: 20
env:
FASTVIDEO_ATTENTION_BACKEND: TORCH_SDPA
TOKENIZERS_PARALLELISM: "false"
MASTER_ADDR: localhost
MASTER_PORT: "29513"
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- uses: astral-sh/setup-uv@v3
- name: Install lightweight MLX smoke dependencies (CPU backend)
run: |
uv pip install --system \
--index-url https://download.pytorch.org/whl/cpu \
torch==2.12.0 torchvision torchaudio
uv pip install --system \
pytest pytest-timeout numpy scipy pillow imageio einops cloudpickle filelock \
PyYAML diffusers huggingface_hub remote-pdb safetensors loguru "mlx[cpu]" \
"ftfy>=6.3.1" "opencv-python>=4.10.0.84" psutil "transformers>=5.0.0"
- name: Run MLX smoke tests (CPU backend)
run: |
python -m pytest \
fastvideo/mlx_runtime/tests/ \
fastvideo/tests/mlx/test_dmd_sampling.py \
fastvideo/tests/mlx/test_memory_limits.py \
fastvideo/tests/mlx/test_quant_capability.py \
fastvideo/tests/mlx/test_mlx_dit_parity.py \
fastvideo/tests/mlx/test_mlx_compile_parity.py \
fastvideo/tests/mlx/test_mlx_checkpoint.py \
fastvideo/tests/mlx/test_mlx_checkpoint_compat.py \
fastvideo/tests/mlx/test_mlx_affine_dq_gemm.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_parity.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa_regressions.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_mode.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_fastwan_benchmark.py \
fastvideo/tests/mlx/test_taehv_decode.py \
fastvideo/tests/mlx/test_frame_upsample.py \
fastvideo/tests/mlx/test_mlx_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_refine.py \
fastvideo/tests/mlx/test_mlx_prompt_enhance.py \
fastvideo/tests/mlx/test_mlx_prompt_to_video_decode.py \
fastvideo/tests/mlx/test_mlx_wan22_prompt_cache_fingerprint.py \
fastvideo/tests/mlx/test_wan22_sample.py \
fastvideo/tests/mlx/test_windowed_attention.py \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_download_unavailable_has_specific_error \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
fastvideo/tests/platforms/test_mps_vsa_error.py \
fastvideo/tests/platforms/test_cpu_sdpa.py \
-v -s --timeout=120 -o faulthandler_timeout=120
+19 -3
View File
@@ -27,14 +27,25 @@ jobs:
ref: ${{ inputs.ref || '' }}
# For PR events, lint the PR head — but keep the hook definitions from
# the base branch so an untrusted PR cannot alter what gets executed.
- name: Save trusted hook config
# The gate scripts are saved too: the self-test step below executes them,
# so it must run the base-branch copies, not the PR head's.
- name: Save trusted hook config and gate scripts
if: github.event_name == 'pull_request_target'
run: cp .pre-commit-config.yaml "$RUNNER_TEMP/trusted-pre-commit-config.yaml"
- uses: actions/checkout@v4
run: |
cp .pre-commit-config.yaml "$RUNNER_TEMP/trusted-pre-commit-config.yaml"
cp -a .github/scripts "$RUNNER_TEMP/trusted-scripts"
echo "GATE_SCRIPTS_DIR=$RUNNER_TEMP/trusted-scripts" >> "$GITHUB_ENV"
# allow-unsafe-pr-checkout acknowledges checkout's pull_request_target
# guard: the head is data for the trusted hooks to lint; nothing from it
# is executed (config and gate scripts are pinned to the base branch
# above) and credentials are not persisted. SHA-pinned to v4.4.0 because
# actionlint's action schema does not know the new input yet.
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
if: github.event_name == 'pull_request_target'
with:
ref: ${{ github.event.pull_request.head.sha }}
persist-credentials: false
allow-unsafe-pr-checkout: true
- name: Restore trusted hook config
if: github.event_name == 'pull_request_target'
run: cp "$RUNNER_TEMP/trusted-pre-commit-config.yaml" .pre-commit-config.yaml
@@ -47,3 +58,8 @@ jobs:
- uses: pre-commit/action@v3.0.1
with:
extra_args: --all-files --hook-stage manual
# After pre-commit so a self-test failure cannot mask lint failures.
# GATE_SCRIPTS_DIR points at the base-branch copy on fork PRs (set above);
# push / workflow_call runs use the checked-out tree directly.
- name: Full-suite gate self-test
run: bash "${GATE_SCRIPTS_DIR:-.github/scripts}/test_gate_full_suite.sh"
+44
View File
@@ -0,0 +1,44 @@
name: Scheduled Full SSIM
on:
schedule:
- cron: "0 5 * * 0"
workflow_dispatch:
permissions:
contents: read
jobs:
trigger:
if: github.repository == 'hao-ai-lab/FastVideo'
runs-on: ubuntu-latest
steps:
- name: Trigger weekly full SSIM on Slinky Slurm
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
SOURCE_SHA: ${{ github.sha }}
SOURCE_BRANCH: ${{ github.event.repository.default_branch }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
set -euo pipefail
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
-H "Content-Type: application/json" \
--data-raw "$(jq -n \
--arg commit "$SOURCE_SHA" \
--arg branch "$SOURCE_BRANCH" \
'{
commit: $commit,
branch: $branch,
message: "Weekly full SSIM on Slinky Slurm",
ignore_pipeline_branch_filters: true,
env: {
TEST_SCOPE: "scheduled",
FULL_SUITE: "false",
TEST_TYPE: "ssim",
PR_NUMBER: "false",
PR_TITLE: "Scheduled full SSIM"
}
}')"
+43 -48
View File
@@ -32,8 +32,7 @@ jobs:
}
core.setOutput('has_write', String(hasWrite));
- name: Add ready label and react
id: label
- name: Add ready label
if: steps.perm.outputs.has_write == 'true'
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
@@ -41,50 +40,33 @@ jobs:
const owner = context.repo.owner;
const repo = context.repo.repo;
const prNumber = context.payload.issue.number;
try { await github.rest.issues.removeLabel({ owner, repo, issue_number: prNumber, name: 'ready' }); } catch {}
await github.rest.issues.addLabels({ owner, repo, issue_number: prNumber, labels: ['ready'] });
- name: React to comment
if: steps.perm.outputs.has_write == 'true'
continue-on-error: true
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
script: |
await github.rest.reactions.createForIssueComment({
owner, repo,
owner: context.repo.owner,
repo: context.repo.repo,
comment_id: context.payload.comment.id,
content: 'rocket',
});
const { data: pr } = await github.rest.pulls.get({ owner, repo, pull_number: prNumber });
core.setOutput('pr_sha', pr.head.sha);
core.setOutput('pr_branch', pr.head.ref);
core.setOutput('pr_number', String(prNumber));
- name: Trigger Full Suite
if: steps.perm.outputs.has_write == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_SHA: ${{ steps.label.outputs.pr_sha }}
PR_BRANCH: ${{ steps.label.outputs.pr_branch }}
PR_NUMBER: ${{ steps.label.outputs.pr_number }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
-H "Content-Type: application/json" \
--data-raw "$(jq -n \
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER} (via /merge)" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
branch: $branch,
message: $message,
ignore_pipeline_branch_filters: true,
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: "full",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring)
}
}')"
trigger-merge-gate:
needs: handle-merge
if: needs.handle-merge.result == 'success'
permissions:
actions: read
contents: read
pull-requests: read
uses: ./.github/workflows/ci-trigger-full-suite.yml
with:
pr_number: ${{ github.event.issue.number }}
secrets:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
parse-command:
if: >-
@@ -125,7 +107,7 @@ jobs:
set -euo pipefail
TEST_NAME=$(echo "$COMMENT" | grep -oP '(?<=/test\s)\S+' | head -1 || true)
VALID="encoder vae transformer kernel unit dreamverse ssim training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval full fastcheck pre-commit"
VALID="encoder vae transformer kernel unit dreamverse ssim golden-gate training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval unit-ci kernel-ci dreamverse-ci ssim-ci golden-gate-ci encoder-ci vae-ci transformer-ci lora-inference-ci lora-training-ci lora-extraction-ci training-ci distillation-ci self-forcing-ci vsa-ci vmoba-ci performance-ci api-ci train-framework-ci eval-ci full fastcheck pre-commit"
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
echo "Unknown test: '$TEST_NAME'. Valid: $VALID"
exit 1
@@ -133,8 +115,18 @@ jobs:
declare -A MAP=(
[encoder]=encoder [vae]=vae [transformer]=transformer
[kernel]=kernel_tests [unit]=unit_test [dreamverse]=dreamverse_app
[ssim]=ssim [training]=training
[kernel]=kernel_tests [unit]=unit_test [unit-ci]=unit_test_ci
[kernel-ci]=kernel_tests_ci [dreamverse-ci]=dreamverse_app_ci
[ssim-ci]=ssim_ci [vmoba-ci]=inference_vmoba_ci
[golden-gate-ci]=golden_gate_ci [training-ci]=training_ci
[encoder-ci]=encoder_ci [vae-ci]=vae_ci [transformer-ci]=transformer_ci
[lora-inference-ci]=inference_lora_ci [lora-training-ci]=training_lora_ci
[lora-extraction-ci]=lora_extraction_ci [distillation-ci]=distillation_dmd_ci
[self-forcing-ci]=self_forcing_ci [vsa-ci]=training_vsa_ci
[performance-ci]=performance_ci [api-ci]=api_server_ci
[train-framework-ci]=train_framework_ci [eval-ci]=eval_ci
[dreamverse]=dreamverse_app
[ssim]=ssim [golden-gate]=golden_gate [training]=training
[lora-inference]=inference_lora [lora-training]=training_lora
[lora-extraction]=lora_extraction
[distillation]=distillation_dmd [self-forcing]=self_forcing
@@ -241,6 +233,7 @@ jobs:
TEST_SCOPE: ${{ needs.parse-command.outputs.test_scope }}
FULL_SUITE: ${{ needs.parse-command.outputs.full_suite }}
TEST_TYPE: ${{ needs.parse-command.outputs.test_type }}
PR_TITLE: ${{ github.event.issue.title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
@@ -257,6 +250,7 @@ jobs:
--arg full_suite "$FULL_SUITE" \
--arg test_type "$TEST_TYPE" \
--arg pr_number "$PR_NUMBER" \
--arg pr_title "$PR_TITLE" \
'{
commit: $commit,
branch: $branch,
@@ -266,8 +260,9 @@ jobs:
pull_request_base_branch: "main",
env: {
TEST_SCOPE: $test_scope,
FULL_SUITE: $full_suite,
TEST_TYPE: $test_type,
PR_NUMBER: $pr_number
}
}')"
FULL_SUITE: $full_suite,
TEST_TYPE: $test_type,
PR_NUMBER: $pr_number,
PR_TITLE: $pr_title
}
}')"
+208 -31
View File
@@ -1,63 +1,232 @@
name: Trigger Full Suite
name: Trigger Merge Gate
on:
pull_request_target:
types: [labeled, synchronize]
workflow_call:
inputs:
pr_number:
description: Pull request number to enter into the merge gate
required: true
type: number
secrets:
BUILDKITE_API_TOKEN:
required: true
permissions:
contents: read
pull-requests: read
concurrency:
group: full-suite-${{ github.event.pull_request.number }}
cancel-in-progress: false
actions: read
jobs:
trigger:
if: >-
(github.event.action == 'labeled' && github.event.label.name == 'ready')
inputs.pr_number > 0
|| (github.event.action == 'labeled' && github.event.label.name == 'ready')
|| github.event.action == 'synchronize'
runs-on: ubuntu-latest
# Job-level concurrency: only this guarded job acquires the group, so an
# unrelated `labeled` event (which skips the job) cannot cancel an in-flight
# gate and then skip its replacement. The newest real trigger (`ready`,
# push, or `/merge`) supersedes the in-flight run, whose Buildkite build the
# cancel step below replaces.
concurrency:
group: merge-gate-${{ inputs.pr_number || github.event.pull_request.number }}
cancel-in-progress: true
# Gate below may wait for cheap checks (up to MAX_WAIT_SECS = 25 min).
timeout-minutes: 35
steps:
- name: Check ready label
id: check
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
env:
CALLED_PR_NUMBER: ${{ inputs.pr_number }}
with:
script: |
const eventPrNumber = context.payload.pull_request?.number;
const calledPrNumber = Number(process.env.CALLED_PR_NUMBER);
const prNumber = eventPrNumber ?? calledPrNumber;
if (!Number.isSafeInteger(prNumber) || prNumber <= 0) {
core.setFailed(`Invalid pull request number: ${process.env.CALLED_PR_NUMBER}`);
return;
}
const { data: pr } = await github.rest.pulls.get({
owner: context.repo.owner,
repo: context.repo.repo,
pull_number: context.payload.pull_request.number,
pull_number: prNumber,
});
if (pr.state !== 'open') {
core.setFailed(`PR #${prNumber} is not open.`);
return;
}
if (pr.base.repo.full_name !== context.payload.repository.full_name
|| pr.base.ref !== context.payload.repository.default_branch) {
core.setFailed(`PR #${prNumber} does not target this repository's default branch.`);
return;
}
const hasReady = pr.labels.some(l => l.name === 'ready');
core.setOutput('has_ready', String(hasReady));
if (!hasReady) core.info('No ready label — skipping Full Suite trigger.');
core.setOutput('changed_files', String(pr.changed_files));
core.setOutput('pr_number', String(pr.number));
core.setOutput('head_sha', pr.head.sha);
core.setOutput('head_ref', pr.head.ref);
core.setOutput('base_sha', pr.base.sha);
core.setOutput('title', pr.title);
if (!hasReady) core.info('No ready label — skipping merge-gate trigger.');
- name: Cancel previous Buildkite builds
# Cancelling stale builds only saves agent time. If it cannot run, the
# merge gate must still be triggered by the steps below, so a failure
# here is reported and stepped over rather than ending the job.
continue-on-error: true
timeout-minutes: 3
if: steps.check.outputs.has_ready == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
run: |
# Find running builds for this branch with TEST_SCOPE=full and cancel them
builds=$(curl -sS -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds?branch=${PR_BRANCH}&state=running,scheduled" \
| jq -r '.[] | select(try (.env.TEST_SCOPE == "full") catch false) | .number')
for build_num in $builds; do
echo "Cancelling Buildkite build #$build_num"
curl -sS -X PUT -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds/${build_num}/cancel"
done
- name: Trigger Buildkite Full Suite
if: steps.check.outputs.has_ready == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_SHA: ${{ github.event.pull_request.head.sha }}
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
PR_NUMBER: ${{ github.event.pull_request.number }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_BRANCH: ${{ steps.check.outputs.head_ref }}
PR_NUMBER: ${{ steps.check.outputs.pr_number }}
run: |
set -euo pipefail
response_file=$(mktemp)
builds_file=$(mktemp)
trap 'rm -f "$response_file" "$builds_file"' EXIT
if [[ ! "$PR_NUMBER" =~ ^[1-9][0-9]*$ ]]; then
echo "::warning::Invalid pull request number; stale Buildkite builds may continue."
exit 1
fi
if ! curl -sS --fail-with-body --connect-timeout 5 --max-time 20 --get \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
--data-urlencode "branch=$PR_BRANCH" \
--data-urlencode "state[]=running" \
--data-urlencode "state[]=scheduled" \
--data-urlencode "state[]=failing" \
--data-urlencode "exclude_jobs=true" \
--data-urlencode "exclude_pipeline=true" \
--output "$response_file" \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds"; then
echo "::warning::Could not list Buildkite builds; stale merge-gate builds may continue."
exit 1
fi
if ! jq -e '
if type != "array" then false
else all(.[];
if type != "object" then false
else
(.number | if type == "number" then . > 0 and floor == . else false end)
and (
(.env? | if . == null then {} else . end) as $env
| if ($env | type) != "object" then false
else
($env.TEST_SCOPE? | . == null or type == "string")
and ($env.PR_NUMBER? | . == null or type == "string")
end
)
end
)
end
' "$response_file" >/dev/null 2>&1; then
# Do not print the response body: it is remote data and may contain
# multiline values that would be interpreted as workflow commands.
echo "::warning::Buildkite returned an invalid build list; stale merge-gate builds may continue."
exit 1
fi
# Match both branch and PR number: forks can reuse the same branch name.
if ! jq -r --arg pr_number "$PR_NUMBER" '
.[]
| select((.env.TEST_SCOPE? == "merge") and (.env.PR_NUMBER? == $pr_number))
| .number
' "$response_file" > "$builds_file"; then
echo "::warning::Could not select stale Buildkite builds; stale merge-gate builds may continue."
exit 1
fi
cancellation_failed=0
while IFS= read -r build_num; do
echo "Cancelling Buildkite build #$build_num"
if ! curl -sS --fail-with-body --connect-timeout 5 --max-time 20 -o /dev/null -X PUT \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds/${build_num}/cancel"; then
echo "::warning::Could not cancel Buildkite build #$build_num; trying remaining builds."
cancellation_failed=1
fi
done < "$builds_file"
if (( cancellation_failed != 0 )); then
exit 1
fi
# Check out the immutable BASE SHA: neither pull_request_target nor the
# privileged slash-command call may run code from the untrusted PR head.
- name: Checkout trusted merge planner
if: steps.check.outputs.has_ready == 'true'
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
ref: ${{ steps.check.outputs.base_sha }}
persist-credentials: false
- name: Collect changed paths
if: steps.check.outputs.has_ready == 'true'
env:
GH_TOKEN: ${{ github.token }}
PR_NUMBER: ${{ steps.check.outputs.pr_number }}
EXPECTED_CHANGED_FILES: ${{ steps.check.outputs.changed_files }}
run: |
set -euo pipefail
changed_json="$RUNNER_TEMP/merge-changed-files.json"
changed_paths="$RUNNER_TEMP/merge-changed-paths.txt"
if gh api --paginate --slurp \
"repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}/files?per_page=100" \
> "$changed_json"; then
observed=$(jq '[.[][] | .filename] | unique | length' "$changed_json")
if [ "$observed" = "$EXPECTED_CHANGED_FILES" ]; then
jq -r '.[][] | .filename, (.previous_filename // empty)' "$changed_json" \
| sort -u > "$changed_paths"
else
echo "::warning::Changed-file API returned $observed of $EXPECTED_CHANGED_FILES paths; selecting all merge lanes."
echo '__FASTVIDEO_CI_PLAN_ALL__' > "$changed_paths"
fi
else
echo "::warning::Changed-file API failed; selecting all merge lanes."
echo '__FASTVIDEO_CI_PLAN_ALL__' > "$changed_paths"
fi
- name: Select minimal merge tests
id: plan
if: steps.check.outputs.has_ready == 'true'
run: |
python3 .github/scripts/plan_merge_ci.py \
--paths-file "$RUNNER_TEMP/merge-changed-paths.txt" \
--github-output "$GITHUB_OUTPUT" \
--summary-file "$GITHUB_STEP_SUMMARY"
- name: Wait for pre-commit and docs build
if: steps.check.outputs.has_ready == 'true'
env:
GH_TOKEN: ${{ github.token }}
PR_SHA: ${{ steps.check.outputs.head_sha }}
PR_NUMBER: ${{ steps.check.outputs.pr_number }}
run: bash .github/scripts/gate_full_suite.sh
- name: Trigger Buildkite merge gate
if: steps.check.outputs.has_ready == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_SHA: ${{ steps.check.outputs.head_sha }}
PR_BRANCH: ${{ steps.check.outputs.head_ref }}
PR_NUMBER: ${{ steps.check.outputs.pr_number }}
PR_TITLE: ${{ steps.check.outputs.title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
MERGE_TEST_PLAN: ${{ steps.plan.outputs.merge_test_plan }}
MERGE_GOLDEN_TESTS: ${{ steps.plan.outputs.merge_golden_tests }}
MERGE_SSIM_TESTS: ${{ steps.plan.outputs.merge_ssim_tests }}
MERGE_PLAN_LABEL: ${{ steps.plan.outputs.merge_plan_label }}
run: |
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
@@ -66,7 +235,11 @@ jobs:
--data-raw "$(jq -n \
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER}" \
--arg message "Merge gate [${MERGE_PLAN_LABEL}] for PR #${PR_NUMBER}" \
--arg pr_title "$PR_TITLE" \
--arg merge_test_plan "$MERGE_TEST_PLAN" \
--arg merge_golden_tests "$MERGE_GOLDEN_TESTS" \
--arg merge_ssim_tests "$MERGE_SSIM_TESTS" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
@@ -76,8 +249,12 @@ jobs:
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: "full",
TEST_SCOPE: "merge",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring)
MERGE_TEST_PLAN: $merge_test_plan,
MERGE_GOLDEN_TESTS: $merge_golden_tests,
MERGE_SSIM_TESTS: $merge_ssim_tests,
PR_NUMBER: ($pr_id | tostring),
PR_TITLE: $pr_title
}
}')"
+4 -4
View File
@@ -38,17 +38,17 @@ jobs:
**How our CI works:**
PRs run a two-tier CI system:
PRs run a three-tier CI system:
1. **Pre-commit** — formatting (yapf), linting (ruff), type checking (mypy). Runs immediately on every PR.
2. **Fastcheck** — core GPU tests (encoders, VAEs, transformers, kernels, unit tests). Runs automatically via Buildkite on relevant file changes (~10-15 min).
3. **Full Suite** — integration tests, training pipelines, SSIM regression. Runs only when a reviewer adds the `ready` label.
2. **Fastcheck** — six core GPU lanes run automatically via Buildkite (~10-15 min).
3. **Merge gate** — a reviewer adds `ready`; changed paths select only the relevant integration, training, golden, or SSIM coverage.
**Before your PR is reviewed:**
- [ ] `pre-commit run --all-files` passes locally
- [ ] You've added or updated tests for your changes
- [ ] The PR description explains what and why
If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and Full Suite results appear in the Checks section below.
If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and merge-gate results appear in the Checks section below.
**Useful links:**
- [Contributing Guide](https://hao-ai-lab.github.io/FastVideo/contributing/overview/)
+41 -7
View File
@@ -13,16 +13,28 @@ on:
required: false
default: false
type: boolean
# Auto-rebuild the CUDA images when their Dockerfile changes on main. The CUDA
# matrix is the only lane that builds from docker/Dockerfile, so a path-scoped
# push trigger is a sufficient change detector on its own -- no separate
# detect-changes/paths-filter job is needed now that there is a single
# in-scope Dockerfile. Dreamverse (apps/dreamverse/docker/Dockerfile) and the
# rocm Dockerfile stay manual-dispatch only.
build_ci_runner_image:
description: 'Build the ARM64 CUDA 13 CI runner image (sm_100)'
required: false
default: false
type: boolean
# Auto-rebuild the CUDA images when a repository-controlled image input
# changes on main. This includes the trusted SM89 kernel artifact's source,
# metadata/key helper, ABI dependency metadata, and build orchestration.
# Dreamverse (apps/dreamverse/docker/Dockerfile) and the ROCm Dockerfile stay
# manual-dispatch only.
push:
branches: [main]
paths:
- '.dockerignore'
- '.github/workflows/_template-build-image.yml'
- '.github/workflows/infra-build-image.yml'
- '.gitmodules'
- 'docker/Dockerfile'
- 'docker/uv-excludes'
- 'fastvideo-kernel/**'
- 'fastvideo/tests/modal/kernel_build_cache.py'
- 'pyproject.toml'
permissions:
@@ -50,7 +62,7 @@ jobs:
# 2.8.3 comes from the architecture-specific prebuilt releases.
build-cuda-images:
# Runs on a manual dispatch when build_cuda_matrix is set, or automatically
# on a push that changed docker/Dockerfile (inputs are null on push). The
# on an in-scope main push (inputs are null on push). The
# repository guard keeps fork syncs from auto-building; manual dispatch
# still works in forks.
if: ${{ (github.event_name == 'push' && github.repository == 'hao-ai-lab/FastVideo') || github.event.inputs.build_cuda_matrix == 'true' }}
@@ -191,6 +203,28 @@ jobs:
docker buildx imagetools create "${TAG_ARGS[@]}" "${IMAGE_REFS[@]}"
docker buildx imagetools inspect "${TAGS[0]}"
# The CI runner is ARM64 like DGX Spark, but targets sm_100 rather than sm_121.
# Publish a single-architecture variant so the self-hosted CI runner can reuse
# the exact prebuilt kernel instead of compiling it in every job.
build-ci-runner-image:
if: ${{ (github.event_name == 'push' && github.repository == 'hao-ai-lab/FastVideo') || github.event.inputs.build_ci_runner_image == 'true' }}
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.12'
dockerfile_path: docker/Dockerfile
tag_suffix: py3.12-cuda13.0.0-sm100
runner: ubuntu-24.04-arm
architecture: arm64
build_args: |
PYTHON_VERSION=3.12
CUDA_VERSION=13.0.0
UV_TORCH_BACKEND=cu130
TORCH_CUDA_ARCH_LIST=10.0
CMAKE_BUILD_PARALLEL_LEVEL=1
FLASH_ATTN_WHEEL_TAG=cu130torch2.12
FLASH_ATTN_WHEEL_RELEASE_ARM64=https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.9.22
secrets: inherit
# Dreamverse matrix: {backend, UI} x {12.6.3, 13.0.0}, Python 3.12. Torch backend
# matches the base CUDA (cu126 / cu130). Keep these images amd64-only until the
# required FA4 dependency stack is available and validated on arm64.
+2
View File
@@ -6,6 +6,7 @@ on:
paths:
- 'docs/**'
- 'examples/**'
- 'scripts/inference/**'
- 'mkdocs.yml'
- 'requirements-mkdocs.in'
- 'requirements-mkdocs.txt'
@@ -16,6 +17,7 @@ on:
paths:
- 'docs/**'
- 'examples/**'
- 'scripts/inference/**'
- 'mkdocs.yml'
- 'requirements-mkdocs.in'
- 'requirements-mkdocs.txt'
+25 -16
View File
@@ -62,12 +62,13 @@ jobs:
cuda-version: '13.0.0'
torch-cuda-short: 'cu130'
platform:
# x86_64 builds the full cu126 + cu130 set (cu130 ships the consumer
# Blackwell sm_120a FP4 kernels).
# x86_64 builds the full cu126 + cu130 set. cu130 ships the
# data-center Blackwell sm_100a/sm_103a VSA and consumer sm_120a FP4
# kernels.
- os: ubuntu-22.04
arch: x86_64
wheel-plat: manylinux_2_35_x86_64
# aarch64 is Blackwell (GB200 sm_100a + DGX Spark / consumer sm_120a), not
# aarch64 is Blackwell (GB200 sm_100a/sm_103a + sm_120a + DGX Spark sm_121a), not
# Hopper, and Blackwell needs CUDA >= 12.8 — so only the cu130 leg applies.
# Added via include so x86 keeps cu126 + cu130 while aarch64 stays cu130-only.
include:
@@ -124,7 +125,7 @@ jobs:
- name: Install dependencies (GCC, Clang, CUDA Paths, Git)
run: |
sudo apt update
sudo apt install -y git patchelf gcc-11 g++-11 clang-11
sudo apt install -y git gcc-11 g++-11 clang-11
sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
# Allow Git to Access Safe Directory
@@ -163,22 +164,24 @@ jobs:
cd fastvideo-kernel
git submodule update --init --recursive # Ensure ThunderKittens submodule is initialized
# Release builds run on GPU-less runners, so set kernels + arch explicitly:
# * aarch64 = Blackwell (GB200 sm_100a + DGX Spark/consumer sm_120a), NOT
# Hopper, so TK (sm_90a wgmma) is OFF. The C++ FP4 (attn_qat_infer, SM120)
# covers sm_120a; turbodiffusion covers sm_100a+sm_120a. The sm_100 FP4
# forward is the FA4 CuTe DSL path in the fastvideo package (PR #1221),
# * aarch64 = Blackwell (GB200 sm_100a/sm_103a + sm_120a + DGX Spark sm_121a), NOT
# Hopper, so TK (sm_90a wgmma) is OFF. The C++ FP4 (attn_qat_infer)
# covers sm_120a+sm_121a; turbodiffusion covers every listed arch. The
# sm_100 FP4 forward is the FA4 CuTe DSL path in the fastvideo package (PR #1221),
# JIT-compiled at runtime — not built into this wheel.
# * x86_64 cu130 = Hopper TK + consumer Blackwell sm_120a FP4.
# * x86_64 cu130 = Hopper TK + data-center Blackwell sm_100a/sm_103a VSA
# + consumer Blackwell sm_120a FP4.
# * x86_64 cu126 = Hopper TK only (older drivers; CUDA < 12.8 has no FP4).
# The per-arch split in CMakeLists pins the FP4 targets to sm_120a and builds
# the main extension for the full arch list. CMAKE_BUILD_PARALLEL_LEVEL caps
# The per-arch split in CMakeLists pins the FP4 targets to requested
# sm_120a/sm_121a and builds the main extension for the full arch list.
# CMAKE_BUILD_PARALLEL_LEVEL caps
# Ninja so heavy CUTLASS/TK template TUs don't OOM the 16 GB runner (exit 143).
if [ "${{ matrix.platform.arch }}" = "aarch64" ]; then
export TORCH_CUDA_ARCH_LIST="10.0a;12.0a"
export TORCH_CUDA_ARCH_LIST="10.0a;10.3a;12.0a;12.1a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=OFF -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON"
export CMAKE_BUILD_PARALLEL_LEVEL=1
elif [ "${{ matrix.torch-cuda.torch-cuda-short }}" = "cu130" ]; then
export TORCH_CUDA_ARCH_LIST="9.0a;12.0a"
export TORCH_CUDA_ARCH_LIST="9.0a;10.0a;10.3a;12.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=ON -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON -DCMAKE_CUDA_ARCHITECTURES=90a"
# A single FP4 TU (attn_qat_infer) can use ~8-12 GB on its own, so serialize.
export CMAKE_BUILD_PARALLEL_LEVEL=1
@@ -194,7 +197,11 @@ jobs:
python -m build --wheel --outdir dist
# Fix the wheel to be manylinux compliant
uv pip install --system auditwheel
# Ubuntu 22.04 ships patchelf 0.14.3, while current auditwheel
# requires at least 0.14.5. Use the stable PyPI binary on both
# x86_64 and aarch64 release runners.
uv pip install --system auditwheel patchelf==0.17.2.4
patchelf --version
# Point auditwheel at torch libs, but do not vendor them into the wheel.
TORCH_LIB_DIR=$(python - <<'PY'
import os
@@ -211,7 +218,8 @@ jobs:
--exclude libtorch.so \
--exclude libc10.so \
--exclude libc10_cuda.so \
--exclude libtorch_python.so
--exclude libtorch_python.so \
--exclude libnccl.so.2
# Move fixed wheels back to dist for upload consistency
rm dist/*.whl
mv fixed_dist/*.whl dist/
@@ -243,7 +251,8 @@ jobs:
- name: Download PyPI wheels
# Publish the cu130 (CUDA 13) wheels to PyPI for both architectures:
# x86_64 — Hopper sm_90a TK + consumer Blackwell sm_120a FP4
# aarch64 — Blackwell: turbodiffusion (sm_100a/sm_120a) + C++ FP4 (sm_120a);
# aarch64 — Blackwell: turbodiffusion (sm_100a/sm_103a/sm_120a/sm_121a)
# + C++ FP4 (sm_120a/sm_121a);
# no TK (Hopper). sm_100 FP4 forward is the FA4 CuTe DSL path in the
# fastvideo package (#1221), shipped/JIT separately.
# The x86_64 cu126 wheel stays available as a build artifact / GitHub-release asset.
+17 -2
View File
@@ -6,6 +6,7 @@ results/
wandb/
*.ipynb
*.jpg
!examples/datasets/lingbotworld2/image.jpg
*.safetensors
*.mp4
*.png
@@ -22,6 +23,7 @@ Miniconda3-latest-Linux-x86_64.sh
*validation/
data/
outputs/
outputs_audio/
outputs_video
checkpoints/
sbatch.sh
@@ -34,6 +36,7 @@ env
*.log
weights/
logs/
/Z-Image/
official_weights/
converted_weights/
@@ -52,6 +55,8 @@ eggs/
# MkDocs documentation
site/
docs/assets/cookbook-serving.json
examples/serving/clients/node_modules/
docs/getting_started/examples/
docs/examples/
docs/inference/examples/
@@ -72,8 +77,8 @@ docs/distillation/examples/
# Python pickle files
*.pkl
# Reference videos
!fastvideo/tests/ssim/reference_videos/**/*.mp4
# Reference videos (negations must come after the catch-all on line below)
!fastvideo/tests/nightly/reference_video_*.mp4
# Static images
!docs/assets/images/**/*.png
@@ -127,8 +132,18 @@ apps/dreamverse/web/.env.production.local
.sisyphus/
openspec/
fastvideo/tests/ssim/reference_videos/**
!fastvideo/tests/ssim/reference_videos/**/*.mp4
!fastvideo/tests/ssim/reference_videos/**/*.png
fastvideo/tests/ssim/.reference_videos_download.lock
# Local H3 MLX kernel / exactness benches (JSON, logs, frames, videos)
.kernel_bench/
# Editor logs and local Python version pins (accidentally committed)
*.nvimlog
.nvimlog
.python-version
/LTX-2-Reference/
/DFDReference/
scripts/benchmarks/minimax_h3_pro6000/headline_results/
fastvideo/tests/ssim/.reference_videos_download.lock
+3
View File
@@ -7,3 +7,6 @@
[submodule "fastvideo/third_party/eval/vbench"]
path = fastvideo/third_party/eval/vbench
url = https://github.com/Vchitect/VBench.git
[submodule "fastvideo/third_party/eval/vqeval"]
path = fastvideo/third_party/eval/vqeval
url = https://github.com/JiusiServe/LongVideoSparseAttention.git
+2 -1
View File
@@ -9,7 +9,7 @@ exclude: |
tests/.*|
scripts/.*|
fastvideo/dataset/.*|
fastvideo/models/.*|
fastvideo/models/(?!wan/(config|vae_config|pipeline_config|definition|__init__)\.py$).*|
^apps/dreamverse/web/.*|
examples/.*|
\.agents/.*|
@@ -22,6 +22,7 @@ repos:
hooks:
- id: yapf
args: [--in-place, --verbose]
language_version: python3.12
additional_dependencies: [toml] # TODO: Remove when yapf is upgraded
- repo: https://github.com/astral-sh/ruff-pre-commit
rev: v0.11.12
+4
View File
@@ -66,14 +66,18 @@ Local guidance lives next to the code. Read the in-scope file before editing:
| `fastvideo/AGENTS.md` | Core package map, public API, registry-driven model dispatch |
| `fastvideo/configs/AGENTS.md` | Arch + pipeline config dataclasses, `param_names_mapping` |
| `fastvideo/models/AGENTS.md` | DiT / VAE / encoder / scheduler / loader layout (pre-commit excluded) |
| `fastvideo/models/wan/AGENTS.md` | Wan family-local transformers, VAE, configs, and the SP sharding invariant |
| `fastvideo/layers/AGENTS.md` | Tensor-parallel linear/attention layer rules for ports |
| `fastvideo/attention/AGENTS.md` | Backend registry + env-var override |
| `fastvideo/pipelines/AGENTS.md` | Stage ABC, `basic/<model>/`, `preprocess/`, presets |
| `fastvideo/pipelines/basic/wan/AGENTS.md` | Wan sampling stages, first-frame conditioning, DMD/causal boundaries |
| `fastvideo/pipelines/basic/magi_human/AGENTS.md` | MagiHuman umbrella repo, lazy-loaded components, packing invariants |
| `fastvideo/training/AGENTS.md` | Legacy monolithic pipelines (frozen for existing models) |
| `fastvideo/train/AGENTS.md` | New modular trainer (methods × models × callbacks, YAML) |
| `fastvideo/tests/AGENTS.md` | Test taxonomy, conftest, pre-commit-excluded path |
| `fastvideo/tests/ssim/AGENTS.md` | GPU SSIM regression authoring + reference video sync |
| `scripts/checkpoint_conversion/AGENTS.md` | Adding a converter for a new HF/official checkpoint |
| `apps/dreamverse/AGENTS.md` | DreamVerse app structure and conventions |
## Critical: Two Training Stacks Coexist
+18 -5
View File
@@ -3,13 +3,19 @@
</div>
<p align="center">
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://github.com/hao-ai-lab/FastVideo/discussions/1097" target="_blank"> <b> WeChat </b> </a> |
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://haoailab.com/FastVideo/cookbook/"><b>Cookbook</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://github.com/hao-ai-lab/FastVideo/discussions/1097" target="_blank"> <b> WeChat </b> </a> |
</p>
**FastVideo is a unified post-training and real-time inference framework for accelerated video generation.**
## NEWS
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), check out the [Blog](https://haoailab.com/blogs/fastwan-qad/).
- `2026/10/06`: FastH3 V2 now runs on a single consumer machine: NVIDIA RTX 5090, RTX 4090 and RTX PRO 6000 GPUs, DGX Spark and Apple Silicon. We also release [FastH3 Trim](https://huggingface.co/FastVideo/FastVideo-FastH3-Trim-8-Step-NVFP4), an experimental pruned model that is 4.2× smaller than base H3 and runs in as little as 8 GB of GPU memory. Get the [models](https://huggingface.co/collections/FastVideo/fastvideo-fasth3) and read the [Blog](https://haoailab.com/blogs/fasth3-rtx/).
- `2026/10/06`: FastVideo now supports [Kandinsky 6](https://x.com/kandinskylab_ai/status/2107374635218055345) from Kandinsky Lab: text- and image-to-video with synchronized audio (base and 10-step distilled pi-Flow checkpoints) plus video super-resolution up to 4x. See the [Kandinsky 6 recipes](https://haoailab.com/FastVideo/cookbook/kandinsky6/).
- `2026/09/15`: Release [FastH3 8-Step V2](https://huggingface.co/FastVideo/FastVideo-FastH3-8-Step-V2), an eight-forward data-free DMD2 checkpoint distilled from MiniMax-H3 with 80% Video Sparse Attention. Run it with `examples/inference/basic/basic_fasth3_8step.py` or the [FastH3 8-Step V2 recipe](https://haoailab.com/FastVideo/cookbook/minimax-h3/).
- `2026/09/01`: FastH3 now runs locally on Apple Silicon through MLX and on NVIDIA DGX Spark through CUDA 13, including two-Spark inference. Follow the [FastH3 recipes](https://haoailab.com/FastVideo/cookbook/minimax-h3/) and read the [Blog](https://haoailab.com/blogs/fasth3-local/).
- `2026/08/27`: [FastH3 Preview v1](https://haoailab.com/blogs/fasth3-preview/) is an open-weight 4-step sparse-distilled MiniMax-H3 model for synchronized video-and-audio generation, developed in collaboration with [Nuva Lab](https://nuvalab.ai/) and the [NVIDIA FastGen team](https://github.com/NVlabs/FastGen). Download the recommended [VSA / Data-Free weights](https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree), or see the [full FastH3 collection](https://huggingface.co/collections/FastVideo/fastvideo-fasth3).
- `2026/08/19`: FastVideo now supports MLX on Apple Silicon with [FastMetal-QAD](https://huggingface.co/collections/FastVideo/fastmetal), a family of 1.3B, 5B, and 14B models optimized for Mac. Follow the [MLX install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mlx/) and read the [Blog](https://haoailab.com/blogs/fastmetal/).
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. See the [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), [Attn-QAT training guide](https://haoailab.com/FastVideo/training/attn_qat/), and [blog](https://haoailab.com/blogs/fastwan-qad/).
- `2026/03/17`: Release demo: Into the Dreamverse: Vibe Directing in FastVideo, check out the [Blog](https://haoailab.com/blogs/dreamverse/).
- `2026/03/13`: Release demo: Create a 5s 1080p Video in 4.5s with FastVideo on a Single GPU, check out the [Blog](https://haoailab.com/blogs/fastvideo_realtime_1080p/).
- `2025/11/19`: Release [CausalWan2.2 I2V A14B Preview](https://huggingface.co/FastVideo/CausalWan2.2-I2V-A14B-Preview-Diffusers) models, [Blog](https://hao-ai-lab.github.io/blogs/fastvideo_causalwan_preview/) and [Inference Code!](https://github.com/hao-ai-lab/FastVideo/blob/main/examples/inference/basic/basic_self_forcing_causal_wan2_2_i2v.py).
@@ -33,7 +39,7 @@ FastVideo has the following features:
- [Sparse distillation](https://hao-ai-lab.github.io/blogs/fastvideo_post_training/) to achieve >50x denoising speedup
- Scalable training with FSDP2, sequence parallelism, and selective activation checkpointing.
- Causal distillation through Self-Forcing
- See this [page](https://hao-ai-lab.github.io/FastVideo/training/overview/) for full list of supported models and recipes.
- See this [page](https://hao-ai-lab.github.io/FastVideo/training/overview/) for the supported training workflows, and the [support matrix](https://hao-ai-lab.github.io/FastVideo/inference/support_matrix/) for supported models.
- State-of-the-art performance optimizations for inference
- Sequence Parallelism for distributed inference
- Multiple state-of-the-art attention backends
@@ -60,7 +66,12 @@ UV_TORCH_BACKEND=cu126 uv pip install fastvideo
```
Use `UV_TORCH_BACKEND=cu130` on CUDA 13. Apple silicon users should follow the
[MPS installation guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
[MLX install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mlx/).
> **On an Apple Silicon Mac?** Install with `uv pip install -e '.[mlx]'` from
> a clone, then pick a recipe in the
> [cookbook](https://haoailab.com/FastVideo/cookbook/). See the
> [MLX install guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mlx/).
Please see our [docs](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/) for more detailed installation instructions.
@@ -78,7 +89,7 @@ Install FastVideo (https://github.com/hao-ai-lab/FastVideo) into a fresh uv virt
https://hao-ai-lab.github.io/FastVideo/getting_started/installation/):
- NVIDIA GPU, x86_64 -> docs/getting_started/installation/gpu.md
- NVIDIA DGX Spark / GB10, aarch64, CUDA 13 -> docs/getting_started/installation/spark.md
- Apple Silicon, macOS -> docs/getting_started/installation/mps.md
- Apple Silicon, macOS -> docs/getting_started/installation/mlx.md
3. Use uv for every step. If a command fails, debug it and tell me what you changed.
4. Verify the result:
python -c "import fastvideo, torch; print('cuda', torch.cuda.is_available())"
@@ -143,6 +154,8 @@ if __name__ == '__main__':
main()
```
`num_gpus=1` runs the worker in-process (weights load once, no extra Python process). On Colab/Kaggle-style machines with ~16GB host RAM, keep `num_gpus=1`; free-tier system memory does not grow with extra T4s, so `num_gpus>1` is likely to OOM.
Run the script with:
```bash
+57 -2
View File
@@ -97,13 +97,33 @@ dreamverse-server --port 8009
dreamverse-mock-server --port 8009
```
### Run Dreamverse with FastH3
Select the VSA data-free FastH3 Preview profile when you start the backend:
```bash
DREAMVERSE_MODEL_ID=fast-h3 dreamverse-server --port 8009
```
The `fast-h3` profile uses four visible GPUs by default. It loads the `MiniMaxAI/MiniMax-H3` base checkpoint and the
`vsa-datafree/adapter_model.safetensors` adapter from
`FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA`. Each request generates a 124-frame, 768×1344 video with
synchronized audio and five sigma-grid points. Dreamverse uses the last frame of each segment as first-frame
conditioning for the following segment.
Set `CUDA_VISIBLE_DEVICES` when you need to choose the four physical GPUs:
```bash
CUDA_VISIBLE_DEVICES=0,1,2,3 DREAMVERSE_MODEL_ID=fast-h3 dreamverse-server --port 8009
```
> **Expect a slow first boot.** With `torch.compile` and startup warmup enabled
> (the default), the backend compiles the segment 1 and segment 2 inference
> paths before it reports ready — this can take **tens of minutes on a cold
> cache**, regardless of how you deploy (local, server, Docker, or Modal).
> `/healthz` responds as soon as the process is up; `/readyz` stays `503` until
> warmup finishes. For a faster, uncompiled startup while testing, set
> `FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend.
> warmup finishes. To defer compilation until the first generated request while
> testing, set `FASTVIDEO_ENABLE_STARTUP_WARMUP=0` before starting the backend.
## Frontend Setup
@@ -138,6 +158,40 @@ dreamverse-server --host 0.0.0.0 --port 8009
The Dreamverse backend defaults to `0.0.0.0:8009` and starts one GPU worker on
the first visible GPU by default.
### Cosmos Predict2.5 DFD continuation (experimental)
Dreamverse can combine two converted Cosmos Predict2.5 2B packages: the
distilled Text2World student creates an unconditioned first segment, then the
Data-Forcing Distillation (DFD) Video2World student conditions each later
segment on the prior terminal frame. Point the runtime at both local converted
packages:
```bash
export DREAMVERSE_MODEL_ID=cosmos25-dfd
export DREAMVERSE_MODEL_PATH=/path/to/Cosmos-Predict2.5-2B-Distilled-TrigFlow-FastVideo
export DREAMVERSE_COSMOS25_DFD_MODEL_PATH=/path/to/Cosmos-Predict2.5-2B-DFD-FastVideo
export ENABLE_TORCH_COMPILE=0
dreamverse-server --host 0.0.0.0 --port 8009
```
The backend loads and warms both model roles before reporting ready. Both use
BF16, Torch SDPA, 704x1280 output, 24 FPS, and four steps. Bootstrap segments
contain 77 frames. DFD segments contain 81 decoded frames, but Dreamverse drops
the repeated conditioning frame before streaming, leaving 80 new frames. An
initial user image selects DFD immediately without treating that first frame as
a cross-segment overlap.
The profile uses a 30-minute session lease because sequential generation on
GB10-class hardware can exceed Dreamverse's five-minute default while the GPU
is still making progress. Deployments can override the lease with
`FASTVIDEO_SESSION_TIMEOUT_SECONDS`.
Cosmos does not produce audio, so the backend supplies duration-matched silent
24 kHz audio for the existing browser streaming contract and trims 1,000 audio
samples with each repeated DFD boundary frame. Runtime LoRA changes are not
supported. Full segments take roughly 145 seconds on GB10, so this profile is a
continuation-quality integration rather than a real-time configuration.
### Check Readiness
In another shell, verify that the backend process is alive:
@@ -219,6 +273,7 @@ selection, and mock-server behavior:
pytest apps/dreamverse/dreamverse/tests/test_config.py \
apps/dreamverse/dreamverse/tests/test_entrypoints.py \
apps/dreamverse/dreamverse/tests/test_gpu_pool.py \
apps/dreamverse/dreamverse/tests/test_minimax_h3_generation.py \
apps/dreamverse/dreamverse/tests/test_mock_server.py -q
```
+12 -1
View File
@@ -139,7 +139,18 @@ session.
- startup warmup
- user join/leave commands
- `USER_STEP` execution for each segment
- continuation state between segments
- generation-command routing and stream-result delivery
Model generation has a separate ownership boundary inside each GPU process:
- `apps/dreamverse/dreamverse/generation_worker.py` selects the backend that the active model profile declares and owns
the backend lifecycle.
- `apps/dreamverse/dreamverse/ltx2_generation.py` owns LTX-2 generator configuration, video and audio continuation, and
runtime LoRA application.
- `apps/dreamverse/dreamverse/minimax_h3_generation.py` owns the VSA data-free FastH3 adapter, FastH3 generator and
request configuration, and last-frame continuation through MiniMax H3 first-frame conditioning.
- `apps/dreamverse/dreamverse/generation_contracts.py` defines the decoded media and stream-trimming result that both
model backends return to `apps/dreamverse/dreamverse/gpu_pool.py`.
`apps/dreamverse/dreamverse/prompt_enhancer.py` manages:
+184
View File
@@ -0,0 +1,184 @@
"""Bounded, runtime-local media library shared by the HTTP and generation APIs."""
from __future__ import annotations
import json
import math
import os
import re
import shutil
import subprocess
import tempfile
import threading
import uuid
from dataclasses import dataclass
from pathlib import Path
from PIL import Image, UnidentifiedImageError
IMAGE_LIMIT = 15 * 1024 * 1024
MEDIA_LIMIT = 100 * 1024 * 1024
STORE_LIMIT = 2 * 1024 * 1024 * 1024
ASSET_LIMIT = 100
MAX_MEDIA_SECONDS = 30
MIME_TYPES = {
"image/png": ("image", ".png"),
"image/jpeg": ("image", ".jpg"),
"image/webp": ("image", ".webp"),
"video/mp4": ("video", ".mp4"),
"video/quicktime": ("video", ".mov"),
"video/webm": ("video", ".webm"),
"audio/mpeg": ("audio", ".mp3"),
"audio/mp4": ("audio", ".m4a"),
"audio/x-m4a": ("audio", ".m4a"),
"audio/wav": ("audio", ".wav"),
"audio/x-wav": ("audio", ".wav"),
"audio/flac": ("audio", ".flac"),
"audio/x-flac": ("audio", ".flac"),
"audio/ogg": ("audio", ".ogg"),
"audio/webm": ("audio", ".webm"),
}
@dataclass(frozen=True)
class StoredAsset:
asset_id: str
kind: str
path: str
name: str
mime_type: str
size: int
def public(self) -> dict:
return {
"asset_id": self.asset_id,
"kind": self.kind,
"name": self.name,
"mime_type": self.mime_type,
"size": self.size,
"url": f"/assets/{self.asset_id}",
}
def validate_media(path: Path, mime_type: str) -> None:
"""Inspect content, not filenames; refuse playlists and non-media uploads."""
kind = MIME_TYPES[mime_type][0]
if kind == "image":
try:
with Image.open(path) as img:
expected = {"image/png": "PNG", "image/jpeg": "JPEG", "image/webp": "WEBP"}[mime_type]
if img.format != expected:
raise ValueError("The image content does not match its file type.")
if img.width * img.height > 16_777_216:
raise ValueError("Images must contain at most 16 megapixels.")
if getattr(img, "is_animated", False):
raise ValueError("Use a still image or upload the animation as a video.")
img.verify()
except (UnidentifiedImageError, OSError, Image.DecompressionBombError) as exc:
raise ValueError("The image could not be decoded. Use PNG, JPEG, or WebP.") from exc
return
probe = shutil.which(os.getenv("FASTVIDEO_FFPROBE_BIN", "ffprobe"))
if not probe:
raise ValueError("This runtime needs ffprobe installed to accept video and audio assets.")
try:
result = subprocess.run(
[
probe, "-v", "error", "-protocol_whitelist", "file,pipe", "-format_whitelist",
"mov,matroska,webm,mp3,wav,flac,ogg", "-show_format", "-show_streams", "-of", "json",
str(path)
],
check=True,
capture_output=True,
timeout=15,
)
info = json.loads(result.stdout)
formats = set(info.get("format", {}).get("format_name", "").split(","))
if not formats.intersection({"mov", "mp4", "matroska", "webm", "mp3", "wav", "flac", "ogg"}):
raise ValueError("Upload a media file, not a playlist or external reference.")
streams = [stream for stream in info.get("streams", []) if stream.get("codec_type") == kind]
if not streams:
raise ValueError(f"The file contains no {kind} stream.")
for stream in info.get("streams", []):
if stream.get("codec_type") == "audio" and int(stream.get("channels", 0)) not in (1, 2):
raise ValueError("H3 references require mono or stereo audio, including video soundtracks.")
duration = float(info.get("format", {}).get("duration", "nan"))
if not math.isfinite(duration) or not 0 < duration <= MAX_MEDIA_SECONDS:
raise ValueError(f"Reference video and audio must be between 0 and {MAX_MEDIA_SECONDS} seconds long.")
for stream in streams:
if kind == "video" and int(stream.get("width", 0)) * int(stream.get("height", 0)) > 8_294_400:
raise ValueError("Reference videos must be 4K or smaller.")
except (subprocess.SubprocessError, json.JSONDecodeError, OSError) as exc:
raise ValueError("The media file could not be decoded. Check its format and try again.") from exc
class AssetStore:
"""Assets live until deletion or runtime exit; pinned generation inputs cannot be deleted."""
def __init__(self) -> None:
self._directory: tempfile.TemporaryDirectory | None = None
self._assets: dict[str, StoredAsset] = {}
self._pins: dict[str, int] = {}
self._lock = threading.RLock()
def staging_path(self, mime_type: str) -> Path:
with self._lock:
if mime_type not in MIME_TYPES:
raise ValueError("Unsupported media type. Use PNG/JPEG/WebP, MP4/WebM/MOV, or WAV/MP3/M4A/FLAC/OGG.")
if len(self._assets) >= ASSET_LIMIT or sum(item.size for item in self._assets.values()) >= STORE_LIMIT:
raise ValueError("The runtime asset library is full. Remove unused assets before uploading more.")
if self._directory is None:
self._directory = tempfile.TemporaryDirectory(prefix="dreamverse-assets-")
return Path(self._directory.name) / f"{uuid.uuid4().hex}{MIME_TYPES[mime_type][1]}"
def add(self, path: Path, name: str, mime_type: str) -> StoredAsset:
validate_media(path, mime_type)
size = path.stat().st_size
if size == 0 or size > (IMAGE_LIMIT if MIME_TYPES[mime_type][0] == "image" else MEDIA_LIMIT):
raise ValueError("The asset is empty or exceeds its upload size limit.")
with self._lock:
if len(self._assets) >= ASSET_LIMIT or size + sum(item.size
for item in self._assets.values()) > STORE_LIMIT:
raise ValueError("The runtime asset library is full. Remove unused assets before uploading more.")
if self._directory is None or path.parent != Path(self._directory.name):
raise ValueError("The asset must be uploaded to this runtime.")
display_name = re.sub(r"[\x00-\x1f\x7f/\\]", "_", name).strip()[:200] or "Untitled asset"
asset = StoredAsset(path.stem, MIME_TYPES[mime_type][0], str(path), display_name, mime_type, size)
self._assets[asset.asset_id] = asset
return asset
def get(self, asset_id: str) -> StoredAsset:
with self._lock:
if not isinstance(asset_id, str) or not re.fullmatch(r"[a-f0-9]{32}", asset_id):
raise ValueError("Invalid asset ID. Upload or select an asset from the library.")
asset = self._assets.get(asset_id)
if asset is None or not Path(asset.path).is_file():
raise ValueError("An asset is no longer available. Upload it again and reselect it.")
return asset
def pin(self, asset_ids: list[str]) -> None:
with self._lock:
for asset_id in asset_ids:
self.get(asset_id)
for asset_id in asset_ids:
self._pins[asset_id] = self._pins.get(asset_id, 0) + 1
def release(self, asset_ids: list[str]) -> None:
with self._lock:
for asset_id in asset_ids:
count = self._pins.get(asset_id, 0)
if count > 1:
self._pins[asset_id] = count - 1
else:
self._pins.pop(asset_id, None)
def delete(self, asset_id: str) -> None:
with self._lock:
asset = self.get(asset_id)
if self._pins.get(asset_id, 0):
raise ValueError("This asset is in use by a generation session. End the session before deleting it.")
Path(asset.path).unlink(missing_ok=True)
del self._assets[asset_id]
asset_store = AssetStore()
@@ -1,6 +1,6 @@
"""Benchmark the LTX-2 generation pipeline driven by the dreamverse Python SDK path.
Mirrors how ``apps/dreamverse/dreamverse/video_generation.py`` constructs
Mirrors how ``apps/dreamverse/dreamverse/ltx2_generation.py`` constructs
``GeneratorConfig`` and calls ``VideoGenerator.generate()``, then
captures per-stage timings via the ``FASTVIDEO_STAGE_LOGGING=1`` log
hooks (same mechanism as ``FastVideo-internal/examples/inference/basic/
@@ -111,7 +111,13 @@ def _build_generator_config(model_path: str, enable_compile: bool, num_gpus: int
mode="max-autotune-no-cudagraphs",
dynamic=False),
use_fsdp_inference=False,
quantization=QuantizationConfig(transformer_quant="NVFP4"),
# The bundled LTX2 model enables a refinement LoRA during the
# first request. NVFP4 otherwise purges the dense weights that
# FastVideo's LoRA merge path requires.
quantization=QuantizationConfig(
transformer_quant="NVFP4",
transformer_retain_original_weights=True,
),
),
pipeline=PipelineSelection(
components=components,
+90 -18
View File
@@ -1,5 +1,6 @@
import os
from pathlib import Path
from typing import cast
_REPO_ROOT = Path(__file__).resolve().parents[1]
_SERVER_ROOT = Path(__file__).resolve().parent
@@ -55,16 +56,65 @@ FRONTEND_STATIC_DIR_CANDIDATES = _resolve_frontend_static_dir_candidates()
MODEL_REGISTRY = {
"fast-ltx2": {
"name": "FastLTX2",
"generation_backend": "ltx2",
"default_sp_size": 1,
"model_path": "FastVideo/LTX2-Distilled-Diffusers",
"config_model_path": "FastVideo/LTX2-Distilled-Diffusers",
"lora_repo": "FastVideo/LTX2-OmniNFT-LoRA",
},
"fast-ltx23": {
"name": "FastLTX23",
"generation_backend": "ltx2",
"default_sp_size": 1,
"model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
"config_model_path": "FastVideo/LTX-2.3-Distilled-Diffusers",
"lora_repo": "FastVideo/LTX-2.3-OmniNFT-LoRA",
},
"fast-h3": {
"name": "FastH3",
"generation_backend": "minimax_h3",
"default_sp_size": 4,
"model_path": "MiniMaxAI/MiniMax-H3",
"adapter_repo": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
"adapter_filename": "vsa-datafree/adapter_model.safetensors",
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
"height": 768,
"width": 1344,
"num_frames": 124,
"num_inference_steps": 5,
"seed": 1000,
},
"full-h3": {
"name": "MiniMax H3 (Full)",
"generation_backend": "minimax_h3",
"default_sp_size": 4,
"model_path": "MiniMaxAI/MiniMax-H3",
"attention_backend": "FLASH_ATTN",
"height": 768,
"width": 1344,
"num_frames": 124,
"num_inference_steps": 50,
"seed": 1000,
"full_checkpoint": True,
},
"cosmos25-dfd": {
"name": "Cosmos Predict2.5 DFD",
"generation_backend": "cosmos25_dfd",
"default_sp_size": 1,
"model_path": "FastVideo/Cosmos-Predict2.5-2B-Distilled-TrigFlow",
"continuation_model_path": "FastVideo/Cosmos-Predict2.5-2B-DFD",
"attention_backend": "TORCH_SDPA",
"height": 704,
"width": 1280,
"bootstrap_num_frames": 77,
"continuation_num_frames": 81,
"fps": 24,
"num_inference_steps": 4,
"seed": 42,
# Six sequential GB10 segments can exceed the legacy five-minute
# DreamVerse lease even though the GPU is making progress.
"session_timeout_seconds": 1800,
},
}
DEFAULT_MODEL_ID = "fast-ltx2"
@@ -76,22 +126,6 @@ if ACTIVE_MODEL_ID not in MODEL_REGISTRY:
# Active model configuration
MODEL_CONFIG = MODEL_REGISTRY[ACTIVE_MODEL_ID]
# Generation limits
SESSION_TIMEOUT_SECONDS = 300
# Frame settings
NUM_FRAMES = 121
FRAME_HEIGHT = 1088
FRAME_WIDTH = 1920
NUM_INFERENCE_STEPS = 5
JPEG_QUALITY = 100
BATCH_SIZE = 3
# Streaming mode:
# - legacy_jpeg: send frame_batch JSON payloads with base64 JPEGs
# - av_fmp4: send muxed fMP4 binary chunks over WebSocket
STREAM_MODE = os.getenv("STREAM_MODE", "av_fmp4").strip().lower()
def _env_int(name: str, default: int) -> int:
value = os.getenv(name)
@@ -168,10 +202,41 @@ def _optional_env(*names: str) -> str | None:
return None
# Generation limits
# Slower backends may own a longer default lease. A profile can set
# ``session_timeout_seconds``; Full H3 loads and generates substantially longer
# than the Preview adapter, which also covers a base/ref pipeline reload inside a
# retained session. An explicit environment override remains available for
# deployment policy: DREAMVERSE_SESSION_TIMEOUT_SECONDS, with
# FASTVIDEO_SESSION_TIMEOUT_SECONDS accepted as an alias.
# Values below 60 seconds are floored so a single segment cannot outlast the session.
_DEFAULT_SESSION_TIMEOUT_SECONDS = cast(
int, MODEL_CONFIG.get("session_timeout_seconds", 7200 if ACTIVE_MODEL_ID == "full-h3" else 300))
SESSION_TIMEOUT_SECONDS = max(
60,
_env_int(
"DREAMVERSE_SESSION_TIMEOUT_SECONDS",
_env_int("FASTVIDEO_SESSION_TIMEOUT_SECONDS", _DEFAULT_SESSION_TIMEOUT_SECONDS),
),
)
# Frame settings
NUM_FRAMES = 121
FRAME_HEIGHT = 1088
FRAME_WIDTH = 1920
NUM_INFERENCE_STEPS = 5
JPEG_QUALITY = 100
BATCH_SIZE = 3
# Streaming mode:
# - legacy_jpeg: send frame_batch JSON payloads with base64 JPEGs
# - av_fmp4: send muxed fMP4 binary chunks over WebSocket
STREAM_MODE = os.getenv("STREAM_MODE", "av_fmp4").strip().lower()
DEVTOOLS_ENABLED = _env_bool("FASTVIDEO_ENABLE_DEVTOOLS", False)
PROMPT_SAFETY_ENABLED = _env_bool("FASTVIDEO_ENABLE_PROMPT_SAFETY", False)
DREAMVERSE_MAX_AUTOTUNE = _env_bool("DREAMVERSE_MAX_AUTOTUNE", True)
DREAMVERSE_SP_SIZE = max(1, _env_int("DREAMVERSE_SP_SIZE", 1))
DREAMVERSE_SP_SIZE = max(1, _env_int("DREAMVERSE_SP_SIZE", cast(int, MODEL_CONFIG["default_sp_size"])))
DREAMVERSE_MODEL_PATH = (os.getenv("DREAMVERSE_MODEL_PATH", "").strip() or None)
if DREAMVERSE_MODEL_PATH:
@@ -181,6 +246,13 @@ if DREAMVERSE_MODEL_PATH:
"config_model_path": DREAMVERSE_MODEL_PATH,
}
DREAMVERSE_COSMOS25_DFD_MODEL_PATH = (os.getenv("DREAMVERSE_COSMOS25_DFD_MODEL_PATH", "").strip() or None)
if DREAMVERSE_COSMOS25_DFD_MODEL_PATH and MODEL_CONFIG.get("generation_backend") == "cosmos25_dfd":
MODEL_CONFIG = {
**MODEL_CONFIG,
"continuation_model_path": DREAMVERSE_COSMOS25_DFD_MODEL_PATH,
}
AVAILABLE_LORAS = {
"pixar": {
"repo": "vrgamedevgirl84/LTX_2.3_Pixar_Toon_Style_LoRa",
@@ -213,7 +285,7 @@ def _resolve_lora_spec(spec: str) -> str | None:
if not spec:
return None
if spec.lower() == "omninft":
return MODEL_CONFIG.get("lora_repo")
return cast(str | None, MODEL_CONFIG.get("lora_repo"))
if spec.lower() in AVAILABLE_LORAS:
return AVAILABLE_LORAS[spec.lower()]["repo"]
return spec
@@ -0,0 +1,252 @@
"""Cosmos Predict2.5 distilled bootstrap and DFD continuation for DreamVerse."""
from __future__ import annotations
import gc
import os
import time
from typing import TYPE_CHECKING, Any
import numpy as np
import torch
from dreamverse.generation_contracts import StepResult
from dreamverse.generation_inputs import GenerationInputs
if TYPE_CHECKING:
from PIL.Image import Image
_SILENT_AUDIO_SAMPLE_RATE = 24_000
def _required_config_str(model_config: dict, field_name: str) -> str:
value = model_config.get(field_name)
if not isinstance(value, str) or not value.strip():
raise ValueError(f"Cosmos Predict2.5 DFD model configuration requires `{field_name}`.")
return value.strip()
class Cosmos25DFDGenerationBackend:
"""Own complementary Cosmos T2W and one-frame-conditioned DFD generators."""
def __init__(self, gpu_id: int):
self.gpu_id = gpu_id
self.bootstrap_generator: Any | None = None
self.continuation_generator: Any | None = None
self.model_config: dict = {}
self.continuation_image: Image | None = None
def _gpu_mem(self) -> str:
allocated_gib = torch.cuda.memory_allocated() / 1024**3
reserved_gib = torch.cuda.memory_reserved() / 1024**3
return f"alloc={allocated_gib:.2f}GiB, reserved={reserved_gib:.2f}GiB"
@staticmethod
def _configure_environment(attention_backend: str) -> None:
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = attention_backend
os.environ.pop("FASTVIDEO_INFERENCE_TORCH_COMPILE", None)
@staticmethod
def _load_generator(model_path: str):
from fastvideo import VideoGenerator
return VideoGenerator.from_pretrained(
model_path,
num_gpus=1,
use_fsdp_inference=False,
dit_cpu_offload=False,
vae_cpu_offload=False,
text_encoder_cpu_offload=True,
pin_cpu_memory=True,
enable_torch_compile=False,
)
def initialize(self, model_config: dict | None = None) -> None:
"""Load both package roles so bootstrap and continuation are ready."""
if model_config is not None:
self.model_config = dict(model_config)
if not self.model_config:
raise ValueError("Cosmos Predict2.5 DFD initialization requires a model configuration.")
self.shutdown()
bootstrap_path = _required_config_str(self.model_config, "model_path")
continuation_path = _required_config_str(self.model_config, "continuation_model_path")
attention_backend = _required_config_str(self.model_config, "attention_backend")
self._configure_environment(attention_backend)
print(f"[GPU {self.gpu_id}] Loading Cosmos T2W bootstrap: {bootstrap_path}")
print(f"[GPU {self.gpu_id}] Before bootstrap load: {self._gpu_mem()}")
self.bootstrap_generator = self._load_generator(bootstrap_path)
print(f"[GPU {self.gpu_id}] Loading Cosmos DFD continuation: {continuation_path}")
self.continuation_generator = self._load_generator(continuation_path)
print(f"[GPU {self.gpu_id}] Cosmos T2W + DFD loaded: {self._gpu_mem()} (warmup pending)")
def shutdown(self) -> None:
"""Release both FastVideo generators and the retained terminal frame."""
self.clear_conditioning()
for attr_name in ("bootstrap_generator", "continuation_generator"):
generator = getattr(self, attr_name)
if generator is not None:
try:
generator.shutdown()
except Exception as exc:
print(f"[GPU {self.gpu_id}] Cosmos generator shutdown warning: {exc}")
setattr(self, attr_name, None)
gc.collect()
if torch.cuda.is_available():
torch.cuda.empty_cache()
def clear_conditioning(self) -> None:
if self.continuation_image is not None:
self.continuation_image.close()
self.continuation_image = None
@staticmethod
def _load_rgb_image(image_path: str) -> Image:
from PIL import Image
with Image.open(image_path) as image:
return image.convert("RGB").copy()
def _select_conditioning_image(
self,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
) -> tuple[Image | None, bool]:
if reset_conditioning:
self.clear_conditioning()
if segment_idx > 1 and self.continuation_image is not None:
return self.continuation_image.copy(), True
if segment_idx > 1 and not reset_conditioning:
raise RuntimeError(f"Cosmos DFD segment {segment_idx} requires a retained continuation frame.")
if segment_idx == 1 and image_path:
return self._load_rgb_image(image_path), False
return None, False
def _sampling_param(self, *, conditioned: bool):
# ``num_cond_frames`` is not yet exposed by the typed SamplingConfig,
# so this backend uses the compatibility request until that field lands.
from fastvideo.api.sampling_param import SamplingParam
num_frames_key = "continuation_num_frames" if conditioned else "bootstrap_num_frames"
return SamplingParam(
negative_prompt="",
save_video=False,
return_frames=True,
height=int(self.model_config["height"]),
width=int(self.model_config["width"]),
num_frames=int(self.model_config[num_frames_key]),
fps=int(self.model_config["fps"]),
num_inference_steps=int(self.model_config["num_inference_steps"]),
guidance_scale=1.0,
seed=int(self.model_config["seed"]),
num_cond_frames=1 if conditioned else 0,
)
def _save_continuation_frame(self, frame: object) -> None:
from PIL import Image
self.clear_conditioning()
if isinstance(frame, Image.Image):
self.continuation_image = frame.convert("RGB").copy()
return
pixels = np.asarray(frame)
self.continuation_image = Image.fromarray(np.ascontiguousarray(pixels)).convert("RGB")
@staticmethod
def _silent_audio(frame_count: int, fps: int) -> torch.Tensor:
sample_count = max(1, int(round((frame_count / float(fps)) * _SILENT_AUDIO_SAMPLE_RATE)))
return torch.zeros(sample_count, dtype=torch.float32)
def generate_step(
self,
prompt: str,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
generation_inputs: GenerationInputs | None = None,
) -> StepResult:
"""Generate a T2W start or DFD continuation and retain its last frame."""
if generation_inputs is not None and (generation_inputs.mode not in (None, "t2va") or generation_inputs.assets):
raise ValueError("Cosmos supports text generation only through the generation mode API.")
if self.bootstrap_generator is None or self.continuation_generator is None:
raise RuntimeError("Cosmos T2W + DFD generators are not initialized.")
conditioning_image, uses_continuation = self._select_conditioning_image(
segment_idx,
image_path,
reset_conditioning,
)
conditioned = conditioning_image is not None
generator = self.continuation_generator if conditioned else self.bootstrap_generator
sampling_param = self._sampling_param(conditioned=conditioned)
started = time.perf_counter()
try:
if conditioned:
sampling_param.pil_image = conditioning_image
result = generator.generate_video(prompt, sampling_param=sampling_param)
finally:
if conditioning_image is not None:
conditioning_image.close()
torch.cuda.synchronize()
generation_ms = (time.perf_counter() - started) * 1000.0
if not isinstance(result, dict):
raise RuntimeError("Cosmos generation did not return one result dictionary.")
frames = result.get("frames")
expected_frames = int(sampling_param.num_frames)
if not isinstance(frames, list) or len(frames) != expected_frames:
actual_frames = len(frames) if isinstance(frames, list) else None
raise RuntimeError(f"Cosmos generation returned {actual_frames} frames; expected {expected_frames}.")
save_started = time.perf_counter()
self._save_continuation_frame(frames[-1])
save_conditioning_ms = (time.perf_counter() - save_started) * 1000.0
fps = int(sampling_param.fps)
timings = {
"generation_ms": generation_ms,
"generation_time_ms": float(result.get("generation_time") or 0.0) * 1000.0,
"save_conditioning_ms": save_conditioning_ms,
"e2e_latency_ms": (time.perf_counter() - started) * 1000.0,
}
trim_frames = 1 if uses_continuation else 0
mode = "DFD continuation" if conditioned else "T2W bootstrap"
print(f"[GPU {self.gpu_id}] Cosmos {mode} segment {segment_idx}: "
f"{len(frames)} frames, gen={generation_ms:.0f}ms, "
f"save_conditioning={save_conditioning_ms:.0f}ms, "
f"e2e={timings['e2e_latency_ms']:.0f}ms")
return StepResult(
frames=frames,
audio=self._silent_audio(len(frames), fps),
audio_sample_rate=_SILENT_AUDIO_SAMPLE_RATE,
timings=timings,
head_trim_frames=trim_frames,
head_trim_audio_frames=trim_frames,
)
def warmup(self, prompt: str) -> dict[str, float]:
"""Exercise both T2W bootstrap and retained-frame DFD request shapes."""
warmup_prompt = (prompt or "").strip()
if not warmup_prompt:
raise RuntimeError("Startup warmup prompt must be non-empty.")
print(f"[GPU {self.gpu_id}] Cosmos startup warmup starting "
"(synthetic segments: T2W bootstrap, DFD continuation)")
started = time.perf_counter()
bootstrap_result = self.generate_step(warmup_prompt, 1, None, True)
continuation_result = self.generate_step(warmup_prompt, 2, None, False)
total_ms = (time.perf_counter() - started) * 1000.0
self.clear_conditioning()
bootstrap_ms = float(bootstrap_result.timings.get("e2e_latency_ms", 0.0))
continuation_ms = float(continuation_result.timings.get("e2e_latency_ms", 0.0))
print(f"[GPU {self.gpu_id}] Cosmos startup warmup complete: "
f"bootstrap={bootstrap_ms:.0f}ms, continuation={continuation_ms:.0f}ms, total={total_ms:.0f}ms")
return {
"warmup_bootstrap_ms": bootstrap_ms,
"warmup_continuation_ms": continuation_ms,
"warmup_total_ms": total_ms,
}
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
del stack
raise RuntimeError("Cosmos Predict2.5 DFD does not support DreamVerse runtime LoRA changes.")
@@ -0,0 +1,49 @@
"""Shared contract between DreamVerse generation backends and GPU workers."""
from __future__ import annotations
from dataclasses import dataclass
from typing import Any, Protocol
from dreamverse.generation_inputs import GenerationInputs
@dataclass
class StepResult:
"""Decoded media and stream-trimming metadata for one DreamVerse segment."""
frames: list
audio: Any
audio_sample_rate: int | None
timings: dict[str, float]
head_trim_frames: int
head_trim_audio_frames: int
class GenerationBackend(Protocol):
"""Model-owned generation operations used by one GPU worker process."""
def initialize(self, model_config: dict | None = None) -> None:
...
def shutdown(self) -> None:
...
def clear_conditioning(self) -> None:
...
def generate_step(
self,
prompt: str,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
generation_inputs: GenerationInputs | None = None,
) -> StepResult:
...
def warmup(self, prompt: str) -> dict[str, float]:
...
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
...
@@ -0,0 +1,102 @@
"""GPU-independent validation for generation modes and ordered asset handles."""
from __future__ import annotations
from dataclasses import dataclass
from PIL import Image, UnidentifiedImageError
from dreamverse.assets import asset_store
GENERATION_MODES = ("t2va", "fl2va", "ref2va")
@dataclass(frozen=True)
class GenerationAsset:
asset_id: str
kind: str
path: str
role: str
@dataclass(frozen=True)
class GenerationInputs:
mode: str | None = None
assets: tuple[GenerationAsset, ...] = ()
@property
def first_frame_path(self) -> str | None:
return next((asset.path for asset in self.assets if asset.role == "first_frame"), None)
@property
def last_frame_path(self) -> str | None:
return next((asset.path for asset in self.assets if asset.role == "last_frame"), None)
@property
def references(self) -> tuple[GenerationAsset, ...]:
return tuple(asset for asset in self.assets if asset.role == "reference")
def supported_generation_modes(model_id: str) -> tuple[str, ...]:
return GENERATION_MODES if model_id in ("full-h3", "mock") else ("t2va", )
def resolve_generation_inputs(payload: dict, model_id: str) -> GenerationInputs:
mode = payload.get("generation_mode")
raw_assets = payload.get("conditioning_assets", [])
if mode is None and "generation_mode" not in payload:
if raw_assets:
raise ValueError("Select a generation mode before attaching conditioning assets.")
return GenerationInputs()
if not isinstance(mode, str) or mode not in GENERATION_MODES:
raise ValueError("Unknown generation mode. Choose T2VA, FL2VA, or Ref2VA.")
if mode not in supported_generation_modes(model_id):
raise ValueError(f"{mode.upper()} requires the Full H3 runtime. This runtime is running {model_id}.")
if payload.get("initial_image") is not None:
raise ValueError("Use asset IDs for generation modes; do not combine them with the legacy initial_image field.")
if not isinstance(raw_assets, list) or len(raw_assets) > 12:
raise ValueError("conditioning_assets must be an ordered list with at most 12 assets.")
if mode == "t2va" and raw_assets:
raise ValueError("T2VA accepts text only. Remove conditioning assets or choose another mode.")
assets: list[GenerationAsset] = []
for item in raw_assets:
if not isinstance(item, dict) or set(item) != {"asset_id", "role"}:
raise ValueError("Each conditioning asset must contain only asset_id and role.")
role = item["role"]
if role not in ("first_frame", "last_frame", "reference"):
raise ValueError("Asset role must be first_frame, last_frame, or reference.")
stored = asset_store.get(item["asset_id"])
assets.append(GenerationAsset(stored.asset_id, stored.kind, stored.path, role))
if mode == "fl2va":
if any(asset.kind != "image" or asset.role == "reference" for asset in assets):
raise ValueError("FL2VA accepts only first-frame and last-frame images.")
if sum(asset.role == "first_frame" for asset in assets) != 1:
raise ValueError("FL2VA requires exactly one first-frame image.")
if sum(asset.role == "last_frame" for asset in assets) > 1:
raise ValueError("FL2VA accepts at most one last-frame image.")
elif mode == "ref2va":
if not assets or any(asset.role != "reference" for asset in assets):
raise ValueError("Ref2VA requires an ordered list of reference assets, without keyframe roles.")
if not any(asset.kind in ("image", "video") for asset in assets):
raise ValueError("Ref2VA requires at least one image or video; audio alone is not supported.")
for kind, limit in (("image", 9), ("video", 3), ("audio", 3)):
if sum(asset.kind == kind for asset in assets) > limit:
raise ValueError(f"Ref2VA accepts at most {limit} {kind} references.")
for asset in assets:
if asset.kind == "image":
try:
with Image.open(asset.path) as image:
if image.width > 4 * image.height or image.height > 4 * image.width:
raise ValueError(
"Ref2VA image aspect ratios must be between 1:4 and 4:1. Crop this image first.")
except (UnidentifiedImageError, OSError, Image.DecompressionBombError) as exc:
raise ValueError("A selected reference image could not be decoded. Upload it again.") from exc
return GenerationInputs(mode, tuple(assets))
def pin_generation_inputs(inputs: GenerationInputs) -> None:
asset_store.pin([asset.asset_id for asset in inputs.assets])
def release_generation_inputs(inputs: GenerationInputs) -> None:
asset_store.release([asset.asset_id for asset in inputs.assets])
@@ -0,0 +1,103 @@
"""Select and own one model-specific generation backend per GPU process."""
from __future__ import annotations
from dreamverse.config import MODEL_CONFIG
from dreamverse.generation_contracts import GenerationBackend, StepResult
from dreamverse.generation_inputs import GenerationInputs
def _create_generation_backend(backend_name: str, gpu_id: int) -> GenerationBackend:
"""Construct the backend that owns the selected model family's behavior."""
if backend_name == "ltx2":
from dreamverse.ltx2_generation import LTX2GenerationBackend
return LTX2GenerationBackend(gpu_id)
if backend_name == "minimax_h3":
from dreamverse.minimax_h3_generation import MiniMaxH3GenerationBackend
return MiniMaxH3GenerationBackend(gpu_id)
if backend_name == "cosmos25_dfd":
from dreamverse.cosmos25_dfd_generation import Cosmos25DFDGenerationBackend
return Cosmos25DFDGenerationBackend(gpu_id)
raise ValueError(f"Unsupported DreamVerse generation backend: {backend_name!r}")
class VideoGenerationWorker:
"""Delegate GPU lifecycle and generation calls to the active model backend."""
def __init__(self, gpu_id: int):
self.gpu_id = gpu_id
self.model_config: dict = dict(MODEL_CONFIG)
self.backend_name: str | None = None
self.backend: GenerationBackend | None = None
def initialize(self, model_config: dict | None = None) -> None:
"""Load the requested model through its generation backend.
Model selection belongs here so the GPU process and streaming layers
use one stable media contract without importing model-specific code.
"""
requested_model_config = dict(model_config) if model_config is not None else dict(self.model_config)
backend_name = requested_model_config.get("generation_backend")
if not isinstance(backend_name, str) or not backend_name:
raise ValueError("DreamVerse model configuration requires `generation_backend`.")
candidate_backend = self.backend
if candidate_backend is None or self.backend_name != backend_name:
if candidate_backend is not None:
candidate_backend.shutdown()
candidate_backend = _create_generation_backend(backend_name, self.gpu_id)
try:
candidate_backend.initialize(requested_model_config)
except Exception:
try:
candidate_backend.shutdown()
except Exception as shutdown_error:
print(f"[GPU {self.gpu_id}] Backend cleanup after initialization failure: {shutdown_error}")
self.backend = None
self.backend_name = None
raise
self.model_config = requested_model_config
self.backend = candidate_backend
self.backend_name = backend_name
def _require_backend(self) -> GenerationBackend:
"""Return the initialized backend or fail before processing a command."""
if self.backend is None:
raise RuntimeError("Generation backend is not initialized.")
return self.backend
def shutdown(self) -> None:
"""Release model resources owned by the selected backend."""
if self.backend is not None:
self.backend.shutdown()
def clear_conditioning(self) -> None:
self._require_backend().clear_conditioning()
def generate_step(
self,
prompt: str,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
generation_inputs: GenerationInputs | None = None,
) -> StepResult:
"""Generate one segment through the selected model backend."""
return self._require_backend().generate_step(
prompt,
segment_idx,
image_path,
reset_conditioning,
generation_inputs=generation_inputs,
)
def warmup(self, prompt: str) -> dict[str, float]:
return self._require_backend().warmup(prompt)
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
return self._require_backend().apply_lora_stack(stack)
+46 -11
View File
@@ -12,7 +12,7 @@ from enum import Enum
from multiprocessing import Process, Queue
from dreamverse.config import (
DEFAULT_MODEL_ID,
ACTIVE_MODEL_ID,
DREAMVERSE_SP_SIZE,
MODEL_REGISTRY,
STARTUP_WARMUP_ENABLED,
@@ -29,6 +29,7 @@ from dreamverse.av_streaming import (
generate_stream_id,
stream_fmp4,
)
from dreamverse.generation_inputs import GenerationInputs, pin_generation_inputs, release_generation_inputs
from dreamverse.worker_ipc import (
CommandPayload,
InitAck,
@@ -54,7 +55,7 @@ from dreamverse.worker_ipc import (
def _parse_requested_gpu_limit() -> int | None:
raw_value = os.getenv("FASTVIDEO_GPU_COUNT", "").strip().lower()
if not raw_value:
return 1
return DREAMVERSE_SP_SIZE
if raw_value == "all":
return None
try:
@@ -164,12 +165,12 @@ def gpu_worker_process(
os.environ["CUDA_VISIBLE_DEVICES"] = cuda_device
os.environ["FASTVIDEO_ATTENTION_BACKEND"] = "FLASH_ATTN"
from dreamverse.video_generation import VideoGenerationWorker
from dreamverse.generation_worker import VideoGenerationWorker
worker = VideoGenerationWorker(gpu_id)
def event_loop(first_cmd: Command = None):
"""Blocking event loop for LTX2; dispatches user commands."""
"""Block on generation commands after the model is initialized."""
print(f"[GPU {gpu_id}] Entering event loop")
def handle_command(cmd: Command):
@@ -189,6 +190,7 @@ def gpu_worker_process(
segment_idx,
image_path=payload.image_path,
reset_conditioning=payload.reset_conditioning,
generation_inputs=payload.generation_inputs,
)
head_trim_frames = step_result.head_trim_frames
head_trim_audio_frames = step_result.head_trim_audio_frames
@@ -432,10 +434,11 @@ class GPUSlot:
self.connected_users: set[str] = set()
self._pending_futures: dict[str, asyncio.Future] = {}
self._stream_queues: dict[str, asyncio.Queue] = {}
self._step_asset_inputs: dict[str, GenerationInputs] = {}
self._response_reader_task: asyncio.Task | None = None
self._active: bool = False
self._reader_lock: asyncio.Lock | None = None
self.current_model_id: str = DEFAULT_MODEL_ID
self.current_model_id: str | None = ACTIVE_MODEL_ID
self.shared_stream_buffer = None
self.shared_stream_buffer_size = SHARED_STREAM_BUFFER_BYTES
@@ -663,6 +666,9 @@ class GPUSlot:
if isinstance(event, (StepComplete, WarmupComplete)):
event.timings["ipc_get_done_ns"] = time.time_ns()
if isinstance(event, (StepComplete, WorkerError)) and event.user_id is not None:
self._release_step_assets(event.user_id)
user_id = event.user_id
if user_id and user_id in self._pending_futures:
future = self._pending_futures.pop(user_id)
@@ -690,7 +696,7 @@ class GPUSlot:
async def join_user(self, user_id: str, model_id: str = None) -> JoinAck:
"""Add a user to this GPU."""
if model_id is None:
model_id = DEFAULT_MODEL_ID
model_id = ACTIVE_MODEL_ID
# Reload model if a different one is requested
if model_id != self.current_model_id and model_id in MODEL_REGISTRY:
@@ -705,16 +711,23 @@ class GPUSlot:
self.connected_users.clear()
model_config = MODEL_REGISTRY[model_id]
reload_response = await self._send_command(Command(CommandType.RELOAD_MODEL,
payload=ReloadModelPayload(model_config=model_config),
user_id="__reload__"),
timeout=600.0)
try:
reload_response = await self._send_command(Command(
CommandType.RELOAD_MODEL,
payload=ReloadModelPayload(model_config=model_config),
user_id="__reload__"),
timeout=600.0)
except Exception:
self.current_model_id = None
raise
match reload_response:
case ReloadAck():
pass
case WorkerError(message=msg):
self.current_model_id = None
raise RuntimeError(f"Model reload failed: {msg}")
case _:
self.current_model_id = None
raise RuntimeError(f"Unexpected reload response: "
f"{type(reload_response).__name__}")
@@ -746,6 +759,7 @@ class GPUSlot:
segment_idx: int = 1,
image_path: str | None = None,
reset_conditioning: bool = False,
generation_inputs: GenerationInputs | None = None,
) -> dict[str, float]:
"""Execute a generation step for a specific user.
@@ -759,9 +773,19 @@ class GPUSlot:
segment_idx=segment_idx,
image_path=image_path,
reset_conditioning=bool(reset_conditioning),
generation_inputs=generation_inputs,
)
if generation_inputs is not None:
if user_id in self._step_asset_inputs:
raise RuntimeError("The previous generation is still using this project's assets.")
pin_generation_inputs(generation_inputs)
self._step_asset_inputs[user_id] = generation_inputs
# Pins intentionally survive a waiter timeout/cancellation: the GPU
# command keeps running. The response reader releases them when the
# worker actually completes (even if that response is now unmatched).
response = await self._send_command_tagged(Command(CommandType.USER_STEP, payload=payload, user_id=user_id),
timeout=1800.0)
self._release_step_assets(user_id)
match response:
case StepComplete(timings=timings):
return timings
@@ -771,6 +795,11 @@ class GPUSlot:
raise RuntimeError(f"Unexpected step response for {user_id[:8]}: "
f"{type(response).__name__}")
def _release_step_assets(self, user_id: str) -> None:
inputs = self._step_asset_inputs.pop(user_id, None)
if inputs is not None:
release_generation_inputs(inputs)
async def apply_lora_stack(
self,
stack: list[tuple[str, float]],
@@ -793,7 +822,9 @@ class GPUSlot:
async def leave_user(self, user_id: str) -> None:
"""Remove a user from this GPU."""
try:
await self._send_command_tagged(Command(CommandType.USER_LEAVE, user_id=user_id), timeout=30.0)
response = await self._send_command_tagged(Command(CommandType.USER_LEAVE, user_id=user_id), timeout=30.0)
if isinstance(response, LeaveAck):
self._release_step_assets(user_id)
except Exception as e:
print(f"[GPU {self.gpu_id}] Leave user error: {e}")
finally:
@@ -830,6 +861,10 @@ class GPUSlot:
except Exception:
pass
if self.process is None or not self.process.is_alive():
for user_id in list(self._step_asset_inputs):
self._release_step_assets(user_id)
for q in (self.command_queue, self.response_queue):
if q is not None:
try:
@@ -1,9 +1,9 @@
"""LTX2 model lifecycle and continuation conditioning.
"""LTX-2 model lifecycle and continuation conditioning.
Runs inside a GPU worker subprocess. Owns the model, the audio
encoder, and the per-session continuation state carried across
segments. Callers must set ``os.environ["CUDA_VISIBLE_DEVICES"]``
before constructing ``VideoGenerationWorker`` — all ``fastvideo.*``
before constructing ``LTX2GenerationBackend`` — all ``fastvideo.*``
imports are deferred to method bodies so nothing touches CUDA at
module import time.
"""
@@ -14,9 +14,6 @@ import gc
import os
import re
import time
from dataclasses import dataclass
from typing import Any
import numpy as np
import torch
@@ -35,6 +32,8 @@ from dreamverse.config import (
DREAMVERSE_LORA_STACK,
_resolve_lora_spec,
)
from dreamverse.generation_contracts import StepResult
from dreamverse.generation_inputs import GenerationInputs
# Multi-frame decoded continuation defaults from
# examples/inference/basic/basic_ltx2_distilled_video_continuation.py.
@@ -80,22 +79,6 @@ def _reset_lora_registry(worker) -> dict:
return {"status": "lora_registry_reset"}
@dataclass
class StepResult:
"""Output of one generation step.
``head_trim_frames`` / ``head_trim_audio_frames`` are derived here
so downstream AV streaming never needs to import conditioning
constants.
"""
frames: list
audio: Any
audio_sample_rate: int | None
timings: dict
head_trim_frames: int
head_trim_audio_frames: int
class ContinuationState:
"""Per-session video + audio conditioning carried across segments."""
@@ -202,7 +185,7 @@ class ContinuationState:
self.audio_latents = latents.detach().clone().cpu()
class VideoGenerationWorker:
class LTX2GenerationBackend:
"""Single-GPU LTX2 generator with continuation state.
Caller must set ``os.environ["CUDA_VISIBLE_DEVICES"]`` before
@@ -308,7 +291,13 @@ class VideoGenerationWorker:
dynamic=False,
),
use_fsdp_inference=False,
quantization=QuantizationConfig(transformer_quant="NVFP4"),
# The bundled LTX2 model enables a refinement LoRA during the
# first request. NVFP4 otherwise purges the dense weights that
# FastVideo's LoRA merge path requires.
quantization=QuantizationConfig(
transformer_quant="NVFP4",
transformer_retain_original_weights=True,
),
),
pipeline=PipelineSelection(
components=components,
@@ -472,8 +461,11 @@ class VideoGenerationWorker:
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
generation_inputs: GenerationInputs | None = None,
) -> StepResult:
"""Execute one generation step; snapshot state for the next segment."""
if generation_inputs is not None and (generation_inputs.mode not in (None, "t2va") or generation_inputs.assets):
raise ValueError("LTX supports text generation only through the generation mode API.")
timings: dict = {}
prompt = self._inject_style_trigger(prompt)
+9
View File
@@ -15,6 +15,7 @@ from dreamverse.gpu_pool import GPUPool, get_available_gpus
from dreamverse.session_logger import SessionEventLogger
from dreamverse.config import (
ACTIVE_MODEL_ID,
AVAILABLE_LORAS,
DEVTOOLS_ENABLED,
FRONTEND_STATIC_DIR_CANDIDATES,
@@ -34,6 +35,8 @@ from dreamverse.routes.presets import (
curated_presets_router,
)
from dreamverse.session.controller import SessionController
from dreamverse.generation_inputs import supported_generation_modes
from dreamverse.routes.assets import router as asset_router
class _HeartbeatAccessLogFilter(logging.Filter):
@@ -92,10 +95,16 @@ app.add_middleware(
app.include_router(build_health_router(lambda: runtime.gpu_pool))
app.include_router(internal_monitor_router)
app.include_router(prompt_config_router)
app.include_router(asset_router)
if DEVTOOLS_ENABLED:
app.include_router(curated_presets_router)
@app.get("/generation-capabilities")
async def generation_capabilities() -> dict:
return {"model_id": ACTIVE_MODEL_ID, "modes": supported_generation_modes(ACTIVE_MODEL_ID), "mock": False}
@app.websocket("/ws")
async def websocket_endpoint(websocket: WebSocket):
controller = SessionController(
@@ -0,0 +1,363 @@
"""Full/Preview H3 lifecycle, conditioning and per-project pipeline selection."""
from __future__ import annotations
import gc
import os
import time
from typing import TYPE_CHECKING, Any
import numpy as np
import torch
from dreamverse.config import DREAMVERSE_SP_SIZE
from dreamverse.generation_contracts import StepResult
from dreamverse.generation_inputs import GenerationInputs
if TYPE_CHECKING:
from PIL.Image import Image
def _required_config_str(model_config: dict, field_name: str) -> str:
"""Read one required non-empty string from a DreamVerse model profile."""
value = model_config.get(field_name)
if not isinstance(value, str) or not value.strip():
raise ValueError(f"FastH3 model configuration requires `{field_name}`.")
return value.strip()
class MiniMaxH3GenerationBackend:
"""Own one H3 pipeline at a time and retain base-pipeline continuation."""
def __init__(self, gpu_id: int):
self.gpu_id = gpu_id
self.generator: Any | None = None
self.model_config: dict = {}
self.continuation_image: Image | None = None
self.pipeline_mode = "base"
def _gpu_mem(self) -> str:
allocated_gib = torch.cuda.memory_allocated() / 1024**3
reserved_gib = torch.cuda.memory_reserved() / 1024**3
return f"alloc={allocated_gib:.2f}GiB, reserved={reserved_gib:.2f}GiB"
@staticmethod
def _configure_environment(attention_backend: str) -> None:
"""Apply the fixed boot-time switches from the FastH3 reference recipe."""
os.environ.update({
"FASTVIDEO_ATTENTION_BACKEND": attention_backend,
"FASTVIDEO_FA4": "1",
"FASTVIDEO_MINIMAX_H3_FUSIONS": "all",
"FASTVIDEO_VSA_SM100A": "0",
})
os.environ.pop("FASTVIDEO_INFERENCE_TORCH_COMPILE", None)
def initialize(self, model_config: dict | None = None) -> None:
"""Load the profile's base pipeline; Ref2VA is loaded on first use."""
if model_config is not None:
self.model_config = dict(model_config)
if not self.model_config:
raise ValueError("FastH3 initialization requires a model configuration.")
self._load_pipeline("base")
def _load_pipeline(self, pipeline_mode: str) -> None:
"""Unload the old executor before loading a base or reference transformer.
GPU worker commands are serialized, so a project boundary never swaps
weights while another request is using them. Keeping one executor also
avoids simultaneously retaining two large H3 transformers in VRAM. A
failed load leaves no executor behind so the next step retries it.
"""
full_checkpoint = bool(self.model_config.get("full_checkpoint", False))
if pipeline_mode == "ref2va" and not full_checkpoint:
raise ValueError("Ref2VA requires the full-h3 model profile.")
if self.generator is not None:
previous_generator = self.generator
self.generator = None
previous_generator.shutdown()
del previous_generator
gc.collect()
torch.cuda.empty_cache()
self.clear_conditioning()
model_path = _required_config_str(self.model_config, "model_path")
attention_backend = _required_config_str(self.model_config, "attention_backend")
self._configure_environment(attention_backend)
from fastvideo import VideoGenerator
from fastvideo.api import (
CompileConfig,
ComponentConfig,
EngineConfig,
GeneratorConfig,
OffloadConfig,
ParallelismConfig,
PipelineSelection,
)
components = ComponentConfig()
if not full_checkpoint:
from huggingface_hub import hf_hub_download
adapter_repo = _required_config_str(self.model_config, "adapter_repo")
adapter_filename = _required_config_str(self.model_config, "adapter_filename")
components.lora_path = hf_hub_download(repo_id=adapter_repo, filename=adapter_filename)
components.lora_strength = 1.0
print(f"[GPU {self.gpu_id}] FastH3 adapter: {adapter_repo}/{adapter_filename}")
if pipeline_mode == "ref2va":
components.override_pipeline_cls_name = "MiniMaxH3Ref2VAModularPipeline"
experimental = {
"attention_backend": attention_backend,
"inference_torch_compile": not full_checkpoint and attention_backend == "FLASH_ATTN",
"vae_parallel_decode": True,
"vae_parallel_decode_strategy": "gather",
}
if attention_backend == "VIDEO_SPARSE_ATTN_H3":
experimental.update({
"VSA_sparsity": 0.9,
"VSA_tile_size": 64,
})
generator_config = GeneratorConfig(
model_path=model_path,
pipeline=PipelineSelection(
workload_type="i2v" if pipeline_mode == "ref2va" else None,
components=components,
experimental=experimental,
),
engine=EngineConfig(
num_gpus=DREAMVERSE_SP_SIZE,
parallelism=ParallelismConfig(tp_size=1, sp_size=DREAMVERSE_SP_SIZE),
offload=OffloadConfig(
dit=False,
dit_layerwise=False,
text_encoder=True,
image_encoder=True,
vae=True,
pin_cpu_memory=not full_checkpoint,
),
compile=CompileConfig(enabled=False, vae_enabled=True),
use_fsdp_inference=full_checkpoint and DREAMVERSE_SP_SIZE > 1,
),
)
print(f"[GPU {self.gpu_id}] Loading H3 model: {model_path} ({pipeline_mode})")
print(f"[GPU {self.gpu_id}] Before model load: {self._gpu_mem()}")
try:
self.generator = VideoGenerator.from_config(generator_config)
except Exception:
# The old executor is already gone; leaving no executor behind lets
# the next step retry this load instead of stranding the GPU slot.
self.generator = None
raise
self.pipeline_mode = pipeline_mode
print(f"[GPU {self.gpu_id}] FastH3 loaded: {self._gpu_mem()} (warmup pending)")
def shutdown(self) -> None:
"""Release the FastVideo generator and cached continuation image."""
self.clear_conditioning()
if self.generator is not None:
self.generator.shutdown()
self.generator = None
def clear_conditioning(self) -> None:
"""Release the first-frame image retained for the next segment."""
if self.continuation_image is not None:
self.continuation_image.close()
self.continuation_image = None
@staticmethod
def _load_rgb_image(image_path: str) -> Image:
"""Load an image into an independent RGB buffer with no open file handle."""
from PIL import Image
with Image.open(image_path) as image:
return image.convert("RGB").copy()
def _select_conditioning_image(
self,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
) -> tuple[Image | None, bool]:
"""Select the initial upload or retained last frame for one segment."""
if reset_conditioning:
self.clear_conditioning()
if segment_idx > 1 and self.continuation_image is not None:
return self.continuation_image.copy(), True
if segment_idx > 1 and not reset_conditioning:
raise RuntimeError(f"FastH3 segment {segment_idx} requires a retained continuation frame.")
if segment_idx == 1 and image_path:
return self._load_rgb_image(image_path), False
return None, False
def _build_request(
self,
prompt: str,
conditioning_image: Image | None,
last_image: Image | None = None,
generation_inputs: GenerationInputs | None = None,
):
"""Build the typed FastVideo request owned by the FastH3 profile."""
from fastvideo.api import GenerationRequest, InputConfig, OutputConfig, SamplingConfig
references = None
if generation_inputs is not None and generation_inputs.mode == "ref2va":
from fastvideo.api import MiniMaxH3Reference
references = [
MiniMaxH3Reference(source=str(asset.path), media_type=asset.kind)
for asset in generation_inputs.references
]
return GenerationRequest(
prompt=prompt,
negative_prompt="",
inputs=InputConfig(pil_image=conditioning_image, last_image=last_image, references=references),
sampling=SamplingConfig(
height=int(self.model_config["height"]),
width=int(self.model_config["width"]),
num_frames=int(self.model_config["num_frames"]),
fps=24,
num_inference_steps=int(self.model_config["num_inference_steps"]),
guidance_scale=1.0,
batch_cfg=False,
seed=int(self.model_config["seed"]),
),
output=OutputConfig(save_video=False, return_frames=True),
)
def _save_continuation_frame(self, frames: list) -> None:
"""Retain the last decoded frame as first-frame conditioning."""
from PIL import Image
self.clear_conditioning()
self.continuation_image = Image.fromarray(np.ascontiguousarray(frames[-1])).convert("RGB")
def generate_step(
self,
prompt: str,
segment_idx: int,
image_path: str | None,
reset_conditioning: bool,
generation_inputs: GenerationInputs | None = None,
) -> StepResult:
"""Generate one synchronized FastH3 segment and retain its last frame.
Later segments use MiniMax H3's first-frame-to-video path. The first
conditioned frame and its matching audio duration are trimmed before
streaming so adjacent segments do not duplicate media.
"""
mode = generation_inputs.mode if generation_inputs is not None else None
if mode not in (None, "t2va", "fl2va", "ref2va"):
raise ValueError(f"Unsupported H3 generation mode: {mode!r}.")
if mode in ("fl2va", "ref2va") and not self.model_config.get("full_checkpoint", False):
raise ValueError(f"{mode.upper()} requires the full-h3 model profile.")
pipeline_mode = "ref2va" if mode == "ref2va" else "base"
if self.generator is None or self.pipeline_mode != pipeline_mode:
# A failed switch leaves no executor behind; reload here so the
# slot recovers on the next step instead of staying broken.
if segment_idx > 1 and not reset_conditioning:
raise ValueError("Generation mode cannot change in the middle of a project.")
self._load_pipeline(pipeline_mode)
conditioning_image = None
last_image = None
uses_continuation = False
if mode == "ref2va":
# The reference pipeline rejects first/last-frame inputs. Preserve
# all original references for every clip and do not trim overlap.
self.clear_conditioning()
else:
if mode == "fl2va" and generation_inputs is not None:
image_path = generation_inputs.first_frame_path
conditioning_image, uses_continuation = self._select_conditioning_image(
segment_idx,
image_path,
reset_conditioning,
)
started = time.perf_counter()
try:
if (mode == "fl2va" and segment_idx == 1 and generation_inputs is not None
and generation_inputs.last_frame_path):
last_image = self._load_rgb_image(generation_inputs.last_frame_path)
request = self._build_request(prompt, conditioning_image, last_image, generation_inputs)
result = self.generator.generate(request)
finally:
if conditioning_image is not None:
conditioning_image.close()
if last_image is not None:
last_image.close()
torch.cuda.synchronize()
generation_ms = (time.perf_counter() - started) * 1000.0
if isinstance(result, list):
raise RuntimeError("FastH3 returned multiple results for one DreamVerse segment.")
frames = result.frames
if not isinstance(frames, list) or not frames:
raise RuntimeError("FastH3 generation did not return decoded frames.")
audio = result.audio
audio_sample_rate = result.audio_sample_rate
if audio is not None and audio_sample_rate is None:
raise RuntimeError("FastH3 returned audio without an audio sample rate.")
save_started = time.perf_counter()
if mode != "ref2va":
self._save_continuation_frame(frames)
save_conditioning_ms = (time.perf_counter() - save_started) * 1000.0
timings = {
"generation_ms": generation_ms,
"generation_time_ms": float(result.generation_time or 0.0) * 1000.0,
"save_conditioning_ms": save_conditioning_ms,
"e2e_latency_ms": (time.perf_counter() - started) * 1000.0,
}
trim_frames = 1 if uses_continuation else 0
print(f"[GPU {self.gpu_id}] FastH3 segment {segment_idx}: "
f"{len(frames)} frames, gen={generation_ms:.0f}ms, "
f"save_conditioning={save_conditioning_ms:.0f}ms, "
f"e2e={timings['e2e_latency_ms']:.0f}ms")
return StepResult(
frames=frames,
audio=audio,
audio_sample_rate=audio_sample_rate,
timings=timings,
head_trim_frames=trim_frames,
head_trim_audio_frames=trim_frames,
)
def warmup(self, prompt: str) -> dict[str, float]:
"""Compile the FastH3 text and first-frame paths before readiness."""
warmup_prompt = (prompt or "").strip()
if not warmup_prompt:
raise RuntimeError("Startup warmup prompt must be non-empty.")
print(f"[GPU {self.gpu_id}] FastH3 startup warmup starting "
"(synthetic segments: text-to-video, first-frame-to-video)")
started = time.perf_counter()
text_result = self.generate_step(
warmup_prompt,
segment_idx=1,
image_path=None,
reset_conditioning=True,
)
first_frame_result = self.generate_step(
warmup_prompt,
segment_idx=2,
image_path=None,
reset_conditioning=False,
)
total_ms = (time.perf_counter() - started) * 1000.0
self.clear_conditioning()
text_ms = float(text_result.timings.get("e2e_latency_ms", 0.0))
first_frame_ms = float(first_frame_result.timings.get("e2e_latency_ms", 0.0))
print(f"[GPU {self.gpu_id}] FastH3 startup warmup complete: "
f"text_to_video={text_ms:.0f}ms, "
f"first_frame_to_video={first_frame_ms:.0f}ms, "
f"total={total_ms:.0f}ms")
return {
"warmup_text_to_video_ms": text_ms,
"warmup_first_frame_to_video_ms": first_frame_ms,
"warmup_total_ms": total_ms,
}
def apply_lora_stack(self, stack: list[tuple[str, float]]) -> tuple[str | None, str | None]:
"""Reject runtime LoRA mutation because FastH3 uses one startup adapter."""
del stack
raise RuntimeError("FastH3 uses its fixed startup adapter and does not support runtime LoRA changes.")
+41 -1
View File
@@ -32,6 +32,14 @@ from fastapi.staticfiles import StaticFiles
from dreamverse._deps import require_dreamverse_runtime_deps
from dreamverse.config import FRONTEND_STATIC_DIR_CANDIDATES, GENERATION_SEGMENT_CAP
from dreamverse.session_init_image import cleanup_session_init_image, persist_session_init_image
from dreamverse.generation_inputs import (
GenerationInputs,
pin_generation_inputs,
release_generation_inputs,
resolve_generation_inputs,
supported_generation_modes,
)
from dreamverse.routes.assets import router as asset_router
LATENCY_MS = 200
SESSION_TIMEOUT_SECONDS = 300
@@ -170,6 +178,12 @@ app.add_middleware(
allow_methods=["*"],
allow_headers=["*"],
)
app.include_router(asset_router)
@app.get("/generation-capabilities")
async def generation_capabilities():
return {"model_id": "mock", "modes": supported_generation_modes("mock"), "mock": True}
@app.get("/healthz")
@@ -290,6 +304,7 @@ async def websocket_endpoint(websocket: WebSocket):
send_lock = asyncio.Lock()
stop_event = asyncio.Event()
session_init_image = None
generation_inputs = GenerationInputs()
async def ws_send_json(payload: dict) -> None:
async with send_lock:
@@ -347,10 +362,13 @@ async def websocket_endpoint(websocket: WebSocket):
generation_paused = bool(initial_rollout_prompt and not single_clip_mode and len(curated_prompts) == 0)
try:
generation_inputs = resolve_generation_inputs(init_data, "mock")
pin_generation_inputs(generation_inputs)
session_init_image = persist_session_init_image(init_data.get("initial_image"))
except ValueError as exc:
await ws_send_json({
"type": "error",
"error_code": "invalid_generation_input",
"message": str(exc),
})
await websocket.close(code=1003, reason="Invalid initial image")
@@ -362,6 +380,8 @@ async def websocket_endpoint(websocket: WebSocket):
"type": "gpu_assigned",
"gpu_id": 0,
"session_timeout": SESSION_TIMEOUT_SECONDS,
"generation_mode": generation_inputs.mode,
"mock": True,
})
raw_prompt_queue: asyncio.Queue[PromptSubmission] = asyncio.Queue()
@@ -486,6 +506,7 @@ async def websocket_endpoint(websocket: WebSocket):
})
async def apply_project_init_payload(payload: dict[str, object], ) -> bool:
nonlocal generation_inputs
nonlocal preset_id
nonlocal preset_label
nonlocal initial_rollout_prompt
@@ -519,10 +540,19 @@ async def websocket_endpoint(websocket: WebSocket):
]
try:
replace_session_image(payload.get("initial_image"))
next_inputs = resolve_generation_inputs(payload, "mock")
pin_generation_inputs(next_inputs)
try:
replace_session_image(payload.get("initial_image"))
except ValueError:
release_generation_inputs(next_inputs)
raise
release_generation_inputs(generation_inputs)
generation_inputs = next_inputs
except ValueError as exc:
await ws_send_json({
"type": "error",
"error_code": "invalid_generation_input",
"message": str(exc),
})
return False
@@ -568,6 +598,7 @@ async def websocket_endpoint(websocket: WebSocket):
return drained
async def enter_project_idle() -> None:
nonlocal generation_inputs
nonlocal seed_prompt_memory
nonlocal curated_prompts
nonlocal curated_idx
@@ -587,6 +618,8 @@ async def websocket_endpoint(websocket: WebSocket):
dropped_raw = drain_queue_nowait(raw_prompt_queue)
dropped_ready = drain_queue_nowait(ready_prompt_queue)
release_generation_inputs(generation_inputs)
generation_inputs = GenerationInputs()
seed_prompt_memory = []
curated_prompts = []
curated_idx = 0
@@ -767,10 +800,16 @@ async def websocket_endpoint(websocket: WebSocket):
continue
try:
if generation_inputs.mode is not None and data.get("initial_image") is not None:
raise ValueError("Choose conditioning assets when starting a project; legacy initial_image "
"cannot replace generation mode inputs.")
if "generation_mode" in data or "conditioning_assets" in data:
raise ValueError("simple_generate cannot change the mode; use project_init_v1.")
replace_session_image(data.get("initial_image"))
except ValueError as exc:
await ws_send_json({
"type": "error",
"error_code": "invalid_generation_input",
"message": str(exc),
})
continue
@@ -1182,6 +1221,7 @@ async def websocket_endpoint(websocket: WebSocket):
finally:
stop_event.set()
cleanup_session_init_image(session_init_image)
release_generation_inputs(generation_inputs)
for static_dir in FRONTEND_STATIC_DIR_CANDIDATES:
@@ -0,0 +1,72 @@
"""Raw, bounded media uploads keep large binary data out of websocket messages."""
from __future__ import annotations
import asyncio
from urllib.parse import unquote
from fastapi import APIRouter, HTTPException, Request, Response
from fastapi.responses import FileResponse
from starlette.concurrency import run_in_threadpool
from dreamverse.assets import IMAGE_LIMIT, MEDIA_LIMIT, MIME_TYPES, asset_store
router = APIRouter()
_upload_lock = asyncio.Lock()
@router.post("/assets", status_code=201)
async def upload_asset(request: Request) -> dict:
mime_type = request.headers.get("content-type", "").split(";", 1)[0].lower()
if mime_type not in MIME_TYPES:
raise HTTPException(415, "Unsupported asset type. Select a supported image, video, or audio file.")
limit = IMAGE_LIMIT if MIME_TYPES[mime_type][0] == "image" else MEDIA_LIMIT
try:
if int(request.headers.get("content-length", "0")) > limit:
raise HTTPException(413, f"Asset exceeds the {limit // (1024 * 1024)} MB upload limit.")
except ValueError as exc:
raise HTTPException(400, "Invalid Content-Length.") from exc
async with _upload_lock:
try:
path = asset_store.staging_path(mime_type)
except ValueError as exc:
raise HTTPException(400, str(exc)) from exc
try:
size = 0
with path.open("xb") as handle:
async for chunk in request.stream():
size += len(chunk)
if size > limit:
raise HTTPException(413, f"Asset exceeds the {limit // (1024 * 1024)} MB upload limit.")
await run_in_threadpool(handle.write, chunk)
asset = await run_in_threadpool(asset_store.add, path,
unquote(request.headers.get("x-asset-name", "Untitled asset")), mime_type)
return asset.public()
except ValueError as exc:
path.unlink(missing_ok=True)
raise HTTPException(400, str(exc)) from exc
except BaseException:
path.unlink(missing_ok=True)
raise
@router.api_route("/assets/{asset_id}", methods=["GET", "HEAD"])
async def get_asset(asset_id: str) -> FileResponse:
try:
asset = asset_store.get(asset_id)
except ValueError as exc:
raise HTTPException(404, str(exc)) from exc
return FileResponse(asset.path, media_type=asset.mime_type, headers={"X-Content-Type-Options": "nosniff"})
@router.delete("/assets/{asset_id}", status_code=204)
async def delete_asset(asset_id: str) -> Response:
try:
asset_store.get(asset_id)
except ValueError as exc:
raise HTTPException(404, str(exc)) from exc
try:
asset_store.delete(asset_id)
except ValueError as exc:
raise HTTPException(409, str(exc)) from exc
return Response(status_code=204)
@@ -26,11 +26,17 @@ from typing import TYPE_CHECKING
from fastapi import WebSocket, WebSocketDisconnect
from dreamverse.gpu_pool import GPUSlot
from dreamverse.generation_inputs import (
GenerationInputs,
pin_generation_inputs,
release_generation_inputs,
resolve_generation_inputs,
)
from dreamverse.session_init_image import cleanup_session_init_image, persist_session_init_image
from dreamverse.worker_ipc import MediaChunk, MediaComplete, MediaInit
from dreamverse.config import (
DEFAULT_MODEL_ID,
ACTIVE_MODEL_ID,
GENERATION_SEGMENT_CAP,
PROMPT_AUTO_SLEEP_MS,
PROMPT_AUTO_TIMEOUT_MS,
@@ -156,6 +162,7 @@ class SessionController:
prompt_worker_task: asyncio.Task | None = None
rewrite_seed_prompts_task: asyncio.Task | None = None
session_init_image = None
generation_inputs: GenerationInputs | None = None
async def session_timeout():
"""Close the session after timeout."""
@@ -191,6 +198,18 @@ class SessionController:
init_data = {}
init_type = init_data.get("type")
try:
next_generation_inputs = resolve_generation_inputs(init_data, ACTIVE_MODEL_ID)
pin_generation_inputs(next_generation_inputs)
generation_inputs = next_generation_inputs
except ValueError as exc:
await ws_send_json({
"type": "error",
"error_code": "invalid_generation_input",
"message": str(exc),
})
await websocket.close(code=1008, reason="Invalid generation inputs")
return
preset_id = init_data.get("preset_id")
preset_label = str(init_data.get("preset_label") or "").strip()
initial_rollout_prompt = str(init_data.get("initial_rollout_prompt") or "").strip()
@@ -264,13 +283,14 @@ class SessionController:
timeout_task = asyncio.create_task(session_timeout())
# Join the engine on this GPU.
await slot.join_user(client_id, model_id=DEFAULT_MODEL_ID)
await slot.join_user(client_id, model_id=ACTIVE_MODEL_ID)
# Notify client they're connected to a GPU.
await ws_send_json({
"type": "gpu_assigned",
"gpu_id": gpu_id,
"session_timeout": SESSION_TIMEOUT_SECONDS,
"generation_mode": generation_inputs.mode,
})
await log_event(
"gpu_assigned",
@@ -361,10 +381,16 @@ class SessionController:
return
try:
if generation_inputs.mode is not None and payload.get("initial_image") is not None:
raise ValueError("Choose conditioning assets when starting a project; legacy initial_image "
"cannot replace generation mode inputs.")
if "generation_mode" in payload or "conditioning_assets" in payload:
raise ValueError("simple_generate cannot change the mode; use project_init_v1.")
replace_session_init_image(payload.get("initial_image"))
except ValueError as exc:
await ws_send_json({
"type": "error",
"error_code": "invalid_generation_input",
"message": str(exc),
})
return
@@ -417,6 +443,7 @@ class SessionController:
})
async def apply_project_init_payload(payload: dict[str, object]) -> bool:
nonlocal generation_inputs
nonlocal preset_id
nonlocal preset_label
nonlocal initial_rollout_prompt
@@ -496,15 +523,26 @@ class SessionController:
})
return False
next_generation_inputs = None
next_inputs_pinned = False
try:
next_generation_inputs = resolve_generation_inputs(payload, ACTIVE_MODEL_ID)
pin_generation_inputs(next_generation_inputs)
next_inputs_pinned = True
replace_session_init_image(payload.get("initial_image"))
except ValueError as exc:
if next_inputs_pinned:
release_generation_inputs(next_generation_inputs)
await ws_send_json({
"type": "error",
"error_code": "invalid_generation_input",
"message": str(exc),
})
return False
release_generation_inputs(generation_inputs)
generation_inputs = next_generation_inputs
initial_rollout_prompt = next_initial_rollout_prompt
enhancement_enabled = next_enhancement_enabled
auto_extension_enabled = next_auto_extension_enabled
@@ -1171,6 +1209,7 @@ class SessionController:
return drained
async def enter_project_idle() -> None:
nonlocal generation_inputs
nonlocal curated_prompts
nonlocal seed_prompt_memory
nonlocal curated_idx
@@ -1221,6 +1260,9 @@ class SessionController:
project_active = False
pending_project_end = False
release_generation_inputs(generation_inputs)
generation_inputs = GenerationInputs()
if project_stream_started:
project_stream_started = False
await ws_send_json({"type": "ltx2_stream_complete"})
@@ -1627,6 +1669,7 @@ class SessionController:
segment_idx=segment_idx,
image_path=step_image_path,
reset_conditioning=step_reset_conditioning,
generation_inputs=generation_inputs,
))
segment_generation_active = True
try:
@@ -1681,10 +1724,10 @@ class SessionController:
print(f"[GPU {gpu_id}] Unknown AV event: "
f"{type(event).__name__}")
if not step_task.done():
step_task.cancel()
else:
timings = await step_task
# A GPU command cannot be cancelled by cancelling its
# asyncio waiter. Await completion before releasing pinned
# asset files or making this GPU available to a new user.
timings = await step_task
finally:
segment_generation_active = False
if not step_task.done():
@@ -1808,3 +1851,5 @@ class SessionController:
await self.gpu_pool.release(client_id)
finally:
cleanup_session_init_image(session_init_image)
if generation_inputs is not None:
release_generation_inputs(generation_inputs)
@@ -3,7 +3,6 @@ from __future__ import annotations
import sys
from pathlib import Path
TESTS_DIR = Path(__file__).resolve().parent
DREAMVERSE_PACKAGE_DIR = TESTS_DIR.parent
DREAMVERSE_APP_DIR = DREAMVERSE_PACKAGE_DIR.parent
+136 -22
View File
@@ -2,14 +2,14 @@ from __future__ import annotations
import importlib.util
from pathlib import Path
from types import ModuleType
import pytest
SERVER_DIR = Path(__file__).resolve().parents[1]
def _load_config_module():
def _load_config_module() -> ModuleType:
spec = importlib.util.spec_from_file_location(
"server_config_test_module",
SERVER_DIR / "config.py",
@@ -21,7 +21,7 @@ def _load_config_module():
return module
def _set_required_prompt_keys(monkeypatch):
def _set_required_prompt_keys(monkeypatch: pytest.MonkeyPatch) -> None:
monkeypatch.setenv("CEREBRAS_API_KEY", "cerebras-key")
monkeypatch.setenv("GROQ_API_KEY", "groq-key")
@@ -53,9 +53,7 @@ def test_config_defaults_to_cerebras_with_parallel_groq_fallback_stage(monkeypat
module = _load_config_module()
assert module.PROMPT_PROVIDER == "cerebras"
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (
("cerebras", "groq"),
)
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (("cerebras", "groq"), )
assert module.PROMPT_PROVIDER_PRIORITY == (
"cerebras",
"groq",
@@ -86,9 +84,7 @@ def test_config_ignores_legacy_groq_primary_override(monkeypatch):
module = _load_config_module()
assert module.PROMPT_PROVIDER == "cerebras"
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (
("cerebras", "groq"),
)
assert module.PROMPT_PROVIDER_RUNTIME_STAGES == (("cerebras", "groq"), )
assert module.PROMPT_PROVIDER_PRIORITY == (
"cerebras",
"groq",
@@ -106,24 +102,17 @@ def test_config_uses_local_overlay_paths_when_devtools_enabled(monkeypatch, tmp_
assert module.DEVTOOLS_ENABLED is True
assert module.FRONTEND_ROOT.as_posix().endswith("apps/dreamverse/web")
assert module.PROMPT_ENHANCE_SYSTEM_PROMPT_PATH.endswith(
"dreamverse/prompts.local/next_segment_system_prompt.md"
)
assert module.PROMPT_ENHANCE_SYSTEM_PROMPT_PATH.endswith("dreamverse/prompts.local/next_segment_system_prompt.md")
assert module.PROMPT_ENHANCE_SYSTEM_PROMPT_FALLBACK_PATH.endswith(
"dreamverse/prompts/next_segment_system_prompt.md"
)
"dreamverse/prompts/next_segment_system_prompt.md")
assert module.PROMPT_REWRITE_USER_SYSTEM_PROMPT_PATH.endswith(
"dreamverse/prompts.local/rewrite_user_system_prompt.md"
)
"dreamverse/prompts.local/rewrite_user_system_prompt.md")
assert module.PROMPT_REWRITE_USER_SYSTEM_PROMPT_FALLBACK_PATH.endswith(
"dreamverse/prompts/rewrite_user_system_prompt.md"
)
"dreamverse/prompts/rewrite_user_system_prompt.md")
assert module.CURATED_PRESETS_FILE_PATH.endswith(
"apps/dreamverse/web/prompts.local/selected_ltx2_continuation_story_presets.json"
)
"apps/dreamverse/web/prompts.local/selected_ltx2_continuation_story_presets.json")
assert module.CURATED_PRESETS_FALLBACK_FILE_PATH.endswith(
"apps/dreamverse/web/prompts/selected_ltx2_continuation_story_presets.json"
)
"apps/dreamverse/web/prompts/selected_ltx2_continuation_story_presets.json")
assert module.FRONTEND_STATIC_DIR_CANDIDATES[:2] == (
str(module.FRONTEND_ROOT / "out"),
str(module.FRONTEND_ROOT / "dist"),
@@ -150,15 +139,140 @@ def test_config_enables_prompt_safety_when_requested(monkeypatch):
def test_config_uses_five_minute_session_timeout(monkeypatch):
_set_required_prompt_keys(monkeypatch)
monkeypatch.delenv("DREAMVERSE_MODEL_ID", raising=False)
monkeypatch.delenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", raising=False)
monkeypatch.delenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", raising=False)
module = _load_config_module()
assert module.SESSION_TIMEOUT_SECONDS == 300
def test_config_uses_thirty_minute_cosmos25_session_timeout(monkeypatch):
_set_required_prompt_keys(monkeypatch)
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "cosmos25-dfd")
monkeypatch.delenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", raising=False)
monkeypatch.delenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", raising=False)
module = _load_config_module()
assert module.SESSION_TIMEOUT_SECONDS == 1800
def test_config_allows_session_timeout_override(monkeypatch):
_set_required_prompt_keys(monkeypatch)
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "cosmos25-dfd")
monkeypatch.delenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", raising=False)
monkeypatch.setenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", "900")
module = _load_config_module()
assert module.SESSION_TIMEOUT_SECONDS == 900
def test_config_prefers_dreamverse_session_timeout_over_alias(monkeypatch):
_set_required_prompt_keys(monkeypatch)
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "cosmos25-dfd")
monkeypatch.setenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", "1200")
monkeypatch.setenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", "900")
module = _load_config_module()
assert module.SESSION_TIMEOUT_SECONDS == 1200
def test_config_rejects_invalid_prompt_provider(monkeypatch):
monkeypatch.setenv("FASTVIDEO_PROMPT_PROVIDER", "unsupported")
_set_required_prompt_keys(monkeypatch)
with pytest.raises(RuntimeError, match="Invalid FASTVIDEO_PROMPT_PROVIDER"):
_load_config_module()
def test_config_registers_vsa_datafree_fasth3_profile(monkeypatch):
"""The FastH3 registry entry owns the complete fixed Preview recipe."""
_set_required_prompt_keys(monkeypatch)
module = _load_config_module()
assert module.MODEL_REGISTRY["fast-h3"] == {
"name": "FastH3",
"generation_backend": "minimax_h3",
"default_sp_size": 4,
"model_path": "MiniMaxAI/MiniMax-H3",
"adapter_repo": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
"adapter_filename": "vsa-datafree/adapter_model.safetensors",
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
"height": 768,
"width": 1344,
"num_frames": 124,
"num_inference_steps": 5,
"seed": 1000,
}
def test_config_uses_fasth3_sequence_parallel_default(monkeypatch):
"""Selecting FastH3 defaults DreamVerse to its four-GPU topology."""
_set_required_prompt_keys(monkeypatch)
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "fast-h3")
monkeypatch.delenv("DREAMVERSE_SP_SIZE", raising=False)
module = _load_config_module()
assert module.ACTIVE_MODEL_ID == "fast-h3"
assert module.MODEL_CONFIG["generation_backend"] == "minimax_h3"
assert module.DREAMVERSE_SP_SIZE == 4
def test_full_h3_profile_has_no_preview_adapter_and_longer_session(monkeypatch):
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "full-h3")
monkeypatch.delenv("DREAMVERSE_SP_SIZE", raising=False)
monkeypatch.delenv("DREAMVERSE_SESSION_TIMEOUT_SECONDS", raising=False)
monkeypatch.delenv("FASTVIDEO_SESSION_TIMEOUT_SECONDS", raising=False)
module = _load_config_module()
assert module.MODEL_CONFIG["full_checkpoint"] is True
assert "adapter_repo" not in module.MODEL_CONFIG
assert module.MODEL_CONFIG["num_inference_steps"] == 50
assert module.DREAMVERSE_SP_SIZE == 4
assert module.SESSION_TIMEOUT_SECONDS == 7200
def test_config_registers_cosmos25_dfd_profile(monkeypatch):
_set_required_prompt_keys(monkeypatch)
module = _load_config_module()
assert module.MODEL_REGISTRY["cosmos25-dfd"] == {
"name": "Cosmos Predict2.5 DFD",
"generation_backend": "cosmos25_dfd",
"default_sp_size": 1,
"model_path": "FastVideo/Cosmos-Predict2.5-2B-Distilled-TrigFlow",
"continuation_model_path": "FastVideo/Cosmos-Predict2.5-2B-DFD",
"attention_backend": "TORCH_SDPA",
"height": 704,
"width": 1280,
"bootstrap_num_frames": 77,
"continuation_num_frames": 81,
"fps": 24,
"num_inference_steps": 4,
"seed": 42,
"session_timeout_seconds": 1800,
}
def test_config_selects_cosmos25_package_roles(monkeypatch, tmp_path):
_set_required_prompt_keys(monkeypatch)
bootstrap_path = tmp_path / "cosmos25-t2w"
continuation_path = tmp_path / "cosmos25-dfd"
monkeypatch.setenv("DREAMVERSE_MODEL_ID", "cosmos25-dfd")
monkeypatch.setenv("DREAMVERSE_MODEL_PATH", str(bootstrap_path))
monkeypatch.setenv("DREAMVERSE_COSMOS25_DFD_MODEL_PATH", str(continuation_path))
monkeypatch.delenv("DREAMVERSE_SP_SIZE", raising=False)
module = _load_config_module()
assert module.ACTIVE_MODEL_ID == "cosmos25-dfd"
assert module.MODEL_CONFIG["generation_backend"] == "cosmos25_dfd"
assert module.MODEL_CONFIG["model_path"] == str(bootstrap_path)
assert module.MODEL_CONFIG["continuation_model_path"] == str(continuation_path)
assert module.DREAMVERSE_SP_SIZE == 1
@@ -0,0 +1,219 @@
from __future__ import annotations
import os
from pathlib import Path
from types import SimpleNamespace
import numpy as np
import pytest
from dreamverse.cosmos25_dfd_generation import Cosmos25DFDGenerationBackend
from dreamverse.generation_inputs import GenerationInputs
COSMOS_CONFIG = {
"name": "Cosmos Predict2.5 DFD",
"generation_backend": "cosmos25_dfd",
"default_sp_size": 1,
"model_path": "/models/cosmos25-t2w",
"continuation_model_path": "/models/cosmos25-dfd",
"attention_backend": "TORCH_SDPA",
"height": 704,
"width": 1280,
"bootstrap_num_frames": 77,
"continuation_num_frames": 81,
"fps": 24,
"num_inference_steps": 4,
"seed": 42,
}
class _RecordingGenerator:
def __init__(self, pixel_value: int = 20) -> None:
self.pixel_value = pixel_value
self.calls: list[dict] = []
self.shutdown_calls = 0
def generate_video(self, prompt, sampling_param):
condition = sampling_param.pil_image
self.calls.append({
"prompt": prompt,
"sampling": sampling_param,
"conditioning_pixels": None if condition is None else np.asarray(condition).copy(),
})
frames = [
np.full((2, 3, 3), self.pixel_value, dtype=np.uint8)
for _ in range(sampling_param.num_frames)
]
frames[-1] = np.full((2, 3, 3), self.pixel_value + 1, dtype=np.uint8)
return {
"frames": frames,
"generation_time": 0.25,
}
def shutdown(self):
self.shutdown_calls += 1
@pytest.fixture
def backend(monkeypatch) -> Cosmos25DFDGenerationBackend:
instance = Cosmos25DFDGenerationBackend(gpu_id=0)
instance.model_config = dict(COSMOS_CONFIG)
instance.bootstrap_generator = _RecordingGenerator(pixel_value=20)
instance.continuation_generator = _RecordingGenerator(pixel_value=40)
monkeypatch.setattr("dreamverse.cosmos25_dfd_generation.torch.cuda.synchronize", lambda: None)
def fake_sampling_param(*, conditioned):
return SimpleNamespace(
negative_prompt="",
save_video=False,
return_frames=True,
height=704,
width=1280,
num_frames=81 if conditioned else 77,
fps=24,
num_inference_steps=4,
guidance_scale=1.0,
seed=42,
num_cond_frames=1 if conditioned else 0,
pil_image=None,
)
monkeypatch.setattr(instance, "_sampling_param", fake_sampling_param)
return instance
def test_initialize_loads_both_package_roles(monkeypatch):
loaded_paths = []
generators = [_RecordingGenerator(), _RecordingGenerator()]
backend = Cosmos25DFDGenerationBackend(gpu_id=0)
def fake_load(model_path):
loaded_paths.append(model_path)
return generators[len(loaded_paths) - 1]
monkeypatch.setattr(backend, "_load_generator", fake_load)
monkeypatch.setattr(backend, "_gpu_mem", lambda: "alloc=0.00GiB, reserved=0.00GiB")
monkeypatch.setattr("dreamverse.cosmos25_dfd_generation.gc.collect", lambda: 0)
monkeypatch.setattr("dreamverse.cosmos25_dfd_generation.torch.cuda.is_available", lambda: False)
monkeypatch.setenv("FASTVIDEO_ATTENTION_BACKEND", "test-attention")
monkeypatch.setenv("FASTVIDEO_INFERENCE_TORCH_COMPILE", "1")
backend.initialize(COSMOS_CONFIG)
assert loaded_paths == [
"/models/cosmos25-t2w",
"/models/cosmos25-dfd",
]
assert backend.bootstrap_generator is generators[0]
assert backend.continuation_generator is generators[1]
assert backend.model_config == COSMOS_CONFIG
assert os.environ["FASTVIDEO_ATTENTION_BACKEND"] == "TORCH_SDPA"
assert "FASTVIDEO_INFERENCE_TORCH_COMPILE" not in os.environ
def test_unconditioned_start_uses_t2w_and_retains_terminal_frame(backend):
result = backend.generate_step("first prompt", 1, None, True)
assert len(backend.bootstrap_generator.calls) == 1
assert backend.continuation_generator.calls == []
sampling = backend.bootstrap_generator.calls[0]["sampling"]
assert sampling.height == 704
assert sampling.width == 1280
assert sampling.num_frames == 77
assert sampling.fps == 24
assert sampling.num_inference_steps == 4
assert sampling.guidance_scale == 1.0
assert sampling.seed == 42
assert sampling.num_cond_frames == 0
assert sampling.pil_image is None
assert result.head_trim_frames == 0
assert result.head_trim_audio_frames == 0
assert result.audio_sample_rate == 24_000
assert result.audio.shape == (77_000, )
assert result.audio.count_nonzero() == 0
assert np.asarray(backend.continuation_image).tolist() == np.full((2, 3, 3), 21).tolist()
def test_retained_frame_uses_dfd_and_trims_repeated_boundary(backend):
backend.generate_step("first prompt", 1, None, True)
result = backend.generate_step("pivot right", 2, None, False)
assert len(backend.continuation_generator.calls) == 1
call = backend.continuation_generator.calls[0]
sampling = call["sampling"]
assert sampling.num_frames == 81
assert sampling.num_cond_frames == 1
assert call["conditioning_pixels"].tolist() == np.full((2, 3, 3), 21).tolist()
assert result.head_trim_frames == 1
assert result.head_trim_audio_frames == 1
assert result.audio.shape == (81_000, )
assert np.asarray(backend.continuation_image).tolist() == np.full((2, 3, 3), 41).tolist()
def test_initial_image_uses_dfd_without_stream_trim(backend, tmp_path: Path):
from PIL import Image
image_path = tmp_path / "initial.png"
Image.fromarray(np.full((2, 3, 3), 7, dtype=np.uint8)).save(image_path)
result = backend.generate_step("animate", 1, str(image_path), True)
assert backend.bootstrap_generator.calls == []
call = backend.continuation_generator.calls[0]
assert call["conditioning_pixels"].tolist() == np.full((2, 3, 3), 7).tolist()
assert result.head_trim_frames == 0
assert result.head_trim_audio_frames == 0
def test_generation_mode_api_accepts_text_only_and_rejects_conditioning_modes(backend):
result = backend.generate_step("first prompt", 1, None, True, generation_inputs=GenerationInputs(mode="t2va"))
assert len(backend.bootstrap_generator.calls) == 1
assert result.head_trim_frames == 0
with pytest.raises(ValueError, match="text generation only"):
backend.generate_step("pivot right", 2, None, False, generation_inputs=GenerationInputs(mode="fl2va"))
assert backend.continuation_generator.calls == []
def test_missing_later_continuation_fails_before_generation(backend):
with pytest.raises(RuntimeError, match="requires a retained continuation frame"):
backend.generate_step("later prompt", 2, None, False)
assert backend.bootstrap_generator.calls == []
assert backend.continuation_generator.calls == []
def test_reset_later_segment_uses_fresh_t2w_bootstrap(backend):
backend.generate_step("first prompt", 1, None, True)
result = backend.generate_step("new scene", 2, None, True)
assert len(backend.bootstrap_generator.calls) == 2
assert backend.continuation_generator.calls == []
assert result.head_trim_frames == 0
def test_warmup_exercises_bootstrap_and_dfd_paths(backend):
timings = backend.warmup("warmup prompt")
assert len(backend.bootstrap_generator.calls) == 1
assert len(backend.continuation_generator.calls) == 1
assert backend.continuation_image is None
assert "warmup_bootstrap_ms" in timings
assert "warmup_continuation_ms" in timings
assert "warmup_total_ms" in timings
def test_shutdown_releases_both_generators_and_conditioning(backend):
bootstrap = backend.bootstrap_generator
continuation = backend.continuation_generator
backend.generate_step("first prompt", 1, None, True)
backend.shutdown()
assert bootstrap.shutdown_calls == 1
assert continuation.shutdown_calls == 1
assert backend.bootstrap_generator is None
assert backend.continuation_generator is None
assert backend.continuation_image is None
@@ -9,6 +9,7 @@ from fastapi.testclient import TestClient
import fastvideo.entrypoints.streaming as streaming_entrypoints
import pytest
def _install_stack03_import_stubs(monkeypatch):
"""Keep entrypoint tests focused while later-stack runtime modules are absent."""
if not hasattr(streaming_entrypoints, "build_health_router"):
@@ -17,6 +18,7 @@ def _install_stack03_import_stubs(monkeypatch):
gpu_pool_stub = types.ModuleType("dreamverse.gpu_pool")
class GPUPool:
def __init__(self, _gpu_ids):
pass
@@ -49,6 +51,7 @@ def _install_stack03_import_stubs(monkeypatch):
controller_stub = types.ModuleType("dreamverse.session.controller")
class SessionController:
def __init__(self, **_kwargs):
pass
@@ -76,13 +79,11 @@ def _run_cli(module, monkeypatch, argv: list[str]) -> list[dict[str, object]]:
uvicorn_stub = types.ModuleType("uvicorn")
def run(app, host: str, port: int) -> None:
calls.append(
{
"app": app,
"host": host,
"port": port,
}
)
calls.append({
"app": app,
"host": host,
"port": port,
})
uvicorn_stub.run = run
monkeypatch.setitem(sys.modules, "uvicorn", uvicorn_stub)
@@ -99,13 +100,11 @@ def test_server_cli_defaults_to_local_web_port(monkeypatch):
server_main = _import_server_main(monkeypatch)
calls = _run_cli(server_main, monkeypatch, ["dreamverse-server"])
assert calls == [
{
"app": server_main.app,
"host": "0.0.0.0",
"port": 8009,
}
]
assert calls == [{
"app": server_main.app,
"host": "0.0.0.0",
"port": 8009,
}]
def test_server_cli_allows_explicit_host_and_port(monkeypatch):
@@ -116,13 +115,11 @@ def test_server_cli_allows_explicit_host_and_port(monkeypatch):
["dreamverse-server", "--host", "127.0.0.1", "--port", "8123"],
)
assert calls == [
{
"app": server_main.app,
"host": "127.0.0.1",
"port": 8123,
}
]
assert calls == [{
"app": server_main.app,
"host": "127.0.0.1",
"port": 8123,
}]
def test_server_does_not_expose_backend_source_as_static_assets(monkeypatch):
@@ -142,13 +139,11 @@ def test_mock_server_cli_defaults_to_local_web_port(monkeypatch):
["dreamverse-mock-server"],
)
assert calls == [
{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8009,
}
]
assert calls == [{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8009,
}]
def test_mock_server_cli_updates_latency(monkeypatch):
@@ -161,13 +156,11 @@ def test_mock_server_cli_updates_latency(monkeypatch):
["dreamverse-mock-server", "--latency", "321", "--port", "8111"],
)
assert calls == [
{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8111,
}
]
assert calls == [{
"app": mock_server.app,
"host": "0.0.0.0",
"port": 8111,
}]
assert mock_server.LATENCY_MS == 321
finally:
mock_server.LATENCY_MS = old_latency_ms
@@ -0,0 +1,248 @@
"""Contract regressions independent of CUDA and model weights."""
import io
import asyncio
import shutil
import subprocess
import pytest
from fastapi import FastAPI
from fastapi.testclient import TestClient
from PIL import Image
from dreamverse import assets, generation_inputs
from dreamverse.generation_inputs import resolve_generation_inputs
from dreamverse.routes import assets as asset_routes
from dreamverse.tests.test_mock_server import _FakeWebSocket
@pytest.fixture
def library(monkeypatch):
store = assets.AssetStore()
monkeypatch.setattr(asset_routes, "asset_store", store)
monkeypatch.setattr(generation_inputs, "asset_store", store)
app = FastAPI()
app.include_router(asset_routes.router)
with TestClient(app) as client:
yield store, client
def upload_image(client, color="red"):
image_bytes = io.BytesIO()
Image.new("RGB", (32, 32), color).save(image_bytes, format="PNG")
response = client.post("/assets", content=image_bytes.getvalue(),
headers={"Content-Type": "image/png", "X-Asset-Name": "frame.png"})
assert response.status_code == 201, response.text
return response.json()
def conditioning(asset, role):
return {"asset_id": asset["asset_id"], "role": role}
def test_assets_validate_content_and_support_head_range_and_delete(library):
store, client = library
asset = upload_image(client)
assert set(asset) == {"asset_id", "kind", "name", "mime_type", "size", "url"}
assert client.head(asset["url"]).status_code == 200
response = client.get(asset["url"], headers={"Range": "bytes=0-7"})
assert response.status_code == 206
assert response.content == b"\x89PNG\r\n\x1a\n"
assert client.post("/assets", content=b"not an image", headers={"Content-Type": "image/png"}).status_code == 400
assert client.post("/assets", content=b"<svg/>", headers={"Content-Type": "image/svg+xml"}).status_code == 415
assert client.post("/assets", content=b"", headers={"Content-Type": "image/png",
"Content-Length": str(assets.IMAGE_LIMIT + 1)}).status_code == 413
with pytest.raises(ValueError, match="Invalid asset ID"):
store.get("../../etc/passwd")
assert client.delete(asset["url"]).status_code == 204
assert client.head(asset["url"]).status_code == 404
def test_generation_pin_prevents_deletion_until_session_releases(library):
_, client = library
asset = upload_image(client)
inputs = resolve_generation_inputs({"generation_mode": "fl2va", "conditioning_assets": [
conditioning(asset, "first_frame")
]}, "full-h3")
generation_inputs.pin_generation_inputs(inputs)
generation_inputs.pin_generation_inputs(inputs)
assert client.delete(asset["url"]).status_code == 409
generation_inputs.release_generation_inputs(inputs)
assert client.delete(asset["url"]).status_code == 409
generation_inputs.release_generation_inputs(inputs)
assert client.delete(asset["url"]).status_code == 204
def test_legacy_init_remains_compatible_but_explicit_t2va_is_text_only(library):
_, client = library
assert resolve_generation_inputs({"initial_image": {"old": "payload"}}, "fast-ltx2").mode is None
assert resolve_generation_inputs({"generation_mode": "t2va"}, "fast-ltx2").mode == "t2va"
with pytest.raises(ValueError, match="legacy initial_image"):
resolve_generation_inputs({"generation_mode": "t2va", "initial_image": {}}, "full-h3")
asset = upload_image(client)
with pytest.raises(ValueError, match="text only"):
resolve_generation_inputs({"generation_mode": "t2va", "conditioning_assets": [
conditioning(asset, "reference")
]}, "full-h3")
@pytest.mark.parametrize("mode", ["unknown", None, 3, [], {}])
def test_unknown_mode_fails_before_assets_are_resolved(mode):
with pytest.raises(ValueError, match="Unknown generation mode"):
resolve_generation_inputs({"generation_mode": mode}, "full-h3")
@pytest.mark.parametrize("model_id", ["fast-h3", "fast-ltx2", "fast-ltx23"])
def test_preview_and_ltx_cannot_advertise_full_h3_modes(model_id):
with pytest.raises(ValueError, match="Full H3"):
resolve_generation_inputs({"generation_mode": "ref2va"}, model_id)
def test_fl2va_first_required_last_optional_and_roles_unique(library):
_, client = library
first = upload_image(client)
last = upload_image(client, "blue")
payload = {"generation_mode": "fl2va", "conditioning_assets": [conditioning(first, "first_frame")]}
inputs = resolve_generation_inputs(payload, "full-h3")
assert inputs.first_frame_path.endswith(".png")
assert inputs.last_frame_path is None
payload["conditioning_assets"].append(conditioning(last, "last_frame"))
assert resolve_generation_inputs(payload, "full-h3").last_frame_path is not None
payload["conditioning_assets"].append(conditioning(first, "first_frame"))
with pytest.raises(ValueError, match="exactly one first-frame"):
resolve_generation_inputs(payload, "full-h3")
with pytest.raises(ValueError, match="exactly one first-frame"):
resolve_generation_inputs({"generation_mode": "fl2va", "conditioning_assets": [
conditioning(last, "last_frame")
]}, "full-h3")
def test_ref_order_is_preserved_and_limits_are_enforced(library):
_, client = library
first, second = upload_image(client), upload_image(client, "blue")
refs = [conditioning(second, "reference"), conditioning(first, "reference")]
payload = {"generation_mode": "ref2va", "conditioning_assets": refs}
inputs = resolve_generation_inputs(payload, "full-h3")
assert [asset.asset_id for asset in inputs.references] == [second["asset_id"], first["asset_id"]]
with pytest.raises(ValueError, match="at most 9 image"):
resolve_generation_inputs({**payload, "conditioning_assets": refs * 5}, "full-h3")
with pytest.raises(ValueError, match="without keyframe roles"):
resolve_generation_inputs({**payload, "conditioning_assets": [conditioning(first, "first_frame")]}, "full-h3")
with pytest.raises(ValueError, match="at most 12"):
resolve_generation_inputs({**payload, "conditioning_assets": refs * 7}, "full-h3")
def test_ref_audio_requires_visual_reference(library, monkeypatch):
store, _ = library
monkeypatch.setattr(store, "get", lambda asset_id: assets.StoredAsset(asset_id, "audio", "/audio.wav", "audio",
"audio/wav", 100))
with pytest.raises(ValueError, match="audio alone"):
resolve_generation_inputs({"generation_mode": "ref2va", "conditioning_assets": [
{"asset_id": "a" * 32, "role": "reference"}
]}, "full-h3")
def test_ref_rejects_extreme_image_aspect_before_gpu(library):
_, client = library
content = io.BytesIO()
Image.new("RGB", (500, 50), "blue").save(content, format="PNG")
response = client.post("/assets", content=content.getvalue(), headers={"Content-Type": "image/png"})
assert response.status_code == 201
with pytest.raises(ValueError, match="aspect ratios"):
resolve_generation_inputs({"generation_mode": "ref2va", "conditioning_assets": [
conditioning(response.json(), "reference")
]}, "full-h3")
def test_ref_reports_undecodable_image_as_invalid_input(library, monkeypatch, tmp_path):
store, _ = library
broken = tmp_path / "broken.png"
broken.write_bytes(b"not an image")
monkeypatch.setattr(store, "get", lambda asset_id: assets.StoredAsset(asset_id, "image", str(broken), "broken.png",
"image/png", 11))
with pytest.raises(ValueError, match="could not be decoded"):
resolve_generation_inputs({"generation_mode": "ref2va", "conditioning_assets": [
{"asset_id": "a" * 32, "role": "reference"}
]}, "full-h3")
def test_audio_upload_rejects_surround_sound(library, monkeypatch):
import json
_, client = library
monkeypatch.setattr(assets.shutil, "which", lambda name: "/usr/bin/ffprobe")
info = {"format": {"format_name": "wav", "duration": "1"},
"streams": [{"codec_type": "audio", "channels": 6}]}
monkeypatch.setattr(assets.subprocess, "run", lambda *args, **kwargs: subprocess.CompletedProcess(
[], 0, stdout=json.dumps(info).encode(), stderr=b""))
response = client.post("/assets", content=b"surround wav", headers={"Content-Type": "audio/wav"})
assert response.status_code == 400
assert "mono or stereo" in response.json()["detail"]
@pytest.mark.parametrize("mime,format_name", [("audio/x-m4a", "mov,mp4,m4a,3gp,3g2,mj2"), ("audio/x-flac", "flac")])
def test_legacy_audio_mime_aliases_are_accepted(library, monkeypatch, mime, format_name):
"""Browsers report x- variants for the M4A and FLAC formats the docs promise."""
import json
_, client = library
monkeypatch.setattr(assets.shutil, "which", lambda name: "/usr/bin/ffprobe")
info = {"format": {"format_name": format_name, "duration": "1"},
"streams": [{"codec_type": "audio", "channels": 2}]}
monkeypatch.setattr(assets.subprocess, "run", lambda *args, **kwargs: subprocess.CompletedProcess(
[], 0, stdout=json.dumps(info).encode(), stderr=b""))
response = client.post("/assets", content=b"audio bytes", headers={"Content-Type": mime})
assert response.status_code == 201, response.text
assert response.json()["kind"] == "audio"
assert response.json()["mime_type"] == mime
@pytest.mark.parametrize("entries", [None, {}, "x", [{"path": "/etc/passwd", "role": "reference"}],
[{"asset_id": "x", "role": "unknown"}]])
def test_malformed_conditioning_is_rejected(entries):
with pytest.raises(ValueError):
resolve_generation_inputs({"generation_mode": "ref2va", "conditioning_assets": entries}, "full-h3")
@pytest.mark.parametrize("mode", ["t2va", "fl2va", "ref2va"])
def test_mock_streams_all_valid_modes_and_releases_assets(library, monkeypatch, mode):
from dreamverse import mock_server
_, client = library
monkeypatch.setattr(mock_server, "MOCK_SEGMENT_BYTES", b"mock-fmp4")
monkeypatch.setattr(mock_server, "LATENCY_MS", 1)
image = upload_image(client)
refs = [] if mode == "t2va" else [conditioning(image, "first_frame" if mode == "fl2va" else "reference")]
ws = _FakeWebSocket([
(0, {"type": "session_init_v2", "generation_mode": mode, "conditioning_assets": refs,
"curated_prompts": ["A bird flies over a lake."], "single_clip_mode": True,
"enhancement_enabled": False}),
(0.15, {"type": "leave"}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
assert not [event for event in ws.sent_json if event["type"] == "error"]
assert any(event["type"] == "media_segment_complete" for event in ws.sent_json)
assert ws.sent_bytes
assert client.delete(image["url"]).status_code == 204
def test_mock_rejects_invalid_mode_before_gpu_assignment(library):
from dreamverse import mock_server
ws = _FakeWebSocket([(0, {"type": "session_init_v2", "generation_mode": "fl2va"})])
asyncio.run(mock_server.websocket_endpoint(ws))
assert not any(event["type"] == "gpu_assigned" for event in ws.sent_json)
errors = [event for event in ws.sent_json if event["type"] == "error"]
assert errors[0]["error_code"] == "invalid_generation_input"
assert "first-frame" in errors[0]["message"]
@pytest.mark.skipif(not shutil.which("ffmpeg") or not shutil.which("ffprobe"), reason="ffmpeg + ffprobe required")
@pytest.mark.parametrize("kind,mime,suffix", [("video", "video/mp4", ".mp4"), ("audio", "audio/wav", ".wav")])
def test_actual_video_and_audio_upload_validation(library, tmp_path, kind, mime, suffix):
_, client = library
media_path = tmp_path / f"sample{suffix}"
source = "testsrc2=size=64x64:rate=24" if kind == "video" else "sine=frequency=440:sample_rate=24000"
command = [shutil.which("ffmpeg"), "-v", "error", "-f", "lavfi", "-i", source, "-t", "0.5", str(media_path)]
subprocess.run(command, check=True, capture_output=True, timeout=30)
response = client.post("/assets", content=media_path.read_bytes(), headers={"Content-Type": mime})
assert response.status_code == 201, response.text
assert response.json()["kind"] == kind
response = client.post("/assets", content=b"#EXTM3U\nhttp://example.com/stream", headers={"Content-Type": mime})
assert response.status_code == 400
@@ -0,0 +1,214 @@
"""CPU contract tests; fake executors do not validate generated-media quality."""
from __future__ import annotations
import importlib.util
import pickle
import sys
from pathlib import Path
from types import ModuleType, SimpleNamespace
from unittest.mock import Mock
import numpy as np
import pytest
from PIL import Image
from dreamverse.config import MODEL_REGISTRY
from dreamverse.generation_inputs import GenerationAsset, GenerationInputs
from dreamverse.generation_worker import VideoGenerationWorker
from dreamverse.minimax_h3_generation import MiniMaxH3GenerationBackend
from dreamverse.worker_ipc import UserStepPayload
@pytest.fixture
def fastvideo_api(monkeypatch):
"""Use the actual lightweight API schema with only GPU execution replaced."""
schema_path = Path(__file__).resolve().parents[4] / "fastvideo/api/schema.py"
spec = importlib.util.spec_from_file_location("dreamverse_test_api_schema", schema_path)
assert spec is not None and spec.loader is not None
schema = importlib.util.module_from_spec(spec)
monkeypatch.setitem(sys.modules, spec.name, schema)
spec.loader.exec_module(schema)
package = ModuleType("fastvideo")
package.__path__ = []
package.VideoGenerator = SimpleNamespace(from_config=Mock())
monkeypatch.setitem(sys.modules, "fastvideo", package)
monkeypatch.setitem(sys.modules, "fastvideo.api", schema)
schema.MiniMaxH3Reference = lambda **kwargs: SimpleNamespace(**kwargs)
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.empty_cache", lambda: None)
return package.VideoGenerator.from_config
class RecordingGenerator:
def __init__(self):
self.requests = []
self.images = []
self.closed = False
def shutdown(self):
self.closed = True
def generate(self, request):
self.requests.append(request)
self.images.append(tuple(None if image is None else np.asarray(image).copy()
for image in (request.inputs.pil_image, request.inputs.last_image)))
return SimpleNamespace(
frames=[np.full((2, 3, 3), 7, dtype=np.uint8), np.full((2, 3, 3), 29, dtype=np.uint8)],
audio=np.zeros((2, 16), dtype=np.float32),
audio_sample_rate=44100,
generation_time=0.1,
)
def prepared_backend(monkeypatch):
backend = MiniMaxH3GenerationBackend(0)
backend.model_config = dict(MODEL_REGISTRY["full-h3"])
backend.generator = RecordingGenerator()
monkeypatch.setattr(backend, "_gpu_mem", lambda: "fake executor")
return backend
def test_ipc_preserves_immutable_ordered_references():
inputs = GenerationInputs("ref2va", (
GenerationAsset("second", "video", "/assets/second.mp4", "reference"),
GenerationAsset("first", "image", "/assets/first.png", "reference"),
))
payload = UserStepPayload("follow the references", 1, None, True, inputs)
restored = pickle.loads(pickle.dumps(payload))
assert restored == payload
assert [asset.asset_id for asset in restored.generation_inputs.references] == ["second", "first"]
def test_worker_passes_conditioning_to_selected_backend():
inputs = GenerationInputs("t2va")
worker = VideoGenerationWorker(0)
worker.backend = Mock()
worker.generate_step("prompt", 1, None, True, inputs)
worker.backend.generate_step.assert_called_once_with("prompt", 1, None, True, generation_inputs=inputs)
def test_full_h3_uses_full_weights_without_preview_lora(monkeypatch, fastvideo_api):
backend = prepared_backend(monkeypatch)
old_generator = backend.generator
fastvideo_api.return_value = RecordingGenerator()
monkeypatch.setattr("dreamverse.minimax_h3_generation.DREAMVERSE_SP_SIZE", 4)
backend.initialize(MODEL_REGISTRY["full-h3"])
config = fastvideo_api.call_args.args[0]
assert old_generator.closed
assert config.pipeline.components.lora_path is None
assert config.pipeline.components.override_pipeline_cls_name is None
assert config.engine.use_fsdp_inference
assert config.engine.num_gpus == 4
assert not config.pipeline.experimental["inference_torch_compile"]
def test_fl2va_maps_endpoints_only_on_initial_segment(monkeypatch, fastvideo_api, tmp_path):
first = tmp_path / "first.png"
last = tmp_path / "last.png"
Image.new("RGB", (3, 2), (10, 20, 30)).save(first)
Image.new("RGB", (3, 2), (40, 50, 60)).save(last)
inputs = GenerationInputs("fl2va", (
GenerationAsset("first", "image", str(first), "first_frame"),
GenerationAsset("last", "image", str(last), "last_frame"),
))
backend = prepared_backend(monkeypatch)
first_result = backend.generate_step("first", 1, None, True, inputs)
later_result = backend.generate_step("later", 2, None, False, inputs)
assert backend.generator.images[0][0][0, 0].tolist() == [10, 20, 30]
assert backend.generator.images[0][1][0, 0].tolist() == [40, 50, 60]
assert backend.generator.images[1][0][0, 0].tolist() == [29, 29, 29]
assert backend.generator.images[1][1] is None
assert first_result.head_trim_frames == 0
assert later_result.head_trim_frames == 1
assert backend.generator.requests[0].sampling.num_inference_steps == 50
def test_ref2va_switches_pipeline_and_preserves_reference_order(monkeypatch, fastvideo_api):
inputs = GenerationInputs("ref2va", (
GenerationAsset("video", "video", "/assets/reference.mp4", "reference"),
GenerationAsset("audio", "audio", "/assets/reference.wav", "reference"),
GenerationAsset("image", "image", "/assets/reference.png", "reference"),
))
backend = prepared_backend(monkeypatch)
base_generator = backend.generator
reference_generator = RecordingGenerator()
def load(config):
assert base_generator.closed, "Old executor must release memory before loading reference weights"
assert config.pipeline.components.override_pipeline_cls_name == "MiniMaxH3Ref2VAModularPipeline"
assert config.pipeline.workload_type == "i2v"
assert config.pipeline.components.lora_path is None
return reference_generator
fastvideo_api.side_effect = load
backend.generate_step("first", 1, None, True, inputs)
result = backend.generate_step("second", 2, None, False, inputs)
assert fastvideo_api.call_count == 1
for request in reference_generator.requests:
assert [(reference.media_type, reference.source) for reference in request.inputs.references] == [
("video", "/assets/reference.mp4"), ("audio", "/assets/reference.wav"), ("image", "/assets/reference.png")
]
assert request.inputs.pil_image is None
assert request.inputs.last_image is None
assert result.head_trim_frames == result.head_trim_audio_frames == 0
assert backend.continuation_image is None
fastvideo_api.side_effect = None
fastvideo_api.return_value = RecordingGenerator()
backend.generate_step("new project", 1, None, True, GenerationInputs("t2va"))
assert reference_generator.closed
config = fastvideo_api.call_args.args[0]
assert config.pipeline.components.override_pipeline_cls_name is None
assert backend.pipeline_mode == "base"
def test_ref2va_pipeline_switch_failure_drops_unloaded_executor(monkeypatch, fastvideo_api):
backend = prepared_backend(monkeypatch)
old_generator = backend.generator
fastvideo_api.side_effect = RuntimeError("checkpoint unavailable")
with pytest.raises(RuntimeError, match="checkpoint unavailable"):
backend.generate_step("prompt", 1, None, True, GenerationInputs("ref2va"))
assert old_generator.closed
assert backend.generator is None
def test_failed_pipeline_switch_reloads_on_the_next_step(monkeypatch, fastvideo_api):
"""A failed base<->ref2va switch must not strand the slot for later steps."""
backend = prepared_backend(monkeypatch)
fastvideo_api.side_effect = RuntimeError("checkpoint unavailable")
with pytest.raises(RuntimeError, match="checkpoint unavailable"):
backend.generate_step("prompt", 1, None, True, GenerationInputs("ref2va"))
fastvideo_api.side_effect = None
fastvideo_api.return_value = RecordingGenerator()
backend.generate_step("retry", 1, None, True, GenerationInputs("ref2va"))
assert fastvideo_api.call_count == 2
assert backend.pipeline_mode == "ref2va"
assert backend.generator is not None
def test_mode_cannot_switch_mid_project(monkeypatch, fastvideo_api):
backend = prepared_backend(monkeypatch)
with pytest.raises(ValueError, match="middle of a project"):
backend.generate_step("prompt", 2, None, False, GenerationInputs("ref2va"))
fastvideo_api.assert_not_called()
@pytest.mark.parametrize("mode", ["fl2va", "ref2va"])
def test_preview_rejects_unsupported_generation_modes(monkeypatch, fastvideo_api, mode):
backend = prepared_backend(monkeypatch)
backend.model_config = dict(MODEL_REGISTRY["fast-h3"])
with pytest.raises(ValueError, match="full-h3"):
backend.generate_step("prompt", 1, None, True, GenerationInputs(mode))
assert backend.generator.requests == []
def test_legacy_h3_continuation_is_preserved(monkeypatch, fastvideo_api):
backend = prepared_backend(monkeypatch)
backend.generate_step("first", 1, None, False)
result = backend.generate_step("second", 2, None, False)
assert backend.generator.requests[0].inputs.pil_image is None
assert backend.generator.images[1][0][0, 0].tolist() == [29, 29, 29]
assert result.head_trim_frames == 1
@@ -0,0 +1,278 @@
"""Session-mode validation and IPC handoff without a GPU worker process."""
from __future__ import annotations
import asyncio
import importlib.util
import sys
from pathlib import Path
from types import ModuleType, SimpleNamespace
from unittest.mock import Mock
import pytest
from dreamverse.generation_inputs import GenerationAsset, GenerationInputs
from dreamverse.worker_ipc import MediaChunk, MediaComplete, MediaInit
@pytest.fixture
def controller_module(monkeypatch):
gpu_pool = ModuleType("dreamverse.gpu_pool")
gpu_pool.GPUSlot = object
monkeypatch.setitem(sys.modules, "dreamverse.gpu_pool", gpu_pool)
path = Path(__file__).resolve().parents[1] / "session/controller.py"
spec = importlib.util.spec_from_file_location("dreamverse_test_session_controller", path)
assert spec is not None and spec.loader is not None
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
monkeypatch.setattr(module, "ACTIVE_MODEL_ID", "full-h3")
monkeypatch.setattr(module, "pin_generation_inputs", Mock())
monkeypatch.setattr(module, "release_generation_inputs", Mock())
return module
class Socket:
def __init__(self):
self.incoming = asyncio.Queue()
self.outgoing = asyncio.Queue()
self.messages = []
self.closed = False
async def accept(self):
pass
async def receive_json(self):
return await self.incoming.get()
async def send_json(self, payload):
self.messages.append(payload)
await self.outgoing.put(payload)
async def send_bytes(self, payload):
pass
async def close(self, **kwargs):
self.closed = True
async def wait_for(self, kind):
while True:
payload = await asyncio.wait_for(self.outgoing.get(), 3)
if payload["type"] == kind:
return payload
class Slot:
def __init__(self):
self.shared_stream_buffer = None
self.queue = asyncio.Queue()
self.calls = []
async def join_user(self, *args, **kwargs):
pass
def register_stream_queue(self, client_id):
return self.queue
def unregister_stream_queue(self, client_id):
pass
async def user_step(self, client_id, **kwargs):
self.calls.append(kwargs)
segment_idx = kwargs["segment_idx"]
await self.queue.put(MediaInit(client_id, segment_idx, "test", "video/mp4", False))
await self.queue.put(MediaChunk(client_id, segment_idx, "test", chunk=b"test"))
await self.queue.put(MediaComplete(client_id, segment_idx, "test", 1))
return {"e2e_latency_ms": 1.0}
class Pool:
def __init__(self):
self.slot = Slot()
self.acquire_count = 0
def get_status(self):
return {"queue_size": 0, "available_gpus": 1, "total_gpus": 1}
async def acquire(self, *args):
self.acquire_count += 1
return 0, self.slot
async def release(self, *args):
pass
def start_controller(module, socket, pool):
enhancer = SimpleNamespace(
resolve_rewrite_model=lambda value: "test-model",
resolve_rewrite_system_prompt=lambda value: "test-system",
resolve_rewrite_temperature=lambda value: 1.0,
)
controller = module.SessionController(socket, pool, enhancer, None, None)
return asyncio.create_task(controller.run())
def test_invalid_initial_mode_does_not_acquire_gpu(controller_module):
async def scenario():
socket, pool = Socket(), Pool()
await socket.incoming.put({"type": "session_init_v2", "generation_mode": "unknown"})
await asyncio.wait_for(start_controller(controller_module, socket, pool), 3)
error = next(message for message in socket.messages if message["type"] == "error")
assert error["error_code"] == "invalid_generation_input"
assert pool.acquire_count == 0
assert socket.closed
controller_module.pin_generation_inputs.assert_not_called()
asyncio.run(scenario())
def test_new_project_replaces_conditioning_and_passes_it_to_gpu(controller_module, monkeypatch):
first = GenerationInputs("t2va")
second = GenerationInputs("fl2va", (GenerationAsset("first", "image", "/assets/first.png", "first_frame"),))
monkeypatch.setattr(controller_module, "resolve_generation_inputs", Mock(side_effect=[first, second]))
async def scenario():
socket, pool = Socket(), Pool()
await socket.incoming.put({
"type": "session_init_v2", "generation_mode": "t2va", "curated_prompts": ["first prompt"],
"enhancement_enabled": False,
})
task = start_controller(controller_module, socket, pool)
try:
await socket.wait_for("media_segment_complete")
await socket.incoming.put({"type": "end_project_keep_session"})
await socket.wait_for("project_idle")
assert first in [call.args[0] for call in controller_module.release_generation_inputs.call_args_list]
await socket.incoming.put({
"type": "project_init_v1", "generation_mode": "fl2va", "curated_prompts": ["second prompt"],
"enhancement_enabled": False,
})
await socket.wait_for("media_segment_complete")
assert [call["generation_inputs"] for call in pool.slot.calls] == [first, second]
assert pool.slot.calls[1]["segment_idx"] == 1
assert pool.slot.calls[1]["reset_conditioning"]
await socket.incoming.put({"type": "leave"})
await asyncio.wait_for(task, 3)
finally:
if not task.done():
task.cancel()
await asyncio.gather(task, return_exceptions=True)
assert [call.args[0] for call in controller_module.pin_generation_inputs.call_args_list] == [first, second]
assert second in [call.args[0] for call in controller_module.release_generation_inputs.call_args_list]
asyncio.run(scenario())
@pytest.mark.parametrize("injection", [
{"initial_image": {"data_url": "not allowed"}},
{"generation_mode": "ref2va"},
{"conditioning_assets": []},
])
def test_simple_generate_cannot_replace_locked_inputs(controller_module, injection):
async def scenario():
socket, pool = Socket(), Pool()
await socket.incoming.put({
"type": "session_init_v2", "generation_mode": "t2va", "single_clip_mode": True,
"enhancement_enabled": False,
})
task = start_controller(controller_module, socket, pool)
try:
await socket.wait_for("gpu_assigned")
await socket.incoming.put({"type": "simple_generate", "prompt": "prompt", **injection})
error = await socket.wait_for("error")
assert error["error_code"] == "invalid_generation_input"
assert "project" in error["message"]
assert pool.slot.calls == []
await socket.incoming.put({"type": "leave"})
await asyncio.wait_for(task, 3)
finally:
if not task.done():
task.cancel()
await asyncio.gather(task, return_exceptions=True)
asyncio.run(scenario())
def test_disconnect_waits_for_worker_before_releasing_assets(controller_module, monkeypatch):
async def scenario():
socket, pool = Socket(), Pool()
worker_started = asyncio.Event()
worker_finished = asyncio.Event()
proceed = asyncio.Event()
async def slow_step(client_id, **kwargs):
worker_started.set()
await proceed.wait()
worker_finished.set()
return {"e2e_latency_ms": 1.0}
pool.slot.user_step = slow_step
await socket.incoming.put({
"type": "session_init_v2", "generation_mode": "t2va", "curated_prompts": ["prompt"],
"enhancement_enabled": False,
})
task = start_controller(controller_module, socket, pool)
try:
await asyncio.wait_for(worker_started.wait(), 3)
await socket.incoming.put({"type": "leave"})
await asyncio.sleep(0.07)
assert not task.done()
controller_module.release_generation_inputs.assert_not_called()
proceed.set()
await asyncio.wait_for(task, 3)
assert worker_finished.is_set()
controller_module.release_generation_inputs.assert_called_once()
finally:
if not task.done():
task.cancel()
await asyncio.gather(task, return_exceptions=True)
asyncio.run(scenario())
@pytest.fixture
def gpu_pool_module(monkeypatch):
streaming = ModuleType("dreamverse.av_streaming")
for name in ("StreamChunk", "StreamComplete", "StreamEvent", "StreamInit", "generate_stream_id", "stream_fmp4"):
setattr(streaming, name, object)
streaming.SHARED_STREAM_BUFFER_BYTES = 1024
streaming.USE_SHARED_STREAM_BUFFER = False
monkeypatch.setitem(sys.modules, "dreamverse.av_streaming", streaming)
path = Path(__file__).resolve().parents[1] / "gpu_pool.py"
spec = importlib.util.spec_from_file_location("dreamverse_test_gpu_pool", path)
assert spec is not None and spec.loader is not None
module = importlib.util.module_from_spec(spec)
monkeypatch.setitem(sys.modules, spec.name, module)
spec.loader.exec_module(module)
monkeypatch.setattr(module, "pin_generation_inputs", Mock())
monkeypatch.setattr(module, "release_generation_inputs", Mock())
return module
def test_gpu_step_timeout_keeps_assets_pinned_until_late_worker_completion(gpu_pool_module):
from dreamverse.worker_ipc import StepComplete
async def scenario():
slot = gpu_pool_module.GPUSlot(0, "0")
inputs = GenerationInputs("ref2va", (GenerationAsset("ref", "image", "/assets/ref.png", "reference"),))
async def timeout(command, timeout):
assert command.payload.generation_inputs == inputs
raise asyncio.TimeoutError
slot._send_command_tagged = timeout
with pytest.raises(asyncio.TimeoutError):
await slot.user_step("user", "prompt", generation_inputs=inputs)
gpu_pool_module.pin_generation_inputs.assert_called_once_with(inputs)
gpu_pool_module.release_generation_inputs.assert_not_called()
def late_response(timeout):
slot._active = False
return StepComplete("user", 1, {})
slot.response_queue = SimpleNamespace(get=late_response)
slot._active = True
await slot._response_reader()
gpu_pool_module.release_generation_inputs.assert_called_once_with(inputs)
assert slot._step_asset_inputs == {}
asyncio.run(scenario())
@@ -0,0 +1,18 @@
from dreamverse.generation_worker import _create_generation_backend
from dreamverse.ltx2_generation import LTX2GenerationBackend
def test_create_generation_backend_ltx2_module_import():
backend = _create_generation_backend("ltx2", gpu_id=3)
assert isinstance(backend, LTX2GenerationBackend)
assert backend.gpu_id == 3
def test_create_generation_backend_cosmos25_dfd_module_import():
from dreamverse.cosmos25_dfd_generation import Cosmos25DFDGenerationBackend
backend = _create_generation_backend("cosmos25_dfd", gpu_id=2)
assert isinstance(backend, Cosmos25DFDGenerationBackend)
assert backend.gpu_id == 2
@@ -7,7 +7,6 @@ from types import SimpleNamespace
import pytest
import dreamverse.gpu_pool as gpu_pool
@@ -64,6 +63,14 @@ def test_get_available_gpus_defaults_to_first_visible_device(monkeypatch):
assert gpu_pool.get_available_gpus() == [3]
def test_get_available_gpus_defaults_to_active_model_sequence_parallel_size(monkeypatch):
monkeypatch.setenv("CUDA_VISIBLE_DEVICES", "0,1,2,3,4")
monkeypatch.delenv("FASTVIDEO_GPU_COUNT", raising=False)
monkeypatch.setattr(gpu_pool, "DREAMVERSE_SP_SIZE", 4)
assert gpu_pool.get_available_gpus() == [0, 1, 2, 3]
def test_get_available_gpus_rejects_invalid_gpu_count(monkeypatch):
monkeypatch.delenv("CUDA_VISIBLE_DEVICES", raising=False)
monkeypatch.setenv("FASTVIDEO_GPU_COUNT", "zero")
@@ -72,6 +79,23 @@ def test_get_available_gpus_rejects_invalid_gpu_count(monkeypatch):
gpu_pool.get_available_gpus()
def test_join_user_failed_reload_marks_model_uninitialized(monkeypatch):
"""A failed model reload forces the next join to reload a model."""
slot = gpu_pool.GPUSlot(gpu_id=0, cuda_device="0")
slot.current_model_id = "fast-ltx2"
async def fake_send_command(command, timeout):
del command, timeout
return gpu_pool.WorkerError(user_id="__reload__", message="load failed")
monkeypatch.setattr(slot, "_send_command", fake_send_command)
with pytest.raises(RuntimeError, match="Model reload failed"):
asyncio.run(slot.join_user("client-id", model_id="fast-h3"))
assert slot.current_model_id is None
def test_send_command_raises_on_worker_death():
"""A worker that consumes a command and exits without replying must
surface as RuntimeError via sentinel detection, not after the long
@@ -85,9 +109,7 @@ def test_send_command_raises_on_worker_death():
cmd_q = ctx.Queue()
resp_q = ctx.Queue()
proc = ctx.Process(
target=_child_consume_and_exit, args=(cmd_q, resp_q)
)
proc = ctx.Process(target=_child_consume_and_exit, args=(cmd_q, resp_q))
proc.start()
# Wait for the spawn child to fully boot. Allow generous time —
@@ -95,9 +117,9 @@ def test_send_command_raises_on_worker_death():
ready = resp_q.get(timeout=30.0)
assert ready == "READY"
async def runner():
async def runner() -> None:
slot = gpu_pool.GPUSlot(gpu_id=0, cuda_device="0")
slot.process = proc
slot.process = proc # type: ignore[assignment]
slot.command_queue = cmd_q
slot.response_queue = resp_q
@@ -1,4 +1,6 @@
import ast
import subprocess
import sys
from pathlib import Path
ALLOWED_PREFIXES = (
@@ -7,7 +9,7 @@ ALLOWED_PREFIXES = (
"fastvideo.entrypoints.video_generator",
"fastvideo.configs",
)
ALLOWED_EXACT = ("fastvideo",)
ALLOWED_EXACT = ("fastvideo", )
FORBIDDEN_PREFIXES = (
"fastvideo.pipelines",
"fastvideo.models",
@@ -17,11 +19,11 @@ FORBIDDEN_PREFIXES = (
)
ALLOWED_INTERNAL_IMPORTS = {
(
"video_generation.py",
"ltx2_generation.py",
"fastvideo.models.audio.ltx2_audio_processing",
),
(
"video_generation.py",
"ltx2_generation.py",
"fastvideo.models.loader.component_loader",
),
}
@@ -38,19 +40,49 @@ def test_dreamverse_server_imports_only_public_fastvideo_surfaces() -> None:
except SyntaxError as task_exc:
raise AssertionError(f"Failed to parse {path}") from task_exc
for node in ast.walk(tree):
names = (
[a.name for a in node.names] if isinstance(node, ast.Import)
else [node.module] if isinstance(node, ast.ImportFrom) and node.module
else []
)
names = ([a.name for a in node.names] if isinstance(node, ast.Import) else
[node.module] if isinstance(node, ast.ImportFrom) and node.module else [])
for name in names:
if not name:
continue
rel_path = str(path.relative_to(root))
if (
name.startswith(FORBIDDEN_PREFIXES)
and (rel_path, name) not in ALLOWED_INTERNAL_IMPORTS
):
if (name.startswith(FORBIDDEN_PREFIXES) and (rel_path, name) not in ALLOWED_INTERNAL_IMPORTS):
bad.append((str(path.relative_to(root)), getattr(node, "lineno", 0), name))
assert bad == [], f"Forbidden internal imports: {bad}"
def test_h3_reference_public_export_is_lazy_and_preserves_type_identity() -> None:
"""Only explicit reference usage should load H3's optional GPU dependencies."""
repo_root = Path(__file__).resolve().parents[4]
# Isolate the import graph: keep the real public API implementation/schema,
# substituting only the unrelated legacy sampling module and heavy H3 leaf.
script = r'''
import importlib
import sys
from pathlib import Path
from types import ModuleType
root = Path(sys.argv[1])
fastvideo = ModuleType("fastvideo")
fastvideo.__path__ = [str(root / "fastvideo")]
sys.modules["fastvideo"] = fastvideo
sampling = ModuleType("fastvideo.api.sampling_param")
sampling.SamplingParam = type("SamplingParam", (), {})
sys.modules[sampling.__name__] = sampling
api = importlib.import_module("fastvideo.api")
assert "MiniMaxH3Reference" in api.__all__
assert "MiniMaxH3Reference" not in vars(api)
assert not any(name.startswith("fastvideo.pipelines") for name in sys.modules)
internal = ModuleType("fastvideo.pipelines.basic.minimax_h3.reference")
internal.MiniMaxH3Reference = type("MiniMaxH3Reference", (), {})
sys.modules[internal.__name__] = internal
from fastvideo.api import MiniMaxH3Reference
assert MiniMaxH3Reference is internal.MiniMaxH3Reference
assert api.MiniMaxH3Reference is internal.MiniMaxH3Reference
assert not hasattr(api, "UnknownReference")
'''
result = subprocess.run([sys.executable, "-c", script, str(repo_root)], capture_output=True, text=True, timeout=30)
assert result.returncode == 0, result.stderr
@@ -0,0 +1,247 @@
from __future__ import annotations
import os
from types import SimpleNamespace
from typing import Any
import numpy as np
import pytest
import dreamverse.generation_worker as generation_worker
from dreamverse.minimax_h3_generation import MiniMaxH3GenerationBackend
FASTH3_MODEL_CONFIG = {
"name": "FastH3",
"generation_backend": "minimax_h3",
"default_sp_size": 4,
"model_path": "MiniMaxAI/MiniMax-H3",
"adapter_repo": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
"adapter_filename": "vsa-datafree/adapter_model.safetensors",
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
"height": 768,
"width": 1344,
"num_frames": 124,
"num_inference_steps": 5,
"seed": 1000,
}
class _RecordingGenerator:
"""Record typed requests and return small synchronized media fixtures."""
def __init__(self) -> None:
self.requests: list[Any] = []
self.conditioning_pixels: list[np.ndarray | None] = []
def generate(self, request):
"""Capture the request and return two tiny video frames with audio."""
self.requests.append(request)
conditioning_image = request.inputs.pil_image
self.conditioning_pixels.append(
None if conditioning_image is None else np.asarray(conditioning_image).copy())
frames = [
np.full((2, 3, 3), 10, dtype=np.uint8),
np.full((2, 3, 3), 20, dtype=np.uint8),
]
return SimpleNamespace(
frames=frames,
audio=np.zeros((2, 16), dtype=np.float32),
audio_sample_rate=44100,
generation_time=0.25,
)
def test_initialize_builds_vsa_datafree_fasth3_generator(monkeypatch):
"""Initialization translates the DreamVerse profile into typed FastVideo config."""
from fastvideo import VideoGenerator
captured = {}
fake_generator = SimpleNamespace(shutdown=lambda: None)
def fake_from_config(config):
captured["config"] = config
return fake_generator
def fake_download(**kwargs):
captured["download"] = kwargs
return f"/models/{kwargs['filename']}"
monkeypatch.setattr("huggingface_hub.hf_hub_download", fake_download)
monkeypatch.setattr(VideoGenerator, "from_config", fake_from_config)
monkeypatch.setattr("dreamverse.minimax_h3_generation.DREAMVERSE_SP_SIZE", 4)
monkeypatch.setenv("FASTVIDEO_ATTENTION_BACKEND", "test-attention")
monkeypatch.setenv("FASTVIDEO_FA4", "0")
monkeypatch.setenv("FASTVIDEO_MINIMAX_H3_FUSIONS", "0")
monkeypatch.setenv("FASTVIDEO_VSA_SM100A", "1")
monkeypatch.setenv("FASTVIDEO_INFERENCE_TORCH_COMPILE", "1")
backend = MiniMaxH3GenerationBackend(gpu_id=0)
monkeypatch.setattr(backend, "_gpu_mem", lambda: "alloc=0.00GiB, reserved=0.00GiB")
backend.initialize(FASTH3_MODEL_CONFIG)
config = captured["config"]
assert captured["download"] == {
"repo_id": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-LoRA",
"filename": "vsa-datafree/adapter_model.safetensors",
}
assert config.model_path == "MiniMaxAI/MiniMax-H3"
assert config.pipeline.components.lora_path.endswith("vsa-datafree/adapter_model.safetensors")
assert config.pipeline.components.lora_strength == 1.0
assert config.pipeline.experimental == {
"attention_backend": "VIDEO_SPARSE_ATTN_H3",
"inference_torch_compile": False,
"vae_parallel_decode": True,
"vae_parallel_decode_strategy": "gather",
"VSA_sparsity": 0.9,
"VSA_tile_size": 64,
}
assert config.engine.num_gpus == 4
assert config.engine.parallelism.tp_size == 1
assert config.engine.parallelism.sp_size == 4
assert config.engine.offload.dit is False
assert config.engine.offload.dit_layerwise is False
assert config.engine.offload.text_encoder is True
assert config.engine.offload.vae is True
assert config.engine.compile.vae_enabled is True
assert config.engine.use_fsdp_inference is False
assert os.environ["FASTVIDEO_ATTENTION_BACKEND"] == "VIDEO_SPARSE_ATTN_H3"
assert os.environ["FASTVIDEO_FA4"] == "1"
assert os.environ["FASTVIDEO_MINIMAX_H3_FUSIONS"] == "all"
assert os.environ["FASTVIDEO_VSA_SM100A"] == "0"
assert "FASTVIDEO_INFERENCE_TORCH_COMPILE" not in os.environ
def test_initialize_selects_declared_generation_backend(monkeypatch):
"""The GPU worker constructs the backend that the active model profile declares."""
from unittest.mock import Mock
selected_backend = Mock()
monkeypatch.setattr(
generation_worker,
"_create_generation_backend",
lambda backend_name, gpu_id: selected_backend,
)
worker = generation_worker.VideoGenerationWorker(gpu_id=3)
worker.initialize(FASTH3_MODEL_CONFIG)
assert worker.backend_name == "minimax_h3"
assert worker.backend is selected_backend
selected_backend.initialize.assert_called_once_with(FASTH3_MODEL_CONFIG)
def test_initialize_failure_clears_backend_ownership(monkeypatch):
"""A failed family change leaves the GPU worker explicitly uninitialized."""
ltx_backend = SimpleNamespace(initialize=lambda config: None, shutdown=lambda: None)
def fail_initialize(config):
del config
raise RuntimeError("load failed")
fasth3_backend = SimpleNamespace(
initialize=fail_initialize,
shutdown=lambda: None,
)
backends = {
"ltx2": ltx_backend,
"minimax_h3": fasth3_backend,
}
monkeypatch.setattr(
generation_worker,
"_create_generation_backend",
lambda backend_name, gpu_id: backends[backend_name],
)
worker = generation_worker.VideoGenerationWorker(gpu_id=3)
worker.initialize({"generation_backend": "ltx2"})
with pytest.raises(RuntimeError, match="load failed"):
worker.initialize(FASTH3_MODEL_CONFIG)
assert worker.backend is None
assert worker.backend_name is None
assert worker.model_config == {"generation_backend": "ltx2"}
def test_generate_step_uses_last_frame_for_continuation(monkeypatch):
"""A later segment receives the prior segment's last decoded frame."""
backend = MiniMaxH3GenerationBackend(gpu_id=0)
backend.model_config = dict(FASTH3_MODEL_CONFIG)
backend.generator = _RecordingGenerator()
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
first_result = backend.generate_step(
"first prompt",
segment_idx=1,
image_path=None,
reset_conditioning=True,
)
second_result = backend.generate_step(
"second prompt",
segment_idx=2,
image_path=None,
reset_conditioning=False,
)
first_request = backend.generator.requests[0]
assert first_request.inputs.pil_image is None
assert first_request.negative_prompt == ""
assert first_request.sampling.height == 768
assert first_request.sampling.width == 1344
assert first_request.sampling.num_frames == 124
assert first_request.sampling.num_inference_steps == 5
assert first_request.sampling.fps == 24
assert first_request.sampling.guidance_scale == 1.0
assert first_request.sampling.batch_cfg is False
assert first_request.sampling.seed == 1000
assert first_request.output.save_video is False
assert first_request.output.return_frames is True
assert backend.generator.conditioning_pixels[1].tolist() == np.full((2, 3, 3), 20).tolist()
assert first_result.head_trim_frames == 0
assert first_result.head_trim_audio_frames == 0
assert second_result.head_trim_frames == 1
assert second_result.head_trim_audio_frames == 1
assert second_result.audio_sample_rate == 44100
def test_generate_step_reset_uses_text_to_video_path(monkeypatch):
"""Resetting continuation produces an unconditioned text-to-video request."""
backend = MiniMaxH3GenerationBackend(gpu_id=0)
backend.model_config = dict(FASTH3_MODEL_CONFIG)
backend.generator = _RecordingGenerator()
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
backend.generate_step("first prompt", 1, None, True)
reset_result = backend.generate_step("reset prompt", 2, None, True)
assert backend.generator.requests[-1].inputs.pil_image is None
assert reset_result.head_trim_frames == 0
assert reset_result.head_trim_audio_frames == 0
def test_generate_step_missing_continuation_frame(monkeypatch):
"""A later segment fails when no reset or retained frame defines its input."""
backend = MiniMaxH3GenerationBackend(gpu_id=0)
backend.model_config = dict(FASTH3_MODEL_CONFIG)
backend.generator = _RecordingGenerator()
with pytest.raises(RuntimeError, match="requires a retained continuation frame"):
backend.generate_step("later prompt", 2, None, False)
assert backend.generator.requests == []
def test_warmup_exercises_text_and_first_frame_paths(monkeypatch):
"""Warmup covers both request shapes used by a DreamVerse session."""
backend = MiniMaxH3GenerationBackend(gpu_id=0)
backend.model_config = dict(FASTH3_MODEL_CONFIG)
backend.generator = _RecordingGenerator()
monkeypatch.setattr("dreamverse.minimax_h3_generation.torch.cuda.synchronize", lambda: None)
timings = backend.warmup("warmup prompt")
assert backend.generator.conditioning_pixels[0] is None
assert backend.generator.conditioning_pixels[1] is not None
assert backend.continuation_image is None
assert "warmup_text_to_video_ms" in timings
assert "warmup_first_frame_to_video_ms" in timings
@@ -6,7 +6,6 @@ import os
from fastapi import WebSocketDisconnect
os.environ.setdefault("CEREBRAS_API_KEY", "dummy")
os.environ.setdefault("GROQ_API_KEY", "dummy")
@@ -14,6 +13,7 @@ import dreamverse.mock_server as mock_server
class _FakeWebSocket:
def __init__(self, messages: list[tuple[float, dict[str, object]]]):
self._messages = messages
self._index = 0
@@ -49,34 +49,34 @@ def test_mock_server_matches_current_single5s_protocol():
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 1
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "simple_prompt_1",
"curated_prompts": ["selected prompt"],
"single_clip_mode": True,
"enhancement_enabled": False,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.01,
{
"type": "simple_generate",
"preset_id": "simple_custom_prompt",
"prompt_id": "simple_custom_prompt",
"prompt": "custom prompt",
"enhancement_enabled": True,
"initial_image": None,
},
),
(0.20, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "simple_prompt_1",
"curated_prompts": ["selected prompt"],
"single_clip_mode": True,
"enhancement_enabled": False,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.01,
{
"type": "simple_generate",
"preset_id": "simple_custom_prompt",
"prompt_id": "simple_custom_prompt",
"prompt": "custom prompt",
"enhancement_enabled": True,
"initial_image": None,
},
),
(0.20, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
@@ -92,24 +92,14 @@ def test_mock_server_matches_current_single5s_protocol():
assert message_types.count("ltx2_stream_complete") == 2
assert "prompt_sources_blocked" not in message_types
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
gpu_assigned_event = next(
payload for payload in ws.sent_json if payload["type"] == "gpu_assigned"
)
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
gpu_assigned_event = next(payload for payload in ws.sent_json if payload["type"] == "gpu_assigned")
assert gpu_assigned_event["session_timeout"] == mock_server.SESSION_TIMEOUT_SECONDS
assert [payload["segment_idx"] for payload in segment_start_events] == [1, 1]
assert segment_start_events[0]["prompt"] == "selected prompt"
assert segment_start_events[1]["prompt"] == "custom prompt"
step_complete_events = [
payload
for payload in ws.sent_json
if payload["type"] == "step_complete"
]
step_complete_events = [payload for payload in ws.sent_json if payload["type"] == "step_complete"]
assert len(step_complete_events) == 2
assert step_complete_events[0]["latency_ms"] == {
"total": 121.0,
@@ -134,29 +124,29 @@ def test_mock_server_regular_cap_waits_for_rewrite_rollout():
mock_server.LATENCY_MS = 1
mock_server.GENERATION_SEGMENT_CAP = 1
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "start a new rollout",
},
),
(0.20, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "start a new rollout",
},
),
(0.20, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
@@ -166,11 +156,7 @@ def test_mock_server_regular_cap_waits_for_rewrite_rollout():
assert "generation_cap_reached" not in message_types
assert "prompt_sources_blocked" not in message_types
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
assert [payload["segment_idx"] for payload in segment_start_events] == [1, 1]
assert segment_start_events[0]["prompt"] == "segment one"
assert segment_start_events[1]["prompt"] == "segment one [start a new rollout]"
@@ -187,54 +173,40 @@ def test_mock_server_rewrite_during_active_segment_restarts_from_first_rewritten
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 100
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one", "segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "restart from rewrite",
},
),
(0.40, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one", "segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(
0.02,
{
"type": "rewrite_seed_prompts",
"rewrite_instruction": "restart from rewrite",
},
),
(0.40, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
assert [payload["prompt"] for payload in segment_start_events[:2]] == [
"segment one",
"segment one [restart from rewrite]",
]
assert all(
payload["prompt"] != "segment two"
for payload in segment_start_events[1:]
)
reset_events = [
payload
for payload in ws.sent_json
if payload.get("type") == "seed_prompts_reset_applied"
]
assert any(
payload.get("reason") == "rewrite_during_generation"
for payload in reset_events
)
assert all(payload["prompt"] != "segment two" for payload in segment_start_events[1:])
reset_events = [payload for payload in ws.sent_json if payload.get("type") == "seed_prompts_reset_applied"]
assert any(payload.get("reason") == "rewrite_during_generation" for payload in reset_events)
finally:
mock_server.MOCK_SEGMENT_BYTES = old_segment_bytes
mock_server.LATENCY_MS = old_latency_ms
@@ -247,24 +219,24 @@ def test_mock_server_supports_initial_custom_rollout_prompt():
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 1
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "custom_editable",
"preset_label": "Custom rollout",
"curated_prompts": [],
"initial_rollout_prompt": "A moonbase corridor thriller with flooding",
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.20, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "custom_editable",
"preset_label": "Custom rollout",
"curated_prompts": [],
"initial_rollout_prompt": "A moonbase corridor thriller with flooding",
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.20, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
@@ -276,15 +248,9 @@ def test_mock_server_supports_initial_custom_rollout_prompt():
assert "ltx2_stream_start" in message_types
assert "prompt_sources_blocked" not in message_types
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
assert segment_start_events
assert segment_start_events[0]["prompt"] == (
"A moonbase corridor thriller with flooding [segment 1]"
)
assert segment_start_events[0]["prompt"] == ("A moonbase corridor thriller with flooding [segment 1]")
finally:
mock_server.MOCK_SEGMENT_BYTES = old_segment_bytes
mock_server.LATENCY_MS = old_latency_ms
@@ -297,35 +263,37 @@ def test_mock_server_can_start_new_project_without_reconnecting():
mock_server.MOCK_SEGMENT_BYTES = b"mock-fmp4-bytes"
mock_server.LATENCY_MS = 40
ws = _FakeWebSocket(
[
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.02, {"type": "end_project_keep_session"}),
(
0.20,
{
"type": "project_init_v1",
"preset_id": "test_preset_2",
"preset_label": "Test Preset 2",
"curated_prompts": ["segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.40, {"type": "leave"}),
]
)
ws = _FakeWebSocket([
(
0.0,
{
"type": "session_init_v2",
"preset_id": "test_preset",
"curated_prompts": ["segment one"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.02, {
"type": "end_project_keep_session"
}),
(
0.20,
{
"type": "project_init_v1",
"preset_id": "test_preset_2",
"preset_label": "Test Preset 2",
"curated_prompts": ["segment two"],
"enhancement_enabled": True,
"auto_extension_enabled": False,
"loop_generation_enabled": False,
},
),
(0.40, {
"type": "leave"
}),
])
asyncio.run(mock_server.websocket_endpoint(ws))
@@ -336,16 +304,11 @@ def test_mock_server_can_start_new_project_without_reconnecting():
project_idle_index = message_types.index("project_idle")
stream_start_indexes = [
index for index, message_type in enumerate(message_types)
if message_type == "ltx2_stream_start"
index for index, message_type in enumerate(message_types) if message_type == "ltx2_stream_start"
]
assert stream_start_indexes[0] < project_idle_index < stream_start_indexes[1]
segment_start_events = [
payload
for payload in ws.sent_json
if payload["type"] == "ltx2_segment_start"
]
segment_start_events = [payload for payload in ws.sent_json if payload["type"] == "ltx2_segment_start"]
assert [payload["prompt"] for payload in segment_start_events[:2]] == [
"segment one",
"segment two",
@@ -6,7 +6,6 @@ import os
import re
import time
os.environ.setdefault("CEREBRAS_API_KEY", "dummy")
os.environ.setdefault("GROQ_API_KEY", "dummy")
@@ -22,6 +21,7 @@ from dreamverse.prompt_enhancer import (
class _FakeResponse:
def __init__(self, payload: dict):
self._payload = payload
@@ -30,6 +30,7 @@ class _FakeResponse:
class _FakeSyncCompletions:
def __init__(self, payload: dict):
self._payload = payload
@@ -38,6 +39,7 @@ class _FakeSyncCompletions:
class _FakeSyncClient:
def __init__(self, payload: dict):
self.chat = type(
"_FakeChat",
@@ -47,6 +49,7 @@ class _FakeSyncClient:
class _DelayedSyncCompletions:
def __init__(self, payload: dict, delay_s: float = 0.0, exc: Exception | None = None):
self._payload = payload
self._delay_s = delay_s
@@ -61,29 +64,26 @@ class _DelayedSyncCompletions:
class _DelayedSyncClient:
def __init__(self, payload: dict, delay_s: float = 0.0, exc: Exception | None = None):
self.chat = type(
"_FakeChat",
(),
{
"completions": _DelayedSyncCompletions(
payload,
delay_s=delay_s,
exc=exc,
)
},
{"completions": _DelayedSyncCompletions(
payload,
delay_s=delay_s,
exc=exc,
)},
)()
def _chat_payload_with_content(content: str) -> dict:
return {
"choices": [
{
"message": {
"content": content,
}
"choices": [{
"message": {
"content": content,
}
]
}]
}
@@ -172,6 +172,7 @@ def _build_staged_enhancer(
class _FakeOpenAIClient:
def __init__(self, **kwargs):
self.kwargs = kwargs
self.chat = type(
@@ -182,6 +183,7 @@ class _FakeOpenAIClient:
class _FakeCerebrasClient:
def __init__(self, **kwargs):
self.kwargs = kwargs
self.chat = type(
@@ -192,16 +194,12 @@ class _FakeCerebrasClient:
def test_parse_json_response_accepts_fenced_json_with_prose():
parsed = _parse_json_response(
"Here is the rewrite:\n```json\n{\"segment_prompts\":[\"A\",\"B\"]}\n```\nThanks."
)
parsed = _parse_json_response("Here is the rewrite:\n```json\n{\"segment_prompts\":[\"A\",\"B\"]}\n```\nThanks.")
assert parsed == {"segment_prompts": ["A", "B"]}
def test_parse_json_response_extracts_first_embedded_object():
parsed = _parse_json_response(
"Model output:\n{\"segment_prompts\":[\"A\",\"B\"]}\n(complete)"
)
parsed = _parse_json_response("Model output:\n{\"segment_prompts\":[\"A\",\"B\"]}\n(complete)")
assert parsed == {"segment_prompts": ["A", "B"]}
@@ -268,16 +266,12 @@ def test_build_client_supports_groq_provider(monkeypatch):
def test_rewrite_prompt_sequence_accepts_segment_prompts_output():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.rollout_id == "preset_a"
@@ -286,15 +280,12 @@ def test_rewrite_prompt_sequence_accepts_segment_prompts_output():
def test_rewrite_prompt_sequence_accepts_legacy_rewritten_prompts_output():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"rewritten_prompts":["A","B"]}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"rewritten_prompts":["A","B"]}'))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.rollout_id == "current_rollout"
@@ -303,19 +294,14 @@ def test_rewrite_prompt_sequence_accepts_legacy_rewritten_prompts_output():
def test_rewrite_prompt_sequence_accepts_segment_dicts_without_top_level_rollout_metadata():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"segments":[{"prompt":"A"},{"text":"B"}]}'
)
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"segments":[{"prompt":"A"},{"text":"B"}]}'))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
preset_id="preset_a",
preset_label="Preset A",
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.rollout_id == "preset_a"
@@ -329,14 +315,12 @@ def test_rewrite_prompt_sequence_accepts_numbered_prose_output():
"The user is asking for a cinematic rewrite.\n\n"
"1. A dog bounds across the moon's dusty surface, kicking up silver regolith as it chases a rabbit beneath the black sky.\n"
"2. The rabbit darts around a crater rim while the dog lunges after it, Earth glowing blue in the distance.\n"
)
)
))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.rollout_id == "current_rollout"
@@ -347,29 +331,27 @@ def test_rewrite_prompt_sequence_accepts_numbered_prose_output():
]
def test_enhance_prompt_prefers_cerebras_before_groq_fallback():
def test_enhance_prompt_uses_groq_when_it_returns_first():
enhancer = _build_staged_enhancer(
cerebras_payload=_chat_payload_with_content('{"prompt":"Cerebras prompt"}'),
groq_payload=_chat_payload_with_content('{"prompt":"Groq prompt"}'),
cerebras_delay_s=0.01,
cerebras_delay_s=0.08,
groq_delay_s=0.01,
)
result = asyncio.run(
enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
)
)
result = asyncio.run(enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
))
assert result.fallback_used is False
assert result.error is None
assert result.provider == "cerebras"
assert result.provider == "groq"
assert result.model == "gpt-test"
assert result.prompt == "Cerebras prompt"
assert result.prompt == "Groq prompt"
assert enhancer.get_provider_success_counts() == {
"cerebras": 1,
"groq": 0,
"cerebras": 0,
"groq": 1,
}
@@ -381,12 +363,10 @@ def test_enhance_prompt_uses_groq_when_cerebras_fails():
groq_delay_s=0.01,
)
result = asyncio.run(
enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
)
)
result = asyncio.run(enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
))
assert result.fallback_used is False
assert result.error is None
@@ -408,13 +388,11 @@ def test_enhance_prompt_can_use_groq_when_cerebras_times_out():
enhancer.http_timeout_ms = 50
enhancer.default_timeout_ms = 50
result = asyncio.run(
enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
timeout_ms=50,
)
)
result = asyncio.run(enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
timeout_ms=50,
))
assert result.fallback_used is False
assert result.error is None
@@ -434,12 +412,10 @@ def test_enhance_prompt_can_use_cerebras_when_it_returns_first():
groq_delay_s=0.08,
)
result = asyncio.run(
enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
)
)
result = asyncio.run(enhancer.enhance_prompt(
"A rainy alley at night",
mode="single_clip",
))
assert result.fallback_used is False
assert result.error is None
@@ -453,15 +429,12 @@ def test_enhance_prompt_can_use_cerebras_when_it_returns_first():
def test_rewrite_prompt_sequence_keeps_raw_output_on_parse_error():
enhancer = _build_test_enhancer(
_chat_payload_with_content("I cannot comply with JSON right now.")
)
enhancer = _build_test_enhancer(_chat_payload_with_content("I cannot comply with JSON right now."))
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is True
assert "No JSON object found in assistant response." in (result.error or "")
assert result.raw_response_text == "I cannot comply with JSON right now."
@@ -473,9 +446,7 @@ def test_rewrite_prompt_sequence_keeps_raw_output_on_parse_error():
def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}'
)
)
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}'))
captured = {
"body": None,
"timeout_seconds": None,
@@ -486,8 +457,7 @@ def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
captured["timeout_seconds"] = timeout_seconds
return (
_chat_payload_with_content(
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}'
),
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}'),
'{"id":"rewritten_rollout","label":"Rewritten Rollout","segment_prompts":["A","B"]}',
)
@@ -502,8 +472,7 @@ def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
rewrite_model="gpt-test",
rewrite_temperature=0.2,
timeout_ms=800,
)
)
))
assert result.fallback_used is False
assert captured["body"]["messages"][0] == {
@@ -512,12 +481,12 @@ def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
}
assert captured["body"]["messages"][1]["role"] == "user"
assert prompt_enhancer_module.json.loads(captured["body"]["messages"][1]["content"]) == {
"mode": "edit_existing_rollout",
"request": (
"Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical."
),
"user_instruction": "make it cinematic",
"mode":
"edit_existing_rollout",
"request": ("Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical."),
"user_instruction":
"make it cinematic",
"current_rollout": {
"id": "preset_a",
"label": "Preset A",
@@ -528,11 +497,8 @@ def test_rewrite_prompt_sequence_uses_current_rollout_payload_shape():
def test_rewrite_prompt_sequence_supports_new_rollout_mode():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}'
)
)
_chat_payload_with_content('{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}'))
captured = {
"body": None,
}
@@ -541,10 +507,8 @@ def test_rewrite_prompt_sequence_supports_new_rollout_mode():
del timeout_seconds
captured["body"] = body
return (
_chat_payload_with_content(
'{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}'
),
_chat_payload_with_content('{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}'),
'{"id":"custom_editable","label":"Custom rollout","segment_prompts":['
'"A","B","C","D","E","F"]}',
)
@@ -560,30 +524,29 @@ def test_rewrite_prompt_sequence_supports_new_rollout_mode():
rewrite_model="gpt-test",
rewrite_temperature=0.2,
timeout_ms=800,
)
)
))
assert result.fallback_used is False
assert result.prompts == ["A", "B", "C", "D", "E", "F"]
assert prompt_enhancer_module.json.loads(captured["body"]["messages"][1]["content"]) == {
"mode": "new_rollout",
"request": (
"Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical."
),
"user_instruction": "A moonbase corridor thriller with flooding and red alarms",
"desired_segment_count": 6,
"rollout_id_hint": "custom_editable",
"rollout_label_hint": "Custom rollout",
"mode":
"new_rollout",
"request": ("Rewrite all segment prompts with improved continuity and cinematic detail. "
"Keep count and ordering identical."),
"user_instruction":
"A moonbase corridor thriller with flooding and red alarms",
"desired_segment_count":
6,
"rollout_id_hint":
"custom_editable",
"rollout_label_hint":
"Custom rollout",
}
def test_rewrite_prompt_sequence_uses_session_override_system_prompt():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.rewrite_all_system_prompt = "shared system prompt"
captured = {
"body": None,
@@ -593,9 +556,7 @@ def test_rewrite_prompt_sequence_uses_session_override_system_prompt():
del timeout_seconds
captured["body"] = body
return (
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
),
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'),
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}',
)
@@ -609,8 +570,7 @@ def test_rewrite_prompt_sequence_uses_session_override_system_prompt():
rewrite_instruction="make it cinematic",
rewrite_model="gpt-test",
system_prompt_override="session specific system prompt",
)
)
))
assert result.fallback_used is False
assert captured["body"]["messages"][0] == {
@@ -621,10 +581,7 @@ def test_rewrite_prompt_sequence_uses_session_override_system_prompt():
def test_resolve_rewrite_new_rollout_system_prompt_uses_dedicated_prompt():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.rewrite_all_system_prompt = "shared rewrite system prompt"
enhancer.rewrite_user_system_prompt = "new rollout rewrite system prompt"
@@ -635,24 +592,17 @@ def test_resolve_rewrite_new_rollout_system_prompt_uses_dedicated_prompt():
def test_resolve_rewrite_new_rollout_system_prompt_prefers_override():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.rewrite_all_system_prompt = "shared rewrite system prompt"
enhancer.rewrite_user_system_prompt = "new rollout rewrite system prompt"
resolved = enhancer.resolve_rewrite_new_rollout_system_prompt(
"session specific system prompt"
)
resolved = enhancer.resolve_rewrite_new_rollout_system_prompt("session specific system prompt")
assert resolved == "session specific system prompt"
def test_generate_auto_prompt_uses_selected_model():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"next_prompt":"Auto next"}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"next_prompt":"Auto next"}'))
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test", "gpt-alt"]
enhancer.rewrite_default_model = "gpt-test"
@@ -680,8 +630,7 @@ def test_generate_auto_prompt_uses_selected_model():
next_segment_idx=2,
model="gpt-alt",
timeout_ms=800,
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.prompt == "Auto next"
@@ -690,9 +639,7 @@ def test_generate_auto_prompt_uses_selected_model():
def test_enhance_prompt_uses_selected_model():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"next_prompt":"Enhanced next"}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"next_prompt":"Enhanced next"}'))
enhancer.enhance_system_prompt = "enhance system prompt"
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test", "gpt-alt"]
@@ -722,8 +669,7 @@ def test_enhance_prompt_uses_selected_model():
next_segment_idx=2,
model="gpt-alt",
timeout_ms=800,
)
)
))
assert result.fallback_used is False
assert result.error is None
assert result.prompt == "Enhanced next"
@@ -732,9 +678,7 @@ def test_enhance_prompt_uses_selected_model():
def test_enhance_prompt_single_clip_uses_auto_extension_prompt_and_prompt_field():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"prompt":"Extended single clip"}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"prompt":"Extended single clip"}'))
enhancer.enhance_system_prompt = "enhance system prompt"
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test", "gpt-alt"]
@@ -764,14 +708,12 @@ def test_enhance_prompt_single_clip_uses_auto_extension_prompt_and_prompt_field(
enhancer._request_content = _fake_request_content # type: ignore[attr-defined]
result = asyncio.run(
enhancer.enhance_prompt(
"short 5s idea",
mode="single_clip",
model="gpt-alt",
timeout_ms=800,
)
)
result = asyncio.run(enhancer.enhance_prompt(
"short 5s idea",
mode="single_clip",
model="gpt-alt",
timeout_ms=800,
))
assert result.fallback_used is False
assert result.error is None
assert result.prompt == "Extended single clip"
@@ -784,17 +726,15 @@ def test_enhance_prompt_single_clip_uses_auto_extension_prompt_and_prompt_field(
"single 5-second LTX-2.3 video clip. Respond with "
'valid JSON only as {"prompt": "..."}.' # noqa: E501
),
"user_prompt": "short 5s idea",
"user_prompt":
"short 5s idea",
}
def test_enhance_prompt_single_clip_rejects_plain_text_response():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
"Medium shot of a woman by a rainy cafe window as she lifts her "
"phone, exhales softly, and the camera makes a slow push in."
)
)
_chat_payload_with_content("Medium shot of a woman by a rainy cafe window as she lifts her "
"phone, exhales softly, and the camera makes a slow push in."))
enhancer.auto_system_prompt = "auto system prompt"
result = asyncio.run(
@@ -803,17 +743,14 @@ def test_enhance_prompt_single_clip_rejects_plain_text_response():
mode="single_clip",
model="gpt-test",
timeout_ms=800,
)
)
))
assert result.fallback_used is True
assert "No JSON object found in assistant response." in result.error
assert result.prompt == ""
def test_enhance_prompt_single_clip_rejects_segment_prompts_json():
enhancer = _build_test_enhancer(
_chat_payload_with_content('{"segment_prompts":["A","B"]}')
)
enhancer = _build_test_enhancer(_chat_payload_with_content('{"segment_prompts":["A","B"]}'))
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test"]
enhancer.rewrite_default_model = "gpt-test"
@@ -824,17 +761,14 @@ def test_enhance_prompt_single_clip_rejects_segment_prompts_json():
mode="single_clip",
model="gpt-test",
timeout_ms=800,
)
)
))
assert result.fallback_used is True
assert result.prompt == ""
assert "Missing prompt string." in (result.error or "")
def test_enhance_prompt_requires_json_and_does_not_fallback_to_raw_text():
enhancer = _build_test_enhancer(
_chat_payload_with_content("A cinematic continuation with slow dolly movement.")
)
enhancer = _build_test_enhancer(_chat_payload_with_content("A cinematic continuation with slow dolly movement."))
enhancer.enhance_system_prompt = "enhance system prompt"
enhancer.rewrite_model_options = ["gpt-test"]
enhancer.rewrite_default_model = "gpt-test"
@@ -846,17 +780,14 @@ def test_enhance_prompt_requires_json_and_does_not_fallback_to_raw_text():
next_segment_idx=2,
model="gpt-test",
timeout_ms=800,
)
)
))
assert result.fallback_used is True
assert result.prompt == ""
assert "No JSON object found in assistant response." in (result.error or "")
def test_generate_auto_prompt_requires_json_and_does_not_fallback_to_raw_text():
enhancer = _build_test_enhancer(
_chat_payload_with_content("A calm, grounded continuation with subtle motion.")
)
enhancer = _build_test_enhancer(_chat_payload_with_content("A calm, grounded continuation with subtle motion."))
enhancer.auto_system_prompt = "auto system prompt"
enhancer.rewrite_model_options = ["gpt-test"]
enhancer.rewrite_default_model = "gpt-test"
@@ -867,34 +798,30 @@ def test_generate_auto_prompt_requires_json_and_does_not_fallback_to_raw_text():
next_segment_idx=2,
model="gpt-test",
timeout_ms=800,
)
)
))
assert result.fallback_used is True
assert result.prompt == ""
assert "No JSON object found in assistant response." in (result.error or "")
def test_rewrite_prompt_sequence_includes_raw_json_when_content_empty():
enhancer = _build_test_enhancer(
{
"choices": [
{
"finish_reason": "length",
"message": {
"content": [],
"refusal": None,
},
}
],
"usage": {"completion_tokens": 0},
}
)
enhancer = _build_test_enhancer({
"choices": [{
"finish_reason": "length",
"message": {
"content": [],
"refusal": None,
},
}],
"usage": {
"completion_tokens": 0
},
})
result = asyncio.run(
enhancer.rewrite_prompt_sequence(
["prompt one", "prompt two"],
rewrite_instruction="make it cinematic",
)
)
))
assert result.fallback_used is True
assert "No rewrite segment prompts found in assistant response." in (result.error or "")
assert isinstance(result.raw_response_text, str)
@@ -903,10 +830,7 @@ def test_rewrite_prompt_sequence_includes_raw_json_when_content_empty():
def test_get_rewrite_model_config_returns_fixed_defaults():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.rewrite_default_model = "gpt-oss-120b"
enhancer.rewrite_model_options = ["gpt-oss-120b"]
@@ -918,10 +842,7 @@ def test_get_rewrite_model_config_returns_fixed_defaults():
def test_get_prompt_config_includes_auto_extension_prompt():
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
enhancer.enhance_system_prompt_path = "/tmp/next.md"
enhancer.auto_system_prompt_path = "/tmp/auto.md"
enhancer.rewrite_all_system_prompt_path = "/tmp/rewrite.md"
@@ -948,19 +869,14 @@ def test_get_prompt_config_includes_auto_extension_prompt():
def test_get_prompt_config_reports_loaded_fallback_prompt_path(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
rewrite_fallback_path = tmp_path / "rewrite_window_system_prompt.md"
rewrite_fallback_path.write_text("rewrite prompt\n", encoding="utf-8")
next_path = tmp_path / "next.md"
next_path.write_text("next prompt\n", encoding="utf-8")
auto_path = tmp_path / "auto.md"
auto_path.write_text("auto prompt\n", encoding="utf-8")
enhancer.rewrite_all_system_prompt_path = str(
tmp_path / "prompts.local" / "rewrite_window_system_prompt.md"
)
enhancer.rewrite_all_system_prompt_path = str(tmp_path / "prompts.local" / "rewrite_window_system_prompt.md")
enhancer.rewrite_all_system_prompt_fallback_path = str(rewrite_fallback_path)
enhancer.enhance_system_prompt_path = str(next_path)
enhancer.auto_system_prompt_path = str(auto_path)
@@ -973,14 +889,9 @@ def test_get_prompt_config_reports_loaded_fallback_prompt_path(tmp_path):
assert config["rewrite_window_system_prompt_path"] == str(rewrite_fallback_path)
def test_reload_system_prompts_falls_back_to_rewrite_window_when_user_prompt_empty(
tmp_path,
):
def test_reload_system_prompts_falls_back_to_rewrite_window_when_user_prompt_empty(tmp_path, ):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite_window_system_prompt.md"
@@ -1007,10 +918,7 @@ def test_reload_system_prompts_falls_back_to_rewrite_window_when_user_prompt_emp
def test_save_prompt_config_updates_auto_extension_prompt(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite.md"
@@ -1021,9 +929,7 @@ def test_save_prompt_config_updates_auto_extension_prompt(tmp_path):
enhancer.auto_system_prompt_path = str(auto_path)
enhancer.rewrite_all_system_prompt_path = str(rewrite_path)
config = enhancer.save_prompt_config(
auto_extension_system_prompt="auto updated",
)
config = enhancer.save_prompt_config(auto_extension_system_prompt="auto updated", )
assert auto_path.read_text(encoding="utf-8").strip() == "auto updated"
assert config["auto_extension_system_prompt"] == "auto updated"
@@ -1031,10 +937,7 @@ def test_save_prompt_config_updates_auto_extension_prompt(tmp_path):
def test_save_prompt_config_updates_rewrite_user_prompt(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite.md"
@@ -1052,9 +955,7 @@ def test_save_prompt_config_updates_rewrite_user_prompt(tmp_path):
enhancer.rewrite_all_system_prompt_fallback_path = None
enhancer.rewrite_user_system_prompt_fallback_path = None
config = enhancer.save_prompt_config(
rewrite_user_system_prompt="rewrite user updated",
)
config = enhancer.save_prompt_config(rewrite_user_system_prompt="rewrite user updated", )
assert rewrite_user_path.read_text(encoding="utf-8").strip() == "rewrite user updated"
assert config["rewrite_user_system_prompt"] == "rewrite user updated"
@@ -1062,10 +963,7 @@ def test_save_prompt_config_updates_rewrite_user_prompt(tmp_path):
def test_save_prompt_config_updates_rewrite_model(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite.md"
@@ -1081,9 +979,7 @@ def test_save_prompt_config_updates_rewrite_model(tmp_path):
enhancer.rewrite_default_model = "gpt-test"
enhancer.rewrite_model_options = ["gpt-test", "gpt-alt"]
config = enhancer.save_prompt_config(
rewrite_model="gpt-alt",
)
config = enhancer.save_prompt_config(rewrite_model="gpt-alt", )
assert enhancer.rewrite_default_model == "gpt-alt"
assert config["rewrite_model"] == "gpt-alt"
@@ -1092,10 +988,7 @@ def test_save_prompt_config_updates_rewrite_model(tmp_path):
def test_save_prompt_config_updates_rewrite_temperature(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite.md"
@@ -1109,9 +1002,7 @@ def test_save_prompt_config_updates_rewrite_temperature(tmp_path):
enhancer.auto_system_prompt_fallback_path = None
enhancer.rewrite_all_system_prompt_fallback_path = None
config = enhancer.save_prompt_config(
rewrite_temperature=1.3,
)
config = enhancer.save_prompt_config(rewrite_temperature=1.3, )
assert enhancer.rewrite_default_temperature == 1.3
assert config["rewrite_temperature"] == 1.3
@@ -1119,10 +1010,7 @@ def test_save_prompt_config_updates_rewrite_temperature(tmp_path):
def test_save_prompt_config_creates_versioned_backup_for_existing_prompt(tmp_path):
enhancer = _build_test_enhancer(
_chat_payload_with_content(
'{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'
)
)
_chat_payload_with_content('{"id":"preset_a","label":"Preset A","segment_prompts":["A","B"]}'))
next_path = tmp_path / "next.md"
auto_path = tmp_path / "auto.md"
rewrite_path = tmp_path / "rewrite_window_system_prompt.md"
@@ -1136,13 +1024,9 @@ def test_save_prompt_config_creates_versioned_backup_for_existing_prompt(tmp_pat
enhancer.auto_system_prompt_fallback_path = None
enhancer.rewrite_all_system_prompt_fallback_path = None
enhancer.save_prompt_config(
rewrite_window_system_prompt="rewrite updated",
)
enhancer.save_prompt_config(rewrite_window_system_prompt="rewrite updated", )
backup_paths = sorted(
tmp_path.glob("rewrite_window_system_prompt.*.bak.md")
)
backup_paths = sorted(tmp_path.glob("rewrite_window_system_prompt.*.bak.md"))
assert rewrite_path.read_text(encoding="utf-8").strip() == "rewrite updated"
assert len(backup_paths) == 1
@@ -27,13 +27,8 @@ try:
except ModuleNotFoundError:
websockets = None # type: ignore[assignment]
DEFAULT_PRESET_FILE = (
Path(__file__).resolve().parents[2]
/ "web"
/ "prompts"
/ "selected_ltx2_continuation_story_presets.json"
)
DEFAULT_PRESET_FILE = (Path(__file__).resolve().parents[2] / "web" / "prompts" /
"selected_ltx2_continuation_story_presets.json")
def utc_now_iso() -> str:
@@ -65,10 +60,7 @@ def safe_percentile(values: list[float], percentile: float) -> float | None:
if lower == upper:
return sorted_values[lower]
fraction = rank - lower
return (
sorted_values[lower]
+ (sorted_values[upper] - sorted_values[lower]) * fraction
)
return (sorted_values[lower] + (sorted_values[upper] - sorted_values[lower]) * fraction)
def summarize_series(values: list[float]) -> dict[str, float | int | None]:
@@ -145,24 +137,16 @@ def load_curated_prompts(
selected_id = str(selected.get("id", "")).strip() or "unknown_preset"
raw_prompts = selected.get("segment_prompts", [])
if not isinstance(raw_prompts, list):
raise ValueError(
f"Preset {selected_id} has invalid segment_prompts (must be list)."
)
raise ValueError(f"Preset {selected_id} has invalid segment_prompts (must be list).")
prompts = [
str(prompt).strip()
for prompt in raw_prompts
if isinstance(prompt, str) and str(prompt).strip()
]
prompts = [str(prompt).strip() for prompt in raw_prompts if isinstance(prompt, str) and str(prompt).strip()]
if not prompts:
raise ValueError(f"Preset {selected_id} has no non-empty prompts.")
limited = prompts[:curated_limit]
if not limited:
raise ValueError(
f"curated_limit={curated_limit} produced no prompts for preset "
f"{selected_id}."
)
raise ValueError(f"curated_limit={curated_limit} produced no prompts for preset "
f"{selected_id}.")
return selected_id, limited, len(prompts)
@@ -224,11 +208,11 @@ async def run_single_session(
try:
async with websockets.connect(
url,
max_size=None,
ping_interval=None,
open_timeout=connect_timeout_s,
close_timeout=2.0,
url,
max_size=None,
ping_interval=None,
open_timeout=connect_timeout_s,
close_timeout=2.0,
) as ws:
connect_finish_monotonic = time.monotonic()
session_data["connect_finish_ts_utc"] = utc_now_iso()
@@ -249,9 +233,7 @@ async def run_single_session(
timeout_remaining = session_timeout_s - elapsed_s
if timeout_remaining <= 0:
session_data["status"] = "timeout"
session_data["error"] = (
f"Session timed out after {session_timeout_s:.1f}s."
)
session_data["error"] = (f"Session timed out after {session_timeout_s:.1f}s.")
break
recv_start_epoch = time.time()
@@ -265,9 +247,7 @@ async def run_single_session(
)
except asyncio.TimeoutError:
session_data["status"] = "timeout"
session_data["error"] = (
"Timed out waiting for websocket message."
)
session_data["error"] = ("Timed out waiting for websocket message.")
break
except Exception as exc:
session_data["status"] = "failed"
@@ -288,20 +268,16 @@ async def run_single_session(
chunk_gap_ms: float | None = None
if last_chunk_finish_monotonic is not None:
chunk_gap_ms = (
recv_finish_monotonic - last_chunk_finish_monotonic
) * 1000.0
chunk_gap_ms = (recv_finish_monotonic - last_chunk_finish_monotonic) * 1000.0
session_data["chunks"].append(
{
"segment_idx": current_segment_idx,
"chunk_idx": session_data["total_chunks"],
"size_bytes": len(message),
"chunk_start_ts_utc": recv_start_iso,
"chunk_finish_ts_utc": recv_finish_iso,
"chunk_gap_ms": chunk_gap_ms,
}
)
session_data["chunks"].append({
"segment_idx": current_segment_idx,
"chunk_idx": session_data["total_chunks"],
"size_bytes": len(message),
"chunk_start_ts_utc": recv_start_iso,
"chunk_finish_ts_utc": recv_finish_iso,
"chunk_gap_ms": chunk_gap_ms,
})
last_chunk_finish_monotonic = recv_finish_monotonic
last_chunk_finish_epoch = recv_finish_epoch
session_data["last_chunk_finish_ts_utc"] = recv_finish_iso
@@ -321,9 +297,7 @@ async def run_single_session(
if msg_type == "gpu_assigned":
session_data["gpu_assigned_ts_utc"] = recv_finish_iso
if connect_finish_monotonic is not None:
session_data["queue_wait_ms"] = (
recv_finish_monotonic - connect_finish_monotonic
) * 1000.0
session_data["queue_wait_ms"] = (recv_finish_monotonic - connect_finish_monotonic) * 1000.0
elif msg_type == "ltx2_stream_start":
if initial_total_segments is None:
parsed_total = parse_int(data.get("total_segments"))
@@ -338,20 +312,13 @@ async def run_single_session(
session_data["media_segments_completed"] += 1
if first_media_segment_complete_epoch is None:
first_media_segment_complete_epoch = recv_finish_epoch
session_data[
"first_media_segment_complete_ts_utc"
] = recv_finish_iso
session_data["first_media_segment_complete_ts_utc"] = recv_finish_iso
elif msg_type == "ltx2_segment_complete":
session_data["segments_completed"] += 1
seg_idx = parse_int(data.get("segment_idx"))
if (
initial_total_segments is not None
and seg_idx is not None
and seg_idx >= initial_total_segments
):
session_data[
"target_segment_complete_ts_utc"
] = recv_finish_iso
if (initial_total_segments is not None and seg_idx is not None
and seg_idx >= initial_total_segments):
session_data["target_segment_complete_ts_utc"] = recv_finish_iso
await asyncio.sleep(post_complete_wait_s)
session_data["leave_sent_ts_utc"] = utc_now_iso()
try:
@@ -362,15 +329,11 @@ async def run_single_session(
break
elif msg_type == "session_timeout":
session_data["status"] = "timeout"
session_data["error"] = str(
data.get("message") or "Backend session timeout"
)
session_data["error"] = str(data.get("message") or "Backend session timeout")
break
elif msg_type == "error":
session_data["status"] = "failed"
session_data["error"] = str(
data.get("message") or "Backend error message"
)
session_data["error"] = str(data.get("message") or "Backend error message")
break
if session_data["status"] == "failed" and session_data["error"] is None:
@@ -379,29 +342,18 @@ async def run_single_session(
session_data["status"] = "failed"
session_data["error"] = f"WebSocket connect/run failed: {exc}"
if (
first_chunk_finish_epoch is not None
and last_chunk_finish_epoch is not None
and session_data["total_chunk_bytes"] > 0
):
if (first_chunk_finish_epoch is not None and last_chunk_finish_epoch is not None
and session_data["total_chunk_bytes"] > 0):
duration_s = last_chunk_finish_epoch - first_chunk_finish_epoch
if duration_s > 0:
session_data["session_goodput_mbps"] = (
session_data["total_chunk_bytes"] * 8.0 / duration_s / 1_000_000.0
)
session_data["session_goodput_mbps"] = (session_data["total_chunk_bytes"] * 8.0 / duration_s / 1_000_000.0)
if (
first_chunk_finish_epoch is not None
and first_media_segment_complete_epoch is not None
):
session_data["first_chunk_before_first_media_complete"] = (
first_chunk_finish_epoch < first_media_segment_complete_epoch
)
if (first_chunk_finish_epoch is not None and first_media_segment_complete_epoch is not None):
session_data["first_chunk_before_first_media_complete"] = (first_chunk_finish_epoch
< first_media_segment_complete_epoch)
session_data["close_ts_utc"] = utc_now_iso()
session_data["duration_ms"] = (
time.monotonic() - session_start_monotonic
) * 1000.0
session_data["duration_ms"] = (time.monotonic() - session_start_monotonic) * 1000.0
return session_data
@@ -412,14 +364,11 @@ async def run_worker_sessions(
config: dict[str, Any],
) -> list[dict[str, Any]]:
tasks = [
asyncio.create_task(
run_single_session(
worker_id=worker_id,
worker_session_idx=idx,
config=config,
)
)
for idx in range(session_count)
asyncio.create_task(run_single_session(
worker_id=worker_id,
worker_session_idx=idx,
config=config,
)) for idx in range(session_count)
]
if not tasks:
return []
@@ -437,29 +386,23 @@ def worker_entry(
try:
ready_queue.put({"worker_id": worker_id, "status": "ready"})
start_event.wait()
sessions = asyncio.run(
run_worker_sessions(
worker_id=worker_id,
session_count=session_count,
config=config,
)
)
result_queue.put(
{
"worker_id": worker_id,
"status": "ok",
"sessions": sessions,
}
)
sessions = asyncio.run(run_worker_sessions(
worker_id=worker_id,
session_count=session_count,
config=config,
))
result_queue.put({
"worker_id": worker_id,
"status": "ok",
"sessions": sessions,
})
except Exception as exc:
result_queue.put(
{
"worker_id": worker_id,
"status": "error",
"error": str(exc),
"traceback": traceback.format_exc(),
}
)
result_queue.put({
"worker_id": worker_id,
"status": "error",
"error": str(exc),
"traceback": traceback.format_exc(),
})
def build_summary(
@@ -517,33 +460,22 @@ def build_summary(
if len(all_chunk_finish_epochs) >= 2 and total_chunk_bytes > 0:
duration_s = max(all_chunk_finish_epochs) - min(all_chunk_finish_epochs)
if duration_s > 0:
global_goodput_mbps = (
total_chunk_bytes * 8.0 / duration_s / 1_000_000.0
)
global_goodput_mbps = (total_chunk_bytes * 8.0 / duration_s / 1_000_000.0)
bucket_throughputs_mbps = [
(bytes_count * 8.0) / 1_000_000.0
for _, bytes_count in sorted(bucket_bytes.items())
]
bucket_throughputs_mbps = [(bytes_count * 8.0) / 1_000_000.0 for _, bytes_count in sorted(bucket_bytes.items())]
bucket_stats = summarize_series(bucket_throughputs_mbps)
chunk_gap_threshold_breaches = [
value for value in chunk_gaps if value >= chunk_gap_threshold_ms
]
chunk_gap_threshold_breaches = [value for value in chunk_gaps if value >= chunk_gap_threshold_ms]
non_success = len(sessions) - status_counts.get("success", 0)
fail_reasons: list[str] = []
if non_success > 0:
fail_reasons.append(
f"{non_success} session(s) did not complete successfully."
)
fail_reasons.append(f"{non_success} session(s) did not complete successfully.")
if not chunk_gaps:
fail_reasons.append("No chunk gap data collected.")
if chunk_gap_threshold_breaches:
fail_reasons.append(
f"{len(chunk_gap_threshold_breaches)} chunk gap(s) were >= "
f"{chunk_gap_threshold_ms:.0f}ms."
)
fail_reasons.append(f"{len(chunk_gap_threshold_breaches)} chunk gap(s) were >= "
f"{chunk_gap_threshold_ms:.0f}ms.")
passed = len(fail_reasons) == 0
progressive_ratio = None
@@ -554,20 +486,18 @@ def build_summary(
"passed": passed,
"fail_reasons": fail_reasons,
"sessions": {
"total": len(sessions),
"success": status_counts.get("success", 0),
"failed": status_counts.get("failed", 0),
"timeout": status_counts.get("timeout", 0),
"protocol_error": status_counts.get("protocol_error", 0),
"other": (
len(sessions)
- (
status_counts.get("success", 0)
+ status_counts.get("failed", 0)
+ status_counts.get("timeout", 0)
+ status_counts.get("protocol_error", 0)
)
),
"total":
len(sessions),
"success":
status_counts.get("success", 0),
"failed":
status_counts.get("failed", 0),
"timeout":
status_counts.get("timeout", 0),
"protocol_error":
status_counts.get("protocol_error", 0),
"other": (len(sessions) - (status_counts.get("success", 0) + status_counts.get("failed", 0) +
status_counts.get("timeout", 0) + status_counts.get("protocol_error", 0))),
},
"chunk_gap_ms": {
**chunk_gap_stats,
@@ -606,51 +536,39 @@ def print_summary(
bucket_bw = bandwidth["bucketed_1s"]
print("=== LTX2 Realtime Stress Test Summary ===")
print(
"Run: "
f"url={run_info['url']} clients={run_info['clients']} "
f"processes={run_info['processes']} "
f"preset={run_info['preset_id']} "
f"curated_limit={run_info['curated_limit']}"
)
print(
"Sessions: "
f"total={sessions['total']} success={sessions['success']} "
f"failed={sessions['failed']} timeout={sessions['timeout']} "
f"protocol_error={sessions['protocol_error']}"
)
print(
"Chunk gap ms: "
f"min={format_num(chunk_gap['min'])} "
f"p50={format_num(chunk_gap['p50'])} "
f"p95={format_num(chunk_gap['p95'])} "
f"p99={format_num(chunk_gap['p99'])} "
f"max={format_num(chunk_gap['max'])} "
f"threshold={format_num(chunk_gap['threshold_ms'])} "
f"breaches={chunk_gap['breach_count']}"
)
print(
"Queue wait ms: "
f"min={format_num(queue_wait['min'])} "
f"p50={format_num(queue_wait['p50'])} "
f"p95={format_num(queue_wait['p95'])} "
f"max={format_num(queue_wait['max'])}"
)
print("Run: "
f"url={run_info['url']} clients={run_info['clients']} "
f"processes={run_info['processes']} "
f"preset={run_info['preset_id']} "
f"curated_limit={run_info['curated_limit']}")
print("Sessions: "
f"total={sessions['total']} success={sessions['success']} "
f"failed={sessions['failed']} timeout={sessions['timeout']} "
f"protocol_error={sessions['protocol_error']}")
print("Chunk gap ms: "
f"min={format_num(chunk_gap['min'])} "
f"p50={format_num(chunk_gap['p50'])} "
f"p95={format_num(chunk_gap['p95'])} "
f"p99={format_num(chunk_gap['p99'])} "
f"max={format_num(chunk_gap['max'])} "
f"threshold={format_num(chunk_gap['threshold_ms'])} "
f"breaches={chunk_gap['breach_count']}")
print("Queue wait ms: "
f"min={format_num(queue_wait['min'])} "
f"p50={format_num(queue_wait['p50'])} "
f"p95={format_num(queue_wait['p95'])} "
f"max={format_num(queue_wait['max'])}")
ratio = progressive["ratio"]
ratio_text = "n/a" if ratio is None else f"{ratio * 100:.2f}%"
print(
"Progressive streaming: "
f"{progressive['success_sessions']}/"
f"{progressive['eligible_sessions']} ({ratio_text})"
)
print(
"Bandwidth Mbps: "
f"per_session_avg={format_num(per_session_bw['avg'])} "
f"per_session_p95={format_num(per_session_bw['p95'])} "
f"global={format_num(bandwidth['global_goodput_mbps'])} "
f"bucket_avg={format_num(bucket_bw['avg_mbps'])} "
f"bucket_peak={format_num(bucket_bw['peak_mbps'])}"
)
print("Progressive streaming: "
f"{progressive['success_sessions']}/"
f"{progressive['eligible_sessions']} ({ratio_text})")
print("Bandwidth Mbps: "
f"per_session_avg={format_num(per_session_bw['avg'])} "
f"per_session_p95={format_num(per_session_bw['p95'])} "
f"global={format_num(bandwidth['global_goodput_mbps'])} "
f"bucket_avg={format_num(bucket_bw['avg_mbps'])} "
f"bucket_peak={format_num(bucket_bw['peak_mbps'])}")
print(f"VERDICT: {'PASS' if summary['passed'] else 'FAIL'}")
if summary["fail_reasons"]:
print("Fail reasons:")
@@ -670,10 +588,8 @@ def distribute_sessions(total_clients: int, process_count: int) -> list[int]:
def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
if websockets is None:
raise RuntimeError(
"Missing dependency: websockets. Install it before running this "
"stress test."
)
raise RuntimeError("Missing dependency: websockets. Install it before running this "
"stress test.")
preset_file = Path(args.preset_file).expanduser().resolve()
selected_preset_id, curated_prompts, total_prompt_count = load_curated_prompts(
@@ -735,13 +651,8 @@ def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
start_event.set()
result_deadline = (
time.monotonic()
+ args.connect_timeout_s
+ args.session_timeout_s
+ args.post_complete_wait_s
+ 180.0
)
result_deadline = (time.monotonic() + args.connect_timeout_s + args.session_timeout_s +
args.post_complete_wait_s + 180.0)
worker_results: list[dict[str, Any]] = []
while len(worker_results) < len(processes):
timeout_s = max(0.1, result_deadline - time.monotonic())
@@ -765,24 +676,20 @@ def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
if result.get("status") == "ok":
sessions.extend(result.get("sessions", []))
else:
worker_errors.append(
{
"worker_id": result.get("worker_id"),
"error": result.get("error"),
"traceback": result.get("traceback"),
}
)
worker_errors.append({
"worker_id": result.get("worker_id"),
"error": result.get("error"),
"traceback": result.get("traceback"),
})
received_workers = {result.get("worker_id") for result in worker_results}
expected_workers = set(range(len(processes)))
missing_workers = sorted(expected_workers - received_workers)
for worker_id in missing_workers:
worker_errors.append(
{
"worker_id": worker_id,
"error": "No worker result received.",
}
)
worker_errors.append({
"worker_id": worker_id,
"error": "No worker result received.",
})
run_end_epoch = time.time()
run_end_iso = iso_from_epoch(run_end_epoch)
@@ -795,9 +702,8 @@ def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
if worker_errors:
summary["passed"] = False
summary["fail_reasons"] = list(summary["fail_reasons"]) + [
f"{len(worker_errors)} worker error(s) occurred."
]
summary["fail_reasons"] = list(
summary["fail_reasons"]) + [f"{len(worker_errors)} worker error(s) occurred."]
output_payload = {
"run_info": {
@@ -833,9 +739,7 @@ def run_stress(args: argparse.Namespace) -> tuple[dict[str, Any], int]:
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(
description="Multiprocess realtime stress test for LTX2 streaming.",
)
parser = argparse.ArgumentParser(description="Multiprocess realtime stress test for LTX2 streaming.", )
parser.add_argument(
"-u",
"--url",
@@ -47,13 +47,11 @@ def test_persist_session_init_image_returns_none_when_missing_data():
def test_persist_session_init_image_rejects_unsupported_mime():
with pytest.raises(ValueError, match="PNG, JPEG, or WebP"):
persist_session_init_image(
{
"name": "frame.gif",
"mime_type": "image/gif",
"data_url": "data:image/gif;base64,R0lGODlhAQABAAAAACw=",
}
)
persist_session_init_image({
"name": "frame.gif",
"mime_type": "image/gif",
"data_url": "data:image/gif;base64,R0lGODlhAQABAAAAACw=",
})
def test_persist_session_init_image_rejects_large_payload(monkeypatch):
@@ -66,10 +64,8 @@ def test_persist_session_init_image_rejects_large_payload(monkeypatch):
monkeypatch.setattr(base64, "b64decode", fake_b64decode)
with pytest.raises(ValueError, match="15 MB or smaller"):
persist_session_init_image(
{
"name": "frame.png",
"mime_type": "image/png",
"data_url": data_url,
}
)
persist_session_init_image({
"name": "frame.png",
"mime_type": "image/png",
"data_url": data_url,
})
File diff suppressed because it is too large Load Diff
+3
View File
@@ -15,6 +15,8 @@ from __future__ import annotations
from dataclasses import dataclass
from dreamverse.generation_inputs import GenerationInputs
# ---- User-scoped events (carry user_id) ------------------------------------
@@ -147,6 +149,7 @@ class UserStepPayload:
segment_idx: int
image_path: str | None
reset_conditioning: bool
generation_inputs: GenerationInputs | None = None
@dataclass(frozen=True)
+7 -7
View File
@@ -70,7 +70,7 @@
<mxCell id="dispatcher" value="command dispatcher&#xa;&#xa;gpu_worker_process() branches on&#xa;CommandType; asserts payload type&#xa;&#xa;INIT / WARMUP / RELOAD_MODEL&#xa;USER_JOIN / USER_STEP / USER_LEAVE&#xa;SHUTDOWN" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe6cc;strokeColor=#d79b00;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="120" y="1120" width="240" height="120" as="geometry"/>
</mxCell>
<mxCell id="do_step" value="VideoGenerationWorker.generate_step()&#xa;video_generation.py:380&#xa;&#xa;reads + updates ContinuationState,&#xa;calls generator" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
<mxCell id="do_step" value="VideoGenerationWorker.generate_step()&#xa;ltx2_generation.py:380&#xa;&#xa;reads + updates ContinuationState,&#xa;calls generator" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
<mxGeometry x="460" y="1120" width="240" height="120" as="geometry"/>
</mxCell>
<mxCell id="stream_av" value="stream_fmp4()&#xa;av_streaming.py:121&#xa;&#xa;trims overlap, pipes to ffmpeg,&#xa;publishes StreamInit / StreamChunk /&#xa;StreamComplete via callback" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#b1d8d7;strokeColor=#23445d;fontSize=11;align=left;spacingLeft=10;spacingTop=8;fontStyle=1;" parent="1" vertex="1">
@@ -79,13 +79,13 @@
<mxCell id="Ot8BU52QTIb4EhyRSe7I-2" value="" style="edgeStyle=none;html=1;" parent="1" source="generator" target="Ot8BU52QTIb4EhyRSe7I-1" edge="1">
<mxGeometry relative="1" as="geometry"/>
</mxCell>
<mxCell id="generator" value="VideoGenerator (fastvideo)&#xa;&#xa;LTX2 DiT + refine upsampler&#xa;FP4 quant, torch.compile&#xa;&#xa;owned by VideoGenerationWorker&#xa;video_generation.py:211" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;" parent="1" vertex="1">
<mxCell id="generator" value="VideoGenerator (fastvideo)&#xa;&#xa;LTX2 DiT + refine upsampler&#xa;FP4 quant, torch.compile&#xa;&#xa;owned by VideoGenerationWorker&#xa;ltx2_generation.py:211" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;" parent="1" vertex="1">
<mxGeometry x="460" y="1300" width="240" height="100" as="geometry"/>
</mxCell>
<mxCell id="ffmpeg" value="ffmpeg subprocess&#xa;&#xa;libx264 / *_nvenc&#xa;fragmented mp4" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=11;" parent="1" vertex="1">
<mxGeometry x="800" y="1300" width="260" height="100" as="geometry"/>
</mxCell>
<mxCell id="caches" value="ContinuationState&#xa;video_generation.py:89&#xa;&#xa;• video_images: list[PIL.Image]&#xa;• audio_latents: torch.Tensor (CPU)&#xa;&#xa;carried across segments" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxCell id="caches" value="ContinuationState&#xa;ltx2_generation.py:89&#xa;&#xa;• video_images: list[PIL.Image]&#xa;• audio_latents: torch.Tensor (CPU)&#xa;&#xa;carried across segments" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#e1d5e7;strokeColor=#9673a6;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxGeometry x="120" y="1300" width="240" height="100" as="geometry"/>
</mxCell>
<mxCell id="e_cp" value="acquire" style="edgeStyle=orthogonalEdgeStyle;rounded=0;html=1;strokeColor=#6c8ebf;endArrow=classic;fontSize=11;exitX=0.5;exitY=1;exitDx=0;exitDy=0;entryX=0.5;entryY=0;entryDx=0;entryDy=0;" parent="1" source="client" target="pool" edge="1">
@@ -250,7 +250,7 @@
<mxPoint x="690" y="880"/>
</Array>
</mxCell>
<mxCell id="legend" value="Legend&#xa;&#xa;■ blue client / external&#xa;■ green main-process pool/slot&#xa; (methods — italic label)&#xa;■ yellow containers (routing state)&#xa;■ red IPC primitives (mp.Queue, mp.RawArray)&#xa;&#xa;Worker subprocess modules:&#xa;■ orange gpu_pool.py (dispatcher)&#xa;■ lavender video_generation.py&#xa;■ teal av_streaming.py&#xa;■ gray worker_ipc.py (shared types)&#xa;&#xa;Flow:&#xa; client → pool → slot&#xa; → _send_command(_tagged) → command_queue&#xa; → dispatcher → generate_step()&#xa; → stream_fmp4() → ffmpeg&#xa; → shared_buf + response_queue&#xa; → _response_reader → futures / stream_queues&#xa; → client awaits (via main.py AV loop)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#f5f5f5;strokeColor=#999999;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxCell id="legend" value="Legend&#xa;&#xa;■ blue client / external&#xa;■ green main-process pool/slot&#xa; (methods — italic label)&#xa;■ yellow containers (routing state)&#xa;■ red IPC primitives (mp.Queue, mp.RawArray)&#xa;&#xa;Worker subprocess modules:&#xa;■ orange gpu_pool.py (dispatcher)&#xa;■ lavender ltx2_generation.py&#xa;■ teal av_streaming.py&#xa;■ gray worker_ipc.py (shared types)&#xa;&#xa;Flow:&#xa; client → pool → slot&#xa; → _send_command(_tagged) → command_queue&#xa; → dispatcher → generate_step()&#xa; → stream_fmp4() → ffmpeg&#xa; → shared_buf + response_queue&#xa; → _response_reader → futures / stream_queues&#xa; → client awaits (via main.py AV loop)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#f5f5f5;strokeColor=#999999;fontSize=11;align=left;spacingLeft=10;spacingTop=8;" parent="1" vertex="1">
<mxGeometry x="39" y="-200" width="270" height="380" as="geometry"/>
</mxCell>
<mxCell id="Ot8BU52QTIb4EhyRSe7I-1" value="FastVideo video_generator" style="whiteSpace=wrap;html=1;fontSize=11;fillColor=#e1d5e7;strokeColor=#9673a6;rounded=1;" parent="1" vertex="1">
@@ -389,10 +389,10 @@
<mxCell id="cw2" value="from fastvideo.entrypoints.video_generator import VideoGenerator&#xa;from fastvideo.models.dits.ltx2 import DEFAULT_LTX2_AUDIO_*&#xa;&#xa;** Dreamverse reaches into fastvideo internals here **" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="695" width="550" height="60" as="geometry"/>
</mxCell>
<mxCell id="cw3" value="on Command(INIT):&#xa; VideoGenerationWorker.initialize() (video_generation.py:247)&#xa; maybe_download_model(model_id)&#xa; VideoGenerator.from_pretrained(path, FP4Config, PipelineConfig)&#xa; load audio VAE, resolve refine upsampler&#xa; resp_q.put(InitAck(success=True))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxCell id="cw3" value="on Command(INIT):&#xa; VideoGenerationWorker.initialize() (ltx2_generation.py:247)&#xa; maybe_download_model(model_id)&#xa; VideoGenerator.from_pretrained(path, FP4Config, PipelineConfig)&#xa; load audio VAE, resolve refine upsampler&#xa; resp_q.put(InitAck(success=True))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="765" width="550" height="95" as="geometry"/>
</mxCell>
<mxCell id="cw4" value="on Command(WARMUP) with WarmupPayload:&#xa; VideoGenerationWorker.warmup(payload.prompt) (video_generation.py:518)&#xa; two synthetic segments prime caches + torch.compile&#xa; resp_q.put(WarmupComplete(timings=...))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxCell id="cw4" value="on Command(WARMUP) with WarmupPayload:&#xa; VideoGenerationWorker.warmup(payload.prompt) (ltx2_generation.py:518)&#xa; two synthetic segments prime caches + torch.compile&#xa; resp_q.put(WarmupComplete(timings=...))" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffffff;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="675" y="870" width="550" height="55" as="geometry"/>
</mxCell>
<mxCell id="cw5" value="enter main worker loop → waits for JOIN_USER / USER_STEP / LEAVE" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#c8e6c9;strokeColor=#388e3c;fontSize=11;fontStyle=1;fontFamily=monospace;" parent="1" vertex="1">
@@ -534,7 +534,7 @@
<mxPoint x="1040" y="1610" as="targetPoint"/>
</mxGeometry>
</mxCell>
<mxCell id="dm11a" value="10a. worker runs:&#xa;VideoGenerationWorker.generate_step()&#xa; (video_generation.py:380)&#xa; → generator.generate_video()&#xa; → updates ContinuationState&#xa;then stream_fmp4() (av_streaming.py:121)&#xa; → ffmpeg (rawvideo+wav → fmp4)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxCell id="dm11a" value="10a. worker runs:&#xa;VideoGenerationWorker.generate_step()&#xa; (ltx2_generation.py:380)&#xa; → generator.generate_video()&#xa; → updates ContinuationState&#xa;then stream_fmp4() (av_streaming.py:121)&#xa; → ffmpeg (rawvideo+wav → fmp4)" style="rounded=1;whiteSpace=wrap;html=1;fillColor=#ffe0b2;strokeColor=#d79b00;fontSize=10;align=left;spacingLeft=8;fontFamily=monospace;" parent="1" vertex="1">
<mxGeometry x="955" y="1640" width="180" height="70" as="geometry"/>
</mxCell>
<mxCell id="dm11" value="10b. resp_q.put(MediaInit / MediaChunk / MediaComplete / StepComplete)" style="endArrow=classic;html=1;strokeColor=#b85450;fontSize=10;labelBackgroundColor=#ffffff;" parent="1" edge="1">
File diff suppressed because one or more lines are too long

Before

Width:  |  Height:  |  Size: 85 KiB

After

Width:  |  Height:  |  Size: 85 KiB

@@ -39,6 +39,17 @@ export FASTVIDEO_GENERATION_SEGMENT_CAP="${FASTVIDEO_GENERATION_SEGMENT_CAP:-6}"
export FASTVIDEO_PROMPT_AUTO_SLEEP_MS="${FASTVIDEO_PROMPT_AUTO_SLEEP_MS:-120}"
export FASTVIDEO_PROMPT_AUTO_TIMEOUT_MS="${FASTVIDEO_PROMPT_AUTO_TIMEOUT_MS:-1800}"
if [[ "${ENABLE_TORCH_COMPILE}" == "1" ]]; then
# Persist Inductor, AOTAutograd, and Triton artifacts across launches.
export DREAMVERSE_TORCH_COMPILE_CACHE_ROOT="${DREAMVERSE_TORCH_COMPILE_CACHE_ROOT:-${HOME}/.cache/dreamverse/torch_compile}"
export TORCHINDUCTOR_CACHE_DIR="${TORCHINDUCTOR_CACHE_DIR:-${DREAMVERSE_TORCH_COMPILE_CACHE_ROOT}/inductor}"
export TRITON_CACHE_DIR="${TRITON_CACHE_DIR:-${DREAMVERSE_TORCH_COMPILE_CACHE_ROOT}/triton}"
export TORCHINDUCTOR_FX_GRAPH_CACHE="${TORCHINDUCTOR_FX_GRAPH_CACHE:-1}"
export TORCHINDUCTOR_AUTOGRAD_CACHE="${TORCHINDUCTOR_AUTOGRAD_CACHE:-1}"
mkdir -p "${TORCHINDUCTOR_CACHE_DIR}" "${TRITON_CACHE_DIR}"
echo "[launch-demo] torch.compile cache: ${DREAMVERSE_TORCH_COMPILE_CACHE_ROOT}"
fi
cd "${DREAMVERSE_ROOT}"
if ! command -v dreamverse-server >/dev/null 2>&1; then
+9 -15
View File
@@ -8,12 +8,10 @@ import modal
IMAGE = os.environ.get("DREAMVERSE_IMAGE")
if not IMAGE:
raise RuntimeError(
"DREAMVERSE_IMAGE is required. Set it to a published SHA-specific Dreamverse image, "
"for example a dreamverse-backend-cuda13.0.0-sha-* tag or a "
"dreamverse-ui-cuda13.0.0-sha-* tag if serving the static UI. "
"CUDA 12 / cu126 images use the corresponding cuda12.6.3 tag."
)
raise RuntimeError("DREAMVERSE_IMAGE is required. Set it to a published SHA-specific Dreamverse image, "
"for example a dreamverse-backend-cuda13.0.0-sha-* tag or a "
"dreamverse-ui-cuda13.0.0-sha-* tag if serving the static UI. "
"CUDA 12 / cu126 images use the corresponding cuda12.6.3 tag.")
# ``@modal.web_server`` invokes ``serve()`` directly and bypasses the image
# ENTRYPOINT (``docker/docker_entrypoint.sh``). That entrypoint normally
@@ -65,14 +63,10 @@ def serve():
# ``or ""`` collapses ``None`` (unset) into an empty string, ``.strip()``
# collapses whitespace-only values (e.g. ``" "``) — both should be
# treated as missing.
missing = [
k for k in _REQUIRED_SECRET_KEYS
if not (os.environ.get(k) or "").strip()
]
missing = [k for k in _REQUIRED_SECRET_KEYS if not (os.environ.get(k) or "").strip()]
if missing:
raise RuntimeError(
"dreamverse-api-keys secret is missing required entries: "
f"{', '.join(missing)}. Add them with `modal secret create "
"dreamverse-api-keys ... --force` and redeploy "
"(see apps/dreamverse/scripts/modal/README.md).")
raise RuntimeError("dreamverse-api-keys secret is missing required entries: "
f"{', '.join(missing)}. Add them with `modal secret create "
"dreamverse-api-keys ... --force` and redeploy "
"(see apps/dreamverse/scripts/modal/README.md).")
subprocess.Popen(["dreamverse-server", "--host", "0.0.0.0", "--port", "8009"])
+107
View File
@@ -0,0 +1,107 @@
# Dreamverse on Slurm
Run Full H3 inside a one-node, four-GPU allocation. The maintained H3 examples
default to four GPUs; this is a starting configuration, not a measured minimum.
The full checkpoint supports T2VA, FL2VA, and Ref2VA. The FastH3 Preview profile
is a separate T2VA configuration.
`launch_backend.sh` checks that it is inside an `srun` step, preserves
`CUDA_VISIBLE_DEVICES`, and replaces itself with the backend process. It does
not allocate GPUs, kill existing processes, or source a personal credentials
file. The local `dreamverse-deploy` helper is not suitable for a shared Slurm
cluster because it kills processes by physical GPU and port.
## Prepare and allocate
Keep the checkout, weights, outputs, and logs on storage visible to the compute
node. Source installation is documented in the [GPU guide](../../../../docs/getting_started/installation/gpu.md).
On ARM64 GB200 use CUDA 13, a matching PyTorch build, and kernels built for
`sm_100`; the DGX Spark `sm_121` kernel image is not the GB200 image.
The repository's image workflow publishes an ARM64 GB200 variant under
`ghcr.io/hao-ai-lab/fastvideo/fastvideo-dev:py3.12-cuda13.0.0-sm100-latest`.
Resolve that tag to a digest for reproducible runs. If your compute nodes use
Pyxis/Enroot, pass the approved image or a prepared SquashFS file to
`srun --container-image`, with explicit mounts for your checkout and model cache.
The Dreamverse-specific Docker images are currently AMD64-only.
For the Slinky customer partition, a bounded allocation is:
```bash
salloc --account=customer --qos=normal --partition=hpc-rack-1 \
--nodes=1 --ntasks=1 --cpus-per-task=72 --gres=gpu:nvidia_gb200:4 \
--mem=800G --time=02:00:00 --job-name=dreamverse
srun --ntasks=1 --pty bash
```
Wait for Slurm to grant the allocation before entering the compute step. A
successful SSH login does not grant GPU resources. Inspect pending capacity
with `squeue -u "$USER" --start`; do not attach to another user's job.
The checkpoint includes duplicate release layouts. Download the diffusers
components needed by both base and reference pipelines, rather than the whole
repository (about 210 GB versus about 498 GB at revision
`42ed227ee7df40d41602854ae760620d6eb651fe`):
```bash
hf download MiniMaxAI/MiniMax-H3 \
--revision 42ed227ee7df40d41602854ae760620d6eb651fe \
--include model_index.json --include modular_model_index.json \
--include 'audio_scheduler/*' --include 'audio_vae/*' \
--include 'processor/*' --include 'scheduler/*' \
--include 'text_encoder/*' --include 'tokenizer/*' \
--include 'transformer/*' --include 'transformer_ref/*' --include 'vae/*' \
--local-dir /path/to/models/MiniMax-H3
```
The GPU environment needs `fastvideo[dreamverse]`, the Dreamverse workspace
package, and FFmpeg with H.264/AAC encoders. In a prepared FastVideo image,
install the checked-out code and its Dreamverse dependencies in that image's
Python environment. Keep its matching CUDA/PyTorch/kernel stack intact.
## Start and connect
From the checked-out repository inside the allocated step:
```bash
export DREAMVERSE_PYTHON=/path/to/environment/bin/python
export DREAMVERSE_MODEL_PATH=/path/to/models/MiniMax-H3
export FASTVIDEO_DREAMVERSE_HOME=/path/to/persistent/dreamverse-state
bash apps/dreamverse/scripts/slurm/launch_backend.sh
```
The default backend binds port 8009 on the private compute node. Connect through
the login node from your laptop, replacing `COMPUTE_NODE_IP` with the allocated
node's `NodeAddr` from `scontrol show node`:
```bash
ssh -N -L 8009:COMPUTE_NODE_IP:8009 USER@LOGIN_NODE
```
In another laptop terminal, run the frontend from your local checkout:
```bash
cd apps/dreamverse/web
BACKEND_HOST=127.0.0.1 BACKEND_PORT=8009 npm run dev
```
Open `http://localhost:5299`. `/healthz` reports the server process; `/readyz`
reports model readiness. Full H3 loads and generates more slowly than the
Preview adapter. Keep prompt enhancement disabled in the UI unless the
runtime has the selected provider's credentials.
## Verify and stop
Check all three modes with small, valid user-owned assets. Capture the selected
mode and assets, WebSocket errors or completion events, the generated video and
audio, and GPU memory usage. Also verify actionable validation errors and
backward compatibility with clients that omit `generation_mode`.
Use the frontend Playwright instructions in the
[Dreamverse development guide](../../../../docs/contributing/dreamverse-development.md)
against the forwarded backend. A mock-server demo validates UI and protocol
behavior; it is not evidence of GPU generation.
Stop the backend with Ctrl-C, exit the compute step, and release your allocation.
For a detached allocation, use `scancel YOUR_JOB_ID`. Cancel a pending demo job
when it is no longer needed; do not leave an unattended reservation queued.
@@ -0,0 +1,49 @@
#!/usr/bin/env bash
# Run inside an existing Slurm step. Slurm owns the GPU visibility and lifetime.
set -euo pipefail
if [[ -z "${SLURM_JOB_ID:-}" || -z "${SLURM_STEP_ID:-}" ]]; then
echo "Run this launcher inside an allocated Slurm step (srun), not on the login node." >&2
exit 2
fi
script_dir="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
repo_root="$(cd -- "${script_dir}/../../../.." && pwd)"
python_bin="${DREAMVERSE_PYTHON:-${repo_root}/.venv/bin/python}"
if [[ ! -x "${python_bin}" ]]; then
echo "Set DREAMVERSE_PYTHON to a Python environment with fastvideo[dreamverse] installed." >&2
exit 2
fi
export DREAMVERSE_MODEL_ID="${DREAMVERSE_MODEL_ID:-full-h3}"
export DREAMVERSE_SP_SIZE="${DREAMVERSE_SP_SIZE:-4}"
export FASTVIDEO_GPU_COUNT="${FASTVIDEO_GPU_COUNT:-${DREAMVERSE_SP_SIZE}}"
export FASTVIDEO_ENABLE_STARTUP_WARMUP="${FASTVIDEO_ENABLE_STARTUP_WARMUP:-0}"
export ENABLE_TORCH_COMPILE="${ENABLE_TORCH_COMPILE:-0}"
export STREAM_MODE="${STREAM_MODE:-av_fmp4}"
export PYTHONPATH="${repo_root}/apps/dreamverse:${repo_root}${PYTHONPATH:+:${PYTHONPATH}}"
export PYTHONUNBUFFERED=1
"${python_bin}" - <<'PY'
import os
import shutil
import torch
expected = int(os.environ["DREAMVERSE_SP_SIZE"])
visible = torch.cuda.device_count()
if expected < 1 or visible < expected:
raise SystemExit(f"The Slurm step exposes {visible} GPUs; DREAMVERSE_SP_SIZE requires {expected}.")
ffmpeg = os.environ.get("FASTVIDEO_FFMPEG_BIN", "ffmpeg")
if not shutil.which(ffmpeg):
raise SystemExit("FFmpeg is missing; install it in the compute environment or set FASTVIDEO_FFMPEG_BIN.")
print(f"Slurm job {os.environ['SLURM_JOB_ID']}: {visible} visible GPUs; using {expected} per worker")
for index in range(expected):
properties = torch.cuda.get_device_properties(index)
print(f" GPU {index}: {properties.name}, {properties.total_memory / 2**30:.1f} GiB")
PY
cd "${repo_root}"
exec "${python_bin}" -m dreamverse.server_entry \
--host "${DREAMVERSE_BIND_HOST:-0.0.0.0}" \
--port "${DREAMVERSE_BACKEND_PORT:-8009}" "$@"
@@ -0,0 +1,126 @@
import { execFileSync } from "node:child_process";
import { readFile } from "node:fs/promises";
import path from "node:path";
import { test, expect } from "@playwright/test";
const imagePath = path.resolve("public/k2.png");
const framePrompt = "A paper fox walks through a sunlit forest, gentle birdsong.";
function makeAudio(sampleRate = 8000, seconds = 1): Buffer {
const sampleCount = sampleRate * seconds;
const bytes = Buffer.alloc(44 + sampleCount * 2);
bytes.write("RIFF", 0); bytes.writeUInt32LE(bytes.length - 8, 4); bytes.write("WAVEfmt ", 8);
bytes.writeUInt32LE(16, 16); bytes.writeUInt16LE(1, 20); bytes.writeUInt16LE(1, 22);
bytes.writeUInt32LE(sampleRate, 24); bytes.writeUInt32LE(sampleRate * 2, 28);
bytes.writeUInt16LE(2, 32); bytes.writeUInt16LE(16, 34); bytes.write("data", 36);
bytes.writeUInt32LE(sampleCount * 2, 40);
for (let i = 0; i < sampleCount; i++) bytes.writeInt16LE(Math.round(Math.sin(i * 440 * 2 * Math.PI / sampleRate) * 1000), 44 + i * 2);
return bytes;
}
test.describe("generation modes through the mock runtime", () => {
for (const mode of ["t2va", "fl2va", "ref2va"] as const) {
test(`${mode} sends validated assets and plays a clearly labeled sample`, async ({ page, request }, testInfo) => {
const response = await request.get("/generation-capabilities");
const capabilities = response.ok() ? await response.json() : {};
test.skip(capabilities.mock !== true, "This test uses the CPU mock runtime; it must not silently allocate a real GPU.");
const sent: Record<string, any>[] = [];
const received: Record<string, any>[] = [];
page.on("websocket", (socket) => {
socket.on("framesent", ({ payload }) => { if (typeof payload === "string") { try { sent.push(JSON.parse(payload)); } catch {} } });
socket.on("framereceived", ({ payload }) => { if (typeof payload === "string") { try { received.push(JSON.parse(payload)); } catch {} } });
});
await page.goto("/");
await expect(page.getByText(/Demo runtime · Sample playback only/)).toBeVisible();
const modeSelect = page.getByRole("combobox", { name: "Generation mode" });
const modeLabel = mode === "ref2va" ? "Ref2VA" : mode.toUpperCase();
await modeSelect.click();
await page.getByRole("option", { name: modeLabel, exact: true }).click();
await expect(modeSelect).toHaveText(modeLabel);
await page.getByLabel("Continuation prompt").fill(framePrompt);
const uploadedIds: string[] = [];
page.on("response", async (uploadResponse) => {
if (uploadResponse.request().method() === "POST" && uploadResponse.url().endsWith("/assets") && uploadResponse.ok()) {
const asset = await uploadResponse.json().catch(() => null);
if (asset?.asset_id) uploadedIds.push(asset.asset_id);
}
});
try {
if (mode === "fl2va") {
await expect(page.getByRole("button", { name: "Generate", exact: true })).toBeDisabled();
await page.locator('input[type="file"]').setInputFiles([
{ name: "first-frame.png", mimeType: "image/png", buffer: await readFile(imagePath) },
{ name: "last-frame.png", mimeType: "image/png", buffer: await readFile(imagePath) },
]);
await expect(page.getByRole("option", { name: "first-frame.png", exact: true }).first()).toBeAttached();
await page.getByRole("combobox", { name: "First frame", exact: true }).selectOption({ label: "first-frame.png" });
await expect(page.getByRole("button", { name: "Generate", exact: true })).toBeEnabled();
await page.getByRole("combobox", { name: "Last frame", exact: true }).selectOption({ label: "last-frame.png" });
}
if (mode === "ref2va") {
const video = execFileSync(process.env.FASTVIDEO_FFMPEG_BIN || "ffmpeg", ["-v", "error", "-f", "lavfi", "-i", "color=c=royalblue:s=64x64:r=8", "-t", "1", "-c:v", "libx264", "-pix_fmt", "yuv420p", "-movflags", "frag_keyframe+empty_moov", "-f", "mp4", "pipe:1"]);
await page.locator('input[type="file"]').setInputFiles([
{ name: "subject.png", mimeType: "image/png", buffer: await readFile(imagePath) },
{ name: "motion.mp4", mimeType: "video/mp4", buffer: video },
{ name: "sound.wav", mimeType: "audio/wav", buffer: makeAudio() },
]);
await expect(page.getByRole("button", { name: "Add sound.wav as reference" })).toBeEnabled();
await page.getByRole("button", { name: "Add sound.wav as reference" }).click();
await expect(page.getByRole("button", { name: "Generate", exact: true })).toBeDisabled();
await page.getByRole("button", { name: "Add subject.png as reference" }).click();
await page.getByRole("button", { name: "Add motion.mp4 as reference" }).click();
await page.getByRole("button", { name: "Move sound.wav down" }).click();
const names = await page.getByRole("list", { name: "Ordered references" }).locator("li p.font-medium").allTextContents();
expect(names).toEqual(["subject.png", "sound.wav", "motion.mp4"]);
}
await page.screenshot({ path: testInfo.outputPath(`${mode}-inputs.png`), fullPage: true });
await page.getByRole("button", { name: "Generate", exact: true }).click();
await expect.poll(() => sent.find((item) => item.type === "session_init_v2")?.generation_mode).toBe(mode);
const init = sent.find((item) => item.type === "session_init_v2")!;
expect(init.conditioning_assets.map((item: any) => item.role)).toEqual(mode === "t2va" ? [] : mode === "fl2va" ? ["first_frame", "last_frame"] : ["reference", "reference", "reference"]);
if (mode === "ref2va") expect(init.conditioning_assets.map((item: any) => item.asset_id)).toEqual([uploadedIds[0], uploadedIds[2], uploadedIds[1]]);
await expect.poll(() => received.find((item) => item.type === "gpu_assigned")?.generation_mode).toBe(mode);
await expect.poll(() => received.some((item) => item.type === "media_segment_complete")).toBe(true);
await expect(page.getByText(/Demo runtime · Sample playback only/)).toBeVisible();
await expect(modeSelect).toHaveCount(0);
await expect.poll(async () => page.locator("video:visible").first().evaluate((element: HTMLVideoElement) => element.readyState)).toBeGreaterThanOrEqual(2);
await page.screenshot({ path: testInfo.outputPath(`${mode}-playback.png`), fullPage: true });
if (mode === "ref2va") {
await page.getByRole("button", { name: "Toggle sidebar" }).click();
await page.getByRole("button", { name: "New project", exact: true }).click();
await expect(modeSelect).toHaveText("T2VA");
await modeSelect.click();
await page.getByRole("option", { name: "FL2VA", exact: true }).click();
await expect(modeSelect).toHaveText("FL2VA");
await page.getByRole("combobox", { name: "First frame", exact: true }).selectOption({ label: "subject.png" });
await expect(page.getByRole("combobox", { name: "Last frame", exact: true })).toHaveValue("");
await page.getByLabel("Continuation prompt").fill("The paper fox explores a new scene.");
await page.getByRole("button", { name: "Generate", exact: true }).click();
await expect.poll(() => sent.find((item) => item.type === "project_init_v1")?.generation_mode).toBe("fl2va");
const secondProject = sent.find((item) => item.type === "project_init_v1")!;
expect(secondProject.conditioning_assets).toEqual([{ asset_id: uploadedIds[0], role: "first_frame" }]);
expect(sent.filter((item) => item.type === "session_init_v2")).toHaveLength(1);
await expect.poll(() => received.filter((item) => item.type === "media_segment_complete").length).toBeGreaterThan(1);
}
} finally {
await page.close();
for (const id of uploadedIds) await request.delete(`/assets/${id}`);
}
});
}
test("proxies a media upload larger than Next's default 10 MiB body limit", async ({ request }) => {
const response = await request.get("/generation-capabilities");
const capabilities = response.ok() ? await response.json() : {};
test.skip(capabilities.mock !== true, "Requires the local mock runtime.");
const audio = makeAudio(192000, 29);
expect(audio.length).toBeGreaterThan(10 * 1024 * 1024);
const upload = await request.post("/assets", {
headers: { "Content-Type": "audio/wav", "X-Asset-Name": "large-proxy-check.wav" },
data: audio,
});
expect(upload.status()).toBe(201);
const asset = await upload.json();
try { expect(asset.size).toBe(audio.length); } finally { await request.delete(`/assets/${asset.asset_id}`); }
});
});
+14
View File
@@ -9,6 +9,8 @@ const configDir = path.dirname(fileURLToPath(import.meta.url));
const staticExport = process.env.NEXT_OUTPUT_EXPORT === '1';
const nextConfig: NextConfig = {
// Next 15.5 name for the dev rewrite-proxy body limit; Next 16 renames it to `proxyClientMaxBodySize`.
experimental: { middlewareClientMaxBodySize: 100 * 1024 * 1024 },
...(staticExport ? { output: 'export' as const } : {}),
...(staticExport ? { images: { unoptimized: true } } : {}),
outputFileTracingRoot: path.join(configDir, '..', '..', '..'),
@@ -38,6 +40,18 @@ const nextConfig: NextConfig = {
source: '/router/:path*',
destination: `${backendUrl}/router/:path*`
},
{
source: '/generation-capabilities',
destination: `${backendUrl}/generation-capabilities`,
},
{
source: '/assets',
destination: `${backendUrl}/assets`,
},
{
source: '/assets/:path*',
destination: `${backendUrl}/assets/:path*`,
},
{
source: '/prompt-system-config',
destination: `${backendUrl}/prompt-system-config`,
@@ -589,6 +589,7 @@ describe.skip('App websocket integration', () => {
});
const initMessage = outbound.find((message) => message.type === 'session_init_v2');
expect(initMessage.generation_mode).toBe('t2va');
expect(initMessage.preset_id).toBe('test_preset');
expect(initMessage.curated_prompts).toEqual(['segment one', 'segment two']);
expect(initMessage.enhancement_enabled).toBe(true);
@@ -597,6 +598,42 @@ describe.skip('App websocket integration', () => {
expect(initMessage.initial_rollout_prompt).toBe('');
});
it('sends the selected generation mode and locks it after session start', async () => {
const outbound: any[] = [];
server.on('connection', (socket) => {
socket.on('message', (rawMessage) => {
outbound.push(JSON.parse(rawMessage as string));
});
});
const user = userEvent.setup();
render(<Page />);
const modeSelect = await screen.findByRole('combobox', { name: 'Generation mode' });
expect(modeSelect).toHaveTextContent('T2VA');
await user.click(modeSelect);
await user.click(await screen.findByRole('option', { name: 'FL2VA' }));
expect(modeSelect).toHaveTextContent('FL2VA');
expect(modeSelect).toHaveAttribute(
'title',
'First/last frames to video + audio. Start from a first frame image. Add an optional last frame to guide the ending.',
);
const generateButton = await screen.findByRole('button', { name: 'Generate' });
await waitFor(() => expect(generateButton).toBeEnabled());
await user.click(generateButton);
await waitFor(() => {
expect(outbound.some((message) => message.type === 'session_init_v2')).toBe(true);
});
const initMessage = outbound.find((message) => message.type === 'session_init_v2');
expect(initMessage.generation_mode).toBe('fl2va');
expect(screen.queryByRole('combobox', { name: 'Generation mode' }))
.not.toBeInTheDocument();
});
it('starts a streaming session from a custom initial prompt without using curated prompts', async () => {
const outbound: any[] = [];
server.on('connection', (socket) => {
+119 -8
View File
@@ -5,6 +5,7 @@ import { Download, Share2 } from "lucide-react";
import DevtoolsShell from "@/components/devtools/DevtoolsShell";
import MonitorPage from "@/components/MonitorPage";
import ChatBar from "@/components/ChatBar";
import AssetList from "@/components/AssetList";
import SessionTimeoutModal from "@/components/SessionTimeoutModal";
import Sidebar from "@/components/Sidebar";
import Header from "@/components/Header";
@@ -13,10 +14,13 @@ import Workspace from "@/components/Workspace";
import { saveProject, saveProjectMetadata, listProjects, loadProjectClips, deleteProject, pruneOldProjects, type StoredProject, type StoredClip } from "@/lib/projectStorage";
import { isInfrastructureError } from "@/lib/ws/reducer";
import { useStore } from "@/hooks/useStore";
import { useAssetLibrary } from "@/hooks/useAssetLibrary";
import { useGenerationCapabilities } from "@/hooks/useGenerationCapabilities";
import { resolveDevtoolsMode } from "@/lib/devtoolsMode";
import { createAvPipeline, DEFAULT_AV_MIME } from "@/lib/media/avPipeline";
import { remuxArchivedFmp4Segments } from "@/lib/media/fmp4Remux";
import { DEFAULT_CUSTOM_PRESET_ID, parseStoryPresets, sanitizePresetId } from "@/lib/presets";
import { DEFAULT_GENERATION_MODE, buildGenerationInitFields, validateGenerationInputs, type GenerationMode, type GenerationInitFields, type GenerationAsset } from "@/lib/generationMode";
import {
buildRewritePromptWindowSnapshot,
buildRewritePromptWindowSnapshotFromPrompts,
@@ -341,6 +345,19 @@ export default function Page() {
const [isMobileShareCapable, setIsMobileShareCapable] = useState(false);
const [videoMuted, setVideoMuted] = useState(true);
const [timeoutModalOpen, setTimeoutModalOpen] = useState(false);
const [generationMode, setGenerationMode] = useState<GenerationMode>(DEFAULT_GENERATION_MODE);
const assetLibrary = useAssetLibrary();
const { capabilities, capabilityNotice, refreshCapabilities } = useGenerationCapabilities();
const joiningRef = useRef(false);
const activeGenerationRef = useRef<{ fields: GenerationInitFields; assets: GenerationAsset[]; mock: boolean } | null>(null);
const generationInputError = validateGenerationInputs(generationMode, assetLibrary.conditioningAssets, assetLibrary.assets);
const generationSupported = capabilities.modes.includes(generationMode);
const generationInputsValid = !generationInputError && generationSupported && !assetLibrary.uploading;
function changeGenerationMode(mode: GenerationMode) {
if (sessionStore.get().sessionStarted || joiningRef.current || !capabilities.modes.includes(mode)) return;
setGenerationMode(mode);
assetLibrary.clearConditioning();
}
useEffect(() => {
setIsMobileShareCapable(typeof navigator.canShare === "function" && window.matchMedia("(pointer: coarse)").matches);
}, []);
@@ -391,7 +408,7 @@ export default function Page() {
// --- Derived values ---
const canStartSession = !projectResetPending && (canJoinSession || Boolean(normalizeInitialPrompt(livePromptDraft as string)));
const canStartSession = generationInputsValid && !projectResetPending && (canJoinSession || Boolean(normalizeInitialPrompt(livePromptDraft as string)));
const currentClipLabel = useMemo(() => {
if ((activeClip as Record<string, any>)?.label) return (activeClip as Record<string, any>).label;
@@ -747,7 +764,11 @@ export default function Page() {
function recoverFailedSessionStart(notice: string) {
const restoredDraft = normalizeInitialPrompt(pendingInitialPromptRef.current);
resetToLobbyState();
if (wsRef.current) {
detachAndCloseWebSocket(wsRef.current);
wsRef.current = null;
}
resetToLobbyState({ preserveSessionNotice: true });
clearPendingProjectPointers();
pendingInitialPromptRef.current = "";
sessionStore.patch({
@@ -1702,6 +1723,10 @@ export default function Page() {
function resetToLobbyState({ preserveSessionNotice = false, preservePlayback = false } = {}) {
setVideoMuted(true);
if (!preserveSessionNotice) {
setGenerationMode(DEFAULT_GENERATION_MODE);
assetLibrary.clearConditioning();
}
clearCountdownInterval();
pendingInitialPromptRef.current = "";
sessionStore.patch({
@@ -1736,6 +1761,8 @@ export default function Page() {
function resetToProjectLobbyState() {
setVideoMuted(true);
setGenerationMode(DEFAULT_GENERATION_MODE);
assetLibrary.clearConditioning();
pendingInitialPromptRef.current = "";
sessionStore.patch({
sessionStarted: false,
@@ -1764,6 +1791,7 @@ export default function Page() {
setSeedPrompts(segmentPrompts);
return {
type,
...(activeGenerationRef.current?.fields || buildGenerationInitFields(generationMode, assetLibrary.conditioningAssets, assetLibrary.assets)),
preset_id: getInitialPresetId(),
preset_label: getInitialPresetLabel(),
curated_prompts: segmentPrompts,
@@ -1812,6 +1840,12 @@ export default function Page() {
return;
}
if (decoded.kind !== "json") return;
if (decoded.data?.type === "error" && sessionStore.get().sessionStarted
&& (!sessionStore.get().gpuAssigned || decoded.data.error_code === "invalid_generation_input")) {
const message = typeof decoded.data.message === "string" ? decoded.data.message : "The generation inputs were rejected. Check the mode and selected assets.";
recoverFailedSessionStart(message);
return;
}
if (decoded.data?.type === "error" && isInfrastructureError(decoded.data)) {
const message = typeof decoded.data?.message === "string" && decoded.data.message.trim()
? decoded.data.message.trim()
@@ -1937,9 +1971,15 @@ export default function Page() {
}
}
function beginProjectLocally({ force = false } = {}) {
function beginProjectLocally({ force = false, mockRuntime = capabilities.mock === true } = {}) {
if (!force && !canStartSession) return;
if (!generationInputsValid) return false;
if (sessionStore.get().sessionStarted || sessionStore.get().projectResetPending) return false;
activeGenerationRef.current = {
fields: buildGenerationInitFields(generationMode, assetLibrary.conditioningAssets, assetLibrary.assets),
assets: assetLibrary.assets.filter((asset) => assetLibrary.conditioningAssets.some((item) => item.asset_id === asset.asset_id)),
mock: mockRuntime,
};
setTimeoutModalOpen(false);
// Unmute during the user gesture so iOS Safari permits audio playback.
setVideoMuted(false);
@@ -1999,12 +2039,43 @@ export default function Page() {
}
async function joinSession({ force = false } = {}) {
if (joiningRef.current || sessionStore.get().sessionStarted) return;
if (generationInputError || assetLibrary.uploading) {
showPreSessionNotice(generationInputError || "Wait for the asset upload to finish.");
return;
}
joiningRef.current = true;
try {
await startGenerationSession({ force });
} finally {
joiningRef.current = false;
}
}
async function startGenerationSession({ force = false } = {}) {
sessionStore.patch({ sessionNotice: "" });
streamStore.patch({ loadingAnimation: true });
const currentCapabilities = await refreshCapabilities();
if (!currentCapabilities.modes.includes(generationMode)) {
streamStore.patch({ loadingAnimation: false });
showPreSessionNotice(`${generationMode.toUpperCase()} is unavailable on this runtime. Connect a full H3 runtime or choose a supported mode.`);
return;
}
const assetProblem = await assetLibrary.verifySelectedAssets();
if (assetProblem) {
streamStore.patch({ loadingAnimation: false });
showPreSessionNotice(assetProblem);
return;
}
if (
wsRef.current
&& wsRef.current.readyState === WebSocket.OPEN
&& sessionStore.get().connected
) {
if (!beginProjectLocally({ force })) return;
if (!beginProjectLocally({ force, mockRuntime: currentCapabilities.mock === true })) {
streamStore.patch({ loadingAnimation: false });
return;
}
sendProjectInitMessage();
return;
}
@@ -2017,7 +2088,7 @@ export default function Page() {
showPreSessionNotice(probe.notice);
return;
}
if (!beginProjectLocally({ force })) {
if (!beginProjectLocally({ force, mockRuntime: currentCapabilities.mock === true })) {
streamStore.patch({ loadingAnimation: false });
sessionStore.patch({ connecting: false });
return;
@@ -2044,6 +2115,10 @@ export default function Page() {
createdAt: currentProjectCreatedAtRef.current || Date.now(),
lastThumbnail: currentThumbnail,
promptEvents: [...(rewriteStore.get().promptEvents as Record<string, unknown>[])],
generationMode: activeGenerationRef.current?.fields.generation_mode || DEFAULT_GENERATION_MODE,
conditioningAssets: activeGenerationRef.current?.fields.conditioning_assets || [],
assets: activeGenerationRef.current?.assets || [],
mock: activeGenerationRef.current?.mock === true,
};
const clips: StoredClip[] = (streamStore.get().completedClips as any[])
.filter((clip: any) => clip?.blob instanceof Blob)
@@ -2478,6 +2553,24 @@ export default function Page() {
// --- Render ---
const conditioningPanel = generationMode !== "t2va" && !sessionStarted && !sessionExpired ? (
<AssetList
mode={generationMode}
assets={assetLibrary.assets}
conditioning={assetLibrary.conditioningAssets}
locked={Boolean(loadingAnimation || projectResetPending)}
uploading={assetLibrary.uploading}
error={assetLibrary.assetError}
validationNotice={generationInputError}
onUpload={assetLibrary.uploadAssets}
onAssign={assetLibrary.assignAsset}
onRemove={assetLibrary.removeAsset}
onUnselect={assetLibrary.removeConditioning}
onMove={assetLibrary.moveConditioning}
onMissing={assetLibrary.checkAssetAvailability}
/>
) : null;
if (!runtimeReady) {
return null;
}
@@ -2501,13 +2594,17 @@ export default function Page() {
enhancementEnabled={enhancementEnabled as boolean}
autoExtensionEnabled={autoExtensionEnabled as boolean}
loopGenerationEnabled={loopGenerationEnabled as boolean}
canJoinSession={canJoinSession as boolean}
canJoinSession={canStartSession}
canSubmitContinuation={canSubmitContinuation}
editableMode={editableMode as boolean}
demoMode={demoMode as boolean}
editableCanJoin={editableCanJoin as boolean}
curatedPromptLimit={curatedPromptLimit as number}
maxCuratedPromptCount={maxCuratedPromptCount as number}
generationMode={generationMode}
supportedGenerationModes={capabilities.modes}
conditioningPanel={conditioningPanel}
onGenerationModeChange={changeGenerationMode}
onPresetChange={handlePresetSelectionChange}
onEnhancementToggle={handleEnhancementToggle}
onCuratedPromptLimitChange={handleCuratedPromptLimitChange}
@@ -2641,7 +2738,12 @@ export default function Page() {
/>
<Header timeLeft={headerTimeLeft} formatTime={formatTime} onToggleSidebar={() => setSidebarOpen((prev) => !prev)} />
<div className="relative flex flex-1 min-h-0 flex-col justify-center px-4 pb-2 sm:px-6 sm:pb-12">
<div className={cn(
"relative flex flex-1 min-h-0 flex-col px-4 pb-2 sm:px-6 sm:pb-12",
!isViewingMode && !showActiveProject && generationMode !== "t2va"
? "justify-start overflow-y-auto pt-4"
: "justify-center",
)}>
{isViewingMode && (
<>
{viewingSelectedClip && (
@@ -2680,6 +2782,8 @@ export default function Page() {
/>
</section>
<motion.div layout="position" className="mx-auto w-full max-w-2xl shrink-0" transition={{ type: "spring", stiffness: 200, damping: 25 }}>
{viewingProject?.project.mock && <p className="mb-2 text-center text-xs text-violet-600 dark:text-violet-300">Demo sample · This saved clip was not generated by an AI model.</p>}
{viewingProject?.project.generationMode && <p className="mb-3 text-center text-xs text-muted-foreground">{viewingProject.project.generationMode.toUpperCase()} · {viewingProject.project.conditioningAssets?.length || 0} saved references. Uploaded originals may expire; your saved video remains available.</p>}
<ChatBar sessionStarted={false} viewingReadOnly={true} onStartNewProject={handleStartNewProject} onBackFromViewing={closeViewingProject} />
</motion.div>
</>
@@ -2759,7 +2863,7 @@ export default function Page() {
</section>
<AnimatePresence>
{!showActiveProject && (
{!showActiveProject && generationMode === "t2va" && (
<motion.div
key="hero-tagline"
initial={{ opacity: 0 }}
@@ -2785,6 +2889,12 @@ export default function Page() {
sessionExpired={sessionExpired as boolean}
sessionNotice={sessionNotice as string}
projectResetPending={projectResetPending as boolean}
generationMode={generationMode}
supportedGenerationModes={capabilities.modes}
generationInputsValid={generationInputsValid}
capabilityNotice={!generationSupported ? `${generationMode.toUpperCase()} is unavailable on this runtime.` : capabilityNotice}
mockRuntime={capabilities.mock}
conditioningPanel={conditioningPanel}
onPresetGenerate={handlePresetGenerate}
onContinuationInput={handleLivePromptInput}
onContinuationKeydown={handleLivePromptKeydown}
@@ -2792,6 +2902,7 @@ export default function Page() {
onSubmitContinuation={submitLivePrompt}
onLeave={leaveSession}
onStartNewProject={handleStartNewProject}
onGenerationModeChange={changeGenerationMode}
onSpeechTranscript={handleLivePromptSpeechTranscript}
onSpeechInterimChange={handleLivePromptSpeechInterim}
/>
@@ -0,0 +1,41 @@
import { render, screen } from "@testing-library/react";
import userEvent from "@testing-library/user-event";
import { describe, expect, it, vi } from "vitest";
import AssetList from "./AssetList";
import type { GenerationAsset } from "@/lib/generationMode";
const frame: GenerationAsset = { asset_id: "first", kind: "image", name: "frame.png", mime_type: "image/png", size: 2000, url: "/assets/first" };
const video: GenerationAsset = { asset_id: "video", kind: "video", name: "motion.mp4", mime_type: "video/mp4", size: 2000, url: "/assets/video" };
function props() {
return { assets: [frame, video], onUpload: vi.fn(), onAssign: vi.fn(), onRemove: vi.fn(), onUnselect: vi.fn(), onMove: vi.fn(), onMissing: vi.fn() };
}
describe("Asset List", () => {
it("uploads to the library and assigns images to endpoint roles", async () => {
const callbacks = props();
const user = userEvent.setup();
render(<AssetList {...callbacks} mode="fl2va" conditioning={[]} />);
await user.selectOptions(screen.getByRole("combobox", { name: "First frame" }), "first");
expect(callbacks.onAssign).toHaveBeenCalledWith("first", "first_frame");
expect(screen.getByRole("combobox", { name: "Last frame" })).toHaveValue("");
expect(screen.queryByRole("option", { name: "motion.mp4" })).not.toBeInTheDocument();
const file = new File(["image"], "new.png", { type: "image/png" });
await user.upload(screen.getByLabelText("Upload assets", { selector: "input" }), file);
expect(callbacks.onUpload).toHaveBeenCalledWith([file]);
});
it("exposes accessible ordering and removal controls for multimodal references", async () => {
const callbacks = props();
const user = userEvent.setup();
render(<AssetList {...callbacks} mode="ref2va" conditioning={[{ asset_id: "first", role: "reference" }, { asset_id: "video", role: "reference" }]} />);
await user.click(screen.getByRole("button", { name: "Move motion.mp4 up" }));
expect(callbacks.onMove).toHaveBeenCalledWith(1, 0);
await user.click(screen.getByRole("button", { name: "Unselect frame.png" }));
expect(callbacks.onUnselect).toHaveBeenCalledWith(0);
expect(screen.getByRole("button", { name: "Move frame.png up" })).toBeDisabled();
});
it("locks uploads and assignments while starting generation", () => {
render(<AssetList {...props()} mode="fl2va" conditioning={[]} locked />);
expect(screen.getByRole("button", { name: "Upload assets" })).toBeDisabled();
expect(screen.getByRole("combobox", { name: "First frame" })).toBeDisabled();
});
});

Some files were not shown because too many files have changed in this diff Show More