Compare commits

...
Author SHA1 Message Date
SolitaryThinkerandClaude Fable 5.1 2d58c9bfde [test]: sm_100a backward parity and run-to-run stability at 65536 tokens
The existing backward tests stop at 16 kv blocks. The TMEM race fixed in
88319b7 never fired there but corrupted dk/dv on every rerun from 32K
tokens, and the nb >= 1024 device-computed work order had no test at all.
This case (1024 kv blocks, 12.5% density) covers both: sm_100a route vs
all-Triton within the existing tolerances, and two sm_100a runs within
summation-order noise (invert_indices orders each kv row's q list with
atomics, so exact bitwise equality is not available through this path;
measured delta 5e-3 rel_max / 7e-6 mean, the race gave 0.6-1.0 / 4e-2).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G4KbxfFobc4FUXLZV3H4QF
2026-09-06 20:05:25 +00:00
SolitaryThinkerandClaude Fable 5.1 62f0bc1daa [misc]: pin the identity-order/chunked-launch invariant; fix set_extension docstring
The identity work order (nullptr remap) cannot be offset per chunk, so the
DQ_L2_KEEP chunked launches must only run with a device-computed order.
Today that holds because the L2 transition (256K tokens) lies above the
order-kernel threshold (64K), but nothing tied the two constants together;
a static_assert now does, plus a runtime guard for external callers of the
launch API. Also drops the reference to a tests/jit_ext.py that does not
exist in the repo.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G4KbxfFobc4FUXLZV3H4QF
2026-09-06 19:59:19 +00:00
SolitaryThinkerandClaude Fable 5.1 60165c7b65 [fix]: restore tcgen05.wait::ld after the TMEM loads in the sm_100a backward
Reverts the effect of d2c9d61. tcgen05.ld is asynchronous and PTX requires
tcgen05.wait::ld before its destination registers are used. ptxas does
scoreboard the register consumers, so three of the four sites were safe in
practice, but the dQ epilogue arrives on empty_bar_dq right after its two
loads and before any consumer: in SASS the SYNCS.ARRIVE carries an empty
wait mask and precedes the first wait on the loads' barriers (B3/B1). That
hands tmem_dq to the MMA warp while the loads may still be reading it; the
next GEMM overwrites the same columns (tmem_dq aliases tmem_dpt), and the
mixed registers are reduce-added into dqaccum as a valid partial. With the
wait restored ptxas emits `NOP wait={B1}` ahead of the arrive.

Measured cost on GB200: none (0.4-0.9% in favour of the build with waits,
within noise); outputs bitwise identical where the race does not fire.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01G4KbxfFobc4FUXLZV3H4QF
2026-09-06 19:59:19 +00:00
rootandClaude Fable 5.1 d2c9d61240 [kernel] sm_100a backward: drop tcgen05.wait::ld after the TMEM loads
The destination registers of tcgen05.ld are scoreboarded, so the consumers wait on their own; the
four waits (dV/dK merge, S^T, dP^T and dQ loads) are removed, the fence stays. Same change as the
reference kernel in books/. Verified: CPU-reference checks incl. ragged variable_block_sizes,
zero-count kv blocks and STRESS_N=5 bitwise reruns; tests/test_block_sparse_bwd_sm100a.py 15
passed, tests/test_block_sparse_sm100a_dispatch.py 10 passed, tests/test_block_sparse_sm100a.py
37 passed (extension rebuilt from this tree). Interleaved A/B 4k..524k: -1.2..+0.9%, neutral.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 16:25:55 -07:00
rootandClaude Fable 5.1 57bb232432 [kernel] sm_100a backward: drop the cached identity order array, nullptr remap = identity
Review finding: the binding's per-device static cache of a torch::arange identity array was
neither stream-safe (a larger call replaced the tensor while an earlier backward on another
stream could still read it; the replacement arange was only ordered with its own stream) nor
thread-safe (unguarded static std::vector).

Fix: no identity array at all. decode_workitem reads
`real_item_id = workitem_remap ? workitem_remap[workitem_id] : workitem_id`; the launch leaves
workitem_remap nullptr below ORDER_MIN_KV_BLOCKS (one launch of every item, so the work id is
the item id) and runs the order kernel into order_workspace from there on. Slice launches
(>= 524288 tokens) always sit in the order-kernel regime and keep passing `work_remap + base`.
block_sparse_bwd_supported requires the workspace only when the order kernel will run. The
binding passes nullptr always and allocates the workspace from ORDER_MIN_KV_BLOCKS on, as
before; identity_remap and the cache are gone.

Verified: CPU-reference checks incl. odd top-k, ragged variable_block_sizes and zero-count kv
blocks on the nullptr path; device order == host sort check at 1024 kv blocks;
tests/test_block_sparse_bwd_sm100a.py 15 passed, tests/test_block_sparse_sm100a_dispatch.py
10 passed, tests/test_block_sparse_sm100a.py 37 passed (extension rebuilt from this tree).
Interleaved A/B vs the previous build: neutral (4k..262k within +-1.2%).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 12:08:40 -07:00
rootandClaude Fable 5.1 88319b75d5 [kernel] sm_100a backward: keep each softmax warp's bf16 overlay inside its own fp32 columns
Review finding: the two softmax warps of a TMEM lane group split the 128-column S^T (and
dP^T) tile into fp32 halves, but the col_half == 1 warp packed its bf16x2 P^T / dS^T into
columns 32-63, inside the fp32 range the col_half == 0 warp was still loading. Nothing but
the shared full_bar_st / full_bar_dpt wait ordered the two, so a delayed col_half == 0 warp
could read bf16 pairs where it expected fp32 scores.

Fix, no extra synchronization: a warp's bf16x2 overlay now starts at the same column as the
64 fp32 columns it loads (ST_QBLOCK_COLS = 64 per q64 block), so it only overwrites data it
has itself consumed; the dV / dK TS MMAs read slot p's bf16 atoms from column
p * ST_QBLOCK_COLS instead of 32 * p. The dQ GEMM already overwrote the dP^T tile only after
the dK GEMM consumed dS^T (MMA issue order), unchanged.

Verified: CPU-reference checks incl. odd top-k, ragged variable_block_sizes and zero-count kv
blocks; tests/test_block_sparse_bwd_sm100a.py 15 passed, tests/test_block_sparse_sm100a_dispatch.py
10 passed, tests/test_block_sparse_sm100a.py 37 passed (extension rebuilt from this tree).
Interleaved A/B vs the previous build: neutral (4k..131k within noise).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-05 11:29:04 -07:00
rootandClaude Fable 5.1 749a9c5d97 [kernel] sm_100a CUDA backward for VSA block-sparse attention (64-token blocks)
Native GB200 (sm_100a) backward pairing with the sm_100a forward: one
warp-specialized tcgen05 kernel per (batch, head, kv64 block) computing dK, dV
and the dQ partials (fp32 accumulator, cp.reduce.async.bulk), plus a preprocess
(Delta, Q^T, dO^T, dqaccum zero, exact-zero dK/dV rows for unselected kv
blocks) and a postprocess (dQ unscramble + sm_scale). Consumes the forward's
lse (Triton M format) and invert_indices' k2q metadata unchanged; returns
(dq, dk, dv) in bf16 with the Triton backward's scaling.

Wiring: with FASTVIDEO_VSA_SM100A=1 the autograd backward of the sm_100a
forward op runs this kernel when block_sparse_attn_bwd_sm100a.is_supported
passes (bf16, head_dim 128, 64-token blocks, even block count, sm_100a device,
op built) and the Triton backward otherwise. No new environment variable.
primitives.cuh gains only the backward's primitives (additions only; the
forward's definitions and SASS are untouched). Built for sm_100a only: the
kernel is validated on GB200, not yet on B300/GB300, so sm_103a devices keep
the Triton backward. The H3 backend's grad gate is unchanged (separate PR).

Semantics: per-row k2q counts may differ arbitrarily (one work item is one kv
row; count 0 is skipped identically by every warp and gets zero dK/dV rows);
variable_block_sizes masks padded kv rows per lane; is_supported reads no
tensor contents. Work-item order: a device kernel sorts by list length from
1024 kv blocks per sequence on; below that a cached identity array is passed.

Tests (GB200, extension built from this branch): tests/test_block_sparse_bwd_sm100a.py
15 passed (fp32 masked-dense autograd reference; ragged vbs, zero-count kv
blocks, top-k 1/2/3/5/7, batch 2), tests/test_block_sparse_sm100a_dispatch.py
10 passed (sm_100a route vs all-Triton grads, ragged vbs, Triton backward
monkeypatched to fail on a supported input), tests/test_block_sparse_sm100a.py
37 passed (forward regression).

Perf (kernels only, same window, B=1 H=8 D=128, 25% density, fp32 accumulator,
TFLOPS on selected blocks) Triton backward vs this kernel:
4k 198/381 (1.93x), 8k 313/575 (1.84x), 16k 389/715 (1.84x), 32k 433/794
(1.84x), 65k 448/821 (1.83x), 131k 452/736 (1.63x), 262k 452/784 (1.73x),
524k 437/771 (1.76x). FastVideo autograd path incl. invert_indices: 1.6-1.7x
at Wan 480P/720P-like shapes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
2026-09-04 19:39:36 -07:00
SYLAR 7bb76b5ec9 [feat] Add native SM103a VSA support (#1812)
Signed-off-by: lishunyang12 <lishunyang12@163.com>
2026-09-03 15:00:28 -07:00
Aryan KumarandAryan Kumar 0bd19a976b [docs]: add one-Spark FastH3 cookbook runtime with a device-count row (#1811)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-02 10:05:59 -07:00
William Lin 40b93784d2 [docs]: add cookbook link to README (#1810) 2026-09-01 12:42:19 -07:00
Aryan Kumar 33d3478bad [docs] Announce local FastH3 support (#1809) 2026-09-01 12:28:04 -07:00
3d8ac9d14b [feat]: collapse cookbook recipe pages into an accordion layout, add … (#1805)
Co-authored-by: Vaish, Ishan <isvaish@UCSD.EDU>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-01 01:51:51 -07:00
aaef49bfc6 [feat]: run FastH3 across two DGX Sparks with Ray sequence parallel (#1803)
Co-authored-by: Kyle <shh075@ucsd.edu>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Satyam Srivastava <srivastavasatyam53@gmail.com>
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-09-01 01:37:23 -07:00
Aryan KumarandAryan Kumar cf6a00b9be [feat]: add opt-in CUDA TAEH3 preview decode for FastH3 (#1795)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 23:34:56 -07:00
Shahrad ZomorrodiandShahrad Zomorrodi 1ae39562dd [bugfix] Write generated documentation as UTF-8 (#1797)
Co-authored-by: Shahrad Zomorrodi <264690209+shahradzomorrodi@users.noreply.github.com>
2026-08-31 22:45:03 -07:00
Aryan KumarandAryan Kumar 26064193e2 [feat] Add an H3 server cookbook and prompt playground (#1798)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 22:14:33 -07:00
William Lin 8446fc003e [docs]: add FastH3 Preview v1 news links (#1804) 2026-08-31 22:12:02 -07:00
Aryan KumarandAryan Kumar a28f2bab4b [feat] Add an optional MLX TAEH3 preview decoder (#1794)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 05:08:06 -07:00
Aryan KumarandAryan Kumar f82d8be4bf [perf] Sequential MiniMax H3 start with GPU-direct DiT load (#1793)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 04:20:21 -07:00
Aryan KumarandAryan Kumar 8e1775183e [perf] Speed up exact MiniMax H3 MLX inference (#1792)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-31 04:14:55 -07:00
Aryan KumarandAryan Kumar 620bc36dc4 [docs]: Cookbook catalog improvements (#1790)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-30 17:50:57 -07:00
Suhaan Khurana 29ff16ec96 [feat] Add MiniMax H3 MLX spatial fast mode (#1789) 2026-08-30 17:47:02 -07:00
Aryan KumarandAryan Kumar 8f9d76a80d [perf]: dispatch wide-M affine H3 MLX linears through dequant plus dense GEMM (#1788)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-30 14:42:01 -07:00
a4d9a75e2c [perf] Add MiniMax H3 MLX VSA and SIMD attention (#1776)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
2026-08-30 13:44:54 -07:00
KyleNeverGivesUp b2db0c0a13 [ci]: seed stable GB10 grad-norm references (#1756) 2026-08-30 03:09:18 -07:00
Kevin Lin 6aa7d8a278 [misc] FastVideo Studio UI Additions (H3 Ref2V support) (#1783) 2026-08-30 03:06:47 -07:00
Aryan KumarandAryan Kumar ccc9014430 [docs] Add model-family inference cookbook (#1787)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-30 02:26:10 -07:00
William Lin a159b63c67 [bugfix] Harden OpenAI serving after post-merge review (#1782) 2026-08-28 22:29:24 -07:00
ac48bb3cd1 [feat] Add MiniMax H3 MLX T2VA inference (#1770)
Co-authored-by: Aryan Kumar <aryank@Aryans-Mac-Studio.local>
Co-authored-by: coderabbitai[bot] <136622811+coderabbitai[bot]@users.noreply.github.com>
Co-authored-by: CodeRabbit <noreply@coderabbit.ai>
2026-08-28 12:54:33 -07:00
William Lin 3987b9ddcd [feat] Align multimodal OpenAI serving APIs (#1781) 2026-08-28 10:09:58 -07:00
William Lin c7da2f5d60 [chore]: release v0.2.1 (#1778) 2026-08-28 02:03:18 -07:00
William Lin 39ae1decc0 [misc] pin fastvideo-kernel to exact 0.3.5 (#1777) 2026-08-28 02:02:27 -07:00
William Lin 1aed667377 [chore] release fastvideo-kernel 0.3.5 (#1775) 2026-08-27 23:18:59 -07:00
William Lin c1612ff397 [bugfix]: pin fastvideo-kernel to Torch 2.12.0 (#1774) 2026-08-27 23:13:54 -07:00
William Linandshaoxiongduan a534ba20a0 [feat] Add MiniMax H3 LoRA inference and preview launchers (#1771)
Co-authored-by: shaoxiongduan <shaoxiongduan@gmail.com>
2026-08-27 14:43:39 -07:00
KyleNeverGivesUp e9bbaca07d [perf] Disable every offload path on unified memory, unblocking MiniMax H3 generation on one GB10 (#1715) 2026-08-26 15:32:56 -07:00
KyleNeverGivesUp 9bfa585448 [perf]: stop holding the whole checkpoint during DiT load, unblocking MiniMax H3 on one GB10 (#1714) 2026-08-26 15:23:51 -07:00
Raghav K b2062556a9 [perf] VSA Triton: widen the autotune num_stages range (the optimum was outside it) (#1706) 2026-08-26 12:30:56 -07:00
KyleNeverGivesUp c9c5585758 [perf]: MiniMax H3 on GB10 - skip text encoder CPU offload on unified memory (5m49s to 30ms) (#1710) 2026-08-26 12:03:14 -07:00
William Lin 9212f4f218 [ci] make GPU validation change-aware (#1747) 2026-08-25 21:26:25 -07:00
Aryan KumarandAryan Kumar 6388db815b [bugfix] FastMetal-QAD MLX support: refuse CUDA QAD trees, use packed mlx_dit config, stream loads (#1736) (#1758)
Co-authored-by: Aryan Kumar <aryan5v@users.noreply.github.com>
2026-08-25 14:51:40 -07:00
lpc0220 7a4285189f [kernel] Route block-sparse VSA to the sm_100a forward behind FASTVIDEO_VSA_SM100A (opt-in) (#1754) 2026-08-24 15:26:56 -07:00
William Lin a837fe841a [docs] Update FastH3 README (#1749) 2026-08-23 05:05:58 -07:00
William Lin f9e3680f11 [perf] Align FastH3 optimized inference profile (#1748) 2026-08-23 02:01:03 -07:00
William Lin 98f761ec45 [bugfix] validation: inherit the trained denoising ladder (#1738) 2026-08-22 23:09:26 -07:00
Shao Duan c041318f2c [perf] Add fused NVLink all-to-all for Ulysses (#1740) 2026-08-22 23:06:24 -07:00
William Lin 604e0205a4 [perf] Keep odd MiniMax-H3 VSA tiles on sm100a (#1745) 2026-08-22 18:29:30 -07:00
William Lin 13213395b4 [perf] Parallelize MiniMax-H3 VAE over sequence ranks (#1744) 2026-08-22 18:05:31 -07:00
William Lin 46afee5998 [bugfix] Classify MiniMax-H3 inference controls in schema inventory (#1743) 2026-08-22 17:17:02 -07:00
William Lin c488fa1211 [perf] Add opt-in packed-varlen FA4 for MiniMax-H3 (#1742) 2026-08-22 17:16:47 -07:00
William Lin d3cff517cd [perf] Add opt-in regional fullgraph compile for DiT inference (#1741) 2026-08-22 12:14:39 -07:00
Junda Su 2f3d407406 [perf] Optimize MiniMax H3 VAE decoding (#1734) 2026-08-21 14:57:32 -07:00
Kaiqin Kong bcffa4026e [perf] Optimize MiniMax-H3 text encoder memory (#1732) 2026-08-21 14:57:06 -07:00
William Lin 6d6a10be7a [feat] FastVideo-Minimax-FastH3-Preview few-step example + 64-token-tile VSA-H3 inference path (#1731) 2026-08-21 12:40:09 -05:00
Kaiqin Kong 73dd105f3d [perf] Add opt-in MiniMax-H3 Sol-Engine fusions (#1735) 2026-08-21 12:39:28 -05:00
362 changed files with 43537 additions and 2580 deletions
+99
View File
@@ -0,0 +1,99 @@
---
name: ci-runner
description: Work on FastVideo's Slurm-only, change-aware GPU CI lanes, static Buildkite graph, trusted ci-runner policy, lane scripts, and GB200 validation.
---
# Slinky Slurm CI lanes
FastVideo's `ci-runner` Buildkite queue is the control plane for all active
GPU CI. A private host-owned dispatcher leases GPUs from the Slinky Slurm tray
and runs the immutable PR SHA inside an isolated Enroot container. Buildkite
pipeline upload and Slurm submission occur on the login plane; every test
payload executes on Slurm compute.
The files under `fastvideo/tests/modal/` and `.buildkite/scripts/pr_test.sh`
are dormant rollback code. Never add an active Buildkite or slash-command
route to them. `pr_test.sh` must continue to reject Buildkite invocations.
The private operator bundle is deliberately outside this repository because
it contains site paths and credentials. See
`docs/contributing/ci_architecture.md`; this skill covers the repository half
and the coordination contract with that bundle.
## Invariants
- `.buildkite/pipeline.yml` contains exactly one static step for every active
GPU lane. Each step pins a unique key and label, a 90-minute timeout, the
trusted `/opt/fastvideo-ci-runner/run-ci` command (`run-unit` is the one
compatibility wrapper), step-level internal `TEST_TYPE`, and
`queue: "ci-runner"`.
- Active CI contains no `pr_test.sh` command, Modal invocation, default queue,
Buildkite plugin, `soft_fail`, or job-controlled artifact glob.
- The six Fastcheck lanes use `:microscope:` labels. Full-Suite-only lanes use
`:test_tube:` or `:bar_chart:` so direct reruns update the right aggregate.
- SSIM and vanilla training request all four GPUs. Keep both in the
`fastvideo/slinky/whole-tray` Buildkite concurrency group with a limit of one
so the second job does not consume an agent or command timeout while waiting
for the same tray.
- `/test full` schedules all twenty lanes. `/merge`, `ready`, and new pushes to
ready PRs use the trusted base-branch planner in
`.github/scripts/plan_merge_ci.py`: automatic Fastcheck remains the universal
six-lane baseline, and the merge build adds only path-relevant integration
lanes. Unknown source/build paths fail closed to all fourteen additive lanes.
The trusted uploader still normalizes and validates the complete static graph
before Buildkite evaluates its plan conditions.
- Focused merge builds may pass allowlisted golden-gate and SSIM test basenames.
The private host validates the lane plan and basenames before staging them,
and the in-container scripts validate them again. Direct `/test ssim`,
explicit `/test full`, and the weekly main-branch schedule run the complete
SSIM matrix.
- The trusted uploader serves exactly three entry pipelines:
`pr-fastcheck` for automatic PR builds, `ci` for slash-command/ready-label
API builds, and `fastvideo-performance-lane` for the weekly schedule. Keep
incoming GitHub webhook processing disabled on `ci` so it cannot duplicate
`pr-fastcheck` on every PR update.
- Test payloads live in `.buildkite/scripts/unit_test.sh` or executable
`.buildkite/scripts/lanes/<lane>.sh`. Backend policy (GPU count, extras,
secrets, kernel build, artifacts) stays in the agent-owned lane table.
- Tests must preserve an inherited `MASTER_PORT`. Packed containers share the
tray network namespace, so the private runner assigns a distinct port range
per GPU lease and the SSIM scheduler assigns task offsets within its range.
- The ARM64 runner image includes the pinned FA4 CuTe overlay validated on
GB200. Keep SSIM at `FASTVIDEO_FA4=1` because its references were seeded with
FA4; keep lanes with FA2 baselines at `FASTVIDEO_FA4=0`. A runner image change
must revalidate both the FA4 import and an actual GB200 forward kernel.
- `fastvideo/tests/ssim/ci_runner.py` is the active four-GPU SSIM scheduler.
New SSIM files are discovered through `REQUIRED_GPUS` and
`*_MODEL_TO_PARAMS`; do not wire them through the dormant Modal scheduler.
- The host policy fail-closes unknown tuples. A repository-side lane change is
inert until the operator updates the private lane table and uploader policy
in the same rollout.
## Adding or changing a lane
1. Read the closest `AGENTS.md` and the domain-specific testing guide.
2. Add or update the executable lane payload under `.buildkite/scripts/`.
Keep it deterministic and free of host-specific paths or credential fetches.
3. Add the static pipeline step and canonical `/test <name>` mapping. Keep the
`<name>-ci` alias only when compatibility requires it.
4. Add its source/test path ownership to `.github/scripts/plan_merge_ci.py`.
Prefer the narrowest correctness-preserving lane set; leave unknown paths
fail-closed. Extend `fastvideo/tests/contract/test_ci_test_collection.py`,
`test_merge_ci_plan.py`, and focused CPU-only scheduler/policy tests.
5. Coordinate the private lane row: GPU count (1-4), wall time, script, scope
pairs, step key, command, HF cache/token, tracking mode, extras, attention
backend policy, kernel policy, and artifact relay. Active training lanes
keep W&B offline and do not stage a W&B credential.
6. Update the trusted pipeline-uploader schema. A mismatch must reject the
pipeline rather than silently skip a lane.
7. Run `pre-commit run --files <changed paths>`, the planner's representative
diff matrix, contract tests, private driver tests, and a real GB200 canary.
Multi-GPU, hardware-reference, training, performance, and SSIM changes need
their own target-hardware evidence.
## Rollback
Rollback the Slurm routing/configuration change or pause the `ci-runner` queue.
Do not silently reactivate Modal. A manual Modal experiment requires the
explicit local opt-in documented in `ci_architecture.md`; returning it to
production CI needs a separate reviewed decision.
@@ -1,6 +1,6 @@
---
name: reseed-ssim-references
description: Re-seed HF reference videos for a single existing SSIM test on Modal L40S. Always backs up current refs locally first, regenerates on Modal, pauses for the user to eyeball before-vs-after quality, then overwrites the targeted `<model_id>` subtree on `FastVideo/ssim-reference-videos` with `--force`. Use when an intentional code change (model port fix, attention backend swap, kernel upgrade, hyperparameter change) has invalidated existing refs and they need to be regenerated. Pairs with `seed-ssim-references`, which is for first-time seeding only.
description: Re-seed HF reference videos for a single existing SSIM test on Modal L40S. Always backs up current refs locally first, regenerates on Modal, pauses for the user to eyeball before-vs-after quality, then overwrites the targeted model subtree on `FastVideo/ssim-reference-videos` with `--force`. Use when an intentional code change (model port fix, attention backend swap, kernel upgrade, hyperparameter change) has invalidated existing refs and they need to be regenerated. Pairs with `seed-ssim-references`, which is for first-time seeding only.
---
# Re-seed SSIM Reference Videos
@@ -13,7 +13,7 @@ on HF — the old refs are overwritten — so the skill always:
1. Confirms intent with a one-liner the user has to type.
2. Downloads the existing refs as a local, timestamped backup.
3. Regenerates on Modal L40S (same code path that CI uses).
3. Regenerates through the manual legacy Modal L40S maintenance path.
4. Pauses for a side-by-side eyeball of backup vs new mp4s.
5. Uploads with `--force`, scoped to the single `--model-id`.
6. Reminds the user to keep the backup until the PR lands.
@@ -51,8 +51,9 @@ harder to recover from than failing closed.
Hardcoded:
- Modal GPU: **L40S** (matches CI; re-seeding from another SKU produces refs
that L40S CI cannot match).
- Modal GPU: **L40S**. This is a manual reference-maintenance target, not the
active Slurm CI compute path; changing the SKU also changes the historical
`L40S_reference_videos` contract.
- Quality tier: **`default`**. `full_quality` is a separate, deliberate
operation.
- HF repo: `FastVideo/ssim-reference-videos` (override via
+4 -2
View File
@@ -35,7 +35,8 @@ The skill is run **manually**, once per new test. Before invoking it, the user
has already sanity-tested the new test locally — it launches `VideoGenerator`
and writes an artefact without crashing (the missing-reference assertion at
the end is expected). The skill does not re-test locally; it goes straight
to Modal L40S (which is what CI uses).
to the manual legacy Modal L40S reference-maintenance target. Active CI runs
on the Slinky Slurm cluster and only consumes the resulting references.
## When to use
@@ -61,7 +62,8 @@ Prompt the user for it if they didn't supply it.
Everything else is fixed:
- Modal runner GPU: **L40S** (hardcoded in `fastvideo/tests/modal/ssim_test.py`).
- Modal maintenance GPU: **L40S** (hardcoded in
`fastvideo/tests/modal/ssim_test.py`; this is not the active CI compute path).
- Device folder: `L40S_reference_videos`.
- Quality tier: `default` (the tier CI runs). The `full_quality` tier is not
seeded by this skill.
+448 -528
View File
@@ -1,7 +1,8 @@
env:
IMAGE_VERSION: "py3.12-latest"
BUILDKITE_CLEAN_CHECKOUT: true
# Buildkite only launches Modal; remote jobs initialize their own submodules.
# Slurm workers clone the immutable commit and initialize submodules inside
# their isolated container. The Buildkite login-plane checkout is a no-op.
BUILDKITE_GIT_SUBMODULES: false
notify:
@@ -10,539 +11,458 @@ notify:
if: build.env("TEST_SCOPE") == "fastcheck" || build.env("TEST_SCOPE") == null
- github_commit_status:
context: "full-suite-passed"
if: build.env("TEST_SCOPE") == "full"
if: build.env("TEST_SCOPE") == "full" || build.env("TEST_SCOPE") == "merge"
- github_commit_status:
context: "direct-test-completed"
if: build.env("TEST_SCOPE") == "direct"
- github_commit_status:
context: "scheduled-ssim-passed"
if: build.env("TEST_SCOPE") == "scheduled"
# This is the complete active GPU CI surface. Every command is a trusted host
# dispatcher, and every test payload executes inside the Slinky Slurm tray.
# fastvideo/tests/modal remains available only for an explicit manual rollback;
# no active pipeline or slash-command route invokes it.
steps:
# ============================================================
# Direct test: triggered by /test <name> slash command.
# Labels match fastcheck/full-suite counterparts so the GitHub
# check status overwrites the original failed check.
# Only ONE step executes per build (gated by TEST_TYPE).
# ============================================================
- label: ":microscope: Encoder Tests"
key: "encoder"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,encoder,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "encoder" || build.env("TEST_TYPE") == "encoder_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "encoder_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
# --- Fastcheck-scope direct tests ---
- label: ":microscope: Encoder Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "encoder"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: VAE Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "vae"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Transformer Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "transformer"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Kernel Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "kernel_tests"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":vertical_traffic_light: Golden-Gate Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "golden_gate"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Unit Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "unit_test"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: DreamVerse App Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "dreamverse_app"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: VAE Tests"
key: "vae"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,vae,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "vae" || build.env("TEST_TYPE") == "vae_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "vae_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
# --- Full-suite-scope direct tests ---
- label: ":bar_chart: SSIM Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "ssim"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "default"
- label: ":test_tube: LoRA Inference Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "inference_lora"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: LoRA Extraction Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "lora_extraction"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Training Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Distillation DMD Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "distillation_dmd"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Self-Forcing Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "self_forcing"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: LoRA Training Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training_lora"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Training Tests VSA"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "training_vsa"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Inference Tests VMoBA"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "inference_vmoba"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Performance Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "performance"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: API Server Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "api_server"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Train Framework Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "train_framework"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":test_tube: Eval Metrics Tests"
if: build.env("TEST_SCOPE") == "direct" && build.env("TEST_TYPE") == "eval"
command: "timeout 90m .buildkite/scripts/pr_test.sh"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "default"
- label: ":microscope: Transformer Tests"
key: "transformer"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,transformer,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "transformer" || build.env("TEST_TYPE") == "transformer_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "transformer_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
# ============================================================
# Fastcheck: Runs on every PR (~10-15 min parallel)
# Core component validation: encoders, VAEs, transformers,
# CUDA kernels, and unit tests.
# ============================================================
- label: "Trigger Fastcheck"
if: build.env("TEST_SCOPE") == "fastcheck" || build.env("TEST_SCOPE") == null
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
plugins:
- monorepo-diff#v1.4.0:
diff: 'git fetch origin "${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}" && git diff --name-only "origin/${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}...HEAD"'
watch:
- path:
- "fastvideo/models/encoders/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/encoders/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: ":microscope: Encoder Tests"
env:
- TEST_TYPE=encoder
agents:
queue: "default"
- path:
- "fastvideo/models/vaes/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/vaes/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: ":microscope: VAE Tests"
env:
- TEST_TYPE=vae
agents:
queue: "default"
- path:
- "fastvideo/models/dits/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/transformers/**"
- "fastvideo/layers/**"
- "fastvideo/attention/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Transformer Tests"
env:
- TEST_TYPE=transformer
agents:
queue: "default"
- path:
- "fastvideo-kernel/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Kernel Tests"
env:
- TEST_TYPE=kernel_tests
agents:
queue: "default"
- path:
- "fastvideo/**"
- ".buildkite/**"
- ".github/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":microscope: Unit Tests"
env:
- TEST_TYPE=unit_test
agents:
queue: "default"
- path:
- "apps/dreamverse/**"
- "pyproject.toml"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":microscope: DreamVerse App Tests"
env:
- TEST_TYPE=dreamverse_app
agents:
queue: "default"
- label: ":microscope: Kernel Tests"
key: "kernel-tests"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,kernel-tests,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "kernel_tests" || build.env("TEST_TYPE") == "kernel_tests_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "kernel_tests_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
# ============================================================
# Full Suite: Runs when TEST_SCOPE=full
# Triggered by adding the 'ready' label (via ci-trigger-full-suite.yml)
# or on-demand via /test full slash command.
# Includes integration tests, SSIM regression, training pipelines,
# and performance benchmarks.
# ============================================================
- label: "Trigger Full Suite"
if: build.env("TEST_SCOPE") == "full"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
plugins:
- monorepo-diff#v1.4.0:
diff: 'git fetch origin "${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}" && git diff --name-only "origin/${BUILDKITE_PULL_REQUEST_BASE_BRANCH:-main}...HEAD"'
watch:
- path:
- "fastvideo/**/*.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":bar_chart: SSIM Tests"
env:
- TEST_TYPE=ssim
retry:
automatic:
- exit_status: 1
limit: 2
agents:
queue: "default"
- path:
- "fastvideo/tests/lora/**"
- "fastvideo/models/loader/**"
- "fastvideo/tests/transformers/**"
- "fastvideo/pipelines/**"
- "fastvideo/layers/lora/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 20m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Inference Tests"
env:
- TEST_TYPE=inference_lora
agents:
queue: "default"
- path:
- "scripts/lora_extraction/**"
- "fastvideo/tests/lora_extraction/**"
- "fastvideo/models/loader/**"
- "fastvideo/training/training_utils.py"
- "fastvideo/layers/lora/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Extraction Tests"
env:
- TEST_TYPE=lora_extraction
agents:
queue: "default"
- path:
- "fastvideo/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 25m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Training Tests"
env:
- TEST_TYPE=training
agents:
queue: "default"
- path:
- "fastvideo/training/*distillation_pipeline.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 25m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Distillation DMD Tests"
env:
- TEST_TYPE=distillation_dmd
agents:
queue: "default"
- path:
- "fastvideo/training/*self_forcing_distillation_pipeline.py"
- "fastvideo/tests/training/self-forcing/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Self-Forcing Tests"
env:
- TEST_TYPE=self_forcing
agents:
queue: "default"
- path:
- "fastvideo/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 25m .buildkite/scripts/pr_test.sh"
label: ":test_tube: LoRA Training Tests"
env:
- TEST_TYPE=training_lora
retry:
automatic:
- exit_status: 1
limit: 2
agents:
queue: "default"
- path:
- "fastvideo/**"
- "fastvideo-kernel/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Training Tests VSA"
env:
- TEST_TYPE=training_vsa
retry:
automatic:
- exit_status: 1
limit: 2
agents:
queue: "default"
- path:
- "fastvideo-kernel/**"
- "fastvideo/attention/backends/vmoba.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 15m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Inference Tests VMoBA"
env:
- TEST_TYPE=inference_vmoba
agents:
queue: "default"
- path:
- "fastvideo/models/dits/**"
- "fastvideo/pipelines/**"
- "fastvideo/attention/**"
- "fastvideo/layers/**"
- "fastvideo/worker/**"
- "fastvideo/entrypoints/**"
- "fastvideo/performance/**"
- "fastvideo/tests/performance/**"
- ".buildkite/performance-benchmarks/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Performance Tests"
env:
- TEST_TYPE=performance
agents:
queue: "default"
- path:
- "fastvideo/entrypoints/openai/**"
- "fastvideo/entrypoints/cli/serve.py"
- "fastvideo/tests/entrypoints/test_openai_api_integration.py"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: API Server Tests"
env:
- TEST_TYPE=api_server
agents:
queue: "default"
- path:
- "fastvideo/train/**"
- "fastvideo/tests/train/models/**"
- "fastvideo/tests/train/fixtures/**"
- "fastvideo/models/dits/**"
- "fastvideo/models/loader/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 30m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Train Framework Tests"
env:
- TEST_TYPE=train_framework
agents:
queue: "default"
- path:
- "fastvideo/eval/**"
- "fastvideo/tests/eval/**"
- "pyproject.toml"
- "docker/Dockerfile"
config:
command: "timeout 90m .buildkite/scripts/pr_test.sh"
label: ":test_tube: Eval Metrics Tests"
env:
- TEST_TYPE=eval
agents:
queue: "default"
- label: ":microscope: Unit Tests"
key: "unit"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,unit,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "unit_test" || build.env("TEST_TYPE") == "unit_test_ci"))
command: "/opt/fastvideo-ci-runner/run-unit"
timeout_in_minutes: 90
env:
TEST_TYPE: "unit_test_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":microscope: DreamVerse App Tests"
key: "dreamverse"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,dreamverse,/) ||
build.env("TEST_SCOPE") == "fastcheck" ||
build.env("TEST_SCOPE") == null ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "dreamverse_app" || build.env("TEST_TYPE") == "dreamverse_app_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "dreamverse_app_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Golden-Gate Tests"
key: "golden-gate"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,golden-gate,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "golden_gate" || build.env("TEST_TYPE") == "golden_gate_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "golden_gate_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":bar_chart: SSIM Tests"
key: "ssim"
if: |
build.env("TEST_SCOPE") == "full" ||
build.env("TEST_SCOPE") == "scheduled" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,ssim,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "ssim" || build.env("TEST_TYPE") == "ssim_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
concurrency: 1
concurrency_group: "fastvideo/slinky/whole-tray"
env:
TEST_TYPE: "ssim_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: LoRA Inference Tests"
key: "lora-inference"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,lora-inference,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "inference_lora" || build.env("TEST_TYPE") == "inference_lora_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "inference_lora_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: LoRA Extraction Tests"
key: "lora-extraction"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,lora-extraction,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "lora_extraction" || build.env("TEST_TYPE") == "lora_extraction_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "lora_extraction_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Training Tests"
key: "training"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,training,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "training" || build.env("TEST_TYPE") == "training_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
concurrency: 1
concurrency_group: "fastvideo/slinky/whole-tray"
env:
TEST_TYPE: "training_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Distillation DMD Tests"
key: "distillation"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,distillation,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "distillation_dmd" || build.env("TEST_TYPE") == "distillation_dmd_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "distillation_dmd_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Self-Forcing Tests"
key: "self-forcing"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,self-forcing,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "self_forcing" || build.env("TEST_TYPE") == "self_forcing_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "self_forcing_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: LoRA Training Tests"
key: "lora-training"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,lora-training,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "training_lora" || build.env("TEST_TYPE") == "training_lora_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "training_lora_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Training Tests VSA"
key: "training-vsa"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,training-vsa,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "training_vsa" || build.env("TEST_TYPE") == "training_vsa_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "training_vsa_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
- exit_status: 1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Inference Tests VMoBA"
key: "inference-vmoba"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,inference-vmoba,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "inference_vmoba" || build.env("TEST_TYPE") == "inference_vmoba_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "inference_vmoba_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Performance Tests"
key: "performance"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,performance,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "performance" || build.env("TEST_TYPE") == "performance_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "performance_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: API Server Tests"
key: "api-server"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,api-server,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "api_server" || build.env("TEST_TYPE") == "api_server_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "api_server_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Train Framework Tests"
key: "train-framework"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,train-framework,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "train_framework" || build.env("TEST_TYPE") == "train_framework_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "train_framework_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
- label: ":test_tube: Eval Metrics Tests"
key: "eval"
if: |
build.env("TEST_SCOPE") == "full" ||
(build.env("TEST_SCOPE") == "merge" &&
build.env("MERGE_TEST_PLAN") =~ /,eval,/) ||
(build.env("TEST_SCOPE") == "direct" &&
(build.env("TEST_TYPE") == "eval" || build.env("TEST_TYPE") == "eval_ci"))
command: "/opt/fastvideo-ci-runner/run-ci"
timeout_in_minutes: 90
env:
TEST_TYPE: "eval_ci"
retry:
automatic:
- exit_status: 128
limit: 3
- exit_status: -1
limit: 2
agents:
queue: "ci-runner"
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the OpenAI-compatible API lane.
set -euo pipefail
exec pytest ./fastvideo/tests/entrypoints/test_openai_api_integration.py -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the distillation-DMD lane.
set -euo pipefail
exec pytest ./fastvideo/tests/training/distill/test_distill_dmd.py -vs
+87
View File
@@ -0,0 +1,87 @@
#!/usr/bin/env bash
# DreamVerse needs a GPU for import-time device resolution, but it does not
# build or exercise fastvideo-kernel. A checksummed Node archive is installed
# in the disposable Slurm container because the shared CI image is
# Python/CUDA focused.
set -euo pipefail
node_version=v22.23.2
case $(uname -m) in
aarch64 | arm64)
node_arch=arm64
node_archive_sha256=013b59cfd2819703a6f4a14ab891fc46fc2a4e3f5bcd92de3fb4929b43e35b30
;;
x86_64 | amd64)
node_arch=x64
node_archive_sha256=b294a556e639d64338823920e5866c21c02741742d2e1529ee1a225c1ec9252a
;;
*)
echo "Unsupported architecture for DreamVerse Node runtime: $(uname -m)" >&2
exit 2
;;
esac
node_archive="node-${node_version}-linux-${node_arch}.tar.gz"
node_runtime_root=$(mktemp -d -t fastvideo-node.XXXXXX)
node_archive_path="${node_runtime_root}/${node_archive}"
node_install_dir="${node_runtime_root}/${node_archive%.tar.gz}"
curl --proto '=https' --tlsv1.2 --retry 5 --retry-all-errors \
--location --fail --silent --show-error \
"https://nodejs.org/dist/${node_version}/${node_archive}" \
--output "$node_archive_path"
printf '%s %s\n' "$node_archive_sha256" "$node_archive_path" | sha256sum --check --status
tar -xzf "$node_archive_path" -C "$node_runtime_root"
export PATH="${node_install_dir}/bin:${PATH}"
node --version
npm --version
export PYTHONPATH="$(pwd)/apps/dreamverse${PYTHONPATH:+:$PYTHONPATH}"
pytest apps/dreamverse/dreamverse/tests -q
cd apps/dreamverse/web
npm ci
npm run typecheck
npm test
machine_arch=$(uname -m)
if [[ $machine_arch =~ ^(aarch64|arm64)$ ]]; then
npx playwright install --with-deps chromium firefox
else
npx playwright install --with-deps chromium webkit firefox
fi
master_port=${MASTER_PORT:-7959}
BACKEND_PORT=${BACKEND_PORT:-$((master_port + 50))}
python -m uvicorn dreamverse.mock_server:app --host 127.0.0.1 --port "$BACKEND_PORT" &
mock_server_pid=$!
cleanup() {
kill "$mock_server_pid" 2>/dev/null || true
wait "$mock_server_pid" 2>/dev/null || true
}
trap cleanup EXIT INT TERM
for _ in {1..30}; do
curl -fsS "http://127.0.0.1:$BACKEND_PORT/healthz" && break
sleep 1
done
curl -fsS "http://127.0.0.1:$BACKEND_PORT/healthz"
if [[ $machine_arch =~ ^(aarch64|arm64)$ ]]; then
# Playwright WebKit traps before opening a page on Linux ARM64, and its
# bundled Chromium lacks the H.264/AAC codecs used by the fMP4 assertions.
# Firefox covers every flow, including streaming. Chromium and its mobile
# profile still cover all codec-independent UI behavior on GB200.
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
npm run e2e -- --project=firefox
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
npm run e2e -- \
--project=chromium \
--project=mobile-chromium \
--grep-invert='streams, plays, and surfaces a downloadable clip|starts a new project and switches back to the prior session|saved projects persist across a page reload'
else
BACKEND_HOST=127.0.0.1 BACKEND_PORT="$BACKEND_PORT" CI=1 \
npm run e2e -- \
--project=chromium \
--project=webkit \
--project=firefox \
--project=mobile-safari \
--project=mobile-chromium
fi
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the encoder lane.
set -euo pipefail
exec pytest ./fastvideo/tests/encoders -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the evaluation lane.
set -euo pipefail
exec pytest ./fastvideo/tests/eval -vs
+35
View File
@@ -0,0 +1,35 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the golden-gate lane. Environment (HF_HOME
# and authentication) is the runner's responsibility.
set -euo pipefail
golden_root=./fastvideo/tests/golden_gate
selected=${FASTVIDEO_GOLDEN_TEST_FILES-}
if [ -z "$selected" ]; then
if [ "${TEST_SCOPE:-}" = merge ]; then
echo "Missing FASTVIDEO_GOLDEN_TEST_FILES for merge scope" >&2
exit 2
fi
selected=all
fi
if [ "$selected" = all ]; then
exec pytest "$golden_root" -vs
fi
[[ $selected =~ ^test_[a-z0-9_]+\.py(,test_[a-z0-9_]+\.py)*$ ]] || {
echo "Invalid FASTVIDEO_GOLDEN_TEST_FILES selection" >&2
exit 2
}
IFS=, read -r -a golden_files <<< "$selected"
golden_paths=()
for golden_file in "${golden_files[@]}"; do
golden_path="$golden_root/$golden_file"
[ -f "$golden_path" ] || {
echo "Selected golden test does not exist: $golden_file" >&2
exit 2
}
golden_paths+=("$golden_path")
done
exec pytest "${golden_paths[@]}" -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the LoRA-inference lane.
set -euo pipefail
exec pytest ./fastvideo/tests/inference/lora/test_lora_inference_similarity.py -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the VMoBA-inference lane.
set -euo pipefail
exec python fastvideo/tests/inference/vmoba/test_vmoba_inference.py
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the custom-kernel lane.
set -euo pipefail
exec pytest fastvideo-kernel/tests/ -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the LoRA-extraction lane.
set -euo pipefail
exec pytest ./fastvideo/tests/lora_extraction/test_lora_extraction.py -vs
+52
View File
@@ -0,0 +1,52 @@
#!/usr/bin/env bash
# Canonical Slurm performance lane. Reports are written outside the checkout
# so the trusted host driver can upload them after untrusted code exits.
set -uo pipefail
export PERFORMANCE_TRACKING_ROOT=/tmp/perf-tracking
export PERF_REPORTS_DIR=/workspace/artifacts/performance
mkdir -p "$PERF_REPORTS_DIR"
if [[ ${BUILDKITE_PULL_REQUEST:-false} =~ ^[1-9][0-9]*$ ]]; then
export PERF_RUN_SOURCE=pr
export PERF_UPLOAD_POLICY=pass
elif [ "${BUILDKITE_BRANCH:-}" = main ] \
&& { [ "${BUILDKITE_SOURCE:-}" = schedule ] || [ "${TEST_SCOPE:-}" = full ]; }; then
export PERF_RUN_SOURCE=scheduled_main
export PERF_UPLOAD_POLICY=always
elif [ "${TEST_SCOPE:-}" = direct ]; then
export PERF_RUN_SOURCE=unknown
export PERF_UPLOAD_POLICY=pass
else
export PERF_RUN_SOURCE=unknown
export PERF_UPLOAD_POLICY=never
fi
nvidia-smi \
--query-gpu=index,timestamp,clocks.sm,clocks.max.sm,power.draw,power.limit,temperature.gpu \
--format=csv -l 10 > "$PERF_REPORTS_DIR/gpu_telemetry.csv" 2>/dev/null &
telemetry_pid=$!
cleanup() {
kill "$telemetry_pid" 2>/dev/null || true
wait "$telemetry_pid" 2>/dev/null || true
}
trap cleanup EXIT INT TERM
pytest ./fastvideo/tests/performance -vs
pytest_rc=$?
compare_rc=0
if [ "$pytest_rc" -eq 0 ] || [ "$PERF_UPLOAD_POLICY" = always ]; then
PERF_PYTEST_RC=$pytest_rc python ./fastvideo/tests/performance/compare_baseline.py
compare_rc=$?
fi
python ./fastvideo/tests/performance/dashboard.py || true
cp -f fastvideo/tests/performance/results/*.json "$PERF_REPORTS_DIR/" 2>/dev/null || true
echo "--- GPU telemetry (clocks.sm vs clocks.max.sm reveals capped hosts) ---"
cat "$PERF_REPORTS_DIR/gpu_telemetry.csv" || true
final_rc=$pytest_rc
if [ "$final_rc" -eq 0 ]; then
final_rc=$compare_rc
fi
exit "$final_rc"
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the self-forcing lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/self-forcing/test_self_forcing.py -vs
+40
View File
@@ -0,0 +1,40 @@
#!/usr/bin/env bash
# Canonical four-GPU SSIM lane for the Slinky Slurm worker.
set -euo pipefail
args=()
if [ "${FASTVIDEO_SSIM_BOOTSTRAP_MODE:-0}" = 1 ]; then
args+=(--bootstrap-mode)
fi
selected=${FASTVIDEO_SSIM_TEST_FILES-}
if [ -z "$selected" ]; then
if [ "${TEST_SCOPE:-}" = merge ]; then
echo "Missing FASTVIDEO_SSIM_TEST_FILES for merge scope" >&2
exit 2
fi
selected=all
fi
if [ "$selected" != all ]; then
[[ $selected =~ ^test_[a-z0-9_]+\.py(,test_[a-z0-9_]+\.py)*$ ]] || {
echo "Invalid FASTVIDEO_SSIM_TEST_FILES selection" >&2
exit 2
}
IFS=, read -r -a ssim_files <<< "$selected"
for ssim_file in "${ssim_files[@]}"; do
args+=(--test-file "$ssim_file")
done
fi
# MoGe's utils3d dependency builds glcontext from source on ARM64. The current
# runner image predates the baked-in X11 headers below, so keep this guarded
# bootstrap until every deployed image digest contains libx11-dev.
if [ ! -f /usr/include/X11/Xlib.h ]; then
apt-get -o Acquire::Retries=5 update
apt-get -o Acquire::Retries=5 install -y --no-install-recommends libx11-dev
rm -rf /var/lib/apt/lists/*
fi
uv pip install git+https://github.com/microsoft/MoGe.git
uv pip install k_diffusion einops_exts alias_free_torch torchsde
exec python fastvideo/tests/ssim/ci_runner.py "${args[@]}"
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the modular training-framework lane.
set -euo pipefail
exec pytest ./fastvideo/tests/train/models ./fastvideo/tests/train/methods -vs
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the legacy vanilla-training lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/Vanilla -srP
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the legacy LoRA-training lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/lora/test_lora_training.py -srP
+6
View File
@@ -0,0 +1,6 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the legacy VSA-training lane.
set -euo pipefail
export WANDB_MODE=offline
exec pytest ./fastvideo/tests/training/VSA -srP
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the transformer lane.
set -euo pipefail
exec pytest ./fastvideo/tests/transformers -vs
+5
View File
@@ -0,0 +1,5 @@
#!/usr/bin/env bash
# Canonical Slurm CI selection for the VAE lane.
set -euo pipefail
exec pytest ./fastvideo/tests/vaes -vs
+13
View File
@@ -1,6 +1,19 @@
#!/bin/bash
set -uo pipefail
# DORMANT ROLLBACK ONLY. Active CI is Slurm-only and pipeline.yml never calls
# this launcher. Refuse every Buildkite invocation even if a stale step or
# operator typo reaches this file; local rollback experiments require an
# explicit opt-in.
if [ -n "${BUILDKITE:-}" ]; then
echo "Legacy Modal CI is disabled; use the Slinky Slurm runner." >&2
exit 2
fi
if [ "${FASTVIDEO_ENABLE_LEGACY_MODAL_CI:-0}" != 1 ]; then
echo "Legacy Modal CI is dormant. Set FASTVIDEO_ENABLE_LEGACY_MODAL_CI=1 only for a manual rollback test." >&2
exit 2
fi
log() {
echo "[$(date '+%Y-%m-%d %H:%M:%S')] $1"
}
+25
View File
@@ -0,0 +1,25 @@
#!/usr/bin/env bash
set -euo pipefail
exec pytest \
./fastvideo/tests/api/ \
./fastvideo/tests/contract/ \
./fastvideo/tests/dataset/ \
./fastvideo/tests/workflow/ \
./fastvideo/tests/entrypoints/ \
./fastvideo/tests/loader/ \
./fastvideo/tests/pipelines/ \
./fastvideo/tests/platforms/ \
./fastvideo/tests/train/ \
./fastvideo/tests/stages/ \
./fastvideo/tests/ops/ \
./fastvideo/tests/worker/ \
./fastvideo/tests/training/test_trackers.py \
./fastvideo/tests/attention/test_sdpa_metadata_mask_contract.py \
./fastvideo/tests/modal/test_kernel_build_cache.py \
./fastvideo/tests/modal/test_pr_test.py \
./fastvideo/tests/modal/test_ssim_test.py \
--ignore=./fastvideo/tests/entrypoints/test_openai_api_integration.py \
--ignore=./fastvideo/tests/train/models \
--ignore=./fastvideo/tests/train/methods \
-vs
+2 -2
View File
@@ -8,10 +8,10 @@ PR TITLE: Must start with a type tag, e.g.:
MERGE WORKFLOW:
1. Ensure pre-commit passes and you have at least 1 approval
2. Comment /merge (or add the "ready" label) to enter the Merge Queue
3. Full Test Suite runs automatically on a staging branch → auto-merge on success
3. A path-aware merge gate runs only relevant integration tests → auto-merge on success
ON-DEMAND TESTING (write access required):
/test full — Full Test Suite /test ssim — SSIM regression
/test full — Explicit all-lane run /test ssim — Full SSIM regression
/test training — Training pipeline /test encoder — Encoder tests
/test transformer — Transformer tests /test vae — VAE tests
/test kernel — CUDA kernel tests /test unit — Unit tests
+10 -10
View File
@@ -1,14 +1,14 @@
#!/usr/bin/env bash
# Gate the expensive Buildkite full suite on the cheap GitHub checks.
# Gate the path-aware Buildkite merge plan on the cheap GitHub checks.
#
# Polls the workflow runs for the PR head commit and only exits 0 once the
# watched cheap workflows (pre-commit, docs build) have succeeded, so the
# 'ready' label cannot burn ~20 GPU lanes on a head that a cheap check has
# already doomed.
# 'ready' label cannot burn path-selected GPU lanes on a head that a cheap
# check has already doomed.
#
# Semantics:
# - watched run completed with a bad conclusion -> exit 1 (fail CLOSED:
# no full suite; the next push re-arms via the 'synchronize' trigger)
# no merge gate; the next push re-arms via the 'synchronize' trigger)
# - watched run cancelled -> still pending: the docs
# workflow's repo-global 'pages' concurrency group cancels runs superseded
# by unrelated pushes, so 'cancelled' is not a verdict on this PR
@@ -29,7 +29,7 @@ set -euo pipefail
: "${PR_NUMBER:?PR_NUMBER (pull request number) is required}"
: "${GITHUB_REPOSITORY:?GITHUB_REPOSITORY is required}"
# Workflow-level `name:` values that must be green before the full suite
# Workflow-level `name:` values that must be green before the merge gate
# may start. "Deploy Documentation" is path-filtered on PRs, so its run may
# legitimately never exist; pre-commit always runs, so it must appear.
WATCHED_NAMES='["pre-commit", "Deploy Documentation"]'
@@ -56,7 +56,7 @@ recheck_ready_label() {
if pr_json=$(gh_api "repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}" 2>/dev/null); then
if ! jq -e '[.labels[]?.name] | index("ready")' <<<"$pr_json" >/dev/null 2>&1; then
echo "::error::PR #${PR_NUMBER} no longer has the 'ready' label —" \
"NOT triggering the Buildkite full suite. Re-add the label to re-arm."
"NOT triggering the Buildkite merge gate. Re-add the label to re-arm."
exit 1
fi
else
@@ -84,7 +84,7 @@ while true; do
| map(.name) | join(", ")' <<<"$state")
if [ -n "$failed" ]; then
echo "::error::Cheap check(s) failed on ${PR_SHA}: ${failed}." \
"NOT triggering the Buildkite full suite. Push a fix (the 'ready'" \
"NOT triggering the Buildkite merge gate. Push a fix (the 'ready'" \
"label re-arms on every push), or re-run the failed check and then" \
"re-run this workflow."
exit 1
@@ -97,7 +97,7 @@ while true; do
if [ "$pending" -eq 0 ]; then
if [ -z "$missing" ]; then
recheck_ready_label
echo "All watched cheap checks are green — full suite may proceed."
echo "All watched cheap checks are green — merge gate may proceed."
exit 0
fi
case "$missing" in
@@ -119,14 +119,14 @@ while true; do
echo "::warning::GitHub API error querying workflow runs for ${PR_SHA} (attempt ${api_fails}/3)."
if [ "$api_fails" -ge 3 ]; then
recheck_ready_label
echo "::warning::FAILING OPEN: cannot query GitHub check status — triggering the full suite WITHOUT the cheap-check gate."
echo "::warning::FAILING OPEN: cannot query GitHub check status — triggering the merge gate WITHOUT the cheap-check gate."
exit 0
fi
fi
if [ "$elapsed" -ge "$MAX_WAIT_SECS" ]; then
recheck_ready_label
echo "::warning::FAILING OPEN: watched checks still pending after $(( MAX_WAIT_SECS / 60 )) min${missing:+ (never appeared: ${missing})} — triggering the full suite anyway."
echo "::warning::FAILING OPEN: watched checks still pending after $(( MAX_WAIT_SECS / 60 )) min${missing:+ (never appeared: ${missing})} — triggering the merge gate anyway."
exit 0
fi
sleep "$POLL_SECS"
+570
View File
@@ -0,0 +1,570 @@
#!/usr/bin/env python3
"""Select the additive GPU integration lanes needed by a PR diff.
Fastcheck is the universal six-lane baseline and is intentionally not repeated
here. This planner selects only the more expensive merge-gate lanes. Unknown
source/build paths fail closed to the complete integration set, while explicit
documentation and repository-metadata paths require no additional GPU work.
"""
from __future__ import annotations
import argparse
import fnmatch
import re
from dataclasses import dataclass, field
from pathlib import Path
from typing import TextIO
MERGE_LANES = (
"golden-gate",
"ssim",
"lora-inference",
"lora-extraction",
"training",
"distillation",
"self-forcing",
"lora-training",
"training-vsa",
"inference-vmoba",
"performance",
"api-server",
"train-framework",
"eval",
)
LANE_SCRIPT_TO_KEY = {
"api_server.sh": "api-server",
"distillation_dmd.sh": "distillation",
"eval.sh": "eval",
"golden_gate.sh": "golden-gate",
"inference_lora.sh": "lora-inference",
"inference_vmoba.sh": "inference-vmoba",
"lora_extraction.sh": "lora-extraction",
"performance.sh": "performance",
"self_forcing.sh": "self-forcing",
"ssim.sh": "ssim",
"train_framework.sh": "train-framework",
"training.sh": "training",
"training_lora.sh": "lora-training",
"training_vsa.sh": "training-vsa",
}
FASTCHECK_LANE_SCRIPTS = {
"dreamverse.sh",
"encoder.sh",
"kernel_tests.sh",
"transformer.sh",
"vae.sh",
}
LEGACY_TRAINING_LANES = (
"training",
"distillation",
"self-forcing",
"lora-training",
"training-vsa",
)
ALL_TRAINING_LANES = (*LEGACY_TRAINING_LANES, "train-framework")
SSIM_SMOKE_TESTS = (
"test_flux_t2i_similarity.py",
"test_wan_t2v_similarity.py",
)
SAFE_PATTERNS = (
"*.md",
"*.rst",
".agents/**",
".claude/**",
".codex/**",
".github/ISSUE_TEMPLATE/**",
".github/PULL_REQUEST_TEMPLATE.md",
".github/dependabot.yml",
".github/mergify.yml",
".github/scripts/**",
".github/workflows/**",
".buildkite/scripts/pre_commit.sh",
".git-blame-ignore-revs",
".gitattributes",
".gitignore",
".pre-commit-config.yaml",
"AGENTS.md",
"CITATION.cff",
"CODE_OF_CONDUCT.md",
"CONTRIBUTING.md",
"LICENSE",
"NOTICE",
"__init__.py",
"collect_env.py",
"SECURITY.md",
"assets/**",
"comfyui/**",
"docs/**",
"examples/**",
"mkdocs.yml",
"requirements-mkdocs.in",
"requirements-mkdocs.txt",
"scripts/**",
"tests/__init__.py",
"tests/local_tests/**",
)
ALL_IMPACT_PATTERNS = (
".buildkite/pipeline.yml",
"docker/**",
"pyproject.toml",
"requirements*.txt",
"setup.cfg",
"setup.py",
"uv.lock",
)
@dataclass(frozen=True)
class FamilyCoverage:
pattern: re.Pattern[str]
golden_tests: tuple[str, ...]
ssim_tests: tuple[str, ...]
FAMILY_COVERAGE = (
FamilyCoverage(
re.compile(r"(^|[/_.-])dreamx(_world)?([/_.-]|$)"),
("test_dreamx.py", ),
("test_dreamx_world_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])flux[_-]?2([/_.-]|$)"),
("test_flux2_klein.py", ),
("test_flux2_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])flux(?![_-]?2)([/_.-]|$)"),
("test_flux.py", ),
("test_flux_t2i_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])(hunyuan)?gamecraft([/_.-]|$)"),
("test_gamecraft.py", ),
("test_gamecraft_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])gen3c([/_.-]|$)"),
("test_gen3c.py", ),
("test_gen3c_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])glm[_-]?image([/_.-]|$)"),
("test_glm_image.py", ),
("test_glm_image_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])kandinsky[_-]?5([/_.-]|$)"),
("test_kandinsky5.py", ),
("test_kandinsky5_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])lingbot([a-z0-9_-]*)([/_.-]|$)"),
("test_lingbot.py", ),
("test_lingbot_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])longcat([/_.-]|$)"),
("test_longcat.py", ),
("test_longcat_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])ltx[_-]?2([/_.-]|$)"),
("test_ltx2.py", ),
("test_ltx2_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])matrixgame[_-]?2([/_.-]|$)"),
("test_matrixgame.py", ),
("test_matrixgame2_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])matrixgame[_-]?3([/_.-]|$)"),
("test_matrixgame.py", ),
("test_matrixgame3_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])minimax[_-]?h3([/_.-]|$)"),
("test_minimax_h3_t2v.py", ),
("test_minimax_h3_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])sd[_-]?3([._-]?5)?([/_.-]|$)"),
("test_sd35.py", ),
("test_sd35_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])stable[_-]?audio([/_.-]|$)"),
("test_stable_audio.py", ),
("test_stable_audio_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])turbo(diffusion)?([/_.-]|$)"),
(),
("test_turbodiffusion_similarity.py", ),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])wan(video)?([/_.-]|$)"),
("test_wan_t2v.py", ),
(
"test_causal_similarity.py",
"test_wan_i2v_similarity.py",
"test_wan_t2v_similarity.py",
),
),
FamilyCoverage(
re.compile(r"(^|[/_.-])z[_-]?image([/_.-]|$)"),
("test_zimage.py", ),
("test_zimage_similarity.py", ),
),
)
@dataclass
class MergePlan:
lanes: set[str] = field(default_factory=set)
golden_tests: set[str] = field(default_factory=set)
ssim_tests: set[str] = field(default_factory=set)
golden_all: bool = False
ssim_all: bool = False
reasons: list[str] = field(default_factory=list)
def add_lanes(self, *lanes: str, reason: str) -> None:
unknown = set(lanes) - set(MERGE_LANES)
if unknown:
raise ValueError(f"Unknown merge lanes: {sorted(unknown)}")
self.lanes.update(lanes)
self.reasons.append(reason)
def add_golden(self, tests: tuple[str, ...], reason: str) -> None:
self.add_lanes("golden-gate", reason=reason)
self.golden_tests.update(tests)
def add_ssim(self, tests: tuple[str, ...], reason: str) -> None:
self.add_lanes("ssim", reason=reason)
self.ssim_tests.update(tests)
def require_all(self, reason: str) -> None:
self.lanes.update(MERGE_LANES)
self.golden_all = True
self.ssim_all = True
self.reasons.append(reason)
def ordered_lanes(self) -> tuple[str, ...]:
return tuple(lane for lane in MERGE_LANES if lane in self.lanes)
def encoded_lanes(self) -> str:
lanes = self.ordered_lanes()
return "," + ",".join(lanes or ("none", )) + ","
def encoded_golden_tests(self) -> str:
if "golden-gate" not in self.lanes:
return "none"
if self.golden_all or not self.golden_tests:
return "all"
return ",".join(sorted(self.golden_tests))
def encoded_ssim_tests(self) -> str:
if "ssim" not in self.lanes:
return "none"
if self.ssim_all or not self.ssim_tests:
return "all"
return ",".join(sorted(self.ssim_tests))
def _matches_any(path: str, patterns: tuple[str, ...]) -> bool:
return any(fnmatch.fnmatchcase(path, pattern) for pattern in patterns)
def _family_coverage(path: str) -> tuple[set[str], set[str]]:
normalized = path.lower()
golden: set[str] = set()
ssim: set[str] = set()
for family in FAMILY_COVERAGE:
if family.pattern.search(normalized):
golden.update(family.golden_tests)
ssim.update(family.ssim_tests)
return golden, ssim
def _select_output_coverage(plan: MergePlan, path: str) -> None:
golden, ssim = _family_coverage(path)
if golden:
plan.add_golden(tuple(sorted(golden)), reason=f"model-family golden coverage: {path}")
else:
plan.golden_all = True
plan.add_lanes("golden-gate", reason=f"shared output golden coverage: {path}")
if ssim:
plan.add_ssim(tuple(sorted(ssim)), reason=f"model-family SSIM coverage: {path}")
else:
plan.add_ssim(SSIM_SMOKE_TESTS, reason=f"shared output SSIM smoke coverage: {path}")
def classify_paths(paths: list[str]) -> MergePlan:
plan = MergePlan()
normalized_paths: list[str] = []
for raw_path in paths:
path = raw_path.strip()
while path.startswith("./"):
path = path[2:]
if path:
normalized_paths.append(path)
normalized_paths = sorted(set(normalized_paths))
if not normalized_paths:
plan.require_all("changed-file list was empty; failing closed")
return plan
for path in normalized_paths:
if path == "__FASTVIDEO_CI_PLAN_ALL__":
plan.require_all("changed-file API failed; failing closed")
continue
if path in {"requirements-mkdocs.in", "requirements-mkdocs.txt"}:
plan.reasons.append(f"documentation dependencies need no GPU integration: {path}")
continue
if _matches_any(path, ALL_IMPACT_PATTERNS):
plan.require_all(f"cross-cutting build/runtime surface: {path}")
continue
lane_script_prefix = ".buildkite/scripts/lanes/"
if path.startswith(lane_script_prefix):
script_name = Path(path).name
lane = LANE_SCRIPT_TO_KEY.get(script_name)
if lane is None:
if script_name in FASTCHECK_LANE_SCRIPTS:
plan.reasons.append(f"covered by automatic Fastcheck lane: {path}")
else:
plan.require_all(f"unknown lane script: {path}")
elif lane == "golden-gate":
plan.golden_all = True
plan.add_lanes(lane, reason=f"golden lane implementation: {path}")
elif lane == "ssim":
plan.ssim_all = True
plan.add_lanes(lane, reason=f"SSIM lane implementation: {path}")
else:
plan.add_lanes(lane, reason=f"lane implementation: {path}")
continue
if path.startswith("fastvideo/tests/golden_gate/"):
name = Path(path).name
if name.startswith("test_") and name.endswith(".py"):
plan.add_golden((name, ), reason=f"changed golden test: {path}")
elif name in {"AGENTS.md", "README.md"}:
plan.reasons.append(f"golden documentation only: {path}")
else:
plan.golden_all = True
plan.add_lanes("golden-gate", reason=f"shared golden harness/reference: {path}")
continue
if path.startswith("fastvideo/tests/ssim/"):
name = Path(path).name
if name.startswith("test_") and name.endswith(".py"):
plan.add_ssim((name, ), reason=f"changed SSIM test: {path}")
elif path.endswith((".py", ".json", ".pt", ".png", ".mp4")):
plan.ssim_all = True
plan.add_lanes("ssim", reason=f"shared SSIM harness/reference: {path}")
continue
if path.startswith("fastvideo/tests/performance/") or path.startswith(".buildkite/performance-benchmarks/"):
plan.add_lanes("performance", reason=f"performance coverage: {path}")
continue
if path.startswith(("fastvideo/performance/", "fastvideo/performance_dashboard/",
"apps/performance_dashboard/")):
plan.add_lanes("performance", reason=f"performance implementation: {path}")
continue
if path.startswith("fastvideo/benchmarks/"):
if "/mlx_" in path or Path(path).name.startswith("mlx_"):
plan.reasons.append(f"covered by the path-filtered macOS MLX workflow: {path}")
else:
plan.add_lanes("performance", reason=f"benchmark implementation: {path}")
continue
if path.startswith("fastvideo/tests/eval/") or path.startswith("fastvideo/eval/"):
plan.add_lanes("eval", reason=f"evaluation coverage: {path}")
continue
if path.startswith("fastvideo/third_party/eval/"):
plan.add_lanes("eval", reason=f"vendored evaluation implementation: {path}")
continue
if path.startswith("fastvideo/tests/lora_extraction/") or path.startswith("scripts/lora_extraction/"):
plan.add_lanes("lora-extraction", reason=f"LoRA extraction coverage: {path}")
continue
if path.startswith("fastvideo/tests/inference/lora/"):
plan.add_lanes("lora-inference", reason=f"LoRA inference coverage: {path}")
continue
if path.startswith("fastvideo/tests/inference/vmoba/"):
plan.add_lanes("inference-vmoba", reason=f"VMoBA inference coverage: {path}")
continue
if path.startswith(("fastvideo/dataset/", "fastvideo/workflow/", "fastvideo/pipelines/preprocess/",
"fastvideo/pipelines/training/")):
plan.add_lanes(*ALL_TRAINING_LANES, reason=f"shared data/training input surface: {path}")
continue
if path.startswith("fastvideo/tests/train/") or path.startswith("fastvideo/train/"):
plan.add_lanes("train-framework", reason=f"modular training coverage: {path}")
continue
if path.startswith("fastvideo/tests/training/"):
lowered = path.lower()
if "/vanilla/" in lowered:
plan.add_lanes("training", reason=f"vanilla training coverage: {path}")
elif "/distill/" in lowered:
plan.add_lanes("distillation", reason=f"distillation coverage: {path}")
elif "/self-forcing/" in lowered:
plan.add_lanes("self-forcing", reason=f"self-forcing coverage: {path}")
elif "/lora/" in lowered:
plan.add_lanes("lora-training", reason=f"LoRA training coverage: {path}")
elif "/vsa/" in lowered:
plan.add_lanes("training-vsa", reason=f"VSA training coverage: {path}")
else:
plan.add_lanes(*LEGACY_TRAINING_LANES, reason=f"shared legacy training coverage: {path}")
continue
if path.startswith("fastvideo/training/"):
lowered = path.lower()
if "self_forcing" in lowered:
plan.add_lanes("self-forcing", reason=f"self-forcing implementation: {path}")
elif "distill" in lowered:
plan.add_lanes("distillation", reason=f"distillation implementation: {path}")
elif "lora" in lowered:
plan.add_lanes("lora-training", reason=f"LoRA training implementation: {path}")
else:
plan.add_lanes(*LEGACY_TRAINING_LANES, reason=f"shared legacy training implementation: {path}")
continue
lowered = path.lower()
if "vmoba" in lowered and path.startswith(("fastvideo/", ".buildkite/")):
plan.add_lanes("inference-vmoba", reason=f"VMoBA implementation: {path}")
plan.add_golden(("test_wan_t2v.py", ), reason=f"VMoBA end-to-end coverage: {path}")
continue
if "lora" in lowered and path.startswith("fastvideo/"):
plan.add_lanes(
"lora-inference",
"lora-extraction",
"lora-training",
reason=f"shared LoRA implementation: {path}",
)
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/entrypoints/") or path.startswith("fastvideo/api/"):
plan.add_lanes("api-server", reason=f"API/entrypoint integration: {path}")
if "openai" not in lowered and "/cli/" not in lowered:
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/worker/"):
plan.add_lanes("api-server", reason=f"worker/API integration: {path}")
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/distributed/"):
plan.add_lanes(
"training",
"train-framework",
reason=f"distributed runtime integration: {path}",
)
_select_output_coverage(plan, path)
continue
if path.startswith(("fastvideo/hooks/", "fastvideo/platforms/", "fastvideo/third_party/")):
_select_output_coverage(plan, path)
continue
if path.startswith(("fastvideo/models/", "fastvideo/pipelines/", "fastvideo/configs/",
"fastvideo/layers/", "fastvideo/attention/")):
_select_output_coverage(plan, path)
continue
if path in {
"fastvideo/fastvideo_args.py",
"fastvideo/forward_context.py",
"fastvideo/image_processor.py",
"fastvideo/registry.py",
"fastvideo/utils.py",
}:
_select_output_coverage(plan, path)
continue
if path.startswith("fastvideo/mlx_runtime/"):
plan.reasons.append(f"covered by the path-filtered macOS MLX workflow: {path}")
continue
if path.startswith("fastvideo/logging_utils/") or path in {
"fastvideo/__init__.py",
"fastvideo/envs.py",
"fastvideo/logger.py",
"fastvideo/profiler.py",
"fastvideo/version.py",
}:
plan.reasons.append(f"covered by automatic Fastcheck: {path}")
continue
if path.startswith(("fastvideo-kernel/", "csrc/")):
plan.add_golden(("test_wan_t2v.py", ), reason=f"kernel integration smoke: {path}")
plan.add_ssim(("test_wan_t2v_similarity.py", ), reason=f"kernel numerical smoke: {path}")
continue
if path.startswith("apps/dreamverse/"):
# DreamVerse is already one of the six automatic Fastcheck lanes.
plan.reasons.append(f"covered by automatic DreamVerse Fastcheck: {path}")
continue
if path.startswith("fastvideo/tests/"):
# The automatic unit/component Fastcheck lanes own the remaining
# package tests. Domain-specific expensive test roots were handled
# above.
plan.reasons.append(f"covered by automatic Fastcheck: {path}")
continue
if path in {".buildkite/scripts/unit_test.sh", ".buildkite/scripts/pr_test.sh"}:
plan.reasons.append(f"covered by automatic unit Fastcheck: {path}")
continue
if _matches_any(path, SAFE_PATTERNS):
plan.reasons.append(f"no additional GPU integration needed: {path}")
continue
plan.require_all(f"unclassified path; failing closed: {path}")
return plan
def _write_github_output(output: TextIO, plan: MergePlan) -> None:
output.write(f"merge_test_plan={plan.encoded_lanes()}\n")
output.write(f"merge_golden_tests={plan.encoded_golden_tests()}\n")
output.write(f"merge_ssim_tests={plan.encoded_ssim_tests()}\n")
output.write(f"merge_plan_label={','.join(plan.ordered_lanes()) or 'none'}\n")
def _write_summary(output: TextIO, plan: MergePlan) -> None:
output.write("## Change-aware merge test plan\n\n")
output.write("| Selection | Value |\n|---|---|\n")
output.write(f"| Additional Slurm lanes | `{','.join(plan.ordered_lanes()) or 'none'}` |\n")
output.write(f"| Golden tests | `{plan.encoded_golden_tests()}` |\n")
output.write(f"| SSIM tests | `{plan.encoded_ssim_tests()}` |\n\n")
output.write("Fastcheck remains the universal six-lane baseline.\n")
def parse_args() -> argparse.Namespace:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--paths-file", type=Path, required=True)
parser.add_argument("--github-output", type=Path)
parser.add_argument("--summary-file", type=Path)
return parser.parse_args()
def main() -> int:
args = parse_args()
paths = args.paths_file.read_text(encoding="utf-8").splitlines()
plan = classify_paths(paths)
print(f"MERGE_TEST_PLAN={plan.encoded_lanes()}")
print(f"MERGE_GOLDEN_TESTS={plan.encoded_golden_tests()}")
print(f"MERGE_SSIM_TESTS={plan.encoded_ssim_tests()}")
for reason in plan.reasons:
print(f"- {reason}")
if args.github_output:
with args.github_output.open("a", encoding="utf-8") as output:
_write_github_output(output, plan)
if args.summary_file:
with args.summary_file.open("a", encoding="utf-8") as output:
_write_summary(output, plan)
return 0
if __name__ == "__main__":
raise SystemExit(main())
+1 -1
View File
@@ -53,7 +53,7 @@ PC_PENDING='{"name": "pre-commit", "id": 1, "status": "in_progress", "conclusion
DOCS_OK='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "success"}'
DOCS_BAD='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "failure"}'
DOCS_CANCELLED='{"name": "Deploy Documentation", "id": 2, "status": "completed", "conclusion": "cancelled"}'
OTHER='{"name": "Trigger Full Suite", "id": 3, "status": "in_progress", "conclusion": null}'
OTHER='{"name": "Trigger Merge Gate", "id": 3, "status": "in_progress", "conclusion": null}'
NULL_NAME='{"name": null, "id": 4, "status": "completed", "conclusion": "failure"}'
PC_OK_RERUN='{"name": "pre-commit", "id": 5, "status": "completed", "conclusion": "success"}'
@@ -190,6 +190,7 @@ jobs:
if: ${{ !inputs.push_by_digest }}
run: |
echo "✅ Python ${{ inputs.python_version }} image successfully built and pushed to ${{ steps.image.outputs.name }}:${{ inputs.tag_suffix }}-sha-${GITHUB_SHA::7}"
echo "Digest: ${{ steps.build-push.outputs.digest }}"
echo "To run tests with this image, manually trigger the 'Run Tests' workflow."
- name: Digest success message
+39 -25
View File
@@ -26,29 +26,48 @@ jobs:
per_page: 100,
});
const bkStatuses = data.statuses.filter(
s => s.context.startsWith('buildkite/ci/')
);
const FASTCHECK_PREFIX = 'buildkite/ci/microscope-';
// Buildkite derives the GitHub context prefix from the label emoji.
// Keep hard Full Suite lanes in test-tube/bar-chart namespaces and
// Fastcheck lanes in microscope so targeted reruns cannot clear the
// wrong aggregate status. Automatic PR jobs use pr-fastcheck while
// slash-command and Full Suite jobs use ci; normalize the suffix
// and keep the newest status for each logical lane.
const FASTCHECK_PREFIXES = [
'buildkite/pr-fastcheck/microscope-',
'buildkite/ci/microscope-',
];
const FULL_SUITE_PREFIXES = [
'buildkite/ci/test-tube-',
'buildkite/ci/bar-chart-',
];
const fastcheck = bkStatuses.filter(
s => s.context.startsWith(FASTCHECK_PREFIX)
);
const fullSuite = bkStatuses.filter(
s => FULL_SUITE_PREFIXES.some(p => s.context.startsWith(p))
);
function newestByLane(prefixes) {
const statuses = new Map();
for (const status of data.statuses) {
const prefix = prefixes.find(p => status.context.startsWith(p));
if (!prefix) continue;
const lane = status.context.slice(prefix.length);
const previous = statuses.get(lane);
if (!previous || Date.parse(status.updated_at) > Date.parse(previous.updated_at)) {
statuses.set(lane, status);
}
}
return statuses;
}
if (
fastcheck.length > 0
&& fastcheck.every(s => s.state === 'success')
) {
const fastcheck = newestByLane(FASTCHECK_PREFIXES);
const fullSuiteOnly = newestByLane(FULL_SUITE_PREFIXES);
const fastcheckPassed =
fastcheck.size === 6
&& [...fastcheck.values()].every(s => s.state === 'success');
const fullSuitePassed =
fastcheckPassed
&& fullSuiteOnly.size === 14
&& [...fullSuiteOnly.values()].every(s => s.state === 'success');
if (fastcheckPassed) {
core.info(
`All ${fastcheck.length} fastcheck tests passed — updating fastcheck-passed`
`All ${fastcheck.size} fastcheck tests passed — updating fastcheck-passed`
);
await github.rest.repos.createCommitStatus({
owner: context.repo.owner,
@@ -56,17 +75,13 @@ jobs:
sha,
state: 'success',
context: 'fastcheck-passed',
description:
`All ${fastcheck.length} fastcheck tests passed`,
description: `All ${fastcheck.size} fastcheck tests passed`,
});
}
if (
fullSuite.length > 0
&& fullSuite.every(s => s.state === 'success')
) {
if (fullSuitePassed) {
core.info(
`All ${fullSuite.length} full suite tests passed — updating full-suite-passed`
'All 20 full suite tests passed — updating full-suite-passed'
);
await github.rest.repos.createCommitStatus({
owner: context.repo.owner,
@@ -74,7 +89,6 @@ jobs:
sha,
state: 'success',
context: 'full-suite-passed',
description:
`All ${fullSuite.length} full suite tests passed`,
description: 'All 20 full suite tests passed',
});
}
+20 -2
View File
@@ -8,6 +8,8 @@ on:
- "fastvideo/mlx_runtime/**"
- "fastvideo/tests/mlx/**"
- "fastvideo/tests/platforms/test_mps_vsa_error.py"
- "fastvideo/tests/platforms/test_cpu_sdpa.py"
- "fastvideo/platforms/cpu.py"
- "fastvideo/platforms/mps.py"
- "fastvideo/platforms/__init__.py"
- "fastvideo/__init__.py"
@@ -78,6 +80,13 @@ jobs:
fastvideo/tests/mlx/test_mlx_dit_parity.py \
fastvideo/tests/mlx/test_mlx_compile_parity.py \
fastvideo/tests/mlx/test_mlx_checkpoint.py \
fastvideo/tests/mlx/test_mlx_checkpoint_compat.py \
fastvideo/tests/mlx/test_mlx_affine_dq_gemm.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_parity.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa_regressions.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_mode.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_fastwan_benchmark.py \
fastvideo/tests/mlx/test_taehv_decode.py \
fastvideo/tests/mlx/test_frame_upsample.py \
@@ -90,7 +99,8 @@ jobs:
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_download_unavailable_has_specific_error \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
fastvideo/tests/platforms/test_mps_vsa_error.py \
-q
fastvideo/tests/platforms/test_cpu_sdpa.py \
-v -s -o faulthandler_timeout=120
# Same tests on MLX's CPU backend. Hosted macOS runners are scarce and
# slower to schedule; this Linux job gives fast PR signal on the identical
@@ -134,6 +144,13 @@ jobs:
fastvideo/tests/mlx/test_mlx_dit_parity.py \
fastvideo/tests/mlx/test_mlx_compile_parity.py \
fastvideo/tests/mlx/test_mlx_checkpoint.py \
fastvideo/tests/mlx/test_mlx_checkpoint_compat.py \
fastvideo/tests/mlx/test_mlx_affine_dq_gemm.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_parity.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_vsa_regressions.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_mode.py \
fastvideo/tests/mlx/test_mlx_minimax_h3_fast_spatial.py \
fastvideo/tests/mlx/test_mlx_fastwan_benchmark.py \
fastvideo/tests/mlx/test_taehv_decode.py \
fastvideo/tests/mlx/test_frame_upsample.py \
@@ -146,4 +163,5 @@ jobs:
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_download_unavailable_has_specific_error \
fastvideo/tests/mlx/test_mlx_rife_interpolation.py::test_rife_backend_regression_is_not_skip_eligible \
fastvideo/tests/platforms/test_mps_vsa_error.py \
-q
fastvideo/tests/platforms/test_cpu_sdpa.py \
-v -s -o faulthandler_timeout=120
+44
View File
@@ -0,0 +1,44 @@
name: Scheduled Full SSIM
on:
schedule:
- cron: "0 5 * * 0"
workflow_dispatch:
permissions:
contents: read
jobs:
trigger:
if: github.repository == 'hao-ai-lab/FastVideo'
runs-on: ubuntu-latest
steps:
- name: Trigger weekly full SSIM on Slinky Slurm
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
SOURCE_SHA: ${{ github.sha }}
SOURCE_BRANCH: ${{ github.event.repository.default_branch }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
set -euo pipefail
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
-H "Content-Type: application/json" \
--data-raw "$(jq -n \
--arg commit "$SOURCE_SHA" \
--arg branch "$SOURCE_BRANCH" \
'{
commit: $commit,
branch: $branch,
message: "Weekly full SSIM on Slinky Slurm",
ignore_pipeline_branch_filters: true,
env: {
TEST_SCOPE: "scheduled",
FULL_SUITE: "false",
TEST_TYPE: "ssim",
PR_NUMBER: "false",
PR_TITLE: "Scheduled full SSIM"
}
}')"
+12 -44
View File
@@ -33,7 +33,6 @@ jobs:
core.setOutput('has_write', String(hasWrite));
- name: Add ready label and react
id: label
if: steps.perm.outputs.has_write == 'true'
uses: actions/github-script@60a0d83039c74a4aee543508d2ffcb1c3799cdea # v7.0.1
with:
@@ -48,47 +47,6 @@ jobs:
comment_id: context.payload.comment.id,
content: 'rocket',
});
const { data: pr } = await github.rest.pulls.get({ owner, repo, pull_number: prNumber });
core.setOutput('pr_sha', pr.head.sha);
core.setOutput('pr_branch', pr.head.ref);
core.setOutput('pr_number', String(prNumber));
core.setOutput('pr_title', pr.title);
- name: Trigger Full Suite
if: steps.perm.outputs.has_write == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_SHA: ${{ steps.label.outputs.pr_sha }}
PR_BRANCH: ${{ steps.label.outputs.pr_branch }}
PR_NUMBER: ${{ steps.label.outputs.pr_number }}
PR_TITLE: ${{ steps.label.outputs.pr_title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
run: |
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
-H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
-H "Content-Type: application/json" \
--data-raw "$(jq -n \
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER} (via /merge)" \
--arg pr_title "$PR_TITLE" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
branch: $branch,
message: $message,
ignore_pipeline_branch_filters: true,
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: "full",
FULL_SUITE: "true",
PR_NUMBER: ($pr_id | tostring),
PR_TITLE: $pr_title
}
}')"
parse-command:
if: >-
@@ -129,7 +87,7 @@ jobs:
set -euo pipefail
TEST_NAME=$(echo "$COMMENT" | grep -oP '(?<=/test\s)\S+' | head -1 || true)
VALID="encoder vae transformer kernel unit dreamverse ssim golden-gate training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval full fastcheck pre-commit"
VALID="encoder vae transformer kernel unit dreamverse ssim golden-gate training lora-inference lora-training lora-extraction distillation self-forcing vsa vmoba performance api train-framework eval unit-ci kernel-ci dreamverse-ci ssim-ci golden-gate-ci encoder-ci vae-ci transformer-ci lora-inference-ci lora-training-ci lora-extraction-ci training-ci distillation-ci self-forcing-ci vsa-ci vmoba-ci performance-ci api-ci train-framework-ci eval-ci full fastcheck pre-commit"
if [ -z "$TEST_NAME" ] || ! echo "$VALID" | grep -qw "$TEST_NAME"; then
echo "Unknown test: '$TEST_NAME'. Valid: $VALID"
exit 1
@@ -137,7 +95,17 @@ jobs:
declare -A MAP=(
[encoder]=encoder [vae]=vae [transformer]=transformer
[kernel]=kernel_tests [unit]=unit_test [dreamverse]=dreamverse_app
[kernel]=kernel_tests [unit]=unit_test [unit-ci]=unit_test_ci
[kernel-ci]=kernel_tests_ci [dreamverse-ci]=dreamverse_app_ci
[ssim-ci]=ssim_ci [vmoba-ci]=inference_vmoba_ci
[golden-gate-ci]=golden_gate_ci [training-ci]=training_ci
[encoder-ci]=encoder_ci [vae-ci]=vae_ci [transformer-ci]=transformer_ci
[lora-inference-ci]=inference_lora_ci [lora-training-ci]=training_lora_ci
[lora-extraction-ci]=lora_extraction_ci [distillation-ci]=distillation_dmd_ci
[self-forcing-ci]=self_forcing_ci [vsa-ci]=training_vsa_ci
[performance-ci]=performance_ci [api-ci]=api_server_ci
[train-framework-ci]=train_framework_ci [eval-ci]=eval_ci
[dreamverse]=dreamverse_app
[ssim]=ssim [golden-gate]=golden_gate [training]=training
[lora-inference]=inference_lora [lora-training]=training_lora
[lora-extraction]=lora_extraction
+66 -13
View File
@@ -1,4 +1,4 @@
name: Trigger Full Suite
name: Trigger Merge Gate
on:
pull_request_target:
@@ -10,7 +10,7 @@ permissions:
actions: read
concurrency:
group: full-suite-${{ github.event.pull_request.number }}
group: merge-gate-${{ github.event.pull_request.number }}
cancel-in-progress: false
jobs:
@@ -34,29 +34,72 @@ jobs:
});
const hasReady = pr.labels.some(l => l.name === 'ready');
core.setOutput('has_ready', String(hasReady));
if (!hasReady) core.info('No ready label — skipping Full Suite trigger.');
core.setOutput('changed_files', String(pr.changed_files));
if (!hasReady) core.info('No ready label — skipping merge-gate trigger.');
- name: Cancel previous Buildkite builds
if: steps.check.outputs.has_ready == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
PR_BRANCH: ${{ github.event.pull_request.head.ref }}
PR_NUMBER: ${{ github.event.pull_request.number }}
run: |
# Find running builds for this branch with TEST_SCOPE=full and cancel them
builds=$(curl -sS -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds?branch=${PR_BRANCH}&state=running,scheduled" \
| jq -r '.[] | select(try (.env.TEST_SCOPE == "full") catch false) | .number')
# Match both branch and PR number: forks can reuse the same branch name.
builds=$(curl -sS --get -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
--data-urlencode "branch=$PR_BRANCH" \
--data-urlencode "state=running,scheduled" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds" \
| jq -r --arg pr_number "$PR_NUMBER" \
'.[] | select((.env.TEST_SCOPE? == "merge") and (.env.PR_NUMBER? == $pr_number)) | .number')
for build_num in $builds; do
echo "Cancelling Buildkite build #$build_num"
curl -sS -X PUT -H "Authorization: Bearer $BUILDKITE_API_TOKEN" \
"https://api.buildkite.com/v2/organizations/${{ vars.BUILDKITE_ORG_SLUG }}/pipelines/${{ vars.BUILDKITE_PIPELINE_SLUG }}/builds/${build_num}/cancel"
done
# Checks out the BASE branch (default for pull_request_target), so PR
# authors cannot tamper with the gate script.
- name: Checkout gate script
# Check out the immutable BASE SHA: pull_request_target must never run a
# planner or gate script from the untrusted PR head.
- name: Checkout trusted merge planner
if: steps.check.outputs.has_ready == 'true'
uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 # v4.2.2
with:
ref: ${{ github.event.pull_request.base.sha }}
persist-credentials: false
- name: Collect changed paths
if: steps.check.outputs.has_ready == 'true'
env:
GH_TOKEN: ${{ github.token }}
PR_NUMBER: ${{ github.event.pull_request.number }}
EXPECTED_CHANGED_FILES: ${{ steps.check.outputs.changed_files }}
run: |
set -euo pipefail
changed_json="$RUNNER_TEMP/merge-changed-files.json"
changed_paths="$RUNNER_TEMP/merge-changed-paths.txt"
if gh api --paginate --slurp \
"repos/${GITHUB_REPOSITORY}/pulls/${PR_NUMBER}/files?per_page=100" \
> "$changed_json"; then
observed=$(jq '[.[][] | .filename] | unique | length' "$changed_json")
if [ "$observed" = "$EXPECTED_CHANGED_FILES" ]; then
jq -r '.[][] | .filename, (.previous_filename // empty)' "$changed_json" \
| sort -u > "$changed_paths"
else
echo "::warning::Changed-file API returned $observed of $EXPECTED_CHANGED_FILES paths; selecting all merge lanes."
echo '__FASTVIDEO_CI_PLAN_ALL__' > "$changed_paths"
fi
else
echo "::warning::Changed-file API failed; selecting all merge lanes."
echo '__FASTVIDEO_CI_PLAN_ALL__' > "$changed_paths"
fi
- name: Select minimal merge tests
id: plan
if: steps.check.outputs.has_ready == 'true'
run: |
python3 .github/scripts/plan_merge_ci.py \
--paths-file "$RUNNER_TEMP/merge-changed-paths.txt" \
--github-output "$GITHUB_OUTPUT" \
--summary-file "$GITHUB_STEP_SUMMARY"
- name: Wait for pre-commit and docs build
if: steps.check.outputs.has_ready == 'true'
@@ -66,7 +109,7 @@ jobs:
PR_NUMBER: ${{ github.event.pull_request.number }}
run: bash .github/scripts/gate_full_suite.sh
- name: Trigger Buildkite Full Suite
- name: Trigger Buildkite merge gate
if: steps.check.outputs.has_ready == 'true'
env:
BUILDKITE_API_TOKEN: ${{ secrets.BUILDKITE_API_TOKEN }}
@@ -76,6 +119,10 @@ jobs:
PR_TITLE: ${{ github.event.pull_request.title }}
BK_ORG: ${{ vars.BUILDKITE_ORG_SLUG }}
BK_PIPELINE: ${{ vars.BUILDKITE_PIPELINE_SLUG }}
MERGE_TEST_PLAN: ${{ steps.plan.outputs.merge_test_plan }}
MERGE_GOLDEN_TESTS: ${{ steps.plan.outputs.merge_golden_tests }}
MERGE_SSIM_TESTS: ${{ steps.plan.outputs.merge_ssim_tests }}
MERGE_PLAN_LABEL: ${{ steps.plan.outputs.merge_plan_label }}
run: |
curl -sS --fail-with-body -X POST \
"https://api.buildkite.com/v2/organizations/${BK_ORG}/pipelines/${BK_PIPELINE}/builds" \
@@ -84,8 +131,11 @@ jobs:
--data-raw "$(jq -n \
--arg commit "$PR_SHA" \
--arg branch "$PR_BRANCH" \
--arg message "Full Suite for PR #${PR_NUMBER}" \
--arg message "Merge gate [${MERGE_PLAN_LABEL}] for PR #${PR_NUMBER}" \
--arg pr_title "$PR_TITLE" \
--arg merge_test_plan "$MERGE_TEST_PLAN" \
--arg merge_golden_tests "$MERGE_GOLDEN_TESTS" \
--arg merge_ssim_tests "$MERGE_SSIM_TESTS" \
--argjson pr_id "$PR_NUMBER" \
'{
commit: $commit,
@@ -95,8 +145,11 @@ jobs:
pull_request_id: $pr_id,
pull_request_base_branch: "main",
env: {
TEST_SCOPE: "full",
TEST_SCOPE: "merge",
FULL_SUITE: "true",
MERGE_TEST_PLAN: $merge_test_plan,
MERGE_GOLDEN_TESTS: $merge_golden_tests,
MERGE_SSIM_TESTS: $merge_ssim_tests,
PR_NUMBER: ($pr_id | tostring),
PR_TITLE: $pr_title
}
+4 -4
View File
@@ -38,17 +38,17 @@ jobs:
**How our CI works:**
PRs run a two-tier CI system:
PRs run a three-tier CI system:
1. **Pre-commit** — formatting (yapf), linting (ruff), type checking (mypy). Runs immediately on every PR.
2. **Fastcheck** — core GPU tests (encoders, VAEs, transformers, kernels, unit tests). Runs automatically via Buildkite on relevant file changes (~10-15 min).
3. **Full Suite** — integration tests, training pipelines, SSIM regression. Runs only when a reviewer adds the `ready` label.
2. **Fastcheck** — six core GPU lanes run automatically via Buildkite (~10-15 min).
3. **Merge gate** — a reviewer adds `ready`; changed paths select only the relevant integration, training, golden, or SSIM coverage.
**Before your PR is reviewed:**
- [ ] `pre-commit run --all-files` passes locally
- [ ] You've added or updated tests for your changes
- [ ] The PR description explains what and why
If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and Full Suite results appear in the Checks section below.
If pre-commit fails, a bot comment will explain how to fix it. Fastcheck and merge-gate results appear in the Checks section below.
**Useful links:**
- [Contributing Guide](https://hao-ai-lab.github.io/FastVideo/contributing/overview/)
+27
View File
@@ -13,6 +13,11 @@ on:
required: false
default: false
type: boolean
build_ci_runner_image:
description: 'Build the ARM64 CUDA 13 CI runner image (sm_100)'
required: false
default: false
type: boolean
# Auto-rebuild the CUDA images when a repository-controlled image input
# changes on main. This includes the trusted SM89 kernel artifact's source,
# metadata/key helper, ABI dependency metadata, and build orchestration.
@@ -198,6 +203,28 @@ jobs:
docker buildx imagetools create "${TAG_ARGS[@]}" "${IMAGE_REFS[@]}"
docker buildx imagetools inspect "${TAGS[0]}"
# The CI runner is ARM64 like DGX Spark, but targets sm_100 rather than sm_121.
# Publish a single-architecture variant so the self-hosted CI runner can reuse
# the exact prebuilt kernel instead of compiling it in every job.
build-ci-runner-image:
if: ${{ (github.event_name == 'push' && github.repository == 'hao-ai-lab/FastVideo') || github.event.inputs.build_ci_runner_image == 'true' }}
uses: ./.github/workflows/_template-build-image.yml
with:
python_version: '3.12'
dockerfile_path: docker/Dockerfile
tag_suffix: py3.12-cuda13.0.0-sm100
runner: ubuntu-24.04-arm
architecture: arm64
build_args: |
PYTHON_VERSION=3.12
CUDA_VERSION=13.0.0
UV_TORCH_BACKEND=cu130
TORCH_CUDA_ARCH_LIST=10.0
CMAKE_BUILD_PARALLEL_LEVEL=1
FLASH_ATTN_WHEEL_TAG=cu130torch2.12
FLASH_ATTN_WHEEL_RELEASE_ARM64=https://github.com/mjun0812/flash-attention-prebuild-wheels/releases/download/v0.9.22
secrets: inherit
# Dreamverse matrix: {backend, UI} x {12.6.3, 13.0.0}, Python 3.12. Torch backend
# matches the base CUDA (cu126 / cu130). Keep these images amd64-only until the
# required FA4 dependency stack is available and validated on arm64.
+14 -7
View File
@@ -62,8 +62,9 @@ jobs:
cuda-version: '13.0.0'
torch-cuda-short: 'cu130'
platform:
# x86_64 builds the full cu126 + cu130 set (cu130 ships the consumer
# Blackwell sm_120a FP4 kernels).
# x86_64 builds the full cu126 + cu130 set. cu130 ships the
# data-center Blackwell sm_100a/sm_103a VSA and consumer sm_120a FP4
# kernels.
- os: ubuntu-22.04
arch: x86_64
wheel-plat: manylinux_2_35_x86_64
@@ -124,7 +125,7 @@ jobs:
- name: Install dependencies (GCC, Clang, CUDA Paths, Git)
run: |
sudo apt update
sudo apt install -y git patchelf gcc-11 g++-11 clang-11
sudo apt install -y git gcc-11 g++-11 clang-11
sudo update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
# Allow Git to Access Safe Directory
@@ -168,7 +169,8 @@ jobs:
# covers sm_120a; turbodiffusion covers sm_100a+sm_120a. The sm_100 FP4
# forward is the FA4 CuTe DSL path in the fastvideo package (PR #1221),
# JIT-compiled at runtime — not built into this wheel.
# * x86_64 cu130 = Hopper TK + consumer Blackwell sm_120a FP4.
# * x86_64 cu130 = Hopper TK + data-center Blackwell sm_100a/sm_103a VSA
# + consumer Blackwell sm_120a FP4.
# * x86_64 cu126 = Hopper TK only (older drivers; CUDA < 12.8 has no FP4).
# The per-arch split in CMakeLists pins the FP4 targets to sm_120a and builds
# the main extension for the full arch list. CMAKE_BUILD_PARALLEL_LEVEL caps
@@ -178,7 +180,7 @@ jobs:
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=OFF -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON"
export CMAKE_BUILD_PARALLEL_LEVEL=1
elif [ "${{ matrix.torch-cuda.torch-cuda-short }}" = "cu130" ]; then
export TORCH_CUDA_ARCH_LIST="9.0a;12.0a"
export TORCH_CUDA_ARCH_LIST="9.0a;10.0a;12.0a"
export CMAKE_ARGS="${CMAKE_ARGS:-} -DFASTVIDEO_KERNEL_BUILD_TK=ON -DFASTVIDEO_KERNEL_BUILD_ATTN_QAT_INFER=ON -DCMAKE_CUDA_ARCHITECTURES=90a"
# A single FP4 TU (attn_qat_infer) can use ~8-12 GB on its own, so serialize.
export CMAKE_BUILD_PARALLEL_LEVEL=1
@@ -194,7 +196,11 @@ jobs:
python -m build --wheel --outdir dist
# Fix the wheel to be manylinux compliant
uv pip install --system auditwheel
# Ubuntu 22.04 ships patchelf 0.14.3, while current auditwheel
# requires at least 0.14.5. Use the stable PyPI binary on both
# x86_64 and aarch64 release runners.
uv pip install --system auditwheel patchelf==0.17.2.4
patchelf --version
# Point auditwheel at torch libs, but do not vendor them into the wheel.
TORCH_LIB_DIR=$(python - <<'PY'
import os
@@ -211,7 +217,8 @@ jobs:
--exclude libtorch.so \
--exclude libc10.so \
--exclude libc10_cuda.so \
--exclude libtorch_python.so
--exclude libtorch_python.so \
--exclude libnccl.so.2
# Move fixed wheels back to dist for upload consistency
rm dist/*.whl
mv fixed_dist/*.whl dist/
+5
View File
@@ -55,6 +55,8 @@ eggs/
# MkDocs documentation
site/
docs/assets/cookbook-serving.json
examples/serving/clients/node_modules/
docs/getting_started/examples/
docs/examples/
docs/inference/examples/
@@ -133,6 +135,9 @@ fastvideo/tests/ssim/reference_videos/**
!fastvideo/tests/ssim/reference_videos/**/*.mp4
!fastvideo/tests/ssim/reference_videos/**/*.png
# Local H3 MLX kernel / exactness benches (JSON, logs, frames, videos)
.kernel_bench/
# Editor logs and local Python version pins (accidentally committed)
*.nvimlog
.nvimlog
+7 -4
View File
@@ -3,12 +3,14 @@
</div>
<p align="center">
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://github.com/hao-ai-lab/FastVideo/discussions/1097" target="_blank"> <b> WeChat </b> </a> |
| <a href="https://hao-ai-lab.github.io/FastVideo"><b>Documentation</b></a> | <a href="https://haoailab.com/FastVideo/cookbook/"><b>Cookbook</b></a> | <a href="https://hao-ai-lab.github.io/FastVideo/inference/inference_quick_start/"><b> Quick Start</b></a> | <a href="https://github.com/hao-ai-lab/FastVideo/discussions/982" target="_blank"><b>Weekly Dev Meeting</b></a> | 🟣💬 <a href="https://join.slack.com/t/fastvideo/shared_invite/zt-3f4lao1uq-u~Ipx6Lt4J27AlD2y~IdLQ" target="_blank"> <b>Slack</b> </a> | 🟣💬 <a href="https://github.com/hao-ai-lab/FastVideo/discussions/1097" target="_blank"> <b> WeChat </b> </a> |
</p>
**FastVideo is a unified post-training and real-time inference framework for accelerated video generation.**
## NEWS
- `2026/09/01`: FastH3 now runs locally on Apple Silicon through MLX and on NVIDIA DGX Spark through CUDA 13, including two-Spark inference. Follow the [FastH3 recipes](https://haoailab.com/FastVideo/cookbook/minimax-h3/) and read the [Blog](https://haoailab.com/blogs/fasth3-local/).
- `2026/08/27`: [FastH3 Preview v1](https://haoailab.com/blogs/fasth3-preview/) is an open-weight 4-step sparse-distilled MiniMax-H3 model for synchronized video-and-audio generation, developed in collaboration with [Nuva Lab](https://nuvalab.ai/) and the [NVIDIA FastGen team](https://github.com/NVlabs/FastGen). Download the recommended [VSA / Data-Free weights](https://huggingface.co/FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree), or see the [full FastH3 collection](https://huggingface.co/collections/FastVideo/fastvideo-fasth3).
- `2026/08/19`: FastVideo now supports MLX on Apple Silicon with [FastMetal-QAD](https://huggingface.co/collections/FastVideo/fastmetal), a family of 1.3B, 5B, and 14B models optimized for Mac—follow the [Apple Silicon guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/) and read the [Blog](https://haoailab.com/blogs/fastmetal/).
- `2026/06/23`: Release FastWan-QAD: 5s of Video generated in 1.8s E2E. See the [FastWan-QAD models](https://huggingface.co/FastVideo/FastWan-QAD-FP8-1.3B), [Attn-QAT training guide](https://haoailab.com/FastVideo/training/attn_qat/), and [blog](https://haoailab.com/blogs/fastwan-qad/).
- `2026/03/17`: Release demo: Into the Dreamverse: Vibe Directing in FastVideo, check out the [Blog](https://haoailab.com/blogs/dreamverse/).
@@ -63,9 +65,10 @@ UV_TORCH_BACKEND=cu126 uv pip install fastvideo
Use `UV_TORCH_BACKEND=cu130` on CUDA 13. Apple silicon users should follow the
[MPS installation guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
> **On an Apple Silicon Mac?** FastVideo runs FastWan text-to-video natively
> through an MLX runtime — a 5-second 480p clip generated locally, no cloud,
> no discrete GPU. Install with `uv pip install -e '.[mlx]'` and follow the
> **On an Apple Silicon Mac?** FastVideo runs FastMetal-QAD through an MLX
> runtime. Install with `uv pip install -e '.[mlx]'`, download
> [`FastVideo/FastMetal-1.3B-QAD`](https://huggingface.co/FastVideo/FastMetal-1.3B-QAD),
> and follow the
> [Apple Silicon guide](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/mps/).
Please see our [docs](https://hao-ai-lab.github.io/FastVideo/getting_started/installation/) for more detailed installation instructions.
+20 -4
View File
@@ -9,6 +9,7 @@ from __future__ import annotations
import contextlib
import logging
import json
import sqlite3
import threading
from pathlib import Path
@@ -46,8 +47,13 @@ DEFAULT_SETTINGS: dict[str, Any] = {
def _sqlite_row_get(row: sqlite3.Row, key: str, default: Any) -> Any:
"""Like dict.get for sqlite3.Row (Row has no .get on Python 3.10)."""
return row[key] if key in row else default # noqa: SIM401
"""Like dict.get for sqlite3.Row (Row has no .get on Python 3.10).
NOTE: `key in row` tests Row *values*, not column names, so the membership
check has to go through .keys() -- otherwise every lookup falls back to the
default and jobs restored from the database lose their stored fields.
"""
return row[key] if key in row.keys() else default # noqa: SIM401, SIM118
def _get_db_path(data_dir: Path) -> Path:
@@ -83,6 +89,9 @@ def _migrate_db(conn: sqlite3.Connection) -> None:
_add_column_if_missing(conn, "jobs", "fps", "INTEGER", "24")
_add_column_if_missing(conn, "jobs", "workload_type", "TEXT", "'t2v'")
_add_column_if_missing(conn, "jobs", "image_path", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "name", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "last_image_path", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "references_json", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "job_type", "TEXT", "'inference'")
_add_column_if_missing(conn, "jobs", "data_path", "TEXT", "''")
_add_column_if_missing(conn, "jobs", "max_train_steps", "INTEGER", "1000")
@@ -242,7 +251,8 @@ class Database:
self._execute(
"""
INSERT INTO jobs (
id, model_id, prompt, workload_type, image_path, job_type, status,
id, model_id, name, prompt, workload_type, image_path,
last_image_path, references_json, job_type, status,
created_at, started_at, finished_at, error, output_path, log_file_path,
num_inference_steps, num_frames, height, width, guidance_scale,
guidance_rescale, fps, seed, num_gpus, dit_cpu_offload,
@@ -254,14 +264,17 @@ class Database:
dmd_use_vsa, dmd_vsa_sparsity, dmd_denoising_steps,
real_score_guidance_scale,
generator_update_interval, real_score_model_path, fake_score_model_path
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
) VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
""",
(
job["id"],
job["model_id"],
job.get("name", ""),
job["prompt"],
job.get("workload_type", "t2v"),
job.get("image_path", ""),
job.get("last_image_path", ""),
json.dumps(job.get("references") or []),
job.get("job_type", "inference"),
job["status"],
job["created_at"],
@@ -540,9 +553,12 @@ def _row_to_job(row: sqlite3.Row) -> dict[str, Any]:
result = {
"id": row["id"],
"model_id": row["model_id"],
"name": _sqlite_row_get(row, "name", "") or "",
"prompt": row["prompt"],
"workload_type": _sqlite_row_get(row, "workload_type", "t2v"),
"image_path": _sqlite_row_get(row, "image_path", "") or "",
"last_image_path": _sqlite_row_get(row, "last_image_path", "") or "",
"references": _sqlite_row_get(row, "references_json", "") or "",
"job_type": _sqlite_row_get(row, "job_type", "inference"),
"status": row["status"],
"created_at": row["created_at"],
+239 -58
View File
@@ -10,7 +10,9 @@ from __future__ import annotations
import atexit
import collections
import contextlib
import copy
import enum
import json
import logging
import logging.handlers
import multiprocessing as mp
@@ -123,10 +125,13 @@ class LogBufferHandler(logging.Handler):
class Job:
id: str
model_id: str
prompt: str
name: str = ""
prompt: str = ""
workload_type: str = "t2v"
job_type: str = "inference"
image_path: str = ""
last_image_path: str = ""
references: list[dict[str, Any]] = field(default_factory=list)
status: JobStatus = JobStatus.PENDING
created_at: float = field(default_factory=time.time)
started_at: float | None = None
@@ -145,6 +150,7 @@ class Job:
negative_prompt: str = ""
num_gpus: int = 1
dit_cpu_offload: bool = False
dit_layerwise_offload: bool = False
text_encoder_cpu_offload: bool = False
vae_cpu_offload: bool = False
image_encoder_cpu_offload: bool = False
@@ -180,10 +186,13 @@ class Job:
return {
"id": self.id,
"model_id": self.model_id,
"name": self.name,
"prompt": self.prompt,
"workload_type": self.workload_type,
"job_type": self.job_type,
"image_path": self.image_path,
"last_image_path": self.last_image_path,
"references": self.references,
"status": self.status.value,
"created_at": self.created_at,
"started_at": self.started_at,
@@ -202,6 +211,7 @@ class Job:
"negative_prompt": self.negative_prompt,
"num_gpus": self.num_gpus,
"dit_cpu_offload": self.dit_cpu_offload,
"dit_layerwise_offload": self.dit_layerwise_offload,
"text_encoder_cpu_offload": self.text_encoder_cpu_offload,
"vae_cpu_offload": self.vae_cpu_offload,
"image_encoder_cpu_offload": self.image_encoder_cpu_offload,
@@ -232,6 +242,75 @@ class Job:
}
MINIMAX_H3_REF2VA_PIPELINE = "MiniMaxH3Ref2VAModularPipeline"
def _build_h3_references(raw: list[dict[str, Any]]) -> list[Any]:
"""Turn the API's reference dicts into MiniMaxH3Reference objects.
Imported lazily so the API server starts without pulling in fastvideo.
"""
from fastvideo.pipelines.basic.minimax_h3 import MiniMaxH3Reference
built = []
for i, ref in enumerate(raw):
source = (ref or {}).get("source")
if not source:
raise ValueError(f"reference {i} has no source")
if not os.path.isfile(source):
raise ValueError(f"reference {i} source not found: {source}")
kwargs: dict[str, Any] = {
"source": source,
"media_type": (ref.get("media_type") or "image"),
}
for opt in ("soundtrack", "fps", "sample_rate"):
if ref.get(opt) not in (None, ""):
kwargs[opt] = ref[opt]
built.append(MiniMaxH3Reference(**kwargs))
return built
JOB_LOG_FILENAME = "out.log"
def _job_log_path(output_dir: str, job_id: str) -> str:
"""Each job's log lives beside its outputs: <output_dir>/<job_id>/out.log."""
return os.path.join(output_dir, job_id, JOB_LOG_FILENAME)
def _decode_references(value: Any) -> list[dict[str, Any]]:
"""Reference lists round-trip through the DB as JSON text."""
if not value:
return []
if isinstance(value, list):
return list(value)
try:
decoded = json.loads(value)
except (TypeError, ValueError):
logger.warning("Could not decode stored references: %r", value)
return []
return list(decoded) if isinstance(decoded, list) else []
def _generator_is_alive(generator: Any) -> bool:
"""True if the generator's worker processes are all still running.
A cached VideoGenerator holds a MultiprocExecutor whose workers are separate
processes; nothing notices when they exit. Probing `proc.is_alive()` is what
the executor itself uses during shutdown. Anything unexpected in the object
graph is treated as alive so a probe failure can never wedge the cache.
"""
executor = getattr(generator, "executor", None)
workers = getattr(executor, "workers", None)
if not workers:
return True
try:
return all(w.proc.is_alive() for w in workers)
except Exception:
logger.debug("Worker liveness probe failed", exc_info=True)
return True
class JobRunner:
"""Manages video generation jobs, their execution, and generator caching."""
@@ -276,7 +355,7 @@ class JobRunner:
"""Populate job's log buffer from its log file if it exists."""
path = job.log_file_path
if not path:
path = os.path.join(self.log_dir, f"{job.id}.log")
path = _job_log_path(self.output_dir, job.id)
if not os.path.isfile(path):
return
try:
@@ -314,10 +393,13 @@ class JobRunner:
job = Job(
id=row["id"],
model_id=row["model_id"],
name=row.get("name", "") or "",
prompt=row["prompt"],
workload_type=row.get("workload_type", "t2v"),
job_type=row.get("job_type", "inference"),
image_path=row.get("image_path", "") or "",
last_image_path=row.get("last_image_path", "") or "",
references=_decode_references(row.get("references")),
data_path=row.get("data_path", "") or "",
max_train_steps=row.get("max_train_steps", 1000),
train_batch_size=row.get("train_batch_size", 1),
@@ -350,6 +432,7 @@ class JobRunner:
negative_prompt=row.get("negative_prompt", "") or "",
num_gpus=row.get("num_gpus", 1),
dit_cpu_offload=row.get("dit_cpu_offload", False),
dit_layerwise_offload=row.get("dit_layerwise_offload", False),
text_encoder_cpu_offload=row.get("text_encoder_cpu_offload", False),
vae_cpu_offload=row.get("vae_cpu_offload", False),
image_encoder_cpu_offload=row.get("image_encoder_cpu_offload", False),
@@ -394,9 +477,12 @@ class JobRunner:
job_id: str,
model_id: str,
prompt: str,
name: str = "",
workload_type: str = "t2v",
job_type: str = "inference",
image_path: str = "",
last_image_path: str = "",
references: list[dict[str, Any]] | None = None,
data_path: str = "",
max_train_steps: int = 1000,
train_batch_size: int = 1,
@@ -422,6 +508,7 @@ class JobRunner:
num_gpus: int = 1,
negative_prompt: str = "",
dit_cpu_offload: bool = False,
dit_layerwise_offload: bool = False,
text_encoder_cpu_offload: bool = False,
vae_cpu_offload: bool = False,
image_encoder_cpu_offload: bool = False,
@@ -435,10 +522,13 @@ class JobRunner:
job = Job(
id=job_id,
model_id=model_id,
name=(name or "").strip(),
prompt=prompt.strip(),
workload_type=workload_type or "t2v",
job_type=job_type or "inference",
image_path=image_path or "",
last_image_path=last_image_path or "",
references=list(references or []),
data_path=data_path or "",
max_train_steps=max_train_steps,
train_batch_size=train_batch_size,
@@ -464,6 +554,7 @@ class JobRunner:
negative_prompt=negative_prompt or "",
num_gpus=num_gpus,
dit_cpu_offload=dit_cpu_offload,
dit_layerwise_offload=dit_layerwise_offload,
text_encoder_cpu_offload=text_encoder_cpu_offload,
vae_cpu_offload=vae_cpu_offload,
image_encoder_cpu_offload=image_encoder_cpu_offload,
@@ -521,6 +612,84 @@ class JobRunner:
logger.info("Deleted job %s", job.id)
return True
CONFIG_FIELDS: tuple[str, ...] = (
"model_id",
"name",
"prompt",
"workload_type",
"job_type",
"image_path",
"last_image_path",
"references",
"negative_prompt",
"num_inference_steps",
"num_frames",
"height",
"width",
"guidance_scale",
"guidance_rescale",
"fps",
"seed",
"num_gpus",
"dit_cpu_offload",
"dit_layerwise_offload",
"text_encoder_cpu_offload",
"vae_cpu_offload",
"image_encoder_cpu_offload",
"use_fsdp_inference",
"enable_torch_compile",
"vsa_sparsity",
"tp_size",
"sp_size",
"data_path",
"max_train_steps",
"train_batch_size",
"learning_rate",
"num_latent_t",
"validation_dataset_file",
"lora_rank",
"dmd_use_vsa",
"dmd_vsa_sparsity",
"dmd_denoising_steps",
"real_score_guidance_scale",
"generator_update_interval",
"real_score_model_path",
"fake_score_model_path",
)
def duplicate_job(self, job_id: str, new_job_id: str) -> Job:
"""Create a new pending job with an existing job's configuration.
Runtime state (status, timings, logs, outputs) is not carried over.
"""
with self._jobs_lock:
source = self._jobs.get(job_id)
if source is None:
raise ValueError(f"Job {job_id} not found")
config = {f: copy.deepcopy(getattr(source, f)) for f in self.CONFIG_FIELDS}
return self.create_job(job_id=new_job_id, **config)
#: Editable exactly when startable: the same set start_job() accepts.
EDITABLE_STATUSES = (JobStatus.PENDING, JobStatus.FAILED, JobStatus.STOPPED)
def update_job_config(self, job_id: str, updates: dict[str, Any]) -> Job:
"""Edit the configuration of a job that has not produced a result."""
with self._jobs_lock:
job = self._jobs.get(job_id)
if job is None:
raise ValueError(f"Job {job_id} not found")
if job.status not in self.EDITABLE_STATUSES:
allowed = ", ".join(s.value for s in self.EDITABLE_STATUSES)
raise ValueError(f"Job is {job.status.value}; only {allowed} jobs can be edited. "
"Duplicate it instead.")
unknown = set(updates) - set(self.CONFIG_FIELDS)
if unknown:
raise ValueError(f"Not editable: {', '.join(sorted(unknown))}")
for field_name, value in updates.items():
setattr(job, field_name, value)
self._save_job(job)
return job
def start_job(self, job_id: str) -> Job:
"""Start (or restart) a pending / stopped / failed job.
@@ -623,6 +792,8 @@ class JobRunner:
workload_type: str,
num_gpus: int,
dit_cpu_offload: bool = False,
dit_layerwise_offload: bool = False,
override_pipeline_cls_name: str | None = None,
text_encoder_cpu_offload: bool = False,
vae_cpu_offload: bool = False,
image_encoder_cpu_offload: bool = False,
@@ -638,6 +809,10 @@ class JobRunner:
workload_type,
num_gpus,
dit_cpu_offload,
dit_layerwise_offload,
# Ref2VA loads different DiT weights (transformer_ref), so the
# override must key the cache or a t2v/i2v generator gets reused.
override_pipeline_cls_name,
text_encoder_cpu_offload,
vae_cpu_offload,
image_encoder_cpu_offload,
@@ -650,8 +825,21 @@ class JobRunner:
# Generators are cached by model_id and configuration parameters
with self._generators_lock:
if cache_key in self._generators:
return self._generators[cache_key]
cached = self._generators.get(cache_key)
if cached is not None:
if _generator_is_alive(cached):
return cached
# Workers can exit while a generator sits idle in the cache;
# reusing it fails every later job with the same config.
logger.warning(
"Cached generator for %s has dead workers; reloading.",
model_id,
)
self._generators.pop(cache_key, None)
try:
cached.shutdown()
except Exception:
logger.debug("Shutdown of the dead generator failed", exc_info=True)
# Import lazily so starting the server is fast even without a GPU.
from fastvideo import VideoGenerator
@@ -677,6 +865,11 @@ class JobRunner:
gen = VideoGenerator.from_pretrained(
model_id,
workload_type=workload_type,
num_gpus=num_gpus,
dit_layerwise_offload=dit_layerwise_offload,
**({
"override_pipeline_cls_name": override_pipeline_cls_name
} if override_pipeline_cls_name else {}),
dit_cpu_offload=dit_cpu_offload,
text_encoder_cpu_offload=text_encoder_cpu_offload,
vae_cpu_offload=vae_cpu_offload,
@@ -706,10 +899,9 @@ class JobRunner:
def _run_training_job(self, job: Job):
"""Run a finetuning, distillation, or LoRA job via subprocess."""
buf = job._log_buf
os.makedirs(self.log_dir, exist_ok=True)
job.log_file_path = os.path.join(self.log_dir, f"{job.id}.log")
job_output_dir = os.path.join(self.output_dir, job.id)
os.makedirs(job_output_dir, exist_ok=True)
job.log_file_path = _job_log_path(self.output_dir, job.id)
if not job.data_path or not os.path.isdir(job.data_path):
job.status = JobStatus.FAILED
@@ -827,8 +1019,8 @@ class JobRunner:
def _run_inference_job(self, job: Job):
buf = job._log_buf
os.makedirs(self.log_dir, exist_ok=True)
job.log_file_path = os.path.join(self.log_dir, f"{job.id}.log")
os.makedirs(os.path.join(self.output_dir, job.id), exist_ok=True)
job.log_file_path = _job_log_path(self.output_dir, job.id)
# Add file handler to persist logs
file_handler = logging.FileHandler(job.log_file_path, mode='w', encoding='utf-8')
@@ -875,62 +1067,45 @@ class JobRunner:
buf.phase = "loading model"
logger.info("Loading model...")
# Run generator creation in a background thread so we
# can poll _stop_event while the (potentially slow)
# model download / load is in progress.
_gen_result: list[Any] = []
_gen_error: list[BaseException] = []
# The generator MUST be created on this thread: building it spawns
# the executor's worker processes, and they are torn down if the
# creating thread exits. Running it in a helper thread (to poll
# _stop_event during load) made every collective_rpc fail with
# ConnectionResetError.
if job._stop_event.is_set():
job.status = JobStatus.STOPPED
job.finished_at = time.time()
self._save_job(job)
logger.warning("Job %s stopped before model loading", job.id)
buf.phase = "stopped"
return
def _load_generator() -> None:
try:
gen = self._get_or_create_generator(
job.model_id,
job.workload_type,
job.num_gpus,
dit_cpu_offload=job.dit_cpu_offload,
text_encoder_cpu_offload=(job.text_encoder_cpu_offload),
vae_cpu_offload=job.vae_cpu_offload,
image_encoder_cpu_offload=(job.image_encoder_cpu_offload),
use_fsdp_inference=job.use_fsdp_inference,
enable_torch_compile=(job.enable_torch_compile),
vsa_sparsity=job.vsa_sparsity,
tp_size=job.tp_size,
sp_size=job.sp_size,
log_queue=log_queue,
)
_gen_result.append(gen)
except BaseException as exc:
_gen_error.append(exc)
loader = threading.Thread(
target=_load_generator,
daemon=True,
generator = self._get_or_create_generator(
job.model_id,
job.workload_type,
job.num_gpus,
dit_cpu_offload=job.dit_cpu_offload,
dit_layerwise_offload=job.dit_layerwise_offload,
override_pipeline_cls_name=(MINIMAX_H3_REF2VA_PIPELINE if job.references else None),
text_encoder_cpu_offload=(job.text_encoder_cpu_offload),
vae_cpu_offload=job.vae_cpu_offload,
image_encoder_cpu_offload=(job.image_encoder_cpu_offload),
use_fsdp_inference=job.use_fsdp_inference,
enable_torch_compile=(job.enable_torch_compile),
vsa_sparsity=job.vsa_sparsity,
tp_size=job.tp_size,
sp_size=job.sp_size,
log_queue=log_queue,
)
loader.start()
while loader.is_alive():
if job._stop_event.is_set():
job.status = JobStatus.STOPPED
job.finished_at = time.time()
self._save_job(job)
logger.warning(
"Job %s stopped during model loading",
job.id,
)
buf.phase = "stopped"
return
loader.join(timeout=0.5)
if _gen_error:
raise _gen_error[0]
generator = _gen_result[0]
buf.phase = "generating"
logger.info("Starting generation for job %s (model=%s)", job.id, job.model_id)
# Without a name FastVideo derives the filename from the prompt.
safe_name = re.sub(r'[\\/:*?"<>|]+', "", job.name).strip().strip(".")
output_target = (os.path.join(job_output_dir, f"{safe_name[:80]}.mp4") if safe_name else job_output_dir)
gen_kwargs: dict[str, Any] = {
"prompt": job.prompt,
"output_path": job_output_dir,
"output_path": output_target,
"save_video": True,
"num_inference_steps": job.num_inference_steps,
"num_frames": job.num_frames,
@@ -945,6 +1120,12 @@ class JobRunner:
}
if job.image_path:
gen_kwargs["image_path"] = job.image_path
if job.references:
gen_kwargs["references"] = _build_h3_references(job.references)
if job.last_image_path:
# _prepare_fl2va requires a PIL image, not a path.
from PIL import Image as _PILImage
gen_kwargs["last_image"] = _PILImage.open(job.last_image_path)
generator.generate_video(**gen_kwargs)
buf.phase = "saving"
@@ -977,7 +1158,7 @@ class JobRunner:
except Exception as exception:
error_msg = str(exception)
logger.error("Critical error in job thread: %s", error_msg)
logger.exception("Critical error in job thread: %s", error_msg)
job.status = JobStatus.FAILED
job.error = f"Critical error ({type(exception).__name__}): {error_msg}"
job.finished_at = time.time()
@@ -1,15 +1,20 @@
# SPDX-License-Identifier: Apache-2.0
"""Request model for creating a job."""
from typing import Any
from pydantic import BaseModel
class CreateJobRequest(BaseModel):
model_id: str
name: str = ""
prompt: str
workload_type: str = "t2v"
job_type: str = "inference"
image_path: str = ""
last_image_path: str = ""
references: list[dict[str, Any]] | None = None
data_path: str = ""
max_train_steps: int = 1000
train_batch_size: int = 1
@@ -28,6 +33,7 @@ class CreateJobRequest(BaseModel):
seed: int = 1024
num_gpus: int = 1
dit_cpu_offload: bool = False
dit_layerwise_offload: bool = False
text_encoder_cpu_offload: bool = False
vae_cpu_offload: bool = False
image_encoder_cpu_offload: bool = False
+92 -2
View File
@@ -18,6 +18,7 @@ import argparse
import contextlib
import logging
import os
import re
import shutil
import signal
import time
@@ -116,7 +117,30 @@ def list_models(workload_type: str | None = None) -> list[dict[str, Any]]:
return _available_models
def _safe_upload_name(filename: str | None, ext: str) -> str:
"""A filesystem-safe version of the client's filename, keeping it readable.
Uploads live under a per-file uuid directory, so the basename does not have
to be unique -- only safe. Keeping the original name means the path stays
self-describing wherever it travels: the database, job logs, and payloads
copied back out to the API.
"""
stem = os.path.basename(filename or "").rsplit(".", 1)[0]
stem = re.sub(r"[^A-Za-z0-9._-]+", "_", stem).strip("._-")
return f"{stem[:80] or 'upload'}{ext}"
def _upload_destination(ext: str, filename: str | None) -> str:
"""<upload_dir>/<uuid4>/<safe original name><ext>"""
directory = os.path.join(upload_dir, uuid.uuid4().hex)
os.makedirs(directory, exist_ok=True)
return os.path.join(directory, _safe_upload_name(filename, ext))
ALLOWED_IMAGE_EXTENSIONS = {".png", ".jpg", ".jpeg", ".webp", ".bmp"}
ALLOWED_VIDEO_EXTENSIONS = {".mp4", ".mov", ".mkv", ".webm", ".avi"}
ALLOWED_AUDIO_EXTENSIONS = {".wav", ".mp3", ".flac", ".m4a", ".ogg"}
ALLOWED_MEDIA_EXTENSIONS = (ALLOWED_IMAGE_EXTENSIONS | ALLOWED_VIDEO_EXTENSIONS | ALLOWED_AUDIO_EXTENSIONS)
@app.post("/api/upload-image")
@@ -136,8 +160,7 @@ async def upload_image(file: Annotated[UploadFile, File()], ) -> dict[str, str]:
f"{', '.join(ALLOWED_IMAGE_EXTENSIONS)}"),
)
os.makedirs(upload_dir, exist_ok=True)
unique_name = f"{uuid.uuid4().hex}{ext}"
dest_path = os.path.join(upload_dir, unique_name)
dest_path = _upload_destination(ext, file.filename)
try:
contents = await file.read()
with open(dest_path, "wb") as f:
@@ -150,6 +173,47 @@ async def upload_image(file: Annotated[UploadFile, File()], ) -> dict[str, str]:
return {"path": os.path.abspath(dest_path)}
@app.post("/api/upload-media")
async def upload_media(file: Annotated[UploadFile, File()], ) -> dict[str, str]:
"""Upload an image, video or audio file for Ref2VA references.
Returns the absolute path plus the media_type MiniMax-H3 expects, so the
caller does not have to re-derive it from the extension.
"""
global upload_dir # noqa: PLW0603
if not upload_dir:
raise HTTPException(
status_code=503,
detail="Upload directory not configured",
)
ext = Path(file.filename or "").suffix.lower()
if ext not in ALLOWED_MEDIA_EXTENSIONS:
raise HTTPException(
status_code=400,
detail=(f"Invalid file type. Allowed: "
f"{', '.join(sorted(ALLOWED_MEDIA_EXTENSIONS))}"),
)
if ext in ALLOWED_VIDEO_EXTENSIONS:
media_type = "video"
elif ext in ALLOWED_AUDIO_EXTENSIONS:
media_type = "audio"
else:
media_type = "image"
os.makedirs(upload_dir, exist_ok=True)
dest_path = _upload_destination(ext, file.filename)
try:
contents = await file.read()
with open(dest_path, "wb") as f:
f.write(contents)
except OSError as e:
raise HTTPException(
status_code=500,
detail=f"Failed to save upload: {e}",
) from e
return {"path": os.path.abspath(dest_path), "media_type": media_type}
ALLOWED_VIDEO_EXTENSIONS = {".mp4", ".webm", ".avi", ".mov", ".mkv"}
@@ -282,10 +346,13 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
job = job_runner.create_job(
job_id=str(uuid.uuid4()),
model_id=req.model_id,
name=req.name or "",
prompt=req.prompt,
workload_type=req.workload_type or "t2v",
job_type=job_type,
image_path=req.image_path or "",
last_image_path=req.last_image_path or "",
references=req.references or [],
data_path=data_path,
max_train_steps=req.max_train_steps,
train_batch_size=req.train_batch_size,
@@ -304,6 +371,7 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
seed=req.seed,
num_gpus=req.num_gpus,
dit_cpu_offload=req.dit_cpu_offload,
dit_layerwise_offload=req.dit_layerwise_offload,
text_encoder_cpu_offload=req.text_encoder_cpu_offload,
vae_cpu_offload=req.vae_cpu_offload,
image_encoder_cpu_offload=req.image_encoder_cpu_offload,
@@ -336,6 +404,28 @@ def create_job(req: CreateJobRequest) -> dict[str, Any]:
return job.to_dict()
@app.post("/api/jobs/{job_id}/duplicate", status_code=201)
def duplicate_job(job_id: str) -> dict[str, Any]:
"""Create a new pending job with the same configuration as an existing one."""
try:
job = job_runner.duplicate_job(job_id, str(uuid.uuid4()))
except ValueError as e:
raise HTTPException(status_code=404, detail=str(e)) from e
return job.to_dict()
@app.patch("/api/jobs/{job_id}")
def update_job(job_id: str, updates: dict[str, Any]) -> dict[str, Any]:
"""Edit a pending job's configuration. Started jobs cannot be edited."""
try:
job = job_runner.update_job_config(job_id, updates)
except ValueError as e:
detail = str(e)
status = 404 if "not found" in detail else 400
raise HTTPException(status_code=status, detail=detail) from e
return job.to_dict()
@app.post("/api/jobs/{job_id}/start")
def start_job(job_id: str) -> dict[str, Any]:
"""Start (or restart) a pending / stopped / failed job."""
@@ -22,15 +22,35 @@ import { useStore } from '@/hooks/useStore';
import { defaultOptionsStore } from '@/stores/defaultOptions';
import {
createJob,
updateJob,
getDatasets,
getModels,
uploadImage,
uploadMedia,
type CreateJobRequest,
type Model,
} from '@/lib/api';
import { getDefaultModelForWorkload } from '@/lib/defaultOptions';
import { WORKLOAD_OPTIONS } from '@/lib/jobConfig';
import type { JobType } from '@/lib/types';
import {
H3_MAX_REFERENCES,
labelReferences,
referencePromptSeed,
validateReferences,
type H3Reference,
} from '@/lib/h3References';
import {
EMPTY_H3_PROMPT_FIELDS,
H3_PROMPT_SECTIONS,
H3_SECTION_HINTS,
H3_SECTION_LABELS,
isEmptyPromptFields,
parseH3Prompt,
serializeH3Prompt,
type H3PromptFields,
} from '@/lib/h3Prompt';
import { jobToFormFields, type JobLike } from '@/lib/jobToFields';
export interface CreateJobModalProps {
isOpen: boolean;
@@ -38,6 +58,10 @@ export interface CreateJobModalProps {
onSuccess: () => void;
jobType: JobType;
workloadType: string;
/** When set, the modal edits this pending job instead of creating a new one. */
editingJob?: JobLike | null;
/** Show the configuration without allowing changes (started/finished jobs). */
readOnly?: boolean;
}
export default function CreateJobModal({
@@ -46,6 +70,8 @@ export default function CreateJobModal({
onSuccess,
jobType,
workloadType,
editingJob,
readOnly = false,
}: CreateJobModalProps) {
const { options } = useStore(defaultOptionsStore);
@@ -55,8 +81,22 @@ export default function CreateJobModal({
const [models, setModels] = React.useState<Model[]>([]);
const [modelId, setModelId] = React.useState('');
const [name, setName] = React.useState('');
const [prompt, setPrompt] = React.useState('');
const [imagePath, setImagePath] = React.useState('');
const [lastImagePath, setLastImagePath] = React.useState('');
const [references, setReferences] = React.useState<H3Reference[]>([]);
const [isUploadingReference, setIsUploadingReference] = React.useState(false);
const [referenceError, setReferenceError] = React.useState<string | null>(null);
const [promptFields, setPromptFields] = React.useState<H3PromptFields>(
EMPTY_H3_PROMPT_FIELDS,
);
const [useGuidedPrompt, setUseGuidedPrompt] = React.useState(true);
const [lastImageFileName, setLastImageFileName] = React.useState('');
const [isUploadingLastImage, setIsUploadingLastImage] = React.useState(false);
const [lastImageUploadError, setLastImageUploadError] = React.useState<
string | null
>(null);
const [imageFileName, setImageFileName] = React.useState('');
const [isUploadingImage, setIsUploadingImage] = React.useState(false);
const [negativePrompt, setNegativePrompt] = React.useState('');
@@ -70,12 +110,50 @@ export default function CreateJobModal({
const [seed, setSeed] = React.useState(1024);
const [numGpus, setNumGpus] = React.useState(1);
const [ditCpuOffload, setDitCpuOffload] = React.useState(false);
const [ditLayerwiseOffload, setDitLayerwiseOffload] = React.useState(false);
const [textEncoderCpuOffload, setTextEncoderCpuOffload] =
React.useState(false);
const [vaeCpuOffload, setVaeCpuOffload] = React.useState(false);
const [imageEncoderCpuOffload, setImageEncoderCpuOffload] =
React.useState(false);
const [useFsdpInference, setUseFsdpInference] = React.useState(false);
// H3 is the only registered model with an end frame or references.
const supportsLastImage = modelId.toLowerCase().includes('minimax-h3');
const usingReferences = supportsLastImage && references.length > 0;
// JobCard re-renders on every job-list poll, so `editingJob` is a fresh
// object each time. Effects must depend on these, never on the object.
const editingJobId = editingJob?.id ?? null;
const editingJobModelId = editingJob?.model_id ?? null;
// Layerwise offload and FSDP compete for the DiT weights and FastVideoArgs
// silently picks a winner (fastvideo_args.py:859); resolve it visibly here.
// dit_cpu_offload is deliberately not interlocked -- it is a modifier, not a
// competing strategy.
const handleDitLayerwiseOffloadChange = React.useCallback((next: boolean) => {
setDitLayerwiseOffload(next);
if (next) {
setUseFsdpInference(false);
}
}, []);
const handleUseFsdpInferenceChange = React.useCallback((next: boolean) => {
setUseFsdpInference(next);
if (next) {
setDitLayerwiseOffload(false);
}
}, []);
const handleNumGpusChange = React.useCallback((next: number) => {
setNumGpus(next);
if (next > 1) {
// Dropping back to one GPU leaves FSDP alone: single-GPU FSDP is a
// valid way to reach its CPU offload (docs/inference/offloading.md).
setUseFsdpInference(true);
setDitLayerwiseOffload(false);
}
}, []);
const [enableTorchCompile, setEnableTorchCompile] = React.useState(false);
const [vsaSparsity, setVsaSparsity] = React.useState(0);
const [tpSize, setTpSize] = React.useState(-1);
@@ -126,6 +204,47 @@ export default function CreateJobModal({
const justOpened = isOpen && !justOpenedRef.current;
justOpenedRef.current = isOpen;
if (!justOpened) return;
if (editingJob) {
// Must not fall through to the defaults below: a partially-seeded
// form silently edits values the user never saw.
const f = jobToFormFields(editingJob);
setModelId(f.modelId);
setName(f.name);
setPrompt(f.prompt);
setNegativePrompt(f.negativePrompt);
setImagePath(f.imagePath);
setImageFileName(f.imagePath.split('/').pop() ?? '');
setLastImagePath(f.lastImagePath);
setLastImageFileName(f.lastImagePath.split('/').pop() ?? '');
setReferences(f.references);
setPromptFields(f.promptFields ?? EMPTY_H3_PROMPT_FIELDS);
setUseGuidedPrompt(f.promptFields !== null);
setNumInferenceSteps(f.numInferenceSteps);
setNumFrames(f.numFrames);
setHeight(f.height);
setWidth(f.width);
setGuidanceScale(f.guidanceScale);
setGuidanceRescale(f.guidanceRescale);
setFps(f.fps);
setSeed(f.seed);
setNumGpus(f.numGpus);
setDitCpuOffload(f.ditCpuOffload);
setDitLayerwiseOffload(f.ditLayerwiseOffload);
setTextEncoderCpuOffload(f.textEncoderCpuOffload);
setVaeCpuOffload(f.vaeCpuOffload);
setImageEncoderCpuOffload(f.imageEncoderCpuOffload);
setUseFsdpInference(f.useFsdpInference);
setEnableTorchCompile(f.enableTorchCompile);
setVsaSparsity(f.vsaSparsity);
setTpSize(f.tpSize);
setSpSize(f.spSize);
setReferenceError(null);
setModelLoadError(null);
setImageUploadError(null);
setLastImageUploadError(null);
setSubmitError(null);
return;
}
const opts = options;
setNumInferenceSteps(opts.numInferenceSteps);
setNumFrames(workloadType === 't2i' ? 1 : opts.numFrames);
@@ -137,6 +256,7 @@ export default function CreateJobModal({
setSeed(opts.seed);
setNumGpus(opts.numGpus);
setDitCpuOffload(opts.ditCpuOffload);
setDitLayerwiseOffload(opts.ditLayerwiseOffload ?? false);
setTextEncoderCpuOffload(opts.textEncoderCpuOffload);
setVaeCpuOffload(opts.vaeCpuOffload);
setImageEncoderCpuOffload(opts.imageEncoderCpuOffload);
@@ -151,8 +271,16 @@ export default function CreateJobModal({
inferenceWorkload as 't2v' | 'i2v' | 't2i',
),
);
setName('');
setImagePath('');
setImageFileName('');
setLastImagePath('');
setLastImageFileName('');
setLastImageUploadError(null);
setReferences([]);
setReferenceError(null);
setPromptFields(EMPTY_H3_PROMPT_FIELDS);
setUseGuidedPrompt(true);
setSelectedDatasetId('');
setSelectedValidationDatasetId('');
setModelLoadError(null);
@@ -168,7 +296,7 @@ export default function CreateJobModal({
setRealScoreModelPath('');
setFakeScoreModelPath('');
}
}, [isOpen, workloadType, inferenceWorkload, options]);
}, [isOpen, workloadType, inferenceWorkload, options, editingJobId]);
// Load the models available for this workload.
React.useEffect(() => {
@@ -188,7 +316,16 @@ export default function CreateJobModal({
opts,
inferenceWorkload as 't2v' | 'i2v' | 't2i',
);
const chosen = ids.includes(defaultId) ? defaultId : (list[0]?.id ?? '');
// When editing, the job's own model wins over the workload default --
// this resolves after the seeding effect, so choosing a default here
// would silently swap the model out from under the user.
const editedId = editingJobModelId;
const chosen =
editedId && ids.includes(editedId)
? editedId
: ids.includes(defaultId)
? defaultId
: (list[0]?.id ?? '');
setModelId(chosen);
if (workloadType === 'dmd_t2v') {
setRealScoreModelPath(chosen);
@@ -210,7 +347,7 @@ export default function CreateJobModal({
return () => {
stale = true;
};
}, [isOpen, inferenceWorkload, workloadType]);
}, [isOpen, inferenceWorkload, workloadType, editingJobModelId]);
// Training jobs need a dataset; load the ready datasets when relevant.
React.useEffect(() => {
@@ -262,6 +399,106 @@ export default function CreateJobModal({
}
}
async function handleLastImageChange(
e: React.ChangeEvent<HTMLInputElement>,
) {
const file = e.target.files?.[0];
if (!file) {
setLastImagePath('');
setLastImageFileName('');
setLastImageUploadError(null);
return;
}
setIsUploadingLastImage(true);
setLastImageFileName(file.name);
setLastImageUploadError(null);
try {
const { path } = await uploadImage(file);
setLastImagePath(path);
} catch (error) {
console.error('Failed to upload end image:', error);
setLastImagePath('');
setLastImageFileName('');
setLastImageUploadError(
error instanceof Error
? `${error.message}. Choose the image again to retry.`
: 'The image could not be uploaded. Choose it again to retry.',
);
} finally {
setIsUploadingLastImage(false);
}
}
async function handleAddReference(
e: React.ChangeEvent<HTMLInputElement>,
) {
const file = e.target.files?.[0];
e.target.value = ''; // allow re-picking the same file
if (!file) return;
setIsUploadingReference(true);
setReferenceError(null);
try {
const { path, media_type } = await uploadMedia(file);
const next: H3Reference[] = [
...references,
{
id: `${Date.now()}-${file.name}`,
source: path,
media_type,
fileName: file.name,
},
];
setReferences(next);
setReferenceError(validateReferences(next));
} catch (error) {
console.error('Failed to upload reference:', error);
setReferenceError(
error instanceof Error ? error.message : 'The file could not be uploaded.',
);
} finally {
setIsUploadingReference(false);
}
}
function removeReference(id: string) {
const next = references.filter((r) => r.id !== id);
setReferences(next);
setReferenceError(validateReferences(next));
}
function seedPromptFields() {
setPromptFields({
...EMPTY_H3_PROMPT_FIELDS,
...referencePromptSeed(references),
});
setUseGuidedPrompt(true);
}
function setPromptField(section: string, value: string) {
setPromptFields((prev) => ({ ...prev, [section]: value }));
}
// Switching between the guided fields and the raw editor keeps whatever was
// typed: serialize on the way out, parse back on the way in.
function toggleGuidedPrompt() {
if (useGuidedPrompt) {
if (!isEmptyPromptFields(promptFields)) {
setPrompt(serializeH3Prompt(promptFields));
}
setUseGuidedPrompt(false);
} else {
const parsed = parseH3Prompt(prompt);
if (parsed) setPromptFields(parsed);
setUseGuidedPrompt(true);
}
}
function clearLastImage() {
setLastImagePath('');
setLastImageFileName('');
setLastImageUploadError(null);
}
function clearImage() {
setImagePath('');
setImageFileName('');
@@ -271,7 +508,16 @@ export default function CreateJobModal({
async function handleSubmit(e: React.FormEvent<HTMLFormElement>) {
e.preventDefault();
if (isInference && workloadType === 'i2v' && !imagePath) return;
if (isInference && workloadType === 'i2v' && !imagePath && !usingReferences)
return;
if (usingReferences && validateReferences(references)) return;
if (
usingReferences &&
useGuidedPrompt &&
isEmptyPromptFields(promptFields) &&
!prompt.trim()
)
return;
// Send the dataset id; the backend resolves it to the on-disk media dir.
const effectiveDataPath = selectedDatasetId ?? '';
if (!isInference && !selectedDatasetId) return;
@@ -285,14 +531,35 @@ export default function CreateJobModal({
try {
const payload: CreateJobRequest = {
model_id: modelId,
prompt,
name: name.trim(),
prompt:
usingReferences && useGuidedPrompt && !isEmptyPromptFields(promptFields)
? serializeH3Prompt(promptFields)
: prompt,
workload_type: workloadType,
job_type: effectiveJobType,
...(isInference
? {
...(workloadType === 'i2v' && imagePath
// Ref2VA and the FL2VA keyframes are mutually exclusive:
// _prepare_ref2va rejects image_path/last_image_path outright
// when references are present.
...(workloadType === 'i2v' && !usingReferences && imagePath
? { image_path: imagePath }
: {}),
...(workloadType === 'i2v' &&
supportsLastImage &&
!usingReferences &&
lastImagePath
? { last_image_path: lastImagePath }
: {}),
...(workloadType === 'i2v' && supportsLastImage && references.length
? {
references: references.map((r) => ({
source: r.source,
media_type: r.media_type,
})),
}
: {}),
negative_prompt: negativePrompt,
num_inference_steps: numInferenceSteps,
num_frames: numFrames,
@@ -304,6 +571,7 @@ export default function CreateJobModal({
seed,
num_gpus: numGpus,
dit_cpu_offload: ditCpuOffload,
dit_layerwise_offload: ditLayerwiseOffload,
text_encoder_cpu_offload: textEncoderCpuOffload,
vae_cpu_offload: vaeCpuOffload,
image_encoder_cpu_offload: imageEncoderCpuOffload,
@@ -334,7 +602,14 @@ export default function CreateJobModal({
: {}),
}),
};
await createJob(payload);
if (editingJob) {
await updateJob(
editingJob.id,
payload as unknown as Record<string, unknown>,
);
} else {
await createJob(payload);
}
onSuccess();
onClose();
} catch (err) {
@@ -356,9 +631,9 @@ export default function CreateJobModal({
const workloadLabel =
WORKLOAD_OPTIONS[jobType]?.find((o) => o.type === workloadType)?.label ?? '';
const title = `New ${jobType.charAt(0).toUpperCase() + jobType.slice(1)} Job${
workloadLabel ? ` (${workloadLabel})` : ''
}`;
const title = `${readOnly ? 'View' : editingJob ? 'Edit' : 'New'} ${
jobType.charAt(0).toUpperCase() + jobType.slice(1)
} Job${workloadLabel ? ` (${workloadLabel})` : ''}`;
return (
<Dialog
@@ -385,6 +660,23 @@ export default function CreateJobModal({
autoComplete="off"
className="flex flex-col gap-3.5"
>
{/* disabled cascades to every control inside; display:contents
keeps the parent's flex layout. */}
<fieldset
disabled={readOnly}
style={{ display: 'contents' }}
className="contents"
>
<FieldRow htmlFor="modal-name" label="Name (optional)">
<Input
id="modal-name"
value={name}
onChange={(e) => setName(e.target.value)}
placeholder="Shown on the job card and used for the output filename"
disabled={isSubmitting}
/>
</FieldRow>
<FieldRow htmlFor="modal-modelId" label="Model">
<NativeSelect
id="modal-modelId"
@@ -429,12 +721,12 @@ export default function CreateJobModal({
type="file"
accept=".png,.jpg,.jpeg,.webp,.bmp"
onChange={handleImageChange}
disabled={isSubmitting || isUploadingImage}
disabled={isSubmitting || isUploadingImage || usingReferences}
aria-describedby={
imageUploadError ? 'modal-image-error' : undefined
}
aria-invalid={imageUploadError ? true : undefined}
required
required={!usingReferences}
className="h-auto py-2 file:mr-3 file:cursor-pointer file:rounded-md file:border-0 file:bg-secondary file:px-2 file:py-1 file:text-sm file:text-secondary-foreground"
/>
{imageFileName && (
@@ -462,24 +754,169 @@ export default function CreateJobModal({
</FieldRow>
)}
<FieldRow
htmlFor="modal-prompt"
label={isInference ? 'Prompt' : 'Description'}
>
<Textarea
id="modal-prompt"
value={prompt}
onChange={(e) => setPrompt(e.target.value)}
rows={isInference ? 3 : 2}
placeholder={
isInference
? 'A curious raccoon peers through a vibrant field of yellow sunflowers…'
: 'Brief description of this training job…'
}
required
disabled={isSubmitting}
/>
</FieldRow>
{isInference && workloadType === 'i2v' && supportsLastImage && (
<FieldRow htmlFor="modal-last-image" label="End Frame (optional)">
<Input
id="modal-last-image"
type="file"
accept=".png,.jpg,.jpeg,.webp,.bmp"
onChange={handleLastImageChange}
disabled={isSubmitting || isUploadingLastImage}
aria-describedby={
lastImageUploadError ? 'modal-last-image-error' : undefined
}
aria-invalid={lastImageUploadError ? true : undefined}
className="h-auto py-2 file:mr-3 file:cursor-pointer file:rounded-md file:border-0 file:bg-secondary file:px-2 file:py-1 file:text-sm file:text-secondary-foreground"
/>
{lastImageFileName && (
<span className="mt-0.5 text-xs text-muted-foreground">
{isUploadingLastImage ? 'Uploading…' : lastImageFileName} ·{' '}
<button
type="button"
onClick={clearLastImage}
disabled={isSubmitting || isUploadingLastImage}
className="text-accent-blue underline-offset-2 hover:underline disabled:cursor-not-allowed disabled:opacity-50"
>
Clear
</button>
</span>
)}
{lastImageUploadError && (
<p
id="modal-last-image-error"
role="alert"
className="text-sm text-destructive"
>
{lastImageUploadError}
</p>
)}
</FieldRow>
)}
{isInference && workloadType === 'i2v' && supportsLastImage && (
<FieldRow htmlFor="modal-reference" label="References (Ref2VA)">
<Input
id="modal-reference"
type="file"
accept=".png,.jpg,.jpeg,.webp,.bmp,.mp4,.mov,.mkv,.webm,.avi,.wav,.mp3,.flac,.m4a,.ogg"
onChange={handleAddReference}
disabled={
isSubmitting ||
isUploadingReference ||
references.length >= H3_MAX_REFERENCES
}
className="h-auto py-2 file:mr-3 file:cursor-pointer file:rounded-md file:border-0 file:bg-secondary file:px-2 file:py-1 file:text-sm file:text-secondary-foreground"
/>
{isUploadingReference && (
<span className="mt-0.5 text-xs text-muted-foreground">
Uploading…
</span>
)}
{references.length > 0 && (
<ul className="mt-1 flex list-none flex-col gap-1 p-0">
{references.map((reference, index) => (
<li
key={reference.id}
className="flex items-center gap-2 text-xs text-muted-foreground"
>
<code className="font-mono text-accent-blue">
{labelReferences(references)[index]}
</code>
<span className="truncate">{reference.fileName}</span>
<button
type="button"
onClick={() => removeReference(reference.id)}
disabled={isSubmitting}
className="ml-auto text-accent-blue underline-offset-2 hover:underline disabled:cursor-not-allowed disabled:opacity-50"
>
Remove
</button>
</li>
))}
</ul>
)}
{references.length > 0 && (
<button
type="button"
onClick={seedPromptFields}
disabled={isSubmitting}
className="mt-1 self-start text-xs text-accent-blue underline-offset-2 hover:underline disabled:cursor-not-allowed disabled:opacity-50"
>
Fill prompt sections from references
</button>
)}
{referenceError && (
<p role="alert" className="mt-0.5 text-xs text-destructive">
{referenceError}
</p>
)}
<span className="mt-0.5 text-xs text-muted-foreground">
Ref2VA replaces the keyframes: up to 9 images, 3 videos, 3 audio
(12 total). Audio needs at least one image or video.
</span>
</FieldRow>
)}
{usingReferences && useGuidedPrompt ? (
/* Six-section format from the model's reference prompt guide. */
<>
{H3_PROMPT_SECTIONS.map((section) => (
<FieldRow
key={section}
htmlFor={`modal-prompt-${section}`}
label={H3_SECTION_LABELS[section]}
>
<Textarea
id={`modal-prompt-${section}`}
value={promptFields[section]}
onChange={(e) => setPromptField(section, e.target.value)}
rows={section === 'detailed_description' ? 5 : 2}
placeholder={H3_SECTION_HINTS[section]}
disabled={isSubmitting}
/>
</FieldRow>
))}
<button
type="button"
onClick={toggleGuidedPrompt}
disabled={isSubmitting}
className="self-start text-xs text-accent-blue underline-offset-2 hover:underline disabled:cursor-not-allowed disabled:opacity-50"
>
Edit as raw prompt
</button>
</>
) : (
<>
<FieldRow
htmlFor="modal-prompt"
label={isInference ? 'Prompt' : 'Description'}
>
<Textarea
id="modal-prompt"
value={prompt}
onChange={(e) => setPrompt(e.target.value)}
rows={isInference ? 3 : 2}
placeholder={
isInference
? 'A curious raccoon peers through a vibrant field of yellow sunflowers…'
: 'Brief description of this training job…'
}
required={!(usingReferences && useGuidedPrompt)}
disabled={isSubmitting}
/>
</FieldRow>
{usingReferences && (
<button
type="button"
onClick={toggleGuidedPrompt}
disabled={isSubmitting}
className="self-start text-xs text-accent-blue underline-offset-2 hover:underline disabled:cursor-not-allowed disabled:opacity-50"
>
Edit as prompt sections
</button>
)}
</>
)}
{isInference && (
<FieldRow htmlFor="modal-negative-prompt" label="Negative Prompt">
@@ -849,6 +1286,13 @@ export default function CreateJobModal({
onChange={setDitCpuOffload}
disabled={isSubmitting}
/>
<ToggleRow
id="modal-dit-layerwise-offload"
label="DiT Layerwise Offload"
checked={ditLayerwiseOffload}
onChange={handleDitLayerwiseOffloadChange}
disabled={isSubmitting}
/>
<ToggleRow
id="modal-text-encoder-cpu-offload"
label="Text Encoder CPU Offload"
@@ -860,7 +1304,7 @@ export default function CreateJobModal({
id="modal-use-fsdp-inference"
label="Use FSDP Inference"
checked={useFsdpInference}
onChange={setUseFsdpInference}
onChange={handleUseFsdpInferenceChange}
disabled={isSubmitting}
/>
<ToggleRow
@@ -891,7 +1335,7 @@ export default function CreateJobModal({
max={8}
step={1}
value={numGpus}
onChange={setNumGpus}
onChange={handleNumGpusChange}
disabled={isSubmitting}
/>
<NumberRow
@@ -906,6 +1350,8 @@ export default function CreateJobModal({
</details>
)}
</fieldset>
<div className="flex flex-col items-start gap-2">
{submitError && (
<p role="alert" className="text-sm text-destructive">
@@ -914,14 +1360,22 @@ export default function CreateJobModal({
)}
<Button
type="submit"
hidden={readOnly}
disabled={
readOnly ||
isSubmitting ||
isUploadingImage ||
!!modelLoadError ||
!!datasetLoadError
}
>
{isSubmitting ? 'Creating…' : 'Create Job'}
{isSubmitting
? editingJob
? 'Saving…'
: 'Creating…'
: editingJob
? 'Save Changes'
: 'Create Job'}
</Button>
</div>
</form>
@@ -3,11 +3,13 @@
import * as React from 'react';
import { Timer } from 'lucide-react';
import CreateJobModal from '@/components/jobs/CreateJobModal';
import { Badge, type BadgeProps } from '@/components/ui/badge';
import { Button } from '@/components/ui/button';
import { useStore } from '@/hooks/useStore';
import {
deleteJob,
duplicateJob,
downloadJobVideo,
startJob,
stopJob,
@@ -62,6 +64,8 @@ export default function JobCard({ job, onJobUpdated }: JobCardProps) {
const isSelected = activeJobId === job.id;
const [isLoading, setIsLoading] = React.useState(false);
const [isEditing, setIsEditing] = React.useState(false);
const [isViewing, setIsViewing] = React.useState(false);
const [currentTime, setCurrentTime] = React.useState(() => Date.now());
const elapsedTime = computeElapsed(job, currentTime);
@@ -103,6 +107,21 @@ export default function JobCard({ job, onJobUpdated }: JobCardProps) {
}
}
async function handleDuplicate(e: React.MouseEvent) {
e.preventDefault();
e.stopPropagation();
if (isLoading) return;
setIsLoading(true);
try {
await duplicateJob(job.id);
onJobUpdated?.();
} catch (err) {
alert(err instanceof Error ? err.message : 'Failed to duplicate job');
} finally {
setIsLoading(false);
}
}
async function handleDelete(e: React.MouseEvent) {
e.preventDefault();
e.stopPropagation();
@@ -156,16 +175,23 @@ export default function JobCard({ job, onJobUpdated }: JobCardProps) {
>
<span className="flex flex-wrap items-center justify-between gap-2">
<span className="text-[0.95rem] font-semibold text-foreground">
{job.model_id}
{job.name?.trim() || job.model_id}
</span>
<Badge variant={BADGE_VARIANTS[job.status] ?? 'secondary'}>
{job.status}
</Badge>
</span>
<span className="max-w-full overflow-hidden text-ellipsis whitespace-nowrap text-sm text-muted-foreground">
{job.prompt}
{job.name?.trim() ? `${job.model_id} · ${job.prompt}` : job.prompt}
</span>
<span className="flex flex-wrap items-center gap-4 text-xs text-muted-foreground">
{/* Short job id; logs and output dirs are keyed on the full UUID. */}
<span
className="font-mono text-muted-foreground/80"
title={job.id}
>
{job.id.slice(0, 8)}
</span>
{job.job_type === 'inference' ? (
<>
<span>{job.num_frames} frames</span>
@@ -226,6 +252,51 @@ export default function JobCard({ job, onJobUpdated }: JobCardProps) {
Download Video
</Button>
)}
{!(
job.status === 'pending' ||
job.status === 'failed' ||
job.status === 'stopped'
) && (
<Button
size="sm"
variant="outline"
onClick={(e) => {
e.preventDefault();
e.stopPropagation();
setIsViewing(true);
}}
disabled={isLoading}
title="View this job's configuration"
>
View
</Button>
)}
{(job.status === 'pending' ||
job.status === 'failed' ||
job.status === 'stopped') && (
<Button
size="sm"
variant="outline"
onClick={(e) => {
e.preventDefault();
e.stopPropagation();
setIsEditing(true);
}}
disabled={isLoading}
title="Edit this job's configuration"
>
Edit
</Button>
)}
<Button
size="sm"
variant="outline"
onClick={handleDuplicate}
disabled={isLoading}
title="Create a new pending job with this configuration"
>
Duplicate
</Button>
<Button
size="sm"
variant="destructive"
@@ -235,6 +306,30 @@ export default function JobCard({ job, onJobUpdated }: JobCardProps) {
Delete
</Button>
</div>
{isViewing && (
<CreateJobModal
isOpen
readOnly
editingJob={job}
jobType={(job.job_type ?? 'inference') as never}
workloadType={job.workload_type ?? 't2v'}
onClose={() => setIsViewing(false)}
onSuccess={() => setIsViewing(false)}
/>
)}
{isEditing && (
<CreateJobModal
isOpen
editingJob={job}
jobType={(job.job_type ?? 'inference') as never}
workloadType={job.workload_type ?? 't2v'}
onClose={() => setIsEditing(false)}
onSuccess={() => {
setIsEditing(false);
onJobUpdated?.();
}}
/>
)}
</article>
);
}
@@ -153,9 +153,16 @@ export default function JobDetailsSidebar({
}}
>
<div className="flex items-center justify-between border-b border-border px-5 py-4">
<h2 className="m-0 text-base font-semibold text-foreground">
Job Details
</h2>
<div className="min-w-0">
<h2 className="m-0 text-base font-semibold text-foreground">
Job Details
</h2>
{/* Full job id: keys ~/h3_studio_logs/<id>.log, the output directory
and every API route, so make it selectable for copy/paste. */}
<code className="mt-0.5 block select-all truncate font-mono text-xs text-muted-foreground">
{job.id}
</code>
</div>
<div className="flex items-center gap-2">
<Button
type="button"
+61
View File
@@ -56,6 +56,8 @@ export function getJobVideoUrl(jobId: string): string {
export interface CreateJobRequest {
model_id: string;
/** Optional label; the card and output filename fall back to the prompt. */
name?: string;
prompt: string;
workload_type?: string;
job_type?: JobType;
@@ -152,6 +154,65 @@ export async function updateSettings(
return response.json();
}
export type MediaType = "image" | "video" | "audio";
/**
* Upload an image, video or audio file for a MiniMax-H3 Ref2VA reference.
* The server derives media_type from the extension and returns it, so callers
* do not have to duplicate that mapping.
*/
/** Create a new pending job with the same configuration as an existing one. */
export async function duplicateJob(jobId: string): Promise<{ id: string }> {
const baseApiUrl = getApiBaseUrl();
const response = await fetch(`${baseApiUrl}/jobs/${jobId}/duplicate`, {
method: "POST",
});
if (!response.ok) {
const err = await response
.json()
.catch(() => ({ detail: "Duplicate failed" }));
throw new Error(err.detail || "Duplicate failed");
}
return response.json();
}
/** Edit a pending job's configuration. Started jobs are rejected by the API. */
export async function updateJob(
jobId: string,
updates: Record<string, unknown>,
): Promise<unknown> {
const baseApiUrl = getApiBaseUrl();
const response = await fetch(`${baseApiUrl}/jobs/${jobId}`, {
method: "PATCH",
headers: { "Content-Type": "application/json" },
body: JSON.stringify(updates),
});
if (!response.ok) {
const err = await response.json().catch(() => ({ detail: "Update failed" }));
throw new Error(err.detail || "Update failed");
}
return response.json();
}
export async function uploadMedia(
file: File,
): Promise<{ path: string; media_type: MediaType }> {
const baseApiUrl = getApiBaseUrl();
const formData = new FormData();
formData.append("file", file);
const response = await fetch(`${baseApiUrl}/upload-media`, {
method: "POST",
body: formData,
});
if (!response.ok) {
const err = await response
.json()
.catch(() => ({ detail: "Upload failed" }));
throw new Error(err.detail || "Upload failed");
}
return response.json();
}
export async function uploadImage(file: File): Promise<{ path: string }> {
const baseApiUrl = getApiBaseUrl();
const formData = new FormData();
@@ -16,6 +16,7 @@ export interface DefaultOptions {
seed: number;
numGpus: number;
ditCpuOffload: boolean;
ditLayerwiseOffload: boolean;
textEncoderCpuOffload: boolean;
vaeCpuOffload: boolean;
imageEncoderCpuOffload: boolean;
@@ -44,6 +45,7 @@ export const DEFAULT_OPTIONS: DefaultOptions = {
seed: 1024,
numGpus: 1,
ditCpuOffload: false,
ditLayerwiseOffload: false,
textEncoderCpuOffload: false,
vaeCpuOffload: false,
imageEncoderCpuOffload: false,
@@ -0,0 +1,65 @@
import { describe, expect, it } from "vitest";
import {
EMPTY_H3_PROMPT_FIELDS,
H3_PROMPT_SECTIONS,
isEmptyPromptFields,
parseH3Prompt,
serializeH3Prompt,
type H3PromptFields,
} from "@/lib/h3Prompt";
const filled: H3PromptFields = {
subject_definitions: "<Subject 1> is the dog in <Picture 1>.",
summary: "[reference generation] The target video shows <Subject 1>.",
retention_analysis: "<Subject 1> (appears in [Shot 1]): fully_preserved - fur retained.",
detailed_description: "[Shot 1] A medium shot establishes <Subject 1>.",
overall_soundscape: "Room tone throughout.",
non_diegetic_music: "N/A",
};
describe("serializeH3Prompt", () => {
it("emits sections in order, flush left, blank-line separated", () => {
const out = serializeH3Prompt(filled);
expect(out).toBe(
[
"subject_definitions:\n<Subject 1> is the dog in <Picture 1>.",
"summary:\n[reference generation] The target video shows <Subject 1>.",
"retention_analysis:\n<Subject 1> (appears in [Shot 1]): fully_preserved - fur retained.",
"detailed_description:\n[Shot 1] A medium shot establishes <Subject 1>.",
"overall_soundscape:\nRoom tone throughout.",
"non_diegetic_music:\nN/A",
].join("\n\n"),
);
// content is never indented
expect(out).not.toMatch(/\n {2}\S/);
});
it("fills blank sections with N/A rather than dropping them", () => {
const out = serializeH3Prompt({ ...EMPTY_H3_PROMPT_FIELDS, summary: "x" });
for (const s of H3_PROMPT_SECTIONS) expect(out).toContain(`${s}:`);
expect(out).toContain("non_diegetic_music:\nN/A");
});
});
describe("parseH3Prompt", () => {
it("round-trips a serialized prompt", () => {
expect(parseH3Prompt(serializeH3Prompt(filled))).toEqual(filled);
});
it("keeps multi-line section bodies", () => {
const p = parseH3Prompt("summary:\nline one\nline two\n\ndetailed_description:\nd");
expect(p?.summary).toBe("line one\nline two");
expect(p?.detailed_description).toBe("d");
});
it("returns null for a plain prompt", () => {
expect(parseH3Prompt("a toy car drives into a plush dog")).toBeNull();
});
});
describe("isEmptyPromptFields", () => {
it("detects blank vs filled", () => {
expect(isEmptyPromptFields(EMPTY_H3_PROMPT_FIELDS)).toBe(true);
expect(isEmptyPromptFields(filled)).toBe(false);
});
});
+94
View File
@@ -0,0 +1,94 @@
/**
* MiniMax-H3 full-reference prompt sections. Serialization follows the worked
* example in the model's VIDEO_PROMPT_WRITING_GUIDE_ref_en.md: `section_name:`
* on its own line, content flush left, one blank line between sections.
*/
export const H3_PROMPT_SECTIONS = [
'subject_definitions',
'summary',
'retention_analysis',
'detailed_description',
'overall_soundscape',
'non_diegetic_music',
] as const;
export type H3PromptSection = (typeof H3_PROMPT_SECTIONS)[number];
export type H3PromptFields = Record<H3PromptSection, string>;
export const EMPTY_H3_PROMPT_FIELDS: H3PromptFields = {
subject_definitions: '',
summary: '',
retention_analysis: '',
detailed_description: '',
overall_soundscape: '',
non_diegetic_music: '',
};
export const H3_SECTION_LABELS: Record<H3PromptSection, string> = {
subject_definitions: 'Subject definitions',
summary: 'Summary',
retention_analysis: 'Retention analysis',
detailed_description: 'Detailed description',
overall_soundscape: 'Overall soundscape',
non_diegetic_music: 'Non-diegetic music',
};
/** Per-section guidance, condensed from the guide's rules for each section. */
export const H3_SECTION_HINTS: Record<H3PromptSection, string> = {
subject_definitions:
'One line per <Subject N>. A subject may draw on several references, e.g. "<Subject 2> is the dog in <Picture 2>, <Picture 3>, and <Picture 4>."',
summary:
'One paragraph, starting with a [task type] prefix: reference generation, keyframe completion, video editing, video continuation, audio reuse, audio reference. Combine with " + ".',
retention_analysis:
'Per subject: "<Subject 1> (appears in [Shot 1], [Shot 2]): fully_preserved - what is retained."',
detailed_description:
'Playback order, using [Shot N] markers, timestamps, (S1) speaker tags and <d>[English] dialogue</d>.',
overall_soundscape: 'Ambience and physical sounds. N/A if none.',
non_diegetic_music:
'Background music audible only to the audience. N/A if none.',
};
/** True when every section is blank. */
export function isEmptyPromptFields(fields: H3PromptFields): boolean {
return H3_PROMPT_SECTIONS.every((s) => !fields[s].trim());
}
/**
* Join the sections into the prompt string the model is given. Blank sections
* become "N/A" rather than being dropped, matching the guide's example.
*/
export function serializeH3Prompt(fields: H3PromptFields): string {
return H3_PROMPT_SECTIONS.map((section) => {
const body = fields[section].trim() || 'N/A';
return `${section}:\n${body}`;
}).join('\n\n');
}
/**
* Split a serialized prompt back into sections, so switching between the
* guided fields and the raw editor does not lose work. Returns null when the
* text is not in section format (a plain prompt, say).
*/
export function parseH3Prompt(text: string): H3PromptFields | null {
const fields = { ...EMPTY_H3_PROMPT_FIELDS };
const headings = new Set<string>(H3_PROMPT_SECTIONS);
let current: H3PromptSection | null = null;
let found = false;
for (const line of text.split('\n')) {
const heading = line.trim().replace(/:$/, '');
if (line.trim().endsWith(':') && headings.has(heading)) {
current = heading as H3PromptSection;
found = true;
continue;
}
if (current) fields[current] += (fields[current] ? '\n' : '') + line;
}
if (!found) return null;
for (const section of H3_PROMPT_SECTIONS) {
fields[section] = fields[section].trim();
}
return fields;
}
@@ -0,0 +1,70 @@
import { describe, expect, it } from "vitest";
import {
labelReferences,
referencePromptSeed,
validateReferences,
type H3Reference,
} from "@/lib/h3References";
const ref = (media_type: H3Reference["media_type"], n: number): H3Reference => ({
id: `${media_type}-${n}`,
source: `/tmp/${media_type}${n}`,
media_type,
fileName: `${media_type}${n}`,
});
describe("labelReferences", () => {
it("numbers each media type independently, in list order", () => {
expect(
labelReferences([ref("image", 1), ref("video", 1), ref("image", 2)]),
).toEqual(["<Picture 1>", "<Video 1>", "<Picture 2>"]);
});
});
describe("validateReferences", () => {
it("accepts an empty list and a normal mix", () => {
expect(validateReferences([])).toBeNull();
expect(validateReferences([ref("image", 1), ref("audio", 1)])).toBeNull();
});
it("rejects audio-only lists", () => {
expect(validateReferences([ref("audio", 1)])).toMatch(/paired/);
});
it("enforces the per-type caps", () => {
const videos = [1, 2, 3, 4].map((n) => ref("video", n));
expect(validateReferences(videos)).toMatch(/At most 3 video/);
});
it("enforces the overall cap", () => {
const many = Array.from({ length: 13 }, (_, i) => ref("image", i));
expect(validateReferences(many)).toMatch(/At most 12/);
});
});
describe("referencePromptSeed", () => {
it("cites the real labels and never indents content", () => {
const seed = referencePromptSeed([ref("image", 1), ref("video", 1)]);
expect(seed.subject_definitions).toContain("<Picture 1>");
expect(seed.subject_definitions).toContain("<Video 1>");
for (const value of Object.values(seed)) {
expect(value).not.toMatch(/^ {2}\S/m);
}
});
it("uses the guide's retention_analysis form", () => {
const seed = referencePromptSeed([ref("image", 1)]);
expect(seed.retention_analysis).toMatch(/fully_preserved - /);
});
it("starts the summary with a bracketed task type", () => {
const seed = referencePromptSeed([ref("image", 1)]);
expect(seed.summary).toMatch(/^\[[a-z +]+\]/);
});
it("describes audio references separately", () => {
const seed = referencePromptSeed([ref("image", 1), ref("audio", 1)]);
expect(seed.subject_definitions).toContain("<Audio 1>");
expect(seed.retention_analysis).toContain("<Audio 1>: reference -");
});
});
@@ -0,0 +1,90 @@
/**
* MiniMax-H3 Ref2VA reference helpers. Labels mirror the per-type counters in
* `build_ref2va_presentation`, so what the user sees is what the model is shown.
*/
import type { MediaType } from "@/lib/api";
export interface H3Reference {
/** stable key for React lists */
id: string;
source: string;
media_type: MediaType;
fileName: string;
}
/** Per-media-type caps enforced by validate_references (reference.py). */
export const H3_REFERENCE_LIMITS: Record<MediaType, number> = {
image: 9,
video: 3,
audio: 3,
};
export const H3_MAX_REFERENCES = 12;
const LABEL_FOR: Record<MediaType, string> = {
image: "Picture",
video: "Video",
audio: "Audio",
};
/** Label each reference the way the pipeline will, e.g. "<Picture 2>". */
export function labelReferences(refs: H3Reference[]): string[] {
const counts: Record<MediaType, number> = { image: 0, video: 0, audio: 0 };
return refs.map((ref) => {
counts[ref.media_type] += 1;
return `<${LABEL_FOR[ref.media_type]} ${counts[ref.media_type]}>`;
});
}
/** Human-readable reason the list is invalid, or null when it is acceptable. */
export function validateReferences(refs: H3Reference[]): string | null {
if (refs.length === 0) return null;
if (refs.length > H3_MAX_REFERENCES) {
return `At most ${H3_MAX_REFERENCES} references (have ${refs.length}).`;
}
const counts: Record<MediaType, number> = { image: 0, video: 0, audio: 0 };
for (const ref of refs) counts[ref.media_type] += 1;
for (const type of Object.keys(counts) as MediaType[]) {
if (counts[type] > H3_REFERENCE_LIMITS[type]) {
return `At most ${H3_REFERENCE_LIMITS[type]} ${type} references (have ${counts[type]}).`;
}
}
if (counts.audio === refs.length) {
return "Audio references must be paired with at least one image or video.";
}
return null;
}
/** Seed the guided prompt fields from the current reference list. */
export function referencePromptSeed(
refs: H3Reference[],
): Record<string, string> {
const labels = labelReferences(refs);
const visual = labels.filter((l) => !l.startsWith("<Audio"));
const audio = labels.filter((l) => l.startsWith("<Audio"));
const subjects = visual.map(
(label, i) =>
`<Subject ${i + 1}> is the subject from ${label}; describe appearance and distinguishing features.`,
);
for (const label of audio) {
subjects.push(`${label} is the audio reference; describe what it provides.`);
}
const retention = visual.map(
(_, i) =>
`<Subject ${i + 1}> (appears in [Shot 1]): fully_preserved - what is retained.`,
);
for (const label of audio) {
retention.push(`${label}: reference - how it guides the audio.`);
}
return {
subject_definitions: subjects.join("\n"),
summary: "[reference generation] Describe the target video and each reference's role.",
retention_analysis: retention.join("\n"),
detailed_description:
"[Shot 1] Describe composition, subjects, environment, lighting, action and camera movement, saying where each reference takes effect.",
overall_soundscape: "",
non_diegetic_music: "",
};
}
@@ -0,0 +1,74 @@
import { describe, expect, it } from "vitest";
import { jobToFormFields, referenceFileName, type JobLike } from "@/lib/jobToFields";
const job: JobLike = {
id: "abc",
model_id: "MiniMaxAI/MiniMax-H3",
prompt: "subject_definitions:\n<Subject 1> is a dog.\n\nsummary:\n[reference generation] x",
workload_type: "i2v",
references: [
{ source: "/uploads/9f/wukong_source.mp4", media_type: "video" },
{ source: "/uploads/2a/MonkeyKing_0.jpg", media_type: "image" },
],
num_frames: 141,
height: 768,
width: 1344,
guidance_scale: 1.0,
guidance_rescale: 0.1,
seed: 0,
num_gpus: 4,
use_fsdp_inference: true,
};
describe("jobToFormFields", () => {
it("carries the settings that differ from form defaults", () => {
const f = jobToFormFields(job);
expect(f.numFrames).toBe(141);
expect(f.height).toBe(768);
expect(f.width).toBe(1344);
expect(f.guidanceScale).toBe(1.0);
expect(f.guidanceRescale).toBe(0.1);
expect(f.seed).toBe(0); // 0 must survive, not fall back to 1024
expect(f.numGpus).toBe(4);
expect(f.useFsdpInference).toBe(true);
});
it("does not let falsy-but-valid values fall through to defaults", () => {
const f = jobToFormFields({ ...job, seed: 0, vsa_sparsity: 0, guidance_rescale: 0 });
expect(f.seed).toBe(0);
expect(f.vsaSparsity).toBe(0);
expect(f.guidanceRescale).toBe(0);
});
it("rebuilds the reference list with readable names", () => {
const f = jobToFormFields(job);
expect(f.references).toHaveLength(2);
expect(f.references[0].fileName).toBe("wukong_source.mp4");
expect(f.references[1].media_type).toBe("image");
expect(new Set(f.references.map((r) => r.id)).size).toBe(2);
});
it("splits a six-section prompt back into fields", () => {
const f = jobToFormFields(job);
expect(f.promptFields?.subject_definitions).toBe("<Subject 1> is a dog.");
expect(f.promptFields?.summary).toBe("[reference generation] x");
});
it("returns null promptFields for a plain prompt", () => {
const f = jobToFormFields({ ...job, prompt: "a toy car" });
expect(f.promptFields).toBeNull();
expect(f.prompt).toBe("a toy car");
});
it("handles a job with no references", () => {
const f = jobToFormFields({ ...job, references: null });
expect(f.references).toEqual([]);
});
});
describe("referenceFileName", () => {
it("takes the basename", () => {
expect(referenceFileName("/a/b/c.mp4")).toBe("c.mp4");
expect(referenceFileName("c.mp4")).toBe("c.mp4");
});
});
@@ -0,0 +1,119 @@
/**
* Map a persisted job back onto the create-job form's fields. Every value the
* form seeds from defaults must be covered here, or edit mode silently shows
* defaults for whatever is missing.
*/
import type { H3Reference } from "@/lib/h3References";
import { parseH3Prompt, type H3PromptFields } from "@/lib/h3Prompt";
export interface JobLike {
id: string;
model_id: string;
name?: string;
prompt: string;
workload_type?: string;
job_type?: string;
image_path?: string;
last_image_path?: string;
references?: { source: string; media_type: string }[] | null;
negative_prompt?: string;
num_inference_steps?: number;
num_frames?: number;
height?: number;
width?: number;
guidance_scale?: number;
guidance_rescale?: number;
fps?: number;
seed?: number;
num_gpus?: number;
dit_cpu_offload?: boolean;
dit_layerwise_offload?: boolean;
text_encoder_cpu_offload?: boolean;
vae_cpu_offload?: boolean;
image_encoder_cpu_offload?: boolean;
use_fsdp_inference?: boolean;
enable_torch_compile?: boolean;
vsa_sparsity?: number;
tp_size?: number;
sp_size?: number;
}
export interface JobFormFields {
modelId: string;
name: string;
workloadType: string;
jobType: string;
prompt: string;
negativePrompt: string;
imagePath: string;
lastImagePath: string;
references: H3Reference[];
promptFields: H3PromptFields | null;
numInferenceSteps: number;
numFrames: number;
height: number;
width: number;
guidanceScale: number;
guidanceRescale: number;
fps: number;
seed: number;
numGpus: number;
ditCpuOffload: boolean;
ditLayerwiseOffload: boolean;
textEncoderCpuOffload: boolean;
vaeCpuOffload: boolean;
imageEncoderCpuOffload: boolean;
useFsdpInference: boolean;
enableTorchCompile: boolean;
vsaSparsity: number;
tpSize: number;
spSize: number;
}
/** Uploads keep their original basename, so this is the display name. */
export function referenceFileName(source: string): string {
return source.split("/").filter(Boolean).pop() ?? source;
}
export function jobToFormFields(job: JobLike): JobFormFields {
const refs: H3Reference[] = (job.references ?? []).map((r, i) => ({
id: `${job.id}-${i}`,
source: r.source,
media_type: r.media_type as H3Reference["media_type"],
fileName: referenceFileName(r.source),
}));
return {
modelId: job.model_id,
name: job.name ?? "",
workloadType: job.workload_type ?? "t2v",
jobType: job.job_type ?? "inference",
prompt: job.prompt ?? "",
negativePrompt: job.negative_prompt ?? "",
imagePath: job.image_path ?? "",
lastImagePath: job.last_image_path ?? "",
references: refs,
// null when the prompt is not in six-section form; the caller then keeps
// the raw editor rather than silently dropping content into fields.
promptFields: parseH3Prompt(job.prompt ?? ""),
numInferenceSteps: job.num_inference_steps ?? 50,
numFrames: job.num_frames ?? 81,
height: job.height ?? 480,
width: job.width ?? 832,
guidanceScale: job.guidance_scale ?? 5.0,
guidanceRescale: job.guidance_rescale ?? 0.0,
fps: job.fps ?? 24,
seed: job.seed ?? 1024,
numGpus: job.num_gpus ?? 1,
ditCpuOffload: job.dit_cpu_offload ?? false,
ditLayerwiseOffload: job.dit_layerwise_offload ?? false,
textEncoderCpuOffload: job.text_encoder_cpu_offload ?? false,
vaeCpuOffload: job.vae_cpu_offload ?? false,
imageEncoderCpuOffload: job.image_encoder_cpu_offload ?? false,
useFsdpInference: job.use_fsdp_inference ?? false,
enableTorchCompile: job.enable_torch_compile ?? false,
vsaSparsity: job.vsa_sparsity ?? 0,
tpSize: job.tp_size ?? -1,
spSize: job.sp_size ?? -1,
};
}
+1
View File
@@ -5,6 +5,7 @@ export type JobType = "inference" | "finetuning" | "distillation";
export interface Job {
id: string;
model_id: string;
name?: string;
prompt: string;
job_type?: JobType;
workload_type?: string;
@@ -0,0 +1,99 @@
# SPDX-License-Identifier: Apache-2.0
"""Duplicating a job's config, and editing one that has not started."""
import uuid
import pytest
from fastvideo_studio.database import Database
from fastvideo_studio.job_runner import JobRunner, JobStatus
@pytest.fixture
def runner(tmp_path):
return JobRunner(
output_dir=str(tmp_path / "out"),
log_dir=str(tmp_path / "logs"),
database=Database(tmp_path / "t.db"),
)
def _make(runner, **over):
kwargs = dict(
job_id=str(uuid.uuid4()),
model_id="MiniMaxAI/MiniMax-H3",
prompt="p",
workload_type="i2v",
num_frames=141,
guidance_scale=1.0,
num_gpus=4,
references=[{"source": "/x/clip.mp4", "media_type": "video"}],
)
kwargs.update(over)
return runner.create_job(**kwargs)
def test_duplicate_copies_config_but_not_runtime_state(runner):
src = _make(runner)
src.status = JobStatus.COMPLETED
src.output_path = "/x/out.mp4"
dup = runner.duplicate_job(src.id, str(uuid.uuid4()))
assert dup.id != src.id
assert dup.status is JobStatus.PENDING
assert dup.output_path is None
for field in ("model_id", "prompt", "workload_type", "num_frames",
"guidance_scale", "num_gpus", "references"):
assert getattr(dup, field) == getattr(src, field)
def test_duplicate_deep_copies_references(runner):
src = _make(runner)
dup = runner.duplicate_job(src.id, str(uuid.uuid4()))
dup.references[0]["source"] = "/changed"
assert src.references[0]["source"] == "/x/clip.mp4"
def test_duplicate_unknown_job(runner):
with pytest.raises(ValueError, match="not found"):
runner.duplicate_job("nope", str(uuid.uuid4()))
def test_edit_pending_job(runner):
job = _make(runner)
updated = runner.update_job_config(job.id, {"num_frames": 192, "seed": 7})
assert updated.num_frames == 192
assert updated.seed == 7
@pytest.mark.parametrize("status", [JobStatus.FAILED, JobStatus.STOPPED])
def test_edit_allows_restartable_jobs(runner, status):
"""Editable exactly when startable: neither has produced an output."""
job = _make(runner)
job.status = status
assert runner.update_job_config(job.id, {"seed": 7}).seed == 7
@pytest.mark.parametrize("status", [JobStatus.COMPLETED, JobStatus.RUNNING])
def test_edit_rejects_jobs_with_or_producing_a_result(runner, status):
job = _make(runner)
job.status = status
with pytest.raises(ValueError, match="can be edited"):
runner.update_job_config(job.id, {"seed": 7})
def test_edit_rejects_unknown_field(runner):
job = _make(runner)
with pytest.raises(ValueError, match="Not editable"):
runner.update_job_config(job.id, {"status": "completed"})
def test_name_is_carried_by_duplicate(runner):
src = _make(runner, name="wukong swap v2")
dup = runner.duplicate_job(src.id, str(uuid.uuid4()))
assert dup.name == "wukong swap v2"
def test_name_is_editable(runner):
job = _make(runner, name="a")
assert runner.update_job_config(job.id, {"name": "b"}).name == "b"
@@ -0,0 +1,36 @@
# SPDX-License-Identifier: Apache-2.0
"""Uploaded files keep a readable basename under a unique directory."""
from __future__ import annotations
import pytest
from fastvideo_studio.server import _safe_upload_name
@pytest.mark.parametrize(
("filename", "ext", "expected"),
[
("wukong_source.mp4", ".mp4", "wukong_source.mp4"),
("MonkeyKing_0.jpg", ".jpg", "MonkeyKing_0.jpg"),
("my clip (final).mp4", ".mp4", "my_clip_final.mp4"),
("../../etc/passwd.png", ".png", "passwd.png"),
("/abs/path/frame.png", ".png", "frame.png"),
("émoji✨.png", ".png", "moji.png"),
("", ".png", "upload.png"),
(None, ".png", "upload.png"),
("...", ".png", "upload.png"),
],
)
def test_safe_upload_name(filename, ext, expected):
assert _safe_upload_name(filename, ext) == expected
def test_long_names_are_capped():
out = _safe_upload_name("x" * 300 + ".png", ".png")
assert out == "x" * 80 + ".png"
def test_no_path_separators_survive():
for bad in ("a/b.png", "a\\b.png", "../x.png"):
assert "/" not in _safe_upload_name(bad, ".png")
assert "\\" not in _safe_upload_name(bad, ".png")
+26 -12
View File
@@ -89,11 +89,29 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
zsh \
vim \
curl \
ffmpeg \
libgl1 \
libglib2.0-0 \
libx11-dev \
gcc-11 \
g++-11 \
clang-11 \
cmake \
pkg-config \
build-essential \
libssl-dev \
&& rm -rf /var/lib/apt/lists/*
# Rust toolchain: some dependencies only ship sdists on aarch64 and need cargo
# to build. The dormant legacy Modal image layers the identical apt set +
# rustup on top of this image (fastvideo/tests/modal/pr_test.py); baking both
# here keeps its manual rollback path reproducible without changing the Slurm
# runner's package surface.
RUN set -o pipefail && \
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y --default-toolchain stable --profile minimal && \
/root/.cargo/bin/cargo --version && /root/.cargo/bin/rustc --version
ENV PATH=/root/.cargo/bin:${PATH}
# Set up C++20 compilers for ThunderKittens
RUN update-alternatives --install /usr/bin/gcc gcc /usr/bin/gcc-11 100 --slave /usr/bin/g++ g++ /usr/bin/g++-11
@@ -138,6 +156,7 @@ RUN --mount=type=cache,target=/opt/uv/cache \
source /opt/venv/bin/activate && \
uv pip install --upgrade pip && \
uv pip install --excludes docker/uv-excludes ".[dev]" && \
python -c "import cv2; print('OpenCV', cv2.__version__)" && \
PYTAG=cp$(echo "${PYTHON_VERSION}" | tr -d .) && \
case "${TARGETARCH:-amd64}" in \
amd64) \
@@ -169,26 +188,21 @@ RUN --mount=type=cache,target=/opt/uv/cache \
# flash_attn/__init__.py), so FA2/varlen/bert_padding stay from the install above;
# rmtree clears the wheel's stale cute files first to avoid an install conflict.
# Then verify both survive so a broken overlay fails the build instead of shipping
# an FA2-less image. x86 only: the FA4 stack (quack-kernels etc.) is unvalidated on
# arm64 / GB10 (sm_121), so there we skip the overlay; FA4 is opt-in
# (FASTVIDEO_FA4=1) and errors if set without the overlay, so leave it unset on
# arm64 and the image runs FA3/FA2 as usual.
# an FA2-less image. The pinned stack is validated on ARM64 GB200 (sm_100) as well
# as x86; FA4 remains opt-in through FASTVIDEO_FA4=1 so lanes with FA2 baselines
# keep their existing numerics.
RUN --mount=type=cache,target=/opt/uv/cache \
source $HOME/.local/bin/env && \
source /opt/venv/bin/activate && \
if [ "${TARGETARCH}" = "arm64" ]; then \
echo "Skipping FA4 cute overlay on arm64 (FA4 stack unvalidated there; do not set FASTVIDEO_FA4)"; \
else \
python -c "import glob, shutil; [shutil.rmtree(d, ignore_errors=True) for d in glob.glob('/opt/venv/lib/python*/site-packages/flash_attn/cute')]" && \
uv pip install "flash-attn-4 @ git+https://github.com/Dao-AILab/flash-attention.git@${FA4_CUTE_REF}#subdirectory=flash_attn/cute" && \
python -c "import flash_attn; assert hasattr(flash_attn, 'flash_attn_func'), 'FA2 was clobbered by the cute overlay'; import flash_attn.cute; print('FA2 + FA4 cute OK')"; \
fi
python -c "import glob, shutil; [shutil.rmtree(d, ignore_errors=True) for d in glob.glob('/opt/venv/lib/python*/site-packages/flash_attn/cute')]" && \
uv pip install "flash-attn-4 @ git+https://github.com/Dao-AILab/flash-attention.git@${FA4_CUTE_REF}#subdirectory=flash_attn/cute" && \
python -c "import flash_attn; assert hasattr(flash_attn, 'flash_attn_func'), 'FA2 was clobbered by the cute overlay'; import flash_attn.cute; print('FA2 + FA4 cute OK')"
COPY . .
# Build immutable FastVideo kernel wheels for the published image. The requested
# architecture remains installed for normal image users; amd64 images also carry
# an SM89 artifact so the predominant L40S Modal lanes can reuse it exactly.
# an SM89 artifact for L40S users and the dormant legacy rollback path.
ARG FASTVIDEO_KERNEL_PREBUILT_DIR=/opt/fastvideo-kernel-prebuilt
RUN --mount=type=cache,target=/opt/uv/cache \
source $HOME/.local/bin/env && \
+726 -13
View File
@@ -1,52 +1,765 @@
{
"version": 8,
"recipes": [
{
"id": "fastwan21-t2v",
"family": "wan",
"stage": "inference",
"task": "Text to video",
"label": "FastWan2.1 1.3B (distilled + VSA)",
"summary": "Generate a video in three denoising steps with the distilled FastWan2.1 1.3B checkpoint and video sparse attention.",
"model": "FastVideo/FastWan2.1-T2V-1.3B-Diffusers",
"source": "scripts/inference/inference_wan_VSA_DMD_1_3B.yaml",
"command": "FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN fastvideo generate --config scripts/inference/inference_wan_VSA_DMD_1_3B.yaml"
"command": "FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN fastvideo generate --config scripts/inference/inference_wan_VSA_DMD_1_3B.yaml",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 videos under outputs_video_dmd_1.3B/"
},
{
"id": "wan22-t2v",
"family": "wan",
"stage": "inference",
"task": "Text to video",
"label": "Wan2.2 A14B",
"summary": "The maintained high-capacity Wan2.2 text-to-video example with CPU offload settings encoded in its checked-in Python source.",
"model": "Wan-AI/Wan2.2-T2V-A14B-Diffusers",
"source": "examples/inference/basic/basic_wan2_2.py",
"command": "python examples/inference/basic/basic_wan2_2.py"
"command": "python examples/inference/basic/basic_wan2_2.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 2, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 videos under video_samples_wan2_2_14B_t2v/"
},
{
"id": "wan21-i2v",
"family": "wan",
"stage": "inference",
"task": "Image to video",
"label": "Wan2.1 14B 480P",
"summary": "Animate an input image at 480P using the maintained Wan2.1 YAML configuration and its recorded offload settings.",
"model": "Wan-AI/Wan2.1-I2V-14B-480P-Diffusers",
"source": "scripts/inference/inference_wan_i2v.yaml",
"command": "fastvideo generate --config scripts/inference/inference_wan_i2v.yaml"
},
{
"id": "turbowan22-i2v",
"task": "Image to video",
"label": "TurboWan2.2 A14B",
"model": "loayrashid/TurboWan2.2-I2V-A14B-Diffusers",
"source": "examples/inference/basic/basic_turbodiffusion_i2v.py",
"command": "python examples/inference/basic/basic_turbodiffusion_i2v.py"
"command": "fastvideo generate --config scripts/inference/inference_wan_i2v.yaml",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 2, "evidence": "source-configured"},
"evidence": "Source-backed"
},
{
"id": "wan22-ti2v",
"family": "wan",
"stage": "inference",
"task": "Text or image to video",
"label": "Wan2.2 TI2V 5B",
"summary": "Use one maintained 5B checkpoint for text-to-video or add an image input to switch the same recipe to image-to-video.",
"model": "Wan-AI/Wan2.2-TI2V-5B-Diffusers",
"source": "examples/inference/basic/basic_wan2_2_ti2v.py",
"command": "python examples/inference/basic/basic_wan2_2_ti2v.py"
"command": "python examples/inference/basic/basic_wan2_2_ti2v.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 videos under video_samples_wan2_2_5B_ti2v/"
},
{
"id": "fastmetal-1-3b-mlx",
"family": "wan",
"stage": "inference",
"task": "Text to video",
"label": "FastMetal 1.3B",
"summary": "Run the released FastMetal 1.3B QAD checkpoint through FastVideo's native Apple Silicon MLX path.",
"model": "FastVideo/FastMetal-1.3B-QAD",
"source": "examples/inference/basic/mlx_wan_prompt_to_video.py",
"command": "hf download FastVideo/FastMetal-1.3B-QAD --local-dir ./FastMetal-1.3B-QAD\npython examples/inference/basic/mlx_wan_prompt_to_video.py --model-root ./FastMetal-1.3B-QAD --mlx-checkpoint ./FastMetal-1.3B-QAD --height 480 --width 832 --num-frames 81 --prompt \"A bird's-eye view of a misty forest valley at dawn.\" --output-path ./outputs/fastmetal_1_3b.mp4",
"gpu_types": ["Apple Silicon"],
"hardware": {
"platform": "mlx",
"accelerator": "Apple M4 Max",
"system_memory": "36 GB unified memory",
"minimum_memory": "16 GB+ unified memory",
"peak_memory": "3.87 GiB peak MLX memory",
"evidence": "validated",
"evidence_url": "https://github.com/hao-ai-lab/FastVideo/pull/1638"
},
"evidence": "Verified",
"expected_artifact": "MP4 video at outputs/fastmetal_1_3b.mp4",
"modes": ["T2V", "temporal --fast", "spatial --fast-spatial", "two-pass --refine"],
"limitations": [
"Native MLX FastMetal T2V. Add --fast for temporal RIFE, --fast-spatial to denoise at half resolution, or --refine for a two-pass upsample. basic_mps.py is the older PyTorch MPS demo."
]
},
{
"id": "fastmetal-5b-mlx",
"family": "wan",
"stage": "inference",
"task": "Text to video",
"label": "FastMetal 5B",
"summary": "Run the released Wan2.2 5B FastMetal checkpoint as MLX T2V with MLX DiT denoising and MLX TAEHV decode. The CUDA Wan2.2 TI2V 5B recipe is the image-capable path.",
"model": "FastVideo/FastMetal-5B-QAD",
"source": "examples/inference/basic/mlx_wan22_generate.py",
"command": "hf download FastVideo/FastMetal-5B-QAD --local-dir ./FastMetal-5B-QAD\npython examples/inference/basic/mlx_wan22_generate.py --mlx-checkpoint ./FastMetal-5B-QAD --text-encoder-root ./FastMetal-5B-QAD --vae-root ./FastMetal-5B-QAD/vae --height 704 --width 1280 --num-frames 81 --prompt \"A cinematic portrait with soft neon lighting and smooth camera motion.\" --output-path ./outputs/fastmetal_5b.mp4",
"gpu_types": ["Apple Silicon"],
"hardware": {
"platform": "mlx",
"accelerator": "Apple M4 Max",
"system_memory": "36 GB unified memory",
"minimum_memory": "16 GB+ unified memory",
"peak_memory": "9.34 GiB peak MLX memory",
"evidence": "validated",
"evidence_url": "https://github.com/hao-ai-lab/FastVideo/pull/1638"
},
"evidence": "Verified",
"expected_artifact": "MP4 video at outputs/fastmetal_5b.mp4",
"modes": ["T2V", "temporal --fast", "spatial --fast-spatial", "two-pass --refine"],
"limitations": [
"The checked-in MLX example is T2V. Image-to-video is not in mlx_wan22_generate.py. Add --fast, --fast-spatial, or --refine on the same script."
]
},
{
"id": "fastmetal-14b-mlx",
"family": "wan",
"stage": "inference",
"task": "Text to video",
"label": "FastMetal 14B",
"summary": "Run the released 14B FastMetal QAD checkpoint through the same Apple Silicon MLX entrypoint as the 1.3B release.",
"model": "FastVideo/FastMetal-14B-QAD",
"source": "examples/inference/basic/mlx_wan_prompt_to_video.py",
"command": "hf download FastVideo/FastMetal-14B-QAD --local-dir ./FastMetal-14B-QAD\npython examples/inference/basic/mlx_wan_prompt_to_video.py --model-root ./FastMetal-14B-QAD --mlx-checkpoint ./FastMetal-14B-QAD --height 480 --width 832 --num-frames 81 --prompt \"A wide cinematic landscape at sunrise.\" --output-path ./outputs/fastmetal_14b.mp4",
"gpu_types": ["Apple Silicon"],
"hardware": {
"platform": "mlx",
"accelerator": "Apple M4 Max",
"system_memory": "36 GB unified memory",
"minimum_memory": "36 GB+ unified memory",
"peak_memory": "21.68 GiB peak MLX memory",
"evidence": "validated",
"evidence_url": "https://github.com/hao-ai-lab/FastVideo/pull/1638"
},
"evidence": "Verified",
"expected_artifact": "MP4 video at outputs/fastmetal_14b.mp4",
"modes": ["T2V", "temporal --fast", "spatial --fast-spatial", "two-pass --refine"],
"limitations": [
"Native MLX FastMetal T2V on 36 GB+ unified memory. Same --fast, --fast-spatial, and --refine flags as the 1.3B script."
]
},
{
"id": "turbodiffusion-wan21-1-3b-t2v",
"family": "turbodiffusion",
"stage": "inference",
"task": "Text to video",
"label": "TurboWan2.1 1.3B",
"summary": "A TurboDiffusion-accelerated Wan2.1 1.3B text-to-video run from its maintained single-GPU example.",
"model": "loayrashid/TurboWan2.1-T2V-1.3B-Diffusers",
"source": "examples/inference/basic/basic_turbodiffusion.py",
"command": "python examples/inference/basic/basic_turbodiffusion.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 videos under video_samples_turbodiffusion/",
"related": ["turbodiffusion-wan21-14b-t2v", "turbowan22-i2v"]
},
{
"id": "turbodiffusion-wan21-14b-t2v",
"family": "turbodiffusion",
"stage": "inference",
"task": "Text to video",
"label": "TurboWan2.1 14B",
"summary": "TurboDiffusion acceleration applied to the 14B Wan2.1 text-to-video checkpoint; the checked-in source is configured for two GPUs.",
"model": "loayrashid/TurboWan2.1-T2V-14B-Diffusers",
"source": "examples/inference/basic/basic_turbodiffusion_14b.py",
"command": "python examples/inference/basic/basic_turbodiffusion_14b.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 2, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 videos under video_samples_turbodiffusion_14B/",
"related": ["turbodiffusion-wan21-1-3b-t2v", "turbowan22-i2v"]
},
{
"id": "turbowan22-i2v",
"family": "turbodiffusion",
"stage": "inference",
"task": "Image to video",
"label": "TurboWan2.2 A14B",
"summary": "A one-to-four-step image-to-video path using TurboDiffusion and the SLA attention backend from its maintained example.",
"model": "loayrashid/TurboWan2.2-I2V-A14B-Diffusers",
"source": "examples/inference/basic/basic_turbodiffusion_i2v.py",
"command": "python examples/inference/basic/basic_turbodiffusion_i2v.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 2, "evidence": "source-configured"},
"evidence": "Source-backed",
"related": ["turbodiffusion-wan21-14b-t2v"]
},
{
"id": "ltx2-distilled-t2v",
"family": "ltx2",
"stage": "inference",
"task": "Text to video",
"label": "LTX-2 distilled",
"summary": "The distilled LTX-2 text-to-video checkpoint with audio, from its maintained example. The source is configured for four GPUs.",
"model": "FastVideo/LTX2-Distilled-Diffusers",
"source": "examples/inference/basic/basic_ltx2_distilled.py",
"command": "python examples/inference/basic/basic_ltx2_distilled.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 4, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 (video with audio) at outputs_video/ltx2_basic/output_ltx2_distilled_t2v.mp4",
"related": ["ltx23-base-t2v"]
},
{
"id": "ltx23-base-t2v",
"family": "ltx2",
"stage": "inference",
"task": "Text to video",
"label": "LTX-2 base (1088p)",
"summary": "Base LTX-2 text-to-video at 1088x1920 using FastVideo default sampling for LTX2 base. The example loads a community Diffusers mirror of the base checkpoint; registered aliases include Lightricks/LTX-2 and FastVideo/LTX2-Diffusers.",
"model": "Davids048/LTX2-Base-Diffusers",
"source": "examples/inference/basic/basic_ltx2.py",
"command": "python examples/inference/basic/basic_ltx2.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 (video with audio) at outputs_video/ltx2_basic/output_ltx2_base_t2v_1088_1920_1.1.mp4",
"limitations": ["The maintained example loads the Davids048/LTX2-Base-Diffusers community mirror rather than a Lightricks upstream ID."],
"related": ["ltx2-distilled-t2v"]
},
{
"id": "hy15-t2v-480p",
"family": "hunyuan",
"stage": "inference",
"task": "Text to video",
"label": "HunyuanVideo 1.5 480P",
"summary": "HunyuanVideo 1.5 text-to-video at 480P with CPU offload enabled in the checked-in source for smaller GPUs.",
"model": "hunyuanvideo-community/HunyuanVideo-1.5-Diffusers-480p_t2v",
"source": "examples/inference/basic/basic_hy15.py",
"command": "python examples/inference/basic/basic_hy15.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 videos under video_samples_hy15/",
"related": ["hy15-1080p-upscale"]
},
{
"id": "hy15-1080p-upscale",
"family": "hunyuan",
"stage": "inference",
"task": "Text to video (upscaled)",
"label": "HunyuanVideo 1.5 1080P upscale",
"summary": "Run HunyuanVideo 1.5 through the 480p to 720p to 1080p upscale chain in one maintained script.",
"model": "weizhou03/HunyuanVideo-1.5-Diffusers-1080p-2SR",
"source": "examples/inference/basic/basic_hy15_1080p.py",
"command": "python examples/inference/basic/basic_hy15_1080p.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 videos under video_samples_hy15_1080p/",
"related": ["hy15-t2v-480p"]
},
{
"id": "cosmos25-t2w",
"family": "cosmos",
"stage": "inference",
"task": "Text to world",
"label": "Cosmos Predict 2.5 2B",
"summary": "Generate a navigable world video from a text prompt with Cosmos Predict 2.5 2B on a single GPU.",
"model": "KyleShao/Cosmos-Predict2.5-2B-Diffusers",
"source": "examples/inference/basic/basic_cosmos2_5_t2w.py",
"command": "python examples/inference/basic/basic_cosmos2_5_t2w.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed"
},
{
"id": "kandinsky5-t2v-lite-sft",
"family": "kandinsky5",
"stage": "inference",
"task": "Text to video",
"label": "Kandinsky 5.0 T2V Lite SFT",
"summary": "Kandinsky 5.0 text-to-video (Lite SFT variant) from the maintained example; alternative Lite/Pro checkpoints are listed in the source.",
"model": "kandinskylab/Kandinsky-5.0-T2V-Lite-sft-5s-Diffusers",
"source": "examples/inference/basic/basic_kandinsky5_t2v.py",
"command": "python examples/inference/basic/basic_kandinsky5_t2v.py",
"gpu_types": ["NVIDIA"],
"hardware": {
"platform": "cuda",
"gpu_count": 1,
"accelerator": "NVIDIA B200",
"evidence": "validated",
"evidence_url": "https://github.com/hao-ai-lab/FastVideo/pull/1471"
},
"evidence": "Verified",
"expected_artifact": "MP4 videos under video_samples_kandinsky5_t2v/",
"related": ["kandinsky5-i2v-pro-distilled"]
},
{
"id": "kandinsky5-i2v-pro-distilled",
"family": "kandinsky5",
"stage": "inference",
"task": "Image to video",
"label": "Kandinsky 5.0 I2V Pro distilled",
"summary": "Animate an input image with Kandinsky 5.0 I2V Pro (distilled) on a single GPU.",
"model": "kandinskylab/Kandinsky-5.0-I2V-Pro-distilled-5s-Diffusers",
"source": "examples/inference/basic/basic_kandinsky5_i2v.py",
"command": "python examples/inference/basic/basic_kandinsky5_i2v.py",
"gpu_types": ["NVIDIA"],
"hardware": {
"platform": "cuda",
"gpu_count": 1,
"accelerator": "NVIDIA B200",
"peak_memory": "10,365.89 MB peak GPU memory",
"evidence": "validated",
"evidence_url": "https://github.com/hao-ai-lab/FastVideo/pull/1471"
},
"evidence": "Verified",
"expected_artifact": "MP4 videos under video_samples_kandinsky5_i2v/",
"related": ["kandinsky5-t2v-lite-sft"]
},
{
"id": "flux2-klein-t2i",
"family": "flux",
"stage": "inference",
"task": "Text to image",
"label": "FLUX.2 Klein 4B",
"summary": "Generate an image in four denoising steps with the distilled FLUX.2 Klein checkpoint.",
"model": "black-forest-labs/FLUX.2-klein-4B",
"source": "examples/inference/basic/basic_flux2_klein.py",
"command": "python examples/inference/basic/basic_flux2_klein.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "PNG image at outputs/flux2/flux2_klein.png",
"related": ["flux2-dev-t2i"]
},
{
"id": "flux2-dev-t2i",
"family": "flux",
"stage": "inference",
"task": "Text to image",
"label": "FLUX.2 dev",
"summary": "Full FLUX.2 dev text-to-image with embedded guidance and the Mistral3 text encoder, from its maintained example.",
"model": "black-forest-labs/FLUX.2-dev",
"source": "examples/inference/basic/basic_flux2.py",
"command": "python examples/inference/basic/basic_flux2.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "PNG image at outputs/flux2/flux2.png",
"related": ["flux2-klein-t2i"]
},
{
"id": "flux1-dev-t2i",
"family": "flux",
"stage": "inference",
"task": "Text to image",
"label": "FLUX.1 dev",
"summary": "FLUX.1 dev text-to-image through the Diffusers-backed pipeline. The example defaults to a local weights directory, so this recipe passes the Hugging Face ID explicitly.",
"model": "black-forest-labs/FLUX.1-dev",
"source": "examples/inference/basic/basic_flux_dev.py",
"command": "python examples/inference/basic/basic_flux_dev.py --model-path black-forest-labs/FLUX.1-dev",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "PNG images under outputs/flux_dev/samples/",
"limitations": ["FLUX.1 is loadable by ID but registers no model_family in fastvideo/registry.py; it is grouped under FLUX for documentation only."]
},
{
"id": "glm-image-t2i",
"family": "glm_image",
"stage": "inference",
"task": "Text to image",
"label": "GLM-Image",
"summary": "GLM-Image text-to-image generation from its maintained example.",
"model": "zai-org/GLM-Image",
"source": "examples/inference/basic/basic_glm_image.py",
"command": "python examples/inference/basic/basic_glm_image.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "PNG image at image_output/landscape.png",
"related": ["glm-image-edit"]
},
{
"id": "glm-image-edit",
"family": "glm_image",
"stage": "inference",
"task": "Image editing",
"label": "GLM-Image editing",
"summary": "Edit an input image with an instruction prompt using GLM-Image, from its maintained editing example.",
"model": "zai-org/GLM-Image",
"source": "examples/inference/basic/edit_glm_image.py",
"command": "python examples/inference/basic/edit_glm_image.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "PNG image at image_output/edited.png (input: assets/images/couple.jpg)",
"related": ["glm-image-t2i"]
},
{
"id": "zimage-turbo-t2i",
"family": "zimage",
"stage": "inference",
"task": "Text to image",
"label": "Z-Image Turbo",
"summary": "Z-Image Turbo text-to-image on a single GPU from its maintained example.",
"model": "Tongyi-MAI/Z-Image-Turbo",
"source": "examples/inference/basic/basic_zimage.py",
"command": "python examples/inference/basic/basic_zimage.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "PNG image at outputs/zimage/zimage_turbo.png"
},
{
"id": "sd35-medium-t2i",
"family": "sd35",
"stage": "inference",
"task": "Text to image",
"label": "Stable Diffusion 3.5 Medium",
"summary": "Stable Diffusion 3.5 Medium text-to-image over a small built-in prompt set, from its maintained example.",
"model": "stabilityai/stable-diffusion-3.5-medium",
"source": "examples/inference/basic/basic_sd35_t2i.py",
"command": "python examples/inference/basic/basic_sd35_t2i.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "PNG images under outputs/sd35/samples/"
},
{
"id": "minimax-h3-t2v",
"family": "minimax_h3",
"stage": "inference",
"task": "Text to video (with audio)",
"label": "MiniMax H3 T2VA",
"summary": "Generate synchronized video and stereo audio from a structured text prompt with the full MiniMax H3 checkpoint.",
"model": "MiniMaxAI/MiniMax-H3",
"source": "examples/inference/basic/basic_minimax_h3_t2v.py",
"command": "python examples/inference/basic/basic_minimax_h3_t2v.py --prompt \"(S1) A presenter says <d>[English] FastVideo runs MiniMax H3.</d>\"",
"gpu_types": ["NVIDIA"],
"hardware": {"platform": "cuda", "gpu_count": 4, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 with synchronized audio at outputs/minimax_h3_t2v/minimax_h3_t2v.mp4",
"modes": ["T2VA"],
"knobs": [
{"key": "num_gpus", "label": "GPUs", "hint": "Sequence-parallel degree", "flag": "--num-gpus", "options": [1, 2, 4, 8], "default": 4}
],
"limitations": ["The checked-in example defaults to four-way sequence parallelism. It does not record a GPU model or memory requirement."]
},
{
"id": "fasth3-preview-cuda",
"group": "fasth3-preview",
"group_label": "FastH3 Preview",
"group_task": "4-step text to video + audio",
"family": "minimax_h3",
"stage": "inference",
"task": "Few-step text to video (with audio)",
"label": "FastH3 Preview on CUDA",
"summary": "Run the DMD2-distilled FastH3 Preview with four DiT forwards, trained H3 sparse attention, compiled decode, and synchronized audio.",
"model": "FastVideo/FastVideo-Minimax-FastH3-Preview-v0.2",
"source": "examples/inference/basic/basic_fasth3.py",
"serving": {
"source": "examples/serving/openai_fasth3.yaml",
"install": "UV_TORCH_BACKEND=cu130 uv pip install -e \".[fasth3]\""
},
"command": "UV_TORCH_BACKEND=cu130 uv pip install -e \".[fasth3]\"\npython examples/inference/basic/basic_fasth3.py --prompt \"(S1) A presenter says <d>[English] FastVideo runs FastH3.</d>\" --profile all",
"gpu_types": ["NVIDIA"],
"hardware": {
"platform": "cuda",
"gpu_count": 4,
"accelerator": "NVIDIA GB200",
"evidence": "validated",
"evidence_url": "https://github.com/hao-ai-lab/FastVideo/pull/1731"
},
"evidence": "Verified",
"expected_artifact": "Warmup and measured MP4 files under outputs/fasth3/",
"modes": ["T2VA", "4-step FastH3"],
"knobs": [
{"key": "num_gpus", "label": "GPUs", "hint": "Sequence-parallel degree", "flag": "--num-gpus", "options": [1, 2, 4, 8], "default": 4},
{"key": "video_decode_backend", "label": "VAE decode", "hint": "Fidelity vs. speed", "flag": "--video-decode-backend", "options": [{"value": "h3-vae", "label": "Full H3 VAE"}, {"value": "taeh3", "label": "TAEH3 preview"}], "default": "h3-vae"}
],
"limitations": ["The default all profile is the measured GB200 performance route and can change floating-point operation order. Use --profile strict --no-inference-torch-compile for the eager strict route."]
},
{
"id": "fasth3-preview-mlx",
"serving": {
"source": "examples/serving/mlx_fasth3.yaml",
"install": "uv pip install -e \".[mlx]\"",
"prepare": "hf download FastVideo/FastVideo-Minimax-FastH3-Preview-v0.2 --local-dir ./FastH3-Preview-v0.2\npython scripts/checkpoint_conversion/convert_minimax_h3_mlx.py --model-root ./FastH3-Preview-v0.2/transformer --out ./FastH3-MLX --formats \"int6\""
},
"group": "fasth3-preview",
"group_label": "FastH3 Preview",
"group_task": "4-step text to video + audio",
"family": "minimax_h3",
"stage": "inference",
"task": "Few-step text to video (with audio)",
"label": "FastH3 Preview on MLX",
"summary": "Run FastH3 Preview on Apple Silicon with a locally converted INT6 DiT, streamed Qwen3-VL conditioning, and native MLX video and audio VAEs.",
"model": "FastVideo/FastVideo-Minimax-FastH3-Preview-v0.2",
"source": "examples/inference/basic/mlx_fasth3.py",
"command": "hf download FastVideo/FastVideo-Minimax-FastH3-Preview-v0.2 --local-dir ./FastH3-Preview-v0.2\npython scripts/checkpoint_conversion/convert_minimax_h3_mlx.py --model-root ./FastH3-Preview-v0.2/transformer --out ./FastH3-MLX --formats \"int6\"\npython examples/inference/basic/mlx_fasth3.py --model-root ./FastH3-Preview-v0.2 --mlx-checkpoint ./FastH3-MLX/int6 --prompt \"(S1) A presenter says <d>[English] FastVideo runs FastH3.</d>\" --height 480 --width 832 --num-frames 124 --seed 2026 --output-path ./outputs/fasth3_int6.mp4",
"gpu_types": ["Apple Silicon"],
"hardware": {
"platform": "mlx",
"accelerator": "Apple M4 Max",
"system_memory": "36 GB unified memory",
"peak_memory": "19.63 GiB peak MLX memory during denoising",
"evidence": "validated",
"evidence_url": "https://github.com/hao-ai-lab/FastVideo/pull/1770"
},
"evidence": "Verified",
"expected_artifact": "MP4 with H.264 video and stereo AAC audio at outputs/fasth3_int6.mp4",
"modes": ["T2VA", "temporal --fast", "spatial --fast-spatial", "opt-in VSA"],
"knobs": [
{"key": "video_decode_backend", "label": "VAE decode", "hint": "Fidelity vs. speed", "flag": "--video-decode-backend", "options": [{"value": "h3-vae", "label": "Full H3 VAE"}, {"value": "taeh3", "label": "TAEH3 preview"}], "default": "h3-vae"}
],
"limitations": [
"The MLX path supports T2VA, optional temporal --fast, optional spatial --fast-spatial, and opt-in VSA on --include-vsa checkpoints. FL2VA, Ref2VA, and two-pass refinement are not wired."
]
},
{
"id": "fasth3-preview-spark",
"group": "fasth3-preview",
"group_label": "FastH3 Preview",
"group_task": "4-step text to video + audio",
"family": "minimax_h3",
"stage": "inference",
"task": "Few-step text to video (with audio)",
"label": "FastH3 Preview on one DGX Spark",
"summary": "Run FastH3 Preview on one GB10 with Triton VSA, FA4 off, and lazy module load. Height, width, frames, and steps in the YAML are examples.",
"model": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree",
"source": "examples/inference/basic/basic_fasth3_spark.yaml",
"serving": {
"source": "examples/serving/openai_fasth3_spark.yaml",
"install": "UV_TORCH_BACKEND=cu130 uv pip install -e .",
"env": "FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3"
},
"command": "FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3 FASTVIDEO_STAGE_LOGGING=1 fastvideo generate --config examples/inference/basic/basic_fasth3_spark.yaml",
"gpu_types": ["NVIDIA"],
"hardware": {
"platform": "cuda",
"device": "spark",
"gpu_count": 1,
"evidence": "source-configured"
},
"evidence": "Source-backed",
"expected_artifact": "MP4 under outputs/fasth3_spark/",
"modes": ["T2VA", "1-Spark"],
"limitations": [
"Install from the DGX Spark guide, not the generic CUDA extra. GB10 has no FA4 / sm_100a VSA kernel; keep FASTVIDEO_FA4=0 and FASTVIDEO_VSA_SM100A=0.",
"Legal num_frames values are 17n+5, capped at 345 (15 s). Native 16:9 sizes include 832x480 and 1344x768.",
"Lazy module load reloads Qwen3-VL and the DiT between phases of each request. Do not pass --no-lazy-module-load on this box.",
"A 345-frame request on one Spark can OOM. Prefer 124 or 243 frames, TAEH3 decode, or two Sparks over QSFP."
]
},
{
"id": "fasth3-spark-pair",
"group": "fasth3-preview",
"group_label": "FastH3 Preview",
"group_task": "4-step text to video + audio",
"family": "minimax_h3",
"stage": "inference",
"task": "Few-step text to video (with audio)",
"label": "FastH3 Preview on two DGX Sparks",
"summary": "Run one FastH3 clip across two GB10s with Ray sequence parallel over QSFP RoCE. Sequential load and lazy module load stay on because SP replicates the DiT on each node.",
"model": "FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree",
"source": "examples/inference/basic/basic_fasth3_spark_pair.yaml",
"command": "source examples/inference/optimizations/spark_pair_env.sh && FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3 FASTVIDEO_VAE_PARALLEL_DECODE=1 fastvideo generate --config examples/inference/basic/basic_fasth3_spark_pair.yaml",
"gpu_types": ["NVIDIA"],
"hardware": {
"platform": "cuda",
"device": "spark",
"gpu_count": 2,
"accelerator": "NVIDIA GB10 (DGX Spark pair)",
"evidence": "validated",
"evidence_url": "https://github.com/hao-ai-lab/FastVideo/pull/1803"
},
"evidence": "Verified",
"expected_artifact": "MP4 under outputs/fasth3_spark_pair/",
"modes": ["T2VA", "2-Spark SP"],
"limitations": [
"Requires a two-node Ray cluster on the QSFP interconnect. There is no cookbook server for this path; use Python / generate.",
"Height, width, frames, and steps in the YAML are examples. Edit them or pass CLI flags. See docs/getting_started/installation/spark_pair.md."
]
},
{
"id": "minimax-h3-fl2va",
"family": "minimax_h3",
"stage": "inference",
"task": "First/last frame to video (with audio)",
"label": "MiniMax H3 FL2VA",
"summary": "Animate a first frame, optionally guide the final frame, and generate synchronized audio with the full MiniMax H3 checkpoint.",
"model": "MiniMaxAI/MiniMax-H3",
"source": "examples/inference/basic/basic_minimax_h3_fl2va.py",
"command": "python examples/inference/basic/basic_minimax_h3_fl2va.py --image path/to/first-frame.png --prompt \"(S1) The subject turns toward the camera and says <d>[English] Hello.</d>\"",
"gpu_types": ["NVIDIA"],
"hardware": {"platform": "cuda", "gpu_count": 4, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 with synchronized audio at outputs/minimax_h3_fl2va/minimax_h3_fl2va.mp4",
"modes": ["FL2VA"],
"knobs": [
{"key": "num_gpus", "label": "GPUs", "hint": "Sequence-parallel degree", "flag": "--num-gpus", "options": [1, 2, 4, 8], "default": 4}
],
"limitations": ["Pass --last-image to constrain the final frame. The checked-in source defaults to four GPUs."]
},
{
"id": "minimax-h3-ref2va",
"family": "minimax_h3",
"stage": "inference",
"task": "Reference media to video (with audio)",
"label": "MiniMax H3 Ref2VA",
"summary": "Condition H3 on an ordered reference video and optional audio reference, then generate a new synchronized video and audio result.",
"model": "MiniMaxAI/MiniMax-H3",
"source": "examples/inference/basic/basic_minimax_h3_ref2va.py",
"command": "python examples/inference/basic/basic_minimax_h3_ref2va.py --reference-video path/to/reference.mp4 --prompt \"Create a new scene that preserves the reference identity and motion language.\"",
"gpu_types": ["NVIDIA"],
"hardware": {"platform": "cuda", "gpu_count": 4, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 with synchronized audio at outputs/minimax_h3_ref2va/minimax_h3_ref2va.mp4",
"modes": ["Ref2VA"],
"knobs": [
{"key": "num_gpus", "label": "GPUs", "hint": "Sequence-parallel degree", "flag": "--num-gpus", "options": [1, 2, 4, 8], "default": 4}
],
"limitations": ["Pass --reference-audio for an additional audio reference. The checked-in source defaults to four GPUs."]
},
{
"id": "fasth3-lora-preview",
"family": "minimax_h3",
"stage": "inference",
"task": "LoRA-adapted few-step video (with audio)",
"label": "FastH3 LoRA Preview",
"summary": "Apply a FastH3 preview adapter at load time while keeping the shared four-forward performance profile and synchronized audio output.",
"model": "MiniMaxAI/MiniMax-H3",
"source": "examples/inference/basic/basic_fasth3_lora_preview.py",
"command": "python examples/inference/basic/basic_fasth3_lora_preview.py --lora-path path/to/adapter.safetensors --prompt \"(S1) A presenter says <d>[English] This is an adapted Fast H3 run.</d>\"",
"gpu_types": ["NVIDIA"],
"hardware": {
"platform": "cuda",
"gpu_count": 4,
"accelerator": "NVIDIA B200",
"evidence": "validated",
"evidence_url": "https://github.com/hao-ai-lab/FastVideo/pull/1771"
},
"evidence": "Verified",
"expected_artifact": "Warmup and measured MP4 files under outputs/fasth3_lora_preview/",
"modes": ["T2VA", "FastH3 LoRA"],
"knobs": [
{"key": "num_gpus", "label": "GPUs", "hint": "Sequence-parallel degree", "flag": "--num-gpus", "options": [1, 2, 4, 8], "default": 4},
{"key": "video_decode_backend", "label": "VAE decode", "hint": "Fidelity vs. speed", "flag": "--video-decode-backend", "options": [{"value": "h3-vae", "label": "Full H3 VAE"}, {"value": "taeh3", "label": "TAEH3 preview"}], "default": "h3-vae"}
],
"limitations": ["Supply a compatible FastH3 adapter. The script infers dense or VSA attention from the adapter payload unless you override it."]
},
{
"id": "longcat-t2v",
"family": "longcat",
"stage": "inference",
"task": "Text to video",
"label": "LongCat Video T2V",
"summary": "LongCat Video text-to-video at 480p (50 steps), with distilled and 720p refinement passes included in the same maintained script.",
"model": "FastVideo/LongCat-Video-T2V-Diffusers",
"source": "examples/inference/basic/basic_longcat_t2v.py",
"command": "python examples/inference/basic/basic_longcat_t2v.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 videos under outputs_video/longcat_t2v_basic/, longcat_t2v_distill/, and longcat_t2v_refine_720p/",
"related": ["longcat-i2v"]
},
{
"id": "longcat-i2v",
"family": "longcat",
"stage": "inference",
"task": "Image to video",
"label": "LongCat Video I2V",
"summary": "LongCat Video image-to-video with optional distilled and refinement passes, from its maintained example.",
"model": "FastVideo/LongCat-Video-I2V-Diffusers",
"source": "examples/inference/basic/basic_longcat_i2v.py",
"command": "python examples/inference/basic/basic_longcat_i2v.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 videos under outputs_video/longcat_i2v_basic/ and longcat_i2v_distill/",
"related": ["longcat-t2v"]
},
{
"id": "stable-audio-open-t2a",
"family": "stable_audio",
"stage": "inference",
"task": "Text to audio",
"label": "Stable Audio Open 1.0",
"summary": "Six-second text-to-audio generation with Stable Audio Open 1.0 from its maintained example; duration and steps are documented knobs in the source.",
"model": "FastVideo/stable-audio-open-1.0-Diffusers",
"source": "examples/inference/basic/basic_stable_audio.py",
"command": "python examples/inference/basic/basic_stable_audio.py",
"gpu_types": ["NVIDIA"],
"hardware": {
"platform": "cuda",
"gpu_count": 1,
"accelerator": "NVIDIA B200",
"evidence": "validated",
"evidence_url": "https://github.com/hao-ai-lab/FastVideo/pull/1260"
},
"evidence": "Verified",
"expected_artifact": "WAV audio at outputs_audio/stable_audio_basic/output_stable_audio.wav",
"limitations": ["Must load the FastVideo converted Diffusers repo; upstream stabilityai monolithic checkpoints are not loader-compatible (see scripts/checkpoint_conversion/stable_audio_to_diffusers.py)."],
"related": ["stable-audio-small-t2a"]
},
{
"id": "stable-audio-small-t2a",
"family": "stable_audio",
"stage": "inference",
"task": "Text to audio",
"label": "Stable Audio Open Small",
"summary": "The smaller Stable Audio Open variant with its own shorter training window, from its maintained example.",
"model": "FastVideo/stable-audio-open-small-Diffusers",
"source": "examples/inference/basic/basic_stable_audio_small.py",
"command": "python examples/inference/basic/basic_stable_audio_small.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"related": ["stable-audio-open-t2a"]
},
{
"id": "mmaudio-v2a",
"family": "mmaudio",
"stage": "inference",
"task": "Video/Text to audio",
"label": "MMAudio large 44k v2",
"summary": "Add synchronized audio to a video (or from a prompt) with MMAudio large 44k v2. The example reads the model path from MMAUDIO_MODEL_PATH; this recipe passes the converted Hugging Face repo explicitly.",
"model": "FastVideo/MMAudio-large-44k-v2-Diffusers",
"source": "examples/inference/basic/basic_mmaudio.py",
"command": "MMAUDIO_MODEL_PATH=FastVideo/MMAudio-large-44k-v2-Diffusers python examples/inference/basic/basic_mmaudio.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"limitations": ["The upstream checkpoint must be converted to Diffusers layout via scripts/checkpoint_conversion/convert_mmaudio_to_diffusers.py unless loaded from the FastVideo converted repo as done here."]
},
{
"id": "matrix-game-2",
"family": "matrixgame",
"stage": "inference",
"task": "Interactive world",
"label": "Matrix Game 2.0",
"summary": "Generate an interactive-world sequence from the maintained Matrix Game 2.0 example.",
"model": "FastVideo/Matrix-Game-2.0-Base-Distilled-Diffusers",
"source": "examples/inference/basic/basic_matrixgame2.py",
"command": "python examples/inference/basic/basic_matrixgame2.py"
"command": "python examples/inference/basic/basic_matrixgame2.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"related": ["matrix-game-3-i2w"]
},
{
"id": "matrix-game-3-i2w",
"family": "matrixgame",
"stage": "inference",
"task": "Interactive world",
"label": "Matrix Game 3.0",
"summary": "Drive Matrix Game 3.0 from an input image plus prompt at 720p, three steps, from its maintained example.",
"model": "FastVideo/Matrix-Game-3.0-Base-Distilled-Diffusers",
"source": "examples/inference/basic/basic_matrixgame3.py",
"command": "python examples/inference/basic/basic_matrixgame3.py",
"gpu_types": ["NVIDIA"],
"hardware": {"gpu_count": 1, "evidence": "source-configured"},
"evidence": "Source-backed",
"expected_artifact": "MP4 videos under video_samples_matrixgame3/",
"related": ["matrix-game-2"]
}
]
}
+711 -44
View File
@@ -1,57 +1,724 @@
(() => {
let recipesPromise;
const PATTERN_CHARS = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789";
const PATTERN_LENGTH = 900;
const SCRAMBLE_MS = 140;
const loadRecipes = (url) => {
recipesPromise ||= fetch(url).then((response) => {
const loadRecipes = (url) =>
fetch(url).then((response) => {
if (!response.ok) throw new Error(`HTTP ${response.status}`);
return response.json();
});
return recipesPromise;
const generatePattern = (length) => {
const chars = new Array(length);
for (let index = 0; index < length; index += 1) {
chars[index] = PATTERN_CHARS.charAt((Math.random() * PATTERN_CHARS.length) | 0);
}
return chars.join("");
};
const motionQuery = () =>
window.matchMedia("(hover: hover) and (pointer: fine) and (prefers-reduced-motion: no-preference)");
const initEvervault = (root) => {
root.querySelectorAll("[data-evervault]").forEach((visual) => {
if (visual.dataset.evervaultReady) return;
visual.dataset.evervaultReady = "true";
const noise = visual.querySelector("[data-cookbook-pattern]");
if (!noise) return;
const query = motionQuery();
if (!query.matches) return;
let rect = null;
let pointerX = 0;
let pointerY = 0;
let frame = 0;
let lastScramble = 0;
let patternReady = false;
const paint = (scramble) => {
visual.style.setProperty("--mouse-x", `${pointerX}px`);
visual.style.setProperty("--mouse-y", `${pointerY}px`);
if (scramble) noise.textContent = generatePattern(PATTERN_LENGTH);
frame = 0;
};
const onEnter = () => {
rect = visual.getBoundingClientRect();
if (!patternReady) {
noise.textContent = generatePattern(PATTERN_LENGTH);
patternReady = true;
}
};
const onMove = (event) => {
if (!rect) rect = visual.getBoundingClientRect();
pointerX = event.clientX - rect.left;
pointerY = event.clientY - rect.top;
if (frame) return;
frame = window.requestAnimationFrame(() => {
const now = performance.now();
const scramble = now - lastScramble >= SCRAMBLE_MS;
if (scramble) lastScramble = now;
paint(scramble);
});
};
const onLeave = () => {
if (frame) window.cancelAnimationFrame(frame);
frame = 0;
rect = null;
pointerX = 0;
pointerY = 0;
visual.style.setProperty("--mouse-x", "50%");
visual.style.setProperty("--mouse-y", "50%");
};
visual.addEventListener("mouseenter", onEnter);
visual.addEventListener("mousemove", onMove);
visual.addEventListener("mouseleave", onLeave);
});
};
let familyPopstate = null;
const bindFamilyPopstate = () => {
if (bindFamilyPopstate.bound) return;
bindFamilyPopstate.bound = true;
window.addEventListener("popstate", () => {
if (typeof familyPopstate === "function") familyPopstate();
});
};
const groupIdFor = (recipe) => recipe.group || recipe.id;
const gpuCountLabel = (hardware, runtimeId) => {
const count = hardware?.gpu_count;
if (count == null) return "";
if (runtimeId === "spark") return `${count} Spark${count === 1 ? "" : "s"}`;
return `${count} GPU${count === 1 ? "" : "s"}`;
};
const knobsFor = (recipe) => recipe.knobs || [];
const knobOptions = (knob) =>
knob.options.map((option) =>
option !== null && typeof option === "object" ? option : { value: option, label: String(option) },
);
const knobDefaultLabel = (knob) => {
const match = knobOptions(knob).find((option) => option.value === knob.default);
return match ? match.label : String(knob.default);
};
// Knob flags are always shown explicitly in the displayed command, even at
// their default value, so the command stays copy-pasteable and precise
// about what it runs -- not just "trust the script's own default".
const appendKnobFlags = (commandText, knobs, knobValues) => {
const flags = knobs
.filter((knob) => knobValues[knob.key] !== undefined)
.map((knob) => `${knob.flag} ${knobValues[knob.key]}`);
if (!flags.length) return commandText;
const lines = commandText.split("\n");
lines[lines.length - 1] = `${lines[lines.length - 1]} ${flags.join(" ")}`;
return lines.join("\n");
};
const runtimeFor = (recipe) => {
const platform = recipe.hardware?.platform || "cuda";
if (platform === "mlx") {
return {
id: "mlx",
label: "Apple Silicon · MLX",
hint:
recipe.hardware?.minimum_memory ||
[recipe.hardware?.accelerator, recipe.hardware?.system_memory].filter(Boolean).join(" · ") ||
"Memory not recorded",
};
}
if (platform === "mps") {
return {
id: "mps",
label: "Apple Silicon · MPS",
hint: recipe.hardware?.minimum_memory || recipe.hardware?.system_memory || "Memory not recorded",
};
}
const hardware = recipe.hardware || {};
if (hardware.device === "spark") {
return {
id: "spark",
label: "NVIDIA DGX Spark",
hint: "GB10 · 128 GB unified memory",
};
}
return {
id: "cuda",
label: "NVIDIA CUDA",
hint: hardware.accelerator
? `${hardware.accelerator} · ${gpuCountLabel(hardware, "cuda")}`
: `${gpuCountLabel(hardware, "cuda")} configured · GPU model not recorded`,
};
};
const runtimeSummary = (recipe) => {
const runtime = runtimeFor(recipe);
const hardware = recipe.hardware || {};
if (runtime.id === "mlx") {
return [hardware.accelerator || "Apple Silicon", hardware.system_memory || hardware.minimum_memory, "MLX"]
.filter(Boolean)
.join(" · ");
}
if (runtime.id === "mps") {
return [hardware.accelerator || "Apple Silicon", hardware.system_memory || hardware.minimum_memory, "PyTorch MPS"]
.filter(Boolean)
.join(" · ");
}
if (runtime.id === "spark") {
return [hardware.accelerator || "NVIDIA GB10", gpuCountLabel(hardware, "spark")].filter(Boolean).join(" · ");
}
if (hardware.accelerator) return [hardware.accelerator, gpuCountLabel(hardware, runtime.id)].join(" · ");
return `NVIDIA CUDA · ${gpuCountLabel(hardware, "cuda")} configured · GPU model and VRAM not recorded`;
};
const renderHardwareEvidence = (container, badge, recipe) => {
const hardware = recipe.hardware || {};
const runtime = runtimeFor(recipe);
const isValidated = hardware.evidence === "validated";
container.classList.toggle("cookbook-hardware-state--verified", isValidated);
container.classList.toggle("cookbook-hardware-state--source", !isValidated);
badge.classList.toggle("cookbook-badge--verified", isValidated);
badge.classList.toggle("cookbook-badge--configured", !isValidated);
badge.textContent = isValidated ? "Recorded run" : "Source config";
const heading = document.createElement("strong");
heading.textContent = isValidated ? "Recorded hardware" : "Source configuration";
const details = document.createElement("span");
if (isValidated) {
const recorded = [
hardware.accelerator,
runtime.id === "mlx" || runtime.id === "mps" ? hardware.system_memory : gpuCountLabel(hardware, runtime.id),
].filter(Boolean);
const statements = [`${recorded.join(" · ")}.`];
if (hardware.minimum_memory) statements.push(`Documented minimum: ${hardware.minimum_memory}.`);
if (hardware.peak_memory) statements.push(`Measured: ${hardware.peak_memory}.`);
if (!hardware.minimum_memory) statements.push("This recorded device is not a minimum requirement.");
details.textContent = ` ${statements.join(" ")}`;
} else if (runtime.id === "spark") {
details.textContent = ` NVIDIA DGX Spark · ${gpuCountLabel(hardware, "spark")}. GB10 has no FA4 / sm_100a VSA kernel; keep Triton VSA and FA4 off.`;
} else if (runtime.id === "cuda") {
details.textContent = ` NVIDIA CUDA · ${gpuCountLabel(hardware, "cuda")}. The source does not record the GPU model or VRAM.`;
} else {
details.textContent = ` ${runtime.label}. The source does not record a device or memory requirement.`;
}
container.replaceChildren(heading, details);
if (hardware.evidence_url) {
const evidenceLink = document.createElement("a");
evidenceLink.href = hardware.evidence_url;
evidenceLink.textContent = "View run evidence";
evidenceLink.setAttribute("aria-label", `View recorded hardware evidence for ${recipe.label}`);
container.append(" ", evidenceLink);
}
};
const compactLifecycle = (root) => {
const lifecycle = root.querySelector(".cookbook-lifecycle");
if (!lifecycle || lifecycle.dataset.compact) return;
lifecycle.dataset.compact = "true";
const stages = [...lifecycle.querySelectorAll(".cookbook-lifecycle__stage")];
const active = stages.find((stage) => stage.classList.contains("cookbook-lifecycle__stage--active"));
const planned = stages
.filter((stage) => stage !== active)
.map((stage) => stage.childNodes[0]?.textContent?.trim())
.filter(Boolean);
const summary = document.createElement("span");
summary.className = "cookbook-lifecycle__summary";
summary.textContent = `Next: ${planned.join(", ")}`;
lifecycle.replaceChildren(...(active ? [active] : []), summary);
};
const initFamilyBuilder = async (root) => {
const family = root.dataset.family;
if (!family) return;
compactLifecycle(root);
const modelOptions = root.querySelector("[data-cookbook-model-options]");
const hardwareOptions = root.querySelector("[data-cookbook-hardware-options]");
const description = root.querySelector("[data-cookbook-description]");
const label = root.querySelector("[data-cookbook-label]");
const model = root.querySelector("[data-cookbook-model]");
const task = root.querySelector("[data-cookbook-task]");
const hardwareValue = root.querySelector("[data-cookbook-gpus]");
const artifact = root.querySelector("[data-cookbook-artifact]");
const evidenceCell = root.querySelector("[data-cookbook-evidence]");
const source = root.querySelector("[data-cookbook-source]");
const modelLink = root.querySelector("[data-cookbook-model-link]");
const command = root.querySelector("[data-cookbook-command]");
const status = root.querySelector("[data-cookbook-status]");
const hardwareState = root.querySelector("[data-cookbook-hardware-state]");
const hardwareBadge = root.querySelector("[data-cookbook-hardware-badge]");
const count = root.querySelector("[data-cookbook-count]");
const result = root.querySelector(".cookbook-result");
const commandBlock = root.querySelector(".cookbook-command");
const servingPanel = root.querySelector("[data-cookbook-serving]");
const usage = root.querySelector("[data-cookbook-usage]");
const servingAvailability = root.querySelector("[data-cookbook-serving-availability]");
const knobsContainer = root.querySelector("[data-cookbook-knobs]");
const deviceRow = root.querySelector("[data-cookbook-device-row]");
const deviceOptions = root.querySelector("[data-cookbook-device-options]");
const deviceCaption = root.querySelector("[data-cookbook-device-caption]");
modelOptions.setAttribute("aria-label", "Recipe");
hardwareOptions.setAttribute("aria-label", "Runtime");
let recipes;
try {
({ recipes } = await loadRecipes(root.dataset.recipes));
} catch (error) {
if (status) status.textContent = "Recipes could not be loaded. Use the maintained examples link below.";
console.error("Failed to load FastVideo cookbook recipes", error);
return;
}
if (!root.isConnected) return;
const familyRecipes = recipes.filter((recipe) => recipe.family === family);
if (!familyRecipes.length) return;
let servingProfiles = {};
let servingLoadFailed = false;
if (servingPanel && familyRecipes.some((recipe) => recipe.serving)) {
try {
const dataUrl = new URL(root.dataset.recipes, document.baseURI);
servingProfiles = await loadRecipes(new URL("cookbook-serving.json", dataUrl));
} catch (error) {
servingLoadFailed = true;
console.error("Failed to load FastVideo serving profiles", error);
}
}
const byId = new Map(familyRecipes.map((recipe) => [recipe.id, recipe]));
const groups = new Map();
familyRecipes.forEach((recipe) => {
const groupId = groupIdFor(recipe);
if (!groups.has(groupId)) groups.set(groupId, []);
groups.get(groupId).push(recipe);
});
if (count) count.textContent = `${familyRecipes.length} maintained recipes`;
modelOptions.replaceChildren();
groups.forEach((groupRecipes, groupId) => {
const representative = groupRecipes[0];
const option = document.createElement("button");
option.type = "button";
option.dataset.recipeGroup = groupId;
option.setAttribute("aria-pressed", "false");
const optionLabel = document.createElement("strong");
optionLabel.textContent = representative.group_label || representative.label;
const optionTask = document.createElement("span");
optionTask.textContent = representative.group_task || representative.task;
option.append(optionLabel, optionTask);
modelOptions.append(option);
});
const query = new URLSearchParams(window.location.search);
const requestedRecipe = query.get("recipe");
const defaultRecipeId = byId.has(root.dataset.defaultRecipe) ? root.dataset.defaultRecipe : familyRecipes[0].id;
let selectedRecipeId = requestedRecipe && byId.has(requestedRecipe) ? requestedRecipe : defaultRecipeId;
let selectedGroupId = groupIdFor(byId.get(selectedRecipeId));
let renderedRuntimeGroup = null;
// Keep previously shared local/openai links working after renaming workflows.
const workflow = (value) => ["local", "python"].includes(value) ? "python" : "server";
let usagePreference = workflow(query.get("use"));
let selectedClient = ["python", "javascript", "curl"].includes(query.get("client")) ? query.get("client") : "curl";
const clientDetails = servingPanel?.querySelector(".cookbook-serving__code");
if (clientDetails && query.has("client")) clientDetails.open = true;
const knobDefs = new Map();
familyRecipes.forEach((recipe) => knobsFor(recipe).forEach((knob) => {
if (!knobDefs.has(knob.key)) knobDefs.set(knob.key, knob);
}));
const knobValues = {};
knobDefs.forEach((knob, key) => {
const fromQuery = query.get(key);
const validValues = knobOptions(knob).map((option) => String(option.value));
const useQueryValue = fromQuery !== null && validValues.includes(fromQuery);
const raw = useQueryValue ? fromQuery : knob.default;
knobValues[key] = typeof knob.default === "number" ? Number(raw) : raw;
});
const renderKnobs = (recipe, hidden) => {
if (!knobsContainer) return;
const knobs = knobsFor(recipe);
const renderedKeys = [...knobsContainer.querySelectorAll("[data-knob-row]")].map((row) => row.dataset.knobRow);
if (renderedKeys.join(",") !== knobs.map((knob) => knob.key).join(",")) {
knobsContainer.replaceChildren();
knobs.forEach((knob) => {
const row = document.createElement("div");
row.className = "cookbook-selection-row";
row.dataset.knobRow = knob.key;
const labelWrap = document.createElement("div");
labelWrap.className = "cookbook-selection-row__label";
const strongLabel = document.createElement("strong");
strongLabel.textContent = knob.label;
const hintLabel = document.createElement("span");
hintLabel.textContent = knob.hint || "";
labelWrap.append(strongLabel, hintLabel);
const grid = document.createElement("div");
grid.className = "cookbook-option-grid cookbook-option-grid--hardware";
grid.setAttribute("role", "group");
grid.setAttribute("aria-label", knob.label);
knobOptions(knob).forEach((option) => {
const optionButton = document.createElement("button");
optionButton.type = "button";
optionButton.dataset.knobKey = knob.key;
optionButton.dataset.knobValue = String(option.value);
optionButton.setAttribute("aria-pressed", "false");
const optionLabel = document.createElement("strong");
optionLabel.textContent = option.label;
optionButton.append(optionLabel);
grid.append(optionButton);
});
row.append(labelWrap, grid);
knobsContainer.append(row);
});
}
knobsContainer.hidden = hidden || knobs.length === 0;
knobsContainer.querySelectorAll("button[data-knob-key]").forEach((optionButton) => {
const selected = String(knobValues[optionButton.dataset.knobKey]) === optionButton.dataset.knobValue;
optionButton.classList.toggle("cookbook-option--selected", selected);
optionButton.setAttribute("aria-pressed", String(selected));
});
};
const recipesForRuntime = (runtimeId, groupRecipes) =>
groupRecipes.filter((item) => runtimeFor(item).id === runtimeId);
const uniqueRuntimeIds = (groupRecipes) => {
const ids = [];
groupRecipes.forEach((item) => {
const id = runtimeFor(item).id;
if (!ids.includes(id)) ids.push(id);
});
return ids;
};
const pickRecipeForRuntime = (runtimeId, preferredCount, groupRecipes) => {
const siblings = recipesForRuntime(runtimeId, groupRecipes);
if (!siblings.length) return null;
if (preferredCount != null) {
const match = siblings.find((item) => item.hardware?.gpu_count === preferredCount);
if (match) return match;
}
return siblings.find((item) => item.hardware?.gpu_count === 1) || siblings[0];
};
const renderRuntimeOptions = () => {
const groupRecipes = groups.get(selectedGroupId) || [];
const runtimeIds = uniqueRuntimeIds(groupRecipes);
const renderedIds = [...hardwareOptions.querySelectorAll("[data-runtime-id]")].map((option) => option.dataset.runtimeId);
if (renderedIds.join(",") === runtimeIds.join(",")) return;
hardwareOptions.replaceChildren();
runtimeIds.forEach((runtimeId) => {
const representative = pickRecipeForRuntime(runtimeId, 1, groupRecipes);
const runtime = runtimeFor(representative);
const option = document.createElement("button");
option.type = "button";
option.dataset.runtimeId = runtime.id;
option.setAttribute("aria-pressed", "false");
const optionLabel = document.createElement("strong");
optionLabel.textContent = runtime.label;
const optionHint = document.createElement("span");
optionHint.textContent = runtime.hint;
option.append(optionLabel, optionHint);
hardwareOptions.append(option);
});
};
const renderDeviceOptions = (recipe) => {
if (!deviceRow || !deviceOptions) return;
const groupRecipes = groups.get(selectedGroupId) || [];
const runtime = runtimeFor(recipe);
const siblings = recipesForRuntime(runtime.id, groupRecipes)
.slice()
.sort((left, right) => (left.hardware?.gpu_count || 0) - (right.hardware?.gpu_count || 0));
const show = siblings.length > 1;
deviceRow.hidden = !show;
if (deviceCaption) {
deviceCaption.textContent = runtime.id === "spark" ? "1 Spark or a QSFP pair" : "GPU count for this runtime";
}
if (!show) {
deviceOptions.replaceChildren();
return;
}
const renderedIds = [...deviceOptions.querySelectorAll("[data-recipe-id]")].map((option) => option.dataset.recipeId);
if (renderedIds.join(",") !== siblings.map((item) => item.id).join(",")) {
deviceOptions.replaceChildren();
siblings.forEach((candidate) => {
const option = document.createElement("button");
option.type = "button";
option.dataset.recipeId = candidate.id;
option.setAttribute("aria-pressed", "false");
const optionLabel = document.createElement("strong");
optionLabel.textContent = gpuCountLabel(candidate.hardware, runtime.id);
const optionHint = document.createElement("span");
optionHint.textContent = runtime.id === "spark" && candidate.hardware?.gpu_count === 2
? "Ray sequence parallel over QSFP"
: runtime.id === "spark"
? "One GB10, local process"
: `${gpuCountLabel(candidate.hardware, runtime.id)} configured`;
option.append(optionLabel, optionHint);
deviceOptions.append(option);
});
}
deviceOptions.querySelectorAll("button").forEach((option) => {
const selected = option.dataset.recipeId === recipe.id;
option.classList.toggle("cookbook-option--selected", selected);
option.setAttribute("aria-pressed", String(selected));
});
};
let notes = root.querySelector("[data-cookbook-notes]");
if (!notes) {
notes = document.createElement("aside");
notes.className = "cookbook-recipe-notes";
notes.dataset.cookbookNotes = "";
notes.hidden = true;
result.insertBefore(notes, commandBlock);
}
const render = ({ groupChanged = false, historyMode = "replace" } = {}) => {
if (!root.isConnected) return;
if (!byId.has(selectedRecipeId)) selectedRecipeId = defaultRecipeId;
let recipe = byId.get(selectedRecipeId);
if (groupChanged || groupIdFor(recipe) !== selectedGroupId) {
const currentRuntime = runtimeFor(recipe).id;
const currentCount = recipe.hardware?.gpu_count;
const groupRecipes = groups.get(selectedGroupId) || [];
recipe = pickRecipeForRuntime(currentRuntime, currentCount, groupRecipes) || groupRecipes[0];
selectedRecipeId = recipe.id;
}
selectedGroupId = groupIdFor(recipe);
if (renderedRuntimeGroup !== selectedGroupId) {
renderRuntimeOptions();
renderedRuntimeGroup = selectedGroupId;
}
renderDeviceOptions(recipe);
const runtime = runtimeFor(recipe);
const profile = servingPanel && servingProfiles[recipe.id];
const useServer = Boolean(profile && usagePreference === "server");
// The measured local profile and the server config have separate evidence.
const activeRecipe = useServer ? { ...recipe, hardware: profile.hardware, evidence: "Source-backed" } : recipe;
const knobs = knobsFor(recipe);
renderKnobs(recipe, useServer);
if (usage) {
usage.querySelectorAll("[data-cookbook-mode]").forEach((option) => {
const selected = option.dataset.cookbookMode === (useServer ? "server" : "python");
option.disabled = option.dataset.cookbookMode === "server" && !profile;
option.classList.toggle("cookbook-option--selected", selected);
option.setAttribute("aria-pressed", String(selected));
});
servingAvailability.textContent = profile
? "The playground and API clients share one server process. Both workflows can run on your own machine."
: servingLoadFailed
? "Server examples could not be loaded. Open the H3 server guide below, or use Python directly."
: "This recipe uses Python directly. For the playground and API clients, choose FastH3 Preview with CUDA, MLX, or one Spark.";
servingPanel.hidden = !useServer;
commandBlock.hidden = useServer;
root.querySelector("[data-cookbook-python-note]").hidden = useServer;
}
if (useServer) {
const isMLX = profile.runtime === "mlx";
const isSpark = runtime.id === "spark";
servingPanel.querySelector("[data-cookbook-server-lifetime]").textContent = isMLX
? "Start once, then change prompts in the playground or your app. MLX reuses its pipeline and prompt cache, but loads and releases model components between phases to limit unified-memory use. It does not keep all weights resident."
: isSpark
? "Start once, then change prompts in the playground or your app. On a DGX Spark, lazy module load still reloads Qwen3-VL and the DiT between phases of each request, so later prompts are not a free hot cache."
: "Start once, then change prompts in the playground or your app. CUDA requests reuse the loaded model. The Python SDK can also reuse a generator within one process.";
servingPanel.querySelector("[data-cookbook-install-guide]").href = isMLX
? "../../getting_started/installation/mps/#run-fasth3-preview"
: isSpark
? "../../getting_started/installation/spark/"
: "../../getting_started/installation/gpu/";
servingPanel.querySelector("[data-cookbook-prepare]").hidden = !profile.prepare;
servingPanel.querySelector("[data-cookbook-server-prepare]").textContent = profile.prepare;
servingPanel.querySelector("[data-cookbook-server-install]").textContent = profile.install;
servingPanel.querySelector("[data-cookbook-server-command]").textContent = profile.command;
servingPanel.querySelector("[data-cookbook-health-command]").textContent = profile.health_command;
servingPanel.querySelector("[data-cookbook-playground]").href = profile.playground_url;
const client = profile.clients[selectedClient];
const filename = client.source.split("/").pop();
servingPanel.querySelector("[data-cookbook-client-install]").textContent = client.install;
const clientCode = servingPanel.querySelector("[data-cookbook-client-code]");
clientCode.className = `language-${selectedClient === "curl" ? "bash" : selectedClient}`;
clientCode.textContent = client.code;
servingPanel.querySelector("[data-cookbook-client-filename]").textContent = filename;
servingPanel.querySelector("[data-cookbook-client-source]").href = `https://github.com/hao-ai-lab/FastVideo/blob/main/${client.source}`;
const runner = { python: "python", javascript: "node", curl: "bash" }[selectedClient];
servingPanel.querySelector("[data-cookbook-client-run]").textContent = `Save as ${filename} and run ${runner} ${filename}. The MP4 is saved with the job ID as its filename.`;
servingPanel.querySelectorAll("[data-cookbook-client]").forEach((option) => {
option.setAttribute("aria-pressed", String(option.dataset.cookbookClient === selectedClient));
});
}
modelOptions.querySelectorAll("button").forEach((option) => {
const selected = option.dataset.recipeGroup === selectedGroupId;
option.classList.toggle("cookbook-option--selected", selected);
option.setAttribute("aria-pressed", String(selected));
});
hardwareOptions.querySelectorAll("button").forEach((option) => {
const selected = option.dataset.runtimeId === runtime.id;
option.classList.toggle("cookbook-option--selected", selected);
option.setAttribute("aria-pressed", String(selected));
});
description.textContent = useServer
? `FastH3 Preview generates video with audio. This server profile uses the checked-in ${runtime.label} configuration.`
: recipe.summary;
label.textContent = useServer ? `${recipe.group_label || recipe.label} · Server` : recipe.label;
model.textContent = recipe.model;
task.textContent = recipe.task;
hardwareValue.textContent = runtimeSummary(activeRecipe);
if (artifact) artifact.textContent = useServer ? "MP4 with audio" : recipe.expected_artifact || "Not yet documented for this recipe.";
if (evidenceCell) {
evidenceCell.textContent = activeRecipe.evidence || "Source-backed";
evidenceCell.classList.toggle("cookbook-badge--verified", activeRecipe.evidence === "Verified");
evidenceCell.classList.toggle("cookbook-badge--source-backed", activeRecipe.evidence !== "Verified");
}
source.href = `https://github.com/hao-ai-lab/FastVideo/blob/main/${useServer ? profile.source : recipe.source}`;
source.textContent = useServer ? "View server configuration" : "Open example source";
modelLink.href = `https://huggingface.co/${recipe.model}`;
command.textContent = useServer ? recipe.command : appendKnobFlags(recipe.command, knobs, knobValues);
renderHardwareEvidence(hardwareState, hardwareBadge, activeRecipe);
const knobCaveats = useServer ? [] : knobs
.filter((knob) => String(knobValues[knob.key]) !== String(knob.default))
.map((knob) => `${knob.label} is set away from its recorded default (${knobDefaultLabel(knob)}). ` +
"The script accepts this value, but it has not been benchmarked here.");
const limitations = [...(useServer ? [
profile.runtime === "mlx"
? "This MLX server config has no recorded hardware run. Measurements from the Python recipe are not server memory requirements. Only text-to-video/audio is wired; reference inputs and fast modes are not exposed here."
: runtime.id === "spark"
? "This Spark server config has no recorded serving benchmark. Lazy module load reloads Qwen3-VL and the DiT between phases of each request. Compilation of the DiT is disabled."
: "This server config has no recorded serving benchmark. Compilation is disabled, unlike the measured Python performance profile.",
`${profile.sampling.width} × ${profile.sampling.height} · ${profile.sampling.num_frames} frames · ${profile.sampling.fps} fps. The server supplies these defaults; the client sends the model and prompt.`,
"Generation is serialized. Job metadata is held in memory and is lost when the server restarts.",
] : recipe.limitations || []), ...knobCaveats];
notes.replaceChildren();
notes.hidden = limitations.length === 0;
if (limitations.length) {
const notesHeading = document.createElement("strong");
notesHeading.textContent = "Know before you run";
const notesList = document.createElement("ul");
limitations.forEach((item) => {
const listItem = document.createElement("li");
listItem.textContent = item;
notesList.append(listItem);
});
notes.append(notesHeading, notesList);
}
const nextQuery = new URLSearchParams(window.location.search);
nextQuery.set("recipe", recipe.id);
nextQuery.set("runtime", runtime.id);
const deviceSiblings = recipesForRuntime(runtime.id, groups.get(selectedGroupId) || []);
if (deviceSiblings.length > 1 && recipe.hardware?.gpu_count != null) {
nextQuery.set("gpus", String(recipe.hardware.gpu_count));
} else {
nextQuery.delete("gpus");
}
knobDefs.forEach((knob, key) => {
if (knobs.some((activeKnob) => activeKnob.key === key)) nextQuery.set(key, String(knobValues[key]));
else nextQuery.delete(key);
});
if (usage) {
nextQuery.set("use", useServer ? "server" : "python");
if (useServer && clientDetails?.open) nextQuery.set("client", selectedClient);
else nextQuery.delete("client");
}
const nextUrl = `${window.location.pathname}?${nextQuery.toString()}${window.location.hash}`;
if (historyMode === "push") window.history.pushState({}, "", nextUrl);
else if (historyMode === "replace") window.history.replaceState({}, "", nextUrl);
const modeSummary = useServer ? " with a persistent server" : "";
status.textContent = `${recipe.label} selected for ${runtime.label}${modeSummary}.`;
};
modelOptions.addEventListener("click", (event) => {
const option = event.target.closest("button[data-recipe-group]");
if (!option) return;
selectedGroupId = option.dataset.recipeGroup;
render({ groupChanged: true, historyMode: "push" });
});
hardwareOptions.addEventListener("click", (event) => {
const option = event.target.closest("button[data-runtime-id]");
if (!option) return;
const groupRecipes = groups.get(selectedGroupId) || [];
const currentCount = byId.get(selectedRecipeId)?.hardware?.gpu_count;
const nextRecipe = pickRecipeForRuntime(option.dataset.runtimeId, currentCount, groupRecipes);
if (!nextRecipe) return;
selectedRecipeId = nextRecipe.id;
selectedGroupId = groupIdFor(nextRecipe);
render({ historyMode: "push" });
});
deviceOptions?.addEventListener("click", (event) => {
const option = event.target.closest("button[data-recipe-id]");
if (!option) return;
selectedRecipeId = option.dataset.recipeId;
selectedGroupId = groupIdFor(byId.get(selectedRecipeId));
render({ historyMode: "push" });
});
usage?.addEventListener("click", (event) => {
const option = event.target.closest("button[data-cookbook-mode]");
if (!option || option.disabled) return;
usagePreference = option.dataset.cookbookMode;
render({ historyMode: "push" });
});
servingPanel?.addEventListener("click", (event) => {
const option = event.target.closest("button[data-cookbook-client]");
if (!option) return;
selectedClient = option.dataset.cookbookClient;
render({ historyMode: "push" });
});
knobsContainer?.addEventListener("click", (event) => {
const option = event.target.closest("button[data-knob-key]");
if (!option) return;
const knob = knobDefs.get(option.dataset.knobKey);
knobValues[option.dataset.knobKey] = typeof knob.default === "number"
? Number(option.dataset.knobValue) : option.dataset.knobValue;
render({ historyMode: "push" });
});
render();
familyPopstate = () => {
if (!root.isConnected) return;
const nextQuery = new URLSearchParams(window.location.search);
const nextRecipe = nextQuery.get("recipe");
selectedRecipeId = nextRecipe && byId.has(nextRecipe) ? nextRecipe : defaultRecipeId;
selectedGroupId = groupIdFor(byId.get(selectedRecipeId));
usagePreference = workflow(nextQuery.get("use"));
selectedClient = ["python", "javascript", "curl"].includes(nextQuery.get("client")) ? nextQuery.get("client") : "curl";
if (clientDetails) clientDetails.open = nextQuery.has("client");
knobDefs.forEach((knob, key) => {
const fromQuery = nextQuery.get(key);
const validValues = knobOptions(knob).map((option) => String(option.value));
if (fromQuery !== null && validValues.includes(fromQuery)) {
knobValues[key] = typeof knob.default === "number" ? Number(fromQuery) : fromQuery;
}
});
render({ historyMode: "none" });
};
bindFamilyPopstate();
};
const init = () => {
document.querySelectorAll("[data-cookbook]").forEach(async (root) => {
initEvervault(document);
document.querySelectorAll("[data-cookbook][data-family]").forEach((root) => {
if (root.dataset.initialized) return;
root.dataset.initialized = "true";
const select = root.querySelector("[data-cookbook-recipe]");
const model = root.querySelector("[data-cookbook-model]");
const source = root.querySelector("[data-cookbook-source]");
const command = root.querySelector("[data-cookbook-command]");
const status = root.querySelector("[data-cookbook-status]");
try {
const { recipes } = await loadRecipes(root.dataset.recipes);
const byId = new Map(recipes.map((recipe) => [recipe.id, recipe]));
const groups = new Map();
select.replaceChildren();
recipes.forEach((recipe) => {
if (!groups.has(recipe.task)) {
const group = document.createElement("optgroup");
group.label = recipe.task;
groups.set(recipe.task, group);
select.append(group);
}
groups.get(recipe.task).append(new Option(recipe.label, recipe.id));
});
const render = () => {
const recipe = byId.get(select.value);
model.textContent = recipe.model;
source.textContent = recipe.source;
source.href = `https://github.com/hao-ai-lab/FastVideo/blob/main/${recipe.source}`;
command.textContent = recipe.command;
status.textContent = `${recipe.label} selected.`;
};
select.addEventListener("change", render);
select.disabled = false;
render();
} catch (error) {
status.textContent = "Recipes could not be loaded. Use the examples link below.";
console.error("Failed to load FastVideo cookbook recipes", error);
}
initFamilyBuilder(root);
});
};
+1674 -38
View File
File diff suppressed because it is too large Load Diff
+26
View File
@@ -0,0 +1,26 @@
# Cookbook logo sources
All marks are official publisher assets from Hugging Face organization pages
(vendored byte-for-byte from each org's public avatar). No imitation or
redrawn logos are used. Families without a publisher-appropriate mark use a
plain typographic tile instead — that tile is a UI placeholder, not a logo.
| Asset | Source | Purpose |
| --- | --- | --- |
| `wan-ai.webp` | [Official Wan-AI Hugging Face organization avatar](https://huggingface.co/Wan-AI) | Wan family card and page header |
| `ltx.webp` | [Official Lightricks Hugging Face organization avatar](https://huggingface.co/Lightricks) | LTX family card and page header |
| `tencent-hunyuan.webp` | [Official Tencent Hunyuan Hugging Face organization avatar](https://huggingface.co/Tencent-Hunyuan) | Hunyuan and GameCraft cards, Hunyuan page header |
| `nvidia.webp` | [Official NVIDIA Hugging Face organization avatar](https://huggingface.co/nvidia) | Cosmos and GEN3C cards, Cosmos page header |
| `kandinsky.webp` | [Official Kandinsky Lab Hugging Face organization avatar](https://huggingface.co/kandinskylab) | Kandinsky 5 family card and page header |
| `black-forest-labs.webp` | [Official Black Forest Labs Hugging Face organization avatar](https://huggingface.co/black-forest-labs) | FLUX family card and page header |
| `minimax.webp` | [Official MiniMax Hugging Face organization avatar](https://huggingface.co/MiniMaxAI) | MiniMax H3 family card and page header |
| `tongyi.webp` | [Official Tongyi MAI Hugging Face organization avatar](https://huggingface.co/Tongyi-MAI) | Z-Image family card and page header |
| `zai.webp` | [Official Z.ai Hugging Face organization avatar](https://huggingface.co/zai-org) | GLM-Image family card and page header |
| `stabilityai.webp` | [Official Stability AI Hugging Face organization avatar](https://huggingface.co/stabilityai) | Stable Diffusion and Stable Audio cards and page headers |
| `meituan-longcat.webp` | [Official Meituan LongCat Hugging Face organization avatar](https://huggingface.co/meituan-longcat) | LongCat family card and page header |
| `fastvideo.webp` | [Official FastVideo Hugging Face organization avatar](https://huggingface.co/FastVideo) | Matrix Game and MMAudio cards (converted weights published by this org), DreamX card |
Typographic tiles (no vendored image): TurboDiffusion ("Turbo"), HY-World
("HY"), LingBot ("LB"), MMAudio page header ("MMA"). These publishers have no
single official mark appropriate for reuse in the catalog; add a licensed
asset here if one becomes available.
Binary file not shown.

After

Width:  |  Height:  |  Size: 2.2 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.9 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.1 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 1.3 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 5.1 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 4.1 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 3.5 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.8 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 7.6 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 6.0 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 6.8 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 2.4 KiB

+41
View File
@@ -6,6 +6,47 @@ Sparse attention mechanism selecting top-k blocks.
VSA is included in the `fastvideo-kernel` package. See the [main Attention page](../index.md) for build instructions.
## Apple Silicon (MiniMax H3 / FastH3)
The native MLX runtime has an inference-only H3 VSA path that is separate from
the CUDA `fastvideo-kernel` package:
- **INT8 / INT6 / INT4** are **weight-only**. They cover linear matrices,
including `attn.to_gate_compress` when you convert with `--include-vsa`.
Attention Q/K/V stay BF16 (or the selected activation dtype). There is no
INT6 Q/K/V attention kernel.
- **Dense-only checkpoints** (the default converter) drop the 50 gate
matrices and keep fused SDPA. They remain valid for dense inference.
- **VSA-capable checkpoints** retain those gates, quantize them on the same
affine grid, and record `vsa.capable` in `mlx_h3_dit.json`. Runtime VSA is
still off until you pass `--vsa`.
- **Tile sizes** 64 `(4, 4, 4)` and 256 `(4, 8, 8)`. Prefix keys can be
`exempt` or `compete`. `--vsa-dense-first-n-steps` and `--vsa-dense-layers`
force dense SDPA on the selected steps or blocks.
- **`--vsa-impl auto`** uses the chunked gather+SDPA **reference** path.
`--vsa-impl simd` runs the SIMD-group Metal kernel (tile 64, head dim 128)
and falls back to reference on unsupported shapes or kernel failure. The
runtime executes a small kernel probe before use and remembers failures
for the process, so later blocks do not retry a broken backend.
It is not the default: 480p four-step generation is faster than reference
but does not yet match reference video. `--vsa-impl reference` is the same
as `auto`.
See the [Apple Silicon guide](../../getting_started/installation/mps.md) for
conversion and `mlx_fasth3.py` flags. Do not enable VSA on a dense-only
checkpoint; reconvert with `--include-vsa` first.
H3 uses fused MLX RMSNorm by default, including dense inference. This can
change BF16 rounding relative to the older explicit normalization path.
`FASTVIDEO_MLX_FAST_NORM` controls Wan normalization only.
The generation report aggregates VSA statistics across all blocks and steps.
`impl_counts` and `fallback_reasons` show mixed execution and fallback.
`video_keep` is the mean number of selected video-key tiles per video query
tile and head, including dense overrides. `achieved_sparsity` is the matching
mean video-tile sparsity; in `compete` mode it measures actual selections,
not the requested top-k budget. These are tile counts, not token-level FLOPs.
## Usage
```python
+208 -76
View File
@@ -6,7 +6,8 @@ lives in [Testing](testing.md).
## Overview
FastVideo splits validation across GitHub Actions, Buildkite, and Modal:
FastVideo splits validation across GitHub Actions, Buildkite, Slinky Slurm,
and Mergify:
```text
PR opened or updated
@@ -16,17 +17,24 @@ PR opened or updated
| style, lint, type, spelling, Markdown, workflow syntax, filenames
|
|-- Tier 2: Fastcheck
| Buildkite orchestrates Modal GPU jobs
| path-filtered component and unit checks
| Buildkite schedules six lanes on the Slinky Slurm cluster
| encoder, VAE, transformer, kernel, unit, DreamVerse
|
|-- /merge, /test full, or ready label
|-- /merge or ready label
|
`-- Tier 3: Full Suite
Buildkite orchestrates Modal GPU jobs
path-filtered integration, SSIM, training, eval, and performance checks
`-- Tier 3: Change-aware merge gate
trusted base-branch planner classifies the complete PR diff
Buildkite adds only relevant integration, quality, training,
API, or performance lanes on Slinky Slurm
|
pass -> Mergify squash-merges when all merge conditions pass
fail -> fix, push, and re-run
|-- /test full
`-- Explicit all-20-lane diagnostic run
`-- weekly schedule on main
`-- Complete four-GPU SSIM matrix
```
CI is not one monolithic job:
@@ -34,10 +42,31 @@ CI is not one monolithic job:
- GitHub Actions owns pre-commit, slash-command handling, aggregate status
updates, docs deployment, image builds, package publishing, and community
automations.
- Buildkite owns the GPU test pipeline and path filtering.
- Modal owns the actual GPU execution environment for test jobs.
- Buildkite owns the GPU test graph, statuses, and trusted dispatch control
plane. Its agent runs on the Slurm login plane; it does not execute test
payloads.
- Slinky Slurm is the only active CI compute backend. A host-owned dispatcher
leases GPUs from a persistent four-GPU allocation and runs each lane in an
isolated Enroot container at the immutable PR SHA.
- Mergify owns merge protection, labeling, and the final squash merge.
The old files under `fastvideo/tests/modal/` are retained as dormant manual
rollback code. `.buildkite/scripts/pr_test.sh` rejects Buildkite invocations,
and no pipeline or slash-command route calls Modal.
Three Buildkite entry pipelines share the validated graph:
| Pipeline | Trigger | Scope |
|---|---|---|
| `pr-fastcheck` | Automatic pull-request webhook | Six Fastcheck lanes |
| `ci` | `/merge`, `ready`, schedules, and `/test` API builds | Change-aware merge gates, scheduled SSIM, explicit Full Suite, Fastcheck reruns, or one direct lane |
| `fastvideo-performance-lane` | Weekly scheduler | Direct performance lane |
Each entry pipeline starts with the same trusted `pipeline-upload` job on the
`ci-runner` queue. The `ci` pipeline's incoming GitHub webhook is disabled;
otherwise it would duplicate the automatic `pr-fastcheck` build. API and
scheduled builds continue to work with webhook processing disabled.
## CI Tiers
### Tier 1: Pre-commit
@@ -70,34 +99,37 @@ debugging a hook implementation.
| Attribute | Value |
|---|---|
| Triggered by | Buildkite PR builds with `TEST_SCOPE=fastcheck` or unset |
| Runner | Buildkite agent that launches Modal GPU jobs |
| Compute | Slinky Slurm (`ci-runner` queue) |
| Definition | `.buildkite/pipeline.yml` |
| Entrypoint | `.buildkite/scripts/pr_test.sh` -> `fastvideo/tests/modal/pr_test.py` |
| Entrypoint | Trusted host driver -> `.buildkite/scripts/unit_test.sh` or `.buildkite/scripts/lanes/*.sh` |
Fastcheck uses Buildkite's `monorepo-diff` plugin. Jobs whose watched paths did
not change are skipped and do not block the aggregate `fastcheck-passed`
status.
Fastcheck always schedules these six lanes: encoder, VAE, transformer, custom
kernels, unit tests, and DreamVerse. Static steps replace the former
host-side path-filter plugin: the login plane never checks out or executes PR
code.
| Buildkite label | `TEST_TYPE` | Main watched paths |
|---|---|---|
| Encoder Tests | `encoder` | `fastvideo/models/encoders/**`, `fastvideo/models/loader/**`, `fastvideo/tests/encoders/**`, `pyproject.toml`, `docker/Dockerfile` |
| VAE Tests | `vae` | `fastvideo/models/vaes/**`, `fastvideo/models/loader/**`, `fastvideo/tests/vaes/**`, `pyproject.toml`, `docker/Dockerfile` |
| Transformer Tests | `transformer` | `fastvideo/models/dits/**`, `fastvideo/models/loader/**`, `fastvideo/tests/transformers/**`, `fastvideo/layers/**`, `fastvideo/attention/**`, `pyproject.toml`, `docker/Dockerfile` |
| Kernel Tests | `kernel_tests` | `fastvideo-kernel/**`, `pyproject.toml`, `docker/Dockerfile` |
| Unit Tests | `unit_test` | `fastvideo/**`, `.buildkite/**`, `.github/**`, `pyproject.toml`, `docker/Dockerfile` |
| DreamVerse App Tests | `dreamverse_app` | `apps/dreamverse/**`, `pyproject.toml` |
### Tier 3: Full Suite
### Tier 3: Change-Aware Merge Gate
| Attribute | Value |
|---|---|
| Triggered by | `/merge`, adding `ready`, `/test full`, or a new push to a PR that already has `ready` |
| Runner | Buildkite agent that launches Modal GPU jobs |
| Triggered by | `/merge`, adding `ready`, or a new push to a PR that already has `ready` |
| Compute | Slinky Slurm only (`ci-runner` queue) |
| Definition | `.buildkite/pipeline.yml` |
| Entrypoint | `.buildkite/scripts/pr_test.sh` -> `fastvideo/tests/modal/pr_test.py` |
| Entrypoint | `/opt/fastvideo-ci-runner/run-ci` (`run-unit` is a compatibility wrapper) |
Full Suite is also path-filtered. It validates broader behavior before Mergify
can merge a PR.
Fastcheck is the universal six-lane baseline. The merge gate does not repeat
those jobs: it classifies every changed path and adds only the relevant lanes
from the fourteen-lane integration set below. Selected jobs are hard gates;
there are no soft-fail hardware lanes. A documentation-only PR can therefore
finish its merge build after the trusted uploader, while model-family changes
typically add focused golden-gate and SSIM files and a training-only change
adds only its owning training lane.
`.github/scripts/plan_merge_ci.py` is the canonical path policy. It runs from
the immutable base SHA under `pull_request_target`; PR code is never executed
on the GitHub runner. The changed-file list includes both sides of renames. An
API failure, truncated response, empty list, unknown build input, or unknown
source path fails closed to all fourteen integration lanes.
A `ready`-labeled PR does not hit Buildkite immediately:
`ci-trigger-full-suite.yml` first runs `.github/scripts/gate_full_suite.sh`,
@@ -106,25 +138,100 @@ head. A red cheap check blocks the suite (fail closed; the next push re-arms
it), while a GitHub outage or a >25 min wait lets it run anyway (fail open).
`/test full` bypasses the gate.
| Buildkite label | `TEST_TYPE` | Main watched paths |
|---|---|---|
| SSIM Tests | `ssim` | `fastvideo/**/*.py`, `pyproject.toml`, `docker/Dockerfile` |
| LoRA Inference Tests | `inference_lora` | LoRA tests, loader, transformer tests, pipelines, LoRA layers |
| LoRA Extraction Tests | `lora_extraction` | LoRA extraction scripts/tests, loader, training utilities, LoRA layers |
| Training Tests | `training` | `fastvideo/**`, `pyproject.toml`, `docker/Dockerfile` |
| Distillation DMD Tests | `distillation_dmd` | `fastvideo/training/*distillation_pipeline.py` |
| Self-Forcing Tests | `self_forcing` | self-forcing distillation pipeline and tests |
| LoRA Training Tests | `training_lora` | `fastvideo/**`, `pyproject.toml`, `docker/Dockerfile` |
| Training Tests VSA | `training_vsa` | `fastvideo/**`, `fastvideo-kernel/**`, `pyproject.toml`, `docker/Dockerfile` |
| Inference Tests VMoBA | `inference_vmoba` | `fastvideo-kernel/**`, `fastvideo/attention/backends/vmoba.py` |
| Performance Tests | `performance` | DiTs, pipelines, attention, layers, worker, entrypoints, performance tests/configs |
| API Server Tests | `api_server` | OpenAI entrypoints, serve CLI, OpenAI API integration test |
| Train Framework Tests | `train_framework` | `fastvideo/train/**`, train model/method tests, model loader, DiTs |
| Eval Metrics Tests | `eval` | `fastvideo/eval/**`, `fastvideo/tests/eval/**`, `pyproject.toml`, `docker/Dockerfile` |
The complete static graph remains available through `/test full`; path
selection never deletes or dynamically invents a Buildkite step.
| Lane | Public `TEST_TYPE` | GPUs | Typical merge trigger |
|---|---|---:|---|
| Encoder | `encoder` | 1 | Universal Fastcheck |
| VAE | `vae` | 1 | Universal Fastcheck |
| Transformer | `transformer` | 1 | Universal Fastcheck |
| Kernel | `kernel_tests` | 1 | Universal Fastcheck |
| Unit | `unit_test` | 1 | Universal Fastcheck |
| DreamVerse | `dreamverse_app` | 1 | Universal Fastcheck |
| Golden gate | `golden_gate` | 1 | Model, pipeline, attention, layer, or output changes |
| SSIM | `ssim` | 4 | Matching model/SSIM paths; focused files when possible |
| LoRA inference | `inference_lora` | 1 | LoRA inference/shared LoRA paths |
| LoRA extraction | `lora_extraction` | 1 | LoRA extraction/shared LoRA paths |
| Vanilla training | `training` | 4 | Legacy vanilla/shared training paths |
| DMD distillation | `distillation_dmd` | 2 | DMD/shared training paths |
| Self-forcing | `self_forcing` | 2 | Self-forcing/shared training paths |
| LoRA training | `training_lora` | 2 | LoRA/shared training paths |
| VSA training | `training_vsa` | 2 | VSA/shared training paths |
| VMoBA inference | `inference_vmoba` | 1 | VMoBA backend/config paths |
| Performance | `performance` | 2 | Performance tests or benchmark policy |
| API server | `api_server` | 1 | API, worker, or server entrypoint paths |
| Modular train framework | `train_framework` | 1 | `fastvideo/train/` and its tests |
| Eval metrics | `eval` | 1 | `fastvideo/eval/` and its tests |
Golden-gate and SSIM selections are basenames, not arbitrary pytest arguments.
The private host checks the comma-separated allowlist before staging, and the
container checks it again before invoking pytest. Shared quality-harness
changes still run the complete owning lane. The full SSIM matrix also runs on
`main` every Sunday at 05:00 UTC through `ci-scheduled-ssim.yml`; `/test ssim`
and `/test full` remain available for deliberate complete runs.
The four Buildkite workers may accept multiple jobs concurrently. The
agent-owned lease broker packs their requested GPU counts onto one persistent
four-GPU Slurm allocation and waits when capacity is full. A four-GPU lane
such as SSIM or vanilla training owns the whole tray; two two-GPU lanes or up
to four one-GPU lanes can overlap without sharing devices. SSIM and vanilla
training also share the `fastvideo/slinky/whole-tray` Buildkite concurrency
group. That keeps the second whole-tray lane in Buildkite instead of consuming
an agent and its command timeout while the first lane waits for four free GPUs.
Because packed Enroot containers share the node network namespace, each GPU
lease receives its own 100-port rendezvous range. Tests preserve the
runner-assigned `MASTER_PORT`, and parallel SSIM tasks use distinct offsets
inside that range.
See [Performance Benchmarks](performance_benchmarks.md) for the performance
lane's thresholds, rolling baseline, artifacts, and reseeding process.
## `/merge` Request Flow
The Buildkite agent and Slurm have deliberately separate responsibilities:
```text
/merge PR comment
-> GitHub verifies write permission and refreshes the `ready` label
-> base-branch `ci-trigger-full-suite` workflow fetches the PR file list,
computes MERGE_TEST_PLAN plus focused golden/SSIM basenames, and gates on
cheap checks
-> sends the PR SHA to pipeline `ci` with TEST_SCOPE=merge and FULL_SUITE=true
-> trusted `pipeline-upload` job on queue `ci-runner`
fetch exact-SHA .buildkite/pipeline.yml
normalize + validate the complete static 20-lane policy and conditions
upload static Buildkite steps
-> each accepted step reaches the host policy hook
validate org/repo/SHA/ref/step key/command/scope/timeout
skip checkout on the login plane
stage a mode-0600 request on Lustre
lease 1-4 GPUs and attach an `srun` step to the Slinky tray
-> Enroot worker container
clone and verify the exact PR SHA
install the lane's project extras and cached kernel
run the repository-owned lane script
write numeric exit status and approved artifacts
-> trusted host returns that status to Buildkite
-> Buildkite publishes `full-suite-passed` only when every selected lane passes
```
The pipeline uploader and dispatcher run on the Slurm login plane, but those
are control-plane operations only. Python tests, model loading, CUDA kernels,
Node/Playwright checks, inference, training, SSIM generation, and performance
benchmarks all execute in Slurm allocations.
PR-controlled values never become host commands. The host policy accepts only
the pinned pipeline uploader or a known lane tuple. It rejects plugins,
artifact globs, shell injection variables, non-immutable commits, and unknown
commands before checkout. Hugging Face credentials are added only for lanes
that declare them, passed through a mode-0600 request file, and removed before
the PR payload starts. Active training lanes keep W&B offline and do not stage
a W&B credential. The ARM64 image includes the pinned FA4 CuTe overlay validated
on GB200. SSIM opts into FA4 to preserve its reference-video numerics; lanes
with FA2 baselines keep `FASTVIDEO_FA4=0`. Performance artifacts are relayed
afterward by the trusted host from an allowlisted directory and extension set.
## Slash Commands
Slash commands are handled by `.github/workflows/ci-slash-commands.yml`.
@@ -132,7 +239,7 @@ Repository write permission is required.
| Command | Effect |
|---|---|
| `/merge` | Adds `ready` and triggers Full Suite for the PR head branch. |
| `/merge` | Adds `ready` and triggers the path-aware merge gate for the PR head. |
| `/test full` | Runs the whole Full Suite with `TEST_SCOPE=full`. |
| `/test fastcheck` | Runs the whole Fastcheck suite with `TEST_SCOPE=fastcheck`. |
| `/test pre-commit` | Re-runs the pre-commit workflow on the PR merge ref. |
@@ -149,6 +256,7 @@ Valid direct test names:
| `/test unit` | `unit_test` |
| `/test dreamverse` | `dreamverse_app` |
| `/test ssim` | `ssim` |
| `/test golden-gate` | `golden_gate` |
| `/test training` | `training` |
| `/test lora-inference` | `inference_lora` |
| `/test lora-training` | `training_lora` |
@@ -162,12 +270,18 @@ Valid direct test names:
| `/test train-framework` | `train_framework` |
| `/test eval` | `eval` |
The temporary `<name>-ci` spellings remain accepted as compatibility aliases;
they select the same Slurm lane and do not identify a second backend.
When a direct test completes successfully, Buildkite posts
`direct-test-completed`. `.github/workflows/ci-aggregate-status.yml` then reads
the latest Buildkite statuses for the commit and updates `fastcheck-passed` or
`full-suite-passed` if all jobs in that group are green.
Skipped path-filtered jobs have no status entry and do not block the aggregate.
Buildkite label emojis define the status namespace used by that aggregation:
`:microscope:` is reserved for the six Fastcheck lanes, while Full-Suite-only
lanes use `:test_tube:` or `:bar_chart:`. Each active lane has exactly one
label and therefore one status context.
## Merge Protection
@@ -177,7 +291,7 @@ Mergify enforces these conditions before it squash-merges to `main`:
|---|---|
| `check-success~=pre-commit` | Tier 1 passed. |
| `check-success=fastcheck-passed` | All triggered Fastcheck jobs passed. |
| `check-success=full-suite-passed` | All triggered Full Suite jobs passed. |
| `check-success=full-suite-passed` | The selected merge gate or explicit Full Suite passed. |
| `#approved-reviews-by>=1` | At least one approving review. |
| Valid title regex | PR title starts with an accepted `[type]` tag. |
| `label=ready` | The PR has entered the merge flow. |
@@ -226,29 +340,40 @@ Process labels:
| Label | Who sets it | Meaning |
|---|---|---|
| `ready` | `/merge` or maintainer action | Triggers/keeps Full Suite active and enables auto-merge. |
| `ready` | `/merge` or maintainer action | Triggers/keeps the change-aware merge gate active and enables auto-merge. |
| `needs-rebase` | Mergify | PR has merge conflicts. |
| `do-not-merge` | Maintainer | Blocks merge regardless of CI status. |
## Modal Test Entrypoints
## Slurm Lane Entrypoints
All Buildkite test jobs go through `.buildkite/scripts/pr_test.sh`, which:
Every active test selection lives in `.buildkite/scripts/unit_test.sh` or a
focused `.buildkite/scripts/lanes/<lane>.sh`. The private, agent-owned lane
table binds each internal `*_ci` type to that script, its GPU count, wall-clock
limit, dependency extras, kernel-build policy, secrets, and artifacts. The
internal suffix is an implementation detail; there is only one active backend.
1. Reads Buildkite secrets for Modal, Hugging Face, and W&B when needed.
2. Selects a Modal function based on `TEST_TYPE`.
3. Passes Buildkite metadata into the Modal container.
4. Runs the selected test command from `fastvideo/tests/modal/pr_test.py` or
`fastvideo/tests/modal/ssim_test.py`.
5. Uploads performance artifacts for `TEST_TYPE=performance`.
SSIM uses `fastvideo/tests/ssim/ci_runner.py` inside a single four-GPU lease.
It discovers `REQUIRED_GPUS` and `*_MODEL_TO_PARAMS` with AST parsing, then
packs independent pytest subprocesses across the visible GPUs with fail-fast
termination. Performance writes reports to a host-mounted artifact directory;
the host uploads only `.md`, `.html`, `.json`, and `.csv` files after the
container exits.
The Modal launchers remain in the repository for manual rollback archaeology,
but they are not CI entrypoints. `pr_test.sh` rejects Buildkite calls and needs
`FASTVIDEO_ENABLE_LEGACY_MODAL_CI=1` even for a local manual invocation.
If you add a new CI test category:
1. Add the Modal function in `fastvideo/tests/modal/pr_test.py` or a focused
companion module.
2. Add the `TEST_TYPE` case in `.buildkite/scripts/pr_test.sh`.
3. Add the Buildkite direct-test step and any Fastcheck/Full Suite path filters
in `.buildkite/pipeline.yml`.
4. Add or update the `/test` mapping in `.github/workflows/ci-slash-commands.yml`.
1. Add an executable `.buildkite/scripts/lanes/<lane>.sh` containing the test
payload only.
2. Add the static Buildkite step in `.buildkite/pipeline.yml`, its source/test
ownership in `.github/scripts/plan_merge_ci.py`, and the `/test` mapping in
`.github/workflows/ci-slash-commands.yml`.
3. Extend `fastvideo/tests/contract/test_ci_test_collection.py`,
`test_merge_ci_plan.py`, and the trusted private runner's lane table plus
uploader policy in the same rollout.
4. Validate on the target GB200 hardware before making the lane a merge gate.
5. Document the lane here and link any domain-specific authoring guide.
## CD And Release Workflows
@@ -282,13 +407,18 @@ the explicit `py3.12-cuda13.0.0-latest` tag. This publication policy does not
change the unparameterized `docker/Dockerfile` build defaults, which remain CUDA
13 and `cu130`.
Published amd64 development images keep their configured Hopper kernel wheel
installed and also carry an immutable SM89 wheel under
`/opt/fastvideo-kernel-prebuilt`. Modal PR and SSIM jobs select the exact
source, ABI, and GPU-architecture match from that directory, so L40S jobs reuse
the trusted image artifact while kernel-changing PRs still build locally. Once
a kernel or artifact-key change reaches `main`, the image workflow republishes
the matching trusted artifact before later jobs consume the updated image tag.
Published development images carry architecture-specific kernel wheels under
`/opt/fastvideo-kernel-prebuilt`. The Slurm worker selects the exact source,
ABI, and GPU-architecture match, so normal lanes reuse the trusted artifact
while kernel-changing PRs still build locally. Once a kernel or artifact-key
change reaches `main`, the image workflow republishes the matching artifact
before later jobs consume the updated image pin.
The same workflow publishes a single-architecture ARM64, CUDA 13, SM100 image
for the self-hosted CI runner under the
`py3.12-cuda13.0.0-sm100-{latest,sha-*}` tags. It carries the matching prebuilt
kernel wheel so runner jobs can validate and install the exact source and ABI
match instead of recompiling it in every lane.
The optional Dreamverse matrix builds backend and UI images for CUDA 12.6 and
CUDA 13 on `amd64`. Dreamverse remains `amd64`-only because its FA4 dependency
@@ -320,15 +450,17 @@ The reusable implementation lives in
| `.github/mergify.yml` | Merge protection, PR title validation, PR labels, conflict labels, auto-merge |
| `.github/workflows/ci-precommit.yml` | Tier 1 pre-commit |
| `.github/workflows/ci-slash-commands.yml` | `/merge` and `/test` handling |
| `.github/workflows/ci-trigger-full-suite.yml` | Full Suite trigger for `ready` PRs and new pushes to ready PRs |
| `.github/workflows/ci-aggregate-status.yml` | Aggregate Fastcheck/Full Suite commit statuses |
| `.buildkite/pipeline.yml` | Buildkite test graph and path filters |
| `.buildkite/scripts/pr_test.sh` | Buildkite-to-Modal test dispatcher |
| `fastvideo/tests/modal/pr_test.py` | Modal functions for most GPU CI lanes |
| `fastvideo/tests/modal/ssim_test.py` | Modal functions and partitioning for SSIM |
| `.github/workflows/ci-trigger-full-suite.yml` | Change-aware merge-gate trigger for `ready` PRs and new pushes |
| `.github/workflows/ci-scheduled-ssim.yml` | Weekly complete SSIM trigger on `main` |
| `.github/scripts/plan_merge_ci.py` | Trusted changed-path to integration-lane planner |
| `.github/workflows/ci-aggregate-status.yml` | Aggregate Fastcheck and explicit Full Suite direct-rerun statuses |
| `.buildkite/pipeline.yml` | Static 20-lane Slurm Buildkite graph |
| `.buildkite/scripts/unit_test.sh`, `.buildkite/scripts/lanes/*.sh` | Active Slurm lane payloads |
| `fastvideo/tests/ssim/ci_runner.py` | Four-GPU Slurm SSIM scheduler |
| `.buildkite/scripts/pr_test.sh`, `fastvideo/tests/modal/*.py` | Dormant manual Modal rollback path (disabled in Buildkite) |
| `.buildkite/performance-benchmarks/tests/*.json` | Performance benchmark configs and thresholds |
| `.github/workflows/infra-docs.yml` | Docs build and GitHub Pages deploy |
| `.github/workflows/infra-build-image.yml` | Automatic CUDA matrix and manual Docker image builds |
| `.github/workflows/infra-build-image.yml` | CUDA matrix, CI runner image, and manual Docker image builds |
| `.github/workflows/publish-fastvideo.yml` | FastVideo PyPI publishing |
| `.github/workflows/publish-kernel.yml` | FastVideo kernel PyPI publishing |
| `.github/workflows/publish-comfyui.yml` | ComfyUI registry publishing |
+10 -8
View File
@@ -23,8 +23,8 @@ It serves three audiences:
pytest fastvideo/tests/performance/ -vs
# Optional: compare against the rolling HF baseline.
# PERF_REPORTS_DIR defaults to /root/data/perf_reports for Modal/CI, so
# override it when running outside the container.
# PERF_REPORTS_DIR defaults to /root/data/perf_reports in a CI container, so
# override it for a local run.
PERF_REPORTS_DIR=/tmp/fastvideo_perf_reports \
python fastvideo/tests/performance/compare_baseline.py
@@ -459,16 +459,18 @@ FlashInfer, Cutlass DSL, SageAttention, Triton, and xFormers when installed.
| `FASTVIDEO_FA4` | `0` | `test_inference_performance.py` | FlashAttention-4 toggle included in `software_profile_id`. |
| `FASTVIDEO_PERFORMANCE_PROFILE_VERSION` | unset | `test_inference_performance.py` | Optional explicit software cohort/profile version included in `software_profile_id`. |
| `IMAGE_VERSION` | unset | `test_inference_performance.py` | CI container image/profile version included in `software_profile_id` when available. |
| `FASTVIDEO_CONTAINER_IMAGE_REF` | unset | `pr_test.py`, `launch_l40s_job.py`, `test_inference_performance.py` | Resolved CI container image ref or digest recorded in `environment_metadata` for audit without changing `software_profile_id`. |
| `FASTVIDEO_CONTAINER_IMAGE_REF` | unset | Slurm runner, `test_inference_performance.py` | Pinned CI container image digest recorded in `environment_metadata` and `software_profile_id`. |
| `FASTVIDEO_STAGE_LOGGING` | set by the pytest test | `test_inference_performance.py` | Enables pipeline stage timing capture for component metrics during benchmark runs. |
## CI integration
The performance step can run on demand with `/test performance` and as part of
the Full Suite (see [CI/CD Architecture](ci_architecture.md)). The Modal entry
point is `fastvideo/tests/modal/pr_test.py:run_performance_tests` and the
Buildkite artifact upload is in
`.buildkite/scripts/pr_test.sh:upload_performance_artifacts`.
The performance step can run on demand with `/test performance`, through a
merge gate when performance tests or benchmark policy changed, and as part of
an explicit `/test full` run (see [CI/CD Architecture](ci_architecture.md)).
The weekly `fastvideo-performance-lane` schedule runs the same Slurm payload.
The active entry point is `.buildkite/scripts/lanes/performance.sh`; the
trusted host dispatcher relays its allowlisted reports to Buildkite after the
isolated container exits.
Each performance build runs pytest first. PR and direct runs only continue to
`compare_baseline.py` when that fixed-threshold phase passes; if pytest fails,
+10 -8
View File
@@ -42,19 +42,20 @@ The important process labels are:
| Label | Meaning |
|---|---|
| `ready` | The PR is ready for Full Suite and auto-merge consideration. |
| `ready` | The PR is ready for the change-aware merge gate and auto-merge consideration. |
| `needs-rebase` | The PR has merge conflicts with `main`. |
| `do-not-merge` | A maintainer has blocked merge. |
## CI Summary
FastVideo has three validation tiers:
FastVideo has three routine validation tiers plus an explicit full diagnostic:
| Tier | Runs when | What it does |
|---|---|---|
| Pre-commit | Pull requests and `/test pre-commit` | Formatting, linting, typing, spelling, Markdown, workflow syntax, filename checks |
| Fastcheck | PR Buildkite builds | Path-filtered component and unit checks on Modal GPU runners |
| Full Suite | `/merge`, `ready`, `/test full`, or new pushes to ready PRs | Path-filtered integration, SSIM, training, eval, API, and performance checks |
| Fastcheck | PR Buildkite builds | Six component, kernel, unit, and app lanes on Slinky Slurm |
| Merge gate | `/merge`, `ready`, or new pushes to ready PRs | Only path-relevant integration lanes on Slinky Slurm; Fastcheck remains the baseline |
| Full Suite | `/test full` | Explicit all-twenty-lane diagnostic run on Slinky Slurm |
See [CI/CD Architecture](ci_architecture.md#ci-tiers) for exact jobs, path
filters, and workflow files.
@@ -66,10 +67,11 @@ filters, and workflow files.
3. Fix pre-commit failures locally with `pre-commit run --all-files`.
4. Wait for at least one approving review.
5. When the PR is approved and ready, comment `/merge`.
6. `/merge` adds `ready` and triggers the Full Suite for the PR branch.
6. `/merge` adds `ready`, waits for cheap checks, and triggers the minimal
path-relevant integration lanes for the PR branch.
7. If all required checks pass, Mergify squash-merges the PR to `main`.
8. If Full Suite fails, fix the regression, push again, and re-run `/merge` or
the failed test.
8. If the merge gate fails, fix the regression, push again, and re-run
`/merge`. Use a targeted `/test` command for diagnosis.
Only contributors with repository write permission can use slash commands. If
you are an external contributor, ask a maintainer to run `/merge` or add
@@ -128,7 +130,7 @@ git push --force-with-lease
Mergify removes `needs-rebase` after conflicts are resolved.
### Full Suite Fails
### Merge Gate Or Full Suite Fails
The failing Buildkite step is the source of truth. Common causes are:
+34 -48
View File
@@ -138,47 +138,31 @@ pytest fastvideo/tests/ssim/ -vs
Use a machine whose GPU and backend match the reference folder you are testing.
## Modal Runs For SSIM
## Slurm CI Runs For SSIM
For CI-like SSIM execution, use `fastvideo/tests/modal/ssim_test.py`:
Comment `/test ssim` on a pull request to run the canonical four-GPU SSIM
lane on the Slinky Slurm cluster. `fastvideo/tests/ssim/ci_runner.py`
discovers the suite without importing test modules, packs independent pytest
processes across the four assigned GPUs, and stops the lane on the first
failure.
The change-aware `/merge` planner may run only the SSIM test basenames owned
by the changed model family. Shared SSIM harness changes still select the
complete lane. Independently, `main` runs the full SSIM matrix every Sunday at
05:00 UTC so infrequently touched model families retain periodic coverage.
For a focused developer run, invoke pytest directly and optionally select one
model from a parameterized test through `FASTVIDEO_SSIM_MODEL_ID`:
```bash
python -m modal run fastvideo/tests/modal/ssim_test.py::run_ssim_tests
pytest fastvideo/tests/ssim/test_wan_t2v_similarity.py -vs
FASTVIDEO_SSIM_MODEL_ID=Wan2.1-T2V-1.3B-Diffusers \
pytest fastvideo/tests/ssim/test_wan_t2v_similarity.py -vs
```
Target specific files or model ids:
```bash
python -m modal run fastvideo/tests/modal/ssim_test.py::run_ssim_tests \
--test-files test_wan_t2v_similarity.py \
--model-ids Wan2.1-T2V-1.3B-Diffusers
```
If `HF_API_KEY`, `HUGGINGFACE_HUB_TOKEN`, or `HF_TOKEN` is not set, the local
entrypoint fails fast.
To export raw generated videos from Modal to the shared volume:
```bash
python -m modal run fastvideo/tests/modal/ssim_test.py::run_ssim_tests \
--sync-generated-to-volume
```
The raw export path is quality-tiered:
- default params: `ssim_generated_videos/default/<subdir>/generated_videos`
- full-quality params: `ssim_generated_videos/full_quality/<subdir>/generated_videos`
The printed `modal volume get` command downloads into
`./generated_videos_modal/<quality-tier>`. Convert those outputs into local
references with `copy-local`:
```bash
python fastvideo/tests/ssim/reference_videos_cli.py copy-local \
--quality-tier full_quality \
--generated-dir ./generated_videos_modal/full_quality/L40S_reference_videos \
--device-folder L40S_reference_videos
```
The files under `fastvideo/tests/modal/` are retained only as a disabled
manual rollback implementation. No active CI trigger invokes them.
### SSIM Bootstrap Mode
@@ -206,28 +190,30 @@ python fastvideo/tests/ssim/reference_videos_cli.py promote-draft \
## CI Integration
FastVideo CI tests are orchestrated by Buildkite and run on Modal GPU
instances. The main files are:
FastVideo GPU CI is orchestrated by Buildkite and runs only on isolated Slinky
Slurm workers. The main files are:
| File | Purpose |
|---|---|
| `.buildkite/pipeline.yml` | Buildkite test graph and path filters. |
| `.buildkite/scripts/pr_test.sh` | Dispatches `TEST_TYPE` to a Modal function. |
| `fastvideo/tests/modal/pr_test.py` | Modal functions for most test lanes. |
| `fastvideo/tests/modal/ssim_test.py` | Modal functions and partitioning for SSIM. |
| `.buildkite/pipeline.yml` | Static, validated 20-lane Slurm test graph. |
| `.github/scripts/plan_merge_ci.py` | Trusted path-to-lane and focused quality-test policy for `/merge`. |
| `.buildkite/scripts/unit_test.sh`, `.buildkite/scripts/lanes/*.sh` | Repository-owned test payloads executed inside Slurm containers. |
| `fastvideo/tests/ssim/ci_runner.py` | Four-GPU SSIM task discovery and scheduling. |
| `.buildkite/scripts/pr_test.sh`, `fastvideo/tests/modal/*.py` | Dormant manual rollback path; rejected in Buildkite. |
For exact tier membership, path filters, slash commands, and aggregate statuses,
For exact tier membership, slash commands, runner isolation, and aggregate statuses,
see [CI/CD Architecture](ci_architecture.md).
### Adding A New CI Test Category
If a new test does not fit an existing lane:
1. Add a Modal function in `fastvideo/tests/modal/pr_test.py` or a focused
companion module.
2. Add a matching `TEST_TYPE` case in `.buildkite/scripts/pr_test.sh`.
3. Add Buildkite direct-test and path-filter entries in `.buildkite/pipeline.yml`.
4. Add the `/test` mapping in `.github/workflows/ci-slash-commands.yml`.
1. Put the test payload in an executable `.buildkite/scripts/lanes/<lane>.sh`.
2. Add its static step to `.buildkite/pipeline.yml`, its changed-path ownership
to `.github/scripts/plan_merge_ci.py`, and extend the CI contract tests.
3. Add the `/test` mapping in `.github/workflows/ci-slash-commands.yml`.
4. Coordinate the matching GPU, timeout, dependency, secret, and artifact
policy in the private Slurm runner allowlist.
5. Document the new category in [CI/CD Architecture](ci_architecture.md) and add
authoring notes here if contributors need them.
+140
View File
@@ -0,0 +1,140 @@
---
hide:
- toc
---
# Cosmos recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="cosmos" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/nvidia.webp" alt="NVIDIA" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>Cosmos inference recipes</h2>
<p>NVIDIA Cosmos Predict 2.5 generates navigable world videos. The maintained example runs the 2B text-to-world checkpoint on a single GPU.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading Cosmos recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/nvidia">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>World-generation prompts work best describing a scene and camera motion; the built-in prompt in the example is a known-good starting point.</li>
<li>Gated or missing checkpoints: run <code>huggingface-cli login</code> and confirm you accepted the model's license on Hugging Face.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>All recipes on this page are <strong>Source-backed</strong>: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are <strong>Unknown</strong> and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.</p>
</div>
</details>
+141
View File
@@ -0,0 +1,141 @@
---
hide:
- toc
---
# FLUX recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="flux" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/black-forest-labs.webp" alt="Black Forest Labs" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>FLUX inference recipes</h2>
<p>Black Forest Labs' FLUX family covers FLUX.1 dev and FLUX.2 (dev and distilled Klein) text-to-image. FLUX.1 registers no `model_family` in the registry and is grouped under FLUX for documentation only.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading FLUX recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/black-forest-labs">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>FLUX.1 dev defaults to a local <code>official_weights/FLUX.1-dev</code> directory in the example; the cookbook command passes the Hugging Face ID explicitly instead.</li>
<li>Image outputs land under <code>outputs/</code>; adjust <code>--output</code> if that path is not writable.</li>
<li>Gated checkpoints (FLUX.1 dev): run <code>huggingface-cli login</code> and accept the license on Hugging Face first.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>All recipes on this page are <strong>Source-backed</strong>: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are <strong>Unknown</strong> and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.</p>
</div>
</details>
+140
View File
@@ -0,0 +1,140 @@
---
hide:
- toc
---
# GLM-Image recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="glm_image" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/zai.webp" alt="Z.ai" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>GLM-Image inference recipes</h2>
<p>GLM-Image from Z.ai supports both text-to-image generation and instruction-based image editing, each with a maintained example.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading GLM-Image recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/zai-org">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>The editing example reads <code>assets/images/couple.jpg</code> from the repository root, so run it from a repo checkout rather than an arbitrary working directory.</li>
<li>Output paths default under <code>image_output/</code>; pass <code>--output</code> to change them.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>All recipes on this page are <strong>Source-backed</strong>: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are <strong>Unknown</strong> and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.</p>
</div>
</details>
+141
View File
@@ -0,0 +1,141 @@
---
hide:
- toc
---
# Hunyuan recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="hunyuan" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/tencent-hunyuan.webp" alt="Tencent Hunyuan" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>Hunyuan inference recipes</h2>
<p>HunyuanVideo 1.5 is Tencent's video generation family. The maintained examples cover 480p text-to-video with CPU offload and a full 480p-to-1080p upscale chain.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading Hunyuan recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/Tencent-Hunyuan">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>Out of memory: <code>basic_hy15.py</code> already enables dit/VAE/text-encoder CPU offload; further tradeoffs are described in <a href="../../inference/optimizations/">Optimizations</a>.</li>
<li><code>pin_cpu_memory</code> errors on low-RAM machines are documented inline in the example; set it to false as the source comment suggests.</li>
<li>Gated or missing checkpoints: run <code>huggingface-cli login</code> and confirm you accepted the model's license on Hugging Face.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>All recipes on this page are <strong>Source-backed</strong>: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are <strong>Unknown</strong> and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.</p>
</div>
</details>
+457 -37
View File
@@ -1,44 +1,464 @@
---
hide:
- toc
---
# Inference Cookbook
Choose a complete recipe maintained in the FastVideo repository. Each command
runs its checked-in source directly, so coupled model, GPU, offload, and
attention settings do not drift into unsupported combinations.
<div class="cookbook-shell cookbook-catalog" data-cookbook-gallery>
<header class="cookbook-hero">
<p class="cookbook-eyebrow">FastVideo inference cookbook</p>
<h2>Choose by output, then by family.</h2>
<p class="cookbook-hero__lede">
Open a family to pick a maintained recipe and a runtime FastVideo
actually supports. Every command runs a checked-in source, so the model,
platform, offload, and attention settings stay tied to that example.
Cards below group video, image, audio, and interactive world models.
Mode chips on each card are the workloads FastVideo maintains for that
family, not a promise that every flag works on every runtime.
</p>
<a class="cookbook-inline-link" href="../inference/support_matrix/">
View the full support matrix <span aria-hidden="true">→</span>
</a>
<a class="cookbook-inline-link" href="./openai-api/">Run FastH3 with a playground and API <span aria-hidden="true">→</span></a>
</header>
The commands expect a local clone:
<section class="cookbook-section" id="video-models" aria-labelledby="video-models-heading">
<div class="cookbook-section__heading">
<h2 id="video-models-heading">Video</h2>
</div>
<div class="cookbook-family-grid">
<a class="cookbook-family-tile cookbook-family-tile--ready cookbook-family-tile--featured" href="./minimax-h3/" aria-label="Open MiniMax H3 recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/minimax.webp" alt="" width="132" height="132" loading="eager">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>MiniMax H3</strong><small>Video + stereo audio</small></span>
<span class="cookbook-count">8 recipes</span>
</span>
<ul class="cookbook-mode-row">
<li>T2VA</li>
<li>FL2VA</li>
<li>Ref2VA</li>
<li>MLX T2VA</li>
<li>DGX Spark</li>
</ul>
</span>
</a>
```bash
git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo
```
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./wan/" aria-label="Open Wan recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/wan-ai.webp" alt="" width="132" height="132" loading="eager">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>Wan</strong><small>FastWan CUDA and FastMetal MLX</small></span>
<span class="cookbook-count">7 recipes</span>
</span>
<ul class="cookbook-mode-row">
<li>T2V</li>
<li>I2V</li>
<li>TI2V</li>
<li>MLX T2V</li>
</ul>
</span>
</a>
<div class="cookbook-picker" data-cookbook data-recipes="../assets/cookbook-recipes.json">
<label for="cookbook-recipe"><strong>Recipe</strong></label>
<select id="cookbook-recipe" data-cookbook-recipe disabled>
<option>Loading recipes…</option>
</select>
<dl>
<dt>Model</dt>
<dd data-cookbook-model>Loading…</dd>
<dt>Source</dt>
<dd><a data-cookbook-source href="../inference/examples/basic/">Browse maintained examples</a></dd>
</dl>
<pre><code class="language-bash" data-cookbook-command>Loading…</code></pre>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
<noscript>
JavaScript is needed for the recipe picker. Browse the
<a href="../inference/examples/examples_inference_index/">inference examples</a>
instead.
</noscript>
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./ltx/" aria-label="Open LTX recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/ltx.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>LTX</strong><small>Video with synchronized audio</small></span>
<span class="cookbook-count">2 recipes</span>
</span>
<ul class="cookbook-mode-row">
<li>T2V</li>
</ul>
</span>
</a>
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./hunyuan/" aria-label="Open Hunyuan recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/tencent-hunyuan.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>Hunyuan</strong><small>480p and 1080p upscale</small></span>
<span class="cookbook-count">2 recipes</span>
</span>
<ul class="cookbook-mode-row">
<li>T2V</li>
</ul>
</span>
</a>
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./kandinsky5/" aria-label="Open Kandinsky 5 recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/kandinsky.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>Kandinsky 5</strong><small>Text and image to video</small></span>
<span class="cookbook-count">2 recipes</span>
</span>
<ul class="cookbook-mode-row">
<li>T2V</li>
<li>I2V</li>
</ul>
</span>
</a>
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./longcat/" aria-label="Open LongCat recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/meituan-longcat.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>LongCat</strong><small>T2V, I2V, optional refine</small></span>
<span class="cookbook-count">2 recipes</span>
</span>
<ul class="cookbook-mode-row">
<li>T2V</li>
<li>I2V</li>
</ul>
</span>
</a>
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./turbodiffusion/" aria-label="Open TurboDiffusion recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<span class="cookbook-family-tile__monogram" aria-hidden="true">Turbo</span>
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>TurboDiffusion</strong><small>Accelerated Wan profiles</small></span>
<span class="cookbook-count">3 recipes</span>
</span>
<ul class="cookbook-mode-row">
<li>T2V</li>
<li>I2V</li>
</ul>
</span>
</a>
</div>
</section>
<section class="cookbook-section" id="image-models" aria-labelledby="image-models-heading">
<div class="cookbook-section__heading">
<h2 id="image-models-heading">Image</h2>
</div>
<div class="cookbook-family-grid">
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./flux/" aria-label="Open FLUX recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/black-forest-labs.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>FLUX</strong><small>FLUX.1 and FLUX.2</small></span>
<span class="cookbook-count">3 recipes</span>
</span>
<ul class="cookbook-mode-row">
<li>T2I</li>
</ul>
</span>
</a>
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./glm-image/" aria-label="Open GLM-Image recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/zai.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>GLM-Image</strong><small>Generate and edit</small></span>
<span class="cookbook-count">2 recipes</span>
</span>
<ul class="cookbook-mode-row">
<li>T2I</li>
<li>Edit</li>
</ul>
</span>
</a>
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./z-image/" aria-label="Open Z-Image recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/tongyi.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>Z-Image</strong><small>Turbo text to image</small></span>
<span class="cookbook-count">1 recipe</span>
</span>
<ul class="cookbook-mode-row">
<li>T2I</li>
</ul>
</span>
</a>
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./stable-diffusion/" aria-label="Open Stable Diffusion recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/stabilityai.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>Stable Diffusion</strong><small>SD 3.5 Medium</small></span>
<span class="cookbook-count">1 recipe</span>
</span>
<ul class="cookbook-mode-row">
<li>T2I</li>
</ul>
</span>
</a>
</div>
</section>
<section class="cookbook-section" id="audio-models" aria-labelledby="audio-models-heading">
<div class="cookbook-section__heading">
<h2 id="audio-models-heading">Audio</h2>
</div>
<div class="cookbook-family-grid">
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./stable-audio/" aria-label="Open Stable Audio recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/stabilityai.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>Stable Audio</strong><small>Open 1.0 and Small</small></span>
<span class="cookbook-count">2 recipes</span>
</span>
<ul class="cookbook-mode-row">
<li>T2A</li>
</ul>
</span>
</a>
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./mmaudio/" aria-label="Open MMAudio recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/fastvideo.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>MMAudio</strong><small>Video or text to audio</small></span>
<span class="cookbook-count">1 recipe</span>
</span>
<ul class="cookbook-mode-row">
<li>V2A</li>
<li>T2A</li>
</ul>
</span>
</a>
</div>
</section>
<section class="cookbook-section" id="world-models" aria-labelledby="world-models-heading">
<div class="cookbook-section__heading">
<h2 id="world-models-heading">World and interactive</h2>
</div>
<div class="cookbook-family-grid">
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./cosmos/" aria-label="Open Cosmos recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/nvidia.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>Cosmos</strong><small>Text to world video</small></span>
<span class="cookbook-count">1 recipe</span>
</span>
<ul class="cookbook-mode-row">
<li>T2W</li>
</ul>
</span>
</a>
<a class="cookbook-family-tile cookbook-family-tile--ready" href="./matrix-game/" aria-label="Open Matrix Game recipes">
<span class="cookbook-family-tile__visual" data-evervault>
<span class="cookbook-evervault" aria-hidden="true">
<span class="cookbook-evervault__gradient"></span>
<span class="cookbook-evervault__noise" data-cookbook-pattern></span>
</span>
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/fastvideo.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>Matrix Game</strong><small>Image-conditioned worlds</small></span>
<span class="cookbook-count">2 recipes</span>
</span>
<ul class="cookbook-mode-row">
<li>I2W</li>
</ul>
</span>
</a>
</div>
</section>
<section class="cookbook-section" id="planned-families" aria-labelledby="planned-families-heading">
<div class="cookbook-section__heading">
<h2 id="planned-families-heading">Pages still to write</h2>
<p>These families already have runnable examples. The cookbook page is not ready, so the cards are not links.</p>
</div>
<div class="cookbook-family-grid">
<article class="cookbook-family-tile cookbook-family-tile--coming" aria-label="GameCraft cookbook page planned; runnable examples exist">
<span class="cookbook-family-tile__visual">
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/tencent-hunyuan.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>GameCraft</strong><small>Game world generation</small></span>
<span class="cookbook-count">Page planned</span>
</span>
</span>
</article>
<article class="cookbook-family-tile cookbook-family-tile--coming" aria-label="GEN3C cookbook page planned; runnable examples exist">
<span class="cookbook-family-tile__visual">
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/nvidia.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>GEN3C</strong><small>Novel-view video</small></span>
<span class="cookbook-count">Page planned</span>
</span>
</span>
</article>
<article class="cookbook-family-tile cookbook-family-tile--coming" aria-label="HY-World cookbook page planned; runnable examples exist">
<span class="cookbook-family-tile__visual">
<span class="cookbook-family-tile__logo-wrap">
<span class="cookbook-family-tile__monogram" aria-hidden="true">HY</span>
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>HY-World</strong><small>Interactive world play</small></span>
<span class="cookbook-count">Page planned</span>
</span>
</span>
</article>
<article class="cookbook-family-tile cookbook-family-tile--coming" aria-label="DreamX cookbook page planned; runnable examples exist">
<span class="cookbook-family-tile__visual">
<span class="cookbook-family-tile__logo-wrap">
<img class="off-glb" src="../assets/logos/fastvideo.webp" alt="" width="132" height="132" loading="lazy">
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>DreamX</strong><small>World generation</small></span>
<span class="cookbook-count">Page planned</span>
</span>
</span>
</article>
<article class="cookbook-family-tile cookbook-family-tile--coming" aria-label="LingBot cookbook page planned; runnable examples exist">
<span class="cookbook-family-tile__visual">
<span class="cookbook-family-tile__logo-wrap">
<span class="cookbook-family-tile__monogram" aria-hidden="true">LB</span>
</span>
</span>
<span class="cookbook-family-tile__footer">
<span class="cookbook-family-tile__footer-top">
<span><strong>LingBot</strong><small>Video and world models</small></span>
<span class="cookbook-count">Page planned</span>
</span>
</span>
</article>
</div>
</section>
</div>
## Customize a recipe
Start from the checked-in source, then change only the settings your model
supports:
- [Configuration](../inference/configuration.md) covers the Python and CLI
config surfaces.
- [Optimizations](../inference/optimizations.md) covers attention backends,
compilation, and memory tradeoffs.
- [Support matrix](../inference/support_matrix.md) lists supported models and
optimizations.
<small class="cookbook-logo-credit">
Catalog marks come from the official model publishers' Hugging Face
organizations; typographic tiles are placeholders, never invented logos. See
<a href="https://github.com/hao-ai-lab/FastVideo/blob/main/docs/assets/logos/SOURCES.md">docs/assets/logos/SOURCES.md</a>.
</small>
+140
View File
@@ -0,0 +1,140 @@
---
hide:
- toc
---
# Kandinsky 5 recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="kandinsky5" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/kandinsky.webp" alt="Kandinsky Lab" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>Kandinsky 5 inference recipes</h2>
<p>Kandinsky 5.0 from the Kandinsky Lab covers text-to-video and image-to-video in Lite and Pro variants, including distilled checkpoints for faster sampling.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading Kandinsky 5 recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/kandinskylab">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>Alternative Lite/Pro checkpoints are listed inline in the maintained examples; swap the model string only after checking its Hugging Face card.</li>
<li>Gated or missing checkpoints: run <code>huggingface-cli login</code> and confirm you accepted the model's license on Hugging Face.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>Both recipes map to checked-in FastVideo examples and recorded single-GPU B200 runs. The image-to-video run also records 10,365.89 MB peak GPU memory. These measurements describe the recorded runs; they are not minimum hardware requirements.</p>
</div>
</details>
+140
View File
@@ -0,0 +1,140 @@
---
hide:
- toc
---
# LongCat recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="longcat" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/meituan-longcat.webp" alt="Meituan LongCat" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>LongCat inference recipes</h2>
<p>LongCat Video from Meituan covers text-to-video and image-to-video, and its maintained examples chain optional distilled and 720p refinement passes.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading LongCat recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/meituan-longcat">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>Each LongCat script runs multiple passes (basic, distilled, refine); total runtime scales accordingly, and every pass prints its own output directory.</li>
<li>Out of memory: the sources already enable VAE and text-encoder CPU offload; further options are covered in <a href="../../inference/offloading/">Offloading</a>.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>All recipes on this page are <strong>Source-backed</strong>: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are <strong>Unknown</strong> and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.</p>
</div>
</details>
+141
View File
@@ -0,0 +1,141 @@
---
hide:
- toc
---
# LTX recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="ltx2" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/ltx.webp" alt="Lightricks" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>LTX inference recipes</h2>
<p>LTX-2 from Lightricks generates video with synchronized audio. FastVideo maintains both a distilled four-GPU path and a base 1088p single-GPU path.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading LTX recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/Lightricks">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>The distilled recipe is source-configured for four GPUs; running it on fewer GPUs is unverified and may fail during distributed setup.</li>
<li>Audio-less output usually means the base checkpoint resolved instead of the distilled LTX-2 checkpoint with audio; check the loaded model ID in the logs.</li>
<li>Gated or missing checkpoints: run <code>huggingface-cli login</code> and confirm you accepted the model's license on Hugging Face.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>All recipes on this page are <strong>Source-backed</strong>: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are <strong>Unknown</strong> and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.</p>
</div>
</details>
+140
View File
@@ -0,0 +1,140 @@
---
hide:
- toc
---
# Matrix Game recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="matrixgame" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/fastvideo.webp" alt="FastVideo org (converted weights)" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>Matrix Game inference recipes</h2>
<p>Matrix Game generates playable interactive worlds from an input image and prompt. Maintained examples cover Matrix Game 2.0 variants and Matrix Game 3.0 at 720p.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading Matrix Game recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/FastVideo">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>Matrix Game 3.0 downloads its reference input image from GitHub; offline machines should pre-download it and edit the <code>IMAGE_URL</code> constant locally.</li>
<li>Streaming variants of Matrix Game 2.0 exist under <code>examples/inference/basic/basic_matrixgame2_streaming.py</code> but are not included as cookbook recipes.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>All recipes on this page are <strong>Source-backed</strong>: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are <strong>Unknown</strong> and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.</p>
</div>
</details>
+283
View File
@@ -0,0 +1,283 @@
---
hide:
- toc
---
# MiniMax H3 recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="minimax_h3" data-default-recipe="fasth3-preview-cuda" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/minimax.webp" alt="MiniMax" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Primary focus · Inference</p>
<h2>MiniMax H3 recipes</h2>
<p>Generate video and audio with H3. Run a server on CUDA, one DGX Spark, or Apple Silicon MLX to iterate on prompts, or call the pipeline directly from Python.</p>
<span class="cookbook-count" data-cookbook-count>8 maintained recipes</span>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<details class="cookbook-modes">
<summary>Compare H3 modes and options</summary>
<h2 id="h3-modes-heading">Supported modes</h2>
<p>
CUDA covers T2VA, FL2VA, and Ref2VA on the full checkpoint, plus FastH3
Preview and FastH3 LoRA. FastH3 Preview also has a DGX Spark runtime with
a 1-Spark or 2-Spark device row. MLX is T2VA only. Temporal <code>--fast</code>,
spatial <code>--fast-spatial</code>, and opt-in VSA are flags on the same
MLX script, not extra recipes.
</p>
<div class="cookbook-modes__table-wrap">
<table>
<thead>
<tr>
<th>Mode</th>
<th>CUDA</th>
<th>MLX FastH3</th>
</tr>
</thead>
<tbody>
<tr>
<td>T2VA</td>
<td>Full H3, FastH3 Preview, FastH3 LoRA</td>
<td>FastH3 Preview after a local DiT conversion</td>
</tr>
<tr>
<td>FL2VA</td>
<td>Full H3</td>
<td>Not wired</td>
</tr>
<tr>
<td>Ref2VA</td>
<td>Full H3</td>
<td>Not wired</td>
</tr>
<tr>
<td>Temporal <code>--fast</code></td>
<td>No cookbook recipe</td>
<td>Shorter video denoise, MLX RIFE back to <code>--num-frames</code>, full-duration audio</td>
</tr>
<tr>
<td>VSA</td>
<td>Trained sparse attention on FastH3 CUDA</td>
<td>Opt-in. Convert with <code>--include-vsa</code> into a new directory such as <code>./FastH3-MLX-vsa</code>, then pass <code>--vsa</code>. Do not overwrite an existing dense export.</td>
</tr>
<tr>
<td>Spatial <code>--fast-spatial</code></td>
<td>No cookbook recipe</td>
<td>Denoise and decode at height/width divided by <code>--fast-spatial-scale</code>, then resample. Composes with <code>--fast</code>.</td>
</tr>
<tr>
<td>Two-pass refine</td>
<td>No cookbook recipe</td>
<td>Not wired</td>
</tr>
<tr>
<td>DGX Spark</td>
<td>FastH3 Preview on one GB10, or two Sparks with Ray sequence parallel (<code>sp_size=2</code>) over QSFP RoCE. Select NVIDIA DGX Spark, then 1 Spark or 2 Sparks.</td>
<td>Not wired</td>
</tr>
</tbody>
</table>
</div>
</details>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick an H3 recipe and runtime</h2>
<p>Choose the result you want, then use a maintained CUDA, DGX Spark, or MLX path.
Device claims stay tied to checked-in sources and recorded runs.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading MiniMax H3 recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<div class="cookbook-selection-row" data-cookbook-device-row hidden>
<div class="cookbook-selection-row__label">
<strong>Devices</strong>
<span data-cookbook-device-caption>1 Spark or a QSFP pair</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-device-options role="group" aria-label="Devices">
</div>
</div>
<div data-cookbook-knobs></div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<div class="cookbook-selection-row" data-cookbook-usage>
<div class="cookbook-selection-row__label">
<strong>Workflow</strong>
<span>Both can run locally</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" role="group" aria-label="How to run this recipe">
<button type="button" data-cookbook-mode="server" aria-pressed="false"><strong>Run a server</strong><span>Playground, cURL, or an API client</span></button>
<button type="button" data-cookbook-mode="python" aria-pressed="false"><strong>Use Python directly</strong><span>Call the model in your own process</span></button>
</div>
</div>
<p class="cookbook-hardware-note" data-cookbook-serving-availability></p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Hardware</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre id="cookbook-local-command"><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<p class="cookbook-hardware-note" data-cookbook-python-note>Running this script again starts a new process and reloads the model. To iterate in Python, create the generator once and reuse it for multiple prompts.</p>
<div class="cookbook-serving" data-cookbook-serving hidden>
<p class="cookbook-serving__intro" data-cookbook-server-lifetime>Start once, then change prompts in the playground or your app. You can run the server and clients on the same machine.</p>
<section class="cookbook-serving__step" aria-labelledby="serving-install-heading">
<h4 id="serving-install-heading"><span aria-hidden="true">1</span> Prepare the machine</h4>
<p>Run from your FastVideo clone in an activated Python environment. See <a data-cookbook-install-guide href="../../getting_started/installation/gpu/">installation requirements</a>.</p>
<div class="cookbook-command"><div class="cookbook-command__bar"><span>GPU machine · Terminal</span></div><pre id="cookbook-server-install"><code class="language-bash" data-cookbook-server-install></code></pre></div>
<details class="cookbook-serving__prepare" data-cookbook-prepare hidden><summary>Download and convert MLX weights once</summary><p>Skip this if the weights are already prepared. Edit the paths in <code>examples/serving/mlx_fasth3.yaml</code> to use your existing files. Install <code>ffmpeg</code> for video and audio output.</p><div class="cookbook-command"><pre id="cookbook-server-prepare"><code class="language-bash" data-cookbook-server-prepare></code></pre></div></details>
</section>
<section class="cookbook-serving__step" aria-labelledby="serving-start-heading">
<h4 id="serving-start-heading"><span aria-hidden="true">2</span> Start the server</h4>
<p>Keep this terminal running while you use the playground or API clients.</p>
<div class="cookbook-command"><div class="cookbook-command__bar"><span>GPU machine · Terminal</span></div><pre id="cookbook-server-command"><code class="language-bash" data-cookbook-server-command></code></pre></div>
<details class="cookbook-serving__check"><summary>Check that the server is ready</summary><p>In another terminal, this returns <code>{"status":"ok"}</code> after startup.</p><div class="cookbook-command"><pre id="cookbook-health-command"><code class="language-bash" data-cookbook-health-command></code></pre></div></details>
</section>
<section class="cookbook-serving__step" aria-labelledby="serving-client-heading">
<h4 id="serving-client-heading"><span aria-hidden="true">3</span> Generate and download a video</h4>
<div class="cookbook-serving__playground">
<div><strong>Try prompts in your browser</strong><p>Edit a prompt, generate, and watch the result. The playground uses the same server as cURL and your app.</p></div>
<a class="cookbook-serving__launch" data-cookbook-playground href="http://127.0.0.1:8000/playground/" target="_blank" rel="noopener">Open playground <span aria-hidden="true">↗</span></a>
</div>
<p class="cookbook-serving__local-hint">Open after the server is ready. On a remote GPU machine, <a href="../openai-api/#connect-your-app">forward port 8000</a> to your computer first. This opens a local page, not a hosted demo.</p>
<details class="cookbook-serving__code"><summary>Use cURL or an SDK</summary>
<p>Each example submits a job, checks its status, and saves the MP4. The Python and JavaScript examples use OpenAI-compatible clients; no OpenAI account is needed.</p>
<div class="cookbook-serving__clients" role="group" aria-label="API client language">
<button type="button" data-cookbook-client="curl" aria-pressed="false">cURL</button>
<button type="button" data-cookbook-client="python" aria-pressed="true">Python</button>
<button type="button" data-cookbook-client="javascript" aria-pressed="false">JavaScript</button>
</div>
<div class="cookbook-command"><div class="cookbook-command__bar"><span>Client dependencies</span></div><pre id="cookbook-client-install"><code class="language-bash" data-cookbook-client-install></code></pre></div>
<div class="cookbook-command cookbook-command--client"><div class="cookbook-command__bar"><span data-cookbook-client-filename>video.py</span><a data-cookbook-client-source href="https://github.com/hao-ai-lab/FastVideo/tree/main/examples/serving/clients">View source</a></div><pre id="cookbook-client-code"><code data-cookbook-client-code></code></pre></div>
<p data-cookbook-client-run></p>
</details>
</section>
<p class="cookbook-serving__boundary">This is a local development server without built-in API-key authentication. The client key <code>local</code> is a placeholder. Keep the server on loopback; use an authenticated TLS proxy before exposing it publicly. Run the JavaScript client in your webapp's backend, not in a browser with a private key.</p>
<a href="../openai-api/">Server guide and API compatibility →</a>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/MiniMaxAI">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../openai-api/">H3 server guide</a> or the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
<p class="cookbook-eyebrow">CUDA</p>
<pre><code>UV_TORCH_BACKEND=cu130 uv pip install -e ".[fasth3]"</code></pre>
<p class="cookbook-eyebrow">Apple Silicon</p>
<pre><code>uv pip install -e ".[mlx]"</code></pre>
<p>Follow the <a href="../../getting_started/installation/mps/#run-fasth3-preview">Apple Silicon guide</a> for the download, conversion, and storage requirements.</p>
<p class="cookbook-eyebrow">NVIDIA DGX Spark</p>
<pre><code>UV_TORCH_BACKEND=cu130 uv pip install -e .</code></pre>
<p>Follow the <a href="../../getting_started/installation/spark/">DGX Spark install guide</a> for ARM64 CUDA 13. One Spark is a local process. Two Sparks need Ray on the QSFP link:</p>
<pre><code>uv pip install ray</code></pre>
<p>Bring up the cluster from <a href="../../getting_started/installation/spark_pair/">pairing two Sparks</a> before selecting 2 Sparks in the builder.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>The full CUDA H3 examples request four GPUs by default. Their sources do not claim a GPU model or memory minimum.</li>
<li>The FastH3 CUDA performance profile was measured on four GB200 GPUs. Use its strict profile when exact operation order matters more than the measured performance configuration.</li>
<li>The MLX source runtime supports T2VA, optional temporal <code>--fast</code>, optional spatial <code>--fast-spatial</code>, and opt-in VSA on <code>--include-vsa</code> checkpoints. FL2VA, Ref2VA, and two-pass refinement are not wired.</li>
<li>GPU count and VAE decode backend are configurable in the builder above for FastH3 CUDA recipes. Only the value shown by default has a recorded run; other supported values are unmeasured here.</li>
<li>DGX Spark is a runtime on FastH3 Preview, not a separate family card. Select NVIDIA DGX Spark, then 1 Spark or 2 Sparks. The CUDA GPU-count knob does not apply to Spark.</li>
<li>GB10 has no FA4 / sm_100a VSA kernel. Keep <code>FASTVIDEO_FA4=0</code> and <code>FASTVIDEO_VSA_SM100A=0</code>. Legal <code>num_frames</code> values are <code>17n+5</code>, capped at 345 (15 s). A 345-frame request on one Spark can OOM.</li>
<li>Gated or missing checkpoints: run <code>huggingface-cli login</code> and confirm you accepted the model's license on Hugging Face.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>Every command, model ID, and flag on this page maps to a checked-in FastVideo source. Recipes marked <strong>Verified</strong> also have a recorded hardware path in linked FastVideo evidence. The full H3 CUDA examples remain <strong>Source-backed</strong> where the source records a GPU count but no GPU model or memory requirement. Unlisted hardware is unknown, not unsupported.</p>
</div>
</details>
+140
View File
@@ -0,0 +1,140 @@
---
hide:
- toc
---
# MMAudio recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="mmaudio" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<span class="cookbook-family-tile__monogram" aria-hidden="true">MMA</span>
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>MMAudio inference recipes</h2>
<p>MMAudio adds synchronized audio to video or generates audio from text. The checkpoint must be in Diffusers layout; the recipe loads the converted FastVideo repo through the `MMAUDIO_MODEL_PATH` environment variable the example reads.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading MMAudio recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/FastVideo">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>If the example reports a missing local <code>converted_weights/mmaudio/large_44k_v2</code> path, the <code>MMAUDIO_MODEL_PATH</code> env var from the cookbook command was not applied; export it in the same shell.</li>
<li>Alternatively convert upstream weights yourself with <code>scripts/checkpoint_conversion/convert_mmaudio_to_diffusers.py</code> and point the env var at the result.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>All recipes on this page are <strong>Source-backed</strong>: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are <strong>Unknown</strong> and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.</p>
</div>
</details>
+250
View File
@@ -0,0 +1,250 @@
# Run H3 with a server and playground
Start FastH3 on CUDA, one DGX Spark, or Apple Silicon MLX, then iterate on prompts in a browser,
with cURL, or from your app. The server and clients can run on the same machine.
The server supports the OpenAI-compatible video-job API. The Python and
JavaScript client examples use that interface, but requests go to FastVideo.
You do not need an OpenAI account or cloud key.
The [H3 recipe selector](minimax-h3.md) provides the same workflow with runtime
selection. This guide covers FastH3 Preview text-to-video/audio. Other H3
recipes keep their direct Python commands.
CUDA requests reuse one loaded `VideoGenerator`. The Python SDK can do the same
when you reuse the generator across `generate()` calls. MLX keeps one
`MiniMaxH3MLXPipeline` and a prompt-embedding cache, but still loads and releases
components between phases to fit unified memory. MLX serving does not keep all
weights resident or remove phase-loading time. Use a server when you want
separate clients to share the process without restarting scripts.
## Install and start the server
### CUDA
Use a FastVideo clone and an activated Python environment. Complete the
[CUDA installation requirements](../getting_started/installation/gpu.md)
before running these commands:
```bash
UV_TORCH_BACKEND=cu130 uv pip install -e ".[fasth3]"
fastvideo serve --config examples/serving/openai_fasth3.yaml --server.host 127.0.0.1
```
The configuration loads `FastVideo/FastVideo-Minimax-FastH3-Preview-v0.2` and
advertises it as `fasth3`. It configures four CUDA GPUs but does not record
a GPU model or VRAM requirement. This is a source-backed server profile, not
the measured GB200 Python performance profile. Compilation is disabled.
Keep the server running. In another terminal, check readiness:
```bash
curl --fail-with-body http://127.0.0.1:8000/health
```
After model loading completes, the response is `{"status":"ok"}`.
### NVIDIA DGX Spark
Complete the [DGX Spark installation](../getting_started/installation/spark.md)
before running these commands. GB10 has no FA4 / sm_100a VSA kernel:
```bash
UV_TORCH_BACKEND=cu130 uv pip install -e .
FASTVIDEO_VSA_SM100A=0 FASTVIDEO_FA4=0 FASTVIDEO_ATTENTION_BACKEND=VIDEO_SPARSE_ATTN_H3 \
fastvideo serve --config examples/serving/openai_fasth3_spark.yaml --server.host 127.0.0.1
```
The configuration loads `FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree`
on one GB10 and advertises it as `fasth3`. Lazy module load still reloads
Qwen3-VL and the DiT between phases of each request. Legal `num_frames` values
are `17n+5`, capped at 345 (15 s); a 345-frame request on one Spark can OOM.
There is no cookbook server for two Sparks; use the generate YAML after
[pairing two Sparks](../getting_started/installation/spark_pair.md).
Keep the server running. In another terminal, check readiness:
```bash
curl --fail-with-body http://127.0.0.1:8000/health
```
After model loading completes, the response is `{"status":"ok"}`.
### Apple Silicon MLX
Complete the [Apple Silicon installation](../getting_started/installation/mps.md#run-fasth3-preview),
including `ffmpeg`. From your FastVideo clone, install the MLX extra:
```bash
uv pip install -e ".[mlx]"
```
Download and convert the weights once. Skip this step if you already have the
snapshot and converted DiT:
```bash
hf download FastVideo/FastVideo-Minimax-FastH3-Preview-v0.2 --local-dir ./FastH3-Preview-v0.2
python scripts/checkpoint_conversion/convert_minimax_h3_mlx.py --model-root ./FastH3-Preview-v0.2/transformer --out ./FastH3-MLX --formats "int6"
```
Edit `generator.model_root` and `generator.mlx_checkpoint` in
`examples/serving/mlx_fasth3.yaml` if your weights are elsewhere. Paths are
relative to the directory where you launch the server. Start it with:
```bash
python -m fastvideo.entrypoints.openai.mlx_server --config examples/serving/mlx_fasth3.yaml
```
This uses the native MLX pipeline, not PyTorch MPS. The default output is
832 × 480, 124 frames at 24 fps, with four DiT forwards and the full H3 VAE.
The HTTP field `num_inference_steps=5` describes five sigma points; the adapter
passes `num_steps=4` to MLX. Temporal/spatial fast modes, VSA, reference inputs,
LoRA selection, and alternate decoders are not exposed by this server adapter.
Unsupported request options return HTTP 400 before a job starts.
The server binds to `127.0.0.1:8000` and advertises `fasth3`, so the playground
and clients below work unchanged. Do not run CUDA, Spark, and MLX servers on the
same port. MLX readiness means the pipeline is initialized; components load during
generation. The first request can take longer than a repeated cached prompt.
The MLX server has no recorded device or unified-memory requirement. The
direct Python recipe's M4 Max measurements are not a server benchmark or a
minimum-memory claim.
## Open the playground
Open [the local H3 playground](http://127.0.0.1:8000/playground/) after startup.
Write a prompt, optionally set a seed, then select **Generate video**. The page
checks the job status and shows the completed video with a download link. Edit
the prompt and generate again without restarting the server.
The playground sends requests to `/v1/videos` on the same server that serves
the page. **Use this prompt with cURL** shows the equivalent submission. Recent
jobs include requests from other clients, so a script and the playground can
use the same model process. Opening the page or copying a command does not
start generation.
The page URL includes the active job ID. Reloading that URL resumes status
checks without resubmitting the prompt. A failed connection or a 30-minute
polling timeout does not cancel execution. Select **Check status** to reconnect.
If submission itself is interrupted, check Recent jobs before submitting again.
This first playground supports H3 text-to-video/audio. Reference-media inputs
remain available through the API, not through the playground. It does not start
or manage a GPU server for you.
## Generate with cURL or an SDK
These examples use the server's resolution, frame count, and sampling defaults.
Do not copy Sora-specific durations or resolutions onto H3. Both server configs
use 124 frames, 24 fps, and the five-point distilled sigma schedule with four
DiT forwards. CUDA and one Spark use 1344 × 768; MLX uses 832 × 480.
Each client submits a job, checks for completion or failure, and downloads an
MP4 named after the job ID. Polling stops after 30 minutes; a timeout does not
cancel GPU execution. Keep the printed job ID to retrieve its status later.
Transport retries are disabled to avoid accidental duplicate submissions.
### OpenAI Python
Install the tested client, then run the checked-in example:
```bash
python -m pip install openai==3.6.0
python examples/serving/clients/video.py
```
```python
--8<-- "examples/serving/clients/video.py"
```
### OpenAI JavaScript
Use Node.js 22 or later on your computer or in your webapp's backend. Do not put a
private server key in browser code.
```bash
npm ci --prefix examples/serving/clients
node examples/serving/clients/video.mjs
```
```javascript
--8<-- "examples/serving/clients/video.mjs"
```
### cURL
Install `curl` and `jq`, then run:
```bash
bash examples/serving/clients/video.sh
```
```bash
--8<-- "examples/serving/clients/video.sh"
```
## Connect your app
Set `FASTVIDEO_BASE_URL` to the FastVideo endpoint, including `/v1`, and
`FASTVIDEO_MODEL` to its advertised model alias. The examples default to
`http://127.0.0.1:8000/v1` and `fasth3`.
For a remote GPU machine, keep the server bound to loopback and forward the
port. Replace `user@gpu-host` with your SSH destination:
```bash
ssh -N -L 8000:127.0.0.1:8000 user@gpu-host
```
Then open `http://127.0.0.1:8000/playground/` on your computer. The same forwarded
address works for the cURL and SDK clients. If local port 8000 is occupied, use
`-L 8001:127.0.0.1:8000`, open port 8001, and set `FASTVIDEO_BASE_URL` to
`http://127.0.0.1:8001/v1` for the example clients.
Your webapp backend can submit a job and return its ID to the browser. Poll
from the backend, then proxy the completed download or store it in your own
artifact store. Do not hold a browser request open for the entire generation.
The client key `local` is a placeholder required by the SDK, not server
authentication. FastVideo's HTTP server has no built-in API-key check. Before
public deployment, put it behind an authenticated TLS proxy with restricted
origins, request limits, and access controls. If your proxy uses bearer tokens,
set `FASTVIDEO_API_KEY` on your backend. Never reuse an OpenAI cloud key here.
## Compatibility and limits
The client examples target the OpenAI video-job API, not Chat Completions.
The compatibility tests cover real HTTP requests with the pinned SDKs and a
fake generator. They establish client and protocol behavior, not GPU generation
quality, latency, or memory use. The server configuration still needs a recorded
H3 hardware run before it can be marked Verified in the cookbook.
| Operation | Endpoint |
| --- | --- |
| List served models | `GET /v1/models` |
| Create a video job | `POST /v1/videos` |
| Retrieve status, including a failed job | `GET /v1/videos/{id}` |
| List jobs | `GET /v1/videos` |
| Download the MP4 | `GET /v1/videos/{id}/content` |
| Delete a job and its artifact | `DELETE /v1/videos/{id}` |
Failed jobs return HTTP 200 with `status: "failed"` and an `error` object so
SDK polling can stop. Invalid requests return HTTP 400; missing jobs return
HTTP 404. The download endpoint supports only the `video` variant, not
thumbnails or spritesheets. Remix, extensions, characters, and an OpenAI Files
store are not implemented. `/v1/videos/sync` is a FastVideo extension and is
not used by these client examples.
OpenAI marks its hosted Sora API as deprecated. FastVideo runs its own models;
the [OpenAI video API reference](https://developers.openai.com/api/reference/python/resources/videos/methods/create)
describes the client interface, not FastVideo model availability. The examples
pin tested SDK versions because future SDKs may change or remove video helpers.
The server serializes generation through one loaded pipeline. Job metadata is
in memory and is lost on restart. Async artifacts remain on disk until deleted
through the API; establish a retention policy for production use. Deleting an
in-progress job does not interrupt an already running CUDA call.
For model-specific request fields and reference media, see the
[HTTP contract](../design/server_contracts/openai.md).
+140
View File
@@ -0,0 +1,140 @@
---
hide:
- toc
---
# Stable Audio recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="stable_audio" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/stabilityai.webp" alt="Stability AI" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>Stable Audio inference recipes</h2>
<p>Stable Audio Open generates audio from text. FastVideo requires the converted Diffusers-format repos published by the FastVideo organization; upstream monolithic checkpoints are not loader-compatible.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading Stable Audio recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/stabilityai">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>Loader errors about monolithic checkpoints mean an upstream <code>stabilityai/stable-audio-open-*</code> ID was used; use the FastVideo converted repos from the recipes.</li>
<li>Duration and step knobs (<code>audio_end_in_s</code>, <code>num_inference_steps</code>) are documented inline in the example source.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>The Stable Audio Open 1.0 recipe maps to a checked-in example and a recorded single-GPU B200 run. The Stable Audio Open Small recipe remains <strong>Source-backed</strong> because the implementation PR did not record a full run for that gated checkpoint. Neither recipe claims a minimum VRAM requirement.</p>
</div>
</details>
+140
View File
@@ -0,0 +1,140 @@
---
hide:
- toc
---
# Stable Diffusion recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="sd35" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/stabilityai.webp" alt="Stability AI" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>Stable Diffusion inference recipes</h2>
<p>Stable Diffusion 3.5 Medium is Stability AI's text-to-image model in this catalog. The maintained example sweeps a small prompt set across seeds.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading Stable Diffusion recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/stabilityai">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>The example writes several PNGs under <code>outputs/sd35/samples/</code>; make sure the output directory is writable.</li>
<li>Gated checkpoints: Stability AI models require accepting the license and <code>huggingface-cli login</code>.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>All recipes on this page are <strong>Source-backed</strong>: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are <strong>Unknown</strong> and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.</p>
</div>
</details>
+140
View File
@@ -0,0 +1,140 @@
---
hide:
- toc
---
# TurboDiffusion recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="turbodiffusion" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<span class="cookbook-family-tile__monogram" aria-hidden="true">Turbo</span>
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>TurboDiffusion inference recipes</h2>
<p>TurboDiffusion profiles accelerate Wan checkpoints with step-distilled sampling and the SLA attention backend. These recipes follow the registry's `turbodiffusion` model family.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading TurboDiffusion recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/loayrashid">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>TurboDiffusion paths load community-published <code>loayrashid/TurboWan*</code> checkpoints; availability is governed by those repos.</li>
<li>The SLA attention backend used by the I2V recipe is selected inside the example source; do not combine it with another <code>FASTVIDEO_ATTENTION_BACKEND</code> override in the same shell.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>All recipes on this page are <strong>Source-backed</strong>: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are <strong>Unknown</strong> and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.</p>
</div>
</details>
+201
View File
@@ -0,0 +1,201 @@
---
hide:
- toc
---
# Wan recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="wan" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/wan-ai.webp" alt="Wan-AI" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>Wan inference recipes</h2>
<p>CUDA covers FastWan and Wan2.1/2.2 text and image recipes. Apple Silicon uses the released FastMetal 1.3B, 5B, and 14B MLX T2V paths. The speed flags below are switches on those same scripts, not extra recipes.</p>
<span class="cookbook-count" data-cookbook-count>7 maintained recipes</span>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<details class="cookbook-modes">
<summary>Compare Wan modes and options</summary>
<h2 id="wan-modes-heading">Supported modes</h2>
<p>
FastMetal MLX is T2V in the checked-in examples. Image-to-video and
TI2V stay on the CUDA recipes. Temporal <code>--fast</code> composes
with either spatial path. <code>--refine</code> and
<code>--fast-spatial</code> cannot run together.
<code>--refine</code> wins if both are set.
<code>basic_mps.py</code> is the older PyTorch MPS demo and is not a
FastMetal recipe.
</p>
<div class="cookbook-modes__table-wrap">
<table>
<thead>
<tr>
<th>Mode</th>
<th>CUDA</th>
<th>MLX FastMetal</th>
</tr>
</thead>
<tbody>
<tr>
<td>T2V</td>
<td>FastWan2.1 1.3B, Wan2.2 A14B</td>
<td>1.3B, 5B, and 14B</td>
</tr>
<tr>
<td>I2V</td>
<td>Wan2.1 14B 480P</td>
<td>Not in the released examples</td>
</tr>
<tr>
<td>TI2V</td>
<td>Wan2.2 TI2V 5B</td>
<td>FastMetal 5B is T2V in <code>mlx_wan22_generate.py</code></td>
</tr>
<tr>
<td>Temporal <code>--fast</code></td>
<td>No cookbook recipe</td>
<td>RIFE. Fewer frames, then interpolate to <code>--num-frames</code></td>
</tr>
<tr>
<td>Spatial <code>--fast-spatial</code></td>
<td>No cookbook recipe</td>
<td>Denoise and decode at half resolution, then upsample. No second denoise</td>
</tr>
<tr>
<td>Two-pass <code>--refine</code></td>
<td>No cookbook recipe</td>
<td>Denoise at base resolution, upsample, re-noise, denoise again. Wins over <code>--fast-spatial</code></td>
</tr>
</tbody>
</table>
</div>
</details>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Choose the result you want, then use a maintained CUDA or native MLX path.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading Wan recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Hardware</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/Wan-AI">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>Out of memory on the A14B recipes: the checked-in sources already enable CPU offload; see <a href="../../inference/configuration/">Configuration</a> for the offload surface before reducing resolution or frames.</li>
<li>The FastWan2.1 recipe requires <code>VIDEO_SPARSE_ATTN</code>; confirm the environment variable in the command was set in the same shell.</li>
<li>FastMetal MLX: install with <code>uv pip install -e ".[mlx]"</code>, then follow the <a href="../../getting_started/installation/mps/">Apple Silicon guide</a>. CUDA FastWan-QAD checkpoints are refused on the MLX runtime.</li>
<li>FastMetal 5B uses <code>mlx_wan22_generate.py</code>. 1.3B and 14B use <code>mlx_wan_prompt_to_video.py</code>.</li>
<li>Gated or missing checkpoints: run <code>huggingface-cli login</code> and confirm you accepted the model's license on Hugging Face.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>Every recipe on this page maps to a checked-in FastVideo source. The FastMetal MLX releases include the recorded M4 Max system memory, documented unified-memory floor, and measured peak MLX memory. CUDA entries remain <strong>Source-backed</strong> where the examples record a GPU count but no exact GPU model or VRAM. Unlisted hardware is unknown, not unsupported.</p>
</div>
</details>
+140
View File
@@ -0,0 +1,140 @@
---
hide:
- toc
---
# Z-Image recipes
<div class="cookbook-shell cookbook-family-page" data-cookbook data-family="zimage" data-recipes="../../assets/cookbook-recipes.json?v=8">
<header class="cookbook-family-header">
<a class="cookbook-back-link" href="../"><span aria-hidden="true">←</span> All model families</a>
<div class="cookbook-family-header__body">
<span class="cookbook-family-header__logo">
<img class="off-glb" src="../../assets/logos/tongyi.webp" alt="Tongyi MAI" width="112" height="112">
</span>
<div>
<p class="cookbook-eyebrow">Maintained family · Inference</p>
<h2>Z-Image inference recipes</h2>
<p>Z-Image Turbo from Tongyi MAI is a fast text-to-image model. The maintained example runs it on a single GPU.</p>
</div>
</div>
<div class="cookbook-lifecycle" aria-label="Lifecycle stages">
<span class="cookbook-lifecycle__stage cookbook-lifecycle__stage--active">Inference <small>live</small></span>
<span class="cookbook-lifecycle__stage">Distillation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Fine-tuning <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Training <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Evaluation <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Optimization <small>planned</small></span>
<span class="cookbook-lifecycle__stage">Deployment <small>planned</small></span>
</div>
</header>
<nav class="cookbook-jumpnav" aria-label="Recipe page sections">
<a href="#recipe-builder">Builder</a>
<a href="#cookbook-setup">Setup</a>
<a href="#cookbook-troubleshooting">Troubleshooting</a>
<a href="#cookbook-evidence">Evidence</a>
</nav>
<section class="cookbook-builder" id="recipe-builder" aria-labelledby="builder-heading">
<div class="cookbook-builder__intro">
<h2 id="builder-heading">Pick a recipe and runtime</h2>
<p>Start with the result you want, then choose one of the runtimes FastVideo actually maintains for it.</p>
</div>
<div class="cookbook-builder__layout">
<div class="cookbook-controls">
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Recipe</strong>
<span>Task and checkpoint</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--models" data-cookbook-model-options role="group" aria-label="Recipe">
<button type="button" disabled>Loading Z-Image recipes...</button>
</div>
</div>
<div class="cookbook-selection-row">
<div class="cookbook-selection-row__label">
<strong>Runtime</strong>
<span>Maintained paths only</span>
</div>
<div class="cookbook-option-grid cookbook-option-grid--hardware" data-cookbook-hardware-options role="group" aria-label="Runtime">
<button type="button" disabled>Loading runtimes...</button>
</div>
</div>
<p class="cookbook-selection-description" data-cookbook-description>Loading recipe details...</p>
<p class="cookbook-hardware-note">Exact device and memory details appear only when a recorded run supports them.</p>
<div class="cookbook-hardware-state" data-cookbook-hardware-state role="status" aria-live="polite">
Reading recipe evidence...
</div>
</div>
<article class="cookbook-result">
<div class="cookbook-result__header">
<h3 data-cookbook-label>Loading...</h3>
<div class="cookbook-result__badges">
<span class="cookbook-badge">Maintained</span>
<span class="cookbook-badge" data-cookbook-evidence>Source-backed</span>
<span class="cookbook-badge cookbook-badge--neutral" data-cookbook-hardware-badge>Source config</span>
</div>
</div>
<dl class="cookbook-result__facts">
<div><dt>Model</dt><dd data-cookbook-model>Loading...</dd></div>
<div><dt>Workload</dt><dd data-cookbook-task>Loading...</dd></div>
<div><dt>Source configuration</dt><dd data-cookbook-gpus>Loading...</dd></div>
<div><dt>Expected output</dt><dd data-cookbook-artifact>Loading...</dd></div>
</dl>
<div class="cookbook-command">
<div class="cookbook-command__bar">
<span>Terminal</span>
</div>
<pre><code class="language-bash" data-cookbook-command>Loading...</code></pre>
</div>
<div class="cookbook-result__footer">
<a data-cookbook-source href="../../inference/examples/basic/">Open example source</a>
<a data-cookbook-model-link href="https://huggingface.co/Tongyi-MAI">View model card</a>
</div>
<p class="cookbook-picker__status" role="status" aria-live="polite" data-cookbook-status></p>
</article>
</div>
<noscript>
<div class="cookbook-noscript">
JavaScript is needed for the guided selector. You can still browse the
<a href="../../inference/examples/examples_inference_index/">maintained inference examples</a>.
</div>
</noscript>
</section>
</div>
<details class="cookbook-collapsible" id="cookbook-setup">
<summary>Setup</summary>
<div class="cookbook-collapsible__body">
<p>The generated commands expect a local clone:</p>
<pre><code>git clone https://github.com/hao-ai-lab/FastVideo.git
cd FastVideo</code></pre>
<p>Use <a href="../../inference/configuration/">Configuration</a> for supported Python and CLI settings, <a href="../../inference/optimizations/">Optimizations</a> for attention and memory tradeoffs, and the <a href="../../inference/support_matrix/">support matrix</a> for the supported model and optimization surface.</p>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-troubleshooting">
<summary>Troubleshooting</summary>
<div class="cookbook-collapsible__body">
<ul>
<li>Output defaults to <code>outputs/zimage/zimage_turbo.png</code>; pass <code>--output</code> to redirect.</li>
<li>Gated or missing checkpoints: run <code>huggingface-cli login</code> and confirm you accepted the model's license on Hugging Face.</li>
</ul>
</div>
</details>
<details class="cookbook-collapsible" id="cookbook-evidence">
<summary>Evidence status</summary>
<div class="cookbook-collapsible__body">
<p>All recipes on this page are <strong>Source-backed</strong>: their commands, model IDs, and flags were validated against the checked-in FastVideo sources listed above (static validation). No runtime GPU validation is recorded for these recipes, so GPU model fit, memory use, throughput, and runtime duration are <strong>Unknown</strong> and deliberately not claimed. Runtime buttons show only the GPU counts configured in checked-in sources.</p>
</div>
</details>
+7 -6
View File
@@ -8,7 +8,7 @@ compute ~3.7×, so the wall-clock drops far more than 2×. RIFE (which estimates
its own optical flow — no motion vectors needed) fills the dropped frames back
in for ~1.4 s, and a light unsharp pass counters its softening.
Measured on the 1.3B INT8 QAD model (fox, 480×832×81, M4): generate 41 + RIFE→81
Measured on the 1.3B INT8 QAD model (480×832×81, M4): generate 41 + RIFE→81
runs in ~35 s of denoise vs ~90 s full, at reconstruction MS-SSIM **0.97**.
Reproduce with `python -m fastvideo.benchmarks.eval_metalfx_rife --mode int8`.
@@ -26,10 +26,11 @@ uv pip install -e ".[mlx]" # RIFE ships vendored; this only needs MLX
```bash
python examples/inference/basic/mlx_wan_prompt_to_video.py \
--mlx-checkpoint <FastWan2.1-T2V-1.3B-INT8-QAD> \
--prompt "A red fox trotting through a snowy pine forest at golden hour, cinematic" \
--model-root ./FastMetal-1.3B-QAD \
--mlx-checkpoint ./FastMetal-1.3B-QAD \
--prompt "A bird's-eye view of a misty forest valley at dawn." \
--num-frames 81 --fast \
--output-path video_samples/fox_fast.mp4
--output-path video_samples/forest_fast.mp4
```
`--num-frames` stays the *target* length; fast mode generates the smallest
@@ -57,9 +58,9 @@ of denoise. It composes with `--fast`; both together run the same clip in
```bash
python examples/inference/basic/mlx_wan_prompt_to_video.py \
--prompt "A red fox trotting through a snowy pine forest at golden hour, cinematic" \
--prompt "A bird's-eye view of a misty forest valley at dawn." \
--height 480 --width 832 --num-frames 81 --fast-spatial \
--output-path video_samples/fox_fast_spatial.mp4
--output-path video_samples/forest_fast_spatial.mp4
```
| Flag | Default | Meaning |
@@ -21,6 +21,8 @@ surfaces:
hsdp_shard_dim: generator.engine.parallelism.hsdp_shard_dim
dist_timeout: generator.engine.parallelism.dist_timeout
lora_path: generator.pipeline.components.lora_path
lora_nickname: generator.pipeline.components.lora_nickname
lora_strength: generator.pipeline.components.lora_strength
dit_cpu_offload: generator.engine.offload.dit
use_fsdp_inference: generator.engine.use_fsdp_inference
dit_layerwise_offload: generator.engine.offload.dit_layerwise
@@ -28,6 +30,7 @@ surfaces:
image_encoder_cpu_offload: generator.engine.offload.image_encoder
vae_cpu_offload: generator.engine.offload.vae
pin_cpu_memory: generator.engine.offload.pin_cpu_memory
lazy_module_load: generator.engine.offload.lazy_module_load
enable_torch_compile: generator.engine.compile.enabled
enable_torch_compile_text_encoder: generator.engine.compile.text_encoder_enabled
enable_torch_compile_vae: generator.engine.compile.vae_enabled
@@ -72,10 +75,18 @@ surfaces:
compatibility_only:
mode: "Legacy multi-mode FastVideoArgs switch; typed inference config should not expose execution mode."
inference_mode: "Legacy boolean mirror of mode; kept only through adapters while FastVideoArgs remains."
lora_nickname: "Legacy adapter-selection surface pending LoRA API cleanup."
lora_target_modules: "Legacy LoRA configuration surface pending dedicated component API."
output_type: "Legacy output formatting surface pending GenerationResult cleanup."
VSA_sparsity: "Model-specific inference optimization not yet represented in the typed public schema."
VSA_tile_size: "VSA-H3 tile geometry request; model-specific optimization not yet represented in the typed public schema."
inference_torch_compile: "Regional inference compile opt-in currently carried through PipelineSelection.experimental rather than CompileConfig."
vae_parallel_decode: "MiniMax-H3 sequence-parallel VAE decode opt-in; model-specific optimization not yet represented in the typed public schema."
h3_sequential_load: "MiniMax-H3 sequential text-encoder then DiT/VAE load; model-specific optimization not yet represented in the typed public schema."
video_decode_backend: "MiniMax-H3 video decoder selection (full VAE vs TAEH3 preview); model-specific optimization not yet represented in the typed public schema."
taeh3_checkpoint: "Optional local TAEH3 safetensors path; model-specific optimization not yet represented in the typed public schema."
taeh3_chunk_size: "TAEH3 temporal chunk length; model-specific optimization not yet represented in the typed public schema."
vae_parallel_encode: "MiniMax-H3 sequence-parallel reference VAE encode opt-in; model-specific optimization not yet represented in the typed public schema."
vae_parallel_decode_strategy: "Chunk-transport collective for vae_parallel_decode; model-specific optimization not yet represented in the typed public schema."
attention_backend: "Process-wide default attention-backend request applied per component at load time; kernel-selection knob not yet represented in the typed public schema."
moba_config_path: "Model-specific MoBA optimization surface not yet represented in the typed public schema."
master_port: "Executor/bootstrap compatibility field; not part of the canonical inference schema."
@@ -119,12 +130,14 @@ surfaces:
vae_precision: "Precision override pending dedicated typed component precision design."
vae_decode_precision: "Decode-only precision override pending dedicated typed component precision design."
image_encoder_precision: "Precision override pending dedicated typed component precision design."
image_encoder_precisions: "Precision overrides pending dedicated typed component precision design."
text_encoder_precisions: "Precision override pending dedicated typed component precision design."
internal_only:
dit_config: "Legacy internal component config object."
upsampler_config: "Legacy internal component config object."
vae_config: "Legacy internal component config object."
image_encoder_config: "Legacy internal component config object."
image_encoder_configs: "Legacy internal component config objects."
text_encoder_configs: "Legacy internal component config object."
preprocess_text_funcs: "Internal text preprocessing hooks."
postprocess_text_funcs: "Internal text postprocessing hooks."
@@ -361,6 +374,30 @@ surfaces:
sources: [fastvideo.configs.pipelines.matrixgame2.MatrixGame2I2V480PConfig]
num_frames_per_block:
sources: [fastvideo.configs.pipelines.matrixgame2.MatrixGame2I2V480PConfig]
duration_s:
sources: [fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig]
spectrogram_frame_rate:
sources: [fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig]
latent_downsample_rate:
sources: [fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig]
clip_frame_rate:
sources: [fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig]
sync_frame_rate:
sources: [fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig]
sync_segment_size:
sources: [fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig]
sync_segment_stride:
sources: [fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig]
sync_downsample_rate:
sources: [fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig]
clip_image_size:
sources: [fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig]
sync_image_size:
sources: [fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig]
clip_batch_size_multiplier:
sources: [fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig]
sync_batch_size_multiplier:
sources: [fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig]
audio_channels:
sources:
- fastvideo.configs.pipelines.stable_audio.StableAudioT2AConfig
@@ -375,6 +412,7 @@ surfaces:
- fastvideo.configs.pipelines.stable_audio.StableAudioOpenSmallConfig
max_audio_duration_s:
sources:
- fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig
- fastvideo.configs.pipelines.stable_audio.StableAudioT2AConfig
- fastvideo.configs.pipelines.stable_audio.StableAudioOpenSmallConfig
sample_size:
@@ -383,6 +421,7 @@ surfaces:
- fastvideo.configs.pipelines.stable_audio.StableAudioOpenSmallConfig
sampling_rate:
sources:
- fastvideo.configs.pipelines.mmaudio.MMAudioV2AConfig
- fastvideo.configs.pipelines.stable_audio.StableAudioT2AConfig
- fastvideo.configs.pipelines.stable_audio.StableAudioOpenSmallConfig
audio_txt_guidance_scale:
@@ -557,15 +596,31 @@ surfaces:
openai_video_request:
kept:
model: "HTTP adapter model-routing field."
user: "OpenAI-compatible caller tracking field."
task: "SGLang-compatible MiniMax-H3 task selector validated against the startup pipeline."
quality: "vLLM-Omni-compatible quality-intent field; model-specific."
lora: "vLLM-Omni-compatible selector for the adapter fixed at server startup."
moved:
prompt: request.prompt
input_reference: request.inputs.image_path
reference_url: request.inputs.image_path
image_reference: request.inputs.image_path,last_image,references
video_reference: request.inputs.video_path,references
audio_reference: request.inputs.references
video_path: request.inputs.video_path
video_url: request.inputs.video_path
video_params: request.sampling.width,height,num_frames,fps
size:
target: request.sampling.width,height
note: "Adapter parses OpenAI size strings as WIDTHxHEIGHT and forwards width then height."
width: request.sampling.width
height: request.sampling.height
fps: request.sampling.fps
num_frames: request.sampling.num_frames
aspect_ratio: request.sampling.width,height
short_edge: request.sampling.width,height
num_outputs_per_prompt: request.sampling.num_videos_per_prompt
n: request.sampling.num_videos_per_prompt
seed: request.sampling.seed
num_inference_steps: request.sampling.num_inference_steps
guidance_scale: request.sampling.guidance_scale
@@ -573,11 +628,21 @@ surfaces:
true_cfg_scale: request.sampling.true_cfg_scale
negative_prompt: request.negative_prompt
enable_teacache: request.runtime.enable_teacache
output_path: request.output.output_path
max_sequence_length: request.sampling.max_sequence_length
boundary_ratio: request.sampling.boundary_ratio
extra_params: request.extensions
compatibility_only:
seconds:
target: request.sampling.num_frames
note: "HTTP adapter duration convenience field. If num_frames is omitted, the adapter computes num_frames = fps * seconds."
start_time_seconds: "vLLM-Omni reference-video offset; rejected by pipelines that cannot represent it."
flow_shift: "vLLM-Omni request field; accepted only when the selected model exposes a matching request parameter."
generate_sound: "vLLM-Omni audio-output intent; accepted only by models with a matching request parameter."
sound_duration: "vLLM-Omni audio-duration intent; accepted only by models with a matching request parameter."
enable_frame_interpolation: "vLLM-Omni post-processing field; unavailable until FastVideo exposes a frame-interpolation stage."
frame_interpolation_exp: "vLLM-Omni post-processing field; unavailable until FastVideo exposes a frame-interpolation stage."
frame_interpolation_scale: "vLLM-Omni post-processing field; unavailable until FastVideo exposes a frame-interpolation stage."
frame_interpolation_model_path: "vLLM-Omni post-processing field; unavailable until FastVideo exposes a frame-interpolation stage."
cli:
notes:
+165 -92
View File
@@ -1,116 +1,189 @@
# OpenAI-compatible HTTP Contract
# OpenAI-compatible HTTP contract
The stateless FastVideo HTTP server lives at
[`fastvideo/entrypoints/openai/`](https://github.com/hao-ai-lab/FastVideo/tree/main/fastvideo/entrypoints/openai).
Launch: `fastvideo serve --config serve.yaml`.
FastVideo exposes one model-agnostic REST engine for image and video models.
Launch it from a typed serve config:
```bash
fastvideo serve --config examples/serving/openai_fasth3.yaml
```
All generation routes share one serialized engine. FastVideo pipelines mutate
per-request sampling state, and some adapters merge weights at load time, so a
single loaded pipeline is never entered concurrently by image and video
requests. HTTP handling and job polling remain asynchronous.
## Endpoints
| Method | Path | Description |
| --- | --- | --- |
| `POST` | `/v1/videos/generations` | Synchronous video generation |
| `GET` | `/v1/videos` | List prior jobs held in the in-memory store |
| `GET` | `/v1/videos/{id}` | Job status / result |
| `GET` | `/v1/videos/{id}/content` | Download the MP4 once ready |
| `POST` | `/v1/images/generations` | Synchronous image generation |
| `GET` | `/v1/models` | Enumerate registered models |
| `GET` | `/v1/models` | List the served model and optional startup adapter |
| `GET` | `/v1/models/{model}` | Retrieve one served model card |
| `POST` | `/v1/videos` | Submit an asynchronous video job |
| `POST` | `/v1/videos/sync` | Generate and return an MP4 response directly |
| `GET` | `/v1/videos` | List in-memory jobs with `after`, `limit`, and `order` |
| `GET` | `/v1/videos/{id}` | Retrieve job status and metadata |
| `GET` | `/v1/videos/{id}/content` | Download a completed MP4 |
| `DELETE` | `/v1/videos/{id}` | Delete a job and its completed artifact |
| `POST` | `/v1/images` | Generate an image |
| `POST` | `/v1/images/edits` | Generate an image from image references |
| `GET` | `/v1/images/{id}/content` | Download a generated image |
| `GET` | `/health` | Liveness probe |
## `VideoGenerationsRequest` shape
`POST /v1/videos/generations` remains an alias for older FastVideo clients.
The OpenAI Python and JavaScript clients can create, retrieve, list, download,
and delete video jobs. Use the [H3 server cookbook](../../cookbook/openai-api.md)
for pinned client versions and executable examples. Download variants other
than `video` return HTTP 400. Remix, extensions, and characters are not implemented.
Mirrors the OpenAI `POST /v1/videos/generations` shape:
An H3 text-to-video/audio server also serves `/playground/`. This same-origin
browser client uses the video routes above and shares their loaded pipeline.
`GET /playground/config` reports the model alias and operator-explicit sampling
defaults. The playground does not add authentication or a second model process.
For native FastH3 MLX, launch
`python -m fastvideo.entrypoints.openai.mlx_server --config examples/serving/mlx_fasth3.yaml`.
This adapter shares the video-job API and executes pipeline calls on one MLX
thread. It rejects unsupported options before media fetching or job creation.
Image routes are not mounted. MLX keeps its existing phase-memory policy, so a
persistent server does not imply persistent residency of all model weights.
## Video requests
The canonical shape follows vLLM-Omni and accepts SGLang's common flat
extensions. Fields that FastVideo cannot represent for the loaded model fail
at admission with HTTP 400 instead of creating a job that later fails.
```json
{
"prompt": "a fox running through snow",
"size": "1024x1536",
"seconds": 5,
"fps": 24,
"num_frames": 121,
"model": "fasth3",
"prompt": "A fox runs through fresh snow.",
"seconds": "5",
"size": "1344x768",
"video_params": {
"fps": 24,
"num_frames": 124
},
"seed": 42,
"num_inference_steps": 8,
"num_inference_steps": 5,
"guidance_scale": 1.0,
"negative_prompt": "blurry, low quality",
"input_reference": "/path/to/init.png"
}
```
SGLang-compatible extensions carried today:
`num_inference_steps`, `guidance_scale`, `guidance_scale_2`,
`true_cfg_scale`, `negative_prompt`, `enable_teacache`, `output_path`.
## Merge precedence
The server builds a `GenerationRequest` each call using three layers,
highest first:
1. **Request body (client-explicit)** — only fields carried in
`request.model_fields_set` (Pydantic v2). Unset fields do not count,
even if the Pydantic model has a schema default for them.
2. **`ServeConfig.default_request` (operator-explicit)** — projected via
[`explicit_request_updates()`](https://github.com/hao-ai-lab/FastVideo/blob/main/fastvideo/api/compat.py);
only fields the operator actually wrote into the YAML count as
defaults. Every other field inherits the schema default rather than
being pinned.
3. **Hardcoded fallback** — e.g. `fps = 24`.
The gate matters: both surfaces carry schema defaults. Without
`model_fields_set` / explicit-path tracking, schema defaults would
masquerade as intent and silently shadow the other side.
See [`video_api.py::_build_generation_kwargs`](https://github.com/hao-ai-lab/FastVideo/blob/main/fastvideo/entrypoints/openai/video_api.py)
for the canonical implementation; the per-request assembly lives there,
not in pipeline code.
## Continuation state
The stateless surface accepts an opaque `ContinuationState` round-trip.
Clients that want continuation pass the prior `state` blob back on the
next request, and receive a new one on the response when
`request.output.return_state = true`.
Shape:
```json
{
"state": {
"kind": "ltx2.v1",
"payload": { "schema_version": 1, "segment_index": 3, ... }
"image_reference": [
{"image_url": "https://example.com/first-frame.png"}
],
"extra_params": {
"vsa_mode": "exempt"
}
}
```
Payload is always JSON-serializable. Large tensors may live in an
opaque blob-store reference the client simply round-trips; see
[`LTX2ContinuationState`](https://github.com/hao-ai-lab/FastVideo/blob/main/fastvideo/pipelines/basic/ltx2/continuation.py).
Resolution precedence matches vLLM-Omni:
Continuation is not yet wired all the way through to
`generator.generate_video(...)` — PR 7.6 (GPU pool upstream) is the
pipeline-level consumer. PR 7 locked the envelope so this surface is
stable ahead of that plumbing.
1. `size`
2. top-level `width` and `height`
3. `video_params.width` and `video_params.height`
## Error codes
Top-level `fps` and `num_frames` similarly take precedence over the nested
block. If `num_frames` is absent, `seconds * fps` is used. FastVideo also keeps
the legacy `input_reference`, `reference_url`, `video_path`, and `video_url`
spellings.
| HTTP | Condition |
| --- | --- |
| `400 Bad Request` | Parse/validation failure (unknown field, type mismatch, incompatible preset/state) |
| `404 Not Found` | `GET /v1/videos/{id}` for an unknown job |
| `409 Conflict` | Job id already exists |
| `500 Internal Server Error` | Pipeline raised; body mirrors upstream OpenAI error envelope |
| `503 Service Unavailable` | No generator loaded, or shutdown in progress |
Reference objects support URL or local-path strings through `image_url`,
`video_url`, and `audio_url`. `file_id` references are schema-compatible but
return HTTP 400 because FastVideo does not provide an OpenAI Files store.
Image URLs, data URLs, local paths, and multipart `input_reference` uploads are
materialized and decoded under the configured output directory during
admission. Invalid media returns HTTP 400 before a job is created.
Errors include a JSON body with
`{"error": {"type": "...", "message": "..."}}` matching the OpenAI
Python SDK's expectation.
## Jobs and synchronous responses
## What does not cross this boundary
An asynchronous submission returns a `video` object in `queued` state. Its
status advances through `in_progress` to `completed` or `failed`. Completed
jobs expose `file_name`, the FastVideo compatibility extension `file_path`,
timings, and peak-memory metadata when the pipeline reports them.
* Flat legacy kwargs (`ltx2_refine_enabled`, `torch_compile_kwargs`,
etc.) — these are init-time, configured via `ServeConfig.generator`,
never per-request.
* Private Dreamverse-only fields — those live in a private adapter on
the Dreamverse side; the public FastVideo surface never promises
backward compatibility for them.
* Raw tensor payloads (`ltx2_audio_clean_latent` et al.) — these are
derived by the pipeline from `ContinuationState`, never shipped as
request fields.
`POST /v1/videos/sync` returns `video/mp4` bytes. It includes
`X-Request-Id`, `X-Model`, `X-Inference-Time-S`, `X-Stage-Durations`, and
`X-Peak-Memory-MB` headers. Its temporary MP4 is removed after the response is
streamed. Asynchronous artifacts remain available until their job is deleted.
Output paths are controlled by the server. Clients cannot choose filesystem
destinations; every video is written beneath `server.output_dir` with a unique
request id.
FastVideo's synchronous CUDA execution cannot be interrupted after launch.
Deleting an in-progress resource removes it from the API immediately; the
engine remains serialized until the call exits and then removes any artifact.
## Model and LoRA selection
`server.served_model_name` controls the public model id. If omitted, the
checkpoint path is used. Requests that name another model fail with HTTP 400.
LoRAs are configured under
`generator.pipeline.components.{lora_path,lora_nickname,lora_strength}`. The
startup adapter is the only model advertised by a LoRA server, and requests can
select it by its model nickname or with a selector:
```json
{
"prompt": "A fox runs through fresh snow.",
"model": "fasth3-dense-datafree",
"lora": {
"name": "fasth3-dense-datafree",
"path": "/models/adapter_model.safetensors",
"scale": 1.0
}
}
```
The selector must match the adapter already loaded at startup. FastH3 adapter
files can contain dense replacement tensors and VSA gates in addition to
low-rank factors, so swapping them inside concurrent requests would corrupt
shared pipeline state. A mismatch is rejected with HTTP 400.
## MiniMax-H3 and FastH3
FastH3 uses the same general routes and adapter. `task` is accepted for
SGLang-compatible H3 clients:
- `t2va` uses text only.
- `fl2va` takes one or two image references.
- `ref2va` takes ordered image, video, and audio references and requires a
server started with `MiniMaxH3Ref2VAModularPipeline`.
The released FastH3 pipeline generates one packed video/audio result per
request, uses 24 fps, requires guidance scale 1, and accepts frame counts on
its causal-VAE grid. The serving examples pin its five-point distilled sigma
schedule (four DiT forwards).
## Defaults and errors
Incoming explicit fields override operator-explicit `default_request` fields,
which override model preset defaults. Pydantic defaults do not masquerade as
client intent; the transport uses `model_fields_set`, while typed config parsing
tracks the exact paths written by the operator.
Errors use the OpenAI envelope:
```json
{
"error": {
"message": "...",
"type": "invalid_request_error",
"param": null,
"code": 400
}
}
```
Parse, model-selection, startup-LoRA, and unsupported-parameter failures are
HTTP 400; missing resources are HTTP 404. Generation failures are stored on
asynchronous jobs. Retrieving a failed job returns HTTP 200 with `status: "failed"`
and a string error code, so OpenAI SDK polling returns the terminal resource
instead of retrying it as a transport error. Synchronous generation failures
still return HTTP 500.
Unknown top-level fields are rejected. `extra_params` accepts only the explicit
request-batch passthrough fields supported by the typed request adapter.
`GET /health` also verifies that the generation engine is open and all local
multiprocess workers are alive. It returns HTTP 503 when the worker pool is no
longer usable.

Some files were not shown because too many files have changed in this diff Show More